Reducing Image Generation cost with AMD and the Luminal Compiler

With the Luminal compiler and AMD’s MI300x, we show that image generation with Flux.2 Klein 9B can be reduced by 47%

Reis McMillan's avatar

Reis McMillan

Sep 17, 2026

4

Share

Introduction

Image generation is a compute-bound problem which requires accelerated hardware to be viable for commercial use cases. Most inference providers look to NVIDIA GPUs, particularly the Blackwell, Hopper, Ampere, and RTX architectures, to serve such models. The high cost of these chips, though, presents a unique opportunity to generate better margins by looking to alternative hardware which offers more FLOPs per dollar, in particular AMD’s Instinct architecture. We show that using the Luminal compiler, competitive inference latency can be retained while reducing inference costs significantly (> 40%.)