{"article":{"slug":"reducing-image-generation-cost-with-amd-and-the-luminal-compiler","title":"Reducing Image Generation cost with AMD and the Luminal Compiler","subtitle":null,"summary":"Luminal engineers show Flux.2 Klein 9B image generation costs cut by up to 47% on AMD MI300X versus an Nvidia H200, using the Luminal compiler.","content_type":"blog_post","language":"en","canonical_url":"https://blog.luminal.com/p/reducing-image-generation-cost-with","author":{"name":"Reis McMillan","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Luminal","url":"https://blog.luminal.com/","listing_slug":"luminal","listing":{"slug":"luminal","name":"Luminal","listing_type":"company","url":"https://listedstartups.com/companies/luminal"}},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"},{"name":"AMD","slug":"amd","url":"https://listedarticles.com/topics/amd"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":3489,"reading_minutes":15,"published_at":"2026-09-21T12:00:00.000Z","added_at":"2026-09-26T03:13:15.759Z","updated_at":"2026-09-26T03:13:15.759Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/reducing-image-generation-cost-with-amd-and-the-luminal-compiler","markdown_url":"https://listedarticles.com/articles/reducing-image-generation-cost-with-amd-and-the-luminal-compiler.md","example":false,"citation":"Reis McMillan, Luminal. \"Reducing Image Generation cost with AMD and the Luminal Compiler.\" 21 Sept 2026. https://blog.luminal.com/p/reducing-image-generation-cost-with (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://blog.luminal.com/p/reducing-image-generation-cost-with"},"body_markdown":"# Reducing Image Generation cost with AMD and the Luminal Compiler\n\n### With the Luminal compiler and AMD’s MI300x, we show that image generation with Flux.2 Klein 9B can be reduced by 47%\n\n[![Reis McMillan's avatar](https://substackcdn.com/image/fetch/$s_!SLjT!,w_36,h_36,c_fill,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadf98603-067c-4942-ba04-cf528668e34e_1365x1365.jpeg)](https://substack.com/@reismcmillan)\n\n[Reis McMillan](https://substack.com/@reismcmillan)\n\nSep 17, 2026\n\n4\n\nShare\n\n[![](https://substackcdn.com/image/fetch/$s_!nNHu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F320ad25b-60d3-452f-851a-94c8216b331b_1448x1086.png)](https://substackcdn.com/image/fetch/$s_!nNHu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F320ad25b-60d3-452f-851a-94c8216b331b_1448x1086.png)\n\n## Introduction\n\nImage generation is a compute-bound problem which requires accelerated hardware to be viable for commercial use cases. Most inference providers look to NVIDIA GPUs, particularly the Blackwell, Hopper, Ampere, and RTX architectures, to serve such models. The high cost of these chips, though, presents a unique opportunity to generate better margins by looking to alternative hardware which offers more FLOPs per dollar, in particular AMD’s Instinct architecture. We show that using the Luminal compiler, competitive inference latency can be retained while reducing inference costs significantly (> 40%.)\n\nThe issue that presents itself revolves around AMD’s software stack. While ROCm has made many improvements since its initial release in 2016, it is still generally regarded as inferior to CUDA. The general consensus is that AMD achieves lower realized TFLOPs at inference time, erasing the margins to be made from the lower cost of AMD chips. This keeps many providers from making the transition to AMD, despite the opportunity for vastly better margins. Furthermore, such a transition usually requires hiring kernel engineers with expertise in a specific chip manufacturer’s software stack, in this case AMD… a nontrivial barrier to diversifying across hardware architectures.\n\nAs we will show, using the Luminal ML graph compiler we can achieve highly competitive performance for image generation on AMD MI300X chips, which translates to significant cost savings for inference providers. The MI300X finishes an image 31% to 16% behind an H200 depending on resolution, while renting for less than half the price. Per image, that works out to 41% to 47% cheaper.\n\n## Background\n\n#### E-graphs\n\nAn e-graph is a data structure for compactly representing a set of equivalent expressions (programs). More concretely, if we think of an ML program as an expression which defines its output(s) in terms of its input(s), then an e-graph compactly represents a set of equivalent expressions.\n\nTwo components form an e-graph: e-nodes and e-classes. E-classes are composed of e-nodes which represent interchangeable terms of an expression, and each e-node points (via an edge) to a subsequent e-class.\n\nThe usefulness of such a representation is that one will always form a valid expression (program) by doing the following:\n\n  1. Begin at the root e-class of an expression.\n\n  2. Select any e-node from the e-class.\n\n  3. Follow the directed edges from that e-node to the e-classes it points at.\n\n  4. Repeat steps 2 and 3 until arriving at leaf nodes.\n\nThe common example that many find useful is to think about the expression `(a*2)/2`. More specifically the reader should think about its representation as a postfix expression: `a 2 * 2 /`.\n\n[![](https://substackcdn.com/image/fetch/$s_!ys9s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93962c30-27ec-4073-9a10-d932709ff74c_904x804.png)](https://substackcdn.com/image/fetch/$s_!ys9s!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93962c30-27ec-4073-9a10-d932709ff74c_904x804.png)The e-graph for (a*2)/2 after equality saturation with the rules in the next section. Solid boxes are e-nodes; dotted boxes are e-classes. The outermost e-class holds the original node, the a*1 node, and a itself, because all three are equivalent. The inner e-class holds both spellings of a*2, and the rightmost holds 2/2 and 1.\n\n#### Equality saturation and `egglog`\n\nGiven some original expression (program) and a set of rules, equality saturation is the process by which an e-graph is populated to form many equivalent expressions.\n\nGoing back to the earlier `(a*2)/2` example, a useful set of rules would be:\n\n  * `x*2 → x«1`\n\n  * `(x*y)/y → x (y≠0)`\n\n  * `x/x → 1 (x≠0)`\n\n  * `x*1 → x`\n\nAn important quality of equality saturation is that original expressions are kept as rewritten expressions are added to the e-graph. For example if we applied a sequential (destructive) rewrite `a*2 → a«1`, we would miss the opportunity to cancel out `2/2 `to arrive at the simplified expression `a`.\n\n`egglog` is the domain-specific programming language that we use to express rewrite rules and perform equality saturation over arbitrary expressions.\n\n### The Luminal compiler\n\nThe core insight at Luminal is that ML graph compilation, the process by which ML models are lowered from a high-level representation like that in PyTorch down to an executable program that runs on heterogeneous hardware, can be treated as a search problem. Our compiler leverages `egglog` and e-graphs to take a generic representation of an ML model and lower it down to optimized kernels for a specific hardware backend.\n\nThe optimizations the Luminal compiler forms are two-fold. The simple rewrites apply basic algebraic properties, constant folding, and layout canonicalization. The more interesting rewrites, however, are the hardware backend-specific rewrites. In the context of AMD, some example rewrites might be conceptualized as follows: “this scatter-multiply-sum operation is a matrix multiply which matches a hipBLASLt kernel”, “this matrix-multiply, scale, softmax, matrix multiple combination is an attention op that can be lowered to this fused flash attention kernel”, or “these primitive operations represent a convolution which can be lowered to this tiled convolution kernel.” Notably, every rule adds an alternative to the e-graph without replacing the pattern a particular rewrite matched on. This enables our genetic search algorithm to progress towards globally optimal expressions (programs.) \n\nLuminal’s compiler approach to inference has two primary advantages. First, optimizing models to run in a commercially-viable manner becomes an automated process with our compiler. With the Luminal compiler, model optimization no longer requires dedicated kernel and inference engineers. Second, the friction of moving inference from one hardware stack to another is greatly reduced with our hardware-agnostic compiler.\n\n### Flux.2 Klein 9B\n\n[![FLUX.2 - Next Generation Image Generation | Black Forest Labs](https://substackcdn.com/image/fetch/$s_!uOw1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff30fc8a0-a73b-4acc-a1a9-11864b5b91f5_1120x1408.png)](https://substackcdn.com/image/fetch/$s_!uOw1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff30fc8a0-a73b-4acc-a1a9-11864b5b91f5_1120x1408.png)\n\nThe Flux.2 family of models provides unified image generation and editing for various commercial applications. In our case, we use Flux.2 Klein 9B and compare model performance on an NVIDIA H200 and an AMD MI300X.\n\nFlux.2 Klein 9B supports text-to-image and image-to-image pipelines; we perform our benchmark using the latter. The image-to-image pipeline consists of four components: a VAE (variational autoencoder) image encoder, a Qwen3 8B text encoder, the multi-modal diffusion transformer itself, and a VAE image decoder. All weights are stored as `bf16`.\n\nThe input image is encoded by the VAE into a latent space which downsamples the height and width 16 times and expands the RGB channels to 128 different channels in the latent space. The image tokens represent non-overlapping patches in this latent space.\n\nThe text encoder takes the hidden states from the ninth, eighteenth, and twenty-seventh layers of the text encoder and concatenates them to form a 12,228 element vector per token.\n\nThe transformer consists of eight double-stream blocks and twenty-four single-stream blocks. (The core difference between a double-stream and single-stream block are the projections before the attention mechanism and the normalization layer and feed-forward network after it. I highly recommend reading [this paper](https://arxiv.org/html/2507.09595v1) to better understand the Flux architecture.) The hidden dimension’s size is 4,096 and the transformer has thirty-two attention heads. \n\nKlein 9B uses guidance-free, step-distilled diffusion to deliver low latency image editing. The diffusion stage of the pipeline runs in four steps, and because the model is distilled there is no classifier-free guidance and therefore no second forward pass per step.\n\n## Setup\n\nFlux.2 Klein 9B’s performance on an H200 and an MI300X was compared using three differently sized input images and accompanying prompts. For each pair of input image and text prompt, we ran five warmup runs and then measured 100 end-to-end runs. Each scenario outputs an image the same size as its input.\n\nWe used the following configurations for comparison:\n\n  1. 512x512; “add a boat to the water”.\n\n  2. 704x864; “add a hot dog stand in the background”.\n\n  3. 1200x800; “give the dog wings”.\n\nBoth setups use the same seed (42), four diffusion steps, a 512-token padded prompt, batch size one, and `bf16` weights throughout. Timing on both sides is measured with a device synchronize at every stage boundary, so the per-stage numbers are wall-clock and nothing is hidden by asynchronous execution. Model loading, compilation, and warmup are excluded from the timed runs on both sides, as is writing the output PNG. A serving process pays those costs once at startup, so folding them into per-image latency would describe a cost production never pays.\n\n### The NVIDIA H200 pipeline\n\nThe Nvidia H200 pipeline uses `torch.compile` with `max-autotune` enabled for the reference encoder, text encoder, and decoder sections. The diffusion transformer is compiled using Nvidia’s `TensorRT` backend to deliver maximum runtime performance, only short of handwriting kernels specific to the Flux.2 Klein 9B diffusion transformer. \n\nTo construct the entire pipeline, we use the Diffusers `Flux2KleinPipeline` and its `Flux2Transformer2DModel`, `AutoencoderKLFlux2`, and Hugging Face `Qwen3ForCausalLM` components, running on PyTorch 2.7 with CUDA 12.6.  \n  \nThe H200 pipeline makes the following optimizations:\n\n  * **A static TensorRT engine for the transformer.** The Diffusers `Flux2Transformer2DModel` is exported to ONNX with PyTorch’s dynamo exporter and built into a `TensorRT` engine. This is the NVIDIA-side analogue of Luminal’s compile-and-search step: a shape-specialised program, with the search done by TensorRT’s tactic selection.\n\n  * **What the engine runs per step.** TensorRT’s own runtime enqueues the engine directly on the PyTorch stream. Each diffusion step is 433 kernel launches: 115 Hopper XMMA GEMMs, 32 fused multi-head attention kernels, and `TensorRT`’s fused elementwise code. Across the run the GEMMs take 58% of GPU time, the attention kernels 23%, and the fused kernels 13%, with a step averaging 252 ms and the GPU busy for about 97% of the run.\n\n  * **Fixed-shape everything.** Prompts are always padded to 512 tokens, width and height must be multiples of 16, and the engine is built for exactly one token count; a different output size derives a different engine file and a rebuild. Static shapes are what make max-autotune, CUDA graphs, and a single-profile engine viable.\n\n  * **Fused attention at every site.** Inside the engine, all 32 transformer attention sites run through TensorRT’s fused attention kernel at about 1.9 ms per call, roughly a quarter of each step. In the text encoder, PyTorch’s scaled_dot_product_attention dispatches to the CUTLASS memory-efficient attention kernel inside the Inductor graph, once per Qwen3 layer. The VAE mid-block attention takes the same path in the encoder and the decoder.\n\n### The AMD MI300X pipeline\n\nThe MI300X examples run using Luminal’s ROCm backend. The different components of the Flux.2 Klein 9B pipeline are written in Luminal’s high-level graph language (the transformer, the Qwen3 text encoder, and the VAE encoder and decoder are all Luminal graphs), and the compiler lowers each of the four stages to a program specific to the MI300X. The Luminal high-level graph is a generic model spelling: there are no references to kernels, libraries, or layouts. The compiler choses all of these itself. The compiler’s ROCm lowers to the following:\n\n  * **hipBLASLt** for matrix multiplies, AMD’s equivalent of cuBLASLt, with bias epilogues fused into the GEMM where the compiler can prove the layout allows it. All `bf16` GEMMs accumulate in `fp32`.\n\n  * **Fused attention.** The compiler identifies fused flash-attention kernels that compute `softmax(QK^T)V` in a tiled fashion, maximizing realized FLOPs. Both the text encoder and diffusion transformer make use of such a kernel.\n\n  * **Tiled convolution kernels.** Similar to the fused attention, tiled convolution kernels are dispatched to, by shape, at runtime. Said kernels are always ahead-of-time compiled based on the shapes the compiler identifies in the ML graph.\n\n  * **Fused RoPE, RMSNorm, and LayerNormModule kernels.** As opposed to launching separate GEMM kernels and reduction kernels, each of these operations is launched as one fused kernel to save on kernel launch overhead.\n\n  * **Generated HIP kernels** for everything else: fused elementwise chains, reductions, gathers, casts. The compiler identifies regions of the ML graph which can be fused together into a singular kernel.\n\n  * **HIP graphs** to batch kernel launches. Each transformer step is 33 HIP graph launches interleaved with 32 attention calls, rather than roughly 1,250 individual kernel launches.\n\nRuntime execution is otherwise uneventful: the four stages run back to back, the Euler sampler integrates the latent on the host in `fp32`, and the reference latent is computed once per run and appended to the noise tokens.\n\nIt should be noted that the entire compilation of the pipeline (effectively four separate compilations), the phase where Luminal performs it’s kernel search, takes a few minutes. While not insignificant, this cost is payed up front across the five warmup and one-hundred timed runs. In a production setting, the cost of compilation would be negligible as it is amortized across inference requests.\n\n## Results\n\nEach configuration shows the input image, the H200 output, and the MI300X output side by side, followed by the per-stage timings on each GPU and the cost comparison. Costs use $4.00 per GPU-hour for the H200 and $1.85 per GPU-hour for the MI300X.\n\n### Configuration 1: 512x512, “add a boat to the water”\n\n[![](https://substackcdn.com/image/fetch/$s_!MPoN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc14b05bb-94c0-4921-9e2d-f7a967e6b2c2_512x512.png)](https://substackcdn.com/image/fetch/$s_!MPoN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc14b05bb-94c0-4921-9e2d-f7a967e6b2c2_512x512.png)Original\n\n[![](https://substackcdn.com/image/fetch/$s_!00-n!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c01c6c4-58d0-4209-bd47-befd64ccc6a6_512x512.png)](https://substackcdn.com/image/fetch/$s_!00-n!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c01c6c4-58d0-4209-bd47-befd64ccc6a6_512x512.png)H200\n\n[![](https://substackcdn.com/image/fetch/$s_!ovXL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc01d2d0-efef-411b-a2d4-2232a63ed73b_512x512.png)](https://substackcdn.com/image/fetch/$s_!ovXL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc01d2d0-efef-411b-a2d4-2232a63ed73b_512x512.png)MI300x\n\n**NVIDIA H200 (PyTorch,**`TensorRT`**):**\n\n[![](https://substackcdn.com/image/fetch/$s_!dZjj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a8f646e-3c2d-4955-b9b6-0addf8686739_1240x587.png)](https://substackcdn.com/image/fetch/$s_!dZjj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a8f646e-3c2d-4955-b9b6-0addf8686739_1240x587.png)\n\n**AMD MI300X (Luminal):**\n\n[![](https://substackcdn.com/image/fetch/$s_!cOQi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51731002-0876-427b-b126-9ccb8d7b5d38_1240x586.png)](https://substackcdn.com/image/fetch/$s_!cOQi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51731002-0876-427b-b126-9ccb8d7b5d38_1240x586.png)\n\nThe MI300X takes 31% longer per image and costs approximately 41% less per image.\n\n[![](https://substackcdn.com/image/fetch/$s_!jR4X!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c80733-a58c-4674-bc73-f650ab2b2750_1240x409.png)](https://substackcdn.com/image/fetch/$s_!jR4X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c80733-a58c-4674-bc73-f650ab2b2750_1240x409.png)\n\n### Configuration 2: 704x864, “add a hot dog stand in the background”\n\n[![](https://substackcdn.com/image/fetch/$s_!r48k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ed93ade-befc-49a0-9b43-1d2f38c6a331_704x864.png)](https://substackcdn.com/image/fetch/$s_!r48k!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ed93ade-befc-49a0-9b43-1d2f38c6a331_704x864.png)Original\n\n[![](https://substackcdn.com/image/fetch/$s_!nwB4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5efbf1e-aeb7-4dbc-a435-318cf6dceb63_704x864.png)](https://substackcdn.com/image/fetch/$s_!nwB4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5efbf1e-aeb7-4dbc-a435-318cf6dceb63_704x864.png)H200\n\n[![](https://substackcdn.com/image/fetch/$s_!4ZS8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8435a33-e9c1-41dc-a33e-e2aa4d586f57_704x864.png)](https://substackcdn.com/image/fetch/$s_!4ZS8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8435a33-e9c1-41dc-a33e-e2aa4d586f57_704x864.png)MI300x\n\n**NVIDIA H200 (PyTorch,**`TensorRT`**):**\n\n[![](https://substackcdn.com/image/fetch/$s_!Nmz3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3787ee67-9bf5-4f5a-8a8d-dbf28dcf1969_1240x587.png)](https://substackcdn.com/image/fetch/$s_!Nmz3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3787ee67-9bf5-4f5a-8a8d-dbf28dcf1969_1240x587.png)\n\n**AMD MI300X (Luminal):**\n\n[![](https://substackcdn.com/image/fetch/$s_!KaoI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F313e2a34-616e-455e-91b8-ea83bf3755d6_1240x587.png)](https://substackcdn.com/image/fetch/$s_!KaoI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F313e2a34-616e-455e-91b8-ea83bf3755d6_1240x587.png)\n\nThe MI300X takes 16% longer per image and costs 47% less per image.\n\n[![](https://substackcdn.com/image/fetch/$s_!_Ewo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe64b2a2b-cc5d-45ba-bc32-4dd90a76f341_1240x408.png)](https://substackcdn.com/image/fetch/$s_!_Ewo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe64b2a2b-cc5d-45ba-bc32-4dd90a76f341_1240x408.png)\n\n### Configuration 3: 1200x800, “give the dog wings”\n\n[![](https://substackcdn.com/image/fetch/$s_!vgT3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88292096-9b3f-481f-a5b7-f3f74acb5dcd_1200x800.png)](https://substackcdn.com/image/fetch/$s_!vgT3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88292096-9b3f-481f-a5b7-f3f74acb5dcd_1200x800.png)Original\n\n[![](https://substackcdn.com/image/fetch/$s_!FKNS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8344760a-d968-4b26-942b-d2eca6527dcf_1200x800.png)](https://substackcdn.com/image/fetch/$s_!FKNS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8344760a-d968-4b26-942b-d2eca6527dcf_1200x800.png)H200\n\n[![](https://substackcdn.com/image/fetch/$s_!kZ0t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ba13541-4bcf-4a37-989a-5c70c531c48b_1200x800.png)](https://substackcdn.com/image/fetch/$s_!kZ0t!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ba13541-4bcf-4a37-989a-5c70c531c48b_1200x800.png)MI300x\n\n**NVIDIA H200 (PyTorch,**`TensorRT`**):**\n\n[![](https://substackcdn.com/image/fetch/$s_!u1Tr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F789bad9a-6abd-4454-bf5d-f66966ef848a_1240x589.png)](https://substackcdn.com/image/fetch/$s_!u1Tr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F789bad9a-6abd-4454-bf5d-f66966ef848a_1240x589.png)\n\n**AMD MI300X (Luminal):**\n\n[![](https://substackcdn.com/image/fetch/$s_!tiVC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc194f27-b687-4cd7-aced-829409b32333_1240x584.png)](https://substackcdn.com/image/fetch/$s_!tiVC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc194f27-b687-4cd7-aced-829409b32333_1240x584.png)\n\nThe MI300X takes 17% longer per image and costs 46% less per image.\n\n[![](https://substackcdn.com/image/fetch/$s_!cJ7q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9eaf291b-a3b6-47b4-8df2-62e9d451cead_1240x410.png)](https://substackcdn.com/image/fetch/$s_!cJ7q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9eaf291b-a3b6-47b4-8df2-62e9d451cead_1240x410.png)\n\n### Comparing Luminal to SGLang Diffusion\n\nWhile the primary concern of our benchmark was to compare the performance the Luminal compiler delivered on an AMD MI300x as compared to an Nvidia H200 using `TensorRT`, we also felt it useful to compare the Luminal compiler against SGLang on the AMD MI300x.  \n  \nWe made the same assessment with SGLang’s multi-modal engine and benchmarked the same three images and corresponding resolutions. We only assessed offline throughput, so no latency introduced by the SGLang inference server was measured. Additionally, SGLang was configured to use `torch.compile`, which the inference engine _does not_ do by default. SGLang _does_ use ROCm’s AITER backend which dispatches out to flash attention kernels during the diffusion stage, but SGLang does not make use of fused LayerNorm+adaLN or RMSNorm+Rope kernels, unlike the Luminal compiler. The result is that the Luminal compiler outperforms SGLang on all three of the image configurations.  \n  \nLuminal achieves latencies of 0.424s, 0.827s, and 1.311s on the 512x512, 704x864, and 1200x800 images respectively. By comparison, SGLang is slower at 1.06s, 1.18s, and 1.44s respectively. (It is interesting to note how the gap narrows as the image becomes larger. As the diffusion stage’s share of the latency increases, SGLang’s relative performance improves as it _does_ optimize the diffusion stage reasonably well.)\n\n[![](https://substackcdn.com/image/fetch/$s_!xBsh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8df8e31a-a457-406b-8f57-d981f21bcfa3_1069x662.png)](https://substackcdn.com/image/fetch/$s_!xBsh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8df8e31a-a457-406b-8f57-d981f21bcfa3_1069x662.png)\n\n### Where the time goes on the MI300X\n\nTo understand what the compiler actually produced, we profiled the 1200x800 run with a rocprofv3 GPU kernel trace, bracketed with ROCTx ranges so that only one steady-state run is recorded and the compile and warm-up are left out.\n\nOne diffusion step on the GPU is 616 kernel executions. Of those, 249 are hipBLASLt GEMM kernels captured inside HIP graphs, 32 are fused attention calls, one per attention site, and the remaining 335 are kernels Luminal generated: 152 fused elementwise kernels, 137 fused normalization kernels, and 46 small index, cast, and constant kernels. Where the GPU time goes within a step:\n\n[![](https://substackcdn.com/image/fetch/$s_!2tCk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95134e78-e775-4d1e-9665-021a27cb3c48_1240x566.png)](https://substackcdn.com/image/fetch/$s_!2tCk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95134e78-e775-4d1e-9665-021a27cb3c48_1240x566.png)\n\nThis is the profile you want for a transformer: two thirds of the time is in vendor-tuned matrix multiplies, a quarter is in fused attention, and the long tail of \"glue\" ops (norms, activations, modulation, RoPE, residual adds) is compressed into under three hundred fused kernels that together take about six percent. The GPU is busy for 97% of each step's wall-clock. The remaining three percent is launch gaps, most of them the ~35 microseconds pause between a HIP graph finishing and the next attention call starting. \n\n### What the compiler found on its own\n\nThe transformer, text encoder, and VAE were written as plain tensor math. Luminal’s front end has no “matmul”, “attention”, “convolution”, or “layer norm” op. It has a small set of primitives: elementwise arithmetic (add, multiply, exponent, reciprocal, and so on), reductions (sum, max), and index manipulation (gather, scatter, iota). A matrix multiply is spelled as a broadcast multiply followed by a sum. Attention is spelled as two of those, a scale, a max, an exponent, another sum, and a divide. A convolution is spelled as a gather that unfolds the input into patches, followed by a multiply and a sum. Every optimization below was discovered by the compiler from those spellings during equality saturation and search; none of it is written into the model.\n\n**Optimized GEMM kernels recognized from multiply-and-sum.** The compiler matches every multiply-then-sum whose strides describe a matrix product and adds a vendor-tuned GEMM kernel to its e-class as an alternative. In the final transformer program every one of the 249 GEMMs per step went to hipBLASLt; the only generic matmuls left in the whole run are 49 tiny projections in the text encoder that together take 0.3 ms.\n\n**Flash attention recognized from seven primitive ops.** The scaled-dot-product pattern (scores, scale, softmax, weighted sum) was matched at every one of the 32 attention sites, in both the double-stream and single-stream block shapes, and each one was replaced with a single fused flash-attention kernel. In the text encoder the same rules also matched the causal-plus-padding mask, which the model spells as an additive bias before the softmax. The VAE’s single mid-block attention was matched the same way in both the encoder and the decoder.\n\n**Row normalizations recognized from a reciprocal square root.** RMSNorm is spelled as a square, a sum, a reciprocal square root, and a multiply; LayerNorm adds a mean and a subtract; the adaLN modulation that follows it is a scale and a shift; and RoPE is a gather of each element’s partner and a multiply-add against cosine and sine tables. Per step that is 80 RMSNorm+RoPE launches and 57 LayerNorm+modulate launches, each about 45 microseconds, and the reductions they replaced are gone from the profile entirely.\n\n**Convolutions recognized from gather-multiply-sum.** The VAE spells each convolution as an unfold followed by a matrix product. The compiler proves the unfold indexing is a convolution and replaces the chain with an optimized convolution kernel: 25 of them in the encoder and 33 in the decoder, accounting for 58% and 61% of those stages’ GPU time. The zero-padding around each convolution, which the front end spells as a chain of scatters and gathers, is rewritten into a single fused select kernel that produces the padded buffer the convolution reads directly.\n\n**Elementwise chains fused into single kernels.** Two fusion rules, one that grows a fused region by pulling a neighboring elementwise op into it and one that merges two adjacent regions, fired hundreds of times in the transformer stage. The result is 17 distinct fused kernels per diffusion step, launched 152 times, covering the SwiGLU activation, the gate-and-residual add after each block, the concatenation of the text and image streams, and the timestep embedding.\n\nThe claim is not that any of these optimizations are exotic or novel. A kernel engineer would think to implement every single on of these optimizations. The key point is that the compiler recognizes and chooses these optimizations from a generic spelling of Flux.2, no kernel engineer has to go in and optimize the stack. No tuning a config, no installing special attention backends, and no installing different ROCm libraries to select fused kernels from. All of this ships with the Luminal ROCm runtime.\n\n### Cost summary\n\nPutting the three configurations together, using $4.00 per GPU-hour for the H200 and $1.85 per GPU-hour for the MI300X:\n\n[![](https://substackcdn.com/image/fetch/$s_!Yp8p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a85f7d3-711f-4060-9ea6-76e6ba996bca_1240x443.png)](https://substackcdn.com/image/fetch/$s_!Yp8p!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a85f7d3-711f-4060-9ea6-76e6ba996bca_1240x443.png)\n\nThe math is simple. The MI300X rents for 46% of the H200’s price, so the break-even point is an MI300X that takes 2.16x as long per image. In practice it takes 1.16x to 1.31x as long, which leaves 41% to 47% of the per-image cost on the table. At one million 1200x800 edits, that is roughly $1,240 on H200s versus $670 on MI300Xs.\n\nIt cannot go without saying: latency matters for interactive products, and the MI300X is slower per image. That being said, for image generation workloads, latency is generally less of a concern. An additional fifth of a second is usually negligible for most products, especially considering the additional latency communication over a network introduces.\n\n## Discussion\n\nThe primary purpose of this endeavor was to assess the viability of AMD chips and the and a properly optimized ROCm stack for image generation workloads in production settings. The issue remains, however, that optimizing a ROCm stack is non-trivial for many inference providers and transitioning to AMD hardware generates a lot of friction for already time-constrained engineers.  \n  \nWe have shown, however, is that the Luminal compiler can absorb a lot of this friction and compile an image generation pipeline, in our case the Flux.2 Klein 9B pipeline, that is competitive from a latency perspective with the same model being served using an optimized Nvidia stack. From a cost perspective, the AMD stack wins every time, and by a significant margin as well (greater than 40% cost savings.)\n\n4\n\nShare","body_html":"<h1 id=\"reducing-image-generation-cost-with-amd-and-the-luminal-compiler\">Reducing Image Generation cost with AMD and the Luminal Compiler</h1>\n<h3 id=\"with-the-luminal-compiler-and-amd-s-mi300x-we-show-that-image-ge\">With the Luminal compiler and AMD’s MI300x, we show that image generation with Flux.2 Klein 9B can be reduced by 47%</h3>\n<p><a href=\"https://substack.com/@reismcmillan\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!SLjT!,w_36,h_36,c_fill,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fadf98603-067c-4942-ba04-cf528668e34e_1365x1365.jpeg\" alt=\"Reis McMillan&#39;s avatar\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a></p>\n<p><a href=\"https://substack.com/@reismcmillan\" rel=\"nofollow ugc noopener\">Reis McMillan</a></p>\n<p>Sep 17, 2026</p>\n<p>4</p>\n<p>Share</p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!nNHu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F320ad25b-60d3-452f-851a-94c8216b331b_1448x1086.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!nNHu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F320ad25b-60d3-452f-851a-94c8216b331b_1448x1086.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a></p>\n<h2 id=\"introduction\">Introduction</h2>\n<p>Image generation is a compute-bound problem which requires accelerated hardware to be viable for commercial use cases. Most inference providers look to NVIDIA GPUs, particularly the Blackwell, Hopper, Ampere, and RTX architectures, to serve such models. The high cost of these chips, though, presents a unique opportunity to generate better margins by looking to alternative hardware which offers more FLOPs per dollar, in particular AMD’s Instinct architecture. We show that using the Luminal compiler, competitive inference latency can be retained while reducing inference costs significantly (&gt; 40%.)</p>\n<p>The issue that presents itself revolves around AMD’s software stack. While ROCm has made many improvements since its initial release in 2016, it is still generally regarded as inferior to CUDA. The general consensus is that AMD achieves lower realized TFLOPs at inference time, erasing the margins to be made from the lower cost of AMD chips. This keeps many providers from making the transition to AMD, despite the opportunity for vastly better margins. Furthermore, such a transition usually requires hiring kernel engineers with expertise in a specific chip manufacturer’s software stack, in this case AMD… a nontrivial barrier to diversifying across hardware architectures.</p>\n<p>As we will show, using the Luminal ML graph compiler we can achieve highly competitive performance for image generation on AMD MI300X chips, which translates to significant cost savings for inference providers. The MI300X finishes an image 31% to 16% behind an H200 depending on resolution, while renting for less than half the price. Per image, that works out to 41% to 47% cheaper.</p>\n<h2 id=\"background\">Background</h2>\n<h4 id=\"e-graphs\">E-graphs</h4>\n<p>An e-graph is a data structure for compactly representing a set of equivalent expressions (programs). More concretely, if we think of an ML program as an expression which defines its output(s) in terms of its input(s), then an e-graph compactly represents a set of equivalent expressions.</p>\n<p>Two components form an e-graph: e-nodes and e-classes. E-classes are composed of e-nodes which represent interchangeable terms of an expression, and each e-node points (via an edge) to a subsequent e-class.</p>\n<p>The usefulness of such a representation is that one will always form a valid expression (program) by doing the following:</p>\n<ol><li>Begin at the root e-class of an expression.</li><li>Select any e-node from the e-class.</li><li>Follow the directed edges from that e-node to the e-classes it points at.</li><li>Repeat steps 2 and 3 until arriving at leaf nodes.</li></ol>\n<p>The common example that many find useful is to think about the expression <code>(a*2)/2</code>. More specifically the reader should think about its representation as a postfix expression: <code>a 2 * 2 /</code>.</p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!ys9s!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93962c30-27ec-4073-9a10-d932709ff74c_904x804.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!ys9s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F93962c30-27ec-4073-9a10-d932709ff74c_904x804.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a>The e-graph for (a<em>2)/2 after equality saturation with the rules in the next section. Solid boxes are e-nodes; dotted boxes are e-classes. The outermost e-class holds the original node, the a</em>1 node, and a itself, because all three are equivalent. The inner e-class holds both spellings of a*2, and the rightmost holds 2/2 and 1.</p>\n<h4 id=\"equality-saturation-and-egglog\">Equality saturation and <code>egglog</code></h4>\n<p>Given some original expression (program) and a set of rules, equality saturation is the process by which an e-graph is populated to form many equivalent expressions.</p>\n<p>Going back to the earlier <code>(a*2)/2</code> example, a useful set of rules would be:</p>\n<ul><li><code>x*2 → x«1</code></li><li><code>(x*y)/y → x (y≠0)</code></li><li><code>x/x → 1 (x≠0)</code></li><li><code>x*1 → x</code></li></ul>\n<p>An important quality of equality saturation is that original expressions are kept as rewritten expressions are added to the e-graph. For example if we applied a sequential (destructive) rewrite <code>a*2 → a«1</code>, we would miss the opportunity to cancel out <code>2/2 </code>to arrive at the simplified expression <code>a</code>.</p>\n<p><code>egglog</code> is the domain-specific programming language that we use to express rewrite rules and perform equality saturation over arbitrary expressions.</p>\n<h3 id=\"the-luminal-compiler\">The Luminal compiler</h3>\n<p>The core insight at Luminal is that ML graph compilation, the process by which ML models are lowered from a high-level representation like that in PyTorch down to an executable program that runs on heterogeneous hardware, can be treated as a search problem. Our compiler leverages <code>egglog</code> and e-graphs to take a generic representation of an ML model and lower it down to optimized kernels for a specific hardware backend.</p>\n<p>The optimizations the Luminal compiler forms are two-fold. The simple rewrites apply basic algebraic properties, constant folding, and layout canonicalization. The more interesting rewrites, however, are the hardware backend-specific rewrites. In the context of AMD, some example rewrites might be conceptualized as follows: “this scatter-multiply-sum operation is a matrix multiply which matches a hipBLASLt kernel”, “this matrix-multiply, scale, softmax, matrix multiple combination is an attention op that can be lowered to this fused flash attention kernel”, or “these primitive operations represent a convolution which can be lowered to this tiled convolution kernel.” Notably, every rule adds an alternative to the e-graph without replacing the pattern a particular rewrite matched on. This enables our genetic search algorithm to progress towards globally optimal expressions (programs.) </p>\n<p>Luminal’s compiler approach to inference has two primary advantages. First, optimizing models to run in a commercially-viable manner becomes an automated process with our compiler. With the Luminal compiler, model optimization no longer requires dedicated kernel and inference engineers. Second, the friction of moving inference from one hardware stack to another is greatly reduced with our hardware-agnostic compiler.</p>\n<h3 id=\"flux-2-klein-9b\">Flux.2 Klein 9B</h3>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!uOw1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff30fc8a0-a73b-4acc-a1a9-11864b5b91f5_1120x1408.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!uOw1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff30fc8a0-a73b-4acc-a1a9-11864b5b91f5_1120x1408.png\" alt=\"FLUX.2 - Next Generation Image Generation | Black Forest Labs\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a></p>\n<p>The Flux.2 family of models provides unified image generation and editing for various commercial applications. In our case, we use Flux.2 Klein 9B and compare model performance on an NVIDIA H200 and an AMD MI300X.</p>\n<p>Flux.2 Klein 9B supports text-to-image and image-to-image pipelines; we perform our benchmark using the latter. The image-to-image pipeline consists of four components: a VAE (variational autoencoder) image encoder, a Qwen3 8B text encoder, the multi-modal diffusion transformer itself, and a VAE image decoder. All weights are stored as <code>bf16</code>.</p>\n<p>The input image is encoded by the VAE into a latent space which downsamples the height and width 16 times and expands the RGB channels to 128 different channels in the latent space. The image tokens represent non-overlapping patches in this latent space.</p>\n<p>The text encoder takes the hidden states from the ninth, eighteenth, and twenty-seventh layers of the text encoder and concatenates them to form a 12,228 element vector per token.</p>\n<p>The transformer consists of eight double-stream blocks and twenty-four single-stream blocks. (The core difference between a double-stream and single-stream block are the projections before the attention mechanism and the normalization layer and feed-forward network after it. I highly recommend reading <a href=\"https://arxiv.org/html/2507.09595v1\" rel=\"nofollow ugc noopener\">this paper</a> to better understand the Flux architecture.) The hidden dimension’s size is 4,096 and the transformer has thirty-two attention heads. </p>\n<p>Klein 9B uses guidance-free, step-distilled diffusion to deliver low latency image editing. The diffusion stage of the pipeline runs in four steps, and because the model is distilled there is no classifier-free guidance and therefore no second forward pass per step.</p>\n<h2 id=\"setup\">Setup</h2>\n<p>Flux.2 Klein 9B’s performance on an H200 and an MI300X was compared using three differently sized input images and accompanying prompts. For each pair of input image and text prompt, we ran five warmup runs and then measured 100 end-to-end runs. Each scenario outputs an image the same size as its input.</p>\n<p>We used the following configurations for comparison:</p>\n<ol><li>512x512; “add a boat to the water”.</li><li>704x864; “add a hot dog stand in the background”.</li><li>1200x800; “give the dog wings”.</li></ol>\n<p>Both setups use the same seed (42), four diffusion steps, a 512-token padded prompt, batch size one, and <code>bf16</code> weights throughout. Timing on both sides is measured with a device synchronize at every stage boundary, so the per-stage numbers are wall-clock and nothing is hidden by asynchronous execution. Model loading, compilation, and warmup are excluded from the timed runs on both sides, as is writing the output PNG. A serving process pays those costs once at startup, so folding them into per-image latency would describe a cost production never pays.</p>\n<h3 id=\"the-nvidia-h200-pipeline\">The NVIDIA H200 pipeline</h3>\n<p>The Nvidia H200 pipeline uses <code>torch.compile</code> with <code>max-autotune</code> enabled for the reference encoder, text encoder, and decoder sections. The diffusion transformer is compiled using Nvidia’s <code>TensorRT</code> backend to deliver maximum runtime performance, only short of handwriting kernels specific to the Flux.2 Klein 9B diffusion transformer. </p>\n<p>To construct the entire pipeline, we use the Diffusers <code>Flux2KleinPipeline</code> and its <code>Flux2Transformer2DModel</code>, <code>AutoencoderKLFlux2</code>, and Hugging Face <code>Qwen3ForCausalLM</code> components, running on PyTorch 2.7 with CUDA 12.6.  </p>\n<p>The H200 pipeline makes the following optimizations:</p>\n<ul><li><strong>A static TensorRT engine for the transformer.</strong> The Diffusers <code>Flux2Transformer2DModel</code> is exported to ONNX with PyTorch’s dynamo exporter and built into a <code>TensorRT</code> engine. This is the NVIDIA-side analogue of Luminal’s compile-and-search step: a shape-specialised program, with the search done by TensorRT’s tactic selection.</li><li><strong>What the engine runs per step.</strong> TensorRT’s own runtime enqueues the engine directly on the PyTorch stream. Each diffusion step is 433 kernel launches: 115 Hopper XMMA GEMMs, 32 fused multi-head attention kernels, and <code>TensorRT</code>’s fused elementwise code. Across the run the GEMMs take 58% of GPU time, the attention kernels 23%, and the fused kernels 13%, with a step averaging 252 ms and the GPU busy for about 97% of the run.</li><li><strong>Fixed-shape everything.</strong> Prompts are always padded to 512 tokens, width and height must be multiples of 16, and the engine is built for exactly one token count; a different output size derives a different engine file and a rebuild. Static shapes are what make max-autotune, CUDA graphs, and a single-profile engine viable.</li><li><strong>Fused attention at every site.</strong> Inside the engine, all 32 transformer attention sites run through TensorRT’s fused attention kernel at about 1.9 ms per call, roughly a quarter of each step. In the text encoder, PyTorch’s scaled_dot_product_attention dispatches to the CUTLASS memory-efficient attention kernel inside the Inductor graph, once per Qwen3 layer. The VAE mid-block attention takes the same path in the encoder and the decoder.</li></ul>\n<h3 id=\"the-amd-mi300x-pipeline\">The AMD MI300X pipeline</h3>\n<p>The MI300X examples run using Luminal’s ROCm backend. The different components of the Flux.2 Klein 9B pipeline are written in Luminal’s high-level graph language (the transformer, the Qwen3 text encoder, and the VAE encoder and decoder are all Luminal graphs), and the compiler lowers each of the four stages to a program specific to the MI300X. The Luminal high-level graph is a generic model spelling: there are no references to kernels, libraries, or layouts. The compiler choses all of these itself. The compiler’s ROCm lowers to the following:</p>\n<ul><li><strong>hipBLASLt</strong> for matrix multiplies, AMD’s equivalent of cuBLASLt, with bias epilogues fused into the GEMM where the compiler can prove the layout allows it. All <code>bf16</code> GEMMs accumulate in <code>fp32</code>.</li><li><strong>Fused attention.</strong> The compiler identifies fused flash-attention kernels that compute <code>softmax(QK^T)V</code> in a tiled fashion, maximizing realized FLOPs. Both the text encoder and diffusion transformer make use of such a kernel.</li><li><strong>Tiled convolution kernels.</strong> Similar to the fused attention, tiled convolution kernels are dispatched to, by shape, at runtime. Said kernels are always ahead-of-time compiled based on the shapes the compiler identifies in the ML graph.</li><li><strong>Fused RoPE, RMSNorm, and LayerNormModule kernels.</strong> As opposed to launching separate GEMM kernels and reduction kernels, each of these operations is launched as one fused kernel to save on kernel launch overhead.</li><li><strong>Generated HIP kernels</strong> for everything else: fused elementwise chains, reductions, gathers, casts. The compiler identifies regions of the ML graph which can be fused together into a singular kernel.</li><li><strong>HIP graphs</strong> to batch kernel launches. Each transformer step is 33 HIP graph launches interleaved with 32 attention calls, rather than roughly 1,250 individual kernel launches.</li></ul>\n<p>Runtime execution is otherwise uneventful: the four stages run back to back, the Euler sampler integrates the latent on the host in <code>fp32</code>, and the reference latent is computed once per run and appended to the noise tokens.</p>\n<p>It should be noted that the entire compilation of the pipeline (effectively four separate compilations), the phase where Luminal performs it’s kernel search, takes a few minutes. While not insignificant, this cost is payed up front across the five warmup and one-hundred timed runs. In a production setting, the cost of compilation would be negligible as it is amortized across inference requests.</p>\n<h2 id=\"results\">Results</h2>\n<p>Each configuration shows the input image, the H200 output, and the MI300X output side by side, followed by the per-stage timings on each GPU and the cost comparison. Costs use $4.00 per GPU-hour for the H200 and $1.85 per GPU-hour for the MI300X.</p>\n<h3 id=\"configuration-1-512x512-add-a-boat-to-the-water\">Configuration 1: 512x512, “add a boat to the water”</h3>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!MPoN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc14b05bb-94c0-4921-9e2d-f7a967e6b2c2_512x512.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!MPoN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc14b05bb-94c0-4921-9e2d-f7a967e6b2c2_512x512.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a>Original</p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!00-n!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c01c6c4-58d0-4209-bd47-befd64ccc6a6_512x512.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!00-n!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c01c6c4-58d0-4209-bd47-befd64ccc6a6_512x512.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a>H200</p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!ovXL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc01d2d0-efef-411b-a2d4-2232a63ed73b_512x512.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!ovXL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc01d2d0-efef-411b-a2d4-2232a63ed73b_512x512.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a>MI300x</p>\n<p><strong>NVIDIA H200 (PyTorch,</strong><code>TensorRT</code><strong>):</strong></p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!dZjj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a8f646e-3c2d-4955-b9b6-0addf8686739_1240x587.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!dZjj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9a8f646e-3c2d-4955-b9b6-0addf8686739_1240x587.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a></p>\n<p><strong>AMD MI300X (Luminal):</strong></p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!cOQi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51731002-0876-427b-b126-9ccb8d7b5d38_1240x586.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!cOQi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F51731002-0876-427b-b126-9ccb8d7b5d38_1240x586.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a></p>\n<p>The MI300X takes 31% longer per image and costs approximately 41% less per image.</p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!jR4X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c80733-a58c-4674-bc73-f650ab2b2750_1240x409.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!jR4X!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06c80733-a58c-4674-bc73-f650ab2b2750_1240x409.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a></p>\n<h3 id=\"configuration-2-704x864-add-a-hot-dog-stand-in-the-background\">Configuration 2: 704x864, “add a hot dog stand in the background”</h3>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!r48k!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ed93ade-befc-49a0-9b43-1d2f38c6a331_704x864.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!r48k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ed93ade-befc-49a0-9b43-1d2f38c6a331_704x864.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a>Original</p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!nwB4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5efbf1e-aeb7-4dbc-a435-318cf6dceb63_704x864.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!nwB4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5efbf1e-aeb7-4dbc-a435-318cf6dceb63_704x864.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a>H200</p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!4ZS8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8435a33-e9c1-41dc-a33e-e2aa4d586f57_704x864.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!4ZS8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff8435a33-e9c1-41dc-a33e-e2aa4d586f57_704x864.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a>MI300x</p>\n<p><strong>NVIDIA H200 (PyTorch,</strong><code>TensorRT</code><strong>):</strong></p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!Nmz3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3787ee67-9bf5-4f5a-8a8d-dbf28dcf1969_1240x587.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!Nmz3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3787ee67-9bf5-4f5a-8a8d-dbf28dcf1969_1240x587.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a></p>\n<p><strong>AMD MI300X (Luminal):</strong></p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!KaoI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F313e2a34-616e-455e-91b8-ea83bf3755d6_1240x587.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!KaoI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F313e2a34-616e-455e-91b8-ea83bf3755d6_1240x587.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a></p>\n<p>The MI300X takes 16% longer per image and costs 47% less per image.</p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!_Ewo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe64b2a2b-cc5d-45ba-bc32-4dd90a76f341_1240x408.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!_Ewo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe64b2a2b-cc5d-45ba-bc32-4dd90a76f341_1240x408.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a></p>\n<h3 id=\"configuration-3-1200x800-give-the-dog-wings\">Configuration 3: 1200x800, “give the dog wings”</h3>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!vgT3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88292096-9b3f-481f-a5b7-f3f74acb5dcd_1200x800.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!vgT3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88292096-9b3f-481f-a5b7-f3f74acb5dcd_1200x800.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a>Original</p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!FKNS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8344760a-d968-4b26-942b-d2eca6527dcf_1200x800.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!FKNS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8344760a-d968-4b26-942b-d2eca6527dcf_1200x800.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a>H200</p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!kZ0t!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ba13541-4bcf-4a37-989a-5c70c531c48b_1200x800.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!kZ0t!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ba13541-4bcf-4a37-989a-5c70c531c48b_1200x800.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a>MI300x</p>\n<p><strong>NVIDIA H200 (PyTorch,</strong><code>TensorRT</code><strong>):</strong></p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!u1Tr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F789bad9a-6abd-4454-bf5d-f66966ef848a_1240x589.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!u1Tr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F789bad9a-6abd-4454-bf5d-f66966ef848a_1240x589.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a></p>\n<p><strong>AMD MI300X (Luminal):</strong></p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!tiVC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc194f27-b687-4cd7-aced-829409b32333_1240x584.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!tiVC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc194f27-b687-4cd7-aced-829409b32333_1240x584.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a></p>\n<p>The MI300X takes 17% longer per image and costs 46% less per image.</p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!cJ7q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9eaf291b-a3b6-47b4-8df2-62e9d451cead_1240x410.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!cJ7q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9eaf291b-a3b6-47b4-8df2-62e9d451cead_1240x410.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a></p>\n<h3 id=\"comparing-luminal-to-sglang-diffusion\">Comparing Luminal to SGLang Diffusion</h3>\n<p>While the primary concern of our benchmark was to compare the performance the Luminal compiler delivered on an AMD MI300x as compared to an Nvidia H200 using <code>TensorRT</code>, we also felt it useful to compare the Luminal compiler against SGLang on the AMD MI300x.  </p>\n<p>We made the same assessment with SGLang’s multi-modal engine and benchmarked the same three images and corresponding resolutions. We only assessed offline throughput, so no latency introduced by the SGLang inference server was measured. Additionally, SGLang was configured to use <code>torch.compile</code>, which the inference engine <em>does not</em> do by default. SGLang <em>does</em> use ROCm’s AITER backend which dispatches out to flash attention kernels during the diffusion stage, but SGLang does not make use of fused LayerNorm+adaLN or RMSNorm+Rope kernels, unlike the Luminal compiler. The result is that the Luminal compiler outperforms SGLang on all three of the image configurations.  </p>\n<p>Luminal achieves latencies of 0.424s, 0.827s, and 1.311s on the 512x512, 704x864, and 1200x800 images respectively. By comparison, SGLang is slower at 1.06s, 1.18s, and 1.44s respectively. (It is interesting to note how the gap narrows as the image becomes larger. As the diffusion stage’s share of the latency increases, SGLang’s relative performance improves as it <em>does</em> optimize the diffusion stage reasonably well.)</p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!xBsh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8df8e31a-a457-406b-8f57-d981f21bcfa3_1069x662.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!xBsh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8df8e31a-a457-406b-8f57-d981f21bcfa3_1069x662.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a></p>\n<h3 id=\"where-the-time-goes-on-the-mi300x\">Where the time goes on the MI300X</h3>\n<p>To understand what the compiler actually produced, we profiled the 1200x800 run with a rocprofv3 GPU kernel trace, bracketed with ROCTx ranges so that only one steady-state run is recorded and the compile and warm-up are left out.</p>\n<p>One diffusion step on the GPU is 616 kernel executions. Of those, 249 are hipBLASLt GEMM kernels captured inside HIP graphs, 32 are fused attention calls, one per attention site, and the remaining 335 are kernels Luminal generated: 152 fused elementwise kernels, 137 fused normalization kernels, and 46 small index, cast, and constant kernels. Where the GPU time goes within a step:</p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!2tCk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95134e78-e775-4d1e-9665-021a27cb3c48_1240x566.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!2tCk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95134e78-e775-4d1e-9665-021a27cb3c48_1240x566.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a></p>\n<p>This is the profile you want for a transformer: two thirds of the time is in vendor-tuned matrix multiplies, a quarter is in fused attention, and the long tail of &quot;glue&quot; ops (norms, activations, modulation, RoPE, residual adds) is compressed into under three hundred fused kernels that together take about six percent. The GPU is busy for 97% of each step&#39;s wall-clock. The remaining three percent is launch gaps, most of them the ~35 microseconds pause between a HIP graph finishing and the next attention call starting. </p>\n<h3 id=\"what-the-compiler-found-on-its-own\">What the compiler found on its own</h3>\n<p>The transformer, text encoder, and VAE were written as plain tensor math. Luminal’s front end has no “matmul”, “attention”, “convolution”, or “layer norm” op. It has a small set of primitives: elementwise arithmetic (add, multiply, exponent, reciprocal, and so on), reductions (sum, max), and index manipulation (gather, scatter, iota). A matrix multiply is spelled as a broadcast multiply followed by a sum. Attention is spelled as two of those, a scale, a max, an exponent, another sum, and a divide. A convolution is spelled as a gather that unfolds the input into patches, followed by a multiply and a sum. Every optimization below was discovered by the compiler from those spellings during equality saturation and search; none of it is written into the model.</p>\n<p><strong>Optimized GEMM kernels recognized from multiply-and-sum.</strong> The compiler matches every multiply-then-sum whose strides describe a matrix product and adds a vendor-tuned GEMM kernel to its e-class as an alternative. In the final transformer program every one of the 249 GEMMs per step went to hipBLASLt; the only generic matmuls left in the whole run are 49 tiny projections in the text encoder that together take 0.3 ms.</p>\n<p><strong>Flash attention recognized from seven primitive ops.</strong> The scaled-dot-product pattern (scores, scale, softmax, weighted sum) was matched at every one of the 32 attention sites, in both the double-stream and single-stream block shapes, and each one was replaced with a single fused flash-attention kernel. In the text encoder the same rules also matched the causal-plus-padding mask, which the model spells as an additive bias before the softmax. The VAE’s single mid-block attention was matched the same way in both the encoder and the decoder.</p>\n<p><strong>Row normalizations recognized from a reciprocal square root.</strong> RMSNorm is spelled as a square, a sum, a reciprocal square root, and a multiply; LayerNorm adds a mean and a subtract; the adaLN modulation that follows it is a scale and a shift; and RoPE is a gather of each element’s partner and a multiply-add against cosine and sine tables. Per step that is 80 RMSNorm+RoPE launches and 57 LayerNorm+modulate launches, each about 45 microseconds, and the reductions they replaced are gone from the profile entirely.</p>\n<p><strong>Convolutions recognized from gather-multiply-sum.</strong> The VAE spells each convolution as an unfold followed by a matrix product. The compiler proves the unfold indexing is a convolution and replaces the chain with an optimized convolution kernel: 25 of them in the encoder and 33 in the decoder, accounting for 58% and 61% of those stages’ GPU time. The zero-padding around each convolution, which the front end spells as a chain of scatters and gathers, is rewritten into a single fused select kernel that produces the padded buffer the convolution reads directly.</p>\n<p><strong>Elementwise chains fused into single kernels.</strong> Two fusion rules, one that grows a fused region by pulling a neighboring elementwise op into it and one that merges two adjacent regions, fired hundreds of times in the transformer stage. The result is 17 distinct fused kernels per diffusion step, launched 152 times, covering the SwiGLU activation, the gate-and-residual add after each block, the concatenation of the text and image streams, and the timestep embedding.</p>\n<p>The claim is not that any of these optimizations are exotic or novel. A kernel engineer would think to implement every single on of these optimizations. The key point is that the compiler recognizes and chooses these optimizations from a generic spelling of Flux.2, no kernel engineer has to go in and optimize the stack. No tuning a config, no installing special attention backends, and no installing different ROCm libraries to select fused kernels from. All of this ships with the Luminal ROCm runtime.</p>\n<h3 id=\"cost-summary\">Cost summary</h3>\n<p>Putting the three configurations together, using $4.00 per GPU-hour for the H200 and $1.85 per GPU-hour for the MI300X:</p>\n<p><a href=\"https://substackcdn.com/image/fetch/$s_!Yp8p!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a85f7d3-711f-4060-9ea6-76e6ba996bca_1240x443.png\" rel=\"nofollow ugc noopener\"><img src=\"https://substackcdn.com/image/fetch/$s_!Yp8p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a85f7d3-711f-4060-9ea6-76e6ba996bca_1240x443.png\" alt=\"\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></a></p>\n<p>The math is simple. The MI300X rents for 46% of the H200’s price, so the break-even point is an MI300X that takes 2.16x as long per image. In practice it takes 1.16x to 1.31x as long, which leaves 41% to 47% of the per-image cost on the table. At one million 1200x800 edits, that is roughly $1,240 on H200s versus $670 on MI300Xs.</p>\n<p>It cannot go without saying: latency matters for interactive products, and the MI300X is slower per image. That being said, for image generation workloads, latency is generally less of a concern. An additional fifth of a second is usually negligible for most products, especially considering the additional latency communication over a network introduces.</p>\n<h2 id=\"discussion\">Discussion</h2>\n<p>The primary purpose of this endeavor was to assess the viability of AMD chips and the and a properly optimized ROCm stack for image generation workloads in production settings. The issue remains, however, that optimizing a ROCm stack is non-trivial for many inference providers and transitioning to AMD hardware generates a lot of friction for already time-constrained engineers.  </p>\n<p>We have shown, however, is that the Luminal compiler can absorb a lot of this friction and compile an image generation pipeline, in our case the Flux.2 Klein 9B pipeline, that is competitive from a latency perspective with the same model being served using an optimized Nvidia stack. From a cost perspective, the AMD stack wins every time, and by a significant margin as well (greater than 40% cost savings.)</p>\n<p>4</p>\n<p>Share</p>","headings":[{"level":1,"text":"Reducing Image Generation cost with AMD and the Luminal Compiler","id":"reducing-image-generation-cost-with-amd-and-the-luminal-compiler"},{"level":3,"text":"With the Luminal compiler and AMD’s MI300x, we show that image generation with Flux.2 Klein 9B can be reduced by 47%","id":"with-the-luminal-compiler-and-amd-s-mi300x-we-show-that-image-ge"},{"level":2,"text":"Introduction","id":"introduction"},{"level":2,"text":"Background","id":"background"},{"level":3,"text":"The Luminal compiler","id":"the-luminal-compiler"},{"level":3,"text":"Flux.2 Klein 9B","id":"flux-2-klein-9b"},{"level":2,"text":"Setup","id":"setup"},{"level":3,"text":"The NVIDIA H200 pipeline","id":"the-nvidia-h200-pipeline"},{"level":3,"text":"The AMD MI300X pipeline","id":"the-amd-mi300x-pipeline"},{"level":2,"text":"Results","id":"results"},{"level":3,"text":"Configuration 1: 512x512, “add a boat to the water”","id":"configuration-1-512x512-add-a-boat-to-the-water"},{"level":3,"text":"Configuration 2: 704x864, “add a hot dog stand in the background”","id":"configuration-2-704x864-add-a-hot-dog-stand-in-the-background"},{"level":3,"text":"Configuration 3: 1200x800, “give the dog wings”","id":"configuration-3-1200x800-give-the-dog-wings"},{"level":3,"text":"Comparing Luminal to SGLang Diffusion","id":"comparing-luminal-to-sglang-diffusion"},{"level":3,"text":"Where the time goes on the MI300X","id":"where-the-time-goes-on-the-mi300x"},{"level":3,"text":"What the compiler found on its own","id":"what-the-compiler-found-on-its-own"},{"level":3,"text":"Cost summary","id":"cost-summary"},{"level":2,"text":"Discussion","id":"discussion"}]}}