The MiniMax H3 video VAE now encodes up to ~2.2x faster and decodes ~1.4-2.7x faster. Together that can roughly halve the time a video workflow spends in the VAE: a 1344x768, 129-frame encode and decode round trip drops from 24.3 to 12.7 seconds.
What changed, the technical details
- A fused encoder kernel, on by default. Between convolutions the encoder normalized each frame, applied an activation and padded the edges, each one a separate pass over hundreds of megabytes. That is now a single pass, writing straight into the memory layout the convolution wants next. ~1.5x faster, with identical output.
- Convolutions that honor
--fast fp16_accumulation*.* PyTorch sends convolutions to cuDNN, NVIDIA’s library, and there is no way through it to ask for fp16 accumulation, so the flag only ever applied to matrix multiplications and the encoder got nothing from it. The new custom convolution does honor it, with the bias and skip connection folded into it. That takes the encoder to ~2.2x. - An int8 decoder. With the int8 VAE file the decoder’s weights are 8-bit, and the normalization, activation and skip connection fold into the matrix multiplications, so intermediate results never reach memory. Attention runs in 8-bit too. 1.4x faster over what int8 VAE used to be.