In our last post, we developed a pixel-space encoder-decoder architecture (JiT-DDT). By throwing away the VAE, we were able to compress tokens more aggressively. In turn, this slashed the cost of attention and directly led to faster training and inference.

While effective, it's unclear how to scale encoder-decoder architectures under fixed parameter budgets. So today, we propose the Pyramid-JiT (P-JiT), a novel decoder-only pixel space architecture that beats our prior models, by predicting the target image at several resolutions along the DiT trunk. It reaches Linum v2's FD-DINOv2 on 11.3× fewer training samples, and trains in 4.3× fewer GPU-hours at 4× the pixels.

Research release

P-JiT code and model weights are available under the Apache 2.0 license. We hope that by sharing our findings with the broader community, we can encourage others to also explore more efficient training methods. This should be treated as a research artifact, not a full model release. Stay tuned for more research checkpoints like this, en route to Linum v3.