Dust: Pretraining Transformers Without Backpropagation

@misc{dahal2026backprop,
  title  = {Dust: Pretraining Transformers Without Backpropagation},
  author = {Dahal, Samip and Mandal, Bishwas and G{\"u}lbahar, Serdar and Vegesna, Akshay},
  year   = {2026},
  url    = {https://qlabs.sh/research/dust}
}

TL;DR

  • We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models. Dust perturbs activations (node perturbation) independently at every token, so each token is a virtual population member and one forward pass evaluates them all in parallel.
  • Dust approximates backprop closely at large population (i.e. substantially more compute) and in multiple settings even exceeds it. This hints that in a compute-rich regime we might be able to surpass backprop.
  • Dust is orders of magnitude more efficient than weight-space ES. From 1M tokens up, Dust is on the order of $10^3$ to $10^4$ times more efficient than a transformer implementation of EGGROLL, a state-of-the-art ES method, based on our extrapolations.
  • Zeroth-order methods are widely believed not to scale to large networks. Strikingly, we find larger models are more population-efficient, not less: a 243M-parameter model outperforms a $120\times$ smaller model at most population sizes.
  • Dust’s gradient estimates align better with backprop’s as population grows, and stay well aligned at every scale we test, up to 1B tokens, which is encouraging for scaling.