Comparing Muon, NorMuon and AdamW for Fine-tuning a Dense Retriever

September 30, 2026

In my last post, I played with late interaction for instruction-following retrieval. This time I want to talk about the other half of every training recipe, the part I usually do not think about very much: the optimizer.

Muon is everywhere right now. Instead of applying the raw momentum, it orthogonalizes the momentum of each hidden weight matrix with a few Newton–Schulz iterations. In LLM pretraining it has been reported to match AdamW with about half the compute (Liu et al., 2025), and NorMuon adds neuron-wise normalization on top. It has also reached the retrieval world: in a small ablation on a NanoBEIR subset, the mxbai-edge-colbert-v0 report found Muon at its best learning rate slightly ahead of the best AdamW run (0.599 vs. 0.592 nDCG@10), though behind that run at every other learning rate, and concluded that Muon appears to be a strong optimizer for ColBERT training.