Comparing Muon, NorMuon and AdamW for Fine-tuning a Dense Retriever
September 30, 2026
In my last post, I played with late interaction for instruction-following retrieval. This time I want to talk about the other half of every training recipe, the part I usually do not think about very much: the optimizer.
Muon is everywhere right now. Instead of applying the raw momentum, it orthogonalizes the momentum of each hidden weight matrix with a few Newton–Schulz iterations. In LLM pretraining it has been reported to match AdamW with about half the compute (Liu et al., 2025), and NorMuon adds neuron-wise normalization on top. It has also reached the retrieval world: in a small ablation on a NanoBEIR subset, the mxbai-edge-colbert-v0 report found Muon at its best learning rate slightly ahead of the best AdamW run (0.599 vs. 0.592 nDCG@10), though behind that run at every other learning rate, and concluded that Muon appears to be a strong optimizer for ColBERT training.
That made me curious. Fine-tuning a retriever is a pretty different regime from pretraining an LLM: the starting model is already contrastively trained, batches are small, training is short, and what we really care about is how well the model transfers to domains it has never seen. Whether Muon’s advantages carry over to fine-tuning a model that was pretrained with Adam is itself an open question. So I asked a narrow version of it: if I give AdamW and Muon the same tuning effort, does Muon give me a better retriever?
The short answer is no, at least not in my setup. Muon and NorMuon fit the training data better in every seed I ran, but the retrievers they produced were no better on BEIR (Figure 1). The details are below.
Setup
I fine-tuned lightonai/DenseOn-unsupervised (Sourty et al., 2026), a 22-layer ModernBERT-style encoder with 768-dimensional embeddings that has already gone through contrastive pretraining. The training data are 500K queries from lightonai/embeddings-fine-tuning, each with one positive and seven mined hard negatives. I use InfoNCE with temperature 0.02 and no in-batch negatives, one epoch (3,907 steps) at batch size 128, 10% warmup followed by linear decay, and bf16. One run takes about 7.8 hours on four GPUs, whichever optimizer I use.
Muon and NorMuon only update the 88 hidden weight matrices (attention and MLP). As the Muon authors recommend, everything else, here the token-embedding matrix and 45 norm weights, is trained by a regular AdamW running alongside. To keep the comparison clean, this AdamW uses exactly the settings of the best AdamW run (learning rate 5e-5, same weight decay), so the only difference between the optimizers is how the hidden matrices are updated. For Muon and NorMuon I use momentum 0.95 and five Newton–Schulz steps, plus β₂ = 0.95 for NorMuon; weight decay is 0.01 everywhere except the norms.
Each optimizer gets six peak learning rates: 1e-6, 3e-6, 1e-5, 3e-5, 5e-5 and 1e-4 for AdamW, and 1e-4, 2e-4, 3e-4, 5e-4, 1e-3 and 3e-3 for Muon and NorMuon. Two rules decide the rest:
- Learning rates are selected by the contrastive loss on 4,096 held-out queries, never by BEIR.
- Evaluation is nDCG@10 on LightOn’s decontaminated versions of 14 BEIR tasks (test queries and documents that also appear in the mGTE pretraining data are removed), on full corpora and with equal task weights. Five of these tasks are in-domain, because their source datasets are part of the training mix (FEVER, FiQA-2018, HotpotQA, MS MARCO and NQ); the other nine are out-of-domain.
Each selected recipe is then retrained with two more seeds (2027 and 3407), without any retuning.
Results
Validation loss picks 5e-5 for AdamW and 5e-4 for both Muon and NorMuon. Both Muon-class optimizers reach a lower validation loss than the best AdamW run, but on BEIR their sweeps from 1e-4 to 1e-3 sit at AdamW’s level rather than above it.
Then I retrained the selected recipes with two more seeds:
| Recipe | Seed 42 | 2027 | 3407 | Mean | vs. AdamW [95% CI] | p |
|---|---|---|---|---|---|---|
| AdamW 5e-5 | 59.18 | 59.13 | 59.00 | 59.10 | – | – |
| Muon 5e-4 | 59.08 | 59.15 | 59.37 | 59.20 | +0.09 [−0.51, +0.70] | 0.57 |
| NorMuon 5e-4 | 59.06 | 59.23 | 59.15 | 59.15 | +0.04 [−0.32, +0.40] | 0.67 |
Neither Muon nor NorMuon beats AdamW. Three seeds cannot show that the recipes are equivalent, since the intervals are 0.7 to 1.2 points wide, but they do bound the effect: for NorMuon, the 95% interval excludes an advantage above 0.40, while for Muon, one favorable seed (3407, +0.37) keeps the upper end at 0.70.
So what does Muon do differently?
That does not mean the two optimizers behave the same. In every seed, the Muon and NorMuon recipes
- end with a lower training loss (0.234 to 0.237, compared with 0.240 for AdamW, averaged over the last 200 steps),
- reach a lower loss on the held-out queries, by 0.002 to 0.005,
- and score higher on average over the five in-domain BEIR tasks, by 0.04 to 0.63.
They also move the hidden weights further away from the pretrained model: a 3.0% relative change, compared with 1.7% for AdamW at seed 42. The nine out-of-domain tasks, however, do not follow. There, the differences range from −0.30 to +0.32 and average −0.06 for Muon and −0.10 for NorMuon.
What I take from this
- No learning rate makes Muon better here. Even if I pick the learning rate by BEIR itself, which my selection rules forbid, the best seed-42 Muon run (3e-4, 59.21) and the best NorMuon run (2e-4, 59.18) are level with AdamW (59.18). The learning rate matters far more than the optimizer.
- Muon did worse on scientific and biomedical retrieval. On all four scientific and biomedical tasks, Muon scores below AdamW in every seed: SciFact, NFCorpus, SCIDOCS and TREC-COVID.
- Held-out loss tunes Muon for fit, not transfer. For a general-purpose retriever, I would try Muon at smaller learning rates and check them on out-of-domain data.
So for fine-tuning an already contrastively pretrained retriever at batch size 128, Muon is not a free upgrade. I would not read this as contradicting the mxbai report, which trained late-interaction models in a different pipeline; it just means that the advantage is not automatic.
My Remaining Questions
- Does Muon help more when the model has more to learn, for example when fine-tuning from a plain MLM checkpoint such as ModernBERT-base instead of an already contrastively trained one?
- Is the setup tilted toward AdamW from the start? Everything before my fine-tuning used Adam-type optimizers.
- Does the picture change with much larger batches?
- Would late-interaction models, like the ones in the mxbai report, behave differently?
- Weight decay is applied in proportion to the learning rate, so at the selected rates Muon’s decay on the hidden matrices is ten times AdamW’s. I did not separate that effect.
That is it for this one. As always, I would really appreciate feedback, corrections, or pointers to related results, especially if you have seen Muon help (or not help) in your own retrieval training.