Muon²: Boosting Muon via Adaptive Second-Moment Preconditioning
1University of California, Santa Barbara2University at Albany, SUNY
Muon² divides Muon's momentum by an Adam-style second moment before orthogonalizing it. That one change gives lower perplexity with 40% fewer Newton–Schulz iterations, and reaches Muon's final loss with 23% fewer GPU-hours.
Show the numbers
| Optimizer | GPU-hours | Training steps |
|---|---|---|
| Muon, 5 steps | 1,041 | 4,796 |
| Muon², 3 steps | 816 | 3,862 |
| Muon², 5 steps | 797 | 3,632 |
Abstract
Muon has emerged as a promising optimizer for large-scale foundation model pre-training by exploiting the matrix structure of neural network updates through iterative orthogonalization. However, the orthogonalization quality of Muon hinges on the number of Newton–Schulz (NS) iterations performed, which poses efficiency challenges due to its non-trivial computation and communication cost. We propose Muon², an extension of Muon, to improve both quality and efficiency by applying Adam-style adaptive second-moment preconditioning before orthogonalization. Our key insight is that the core challenge of polar approximation in Muon lies in the ill-conditioned momentum matrix, of which the spectrum is substantially improved by Muon², leading to faster convergence toward a practically sufficient orthogonalization. We further characterize the practical orthogonalization quality via directional alignment, under which Muon² demonstrates dramatic improvement over Muon at each polar step. Across GPT, LLaMA, and Mixture-of-Experts pre-training experiments up to 13B parameters, Muon² (and its memory-efficient variant Muon²-F that preserves most of its benefits) consistently outperforms Muon and its variants while reducing NS iterations by 40%, and saves up to 1/4 training time over Muon when achieving the same loss.
One extra step before Newton–Schulz
Muon replaces the momentum with an approximately orthogonal matrix computed by a few Newton–Schulz (NS) iterations. Each iteration adds matrix multiplications on the full matrix, which distributed training keeps sharded, so an implementation either duplicates the work on every device or communicates in every iteration. Cutting iterations hurts Muon's orthogonalization; Muon² needs fewer because it keeps an Adam-style second moment and scales the momentum element-wise before orthogonalizing it. Everything else is Muon.
Why it helps: a better-conditioned input
NS pushes every singular value of its input toward 1, and tiny singular values need many iterations to get there. With Muon's usual five steps, values below about 10−3 never reach the target range. Early in training, a large share of Muon's normalized momentum sits there. Second-moment scaling moves the whole spectrum about ten times higher and tightens it.
A random-matrix view
Model the momentum as Mij = sij ξij, where sij is a fixed per-entry scale and the ξij are i.i.d. with zero mean and unit variance. Uneven scales push many singular values toward zero. Muon²'s second-moment preconditioning divides each entry by an estimate of sij, so the NS input becomes approximately the i.i.d. matrix ξ:
For rectangular layers (λ < 1), this support is bounded away from zero, so the fraction of singular values in the dead zone vanishes regardless of the scale profile. For the 512 × 1344 MLP projections of LLaMA-60M (λ ≈ 0.38), the lower edge after Frobenius normalization is about 1.7 × 10−2, an order of magnitude above the dead-zone boundary of 10−3. The conclusion still holds when the momentum also carries a signal: a low-rank signal of bounded strength only lowers this edge by a constant factor, and a weak signal of any rank leaves it unchanged.
Measuring orthogonalization by direction
The optimizer only uses the direction of the orthogonalized update; any overall scale is absorbed by the learning rate. So instead of the orthogonality error, Muon² is judged by the cosine similarity between the NS output and the exact polar factor. With 3 steps, Muon² reaches 0.916, close to Muon with 5 steps at 0.931. Cutting Muon to 3 steps drops it to 0.808. The same analysis on GPT-2 Small agrees: Muon² keeps about 4× fewer singular values in the dead zone, and with 3 NS steps it reaches the alignment Muon needs 5 steps for (0.84 against 0.86).
Show the numbers
| Optimizer | 3 NS steps | 5 NS steps |
|---|---|---|
| Muon | 0.808 | 0.931 |
| Muon²-F | 0.902 | 0.971 |
| Muon² | 0.916 | 0.975 |
Muon²-F keeps Muon's memory footprint
Storing the full second moment costs one extra matrix per weight. Muon²-F keeps only Adafactor-style row and column statistics and rebuilds the second moment from their outer product. Its memory stays at Muon's level, and it keeps most of the gain.
Measured on one H100 94GB in bf16 with sequence length 4096 and micro batch size 1. The dashed line marks Muon.
Results
GPT models pre-trained on FineWeb (3.0B to 15.5B tokens), LLaMA models on C4 (60M to 13B parameters), and a 7B mixture-of-experts model with 1B active parameters, each at 3 and 5 NS steps. Every learning rate was swept for both optimizers.
Fewer steps, lower perplexity
With 3 NS steps, Muon² has lower validation perplexity than Muon with 5 steps on every model. At equal steps the gap is wider still.
| Model | 3 NS steps | 5 NS steps | ||||
|---|---|---|---|---|---|---|
| Muon | Muon² | Muon²-F | Muon | Muon² | Muon²-F | |
| GPT-Small | 32.70 | 28.12 | 28.30 | 29.51 | 27.95 | 27.93 |
| GPT-Base | 24.69 | 20.39 | 21.17 | 21.47 | 19.96 | 20.58 |
| GPT-Large | 21.13 | 16.99 | 17.69 | 17.56 | 16.52 | 16.55 |
| LLaMA-60M | 26.37 | 24.59 | 24.68 | 24.98 | 24.60 | 24.66 |
| LLaMA-350M | 14.91 | 13.44 | 13.55 | 14.03 | 13.46 | 13.44 |
| LLaMA-1B | 11.63 | 10.42 | 10.49 | 10.62 | 10.21 | 10.21 |
| LLaMA-13B | 10.54 | 9.11 | n/a | 9.24 | 8.71 | n/a |
| MoE-7B-A1B | 11.93 | 9.85 | n/a | 10.61 | 9.66 | n/a |
Not tied to one learning rate
On GPT-Large, Muon² is better at every learning rate in the sweep, with 3 steps and with 5.
Against other Muon variants
PolarExpress and Turbo-Muon change the polar approximation; NorMuon and AdaMuon add second-moment statistics to the update itself. With tuned hyper-parameters, none of them matches Muon², even when they run more NS steps than it does.
| GPT-Small3 steps | GPT-Small5 steps | GPT-Base3 steps | GPT-Base5 steps | |
|---|---|---|---|---|
| Muon² | 28.12 | 27.95 | 20.39 | 19.96 |
| Muon | 32.70 | 29.51 | 24.69 | 21.47 |
| PolarExpress | 30.01 | 29.42 | 22.74 | 21.16 |
| Turbo-Muon | 29.70 | 29.66 | 23.46 | 21.93 |
| NorMuon | 30.35 | 28.40 | 23.33 | 21.27 |
| AdaMuon | 31.20 | 29.30 | 26.07 | 22.42 |
Downstream
On eight tasks with the trained LLaMA-1B, Muon² averages 47.83 with 5 steps against 45.36 for Muon, and with 3 steps (45.89) it is still ahead of Muon with 5. At this scale ARC-Challenge, WinoGrande and MMLU sit close to random guessing when evaluated zero-shot, so every model is fine-tuned on those three tasks with the same protocol before evaluation. GPT-Large shows the same ordering.
Show all eight tasks
| ARC-c† | ARC-e | OBQA | HellaSwag | PIQA | WinoGrande† | LAMBADA | MMLU† | Avg | |
|---|---|---|---|---|---|---|---|---|---|
| LLaMA-1B, Muon, 3 steps | 27.08 | 42.80 | 32.00 | 44.37 | 71.22 | 56.75 | 36.52 | 29.75 | 42.56 |
| LLaMA-1B, Muon, 5 steps | 30.01 | 46.09 | 33.20 | 49.75 | 72.25 | 59.91 | 40.99 | 30.71 | 45.36 |
| LLaMA-1B, Muon², 3 steps | 31.91 | 47.81 | 31.80 | 50.23 | 72.20 | 59.88 | 42.50 | 30.79 | 45.89 |
| LLaMA-1B, Muon², 5 steps | 33.48 | 49.03 | 35.20 | 52.61 | 74.05 | 61.90 | 45.08 | 31.31 | 47.83 |
| GPT-Large, Muon, 3 steps | 27.99 | 44.23 | 30.80 | 39.92 | 68.82 | 54.56 | 32.19 | 27.02 | 40.69 |
| GPT-Large, Muon, 5 steps | 28.21 | 44.07 | 30.40 | 42.50 | 69.80 | 59.12 | 37.07 | 25.41 | 42.07 |
| GPT-Large, Muon², 3 steps | 28.30 | 44.91 | 29.40 | 43.72 | 69.70 | 60.77 | 40.03 | 26.38 | 42.90 |
| GPT-Large, Muon², 5 steps | 30.26 | 46.93 | 31.60 | 45.01 | 69.80 | 60.38 | 40.07 | 25.52 | 43.70 |
† fine-tuned on the task before evaluation, with the same protocol for every optimizer; the other tasks are zero-shot. On GPT-Large all four models stay at chance on MMLU even after fine-tuning.
Citation
@inproceedings{liu2026muon2,
title = {Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning},
author = {Liu, Ziyue and Zhang, Ruijie and Wang, Zhengyang and Zhao, Yequan and
Su, Yupeng and Yang, Zi and Zhang, Zheng},
booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
year = {2026}
}