Ziyue (Alvin) Liu
EMNLP 2026Oral

Muon²: Boosting Muon via Adaptive Second-Moment Preconditioning

Ziyue Liu1, Ruijie Zhang1, Zhengyang Wang1, Yequan Zhao1, Yupeng Su1, Zi Yang2, Zheng Zhang1

1University of California, Santa Barbara2University at Albany, SUNY

In short

Muon² divides Muon's momentum by an Adam-style second moment before orthogonalizing it. That one change gives lower perplexity with 40% fewer Newton–Schulz iterations, and reaches Muon's final loss with 23% fewer GPU-hours.

GPU-hours to reach a loss of 2.36 on LLaMA-1B
2.36 is where Muon with 5 Newton–Schulz steps ends; fewer is better
Muon, 5 steps
1,041
Muon², 3 steps
816−22%
Muon², 5 steps
797−23%
03006009001,200
Show the numbers
OptimizerGPU-hoursTraining steps
Muon, 5 steps1,0414,796
Muon², 3 steps8163,862
Muon², 5 steps7973,632
40%
fewer Newton–Schulz iterations: with 3, Muon² beats Muon with 5 in every setting tested
23%
fewer GPU-hours to reach Muon's final loss on LLaMA-1B
13B
largest model, across GPT, LLaMA and mixture-of-experts pre-training

Abstract

Muon has emerged as a promising optimizer for large-scale foundation model pre-training by exploiting the matrix structure of neural network updates through iterative orthogonalization. However, the orthogonalization quality of Muon hinges on the number of Newton–Schulz (NS) iterations performed, which poses efficiency challenges due to its non-trivial computation and communication cost. We propose Muon², an extension of Muon, to improve both quality and efficiency by applying Adam-style adaptive second-moment preconditioning before orthogonalization. Our key insight is that the core challenge of polar approximation in Muon lies in the ill-conditioned momentum matrix, of which the spectrum is substantially improved by Muon², leading to faster convergence toward a practically sufficient orthogonalization. We further characterize the practical orthogonalization quality via directional alignment, under which Muon² demonstrates dramatic improvement over Muon at each polar step. Across GPT, LLaMA, and Mixture-of-Experts pre-training experiments up to 13B parameters, Muon² (and its memory-efficient variant Muon²-F that preserves most of its benefits) consistently outperforms Muon and its variants while reducing NS iterations by 40%, and saves up to 1/4 training time over Muon when achieving the same loss.

One extra step before Newton–Schulz

Muon replaces the momentum with an approximately orthogonal matrix computed by a few Newton–Schulz (NS) iterations. Each iteration adds matrix multiplications on the full matrix, which distributed training keeps sharded, so an implementation either duplicates the work on every device or communicates in every iteration. Cutting iterations hurts Muon's orthogonalization; Muon² needs fewer because it keeps an Adam-style second moment and scales the momentum element-wise before orthogonalizing it. Everything else is Muon.

1Gt = ∇W ℒ(Wt)gradient
2Mt = β1Mt−1 + (1 − β1) Gtmomentum
3Vt = β2Vt−1 + (1 − β2) Gt ⊙ Gtadded: second moment
4M̃t = Mt ⊘ (√Vt + ε)added: precondition
5Ot = NewtonSchulz(M̃t, K)orthogonalize
6Wt+1 = Wt − η √(m/n) Otupdate

Why it helps: a better-conditioned input

NS pushes every singular value of its input toward 1, and tiny singular values need many iterations to get there. With Muon's usual five steps, values below about 10−3 never reach the target range. Early in training, a large share of Muon's normalized momentum sits there. Second-moment scaling moves the whole spectrum about ten times higher and tightens it.

Singular values entering Newton–Schulz
LLaMA-60M, normalized momentum. Values in the dead zone are still far from 1 after five steps.
Muon²Muon
Early in training, Muon's input spreads from 10−4 to 1 and centers near 10−3, with almost half of it in the dead zone. Muon²'s input centers near 10−2, ten times higher, and mostly falls in the transition zone.

A random-matrix view

Model the momentum as Mij = sij ξij, where sij is a fixed per-entry scale and the ξij are i.i.d. with zero mean and unit variance. Uneven scales push many singular values toward zero. Muon²'s second-moment preconditioning divides each entry by an estimate of sij, so the NS input becomes approximately the i.i.d. matrix ξ:

M̃ij = Mij / sij = ξijsingular values of M̃ / √n → Marchenko–Pastur law on [1 − √λ, 1 + √λ], λ = m / n

For rectangular layers (λ < 1), this support is bounded away from zero, so the fraction of singular values in the dead zone vanishes regardless of the scale profile. For the 512 × 1344 MLP projections of LLaMA-60M (λ ≈ 0.38), the lower edge after Frobenius normalization is about 1.7 × 10−2, an order of magnitude above the dead-zone boundary of 10−3. The conclusion still holds when the momentum also carries a signal: a low-rank signal of bounded strength only lowers this edge by a constant factor, and a weak signal of any rank leaves it unchanged.

Measuring orthogonalization by direction

The optimizer only uses the direction of the orthogonalized update; any overall scale is absorbed by the learning rate. So instead of the orthogonality error, Muon² is judged by the cosine similarity between the NS output and the exact polar factor. With 3 steps, Muon² reaches 0.916, close to Muon with 5 steps at 0.931. Cutting Muon to 3 steps drops it to 0.808. The same analysis on GPT-2 Small agrees: Muon² keeps about 4× fewer singular values in the dead zone, and with 3 NS steps it reaches the alignment Muon needs 5 steps for (0.84 against 0.86).

Alignment with the exact orthogonalized update
cosine similarity at the end of training
3 steps5 steps
Muon
0.8080.931
Muon²-F
0.9020.971
Muon²
0.9160.975
0.800.850.900.951.00
Show the numbers
Optimizer3 NS steps5 NS steps
Muon0.8080.931
Muon²-F0.9020.971
Muon²0.9160.975

Muon²-F keeps Muon's memory footprint

Storing the full second moment costs one extra matrix per weight. Muon²-F keeps only Adafactor-style row and column statistics and rebuilds the second moment from their outer product. Its memory stays at Muon's level, and it keeps most of the gain.

LLaMA-1B
peak training memory, GB
AdamW
Muon
20.39
Muon
18.56
Muon²
20.36
Muon²-F
18.57
0510152025
LLaMA-7B
peak training memory, GB
AdamW
Muon
59.31
Muon
47.29
Muon²
60.66
Muon²-F
47.53
0204060

Measured on one H100 94GB in bf16 with sequence length 4096 and micro batch size 1. The dashed line marks Muon.

Results

GPT models pre-trained on FineWeb (3.0B to 15.5B tokens), LLaMA models on C4 (60M to 13B parameters), and a 7B mixture-of-experts model with 1B active parameters, each at 3 and 5 NS steps. Every learning rate was swept for both optimizers.

Fewer steps, lower perplexity

With 3 NS steps, Muon² has lower validation perplexity than Muon with 5 steps on every model. At equal steps the gap is wider still.

Muon² with 3 steps against Muon with 5 steps
lower validation perplexity, relative
GPT-Small
−4.7%
GPT-Base
−5.0%
GPT-Large
−3.2%
LLaMA-60M
−1.6%
LLaMA-350M
−4.2%
LLaMA-1B
−1.9%
LLaMA-13B
−1.4%
MoE-7B-A1B
−7.2%
0%2%4%6%8%
Model3 NS steps5 NS steps
MuonMuon²Muon²-FMuonMuon²Muon²-F
GPT-Small32.7028.1228.3029.5127.9527.93
GPT-Base24.6920.3921.1721.4719.9620.58
GPT-Large21.1316.9917.6917.5616.5216.55
LLaMA-60M26.3724.5924.6824.9824.6024.66
LLaMA-350M14.9113.4413.5514.0313.4613.44
LLaMA-1B11.6310.4210.4910.6210.2110.21
LLaMA-13B10.549.11n/a9.248.71n/a
MoE-7B-A1B11.939.85n/a10.619.66n/a
Validation perplexity, lower is better; the best in each group is in bold. Muon²-F is reported up to LLaMA-1B.

Not tied to one learning rate

On GPT-Large, Muon² is better at every learning rate in the sweep, with 3 steps and with 5.

3 Newton–Schulz steps
GPT-Large, validation perplexity
Muon²Muon
5 Newton–Schulz steps
GPT-Large, validation perplexity
Muon²Muon

Against other Muon variants

PolarExpress and Turbo-Muon change the polar approximation; NorMuon and AdaMuon add second-moment statistics to the update itself. With tuned hyper-parameters, none of them matches Muon², even when they run more NS steps than it does.

GPT-Small3 stepsGPT-Small5 stepsGPT-Base3 stepsGPT-Base5 steps
Muon²28.1227.9520.3919.96
Muon32.7029.5124.6921.47
PolarExpress30.0129.4222.7421.16
Turbo-Muon29.7029.6623.4621.93
NorMuon30.3528.4023.3321.27
AdaMuon31.2029.3026.0722.42
Validation perplexity on GPT-Small and GPT-Base.

Downstream

On eight tasks with the trained LLaMA-1B, Muon² averages 47.83 with 5 steps against 45.36 for Muon, and with 3 steps (45.89) it is still ahead of Muon with 5. At this scale ARC-Challenge, WinoGrande and MMLU sit close to random guessing when evaluated zero-shot, so every model is fine-tuned on those three tasks with the same protocol before evaluation. GPT-Large shows the same ordering.

Downstream accuracy
average over eight tasks, %
3 steps5 steps
LLaMA-1B
Muon
42.5645.36
Muon²
45.8947.83
GPT-Large
Muon
40.6942.07
Muon²
42.9043.70
4042444648
Show all eight tasks
ARC-c†ARC-eOBQAHellaSwagPIQAWinoGrande†LAMBADAMMLU†Avg
LLaMA-1B, Muon, 3 steps27.0842.8032.0044.3771.2256.7536.5229.7542.56
LLaMA-1B, Muon, 5 steps30.0146.0933.2049.7572.2559.9140.9930.7145.36
LLaMA-1B, Muon², 3 steps31.9147.8131.8050.2372.2059.8842.5030.7945.89
LLaMA-1B, Muon², 5 steps33.4849.0335.2052.6174.0561.9045.0831.3147.83
GPT-Large, Muon, 3 steps27.9944.2330.8039.9268.8254.5632.1927.0240.69
GPT-Large, Muon, 5 steps28.2144.0730.4042.5069.8059.1237.0725.4142.07
GPT-Large, Muon², 3 steps28.3044.9129.4043.7269.7060.7740.0326.3842.90
GPT-Large, Muon², 5 steps30.2646.9331.6045.0169.8060.3840.0725.5243.70

† fine-tuned on the task before evaluation, with the same protocol for every optimizer; the other tasks are zero-shot. On GPT-Large all four models stay at chance on MMLU even after fine-tuning.

Citation

BibTeX
@inproceedings{liu2026muon2,
  title     = {Muon$^2$: Boosting Muon via Adaptive Second-Moment Preconditioning},
  author    = {Liu, Ziyue and Zhang, Ruijie and Wang, Zhengyang and Zhao, Yequan and
               Su, Yupeng and Yang, Zi and Zhang, Zheng},
  booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
  year      = {2026}
}