Implement Muon from Scratch in PyTorch
Hardauto-graded
Implement the Muon optimizer from scratch on top of torch.optim.Optimizer: momentum orthogonalized by Newton–Schulz for matrix parameters, with an AdamW fallback for 1D parameters.
Solve it
Check your answer
The grader verifies properties of your implementation, so a correct solution written differently from ours still passes.
pip install torchleet
from torchleet import check
check("muon", newton_schulz, MyMuon)