Implement Muon from Scratch in PyTorch

Hardauto-graded

Implement the Muon optimizer from scratch on top of torch.optim.Optimizer: momentum orthogonalized by Newton–Schulz for matrix parameters, with an AdamW fallback for 1D parameters.

Solve it

Check your answer

The grader verifies properties of your implementation, so a correct solution written differently from ours still passes.

pip install torchleet

from torchleet import check
check("muon", newton_schulz, MyMuon)

Company tags

How these tags are sourced