Implement Mixture of Experts Layer in PyTorch
Hardmodern-architecturesauto-graded
Build a gated MoE with top-k routing, load balancing loss, and expert capacity constraints for sparse computation.
Solve it
Check your answer
The grader verifies properties of your implementation, so a correct solution written differently from ours still passes.
pip install torchleet
from torchleet import check
check("mixture-of-experts", Expert, MoELayer)