Implement AdamW from Scratch in PyTorch
Hardauto-graded
Implement the AdamW optimizer from scratch on top of torch.optim.Optimizer: Adam's update with weight decay decoupled from the gradient, applied directly to the parameters.
Solve it
Check your answer
The grader verifies properties of your implementation, so a correct solution written differently from ours still passes.
pip install torchleet
from torchleet import check
check("adamw", MyAdamW)