Write a Transformer from Scratch in PyTorch

Hardauto-graded

Build the full transformer architecture — multi-head attention, positional encoding, encoder-decoder stacks — from raw tensors.

Solve it

Check your answer

The grader verifies properties of your implementation, so a correct solution written differently from ours still passes.

pip install torchleet

from torchleet import check
check("transformer", PositionalEncoding, MultiHeadSelfAttention, FeedForward, TransformerEncoderLayer, TransformerModel)