Implement Vision Transformer + MAE Pretraining in PyTorch

Hardmodern-architecturesauto-graded

Build ViT with masked autoencoder pretraining — randomly mask patches, encode visible ones, decode to reconstruct.

Solve it

Check your answer

The grader verifies properties of your implementation, so a correct solution written differently from ours still passes.

pip install torchleet

from torchleet import check
check("vit-mae", PatchEmbedding, ViT, MAE)

Company tags

How these tags are sourced