Implement KV Cache for Autoregressive Generation in PyTorch
Mediumllm-inferenceauto-graded
Cache key and value tensors from previous timesteps so each new token only computes attention over one new position.
Solve it
Check your answer
The grader verifies properties of your implementation, so a correct solution written differently from ours still passes.
pip install torchleet
from torchleet import check
check("kv-cache", KVCache, CachedAttention)