Implement KV Cache for Autoregressive Generation in PyTorch

Mediumllm-inferenceauto-graded

Cache key and value tensors from previous timesteps so each new token only computes attention over one new position.

Solve it

Check your answer

The grader verifies properties of your implementation, so a correct solution written differently from ours still passes.

pip install torchleet

from torchleet import check
check("kv-cache", KVCache, CachedAttention)

Company tags

How these tags are sourced