Implement Continuous Batching for LLM Inference in PyTorch
Hardllm-inferenceauto-graded
Dynamically slot sequences in and out of a running batch as they finish, maximizing throughput without padding waste.
Solve it
Check your answer
The grader verifies properties of your implementation, so a correct solution written differently from ours still passes.
pip install torchleet
from torchleet import check
check("continuous-batching", Request, ContinuousBatchScheduler)