Implement Continuous Batching for LLM Inference in PyTorch

Hardllm-inferenceauto-graded

Dynamically slot sequences in and out of a running batch as they finish, maximizing throughput without padding waste.

Solve it

Check your answer

The grader verifies properties of your implementation, so a correct solution written differently from ours still passes.

pip install torchleet

from torchleet import check
check("continuous-batching", Request, ContinuousBatchScheduler)

Company tags

How these tags are sourced