Write a Fused Softmax Kernel in Triton in PyTorch
Expertgpu-systemsauto-graded
Write a GPU kernel in Triton that fuses the softmax computation into a single pass, eliminating intermediate memory reads.
Solve it
Check your answer
The grader verifies properties of your implementation, so a correct solution written differently from ours still passes.
pip install torchleet
from torchleet import check
check("triton-fused-softmax", online_softmax_pytorch)