Write a Fused Softmax Kernel in Triton in PyTorch

Expertgpu-systemsauto-graded

Write a GPU kernel in Triton that fuses the softmax computation into a single pass, eliminating intermediate memory reads.

Solve it

Check your answer

The grader verifies properties of your implementation, so a correct solution written differently from ours still passes.

pip install torchleet

from torchleet import check
check("triton-fused-softmax", online_softmax_pytorch)

Company tags

How these tags are sourced