Implement PPO for RLHF in PyTorch

Hardalignment-trainingauto-graded

Build Proximal Policy Optimization with clipped surrogate objective, value function baseline, and KL penalty for RLHF.

Solve it

Check your answer

The grader verifies properties of your implementation, so a correct solution written differently from ours still passes.

pip install torchleet

from torchleet import check
check("ppo-rlhf", generate_sequences, compute_gae, ppo_step)

Company tags

How these tags are sourced