Proximal Policy Optimization
Clipping probability ratio updates to achieve stable, sample-efficient policy gradient updates in RL and LLM alignment.
Why PPO Was Created: Trust Region Motivation
In supervised learning, if a gradient step is too large, the next mini-batch corrects it.
In Reinforcement Learning, taking an overly aggressive policy step ruins data collection: the degraded policy generates bad trajectory data, causing irrecoverable policy collapse.
- TRPO (Trust Region Policy Optimization): Enforces strict KL constraint using complex Second-Order Conjugate Gradient optimization.
- PPO (Proximal Policy Optimization): Achieves the stability of TRPO using First-Order SGD with a Clipped Loss Objective!
The Clipped Surrogate Objective
Let the Probability Ratio be:
Standard hyperparameter setting: (Clips ratio ).
Positive Advantage (A_t > 0) Negative Advantage (A_t < 0)
L_CLIP L_CLIP
│ │
1+ε├─────────────── Upper Bound │
│ / │ /
│ / 1-ε├──────────────/─ Lower Bound
│ / │ /
1.0├───────────/ 1.0├────────────/
└───────────┴─────────────► r_t └────────────┴──────────────► r_t
1.0 1+ε 1-ε 1.0
- Positive Advantage (): Wants to increase . Clipped at so policy cannot gain extra reward by pushing .
- Negative Advantage (): Wants to decrease . Clipped at so policy cannot gain extra reward by pushing .
Complete PPO Combined Loss
- : Clipped policy surrogate loss.
- : Critic Squared Error Value Loss.
- : Entropy Bonus encouraging exploration.
Say this out loud
"PPO stabilizes policy gradient updates using a Clipped Surrogate Objective L_CLIP(θ) = E [ min( r_t A_t, clip(r_t, 1-ε, 1+ε) A_t ) ]. By capping probability ratio r_t = π_θ / π_old within [1-ε, 1+ε] (typically [0.8, 1.2]), PPO prevents destructive large policy steps, providing smooth, monotonic policy improvements in RL and LLM RLHF."
Follow-ups to expect
- How is PPO used in LLM RLHF alignment? The Policy LLM generates responses for prompts . PPO updates LLM parameters using Advantage .
- What is PPO-Penalty vs PPO-Clip? PPO-Penalty enforces KL divergence as an adaptive loss penalty term . PPO-Clip uses hard ratio clipping. PPO-Clip is almost universally preferred due to superior empirical stability.
Check yourself
Why is standard un-clipped Policy Gradient (REINFORCE) dangerously unstable during deep neural network training?