GRPO & Group-Relative Methods
Eliminating separate Critic value networks in LLM alignment by computing group-relative advantage across sampled responses.
The Paradigm Shift: PPO vs GRPO
TRADITIONAL PPO RLHF:
Prompt x ──► Policy LLM ──► Response y ──► Reward Model r(x,y)
│
▼
Requires CRITIC MODEL V_ϕ(x) to compute Advantage: A(x,y) = r(x,y) - V_ϕ(x)
(Massive VRAM Overhead: Actor + Critic + Ref + Reward = 4 Models!)
GROUP RELATIVE POLICY OPTIMIZATION (GRPO):
Prompt x ──► Policy LLM ──► Samples Group of G = 4 Responses: {y_1, y_2, y_3, y_4}
│
▼
Scalar Rewards: {r_1, r_2, r_3, r_4}
Group Mean: μ_r = mean(r)
Group Std: σ_r = std(r)
Group Advantage: A_i = (r_i - μ_r) / (σ_r + ε) ──► NO CRITIC MODEL NEEDED!
(50% Lower VRAM: Actor + Ref + Reward = 3 Models!)
The GRPO Mathematical Objective
For prompt and group of sampled outputs :
Group Advantage :
Memory & VRAM Breakdown
| Resource Component | Standard PPO RLHF | GRPO (Group Relative) |
|---|---|---|
| Policy Network () | Active (Trainable) | Active (Trainable) |
| Critic Network () | Active (Trainable - 100% Params) | ELIMINATED (0 Params!) |
| Reference Model () | Frozen | Frozen |
| Reward Model () | Frozen | Frozen |
| VRAM Footprint Reduction | Baseline | ~50% Lower Memory Overhead |
Why GRPO Excels in Reasoning Models (DeepSeek-R1)
For math and coding tasks, binary rule-based verifiers provide exact execution rewards ( if code passes unit tests; if it fails).
Sampling candidate reasoning chains per prompt creates a natural intra-prompt tournament:
- Responses that solve the math problem receive positive normalized Advantage .
- Responses that fail receive negative normalized Advantage .
Grades update the Policy LLM without needing any neural value estimator!
Say this out loud
"GRPO eliminates the memory-heavy Critic Value Model in RLHF. For each prompt, GRPO samples a group of G responses, computing baseline value dynamically as group average reward. Advantage for response i is normalized against group mean and std: A_i = (r_i - μ_r) / σ_r. DeepSeek-R1 used GRPO to cut VRAM memory by 50% while scaling reasoning performance."
Follow-ups to expect
- What is the optimal Group Size G in GRPO? DeepSeek-Math empirically evaluated group sizes , finding provides optimal variance reduction without inflating generation rollout time.
- How does GRPO handle KL Divergence? Instead of adding KL penalty directly into reward , GRPO adds an explicit KL divergence loss term directly to the main loss function.
Check yourself
How does GRPO (Group Relative Policy Optimization) eliminate the need for training a separate Critic Value Model V_ϕ(s) in LLM RLHF?