Value-Based vs Policy-Based Methods
Comparing the two core families of Reinforcement Learning algorithms: learning action values vs directly optimizing policy parameters.
Comparison Matrix
VALUE-BASED (DQN) POLICY-BASED (REINFORCE) ACTOR-CRITIC (PPO)
Learns Q(s, a) Learns π_θ(a|s) Learns BOTH!
Policy: a = argmax_a Q(s, a) Optimizes θ via ∇_θ J(θ) Actor π_θ(a|s) + Critic V_ϕ(s)
| Property | Value-Based (DQN, SARSA) | Policy-Based (REINFORCE) | Actor-Critic (PPO, SAC, A2C) |
|---|---|---|---|
| Core Target Learned | Action-Value | Policy Distribution | Both Policy & Value |
| Action Space | Discrete (Buttons, Moves) | Continuous & Discrete | Continuous & Discrete |
| Policy Type | Deterministic (Greedy ) | Stochastic / Continuous Gaussian | Stochastic / Continuous |
| Sample Efficiency | High (Off-policy Replay Buffer) | Low (On-policy Monte Carlo) | Moderate to High |
| Variance | Low (Bootstrapping) | High (Monte Carlo Return ) | Low (Advantage ) |
Mathematical Objectives
1. Value-Based Loss (DQN)
Minimizes Bellman Mean Squared Error over replay buffer :
2. Policy-Based Objective (Policy Gradient Theorem)
Maximizes expected trajectory return via gradient ascent:
3. Actor-Critic Objective
Replaces Monte Carlo return with Advantage Function :
Say this out loud
"Value-based methods (DQN) learn action values Q(s,a) for discrete actions, deriving greedy policies via argmax_a Q(s,a). Policy-based methods (REINFORCE) parameterize policy π_θ(a|s) directly using gradient ascent, handling continuous action spaces naturally. Actor-Critic methods (PPO) combine both: the Actor learns policy π_θ and the Critic learns value baseline V_ϕ(s) to reduce gradient variance."
Follow-ups to expect
- What is On-Policy vs Off-Policy learning? On-policy methods (PPO, REINFORCE) evaluate and update the exact policy currently generating environment samples. Off-policy methods (DQN, SAC) evaluate one policy while using samples generated by an older/different exploration policy stored in a Replay Buffer.
- What is Soft Actor-Critic (SAC)? Off-policy maximum-entropy actor-critic algorithm that adds entropy bonus to the reward objective, encouraging exploration in continuous control tasks.
Check yourself
Why do Value-Based methods (Q-Learning / DQN) struggle with continuous action spaces (e.g. steering wheel angle from -180° to +180°)?