SARSA vs Q-Learning (On vs Off Policy)
Comparing On-Policy temporal difference control (SARSA) against Off-Policy greedy control (Q-Learning).
SARSA and Q-Learning are two foundational Temporal Difference (TD) Reinforcement Learning control algorithms. SARSA is On-Policy: its name derives from tuple (S_t, A_t, R_{t+1}, S_{t+1}, A_{t+1}), updating Q(s,a) using the actual next action a_{t+1} selected by the exploratory behavior policy π. Q-Learning is Off-Policy: it updates Q(s,a) using the maximum greedy action max_{a'} Q(s',a') ignoring the actual exploratory action taken. SARSA learns safer policies when exploration risks (e.g. falling off a cliff) carry high negative penalties.
The Core Mathematical Difference
SARSA (On-Policy TD Control):
Q(s_t, a_t) ← Q(s_t, a_t) + α [ R_{t+1} + γ Q(s_{t+1}, A_{t+1}) - Q(s_t, a_t) ]
▲
USES ACTUAL EXECUTED A_{t+1}! (Includes ε exploration risk)
Q-LEARNING (Off-Policy TD Control):
Q(s_t, a_t) ← Q(s_t, a_t) + α [ R_{t+1} + γ max_{a'} Q(s_{t+1}, a') - Q(s_t, a_t) ]
▲
USES GREEDY OPTIMAL max_{a'}! (Assumes zero future errors)
The Cliff Walking Experiment (Sutton & Barto)
Imagine a grid world where walking along the edge of a cliff yields high immediate reward, but stepping off the cliff incurs a catastrophic penalty ( reward):
Start [S] ──► [ . ][ . ][ . ][ . ][ . ][ . ][ . ][ . ][ . ] ──► Goal [G]
══════════════ CLIFF (-100) ═════════════════
- Q-Learning Path: Learns the mathematically optimal shortest path directly along the cliff edge. However, during execution with exploration, occasional random steps fall off the cliff!
- SARSA Path: Factors the 10% random exploration chance into , learning a safer, longer path 2 rows away from the cliff, maximizing total real-world expected reward during exploration.
Comparison Summary
| Dimension | SARSA | Q-Learning |
|---|---|---|
| Policy Type | On-Policy | Off-Policy |
| Update Target | ||
| Exploration Risk | Accounts for exploratory action risks | Ignores exploration risk in target calculation |
| Learned Path | Conservative, safe path under active | Mathematically optimal shortest path |
| Replay Buffers | Hard to use (requires current policy samples) | Easy to use (Stores arbitrary historical transitions) |
Say this out loud
"SARSA is On-Policy TD control updating Q(s,a) using the actual executed next action A_t+1: Q(s,a) ← Q(s,a) + α [R + γ Q(s', A') - Q(s,a)]. Q-Learning is Off-Policy updating with max_a' Q(s', a'). Because SARSA incorporates exploratory action risks into its TD target, it learns safer conservative paths in environments with high penalty risks."
Follow-ups to expect
- What is Expected SARSA? Replaces the single sample action with the expected action-value under the policy: . Combines the stability of SARSA with reduced variance.
- Can SARSA use Replay Buffers? Standard SARSA cannot easily use Replay Buffers because old transitions generated by past policies violate the on-policy assumption. Expected SARSA or Importance Sampling is required.
Check yourself
Question 1 of 3
What tuple of experience components gives SARSA its name?