Actor–Critic Methods
Combining policy gradients and value bootstrapping to achieve low-variance, sample-efficient reinforcement learning.
Actor-Critic methods combine Policy-Based (Actor) and Value-Based (Critic) reinforcement learning. The Actor parameterizes policy π_θ(a|s), selecting actions; the Critic parameterizes value function V_ϕ(s) or Q_ϕ(s,a), evaluating action quality. By replacing high-variance Monte Carlo trajectory returns G_t with the Advantage function A(s,a) = Q(s,a) - V(s) estimated via 1-step TD bootstrapping r + γ V(s') - V(s), Actor-Critic architectures achieve significantly lower gradient variance and higher sample efficiency.
Actor-Critic Architecture
ENVIRONMENT (State s_t)
│
┌───────────────────────┴───────────────────────┐
▼ ▼
[ ACTOR NETWORK π_θ(a|s) ] [ CRITIC NETWORK V_ϕ(s) ]
Generates Action Distribution Estimates State Value
│ │
▼ ▼
Action a_t ──► Environment Step ──► Reward r_{t+1}, Next State s_{t+1}
│
▼
[ ADVANTAGE ESTIMATOR ]
TD Error δ_t = r + γ V_ϕ(s') - V_ϕ(s)
│
┌───────────────────────────────────────────────┘
▼ (Updates BOTH networks!)
- Actor Update: θ ← θ + α_actor · ∇_θ ln π_θ(a_t|s_t) · δ_t
- Critic Update: ϕ ← ϕ - α_critic · ∇_ϕ ( δ_t )²
The Advantage Function
Using 1-step TD Bootstrapping, :
Notice that TD Error acts as an unbiased sample estimate of Advantage !
Advantages over Pure Methods
- Vs Pure Policy Gradient (REINFORCE): Substantially lower gradient variance due to TD bootstrapping baseline .
- Vs Pure Value-Based (DQN): Handles continuous action spaces naturally and learns stochastic policies.
Asynchronous Advantage Actor-Critic (A3C / A2C)
- A3C (Asynchronous): Multiple CPU worker threads run parallel environments independently, executing asynchronous gradient updates to a shared central model.
- A2C (Synchronous): Wait for all parallel environment workers to finish step , stacking mini-batches into a single GPU tensor for synchronous parallel forward/backward passes (faster on GPUs!).
Say this out loud
"Actor-Critic methods combine policy optimization and value estimation. The Actor π_θ(a|s) proposes actions; the Critic V_ϕ(s) evaluates state values to compute Advantage A(s,a) = r + γ V(s') - V(s). TD bootstrapping replaces noisy Monte Carlo returns with smooth value baselines, drastically reducing gradient variance while supporting continuous action spaces."
Follow-ups to expect
- What is Generalized Advantage Estimation (GAE)? A trade-off hyperparameter that smoothly interpolates between 1-step TD Advantage (, low variance, high bias) and full Monte Carlo Advantage (, high variance, zero bias): .
- How does A2C handle multi-task parallel environments? Uses parallel environment instances (e.g. 16 vectorized games), stepping all 16 environments simultaneously to generate a batch of 16 tuples per iteration for GPU matrix multiplication.
Check yourself
Question 1 of 3
What are the distinct responsibilities of the Actor and the Critic in an Actor-Critic architecture?