Policy Gradients & REINFORCE
Directly optimizing policy parameters θ via gradient ascent on expected trajectory returns.
The Policy Objective Function
We parameterize a policy using neural network weights .
The goal is to maximize total expected discounted return over trajectories :
Where trajectory probability .
The Log-Derivative Trick & Policy Gradient Theorem
Taking the gradient with respect to :
Applying Log-Derivative Trick ():
Notice that .
Taking derivatives eliminates unknown environment terms completely!
Intuition of Policy Gradient Step
θ ← θ + α · ∇_θ ln π_θ(a_t | s_t) · G_t
If Trajectory Return G_t > 0 ──► INCREASE probability of taking action a_t in state s_t!
If Trajectory Return G_t < 0 ──► DECREASE probability of taking action a_t in state s_t!
The REINFORCE Algorithm (Monte Carlo Policy Gradient)
# REINFORCE Update Loop
for episode in range(num_episodes):
trajectory = generate_trajectory(policy_net) # Collect full episode
for t in range(len(trajectory)):
G_t = sum(gamma**k * r for k, r in enumerate(rewards[t:])) # Monte Carlo Return
loss = -log_prob(actions[t]) * G_t # Policy Gradient Loss
loss.backward()
optimizer.step()
Say this out loud
"Policy Gradient methods optimize policy parameters θ directly via gradient ascent on expected return J(θ). The Policy Gradient Theorem uses the log-derivative trick to eliminate unknown environment transition dynamics, deriving ∇_θ J(θ) = E [ ∇_θ ln π_θ(a|s) · G_t ]. REINFORCE uses Monte Carlo episode returns G_t as unbiased estimators, but suffers from high variance."
Follow-ups to expect
- How do you reduce variance in REINFORCE? Subtract a state-dependent Baseline from return : . Setting yields the Advantage Function , significantly lowering variance without introducing bias.
- What is Continuous Action Policy Parametrization? For continuous actions , the policy network outputs Mean and Standard Deviation of a Gaussian distribution , sampling actions via reparameterization.
Check yourself
Why is the Policy Gradient Theorem mathematically remarkable regarding environment transition dynamics P(s'|s,a)?