Markov Decision Processes
The formal mathematical framework for modeling sequential decision making under uncertainty.
A Markov Decision Process (MDP) is the mathematical foundation of Reinforcement Learning. An MDP is defined as a 5-tuple (S, A, P, R, γ): States S, Actions A, Transition Probability P(s'|s,a), Reward function R(s,a,s'), and Discount factor γ ∈ [0, 1). The core assumption is the Markov Property: future state s_{t+1} depends ONLY on current state s_t and action a_t, independent of past history.
The 5-Tuple Definition
REINFORCE ENVIRONMENT LOOP
┌──────────────────┐
Action │ │ State s_{t+1}
┌─────────┤ ENVIRONMENT ├─────────┐
│ a_t │ │ reward r_{t+1}
▼ └──────────────────┘ ▼
┌─────────┴────────┐ ┌────────┴─────────┐
│ AGENT │ │ AGENT │
│ Policy π(a|s) │ │ Value V(s) │
└──────────────────┘ └──────────────────┘
- State Space : Set of all valid environment states.
- Action Space : Set of all valid agent actions.
- Transition Dynamics : .
- Reward Function : Scalar feedback signal .
- Discount Factor : Weighs future rewards relative to immediate rewards.
Cumulative Discounted Return
The goal of the agent is to maximize expected cumulative discounted return :
- : Myopic agent (cares only about immediate next reward ).
- : Far-sighted agent (weighs long-term future rewards heavily).
Say this out loud
"An MDP is a 5-tuple (S, A, P, R, γ) defining sequential decision making under the Markov property: future state s_{t+1} depends strictly on current state s_t and action a_t. The agent optimizes policy π(a|s) to maximize expected cumulative discounted return G_t = ∑ γ^k R_{t+k+1}."
Follow-ups to expect
- What is a Partially Observable MDP (POMDP)? An MDP extension where the agent does not observe true state directly, but receives noisy observation , maintaining a probability belief state over hidden states.
- What is the difference between State-Value V(s) and Action-Value Q(s,a)? is expected return starting from state following policy . is expected return starting from state , taking explicit action , and thereafter following policy .
Check yourself
Question 1 of 3
What does the Markov Property state regarding state transitions in a Markov Decision Process?