MAP vs MLE
Frequentist likelihood maximization vs Bayesian posterior maximization with prior regularization.
Maximum Likelihood Estimation (MLE) maximizes data likelihood P(X | θ) ignoring prior beliefs. Maximum A Posteriori (MAP) incorporates a prior distribution P(θ), maximizing posterior P(θ | X) ∝ P(X | θ) P(θ). MAP acts as regularized MLE: assuming a zero-mean Gaussian prior P(θ) yields L2 Regularization (Ridge), while a Laplacian prior yields L1 Regularization (Lasso). As dataset size N → ∞, MAP converges to MLE.
The Derivation Bridge
Applying Bayes' Theorem to parameter estimation:
MLE Objective: arg max_θ [ Log-Likelihood: ln P(X | θ) ]
MAP Objective: arg max_θ [ Log-Likelihood: ln P(X | θ) + Log-Prior: ln P(θ) ]
│
Acts as Regularizer!
How MAP Derives L1 and L2 Regularization
Given Linear Regression with Gaussian noise :
1. Gaussian Prior L2 Regularization (Ridge)
Assume prior :
2. Laplace Prior L1 Regularization (Lasso)
Assume prior :
Comparison Matrix
| Property | MLE (Maximum Likelihood) | MAP (Maximum A Posteriori) |
|---|---|---|
| Formulation | ||
| Philosophical View | Frequentist (Parameters are fixed unknown constants) | Bayesian (Parameters have probability distributions) |
| Priors | No prior knowledge used | Incorporates explicit prior |
| Overfitting Risk | High on small datasets | Low (Prior acts as explicit regularizer) |
| Large Data Limit () | Identical | Converges to MLE (Data dominates prior) |
Say this out loud
"MLE maximizes data likelihood P(X|θ) without priors, making it prone to overfitting on small datasets. MAP incorporates a prior P(θ), maximizing log P(X|θ) + log P(θ). MAP bridges statistics and ML regularization: a Gaussian prior over weights derives L2 regularization (Ridge), while a Laplacian prior derives L1 regularization (Lasso). As sample size N grows, MAP converges to MLE."
Follow-ups to expect
- What is Full Bayesian Inference vs MAP? MAP outputs a single point estimate (the mode of the posterior). Full Bayesian inference computes the full posterior distribution , integrating over all parameters for predictions: .
- What is Conjugate Prior? A prior is conjugate to likelihood if the resulting posterior belongs to the same probability distribution family as the prior (e.g. Beta prior + Binomial likelihood = Beta posterior).
Check yourself
Question 1 of 3
What happens to MAP estimation when the prior P(θ) is chosen to be a uniform flat distribution (uninformative prior)?