Maximum Likelihood Estimation
Finding the parameter values that maximize the probability of observing your collected sample data.
Maximum Likelihood Estimation (MLE) estimates parameters θ of a probability distribution by maximizing the Likelihood function L(θ) = P(X | θ). For i.i.d. data, we maximize Log-Likelihood log L(θ) = ∑ log P(x_i | θ) because log turns products into computationally stable sums. MLE forms the statistical foundation for parameter fitting in OLS Linear Regression (under Gaussian noise), Logistic Regression, and Neural Networks (via Cross-Entropy).
The Likelihood Function
Given i.i.d. observations from distribution :
To find , solve score equation .
MLE Equivalence in ML Models
- Linear Regression (Gaussian Noise): (MSE Loss).
- Logistic Regression (Bernoulli Noise): (Binary Cross-Entropy).
- Multi-Class Networks (Multinomial Noise): (Categorical Cross-Entropy).
Asymptotic Properties of MLE
- Consistency: As sample size , MLE (converges to true parameter).
- Asymptotic Normality: As , , where is the Fisher Information Matrix.
- Efficiency: Achieves the Cramér-Rao lower bound for variance among unbiased estimators.
Say this out loud
"Maximum Likelihood Estimation finds parameter values that maximize the probability of observing the data. We maximize Log-Likelihood because log converts probability products into stable additions. Under Gaussian noise assumptions, maximizing MLE is mathematically identical to minimizing MSE loss in linear regression, while under Bernoulli assumptions it corresponds to Cross-Entropy."
Follow-ups to expect
- What is the difference between Likelihood and Probability? Probability evaluates the chance of data for fixed parameter . Likelihood evaluates the plausibility of parameter for observed fixed data .
- How does MLE handle small datasets? MLE overfits small datasets because it relies entirely on observed sample frequencies without incorporating prior domain knowledge (MAP/Bayesian estimation solves this).
Check yourself
Question 1 of 3
Why do we maximize Log-Likelihood log L(θ) instead of raw Likelihood L(θ) = ∏ P(x_i | θ) in practice?