Gaussian Mixtures & EM
Probabilistic soft-clustering using mixtures of Gaussians optimized via Expectation-Maximization.
Model Formulation
Data is generated from a mixture of Gaussian components:
- Mixing Coefficients (): .
- Component Means (): Center of Gaussian component .
- Covariance Matrices (): Shape, orientation, and spread of component .
k-Means Hard Clusters GMM Soft Elliptical Clusters
(Spherical Voronoi Cells) (Overlapping Probabilities)
┌──────┬──────┐ . . . . . .
│ * │ * │ ( * ) ( * ) γ_ik = 0.85
│ │ │ . . . . . .
└──────┴──────┘
The Expectation-Maximization (EM) Algorithm
EM finds local maximum of log-likelihood :
1. Expectation Step (E-step): Compute Responsibilities
Compute posterior probability that point belongs to component :
2. Maximization Step (M-step): Re-estimate Parameters
Update parameters using weighted responsibilities:
Repeat E and M steps until log-likelihood converges.
Covariance Matrix Constraints in Scikit-Learn
covariance_type='full': Each component has its own general covariance matrix (Elliptical, arbitrary orientation).covariance_type='tied': All components share the exact same covariance matrix.covariance_type='diag': Axis-aligned elliptical clusters ( is diagonal).covariance_type='spherical': Spherical clusters () — equivalent to soft k-Means!
Say this out loud
"GMM is a probabilistic clustering model representing data as a sum of K Gaussians. Unlike k-Means hard assignments, GMM outputs soft posterior probabilities P(k|x) and fits full covariance matrices Σ_k for elliptical clusters. Parameters are optimized via Expectation-Maximization: the E-step calculates component responsibilities γ_ik, and the M-step updates means, covariances, and mixing weights."
Follow-ups to expect
- Does EM guarantee global convergence? No. EM guarantees monotonic log-likelihood improvement at every step (), but can converge to local maxima depending on initialization (use k-Means++ to seed initial GMM means).
- What is Singular Covariance in GMM? If a Gaussian component collapses onto a single data point, its variance , causing likelihood to approach . Fix by adding small regularization value to diagonal of .
Check yourself
How does Gaussian Mixture Model (GMM) clustering differ fundamentally from k-Means clustering?