LDA for Dimensionality Reduction
Projecting supervised data into lower dimensions while maximizing between class separation and minimizing within class variance.
PCA vs LDA: Unsupervised vs Supervised
Both PCA and LDA project data into a lower-dimensional subspace.
However, their underlying goals are completely different:
PCA (Unsupervised - Ignores Class Labels): LDA (Supervised - Uses Class Labels):
Finds directions of MAXIMUM TOTAL VARIANCE. Finds directions of MAXIMUM CLASS SEPARATION.
▲ ▲
│ (o) (o) │ (o) (o) Class 1
│ (o) (x) │ ───────► MAXIMUM CLASS SEPARATION!
│ (x) (x) │ (x) (x) Class 2
───┴─────────────────► Feature X1 └───┴─────────────────► Feature X1
If you apply PCA to data where class separation lies along a low-variance direction, PCA will accidentally wipe out your class signals!
LDA uses target labels to preserve classification signals.
Fisher's Linear Discriminant Criterion
Fisher (1936) defined the optimal projection direction as the vector that maximizes between-class variance while minimizing within-class variance:
- Between-Class Scatter Matrix (): Measures the distance between class mean vectors ().
- Within-Class Scatter Matrix (): Measures the variance spread of data points within each individual class.
Maximizing forces data points in the same class to cluster tightly together () while pushing class means far apart ().
Maximizing J(w) = ( Between-Class Distance )^2 / ( Within-Class Variance Spread )
Dimensionality Constraint: At Most Components
Because LDA relies on class means to construct the Between-Class Scatter matrix :
For a classification problem with classes, has rank at most .
Therefore, LDA can project data into at most dimensions:
- Binary Classification (): Reduces data to 1 single dimension.
- 10-Class Problem (): Reduces data to at most 9 dimensions.
Assumptions of LDA
- Gaussian Distribution: Each class features are normally distributed.
- Homoscedasticity: All classes share the same covariance matrix . If classes have significantly different covariance matrices, use Quadratic Discriminant Analysis (QDA) instead.
Say this out loud
LDA is a supervised dimensionality reduction algorithm that projects data onto axes maximizing between-class variance while minimizing within-class variance using Fisher's criterion. Unlike unsupervised PCA, LDA uses target labels to preserve class separability. For C classes, LDA can extract at most C minus 1 linear discriminant components.
Followups to expect
- What is Quadratic Discriminant Analysis (QDA)? A non linear extension of LDA that drops the assumption of shared covariance matrices, allowing each class to estimate its own covariance matrix , producing quadratic decision boundaries.
- What happens if feature count d is larger than sample count N (Small Sample Size problem)? Within-class scatter matrix becomes singular and non-invertible. Apply PCA first to reduce before running LDA.
Check yourself
What is the primary difference between PCA and LDA for dimensionality reduction?