Multicollinearity & VIF
Detecting and resolving linear dependence among predictor features using Variance Inflation Factor.
What is Multicollinearity?
Multicollinearity occurs when two or more input features in a linear regression model are highly correlated with one another.
Example
Predicting Home Price () using two features:
- : Square Footage in Feet.
- : Square Footage in Meters.
Because and carry identical information (), the model cannot determine whether or is driving home prices!
Matrix Inversion in OLS: beta = (X^T X)^-1 X^T Y
When features are collinear ──► X^T X is NEAR SINGULAR (Determinant ~ 0!)
──► Inverting (X^T X) produces HUGE INFLATED VARIANCE on coefficients!
Symptoms of Multicollinearity
- Inflated Standard Errors: Regression coefficients have huge standard errors, producing wide confidence intervals and non-significant p-values () despite a high overall score!
- Unstable Coefficient Weights: Adding or removing a few data rows causes coefficient values to swing wildly or flip signs (e.g. positive correlation feature getting a negative ).
- Loss of Interpretability: You cannot isolate the individual effect of feature holding constant.
Note: Multicollinearity does NOT reduce the overall predictive accuracy or of the model on the training data! It destroys feature interpretability and statistical inference.
Diagnosing Multicollinearity: Variance Inflation Factor (VIF)
The Variance Inflation Factor (VIF) measures how much the variance of coefficient is inflated due to collinearity with other features:
- : score obtained by regressing feature against all other remaining predictor features.
┌──────────────────────────┬──────────────────────────┐
│ VIF VALUE │ INTERPRETATION │
├──────────────────────────┼──────────────────────────┤
│ VIF = 1.0 │ Zero Multicollinearity. │
│ 1.0 < VIF < 5.0 │ Moderate Correlation. │
│ VIF > 5.0 or 10.0 │ SEVERE MULTICOLLINEARITY!│
└──────────────────────────┴──────────────────────────┘
How to Fix Multicollinearity
┌──────────────────────────┬──────────────────────────┬──────────────────────────┐
│ 1. DROP FEATURES │ 2. L2 RIDGE REGRESSION │ 3. PCA REDUCTION │
├──────────────────────────┼──────────────────────────┼──────────────────────────┤
│ Remove one of the highly │ Adds λI to X^T X matrix, │ Transforms correlated │
│ correlated feature pairs │ making matrix inversion │ features into orthogonal │
│ (e.g. drop SqMeters). │ numerically stable. │ principal components. │
└──────────────────────────┴──────────────────────────┴──────────────────────────┘
Say this out loud
Multicollinearity occurs when predictor features are highly correlated. It inflates coefficient standard errors and flips parameter signs without reducing overall R squared. Diagnosed using Variance Inflation Factor (VIF > 5), multicollinearity is resolved by dropping redundant features, performing PCA, or applying L2 Ridge Regularization.
Followups to expect
- Does Multicollinearity impact Decision Trees or Random Forests? No! Tree based algorithms select single features sequentially at orthogonal split nodes, making them immune to multicollinearity during prediction.
- What is Structural Multicollinearity vs Data Multicollinearity? Structural multicollinearity is created by adding polynomial features ( and ). Fix by centering variables () before squaring.
Check yourself
What primary numerical artifact indicates severe Multicollinearity in linear regression outputs?