Bias, Variance, Overfitting and Validation
This is the most-asked theory topic in ML interviews, because every modelling decision is a position on it. You should be able to explain the decomposition, show how to diagnose it from learning curves, name the levers for each side, and run a clean validation procedure.
1. The decomposition
Suppose the data are generated as with noise of variance . Train the same algorithm on many different training sets. For squared error, the expected prediction error at a point splits into three parts:
- Bias is systematic error: even averaged over all possible training sets, the model is wrong because it cannot represent the truth (a line fit to a curve).
- Variance is sensitivity: the fitted function changes a lot when the training sample changes (a high-degree polynomial chasing noise).
- Noise is the floor you cannot beat.
Simple models have high bias and low variance. Flexible models have low bias and high variance. Total error is U-shaped in complexity, and you aim for the bottom.
<!--fig:biasvar-->2. Measuring it
You can estimate bias and variance directly by simulation: draw many training sets, fit a model on each, and look at the predictions at a fixed point.
import numpy as np
rng = np.random.default_rng(1)
f = lambda x: np.sin(2 * np.pi * x)
x_test = 0.3
true_val = f(x_test)
def fit_predict(degree, n=20, sigma=0.3):
x = rng.uniform(0, 1, n)
y = f(x) + rng.normal(0, sigma, n)
coefs = np.polyfit(x, y, degree)
return np.polyval(coefs, x_test)
results = {}
for degree in (1, 3, 9):
preds = np.array([fit_predict(degree) for _ in range(2000)])
bias2 = (preds.mean() - true_val) ** 2
var = preds.var()
results[degree] = (bias2, var)
assert results[1][0] > results[3][0] # a line cannot follow a sine wave: high bias
assert results[9][1] > results[3][1] # a degree-9 polynomial on 20 points: high variance
assert results[3][0] + results[3][1] < results[1][0] + results[1][1]
assert results[3][0] + results[3][1] < results[9][0] + results[9][1]
Degree 3 wins because it balances the two. This experiment is a good answer to "show me bias and variance concretely".
3. Overfitting and underfitting
| Training error | Validation error | Gap | |
|---|---|---|---|
| Underfitting (high bias) | high | high | small |
| Good fit | low | slightly higher | small |
| Overfitting (high variance) | very low | high | large |
Learning curves
Plot training and validation error as the training set grows.
- If both curves converge to a high error, you have high bias. More data will not help. Use a richer model or better features.
- If there is a large gap that narrows as data grows, you have high variance. More data helps, as do regularisation and simpler models.
This is the most useful diagnostic in practice, and it tells you where to spend effort.
4. Levers
| To reduce variance | To reduce bias |
|---|---|
| More training data | A more flexible model (more depth, more features) |
| Regularisation (L1, L2, dropout, weight decay) | Better features or interactions |
| Simpler model, fewer features | Less regularisation |
| Early stopping | Train longer, higher capacity |
| Bagging and ensembling (random forests) | Boosting (reduces bias sequentially) |
| Data augmentation | Reduce noise in the labels or features |
Interviewers like the contrast: bagging reduces variance by averaging independent-ish models; boosting reduces bias by fitting residuals sequentially.
5. Modern wrinkle: double descent
Classical theory says test error rises forever past the interpolation point. Very large neural networks show double descent: test error rises near the point where the model can just barely fit the training data exactly, then falls again as capacity grows further, helped by the implicit regularisation of gradient descent. Mention this to show awareness, but do not use it as a reason to ignore regularisation on small tabular problems.
6. Validation procedures
Hold-out
One split, such as 70/15/15. Fast, but noisy on small data.
k-fold cross-validation
Split into folds (often 5 or 10). Train on , validate on the remaining one, rotate, and average. It uses all data for both roles and gives an estimate of the spread across folds.
import numpy as np
def kfold_indices(n, k, seed=0):
idx = np.random.default_rng(seed).permutation(n)
return np.array_split(idx, k)
def cv_mse(x, y, degree, k=5):
folds = kfold_indices(len(x), k)
errs = []
for i in range(k):
val = folds[i]
train = np.concatenate([folds[j] for j in range(k) if j != i])
coefs = np.polyfit(x[train], y[train], degree)
errs.append(np.mean((np.polyval(coefs, x[val]) - y[val]) ** 2))
return np.mean(errs)
rng = np.random.default_rng(2)
x = rng.uniform(0, 1, 60)
y = np.sin(2 * np.pi * x) + rng.normal(0, 0.2, 60)
scores = {d: cv_mse(x, y, d) for d in (1, 3, 5, 12)}
best = min(scores, key=scores.get)
assert best in (3, 5) # the model neither too simple nor too flexible wins
assert scores[1] > scores[best] and scores[12] > scores[best]
Variants worth naming
- Stratified k-fold: preserve class proportions in each fold (essential for imbalanced classes).
- Group k-fold: keep all rows of one group (user, patient, document) in the same fold.
- Time-series split: train on the past, validate on the next block, expanding the window.
- Leave-one-out: . Low bias, high compute and high variance of the estimate.
- Nested cross-validation: an outer loop for estimating performance and an inner loop for tuning, so tuning does not inflate the estimate.
7. Hyperparameter search
| Method | Idea | Note |
|---|---|---|
| Grid search | Try every combination | Wasteful in high dimensions |
| Random search | Sample combinations | Often better than a grid with the same budget, because only a few hyperparameters matter |
| Bayesian optimisation | Model the score surface and pick promising points | Good for expensive training |
| Successive halving / Hyperband | Drop poor configurations early | Saves compute |
Always tune on validation (or cross-validation), then evaluate the chosen configuration once on the test set.
8. Data leakage: the silent overfitter
Leakage is information in training that will not be available at prediction time, or information from the validation set that leaks into training. It produces beautiful validation scores and a failed deployment.
- Fitting a scaler, imputer or encoder on the whole dataset before splitting.
- Using a feature computed from the label or from the future (for example "number of refunds" when predicting churn that already happened).
- Duplicates or near-duplicates across train and test.
- Tuning repeatedly against the test set.
The rule: everything learned from data (including preprocessing) is fitted on the training fold only, then applied to the others.
import numpy as np
rng = np.random.default_rng(3)
X = rng.normal(5, 2, size=(100, 1))
train, test = X[:80], X[80:]
mu, sd = train.mean(), train.std() # fit the scaler on training data only
test_scaled = (test - mu) / sd # apply to test with the training statistics
leaky_mu = X.mean() # a leaky scaler would use all rows
assert mu != leaky_mu # the statistics differ, which is exactly the leak
9. Common mistakes
- Reporting training error or a score on data used for tuning.
- Adding data when the problem is bias, or adding complexity when the problem is variance.
- Random splits on temporal or grouped data.
- Treating cross-validation scores as exact. Report the mean and the spread.
- Preprocessing before splitting.
- Forgetting the noise floor: if labels are noisy, a large train/validation gap may be unavoidable.
10. Practice questions
- Derive the bias-variance decomposition for squared error (sketch the steps).
- Your training accuracy is 99% and validation is 82%. List four things you would try, in order.
- Training and validation errors are both high and close together. What now?
- How do bagging and boosting each interact with bias and variance?
- Why is k-fold cross-validation better than a single split on small data? When is it unsuitable?
- How would you validate a model that predicts next month's demand?
- Describe three forms of data leakage you have seen or could see.