Linear and Logistic Regression
Linear and logistic regression are the most common warm-up and the most common baseline. They are simple enough to derive in an interview, and a surprising amount of deeper material (regularisation, loss functions, neural networks) builds directly on them. Know the derivations, the assumptions, the interpretation of coefficients, and how to implement both from scratch.
1. Linear regression
Model: . Fit by minimising mean squared error (MSE):
(absorb the bias by adding a column of ones to ).
Closed form: the normal equation
Set the gradient to zero:
This needs to be invertible, meaning no perfectly collinear features. Cost is roughly , fine for thousands of features and impractical for millions, where gradient descent is used. In numerical code you use lstsq or a QR/SVD decomposition rather than an explicit inverse, for stability.
import numpy as np
rng = np.random.default_rng(0)
n, d = 500, 3
X = rng.normal(size=(n, d))
true_w = np.array([2.0, -1.0, 0.5])
y = X @ true_w + 4.0 + rng.normal(0, 0.1, n)
Xb = np.column_stack([X, np.ones(n)])
w_normal = np.linalg.inv(Xb.T @ Xb) @ Xb.T @ y
w_lstsq = np.linalg.lstsq(Xb, y, rcond=None)[0]
assert np.allclose(w_normal, w_lstsq, atol=1e-8)
assert np.allclose(w_normal[:3], true_w, atol=0.05) and abs(w_normal[3] - 4.0) < 0.05
Why squared error? A probabilistic view
If with , the maximum-likelihood estimate of minimises the sum of squared errors. So MSE is not arbitrary: it is the right loss when noise is Gaussian. With heavy-tailed noise or outliers, absolute error or Huber loss is more robust.
Assumptions (and what breaks)
| Assumption | If violated |
|---|---|
| Linear relationship | Biased predictions. Add transformations or interactions. |
| Independent errors | Standard errors are wrong (time series, grouped data). |
| Constant error variance (homoscedasticity) | Inference unreliable; consider weighting or transforming . |
| No severe multicollinearity | Coefficients become unstable and hard to interpret, though predictions can stay fine. |
| (For inference) roughly normal errors | Confidence intervals off in small samples. |
For prediction you mainly need the model to fit well on held-out data. For inference (what is the effect of a feature?) the assumptions matter much more.
Interpreting coefficients
A coefficient is the expected change in for a one-unit change in that feature holding the others fixed. Standardising features lets you compare magnitudes. Beware reading causal meaning into correlational data, and beware correlated features that share credit arbitrarily.
is the fraction of variance explained relative to predicting the mean. It never decreases when you add features, so use adjusted or validation error to compare models.
2. Gradient descent for linear regression
When or is large, iterate: . Three variants:
- Batch: use all data each step. Stable, slow per step.
- Stochastic (SGD): one example per step. Noisy, cheap, can escape shallow traps.
- Mini-batch: a small batch (32 to 1024). The practical default.
Scale features first. With very different scales, the loss surface is a narrow valley and gradient descent zig-zags or diverges. Standardisation to zero mean and unit variance fixes this.
3. Regularised variants
| Model | Penalty | Effect |
|---|---|---|
| Ridge (L2) | Shrinks coefficients smoothly; handles collinearity; closed form | |
| Lasso (L1) | Drives some coefficients exactly to zero: feature selection | |
| Elastic net | both | Sparsity with stability under correlated features |
Ridge as a closed form shows why it helps: adding makes the matrix invertible even when features are collinear. Geometrically, the L1 constraint region has corners on the axes, so the optimum often lands where some weights are zero. From a Bayesian view, L2 is a Gaussian prior on weights and L1 a Laplace prior.
import numpy as np
rng = np.random.default_rng(1)
n = 50
x1 = rng.normal(size=n)
x2 = x1 + rng.normal(0, 1e-3, n) # almost perfectly collinear with x1
X = np.column_stack([x1, x2])
y = 3 * x1 + rng.normal(0, 0.1, n)
ols = np.linalg.lstsq(X, y, rcond=None)[0]
lam = 1.0
ridge = np.linalg.solve(X.T @ X + lam * np.eye(2), X.T @ y)
assert np.abs(ols).max() > np.abs(ridge).max() # OLS coefficients are large and unstable
assert abs(ridge.sum() - 3) < 0.2 # ridge splits the effect sensibly, total near 3
4. Logistic regression
For binary classification, pass the linear score through the sigmoid:
Equivalently, the log-odds are linear: . A coefficient is the change in log-odds per unit change in the feature; is the odds ratio.
The loss: cross-entropy (log loss)
Maximum likelihood for Bernoulli outputs gives
Why not squared error? With a sigmoid, squared error is non-convex in and has vanishing gradients when the model is confidently wrong. Cross-entropy is convex for logistic regression and its gradient is clean:
There is no closed-form solution, so you use gradient descent or Newton's method (iteratively reweighted least squares).
import numpy as np
def sigmoid(z):
return 1 / (1 + np.exp(-z))
rng = np.random.default_rng(2)
n = 1000
X = rng.normal(size=(n, 2))
true_w = np.array([2.0, -1.5]); true_b = 0.3
y = (rng.uniform(size=n) < sigmoid(X @ true_w + true_b)).astype(float)
Xb = np.column_stack([X, np.ones(n)])
w = np.zeros(3)
for _ in range(3000):
p = sigmoid(Xb @ w)
w -= 0.5 * Xb.T @ (p - y) / n # gradient of the mean log loss
p = sigmoid(Xb @ w)
loss = -np.mean(y * np.log(p) + (1 - y) * np.log(1 - p))
assert abs(w[0] - 2.0) < 0.5 and abs(w[1] + 1.5) < 0.5 # recovers the true direction
assert ((p > 0.5) == y).mean() > 0.8
assert loss < np.log(2) # beats the "always 0.5" model
Decision boundary, thresholds and calibration
The boundary is linear. The 0.5 threshold is a default, not a law: choose it from the costs of false positives and false negatives (see the metrics chapter). Logistic regression tends to give well-calibrated probabilities when it is the right model, which is one reason it is favoured in credit and medicine.
Multiclass: softmax
For classes, with , trained with categorical cross-entropy. Subtract the maximum from before exponentiating for numerical stability.
import numpy as np
def softmax(z):
z = z - z.max(axis=-1, keepdims=True) # stability: does not change the result
e = np.exp(z)
return e / e.sum(axis=-1, keepdims=True)
p = softmax(np.array([1000.0, 1001.0, 1002.0])) # would overflow without the shift
assert np.isclose(p.sum(), 1.0) and p.argmax() == 2
assert np.allclose(softmax(np.array([1.0, 2.0, 3.0])), softmax(np.array([101.0, 102.0, 103.0])))
Separable data
If one feature perfectly separates the classes, the unregularised maximum-likelihood weights diverge to infinity. Regularisation (an L2 penalty) keeps them finite. This is a classic trap question.
5. Practical checklist
- Standardise numeric features. One-hot encode low-cardinality categoricals.
- Check for multicollinearity (correlation matrix, variance inflation factor).
- Use regularisation by default; tune by cross-validation.
- Look at residual plots (linear) and calibration plots (logistic).
- Handle class imbalance with class weights or resampling, and judge with PR-based metrics.
- For nonlinearity, add polynomial terms or interactions, or switch to trees or boosting.
6. Common mistakes
- Using the normal equation with collinear features and getting garbage or a singular matrix.
- Not scaling features before regularisation, so the penalty treats features unequally.
- Reading coefficients causally, or comparing raw coefficients across different units.
- Using accuracy with imbalanced classes.
- Fitting squared loss for classification and thresholding the output.
- Forgetting the intercept is usually not regularised.
7. Practice questions
- Derive the normal equation for linear regression.
- Why is cross-entropy preferred to MSE for logistic regression? Derive the gradient.
- What does an L1 penalty do geometrically, and why does it give sparsity?
- What happens to logistic regression on perfectly separable data?
- Two features are almost perfectly correlated. What happens to the OLS coefficients, and what would you do?
- How do you interpret the coefficient of a feature in a logistic regression?
- Implement logistic regression with gradient descent from memory.
- When would you prefer Huber loss to squared loss?