Intermediate to senior

Machine Learning Interview Prep

Fifteen chapters from the learning problem and bias-variance to trees, neural networks, transformers, recommenders and ML system design, with tested NumPy code and diagrams.

Chapter 10 of 15Training and evaluation · Feature Engineering, Data Quality and Leakage

Feature Engineering, Data Quality and Leakage

On real projects the data decides the ceiling, and the model only decides how close you get. Interviewers use feature and data questions to separate people who have shipped models from people who have only run notebooks. This chapter covers encoding, scaling, missing values, feature construction, selection, and the leakage traps that ruin deployments.

1. Look at the data first

Before modelling:

  • Shape and types: rows, columns, which are numeric, categorical, text, timestamps, IDs.
  • Target: its distribution, class balance, how it was created, and how noisy it is.
  • Missingness: how much, and whether it is random or informative.
  • Duplicates and inconsistencies: repeated rows, mixed units, impossible values (negative ages, future dates).
  • Outliers: errors, or rare but real events?
  • Time: is there a temporal order or drift?
  • Relationships: correlations, and relationships with the target.

A short exploration often finds the single biggest accuracy gain, such as a unit mix-up or a label defect.

2. Encoding categorical variables

MethodHowGood forWatch out
One-hota binary column per categorylow cardinality, linear modelsexplodes with many categories
Ordinalintegers by ordertruly ordered categories (small, medium, large)invents order if none exists
Frequency / countcategory frequencyhigh cardinality, treescollisions between equally frequent categories
Target (mean) encodingmean target per categoryhigh cardinalityleaks the label unless computed out-of-fold; smooth toward the global mean
Hashinghash into fixed bucketshuge, streaming vocabulariescollisions
Embeddingslearned dense vectorsdeep models, very high cardinalityneeds enough data

Target encoding done carelessly is a classic leak: each row's own label contributes to its own encoding. The remedy is out-of-fold encoding (compute a row's encoding from folds that exclude it) with smoothing for rare categories.

import numpy as np

def smoothed_target_encoding(cats, y, global_mean, m=10):
    """Per-category mean shrunk toward the global mean; m controls how many rows we trust."""
    out = {}
    for c in set(cats):
        mask = np.array([ci == c for ci in cats])
        n, mean = mask.sum(), y[mask].mean()
        out[c] = (n * mean + m * global_mean) / (n + m)
    return out

cats = ["a"] * 100 + ["b"] * 2
y = np.array([1] * 60 + [0] * 40 + [1, 1], dtype=float)
gm = y.mean()
enc = smoothed_target_encoding(cats, y, gm, m=10)
assert abs(enc["a"] - 0.6) < 0.01                        # plentiful data: close to the raw mean
assert enc["b"] < 1.0 and enc["b"] > gm                  # two rows: raw mean is 1.0, shrunk toward the global mean

3. Scaling and transforming numeric features

TransformUse
Standardisation linear models, SVMs, neural nets, k-NN, PCA
Min-max scalingbounded inputs, image pixels
Log or Box-Coxright-skewed values (income, counts, prices)
Clipping / winsorisingtame extreme outliers
Binningcapture non-monotone effects; robust to outliers
Robust scaling (median, IQR)outlier-heavy data

Tree models do not need scaling because they depend only on the order of values. Distance-based and gradient-based methods do.

Fit all transformers on the training data only, then apply them to validation and test sets.

4. Missing values

First ask why it is missing:

  • MCAR (completely at random): no pattern; safe to impute.
  • MAR (at random given other features): imputation using other features works.
  • MNAR (not at random): missingness itself carries information (income left blank by high earners). Add a missing-indicator column.

Options: drop rows or columns (only if little is lost), impute with mean, median or mode, impute with a model (k-NN, iterative), or use algorithms that handle missing values natively (XGBoost, LightGBM). Impute inside the cross-validation loop.

import numpy as np

x = np.array([1.0, 2.0, np.nan, 100.0, np.nan])
median = np.nanmedian(x)
mean = np.nanmean(x)
assert median == 2.0 and abs(mean - 103 / 3) < 1e-12     # the outlier drags the mean far from the typical value
filled = np.where(np.isnan(x), median, x)
indicator = np.isnan(x).astype(int)
assert filled.tolist() == [1.0, 2.0, 2.0, 100.0, 2.0] and indicator.tolist() == [0, 0, 1, 0, 1]

5. Constructing features

Good features encode domain knowledge the model cannot easily discover.

  • Ratios and differences: debt-to-income, price per square foot, change since last month.
  • Aggregates over history: a user's purchases in the last 7 and 30 days, rolling means, counts, recency.
  • Interactions: products of features for linear models (trees find these themselves).
  • Date parts: hour, day of week, month, is-holiday. For cyclical values (hour of day) use sine and cosine so that 23:00 is near 00:00.
  • Text: TF-IDF, n-grams, or pretrained embeddings.
  • Geography: distance to a landmark, clusters, regional aggregates.
  • Domain transformations: log returns for prices, per-capita normalisation.
import numpy as np

hours = np.array([0, 6, 12, 18, 23])
s, c = np.sin(2 * np.pi * hours / 24), np.cos(2 * np.pi * hours / 24)

def dist(i, j):
    return np.hypot(s[i] - s[j], c[i] - c[j])

assert dist(0, 4) < dist(0, 2)                           # 23:00 is close to 00:00, far from 12:00
assert abs(hours[4] - hours[0]) > abs(hours[2] - hours[0])   # on the raw scale, 23:00 looks farther from 00:00 than 12:00 does

(On the raw integer scale, 23 and 0 look 23 apart. The sine and cosine pair places them next to each other on a circle.)

6. Feature selection

Why select: less overfitting, faster training and serving, easier interpretation, lower cost of data collection.

  • Filter methods: correlation, mutual information, chi-square. Fast, ignore interactions.
  • Wrapper methods: recursive feature elimination, forward selection. Expensive.
  • Embedded methods: L1 regularisation, tree importances.
  • Permutation importance on a validation set, then drop features whose removal does not hurt.

Remove near-constant columns and one of each highly correlated pair. For production, favour features that are cheap, stable and available at prediction time.

7. Data leakage

Leakage means the model sees information at training time that it will not have in production. It produces impressive validation numbers and then fails.

Kinds of leakage

  1. Target leakage. A feature that is derived from, or recorded after, the outcome. Examples: "days since last payment" for a loan default model where the field is updated after default; "discharge diagnosis" to predict admission; a refund_issued flag when predicting a return.
  2. Train-test contamination. Fitting scalers, imputers, encoders, PCA or feature selection on the full data before splitting. Resampling (SMOTE or oversampling) before the split. Duplicates present in both sets.
  3. Temporal leakage. Random splitting of time-ordered data so training sees the future. Features aggregated over a window that extends past the prediction time.
  4. Group leakage. The same user, patient, device or document appears in both train and test, so the model memorises the entity.
  5. Tuning leakage. Repeatedly selecting models by test performance.

Detecting it

  • A feature with suspiciously high importance or a single feature that nearly predicts the target.
  • Validation performance far better than any reasonable baseline, then collapsing in production.
  • Ask of every feature: "At the moment of prediction, would I have this value, and would it have the same meaning?"
  • Rebuild features as of the prediction timestamp (point-in-time correctness), especially with feature stores and joins on history.
import numpy as np

rng = np.random.default_rng(0)
# a leaky feature: it is the label plus a little noise (recorded after the outcome)
y = rng.integers(0, 2, 1000)
leaky = y + rng.normal(0, 0.1, 1000)
honest = rng.normal(0, 1, 1000)

def auc(y, s):
    pos, neg = s[y == 1], s[y == 0]
    return ((pos[:, None] > neg[None, :]).mean())

assert auc(y, leaky) > 0.99                              # "wonderful" performance that will vanish in production
assert abs(auc(y, honest) - 0.5) < 0.05                  # an honest uninformative feature is at chance

8. Training-serving skew

The features in production must be computed the same way as in training. Common causes of skew: separate code paths in notebooks and services, different rounding or time zones, missing-value handling that differs, and features that are late-arriving in production. Use one shared feature definition (a feature store or a single library) and log served features to compare against training distributions.

9. Labels matter as much as features

Check how labels were produced: human annotators (inter-annotator agreement), heuristics, delayed outcomes, or a previous model's decisions (which creates selection bias and feedback loops). Label noise caps achievable accuracy; cleaning labels often beats tuning models.

10. Common mistakes

  • Spending all the time on models and none on data.
  • Fitting preprocessing on all data before the split.
  • Target encoding without out-of-fold computation.
  • Dropping rows with missing values without checking what was lost.
  • Treating IDs as numbers (a user ID has no order).
  • Building features unavailable at serving time.
  • Ignoring distribution shift between training and production.

11. Practice questions

  1. Compare one-hot, target and frequency encoding. When does target encoding leak?
  2. How do you handle missing values? When is a missing-indicator useful?
  3. Which models need feature scaling, and why?
  4. Name four kinds of data leakage with an example of each.
  5. How would you encode the hour of the day?
  6. How do you prevent training-serving skew?
  7. A model has 0.99 AUC in validation. What do you check before celebrating?
  8. How do you select features for a model that must run in 10 ms?
Header Logo