Intermediate to senior

Machine Learning Interview Prep

Fifteen chapters from the learning problem and bias-variance to trees, neural networks, transformers, recommenders and ML system design, with tested NumPy code and diagrams.

Chapter 15 of 15Deep learning and applications · ML System Design and MLOps

ML System Design and MLOps

An ML system design interview asks you to take a vague product goal and design the whole pipeline: framing, data, features, model, evaluation, serving, monitoring and iteration. The model is typically a small part of the answer. What impresses is structure, trade-offs and awareness of what goes wrong in production.

1. A framework for any ML design question

Work through these steps out loud and keep each short.

<!--fig:mlloop-->
1. Frame the problemgoal, metric, baseline 2. Data and labelssources, quality, privacy 3. Featurespoint-in-time, shared with serving 4. Train and evaluatesplits, slices, error analysis 5. Deployshadow, canary, A/B 6. Monitordrift, quality, business metric Monitoring feeds back into data, features and retraining: the loop never ends. Figure 1. The ML lifecycle is a loop, not a line.
  1. Clarify the goal. What decision does the model support? Who uses it? What does success look like for the business?
  2. Frame it as an ML problem. Classification, regression, ranking, generation, anomaly detection, or maybe no ML at all. Define inputs, outputs and the prediction moment.
  3. Metrics. An offline model metric tied to an online business metric, plus guardrails.
  4. Data. Sources, labels (how, how fast, how noisy), volume, freshness, privacy, bias.
  5. Features. Signals, how they are computed, and availability at serving time.
  6. Baseline, then model. A simple heuristic and a simple model first. Justify complexity.
  7. Training and evaluation. Splits, validation, error analysis, fairness slices.
  8. Serving. Batch or online, latency, throughput, cost, fallbacks.
  9. Launch. Shadow mode, A/B test, gradual rollout, rollback plan.
  10. Monitoring and iteration. Data and concept drift, retraining, feedback loops, ownership.

Spend the first few minutes on 1 to 3. Most weak answers start with a model architecture.

2. Is ML even needed?

Ask whether rules or analytics solve the problem. Use ML when the pattern is too complex for rules, data exists, the cost of errors is manageable, and the pattern changes over time. If a decision table or a simple threshold gets 90% of the value, ship that first and measure.

3. Choosing metrics

Connect three layers:

LayerExample for a fraud system
Businessfraud losses prevented, customer friction, cost of manual review
Onlinechargeback rate, false-decline rate, review queue size
Offlinerecall at a fixed precision, PR-AUC, calibration

State guardrails: latency, fairness across groups, complaint rate. A model that improves the offline metric but increases customer friction is a loss.

4. Batch versus online prediction

BatchOnline (real-time)
Triggerschedule, such as nightlyper request
Latencyminutes to hoursmilliseconds
Costcheap, efficient hardware usealways-on capacity
Freshnessstale between runsfresh features and context
Examplesweekly churn scores, email targetingfraud checks, search ranking, autocomplete
Complexitylowerhigher: feature serving, caching, timeouts

A common hybrid: precompute heavy parts (embeddings, long-term features) in batch and combine with real-time context at request time. Streaming features (counts in the last 5 minutes) need a stream processor.

5. Components of a production ML system

  • Data pipelines: ingestion, validation, cleaning, labelling, versioning.
  • Feature store: one definition of each feature for both training (point-in-time correct, offline) and serving (low-latency, online), preventing training-serving skew.
  • Training pipeline: reproducible, scheduled or triggered, with experiment tracking (parameters, data version, code version, metrics, artefacts).
  • Model registry: versioned models with metadata, approval and rollback.
  • Serving layer: an API or batch job, with autoscaling, caching, timeouts and fallbacks (a default model or a rule).
  • Monitoring: system health, data quality, prediction distribution, and business metrics.
  • Orchestration and CI/CD: tests for data, code and models; automated deployment with gates.

6. Serving and performance

  • Latency budget: break the end-to-end budget across feature fetch, model inference and post-processing. Measure the 95th and 99th percentiles, not only the mean.
  • Throughput and cost: batching requests raises GPU efficiency but adds queuing delay.
  • Model optimisation: quantisation (lower-precision weights), pruning, distillation (a small student mimics a large teacher), compilation and operator fusion, caching frequent predictions.
  • Hardware: CPU for tree models and small nets; GPU or accelerators for large neural networks.
  • Safe fallbacks: if the model times out, serve a cached or rule-based answer rather than failing.
import numpy as np

# a latency budget check: p99 matters more than the mean for tail-sensitive services
rng = np.random.default_rng(0)
lat = rng.lognormal(mean=np.log(30), sigma=0.6, size=100000)       # milliseconds, right-skewed like real traffic
mean, p50, p99 = lat.mean(), np.percentile(lat, 50), np.percentile(lat, 99)
assert p50 < mean < p99
assert p99 > 3 * p50                                                # the tail is several times the typical request

7. Deployment strategies

StrategyHowWhy
Shadow modenew model runs on live traffic, outputs logged but not usedcompare with the old model with no user risk
Canarysend a small share of traffic to the new modelcatch problems early
A/B testrandomised comparison on the business metriccausal evidence
Gradual rollout and rollbackramp up with automated checkslimit blast radius
Interleaving (ranking)mix results from two rankers in one listsensitive, fast comparison
Banditsallocate traffic adaptivelywhen exploration cost matters

8. Monitoring and drift

Models decay because the world changes.

TypeWhat changesDetect with
Data drift (covariate shift)input distribution feature statistics, PSI, KS test
Concept driftthe relationship performance on fresh labels, proxy metrics
Label shiftclass prior prediction distribution, delayed labels
Data quality issuesnulls, schema, units, upstream bugsvalidation checks, anomaly alerts

The population stability index compares a feature's binned distribution in training with production: with the expected share and the actual share per bin. Rules of thumb: below 0.1 stable, 0.1 to 0.25 moderate shift, above 0.25 significant.

import numpy as np

def psi(expected, actual, bins=10):
    edges = np.quantile(expected, np.linspace(0, 1, bins + 1))
    edges[0], edges[-1] = -np.inf, np.inf
    e = np.histogram(expected, edges)[0] / len(expected)
    a = np.histogram(actual, edges)[0] / len(actual)
    e, a = np.clip(e, 1e-6, None), np.clip(a, 1e-6, None)
    return float(np.sum((a - e) * np.log(a / e)))

rng = np.random.default_rng(0)
train = rng.normal(0, 1, 50000)
same = rng.normal(0, 1, 50000)
shifted = rng.normal(0.8, 1.2, 50000)
assert psi(train, same) < 0.02
assert psi(train, shifted) > 0.25

Labels often arrive late (a loan defaults months later). Monitor proxy metrics meanwhile: prediction distribution, feature drift, upstream data health, and online metrics. Plan retraining: scheduled, triggered by drift or performance, or continuous, with evaluation gates before promotion.

9. Testing ML systems

  • Data tests: schema, ranges, null rates, uniqueness, freshness, label distribution.
  • Unit tests for feature code and preprocessing.
  • Model tests: performance above a baseline, behaviour on slices, invariance and directional tests (increasing income should not lower a credit score), robustness to noise.
  • Pipeline tests: reproducible runs, training-serving parity.
  • Integration and load tests on the serving path.
  • Fairness checks across relevant groups.

10. Responsible ML

  • Bias: measure error rates by group; consider equalised odds, demographic parity, and the trade-offs between them (they cannot generally all hold at once).
  • Privacy: minimise data, anonymise, access controls, differential privacy or federated learning where appropriate, and legal compliance (India's DPDP Act, GDPR).
  • Explainability: simple models, SHAP, counterfactual explanations, documentation (model cards).
  • Human oversight: review for high-stakes decisions, appeal paths.
  • Security: poisoning, adversarial inputs, model extraction, prompt injection for LLM systems.

11. A worked example outline: "Design a spam detector for a messaging app"

  1. Goal: reduce spam reaching users without blocking real messages. Cost asymmetry: blocking a real message is worse than showing spam, so aim for very high precision at useful recall.
  2. Framing: binary classification at message arrival; later user reports provide labels.
  3. Data: reports, blocks, deletion behaviour, labelled samples, sender graphs; careful about label delay and adversarial adaptation.
  4. Features: text and URL features, sender account age and message rate, recipient diversity, reputation, device signals. Privacy: use on-device or hashed signals where end-to-end encryption applies.
  5. Models: rules for known patterns, gradient-boosted trees on behavioural features, a text classifier for content.
  6. Serving: real-time, tens of milliseconds, with a cheap first-pass filter and an expensive second stage for borderline cases.
  7. Evaluation: recall at 99.9% precision, user report rate, false-block appeals.
  8. Adversaries: spammers adapt, so retrain frequently, monitor for drift, and keep the rules layer fast to update.
  9. Rollout: shadow, canary, A/B with guardrails.

12. Common mistakes

  • Starting with the architecture instead of the goal and metric.
  • Ignoring serving constraints such as latency and feature availability.
  • No baseline and no plan to prove value.
  • No monitoring or retraining plan. The model is deployed once and forgotten.
  • Training-serving skew from separate code paths.
  • Ignoring feedback loops where the model's decisions shape its own future training data.
  • Skipping fallbacks and failure modes.
  • Treating fairness and privacy as afterthoughts.

13. Practice questions

  1. Walk me through designing an ML system end to end for [fraud, search ranking, notifications, ETA prediction].
  2. How would you decide between batch and real-time prediction?
  3. What is a feature store, and which problem does it solve?
  4. Explain data drift versus concept drift and how you would detect each.
  5. How do you roll out a new model safely?
  6. Labels arrive 60 days late. How do you monitor model quality now?
  7. How would you reduce the serving cost of a large model by 5x?
  8. How do you test an ML pipeline, beyond unit tests?
Header Logo