Probability, Statistics and A/B Testing for ML
Probability and statistics underpin both the theory (maximum likelihood, Bayesian priors, generalisation) and the practice (evaluating models, running experiments) of ML. Interviewers ask short puzzles and conceptual questions to check that your foundations are firm. This chapter covers the core results, with simulations you can run to check your answers.
1. Probability essentials
- Conditional probability: .
- Independence: .
- Law of total probability: over a partition .
- Bayes' theorem: .
- Expectation is linear: always. Variance adds only for independent (or uncorrelated) variables: .
The base-rate trap
A test for a disease that affects 1% of people is 95% sensitive and 90% specific. If you test positive, what is the chance you are ill?
About 9%, not 95%. This is the same reason a fraud model with a low false-positive rate can still flag mostly innocent transactions when fraud is rare.
import numpy as np
sens, spec, prev = 0.95, 0.90, 0.01
posterior = sens * prev / (sens * prev + (1 - spec) * (1 - prev))
assert abs(posterior - 0.0876) < 5e-4
# simulate it
rng = np.random.default_rng(0)
n = 2_000_000
ill = rng.uniform(size=n) < prev
positive = np.where(ill, rng.uniform(size=n) < sens, rng.uniform(size=n) < 1 - spec)
assert abs(ill[positive].mean() - posterior) < 0.005
Classic puzzles
- Monty Hall: switching wins with probability 2/3.
- Two children: given at least one boy, the probability both are boys is 1/3; given the older is a boy, it is 1/2. The wording of the condition matters.
- Birthday problem: among 23 people, the chance of a shared birthday exceeds 50%.
import numpy as np
rng = np.random.default_rng(1)
n = 200_000
# Monty Hall: always switch
car = rng.integers(0, 3, n); pick = rng.integers(0, 3, n)
switch_wins = (pick != car).mean() # switching wins exactly when the first pick was wrong
assert abs(switch_wins - 2 / 3) < 0.01
# birthday problem: exact calculation
p_no_match = 1.0
for i in range(23):
p_no_match *= (365 - i) / 365
assert 0.50 < 1 - p_no_match < 0.51
2. Distributions you must know
| Distribution | Models | Key facts |
|---|---|---|
| Bernoulli / Binomial | success/failure; number of successes in trials | mean , variance |
| Poisson | counts of rare independent events in a fixed interval | mean = variance = |
| Normal (Gaussian) | sums of many small effects; measurement error | 68-95-99.7 rule; closed under addition |
| Exponential | waiting time between Poisson events | memoryless; mean |
| Uniform | equal likelihood on an interval | |
| Beta | a probability between 0 and 1; conjugate prior for Bernoulli | |
| Categorical / Multinomial | one of outcomes | softmax outputs parameterise it |
| Power law / heavy tail | file sizes, wealth, popularity | mean may be dominated by rare extremes |
Central limit theorem (CLT)
The mean of independent, identically distributed samples with finite variance is approximately normal with standard deviation , whatever the original distribution. That is why you need four times the data to halve the uncertainty.
import numpy as np
rng = np.random.default_rng(2)
means = rng.exponential(1.0, size=(20000, 100)).mean(axis=1) # means of 100 draws from a skewed distribution
assert abs(means.mean() - 1.0) < 0.01
assert abs(means.std() - 1.0 / np.sqrt(100)) < 0.005 # standard error = sigma / sqrt(n)
3. Maximum likelihood and Bayesian estimation
Maximum likelihood estimation (MLE): choose parameters that make the observed data most probable. For Bernoulli data with successes in trials, . For Gaussian data, the MLE of the mean is the sample mean. Minimising squared error is MLE under Gaussian noise; minimising cross-entropy is MLE for Bernoulli/categorical outputs.
Maximum a posteriori (MAP): MLE with a prior: . A Gaussian prior on weights gives L2 regularisation; a Laplace prior gives L1.
Full Bayesian inference keeps the whole posterior and averages predictions over it, giving uncertainty, at greater cost.
A Beta prior with Bernoulli data stays Beta (conjugacy): with prior Beta() and successes in trials, the posterior is Beta(, ). This is the engine behind Thompson sampling for bandits.
import numpy as np
alpha, beta = 1, 1 # uniform prior
k, n = 3, 10 # 3 successes in 10 trials
post_alpha, post_beta = alpha + k, beta + n - k
post_mean = post_alpha / (post_alpha + post_beta)
assert abs(post_mean - 4 / 12) < 1e-12 # shrunk from the MLE 0.3 toward the prior mean 0.5
mle = k / n
assert mle == 0.3 and 0.3 < post_mean < 0.5
4. Hypothesis testing
- State a null hypothesis (no effect) and an alternative.
- Choose a significance level (commonly 0.05): the accepted rate of false positives (Type I errors).
- Compute a test statistic and its p-value: the probability, assuming is true, of a result at least as extreme as the one observed.
- Reject if .
What a p-value is not: the probability that is true, or the probability the result is due to chance. And "not significant" does not mean "no effect": it may mean too little data.
| Term | Meaning |
|---|---|
| Type I error | rejecting a true null (false positive), rate |
| Type II error | failing to reject a false null (false negative), rate |
| Power | : the chance of detecting a real effect of a given size |
| Confidence interval | a range computed so that, over repeated experiments, a stated fraction (such as 95%) contain the true value |
| Effect size | the magnitude of the difference, as opposed to whether it is detectable |
Multiple comparisons
Testing 20 hypotheses at yields about one false positive by chance alone. Correct with Bonferroni (use ), Holm, or control the false discovery rate (Benjamini-Hochberg). This also applies to checking many metrics and many segments in an experiment.
import numpy as np
rng = np.random.default_rng(3)
m, trials = 20, 5000
# all nulls true: p-values are uniform
p = rng.uniform(size=(trials, m))
any_false_positive = (p < 0.05).any(axis=1).mean()
assert abs(any_false_positive - (1 - 0.95 ** 20)) < 0.03 # about 64 % of experiments "find" something
bonferroni = (p < 0.05 / m).any(axis=1).mean()
assert bonferroni < 0.07 # corrected back to roughly 5 %
5. A/B testing
An A/B test randomly assigns users to control (A) or treatment (B) and compares a metric. Randomisation makes the groups comparable, so a difference can be attributed to the change.
Designing an experiment
- Hypothesis and primary metric, chosen in advance, with guardrail metrics (latency, errors, revenue, complaints).
- Unit of randomisation: usually the user, not the page view, to avoid contamination and repeated measures.
- Sample size from the baseline rate, the minimum detectable effect (MDE), and the desired power.
- Duration: run for full weekly cycles, long enough to cover novelty effects.
- Check validity: sample ratio mismatch (is the split what you designed?), pre-experiment balance.
- Analyse as planned, correcting for multiple metrics, and decide against pre-agreed criteria.
For a conversion-rate test, a rough per-group sample size is
where is the absolute lift to detect and the baseline rate. With (two-sided, ) and power 80% (), the numerator constant is .
import numpy as np
def sample_size_per_group(p, delta, z_alpha=1.96, z_beta=0.8416):
return 2 * (z_alpha + z_beta) ** 2 * p * (1 - p) / delta ** 2
n = sample_size_per_group(0.10, 0.01) # 10 % baseline, detect +1 percentage point
assert 14000 < n < 15000
assert sample_size_per_group(0.10, 0.02) < n / 3.5 # halving the detectable effect quadruples the sample, so doubling it cuts to a quarter
Pitfalls
- Peeking: stopping the moment inflates false positives. Use a fixed horizon, or sequential methods designed for it.
- Underpowered tests: inconclusive results and exaggerated effect sizes among the "winners".
- Many metrics and segments without correction.
- Novelty and primacy effects: early behaviour is not long-run behaviour.
- Interference: in marketplaces and social networks, treated users affect controls. Randomise by cluster or use switchback tests.
- Simpson's paradox: an aggregate trend reverses within segments when group sizes differ. Always check segment mixes.
- Survivorship and selection bias: analysing only users who stayed.
- Sample ratio mismatch: a sign of broken assignment or logging; do not trust the result.
- Statistical versus practical significance: a tiny effect can be significant with large and still not worth shipping.
6. Correlation and causation
Correlation measures association, not cause. A confounder can drive both variables. Randomised experiments are the gold standard; when impossible, use quasi-experimental methods (difference-in-differences, regression discontinuity, instrumental variables, propensity matching) with explicit assumptions. In interviews, ask "how was the treatment assigned?" before interpreting any effect.
7. Common mistakes
- Reading p-values as the probability the null is true.
- Adding variance when variables are not independent.
- Ignoring base rates in classification and tests.
- Stopping experiments early because the number looks good.
- Sample sizes picked after the fact.
- Claiming causation from observational data.
- Treating "not significant" as "no difference".
8. Practice questions
- A disease affects 0.5% of people; the test has 99% sensitivity and 95% specificity. What is the chance a positive result is real?
- State the central limit theorem and say why it matters for experiments.
- Show that MSE is the MLE loss for Gaussian noise.
- How is MAP estimation related to regularisation?
- Explain a p-value in plain words. What does it not mean?
- How do you size an A/B test? What is the effect of halving the detectable lift?
- Why is peeking at A/B results harmful?
- Give an example of Simpson's paradox and how you would detect it.
- How would you test a change in a two-sided marketplace where users interact?