Biostatistics & Epidemiology
1. What this chapter covers, and how NEET PG actually tests it
This is the most calculable chapter in PSM, and the questions are almost always a two-by-two table, a study design or a choice of statistical test.
The organising principle has two halves.
Every epidemiological measure answers one of three questions: how much disease is there, does the exposure matter, and could the finding be an artefact.
Every statistical test follows from two facts about the data: what type it is, and how many groups are being compared.
| Question | Measures |
|---|---|
| How much disease? | Incidence, prevalence, attack rate, mortality |
| Does the exposure matter? | Relative risk, odds ratio, attributable risk |
| Could this be an artefact? | Bias, confounding, chance |
2. Measuring disease frequency
2.1 Incidence and prevalence
Incidence counts new cases arising in a defined population during a defined period, so it measures risk.
Prevalence counts all existing cases at a point or over a period, so it measures burden.
The relationship between them is the single most useful formula in the chapter.
A disease can therefore have high prevalence for two quite different reasons: many new cases, or few new cases that last a long time.
This explains a common source of confusion. A treatment that prevents death without curing the disease raises prevalence, because it lengthens duration, even though it is unambiguously beneficial.
Incidence is the measure of choice for studying causation and for evaluating prevention; prevalence is the measure for planning services.
2.2 Rates in outbreaks and mortality
Attack rate is the incidence during an epidemic, expressed as a percentage of those exposed over the epidemic period.
Secondary attack rate is the proportion of susceptible contacts of a case who develop disease within the incubation period, and it measures infectivity.
Case fatality rate is deaths among diagnosed cases, and it measures virulence.
| Rate | Numerator | Denominator |
|---|---|---|
| Crude death rate | All deaths in a year | Mid-year population |
| Infant mortality rate | Deaths under 1 year | Live births |
| Maternal mortality ratio | Maternal deaths | 100,000 live births |
| Case fatality rate | Deaths from a disease | Diagnosed cases of it |
Note that the maternal mortality ratio uses live births, not pregnancies, as its denominator, which is why it is a ratio rather than a true rate.
2.3 Natural history and levels of prevention
Every intervention is placed by where it acts on the natural history of the disease, and the exam tests the placement rather than the definition.
| Level | Acts | Example |
|---|---|---|
| Primordial | Before the risk factor appears | Discouraging children from taking up tobacco |
| Primary | Before disease begins | Immunisation, sanitation, seat belts |
| Secondary | Early disease, before symptoms | Screening, early treatment |
| Tertiary | Established disease | Preventing disability, rehabilitation |
Health promotion and specific protection are the two components of primary prevention, and the distinction is that promotion is non-specific while protection targets one agent.
The iceberg of disease describes the far larger submerged portion of subclinical infection, carriers and undiagnosed cases lying beneath the visible clinical tip.
The ratio of subclinical to clinical infection varies enormously by organism, being very high in poliomyelitis and low in measles, and it determines whether case-finding alone can control an outbreak.
2.4 Investigating an outbreak
The steps run in a fixed order: verify the diagnosis, confirm that an epidemic exists, define a case, describe by time, place and person, formulate and test a hypothesis, then act.
Describing by time, place and person comes before any hypothesis, because a hypothesis formed too early narrows the investigation prematurely.
A point source epidemic curve rises and falls sharply within one incubation period, while a propagated curve shows successive waves separated by roughly one incubation period each.
Herd immunity is the protection of susceptible individuals by a sufficiently immune surrounding population, and it works only for diseases transmitted person to person.
That is why tetanus has no herd immunity, since the reservoir is soil rather than other people, and why every individual must be immunised personally.
The basic reproduction number is the average number of secondary cases arising from one case in a fully susceptible population, and the herd immunity threshold rises as it rises.
3. Study designs
3.1 The hierarchy and what each yields
| Design | Direction | Measure obtained |
|---|---|---|
| Case report or series | None | Description only |
| Cross-sectional | Snapshot | Prevalence |
| Case-control | Backward from disease | Odds ratio |
| Cohort | Forward from exposure | Relative risk, incidence |
| Randomised controlled trial | Forward with allocation | Efficacy, relative risk |
Systematic review with meta-analysis sits above all of these, because it pools studies rather than generating new data.
3.2 Case-control studies
Cases with the disease are compared with controls without it, and past exposure is compared between them.
They are quick, cheap, need few subjects, and are the only practical design for a rare disease or a long latent period.
They cannot measure incidence, because the investigator chose how many cases and controls to include, so the odds ratio is used as an approximation of relative risk.
Recall bias is their characteristic weakness, since people with a disease search their memory for explanations more thoroughly than healthy controls do.
3.3 Cohort studies
A group defined by exposure is followed forward to see who develops disease.
Cohorts yield incidence directly and therefore relative risk, and they establish the correct temporal sequence.
They are the design of choice for a rare exposure, and they can study many outcomes of a single exposure at once.
Their weaknesses are cost, duration and loss to follow-up, and they are impractical for rare diseases because the required sample becomes enormous.
3.4 Randomised controlled trials
Randomisation is what distinguishes a trial, because it distributes both known and unknown confounders evenly between groups.
Allocation concealment prevents the person enrolling participants from knowing the next assignment, and it is a different safeguard from blinding.
Blinding prevents knowledge of allocation from influencing behaviour, assessment or reporting after enrolment.
Intention-to-treat analysis keeps participants in the group to which they were randomised regardless of what they actually received, because analysing by treatment received destroys the benefit of randomisation.
4. Measures of association
4.1 Risk and odds
In the standard two-by-two table, a is exposed with disease, b is exposed without, c is unexposed with disease and d is unexposed without.
Relative risk is the ratio of incidence in the exposed to incidence in the unexposed, and a value of one means no association.
The odds ratio approximates the relative risk closely only when the disease is rare, because when disease is common the terms a and c are no longer small relative to their denominators.
4.2 Attributable measures
Attributable risk is the incidence in the exposed minus the incidence in the unexposed, and it states how much of the risk the exposure adds.
Attributable risk percent divides that difference by the incidence in the exposed, giving the proportion of disease in exposed people that is due to the exposure.
Population attributable risk applies the same logic to the whole population and therefore depends on how common the exposure is.
Relative risk answers a question about biology; attributable risk answers a question about public health. A strong association with a rare exposure may matter less to a population than a weak association with a very common one.
Number needed to treat is the reciprocal of the absolute risk reduction.
5. Bias, confounding and chance
5.1 Bias
Bias is a systematic error that distorts the result in a consistent direction, and no increase in sample size will correct it.
| Bias | Nature |
|---|---|
| Selection bias | Those studied differ systematically from those not studied |
| Recall bias | Cases remember exposure differently from controls |
| Berkson bias | Hospital cases and controls differ because admission itself is selective |
| Neyman bias | Cases who died or recovered quickly are missed in a prevalence study |
| Observer bias | The assessor's expectation influences measurement |
| Lead-time bias | Screening advances diagnosis without postponing death |
| Length-time bias | Screening preferentially detects slowly progressive disease |
5.2 Confounding
A confounder is associated with the exposure, is an independent risk factor for the outcome, and is not on the causal pathway between them.
The classic example is that alcohol appears associated with lung cancer only because drinkers smoke more.
Confounding is controlled at the design stage by randomisation, restriction or matching, and at the analysis stage by stratification or multivariable adjustment.
Randomisation is uniquely powerful because it controls unknown confounders as well as known ones, which no analytical method can do.
5.3 Causation
Statistical association is not causation, and the Bradford Hill considerations set out what strengthens the case.
Temporality is the only one that is strictly necessary; the exposure must precede the outcome.
The others, including strength of association, dose-response relationship, consistency across studies, biological plausibility and reversibility on removing the exposure, add weight without being individually decisive.
6. Screening tests
6.1 The four measures
Sensitivity is the ability to identify those who have the disease, and a highly sensitive test is used to rule disease out when negative.
Specificity is the ability to identify those who do not, and a highly specific test is used to rule disease in when positive.
Sensitivity and specificity are properties of the test and do not change with prevalence; predictive values are properties of the population and change with it.
Positive predictive value is the proportion of positive results that are true positives, and it falls sharply when a test is applied to a low-prevalence population.
This is why mass screening of a healthy population with even an excellent test generates large numbers of false positives.
6.2 Choosing a cut-off and judging a programme
Raising the cut-off of a continuous test raises specificity and lowers sensitivity, and lowering it does the reverse.
The receiver operating characteristic curve plots sensitivity against one minus specificity, and the area under it summarises overall accuracy.
A screening programme requires more than a good test: the disease must be important and have a recognisable latent stage, an accepted treatment must exist, facilities must be available, and the process must be continuous rather than a one-off exercise.
Lead-time and length-time bias both make screening look more effective than it is, which is why mortality rather than survival is the correct outcome measure for a screening programme.
7. Biostatistics
7.1 Types of data and summary measures
Nominal data are unordered categories, ordinal data are ordered categories, and interval or ratio data are numerical.
Mean, median and mode coincide in a symmetrical distribution; in a skewed distribution the mean is pulled toward the tail, so the median is the better summary.
Standard deviation measures the scatter of individual observations, whereas standard error of the mean measures the precision of the estimate of the mean.
Because the standard error shrinks with the square root of the sample size, quadrupling the sample halves the standard error, which is the reason large studies give narrow confidence intervals.
In a normal distribution roughly 68 per cent of observations lie within one standard deviation of the mean, 95 per cent within about two and 99.7 per cent within three.
7.2 Hypothesis testing
The null hypothesis states that there is no difference, and the p-value is the probability of observing a result at least as extreme as the one obtained if the null hypothesis were true.
A p-value is not the probability that the null hypothesis is true, and this misinterpretation is examined directly.
| Error | Meaning | Denoted |
|---|---|---|
| Type I | Rejecting a true null hypothesis, a false positive | Alpha |
| Type II | Failing to reject a false null hypothesis, a false negative | Beta |
Power is one minus beta, the probability of detecting a true difference, and it rises with sample size, with effect size and with lower variability.
A confidence interval is more informative than a p-value because it shows the size and precision of the effect, and an interval for a ratio that includes one indicates no significant difference.
7.3 Choosing a test
| Situation | Test |
|---|---|
| Two independent group means | Unpaired t-test |
| Paired measurements | Paired t-test |
| More than two group means | Analysis of variance |
| Proportions or categorical data | Chi-square test |
| Small expected cell counts | Fisher exact test |
| Non-normal or ordinal data | Mann-Whitney or Wilcoxon |
| Relationship between two continuous variables | Correlation and regression |
Correlation describes the strength and direction of a linear relationship, while regression predicts one variable from another.
A correlation coefficient runs from minus one to plus one, and a value near zero excludes a linear relationship but not a curved one.
Choosing between parametric and non-parametric tests turns on whether the data are approximately normally distributed, not on sample size alone.
8. Worked examples
Example 1. A new treatment prevents deaths from a chronic disease without curing it. What happens to incidence and prevalence?
Incidence is unchanged, because no new cases are prevented. Prevalence rises, because prevalence equals incidence multiplied by duration and the treatment has lengthened duration.
Example 2. A test with 99 per cent sensitivity and 99 per cent specificity is applied to a population where the disease prevalence is one in ten thousand. What is the problem?
Nearly all positives will be false. With so few true cases, the one per cent false positive rate applied to the enormous disease-free majority swamps the true positives, so positive predictive value collapses.
Example 3. A case-control study of a rare cancer reports an odds ratio of 4.2. Can this be read as a relative risk?
Yes, approximately. The odds ratio approximates relative risk closely when disease is rare, which is satisfied here. A case-control study cannot measure incidence directly, so the odds ratio is the only available estimate.
Summary
Ask how much disease there is, whether the exposure matters, and whether the finding could be an artefact.
Incidence measures risk and prevalence measures burden, and prevalence equals incidence multiplied by duration.
A treatment that prolongs life without curing raises prevalence, which is a benefit rather than a failure.
Secondary attack rate measures infectivity and case fatality rate measures virulence.
Maternal mortality ratio uses live births as its denominator, which is why it is a ratio.
Case-control studies go backward from disease and yield an odds ratio; cohort studies go forward from exposure and yield relative risk and incidence.
Case-control is the design for a rare disease, cohort for a rare exposure.
Recall bias is the characteristic flaw of case-control studies, and loss to follow-up of cohorts.
Randomisation controls unknown as well as known confounders, which no analytical adjustment can do.
Allocation concealment and blinding are different safeguards operating at different stages.
Intention-to-treat analysis preserves the benefit of randomisation and is the correct primary analysis.
The odds ratio approximates relative risk only when the disease is rare.
Relative risk speaks to biology and attributable risk to public health, and population attributable risk depends on how common the exposure is.
Number needed to treat is the reciprocal of the absolute risk reduction.
Bias is systematic and is not fixed by a larger sample; chance is random and is.
A confounder is linked to the exposure, independently causes the outcome, and is not on the causal pathway.
Temporality is the only Bradford Hill criterion that is strictly necessary.
Sensitivity and specificity belong to the test; predictive values belong to the population and move with prevalence.
Lead-time and length-time bias flatter screening, which is why mortality and not survival is the correct outcome.
Standard deviation describes scatter and standard error describes precision, falling with the square root of sample size.
A p-value is not the probability that the null hypothesis is true.
Type I error is a false positive and type II a false negative, and power is one minus beta.
Choose the test from the data type and the number of groups, and prefer a confidence interval to a bare p-value.
Place every intervention on the natural history: primordial before the risk factor, primary before disease, secondary before symptoms, tertiary after disability.
Investigate an outbreak in order, describing time, place and person before forming any hypothesis.
A point source curve rises and falls within one incubation period; a propagated curve shows successive waves.
Herd immunity requires person-to-person transmission, which is why tetanus has none.
