Skip to main content
Quant Research

Berkson's Paradox: How Selecting Your Sample Can Manufacture a Correlation That Isn't Really There

Two genuinely unrelated factors can appear negatively correlated purely because of how the specific sample being studied was selected, with no real relationship between them in the broader population at all.

Key Takeaways
  • Berkson's paradox describes how selecting a sample based on a criterion related to two otherwise independent variables can create an artificial correlation between those variables within the selected sample
  • The classic illustration involves hospital admissions, where studying only hospitalized patients can produce an apparent negative correlation between two diseases that are actually unrelated in the general population
  • This occurs because the selection criterion (hospital admission) can be satisfied by either variable alone, meaning within the selected sample, having one reduces the statistical need for the other to explain the selection
  • Recognizing when a sample has been selected based on a variable related to both factors being studied is the direct way to avoid drawing a spurious correlational conclusion from this pattern

A study examining only hospitalized patients finds that two entirely unrelated diseases appear negatively correlated with each other — patients with one condition seem less likely to have the other. In the general population, outside the hospital-admitted sample, no such relationship exists between the two diseases at all. This is Berkson's paradox, a specific and genuinely counterintuitive statistical phenomenon where the very act of selecting a sample based on a criterion related to two variables can manufacture an apparent correlation between them that doesn't reflect any real relationship in the broader population.

The classic hospital admission illustration

Imagine two unrelated diseases, each independently capable of causing a person to be admitted to the hospital, with hospital admission requiring either condition to be present in sufficiently severe form. Within the general population, having one disease tells you nothing about your likelihood of having the other, since they're genuinely unrelated. Within the hospitalized sample specifically, however, a patient who's there because of the first disease had less "need" for the second disease to also be present in order to explain their hospitalization — meaning among hospitalized patients specifically, having one disease is statistically associated with a lower likelihood of having the other, purely as a consequence of how the sample was selected, not because of any genuine biological relationship between the two conditions.

Why this happens mathematically

When a sample is selected based on a criterion that can be satisfied by either of two otherwise independent factors, conditioning on that selection criterion introduces a statistical dependence between the two factors within the selected sample that didn't exist in the unselected, general population — a pattern statisticians sometimes describe through the related concept of collider bias, where conditioning on a shared downstream consequence (here, hospital admission) of two independent upstream causes creates an artificial statistical relationship between those causes specifically within the conditioned, selected group.

Why this is a genuine, practical risk beyond the classic hospital example

Any research design that studies a sample selected based on some criterion related to both of the variables under investigation carries this same risk — a study of successful applicants to a competitive program, examining the relationship between two independent qualifications both used in the selection process, can find a spurious negative correlation between those qualifications purely because of how the selection process operated, entirely independent of whether any genuine relationship between the two qualifications exists in the broader applicant population.

What actually protects against drawing a spurious conclusion from this pattern

Examining whether the sample was selected based on a criterion plausibly related to both variables being studied, and specifically checking whether the same relationship holds in a broader, less selectively conditioned population, directly tests whether an observed correlation reflects a genuine relationship or a Berkson's paradox artifact of the specific selection process used to construct the sample being analyzed.

What this means for interpreting correlations found within a selected sample

  • Ask explicitly how the sample being studied was selected, and whether that selection criterion is plausibly related to both variables whose correlation is being examined
  • Be specifically skeptical of a correlation found only within a selected subgroup (hospitalized patients, admitted applicants, hired candidates) without checking whether it holds in the broader, unselected population
  • Recognize collider bias as the more general statistical mechanism behind this specific paradox, relevant well beyond the classic hospital admission example
  • Treat an unexpected, counterintuitive correlation found within a selectively sampled population as a candidate explanation for Berkson's paradox before concluding a genuine underlying relationship exists

Berkson's paradox is a reminder that the process used to select a sample isn't a neutral backdrop to the analysis performed on it — it can actively manufacture statistical relationships that have nothing to do with the actual, underlying reality the analysis is trying to describe.

Berkson's paradoxselection bias correlation artifactcollider biasstatisticiansspurious negative correlation