An analysis finds that a company's overall customer satisfaction has improved year over year across the aggregate dataset. Breaking that same dataset into every individual customer segment shows satisfaction actually declined within every single segment over the same period. Both findings are correct, computed from the identical underlying data — this is Simpson's Paradox, a genuine and well-documented statistical phenomenon, not a data error, a coding mistake, or a contradiction that needs to be resolved by figuring out which number is wrong.
How both statements can be true at once
Simpson's Paradox typically arises when the composition of the aggregate population shifts between the two time points or groups being compared — in the example above, if a segment with historically lower satisfaction scores shrank as a share of the total customer base while a segment with historically higher scores grew, the aggregate average can rise even while every individual segment's own satisfaction score fell, purely because the mix of segments contributing to the aggregate shifted toward the historically higher-scoring group. The aggregate trend and the subgroup trends are each measuring something real and specific — they're just measuring different things, and the difference is driven by a shift in the underlying composition that the aggregate number alone doesn't reveal.
A confounding variable is usually the actual mechanism
Simpson's Paradox typically emerges because a third variable — group composition, in the example above — is unevenly distributed across whatever comparison is being made, and that uneven distribution is doing enough work to overturn the aggregate direction of an effect that holds consistently within every properly controlled subgroup. Identifying and adjusting for this confounding variable is exactly what resolves the apparent paradox intellectually, even though both the aggregate and subgroup numbers remain, individually, entirely accurate calculations from the data.
Why this is more than a statistical curiosity
Simpson's Paradox has shown up in real, consequential analyses — including a well-known historical case involving graduate school admissions data, where aggregate admission rates appeared to favor one applicant group, while admission rates within nearly every individual academic department favored the other group, once department-level application patterns were accounted for. Decisions made based on the aggregate number alone, without checking whether a confounding variable like department choice was driving the reversal, would have reached a conclusion opposite to what the more granular, department-level data actually supported.
Which level of analysis is actually the right one to trust
Neither the aggregate view nor the subgroup view is universally the correct one to rely on — the right choice depends entirely on the specific causal or decision-relevant question being asked. If the question is genuinely about the population as a whole, in its actual current composition, the aggregate figure is the relevant one. If the question is about whether some underlying factor caused an improvement or decline independent of compositional shifts, the subgroup-level, composition-adjusted view is the one that actually answers it. Treating one level as automatically more sophisticated or more correct than the other, without asking which question is actually being addressed, is itself a common source of misinterpretation.
What this means for anyone analyzing data across subgroups
- Whenever comparing an aggregate trend across time or groups, check whether the composition of the underlying subgroups shifted between the points being compared
- Treat a mismatch between an aggregate trend and every individual subgroup's trend as a signal to look for a confounding compositional variable, not as an error to be dismissed
- Decide explicitly which level of analysis actually answers the specific question at hand, rather than defaulting to whichever number is more readily available or more convenient
- Present both the aggregate and relevant subgroup-level views together when a real risk of Simpson's Paradox exists, rather than reporting only one level as if it were the complete picture
Simpson's Paradox is a reminder that a single summary number, however cleanly calculated, can genuinely obscure the underlying reality — not through any error, but through the ordinary mathematics of how aggregation and composition interact.