A wellness clinic reports that patients who began a course of treatment while experiencing unusually severe symptoms showed meaningful average improvement four weeks later — presented, understandably, as evidence the treatment works. Some or even all of that observed improvement can occur for reasons that have nothing to do with the treatment's actual effectiveness, through a well-documented statistical pattern called regression to the mean, and a simple before-and-after comparison has no way to distinguish a real treatment effect from this pattern on its own.
What regression to the mean actually describes
Any measurement that fluctuates naturally over time — including symptom severity, which varies day to day for reasons unrelated to any specific treatment — will tend to show an extreme reading followed, on average, by a reading closer to that measurement's typical range, purely as a statistical consequence of natural variation around a stable average, with no causal intervention required to produce this pattern. A patient measured at an unusually bad moment is, for purely statistical reasons, more likely to be measured somewhat better next time than someone measured at a typical moment, independent of anything done in between.
Why this specifically distorts clinical and wellness outcome data
Patients overwhelmingly tend to seek treatment, or specifically decide to start a new treatment, when their symptoms are unusually severe relative to their own typical baseline — not on an average day, but on a particularly bad one. This means the intake measurement that anchors a before-and-after comparison is systematically drawn from an unusually severe moment, which makes some average improvement at the next measurement point statistically likely regardless of whether the treatment itself does anything at all, purely because of when patients typically choose to start treatment.
Why a before-and-after design alone can't distinguish the two explanations
A before-and-after comparison, without a control group of comparable patients who didn't receive the treatment, has no way to isolate how much of the observed improvement is attributable to the treatment itself versus how much would have occurred anyway, purely from regression to the mean and the natural course of the condition. Both explanations predict the exact same observable pattern — symptom severity measured at an unusually bad moment, followed by improvement — which is precisely why this design can't distinguish between them, no matter how large or statistically clean the before-and-after difference appears.
What a proper comparison actually reveals
A control or comparison group — patients with similarly severe symptoms at intake who don't receive the treatment, or who receive a different established treatment — experiences the same regression-to-the-mean effect, since they were also measured at an unusually severe moment. Comparing improvement in the treatment group against improvement in this comparison group isolates whatever additional effect the treatment itself contributes, above and beyond the improvement that would have occurred from natural fluctuation and regression to the mean alone. Without this comparison, a treatment that does absolutely nothing beyond a placebo effect can still show what looks like a strong, clinically meaningful before-and-after result.
What this means for how clinics and wellness practices should evaluate their own outcomes
- Treat before-and-after outcome data, without any comparison group, as suggestive at best — not evidence the treatment itself caused the observed improvement
- Recognize that intake severity is systematically biased toward unusually bad moments, which alone predicts some average subsequent improvement
- Seek out or reference controlled studies of a treatment's effectiveness rather than relying on internal before-and-after outcome tracking as the primary evidence
- Be appropriately cautious of marketing claims built entirely on before-and-after patient outcome data, since the same pattern would appear even for an ineffective treatment
None of this means a given treatment doesn't work — it means before-and-after data alone, however compelling it looks, cannot answer that question, and a comparison group is what actually separates a genuine treatment effect from a well-understood statistical artifact of when people choose to seek treatment in the first place.