Skip to main content
Quant Research

Regression Discontinuity Design: Turning an Arbitrary Cutoff Into a Natural Experiment

People just above and just below a strict eligibility cutoff are otherwise nearly identical, which means comparing outcomes right at that boundary approximates a genuine randomized comparison without ever randomizing anything.

Key Takeaways
  • Regression discontinuity design exploits a strict, arbitrary cutoff rule that determines eligibility for a program, treatment, or benefit, based on some continuous measure
  • People scoring just above and just below the cutoff are assumed to be otherwise very similar on average, differing systematically only in whether they received the treatment
  • Comparing outcomes for those just above versus just below the cutoff approximates a randomized comparison, since which side of an arbitrary threshold someone falls on is close to random for people near the boundary
  • The design's credibility depends on the cutoff being genuinely strictly enforced and not manipulable, and results are most trustworthy specifically near the threshold, not far from it

A scholarship program awards funding strictly to students scoring above a specific threshold on a standardized test, with no funding for students scoring below it, however narrow the gap. Students scoring 71 and 69 on this test are, in every meaningful sense relevant to the outcome being studied, nearly identical — yet one receives the scholarship and the other doesn't, purely because of an arbitrary point on a continuous scale. Regression discontinuity design exploits exactly this feature, treating the sharp cutoff as something close to a naturally occurring randomized assignment for people near the boundary.

Why people right around the cutoff are the credible comparison group

The core logic depends on the reasonable assumption that whatever underlying factors determine a person's exact score near a threshold — measurement error, small variations in test-day performance, minor scoring differences — are essentially unrelated to any other characteristic that would independently affect the outcome being studied, meaning a student scoring 69 versus 71 is unlikely to differ systematically on any dimension other than the tiny difference reflected in the two points, apart from whether they happened to cross the specific threshold.

What this design allows researchers to estimate

Comparing outcomes for students just above and just below the cutoff isolates something close to the scholarship's actual causal effect specifically for students near that threshold, since the comparison holds essentially constant the underlying academic characteristics that would otherwise confound a simple comparison between scholarship recipients broadly and non-recipients broadly, many of whom would differ substantially in ways unrelated to the scholarship itself.

Why the design's credibility depends specifically on the cutoff being strictly enforced

The core logic breaks down if the cutoff can be manipulated or gamed — if, for instance, students or administrators could influence which side of the threshold a specific student's score fell on, the assumption that assignment near the cutoff is essentially random would no longer hold, since people who successfully manipulated their way across the threshold might differ systematically from those who didn't attempt or couldn't achieve such manipulation. Checking for a suspicious clustering of scores just above the cutoff, more than would be expected from a smooth underlying distribution, is a standard diagnostic check for this specific threat to the design's validity.

Why results near the cutoff don't necessarily generalize further away

The design's credible causal estimate applies specifically to people near the threshold — it doesn't directly tell you what effect the scholarship would have on a student scoring considerably higher or considerably lower than the cutoff, since those students may differ from the near-threshold comparison group in ways relevant to how much the scholarship would actually help them, a genuine limitation on how broadly a regression discontinuity finding can be generalized beyond the specific population right around the boundary being studied.

What this means for evaluating research using this design

  • Check whether the cutoff rule is genuinely strict and non-manipulable, since gaming the threshold undermines the design's core logic
  • Look for evidence the researchers checked for suspicious score clustering just above the cutoff, a standard diagnostic for manipulation
  • Treat the resulting causal estimate as most credible specifically for people near the threshold, not as a general claim about the treatment's effect across the full range of the underlying measure
  • Recognize regression discontinuity as a genuinely credible quasi-experimental method precisely because it turns an existing, arbitrary administrative rule into something approximating random assignment

Regression discontinuity design is one of the more elegant tools in applied causal inference specifically because it doesn't require the researcher to create anything new — it finds a genuine, credible natural experiment hiding inside an arbitrary administrative rule that already exists.

regression discontinuity designnatural experiment cutoffcausal inference thresholdstatisticiansquasi-experimental methods