Skip to main content
Assessments & Testing

Cut Scores Are a Policy Decision Dressed Up as a Statistical One

The line between pass and fail on a certification exam looks like an objective statistical output. It's actually the product of a series of judgment calls made by a panel of people.

Key Takeaways
  • A cut score determines the pass/fail threshold on a certification exam, and is frequently presented with a precision that implies it was derived objectively from the data alone
  • Standard-setting methods like the Angoff method are built on structured expert judgment about item difficulty for a minimally competent candidate, not a statistical calculation performed on test results
  • Different, equally defensible panels of subject-matter experts following the same standard-setting method can and do produce somewhat different cut scores
  • None of this makes cut scores arbitrary — it means they're a defensible policy judgment, and should be evaluated and communicated as one, not presented as a purely objective statistical fact

A certification body announces a cut score of 72% required to pass — a specific number, presented with the apparent precision of a scientifically derived measurement. In most standard-setting methodologies actually used to establish that number, including the widely used Angoff method, the cut score is the output of a structured process of expert human judgment, not a statistical calculation performed directly on candidate response data. This distinction matters more than it usually gets credit for.

How a cut score is actually set, in outline

In the Angoff method and its variants, a panel of subject-matter experts reviews each exam item and estimates the probability that a hypothetical "minimally competent candidate" — someone who just barely meets the standard the certification is meant to represent — would answer that item correctly. These item-level probability estimates are aggregated across the panel and across items to produce the recommended cut score. Every step of this process — who sits on the panel, how "minimally competent" is defined and communicated to them, how their individual estimates are aggregated — involves a genuine judgment call, made by people, not an output derived mechanically from test data.

This is a genuinely different kind of process than it's often assumed to be

A cut score is frequently discussed as though it were discovered in the data — as if some inherent property of the test separates competent from incompetent candidates at a specific, statistically determined point. What standard-setting methods like Angoff actually do is operationalize a policy decision (what does "minimally competent" mean for this credential, and how demanding should the standard be) through a structured, defensible process of expert judgment applied to the test's content. The process is rigorous and well-established specifically because it makes the judgment involved explicit and structured — not because it removes judgment from the process entirely.

Different panels can produce different, equally defensible cut scores

A second, equally qualified panel of subject-matter experts, following the identical Angoff methodology on the identical exam, can and does produce a somewhat different recommended cut score than the first panel — a well-documented property of standard-setting research, not a sign of a flawed or poorly executed process. This variability exists precisely because the underlying task involves genuine judgment (what does minimal competence actually look like on this specific item), and reasonable experts can and do arrive at somewhat different, still defensible, estimates of that judgment.

Why this doesn't make cut scores arbitrary or illegitimate

None of this means a properly conducted standard-setting study is unreliable or that cut scores could reasonably be set anywhere — a well-run Angoff study with a properly selected, well-trained panel is a rigorous, defensible way to operationalize a genuine policy decision about what a credential should certify. The point isn't that the process is flawed; it's that the resulting number is a policy judgment made through a structured expert process, and should be understood, communicated, and defended as exactly that — not presented with a false precision that implies it was extracted objectively from the data itself.

What this means for organizations relying on certification cut scores

  • Ask what standard-setting method was used to establish a cut score, not just what the resulting number is
  • Understand that a cut score can be re-examined and legitimately adjusted through a new standard-setting study without implying the original process was flawed
  • Communicate cut scores to stakeholders and candidates in terms that reflect their actual nature — a defensible expert-judgment-based standard, not an objectively discovered statistical fact
  • Be appropriately skeptical of a certification body that can't describe or produce documentation of the standard-setting method actually used to establish its cut score

A rigorous cut score and an arbitrary one are genuinely different things — the distinction lies in whether a defensible standard-setting process was actually followed, not in whether human judgment was involved at all, since it always is.

cut scorestandard settingAngoff methodcertification examscertification bodies