Skip to main content
Assessments & Testing

The Bookmark Method: A Different Way to Set a Passing Score Than the Classic Angoff Approach

Rather than asking expert judges to estimate item-by-item difficulty for a hypothetical minimally competent candidate, the bookmark method has judges work through items ordered by actual difficulty and mark where competence begins.

Key Takeaways
  • The bookmark method for setting a test's cut score has expert judges work through a booklet of items ordered from easiest to hardest, based on item response theory difficulty estimates, and mark the point where a minimally competent candidate would likely start struggling
  • This is a genuinely different judgment task than the Angoff method discussed elsewhere, which asks judges to estimate, item by item, the probability a minimally competent candidate would answer each specific item correctly
  • The bookmark method's ordered-booklet format is thought by some researchers to be a more natural, intuitive judgment task for expert panelists than Angoff's item-by-item probability estimation
  • Both methods require genuine expert judgment and carry their own specific strengths and limitations, meaning the choice between them is a genuine methodological decision, not a matter of one method being simply superior

A certification body setting a passing score for a new exam has a genuine choice between standard-setting methodologies, including the widely used Angoff method discussed elsewhere and an alternative called the bookmark method, which asks expert judges to work through a specially ordered booklet of test items, arranged from easiest to hardest based on item response theory difficulty estimates, and place a literal bookmark at the point where they believe a minimally competent candidate would begin to struggle.

How the bookmark method's specific judgment task differs from Angoff's

Rather than asking judges to estimate, item by item, the specific probability that a minimally competent candidate would answer each individual item correctly — the Angoff method's core task — the bookmark method presents items already ordered by their actual, empirically estimated difficulty and asks judges to identify a single transition point in that ordered sequence, a genuinely different cognitive task that some standard-setting researchers argue is more natural and intuitive for expert panelists to perform reliably.

Why the item ordering itself requires item response theory data

The bookmark method's ordered-booklet format depends on having genuine, empirically estimated item difficulty parameters available in advance, typically derived from item response theory modeling of prior candidate response data — meaning the bookmark method specifically requires this kind of psychometric infrastructure already in place, a requirement the Angoff method's item-by-item judgment task doesn't share in the same way.

Why some standard-setting researchers consider the bookmark method's task more intuitive

Estimating a specific numeric probability for each individual item, as the Angoff method requires, is a demanding and somewhat unnatural cognitive task for many expert judges, who may not have strong, well-calibrated intuitions for translating their sense of an item's difficulty into a precise probability estimate — identifying a single transition point in an already-difficulty-ordered sequence of items is argued by proponents to more closely match how judges naturally think about a minimally competent candidate's likely performance, without requiring the same kind of precise numeric probability estimation.

Why the choice between these methods remains a genuine methodological decision, not a settled preference

Both methods have documented strengths and limitations, and the broader standard-setting research literature doesn't present either as simply, universally superior to the other — the Angoff method's item-by-item structure provides more granular diagnostic information about specific items, while the bookmark method's ordered-booklet format may better suit judges' natural intuitions, and the appropriate choice depends on a testing program's specific available data, judge training resources, and practical constraints.

Why understanding this choice matters for anyone evaluating a certification program's standard-setting rigor

A testing program that has thoughtfully selected and properly implemented either the Angoff or bookmark method, with well-trained expert panelists and a defensible process, demonstrates a genuinely rigorous approach to standard setting — the specific method chosen matters less than whether it was implemented with genuine methodological rigor, adequate panelist training, and appropriate documentation of the process actually followed.

What this means for evaluating a testing program's standard-setting methodology

  • Recognize the bookmark method as a legitimate, well-established alternative to the Angoff method, not a lesser or less rigorous approach
  • Ask which specific standard-setting method a testing program used and whether it was implemented with proper panelist training and documentation
  • Understand that the bookmark method specifically requires item response theory-based difficulty data already in place, a real prerequisite not every testing program has available
  • Focus evaluation on implementation rigor and documentation quality rather than which specific method was chosen, since both are legitimate approaches when properly executed

The bookmark method offers a genuinely different, well-established path to the same fundamental goal Angoff's method addresses — translating expert judgment about minimal competence into a specific, defensible cut score — and understanding both methods helps clarify that a cut score's legitimacy rests on the rigor of whichever defensible process was actually followed, not on any single universally correct methodology.

bookmark method standard settingcut score methodology comparisonitem response theory ordered bookletcertification bodiesAngoff method alternative