A certification body setting a passing score for a new exam has a genuine choice between standard-setting methodologies, including the widely used Angoff method discussed elsewhere and an alternative called the bookmark method, which asks expert judges to work through a specially ordered booklet of test items, arranged from easiest to hardest based on item response theory difficulty estimates, and place a literal bookmark at the point where they believe a minimally competent candidate would begin to struggle.
How the bookmark method's specific judgment task differs from Angoff's
Rather than asking judges to estimate, item by item, the specific probability that a minimally competent candidate would answer each individual item correctly — the Angoff method's core task — the bookmark method presents items already ordered by their actual, empirically estimated difficulty and asks judges to identify a single transition point in that ordered sequence, a genuinely different cognitive task that some standard-setting researchers argue is more natural and intuitive for expert panelists to perform reliably.
Why the item ordering itself requires item response theory data
The bookmark method's ordered-booklet format depends on having genuine, empirically estimated item difficulty parameters available in advance, typically derived from item response theory modeling of prior candidate response data — meaning the bookmark method specifically requires this kind of psychometric infrastructure already in place, a requirement the Angoff method's item-by-item judgment task doesn't share in the same way.
Why some standard-setting researchers consider the bookmark method's task more intuitive
Estimating a specific numeric probability for each individual item, as the Angoff method requires, is a demanding and somewhat unnatural cognitive task for many expert judges, who may not have strong, well-calibrated intuitions for translating their sense of an item's difficulty into a precise probability estimate — identifying a single transition point in an already-difficulty-ordered sequence of items is argued by proponents to more closely match how judges naturally think about a minimally competent candidate's likely performance, without requiring the same kind of precise numeric probability estimation.
Why the choice between these methods remains a genuine methodological decision, not a settled preference
Both methods have documented strengths and limitations, and the broader standard-setting research literature doesn't present either as simply, universally superior to the other — the Angoff method's item-by-item structure provides more granular diagnostic information about specific items, while the bookmark method's ordered-booklet format may better suit judges' natural intuitions, and the appropriate choice depends on a testing program's specific available data, judge training resources, and practical constraints.
Why understanding this choice matters for anyone evaluating a certification program's standard-setting rigor
A testing program that has thoughtfully selected and properly implemented either the Angoff or bookmark method, with well-trained expert panelists and a defensible process, demonstrates a genuinely rigorous approach to standard setting — the specific method chosen matters less than whether it was implemented with genuine methodological rigor, adequate panelist training, and appropriate documentation of the process actually followed.
What this means for evaluating a testing program's standard-setting methodology
- Recognize the bookmark method as a legitimate, well-established alternative to the Angoff method, not a lesser or less rigorous approach
- Ask which specific standard-setting method a testing program used and whether it was implemented with proper panelist training and documentation
- Understand that the bookmark method specifically requires item response theory-based difficulty data already in place, a real prerequisite not every testing program has available
- Focus evaluation on implementation rigor and documentation quality rather than which specific method was chosen, since both are legitimate approaches when properly executed
The bookmark method offers a genuinely different, well-established path to the same fundamental goal Angoff's method addresses — translating expert judgment about minimal competence into a specific, defensible cut score — and understanding both methods helps clarify that a cut score's legitimacy rests on the rigor of whichever defensible process was actually followed, not on any single universally correct methodology.