A classical test theory reliability coefficient reports that a performance assessment has a certain overall level of measurement error, without distinguishing whether that error comes primarily from inconsistency between different raters, from the specific items or tasks sampled, from the specific occasion the assessment was administered, or from some combination of all three — a genuine limitation that generalizability theory, a statistical framework decomposing measurement error into its separately estimated component sources, was specifically developed to address.
What classical test theory's single reliability coefficient actually obscures
Classical test theory treats all measurement error as a single undifferentiated quantity, producing one reliability coefficient summarizing overall measurement consistency — this single number can't distinguish whether error stems mainly from raters scoring inconsistently, from the specific sample of items or tasks used, from occasion-to-occasion variability, or from some specific combination, even though these different sources of error call for genuinely different practical remedies.
How generalizability theory's decomposition approach actually works
Generalizability theory uses a statistical design — typically a generalizability study varying multiple facets like rater, item, and occasion systematically across the same examinees — to separately estimate how much of total measurement error is attributable to each specific facet and to the interactions between them, producing a genuinely more detailed picture of where a test's measurement inconsistency is actually coming from, rather than a single undifferentiated summary number.
Why this decomposition carries real, practical diagnostic value
If a generalizability study finds that rater variance is the dominant source of measurement error in a specific performance assessment, the practical fix is improved rater training, calibration, or increasing the number of raters scoring each performance — if item variance dominates instead, the fix is a larger or more carefully constructed item pool. A single classical reliability coefficient can't distinguish which of these genuinely different fixes is actually needed, while a generalizability theory decomposition points specifically toward the source requiring practical attention.
Why this is particularly valuable for performance assessments involving human raters
Performance assessments — essay scoring, clinical skills evaluations, structured interview ratings — involve rater judgment as a genuine, distinct potential source of measurement error that multiple-choice testing formats don't share in the same way, making the ability to specifically isolate and quantify rater variance, separate from item and occasion variance, especially valuable for exactly this kind of rater-dependent assessment format.
Why a decision study can then optimize an assessment's actual design using this decomposed information
Once a generalizability study has decomposed measurement error into its component sources, a related decision study can use these estimates to model how reliability would change under different practical design choices — using more raters per performance, more items per assessment, more testing occasions — allowing an assessment program to make an evidence-based, cost-aware decision about where to invest limited resources for the most meaningful actual gain in overall measurement reliability.
What this means for evaluating or designing rigorous performance-based assessment programs
- Look for evidence of a generalizability theory decomposition, not just a single classical reliability coefficient, for any assessment involving rater judgment
- Ask specifically which error source — rater, item, occasion — a testing program's decomposition identifies as dominant, and whether the program has addressed it accordingly
- Recognize that a high classical reliability coefficient alone doesn't reveal whether an assessment is robust across raters specifically, a distinct question generalizability theory directly addresses
- Treat generalizability theory as the more diagnostically useful framework specifically for performance assessments and other rater-dependent testing formats
Generalizability theory doesn't produce a fundamentally different verdict about whether a test is reliable — it produces considerably more actionable information about exactly why a test's reliability sits where it does, which specific source of measurement inconsistency deserves practical attention, and where limited improvement resources should actually be spent.