A test with a reliability coefficient of 0.90 — a genuinely strong figure by most standards — is used to compare two candidates who scored 78 and 82 respectively, with the four-point difference treated as a meaningful distinction between them. Translating that same 0.90 reliability coefficient into a standard error of measurement for this specific test often reveals that a four-point gap falls well within the range of measurement error expected for any individual score, meaning the two candidates may not differ meaningfully in underlying ability at all, despite the test's strong overall reliability rating.
Why reliability and standard error of measurement answer different practical questions
A reliability coefficient is a single number summarizing how consistently a test ranks an entire population of test-takers relative to each other across repeated administrations — a genuinely useful, but population-level, summary statistic. The standard error of measurement translates that same reliability information into a specific, practically usable estimate of how much random measurement error is likely present around any one individual's particular score, answering the more directly relevant question a practitioner interpreting an individual result actually needs answered: how much could this specific person's observed score plausibly differ from their true underlying ability, just due to ordinary measurement imprecision.
How the two relate mathematically, in practical terms
The standard error of measurement is calculated directly from a test's reliability coefficient and the standard deviation of scores in the population, and it can be used to construct a practical confidence band around any individual score — a commonly used rule of thumb suggests that two scores differing by less than roughly two standard errors of measurement shouldn't be treated as a reliably real, meaningful difference, since a gap that small falls within the range plausibly attributable to ordinary measurement error alone.
Why this specifically matters for close-score decisions
A hiring or admissions decision, or a certification pass/fail cutoff, that treats a small gap between two individual scores as a meaningful, decision-relevant difference is often making a distinction the test's own precision doesn't actually support — a four-point gap between two candidates on a test with a standard error of measurement of three points, for instance, is well within the range where the true underlying difference between the two candidates could plausibly be zero, or could even run in the opposite direction from what the observed scores suggest.
Why a single, precise-looking number can mislead more than a confidence band would
Reporting a test score as a single precise number implicitly presents more certainty than the test's own measurement properties actually justify, while reporting the same score alongside its standard error of measurement, or as an explicit range or confidence band, communicates the genuine uncertainty involved directly and transparently, which is a more honest and more practically useful way to convey what an individual test result actually tells you.
What this means for interpreting and reporting individual test scores
- Calculate and report the standard error of measurement alongside any individual test score, not just the overall reliability coefficient for the test as a whole
- Treat score differences smaller than roughly two standard errors of measurement as not reliably meaningful, particularly in close-call decisions
- Communicate individual scores as a range or confidence band where the decision context allows, rather than as a single, falsely precise number
- Be specifically cautious of high-stakes cutoff decisions (pass/fail, hire/no-hire) that fall within a standard error of measurement's width of the actual cutoff score
A test's reliability coefficient tells you it's a generally trustworthy instrument across a population — the standard error of measurement is what actually tells you how much to trust, or distrust, the precision of any one specific number that instrument produces.