Skip to main content
Assessments & Testing

Test-Retest Reliability Doesn't Mean What Most People Assume It Means

A high test-retest correlation shows a measure produces consistent scores over time. It says nothing, on its own, about whether those consistent scores are actually accurate.

Key Takeaways
  • Test-retest reliability measures whether a test produces consistent scores when the same person takes it again — it's a measure of consistency, not accuracy
  • A test can have excellent test-retest reliability while measuring something entirely different from what it claims to measure, as long as it measures that wrong thing consistently
  • Reliability is a necessary but far from sufficient condition for validity — a test must be reliable to be valid, but being reliable doesn't make it valid
  • Evaluating a psychometric instrument requires checking both properties separately, since a strong report on one says very little about the other

A psychometric instrument reports a test-retest reliability coefficient of 0.88 — a genuinely strong number, indicating that people who take the test twice, a few weeks apart, tend to receive very similar scores both times. This is presented, often implicitly, as evidence the test is a good, trustworthy measure. It's evidence the test is consistent. Whether it's actually measuring what it claims to measure is an entirely separate question that a reliability coefficient, however strong, cannot answer on its own.

What test-retest reliability actually measures

Test-retest reliability quantifies the correlation between a person's scores on two separate administrations of the same test, typically separated by a period short enough that the underlying trait being measured shouldn't have genuinely changed. A high coefficient means the test produces stable, repeatable scores for the same person over that period — a genuinely important and necessary property for any measurement instrument, since a test that gives wildly different results to the same person on different days would be useless regardless of anything else about it.

Why consistency and accuracy are completely separate properties

A broken bathroom scale that reliably reads exactly ten pounds heavier than a person's actual weight, every single time they step on it, has excellent test-retest reliability — the same person gets the same (wrong) reading consistently — and is nonetheless a systematically inaccurate measure of weight. This isn't a hypothetical edge case invented to make a point — it's a precise illustration of the actual logical relationship between reliability and validity: a measure can be perfectly consistent while consistently measuring the wrong thing, or measuring the right thing with a consistent, systematic error.

How this plays out with real psychometric instruments

A personality or aptitude instrument can achieve strong test-retest reliability simply because respondents tend to answer self-report items consistently with their own prior answers, or because certain response styles (a tendency to agree, a tendency toward extreme responses) are themselves stable traits being inadvertently measured — none of which requires the instrument to actually be measuring the underlying construct it claims to assess. A test can be highly reliable specifically because it's measuring something stable and real about the respondent, while that something isn't the construct the test was designed and marketed to measure at all.

Reliability is necessary but not sufficient for validity

An unreliable test — one producing wildly inconsistent scores for the same person — cannot be valid, because validity requires the test to consistently measure the intended construct, and an unreliable test isn't measuring anything consistently in the first place. This one-directional relationship is where the common confusion comes from: reliability is a genuine precondition for validity, which makes it tempting to treat a strong reliability coefficient as if it were most of the way to establishing validity, when it's actually establishing something logically prior and considerably less demanding.

What actually needs to be checked, separately, before trusting an instrument

  • Confirm test-retest reliability as a baseline requirement, but never treat it as evidence of validity on its own
  • Look specifically for criterion validity evidence — does the test's score actually correlate with independent, real-world evidence of the construct it claims to measure
  • Look for construct validity evidence — does the test relate to other measures the way theory would predict it should, and fail to relate to things it theoretically shouldn't
  • Be specifically skeptical of any psychometric instrument marketed primarily on the strength of its reliability coefficient alone, without separate validity evidence presented

A reliable test is a test you can trust to give you the same answer twice. A valid test is a test giving you the right answer in the first place. Neither property substitutes for the other, and a well-marketed reliability number is often doing work that only a validity study can actually justify.

test-retest reliabilitypsychometric validitymeasurement consistencypsychometric researchersassessment reliability