A pre-employment test is validated by correlating its scores against supervisor performance ratings for a sample of current employees, producing a strong, statistically significant correlation presented as evidence the test predicts job performance. This validation is only as sound as the criterion it's measured against — and supervisor performance ratings, the most commonly used criterion in this kind of study, are themselves well documented to be affected by a range of rating biases unrelated to genuine job performance, a problem called criterion contamination that can undermine an otherwise carefully conducted validation study.
What criterion contamination specifically describes
A validation study's entire logic depends on the criterion measure — the outcome the test is being validated against — genuinely and accurately representing the underlying construct of interest, typically actual job performance. Criterion contamination occurs when that criterion measure is influenced by factors unrelated to the genuine underlying construct, meaning a test correlating strongly with the contaminated criterion is demonstrating something other than what the validation study claims to be showing.
Why supervisor ratings specifically carry this risk
Supervisor performance ratings are documented, across substantial research, to be influenced by halo effects (an overall favorable impression bleeding into every specific rated dimension), recency bias (recent performance disproportionately shaping the rating), and similar-to-me bias (rating employees who share the supervisor's own background or communication style more favorably) — none of which reflect genuine underlying job performance, and all of which can introduce systematic noise or bias into a criterion measure built entirely from these ratings.
Why this specifically threatens the validation study's core claim
If a test happens to correlate with the same underlying factors that bias supervisor ratings — for instance, if a personality test measures traits that happen to correlate with the kind of interpersonal similarity that drives similar-to-me bias in raters — the test can show a strong correlation with the contaminated criterion that has little to do with the test's actual ability to predict genuine job performance, and everything to do with both the test and the criterion being influenced by the same irrelevant, contaminating factor.
What actually reduces this risk in validation research
Using multiple, independent performance criteria — objective output measures where available, alongside ratings from multiple different supervisors or peers rather than a single rater — reduces the influence any single contaminated source can have on the overall validation result, since genuine job performance should show up consistently across multiple, independently measured criteria, while contamination specific to one rater or one measure is less likely to appear consistently across all of them. Training raters specifically on common rating biases, and using structured, behaviorally anchored rating scales rather than open-ended overall impressions, can reduce (though not eliminate) the contamination present in any single criterion measure used.
What this means for evaluating a test validation study
- Check what criterion was actually used to validate a test, and specifically whether it was a single supervisor rating or a more robust, multi-source performance measure
- Be specifically skeptical of validation studies relying entirely on a single rater's overall impression as the criterion, given the well-documented biases affecting this kind of measure
- Look for validation studies using objective performance data alongside or instead of purely subjective ratings, where the nature of the job allows for it
- Recognize that a strong correlation with a contaminated criterion demonstrates the test predicts that specific, flawed measure — a genuinely different and weaker claim than demonstrating it predicts true underlying job performance
A validation study is only as trustworthy as the criterion it's built on — and when that criterion is a single supervisor's subjective rating, well-documented rating biases mean the validation may be demonstrating something considerably narrower and less useful than what it's typically presented as showing.