A recruitment test is rigorously validated before launch — a proper predictive validity study, a defensible adverse impact analysis, solid psychometric properties. Two years later, the same test is still officially validated, still cited in the same terms, and quietly measuring something closer to who had access to a shared answer key than the skill it was originally built to assess. The validation study measured the test's properties at one point in time. Nothing about how that study was conducted continues to guarantee those properties hold as item exposure accumulates.
A validation study is a snapshot, not an ongoing guarantee
Predictive validity is established against a specific group of candidates taking a specific set of items for, functionally, the first time — none of them had prior exposure to those exact questions, because the test was new. This condition is baked into the validation study's design and is exactly the condition that erodes as the same test is administered to more and more candidates over time, some meaningful share of whom will have encountered specific items beforehand through shared prep materials, forums, or informal networks among applicants.
Item exposure changes what the test is actually measuring
Once specific items and their correct answers circulate informally among candidates, the test increasingly measures which candidates had access to that information rather than the underlying skill or ability it was designed to assess — a candidate who memorized a leaked answer key will score well regardless of whether they possess the ability the test claims to measure. This isn't a hypothetical edge case; item exposure through informal candidate networks, online forums, and test-prep communities is a well-documented, ordinary phenomenon for any test used repeatedly on the same general candidate population over an extended period.
The test can look validated on paper long after this has happened
Nothing about accumulating item exposure automatically triggers a new validation study or changes the test's official documentation — a company or vendor citing the original validity coefficient two or three years after launch is citing a number established under conditions that increasingly don't reflect how the test is actually being taken by that point. This is precisely why the decay is dangerous: it doesn't show up as an obvious failure, it shows up as a slow, invisible drift between the test's paper credentials and what it's actually measuring in practice.
What actually slows this decline
Item rotation — regularly retiring exposed items and introducing new, unexposed ones, validated against the same construct — directly addresses the mechanism causing the decay, rather than hoping exposure doesn't accumulate. Monitoring for unusual response patterns (implausibly fast completion times, response patterns matching known leaked answer sets) can flag when exposure has become a meaningful problem for a specific item pool. Periodic re-validation, not treated as a one-time launch requirement but as an ongoing practice, is what actually confirms a test still measures what it claims to measure after years of real-world use.
What this means for anyone relying on a recruitment test
- Ask when the test's items were last rotated, not just when the original validation study was conducted
- Treat a validity coefficient from a launch study as time-bound evidence, not a permanent property of the test
- Look for evidence of exposure monitoring or unusual response-pattern detection as part of the vendor's ongoing test-security practice
- Be specifically cautious with tests used unchanged for long periods against large, overlapping candidate populations, where item exposure has the most opportunity to accumulate
A test's validity isn't a fact established once and then permanently true — it's a property that has to be actively maintained against a very ordinary, very human process of information spreading among the people being tested.