An assessment built to screen job candidates or evaluate employees for promotion is, legally, an employment test — whether it's a 200-item psychometric instrument or a ten-question scorecard someone built last quarter. That status brings requirements most builders of quick internal assessments never learn until a claim is filed.
Reliability: does it give the same answer twice?
Reliability measures consistency: would the same person, assessed under similar conditions, get a similar score? An assessment with poor reliability (commonly measured via internal consistency, such as Cronbach's alpha, or test-retest correlation) is unstable enough that a candidate's outcome partly reflects noise rather than the trait being measured. Low reliability alone doesn't just weaken the assessment's usefulness — it undermines any validity claim built on top of it, since a measure can't validly predict an outcome if it can't consistently measure anything in the first place.
Validity: does it measure what it claims to?
Validity is the harder, more consequential property, and it comes in a few recognized forms relevant to employment testing:
- Content validity — does the assessment's content actually sample the skills or knowledge required for the job, rather than tangentially related material?
- Criterion-related validity — do scores actually correlate with a real job-performance outcome, demonstrated with data, not assumption?
- Construct validity — does the assessment measure the underlying trait it claims to (e.g., "leadership potential"), rather than something else that happens to correlate with it?
An assessment can be reliable without being valid — consistently measuring the wrong thing is still consistent. Validity is what connects the score to the decision being made.
Adverse impact: intent doesn't matter, outcome does
Under the framework used by U.S. employment regulators (and echoed in similar forms elsewhere), an assessment can be legally challenged based purely on its outcome pattern, regardless of intent. The commonly cited rule of thumb is the "four-fifths rule": if one group's selection rate is less than 80% of the rate for the highest-selected group, that's a signal of adverse impact worth investigating, even if no single question appears discriminatory. An assessment with strong content validity can still generate adverse impact if it wasn't checked for it.
What this means practically for anyone building an assessment
- Pilot the assessment on a real sample before using it for actual decisions, and check reliability statistics before trusting any score it produces
- Tie assessment content directly and demonstrably to the job requirements, not to generic "soft skill" categories
- Run an adverse impact check across demographic groups before deployment, not after a complaint
- Keep documentation of the validation process itself — in a dispute, the paper trail showing due diligence matters as much as the statistics
None of this is optional once an assessment starts influencing who gets hired, promoted, or let go. The rigor is the product, not overhead on top of it.