Topic
Assessments & Testing
For psychometric researchers, skills-assessment companies, recruitment testers, and certification bodies.
A test score gain from targeted coaching on the test's specific format and question patterns isn't the same as a gain in the underlying ability the test exists to measure — and the two can diverge considerably.
A candidate who codes well but types slowly, or who's unfamiliar with a specific IDE's quirks, can score poorly on a coding assessment for reasons that have nothing to do with coding ability.
"We had subject-matter experts confirm the content covers everything relevant" answers a narrower and easier question than "this test measures the actual ability we care about."
A validation study showing a test correlates well with supervisor ratings has only shown the test predicts supervisor ratings — which is a genuinely different and weaker claim than showing it predicts actual job performance.
A cut score is often presented with the precision of a measurement — but standard-setting methods like the Angoff method are built on structured expert judgment, not a statistically derived threshold.
Showing a test's correlation with performance is similar across two groups answers a narrower question than showing the test predicts the identical performance level from the identical score across those same two groups.
Stakeholders judge an assessment by whether it feels credible. That judgment has almost no relationship to whether the assessment actually measures what it claims to.
The same person answers an identical personality inventory differently depending on whether they're taking it for genuine self-insight or to get a job — and the job-context version is measurably more favorable.
A single reliability coefficient tells you a test is reasonably consistent overall — it doesn't tell you whether the inconsistency that does exist comes mainly from different raters, different testing occasions, or different item sets.
Knowing a test has a certain reliability coefficient tells you less than knowing specifically how much of its error comes from raters, how much from item sampling, and how much from occasion to occasion variability.
A guessing-correction formula treats every unanswered or wrong response the same way, when test-takers actually arrive at those outcomes through very different mixes of knowledge and risk tolerance.
Two candidates sharing an unusually high number of identical wrong answers, in an unusually similar pattern, is exactly the kind of statistical anomaly detection methods are built to catch — regardless of whether any proctor ever noticed anything.
A test built under classical test theory can only be as good as the specific sample it was validated on — item response theory was developed specifically to escape that dependency.
Five questions about the same reading passage don't provide five independent pieces of evidence about a candidate's ability — getting the passage's main idea wrong tends to drag down performance on every related question together.
A validity study run entirely on current employees, measured at a single point in time, is answering an easier and different question than the one a hiring decision actually needs answered.
A validity study run on already-hired employees is missing the full range of scores the test would show across all applicants, and that missing range quietly deflates the reported validity coefficient.
A respondent's tendency to favor extreme or midpoint response options is often a stable personal style, independent of their actual attitude toward the specific content being asked about.
A test can have excellent, well-validated items and still produce distorted results if a meaningful share of candidates never actually reach a substantial portion of them due to time constraints.
Two candidates scoring 78 and 82 on the same test might not be meaningfully different at all, once the standard error of measurement around each individual score is actually taken into account.
Two candidates who took the same certification exam two years apart, with genuinely different item sets, need their scores to mean the same thing — equating is the specific statistical work that makes that claim defensible.
A psychometric instrument can have excellent test-retest reliability and still be measuring the wrong thing entirely, consistently — reliability and validity answer genuinely different questions.
A personality report full of statements vague and flattering enough to apply to almost anyone will feel personally accurate to nearly every reader — which says nothing about whether the underlying assessment measured anything real.
The bookmark method asks expert judges to place a literal bookmark in a booklet of items already ordered from easiest to hardest, a genuinely different judgment task than the item-by-item probability estimates the Angoff method requires.
A test normed twenty years ago is quietly measuring against an outdated baseline, since average scores have shifted meaningfully over that time in most populations that have been studied.
The four-fifths rule is a practical, widely used screening threshold for adverse impact — treating it as a definitive legal safe harbor in either direction misunderstands what it was actually designed to do.
A performance review where every competency score moves together, uniformly high or uniformly low, is itself a signal worth questioning — genuine performance is rarely that evenly distributed across a dozen distinct dimensions.
It's a statistical impossibility for every school district to be genuinely above the national average on the same norm-referenced test — when that's exactly what gets reported, stale norms are the far more likely explanation.
A test's construct validity requires more than a single supportive correlation — it requires showing convergence with genuinely related measures and, just as importantly, divergence from measures it should be unrelated to.
A vertically scaled testing program can show whether a given student's growth from one grade to the next outpaced, matched, or lagged typical growth, a comparison raw grade-level scores alone can't support.
An assessment used for hiring or promotion decisions is a legal instrument, not just a questionnaire. Here's what that actually requires.
The claim that a test predicts performance is a common and incomplete defense of a skills assessment — it answers the validity question but not the job-relatedness and business-necessity questions regulators actually ask.
A test that was rigorously validated before launch doesn't stay validated indefinitely — item exposure degrades a test's usefulness in a way validation studies don't measure or catch.
A test can report a respectable overall average measurement error while still measuring quite imprecisely in the specific score region right around its actual pass-fail cut score, where precision matters most.
A participant rating their own skill level before training doesn't yet know what genuine competence in that skill actually requires, making their pre-training self-rating a less informed and differently calibrated measure.
Item response theory's standard models assume each item is answered independently, and a group of questions sharing one reading passage violates this assumption in a specific, well-documented way.
Continuing education completion data shows a heavy, predictable concentration in the final weeks before a certification's renewal deadline, directly consistent with how the goal gradient hypothesis predicts motivation actually develops.