Skip to main content

Topic

Assessments & Testing

For psychometric researchers, skills-assessment companies, recruitment testers, and certification bodies.

Coaching Effects: When Test Prep Improves the Score More Than It Improves the Underlying Ability

A test score gain from targeted coaching on the test's specific format and question patterns isn't the same as a gain in the underlying ability the test exists to measure — and the two can diverge considerably.

Construct-Irrelevant Variance: When a Coding Test Measures Typing Speed Instead of Coding Skill

A candidate who codes well but types slowly, or who's unfamiliar with a specific IDE's quirks, can score poorly on a coding assessment for reasons that have nothing to do with coding ability.

Content Validity and Construct Validity Are Not the Same Thing, and Testing Vendors Conflate Them Constantly

"We had subject-matter experts confirm the content covers everything relevant" answers a narrower and easier question than "this test measures the actual ability we care about."

Criterion Contamination: When the Yardstick Used to Validate a Test Is Itself Corrupted

A validation study showing a test correlates well with supervisor ratings has only shown the test predicts supervisor ratings — which is a genuinely different and weaker claim than showing it predicts actual job performance.

Cut Scores Are a Policy Decision Dressed Up as a Statistical One

A cut score is often presented with the precision of a measurement — but standard-setting methods like the Angoff method are built on structured expert judgment, not a statistically derived threshold.

Differential Prediction and Differential Validity Are Two Distinct Fairness Questions, Not One

Showing a test's correlation with performance is similar across two groups answers a narrower question than showing the test predicts the identical performance level from the identical score across those same two groups.

Face Validity Is Not Validity: A Field Guide for Assessment Builders

Stakeholders judge an assessment by whether it feels credible. That judgment has almost no relationship to whether the assessment actually measures what it claims to.

Faking Good: Why Job Applicants Answer Personality Inventories Differently Than They'd Answer for Themselves

The same person answers an identical personality inventory differently depending on whether they're taking it for genuine self-insight or to get a job — and the job-context version is measurably more favorable.

Generalizability Theory: A More Complete Way to Ask Why Test Scores Vary Than a Single Reliability Coefficient

A single reliability coefficient tells you a test is reasonably consistent overall — it doesn't tell you whether the inconsistency that does exist comes mainly from different raters, different testing occasions, or different item sets.

Generalizability Theory: Splitting Test Error Into Its Actual Separate Sources

Knowing a test has a certain reliability coefficient tells you less than knowing specifically how much of its error comes from raters, how much from item sampling, and how much from occasion to occasion variability.

Guessing Correction: Why Penalizing Wrong Answers on a Multiple-Choice Test Is More Complicated Than It Sounds

A guessing-correction formula treats every unanswered or wrong response the same way, when test-takers actually arrive at those outcomes through very different mixes of knowledge and risk tolerance.

How Testing Organizations Actually Catch Cheating: Statistical Anomaly Detection, Not Just Proctors

Two candidates sharing an unusually high number of identical wrong answers, in an unusually similar pattern, is exactly the kind of statistical anomaly detection methods are built to catch — regardless of whether any proctor ever noticed anything.

Item Response Theory vs. Classical Test Theory: Why Modern Testing Programs Made the Switch

A test built under classical test theory can only be as good as the specific sample it was validated on — item response theory was developed specifically to escape that dependency.

Local Item Dependence: When Several Questions About the Same Reading Passage Aren't Actually Independent Evidence

Five questions about the same reading passage don't provide five independent pieces of evidence about a candidate's ability — getting the passage's main idea wrong tends to drag down performance on every related question together.

Predictive Validity and Concurrent Validity Answer Genuinely Different Questions, and Vendors Blur Them Constantly

A validity study run entirely on current employees, measured at a single point in time, is answering an easier and different question than the one a hiring decision actually needs answered.

Range Restriction: Why a Test's Real-World Validity Coefficient Looks Weaker Than It Actually Is

A validity study run on already-hired employees is missing the full range of scores the test would show across all applicants, and that missing range quietly deflates the reported validity coefficient.

Response Style Bias: Why Some People Consistently Pick the Extremes and Others Consistently Pick the Middle

A respondent's tendency to favor extreme or midpoint response options is often a stable personal style, independent of their actual attitude toward the specific content being asked about.

Speededness: When a Timed Test Measures Working Speed Instead of the Skill It Claims to Assess

A test can have excellent, well-validated items and still produce distorted results if a meaningful share of candidates never actually reach a substantial portion of them due to time constraints.

Standard Error of Measurement: The Number That Tells You How Much a Test Score Could Bounce Around

Two candidates scoring 78 and 82 on the same test might not be meaningfully different at all, once the standard error of measurement around each individual score is actually taken into account.

Test Equating: How a Score on Last Year's Exam Gets Made Comparable to a Score on This Year's

Two candidates who took the same certification exam two years apart, with genuinely different item sets, need their scores to mean the same thing — equating is the specific statistical work that makes that claim defensible.

Test-Retest Reliability Doesn't Mean What Most People Assume It Means

A psychometric instrument can have excellent test-retest reliability and still be measuring the wrong thing entirely, consistently — reliability and validity answer genuinely different questions.

The Barnum Effect: Why Vague Personality Descriptions Feel Eerily Accurate to Almost Everyone

A personality report full of statements vague and flattering enough to apply to almost anyone will feel personally accurate to nearly every reader — which says nothing about whether the underlying assessment measured anything real.

The Bookmark Method: A Different Way to Set a Passing Score Than the Classic Angoff Approach

The bookmark method asks expert judges to place a literal bookmark in a booklet of items already ordered from easiest to hardest, a genuinely different judgment task than the item-by-item probability estimates the Angoff method requires.

The Flynn Effect: Why the Same Raw Test Score Meant Something Different a Generation Ago

A test normed twenty years ago is quietly measuring against an outdated baseline, since average scores have shifted meaningfully over that time in most populations that have been studied.

The Four-Fifths Rule: A Useful Screening Tool for Adverse Impact That Gets Treated as a Legal Finish Line

The four-fifths rule is a practical, widely used screening threshold for adverse impact — treating it as a definitive legal safe harbor in either direction misunderstands what it was actually designed to do.

The Halo Effect in Performance Ratings: When One Strong Impression Colors Every Other Rated Dimension

A performance review where every competency score moves together, uniformly high or uniformly low, is itself a signal worth questioning — genuine performance is rarely that evenly distributed across a dozen distinct dimensions.

The Lake Wobegon Effect: When Every District Reports Above-Average Scores, Someone's Norms Are Stale

It's a statistical impossibility for every school district to be genuinely above the national average on the same norm-referenced test — when that's exactly what gets reported, stale norms are the far more likely explanation.

The Multitrait-Multimethod Matrix: How Psychometricians Actually Prove a Test Measures What It Claims To

A test's construct validity requires more than a single supportive correlation — it requires showing convergence with genuinely related measures and, just as importantly, divergence from measures it should be unrelated to.

Vertical Scaling: Putting Test Scores From Different Grade Levels on One Comparable Scale

A vertically scaled testing program can show whether a given student's growth from one grade to the next outpaced, matched, or lagged typical growth, a comparison raw grade-level scores alone can't support.

What Makes a Psychometric Assessment Legally Defensible

An assessment used for hiring or promotion decisions is a legal instrument, not just a questionnaire. Here's what that actually requires.

Why a Skills Test That Correlates with Job Performance Can Still Be Illegal to Use

The claim that a test predicts performance is a common and incomplete defense of a skills assessment — it answers the validity question but not the job-relatedness and business-necessity questions regulators actually ask.

Why a Test That Works Great in Validation Can Still Fail Once Everyone Knows the Answers

A test that was rigorously validated before launch doesn't stay validated indefinitely — item exposure degrades a test's usefulness in a way validation studies don't measure or catch.

Why a Test's Reliability Near the Passing Score Matters More Than Its Overall Average Reliability

A test can report a respectable overall average measurement error while still measuring quite imprecisely in the specific score region right around its actual pass-fail cut score, where precision matters most.

Why a Training Program's Self-Reported Pre-Post Improvement Can Be Partly an Illusion

A participant rating their own skill level before training doesn't yet know what genuine competence in that skill actually requires, making their pre-training self-rating a less informed and differently calibrated measure.

Why Grouping Test Items Around One Reading Passage Can Quietly Break a Key Statistical Assumption

Item response theory's standard models assume each item is answered independently, and a group of questions sharing one reading passage violates this assumption in a specific, well-documented way.

Why Professionals Delay Continuing Education Requirements Until the Very Last Possible Moment

Continuing education completion data shows a heavy, predictable concentration in the final weeks before a certification's renewal deadline, directly consistent with how the goal gradient hypothesis predicts motivation actually develops.