A reading comprehension test presents test-takers with a passage followed by five separate questions about that passage, and psychometric analysis reveals that a test-taker who misses one of these five items is measurably more likely to miss the others in the same group too, beyond what their overall estimated ability level alone would predict — a specific, well-documented violation of local independence, a foundational assumption behind standard item response theory models, called the testlet effect.
What the local independence assumption actually requires
Standard item response theory models assume that once a test-taker's underlying ability level is accounted for, their response to any given item is statistically independent of their response to any other specific item — meaning ability alone should fully explain any correlation between how a test-taker performs on different items, with no leftover, item-specific correlation remaining once ability is properly accounted for.
Why a group of items sharing a common passage tends to violate this assumption
Items grouped around a shared reading passage, shared data set, or other common stimulus can share sources of correlated performance beyond the test-taker's general underlying ability — a test-taker's specific comprehension of that particular passage, their familiarity with its particular subject matter, or their success navigating that specific stimulus's particular format all create a shared, passage-specific performance factor that standard local independence assumes shouldn't exist beyond general ability alone.
Why this specific dependence is called the testlet effect
A group of items sharing a common stimulus and exhibiting this kind of correlated, non-independent performance is called a testlet, and the resulting violation of local independence is called the testlet effect — a well-documented and specifically named phenomenon in psychometric research precisely because it shows up so consistently and predictably whenever items are grouped around any kind of shared common stimulus.
Why ignoring this dependence and applying standard independence-assuming models produces overstated reliability
Standard item response theory models, when applied to testlet-based items without accounting for their actual local dependence, treat the correlated performance within a testlet as if it reflects genuine additional independent information about the test-taker's ability, when it actually partly reflects a redundant, shared passage-specific effect — this leads these standard models to overstate a test's actual measurement precision and reliability, since the testlet's items aren't contributing as much genuinely independent information as the model assumes.
How testlet-specific psychometric models address this directly
Specialized testlet response models incorporate an additional parameter or structure specifically accounting for the shared, testlet-specific dependence among grouped items, properly separating the genuinely independent ability-related information each item contributes from the redundant, shared passage-specific effect — producing considerably more accurate estimates of a test's actual measurement precision than standard models that ignore this dependence entirely.
What this means for evaluating tests that group items around shared passages or stimuli
- Ask specifically whether a testing program has evaluated and modeled local item dependence for any testlet-structured items, such as passage-based reading questions
- Be skeptical of reliability estimates for testlet-based assessments that were calculated using standard models assuming full local independence
- Recognize the testlet effect as a well-documented, specifically named phenomenon, not a rare or unusual psychometric edge case
- Look for evidence of testlet-specific modeling approaches in the technical documentation of any assessment using passage-based or similarly grouped item structures
The testlet effect is a genuine, well-documented reminder that a test's actual item structure matters directly to whether foundational psychometric assumptions actually hold — items that look independent on the surface can share hidden, common sources of correlated performance that a rigorous measurement model needs to explicitly account for, not simply assume away.