Skip to main content
Assessments & Testing

Why Grouping Test Items Around One Reading Passage Can Quietly Break a Key Statistical Assumption

Several questions attached to the same reading passage or shared stimulus don't behave like independent items — getting one wrong measurably raises the chance of getting the others wrong too, beyond what ability alone predicts.

Key Takeaways
  • Standard item response theory models assume local independence — that a test-taker's response to any given item depends only on their underlying ability, not on their response to any other specific item
  • A group of items sharing a common reading passage or other stimulus, called a testlet, often violates this assumption, since getting one item in the group wrong tends to raise the chance of getting other items in the same group wrong too, beyond what ability alone predicts
  • This local item dependence, sometimes called the testlet effect, happens because shared passage comprehension or other factors specific to that particular stimulus create a common source of correlated performance across the grouped items
  • Testlet-specific psychometric models account for this dependence directly, while ignoring it and applying standard independence-assuming models to testlet-based items can produce measurably overstated test reliability and precision

A reading comprehension test presents test-takers with a passage followed by five separate questions about that passage, and psychometric analysis reveals that a test-taker who misses one of these five items is measurably more likely to miss the others in the same group too, beyond what their overall estimated ability level alone would predict — a specific, well-documented violation of local independence, a foundational assumption behind standard item response theory models, called the testlet effect.

What the local independence assumption actually requires

Standard item response theory models assume that once a test-taker's underlying ability level is accounted for, their response to any given item is statistically independent of their response to any other specific item — meaning ability alone should fully explain any correlation between how a test-taker performs on different items, with no leftover, item-specific correlation remaining once ability is properly accounted for.

Why a group of items sharing a common passage tends to violate this assumption

Items grouped around a shared reading passage, shared data set, or other common stimulus can share sources of correlated performance beyond the test-taker's general underlying ability — a test-taker's specific comprehension of that particular passage, their familiarity with its particular subject matter, or their success navigating that specific stimulus's particular format all create a shared, passage-specific performance factor that standard local independence assumes shouldn't exist beyond general ability alone.

Why this specific dependence is called the testlet effect

A group of items sharing a common stimulus and exhibiting this kind of correlated, non-independent performance is called a testlet, and the resulting violation of local independence is called the testlet effect — a well-documented and specifically named phenomenon in psychometric research precisely because it shows up so consistently and predictably whenever items are grouped around any kind of shared common stimulus.

Why ignoring this dependence and applying standard independence-assuming models produces overstated reliability

Standard item response theory models, when applied to testlet-based items without accounting for their actual local dependence, treat the correlated performance within a testlet as if it reflects genuine additional independent information about the test-taker's ability, when it actually partly reflects a redundant, shared passage-specific effect — this leads these standard models to overstate a test's actual measurement precision and reliability, since the testlet's items aren't contributing as much genuinely independent information as the model assumes.

How testlet-specific psychometric models address this directly

Specialized testlet response models incorporate an additional parameter or structure specifically accounting for the shared, testlet-specific dependence among grouped items, properly separating the genuinely independent ability-related information each item contributes from the redundant, shared passage-specific effect — producing considerably more accurate estimates of a test's actual measurement precision than standard models that ignore this dependence entirely.

What this means for evaluating tests that group items around shared passages or stimuli

  • Ask specifically whether a testing program has evaluated and modeled local item dependence for any testlet-structured items, such as passage-based reading questions
  • Be skeptical of reliability estimates for testlet-based assessments that were calculated using standard models assuming full local independence
  • Recognize the testlet effect as a well-documented, specifically named phenomenon, not a rare or unusual psychometric edge case
  • Look for evidence of testlet-specific modeling approaches in the technical documentation of any assessment using passage-based or similarly grouped item structures

The testlet effect is a genuine, well-documented reminder that a test's actual item structure matters directly to whether foundational psychometric assumptions actually hold — items that look independent on the surface can share hidden, common sources of correlated performance that a rigorous measurement model needs to explicitly account for, not simply assume away.

local item dependence testingtestlet effect psychometricsshared passage item clusteringpsychometriciansitem response theory assumptions