A reading comprehension test includes five separate questions, all based on the same shared reading passage, each scored and initially treated in the underlying statistical model as though it provided independent evidence of the candidate's ability, alongside every other item on the test. This violates a specific assumption most standard psychometric models rely on — that each item's performance is statistically independent given the candidate's underlying ability — since a candidate who fundamentally misunderstands a specific aspect of the shared passage is likely to perform poorly on multiple related questions together, not independently across each one.
What local item dependence specifically describes
Local item dependence occurs when items sharing a common stimulus, context, or testlet structure show correlated performance beyond what the underlying ability model, assuming full item independence, would predict — a candidate's specific misunderstanding of one aspect of a shared reading passage can plausibly affect their performance on several separate questions tied to that same passage simultaneously, producing a genuine statistical dependence among those items that a model assuming full independence doesn't account for.
Why this specifically inflates apparent score precision when ignored
Standard reliability and precision calculations, built on the assumption that each item contributes independent evidence, treat five items from the same testlet as though they were five fully independent sources of information about a candidate's ability — when local item dependence is actually present, these five items are providing meaningfully less than five independent pieces of evidence, since some of their correlated variance reflects the shared stimulus's specific effect rather than five genuinely separate, independent measurements of the underlying ability.
Why this produces an overstated sense of the test's actual statistical reliability
A test's calculated reliability coefficient, computed under the assumption that all items are independent, will appear higher than the test's true, effective reliability once local item dependence within testlets is properly accounted for — meaning tests built substantially from testlets, without correcting for this dependence, can present themselves as more statistically precise and reliable than they actually are, a genuine measurement problem distinct from and in addition to any concern about the test's actual content validity.
What actually addresses this specific problem
Testlet response theory, a specialized extension of standard item response theory models developed specifically to account for shared local dependence within a testlet, models the testlet as a unit alongside individual item parameters, producing more accurate estimates of the test's true, effective reliability and score precision than a standard model assuming full item independence would provide. Treating an entire testlet as a single scored unit, rather than as several separately scored independent items, is a simpler, more conservative alternative that avoids the dependence problem by not treating the testlet's items as independent in the first place.
What this means for evaluating and designing tests that include testlets
- Ask whether a test's reliability and precision estimates account for local item dependence within any testlets, or assume full item independence throughout
- Recognize testlets — items sharing a reading passage, a case study, a shared data set — as a specific, common source of this statistical dependence problem
- Favor testlet response theory or unit-scoring approaches for tests built substantially from testlets, over standard models assuming full item independence
- Be specifically skeptical of a test's stated reliability if it includes substantial testlet content without any stated adjustment for local item dependence
Local item dependence is a specific, well-documented statistical issue that's easy to overlook precisely because testlets feel intuitively reasonable as an assessment format — the problem isn't the testlet structure itself, it's treating each item within it as though it provided the same independent evidence a genuinely separate, unrelated item would.