Skip to main content
Education & Training

The Modality Effect: Why Pairing Narration With Visuals Beats Pairing On-Screen Text With Visuals

Splitting information across the ears and eyes simultaneously — spoken narration plus a visual — uses working memory more efficiently than routing both text and visual information through the eyes alone.

Key Takeaways
  • The modality effect describes how pairing spoken narration with a visual produces better learning than pairing the equivalent information as on-screen text with the same visual
  • This happens because narration and visuals are processed through separate channels, auditory and visual, while on-screen text and a visual both compete for the same limited visual-processing channel
  • This is distinct from the redundancy effect, which concerns identical information presented simultaneously through two channels — the modality effect concerns complementary, non-redundant information split across channels
  • Applying this effect well means using narration to explain a visual rather than on-screen text, specifically when the two elements are meant to be processed together rather than sequentially

A complex diagram explained through spoken narration produces better comprehension and retention than the identical diagram explained through on-screen text, even when the actual words used are nearly identical between the two versions — a pattern called the modality effect, explained by how narration and visual information get processed through separate channels in working memory, while on-screen text and a visual both compete for the same limited visual-processing channel simultaneously.

Why routing information through two separate channels is more efficient

Working memory research distinguishes between a visual-spatial processing channel and a separate auditory-verbal processing channel, each with its own limited capacity — when a visual is explained through spoken narration, the visual information and the explanatory information are processed through these two separate channels simultaneously, each within its own capacity, rather than both competing for the single visual channel's limited capacity, which is exactly what happens when the same explanatory information is presented as on-screen text alongside the visual.

Why on-screen text specifically creates this channel competition

A learner trying to simultaneously read on-screen text and examine an accompanying diagram has to split visual attention between the two, or resolve them sequentially rather than truly simultaneously, since both are competing for the same visual-processing capacity — narration removes this specific competition entirely, since the explanatory content no longer needs to occupy any of the visual channel's capacity, freeing that capacity to be devoted entirely to processing the visual itself.

Why this is a genuinely distinct finding from the related redundancy effect

The modality effect specifically concerns complementary information — a visual and an explanation of that visual, meant to be integrated and processed together, split across two different channels to reduce channel competition. The redundancy effect, discussed elsewhere, specifically concerns identical, fully duplicative information presented simultaneously through two channels, which overloads processing through unnecessary duplication rather than through channel competition — the two effects point toward different underlying mechanisms and shouldn't be conflated, even though both concern the interaction between narration, text, and visuals in multimedia learning.

Why this specifically favors narration over on-screen text for explaining visuals

Given the choice between explaining a visual through spoken narration or through on-screen text, the modality effect research favors narration specifically when the visual and the explanation are meant to be processed together and integrated in real time, since narration avoids the visual-channel competition that on-screen text specifically creates in exactly this situation.

Why this doesn't mean text should never be used in instructional materials

The modality effect specifically applies to situations where a visual and explanatory content need to be processed simultaneously and integrated together — text remains genuinely useful and often preferable for content meant to be read and processed independently of any simultaneous visual, or for learners who need to review material at their own pace in a way narration doesn't easily accommodate, meaning the practical takeaway is choosing modality deliberately based on how the content is actually meant to be processed, not eliminating text from instructional design altogether.

What this means for designing multimedia instructional content

  • Use spoken narration rather than on-screen text specifically when explaining a visual meant to be processed simultaneously alongside that explanation
  • Distinguish the modality effect (complementary information split across channels) clearly from the redundancy effect (identical information duplicated across channels), since they call for different design responses
  • Retain on-screen text for content meant to be read independently or reviewed at a learner's own pace, where the modality effect's specific channel-competition concern doesn't directly apply
  • Test multimedia formats directly where feasible, since the size of the modality effect can vary depending on the specific content and visual complexity involved

The modality effect offers a specific, mechanistically grounded reason to favor narration over on-screen text for explaining visuals processed simultaneously — not a general preference for audio over text, but a precise application of how working memory's separate processing channels can be used more efficiently when information is split across them deliberately.

modality effect multimedia learningnarration versus text visualsworking memory channel separationcourse creatorscognitive load multimedia design