A thousand open-ended customer feedback comments get coded into themes by an analyst, producing a summary breakdown like "32% mentioned pricing concerns, 21% mentioned support quality." A second, equally competent analyst working independently from the same raw comments, without access to the first analyst's specific coding decisions, can plausibly arrive at a noticeably different breakdown — not because either analyst did careless work, but because coding open-ended verbatim responses into discrete themes involves genuine interpretive judgment at multiple points, and different reasonable judgments produce different results.
Where the genuine interpretive judgment actually enters
Deciding whether a comment mentioning both price and a specific feature gets coded under one theme, both themes, or a combined theme; deciding how granular or broad a given theme category should be; deciding how to categorize an ambiguous or multi-part comment — each of these represents a genuine judgment call that reasonable, competent analysts can make differently, absent an explicit, shared framework specifying how to handle these situations consistently.
Why this variability is a normal feature, not a fatal flaw
The interpretive judgment inherent in qualitative coding doesn't make thematic analysis unreliable or unscientific — it means the analysis requires explicit process discipline to be defensible and reproducible, in the same way that quantitative analysis requires explicit, stated methodological choices (which statistical test, which threshold) to be defensible and reproducible, rather than assuming that any two competent researchers would automatically arrive at identical conclusions from raw numbers alone.
What an explicit codebook actually does
Developing a codebook — an explicit, written definition of each theme category, with clear rules and example comments illustrating how ambiguous or borderline cases should be classified — converts what would otherwise be a series of ad hoc individual judgment calls into a documented, consistent, and reproducible classification system that a second analyst can actually apply the same way the first one did, rather than reinventing the classification logic independently from scratch.
Why inter-rater reliability checking is the direct test of whether this actually worked
Having two independent analysts code a meaningful subset of the same raw comments using the shared codebook, then formally calculating the level of agreement between their independent codings (using an established statistic like Cohen's kappa), provides a direct, quantified check on whether the codebook actually achieved the consistency it was designed to produce, rather than simply assuming that writing a codebook down automatically guarantees consistent application of it.
What this means for producing defensible qualitative research
- Develop an explicit, written codebook with clear category definitions and example cases before beginning full-scale coding, rather than coding first and defining categories afterward
- Have more than one analyst code a meaningful subset of the data independently and formally check inter-rater reliability
- Treat a codebook as a living document to be refined once initial disagreements between coders reveal genuine ambiguity in the original category definitions
- Be specifically skeptical of thematic analysis results presented without any description of the coding process or reliability checking used to produce them
Verbatim coding's genuine interpretive element isn't a weakness to be embarrassed about — it's a feature of the method that specifically requires explicit process discipline to manage well, and that discipline is exactly what distinguishes a defensible thematic analysis from an impressionistic one dressed up with percentages.