You have forty interviews. Transcribed, that is around nine hundred pages.
There are two ways to proceed. One is to read them until something occurs to you, then go back and find quotations that support it. This is what a great deal of published qualitative research actually consists of , and its results are unfalsifiable because nothing in the process could have produced a different answer.
The other is a set of procedures for working through the material systematically, so that the analysis is shaped by the data rather than by what you remembered from interview three. This lesson is about those procedures — what they are, what they are called, what the most-cited method in the field actually requires, and where its central concept is currently under attack.
One paragraph, coded twice.
Return to the nurses from 7.5.2. Here is a fragment of one transcript.
"I'd been doing that ward eleven years. Eleven years. And the new manager comes in and the first thing she does is change the handover — no discussion, just an email, we're doing it at quarter past now. And I remember thinking, right, so nobody's going to ask me. And it sounds so small when I say it out loud. It's fifteen minutes. But it's — you get to a point where you think, they don't actually know what I do."
A first pass — open coding — labels what is going on, staying close to the data.
Tenure and investment. Change imposed. No consultation. Communication by email. The realisation moment ("I remember thinking"). Anticipating the interviewer's judgement ("it sounds so small"). Minimising her own reaction. Not being known ("they don't actually know what I do").
Some of those are in vivo codes — the participant's own words used as the label, because her phrasing is doing analytic work that a paraphrase would lose. "They don't actually know what I do" is one.
Now the second pass, and this is where analysis starts rather than filing.
The researcher writes a memo — not a code, a paragraph of thinking:
This is the third participant to apologise for the smallness of their trigger before describing it. Each time it is a procedural change made without asking. And each time the account moves immediately from the procedure to a claim about being unknown. So the fifteen minutes is not the grievance; it is the evidence for the grievance. Is the mechanism here about recognition rather than workload? If so, I should expect the people who stayed to describe someone who did know what they did. Check against the two participants who considered leaving and didn't.
That memo does four things at once. It notices a pattern across cases. It reinterprets a code (the trigger is evidence, not cause). It generates a hypothesis. And it specifies what would test it, including who to talk to next.
Then the next interview is chosen because of that memo — someone who stayed — rather than because they were next on the list. That is theoretical sampling (see 7.2.3), and it is what separates grounded theory from coding.
And notice what would have happened without any of this. The researcher would have remembered the vivid quote about the handover, written a section called "Communication failures", and illustrated it with three examples. The recognition mechanism — which is the actual finding — would never have surfaced , because nothing in that process compares cases.
How do you work through nine hundred pages without either drowning or cherry-picking?
The drowning failure : coding everything into four hundred codes, ending with a filing system and no argument. Fragmentation — the data cut so finely that individual people disappear and only decontextualised snippets remain.
The cherry-picking failure : forming an impression early and then reading for confirmation. This one is invisible from inside , because the supporting quotations really are in the data.
The procedures in this lesson exist to sit between those two. And they share a single principle, which is worth stating before any of the vocabulary: compare. Incident with incident, case with case, code with code, the cases that fit with the cases that do not. Analysis in qualitative work is comparison, and everything else is bookkeeping.
Codes, themes, and the difference
A code is a label attached to a segment of data. It is not a finding.
Descriptive or topic codes name what a segment is about. Pay. Handover. Childcare. Useful for retrieval, analytically inert on their own.
Analytic or conceptual codes name what is going on in it, in the researcher's terms. Anticipatory minimising. Recognition claim. Evidence of the trigger. These are the ones that build towards an argument.
In vivo codes use the participant's own words, preserving a formulation that carries something a paraphrase would lose.
And a theme is not a code, nor a collection of codes on the same topic.
Braun and Clarke's argument here is worth taking seriously because it corrects the commonest failure in qualitative writing. A theme is a pattern of shared meaning organised around a central concept. "Participants talked about pay" is not a theme; it is a domain summary — a bucket labelled with a topic. A theme says something : pay is described not as compensation but as an index of whether the institution values the work .
The test is simple: can the theme be stated as a claim? If the heading works as a sentence with a verb in it, it is a theme. If it is a noun phrase naming a subject area, it is a filing category, and the analysis has not happened yet.
And they are equally firm about the passive voice that pervades this literature. "Three themes emerged from the data." Themes do not emerge. They are constructed by an analyst, with a question, a position and a theoretical commitment. The passive construction conceals the analyst precisely where 7.1.2 says they should be visible.
Grounded theory: what it requires, and what it usually means
The real method, in its parts.
Glaser and Strauss proposed in 1967 that theory should be generated from systematically analysed data rather than deduced in advance and tested — a direct challenge to the grand theorising of the period (see 5.1.1). The method has five components, and all five are required.
Initial (open) coding , close to the data, often line by line at the start, using gerunds — minimising, justifying, withdrawing — because verbs keep the analysis on processes rather than topics.
Constant comparison. Every new incident is compared with those already coded; every code with the other codes; every category with the others. This is the engine , and it is what stops a code from quietly becoming a container.
Memo-writing. Continuous, throughout, from the first interview. Memos are where the analysis is actually done — the coding is preparation for them. A study with no memos has coding and no theory.
Theoretical sampling. Later cases chosen by what the developing analysis needs: the contrasting case, the case that should falsify, the setting where the mechanism should be absent. Not a sample designed in advance (see 7.3.1).
Theoretical saturation as the stopping rule, discussed below.
And the tradition split, publicly and acrimoniously. Strauss, with Corbin, later published a more prescriptive procedure with a fixed coding paradigm — conditions, actions, consequences — and Glaser accused them of forcing data into a template rather than letting analysis develop. Kathy Charmaz's constructivist version is now the most widely used , and it accepts what the original framing did not: the researcher does not discover categories lying in the data, but constructs them from a position (see 7.1.3).
Now the honest observation about how the term is used. Grounded theory is among the most cited methodologies in the social sciences, and in a large share of papers claiming it, what was done was: some interviews were coded, and themes were reported. No theoretical sampling, no memos, no constant comparison, no theory produced. That is thematic analysis — which is a perfectly good method with its own literature — described with someone else's name. The mislabelling matters because it lets a study claim the authority of a demanding procedure it did not follow.
Saturation: the concept, and the serious case against it.
Theoretical saturation , in the original sense, means that a category is fully developed — its properties and dimensions are worked out, and new data adds nothing to it. It is a claim about a category, not about a dataset.
What the term usually means in practice is different : no new codes appeared in the last few interviews. This is data or thematic saturation , and it is much weaker.
The empirical work is genuinely useful. A well-known study coding sixty interviews found that the great majority of codes — around nine in ten — appeared within the first twelve, and that the basic structure of the analysis was stable well before the end. The finding is real and it is conditional : those interviews were with a relatively homogeneous population on a fairly focused topic. Widen the population or the question and the number rises steeply , and no general number exists.
The case against the concept is stronger than most textbooks admit. Braun and Clarke and others argue that saturation is incoherent within an interpretive framework: if meaning is constructed by the analyst rather than extracted from the data, there is no point at which the data is exhausted — a more experienced or differently positioned analyst would keep finding more. On this view, saturation borrows the vocabulary of a realist, information-extraction model of data and applies it where that model does not hold.
And the practical criticism is unanswerable : saturation is almost always asserted rather than demonstrated. "Saturation was reached at interview 24" appears with no evidence whatever.
The workable position, which respects both sides. Decide the sample size for stated, defensible reasons — the diversity of the population, the breadth of the question, what the analysis requires, and what resources exist. If you claim saturation, say which kind, and show something : a record of when new codes stopped appearing, or an account of which categories were developed to the point where further cases added no properties. And if your framework does not support the concept, say that instead, and justify the sample another way. Either is defensible. The assertion without evidence is not.
The wider landscape
Six approaches, so you can recognise what you are reading.
Reflexive thematic analysis. Braun and Clarke's programme: familiarisation, coding, generating initial themes, developing and reviewing, refining and naming, writing. Explicitly analyst-centred , rejecting both "emergence" and intercoder reliability as inappropriate to its epistemology. The most widely used named approach, and the one most often used well.
Framework analysis. Developed for applied and policy research: material is charted into a matrix with cases as rows and themes as columns. Its great strength is that the case stays visible — you can read across a row and see one person's whole account — which directly addresses fragmentation. Transparent, auditable, and well suited to teams and to research with a client.
Narrative analysis. Treats accounts as stories with structure — a beginning, a turning point, a moral — and asks what work the structure does. Right when the object is how people make their lives coherent , and it resists chopping accounts into codes at all.
Interpretative phenomenological analysis. Small samples, intensive case-by-case analysis of lived experience, cases analysed fully before comparison. Idiographic first, comparative second.
Qualitative content analysis. Systematic categorisation, often with a partly pre-specified frame and reported reliability — closer to 7.5.3's tradition, and appropriate where the codes are meant to be a shared instrument.
Analytic induction. The most demanding: a provisional explanation, then an active hunt for the case that breaks it, then reformulation, repeated (see 7.2.3).
These are not interchangeable , and choosing between them is the ontology-and-epistemology decision from 7.1.3, made at the analysis stage.
Software, and what it will not do.
Qualitative analysis packages — NVivo, ATLAS.ti, MAXQDA, Dedoose, and free tools such as Taguette — manage the mechanics: attaching codes, retrieving every segment with a given code, cross-tabulating codes by case attributes, keeping memos linked to data, and handling audio, images and video.
What they do well is retrieval and organisation at scale , and they make an audit trail almost automatic.
What they do not do is think. And there is a documented occupational hazard: the ease of code-and-retrieve encourages treating the coded output as the analysis. You end up with a report structured by the code list, which is a description of your filing system rather than an argument.
Two guards. Write memos in the software as you code, so thinking accumulates alongside labelling. And produce case-level summaries as well as code-level ones , so that a whole person can still be seen after the transcript has been cut into pieces.
On automated coding with language models , which is arriving quickly: the rule from 7.5.3 applies without modification. These systems can accelerate initial coding of large corpora and cannot validate their own output. Anything they produce must be checked against human coding of a sample, the prompt is a coding instruction and should be reported as one, and the interpretive judgement — which is the part that constitutes the analysis — is not the part that can be delegated.
Five failure modes.
Coding as filing. Four hundred descriptive codes, no analytic ones, no memos. The output is a well-organised set of quotations and no claim.
Fragmentation. The data cut so finely that context is lost and no individual case survives. The framework matrix and case summaries exist to prevent this.
Counting codes as importance. "The most frequent theme was pay." Frequency in a transcript reflects what you asked about, who was talkative and how you coded — not what mattered. Something mentioned once, with difficulty, may be the most important thing in the study.
Themes that are topics. Headings that are noun phrases. If the heading cannot be written as a claim, it is a domain summary.
And the missing negative case. An analysis that fits every participant is either a genuine finding or a failure to look. Report who did not fit and what you concluded from them (see 7.2.3, 7.5.1).
Because the difference between analysis and illustration is not visible in the finished text unless the author shows their working.
Four things a reader can check.
Is there an account of the analytic process — how coding was done, by whom, whether a codebook exists, whether memos were kept? "Data were analysed thematically" is not an account.
Are the themes claims or topics?
Does anything in the paper complicate the argument — a participant who did not fit, a case that changed the analysis?
Is the case still visible , or only fragments?
And for anyone doing this work, one sentence carries most of the method. Write memos from the first interview. Not notes on what was said — thinking about what it might mean, what pattern it might belong to, and what would test it.
The coding is not the analysis. The memos are the analysis , and the coding exists to make them possible.
Two failure modes bracket qualitative analysis : drowning in an unusable filing system, and forming an impression early and reading for confirmation. The principle that sits between them is comparison — incident with incident, case with case, the cases that fit with the cases that do not.
A code is a label, not a finding. Descriptive codes name topics; analytic codes name what is going on; in vivo codes preserve the participant's formulation. A theme is a pattern of shared meaning organised around a central concept, and it can be stated as a claim — a noun-phrase heading is a domain summary, and the analysis has not happened. Themes do not emerge; analysts construct them.
Grounded theory requires five things together : initial coding in gerunds, constant comparison, continuous memo-writing, theoretical sampling of later cases by what the analysis needs, and theoretical saturation. The tradition split — Strauss and Corbin's prescriptive paradigm, Glaser's objection that it forces data, Charmaz's constructivist version that accepts the analyst's position. And a large share of papers claiming grounded theory did thematic analysis under another name.
Saturation has two senses — a fully developed category, or no new codes — and the empirical work showing most codes appear in the first dozen interviews is real and conditional on a homogeneous sample and a focused question. The interpretive critique is serious : if meaning is constructed, there is no point of exhaustion. The practical rule: justify your sample size for stated reasons, and if you claim saturation, say which kind and show evidence.
Six approaches to recognise : reflexive thematic analysis, framework analysis (which keeps the case visible in a matrix), narrative analysis, interpretative phenomenological analysis, qualitative content analysis, and analytic induction.
Software retrieves and organises; it does not think — and code-and-retrieve can quietly become the whole analysis. Automated coding must be validated against human coding, and the prompt is a coding instruction.
Five failures : coding as filing, fragmentation, treating code frequency as importance, themes that are topics, and the missing negative case.
Write memos from the first interview. The memos are the analysis.
Code — a label attached to a segment of data. Descriptive / analytic / in vivo — naming the topic, the process, or using the participant's own words.
Theme — a pattern of shared meaning organised around a central concept, statable as a claim.
Domain summary — a topic-labelled bucket masquerading as a theme.
Memo — written analytic thinking produced alongside coding; where the analysis happens.
Constant comparison — systematically comparing incidents, codes and categories against each other.
Theoretical sampling — selecting later cases by what the developing analysis requires.
Theoretical saturation — a category developed to the point where new data adds no properties. Data saturation — no new codes appearing.
Grounded theory — the Glaser and Strauss programme; its Straussian, Glaserian and constructivist variants.
Reflexive thematic analysis — Braun and Clarke's analyst-centred approach.
Framework analysis — matrix-based charting that keeps whole cases visible.
Fragmentation — loss of case and context through fine-grained coding.
CAQDAS — qualitative analysis software; retrieval and organisation, not interpretation.
Audit trail — the documented record of analytic decisions.
One — code a paragraph twice. Take any interview transcript, blog post or letter. Code it descriptively, then code it again analytically — what is the person doing in this passage? Compare the two lists. Only the second could build an argument.
Two — write a memo. After coding, write two hundred words that are not codes: what pattern might this be part of, what would confirm it, and who would you want to talk to next.
Three — convert a domain summary. Find a qualitative paper and look at its theme headings. Rewrite each as a sentence with a verb. Some will not survive , and those were topics.
Four — hunt the negative case. Take a claim from any qualitative study you have read and ask whether the paper reports anyone who did not fit. If not, note what the analysis cannot rule out.
Five — check the saturation claim. Find a paper that says saturation was reached. Look for the evidence. You will usually find an assertion and a number.
Coding produces an analysis. The remaining question is the one qualitative researchers are asked most often and answer least well: how do you know it is right?
7.7.2 — Validity in Qualitative Work covers the criteria that actually apply — credibility, transferability, dependability, confirmability — what triangulation can and cannot deliver, why member checking is weaker than its reputation, the role of the negative case, and how to judge a qualitative study without borrowing standards that were built for something else.