This lesson covers three tools that look unrelated and are not.
A focus group produces material that individual interviews cannot, because people are talking to each other rather than to you.
A document is not a container of information; it is something written by someone, for someone, in order to do something.
Content analysis turns text into countable data — and, now that a laptop can process a century of newspapers, has become one of the fastest-moving areas in the discipline.
What unites them is a single question: who produced this, in front of whom, and what were they doing by producing it?
One redundancy round, three kinds of evidence, three different truths.
A researcher is studying a restructuring at a manufacturing firm. Ninety jobs went.
She starts with individual interviews , and gets careful, guarded accounts. People are still employed there. Several say the process was "handled properly, all things considered". One says, after a pause, that it was "difficult but I understand why they did it".
Then she runs a focus group with eight people from the same shop floor who all knew each other before the study.
The first ten minutes reproduce the interviews almost exactly — the same phrases, the same care. Then a man who has said nothing says: "That's not what happened, though, is it."
And the conversation changes. Someone remembers the date the "consultation" started and points out it was after the agency staff had already been told. Another disputes it — no, that was the week after — and a third settles it by reference to her daughter's birthday. They correct each other, in public, using shared reference points the researcher does not have. Within twenty minutes there is a collective account of the sequence that no individual had given her, and that is more accurate on dates than any of them alone.
And something else appears that individual interviews had entirely hidden. When one woman says she thought the selection criteria were fair, the temperature in the room drops. Nobody contradicts her. The silence is the data — it tells the researcher that this is a position with a cost, in this group, which is a fact about the shop floor that no interview would have shown her.
Then she reads the documents.
The restructuring rationale paper — written for the board, and readable as an argument that a decision already made was prudent.
The consultation minutes — written by HR, recording that views were "invited and considered", in language visibly constructed to withstand an employment tribunal.
The selection matrix — a scoring grid with weighted criteria: skills, flexibility, attendance, disciplinary record. A beautifully objective-looking instrument , and the researcher notices that "flexibility" is scored by the line manager, unmoderated, with no definition attached.
Three sources. Three different objects. The interviews gave individual accounts, produced for her. The group gave a collective reconstruction, and a map of what is sayable. The documents gave the organisation's account of itself, produced for an audience that might one day be a court — which is not a lesser kind of evidence, provided you read it as what it is.
Each of these tools produces evidence about something other than what it appears to be about.
A focus group is not a cheap way to interview eight people at once. It produces group data: what can be said in front of these people, what gets corrected, what gets laughed at, what nobody will touch. The interaction is the finding, not the packaging.
A document is not a record of what happened. It is an artefact produced in a process, for a purpose, by someone with a position. Minutes record what it was decided to record.
And a content analysis does not measure what a text "says". It measures what a coding scheme, applied by particular people to a particular corpus, produced. The scheme is the theory; the counting is arithmetic (see 7.2.2).
Read each of them for what generated them, and all three become powerful. Read them as transparent, and all three mislead.
Focus groups
What they are for, which is narrower and more interesting than "cheap interviews".
The method descends from Merton, Fiske and Kendall's wartime focused interview work, developed to study audience reactions to radio broadcasts and propaganda films — and it has been shaped ever since by the tension between its academic origins and its enormous commercial use in marketing.
Five things a group does that an individual cannot.
It reveals what is sayable. In front of these people, in this setting, which views require hedging, which produce silence, which get laughter. This is normative data of a kind no questionnaire reaches.
It generates correction. Participants who share a world hold each other to account on facts, dates and sequence, using reference points the researcher does not possess.
It surfaces vocabulary. How a community actually talks about a thing, including the words they use with each other and not with outsiders. This is the ideal preparation for building a questionnaire (see 7.4.1 on why closed questions should be built from open material).
It shows opinion being formed and defended , rather than reported. You watch someone shift position under challenge, or dig in, and see what arguments do the work.
And it can create safety in numbers. On some sensitive topics, a group of similar people makes a subject discussable that a one-to-one interview would make too exposed.
Five things it does badly. Individual attitudes (contaminated by the group). Prevalence (never count individuals across groups as though they were a sample). Personal disclosure of anything the participant would not want the others to know. Anything where confidentiality cannot be guaranteed — because you can promise your own silence and not theirs. And any topic where a strong social norm will simply produce the norm.
Design decisions, each of which changes the data.
Size. Six to ten is the usual range. Fewer feels like a stilted conversation; more and quiet participants disappear entirely.
Number of groups. One group is an anecdote. The standard approach is segmentation — several groups within each category of interest (by age, role, seniority, place), with at least two or three per segment so that a peculiarity of one group can be recognised as such.
Homogeneous or mixed. Homogeneous groups talk more freely and reveal within-group norms. Mixed groups reveal the fault lines — what happens when managers and staff, or men and women, are in the room together — at the cost of candour. Never mix people with power over each other unless the power relation is your subject.
Strangers or an existing group. Strangers give more candour and less context. An existing group — a shift, a class, a congregation — gives you the real norms in operation , with the real history, and the real reasons certain things do not get said. The redundancy story shows both effects at once.
The moderator's job , which is not to be neutral wallpaper: keep one person from owning the room, protect the minority position, and actively seek disagreement — does anyone see it differently? — because a group's natural tendency is towards a comfortable consensus that represents nobody's actual view.
And stimulus material — a news clipping, a policy extract, a photograph, a scenario — often works better than a question, because it gives the group something to argue about rather than something to answer.
Three dynamics that will wreck a group if unmanaged.
The dominant speaker. One confident participant sets the frame in the first five minutes and everything afterwards is a response to them. The moderator's countermeasures are structural — going round the table on the opening question, direct invitations to those who have not spoken, and splitting into pairs.
Conformity. People converge on what appears to be the group's view, and the convergence is faster where status differences exist. The absence of disagreement is not agreement , and a transcript with no conflict in it should be read with suspicion rather than satisfaction.
And the analysis error that follows from all of this: treating the group as eight interviews. Counting how many participants said something, across groups, as though they were a sample, is the single most common misuse of the method. The unit of analysis is the group. What you can report is what emerged, what was contested, what was silenced, and how positions moved — not proportions.
Documents
John Scott's four questions, which are the best checklist in the field.
Authenticity — is it what it purports to be, by whom it claims, of the date it claims? Forgeries exist; so do backdated files, minutes written afterwards, and documents assembled for an inquiry.
Credibility — is it free from error and distortion? Who wrote it, from what position, with what knowledge, and with what interest in the account? A meeting note taken by a participant with a stake differs from one taken by a clerk.
Representativeness — is it typical of its kind, and if not, how does it differ? This is where archives are dangerous (below).
Meaning — is it clear and comprehensible? Not just literally: what did these words mean in this institution, at this time, to these readers? Bureaucratic language is a dialect , and "concerns were noted" carries information to insiders that it does not carry to you.
Documents do things. That is what they are for.
Lindsay Prior's argument is the shift that makes documentary research sociological rather than merely archival: stop treating a document as a container of content and start treating it as a participant in a process.
A policy document does not describe a practice; it authorises one, and defends it. A case file does not record a client; it constructs a client, in the categories the agency acts on, for the next professional who will read it and for a possible future inquiry. A selection matrix does not measure employees; it converts a decision into a defensible form. A form does not collect information; it determines what information exists.
Three questions to ask of any document, and they are the whole method.
Who produced it, in what role?
For whom — who was the intended reader, and who was the imagined reader? The second is often more revealing: much organisational writing is produced for a hypothetical auditor, regulator or court that never appears, and the anticipation shapes every sentence.
What does it do? Authorise, record, defend, allocate, exclude, comply, warn.
And then the most productive question of all: what does it not say, and would its absence be noticed? A minute recording that a decision was taken unanimously, in an organisation where dissent is normally minuted, is telling you something. Categories that a form does not have are situations the organisation cannot officially perceive (see 7.2.2).
Archives are survivor samples, and the survival was not random.
Everything in 7.3.2 applies to documents with extra force.
What survives is what someone kept , and organisations keep what they are required to keep, what protects them, and what nobody thought to destroy. Personal papers survive from people who had the space, the literacy and the sense that their papers mattered. The poor, the mobile, the illiterate and the deliberately unrecorded are absent by construction.
Michel-Rolph Trouillot's analysis of how silences enter the historical record is the sharpest statement of this , and it identifies four distinct moments where it happens: at fact creation (what gets written down at all), at fact assembly (what enters an archive), at fact retrieval (what gets found and used in narratives), and at retrospective significance (what is treated as history worth telling). Silences compound across the four , which is why absence from the record cannot be read as absence from the world.
And the modern version is digitisation. What has been scanned, OCR'd and made searchable is not a random sample of what exists — it favours major newspapers, national collections, English-language material and institutions with funding. A computational study of "the discourse" is a study of the digitised discourse , and the gap is systematic (see 7.4.4).
Content analysis, from index cards to language models
Classical content analysis: the coding frame is the theory.
The traditional method is straightforward and its discipline is in the details.
Define the corpus — which sources, which period, sampled how. This is a sampling problem and deserves 7.3.1's attention: choosing three national newspapers is a decision about whose discourse counts.
Define the units. The sampling unit (an issue, an article), the coding unit (an article, a paragraph, a sentence, a mention), and the context unit (how much surrounding text a coder may use to decide).
Build the coding frame — the categories, defined precisely enough that two people apply them the same way, with rules for the awkward cases. This is where the analysis actually happens. Everything after it is counting.
Establish intercoder reliability. Two or more coders independently code a subset, and agreement is reported with a chance-corrected statistic — Cohen's kappa or Krippendorff's alpha (see 7.2.2). Raw percentage agreement is not acceptable , because two coders applying a common code will agree often by luck alone.
And the manifest/latent distinction matters. Manifest content is what is literally there — the word "immigrant" appears, the article is 600 words, the source quoted is a police officer. It is countable with high reliability. Latent content is what the text implies, frames or evokes. It is more interesting and much harder to code reliably , and the trade-off between the two is permanent.
Computational text analysis: four families, four failure modes.
Dictionary and lexicon methods. Count words from a predefined list — sentiment lexicons, moral-language dictionaries, topic word lists. Fast, transparent, reproducible. Failure mode : no context. Negation, irony, quotation and domain shift all defeat them. A dictionary built on product reviews applied to parliamentary debate is measuring something, and it is not what it says on the label.
Supervised classification. Humans label a training sample; a model learns to reproduce the labels at scale. Excellent when the category is well defined. Failure mode : the model inherits the labellers' judgements exactly, including their inconsistencies and their blind spots — and then applies them a million times with perfect confidence. Automating a coding frame does not validate it.
Unsupervised topic models. Methods such as latent Dirichlet allocation discover clusters of co-occurring words without being told what to look for. Genuinely useful for exploring a corpus you have not read. Failure modes : the number of topics is chosen by the analyst and changes the results; topics are unstable across runs and settings; and the topics have no names until a human names them — which is an interpretive act, done after seeing the output, with all the risks of 7.2.1's HARKing.
Embeddings and large language models. Representations that capture contextual similarity, now including models that can be asked to classify or summarise text directly. Powerful, and the newest. Failure modes : opacity about why a case was classified as it was; sensitivity to prompt wording, which is 7.4.1's question-wording problem transplanted into a pipeline; drift as models are updated, breaking reproducibility; and the well-documented tendency of such systems to encode the associations present in their training material — which, when the research question is about social bias, is a confound sitting in the instrument.
And one rule governs all four , stated most memorably by Grimmer and Stewart: validate, validate, validate. Every computational measure must be checked against human coding of a sample. A method that scales a coding decision to a million documents has not removed the decision — it has removed your ability to notice it was wrong.
The alternative tradition, and the honest tension between them.
Discourse analysis and critical discourse analysis take the opposite approach: close reading of a small number of texts, attending to how language constructs categories, positions, agency and legitimacy. Who is the subject of the sentence; what is in the passive voice; what is presupposed rather than argued; what is nominalised so that no one is doing it. "Mistakes were made."
Fairclough's programme connects the textual level to the practices of production and consumption and to wider social structures — which is what makes it sociological rather than linguistic.
And the criticism is serious and should be stated. Critics — Widdowson's is the best known — have argued that close reading of selected texts too easily finds what the analyst went in believing, that the selection of texts is rarely justified, and that alternative readings are not tested.
The productive response has been methodological rather than defensive : corpus-assisted discourse analysis, where a large corpus is used to establish what is typical — which words habitually occur together, how a term's associations have changed — and close reading is then applied to material that has been shown to be representative rather than merely striking. This is the best current answer to the "you found what you looked for" objection , and it is a good example of two traditions improving each other rather than arguing.
Because most of the record of social life is text, and it is now searchable.
Four habits.
For a focus group finding: ask what the group was, who was in the room, and whether disagreement appears anywhere in the report. If eight strangers agreed about everything, you are reading a norm rather than a view.
For a document: ask who wrote it, for whom, and what it was doing. Then ask what a document of this type is required to say regardless of the facts.
For a content analysis: ask what the corpus was and how it was assembled. A study of "media coverage" that used three national outlets available in a database has measured three national outlets available in a database.
And for anything computational: ask what the validation was. A paper reporting a topic model with no human check on what the topics contain has reported the output of a procedure, not a finding about the world.
The unifying question, one more time, because it is the whole lesson : who produced this, in front of whom, and what were they doing by producing it? Applied to a group transcript, an organisational memo and a corpus of a million tweets, it is the same question — and it is the question that turns each of them from a source of quotations into evidence.
A focus group produces group data : what is sayable, what gets corrected, what silence follows. It is excellent for norms, vocabulary, collective reconstruction of events, and watching opinion being defended — and bad for individual attitudes, prevalence, personal disclosure and anything requiring confidentiality between participants. Segment into several groups per category; choose homogeneous groups for candour and mixed for fault lines; never mix people with power over each other; and the moderator's job is to seek disagreement. The unit of analysis is the group, and counting participants across groups is the method's commonest misuse.
A document is not a record but an artefact that does something — authorises, defends, allocates, complies. Scott's four questions : authenticity, credibility, representativeness, meaning. Prior's shift : ask who produced it, for whom (including the imagined auditor), and what it does — then ask what it does not say and whether the absence would be noticed.
Archives are survivor samples. Trouillot's four moments of silencing — fact creation, fact assembly, fact retrieval, retrospective significance — compound, so absence from the record is not absence from the world. Digitisation is the modern version of the same filter.
Classical content analysis puts the theory in the coding frame : corpus, sampling and coding units, defined categories, and chance-corrected intercoder reliability. Manifest content is reliable; latent content is interesting; the trade-off is permanent.
Four computational families with four failure modes : dictionaries (no context), supervised classification (inherits the labellers' judgements at scale), topic models (analyst-chosen, unstable, and named after the fact), and embeddings and language models (opaque, prompt-sensitive, drifting, and carrying their training material's associations). The rule for all four: validate against human coding, because scaling a decision removes your ability to see it was wrong.
And critical discourse analysis is the close-reading alternative , with a real vulnerability — finding what you looked for — whose best answer is corpus-assisted work that establishes typicality before reading closely.
Focus group / focused interview — group discussion producing data about shared norms, vocabulary and contestation.
Segmentation — running several groups within each category so that a peculiar group can be recognised.
Natural versus stranger groups — existing groups reveal real norms; strangers give more candour.
Scott's four criteria — authenticity, credibility, representativeness, meaning.
Documents in action — Prior: treating documents as participants in processes rather than containers of content.
The imagined reader — the hypothetical auditor or court that shapes organisational writing.
Trouillot's four silences — fact creation, fact assembly, fact retrieval, retrospective significance.
Sampling, coding and context units — what is selected, what is coded, and how much surrounding text may inform the coding.
Coding frame — the defined categories; where the analysis actually happens.
Intercoder reliability — chance-corrected agreement between independent coders (kappa, Krippendorff's alpha).
Manifest / latent content — what is literally present; what is implied or framed.
Dictionary methods / supervised classification / topic models / embeddings — the four computational families.
Validation — checking any automated measure against human coding of a sample. Not optional.
Critical discourse analysis — close reading of how language constructs categories, agency and legitimacy.
Corpus-assisted discourse analysis — establishing typicality across a large corpus before reading closely.
One — read a document as an act. Take any policy, minute or form from your own work or life. Write down who wrote it, for whom, what it authorises or defends, and one thing it cannot record.
Two — find the imagined reader. In the same document, mark the sentences that exist because someone might one day complain. They are usually the most carefully written ones.
Three — build a coding frame. Take twenty headlines about one topic and construct three categories that another person could apply identically. Then have someone apply them and compare. The disagreements will show you what your definitions were missing.
Four — test a dictionary. Take a sentiment word list and apply it by hand to ten sentences including irony, negation and quotation. Count how many it gets wrong.
Five — look for the silence. Take a historical topic you know something about and ask which participants left no documents. Then ask what would be concluded by someone reading only the surviving record.
The last lesson in this Topic addresses the questions Parts 4, 5 and 6 of this course were made of — why revolutions happen where they do, why welfare states took different shapes, why one region industrialised and another did not.
These cannot be surveyed, randomised or observed. They have small numbers of cases, enormous numbers of differences between them, and outcomes that unfolded over centuries.
7.5.4 — Comparative-Historical Method covers how the discipline reasons about causes when there are eight cases and a hundred variables, what Mill's methods can and cannot do, and why sequence and timing are treated as evidence in their own right.