Qualitative research gets judged badly in two opposite directions, and both do damage.
It is dismissed by people applying criteria built for measurement — what was your sample size, is it representative, what is the effect size, did you compute intercoder reliability — questions that a good ethnography or interview study is not attempting to answer and should not be expected to.
And it is accepted uncritically by people who, lacking any other criteria, treat the presence of rich quotations and a reflexive paragraph as sufficient. Which means bad qualitative work is often unchallengeable , because its readers have no vocabulary for saying what is wrong with it.
This lesson supplies the vocabulary. There are real standards, they are demanding, and they are not the statistical ones.
The reviewer who could not say what was wrong.
A funding panel is assessing two proposals.
The first is an interview study of how families make decisions about care for an ageing parent. Thirty-five interviews across fifteen families, including more than one member of each family where possible, recruited through three different routes — a carers' organisation, two GP practices and a snowball from each — deliberately including families who had decided differently. The proposal specifies what would count as a case that challenges the emerging account, and says the analysis will report them. It names the analytic approach, includes the topic guide, and says who will code and how disagreements will be handled.
The second proposes "in-depth qualitative interviews with approximately 20 participants, analysed thematically to explore lived experiences", recruited through a support group, with themes to be reported.
A panel member objects to the first : thirty-five is too small to generalise, there is no control group, and the recruitment is not random.
Every one of those objections is a category error (see 7.3.1). The study is not estimating a prevalence, has no counterfactual to construct, and its sampling logic is theoretical rather than statistical.
And the same panel member has no objection to the second — which is, on the evidence provided, a proposal to talk to twenty self-selected people from one organisation and report what they said.
The failure is symmetrical, and it comes from the same source. Having only quantitative criteria, the reviewer applied them where they did not belong and had nothing at all to apply where standards genuinely were missing.
The second proposal is not weak because twenty is too few. It is weak because it does not say who those twenty will be or why, does not describe an analytic procedure, cannot say what would count as being wrong, and has recruited entirely from a population that has already sorted itself by having sought support. All four of those are stateable criticisms with proper names, and this lesson is where they come from.
Reliability and validity were built for instruments. What replaces them?
Reliability asks whether repeating the procedure gives the same result. In interpretive work the expectation is wrong on its own terms : a second researcher, with a different position and a different relationship to participants, would not produce an identical analysis, and should not be expected to. That is not error; it is what 7.1.2 said about situated knowledge, arriving as a practical problem.
Validity asks whether the instrument measures the intended construct. In qualitative work there is often no instrument and no construct being measured — the object is meaning, process or practice, and the researcher is the instrument.
So the question has to be reframed. Not did you measure it accurately but: is there good reason to believe this account, could it have come out otherwise, and can I see enough of how it was produced to judge?
Three replacements follow , and everything below elaborates them. Could the analysis have been wrong and would you have noticed? Can I see the working? And where does this apply?
The standard framework, and its critics
Lincoln and Guba's four criteria, which remain the most used.
Credibility — is the account a defensible representation of what it describes? The practices that support it: prolonged engagement (enough time to get past the front stage); persistent observation (depth on what matters, not just breadth); triangulation ; peer debriefing (a colleague outside the study interrogating the analysis); negative case analysis ; and member checking , with the limits set out below.
Transferability — not generalisation, but whether the findings apply elsewhere. The move is to shift the burden to the reader : provide description thick enough (see 7.5.1) that someone who knows another setting can judge whether it is similar in the ways that matter. The researcher's obligation is to supply the material for that judgement, not to make it.
Dependability — is the process traceable and consistent? The instrument is an audit trail : the decisions taken, when and why, the codebook and its changes, the memos, the sampling choices. The claim is not that a second researcher would reproduce the analysis, but that they could follow how this one was reached.
Confirmability — are the findings grounded in the data rather than in the researcher's predispositions? Again by audit: can each claim be traced back to material?
And the authenticity criteria added later push in a different direction, asking whether the research represented the range of views fairly, and whether it produced any benefit or understanding for those studied — which imports 7.1.2's questions about whose purposes research serves.
And the objection to all of it, which deserves a hearing.
The four criteria above are deliberate parallels to internal validity, external validity, reliability and objectivity. Critics argue that the parallel is itself the problem — that qualitative inquiry is being licensed by demonstrating that it has equivalents of standards developed for something else, and thereby accepting that the other framework sets the terms (which is 6.8.1's argument, transposed).
The deeper version is the objection to "criteriology" — to the idea that the quality of interpretive work can be secured by a checklist at all. A list of procedures can be performed without the judgement they were meant to serve : a study can triangulate, member-check, keep an audit trail and produce a shallow, unilluminating account, while a study that did none of those can be revelatory. Judgement cannot be proceduralised , and a checklist creates the appearance that it has been.
Alternative frameworks exist that are broader and less parallel. Tracy's "big-tent" criteria are the most used: a worthy topic; rich rigour (sufficient and appropriate data, time and care); sincerity (about position and about limitations); credibility ; resonance (whether it moves or transfers); significant contribution ; ethics ; and meaningful coherence — whether the study achieves what it claims and connects its literature, question, method and findings into one thing.
Where this lands, and it is a real position rather than a compromise. Use the criteria as prompts for judgement, not as a scorecard. Their function is to make you ask the right questions — could this have been wrong; can I see the working; how far does it reach — and the answers require reading the study, not counting its features. A methods section that lists seven quality procedures is not thereby a good study , and a great many are exactly that.
Triangulation, and what it can honestly deliver
Four kinds, and one persistent misunderstanding.
Data triangulation — different sources, times, places, groups. Investigator triangulation — more than one researcher. Theory triangulation — interpreting the same material through different theoretical frames. Methodological triangulation — different methods on the same question.
The misunderstanding is in the metaphor. In surveying, triangulation fixes a single true position by taking bearings from two known points. The metaphor implies there is one true answer that multiple methods converge upon , and that convergence therefore validates.
Two problems with that.
Different methods measure different things. A survey measure of trust and an ethnographic account of trust are not two bearings on one object; they are two objects (see 7.1.3). When they disagree, it does not follow that one is wrong — and treating disagreement as a validation failure discards the most interesting result.
And convergence can be spurious. Two methods sharing a bias converge on the same error. Interviews and focus groups with the same self-selected population will agree with each other beautifully (see 7.3.2).
The honest version of what triangulation does. It produces a fuller account — different methods reach different aspects. It locates discrepancies , and the discrepancy is often the finding: the gap between what people say in a survey, say in an interview and do in a setting is a substantive result about the situation, not a measurement problem (see 7.4.1 on the attitude–behaviour gap). And where methods with genuinely different weaknesses agree, confidence does rise — which is 7.6.2's triangulation-across-designs argument.
What it does not do is validate. "We triangulated" is not a quality claim, and the word is used as one constantly.
Member checking is weaker than its reputation, and the distinction that saves it is simple.
Member checking of data — returning a transcript or a factual account. This works. It catches mishearings, wrong dates, misunderstood terminology, and gives participants a chance to correct or withdraw material. It is also good ethical practice.
Member checking of interpretation — returning the analysis and treating agreement as validation. This is much more problematic, for five reasons.
Participants may reject a correct analysis. An account of how a hierarchy operates will not be endorsed by those at the top of it; an analysis of unacknowledged constraint may be rejected precisely by the people whose constraint it describes. Requiring participants' assent gives every group a veto over unflattering findings , which would abolish most critical sociology at a stroke (see 7.5.1 on Winch).
They may agree out of politeness or deference. The power relation that shaped the interview shapes the check.
They have moved on. Six months later, people's accounts of their own past have been revised by what happened since.
They are not positioned to assess a cross-case claim. A participant knows their own case; the analysis is about the pattern across thirty-five.
And the framing of the question determines the answer. "Does this seem right to you?" asked by someone who has spent a year on it, of someone who spent an hour, will usually get a yes.
So: use it, and be precise about what it establishes. Factual accuracy, terminological fit, ethical accountability — yes. Validation of an interpretation — no , and a study claiming its analysis is valid because participants agreed has confused two different things.
The strongest practice, and it is not on most checklists prominently enough.
Deliberately looking for the case that breaks your account, and reporting what you found (see 7.2.3, 7.5.1, 7.7.1).
Why it outranks the others : every other practice is compatible with confirmation-seeking. You can triangulate towards what you already think, member-check with people who share your view, and keep a beautiful audit trail of a one-sided analysis. Negative case analysis is the only routine practice whose whole purpose is to give the analysis a chance to fail — which is Mayo's severity requirement (see 7.2.3) in qualitative form.
What it looks like in a finished study. Twenty-eight of thirty-five accounts fit this pattern. Four did not, and here is what they had in common. Three of those four were in the one family where the parent had their own income, which suggests the mechanism is conditional on financial dependence — so we sought two further families in that position, and here is what we found.
That paragraph does more for a reader's confidence than any list of procedures , and it is rare.
Transparency, and the argument the field is currently having.
The quantitative version of open research — sharing data, code and pre-registered plans — has a qualitative counterpart, and it is contested.
The straightforward parts are not controversial and should be routine. Publish the topic guide. Publish the coding frame. State the sampling logic and the recruitment routes, including the ones that failed. State who coded and how disagreements were resolved. Report the number of participants who expressed a view when you write "participants felt" (see 7.5.2).
Sharing the data itself is harder, and the objections are serious. Qualitative data is often deeply identifying and cannot be reliably anonymised — remove the names and the setting is still recognisable to anyone in it. Consent was given for a relationship, not for an archive. And the meaning of a transcript depends on context the reader does not have: the tone, the preceding hour, the setting, what the interviewer already knew.
When transparency requirements were pushed hard in political science, qualitative researchers objected on exactly these grounds , and the resulting deliberations produced a more differentiated position: transparency about process and analytic procedure as a general obligation, with data sharing handled case by case, subject to consent and to the risks in the particular setting.
That is the right shape of answer. The demand to see how a conclusion was reached is legitimate and applies to everyone. The demand to see the raw material is not always compatible with the obligations under which the material was obtained (see 7.8.1).
Five ways this goes wrong.
The checklist as a substitute. Seven procedures listed in a methods section, none of which changed the analysis.
"Triangulation" as a magic word , meaning "we did two things".
Immersion as a warrant. "Having spent eighteen months in the field, it became clear that..." Time in the field is a precondition, not evidence (see 7.5.1 on the unverifiability problem).
Over-claiming reach. A study of one organisation whose conclusions are written about "organisations". The remedy is not hedging but specification : say what conditions the finding depends on, so a reader can judge transfer.
And under-claiming. Endless qualification until nothing is asserted. "This small exploratory study suggests that experiences may vary" says nothing, protects the author, and wastes the participants' time. Timidity is not rigour , and a study that will not commit to a claim cannot be wrong — which by 7.2.3's standard means it cannot be right either.
Because you now need to be able to judge a qualitative study on its own terms — and this is how.
Six questions.
Who was studied, and why those people? A stated logic — theoretical, purposive, comparative — or convenience with no account.
How much contact, over how long, in what role?
What was the analytic procedure? Named, described, and specific enough that you can picture it.
Could this have come out differently — and does the study show you a case where it nearly did?
How far does the author claim it reaches, and have they given you what you need to judge that yourself?
And is the account illuminating? This one is not proceduralisable and it is not optional. A study can satisfy every criterion and tell you nothing. The best qualitative work makes you see something you had been looking at without noticing — the bowling scores in 7.5.1, the recognition mechanism in 7.7.1 — and that quality is the point of the method, not a bonus.
The two failures at the start of this lesson are the same failure : not knowing what the work is for. Judge measurement by measurement standards and meaning by meaning standards, and both traditions get better.
Qualitative work is judged badly in two directions — dismissed by criteria built for measurement, and accepted uncritically for lack of any criteria at all. Both come from not knowing what the method is for.
Reliability and validity need reframing , not translation: a second researcher would not reproduce an interpretive analysis and should not be expected to. The three real questions are: could this have been wrong and would you have noticed; can I see the working; and where does it apply?
Lincoln and Guba's four criteria remain the standard frame. Credibility — prolonged engagement, persistent observation, triangulation, peer debriefing, negative cases, member checks. Transferability — thick description that lets the reader judge fit, with the burden shifted to them. Dependability — an audit trail making the process traceable, not replicable. Confirmability — claims traceable to data.
And the objection is real : the criteria are parallels to a framework built elsewhere, and "criteriology" can substitute procedure for judgement. Tracy's broader criteria — worthy topic, rich rigour, sincerity, credibility, resonance, contribution, ethics, meaningful coherence — are the main alternative. Use criteria as prompts for judgement, not a scorecard.
Triangulation comes in four kinds and does not validate. The surveying metaphor is misleading: different methods reach different objects, convergence can be spurious through shared bias, and divergence is often the finding — the gap between what people say and do is a result, not an error.
Member checking of data works; member checking of interpretation does not validate — participants may reject a correct analysis, agree out of deference, have revised their own past, or not be positioned to assess a cross-case claim.
Negative case analysis is the strongest practice , because it is the only routine one whose purpose is to give the analysis a chance to fail.
Transparency about process and procedure is a general obligation ; sharing raw qualitative data is genuinely contested, because it is identifying, context-dependent, and consented to for a relationship rather than an archive.
And over-claiming and under-claiming are both failures. Timidity is not rigour.
Trustworthiness — Lincoln and Guba's overarching frame.
Credibility / transferability / dependability / confirmability — defensible representation; applicability judged by the reader; a traceable process; claims grounded in data.
Prolonged engagement / persistent observation — enough time to get past the front stage; depth on what matters.
Peer debriefing — an outside colleague interrogating the analysis.
Audit trail — the documented record of analytic decisions.
Thick description — description rich enough for a reader to judge transfer.
Triangulation — data, investigator, theory and methodological; produces fullness and locates discrepancy, does not validate.
Member checking — of data (works) versus of interpretation (does not validate).
Negative / deviant case analysis — deliberately seeking and reporting cases that break the account.
Authenticity criteria — fairness of representation and benefit to those studied.
Big-tent criteria — Tracy's eight: worthy topic, rich rigour, sincerity, credibility, resonance, contribution, ethics, meaningful coherence.
Criteriology — the objection that quality in interpretive work cannot be secured by checklist.
One — judge one properly. Take a qualitative paper and answer the six questions at the end of this lesson. Write one sentence on each. You will have a better review than most published referee reports.
Two — find the negative case. In the same paper, search for anyone who did not fit. Note whether the author looked, found, and reported.
Three — test a triangulation claim. Find a study saying it triangulated. Ask whether the methods had different weaknesses or the same ones, and whether any divergence is reported.
Four — write your own transfer conditions. Take a qualitative finding you find convincing and write the three conditions a different setting would have to share for it to apply there. This is what transferability actually asks of a reader.
Five — spot the hedge. Find a conclusion so qualified it asserts nothing. Rewrite it as a claim that could be wrong. Then decide whether the study's evidence supports it — which is the question the hedging was avoiding.
Topic 7.7 is done. Topic 7.8 turns to the things that go wrong across the whole enterprise, whatever the method — and it begins with the constraint that shapes what can be studied at all.
7.8.1 — Ethics covers consent and where it breaks down, the studies that produced the modern rules, harm and its distribution, the specific problems of covert work, deception, vulnerable participants and digital data, and the standing tension between protecting participants and protecting institutions from scrutiny.