Almost every hard question in this course needs more than one method.
How much requires a sample. Why requires getting inside a process. Whether it caused anything requires a counterfactual. What it means to the people involved requires asking them. No single design supplies more than one of these , and a study that claims all four from one method has over-reached.
So combining methods should be routine. In practice it is done badly far more often than it is done well, and the difference is entirely about integration — whether the methods speak to each other, or merely appear in the same document.
What employers said, and what the same employers did.
Recall Devah Pager's audit study from 7.4.2. Matched pairs of testers, identical résumés, a randomly rotated criminal record, sent to real employers in Milwaukee. White applicant, no record: 34 per cent callback. White with a record: 17. Black without: 14. Black with a record: 5.
That establishes magnitude . It does not establish why , and the standard explanations were incompatible with each other: conscious prejudice, statistical discrimination, unexamined assumption, or the mechanical operation of screening rules nobody had examined.
So Pager and Lincoln Quillian went back to the same employers with a survey , asking whether they would be willing to hire someone with a criminal record like the one described.
Two findings, and the second is the important one.
Employers said yes far more often than they behaved as if yes. Stated willingness to hire an ex-offender vastly exceeded the actual callback rate — the attitude–behaviour gap from 7.4.1, measured within the same firms rather than inferred across studies.
And there was essentially no relationship between what an individual employer said and what that employer did. The employers who expressed the greatest willingness were not the ones who called back. Stated attitudes did not predict behaviour at the level of the individual firm at all.
Now see what that combination establishes that neither method could alone.
The audit alone gives a number with no mechanism. The survey alone would have produced a badly wrong picture — it would have shown employers as broadly willing, and would have been reported as evidence of declining discrimination. And the survey's error is not random : it is systematically optimistic in exactly the direction that makes the problem look smaller.
Together they establish something neither contains : that the mechanism is not conscious, stated unwillingness, and that survey evidence on this topic is not merely noisy but structurally misleading. The divergence between the methods is the finding , and a researcher who had treated the disagreement as a validity problem to be reconciled would have thrown away the result.
Methods answer different questions, so combining them is not addition.
If a survey and an ethnography were two measurements of one thing, mixing would be simple: average them, or trust the better one.
They are not. As 7.1.3 established, they frequently have different objects. A survey measure of trust and an ethnographic account of trust are not two readings of the same quantity.
Which means the interesting design questions are three , and most published mixed-methods work answers only the first.
Which method leads, and which serves? Priority is a decision, and leaving it implicit usually means the quantitative component leads by default.
In what order? Sequence determines what each component can do — a qualitative phase before measurement builds the instrument; after analysis, it explains the pattern.
And where do they meet? This is the one that decides whether a study is mixed methods or two studies in one document.
The designs
Five shapes, each for a different job.
Sequential exploratory: qualitative first, then quantitative. Interviews or fieldwork discover the categories, the vocabulary and the range of experience; the survey is then built from what was found rather than from what the researcher assumed. This is the correct way to construct a questionnaire (see 7.4.1 on why closed questions must be built from open material) and it is skipped constantly.
Sequential explanatory: quantitative first, then qualitative. The analysis establishes a pattern; the qualitative phase explains it. Its highest-value use is the deviant case : identify statistically the cases that the model fits worst — the neighbourhoods that should have high crime and do not, the pupils who should have failed and did not — and go and find out why. This turns a residual into a research question , which is the opposite of what 7.4.2's Team A had to do.
Convergent (concurrent): both at once, analysed separately, compared at the end. Used for completeness and for triangulation. Its whole value is in the comparison , so a convergent design that does not systematically compare has not been executed.
Embedded: one method nested inside the other. The standard case is a process evaluation inside a trial — the experiment establishes whether the intervention worked; the embedded qualitative work establishes what was actually delivered, how it was received, and why it worked where it did. This directly addresses 7.4.2's black-box problem , and it is the single most useful mixed design in applied research.
Multiphase : several linked studies over time, each informing the next.
And there is a notation worth knowing , because you will see it: capitals for the dominant component and lower case for the subordinate one, with an arrow for sequence and a plus for concurrency. QUANT → qual is a quantitative study with a qualitative follow-up; QUAL + quant is qualitatively-led with a concurrent quantitative element. The notation forces the priority decision into the open , which is its point.
Why mix at all: five distinct purposes, and they are not interchangeable.
The classic statement identifies five reasons, and a study should know which it is pursuing.
Triangulation — seeking convergence on the same question, to increase confidence.
Complementarity — using each method to illuminate a different facet, producing a fuller account rather than a confirmation.
Development — using one method to build the other: fieldwork to construct a survey instrument, or survey results to select interview cases.
Initiation — deliberately seeking contradiction , in order to generate new questions. This is the most under-used and the most productive , and the Pager and Quillian result is what it looks like.
Expansion — extending the scope of a study by using different methods for different components: the trial for the effect, the ethnography for the implementation.
Naming the purpose in advance disciplines the design. A study mixing for complementarity should not be embarrassed that its methods do not converge; a study mixing for triangulation must say what convergence would look like before it looks.
Integration: where mixed methods lives or dies
The commonest failure is not doing it badly. It is not doing it at all.
The recognisable shape : a quantitative results section, a qualitative results section, and a discussion saying both were valuable. Two studies stapled together. The reader can remove either without changing the other, which is the diagnostic.
Integration can happen at four points , and good studies use more than one.
At design. The sampling for one component is determined by the other; the interview guide is written from the survey findings; the survey items come from the fieldwork.
At data collection. The same participants, so that individual-level linkage is possible — which is what made the employer study work. This is a decision that must be taken early and usually cannot be retrofitted.
At analysis. Qualitative themes converted into variables and tested; quantitative groupings used to structure qualitative comparison; cases selected for interview because of where they fall in the model.
At interpretation. The meta-inference — the claim that neither component supports alone.
And the practical instrument that does most of the work is the joint display : a table or matrix with the quantitative findings on one axis and the qualitative findings on the other, cell by cell, showing where they agree, where they diverge and where one is silent. It is unglamorous, it takes an afternoon, and it forces every comparison to be made explicitly rather than gestured at in a discussion section.
A related technique deserves naming because it is so useful in comparative work: nested analysis. Use a large-N statistical analysis to identify the pattern, then select cases for intensive study on the basis of where they sit relative to that pattern — a case the model predicts well, to trace the mechanism, and a case it predicts badly, to find what is missing. The statistical model chooses the cases; the cases interrogate the model (see 7.5.4).
This is the interesting case, and it is usually fudged.
Four possibilities, and they require different responses.
One — one component is wrong. The sample was biased, the measure invalid, the fieldwork thin. Check this first , honestly, on both components rather than on the one you liked less.
Two — they measure different constructs. The survey measured stated attitude; the observation measured behaviour; the interview measured accounts. These are three different objects (see 7.1.3), and "disagreement" between them is a category error.
Three — they operate at different levels. An area-level pattern and an individual-level account need not agree, and 7.6.4 explains why they need not.
Four — the divergence is the finding.
The fourth is the one to look for, and the examples across this Part are consistent. People's stated reasons versus their reconstructed sequences (7.5.2). Official statistics versus victimisation surveys (7.4.4). Stated willingness versus callback behaviour. In every case the gap is a substantive result about legitimation, recording or self-presentation — and in every case the standard move of "reconciling" the methods would have destroyed it.
The rule: never average across a disagreement. Explain it. A mixed-methods study whose components disagree and which reports why is worth more than one whose components agree , because agreement is also what you get when both methods share a bias (see 7.7.2).
The paradigm objection, and how much of it survives.
The incompatibility thesis held that quantitative and qualitative research rest on incompatible ontologies and epistemologies, so combining them is incoherent — you cannot simultaneously treat social reality as external and measurable and as constructed in interaction (see 7.1.1, 7.1.3). This was argued seriously during the paradigm wars and is not simply a confusion.
Three responses, of increasing strength.
Pragmatism : the philosophical question is not the practical one, and methods are tools selected by the question (see 7.1.1). Effective and slightly evasive — it declines the argument rather than answering it.
Critical realism : a stratified reality of mechanisms operating as tendencies in open systems positively requires both. Measurement establishes the pattern the mechanism produced; interpretation and observation establish the mechanism. This is the most coherent philosophical home for mixed methods, and it is why 7.1.1 spent time on it.
The dialectical position : do not dissolve the tension. Hold both frameworks, let their different assumptions generate different readings, and treat the friction between them as a source of insight rather than a problem to be resolved.
And now the concession the objection has earned. In a great deal of published work, the qualitative component is reduced to illustration — a few quotations decorating a regression table, contributing nothing that would change the conclusion. That practice does exactly what the objection warned about : it subordinates one framework to another while claiming to combine them.
The test is simple and it is worth applying to every mixed-methods paper you read. Could the qualitative component have changed the conclusion? If the answer is no — if it could only have illustrated whatever the numbers said — then it was decoration, and the study is quantitative research with quotations in it. The same test applies in reverse, and is failed less often only because quantitative components are rarely subordinate.
Quality criteria, and the practical obstacles.
Each component must meet its own standards. The quantitative element is judged by sampling, design, measurement and analysis (7.3, 7.4, 7.6). The qualitative element by credibility, transferability, dependability and confirmability (7.7.2). A weak component is not rescued by a strong one — and mixed-methods work is disproportionately likely to contain one weak component, because it is usually done by someone trained in one tradition who has added the other.
Then two further criteria specific to mixing. Is the rationale stated — which of the five purposes, and why does the question require it? And is there integration , at a nameable point, producing an inference neither component supports alone?
And the obstacles are real rather than intellectual. Mixed-methods work costs more, takes longer, and needs either a rare individual or a team — and a team introduces its own problem, since the components must be genuinely co-designed rather than subcontracted. Journals have word limits that force one component into a paragraph. Reviewers are usually competent in one tradition and assess the other by the wrong criteria (see 7.7.2's opening). And careers are built within traditions , so the person who does both is assessed by two sets of specialists, each unimpressed.
None of this is an argument against mixing. It is an explanation of why so much of it is done badly , and it is a sociological account of a methodological problem — which is the appropriate note for this Part to end its toolkit on.
Because the questions that matter most cannot be answered any other way.
Did the programme work, and why, and for whom, and will it transfer? That is a trial, plus a process study, plus subgroup analysis, plus enough contextual knowledge to reason about transfer (see 7.4.2).
How much discrimination is there, and through what mechanism? An audit study, plus the accounts of the people doing the deciding, plus attention to the gap between them.
What is happening to this community, and is it happening elsewhere? Ethnography plus comparative and statistical work — which is the design implied by every Part of this course from 2 to 6.
Three questions for reading any mixed-methods study.
What was the purpose of mixing — triangulation, complementarity, development, initiation, expansion?
Where did the components meet? Design, collection, analysis or interpretation — and can you point to the sentence?
And could either component have overturned the other's conclusion?
And one for anyone designing. Decide the point of integration before collecting anything, because the most valuable form — linking the same individuals across components — is a sampling decision that cannot be made afterwards.
The reason to combine methods is not thoroughness or diplomacy between traditions. It is that the world does not divide along the lines the methods divide along , and any single method's picture is partial in a way that method cannot itself detect.
Combining methods is a design problem, not extra work. The employer study is the model: an audit established the magnitude of discrimination; a survey of the same employers established that stated willingness to hire was far higher than actual behaviour and did not predict it at the individual firm level. Neither method alone could have produced that, and the survey alone would have been systematically misleading in the reassuring direction.
Three design decisions : which component leads, in what order, and — decisively — where they meet.
Five designs : sequential exploratory (qualitative first, to build the instrument); sequential explanatory (quantitative first, with the deviant case as its highest-value use); convergent; embedded (a process evaluation inside a trial , which addresses the black-box problem directly); and multiphase. The QUANT/qual notation exists to force the priority decision into the open.
Five purposes : triangulation, complementarity, development, initiation — deliberately seeking contradiction, the most under-used — and expansion.
Integration is where mixed methods lives or dies , and the commonest failure is its complete absence: two studies stapled together, either removable without affecting the other. Integrate at design, at collection (the same participants, a decision that cannot be retrofitted), at analysis, and at interpretation — with the joint display as the workhorse and nested analysis as the comparative version.
When methods disagree, there are four possibilities : one is wrong; they measure different constructs; they operate at different levels; or the divergence is the finding. Look for the fourth. Never average across a disagreement — and agreement is also what shared bias produces.
The incompatibility thesis has a real answer in critical realism, a pragmatic evasion, and a dialectical alternative — and it has earned its concession , because qualitative components are so often reduced to illustration. The test: could the qualitative component have changed the conclusion?
Each component must meet its own tradition's standards , plus a stated rationale and demonstrable integration — and the obstacles to doing this well are institutional rather than intellectual.
Sequential exploratory / explanatory — qualitative then quantitative; quantitative then qualitative.
Convergent design — both concurrently, compared systematically.
Embedded design — one method nested inside another, as in a process evaluation within a trial.
Nested analysis — using a statistical model to select cases for intensive study by where they sit relative to the pattern.
Deviant case follow-up — investigating qualitatively the cases the model fits worst.
QUANT/qual notation — capitals for the dominant component, arrow for sequence, plus for concurrency.
Five purposes of mixing — triangulation, complementarity, development, initiation, expansion.
Integration — the point at which components inform each other: design, collection, analysis or interpretation.
Joint display — a matrix comparing quantitative findings against qualitative findings cell by cell.
Meta-inference — the conclusion neither component supports alone.
Incompatibility thesis — the claim that the paradigms cannot legitimately be combined.
Illustration failure — a qualitative component that could not have changed the conclusion.
One — apply the removal test. Take a mixed-methods paper and ask whether either results section could be deleted without changing the conclusion. If yes, it was not integrated.
Two — design a joint display. For any study with two kinds of evidence, draw the matrix: findings from one method down the side, from the other across the top. Fill in the cells where they agree, disagree and are silent.
Three — find a productive disagreement. Identify a case where survey evidence and observed behaviour differ on the same question. Write two sentences on what the gap itself establishes.
Four — pick the deviant cases. Take any statistical relationship you know of and describe which cases you would interview: the ones that fit best, the ones that fit worst, or both — and what each would tell you.
Five — state the purpose. For a question you care about, sketch a two-method design and name which of the five purposes it serves. Then write what you would do if the two methods disagreed.
Part 7 has one lesson left, and it is the one everything has been building towards.
You now have the philosophy, the design principles, the sampling logic, the quantitative and qualitative toolkits, the analytical machinery, the fallacies, the ethics and the crisis. The remaining task is to turn all of it into something you can do in twenty minutes with a study in front of you.
7.9.1 — How to Read a Study Without Believing It or Dismissing It is a working procedure: what to read first, what the sections actually tell you, the questions that do most of the work, the specific tells of weak research, and how to reach a proportionate judgement rather than a verdict.