The biggest questions in this course cannot be surveyed, randomised or observed.
Why did some societies industrialise first? Why did revolutions succeed in France, Russia and China and not elsewhere? Why did some rich democracies build expansive welfare states and others did not? Why did democracy consolidate here and collapse there?
There are perhaps a dozen relevant cases, they differ from each other in a thousand ways, the outcomes unfolded over centuries, and the evidence is whatever survived (see 7.5.3).
And yet the discipline's most enduring work is of exactly this kind. This lesson is about how that reasoning is done, what it can establish, and where it goes wrong.
Three revolutions and the cases that did not have one.
In 1966 Barrington Moore published a comparison of the long transitions to the modern world in England, France, the United States, Germany, Japan, China and India. His conclusion was that there were three routes — a commercial-bourgeois path leading to liberal democracy, a "revolution from above" in which a landed elite modernised in alliance with a weak bourgeoisie and which led towards fascism, and a peasant revolution leading to communism — and that what determined the route was the relationship between landlords, peasants, the crown and the towns during commercialisation. His compressed slogan: "No bourgeois, no democracy."
Thirteen years later Theda Skocpol did something more disciplined with a narrower question.
She asked what produces a social revolution — not merely a change of regime, but the rapid transformation of a state and its class structure, carried through by revolt from below. She identified three cases: France from 1789, Russia from 1917, and China from 1911.
Her explanation was deliberately anti-heroic. Revolutions, on her account, are not produced by revolutionaries, ideologies or vanguards. They occur when two structural conditions coincide : a state that has been placed under intolerable pressure from more developed competitors abroad and cannot reform because a landed upper class blocks it, so that its administrative and military machinery breaks down — and agrarian structures that give peasants sufficient solidarity and autonomy from landlords to revolt when the state's coercive grip loosens. She endorsed the old line that revolutions are not made; they come.
And here is the methodological move that made the book a landmark.
She did not study only revolutions. She systematically compared her three cases with cases where the same pressures existed and no social revolution followed — England, Prussia and Germany, Japan — asking in each what was present or absent.
Japan faced acute external pressure and its state did not break down; its elite carried through a reform from above. Absent condition: state breakdown. Prussia faced defeat and pressure and its agrarian structure, with strong landlord control over an unfree peasantry, gave peasants no capacity for autonomous revolt. Absent condition: peasant capacity. England commercialised early and its state was not caught in the same trap. Absent condition: the reform blockage.
Three positive cases, three negative cases, and two conditions that appear together in the first set and never together in the second.
The criticisms are as important as the argument , and this course grades every tradition the same way. William Sewell argued that her structural account left ideology and culture out of a phenomenon that participants experienced as meaning-laden, and that the French case in particular cannot be told without the ideological transformation. Others noted that six cases cannot support the number of conditions being considered. Michael Burawoy compared her method to Trotsky's and argued that a research programme should be judged by whether it improves under anomaly rather than by whether its comparisons are clean (see 7.2.3 on Lakatos).
Skocpol replied that culture matters and that her question was about outcomes rather than experiences. The exchange is one of the best in the discipline's history, and it is unresolved — which is normal, and not a failure.
Eight cases, a hundred differences. Why is that evidence of anything?
The statistical answer to confounding is to have enough cases that the confounders wash out (see 7.4.2). Comparative-historical work never has enough cases. France and Prussia differ in language, religion, geography, empire, legal tradition, literacy, urbanisation and a thousand other things. On the face of it, any of them could be the cause.
The tradition's answer has three parts, and all three are needed.
One — choose the cases so that most differences are held constant. Comparison is not taking whatever is available; it is a design.
Two — do not treat a case as a single observation. A case is a long sequence of events, decisions, timings and mechanisms, and the internal evidence about how the outcome came about is where most of the causal leverage actually lives.
Three — take timing and sequence as evidence. If the proposed cause occurred after the outcome began, it is not the cause. If the same factor produced different results depending on when it arrived, that is a finding rather than a nuisance.
These are not weak substitutes for a regression. They are a different logic , and where they are done well they can be more convincing than a regression on 150 countries whose observations are not independent and whose variables mean different things in each.
The classical machinery, and its limits
Mill's methods, which everyone invokes and Mill himself thought unusable here.
John Stuart Mill set out canons of causal inference in 1843, and comparative sociology has used them ever since.
The method of agreement. Cases sharing an outcome are examined for what else they share. If revolutions occurred in France, Russia and China, and the only condition present in all three is state breakdown under external pressure, that condition is a candidate cause.
The method of difference. Cases alike in most respects but differing in outcome are examined for what differs. Prussia and France, both agrarian monarchies under military pressure, one with a revolution and one without: what differed?
The joint method uses both, which is exactly Skocpol's design.
Concomitant variation — more of the cause, more of the outcome.
And Mill himself, in the sixth book of the Logic , argued that these methods do not work for social phenomena , for two reasons that have not gone away.
Plurality of causes : the same outcome can be produced by different routes, so a condition absent in one positive case is not thereby eliminated. Complex interaction : causes in social life combine, and a factor that matters only in combination with another will not show up as a shared feature.
So the classical methods are a starting discipline, not a proof procedure , and the modern tradition has built its real machinery on top of them.
Two designs, and the trade they involve.
Most similar systems design. Choose cases that resemble each other on as much as possible and differ in the outcome. Any remaining difference is a candidate cause , and the closer the match, the shorter the candidate list. Neighbouring countries, adjacent regions, similar organisations, the same country before and after — the logic behind the natural experiments of 7.4.3, applied by hand.
Most different systems design. Choose cases with almost nothing in common that nonetheless share the outcome. Any factor they do share is a candidate , and the more different the cases, the more impressive a shared factor is. This is the design that guards against a finding being an artefact of one region's peculiarities — and it is exactly what 6.8.1 was asking for.
The trade is straightforward. Most-similar designs give precise comparisons over narrow ranges, and risk finding that everything relevant was held constant. Most-different designs give wide scope and weak control. Serious work uses both , and the strongest comparative claims are those that survive each.
Where the real leverage is: inside the case
Process tracing — the modern tradition's central contribution.
The objection "you only have six cases" assumes each case contributes one observation. It does not.
Inside a single case there is a sequence: who did what, when, in what order, saying what, in response to what. A causal claim implies that certain things should be findable in that sequence and certain things should not. Process tracing is the systematic examination of within-case evidence for the fingerprints a proposed mechanism would have left.
Four kinds of test, distinguished by what passing and failing each one implies.
A straw-in-the-wind test. Passing slightly supports the hypothesis; failing slightly weakens it. Neither is decisive. Most evidence is of this kind.
A hoop test. The hypothesis must pass to remain viable, but passing does not confirm it. If the minister was demonstrably abroad on the day of the meeting, the hypothesis that he attended it is dead — while his being in the country proves nothing. Hoop tests eliminate.
A smoking-gun test. Passing strongly confirms; failing does not eliminate. A memo in the minister's hand instructing the action confirms it ; the absence of such a memo shows only that nothing was written down.
A doubly-decisive test. Passing confirms and eliminates the rivals at once. Rare, and worth an enormous amount when available.
This is what converts a small-N study from an illustration into an inference. The distinction — made most clearly in the debate between the statistically-minded programme of Designing Social Inquiry and its critics in Rethinking Social Inquiry — is between data-set observations (one row per case) and causal-process observations (individual pieces of evidence about the mechanism, of which a single case may supply hundreds). A well-chosen causal-process observation can be worth more than a hundred rows , and the two kinds of evidence are not interchangeable.
Path dependence: why sequence and timing are causes.
Some processes have the property that early events constrain later ones , so that the order in which things happened is itself part of the explanation.
The core mechanism is increasing returns. Once a path is taken, the returns to continuing on it rise — through sunk investment, coordination effects (it pays to do what others do), learning effects, and adaptive expectations. The result is lock-in : an arrangement persists not because it is best but because the cost of switching has grown, and because the actors who benefit from it have become powerful enough to defend it.
A critical juncture is a relatively brief period in which several outcomes were genuinely possible and the choice made had lasting consequences. A reactive sequence is a chain in which each step is a response to the last, so that the eventual outcome cannot be read off the initial conditions at all.
The standard illustration is the QWERTY keyboard layout — allegedly designed to slow typists, retained through coordination and training effects long after the mechanical reason disappeared. And the honest treatment requires the correction : economists Liebowitz and Margolis reviewed the evidence and disputed both the claim that QWERTY was significantly inferior to alternatives and the quality of the studies used to support it. The example is contested; the mechanism is not. Better-evidenced instances abound in institutional history — the persistence of electoral systems, of pension architectures, of legal traditions, of railway gauges and of national administrative boundaries long outliving their reasons.
Why this matters methodologically : if a process is path dependent, then a variable's effect depends on when it arrived , and a method that treats each case as a bundle of simultaneous attributes cannot represent that at all. Sequence is data.
Multiple and conjunctural causation, and Ragin's answer to it.
Comparative-historical work almost always finds that several different combinations of conditions produce the same outcome (equifinality), and that conditions matter only in combination (conjunctural causation). Neither fits comfortably into a framework of independent additive effects.
Charles Ragin's qualitative comparative analysis was built for exactly this. It treats cases as configurations rather than as bundles of variables, uses set theory to express necessary and sufficient conditions, and assembles a truth table of which combinations of conditions are associated with which outcomes, then reduces it logically to the simplest expression consistent with the evidence. Fuzzy-set QCA extends this to degrees of membership rather than crisp yes/no.
Its real appeal is that it formalises what comparative historians were already doing , and makes the reasoning explicit and checkable.
And the criticisms are substantial and should be held alongside it. Results can be sensitive to how cases are calibrated into set memberships, a step involving judgement that then disappears into the algebra. Limited diversity — many logically possible combinations do not occur in the real world — forces assumptions about unobserved configurations, and different assumptions yield different solutions. And with few cases and many conditions, a solution that fits perfectly may be fitting noise. QCA makes reasoning transparent, which is its main virtue; it does not make small numbers of cases into large ones.
Five ways comparative-historical work goes wrong.
One — selecting on the dependent variable. Studying only revolutions, only successful firms, only collapsed democracies. The classic critique is that this cannot establish that anything causes the outcome, because you never observe the same conditions failing to produce it. The defence is partial and real : cases with the outcome can establish necessary conditions (a condition absent in a positive case is not necessary), and Skocpol's inclusion of negative cases is exactly the standard remedy. What it can never establish alone is sufficiency.
Two — Galton's problem. The cases are not independent. States copy each other, colonial powers imposed the same institutions across continents, professionals trained in the same places carry the same models home. What looks like six independent confirmations may be one event that spread , and diffusion is a rival explanation that must be addressed rather than assumed away.
Three — over-fitting the narrative. With enough historical detail, any sequence can be told as the inevitable working-out of the proposed cause. The protection is to state in advance what the mechanism implies should be findable, and to report what was not found (see 7.2.3 on severity).
Four — the sources problem. Comparative historians work largely from secondary literature — historians' accounts, themselves contested. The critique, made most sharply by John Goldthorpe, is that the sociologist is choosing between historians' interpretations without the specialist competence to judge , and that the choices may be made by what fits the argument. The honest practice is to say where the historiography is disputed and what turns on it — and this course applies the same standard to itself.
Five — the unit is usually the nation-state, by default. Nations are the units because that is how data and historiography are organised, not because social processes respect borders. Methodological nationalism — treating the national container as the natural unit of society — imports an assumption that empires, regions, diasporas, trade networks and colonial relations all violate (see 6.5.1, 6.7.1).
Because the questions that matter most are almost all of this kind, and the alternative to doing them carefully is not doing them at all.
Nobody will ever randomise a revolution, a welfare state, an empire or an industrial transition. The choice is between disciplined comparison and confident storytelling , and the storytelling will happen either way — in politics, in the press, in national mythology.
Four questions for reading any comparative or historical argument.
Which cases, and why those? If they all share the outcome, what is being claimed — necessity, or something the design cannot support?
Are the negative cases there? The places where the cause was present and the outcome did not follow. Their presence is the single best indicator of seriousness.
Is there within-case evidence, or only the pattern across cases? A claim about a mechanism with no evidence of the mechanism operating is a correlation with six observations.
Are the cases independent, or did they copy each other?
And one for the reader of this course. Parts 4, 5 and 6 were full of comparative-historical claims — Weber on the Protestant ethic and the world religions, Moore on the routes to modernity, dependency theory on Latin America and East Asia, Mamdani on the bifurcated state. You now have the criteria to grade them , and you should. That is what the whole of Part 7 has been for.
Comparative-historical method reasons about causes with few cases, many differences and outcomes that took centuries. Moore's three routes to the modern world and Skocpol's structural account of social revolution are the exemplars — and Skocpol's methodological move was to include the cases where the outcome did not occur , showing that in each, one of the two conditions was absent. The criticisms — culture left out, too many conditions for six cases — are part of the record.
Mill's methods (agreement, difference, joint, concomitant variation) are the starting discipline, and Mill himself judged them inapplicable to social life because of plurality of causes and complex interaction.
Two designs : most similar systems (hold everything constant, examine what differs) and most different systems (share almost nothing, examine what is shared). Strong claims survive both.
The real leverage is within the case. Process tracing looks for the fingerprints a mechanism would have left, using straw-in-the-wind, hoop (eliminates), smoking-gun (confirms) and doubly-decisive tests. Causal-process observations are a different kind of evidence from data-set observations, and one case supplies many.
Path dependence makes sequence a cause : increasing returns produce lock-in; critical junctures make brief moments consequential; reactive sequences make outcomes unreadable from initial conditions. The QWERTY illustration is contested; the mechanism is well evidenced elsewhere.
Equifinality and conjunctural causation are the norm, and Ragin's QCA formalises configurational reasoning with necessary and sufficient conditions — its transparency is its virtue, and calibration judgements and limited diversity are its real limitations.
Five failure modes : selecting on the dependent variable (can establish necessity, never sufficiency); Galton's problem (cases copy each other, so diffusion is a rival explanation); over-fitted narratives; reliance on contested secondary historiography; and methodological nationalism, which makes the nation-state the default unit because that is how the archives are filed.
Comparative-historical analysis — causal reasoning across a small number of cases over long time spans.
Mill's methods — agreement, difference, joint method, concomitant variation.
Most similar / most different systems design — matching cases to isolate a difference; contrasting cases to isolate a commonality.
Process tracing — examining within-case evidence for the traces a proposed mechanism would leave.
Straw-in-the-wind / hoop / smoking-gun / doubly-decisive tests — evidence that weakly bears, eliminates, confirms, or does both.
Causal-process observation — a piece of within-case evidence about mechanism, distinct from a row of data.
Path dependence — early events constraining later ones through increasing returns and lock-in.
Critical juncture / reactive sequence — a brief period of open possibility; a chain in which each step responds to the last.
Equifinality — several distinct combinations of conditions producing the same outcome.
Conjunctural causation — conditions that matter only in combination.
QCA / fuzzy sets / calibration / limited diversity — Ragin's configurational method; degrees of set membership; the judgement converting cases into memberships; the absence of many logically possible combinations from the real world.
Selection on the dependent variable — choosing cases by their outcome; can support necessity claims, never sufficiency.
Galton's problem — non-independence of cases through diffusion and imitation.
Methodological nationalism — treating the nation-state as the natural unit of social analysis.
One — build a most-similar comparison. Take two places, organisations or periods that resemble each other closely and differ in an outcome you care about. List everything they share, then everything that differs. The second list is your candidate causes, and it should be short.
Two — add the negative cases. Take any explanation you have read for why something happened. Name two cases where the proposed cause was present and the outcome did not follow. What was different?
Three — design a hoop test. For a historical claim you hold, write down one thing that must be true if it is correct, and where you would look. Then write one thing that would confirm it if found but proves nothing by its absence.
Four — find a lock-in. Identify an arrangement in your own field, workplace or country that persists although nobody would design it that way now. Trace what made switching progressively more expensive, and who benefits from it continuing.
Five — check independence. Take any claim that "countries which did X all got Y". Ask whether those countries adopted X independently, or copied one another, or had it imposed by the same power. The answer usually reduces the number of real cases sharply.
Topic 7.5 is complete. You have the qualitative toolkit — immersion, asking, groups, documents and comparison across cases and centuries.
Topic 7.6 turns to the numbers , and it starts somewhere unfashionable and important. Before anything can be inferred, it has to be described — and a great deal of what passes for sophisticated analysis is built on descriptions nobody checked.
7.6.1 — Description Before Inference covers what an average conceals, why the shape of a distribution usually matters more than its centre, how to read a graph, and why careful description is not the poor relation of causal analysis but its precondition.