You cannot randomise a war, a law, a border, a recession or the year someone was born.
But something else can. A lottery. A policy that began on a Tuesday. A rule that applied to children born in September and not August. An administrative threshold at 40 pupils. A river.
When the assignment of a cause is made by something unrelated to the outcome, you can recover much of the logic of an experiment without running one. This lesson covers how, and — more importantly — how to tell a real natural experiment from a hopeful piece of branding , which is a skill you will use constantly.
Two water companies, one street, and the first natural experiment.
By the 1850s John Snow was convinced that cholera was carried in water and not in bad air. The medical establishment disagreed, and Snow's famous work around the Broad Street pump — where removing a pump handle coincided with a fall in deaths — was suggestive but not decisive: the outbreak was already declining, and people had fled the district.
His second study is the one that should be famous.
South London was supplied by two competing water companies whose pipes ran down the same streets. In many neighbourhoods, adjacent houses were served by different companies — the result of a competitive scramble years earlier, in which each company had signed up whichever householders would sign.
In 1852, one of them — the Lambeth Company — moved its intake upstream, above the point where London's sewage entered the Thames. The other, Southwark and Vauxhall, continued to draw water from the tidal river below the sewage outfalls.
Snow saw what he had. He wrote that the mixing of supply in the same streets, among households of the same class, occupation and condition, meant that
"no experiment could have been devised which would more thoroughly test the effect of water supply on the progress of cholera than this."
He set out to determine, house by house, which company supplied each cholera death — going to the trouble of a chemical test on the water, because tenants often did not know who their landlord paid.
The result was overwhelming. Houses supplied by Southwark and Vauxhall had cholera death rates roughly eight to nine times those supplied by Lambeth — of the order of 315 deaths per 10,000 houses against 37.
Now see why this works as evidence. The households did not choose their water company on any basis related to their risk of cholera; a long-forgotten commercial rivalry had assigned them. They lived on the same streets, in the same conditions, breathing the same air. The one thing that differed was the water — and the assignment of the water was, for these purposes, as good as random.
That is the whole idea of this lesson, and it was already fully formed in 1855.
Every design below answers one question: what assigned the treatment?
Randomisation works because assignment is independent of everything (see 7.4.2). Observational comparison fails because assignment is a choice , and choices are made for reasons that predict outcomes (see 7.3.2).
A natural experiment is a situation where assignment was made by something that is plausibly unrelated to the outcome. A lottery number. An arbitrary date. A rule with a sharp cutoff. An administrative boundary. A commercial history nobody remembers.
So the question you ask of any quasi-experimental study — and the only one that really matters — is: what mechanism assigned people to treatment, and is there any reason that mechanism is related to the outcome?
Everything technical below is a way of exploiting a particular kind of answer. And every failure below is a case where the answer was assumed rather than established.
The four workhorses
One — difference-in-differences.
The situation : a policy applies to one group and not another, at a known moment.
The move : measure both groups before and after. The treated group's change, minus the comparison group's change, is the estimate. The second subtraction removes anything that happened to everyone.
The most argued-about example in modern economics. In 1992 New Jersey raised its minimum wage; neighbouring Pennsylvania did not. Card and Krueger surveyed hundreds of fast-food restaurants on both sides before and after, and found no employment decline in New Jersey — if anything a small rise. This contradicted a standard textbook prediction and set off thirty years of argument.
And the argument is instructive. Neumark and Wascher re-examined the question using payroll records rather than the telephone survey and found employment falls. Card and Krueger responded using administrative unemployment insurance data and again found no significant negative effect. The dispute turned substantially on which data source measured employment better — which is 7.2.2's lesson appearing in a famous debate, and it is not fully resolved.
The assumption everything rests on: parallel trends. In the absence of the policy, the two groups would have moved in parallel. This is not testable directly — it concerns a counterfactual — but it can be probed: plot several periods before the change and see whether the two lines actually moved together. A study that shows only one before-period has assumed the entire identifying assumption and shown you nothing about it.
Where it breaks : when the treated group was selected because it was already changing. Places that raise minimum wages tend to be places with tightening labour markets. The policy is a response to a trend, and the trend is then attributed to the policy.
Two — regression discontinuity.
The situation : a treatment is assigned by a sharp threshold on some continuous measure. Scoring above 60 gets you the scholarship. Income below a line gets you the benefit. Winning an election by one vote makes you the incumbent.
The move : compare cases just above the threshold with cases just below . Someone with 59.8 and someone with 60.2 are, on everything except the treatment, essentially identical — and which side of the line they landed on is close to a coin flip.
The classic sociological application is Angrist and Lavy's use of a rule in Israeli schools derived from Maimonides: a class must split when enrolment exceeds 40. A cohort of 40 sits in one class of 40; a cohort of 41 sits in two classes of about 20. One extra enrolment halves class size, and nobody engineered it. Comparing achievement across that boundary gives a credible estimate of the effect of class size — a question that decades of ordinary regression had failed to settle, because small classes are given to particular kinds of pupil in particular kinds of school.
Two things you must check.
Manipulation of the running variable. If people can precisely control which side they land on, the design is dead. Look for a suspicious bunching of cases just on the favourable side — the standard density test. Where the threshold is a test score marked by a person who knows the consequence, expect bunching. Where it is a birth date or a vote count, expect none.
And what it estimates. A local effect, at the threshold, for cases near it. The effect of class size around 40 tells you little about the effect around 15. Regression discontinuity buys extremely high credibility in exchange for extremely narrow scope , and that trade is the honest description of it.
Three — instrumental variables.
The situation : you cannot assign the treatment, but you can find something that nudges people into it for reasons unrelated to the outcome.
The move : use the nudge — the instrument — to isolate the part of the treatment's variation that is as-good-as-random, and use only that part.
The cleanest example is the Vietnam draft lottery. Whether an American man served in Vietnam was not random — volunteers and avoiders differed in every way that predicts earnings. But the draft lottery number, assigned by date of birth, was random , and it strongly raised the chance of serving. Angrist used it to estimate the effect of military service on later earnings, finding a substantial long-run earnings penalty for veterans.
Three requirements, and the middle one is the problem.
Relevance — the instrument really does shift the treatment. Testable, and it must be strong : weak instruments produce estimates that are badly biased and deceptively precise.
The exclusion restriction — the instrument affects the outcome only through the treatment. This is not testable. It is an argument , and it is where instrumental-variable studies live or die.
Independence — the instrument is unrelated to everything else that matters.
The cautionary tale is famous. Angrist and Krueger used quarter of birth as an instrument for years of schooling — compulsory-schooling laws let those born earlier in the year leave with less education. Elegant. Bound, Jaeger and Baker then showed that the instrument was very weak, that with weak instruments the estimator drifts towards the biased one it was meant to fix, and — devastatingly — that they could produce similar-looking results using entirely random numbers as instruments. They also noted that season of birth is itself related to family background and to health, which would violate the exclusion restriction directly.
And what it estimates is not the average effect. It is the local average treatment effect (LATE) — the effect among compliers , the people whose treatment status was actually changed by the instrument. The men who served because their lottery number was low are not all men , and the effect for them need not be the effect for volunteers.
Four — matching, propensity scores and fixed effects.
Matching pairs each treated case with an untreated case that looks the same on measured characteristics. Propensity score matching compresses those characteristics into a single estimated probability of being treated, and matches on that.
What this buys : a comparison that is balanced on the things you measured, without imposing a particular functional form.
What it does not buy — and this is the point — is any protection against unmeasured confounders. Matching is not a substitute for randomisation; it is a tidier form of controlling. If motivated people enrol and motivation is unrecorded, matched groups differ on motivation exactly as unmatched ones did.
Fixed effects designs compare each unit to itself over time, which removes every stable characteristic — measured or not. Powerful, and its limit is precise : it handles unobserved things that do not change, and nothing about things that change at the same time as the treatment.
Sibling, cousin and twin designs are the same idea using families: comparing siblings removes everything shared in the family. Discordant identical twins remove genetics as well. Their limits are equally precise : whatever made the siblings differ may itself explain the outcome, and the reason one twin took a different path is unlikely to be random.
And synthetic control is the most elegant recent addition. When one unit is treated — a state, a country, a city — construct a weighted combination of untreated units that reproduces the treated unit's pre-treatment trajectory , and use that composite as the counterfactual. Abadie and colleagues' study of California's tobacco control programme built a "synthetic California" from other states and tracked the divergence in cigarette consumption after 1988. The method's discipline is that the weights are chosen to fit the pre-period, before the outcome is examined.
Telling a real one from a hopeful one
Five things a credible quasi-experimental study does, and a weak one omits.
One — it names the assignment mechanism explicitly and defends it. Which company piped your street. Your lottery number. Your birth month. Whether your cohort had 40 or 41 pupils. If you cannot state what did the assigning in one sentence, it is not a natural experiment , however the abstract describes it.
Two — it shows pre-treatment balance. The groups should look alike on characteristics measured before the treatment. If they differ substantially beforehand, the "as good as random" claim is already refuted.
Three — it shows pre-trends. For difference-in-differences, several periods before the change, moving together. This is the single most informative graph in the applied social sciences , and its absence is the single most common tell.
Four — it runs placebo and falsification tests. Apply the same design where there should be no effect: a fake treatment date, an outcome the policy could not have touched, a group that was not treated. If the "effect" shows up there too, the design is picking up something else.
Five — it shows the result is not an artefact of one specification. Bandwidth choices, control sets, functional forms, sample restrictions. A finding that appears at one bandwidth and vanishes at the next is a finding about the bandwidth. The disciplined version is to report the whole family of reasonable specifications — a specification curve — rather than the one that worked.
Four ways these designs are abused, and they are extremely common.
"Natural experiment" as branding. An ordinary before-and-after comparison, or a comparison of places that differ in a policy, described in experimental language with no assignment mechanism identified. The phrase has become a claim to credibility rather than a description of a design.
Parallel trends asserted, never shown. One before-period, one after-period, two groups, a confident causal conclusion. You cannot check the assumption and neither did they.
A running variable someone controls. Regression discontinuity at a threshold where the person allocating knows the consequence — an examiner marking a borderline script, an official assessing income against an eligibility line. The cases near the line are precisely the manipulated ones.
An instrument that plausibly affects the outcome directly. Rainfall as an instrument for economic conditions — but rainfall also affects health, mood, movement and violence directly. Distance to a facility as an instrument for using it — but people choose where to live. Every such study rests on an argument that the direct path does not exist, and that argument deserves at least as much attention as the statistics.
And the failure that spans all four: the garden of forking paths. With many defensible choices — which years, which comparison group, which controls, which bandwidth, which outcome — something will be significant somewhere. Nothing in the reported analysis reveals how many paths were walked (see 7.6.3, 7.8.2).
Because most of what you will be told about the effects of policies, laws and institutions comes from these designs, not from trials.
Four questions, in order, and they will take you further than any statistical training.
One — what assigned the treatment? Answerable in one sentence, or the study is observational with better vocabulary.
Two — is there any reason that assignment mechanism relates to the outcome? Not "is it random" — almost nothing is — but is the story of how it was assigned free of the outcome .
Three — what did they show me about the assumption? Pre-trends, balance, placebo tests, density at the cutoff. A paper confident about its identification and silent about its assumptions has told you which it cared about.
Four — for whom is this the effect? RD gives you the threshold. IV gives you the compliers. Both are local, and both are routinely reported as though they were general.
And the constructive half. These designs are the reason sociology can say anything credible about class size, incarceration, minimum wages, immigration, health insurance and neighbourhood effects — questions where randomisation is impossible and where, thirty years ago, the discipline had ideology and correlations. The credibility revolution was real. What it did not do is remove the need for judgement: every one of these designs replaces an untestable assumption with a different, more plausible untestable assumption , and the whole skill is in assessing how plausible.
A natural experiment is a situation in which something other than a researcher assigned the treatment, for reasons unrelated to the outcome. Snow's South London study is the founding example: two water companies with intermingled pipes on the same streets, one of which moved its intake upstream — cholera death rates differing roughly eight- or nine-fold, with the assignment made by a forgotten commercial rivalry.
Difference-in-differences subtracts the comparison group's change from the treated group's change. It stands or falls on parallel trends , which cannot be tested directly but can be probed with several pre-periods — and its classic case, Card and Krueger on the New Jersey minimum wage, remains contested largely over which data measured employment properly.
Regression discontinuity compares cases just either side of a sharp threshold — Maimonides' rule splitting Israeli classes at 40 pupils. Check for manipulation of the running variable, and remember it estimates a local effect at the cutoff only.
Instrumental variables use something that nudges people into treatment for unrelated reasons — the Vietnam draft lottery being the cleanest. Relevance is testable; the exclusion restriction is an argument, not a test ; weak instruments are badly biased and deceptively precise, as the quarter-of-birth critique demonstrated. The estimate is a LATE — the effect on compliers.
Matching and propensity scores handle observed confounders only. Fixed effects remove stable unobservables and nothing time-varying. Sibling and twin designs remove what is shared and leave whatever made them differ. Synthetic control builds a weighted counterfactual that reproduces the pre-treatment path.
A credible study names the assignment mechanism, shows pre-treatment balance, shows pre-trends, runs placebo tests, and demonstrates the result survives reasonable alternative specifications. The abuses are branding without a mechanism, parallel trends asserted, a manipulable cutoff, an instrument with a direct path — and the forking paths that hide behind all of them.
Natural experiment — assignment made by a process outside the researcher's control and plausibly unrelated to the outcome.
Difference-in-differences — the treated group's before-after change minus the comparison group's.
Parallel trends — the assumption that the groups would have moved together absent the treatment; probed with pre-periods.
Regression discontinuity — comparing cases just above and below an assignment threshold. Running variable — the measure the threshold is applied to.
Density / manipulation test — checking for bunching just on the favourable side of a cutoff.
Instrumental variable — something that shifts treatment but affects the outcome only through it.
Relevance / exclusion restriction / independence — the three IV requirements; the second is untestable.
Weak instrument — one that shifts treatment only slightly, producing biased and falsely precise estimates.
LATE / compliers — the local average treatment effect, among those whose treatment the instrument actually changed.
Matching / propensity score — balancing on measured characteristics; no protection against unmeasured ones.
Fixed effects — comparing units to themselves, removing stable unobserved differences only.
Synthetic control — a weighted combination of untreated units built to reproduce the treated unit's pre-treatment path.
Placebo / falsification test — applying the design where no effect should exist.
Garden of forking paths — the many defensible analytic choices that make some significant result nearly certain.
One — name the assigner. Take any study described as a natural experiment and write, in one sentence, what did the assigning. If you cannot, that is the finding.
Two — demand the pre-trend. Find a before-and-after policy comparison in the news. Ask whether the outcome was already moving before the policy. This alone overturns a surprising share of confident claims.
Three — find a discontinuity. Identify a threshold in your own life or work — an age cutoff, a score, an income line, a size rule that changes what an organisation must do. Sketch how you would use it, and whether anyone can manipulate which side they land on.
Four — break an instrument. Take any instrumental-variable study and try to think of one way the instrument could affect the outcome other than through the treatment. If you can think of a plausible one in two minutes, so could the referees, and you should check what the authors said about it.
Five — ask "for whom". For any quasi-experimental result you have read, write down which subgroup the estimate actually applies to. Then check whether the abstract said so.
So far every design has assumed you are collecting your own data. Most social research does not. It uses data somebody else produced — censuses, panel studies, administrative records, registers, and the enormous exhaust of digital platforms.
That data is cheap, vast and already there. It was also produced by an institution, for a purpose, about the people that institution processes — and every one of those facts leaves a mark on what you can conclude.
7.4.4 — Secondary Data, Official Statistics and Big Data covers what these sources are excellent for, how their categories were made, and the specific ways they mislead people who forget where they came from.