Everything so far has been about description and association. This lesson is about cause — and about the one design that answers causal questions cleanly, why it works, and the large territory where it cannot go.
The logic of the experiment matters even to people who will never run one , because it is the standard against which every other causal claim in social science is judged, including all the claims that cannot be tested this way.
Five thousand job applications, and the names on them.
Employers, asked whether they discriminate, say no. Applicants, asked whether they were discriminated against, cannot know — they were not told why they were rejected, and neither was anyone else. Both accounts are unfalsifiable, which is why the argument had run for decades without moving.
So Marianne Bertrand and Sendhil Mullainathan built the counterfactual by hand.
They wrote roughly five thousand fictitious résumés and sent them to help-wanted advertisements in Boston and Chicago. The résumés varied in quality, experience and skills. And to each one they randomly assigned a name — either one strongly associated in that context with white applicants (Emily, Greg) or one strongly associated with African-American applicants (Lakisha, Jamal).
Randomly. The same résumé, on a different draw, would have carried a different name.
The white-sounding names received about 50 per cent more callbacks. Roughly one callback in ten, against one in fifteen. To put it in a currency employers understand: the size of the name penalty was comparable to about eight additional years of experience.
And a second finding, quieter and more damning: improving the résumé raised callbacks substantially for the white-sounding names and barely at all for the others. The return on qualifications was not the same.
Devah Pager's study, published a year earlier, did the same thing with people rather than paper. She trained matched pairs of young men — similar in age, appearance and manner, differing in race — sent them to apply for entry-level jobs in Milwaukee with identical résumés, and randomly rotated which of them reported a criminal conviction.
The callback rates: white applicant with no record, 34 per cent. White applicant with a record, 17 per cent. Black applicant with no record, 14 per cent. Black applicant with a record, 5 per cent.
Read the middle two numbers again. A white man with a felony conviction was called back at least as often as a black man with a clean record.
And notice what makes these studies unanswerable in a way that surveys are not. Nobody was asked what they would do. Nothing was controlled for statistically, because nothing needed to be : the résumés were identical, and the only thing that varied was assigned by chance. There is no third variable. There cannot be one.
The fundamental problem: the comparison you need does not exist.
To say that something caused an outcome is to say that without it, the outcome would have been different. That is a claim about a world that did not happen — a counterfactual .
Suppose Lakisha's application was rejected. The causal question is: would this identical application, from this identical person, have been accepted if the name had been Emily?
We cannot observe that. The application was sent with one name. The other version of the world is unavailable. Paul Holland named this the fundamental problem of causal inference: for any single unit, we can observe the outcome under treatment or the outcome without it, never both.
And every causal method in existence is a strategy for manufacturing a substitute for the missing observation.
Compare the person to themselves earlier — but other things changed too. Compare them to someone similar — but "similar" is doing enormous work, and the ways people differ that you cannot measure are exactly the ways that matter (see 7.3.2). Compare groups statistically — but you can only control for what you thought of and measured.
Randomisation is the one strategy that does not require you to have thought of the confounders. That is its entire claim to superiority, and it is a large one.
What randomisation does, stated precisely
It does not make the groups identical. It makes the procedure unbiased.
This distinction is where most misunderstanding lives, so it is worth doing slowly.
Assign a large group at random to treatment and control. On average across all possible random assignments, the two groups have the same distribution of everything — age, education, motivation, health, family background, and every unmeasured thing you never thought about, including things nobody has a name for.
On average. In any particular assignment, they will differ somewhat by chance. That is not a flaw; it is exactly what the statistical machinery is built to account for , and it is what the confidence interval around an experimental estimate is measuring.
So the correct statement is: randomisation makes the treatment assignment independent of everything that came before it. Any difference in outcome must therefore be caused by the treatment or by chance — and the size of the "or by chance" part is calculable, which is the whole trick.
Compare that with statistical control. Controlling for measured confounders handles the confounders you measured. It does nothing about the ones you did not, and in social life the decisive ones are typically unmeasured: motivation, family support, information, health you did not record, the reason someone signed up at all. Randomisation handles the unknown ones without knowing them. No other method does.
Random assignment ≠ random sampling. They do entirely different jobs and are constantly confused.
Random sampling determines who is in the study , and governs generalisation to a population (see 7.3.1).
Random assignment determines who gets the treatment , and governs causal inference within the study.
A study can have either, both or neither. A lab experiment on 200 undergraduates has excellent random assignment and no random sampling: strong internal validity, unknown generalisation. A national survey has excellent sampling and no assignment: strong description, weak causal claims.
Average treatment effect (ATE) — the average difference in outcome between everyone being treated and everyone not.
Internal validity — whether the causal claim holds within the study . External validity — whether it holds elsewhere, for other people, at other times, at other scales.
The anatomy of a good experiment
Seven design elements, and what each defends against.
A control group. Without one, you are comparing to nothing, and every threat below applies. "Participants improved after the programme" is not evidence the programme worked — people improve for many reasons.
Random assignment , done by a mechanism nobody can influence. The commonest corruption in real field trials is assignment being quietly overridden — a caseworker who thinks this family really needs the programme.
Blinding. Participants unaware of their assignment (single-blind); those assessing the outcome unaware too (double-blind). Often impossible in social interventions — people know whether they received job training — which is why the outcome measure should be as objective as possible when blinding cannot be achieved.
A placebo or active control , so that receiving something is held constant and the effect of attention alone is not counted as the effect of the intervention (see 7.2.2 on reactivity).
Pre-registration of hypotheses, outcomes and analysis, before the data exist. In a trial with fifteen outcome measures, something will be significant — pre-specifying the primary outcome is what prevents that from being reported as the finding (see 7.6.3).
Intention-to-treat analysis. Analyse everyone by the group they were assigned to, whatever they actually did. Analysing only those who completed the programme destroys the randomisation , because completing is a choice, and it reintroduces exactly the selection the design was built to remove (see 7.3.2).
Adequate power , decided in advance. An underpowered trial has two failure modes, and the second is less known: it usually finds nothing when there is something, and when it does find something, the estimate is inflated — because only large fluctuations clear the threshold in a small sample.
Four kinds of experiment sociologists actually run.
Laboratory experiments. Tight control, precise manipulation, clean measurement. Strong internal validity, questionable ecological validity — behaviour in a room with a researcher and a small stake is not behaviour in a life. Best for demonstrating that a mechanism can operate.
Survey experiments. Randomisation inside a questionnaire, and now the workhorse of attitude research (see 7.4.1). Vignette or factorial designs present a described situation with randomly varied attributes — a welfare claimant's age, work history and family situation independently randomised — so the effect of each on the judgement can be estimated separately. List experiments estimate the prevalence of sensitive attitudes without any individual disclosing anything. Conjoint designs present pairs of profiles with many randomised attributes and ask the respondent to choose, recovering the weight given to each.
Field experiments. Randomisation in a real setting, with real consequences, usually among people who do not know they are in a study. The audit and correspondence studies above are the outstanding sociological example , and there is no other way to measure discrimination in the act.
Policy trials. Randomised evaluation of a real programme. Often cluster-randomised — schools, clinics, villages assigned rather than individuals, because the treatment operates at that level and would spill over otherwise. Sometimes stepped-wedge , where everyone eventually receives the programme but the order of rollout is randomised, which is often the only politically acceptable design.
Moving to Opportunity, and what a good experiment can teach you by disappointing you.
In the 1990s, thousands of families in high-poverty American public housing were randomly offered housing vouchers, some restricted to low-poverty neighbourhoods. The question was whether the neighbourhood itself shapes life chances — a question the discipline had argued about for decades using observational data that could never settle it, because who lives where is not random.
The first wave of results, roughly a decade on, was widely read as a disappointment. Adults who moved showed no significant improvement in employment or earnings. They did show substantial improvements in mental and physical health — large reductions in distress and in obesity and diabetes markers — and effects on girls' wellbeing that did not appear for boys.
Then, roughly twenty years after the moves, a re-analysis using tax records changed the picture. Children who moved to a lower-poverty neighbourhood before about age 13 had markedly higher college attendance and substantially higher adult earnings. Children who moved as adolescents did not benefit and may have been slightly harmed.
Three lessons, and they are the reasons this example is in the lesson.
Average effects can conceal everything that matters. The headline "no effect" was true of the average and false of most of the people in it, in opposite directions.
Timing decides the answer. Measured at ten years, the programme looked ineffective on economic outcomes. The mechanism — childhood exposure — could not have produced a measurable earnings effect by then, because the children were not yet earning. Many programmes are evaluated on a horizon shorter than their mechanism.
And a good experiment is valuable even when it disconfirms. It killed some plausible stories, established that neighbourhood effects are real and operate through childhood exposure, and did all of this on evidence that observational studies had been arguing over inconclusively for thirty years.
The threats, and what experiments cannot reach
The classic threats to internal validity — which are exactly what a control group defends against.
History — something else happened during the study, to everyone. Maturation — participants changed simply through time. Testing — being measured changed them. Instrumentation — the measure or the measurer drifted. Regression to the mean — cases selected because they were extreme move towards average on their own, which makes any intervention aimed at the worst-performing schools, hospitals or offenders look effective. Selection — the groups differed to begin with. Attrition — differential dropout after assignment (see 7.3.2).
A randomised design with a control group neutralises nearly all of these at once , because both groups experience history, maturation, testing and regression equally. That is the source of its power, and it is worth stating plainly: the control group is doing most of the work, and randomisation is what makes the control group fair.
Six things randomisation cannot do, and they are large.
One — most sociological causes cannot be assigned. You cannot randomise someone's caste, race, gender, nationality, class origin, religion or the century they were born in. You cannot randomise poverty, incarceration, war or a revolution. Audit studies get round this brilliantly by randomising the signal of a category rather than the category — which is why they are so celebrated, and it should be noted that what they measure is the effect of a name or a stated record, which is related to but not identical to the effect of being that person.
Two — you cannot randomise macro-structures. Institutions, states, markets, legal systems, historical sequences. The questions in Parts 4, 5 and 6 of this course are almost entirely outside experimental reach , which is why comparative-historical method exists (see 7.5.4).
Three — external validity is not delivered by randomisation at all. An experiment establishes an effect for those participants, in that setting, at that time, at that scale, delivered by those people. Nothing about random assignment tells you it will hold elsewhere. Extrapolating requires theory about why it worked — which is an argument, not a result.
Four — scale changes effects. A job-training programme that helps its graduates compete for jobs may do nothing when everyone receives it, because the jobs are the constraint. General equilibrium effects are invisible to a trial that treats a small fraction of a market, and they are the difference between a promising pilot and a disappointing rollout.
Five — the black box. An experiment can establish that something worked without any account of why . Two programmes with the same average effect may operate through completely different mechanisms, transfer to different places, and require different things to keep working. This is why the strongest evaluations combine a trial with qualitative work on mechanism (see 7.8.3).
Six — the average may describe nobody. An ATE of zero is consistent with helping half and harming half — which is Moving to Opportunity's lesson, and it is common.
The serious critique, stated fairly.
Angus Deaton and Nancy Cartwright's argument is the one to know, because it is made by people who understand the machinery rather than by people uncomfortable with numbers.
Their points, compressed. Randomisation guarantees balance only in expectation , over the space of possible assignments — not in your trial, and with modest samples imbalance on important variables is common. The ATE is a particular quantity that may not be the one anyone needs. Trial results do not transport by themselves : using a result somewhere else requires knowing which features of the original setting mattered, which is theoretical knowledge that the trial did not produce. And a hierarchy of evidence that automatically places RCTs above everything else discards information — including prior knowledge, mechanism, and well-identified observational work.
Their conclusion is not that trials are bad. It is that an RCT is one piece of evidence, whose interpretation requires exactly the theoretical and contextual knowledge that its enthusiasts sometimes present it as replacing.
And there is a real cost to the alternative view. When a discipline rewards questions that can be randomised, questions that cannot be randomised get asked less — which quietly redistributes attention from structures to interventions , from causes that are large and immovable to causes that are small and assignable. That is a sociological observation about a methodological fashion, and it belongs in a sociology course.
The ethics are not a footnote here (see 7.8.1).
Field experiments impose real costs on people who did not consent. Audit studies occupy employers' time with applications from people who do not exist. Landlords show flats to fictitious tenants. Real applicants may be affected.
The defence is serious and it is not a formality : there is no other way to measure discrimination as it happens, the knowledge is of substantial public value, the burden per employer is small, and no individual is identified or harmed. Ethics committees weigh exactly this , and the studies above were approved on exactly these grounds.
And randomised policy trials have their own version : someone is assigned to the control group, and if the programme works, they did not get it. The justification is that we did not know it worked — which is honest when true, and is the whole reason for running the trial. It stops being honest when a trial is run on something already well evidenced.
Because "compared to what?" is the question that dismantles most causal claims you will meet.
A programme reports 70 per cent of participants found work. Compared to what? What share of similar people who did not participate found work? Without that number the 70 means nothing — and if the comparison group is people who did not enrol, it means less than nothing, because enrolling is a choice (see 7.3.2).
A school's results improved after a new head. Compared to what — other schools that year, or the same school's own trajectory? And was the head appointed because results were unusually bad, in which case regression to the mean predicts improvement with no cause at all?
A country changed a law and the rate fell. Compared to what — countries that did not change it? Was it already falling?
Three reading habits follow.
Find the counterfactual, or note that there isn't one.
Ask whether assignment was chosen or imposed — the single most informative fact about any evaluation.
When you see an average effect, ask who it might be concealing.
And the honest caution against over-correcting. Most of what sociology knows was not learned from experiments and could not have been. The right response to the experimental standard is not to dismiss everything that fails it, but to ask what each design does to substitute for the missing counterfactual — which is the subject of the next lesson.
Causation is a claim about a counterfactual, and the counterfactual is never observed — Holland's fundamental problem. Every causal method is a strategy for substituting for it.
Randomisation is the only strategy that handles confounders you never thought of , because it makes assignment independent of everything prior. It does not make the groups identical ; it makes the procedure unbiased and the residual uncertainty calculable. Random assignment governs causal inference; random sampling governs generalisation; they are constantly confused.
Audit and correspondence studies are sociology's outstanding use of the design. Identical résumés with randomly assigned names produced about 50 per cent more callbacks for white-sounding names, a gap worth roughly eight years of experience — and the return on a better résumé was far larger for those names. Pager's matched-pair audit found a white applicant with a criminal record called back at least as often as a black applicant without one.
Good design elements : control group, uncorruptible assignment, blinding where possible, placebo or active control, pre-registration, intention-to-treat, adequate power — and an underpowered trial that does find an effect has overestimated it.
Randomisation neutralises history, maturation, testing, instrumentation, regression to the mean and selection at once — mostly through the control group, which randomisation makes fair.
What it cannot do : assign most sociological causes; touch macro-structures; deliver external validity; anticipate general equilibrium effects at scale; explain mechanism; or prevent an average from describing nobody. Moving to Opportunity demonstrated the last two — an apparent null on adult economics concealed large gains for children who moved young, invisible on a ten-year horizon.
And Deaton and Cartwright's critique stands : balance only in expectation, an ATE that may not be the quantity needed, no automatic transportability, and a hierarchy that discards mechanism and prior knowledge. A trial is a strong piece of evidence that still requires theory to use.
Counterfactual — what would have happened to the same unit without the treatment.
Fundamental problem of causal inference — only one of the two potential outcomes is ever observed for any unit.
Random assignment / random sampling — who gets treated; who is in the study. Causal inference; generalisation.
Average treatment effect (ATE) — the mean outcome difference between treated and untreated states.
Internal / external validity — whether the causal claim holds within the study; whether it holds elsewhere.
Intention to treat — analysing by assigned group regardless of compliance.
Blinding / placebo / active control — concealing assignment; holding constant the effect of receiving something.
Regression to the mean — extreme cases moving towards average without any cause, making interventions on the worst look effective.
Correspondence / audit study — randomly varying an applicant signal or using matched testers to measure discrimination directly.
Vignette, factorial, list and conjoint designs — randomisation embedded within a survey.
Cluster randomisation / stepped wedge — assigning groups; randomising the order of a universal rollout.
General equilibrium effects — effects that appear only when an intervention operates at scale.
Heterogeneous treatment effects — differing effects across subgroups that an average conceals.
One — ask the question. Take three claims that something worked — a policy, a course, a diet, a management change. For each, write down what the comparison group is. Where there isn't one, note what the claim rests on instead.
Two — find the assignment. For any evaluation you read, determine whether people chose to be in the treated group or were placed there. This one fact predicts most of what is wrong with most evaluations.
Three — design an audit study. Pick a setting where you suspect unequal treatment — rental enquiries, service response, quotation requests. Sketch what you would randomise and what you would hold constant. Then list the ethical objections and how you would answer them. Both halves are the exercise.
Four — hunt regression to the mean. Find an intervention targeted at the worst-performing cases — schools, hospitals, teams, offenders. Ask what would have happened to the worst performers with no intervention at all.
Five — split an average. Take a reported null result from an evaluation and write down two subgroups for whom the effect might run in opposite directions, with a reason for each. Then check whether the study looked.
Most of the causes sociology cares about cannot be assigned by a researcher. But they are sometimes assigned by something else — a lottery, a policy that started on a particular date, an administrative cutoff, a border, a rule that applied to people born one month and not the next.
When the world does the randomising, you can recover much of an experiment's logic without running one.
7.4.3 — Natural Experiments and Quasi-Experimental Designs covers difference-in-differences, regression discontinuity, instrumental variables and matching — what each requires to be believable, and how to tell a real natural experiment from a hopeful one.