Everything in this Part has been about how a single study can go wrong. This lesson is about what happened when several fields stopped assessing studies one at a time and asked a different question:
If we run these again, do we get the same answers?
The results, from about 2011 onwards, reorganised how a generation of researchers thinks about evidence. They also produced an over-correction that is now its own problem , and this lesson tries to hold both.
One hundred studies, run again.
Beginning in 2011, a large collaboration of researchers — eventually more than 250 people — set out to replicate 100 published studies from three of psychology's most prominent journals. Each replication used the original materials wherever possible, was reviewed in advance by the original authors where they were willing, and was designed with higher statistical power than the original.
Of the 100 original studies, 97 had reported statistically significant results.
Of the 100 replications, 36 did.
And the effect sizes were on average about half the size of those originally reported. Where the replications did find effects, they found smaller ones.
The findings were published in 2015 and produced an immediate and useful argument. A group of critics responded that the replications were not always faithful to the originals — that some had been run on different populations in different countries with altered materials — and that once the statistical uncertainty of the replications themselves was accounted for, the results were less damning than the headline suggested. The original team replied that the critique's own assumptions were doing a great deal of work.
The exchange is worth knowing about , because it is what a healthy field looks like. And it did not overturn the central result , which has since been reproduced in adjacent efforts.
In experimental economics , a systematic replication of 18 laboratory studies from top journals found around 11 replicated. In a project targeting social science studies published in Nature and Science , 13 of 21 replicated, again with effect sizes roughly half the originals.
And the multi-laboratory projects added the most useful nuance of all. When the same set of effects is tested simultaneously in dozens of laboratories around the world, some replicate robustly everywhere , some replicate in some places and not others, and some show essentially nothing anywhere. The picture is not "psychology is broken". It is that the published literature did not distinguish between these three kinds of finding, and the field could not tell them apart from the inside.
Some individually famous casualties. The finding that adopting an expansive posture changes hormones and risk-taking failed to replicate on the physiological measures — and one of the original authors subsequently stated publicly that she no longer believed the effect was real , which is among the more admirable acts in recent social science. Large pre-registered multi-site tests of willpower depletion found an effect close to zero. A pre-registered multi-lab test of the claim that holding a pen in the teeth makes cartoons funnier found nothing. And the celebrated demonstration that priming people with words about old age made them walk more slowly failed when the experimenters were kept blind to condition — which points directly at where the original effect may have come from.
What does it actually mean that a finding "failed to replicate"?
Three different things are called replication and they are not the same.
Reproducibility — same data, same analysis, same result. This checks the arithmetic and the code. It fails more often than anyone expects : one systematic attempt in economics could reproduce the published results of only about a third of papers using the authors' own data without contacting them, and around half with author assistance.
Replicability — new data, same procedure, same result. This is what the projects above tested.
Generalisability — different population, different setting, different time. A finding can be perfectly replicable in the original population and not travel at all (see 7.4.2 on external validity), which is a fact about the world rather than a defect in the study.
And a failed replication does not establish that the original was wrong. There are six reasons a replication fails, and only some of them indict anybody.
One — the original was a false positive. Produced by the mechanisms of 7.6.3: analytic flexibility, low power, publication selection.
Two — the original effect is real but was exaggerated. The winner's curse; the replication finds the true, smaller effect and misses significance.
Three — the replication was itself underpowered , or badly executed.
Four — the effect is real and conditional. A hidden moderator: the effect exists in that population, that period, that culture, that laboratory. This is the defenders' standard argument, it is sometimes correct, and it is unfalsifiable unless the moderator is specified in advance and tested. A theory that acquires a new boundary condition every time it fails is Lakatos's degenerating programme (see 7.2.3).
Five — procedural differences. Different stimuli, different instructions, an online sample instead of a laboratory.
Six — fraud. Rare, real, and not the main story. The largest single case in social psychology involved dozens of retracted papers built on fabricated data ; and investigations into prominent behavioural work on honesty have led to retractions after evidence of data fabrication, which is a coincidence with an unpleasant symmetry. But fraud does not account for a 36 per cent replication rate. Ordinary practice does.
Why it happened: the machine that produces it
Four causes, none of which requires anyone to behave badly.
One — publication bias. Journals prefer positive, clean, surprising findings. Null results are harder to publish, so researchers do not write them up, so they sit in the file drawer. The published literature is therefore a biased sample of the research conducted (see 7.3.2 — it is survivorship bias applied to studies).
The most striking single piece of evidence for this comes from clinical trials. When American regulators began requiring large publicly funded trials to be pre-registered with their primary outcomes stated in advance, the proportion of large cardiovascular trials reporting a positive result for the primary outcome fell from around 57 per cent before the requirement to around 8 per cent after it.
Read that again. The trials did not get worse. The freedom to decide afterwards what the outcome had been was removed, and most of the positive findings went with it.
Two — low power. Typical studies are too small to detect the effects they study (see 7.6.3), which means both that real effects are missed and — the crucial part — that the effects which do clear significance are exaggerated. A literature built from underpowered significant results is a literature of overestimates.
Three — analytic flexibility. The garden of forking paths. Not fraud: ordinary, defensible choices, made after seeing the data, in a context where a significant result is the route to publication.
Four — incentives. Careers are built on publications in prominent journals, which want novelty. Nobody is promoted for a replication , few are funded to do one, and until recently most journals would not publish one. A field can therefore produce and reward findings with no mechanism for discovering that they are wrong — which is precisely what happened.
And now the part that concerns this course directly: what about sociology?
Sociology has been less visibly affected, and the reasons are not entirely flattering.
The honest defence. Much of the discipline is not experimental, and does not depend on small laboratory effects with large analytic flexibility. Large secondary datasets are shared and re-analysed, which is a form of checking. Qualitative work is not in the same position at all, since it makes different claims (see 7.7.2). And demographic and stratification research works with big, well-measured quantities and long series.
The uncomfortable part. Most sociological findings have never been re-tested by anybody. A discipline with few replications has not passed a test; it has not taken one. The absence of a visible crisis is not evidence of health — it is 7.4.4's lesson about a field with only one measure.
And where sociology has been checked, the results are sobering in a distinctively sociological way.
Two large crowdsourced studies asked many independent research teams to answer the same question using the same data. In one, 29 teams were given identical data and asked whether football referees are more likely to give red cards to darker-skinned players. They produced widely varying estimates — about two-thirds found a statistically significant positive effect and about a third did not, with effect sizes ranging from negligible to very large. All were using the same data.
In the second and larger project, 73 teams analysed the same data to test whether immigration reduces support for social policy — a central question in political sociology. The estimates were dispersed across the entire plausible range, including a substantial number in each direction , and the researchers could not explain the variation by teams' methodological choices, their expertise, or their prior beliefs. A great deal of it was simply the accumulation of small, defensible decisions.
This is not a replication problem in the psychological sense. Nobody's data were noisy and nobody's sample was small. It is analytic flexibility in observational research , and it says that the answer to a well-posed sociological question, on fixed data, is substantially indeterminate across competent analysts.
That finding deserves to unsettle you more than the psychology results do , because it applies to the kind of research this discipline mostly does.
What has been done, and whether it works
Six reforms, in rough order of impact.
Pre-registration. Stating hypotheses, sample size, outcomes and analysis before data collection, in a public time-stamped record. It closes the garden of forking paths by making exploratory and confirmatory analysis distinguishable — and it does not forbid exploration, only mislabelling it.
Registered reports. The stronger version: the introduction and method are peer-reviewed and accepted before the data exist , with publication guaranteed regardless of the result. This removes the incentive to produce a positive finding entirely , and the evidence on what that does is remarkable: studies published as registered reports report positive results at rates far below the standard literature — on the order of 40-something per cent, against something like 95 per cent in conventional papers. The gap is a direct measurement of what the incentive structure was doing.
Open data and code. Deposit what is needed to check the analysis. This addresses reproducibility, which as noted is a real and separate failure , and it is now required by many journals and funders — with the qualitative caveats from 7.7.2.
Larger samples and multi-site collaboration. Many laboratories running the same protocol simultaneously, with a pre-agreed analysis. This is the strongest design available for establishing whether an effect exists , and it also measures how much it varies by site — which converts "hidden moderators" from an excuse into a finding.
Post-publication review and retraction. Public commenting platforms, statistical forensic methods that detect impossible or improbable data patterns, and — slowly — journals more willing to retract.
And structural change : journals publishing replications, funders financing them, and hiring and promotion criteria that credit reliability rather than only novelty. This is the slowest and the most important , because it addresses the cause rather than the symptom.
Meta-analysis is not the answer people think it is.
The intuitive response to unreliable individual studies is to pool them. And meta-analysis is valuable — it increases precision, tests whether effects vary systematically, and imposes a discipline of systematic rather than selective literature searching.
But it inherits publication bias entirely. A meta-analysis of a biased literature produces a precise estimate of a biased quantity. Garbage in, authoritative-looking garbage out.
Detection and correction methods exist and are imperfect. Funnel plots — plotting effect size against precision, and looking for the asymmetry that selective publication produces. Methods that estimate and adjust for the missing studies. Selection models that formalise how publication depends on the result. All of them rest on assumptions that are frequently violated , and several have been shown to perform poorly under realistic conditions.
The strongest evidence remains a large pre-registered study or a coordinated multi-site replication — not a synthesis of a literature that was selected for its conclusions. When a meta-analysis and a large registered replication disagree, the current consensus favours the replication , and this has happened repeatedly.
Six rules, and they will serve you for the rest of your life.
One — never update strongly on a single study. Whatever the journal, whatever the p-value. A single study is a piece of evidence, not a fact.
Two — be most suspicious of the findings you most enjoy. Surprising, counterintuitive, media-friendly results are exactly the ones selection favours. The delight you feel on reading a striking finding is a signal about the selection process, not about the finding.
Three — prefer pre-registered, multi-site, high-powered work , and check whether the study you are reading is any of those. It usually takes thirty seconds.
Four — ask whether the famous version has been tested. For any well-known effect you can name from popular science, search for the replication. Some hold up completely; some do not exist.
Five — treat effect sizes from small significant studies as ceilings , not estimates.
And six — do not over-correct. "Nothing replicates" is as lazy as "it's published, so it's true", and it is becoming fashionable. A great deal does replicate : the multi-lab projects found effects that held up everywhere they were tested. Large, well-measured demographic and economic relationships are robust. The crisis is uneven, and it is worst where effects are small, samples are small, measurement is noisy and analytic freedom is large — which is a diagnosis you can apply to any specific literature.
Blanket cynicism is the same failure as blanket credulity : both are ways of not doing the work of assessing a particular claim on its particular merits, which is what this entire Part has been training you to do.
Because it changes what "the research shows" can mean.
For a citizen : the finding you heard about on the radio may be one underpowered study with a good press office, and knowing that is not cynicism but calibration.
For a professional using evidence — teaching, medicine, management, policy — the question is no longer "is there a study" but "how much of this literature would survive a coordinated replication". For many well-known behavioural interventions the honest answer is: less than you would hope.
And for the discipline this course is about , the crowdsourced analysis results are the ones to carry. They say that the choices made in observational analysis are consequential enough that competent researchers using identical data reach different conclusions — which is not a scandal about individuals but a fact about the method, and it makes the transparency arguments of 7.1.2 and 7.7.2 practical rather than moral.
The reforms all point the same way : state in advance what you will do, show the working, make it possible for someone who would enjoy refuting you to try. That is the same sentence this Part has arrived at from five different directions , and this lesson is the empirical demonstration that it was needed.
When 100 psychology studies from top journals were replicated with higher power, 97 originals had reported significant results and 36 replications did , with effect sizes about half the size. A serious critique followed and did not overturn the central result , which has been echoed in experimental economics and in replications of social science published in the most prominent general journals.
Multi-laboratory projects added the crucial nuance : some effects replicate everywhere, some vary by site, some show nothing — and the published literature could not distinguish the three.
Three distinct things are called replication : reproducibility (same data, same analysis — which fails surprisingly often), replicability (new data, same procedure), and generalisability (different setting, and its failure is a fact about the world).
Six reasons a replication fails : the original was a false positive; the original was real but exaggerated; the replication was weak; the effect is genuinely conditional (unfalsifiable unless the moderator is specified in advance ); procedural differences; and fraud, which is real and is not the main story.
Four causes, none requiring bad behaviour : publication bias — demonstrated most starkly by large cardiovascular trials, where positive primary results fell from around 57 per cent to around 8 per cent after pre-registration became mandatory; low power, which exaggerates the effects that survive; analytic flexibility; and incentives that reward novelty and never reward checking.
Sociology has been less visibly affected largely because it has been less checked — and where it has, the crowdsourced analyses are sobering: 29 teams on the same red-card data and 73 teams on the same immigration-and-social-policy data produced estimates dispersed across the plausible range , unexplained by expertise or method. That is analytic flexibility in observational research, and it applies to what this discipline mostly does.
The reforms : pre-registration; registered reports, which produce positive results at roughly half the rate of the conventional literature; open data and code; multi-site collaboration; post-publication review; and slow structural change in what careers reward. Meta-analysis inherits publication bias, and its corrections are imperfect — a large registered replication beats a synthesis of a selected literature.
And do not over-correct. The crisis is worst where effects are small, samples small, measurement noisy and analytic freedom large. Blanket cynicism is the same failure as blanket credulity.
Reproducibility / replicability / generalisability — same data and analysis; new data and same procedure; different population or setting.
Publication bias / file drawer — the selective publication of positive results, making the literature a biased sample of the research done.
Winner's curse — the exaggeration of effects that clear a significance threshold in underpowered studies.
Hidden moderator — a claimed boundary condition explaining a failed replication; legitimate only if specified in advance and tested.
Pre-registration — time-stamped public statement of hypotheses, sample and analysis before data collection.
Registered report — peer review and acceptance of the design before results exist, with publication guaranteed either way.
Multi-site / many-labs replication — the same protocol run simultaneously across laboratories with a pre-agreed analysis.
Crowdsourced analysis — many independent teams analysing identical data to measure analytic variability.
Funnel plot — a diagnostic for publication bias, plotting effect size against precision.
Post-publication review — public scrutiny and statistical forensics after publication.
Retraction — formal withdrawal of a published paper.
One — check a famous finding. Pick any well-known effect from popular psychology or behavioural economics and search its name with "replication". Note what you find, including if the answer is nothing.
Two — look for the pre-registration. For any recent study you read about, check whether a pre-registration exists and whether the reported outcomes match it. The mismatch rate in the literature is not small.
Three — build the incentive. Write down what would have to change for a researcher in your country to benefit professionally from publishing a null result. Then ask whether any of it is likely.
Four — apply the diagnosis. For a literature you care about, rate it on the four risk factors: small effects, small samples, noisy measurement, analytic freedom. That rating is a better guide to reliability than any individual paper's p-value.
Five — sit with the crowdsourced result. Take the immigration-and-social-policy finding: 73 teams, one dataset, answers across the range. Write two sentences on what you think that means for reading any single observational study — including one whose conclusion you agree with.
The last lesson of the toolkit topics returns to something that has been implied throughout: almost every hard question needs more than one method , and combining them well is a design problem rather than a matter of doing extra work.
7.8.3 — Mixed Methods and Triangulation covers the main designs and what each is for, how to decide which method leads, what to do when the methods disagree — which is the interesting case and the one most often fudged — and why the paradigm-war objection to mixing is weaker than it once seemed.