In 1966 a study was published that was expected to demonstrate one thing and was read as demonstrating almost the opposite.
It is still cited, still misquoted, and still shapes what people believe about whether schools matter. This lesson is about what it actually found, what it could not have found given its design, and what six decades of better evidence have since established.
The short answer, given now so that the rest of the lesson can be read properly: schools matter a great deal, differences between schools matter much less than people assume, and both of those statements are compatible.
The report that was supposed to prove one thing.
The American Civil Rights Act of 1964 required a survey. Congress wanted documentation of the inequality of educational opportunity by race, and the expectation — shared by essentially everyone who commissioned it — was that it would show large gaps in school resources between Black and white schools, and that those gaps would account for the gap in achievement.
The study that resulted was one of the largest ever conducted in social science : something like six hundred thousand students, sixty thousand teachers, four thousand schools.
It produced two findings that nobody was expecting.
The first : resource differences between schools serving Black and white students were smaller than anticipated — considerably smaller — on measures such as expenditure per pupil, class size, library books, laboratory facilities and teacher qualifications.
The second, and the one that made the report famous : when achievement was analysed, differences in school resources accounted for far less of the variation between students than family background did. Measures of a family's education, possessions and circumstances predicted achievement strongly. Measures of a school's resources predicted it weakly.
The finding entered public discussion in a compressed form: schools don't matter.
That is not what it said, and the compression matters because it licensed a policy conclusion — that spending on schools is futile — that the report's own author rejected.
Three things about the study are usually omitted.
One — its strongest school-level finding was about peers, not resources. The composition of the student body — the backgrounds and attainment of the other children in the school — predicted achievement more than any resource variable did. That is a school effect. It is not a resource effect, and it is the basis on which the report's author subsequently argued for desegregation as an educational policy rather than only a legal one.
Two — the design could not detect what people took it to have refuted. It was cross-sectional : it measured achievement at one moment, with no measure of what students had known when they arrived. So it could not distinguish what a school added from what it received (see 7.4.2). A school with excellent teaching and a disadvantaged intake looks identical to a poor school with the same intake.
Three — the method understates school effects by construction when schools are similar. A variance decomposition asks how much of the variation between students is associated with variation between schools. If the schools do not differ much on the measured inputs, the associated variance is small — regardless of how much schooling as such is doing. The report had found, in its first finding, that the schools did not differ much. The second finding is partly a consequence of the first.
None of this makes the report wrong. It makes it a study of the effect of differences between schools on achievement levels , which is a much narrower thing than the effect of schooling.
"Does school make a difference" is four questions with four different answers.
One — how much of the variation between individual students is associated with which school they attend? Answer: not much, in most systems. This is the Coleman question, and it is a question about between-school variance.
Two — does attending school at all make a difference? Answer: enormously. And it is a completely separate question, because a study comparing schools cannot address it — every child in it went to school.
Three — do some schools produce more learning than others, once intake is accounted for? Answer: yes, measurably, and by less than league tables suggest.
Four — can specific interventions improve outcomes? Answer: yes, some of them, by known amounts.
The confusion between the first and the second is the source of nearly all the misreading , and the distinction is exactly 7.6.4's: a variance decomposition tells you how much of the variation a factor is associated with in a particular population, not how much the factor does. In a country where every child attends a broadly similar school, the between-school variance will be small even if schooling is transformative.
Does schooling as such do anything? Two natural experiments
One — the seasonal comparison, and the caveat that came later.
A clever design asks what happens to achievement gaps during the school year and during the summer holiday. If schools amplify inequality, gaps should widen faster in term time. If schools compensate, gaps should widen faster when school is out.
The finding, from American data measuring children twice a year, was that gaps by family background grew substantially faster during the summer than during the school year — and on some measures barely grew during term at all.
The conclusion drawn was that schools are compensatory rather than amplifying : not that they eliminate the gradient, but that they slow its growth relative to what happens when children are at home.
And the caveat, which this course reports because it matters. The seasonal comparison depends on the assumption that a test's scale means the same thing at different points — that a given number of points represents the same amount of learning at the start of a year as at the end. Later work showed that some seasonal findings are sensitive to how the tests were scaled , and that with different defensible scalings the pattern changes.
So: the compensatory finding is suggestive and less secure than it was once presented as being. Which is why the second natural experiment matters so much more.
Two — the school closures, which produced the cleanest evidence anyone will ever get.
In 2020 schools closed. For social science this was an unwanted natural experiment of a kind nobody could have designed (see 7.4.3).
The cleanest study came from a country with near-universal broadband, high computer access, and — decisively — national standardised tests administered at fixed points in the year, so that a cohort had been tested shortly before the closure and was tested again shortly after. The closure lasted about eight weeks.
The result: learning loss of about one-fifth of a standard deviation — which, translated into the study's own terms, amounted to roughly the whole of what would have been learned in the eight weeks. Progress during the closure was, on average, close to zero.
And the loss was around 60 per cent larger for children from less-educated households.
Read what this establishes. In one of the best-resourced systems in the world, with remote provision, near-universal internet and highly educated parents, removing school produced learning loss equal to the time removed, and it fell disproportionately on the disadvantaged.
That is the answer to the second question , and it is not compatible with any reading of the 1966 report as showing that schools do not matter.
Do some schools do better than others?
Value-added, what it measures, and how large the differences are.
The measure that answers the third question compares a school's or a teacher's students' outcomes to what would be predicted from their prior attainment and characteristics. The difference is the value added .
Two things it fixes. It removes the intake problem that made the 1966 design uninterpretable. And it removes most of the reason that raw league tables are useless: a raw ranking of schools by results is very largely a ranking of their intakes , and reordering by value added moves schools around dramatically.
How big are the differences?
Between teachers , credible estimates put the difference between a teacher at the 75th percentile of effectiveness and one at the median at somewhere around a tenth to a fifth of a standard deviation of achievement per year. That is a real and non-trivial effect , and it accumulates across years.
And there is evidence that it persists into adult outcomes. Large administrative studies linking students to teachers and then to tax records have found that being taught by higher value-added teachers is associated with higher rates of college attendance and higher earnings in adulthood.
And the criticism is substantial and must be stated. A prominent falsification test asked whether value-added measures predict prior achievement — which they cannot legitimately do, since a teacher cannot affect what happened before they taught. In some data they did , which indicates that students are not randomly assigned to teachers in ways the model captures: better teachers get, or are given, systematically different students. The subsequent exchange between the researchers involved is worth reading as a model of how such a dispute proceeds , and it has not been fully resolved.
The proportionate position : teacher effects are real, they are moderate in size, they persist, and the individual value-added estimates are noisy and partly contaminated by sorting — which means they are defensible for research on averages and much weaker as an instrument for evaluating individual teachers, which is what they are often used for.
Between schools , once intake is accounted for, the differences in most systems are modest. And the amount of between-school variation differs enormously by country in a way that is itself the finding — it is high in systems that sort children into different school types early, and low in systems that keep them together. Which is 9.1.3's subject , and it means "how much do schools differ" is partly a question about how a system is designed.
What actually works, with magnitudes.
Class size. A large randomised experiment assigning early-grade children to small and regular classes found real effects — on the order of a fifth of a standard deviation — larger for disadvantaged and minority children, and detectable years later. Quasi-experimental evidence using enrolment caps points the same way (see 7.4.3).
It is a genuine effect, and it is expensive. Reducing class sizes across a system costs approximately proportionally, which is why the policy debate is about cost-effectiveness rather than about whether it works.
Tutoring, which has the strongest evidence base in the field. Meta-analyses of small-group and one-to-one tutoring find average effects around a third of a standard deviation — large by the standards of educational interventions — with the strongest results for tutoring delivered frequently, during the school day, by trained tutors, in small groups.
Teacher effectiveness , as above.
Early years provision , which has real effects with a specific pattern discussed in 9.1.4.
And a warning about magnitudes. An effect of 0.2 standard deviations is a serious educational effect and it is not transformative. Interventions that claim much larger effects are usually small studies (see 7.6.3 on the winner's curse), and the effects of many programmes fade over subsequent years even where longer-run outcomes are affected.
Three errors this literature invites.
Reading a variance decomposition as an effect size. "Schools explain 15 per cent of the variance" does not mean schooling produces 15 per cent of achievement. It means that in this population, with these schools, 15 per cent of the differences between students are associated with which school they attend — a quantity that depends on how similar the schools are (see 7.6.4).
Comparing raw results. A school's raw examination results are dominated by its intake. Any comparison, ranking or inspection judgement based on them is largely a measure of catchment — which, given 8.3.3, is largely a measure of house prices.
And treating value added as a measure of teaching. It is a measure of everything that happened to a class in a year, including the other students, the school's stability, disruption, illness, and chance. Its noisiness is well documented , and a single year's estimate for an individual teacher is a weak signal.
Because the misreading of 1966 has had sixty years of policy consequences, in both directions.
In one direction it licensed the conclusion that spending on schools is pointless , because family background dominates. The closure evidence refutes this decisively : removing school for eight weeks removed eight weeks of learning, unequally.
In the other direction it has been used to argue that schools are the primary lever on inequality , and that a sufficiently good school can close a gap that everything in Part 8 shows to be produced elsewhere. The between-school variance findings do not support that either.
The position the evidence supports is specific and satisfies neither camp.
Schooling as such is powerful and compensatory. Gaps widen faster when school is absent.
Differences between schools are real, moderate, and much smaller than raw results imply — most of what looks like school quality is intake.
Teacher effects are real, moderate, persistent, and badly measured at the individual level.
Some interventions work at known and modest magnitudes , with tutoring currently the best-evidenced.
And none of this closes the gradient , because the gradient is produced by seven channels of which school is one (see 8.3.3) — and because advantage relocates within the system at every branch point.
Two questions to carry.
When a school is described as good or failing, ask whether the judgement is adjusted for intake. In most public reporting it is not.
And when an educational intervention claims a large effect, ask for the standard deviation figure and the sample size. Anything above half a standard deviation from a small study should be treated as a hypothesis (see 7.9.1).
A 1966 study of six hundred thousand students found that school resource differences accounted for much less of the variation in achievement than family background did , and was compressed into "schools don't matter". Three omissions matter. Its strongest school-level finding was about peer composition , not resources. Its cross-sectional design could not separate what a school added from what it received. And a variance decomposition understates school effects when schools are similar — which the study's own first finding established they were.
Four questions with four answers. How much variation is associated with which school — not much in most systems. Does attending school matter — enormously, and it is a different question. Do schools differ once intake is accounted for — yes, moderately. Do interventions work — some, by known amounts.
The seasonal comparison suggested schools are compensatory, gaps growing faster in summer than in term — and it is less secure than once presented , being sensitive to test scaling.
The 2020 closures gave the cleanest evidence available. In a well-resourced system with national tests either side of an eight-week closure, learning loss amounted to approximately the whole of the period removed , and was around 60 per cent larger for children from less-educated households.
Value-added removes the intake problem , and reveals that raw league tables are largely rankings of catchments. Teacher effects run at roughly a tenth to a fifth of a standard deviation per year and persist into college attendance and earnings — with a serious falsification-test criticism showing that value-added measures can predict prior achievement, indicating sorting the models do not capture. Defensible for research on averages; weak for evaluating individuals.
Between-school variance differs enormously by country , being high in early-sorting systems and low in comprehensive ones — which makes it partly a question about system design.
What works, with magnitudes : class size reduction at around a fifth of a standard deviation, larger for the disadvantaged and expensive; tutoring at around a third of a standard deviation, the strongest evidence base in the field ; teacher effectiveness; and early years provision. Effects above half a standard deviation from small studies should be treated as hypotheses.
Variance decomposition — how much of the variation between students is associated with variation between schools; not an effect size.
Peer composition effect — the influence of the other students' characteristics, distinct from resources.
Cross-sectional design — measurement at one point, unable to separate what a school adds from what it receives.
Seasonal comparison — using term-time and holiday learning rates to test whether schools compensate or amplify.
Learning loss — measured decline relative to expected progress, established most cleanly by the 2020 closures.
Value added — outcomes relative to what prior attainment and characteristics predict.
Falsification test for value added — checking whether the measure predicts prior achievement, which it should not.
Standard deviation effect size — the standard unit of educational effects; 0.2 is a serious effect and 0.5 from a small study is a hypothesis.
Fadeout — the decline of measured effects in the years after an intervention, sometimes with longer-run outcomes still affected.
One — reorder a league table. Find raw results for schools in one area and a measure of intake — free school meal eligibility, or a deprivation index. Plot one against the other. The correlation is the league table.
Two — check an adjustment. For any claim that a school is outstanding or failing, find out whether the judgement accounted for intake, and how.
Three — do the closure arithmetic. Take the finding that eight weeks of closure produced eight weeks of lost learning. Work out what that implies about the value of a school year , and compare it with what the 1966 report is usually said to have shown.
Four — size an intervention. Take any educational programme advertised in your country and find its effect size in standard deviations. If none is given, that is the finding.
Five — separate the four questions. Next time someone says schools do or don't make a difference, work out which of the four they mean. Usually they mean the first and are concluding about the second.
The between-school variance figures differ enormously between countries, and the reason is that systems are designed differently.
9.1.3 — Selection, Tracking and Choice covers what happens when children are sorted into different schools or streams, at what age, and by what criteria; the evidence on early selection and its effects on inequality and on average attainment; what school choice does where it has been tried; and why a system that sorts by ability sorts by something else as well.