Every education system sorts children. The interesting variables are when, on what, and how reversibly — and countries differ on all three in ways that produce measurably different outcomes.
This lesson is about what the sorting does. And it starts with the clearest available demonstration that a selection system does not sort on what it thinks it is sorting on.
Two children, four weeks apart, eighteen years of consequences.
In a system where the school year begins in September, a child born on 1 September and a child born on 31 August start school on the same day.
One of them is eleven months and thirty days older than the other. At age four, that is a quarter of a lifetime — and it shows in vocabulary, in physical coordination, in attention span, in the ability to sit still, and in almost everything a reception teacher observes.
Nothing about this is contentious. What is striking is how long it lasts and how far it reaches.
The younger child in the year is measurably less likely to reach the expected standard at each early assessment. The gap narrows with age and does not close for a long time.
They are substantially more likely to be identified as having a special educational need — the same child, assessed against the same criteria, being younger.
They are more likely to be placed in a lower set , and set placement determines which examination course they are entered for, which determines what they can study next (see 8.3.3 on branch points).
They are less likely to be picked for representative sports teams — an effect first documented in ice hockey, where the birth months of elite players cluster heavily in the months immediately after the age-group cutoff, and subsequently found across many sports and countries. The selected players get more coaching, which makes them better, which makes them more likely to be selected again.
And the effect is still detectable in university attendance — a difference of a percentage point or two in the probability of attending, between children born a month apart, across whole national cohorts.
Now the second half of the story.
In systems that selected children for different types of secondary school by examination at eleven, the month-of-birth pattern appeared in the results. Autumn-born children were disproportionately selected; summer-born children disproportionately were not.
The examination was designed to measure aptitude. It measured aptitude, plus eleven months of development, plus everything else correlated with the two — and it allocated children to different schools, different curricula and different destinations on the result.
This is the general principle of the lesson, in its cleanest form. A selection system sorts on its criterion and on everything correlated with its criterion — and the correlates include the month you were born, the language spoken at home, the vocabulary you arrived with, and how you behave in a room (see 9.1.1 on what is actually rewarded).
Sorting is not one thing, and the four kinds have different consequences.
Between schools : children are allocated to institutionally distinct schools with different curricula and different destinations. The key variable is the age at which this happens.
Within schools : children attend the same school and are placed in different streams or sets, sometimes for all subjects, sometimes subject by subject.
Within classes : grouping by attainment for particular activities.
And by choice : children are not allocated but families select, which produces sorting through a different mechanism and with different distributive consequences.
Two questions have to be asked separately of each.
What does it do to the average? Does the system as a whole produce more learning?
And what does it do to the distribution? Does it widen or narrow the gap between children from different backgrounds?
These can move in opposite directions , and the strongest evidence in this area addresses the second more cleanly than the first.
Early selection between schools
The comparative evidence, and the design that makes it credible.
Countries differ enormously in when they first sort children into different school types — from around age ten in some systems to sixteen or never in others. And the amount of variation in achievement that lies between schools rather than within them tracks that difference closely : it is high where selection is early, low where children stay together (see 9.1.2).
The obvious comparison — do early-tracking countries do worse? — is confounded by everything else that differs between countries (see 7.5.4 on Galton's problem).
The design that gets round it is elegant. International assessments test children both before the age of selection (in primary school) and after it. So each country can be compared to itself : the inequality of outcomes at primary, and the inequality of outcomes at secondary, with the change compared across countries that do and do not select in between. This is a difference-in-differences design (see 7.4.3), and it removes everything about a country that is constant.
The finding: early tracking increases the inequality of outcomes. Countries that sort earlier show a larger increase in the dispersion of achievement between primary and secondary than countries that sort later.
And the effect on the mean is ambiguous — estimates are small, and where they point anywhere they point slightly negative. So the trade-off that early selection is usually defended by — more inequality in exchange for higher average performance — is not supported. There is inequality, and there is little evidence of the compensating gain.
Two national reforms give the same answer with a stronger design.
Finland replaced a selective system with a comprehensive one between 1972 and 1977, rolling the reform out region by region over five years. That staggered rollout is a natural experiment (see 7.4.3): cohorts of the same age in different regions were educated under different systems for reasons of administrative scheduling.
Analysis of the rollout found a substantial reduction in the intergenerational income elasticity — the association between fathers' and sons' incomes fell for cohorts educated under the comprehensive system, by a meaningful proportion.
Sweden's comparable reform, also rolled out unevenly, produced a similar picture : earnings gains concentrated among children of less-educated parents.
Both are cases where a change in system design measurably changed the transmission of advantage — which is 8.8.1's argument, in the institution 8.3.3 identified as the central channel.
And the case of remaining selective systems within a country.
Where selective schools survive alongside comprehensive ones in the same country, a cleaner comparison is available — and the finding is consistent.
Children who attend selective schools do somewhat better than comparable children who do not. Children who attend non-selective schools in areas that have selection do somewhat worse than comparable children in fully comprehensive areas — because the selective schools have removed the higher-attaining peers (see 9.1.2 on composition).
The two effects are broadly offsetting, so the average across a selective area is close to that of a comprehensive one.
And the distributional effect is not offsetting at all. The attainment gap by family income is substantially larger in selective areas, and the proportion of selective school places going to children from the poorest backgrounds is very low — far below their share of the population, and far below their share of the highest-attaining children at the point of selection.
The last figure is the one to hold , because it separates two explanations. If selective schools simply took the highest attainers, their intake would reflect the composition of the highest-attaining group. It does not — which means that at a given level of prior attainment, children from advantaged families are considerably more likely to be selected. The mechanisms are 8.3.3's: preparation, tutoring, information, application, and the choice of whether to enter at all.
Within-school sorting
Setting and streaming: smaller effects, redistributive in the same direction, and considerable misallocation.
The average effect of ability grouping on overall attainment is small — meta-analyses generally find something close to zero for the system as a whole.
The distributional effect is not zero. The consistent pattern is that higher groups gain modestly and lower groups lose , so that grouping widens the spread without much changing the mean.
Three mechanisms are documented. Curriculum differentiation: lower sets are taught less content, more slowly, with more repetition — so the gap that justified the grouping is subsequently produced by it. Teacher allocation: more experienced teachers are disproportionately assigned to higher sets. And expectations, which is discussed below.
And there is a great deal of misallocation. Studies matching set placement against prior attainment consistently find a substantial minority of pupils placed in a set inconsistent with their measured attainment — and the misallocation is patterned: by month of birth, as the story showed, and by family background and behaviour.
Which is the lesson's principle again. A grouping system that is intended to sort on attainment sorts on attainment plus everything correlated with attainment plus everything correlated with the judgement of the person doing the sorting.
Teacher expectations: the famous study, and what the evidence actually supports.
The best-known experiment in this area told teachers that certain randomly selected children were expected to bloom intellectually during the coming year , and reported that those children subsequently gained more on an intelligence test than their classmates. The result — that a teacher's expectation can produce the reality it expects — became one of the most cited findings in education.
It has serious methodological problems , and this course reports them. The test used was of questionable properties at the youngest ages where the effects were largest; the effects were concentrated in a small number of classes; and replication attempts have produced inconsistent results.
What the more careful subsequent literature supports is more modest and still worth knowing.
Expectancy effects are real and small. Teachers' expectations do influence outcomes, by amounts that are detectable and not transformative, and they do not generally accumulate into the runaway self-fulfilling spiral the popular version describes.
And they are systematically patterned. The more consequential finding is not that expectations create ability but that they are formed partly on characteristics unrelated to attainment — and this shows up cleanly in comparisons between teacher-assessed and externally, anonymously marked work, where the pattern of gaps between groups differs.
That comparison is the useful design , because it holds the child constant and varies whether the assessor knows who they are — which is an audit study inside a school (see 7.4.2).
Choice
What happens when families select rather than being allocated.
The theoretical case is competition : schools that must attract families will improve, and families will move away from poor provision.
The evidence, across several national experiments, is more mixed than either advocates or opponents present it.
A large national voucher system in Chile, introduced in 1981, produced no detectable improvement in average outcomes. What it produced was sorting : the more advantaged families moved to the private-voucher sector, leaving the public sector with a more disadvantaged intake — so measured differences between sectors grew while the overall level did not.
Sweden's introduction of independently run, publicly funded schools produced small positive effects on longer-run outcomes in some analyses, alongside increased between-school segregation and evidence of grade inflation in a system where schools set the grades that determine progression.
And the American charter school evidence is the most instructive, because its central finding is about variance. Averaged across all charter schools, effects are close to zero. But the average conceals an enormous spread : some urban charter schools, evaluated using admission lotteries — which is a genuine randomised design (see 7.4.2) — show substantial gains for disadvantaged students , while others do worse than the schools their students left.
So the honest statement is: "does school choice work" is the wrong question , because the sector contains schools doing very different things. What the lottery studies establish is that certain specific practices, in certain settings, produce large gains for disadvantaged children — which is a finding about those practices rather than about choice as a mechanism.
And three things about choice as a mechanism are consistent across systems.
Choice is exercised unequally. Understanding the system, evaluating options, arranging transport, completing applications on time and — decisively — moving house are all unequally distributed (see 8.3.3 on capitalisation).
Schools choose too. Wherever schools are oversubscribed and have any discretion, selection operates in both directions; and where funding follows pupils, there are incentives to attract those cheapest to educate.
And the result is that choice systems generally increase sorting between schools , which — given the peer composition evidence from 9.1.2 — has consequences independent of anything about school quality.
Because the design of a sorting system is one of the few levers that has demonstrably moved the transmission of advantage.
Most of Part 8 identified mechanisms that are hard to change. The comprehensive reforms are among the clearest cases where a deliberate institutional change measurably reduced intergenerational persistence — which is why they matter out of proportion to their apparent technicality.
Four questions for any sorting arrangement.
At what age does the first irreversible sort happen? Earlier means more inequality, with no established compensating gain in the mean.
How reversible is it? A system where movement between tracks is common has different consequences from one where the allocation at eleven determines the destination.
What is it actually sorting on? The criterion, plus the correlates — month of birth, vocabulary, behaviour, preparation, and whether a family knew the system existed.
And who is doing the choosing — the family, the school, or both?
One closing observation, which is the argument of this whole Topic in a sentence. The month-of-birth effect exists because an administrative convenience — a single annual intake — became a criterion of selection without anybody deciding it should be.
Nobody believes children born in September are more able. The system does not believe it either. It simply sorts, and what it sorts on includes things nobody chose — which is exactly what happened in the Tuesday of 9.1.1, and it is why an institution's effects have to be studied rather than read off its purposes.
A child born just before the cutoff and one born just after start school on the same day eleven months apart. The younger is less likely to reach early standards, more likely to be identified with a special educational need , more likely to be placed in a lower set, less likely to be selected for teams — an effect first documented in ice hockey and found across sports — and still measurably less likely to attend university. Where selection at eleven existed, the month-of-birth pattern appeared in the results.
Four kinds of sorting — between schools, within schools, within classes, and by choice — each requiring two separate questions : what it does to the mean, and what it does to the distribution.
On early between-school selection, the strongest design compares countries to themselves before and after the age of selection, using international tests at primary and secondary. Early tracking increases the inequality of outcomes, and the effect on the mean is ambiguous or slightly negative — so the claimed trade-off is not supported. And two national comprehensive reforms rolled out region by region provide natural experiments: both reduced intergenerational persistence, with gains concentrated among children of less-educated parents.
Where selective schools survive within a country , selected children do somewhat better, unselected children in selective areas do somewhat worse, the means roughly offset — and the attainment gap by income is substantially larger. The decisive figure is that the share of selective places going to the poorest children is far below their share even of the highest-attaining children, which means selection at a given level of attainment favours the advantaged.
Within-school setting has near-zero average effects and redistributive ones — higher sets gain modestly, lower sets lose — through curriculum differentiation, teacher allocation and expectations. And misallocation is substantial and patterned , by month of birth, background and behaviour.
The famous teacher-expectancy experiment has serious methodological problems and inconsistent replication. What survives: expectancy effects are real, small, and systematically patterned — visible most cleanly in comparisons of teacher-assessed against anonymously marked work.
On choice : a national voucher system produced sorting rather than gains; an independent-school reform produced small long-run effects alongside increased segregation and grade inflation; and the charter evidence's central finding is variance — near-zero on average, with lottery-based studies showing substantial gains for specific practices in specific settings. Choice is exercised unequally, schools choose too, and choice systems increase sorting.
Relative age effect — advantages accruing to children older within their school-year cohort, persisting into later outcomes.
Tracking / streaming / setting — allocation to different school types; to different classes across subjects; to different groups subject by subject.
Age of first selection — the key comparative variable, strongly related to between-school variance.
Difference-in-differences across school stages — comparing inequality before and after the age of selection, within countries.
Comprehensivisation — replacing selective allocation with common schools; rolled out unevenly in two countries, providing natural experiments.
Peer composition — the effect of who else is in the school, which selective systems redistribute.
Misallocation — placement inconsistent with measured attainment, patterned by birth month, background and behaviour.
Expectancy effect — the influence of a teacher's expectation on outcomes; real, small, and unevenly formed.
Blind versus teacher assessment — the design isolating assessor knowledge of the pupil.
Cream-skimming — selection by schools of the pupils cheapest or easiest to educate.
Voucher and lottery evidence — publicly funded choice systems, and the randomised admission designs that evaluate oversubscribed schools.
One — check your own month. Find out when your school system's cutoff falls and where your birthday sits. Then look up the relative age evidence for your country.
Two — find the age of first selection. For three countries, establish at what age children are first placed in institutionally different schools. Then look up their between-school variance in an international assessment.
Three — audit a set. For any school you know that groups by attainment, ask how placement is decided, how often it is reviewed, and what proportion of pupils move sets in a year.
Four — check the composition claim. For a selective system, find the proportion of selected places going to children eligible for free school meals, and the proportion of high-attaining eleven-year-olds who are. The gap between those two figures is the finding.
Five — apply the correlate principle. Take any selection process you know — for a job, a course, a team. List the criterion, then list five things correlated with it that the process is also sorting on.
The final lesson of this Topic follows the sorting to its destination: the labour market that the credentials are for.
9.1.4 — Credentialism and the Graduate Premium covers what a degree is actually worth and how much that varies, why the premium has held up despite enormous expansion, what happens to the marginal graduate, the evidence on over-education and skills mismatch, and whether the returns are to skill, to signal, or to the closure of an occupation.