These errors survive everything in the last three lessons. The sample can be perfect, the design clean, the confounders handled and the p-value irrelevant — and the conclusion can still be about a different level of reality from the evidence.
A relationship between countries need not hold between people. A relationship between people need not hold within a person over time. And a relationship that holds in every single subgroup can reverse when the subgroups are added together.
That last sentence is not a figure of speech. It happens, it is not rare, and it is the most disorienting fact in this Part.
More immigrants, more literacy — and immigrants who could not read.
In 1950 the sociologist William Robinson published a short paper that changed how the discipline reads aggregate data.
He took the 1930 United States census and computed, across the 48 states, the correlation between the proportion of the population who were foreign-born and the proportion who were literate.
The correlation was +0.53. States with more immigrants had markedly higher literacy. On the face of it, immigration and literacy went together strongly.
Then he computed the same relationship at the level of individuals — using the same census, asking whether a given person being foreign-born was associated with that person being literate.
The correlation was −0.11. Foreign-born individuals were, on average, less likely to be literate than the native-born.
A strong positive relationship between places. A weak negative relationship between people. From the same data.
And the explanation is entirely mundane. Immigrants settled disproportionately in the industrial, urban states of the north-east and the upper midwest — states which, for reasons having nothing to do with immigration, had higher literacy among their native-born populations than the rural south did. The states were doing the correlating, not the people.
Robinson's point was not that the aggregate figure was wrong. It is a correct fact about states. His point was that it licenses no inference whatever about individuals — and that a great deal of social research was quietly making exactly that inference.
One more, because the stakes can be much higher.
For decades it was a commonplace that unemployment in Weimar Germany drove voters to the Nazis — an inference from the fact that regions with high unemployment tended to have high Nazi vote shares. Detailed analysis of the electoral evidence has substantially overturned this. The unemployed disproportionately supported the Communists; Nazi gains came heavily from Protestant middle-class, small-town and rural constituencies, including previous non-voters. The regional correlation was real. The individual-level story read off it was wrong , and it shaped popular historical understanding for two generations.
Every finding lives at a level, and there are four to keep apart.
Between areas — states, districts, countries. What varies here is the place.
Between individuals — people, at a moment. What varies here is the person.
Within an individual over time — the same person, different days or years. What varies here is the occasion.
And between groups defined by something else entirely — departments, schools, firms — where composition can do all the work.
No relationship at any of these levels implies a relationship at any other , and the failures have names because they are so common.
The ecological fallacy infers about individuals from aggregates. The individualistic (or atomistic) fallacy infers about aggregates from individuals. Simpson's paradox is what happens when combining groups reverses a relationship. And within-person and between-person relationships can differ completely without anybody having made an error at all.
The two directions
Downwards: the ecological fallacy.
Aggregate data is abundant, cheap and public (see 7.4.4), and area-level relationships are easy to compute. They are relationships between areas.
The mechanism that produces the trap is always the same: composition. An area's average reflects who lives there, and people do not distribute themselves randomly across areas — which is 7.3.2's selection problem operating geographically.
Examples with real consequences. Districts with more migrants have higher crime rates — used constantly, and consistent with migrants having lower individual offending rates if migrants settle in places that were already high-crime for reasons of poverty and housing. Countries with more chocolate consumption per head have more Nobel laureates — a genuinely published correlation, and an entirely spurious one produced by national wealth. Regions with more Protestants had higher suicide rates — the foundation of one of the discipline's founding studies, and the individual-level inference was never directly established (see 4.2.4). Part of the difference has since been attributed to how deaths were classified in Catholic and Protestant jurisdictions , which makes it 7.4.4's recording problem as well.
Statistical methods for ecological inference exist — the older approach of regressing area outcomes on area compositions, and more elaborate modern methods that add bounds and assumptions. They can help and they are not a solution. All require assumptions about how the unobserved individual-level relationship behaves across areas, and where those assumptions fail the estimates fail with them. The honest position is that ecological inference is hazardous and should be presented as an estimate with strong assumptions, not as a finding.
Upwards: the individualistic fallacy, and Rose's argument.
The reverse error is less discussed and more interesting for sociology.
Knowing why some individuals in a population have an outcome does not tell you why that population has the rate it does.
The epidemiologist Geoffrey Rose stated this most clearly. Ask why some individuals in Britain have high blood pressure, and the answer involves genetics, weight, alcohol and salt intake relative to their neighbours. Ask why Britain as a whole has far higher average blood pressure than a population living traditionally in rural Kenya, and none of those individual-level answers will do the work — because within each population the individual variation is the same kind of variation, while the whole distributions sit in different places.
Rose's conclusion: the causes of cases within a population and the causes of differences between populations are different questions with different answers. Individual risk factors explain who, within a given society, gets the outcome. Population-level causes explain why that society's whole distribution sits where it does.
And this is Durkheim's argument, arriving from epidemiology a century later (see 4.2.1). Why this person killed themselves and why this society has that suicide rate are not the same question, and answering the first exhaustively would leave the second untouched.
The practical version, which matters for policy : an individual-level relationship does not tell you the effect of changing the whole distribution. If more education raises an individual's earnings partly by improving their position relative to others, educating everyone does not raise everyone's earnings by that amount — this is 7.4.2's general equilibrium problem, and it is the individualistic fallacy in its most expensive form.
Simpson's paradox
The relationship holds in every subgroup and reverses in the total.
The classic case is the University of California, Berkeley graduate admissions of 1973.
In aggregate, the figures looked stark. Roughly 44 per cent of male applicants were admitted, against roughly 35 per cent of women — a large gap across thousands of applications.
Department by department, the picture changed. In most departments women were admitted at rates equal to or slightly higher than men. The aggregate gap came from where people applied : women applied disproportionately to departments with low admission rates for everyone, and men to departments with high ones.
The arithmetic that makes this possible is simple and it feels impossible. Combining groups with different base rates and different group sizes can reverse a relationship, and no amount of staring at the totals reveals it.
And now the part that is usually left out, which is what makes this a sociology lesson rather than a statistics puzzle. The analysis did not show there was no problem. It relocated the question. Why were the departments women applied to the ones with the least funding, the most applicants per place and the lowest admission rates? Why had subject choice been so strongly patterned by sex long before anyone applied? Those are excellent questions, and they are the ones the disaggregation makes visible.
The general lesson : when an aggregate difference disappears on disaggregation, the correct response is not "so there was no effect" but "so the effect operates through the variable I just controlled for" — which is 7.2.2's mediator problem in another guise, and 7.6.2's warning about controlling away your own finding.
Two spatial versions, and one is a political weapon.
The modifiable areal unit problem. Results computed on areas depend on how the areas were drawn, in two distinct ways. The scale effect : the same data aggregated into large units or small units gives different correlations, usually stronger at larger scales because averaging removes variation. The zoning effect : at the same scale, different boundary configurations give different results.
This is not a small technical wrinkle. Neighbourhood effects, segregation indices, area deprivation measures and school catchment analyses are all sensitive to it, and the boundaries were drawn by administrators for administrative reasons (see 7.4.4).
And gerrymandering is the deliberate exploitation of the zoning effect. The same votes, differently partitioned, produce different representation — which is the modifiable areal unit problem used as an instrument of power rather than encountered as a nuisance.
The practical response : where possible, report results at more than one scale, and treat a finding that appears only at one particular aggregation with suspicion.
Between people and within people: the level nobody checks.
Across people , those who exercise more are healthier. Within a person , on days when you are unwell you exercise less — so the within-person relationship over time may run the other way, or be dominated by reverse causation.
Across people , those who work longer hours earn more. Within a person , the relationship between this month's hours and this month's pay may be entirely different.
Across countries , richer ones report higher happiness. Within a country over time , the relationship between growth and reported happiness has been argued about for fifty years precisely because it is a different question.
A between-person relationship does not license a within-person claim , and almost all advice given to individuals on the basis of cross-sectional research makes exactly that leap.
The methodological response is longitudinal data with within-person comparison — fixed effects designs, which compare each person to themselves and thereby remove everything stable about them (see 7.4.3). What they buy is protection against stable confounders; what they cost is that they can only speak about things that change.
Three more fallacies that belong in the same family.
The regression fallacy. Mistaking regression to the mean for an effect (see 7.4.2). Any intervention applied to the worst-performing cases will appear to work , because extremes move towards the average on their own. The version worth remembering because it inverts a common belief: praise appears to be followed by worse performance and criticism by better, purely because exceptional performances are followed by more ordinary ones.
The base-rate fallacy, and its legal form. Confusing the probability of the evidence given innocence with the probability of innocence given the evidence. These are different quantities and the difference depends on how common the thing is to begin with.
The most consequential instance is the case of Sally Clark , convicted in England in 1999 of murdering her two infant sons. An expert witness testified that the chance of two cot deaths in such a family was about one in 73 million — a figure obtained by squaring a single-death probability, which assumes the two deaths are independent when there are strong reasons, genetic and environmental, to think they are not. The Royal Statistical Society issued a public statement criticising the reasoning , noting both the invalid independence assumption and, more fundamentally, that the relevant question is not how unlikely two cot deaths are but how that compares with the likelihood of two murders in the same family — which is also very rare. Her conviction was overturned on appeal in 2003, principally on the ground of undisclosed medical evidence, and the statistical testimony was central to the surrounding controversy. She never recovered and died in 2007.
The Texas sharpshooter. Firing at a barn and then drawing the target around the tightest cluster of holes. Any dataset contains clusters , and finding one and then testing it on the same data is not a test (see 7.6.3 on the forking paths).
Because these errors are invisible from inside, and the corrective is a single question asked before anything else.
"What is a case here?"
If the cases are countries, the finding is about countries. If they are people, it is about people. If they are person-years, it is about person-years — which is a different thing again, and is what most panel analysis actually studies.
Four follow-up questions.
Am I about to say something about individuals using data about places? The ecological fallacy, and it is the commonest misuse of official statistics anywhere.
Am I about to say what would happen if everyone changed, using data about how people differ? The individualistic fallacy, and it is how policies are oversold.
Would this relationship survive disaggregation — and if it reverses, what does the grouping variable represent? Not an answer, a relocation of the question.
And is this a claim about people, or about the same person over time?
A final word about what this Topic has been for. Numbers do not speak. They are answers to questions that somebody framed, at a level that somebody chose, about units that somebody defined (see 7.1.3, 7.2.2). Everything in 7.6 has been a way of recovering the question from the answer.
Do that, and the numbers become useful. Skip it, and they become authoritative — which is worse than useless, because authority does not invite checking.
Every finding lives at a level: between areas, between individuals, within an individual over time, or between groups where composition does the work. No relationship at one level implies a relationship at another.
Robinson's demonstration : across the 48 US states in 1930, the proportion foreign-born correlated +0.53 with literacy; across individuals, being foreign-born correlated −0.11 with being literate. The states were doing the correlating. The Weimar case shows the same error with historical consequences — regional unemployment and Nazi voting correlated, and the unemployed disproportionately voted Communist.
The ecological fallacy infers downwards from aggregates to individuals, and composition is always the mechanism. Ecological inference methods exist and rest on strong assumptions ; they produce estimates, not findings.
The individualistic fallacy infers upwards. Rose's argument: the causes of cases within a population and the causes of differences between populations are different questions. This is Durkheim's argument arriving from epidemiology — and its expensive policy form is assuming that an individual-level relationship tells you the effect of changing everyone.
Simpson's paradox : a relationship in every subgroup can reverse in the total, as in the Berkeley admissions case, where a large aggregate gap came from which departments applicants applied to. When a difference disappears on disaggregation, the effect operates through the variable you disaggregated by — which relocates the question rather than dissolving it.
The modifiable areal unit problem : results depend on the scale and the drawing of boundaries — and gerrymandering is that effect used deliberately.
Between-person and within-person relationships can differ entirely , and almost all individual advice derived from cross-sectional research makes the leap without noticing.
And three relatives : the regression fallacy (extremes move to the average, so anything applied to the worst appears to work), the base-rate fallacy in its legal form — the Sally Clark case, where an invalid independence assumption and a missing comparison to the rarity of double murder contributed to a wrongful conviction — and the Texas sharpshooter, drawing the target after the shots.
One question first, every time: what is a case here?
Ecological fallacy — inferring individual relationships from aggregate data.
Individualistic (atomistic) fallacy — inferring population-level relationships from individual data.
Rose's distinction — the causes of cases within a population differ from the causes of differences between populations.
Simpson's paradox — a relationship reversing when subgroups are combined.
Composition effect — an aggregate difference produced by who is in each group rather than by anything happening to them.
Modifiable areal unit problem — the dependence of area-level results on the scale and configuration of boundaries. Scale effect / zoning effect.
Ecological inference — statistical estimation of individual relationships from aggregate data; assumption-heavy.
Between-person / within-person — variation across people at a moment; variation within one person over time.
Fixed effects — comparing each unit to itself, removing stable characteristics.
Regression fallacy — mistaking regression to the mean for an effect.
Base-rate fallacy / prosecutor's fallacy — confusing the probability of the evidence given a hypothesis with the probability of the hypothesis given the evidence.
Texas sharpshooter fallacy — identifying a cluster in data and then "testing" it on the same data.
One — name the case. Take five quantitative claims you meet this week and write, for each, what one row of the underlying data would be: a person, a place, a year, a person-year, an event.
Two — do the Robinson check. Find any claim about individuals supported by an area-level statistic — crime and migration, deprivation and health, education and voting. Ask what composition could produce it with the opposite individual relationship.
Three — disaggregate something. Take an aggregate difference between two groups and think of one variable that differs between them. Ask whether the difference could vanish or reverse within levels of that variable.
Four — test the level of your own advice. Take a piece of health, career or lifestyle advice you believe. Ask whether the evidence for it compares different people or the same person over time.
Five — find a base-rate error. Take any claim of the form "this result would be very unlikely if they were innocent / if the drug did not work / if there were no effect". Ask what the alternative's probability is. The comparison, not the small number, is the whole question.
Topic 7.6 is finished. Topic 7.7 does the same job for qualitative material: what it means to analyse hundreds of pages of transcript and fieldnotes systematically, rather than reading them until something occurs to you.
7.7.1 — Coding, Grounded Theory and Saturation covers what coding is and is not, the difference between a code and a theme, how grounded theory actually works as against how it is cited, what saturation means and why the concept is under attack, and how to tell analysis from illustration.