The last lesson explained why a well-drawn sample can speak for millions.
This one is about all the ways a sample that looks well-drawn produces an answer that is confidently, systematically, unrecoverably wrong — and about the single idea that unifies them , which is worth more than the whole list.
The idea is this: the people who are not in your data are not a random selection of the people who could have been. They are missing for reasons , and those reasons are very often the same reasons that determine the answer.
The armour on the bombers.
During the Second World War, American bombers were coming back from raids over Europe full of holes, and the question was where to put more armour. Armour is heavy; a plane cannot be armoured everywhere; the weight has to be spent where it will save the most aircraft.
So the analysis began sensibly. Examine the returning aircraft, map every bullet hole, and find where the damage concentrates. The maps came back consistent: heavy damage on the wings, the fuselage, the tail gunner's position. Very little on the engines. Very little on the cockpit.
The obvious conclusion: armour the wings and fuselage, where the planes are being hit.
Abraham Wald , a statistician working with the Statistical Research Group at Columbia, pointed out that this reasoning was inverted.
The data consisted entirely of planes that came back.
Bullets do not preferentially avoid engines. The engines were being hit at about the rate everything else was. The planes hit in the engines were not in the sample, because they were in the sea.
So the map of damage on survivors is a map of where a plane can be hit and still survive — and the places with no damage in the data are precisely the places where damage is fatal. Put the armour where the holes aren't.
A note on the story, in keeping with this course's rules. Wald's memoranda from 1943 are real and are exactly about this inference — estimating a plane's vulnerability from the damage on survivors. The vivid version usually told, with a room of generals and a single dramatic correction, is a later popularisation. The logic is Wald's; the anecdote has been polished by retelling.
And the logic is everywhere. Every time you look at a set of outcomes and reason about causes, ask what was removed from the set before you saw it.
Every study has a hidden population: the people who could have been in it and are not.
They were never on the list. They were on the list and could not be contacted. They were contacted and refused. They started and dropped out. They answered the easy questions and skipped the sensitive one. They died, moved, closed down, or were removed from the register.
If those departures were random, they would cost you precision and nothing else. A sample of 400 instead of 1,000 is wider but not wrong.
They are almost never random. People refuse for reasons. People drop out of programmes for reasons. Firms go out of the register for reasons. And in the cases that matter most, the reason is related to the very thing you are measuring.
Which gives the one equation you should carry out of this lesson , not as algebra but as a habit of mind:
Bias ≈ the proportion missing × how different the missing are.
Both terms. A 90 per cent non-response rate does no damage if the missing are just like the present. A 5 per cent non-response rate is fatal if the missing 5 per cent are the entire phenomenon.
And this is why "what was the response rate?" is only half a question. The full question is: is there a reason the missing would answer differently?
The five errors, of which the famous one is the smallest
Total survey error — the framework that puts the margin of error in its place.
Coverage error — people not on the frame at all (see 7.3.1). Invisible from inside the data.
Sampling error — the luck of which frame members were drawn. This is the only one the ± figure covers , and in a well-run survey it is frequently the smallest of the five.
Non-response error — selected people who did not take part.
Measurement error — wrong answers: misunderstanding, misremembering, lying, question wording, interviewer effects, mode effects (see 7.2.2).
Processing error — coding, editing, weighting and data-entry mistakes. Unglamorous and not rare.
Hold on to the proportions. A survey report says "±3 points". The coverage gap might be worth two points, the non-response bias four, the wording effect five. The number in the report is the one error the researchers could calculate, not the one most likely to have misled them — which is a general feature of quantified uncertainty and worth remembering well beyond surveys.
Non-response, and the puzzle of why modern surveys still work
Response rates have collapsed, and the sky has not fallen. Understanding why is the point.
Unit non-response — the person never takes part at all. Item non-response — they take part and skip questions, disproportionately on income, sensitive behaviour, and anything they think reflects badly on them.
The scale of the collapse is dramatic. Pew Research Center's telephone surveys had response rates around 36 per cent in the late 1990s; by 2018 the comparable figure was around 6 per cent, with similar declines across the industry. Ninety-four refusals for every participant.
And yet — this is the genuinely interesting part — their methodological studies repeatedly found that estimates from these very low-response surveys remained close to benchmarks from high-quality government surveys on a wide range of measures.
Why? Because of the equation. For most attitudes and demographics, people who answer surveys are not systematically different in the relevant way, once weighted. The rate is high, the difference is small, so the bias is small.
But not for everything, and the exceptions are patterned. The same research found that survey respondents consistently overstate civic and political engagement — volunteering, contacting officials, group membership, trust in neighbours. The reason is exactly the one you would guess: people who agree to spend twenty minutes helping a stranger with a survey are, on average, more the sort of person who volunteers.
Which is the general rule. Non-response bias is not a global property of a survey. It is specific to each measure , and it is largest for measures correlated with the disposition to participate. So the question to ask of any given finding is: is this the kind of thing that co-operative people do more of?
Self-selection: when there was no sample at all, only volunteers.
The strongest form of non-response is when nobody was selected in the first place and everyone chose themselves.
Open online polls. "Click here to have your say." A measure of who was motivated and organised enough to click.
Customer reviews. Ratings distributions are famously J-shaped — piles of five stars, a smaller pile of one star, almost nothing in between. This is not what satisfaction looks like; it is what the motivation to write looks like. People who are delighted or furious write; the mildly satisfied majority does not. The average rating is not average satisfaction.
Phone-ins, petitions, consultations, complaint data, and "we asked our members". Each measures intensity of feeling combined with organisational capacity, not prevalence of opinion.
And the courteous killer: "the survey was sent to all staff and 340 responded." Sent to all is not a sample; who answered is.
The tell in every case : the participants chose, rather than being chosen. When you see that, no size rescues it — this is precisely the Literary Digest mechanism from 7.3.1, and Meng's result explains why more of it makes matters worse rather than better.
Survivorship, attrition and selection: the same error in three costumes
Survivorship bias — the sample was filtered by the outcome before you saw it.
Investment funds. Performance tables of "funds available today" exclude the funds that closed, which closed because they did badly. The surviving average overstates what an investor would have experienced.
Successful people's habits. Study a hundred founders who succeeded and find they were persistent, risk-tolerant and dropped out of university. The thousands who did exactly the same and failed are not in the book , because nobody writes books about them. The trait may be as common among failures — in which case it explains nothing at all.
"They built things better in the old days." The buildings from 1780 that you can visit are the ones that lasted. The badly built ones are gone. You are looking at a survivor sample of two centuries of selection.
Company practices. Studies of "what excellent companies do" sample on the outcome. Some of the most celebrated examples subsequently collapsed, which is what you would expect if the practices were not the cause.
Medicine and crime and social programmes alike : case series of people who reached the clinic, the court or the completion certificate are filtered populations, and the filter is usually related to the outcome.
The single diagnostic question : were the cases selected on the outcome? If the answer is yes, you can describe them but you cannot explain them, because the explanation requires the ones that dropped out.
Attrition — survivorship happening inside your own study.
Panel and longitudinal studies follow the same people over years, which is the only way to observe change properly (see 7.4.4) — and they lose people continuously.
The losses are patterned. People who move house, who have unstable housing, who are in poor health, who are in prison, who are poorer, who are younger, and who are recent migrants all drop out at higher rates. Long-running panels often retain well under half of the original sample after two decades.
And the effect on findings is systematic rather than random. A study of employment trajectories loses disproportionately those with the most chaotic trajectories. So it will report that careers are more stable than they are — not because anyone lied, but because instability and dropping out have a common cause.
Attrition is also how promising programme evaluations become less promising. If the participants who were doing badly stop attending and stop being measured, the remaining group improves without any individual improving. The correct analysis counts everyone from the moment of assignment, whatever happened afterwards — the principle known in trials as intention to treat , and it exists precisely for this.
Selection into treatment — the version that wrecks causal claims.
People who take part in things are different from people who do not, in ways that predict the outcome.
A job-training scheme reports that graduates earn more than non-participants. Who enrols? People who sought it out, were assessed as suitable, could attend, had childcare, had transport. Every one of those is a reason they would have done better anyway.
Private schools, gym memberships, therapy, mentoring, university courses, migration, marriage — each is entered into by people who differ systematically from those who do not. Compare participants to non-participants and you have measured who joins plus what joining does , with no way to separate them.
This is the central problem of causal inference in the social world , and it is what the whole of Topic 7.4 is about — because randomisation exists to destroy exactly this, and natural experiments exist for when you cannot randomise.
Missing data: the three regimes
Rubin's classification, in plain words, because the labels are unhelpfully abstract.
Missing Completely At Random (MCAR). Missingness has nothing to do with anything — a lab freezer failed, a batch of forms was lost in the post. Then dropping the incomplete cases costs precision only. It is the assumption most software makes by default, and it is rarely true.
Missing At Random (MAR) — a badly named category. It means missingness depends on things you did observe. Younger respondents skip the income question more often, and you know everyone's age. Then the missingness is predictable from your data and can be handled , most respectably with multiple imputation — generating several plausible completed datasets from the observed relationships and combining the results, so the extra uncertainty is carried through rather than hidden.
Missing Not At Random (MNAR). Missingness depends on the missing value itself. The highest earners decline to state their income because it is high. People with the most stigmatised experience skip the question about it. Those who dropped out of the programme are unreachable because it went badly.
MNAR cannot be fixed by any statistical technique , because the information needed is precisely what is absent. It can only be addressed by design — a follow-up push on non-responders, administrative linkage, a sensitivity analysis showing how bad the assumption would have to be to overturn the finding — or by honesty.
And the crucial practical point: you can rarely tell which regime you are in from the data , because distinguishing them requires knowing what the missing values are. It is a judgement about the world, argued for, not a test result — which is 7.1.3's lesson arriving in the analysis stage.
Weighting: what it can and cannot repair.
The standard remedy is to weight the sample so it matches known population totals — by age, sex, region, education, and often more. Raking adjusts iteratively across several such margins.
What weighting does : corrects imbalance on the variables you weight on.
What weighting assumes : that within each weighting cell, the people who responded are like the people who did not. Weight up young men because too few answered, and you are assuming the young men who answered represent the young men who did not. If the ones who answered are the joiners, the volunteers and the politically engaged, you have multiplied the bias rather than removed it.
So weighting fixes composition and not disposition — and disposition is what non-response selects on.
The strongest modern version is worth knowing about, because it is genuinely impressive and genuinely limited. Multilevel regression and post-stratification (MRP) models the outcome within many small demographic cells, borrowing strength across them, then re-weights to census totals. A well-known demonstration used an Xbox gaming panel — respondents overwhelmingly young and male, about as unrepresentative as a sample can be — and after MRP adjustment produced election estimates that tracked the outcome well.
Read that correctly in both directions. It shows that a badly skewed sample can be rescued when the skew is on things you can model and adjust for. It does not show that representativeness stopped mattering — it depends on the model being right, on having good population totals, and on the same assumption as all weighting: that within a cell, respondents resemble non-respondents. When the missingness is MNAR, no amount of modelling reaches it.
Four smaller lies, which are common enough to name.
Recall error and telescoping. People misremember, and they systematically pull memorable events forward in time — reporting something from eighteen months ago as having happened within the last year. This inflates reported rates in any "in the past twelve months" question , and it inflates them more for vivid events, which is precisely the category most surveys ask about.
Proxy reporting. One household member answers for everyone. Their account of others' earnings, hours, health and behaviour is systematically less accurate and biased in predictable directions.
Panel conditioning. Being surveyed repeatedly changes people. They learn that saying yes triggers a long module and start saying no. They become more attentive to the topic. The panel becomes progressively less like the population it was drawn from — a Hawthorne effect distributed over years (see 7.2.2).
The wrong denominator. A rate is a fraction, and the bottom of it is chosen. Crime "per resident" in a city centre with a daytime population five times its resident population; accident rates per vehicle versus per mile travelled; mortality per case where cases are only those tested. Two studies can report opposite trends from identical numerators.
Because this is the failure mode that no amount of statistical sophistication corrects, and it is the one most often ignored in the reporting.
Four questions, in order, for anything you read.
One — who could not be here? Not on the frame, unreachable, institutionalised, unregistered.
Two — who chose not to be here, and would they have answered differently? Not the rate: the reason.
Three — was anything filtered by the outcome before I saw it? Survivors, completers, the still-open, the still-traceable.
Four — if people left partway through, who left?
And one for the researcher. If you cannot answer these, the honest report says so. A study that states its likely direction of bias is more useful than one that reports a margin of error and stops — because a reader who knows which way an estimate is probably wrong can still use it, and a reader given false precision cannot.
Wald's principle, in its most portable form : look at what is not in the data, and ask why it isn't. It is the most consistently valuable analytic habit in this entire Part , and it costs nothing.
The missing are not missing at random. Wald's bombers: the damage map of survivors shows where a plane can be hit and survive , so the undamaged areas are the fatal ones. Every filtered sample has this structure.
Bias ≈ how many are missing × how different they are. Both terms matter, which is why a 6 per cent response rate can be nearly unbiased for some measures and badly biased for others. Pew's finding : as telephone response rates fell from ~36 per cent to ~6 per cent, most estimates stayed close to benchmarks — except measures related to the disposition to participate , where civic engagement is consistently overstated.
Total survey error has five components — coverage, sampling, non-response, measurement, processing — and the margin of error covers only the second , often the smallest.
Self-selection is the extreme case : online polls, reviews with their J-shaped distributions, phone-ins, "sent to all staff". Participants chose rather than being chosen, and size makes it worse.
Survivorship, attrition and selection into treatment are one error in three costumes : the sample was filtered by the outcome — before the study, during it, or at the point of joining. Intention-to-treat analysis exists because of the second.
Missing data comes in three regimes : MCAR (harmless, rarely true), MAR (predictable from observed data, handled by multiple imputation), MNAR (depends on the missing value itself, unfixable statistically). Which regime you are in is a judgement about the world, not a test result.
Weighting fixes composition, not disposition , and assumes respondents resemble non-respondents within cells. MRP is the strongest version and rests on the same assumption.
And the habit that carries all of it: look at what is not in the data, and ask why it isn't.
Total survey error — coverage, sampling, non-response, measurement and processing error together.
Unit / item non-response — not taking part at all; skipping particular questions.
Non-response bias — the product of the missing proportion and the difference between missing and present. Measure-specific, not survey-wide.
Self-selection — participants choosing themselves; the mechanism behind open polls and review distributions.
Survivorship bias — the sample was filtered by the outcome before observation.
Attrition — patterned loss of participants over a study's life.
Intention to treat — analysing everyone from the moment of assignment, regardless of what they did afterwards.
Selection into treatment — people who take part differ in ways that predict the outcome; the central obstacle to causal inference.
MCAR / MAR / MNAR — missingness unrelated to anything; predictable from observed data; dependent on the missing value itself.
Multiple imputation — generating several plausible completed datasets and combining results, carrying the extra uncertainty through.
Weighting / raking / post-stratification — adjusting the sample to known population totals.
MRP — multilevel regression and post-stratification; models outcomes in small cells then re-weights.
Telescoping — pulling remembered events forward in time, inflating recent-period rates.
Panel conditioning — repeated surveying changing the respondents.
One — do Wald on something. Take any set of successful examples you have been shown — companies, careers, schools, treatments — and describe the cases that would have been removed from that set before you saw it. Then ask whether the "lesson" survives their inclusion.
Two — find the J. Look at the star-rating distribution of any product or restaurant with many reviews. Note the shape. Then estimate how many of the people who bought it wrote anything at all.
Three — apply the disposition test. Take a survey finding about behaviour — volunteering, reading, exercising, attending, giving. Ask whether that behaviour is correlated with the willingness to answer a survey. If it is, expect the estimate to be too high.
Four — find a denominator. Take a rate quoted in the news and identify what is on the bottom. Then construct a different, equally defensible denominator and consider whether the comparison would still hold.
Five — write a bias direction. Take a study you find persuasive and write one sentence: if this estimate is wrong because of who is missing, it is probably wrong in this direction, by roughly this much. Being able to write that sentence is a more useful skill than being able to compute a confidence interval.
Topic 7.3 is done. You know who a study is looking at and what could have gone wrong before a single question was asked.
Topic 7.4 opens the quantitative toolkit , and it starts with the instrument that produces most of what the public knows about society — and that fails in ways almost nobody notices, because the failures are in the wording.
7.4.1 — Surveys covers what question wording does, why the order of questions changes the answers, what an attitude measurement actually captures, and how to tell a well-built questionnaire from one that manufactures its findings.