A survey of 1,200 people is reported as telling you what a country of two hundred million thinks. That should strike you as an outrageous claim, and it is worth holding on to the outrage for a moment , because the reasons it usually works — and the specific conditions under which it collapses completely — are among the most useful things in this Part.
The heart of it is a fact that almost nobody believes on first hearing: the size of the sample matters far less than how it was chosen.
1936: two million answers that were wrong, and fifty thousand that were right.
The Literary Digest was one of the most respected magazines in the United States, and it had correctly called the presidential election five times running using enormous postal straw polls.
For the 1936 contest between Franklin Roosevelt and Alf Landon it went further than ever before. It mailed roughly ten million ballot cards. About 2.4 million came back.
The result was emphatic: Landon would win, 57 per cent to 43.
Roosevelt won 61 per cent of the popular vote and carried every state but two — an electoral college margin of 523 to 8. The magazine's reputation did not survive it, and it ceased publication within two years.
Meanwhile George Gallup, with a sample a fraction of the size — tens of thousands — called it correctly. He also did something more impressive: before the Digest's poll closed, he predicted what the Digest was going to get wrong, and by roughly how much. He could do this because he understood something the Digest did not.
What went wrong has two parts, and both matter.
The list. The Digest assembled its addresses from telephone subscribers, automobile registrations, club memberships and its own subscriber list. In 1936 — in the depths of the Depression — owning a telephone or a car meant something. The list was systematically richer than the electorate, and wealth was strongly related to the vote.
And the returns. Later analysis found the second problem was, if anything, larger. Only about a quarter of those who received a card sent it back , and Landon's supporters — motivated, and hostile to the incumbent — returned them at a higher rate than Roosevelt's. So even from a skewed list, the people who chose to answer skewed it further.
Ten million cards could not rescue it. Nothing could have. A bigger sample from a biased process is a more precise estimate of the wrong number.
And now the part that keeps everybody honest. Twelve years later, in 1948, Gallup confidently predicted that Dewey would beat Truman — and newspapers printed it. He was wrong, for different reasons: quota sampling that left interviewers free to choose whom to approach within quotas, and stopping polling weeks before the election as though opinion had settled.
Two lessons, and you need both. Sample size is not the thing. And the person who understands sampling can still be wrong , which is why the rest of this lesson is about mechanisms rather than about heroes.
Why should a spoonful tell you about a pot — and why does a bigger spoon not help?
The cooking analogy is genuinely the right one, and it repays being taken seriously.
You taste soup with a spoon, not a ladle, and never the whole pot. One spoonful tells you about the whole thing — provided it was stirred.
If it was not stirred, no amount of tasting from the top helps. A bigger spoon from the unstirred top is not closer to the truth; it is only a more confident report about the top. That is the Digest.
And here is the part that surprises everyone: once it is stirred, a bigger pot does not need a bigger spoon. The same spoonful tells you about a litre or a bathtub. This is exactly true of sampling , and it is why 1,200 people can describe a country: the precision of an estimate depends on the absolute number sampled, not on what fraction of the population that is.
Stirring is the whole thing. Stirring is randomisation.
The vocabulary, which is where confusion starts
Five terms, and one of them is the villain of most surveys.
The population (or target population ) — everyone or everything you want to describe. All adults in the country. All firms with more than fifty employees. All arrests in a city in a year.
The sampling frame — the actual list, register or procedure from which you draw. The electoral roll. A telephone number bank. A business register. Everyone who walks past a particular door.
The sample — those actually selected.
The respondents — those who actually take part, which is not the same thing (see 7.3.2).
The unit of analysis — what each case is : a person, a household, a firm, a neighbourhood, an event (see 7.2.1).
The frame is where studies go wrong, and it is the least discussed part of any report.
You do not sample the population. You sample the frame. Every difference between them is coverage error , and it is invisible from inside the study, because the people missing from the frame cannot appear anywhere in the data — not even as non-responders.
Who falls out of ordinary frames : people without fixed addresses; people in institutions; people without phones or with only work numbers; people who moved recently; people not on the register they were supposed to be on; unregistered businesses and informal work; and — in almost every country — the very poor and the very rich at once, for opposite reasons.
And notice : the Digest's frame problem is not a historical curiosity. Every frame encodes the administrative reach of some institution , which is 6.1.1's argument about the census, arriving as a technical parameter.
Probability sampling, and why it earns its authority
The defining property: every element has a known, non-zero chance of selection.
Not equal — known. That is the whole basis of the inference, because if you know the probability with which each person could have been chosen, you can reason backwards from the sample to the population and put a number on your uncertainty. Without known selection probabilities, that reasoning is not available — which is the real difference between probability and non-probability sampling, and it is a difference in kind, not in quality.
Simple random sampling. Every element has the same chance; every combination is equally likely. The reference case, rarely used directly because you need a complete list.
Systematic sampling. Take every k -th element from a list, with a random start. Practical and usually equivalent to simple random — unless the list has a periodicity that matches your interval , which is a real and occasionally spectacular failure mode (take every 10th house in blocks of ten and you may select every corner house in the city).
Stratified sampling. Divide the population into strata — region, sector, age band — and sample within each. Two things this buys you. It guarantees that each stratum appears rather than leaving it to luck. And if the strata differ from each other and are internally similar, it produces more precision than simple random sampling of the same size , which is free accuracy for a little prior knowledge.
Disproportionate stratification deliberately over-samples small groups so there are enough of them to analyse — a minority population, very large firms — and then weights them back down for population estimates. This is not cheating; it is the correct way to study small groups , and the weights are what make it legitimate.
Cluster sampling. Sample groups, then everyone or a subset within them. Select 60 villages, then 20 households in each. Done for cost , because interviewers travel.
Multi-stage sampling. Clusters within clusters: districts, then wards, then blocks, then households, then a person within the household. Almost every serious national survey in the world is a stratified multi-stage cluster design , and the phrase "a nationally representative sample" nearly always means this rather than simple random sampling.
Probability proportional to size (PPS). When clusters differ greatly in size, select them with probability proportional to their size, so that every individual retains an equal overall chance. This is why a big city and a small town are not sampled as though they were equivalent units.
Clustering has a price, and it is called the design effect.
People in the same village, school or workplace resemble each other more than two people drawn at random. So each additional person from an already-sampled cluster adds less new information than a fresh independent draw would.
The design effect quantifies this, and the effective sample size is the real sample size divided by it. A clustered survey of 4,000 with a design effect of 2 has the precision of a simple random sample of 2,000.
This matters for reading : a headline sample size means less than it appears if the design was clustered, and it is a common source of overstated confidence in reports that quote n without quoting the design.
Why it works, and the fact that surprises everyone
The machinery, in words, with no algebra you need to do.
Imagine drawing a random sample of 1,000 and computing the average. Then doing it again. And again — thousands of times.
Those averages would not all be the same, but they would cluster around the true population value, in a predictable bell-shaped spread. That spread is the sampling distribution , and its width is the standard error . This clustering happens regardless of the shape of the underlying population — the central limit theorem — which is why the method works on income, which is wildly skewed, as well as on height, which is not.
The standard error shrinks with the square root of the sample size. Which has two enormous practical consequences.
Consequence one: precision is expensive. To halve your margin of error you must quadruple your sample. Going from 1,000 to 2,000 buys a modest improvement; going from 1,000 to 10,000 costs ten times as much for a roughly threefold gain.
Consequence two — and this is the one nobody believes: the population size is almost irrelevant. The formula contains the sample size and not the population size. A random sample of 1,000 gives you the same precision for a town of 40,000 as for a country of 400 million.
(The one exception: when your sample is a large fraction of the population — say more than about 5 per cent — a finite population correction makes your estimate better than the formula suggests, not worse. Sampling a fifth of a small town is more precise than the standard formula implies, not less.)
So the standard poll's ±3 points at n≈1,000 — from the arithmetic for a proportion near 50 per cent, at 95 per cent confidence — is ±3 points whether the country has five million people or a billion.
Stop and sit with that. It is the single most counterintuitive true thing in survey research, and it is why "they only asked a thousand people!" is not the objection people think it is. The objection worth making is always about the frame, the selection and the non-response — never about the size.
What "95 per cent confidence" actually means, since almost everyone gets this wrong.
It does not mean there is a 95 per cent chance the true value lies in this particular interval. The true value is a fixed number; it is either in the interval or it is not.
It means: the procedure that generated this interval would capture the true value 95 times out of 100, if repeated. The confidence is a property of the method , not of this one result.
Two things follow that matter practically.
One in twenty intervals misses. In a report with forty published estimates, expect a couple to be off — and you cannot tell which.
And the margin of error only covers sampling variability. It says nothing about coverage error, non-response, question wording, mode effects or interviewer effects (see 7.2.2, 7.3.2). Those are usually the larger errors, and they never appear in the ± figure. A poll reported as "±3 points" may easily be six points wrong from causes the interval was never designed to include.
Non-probability sampling, and the logic that is not a poor relation
When you cannot use probability sampling, you have two very different situations.
The first: you wanted probability sampling and could not get it. Then you have a weaker version of the same task, and you should say so.
Convenience sampling — whoever is available. Students, passers-by, an online panel. Cheap, and generalisation is unsupported.
Quota sampling — filling targets so the sample matches the population on chosen characteristics: so many women, so many over-65s, so many in each region. It looks like stratification and it is not , because the selection within each quota is left to whoever is recruiting, and interviewers approach approachable people. This is what beat Gallup in 1948. Modern online panels with quotas and statistical adjustment are considerably more sophisticated, and their errors remain of this family.
Snowball sampling — respondents refer others. Often the only way to reach hidden populations. It over-represents the well-connected, because the more ties you have the more likely you are to be referred.
Respondent-driven sampling — Heckathorn's refinement of snowballing: structured referral coupons, limits on how many each person may pass on, and recording of network size, which under stated assumptions permits estimates with quantified uncertainty. A genuine improvement, and its assumptions are strong.
The second situation is entirely different, and it is not a compromise: you never wanted a statistical generalisation in the first place.
Purposive sampling — cases chosen for a reason. The extreme case, the deviant case, the typical case, the critical case ("if it fails here , it fails anywhere"), the pair matched on everything but one thing.
Theoretical sampling — cases chosen by what the emerging theory needs next (see 7.2.3, 7.7.1).
Statistical generalisation and theoretical generalisation are different operations, and only one of them needs a random sample.
Statistical generalisation infers from a sample to a population: about 34 per cent of adults, ±3. It requires probability sampling and its authority comes from the selection mechanism.
Theoretical (or analytic) generalisation infers from a case to a proposition : this mechanism operates in this way under these conditions. Its authority comes from the logic of the case, not from its representativeness — which is why "but it's only one factory" is often not an objection at all.
Mario Small's argument in "How many cases do I need?" is the cleanest statement : qualitative researchers who defend small samples in the language of statistical representativeness are fighting on ground they cannot win and did not need. A case study is not a small survey. The proper logic is case-based — sequential selection, each case chosen for what it can establish, saturation rather than sample size (see 7.5.1).
And the two can refute each other in one direction only. A single well-documented case cannot tell you a rate — but it can decisively refute a universal claim. One instance of a thing happening is enough to kill "that never happens." This asymmetry is why single cases sometimes matter enormously and sometimes not at all , and knowing which is a matter of what claim is on the table.
Big data does not solve this, and it makes one part of it worse.
The intuition that vast datasets escape sampling problems is wrong, and the mathematics of why is one of the most useful results of the last decade.
Xiao-Li Meng showed that the error in an estimate from a large non-random dataset depends on three things : how big the population is, how large your dataset is relative to it, and — decisively — how strongly the probability of a record being in your data correlates with the thing being measured.
That third term is the killer, because it is multiplied by the population size. A tiny systematic relationship between "being in the data" and "the answer" destroys an enormous dataset. Meng's demonstration is worth remembering: a survey of over two million respondents, with a small selection-outcome correlation, had the effective accuracy of a simple random sample of a few hundred. He called it the big data paradox — the more data you have, the more confident and the more wrong you can be.
Which is the Literary Digest, restated in modern algebra. Ten million cards, correlated non-response, catastrophic error.
The famous applied example is Google Flu Trends, which estimated influenza prevalence from search queries, worked impressively at first, and then substantially overestimated flu for an extended period — with the analysis by Lazer and colleagues attributing this to changes in the search platform itself and to the model fitting seasonal patterns rather than flu. The data were enormous. Nobody knew who was in them or why.
The rule to keep : whose behaviour generates this data, and what makes someone appear in it? Platform data covers platform users doing platform-visible things. Administrative data covers people an institution has processed — and being processed is not random (see 7.4.4).
Because "how many people did they ask?" is the wrong question, and it is the one everybody asks.
Three better questions, in order of importance.
One — what list did they draw from, and who is not on it? This is coverage error, it is invisible in the results, and it destroyed the Digest.
Two — how were people selected from that list, and did they choose to be in it? Self-selection is the mechanism that turns a big sample into a confident error.
Three — who ended up answering, and are they different from those who did not? Which is the next lesson.
Only after those three does size matter at all — and by then it usually does not decide anything.
And on the other side : when you meet a study of eleven people, do not reach for "too small". Ask what claim it is making. If it says this is what happens generally , the objection is right. If it says here is a mechanism, documented in operation, with conditions specified , the objection is a category error — and the study may be worth more than a survey of four thousand.
Precision comes from how the sample was chosen, not from how big it is. The Digest's 2.4 million returns were wrong because of a frame skewed towards telephone and car owners and differential non-response; Gallup's tens of thousands were right — and twelve years later his quota sampling failed too.
You sample the frame, not the population , and the gap between them is coverage error — invisible from inside the data. Every frame encodes some institution's administrative reach.
Probability sampling means known, non-zero selection probabilities — not equal ones. Simple random, systematic, stratified (including disproportionate with weights), cluster, multi-stage and PPS. Clustering costs precision , measured by the design effect and expressed as an effective sample size.
The standard error falls with the square root of n , so halving the margin of error requires quadrupling the sample — and the population size barely enters , which is why 1,000 people give ±3 points whether the country has five million or a billion. "They only asked a thousand people" is not the objection; the frame is.
95 per cent confidence is a property of the procedure , not of the particular interval; and the margin of error covers only sampling variability, not coverage, non-response, wording or mode — which are usually larger.
Non-probability sampling comes in two kinds : a compromise (convenience, quota, snowball, respondent-driven), or a different logic entirely. Statistical generalisation infers to a population and needs random selection; theoretical generalisation infers to a proposition and needs a well-chosen case. A single case cannot establish a rate; it can decisively refute a universal claim.
And big data does not escape any of this. Meng's result: when selection into the data correlates even slightly with the outcome, the error scales with population size — a two-million-person survey with the accuracy of a few hundred. The more data, the more confidently wrong.
Population / sampling frame / sample / respondents — who you want to describe; the list you draw from; who is selected; who takes part.
Coverage error — the gap between frame and population.
Probability sampling — every element has a known, non-zero chance of selection.
Simple random / systematic / stratified / cluster / multi-stage / PPS — the standard designs.
Weighting — adjusting for unequal selection probabilities or known population totals.
Design effect / effective sample size — the precision cost of clustering, and the equivalent simple-random size.
Sampling distribution / standard error — the spread of estimates across hypothetical repeated samples, and its width.
Central limit theorem — why sample means cluster in a bell shape regardless of the population's shape.
Margin of error / confidence interval — the sampling-variability range; the confidence is a property of the procedure.
Convenience / quota / snowball / respondent-driven sampling — non-probability designs, in ascending order of discipline.
Purposive and theoretical sampling — cases chosen for what they can establish or for what the theory needs.
Statistical vs theoretical generalisation — inference to a population; inference to a proposition.
Big data paradox — Meng: when selection correlates with the outcome, larger non-random datasets can be more confidently wrong.
One — find the frame. Take any survey finding reported this month. Locate the methodology note and identify the frame. Write down two kinds of person who cannot be in it.
Two — do the square-root arithmetic. A study of 400 has roughly twice the margin of error of a study of 1,600. Check this against two published studies of different sizes and see whether their reported precision matches.
Three — kill the size objection. Next time you hear "they only asked a thousand people", write down what the actual weakness of that poll is. It will be the frame, the response rate or the wording, every time.
Four — audit a big dataset. Take any claim based on platform or administrative data — search trends, transactions, app usage, service records. Ask: what makes a person appear in this data, and could that be related to the thing being measured? If yes, the size is irrelevant.
Five — classify a claim. Find a qualitative study and decide whether its claim is statistical or theoretical. Then decide whether "the sample is too small" is a real objection or a category error. Do this three times and the distinction will be permanent.
Everything so far assumed the people you selected actually took part. In modern survey research, most of them do not.
Response rates that were 70 or 80 per cent in the mid-twentieth century are now, for many telephone surveys, in the single digits. And yet a great deal of that research remains reasonably accurate , which is a puzzle in itself — until you understand what determines whether non-response causes bias, which is not the rate.
7.3.2 — When Samples Lie covers non-response, self-selection, survivorship, attrition, and the everyday cases where a sample that looks fine produces an answer that is confidently, systematically wrong.