Most social research does not collect its own data.
It uses censuses, national surveys, panel studies, tax records, school records, hospital records, police records, registers, archives and the enormous digital residue of everyday life. This is a good thing — the data are vast, long-running and often better than any individual researcher could produce.
And every one of these sources was made by an institution, for a purpose that was not yours, about the people that institution deals with. Those three facts leave marks on everything you can conclude, and this lesson is about reading them.
A crime rate that fell without any crime disappearing.
For years in England and Wales, two entirely different official measures of crime existed side by side, and they told different stories.
The first was police recorded crime : the count of offences recorded by police forces.
The second was the Crime Survey for England and Wales : a large, long-running survey that asks a representative sample of households what has actually happened to them, whether or not they reported it.
The survey exists because of a simple observation: most crime is never reported, and what is reported is not always recorded. The gap between crimes committed and crimes recorded has a name — the dark figure of crime — and its size varies enormously by offence. Vehicle theft is reported at very high rates, because insurance requires it. Sexual offences and domestic abuse are reported at low rates, for reasons that have nothing to do with how often they occur.
And in the early 2010s the two measures diverged in a way that could not be explained by crime.
The inspectorate of constabulary investigated and found systematic under-recording. Its 2014 review concluded that a substantial share of crimes reported to the police — of the order of one in five — were not being recorded as crimes at all. Incidents were "no-crimed", downgraded, filed as anti-social behaviour, or dealt with informally. Some of this reflected pressure to show falling crime figures, which is Goodhart's law arriving exactly as predicted (see 7.2.2).
In January 2014 the UK Statistics Authority removed the National Statistics designation from police recorded crime — a formal statement that a headline official statistic was no longer trustworthy as a measure of what it appeared to measure.
Sit with the shape of that. For a period, the country's crime rate — quoted in Parliament, argued over in elections, used to allocate resources — was substantially a measure of police recording practice.
And now the deeper point, which is not about scandal. Even when nobody is manipulating anything, police recorded crime was never a measure of crime. It is a measure of crimes that were witnessed or suffered, defined by the victim as criminal, reported to police, and recorded by an officer as an offence. Every one of those five steps is a social process with its own determinants — and the biggest of them, whether people report at all, depends on trust in the police, which is exactly what the crime rate is often used to argue about.
Administrative data records institutional processing, not events.
This is the single idea that organises the whole lesson, and it generalises far beyond crime.
Hospital records record illnesses that reached a hospital, in people who sought care, coded by a clinician for a system that pays by code.
School records record what schools are required to record, by definitions set by a ministry, about pupils who are enrolled and attending.
Tax records record declared income of people inside the tax system.
Welfare records record claims made by people who knew about the benefit, believed themselves eligible, and completed the process.
Migration statistics record crossings that were registered.
Kitsuse and Cicourel put it with permanent clarity in 1963: official rates are produced by the rate-producing agencies. They are outputs of an organisation's activity. A change in a rate is therefore always ambiguous between three things — the world changed, the reporting changed, or the recording changed — and you cannot tell which from the series alone.
This is not a reason to distrust official statistics. They are usually the best data in existence and produced by serious people under professional standards. It is a reason to read them as what they are : traces of institutional activity, from which facts about the world can be inferred with care.
What is out there
Five families of secondary data, and what each is good for.
One — official statistics. Censuses; vital registration of births, marriages and deaths; national accounts; labour force surveys; price indices; health and education statistics. Their strengths are coverage, continuity and professional standards. Their weakness is that their categories serve administration (see 6.1.1 on the colonial census, which is the same point at maximum volume).
Two — major academic surveys and panels. Repeated cross-sections like the General Social Survey, the European Social Survey, the World Values Survey and the International Social Survey Programme; and true panels that follow the same people for decades — the Panel Study of Income Dynamics, the British household panels now running as Understanding Society, the national longitudinal cohort studies, the Demographic and Health Surveys across many countries, and the India Human Development Survey.
The panel design is the important one , because it is the only way to observe the same people changing . A repeated cross-section shows you that poverty was 20 per cent in both years; a panel shows you whether it was the same people — and it is usually not, which changes everything about what poverty is (see 7.1.3).
Three — administrative and register data. The Nordic countries are the extreme case: population registers with unique identifiers permitting linkage of education, income, health, employment and family records across a whole population and across generations. Research possible there is not possible anywhere else — whole-population studies with no sampling error and no non-response.
Four — archives. Historical records, government papers, organisational records, newspapers; and increasingly qualitative data archives holding deposited interview transcripts and fieldnotes for re-analysis (which raises its own consent questions — see 7.8.1).
Five — digital trace data. Platform logs, search queries, transactions, mobile-phone location records, sensor data, and enormous text corpora. Sometimes called "found" data , as against "designed" data — and the distinction is exactly right.
Why secondary data is often the better choice, stated fairly.
Scale and power. Sample sizes no project could afford, which makes it possible to study small subgroups and rare events at all.
Time. Series running for fifty years, which no individual career can produce. Nearly everything known about long-run trends in mobility, fertility, attitudes and inequality comes from this.
Quality. National statistical offices and major panel studies employ people who do nothing but design questions, manage fieldwork and handle non-response. A lone researcher's bespoke survey is usually worse.
Replication. Public datasets can be re-analysed by anyone, which is the mechanism from 7.1.2 that makes objectivity a community property rather than a personal virtue.
And ethics. No new burden on respondents, and often the only way to study populations too small, too dispersed or too vulnerable to survey directly.
The specific ways it misleads
Six problems, each with a diagnostic.
One — the concepts are theirs, not yours. The dataset measures what the agency needed. You wanted precarity; they measured contract type. You wanted household deprivation; they measured income at a moment (see 7.2.2). The diagnostic: read the actual question wording in the codebook before you use the variable. Not the variable name — the wording. Variable names lie by compression.
Two — definitions change, and the series breaks. Unemployment definitions, disability categories, ethnic classifications, the boundary of a city, what counts as a household, what counts as a hate crime, the threshold for a "serious" incident. A jump in a long series is more often a change in definition, coverage or collection method than a change in the world. The diagnostic: before interpreting any discontinuity, look for a methodological note at that date. Statistical agencies publish them; almost nobody reads them.
Three — the unit is theirs too. "Household" is a definition, and it differs across countries in ways that make comparisons of household income, household size and household poverty partly artefacts (see 7.2.2 on measurement invariance).
Four — who is missing, and why. Administrative data covers people the institution processes. Undocumented migrants, informal workers, people outside the tax system, people who never present at services, and those the system has excluded are absent — and their absence is systematic, not random (see 7.3.2).
Five — cross-national comparison is harder than it looks. The most instructive example: homicide is the only crime that compares reasonably well internationally , because a body is difficult to conceal and most countries investigate deaths. Everything else varies with legal definitions, reporting propensity and recording practice, so international "crime rates" other than homicide are largely comparisons of criminal justice systems.
Six — aggregate data invites a specific fallacy. Much secondary data comes at area level — districts, states, countries — and a relationship between area averages need not hold between individuals at all. This is the ecological fallacy, it has its own lesson (see 7.6.4), and it is the single most common error made with official statistics.
Big data: three genuine advantages and seven real problems
Matthew Salganik's framing is the most useful available, and it is worth carrying whole.
The genuine advantages.
Big. Enough observations to study rare events and small subgroups, and to detect small effects — though as 7.3.1 showed, size does nothing about bias, and can make it worse.
Always-on. Continuous rather than periodic, which makes it possible to study unexpected events — a disaster, a strike, a sudden policy — with a genuine before-period, because the data were already being collected.
Non-reactive. People are not answering a researcher's question, so social desirability bias does not operate in the same way (see 7.4.1). What people search for differs from what they will say.
The problems, and each is serious.
Incomplete. Enormous on some dimensions and empty on the ones you need — usually demographics, income, and anything the platform had no commercial reason to record.
Inaccessible. Held by companies, released selectively, withdrawn without notice. APIs close, and published findings become permanently unreproducible.
Non-representative. Users of a platform doing platform-visible things (see 7.3.1). The data-defect problem means scale amplifies rather than dilutes this.
Drifting. The user population changes, the platform changes its interface, the algorithm changes what people see and do. A five-year trend in platform behaviour is partly a five-year history of product decisions.
Algorithmically confounded. The system generating the data is itself intervening. Recommendation systems create the correlations you then observe; "people who liked X liked Y" may be true because the system showed them Y. This is Hacking's looping effect (see 7.1.3) implemented in software and running continuously.
Dirty. Bots, duplicates, spam, testing traffic, commercial manipulation. Any measure of "public sentiment" from open platform data includes an unknown quantity of paid and automated activity.
Sensitive. Data collected for one purpose, used for another, about people who never consented and often cannot be meaningfully de-identified.
The Google Flu Trends case from 7.3.1 combines several of these at once , which is why it has become the standard parable: a large, always-on, non-reactive dataset that worked well and then degraded, through platform change and a model that had partly learned seasonality rather than influenza.
Linkage and privacy: the trade-off nobody can avoid.
Record linkage — joining tax records to health records to school records — is the most powerful thing in modern quantitative social science, and the Nordic register studies show what it makes possible.
It is also the point at which research becomes surveillance-adjacent , and the safeguards are real: secure data enclaves, approval processes, output checking, and statistical disclosure control — suppressing or perturbing cells small enough to identify someone.
And there is no free option. The most sophisticated approach, differential privacy , adds calibrated statistical noise so that no individual's presence in the data can be detected from published outputs, with a mathematically stated privacy guarantee. The United States adopted it for the 2020 census, and the controversy that followed is exactly the right thing to understand : demographers, tribal governments and redistricting analysts objected that the injected noise degraded counts for small geographies and small populations — the very groups least able to absorb the error.
Both sides were right about something. Detailed cross-tabulations of a full population genuinely can re-identify individuals. Noise genuinely does damage the small-area statistics that minority and rural communities depend on for representation and funding. There is no arrangement that gives full privacy and full accuracy , and the choice between them is a value judgement made inside a technical procedure — 7.1.2's lesson, in a census bureau.
Six habits, in the order you should do them.
One — read the documentation before the data. The codebook, the methodology report, the questionnaire, the sample design, the weighting scheme. Hours here save months.
Two — find the exact question wording for every variable you plan to use, and check that it measures what your concept needs (see 7.2.2).
Three — check the sample design and use it. Weights, strata, clusters, primary sampling units. Analysing a complex survey as though it were a simple random sample produces confidence intervals that are too narrow — routinely, by a lot (see 7.3.1 on design effects).
Four — look for breaks before interpreting trends. Definition changes, mode changes, coverage changes, boundary changes.
Five — triangulate with an independent source. The crime story above was uncovered precisely because two independent measures existed. Where only one measure exists, an error can persist indefinitely — which is an argument for maintaining expensive parallel measures that always look redundant until the day they are not.
Six — cite the version. Datasets are revised. A result from the 2019 release may not reproduce in the 2023 release, and the difference will not be your analysis.
Because these are the numbers that public argument runs on, and almost nobody asks where they came from.
Three questions, applicable to any statistic you meet.
One — what process generated this number? Not "is it accurate" but: what had to happen for a case to appear in this count? Reported, presented, diagnosed, registered, coded, filed. Each step has determinants of its own, and each is a candidate explanation for a change.
Two — has anything about that process changed? Before concluding the world changed.
Three — who never enters the process at all?
And for the researcher, one more. Secondary data is cheap in money and expensive in a different currency: you are inheriting somebody else's decisions about what exists (see 7.1.3). The variables available quietly become the concepts studied, the categories used become the categories reproduced, and a discipline that overwhelmingly analyses administrative data will produce a picture of society shaped like the state's filing system.
That is not an argument against using it. It is an argument for knowing whose filing system you are looking at.
Administrative data records institutional processing, not events — Kitsuse and Cicourel: official rates are produced by the rate-producing agencies. Police recorded crime measures crimes witnessed, defined as criminal, reported and recorded; the UK Statistics Authority stripped it of its National Statistics designation in 2014 after inspectors found around one in five reported crimes were not being recorded. The divergence was only visible because an independent victimisation survey existed.
Five families : official statistics; academic surveys and panels — with panels the only way to see the same people change ; administrative and register data, of which the Nordic linked registers are the extreme case; archives, including qualitative data archives; and digital trace data.
The advantages are real : scale, long time series, professional quality, replication by others, and no new burden on respondents.
Six problems with diagnostics : the concepts are the agency's (read the wording, not the variable name); definitions change and break series (look for a methodological note at every discontinuity); the units are theirs; missing populations are systematically missing; cross-national comparison is largely unsafe outside homicide; and area-level data invites the ecological fallacy.
Big data is big, always-on and non-reactive — and incomplete, inaccessible, non-representative, drifting, algorithmically confounded (the system generating the data is intervening in it), dirty and sensitive.
Linkage is the most powerful tool in quantitative social science and the point where research meets surveillance. Differential privacy states the trade-off exactly: the 2020 US census dispute was a genuine conflict between re-identification risk and the small-area accuracy that minority and rural communities depend on. There is no option that gives both.
And the habit: ask what had to happen for a case to appear in this count.
Secondary data — data collected by someone else, for another purpose.
Dark figure — events that occur but never enter official records.
Rate-producing agencies — Kitsuse and Cicourel's point that official rates are outputs of organisational activity.
Repeated cross-section / panel — new samples each wave; the same units followed over time.
Register data / record linkage — population administrative records; joining them across domains by identifier.
Codebook / methodology report — the documentation defining every variable and the sample design. Read first.
Trend break — a discontinuity produced by a change in definition, coverage or collection method.
Complex survey design — strata, clusters and weights that must be used in analysis or standard errors will be too small.
Found vs designed data — data that accumulated as a by-product; data produced to answer a question.
Algorithmic confounding — the platform's own systems producing the patterns observed in its data.
Statistical disclosure control / differential privacy — suppressing or perturbing outputs to prevent identification; a formal guarantee purchased with accuracy.
Ecological fallacy — inferring individual-level relationships from area-level data.
One — trace a number. Take one official statistic you have quoted or heard quoted. Find the agency's methodology note and write down the five steps a case must pass through to be counted.
Two — find a break. Take any long official series and look for a jump. Then find the methodological note at that date. The first time you find that a "dramatic rise" was a definition change, this lesson becomes permanent.
Three — read a codebook. Download the documentation for any major public survey and find the exact wording of a variable you would want to use. Compare it with what you assumed the variable meant.
Four — audit a platform claim. Take any statement based on social media or search data. Ask who is on that platform, what the algorithm shows them, and how much of the activity might not be human.
Five — check for a second measure. For any contested social trend — crime, hate incidents, mental health, homelessness, migration — find out whether an independent measure exists. Where there is only one, note that no error in it could ever be detected.
Topic 7.4 is finished. You have surveys, experiments, quasi-experiments and secondary data — the machinery for establishing how much, how many and what causes what.
None of it can tell you what something is like, what it means to the people doing it, or how a process actually unfolds. For that the discipline has an entirely different toolkit, with its own standards of rigour that are frequently mistaken for an absence of standards.
Topic 7.5 opens with the most demanding method in sociology: 7.5.1 — Ethnography and Participant Observation , where the instrument is a person, the fieldwork takes years, and the central problem is that you are changing what you are there to see.