"Merely descriptive" is used in the social sciences as a mild insult, applied to work that reports what is the case without explaining why.
It is the wrong attitude and it causes real damage. Most of what anyone actually needs to know about a society is descriptive: how many, how distributed, who, where, changing how fast. And every causal analysis rests on descriptive facts that somebody had to establish first — usually facts that, when checked, turn out to be more complicated than the analysis assumed.
This lesson is about reading and producing description honestly, which is harder than it sounds, because a summary is a compression, and every compression throws something away.
The average wait was four weeks. Almost nobody waited four weeks.
A health service publishes its performance: the mean waiting time for a specialist appointment is 28 days. It is within target. It has been broadly stable for three years.
A researcher asks for the underlying distribution, plots it, and finds something that no summary had shown.
The distribution has two humps.
One large group is seen within three to five days. These are patients referred through an urgent pathway, or with conditions on a fast-track list, or referred by practices that share a building with the clinic.
A second, smaller group waits between four and nine months. These are routine referrals from practices without those arrangements, in the outer part of the district, disproportionately for conditions that are chronic rather than acute.
Almost nobody waits four weeks. The mean sits in the empty valley between the two humps, describing a patient who does not exist.
And the two humps are not variation. They are two different processes — two referral systems, two sets of practices, two populations — averaged together and reported as one system with a typical performance.
Everything follows from seeing the shape. The target is met and is meaningless. The stability over three years conceals the possibility that the two groups are diverging. And the policy question is not "how do we reduce the average" but "why are there two systems" — a question the summary statistic made unaskable.
Now the general version of the same point.
In 1973 the statistician Francis Anscombe published four small datasets. Each has the same mean of x , the same mean of y , the same variances, the same correlation coefficient, and the same fitted regression line to two decimal places.
Plotted, they look nothing like each other. One is a clean linear relationship. One is a perfect curve — a strong relationship that is simply not straight. One is a tight straight line with a single wild outlier dragging the fit. One is a vertical stack of points at a single x value plus one distant point that alone creates the entire apparent relationship.
Four different worlds, one set of summary statistics.
The modern extension makes it funnier and no less serious: a set of datasets constructed to share summary statistics to two decimal places, one of which, when plotted, is a picture of a dinosaur.
Look at the data.
Which losses matter, and how would you know?
Every descriptive statistic answers one question and silences others.
A mean answers "if the total were shared equally, how much each?" — which is exactly the right question for a budget and exactly the wrong one for a typical experience.
A median answers "what is the middle person's value?" — right for typicality, and blind to everything happening at the extremes, which for income and wealth is where the action is.
A total answers "how much altogether?" and hides distribution entirely.
A percentage change answers "how much bigger?" and hides the base, which is how a rise from two cases to three becomes "a 50 per cent increase".
None of these is dishonest. Each is a choice about which question to answer , made by whoever produced the number — and usually not stated. The reader's task is to reconstruct the question, and then to ask what the other ones would have shown.
The four things you need about any distribution
Centre, spread, shape, position. Most reporting gives you one of the four.
Centre.
The mean is the balance point; it uses every value, which is its strength and its weakness — one extreme value moves it, and social distributions are full of extreme values.
The median is the middle; half above, half below. Unaffected by how extreme the extremes are , which makes it the right default for income, wealth, house prices, waiting times, firm sizes and city populations — all strongly right-skewed.
The mode is the most common value, and it is neglected. For a bimodal distribution the mode is the only measure of centre that tells you the truth : that there are two.
The diagnostic : when the mean is much larger than the median, the distribution has a long right tail — a minority with very high values. When someone reports a mean for a quantity you know to be skewed, ask why.
Spread.
Range (max minus min) is fragile — two values decide it. The interquartile range — the span of the middle half — is robust and underused. The standard deviation is the standard measure and assumes a roughly symmetric distribution to be interpretable.
And for many social questions spread is the finding. Two countries with the same median income and different spreads are different societies. Two schools with the same average result and different spreads are doing different things.
Shape.
Skew — a long tail on one side. Modality — how many humps. And bimodality almost always means two populations or two processes mixed together , which is the most useful single diagnostic in this lesson.
Position.
Percentiles and deciles. Where does a particular case sit? And ratios between positions — the ninetieth percentile divided by the tenth is a far more legible measure of dispersion than a standard deviation, and can be reported to anybody.
Inequality measures answer different questions, and choosing one is a value judgement.
The Gini coefficient is the standard single number, running from 0 (everyone equal) to 1 (one person has everything). It is convenient and it has a specific insensitivity: it is most responsive to changes around the middle of the distribution and least responsive at the extremes. A large transfer from the very poor to the very rich can move it remarkably little.
The Palma ratio — the share of the top 10 per cent divided by the share of the bottom 40 per cent — was proposed precisely because those are the parts that move, while the middle 50 per cent's share is strikingly stable across countries and time.
Top income shares — the share going to the top 1 per cent or 0.1 per cent — became central to the modern literature because they can be measured from tax records rather than surveys , and surveys systematically miss the top (see 7.3.2). Much of what is now known about long-run inequality comes from that source switch, which is 7.4.4's lesson producing a research revolution.
The Theil index has a property the others lack: it decomposes , so total inequality can be split into the part between groups (regions, sectors, ethnic groups) and the part within them. That decomposition is often the sociologically interesting result.
So there is no neutral measure of inequality. Each embeds a judgement about which differences matter most, and a country's ranking can change depending on the choice. The honest practice is to report more than one — and when a single number is quoted at you, to ask which part of the distribution it is sensitive to.
Rates, comparisons and change
Four places where description goes wrong before any analysis begins.
One — the denominator. A rate is a fraction and the bottom is a choice (see 7.3.2). Crime per resident or per person present; accidents per vehicle or per mile; infections per capita or per test. Choose the exposure, not the convenient population.
Two — standardisation. Country A has a higher crude death rate than country B. Country A may simply be older. Comparing crude rates across populations with different age structures is close to meaningless, and the remedy — age standardisation , applying each population's age-specific rates to a common reference age structure — is routine in demography and routinely missing everywhere else.
The same logic applies far beyond mortality. Comparing schools' raw results without accounting for intake, hospitals' outcomes without case-mix, or firms' accident rates without the composition of their work, is the same error wearing different clothes.
Three — absolute versus relative change. "Risk doubles" is uninformative without the baseline: from one in a million to two in a million, or from one in five to two in five. Relative change makes small things sound large, and it is the single most common device in health and crime reporting. Always ask for both.
And the related confusion: percentage points versus per cent. A rise from 4 per cent to 6 per cent is two percentage points and a 50 per cent increase . Both are true; only one is usually intended; and the ambiguity is exploited constantly.
Four — real versus nominal, and the base year. Money figures across time must be adjusted for prices or they are not comparable. Index numbers depend on the base year, and choosing a base year at an unusual moment makes everything afterwards look dramatic. Whenever you see an index, ask what year equals 100, and why that year.
Looking at it
Tukey's imperative, and what a good graph does.
John Tukey's programme of exploratory data analysis was an argument that looking at data is not preliminary to the real work; it is work, and skipping it is how people fit models to shapes the model cannot represent.
The order of operations that experienced analysts follow: plot before you model, plot the residuals after you model, and plot the sub-groups before you believe the whole.
What a good graph does.
Shows the distribution, not only the summary. A bar chart of means hides everything in this lesson. A box plot shows quartiles and outliers; a violin plot shows the shape; and plotting the individual points — where numbers permit — shows the truth, including the bimodality, the ceiling effects and the impossible values.
Uses small multiples for comparison. The same chart repeated across groups, on identical axes, lets the eye do the comparison that a single crowded chart cannot.
Encodes with position and length , which people read accurately, rather than area and angle, which they do not. This is the case against pie charts and bubble charts — human judgement of area is systematically compressed, so differences look smaller than they are.
Uses a log scale when the data span orders of magnitude — incomes, city sizes, firm sizes, epidemic growth — and says so clearly , because a log scale makes exponential growth look linear and unlabelled it is genuinely misleading.
And is honest about the axis. A truncated y -axis magnifies trivial differences into apparent drama. It is sometimes legitimate — with a variable that never approaches zero, forcing a zero baseline wastes the whole plot — and it must be visible.
Data cleaning is not clerical work. It is where the substantive errors hide.
Missing-value codes read as data. Datasets encode missingness as −9, −99, 999 or 9999. A researcher who does not read the codebook computes an average age that includes several respondents aged ninety-nine hundred — and the plot would have shown it instantly (see 7.4.4).
Top-coding. Many public datasets cap the highest values to protect anonymity, so every income above a threshold is recorded at that threshold. Any analysis of the top of the distribution using top-coded data is measuring the cap.
Structural zeros and true zeros. A zero that means "none" and a zero that means "not applicable" are different, and combining them corrupts every average.
Impossible and implausible values. Ages of 200, incomes of zero for full-time employees, dates before the person's birth. These are informative : they tell you something about how the data were collected, and where they cluster tells you which field or which office.
And the honest rule: every cleaning decision is a decision about what exists, and should be documented. Dropping the outliers changes the answer. So does keeping them. What is not acceptable is doing either without saying so.
Some of the most consequential social science ever done was descriptive, and this is not a consolation prize.
Booth's survey of London and Rowntree's of York established, by counting, that the majority of poverty had structural causes — low wages, old age, sickness, large families at a particular life stage — rather than idleness. That was a description, and it reorganised British social policy.
Du Bois's The Philadelphia Negro (see 4.5.1) was a house-by-house description of a population that had been discussed for decades entirely through assertion.
The reconstruction of long-run top income shares from tax records transformed what is known about twentieth-century inequality — not by explaining anything, but by measuring what surveys could not reach.
The mapping of intergenerational mobility down to small local areas changed the questions people ask about opportunity, because it showed variation within countries that national averages had hidden completely.
In every case the contribution was to establish what is the case , against a background of confident assertion. And in every case the description made the causal questions askable — you cannot explain a pattern nobody has established.
The bimodal waiting times in the story are the small version of the same thing. The description is the finding, and it changes the question.
Because most misreading of quantitative claims happens at the descriptive stage, before any statistics are involved.
Six questions, none requiring any technical knowledge.
Was that a mean or a median, and what would the other one have been?
What does the distribution look like — and is there any reason to suspect two populations?
What is the denominator, and is it the right exposure?
Is this comparison standardised for the obvious composition differences?
Is that a percentage change, and of what base?
Where does the y-axis start?
And one for producers. Before you fit anything, plot it — the distribution of every variable you will use, the relationship between the ones you care about, and the same plots broken down by the groups that might differ. Anscombe's four datasets exist to make one point , and it has not become less true with better software: summary statistics are not a substitute for looking, and the modern version of the demonstration is a dinosaur.
A summary is a lossy compression, and the losses are chosen by whoever produced it. The bimodal waiting-time distribution has a mean describing nobody, produced by two systems averaged together — and Anscombe's four datasets share every summary statistic and look nothing like each other.
Four properties of any distribution : centre (mean uses everything and is dragged by extremes; median is robust and is the right default for skewed social quantities; mode is the only honest centre for a bimodal distribution ), spread (interquartile range is robust; for many questions spread is the finding ), shape (bimodality almost always means two populations mixed ), and position (percentiles, and p90/p10 ratios that anyone can read).
Inequality measures are not neutral. The Gini is least sensitive at the extremes; the Palma ratio was built for the parts that actually move; top shares came from tax records because surveys miss the top; the Theil index decomposes into between-group and within-group inequality. Report more than one.
Four places description fails before analysis : the denominator; the absence of standardisation for composition (age, intake, case-mix); relative change quoted without the baseline, and percentage points confused with per cent; and unadjusted money figures with an unstated base year.
Look at the data. Plot before modelling, plot residuals after, plot subgroups before believing the whole. Show distributions rather than bars of means; use small multiples; encode with position and length rather than area; label log scales; make truncated axes visible.
Cleaning is substantive : missing-value codes read as data, top-coding that caps the top of the distribution, structural versus true zeros, and implausible values that tell you how collection worked. Every cleaning decision changes the answer and belongs in the write-up.
And "merely descriptive" is a bad insult. Booth and Rowntree, Du Bois in Philadelphia, top income shares from tax data, and small-area mobility maps all changed what could be asked — because you cannot explain a pattern nobody has established.
Mean / median / mode — balance point; middle value; most common value.
Skew — asymmetry; a long right tail makes the mean exceed the median.
Bimodality — two humps, almost always indicating two populations or processes.
Interquartile range — the span of the middle half; robust measure of spread.
Percentile / decile / p90–p10 ratio — position in the distribution and a legible measure of dispersion.
Gini coefficient — single-number inequality measure, least sensitive at the extremes.
Palma ratio — top 10 per cent share divided by bottom 40 per cent share.
Theil index — decomposable into between-group and within-group inequality.
Standardisation — applying a population's group-specific rates to a common reference structure so comparisons are not composition effects.
Absolute vs relative change; percentage points vs per cent — the difference between two-in-a-million and doubling.
Real vs nominal; base year — price-adjusted figures; the year an index equals 100.
Exploratory data analysis — Tukey's insistence that looking at the data is the work, not a preliminary.
Small multiples — the same chart repeated across groups on identical axes.
Top-coding — capping the highest recorded values, which makes the top of the distribution unmeasurable.
Missing-value codes — sentinel numbers such as −9 or 99 that become nonsense if treated as values.
One — find the mean–median gap. For any quantity you can look up — income, house prices, waiting times, salaries in a field — find both. The size of the gap tells you the shape without a plot.
Two — hunt a bimodal distribution. Think of a service or process you know well and ask whether there are really two routes through it. Then ask what its average conceals.
Three — restate a scary statistic. Take a health or crime claim reported as a relative change and find the absolute numbers. Do this five times and it becomes automatic.
Four — check an axis. Find a chart in the news with a truncated y-axis and redraw it in your head from zero. Note how much of the story survives.
Five — pick your inequality measure. For a country you care about, find its Gini and its top 1 per cent share over the same period. Ask whether the two tell the same story ; frequently they do not, and the reason is in this lesson.
Description establishes what the pattern is. The next question is the one everyone jumps to immediately and almost nobody handles well: does A cause B?
7.6.2 — Correlation, Causation and the Confounder covers what a correlation is and is not, the four things that produce one besides causation, why "controlling for" variables is far less powerful than it appears, what a regression coefficient actually means, and how to tell a causal claim that has earned its verb from one that has borrowed it.