"Correlation is not causation" is the one methodological slogan everybody knows, and knowing it turns out to help almost nobody.
People repeat it and then reason as though a large, careful, well-controlled correlation were causation after all — because surely, with a big enough sample and enough variables controlled for, what else could it be?
This lesson answers that question precisely: here are the four other things it could be , here is why "controlling for" is far weaker than it sounds, and here is what a regression coefficient actually means.
The medicine that protected millions of women's hearts, until it didn't.
Through the 1980s and 1990s an impressive body of evidence accumulated on hormone replacement therapy after the menopause. Large prospective cohort studies — including some of the best-run epidemiological studies ever conducted, following tens of thousands of women for years — consistently found that women taking hormone therapy had substantially lower rates of coronary heart disease , with reductions in the region of a third to a half.
This was not a small or careless literature. The studies were large. They followed people forwards in time, so the exposure preceded the outcome. They controlled for the obvious things : age, smoking, blood pressure, cholesterol, body mass, diabetes, exercise, family history. The association survived every adjustment. It was biologically plausible, with mechanisms proposed for how oestrogen would protect arteries. Clinical guidelines came to recommend hormone therapy partly for cardiac protection, and millions of women took it for that reason.
Then the Women's Health Initiative randomised women to hormone therapy or placebo. In 2002 the combined oestrogen-plus-progestin arm was stopped early. The treated group had more coronary events, not fewer , along with more strokes, more clots and more breast cancer — set against fewer fractures and less colorectal cancer.
What had gone wrong is instructive, because nothing was fraudulent and nobody was careless.
Who took hormone therapy was not random. In that period, women who received it were on average better off, better educated, more likely to have a regular doctor, more likely to attend screening, thinner, more physically active, and more likely to comply with any treatment given to them. They were, in the epidemiological phrase, healthy users. They would have had fewer heart attacks whatever pills they were given.
Every one of those characteristics is a confounder , and the studies had controlled for several of them — imperfectly, using crude measures, on variables that are themselves proxies for something deeper. Adjustment removed part of the bias and left enough to reverse the sign of the finding.
And the honest coda, because this course does not simplify. The story is not simply "observational studies bad, trial good". The WHI participants were on average considerably older than typical new users of hormone therapy, and later analysis of the timing of initiation suggested the cardiac picture differs for women starting close to menopause. The trial and the cohort studies were partly studying different populations — which is 7.4.2's external-validity problem arriving in the middle of the most consequential methodological lesson in modern epidemiology.
Both things are true. The observational estimate was badly confounded, and the trial does not settle every version of the question.
A and B move together. Name every reason that could be true.
One — A causes B. The interesting case, and one of five.
Two — B causes A. Reverse causation. Does social support improve health, or does being well enough to leave the house produce social support? Does religious participation reduce offending, or do offenders stop attending? Does reading to children raise attainment, or do children who are progressing well make reading to them enjoyable? In each case the arrow plausibly runs both ways , and cross-sectional data cannot distinguish them.
Three — something causes both. Confounding. The healthy-user problem above.
Four — you selected on a common effect. Collider bias — the least known, the most counterintuitive, and increasingly recognised as widespread. Explained below.
Five — chance, or shared trend. With enough comparisons something will correlate (see 7.6.3). And two variables that both rise over time will correlate with each other regardless of any connection — which is why time-series correlations between unrelated national statistics can be produced by the thousand as a party trick.
A study earns its causal verb by arguing against the other four. Not by having a large sample — a large sample makes a confounded estimate more precise, not less wrong (see 7.3.1).
What a correlation is, exactly
A number with three limitations built into it.
The Pearson correlation coefficient runs from −1 to +1 and measures how well the relationship between two variables is captured by a straight line. Zero means no linear relationship.
Limitation one: it only sees straight lines. A perfect U-shaped relationship — strong, deterministic, entirely real — can have a correlation of zero. This is Anscombe's second dataset from 7.6.1 , and it is why you plot first. Rank-based correlations pick up any consistently increasing or decreasing relationship, curved or not, and are the right tool when the relationship is monotonic but not straight.
Limitation two: strength is not the same as importance, and the intuition about size is bad. Squaring the correlation gives the share of variance accounted for: a correlation of 0.3 accounts for 9 per cent of the variation. That sounds negligible and can be enormously important in a population — small effects on common exposures produce large numbers of cases. Conversely a correlation of 0.9 between two measures of nearly the same thing means very little.
Limitation three: it says nothing whatever about direction of influence. The coefficient of A with B is identical to the coefficient of B with A. The arrow is supplied entirely by your argument, never by the data.
Confounding, and why control is weaker than it sounds
Draw the arrows first. It changes everything.
The most useful development in causal reasoning of the last thirty years is embarrassingly simple: before analysing anything, draw a diagram of what you think causes what. Variables as nodes, arrows as causal claims. This is a causal diagram — a directed acyclic graph, in the technical vocabulary — and its virtue is that it forces you to state assumptions you would otherwise leave implicit, and then tells you what to do.
A confounder is a variable with arrows pointing to both your cause and your outcome. It opens a back-door path between them — a route by which they are associated without one causing the other. Blocking that path, by controlling for the confounder, is what adjustment is for.
And the diagram immediately shows you three things that regression alone never will.
A mediator has an arrow from the cause to the outcome. It is part of the effect. Controlling for it removes the pathway and shrinks the estimate towards zero (see 7.2.2). Education raises earnings partly by changing occupation; control for occupation and much of the effect disappears — not because it was not there, but because you deleted it.
A collider is a variable that both the cause and the outcome point into . It is a common effect , not a common cause. Conditioning on a collider creates an association that does not exist in the population — which is the exact opposite of what people expect from adding a control.
And post-treatment variables of any kind are dangerous , because anything measured after the cause may be a mediator or a collider, and controlling for it can bias the estimate in either direction.
Collider bias, made concrete, because nobody believes it until they see it.
The dating example. Suppose, in the population, kindness and physical attractiveness are entirely unrelated. Now suppose people will date someone who is either kind or attractive, but not someone who is neither.
Among people you would date, the two are now negatively correlated. If someone unattractive is in your pool, they must be kind — otherwise they would not be there. The correlation is real within the pool and does not exist in the world , and it was created entirely by selecting on a common effect. "Why are attractive people so often unpleasant?" is a question generated by the sampling.
Berkson's original medical version. Two unrelated diseases will appear correlated among hospital patients, because having either raises the chance of admission. Study hospital patients and you find associations between conditions that do not exist in the population.
And the version that matters most for social science: any study of a selected group. University students. People in employment. Users of a service. Volunteers for a study. Anything that affects selection into your sample can generate correlations inside it that are pure artefact — and this is the deep link back to 7.3.2, because non-response and attrition are themselves colliders when the cause and the outcome both influence participation.
The practical warning: adding a control variable is not automatically an improvement. Controlling for a confounder reduces bias. Controlling for a collider introduces it. And you cannot tell which a variable is from the data — only from your account of how the world works (see 7.1.3).
What a regression coefficient actually means.
Say it precisely, because the precise version is much weaker than the usual reading.
A coefficient on X is the average difference in Y between cases that differ by one unit on X and have the same values on the other included variables.
It is a conditional comparison. It becomes an effect only if you can argue that, within the levels of those controls, assignment to X is as good as random (see 7.4.3). The regression does not supply that argument, and no diagnostic statistic tests it.
Four consequences that are routinely ignored.
Omitted variable bias, and its direction is reasonable to think about. If an omitted confounder is positively related to both X and Y, the coefficient on X is inflated. You can often say which way an unmeasured confounder would push, which is more useful than declaring the problem unknowable.
Measurement error changes the picture more than people realise. Error in a predictor pulls its coefficient towards zero. Which produces a nasty consequence: when a confounder is measured badly — and social confounders like "socioeconomic status" always are — controlling for it removes only part of the confounding. The estimate is then reported as "adjusted for socioeconomic status" and remains substantially confounded. This is a large part of what happened in the hormone therapy story.
More controls is not better. The right set is determined by the causal diagram, not by availability. Throwing in every variable in the dataset will include mediators, colliders and post-treatment variables, in unknown proportions.
And the table 2 fallacy. A regression is usually built to estimate one effect. The other coefficients in the same table are not each valid estimates of their own effects , because the control set that is correct for the main exposure is generally wrong for them — some of their confounders are missing, and some of their mediators are included. Yet they are read across the row and interpreted as a list of causes, constantly.
The other classic: the drinking curve.
For decades, studies found that people who drank moderately had lower mortality than those who abstained — the famous J-shaped curve, and one of the most repeated health findings in the world.
The problem is the reference group. The "abstainers" included people who had stopped drinking because they were ill — the so-called sick-quitter effect — as well as lifelong abstainers, who are themselves an unusual group in cultures where drinking is normal, and who differ on health, income, religion and sociability.
When studies separate lifelong abstainers from former drinkers, the protective effect shrinks substantially , and meta-analyses adjusting for these biases find much of it disappears. Studies using genetic variants associated with alcohol metabolism as instruments — an application of 7.4.3's logic that is much less vulnerable to lifestyle confounding — have generally not supported a cardioprotective effect at low levels.
The finding is still argued about, and the direction of the correction is clear. What makes it a good teaching case is that the bias was not in the exposure at all. It was in who was in the comparison group — which is the question 7.4.2 says to ask of everything: compared to what?
What to do instead of hoping.
One — state the diagram. What you think causes what, including the unmeasured things, drawn before the analysis. A study that does this is legible; one that does not is asking you to trust its variable list.
Two — reason about the direction of remaining bias. Not "there may be unmeasured confounding" as a ritual limitation, but the likely unmeasured confounder is X, it would push the estimate up, so the true effect is probably smaller than reported.
Three — quantify how much confounding it would take. Sensitivity analysis asks: how strongly would an unmeasured confounder have to be related to both the exposure and the outcome to explain away this result? Where the answer is "more strongly than any known risk factor", the finding is robust. Where it is "about as strongly as income", it is fragile. This converts an unanswerable objection into a number.
Four — triangulate across designs with different weaknesses. A cohort study, a sibling comparison, a natural experiment and an instrumental-variable analysis are vulnerable to different biases. When they converge, the case is strong; when they diverge, the divergence itself locates the problem. This is the most reliable route to causal knowledge outside a trial, and the hormone therapy episode is what happens when a field has only one design.
Because the causal verb does most of the work in public argument, and it is usually unearned.
Six questions, in order, for any claim that X affects Y.
Could Y cause X? Ask it every time; the answer is yes more often than people expect.
What would cause both? Name a specific candidate, not "other factors".
Was the sample selected in a way that could manufacture this?
What did they control for — and is anything in that list a mediator or a collider?
How well were the confounders measured? Badly measured controls leave residual confounding and the label "adjusted" conceals it.
And does any other design point the same way?
And a note on tone that this course has earned by now. The point of all this is not to be able to dismiss things. Anyone can say "correlation isn't causation" and stop thinking. The skill is to say which of the five explanations is live in this particular case, how large it could plausibly be, and what evidence would separate them.
That is the difference between scepticism and analysis , and it is the difference between a reader who cannot be fooled and a reader who cannot be persuaded of anything.
Two things move together for five reasons : A causes B; B causes A; something causes both; you selected on a common effect; or chance and shared trend. A causal claim earns its verb by arguing against the other four, and a large sample does not help — it makes a confounded estimate more precise, not less wrong.
The hormone therapy case is the canonical demonstration. Large, careful, prospective, well-adjusted cohort studies found a large protective effect on heart disease; a randomised trial found the opposite sign. The mechanism was healthy-user confounding, partially adjusted for, using imperfect measures. And honestly: the trial population was older, so the two literatures were also partly studying different people.
Correlation sees only straight lines, says nothing about direction, and its size is easy to misread — r = 0.3 accounts for 9 per cent of variance and can still matter enormously across a population.
Draw the causal diagram first. A confounder points into both cause and outcome and should be controlled. A mediator lies on the path and must not be — controlling for it deletes the effect. A collider is a common effect, and conditioning on it creates association that does not exist: the dating-pool example, Berkson's hospital patients, and any selected sample at all.
A regression coefficient is a conditional comparison, not an effect. Omitted-variable bias has a reasonable direction; measurement error in a confounder leaves residual confounding behind the word "adjusted" ; more controls is not better; and the table 2 fallacy — reading every coefficient in a model as its own causal estimate — is everywhere.
The drinking J-curve shows the same lesson from the comparison group: abstainers include sick quitters, and designs less vulnerable to lifestyle confounding do not support the protective effect.
What to do: state the diagram, reason about the direction of residual bias, quantify how strong an unmeasured confounder would have to be, and triangulate across designs with different weaknesses.
Correlation (Pearson r) — linear association, from −1 to +1; blind to curves and to direction of influence.
Reverse causation — the outcome influencing the supposed cause.
Confounder — a common cause of both exposure and outcome, opening a back-door path.
Causal diagram (DAG) — nodes and arrows stating what causes what, drawn before analysis.
Mediator — a variable on the causal path; controlling for it removes the effect.
Collider — a common effect of both; conditioning on it creates spurious association.
Berkson's paradox — associations manufactured by selection into a group, as among hospital patients.
Healthy user bias — those who take a treatment differ systematically in health-protective ways.
Sick-quitter effect — former users classified as never-users, contaminating the reference group.
Omitted variable bias — distortion from an unmeasured common cause, with a reasonable direction.
Residual confounding — what remains after adjusting for a badly measured confounder.
Table 2 fallacy — interpreting every coefficient in a model as a valid causal estimate.
Sensitivity analysis — quantifying how strong unmeasured confounding would need to be to explain a result away.
Triangulation — convergence across designs with different biases.
One — reverse every arrow. Take five claims of the form "X affects Y" from this week's news and write the reverse-causation version of each. Note how many are at least as plausible.
Two — name the confounder. For the same five, name one specific variable that would cause both. Not "other factors" — a named thing.
Three — draw a diagram. For a question you care about, draw the arrows: cause, outcome, three confounders, one mediator. Then write which of them you would control for and which you would not, and why.
Four — find a collider. Identify a group you belong to that people had to pass some filter to join. Then think of two qualities that would each independently help someone pass it. Those two qualities will appear negatively related inside the group and need not be related outside it.
Five — check the adjustment list. Take a published study reporting an adjusted association. Look at the list of controls and ask, for each: confounder, mediator, or collider? You will usually find at least one that is not a confounder.
Suppose the design is sound and the confounders are handled. There is still the question everyone gets wrong, including a great many researchers: what does it mean when a result is "statistically significant", and what does it not mean?
7.6.3 — Significance, Effect Size and What p Actually Means covers the definition of a p-value in plain words, the four things it is routinely mistaken for, why significance says nothing about importance, what statistical power is and why most social research does not have enough of it, and the analytic freedom that makes significant findings far easier to produce than they should be.