There is a difference between a theory that survived a test and a theory that was decorated with evidence, and it is not visible in the results table. Both look like confirmation.
The difference is in what the theory risked — what result would have counted against it, and how likely that result was if the theory were wrong.
This lesson is about that difference, in both directions: testing theories you already have, and building theories you do not yet have. They are not rivals. They are two halves of a cycle , and doing either while pretending to do the other is where most of the trouble lives.
A prophecy, a deadline, and a prediction that ran backwards.
In 1954 a small group formed around a woman in the American Midwest who reported receiving messages from beings on another world. The messages named a date: before dawn on 21 December, a great flood would destroy much of the continent , and the faithful would be collected by a spacecraft at midnight.
Members had acted on it. Some had left jobs. Some had given away possessions. One had left a marriage.
Three social scientists — Leon Festinger, Henry Riecken and Stanley Schachter — joined the group as observers before the date , and they went in with a prediction that sounds, on first hearing, absurd.
Common sense says that when a prophecy fails publicly and completely, believers will be humiliated and the group will dissolve.
Their theory predicted the opposite. A person who holds a belief with conviction, has taken irrevocable action on it, and then meets unambiguous disconfirmation, is in a state of unbearable inconsistency — cognitive dissonance . They cannot undo the action. They cannot easily unbelieve. But if more people can be brought to believe, the belief becomes more bearable. So the predicted response to disconfirmation is increased proselytising — a group that had been secretive and uninterested in recruitment becoming, at the moment of maximum embarrassment, evangelical.
Festinger specified the conditions under which this should occur: conviction, commitment through irrevocable action, a belief specific enough to be clearly disconfirmable, and social support — other believers present at the moment of failure.
Midnight passed. No spacecraft. Dawn came. No flood.
At about 4.45 a.m. a new message arrived: the group's faith had been so strong that God had spared the world. The group, which had refused press calls for weeks, began telephoning newspapers.
Now notice what makes this a real test rather than a good story. The prediction was counterintuitive — the opposite of what common sense expects. It was specific — this behaviour, in this window, under these stated conditions. And it was disconfirmable in an obvious way : if the group had quietly dispersed in shame, as everybody would have guessed, the theory would have been in serious trouble.
And it must be graded honestly, as everything in this course is. The study has real defects. The group was tiny. The observers made up a substantial fraction of those present and may themselves have propped it up — one was taken for a genuine believer and treated as confirmation. The increase in proselytising was modest and short-lived. Later work on failed prophecies finds the response is far from uniform.
Which is exactly the right note. The prediction was risky and it is that riskiness that gives the finding its weight — and the execution had weaknesses that limit how much weight it can carry. Both things, at once, is what reading research honestly consists of.
Any theory can be made to fit a finding after the fact. So what makes a test a test?
Consider a claim: organisations under pressure become more rule-bound.
Suppose we observe an organisation under pressure that becomes more rule-bound. Confirmation. Suppose instead it becomes chaotic and improvisational. A defender says: pressure exceeded the threshold at which formalisation is possible — the theory still holds. Suppose it does not change. A defender says: it had already formalised.
A claim that accommodates every outcome has told you nothing about the world , however true it feels. It is unfalsifiable in Popper's sense (see 7.1.1), and its comfort is exactly the problem.
The remedy is not to demand certainty. It is to demand risk. A hypothesis earns its keep by ruling something out in advance — by naming an observation that, if it occurred, the theory could not survive.
And the sharpest formulation of that demand comes from Deborah Mayo: a claim is supported by evidence only if it has passed a test it would probably have failed , had it been false. Read that twice. It is the whole of good testing in one sentence, and it explains why so much apparent confirmation is worth so little: the test was one the claim would have passed either way.
What a hypothesis is, and what it must contain
A hypothesis is a statement about the world, derived from a theory, that could turn out to be false.
It is not a guess, and it is not a research question. The question asks; the hypothesis commits.
A usable hypothesis contains four things.
A relation — between specified variables, defined as in 7.2.2. Not "social media affects wellbeing" but a stated relation between a stated measure of one and a stated measure of the other.
A direction — which way. "There is a relationship" is barely a claim.
A population and conditions — for whom, where, when. A hypothesis without a scope condition is either false somewhere or untestable.
Ideally, a magnitude — how big, or at least how big it must be to count as support. This is the most commonly missing element and it does more damage than any other , because without it any non-zero result in the predicted direction is claimed as confirmation.
And there is a fifth thing, which lives outside the sentence: the theory it came from. A prediction that does not follow from anything is a hunch. If the theory is wrong, the prediction should fail — that connection is what makes the test informative about the theory rather than only about the data.
Meehl's paradox, which ought to be taught in every first methods class and almost never is.
Paul Meehl noticed something disturbing about the soft sciences in 1967, and forty years of replication problems have made it look prophetic.
In physics, better measurement makes theories harder to confirm. Theories make point predictions — this quantity, this value — and greater precision means more ways to be wrong.
In the soft sciences, better measurement makes theories easier to confirm. Because the typical prediction is only directional — group A will score higher than group B — and the thing it is tested against is the null hypothesis of exactly zero difference , which in social data is almost never true. Everything is faintly correlated with everything: what Meehl called the crud factor . So with a large enough sample, almost any directional prediction will reach statistical significance.
Which means the prior probability of "confirming" a directional hypothesis is roughly one in two, before the theory has contributed anything. A theory that survives such a test has survived almost nothing.
The remedies all point the same way. Predict a magnitude, not just a direction. Predict a pattern — this order across five groups, this effect present here and absent there. Predict something the obvious rival explanation does not also predict. Design tests your theory could plausibly fail.
And use this as a reader. When a paper reports that its hypotheses were supported, ask: what was the chance of that, if the theory were false? If the answer is "about half", you have learned very little, no matter how small the p-value.
The null hypothesis, and its strange backwards logic.
Because it is impossible to prove a universal claim by observation, the standard testing framework works by elimination.
The null hypothesis (H₀) is the statement of no relation, no difference, no effect. The alternative (H₁) is your claim. You then ask: if H₀ were true, how likely is data as extreme as what I observed? If that is sufficiently unlikely, you reject the null .
Three consequences follow, and all three are routinely mishandled.
You never prove H₁. You fail to find the data compatible with H₀. That is weaker, and it is the correct description of almost every result you have ever read.
Failing to reject H₀ is not evidence of no effect. It may mean there is no effect; it may equally mean the sample was too small, the measure too noisy, or the design too weak to detect one. "No significant difference" is not "no difference" — the distinction between absence of evidence and evidence of absence, which needs a different analysis to establish (see 7.6.3).
And the whole framework tests something nobody believes. Exact zero is a straw man in social data, which is Meehl's point arriving again from another direction.
Two directions: testing and building
Deduction and induction, and why the argument about which is superior is a waste of time.
Deductive research (theory-testing) runs downwards. Start from theory. Derive a hypothesis. Design a study that could disconfirm it. Collect data. Test. Revise the theory. This is the hypothetico-deductive model.
Inductive research (theory-building) runs upwards. Start from observation. Look for patterns. Build concepts and propositions that account for them. Produce a theory that can then be tested — usually on different data, by someone else.
Walter Wallace's "wheel of science" is the right image : theory → hypotheses → observations → empirical generalisations → theory. You can enter the circle at any point. Nobody goes round it alone , and a discipline that only tested would run out of things to test, while a discipline that only built would accumulate ungraded theories forever.
And there is a third movement , from 7.1.1, which is what researchers actually do most of the time: abduction — encountering something surprising and asking what would have to be true for it to make sense. That is how the interesting hypotheses arrive. It is not a form of proof and does not pretend to be; it is the engine of discovery, and its products then have to face a test.
How theory gets built, properly.
Theory-building is not "having no hypotheses". It has its own discipline, and there are four main routes.
Grounded theory — Glaser and Strauss's programme: systematic coding of data, concepts developed from it, comparison between cases, and theoretical sampling — choosing the next case because of what the emerging theory needs, not by a sampling frame decided in advance. Continue until new cases stop changing the categories (see 7.7.1).
Analytic induction — the older and more demanding method associated with Znaniecki. Formulate an explanation from a few cases, then actively hunt for a case that contradicts it , and on finding one, either reformulate the explanation or redefine the phenomenon so that the case falls outside it. Repeat until you cannot find a contradiction. It has an obvious weakness — redefinition can be used to make the theory true by construction — and an outstanding strength: negative cases are sought rather than avoided.
Typology construction — identifying the kinds of a thing and what distinguishes them. Weber's ideal types (see 4.3.8); the welfare-regime typologies; Merton's typology of adaptations to strain. A good typology is a theory in compressed form , because it claims that the dimensions it uses are the ones that matter.
The extended case method — Michael Burawoy's programme, and the most explicitly ambitious. Take a case not as a sample of a population but as a site for reconstructing an existing theory . You go in with the theory, you find where it fails, and you use the failure to improve it. The case is chosen precisely because it should be awkward for the theory — which makes it a test and a building exercise at once.
Duhem, Quine, and the reason a failed prediction never straightforwardly kills a theory.
Popper's picture — theory forbids something, the something happens, theory dies — is too clean, and the reason has a name.
You never test a hypothesis on its own. You test it bundled with a crowd of auxiliary assumptions: that the measure is valid, that the sample is adequate, that the manipulation worked, that the setting is one where the theory should apply, that the analysis is right. When the prediction fails, logic tells you that something in the bundle is wrong — and not which thing. This is the Duhem–Quine problem.
So a defender of any theory can always relocate the blame: the measure was crude, the sample unusual, the context inappropriate, the dose too low. Every one of those defences is sometimes correct , which is why this is a real difficulty and not just bad faith.
Imre Lakatos gave the best available answer. Theories come in research programmes with a hard core of central claims that adherents will not abandon, and a protective belt of auxiliary assumptions that get adjusted when trouble comes. Adjusting the belt is legitimate — it is how science works.
The test is what the adjustment does next. A programme is progressive if its modifications predict novel facts that are then found. It is degenerating if its modifications only accommodate the anomalies that prompted them, absorbing each new problem and predicting nothing.
This is a usable criterion and you should carry it. Ask of any body of theory: in the last twenty years, has it predicted anything new that turned out to be true — or has it only explained, after the event, whatever happened? Some very prestigious traditions do badly on that question , and applying it evenly is the symmetry rule from 6.8.1 in operation.
Merton's middle range, and why it is the practical answer.
Grand theory — a general account of society as such — is nearly impossible to test, because it does not forbid enough. Raw empiricism produces findings that connect to nothing.
Merton's middle-range theories sit between (see 5.1.2): accounts limited in scope, addressing a specific class of phenomena, from which testable propositions follow. Reference group theory. The self-fulfilling prophecy. Strain theory. Opportunity structures.
Each is small enough to be wrong in a specifiable way, and general enough to travel between settings. That combination is what makes cumulative research possible, and it is the standard against which the theories in Part 5 look very different from one another — a point worth revisiting now that you have the criterion.
Five ways hypothesis-testing goes wrong in practice.
One — the vague hypothesis. "There will be a relationship between X and Y." Direction unstated, magnitude unstated, scope unstated. Any outcome confirms it, and it will be reported as supported.
Two — HARKing. Presenting a post-hoc finding as a prediction (see 7.2.1). It looks like a strong confirmation and is in fact an exploratory result with its uncertainty removed.
Three — the unfalsifiable defence. Every disconfirmation absorbed by an adjustment that predicts nothing new. Lakatos's degenerating programme, and it can persist for decades because it always has an answer.
Four — confirmation-seeking design. Looking only where the effect would be, choosing the comparison that flatters, dropping the site where it did not work. Not fraud; just never asking the question the other way round.
Five — accepting the null. "We found no significant difference, therefore the intervention does not work." With eighty participants and a noisy measure, this claim has no basis. The right report is: this study could not detect an effect of the size we looked for, and here is what size that was.
Because it gives you one question that cuts through any paper, any argument, any confident explanation.
"What would you accept as evidence that you are wrong?"
Ask it of a study: what result would the authors have reported as disconfirming? If the design could not have produced one, the confirmation is empty.
Ask it of a theory: what has it predicted, in advance, that was not already known? If the honest answer is nothing in twenty years, you are looking at a degenerating programme, whatever its prestige.
Ask it of yourself, about the things you believe about society. This is the hardest version, and it is the one this whole Part is for. Most people, asked what would change their mind about a strongly held social belief, discover they cannot answer — and the discovery is more valuable than any finding in this lesson.
And the constructive half matters just as much. Not every study should test a hypothesis. Some of the most important work in the discipline's history built the concepts everyone else later tested — Goffman had no hypotheses; neither did Du Bois in Philadelphia. What separates good theory-building from storytelling is discipline : systematic comparison, deliberately sought negative cases, categories that could have come out differently, and honesty about which it was.
The failure mode is not doing one or the other. It is doing one and claiming the other's authority.
A theory is supported only if it passed a test it would probably have failed had it been false — Mayo's severity requirement, and the standard everything else in this lesson serves.
A usable hypothesis states a relation, a direction, a population and scope, and ideally a magnitude — and follows from a theory, so that its failure would be informative about the theory.
Meehl's paradox : because everything in social data is faintly correlated with everything (the crud factor) and the null of exactly zero is nearly always false, a merely directional prediction has about a one-in-two chance of "confirmation" before the theory contributes anything. Predict magnitudes, patterns, and things the obvious rival does not predict.
Null-hypothesis logic runs backwards : you never prove your claim, only fail to find the data compatible with no-effect; and failing to reject is not evidence of absence.
Two directions, one wheel. Deductive testing and inductive building are halves of a cycle, with abduction — inference to the best explanation of something surprising — as the engine of discovery. Building has its own discipline : grounded theory with theoretical sampling; analytic induction, which hunts negative cases; typology construction; and Burawoy's extended case method, which selects the case that should be awkward for the theory.
Duhem–Quine : hypotheses are never tested alone, so a failed prediction never says which assumption failed. Lakatos's answer — a research programme is progressive if its adjustments predict novel facts, degenerating if they only absorb anomalies. Ask of any tradition: what has it predicted in advance that turned out to be true?
And Merton's middle range is the practical zone : small enough to be wrong in a specifiable way, general enough to travel.
Hypothesis — a derived, directional, scoped, potentially false statement about a relation between measured things.
Severity (Mayo) — evidence supports a claim only if the claim passed a test it would probably have failed if false.
Null hypothesis (H₀) / alternative (H₁) — the statement of no effect against which the test is run; and the claim of interest.
Crud factor / Meehl's paradox — in social data everything correlates weakly with everything, so directional predictions are confirmed at roughly chance rates before any theory contributes.
Hypothetico-deductive model — theory → hypothesis → test → revision.
Wheel of science — Wallace's cycle: theory, hypotheses, observations, empirical generalisations, theory.
Abduction — inference from a surprising observation to the best available explanation.
Grounded theory / theoretical sampling — building theory from systematically coded data, choosing later cases by what the emerging theory requires.
Analytic induction — reformulating an explanation until no contradicting case can be found; negative cases actively sought.
Extended case method — Burawoy: using a deliberately awkward case to reconstruct an existing theory.
Duhem–Quine problem — hypotheses are tested only in bundles with auxiliary assumptions, so failure does not identify the culprit.
Research programme, hard core, protective belt — Lakatos's anatomy of a theoretical tradition.
Progressive / degenerating programme — modifications that predict novel facts, versus modifications that only accommodate anomalies.
Middle-range theory — Merton: limited in scope, generating testable propositions, able to travel.
One — sharpen a vague hypothesis. Take any claim you have read this week about why people behave a certain way. Rewrite it with a direction, a population, a scope condition and a magnitude. Note how much more it now forbids.
Two — apply severity. For that sharpened hypothesis, describe a study in which it would probably fail if it were false. If you cannot design one, the claim is not yet testable.
Three — grade a research programme. Take a theoretical tradition from Part 5 or Part 6 and ask Lakatos's question: what has it predicted in advance that was subsequently found? Do this for one you like and one you dislike, and compare how hard you looked in each case.
Four — find a degenerating defence. Identify an argument — political, professional, personal — that has absorbed every piece of contrary evidence with an adjustment. Write down what it now forbids. Often the answer is nothing.
Five — turn it on yourself. Write down one strongly held belief you have about society, and then the specific observation that would make you abandon it. If you cannot complete the second sentence, you have found something worth knowing about the first.
Design settled, you now face the question every study answers whether or not it admits it: who, exactly, are you looking at, and what do they stand for?
A study of 1,200 people is used to describe a nation of 200 million. That should be astonishing, and it works — under conditions that are specific, well understood, and frequently not met.
7.3.1 — Populations, Samples and Why Sampling Works explains why a small sample can speak for a huge population, what the machinery actually requires, and why the size of the sample matters far less than almost everyone thinks.