Most failed research did not fail at the data stage. It failed months earlier, in a sentence.
The question was too big, or it was not a question, or it had already been answered, or no evidence in the world could have settled it, or it quietly assumed the thing it claimed to be investigating. None of that is visible when you are excited about a topic. All of it is fatal.
This lesson is about the single most undertaught skill in the discipline: turning something you care about into something that can be found out.
Three people, three sentences, three different problems.
A supervisor meets three students in the same afternoon. Each has a topic they are genuinely gripped by. Each has written one sentence.
The first has written : "I want to study social media and mental health."
The supervisor asks: which platforms, whose mental health, over what period, compared to what, and measured how? The student has answers to none of these, and the honest reason is that the sentence is not a question. It is a pair of nouns. It marks out a territory. It does not tell anyone what would count as an answer.
The second has written : "Does poverty cause crime?"
Better — it has a verb, and it could be false. But the supervisor asks: what would surprise you? The student says nothing would; poor areas have more recorded crime and everybody knows it. The question is not wrong. It is closed. A question whose answer you already accept, and that a hundred studies have addressed, will produce a literature review with a shrug at the end. And the interesting thing — some poor places have very little crime, and one poor place a mile from another can differ by a factor of five — is sitting right there unasked.
The third has written : "Is inequality too high in this country?"
The supervisor asks what evidence would show that it was not. The student begins to answer and stops. There is none, because the word doing the work is "too". How much inequality there is, is a factual question with an answer. Whether that amount is too much depends on what you think a society owes its members, and no dataset settles it (see 7.1.2 on the warring gods).
Three failures : a topic mistaken for a question, a question with no puzzle in it, and a value judgement wearing a question's clothes.
And each is one afternoon's work away from being an excellent project.
What has to be true of a sentence before evidence can answer it?
That is the whole lesson, and the answer is a checklist of six things. A sentence that fails any one of them will consume months and produce nothing.
It has to be a question , not a territory.
It has to be empirical — settled by evidence, not by definition or by values.
It has to be open — a real possibility of coming out either way.
It has to be specified — about someone, somewhere, in some period, compared to something.
It has to be doable — with the access, time and money that actually exist.
And it has to matter — there has to be an answer to so what .
Six is a lot to hold, so the rest of the lesson takes them in the order they cause trouble.
From topic to puzzle to question
The missing step is the puzzle, and this is the heart of it.
A topic is a region of the world: social media, caste, migration, burnout, gentrification. Everyone starts here and there is nothing wrong with it.
A puzzle is a tension — a place where what you observe does not fit what you would expect. This is the step almost everyone skips , and it is the step that turns a territory into a project.
A puzzle has a shape, and there are only a few shapes.
The unexpected variation. Things that ought to be the same are not. Two neighbourhoods with the same poverty rate, the same age profile and the same housing stock have very different crime rates. Two states with the same income have very different literacy. Why?
The unexpected constancy. Things that ought to differ do not. The policy changed dramatically and the outcome did not move. The technology transformed the industry and the gender composition of its senior ranks is what it was in 1985. Why not?
The failed prediction. A well-established theory says X and the case shows not-X. Secularisation theory predicted religion would recede with modernity; in much of the world it did not (see 9.5). This is the most productive shape of all , because whatever you find bears on the theory.
The unnoticed ordinary. Something everyone does and nobody has explained. Why do people queue in some places and not others? Why is tipping obligatory here and insulting there? C. Wright Mills's advice was to make the familiar strange — a habit this course opened with (see 1.1.2).
The rate that changed. Something moved sharply at a datable moment. Why then? Timing is the friendliest evidence there is, because it rules out everything that did not also change then.
The account that cannot be right. The official explanation for something is available and does not survive arithmetic. Berger's debunking motif (see 1.5.2), converted into a research design.
Merton called this "the specification of ignorance" — the discipline advances not by listing what is unknown but by stating precisely what it is that we do not know and why it matters that we don't.
The second student's question, rebuilt in ten minutes.
"Does poverty cause crime?" → not open, and the answer is a shrug.
Insert a puzzle. Among neighbourhoods matched on poverty rate, unemployment and housing type, recorded violent crime varies by a factor of several. That is unexpected variation.
Now the question : Among neighbourhoods with similar economic disadvantage, what accounts for the differences in violent crime rates?
Now it is open — several candidate answers exist and they disagree. Collective efficacy, in Sampson, Raudenbush and Earls's sense: whether residents trust each other enough and expect each other to intervene. Residential stability, which gives people time to build those ties. Concentration and design of licensed premises. The nature of policing and, crucially, whether people are willing to report — which changes the measured rate without changing the actual one (see 7.4.4). Local organisations and what they do.
Now it is specifiable : which city, which years, which crimes, which measure of the candidate mechanisms, and what comparison.
Now it matters : if the answer is collective efficacy, the policy implication is about stability and organisation; if it is reporting, the "crime rate" everyone argues about is partly a measure of trust in the police rather than of crime.
Same topic. Same student. Same week. Entirely different project — and the change was inserting a puzzle.
The six requirements, and how to test each one
One — is it a question?
The test: does it end in a question mark and would two competent people know what would settle it? "Social media and mental health" fails. "Does time spent on image-based platforms predict later depressive symptoms among adolescent girls, net of prior symptoms?" passes.
Two — is it empirical?
The test: is there any observation that could settle it? Questions containing should , ought , too much , fair , deserve or legitimate are usually not empirical as posed.
But they can nearly always be converted, and the conversion is not a defeat. "Is inequality too high?" becomes a family of answerable questions: how high is it, by which measures, compared to when and where; what do people in this society believe it to be, and how wrong are they; what do they say it should be; what happens to health, trust, mobility and political participation as it changes; and what did specific interventions do. Answer those and you have not settled the value question — you have equipped the person who has to. That is the correct division of labour from 7.1.2.
Three — is it open?
The test: name the two or three answers you might get, and check that you would believe any of them. If only one is credible to you, you have a demonstration, not a study. Also check the other direction: if every answer is equally credible and nothing constrains them, the question may be too unformed to guide a design.
Four — is it specified?
A specified question names four things:
Who — the population and the unit of analysis (individuals? households? firms? neighbourhoods? countries? years? events?).
What — the outcome, precisely enough that someone else would measure the same thing.
Compared to what — the contrast that gives the answer meaning. "Is X high?" is unanswerable without a comparison ; higher than last decade, than the neighbouring state, than a matched group, than the rate before the law changed.
When and where — period and setting, because social findings are usually conditional on both.
Five — is it doable?
The test: write, in one sentence, the data you would need, then ask who has it and whether they would give it to you. A great many excellent questions require records that no one keeps, access no one will grant, or an experiment that would be unethical (see 7.8.1). Discovering this at the start is a success; discovering it in month nine is not.
Six — does it matter?
The "so what" test, run against three audiences. What changes in the literature if the answer is A rather than B? What changes in the world? What theory is put under strain? A question that survives the first five and fails this one produces the most demoralising kind of finished work : correct, competent, unarguable, and of interest to nobody, including its author.
Six ways a question sabotages itself, in roughly ascending order of subtlety.
The topic in disguise. "An investigation into X." "The role of Y in Z." No question, and therefore no possible failure — which sounds safe and is why so much of it exists.
The presupposition. "Why does poverty cause crime?" — the "why" has smuggled in the causal claim as settled and the study can now only elaborate it. Any question beginning "why does X cause Y" should be split : does it, and if so how. The same trap: "How has neoliberalism destroyed community?" and "Why are young people more anxious than previous generations?" Each contains an unexamined empirical claim in its premise , and once it is in the premise no finding can dislodge it.
The double-barrel. Two questions joined by "and", which will need different designs and will produce two half-studies.
The infinite question. "How does globalisation affect culture?" No population, no outcome, no comparison, no boundary. Not too ambitious — too unbounded to be ambitious about.
The data-led question. The dataset arrived first and a question was reverse-engineered to fit it. Sometimes unavoidable and often productive — but it changes what the finding can bear, because with enough variables some relationship will always be significant (see 7.6.3). The honest version says so: this is exploratory, and here is what would confirm it.
HARKing — Hypothesising After the Results are Known , and the difference between it and legitimate discovery is entirely a matter of disclosure. Finding something unexpected and pursuing it is how discovery works. Presenting it afterwards as though you had predicted it is misconduct , because it converts a hypothesis-generating result into a hypothesis-testing one and destroys the reader's ability to judge the odds (see 7.8.2).
Five kinds of question, and what each obliges you to do.
The kind of question you ask determines the design; not the other way round.
Descriptive — how many, how much, how distributed, what changed? Obliges you to worry about who is counted and who is missing . Underrated: a large share of the most valuable social research is careful description, and it is dismissed as "merely descriptive" by people whose causal claims rest on nobody having done it (see 7.6.1).
Relational — does X go with Y, and how strongly? Obliges you to worry about confounding, and to be disciplined about not sliding into causal language.
Causal — does X make Y happen; what would have happened otherwise? Obliges you to construct a credible counterfactual, which is the hardest thing in social science and the subject of most of Topic 7.4.
Mechanistic — how does X produce Y; through what steps? Obliges you to get inside the process, usually with qualitative or mixed work, and to specify what should be observable at each step if the mechanism is real.
Interpretive — what does this mean to the people doing it; how do they make sense of it? Obliges you to establish that you have the participants' frame and not your own (see 7.1.1, and 7.7.2 on validity).
And a sixth, which is really a family : evaluative — did this intervention work, for whom, under what conditions? It is causal plus mechanistic plus normative-adjacent, which is why evaluation is so often contested (see 12.2).
The commonest single error in the discipline is a design built for one kind and a conclusion written for another: relational evidence, causal conclusion. Watch for it in everything you read, including your own drafts.
One exercise that catches almost everything, before you spend anything.
Before collecting a single piece of data, write two sentences:
If the answer is yes, I expect to observe ________.
If the answer is no, I expect to observe ________.
If you cannot fill both blanks with different content, stop. Either the question is not empirical, or the design cannot bear on it, or you have already decided.
This is Popper's falsifiability made practical (see 7.1.1) — the requirement that a claim forbid something. And it is the cheapest quality check in existence: it costs five minutes, and it is routinely skipped, which is why so many studies conclude that their expectations were confirmed by evidence that could not have disconfirmed them.
Two refinements make it stronger.
Add a magnitude : not just "a difference" but how big a difference would count as supporting the claim, and how small would count against it . Decide before you look.
Add the rival : name the most plausible competing explanation, and write what you would observe if it were true instead. A design that cannot distinguish your explanation from the obvious rival has not tested anything — it has only shown that your explanation is compatible with the data, which weak explanations also are.
Questions change during research, and that is legitimate. Here is the line.
Nobody's final question is their first one. You learn that the outcome cannot be measured, that the interesting variation is somewhere else, that the participants think you are asking about the wrong thing entirely — and the question moves. Qualitative traditions build this in explicitly : grounded theory treats question refinement as the method (see 7.7.1), and an ethnographer who returns with the question they left with was probably not listening.
The line is not "did the question change". It is "does the reader know".
Legitimate : We began asking A. Early fieldwork showed B was the operative question, for these reasons. Here is what we then did. The reader can now judge everything correctly.
Not legitimate : presenting the final question as the original one, and the exploratory findings as confirmations of a prediction nobody made. Same sequence of events; different account of it; and the difference is the whole of the reader's ability to weigh the result.
Because the question determines what the study is capable of, and no amount of later rigour repairs it.
A brilliant analysis of the wrong question is worth less than a rough analysis of the right one. Technique cannot rescue a question , and this is the asymmetry people learn too late: it is always possible to add a robustness check, and never possible to add a comparison you did not build in.
And for a reader, this is the fastest diagnostic there is. Before you evaluate a study's methods, find its question — the real one, which is often not the one in the title. Then ask: is it open, is it specified, is it the kind of question this design can answer? A great deal of published work fails at this level , and you can see it without knowing any statistics at all.
The final thing, which matters for anyone who wants to do this rather than only read it. Mills's advice in "On Intellectual Craftsmanship" was to keep a file — notes, fragments, overheard things, figures that surprised you, personal experiences treated as data — and to rearrange it periodically, deliberately, looking for juxtapositions. Questions do not arrive. They are assembled , out of things noticed over months by someone in the habit of noticing.
A topic is a territory; a question is answerable; and the step between them is a puzzle — a tension between what you observe and what you would expect. The productive shapes: unexpected variation, unexpected constancy, a failed prediction, the unnoticed ordinary, a rate that changed, an official account that does not survive arithmetic. Merton: the specification of ignorance.
Six requirements. It must be a question; empirical (settled by evidence, not by values — and value questions convert into families of answerable ones); open (you could believe more than one answer); specified (who, what, compared to what, when and where); doable with the access that exists; and it must survive so what against the literature, the world and the theory.
Five kinds of question oblige different things : descriptive (who is missing from the count), relational (confounding), causal (a credible counterfactual), mechanistic (the observable steps), interpretive (the participants' frame) — plus evaluation, which combines them. The commonest error in the discipline is relational evidence with a causal conclusion.
The traps : the topic in disguise, the presupposition ("why does X cause Y"), the double-barrel, the unbounded question, the data-led question, and HARKing — which differs from legitimate discovery only by disclosure.
And the cheapest quality check ever devised : write what you would observe if the answer is yes and what you would observe if it is no, with a magnitude and a named rival explanation. If both blanks cannot be filled differently, nothing has been asked.
Topic / puzzle / question — a territory; a tension between observation and expectation; a sentence evidence can settle.
Specification of ignorance — Merton's term for stating precisely what is not known and why it matters.
Unit of analysis — the kind of thing each case in your data is: a person, a household, a firm, a neighbourhood, a year, an event.
Comparison / contrast case — the "compared to what" without which magnitude claims are meaningless.
Open question — one where more than one answer is credible to the researcher in advance.
Answer-shape test — writing in advance what would be observed under each answer, with a magnitude and a named rival.
Presupposition — an empirical claim hidden in the premise of a question, where no finding can reach it.
HARKing — hypothesising after the results are known and presenting it as prediction.
Descriptive / relational / causal / mechanistic / interpretive / evaluative questions — the six families, each obliging a different design.
The "so what" test — what changes in the literature, in the world, and for theory, depending on the answer.
One — do the conversion. Take a topic you care about and write it as a territory. Then find a puzzle in it using one of the six shapes. Then write the question. Keep all three lines ; comparing them is the lesson.
Two — run the answer-shape test on someone else. Find any study and write the two sentences for it: what would have been observed if its claim were false? If you cannot construct that observation from what they did, you have found the study's real limitation.
Three — hunt presuppositions. Take five headlines or article titles containing "why". Rewrite each as the two questions it has fused: does it, and if so how. Note how often the first is unestablished.
Four — convert a value question. Take something you have strong views about — "is X exploitative", "is Y unfair" — and write four empirical questions that would inform someone's judgement without settling it. This is the 7.1.2 division of labour, done by hand.
Five — start the file. Keep a single document for four weeks. Put in it anything that surprised you, any number you did not expect, any moment when someone's explanation for their own behaviour did not fit what they did. Read it at the end of the month looking only for tensions. That is where questions come from.
You have a question. It contains words like trust , inequality , radicalisation , wellbeing , precarity — words everyone uses and no two people mean identically.
Before anything can be measured, those words have to be turned into something specific enough that another researcher could produce the same thing. That translation is where a great deal of research quietly goes wrong, because it looks technical and it is actually the point at which you decide what the thing is (see 7.1.3).
7.2.2 — Concepts, Variables and Operationalisation is about doing it honestly, and about spotting when it has been done badly.