The survey is the instrument that produces most of what the public believes it knows about society. Unemployment, trust, religiosity, attitudes to migration, satisfaction with the health service, how many people voted and why — nearly all of it comes from someone asking a question.
And a survey question is not a window. It is a stimulus. It is a piece of language delivered to a person in a social situation, and the answer is produced by all of that together — the words, the order, the alternatives offered, the person asking, and the respondent's guess about what is wanted.
Which means, more often than anyone is comfortable with, the wording is the finding.
The same question, two words apart, two different countries.
In 1940 the researcher Donald Rugg ran a split-ballot experiment: two random halves of a sample, two versions of one question.
Version A : Do you think the United States should allow public speeches against democracy?
Version B : Do you think the United States should forbid public speeches against democracy?
Logically these are the same question. To not allow is to forbid . Anyone answering consistently should refuse to allow in the same proportion that they agree to forbid.
They did not. Only about a fifth said the speeches should be allowed — implying roughly four-fifths opposed. But only about two-fifths said they should be forbidden — implying a majority against prohibition.
A twenty-point swing, produced by one word. The finding has been replicated repeatedly since, in different decades and different countries. Something about the word forbid is harsher; agreeing to it feels like an act rather than a preference; and a substantial slice of the population sits in the gap, unwilling to permit and unwilling to prohibit.
A second, larger one, from the American General Social Survey. Two versions of a spending question:
Are we spending too much, too little, or about the right amount on welfare ?
Are we spending too much, too little, or about the right amount on assistance to the poor ?
The gap is enormous and it has persisted for decades. Only a small minority say too little is spent on "welfare"; a large majority say too little is spent on "assistance to the poor". Same budget. Same respondents. Different word.
Now consider what this means for the phrase "the public thinks". A newspaper reporting either result would write a confident sentence about public opinion. Both sentences would be true of the question asked. They cannot both be true of the public.
If a word can move the answer twenty points, was there an opinion there at all?
This is not a technical complaint about sloppy questionnaires. It is a substantive question about what an attitude is , and two answers have shaped the field.
Philip Converse's answer, in 1964, was blunt. Studying panel data where the same people were asked the same policy questions at intervals, he found the responses of a large part of the public were remarkably unstable — many people's answers over time looked close to what you would get from a coin. He proposed that on many issues there is no attitude to measure: people confronted with a question they have never thought about generate an answer on the spot , because refusing to answer is socially awkward. He called these non-attitudes .
John Zaller's answer, in 1992, is subtler and is now the standard one. People are not empty and they are not fixed. They carry a stock of considerations — half-formed thoughts, images, values, things heard — that bear on an issue, often pointing in opposite directions. When asked, they do not consult a stored opinion. They sample from whatever considerations are accessible at that moment , and average them into an answer.
So the question wording, the preceding questions, the news that week, and the interviewer's presence all change which considerations come to mind — and therefore change the answer, without anybody being inconsistent or dishonest.
This reframes everything that follows. Wording effects are not noise contaminating a signal. They are evidence about how opinion actually works — which is a genuinely sociological finding, and a more interesting one than the number that gets reported.
What the survey interview really is
A conversation in which one party is forbidden to behave conversationally.
Ordinary conversation runs on cooperation. We assume our partner is being relevant and informative, we ask when we do not understand, we adjust to each other, we treat what was just said as context for what comes next.
The standardised survey interview suspends most of this by design. The interviewer must read the question exactly as printed, must not explain what a word means, must not react to the content, and must record the answer in one of the boxes. This is done for a good reason — if every interviewer improvises, differences between respondents are confounded with differences between interviewers — and it has a cost.
Suchman and Jordan's analysis of recorded survey interviews showed what the cost looks like: respondents trying to work out what a question means, offering qualified answers ("well, it depends on whether you count..."), and interviewers pushing them into a code. The uncertainty is real information, and the instrument is built to discard it.
Two consequences follow that you can see in almost any survey.
Respondents infer meaning from the instrument itself. Since they cannot ask, they use whatever cues are available — the response options, the previous questions, the sponsor's name — to work out what is being asked about. The questionnaire is not just measuring; it is communicating.
And respondents try to be good participants. They answer questions they have no view on, because that is what the situation asks of them.
The wording rules, and what each one prevents
Ten rules. Each exists because of a documented failure.
One — one idea per question. "Should the government increase spending on schools and hospitals?" Someone who wants one and not the other has no answer. Double-barrelled questions are the commonest defect in amateur questionnaires , and the and is the tell.
Two — no leading. "Do you agree that the council has failed to..." Prefaces that supply a position get it endorsed.
Three — watch loaded words. Welfare, elite, regime, reform, natural, illegal, extremist. Each carries an evaluation. You cannot remove all connotation; you can notice which one you chose , and where it matters, run both.
Four — avoid negatives, and never use two. "Do you disagree that teachers should not be required to..." Every double negative produces a measurable share of respondents answering the opposite of what they mean.
Five — define terms and specify the reference. "Regularly" means once a week to one person and once a month to another. "Your household" — does the lodger count? "Income" — before or after tax, yours or the household's?
Six — specify a time frame, and keep it short. "In the last twelve months" invites telescoping (see 7.3.2). "In the last four weeks" is remembered better.
Seven — balance the alternatives. "Do you favour X?" pulls more agreement than "Do you favour X, or do you oppose it?" — because the second reminds the respondent that opposing is a normal thing to do.
Eight — no presuppositions. "How often do you argue with your partner?" presupposes both a partner and arguments. Use a filter question first.
Nine — match the respondent's vocabulary, not the researcher's. Terms of art — precarious employment , food insecurity , social capital — do not survive contact with the public, and respondents will guess rather than admit not knowing.
Ten — do not require impossible arithmetic. "How many times in the past year did you..." for a frequent behaviour. People do not count; they estimate from a rate, badly. Ask about a short recent period instead.
Open and closed questions do not measure the same thing, and the difference is bigger than it looks.
Ask "What is the most important problem facing the country?" as an open question and let people answer freely: you get a wide spread, a long tail, and a fair number of concerns nobody expected.
Ask it as a closed question with eight options: the answers concentrate heavily on the list. Concerns not offered are largely lost — mentioned by a handful under "other", if that option is even provided.
The list is not a neutral container. It is a statement about what the answers are , and respondents take it as such.
So: open questions are better for discovering what is in people's heads; closed questions are better for comparing across people and over time. The right procedure is to use open questions in development, discover the real range, and then build the closed instrument from what was found — which is why questionnaire construction properly begins with qualitative work.
Order, context, and the scale that answers the question for you
Four effects that operate below anyone's awareness, including the researcher's.
One — question order changes the meaning of later questions.
The classic demonstration: students asked "How happy are you with your life as a whole?" and "How many dates did you have last month?" When the general question came first, the correlation between the two was weak. When the dating question came first, the correlation rose dramatically — because it made dating life salient, and respondents used it to construct their overall answer.
Nothing about the students' lives differed. The question order told them what "life as a whole" meant.
Two — reciprocity and consistency effects.
In 1950 Hyman and Sheatsley found that Americans were far more willing to say that communist reporters should be allowed into the United States if they had first been asked whether American reporters should be allowed into the Soviet Union. Having endorsed the general principle, respondents applied it — a consistency effect. The reverse order produced far less support.
Three — the response scale is itself information.
Schwarz and colleagues asked how much television people watched, with two versions of the answer scale. One ran in low bands (up to half an hour, half to one hour... more than two and a half hours). The other ran in high bands (up to two and a half hours ... more than four and a half).
With the high-frequency scale, roughly twice as many respondents reported heavy viewing. And there is a further twist: respondents given the high scale also rated television as less important in their lives — because the scale told them their own viewing was around average, and they adjusted their self-assessment accordingly.
The scale had communicated a norm. Respondents inferred that the middle of the range was typical, located themselves relative to it, and reported accordingly.
Four — the order of response options.
When options are read aloud , later options are favoured (recency ). When options are read on a page or screen , earlier ones are favoured (primacy ). The standard remedy is to rotate the order randomly across respondents — and a survey that has not done so has an unmeasured bias in every list-based question.
Acquiescence and satisficing: two ways respondents give you an answer that is not their answer.
Acquiescence — "yea-saying". Some respondents agree with whatever is proposed, and the tendency is stronger among those with less education, in interviewer-administered modes, and where the respondent perceives a status difference. A battery of items all worded in the same direction will therefore produce a spuriously coherent-looking attitude.
The standard remedy is to reverse the wording of some items so that agreement means the opposite. It has its own cost — reversed items are cognitively harder, some respondents miss the reversal, and the "reverse-worded factor" that appears in many scale analyses is often an artefact of exactly this.
Satisficing. Krosnick's account: answering a survey question properly requires interpreting it, retrieving relevant information, integrating it and mapping it onto the response options. That is work, and respondents who are tired, unmotivated or under-equipped take shortcuts.
What satisficing looks like in data : choosing the first acceptable option rather than the best one; straightlining — picking the same point down a whole grid; failing to differentiate between items; endorsing the status quo; agreeing regardless; and selecting "don't know" as an exit rather than as a report.
And the design conditions that produce it are entirely under the researcher's control : long questionnaires, long grids, complex questions, low motivation, late position in the instrument. A survey's last section is measurably worse data than its first.
Modes, and the trade-offs between them.
Face-to-face. Best response rates, best for long and complex instruments, shows cards and visual aids, allows observation of the setting. Enormously expensive , and the interviewer's presence suppresses sensitive disclosures.
Telephone. Cheaper, fast, once excellent — now with response rates in the single digits in many countries, and coverage problems as landlines vanish (see 7.3.2).
Postal. Cheap, no interviewer effects, good for sensitive items. Slow, low response, no control over who in the household fills it in or in what order they read it.
Web. Cheapest, fast, allows randomisation and complex routing automatically, best for sensitive topics because no human is present. Coverage and self-selection are the problem — panels are recruited, and the recruited differ.
Mixed-mode is now standard practice, and it introduces a difficulty that is easy to miss: if some respondents answer online and some by phone, differences between them are partly mode effects. And when a long-running survey changes mode, a genuine trend break appears in the series that has nothing to do with the world.
What surveys measure, and the gap they cannot see.
In 1934 Richard LaPiere published a study that has been argued about ever since. He travelled across the United States with a young Chinese couple, stopping at some 250 hotels, restaurants and other establishments. They were refused service once.
Six months later he sent questionnaires to the same establishments asking whether they would "accept members of the Chinese race as guests". Of those who replied, over 90 per cent said they would not.
The apparent finding: what people say and what people do are two different things , and surveys measure the first.
And the study has serious defects, which must be stated. Only about half responded, and non-response was not random. The person who answered the questionnaire was often not the person who had served them. The couple were young, well-dressed, fluent in English and accompanied by a white American professor — the abstract "member of the Chinese race" in the questionnaire is not the specific person who appeared at the desk. And six months is a long gap.
So the study does not show that attitudes are irrelevant to behaviour. What it does show, and what the far more careful literature since has confirmed, is that the relationship between a stated general attitude and a specific act in a specific situation is weak — and gets much stronger when the attitude is measured at the same level of specificity as the behaviour.
The rule to keep : a survey measures what people report, in the situation of being surveyed. For facts they know and will state — age, employment, housing, voting — that is close to what you want. For behaviour that is embarrassing, forgotten or socially loaded, and for predictions of what people would do, it is a different quantity, and the difference is systematic rather than random.
Three procedures that separate professional survey research from questionnaires.
Cognitive interviewing. Sit with a small number of people, have them answer while thinking aloud , and probe: what did that question mean to you? how did you arrive at that? what were you including? This finds the misunderstandings that no amount of expert review will , because researchers cannot un-know what they meant.
Piloting. Run the whole instrument on a small sample under real conditions. It reveals routing errors, timing, fatigue points, and the questions where everyone gives the same answer — which measure nothing.
Split-ballot experiments. Randomly assign versions of a question and compare. This is the only way to know how large a wording effect is , and the best survey organisations build them in permanently, which is how everything in this lesson is known.
Because a poll result without its wording is not information.
Four habits, and they cost nothing.
One — always find the exact question. Reputable organisations publish the full wording. If it is not published, treat the result as an advertisement , because the most common reason for withholding wording is that the wording did the work.
Two — check what preceded it. A question about immigration asked after four questions about crime is a different question.
Three — never compare across differently worded surveys. "Support has risen from 42 to 51 per cent" is meaningless if the two figures come from different houses asking different questions. Compare within a series, from one organisation, with stable wording — and check whether the mode changed.
Four — remember the margin of error covers none of this (see 7.3.1). The ±3 describes sampling variability alone. The wording effect in the story above was twenty points.
And for anyone building a survey: the questionnaire is where your ontology is enacted (see 7.1.3). Every option you offer is a claim that these are the kinds of answer there are. Respondents will take you at your word — because in the survey situation, they have nothing else to go on.
A survey question is a stimulus, not a window. Rugg's allow versus forbid split moved answers by around twenty points; the welfare versus assistance to the poor gap has persisted for decades. Both results are true of the question; neither is true of "the public".
Which raises a real question about what an attitude is. Converse proposed non-attitudes — answers manufactured on the spot to avoid the awkwardness of refusing. Zaller's account is now standard : people hold conflicting considerations and sample from whichever are accessible, so wording and context change the answer without anyone being inconsistent. Wording effects are evidence about opinion, not noise obscuring it.
The standardised interview suspends ordinary conversational cooperation for good reasons and at a cost: respondents cannot ask what a question means, so they read the instrument for clues, and their genuine uncertainty is coded away.
Ten wording rules , each preventing a documented failure — one idea per question, no leading, watch loaded words, no double negatives, define terms, short time frames, balanced alternatives, no presuppositions, the respondent's vocabulary, no impossible arithmetic. And open and closed questions do not measure the same thing : closed answers concentrate on the offered list.
Four sub-awareness effects : question order changes what later questions mean (dating before happiness); consistency and reciprocity effects (American and communist reporters); the response scale communicates a norm (the television-hours study, where the scale changed both the reports and the self-assessments); and option order produces primacy on paper, recency by voice.
Acquiescence produces spurious coherence in same-direction batteries; satisficing — straightlining, first-acceptable-option, don't-know as exit — is produced by long, complex, late, unmotivating instruments.
And the survey measures what people report in the situation of being surveyed. LaPiere's study is the classic demonstration of the gap and has real defects; the durable finding is that general attitudes predict specific acts weakly, and predict much better when measured at the same specificity.
Never accept a poll result without its wording.
Split-ballot experiment — randomly assigning question versions to measure a wording effect.
Non-attitudes — Converse: answers generated on the spot where no opinion exists.
Considerations / RAS model — Zaller: people sample from conflicting accessible thoughts rather than consulting a stored attitude.
Double-barrelled question — two questions fused, with no coherent single answer.
Presupposition / filter question — a question assuming a fact; the prior question that establishes it.
Open / closed question — free response versus selection from a supplied list.
Question order effect — earlier questions changing the meaning or salience of later ones.
Assimilation / contrast — later answers pulled towards or pushed away from what was made salient.
Response scale effects — the range offered communicating a norm and shifting reports.
Primacy / recency — favouring early options on paper, late options when read aloud; remedied by rotation.
Acquiescence bias — agreeing regardless of content.
Satisficing — shortcut answering: straightlining, non-differentiation, first-acceptable-option, don't-know as exit.
Mode effect — the influence of administration method on answers; and the trend break created by changing mode.
Cognitive interviewing — think-aloud pretesting to discover how respondents actually understand a question.
Attitude–behaviour gap — the weak relation between general stated attitudes and specific situated acts.
One — find the wording. Take a poll result reported this week and locate the exact question. Then write a differently worded version, equally defensible, that you think would move the number. This is the whole skill.
Two — run your own split ballot. Ask ten people the allow version and ten the forbid version of any policy question you like. The effect is large enough to appear in twenty people.
Three — diagnose a bad questionnaire. Find any customer, employee or membership survey and mark every violation of the ten rules. Most will have several in the first five questions.
Four — spot the scale. Next time you meet a frequency question with banded answers, note where the middle of the range sits, and ask what it implies is normal.
Five — test the gap. Ask someone what they would do in a specific situation, in the abstract. Then ask them what they actually did the last time something like it happened. Note the distance, and note that only one of those is behaviour.
Surveys tell you what is associated with what. They almost never tell you what causes what, because the people who differ on the thing you are interested in differ in a hundred other ways too (see 7.3.2).
The instrument built to solve exactly that problem is the experiment — and its logic, once you see it clearly, is the standard against which every other causal claim in social science gets judged, including the ones that cannot use it.
7.4.2 — Experiments and the Logic of the Counterfactual covers what randomisation actually does, why it is so powerful, what it still cannot deliver, and the field experiments that have taught sociology the most.