Your question contains words. Poverty. Trust. Precarity. Radicalisation. Wellbeing. Integration. Discrimination.
Everyone knows what those words mean, which is the problem — everyone knows something slightly different, and none of them can be counted. Between the word and the evidence there is a chain of decisions, and each link is a place where a study quietly becomes about something other than what it says it is about.
This is the least glamorous topic in Part 7 and the one that most often decides the answer.
One district, three definitions of poverty, three different lists of the poor.
An organisation has money to support the poorest households in a district of about 12,000 families. It needs a list. It hires three assessors, who work independently.
The first uses an income line. She collects reported household income, divides by household size, and takes everyone below the national threshold. She gets 3,100 households.
The second uses consumption. He argues that reported income is unreliable where work is casual and seasonal — people genuinely do not know what they earned last year — and that what a household actually spends is a steadier signal. He administers a consumption module and takes the bottom equivalent share. He gets 3,100 households, and about 1,900 of them are also on the first list.
The third uses a multidimensional measure , in the style of the Alkire–Foster method used for the global Multidimensional Poverty Index: not money at all, but deprivations — years of schooling in the household, child attendance, child mortality, nutrition, cooking fuel, sanitation, drinking water, electricity, housing, assets. A household counts as poor if it is deprived in a weighted third or more. He gets 3,100 households, of which about 1,700 are on the first list.
Now look at who moves.
A household with a decent seasonal income, no toilet, no electricity and a child out of school is poor on the third measure and not on the first.
A pensioner living alone on a tiny income in a solid house with piped water is poor on the first and not on the third.
A family that had a bad year — illness, a failed crop — is poor on income now and not deprived yet, because the assets have not gone yet. Give it two more bad years and it will be on every list.
All three assessors are competent. None has made an error. They have measured three different things, all of which are called poverty in ordinary speech.
And around a thousand households will receive support or not receive it depending on which one the organisation adopts — a decision that will be made in twenty minutes, by someone who was not in the field.
The word is not the thing, and the measure is not the word.
There are three separate objects here and they are routinely conflated:
The concept — poverty, as an idea about a condition in the world.
The definition — the specific account you commit to: lacking the resources for a customary standard of living; or, being deprived on multiple dimensions simultaneously.
The measure — the actual operation performed: this question, this threshold, this index.
Every step down that ladder loses something, and every step is a choice that could have gone otherwise. The finished paper reports the measure and discusses the concept, and the reader is invited to forget there were two intervening decisions.
So the question that unlocks nearly every study you will ever read is four words long: how was that measured? It sounds pedantic. It is the single most powerful thing you can ask.
The chain, one link at a time
Lazarsfeld's four steps, which are still the clearest account of how this works.
Paul Lazarsfeld, who built much of the machinery of modern survey research, described the process as four movements.
One — the imagery. You begin with a vague but real sense of a phenomenon: some workplaces feel insecure in a way that goes beyond the contract . It is not yet a concept. It is a perception looking for one.
Two — concept specification. You break the imagery into dimensions — distinguishable components. Job insecurity might have: contractual insecurity (how easily can they end it), income insecurity (how predictable is the amount), schedule insecurity (do you know when you work), and subjective insecurity (do you expect to lose it). These are not the same and they do not move together ; a tenured academic on a stable salary can have high subjective insecurity, and a well-paid contractor can have none.
Three — selection of indicators. For each dimension, observable things that would be evidence of it. For schedule insecurity: notice period for shifts, variation in weekly hours, whether hours are guaranteed.
Four — formation of indices. Combining the indicators into a measure. Which is where the weighting problem appears , and it is unavoidable: is one week's notice worth the same as a twenty-hour swing in weekly hours? Any answer is a judgement, including the judgement to weight everything equally — that is a claim too, and usually an unexamined one.
The honest way to present this is to show the chain and defend each link. The common way is to name the concept in the introduction, cite an existing scale in the methods, and never mention the dimensions again.
Conceptual (or nominal) definition — what you mean by the term, stated in words.
Operational definition — the exact procedure by which you will produce a value: the question asked, the records consulted, the threshold applied.
Indicator — an observable thing taken as evidence of something not directly observable.
Latent variable — the underlying thing you cannot observe (trust, ability, prejudice). Manifest variable — what you actually record. The indicators are manifest; the concept is usually latent , and the entire art is in the relation between them.
Index / scale — a composite built from several indicators. Loosely, an index sums components that each contribute independently (deprivations); a scale combines items taken to reflect one underlying latent thing (attitude items).
The vocabulary of variables, which you need in order to read anything.
Independent variable — the presumed cause; the thing whose variation you are interested in.
Dependent variable — the outcome.
Mediator — a variable on the path between them: X changes M, and M changes Y. Mediators are mechanisms (see 7.2.3).
Moderator — a variable that changes the size or direction of the X→Y relation. Not on the path; it conditions the path. "The effect of scholarships is larger for first-generation students" is a moderation claim.
Confounder — a variable that causes both X and Y, producing an association between them that is not causal (see 7.6.2).
Control variable — something held constant statistically.
And the distinction that trips up almost everyone : you should control for confounders and you should not control for mediators. Controlling for a mediator removes the very pathway by which the cause operates, and the effect obediently disappears — which the researcher then reports as an absence of effect.
Example : does education raise earnings? Control for occupation, and much of the effect vanishes — because getting a different occupation is how education raises earnings . The control did not remove bias. It removed the finding.
This is why 7.1.3 ended where it did. Whether a variable is a confounder or a mediator is not a statistical question. It is a question about how you think the world works, answered before the analysis, and the analysis cannot answer it for you.
Levels of measurement, and what each one lets you do.
Stevens's 1946 classification, still standard:
Nominal — categories with no order. Religion, region, occupation type. You can count them and compare proportions. Nothing else.
Ordinal — ordered categories with unknown spacing. Education levels; "strongly agree" to "strongly disagree"; social class schemes. You can rank and take medians.
Interval — equal spacing, arbitrary zero. Temperature in Celsius; many constructed scale scores.
Ratio — equal spacing and a true zero. Income, age, hours, counts. Everything is permitted, including "twice as much".
The everyday violation : treating ordinal as interval. Averaging Likert responses assumes the distance from "strongly agree" to "agree" equals the distance from "neutral" to "disagree", which is not established and is often false. It is done constantly, and there is a defensible pragmatic case for it with enough categories and a symmetric scale — but it is a decision, not a fact , and it should be visible.
And a related one : making a continuous variable into categories. Splitting a continuous measure at the median throws away information, and choosing where to split after seeing the results is one of the most effective ways to manufacture a significant finding (see 7.6.3).
Is the measure any good? Two different questions
Reliability is consistency. Validity is aim. You need both, and they are unrelated.
The dartboard picture. A tight cluster in the wrong corner is reliable and invalid . Scattered all over but centred on the bullseye is valid on average and unreliable . Reliable and invalid is the dangerous one , because it looks like precision and nobody questions it.
Reliability — will it give the same answer again?
Test–retest : same people, later, unchanged conditions, same answer?
Inter-rater : two coders reading the same interview, do they code it the same? Reported as Cohen's kappa , which corrects for agreement that would occur by chance — and this correction matters, because two coders who both apply a code to 90% of cases will agree 82% of the time by luck alone.
Internal consistency : do the items in a scale move together? Reported as Cronbach's alpha . Two standing warnings : alpha rises mechanically as you add items, so a high alpha on a twenty-item scale is not impressive; and alpha measures consistency, not unidimensionality — a scale measuring two distinct things that happen to correlate can have an excellent alpha.
Validity — is it measuring the thing you claim?
Face validity : does it look plausible? The weakest, and not nothing.
Content validity : does it cover the whole concept, or only one dimension? A measure of "civic participation" that asks only about voting has a content problem.
Criterion validity : does it agree with an accepted benchmark (concurrent ) or predict what it should (predictive )?
Construct validity : does it behave the way the concept should? Two halves — convergent (it correlates with things it ought to) and discriminant (it does not correlate with things it ought not to). The second half is skipped far more often than the first , and it is where most measures fail: a "social capital" scale that correlates equally with income, education, health and happiness may be measuring general advantage.
Ecological validity : does it hold outside the setting where it was established? Behaviour in a lab, a survey or a training exercise is not automatically behaviour in a life (see 7.4.2).
Measurement invariance — the problem that eats comparisons.
A measure can be excellent in one group and mean something different in another, and then every comparison between them is an artefact.
Ask people in twenty countries how much they trust "most people". The word translates; the referent may not. In one setting "most people" means neighbours; in another it means strangers on the street; in another the question is heard as being about the state.
Ask about depressive symptoms across cultures where distress is somatised differently — reported as headaches and fatigue rather than as sadness — and a symptom checklist built in one place undercounts in another.
Ask about "household income" where households are nuclear and where they are joint and multi-generational, and you are not counting the same object.
The technical name is measurement invariance , and there are formal tests for it. The important part is the habit: before you accept that country A is more trusting than country B, ask whether the instrument means the same thing in both. Frequently the finding is entirely in the instrument — which is 6.8.1's argument arriving in the form of a statistical diagnostic.
When measuring changes the thing
Three effects that have no equivalent in the physical sciences.
Reactivity — people behave differently when observed. The Hawthorne effect is the famous name, from the Western Electric studies of the 1920s and 1930s where productivity appeared to rise under every change in lighting, apparently because of the attention itself.
And the legend needs a correction. When Levitt and List recovered and reanalysed the original illumination experiment data, they found the evidence much weaker and more ambiguous than the textbook story — some of the pattern is explained by the fact that lighting changes were made on Sundays, with productivity always rising at the start of the week. Reactivity is real and well demonstrated elsewhere; the specific canonical study is much shakier than its fame suggests. It belongs in this course as an example of a second thing: how a tidy story survives in textbooks long after the evidence for it has been questioned.
Social desirability bias — people report what it is respectable to report. Reported voting exceeds actual voting. Reported charitable giving exceeds receipts. Reported prejudice has fallen faster than behavioural measures of it. The remedies are design remedies : self-completion rather than interviewer administration; indirect questions; and the list experiment , where respondents are shown a list and asked only how many items apply to them, with a randomly assigned half seeing one extra sensitive item — the difference in averages estimates its prevalence without any individual ever revealing anything.
Mode effects — the same question produces different answers by telephone, face-to-face, on paper and online. Sensitive items get higher endorsement without a human present; agreement bias is stronger when someone is asking aloud. A trend that spans a change of survey mode is partly a measurement of the change of mode.
Goodhart's law, and why every good indicator eventually stops being one.
Charles Goodhart , writing about monetary policy in 1975: any observed statistical regularity tends to collapse once pressure is placed on it for control purposes. Marilyn Strathern's compression is the one everybody quotes: when a measure becomes a target, it ceases to be a good measure. Donald Campbell stated a social-science version at almost the same time.
The mechanism is not cheating, though cheating occurs. It is that people optimise the measured thing, and the measured thing was only ever a proxy for the unmeasured thing you cared about.
Examples across systems : exam results as a proxy for learning, followed by teaching to the test and entering weak candidates for easier qualifications. Ambulance response times measured from a defined start point, followed by intense management of when the clock starts. Hospital waiting-time targets met by reclassifying the point at which the wait begins. Police detection rates improved by recording practice — the same underlying incidents, differently classified.
And the sociological point that outlasts the examples : this is Hacking's looping effect from 7.1.3, in institutional form. A measure applied to people with consequences attached changes the people, and then no longer measures what it measured before it was applied. Which is why time-series of administrative indicators must always be read with the question: did the world change, or did the recording change?
Unemployment, which almost nobody has ever looked up.
The standard international definition — the ILO one, used with variations by most national statistical offices — counts someone as unemployed if they are without work, currently available for work, and actively seeking work in a specified recent period.
Look at what each clause does.
Actively seeking excludes the discouraged worker — the person who stopped looking because there was nothing. In a deep recession, the unemployment rate can fall as conditions worsen, because people exit the search.
Without work usually means without work for at least one hour in the reference week. Someone doing four hours of casual labour is employed.
Available for work excludes people whose caring responsibilities make immediate availability impossible, whatever they would do with a real offer.
None of this is a scandal ; every alternative has its own problems, which is why statistical agencies publish several supplementary measures alongside the headline. The scandal is that public argument treats the headline number as "how many people cannot find work", which is not what it measures , and that people arguing about it have usually not read the definition.
Do this once, for one indicator you care about, and it will change how you read everything.
Because the measurement decision is where the ontology of 7.1.3 gets made, silently, by someone junior, under time pressure.
Nobody convenes a seminar to decide whether poverty is a state of a household at a moment or a trajectory over years. Somebody chooses an income threshold at a single time point, and the question has been answered.
For a reader, this yields three habits that will serve you for life.
One — always ask how it was measured, before evaluating anything else. A surprising number of astonishing findings are ordinary findings about an unusual measure.
Two — when a trend appears, ask whether the definition or the recording changed. Rises in recorded crime, hate crime, mental illness diagnoses and homelessness have all, at various times and places, been substantially artefacts of changed recording — and have also, at other times and places, been real. The point is not that they are always artefacts. It is that you cannot tell without checking, and the reporting almost never checks.
Three — when a difference between groups or countries is reported, ask whether the instrument means the same thing in both.
And for anyone doing research: write down your conceptual definition before you go looking for a measure. If you choose the measure first — because it is available, because everyone uses it, because it is in the dataset — you will find that your concept has quietly become whatever the measure captures , and you will not notice, because the word in your title will not have changed.
Three distinct objects: the concept, the definition, the measure. Every step down the chain loses something, and each is a choice. Lazarsfeld's four steps — imagery, concept specification into dimensions, selection of indicators, formation of indices — with the weighting problem unavoidable at the end.
Variable vocabulary : independent, dependent, mediator (on the path — a mechanism), moderator (conditions the path), confounder (causes both), control. Control for confounders, never for mediators — controlling for a mediator deletes the pathway and reports the deletion as an absence of effect.
Levels of measurement — nominal, ordinal, interval, ratio — determine what operations are legitimate; treating ordinal as interval is a defensible decision but a decision.
Reliability is consistency; validity is aim; they are independent, and reliable-but-invalid is the dangerous combination. Reliability: test–retest, inter-rater (kappa, which corrects for chance agreement), internal consistency (alpha, which rises with item count and does not establish unidimensionality). Validity: face, content, criterion, construct — with discriminant validity the most-skipped and most-failed test — and ecological.
Measurement invariance : a comparison between groups is only as good as the instrument's meaning the same thing in both.
Measuring changes the measured. Reactivity (with the Hawthorne studies themselves much weaker than the legend), social desirability bias (remedied by design, including list experiments), and mode effects. And Goodhart's law : when a measure becomes a target it ceases to be a good measure — the looping effect, institutionalised.
Four words, applied to everything: how was that measured?
Conceptual definition / operational definition — what you mean; the exact procedure that produces a value.
Indicator — an observable taken as evidence of something unobservable. Latent / manifest — the underlying thing; what is recorded.
Index / scale — composites of several indicators, summed or taken to reflect one latent thing.
Mediator — a variable on the causal path; a mechanism. Moderator — a variable that changes the size or direction of an effect. Confounder — a common cause of both.
Nominal / ordinal / interval / ratio — Stevens's levels of measurement, determining permissible operations.
Reliability — consistency: test–retest, inter-rater (Cohen's kappa), internal consistency (Cronbach's alpha).
Validity — face, content, criterion (concurrent and predictive), construct (convergent and discriminant), ecological.
Measurement invariance — whether an instrument means the same thing across groups or settings.
Reactivity / Hawthorne effect — behaving differently when observed.
Social desirability bias — reporting what is respectable. List experiment — an indirect design estimating prevalence without individual disclosure.
Mode effect — the influence of how a question is administered on the answer.
Goodhart's law / Campbell's law — a measure under pressure as a target stops measuring what it did.
One — read a definition. Pick one official statistic you have opinions about — unemployment, poverty, literacy, crime, homelessness — and find the actual definition your country's statistical agency uses. Write down two kinds of person it excludes who you assumed were included.
Two — build the chain. Take an abstract word from your own question. Write the conceptual definition, then two or three dimensions, then two indicators for each, then how you would combine them and why with those weights. The weighting step is where you will feel the difficulty; that feeling is the lesson.
Three — hunt a mediator control. Find a study reporting that an effect "disappears after controlling for" something. Ask whether that something is a cause of both, or one of the ways the effect happens. If it is the second, the study has deleted its own finding.
Four — run the discriminant test. Take any scale you meet — resilience, wellbeing, engagement, social capital. Ask what it should not correlate with if it measures what it claims. Then check whether the paper reports that.
Five — find a Goodhart. Identify one measure in your workplace, school or local services that has consequences attached. Describe how people have adapted to it, and what the measure no longer captures.
You have a question and you have concepts you can actually measure. The last piece of design is the statement that connects them — and it is where the difference between testing an idea and decorating one gets settled.
7.2.3 — Hypotheses, Theory-Testing and Theory-Building covers what a hypothesis has to risk, the two directions research can run in, the strange logic of the null hypothesis, and why the most valuable studies are often the ones designed so that the author's favourite explanation could lose.