"Is there discrimination in hiring?" is a causal question about a counterfactual: would this person have been treated differently if they had belonged to another group, everything else being identical? (see 7.4.2).
Most public argument about it proceeds entirely by correlation , which cannot settle it (see 7.6.2). Both sides then produce evidence, disagree, and conclude the other is motivated.
There is better evidence than that , and this lesson is about which designs can carry the claim, which cannot, what they have found, and where the honest uncertainty lies.
One question, three methods, three answers.
Method one: ask.
Survey employers about whether they would hire someone from a given group, or with a given history. They say yes, in large majorities.
And 7.8.3 showed exactly what that is worth. When the same employers who had been through an audit study were surveyed about willingness to hire an applicant with a criminal record, stated willingness vastly exceeded actual callback behaviour — and, at the level of the individual firm, did not predict it at all.
Verdict from method one: little discrimination. And the method is known to be structurally misleading on this question.
Method two: control for things.
Take a large dataset of workers, regress earnings or employment on group membership, and add controls: education, experience, region, industry, occupation, hours, firm size. Watch the gap shrink as controls are added. Report what remains as the "unexplained" portion, and describe it as an estimate of discrimination.
With a rich enough dataset, the gap can usually be reduced to something small.
Verdict from method two: it depends entirely on what you controlled for — and that is not a technicality, it is the whole result.
Method three: make the counterfactual.
Send matched applications that are identical in everything except a signal of group membership, to real vacancies, and count callbacks (see 7.4.2).
Verdict from method three: substantial, replicated, and unarguable in a way the others are not — because there is nothing to control for. The applications were the same.
Three methods. Three verdicts. And they are not equally good.
Method two's problem is the deepest, and it is 7.6.2's bad-control problem with political consequences.
Suppose discrimination operates partly by limiting access to education, partly by channelling people into lower-paying occupations, and partly by paying less within the same job.
Now control for education and occupation. You have removed two of the three channels from the estimate. The remaining "unexplained" gap measures only the third , and the more thoroughly you control, the smaller the discrimination you can find.
In the limit, a sufficiently rich set of controls estimates discrimination as zero — not because there is none, but because every pathway through which it operates has been held constant. This is controlling for a mediator (see 7.2.2), and it is the single most common error in the empirical literature on group gaps.
And the error runs in the other direction too. The unexplained residual also contains everything relevant that was not measured . If groups differ in something productive that the data does not record, that difference sits in the residual and is reported as discrimination.
So the decomposition has two biases pointing opposite ways, and no way to size either. It is genuinely informative about how much of a gap is associated with measured characteristics — which is worth knowing. It cannot identify discrimination , and reporting the residual as a discrimination estimate is a claim the method cannot support in either direction.
What is being measured
Three kinds of discrimination, which the designs distinguish unevenly.
Taste-based discrimination — the decision-maker prefers not to deal with members of a group, and accepts a cost to avoid it. The classic economic prediction is that competition should erode it , since a firm that discriminates forgoes productive workers and pays for the privilege. The prediction has not been borne out , which is itself a finding requiring explanation.
Statistical discrimination — the decision-maker uses group membership as a proxy for something unobserved, because individual information is costly. No animus is required , and this is the version employers most readily describe when interviewed.
Two things must be said about it, and they usually are not.
It is still discrimination , legally and in effect: an individual is treated according to a group average rather than their own characteristics, and is harmed by it.
And it is self-reinforcing. If employers expect less from a group and therefore invest less in them, train them less, and promote them less, the group's observed average performance falls, which appears to confirm the expectation. The formal version of this argument shows that a stereotype can be self-fulfilling even when it started as false — which means "statistical" discrimination does not require the statistics to have been true to begin with.
Structural or institutional discrimination — unequal outcomes produced by rules and arrangements without any discriminating decision-maker. Recruitment through referrals in a segregated network; criminal-record filters where enforcement is unequal; catchment-based admission where residence is segregated; credential requirements that exceed the job's demands. No individual does anything, and the outcome is systematically unequal.
The audit design detects unequal treatment , whatever the motive. It cannot separate taste from statistics, and — this is often missed — it is largely blind to the third kind entirely , because structural discrimination operates through rules the design holds constant.
Audit and correspondence studies: what they establish and what they do not
What they establish is a clean causal estimate at one point in one process.
Identical applications, random assignment of the signal, real vacancies, real decisions (see 7.4.2). The internal validity is close to experimental , and this is why the design transformed the field.
Six limits, and each bounds the claim.
One — they measure the callback stage of advertised, entry-level vacancies. They cannot see discrimination in promotion, in pay setting, in assignment to projects, in dismissal, or in the large share of jobs never advertised — which is exactly where 8.3.3's networks operate. The estimate is a lower bound on total labour-market discrimination.
Two — they measure the effect of a signal, not of a person. A name conveys race and also conveys other things, which is why careful designs vary the signal in several ways and check that the results agree.
Three — the unobservables critique , which is technical and genuine. If the variance of unobserved suitability differs between the groups being compared, then the same threshold applied to both can generate different callback rates without any differential treatment of an equivalent candidate. This is a real identification problem, not a rhetorical objection.
The response is a design response , and it is the right one: vary the quality of the applications systematically and see whether the gap behaves as the variance explanation predicts. Studies doing this generally find patterns inconsistent with the artefact account — and, in one influential case, found something the artefact story cannot explain at all: the return to a stronger résumé was substantially larger for the majority-signalled applications than for the minority-signalled ones (see 7.4.2). A difference in the slope of returns to quality is not what unequal variance produces.
Four — the effect is a rate, not an experience. A callback gap of the size typically found means that most applications from both groups are rejected, and the difference accumulates over a job search rather than appearing in any single decision. This is why individual experience and aggregate evidence so often seem to conflict : neither applicant can see the pattern from inside it.
Five — ethics (see 7.8.1). Real employers spend real time on applications from people who do not exist.
And six — the design is easiest where hiring is formal and advertised , which means it works best in exactly the labour markets that are already most regulated, and worst where informality is greatest.
What the accumulated audit evidence shows.
The most informative single result comes from meta-analysis — pooling every field experiment of hiring discrimination conducted in the United States over a quarter of a century, covering tens of thousands of applications.
Discrimination against African American applicants showed no decline over the period. The callback disadvantage measured in studies from the late 1980s was statistically indistinguishable from that measured in studies from the 2010s. Discrimination against Latino applicants showed some decline.
This is the finding to hold , because it is a claim about a trend , established from a design whose weaknesses are the same at both ends of the period, and it is what the surveys of stated attitudes over the same decades would not have led anyone to expect — attitudes measured by survey improved substantially and continuously (see 7.4.1 on the attitude–behaviour gap).
And the cross-national comparison matters as much. Applying the same design across many countries reveals wide and systematic variation in the size of hiring discrimination between national labour markets , with some countries showing several times the disadvantage of others for comparable groups.
That variation is the most important thing in this lesson , and it is what 8.1.2 said too: if the magnitude differs this much between countries at similar levels of development, it is a product of institutions rather than an inevitability.
Other designs, including three that produced counterintuitive results.
Blind evaluation. The best-known example concerns orchestral auditions conducted behind a screen, with evidence that the practice increased the proportion of women advancing. The finding is widely cited and its statistical precision has been questioned — the samples are small and some of the estimates are imprecise. It is suggestive rather than decisive , and it is quoted far more confidently than the underlying numbers support.
Anonymised applications, where the result went the other way. A large French field experiment randomly assigned firms to receive anonymised CVs. Anonymisation reduced the interview rate for minority candidates. The explanation offered — supported by the pattern of results — is that in the observed setting, employers who were inclined to consider minority candidates had been using contextual information to interpret gaps in a CV, and anonymisation removed their ability to do so, while doing nothing about employers who were not so inclined.
Removing information can make things worse when some decision-makers were using it favourably.
And the same lesson from a different intervention. Policies removing criminal-record questions from job applications, intended to help applicants with records, were found in a field experiment to increase the racial gap in callbacks — because employers deprived of the specific information fell back on group-level assumptions about who was likely to have a record. Statistical discrimination filling an information vacuum, exactly as the theory predicts.
Both results should be held carefully. They are single studies in specific contexts and should not be over-generalised (see 7.9.1). What they establish reliably is a mechanism : information removal does not reliably reduce discrimination, because it changes what decision-makers use instead.
Implicit bias measures: what the evidence actually supports.
The implicit association test measures the speed of sorting words and images into paired categories, on the reasoning that faster pairing indicates stronger mental association. It became enormously influential , and a large industry of implicit bias training grew from it.
The evidence has not been kind, and this course reports it evenly.
Meta-analyses have found the test's ability to predict discriminatory behaviour to be weak — correlations small enough that its use to assess individuals is not defensible, a point its own developers have acknowledged regarding individual diagnostic use.
A large meta-analysis of intervention studies found that procedures which do change implicit measures do not reliably produce corresponding changes in behaviour — which severs the link the whole applied enterprise depends on.
And evaluations of implicit bias training have generally found small or short-lived effects on outcomes , with some evidence that poorly designed versions can provoke backlash.
Now the two corrections that matter, in both directions.
This is not evidence that discrimination is not real. The audit evidence is a completely separate body of work using a completely different design, and it is strong. A weak psychological measure does not undo a field experiment.
And it is not evidence that unconscious processes play no part. It is evidence that this instrument predicts behaviour weakly and that interventions targeting it do not reliably change outcomes.
The proportionate conclusion : discrimination in behaviour is well established; the popular measure of the psychology behind it is a poor predictor; and interventions aimed at individual mental states have a weaker record than interventions that change procedures — structured criteria, blind stages where appropriate, accountability for decisions, and audit. The evidence points away from changing minds and towards changing processes , which is a sociological conclusion arrived at through psychology's disappointments.
Because the question is answerable, and most of the argument is conducted with the methods that cannot answer it.
Five questions for any claim about discrimination.
Is this a gap or an estimate of differential treatment? A gap is a description. Differential treatment requires a counterfactual.
If controls were used, what was controlled? Education, occupation and experience are pathways, not confounders, and controlling for them measures only what remains after removing most of the effect (see 7.6.2).
If it is an audit, what stage does it capture? Callbacks are one point in a long process, so it is a lower bound.
Is the finding a single study or a body? For hiring discrimination, there is a large meta-analysed body with a clear trend result. For most other questions there is not , and single-study findings should be held as single-study findings (see 7.9.1).
And does the argument distinguish between the claim that discrimination exists and the claim that it explains a particular gap? These are different claims , and both sides routinely conflate them: demonstrating discrimination does not establish that it accounts for the whole of an observed difference, and demonstrating that other factors matter does not establish that discrimination does not.
Two things to carry out of this lesson.
The trend result. Measured attitudes improved for decades; measured hiring discrimination against African American applicants did not decline over a quarter of a century. Whatever changed, it was not the behaviour at the point of decision.
And the cross-national variation. The magnitude differs several-fold between comparable countries — which means it is institutionally produced, and therefore institutionally changeable. That is the same conclusion 8.1.2 reached about the distribution itself, and it is the most consequential finding in this Part.
Three methods, three verdicts. Asking employers understates it and does not predict their own behaviour. Controlling for characteristics can reduce any gap to nothing , because education, occupation and experience are pathways through which discrimination operates — controlling for a mediator (7.2.2) — while the residual also absorbs everything unmeasured. Two biases pointing opposite ways and no way to size either: the decomposition cannot identify discrimination. Audit designs construct the counterfactual directly and have nothing to control for.
Three kinds of discrimination. Taste-based — which competition was predicted to erode and has not. Statistical — no animus required, still discrimination in law and effect, and self-reinforcing , since expecting less and investing less produces the performance that appears to confirm the expectation. And structural — unequal outcomes from rules with no discriminating actor, which the audit design is largely blind to because it holds those rules constant.
Six limits on audit studies : they capture the callback stage of advertised vacancies only, so they are a lower bound ; they measure a signal rather than a person; the unobservables critique is genuine and is addressed by varying application quality — with the finding that returns to a stronger résumé differ by group, which unequal variance cannot explain; the effect is a rate that accumulates across a search rather than appearing in any one decision; the ethics are real; and the design works worst where hiring is informal.
The accumulated evidence. Pooling a quarter-century of American field experiments: no decline in discrimination against African American applicants , some decline against Latino applicants — over a period in which survey-measured attitudes improved continuously. And cross-nationally, the magnitude varies several-fold between comparable countries , which makes it institutional rather than inevitable.
Three counterintuitive design results. Blind auditions are suggestive and less statistically robust than their fame implies. Anonymised applications reduced minority interview rates in a large French experiment, by removing information that favourably-inclined employers had been using. And removing criminal-record questions widened racial callback gaps , because employers fell back on group assumptions — statistical discrimination filling an information vacuum.
On implicit measures : the test predicts behaviour weakly, interventions that shift it do not reliably shift behaviour, and training evaluations are disappointing. This does not undo the audit evidence, which is separate and strong — and the practical implication points away from changing minds and towards changing procedures.
Taste-based discrimination — a preference not to deal with a group, borne at a cost.
Statistical discrimination — using group membership as a proxy for unobserved individual characteristics; still discrimination, and self-reinforcing.
Self-fulfilling stereotype — expectations producing the behaviour or performance that appears to confirm them.
Structural discrimination — unequal outcomes generated by rules and arrangements without a discriminating decision-maker.
Decomposition / unexplained residual — the portion of a gap not accounted for by measured characteristics; not an identification of discrimination.
Bad control — conditioning on a variable that lies on the causal pathway.
Correspondence / audit study — matched applications or testers with a randomly assigned group signal.
Unobservables critique — the argument that differing variance in unobserved suitability can generate apparent discrimination; addressed by varying application quality.
Lower bound — what an audit estimate is with respect to total labour-market discrimination.
Implicit association test — a reaction-time measure of mental association; a weak predictor of individual behaviour.
One — audit the controls. Find a study reporting an "unexplained" gap. List its controls and mark each as a confounder or a pathway. Then reconsider what the residual measures.
Two — separate the two claims. Take an argument about a group gap and identify whether the claim is "discrimination exists" or "discrimination explains this gap". Then check whether the evidence offered supports the one being made.
Three — find the stage. For any discrimination estimate, work out which point in the process it measures — application, callback, offer, pay, promotion, dismissal. Then list the stages it says nothing about.
Four — test the information intuition. Before reading further about it, predict what removing names from applications would do. Then reconcile your prediction with the French result , and write down what mechanism you had not considered.
Five — check a trend claim. Find a statement that discrimination has declined. Ask whether the evidence is attitudinal or behavioural. The two have moved differently for decades.
Measurement establishes that unequal treatment occurs and that its magnitude is institutionally produced. It does not explain why the arrangement exists, why it has proved so durable, or why it survives the disappearance of the beliefs that created it.
8.4.3 — Racial Capitalism and Structural Accounts covers the theories that treat racial inequality as produced by economic and institutional arrangements rather than by attitudes: the split labour market, racial capitalism, the wages of whiteness, colour-blind racism, and the empirical claims each makes — including the ones that have not held up.