Abstract
Implicit bias is an evaluation or stereotype that operates without conscious intention or awareness and that can diverge from a person's deliberately held beliefs. It is inferred indirectly — most often from the speed with which people pair social categories with good or bad attributes on the Implicit Association Test — rather than from what they report. The construct grew from the theory of implicit social cognition, which holds that attitudes and stereotypes leave traces on judgment that introspection cannot reach. This article covers what implicit bias is, how it is measured, the dual-process models that explain why implicit and explicit measures diverge, the vigorous debate over how well implicit measures predict behavior, the evidence on whether such bias can be reduced, and the reframing of implicit bias as a property of situations and populations rather than fixed individuals.
Keywords: implicit bias, implicit association test, attitudes, stereotyping, dual-process models
Few ideas from social cognition have moved as quickly from the laboratory into public life as implicit bias. It began as a technical proposal — that attitudes and stereotypes could influence behavior through channels that self-report cannot capture — and became the rationale for diversity trainings, courtroom arguments, and policy debates the world over (Greenwald & Banaji, 1995). That reach has made precision unusually important, because the construct is easy to overstate. Implicit bias is not a hidden prejudice that reveals a person's “true self,” nor a diagnosis of character; it is a measured tendency for automatic associations to shape judgment, one that varies across situations, predicts behavior modestly and unevenly, and resists lasting change (Payne, Vuletich, & Lundberg, 2017).
- Implicit bias is an automatic evaluation or stereotype that can operate outside awareness and diverge from consciously endorsed beliefs.
- It is measured indirectly, chiefly by the Implicit Association Test, which infers association strength from category-sorting speed.
- Dual-process models explain the frequent gap between implicit and explicit measures as the output of distinct associative and propositional processes.
- Implicit measures predict individual behavior only weakly, and their predictive validity is one of the field's most contested questions.
- Interventions shift implicit measures briefly but rarely produce durable change, prompting a reframing of bias as a property of situations and populations.
What Implicit Bias Is
Implicit bias is defined by two features: the evaluation is automatic — it is activated quickly, without intention, and is hard to suppress — and it is implicit in the technical sense that the person may not be aware of it, or may not recognize its influence on their judgment. The theoretical foundation is the concept of implicit social cognition, the claim that traces of past experience mediate attitudes and stereotypes in ways that introspection cannot accurately report (Greenwald & Banaji, 1995). On this view, a person can sincerely reject a stereotype while still carrying an association — learned from a lifetime of cultural exposure — that biases a split-second judgment before deliberate thought intervenes.
The construct has a direct ancestor in social cognition research on automaticity. A foundational study distinguished the automatic activation of a stereotype, which occurred for everyone who knew the cultural content regardless of their personal prejudice, from the controlled processes that low-prejudice people could recruit to override it (Devine, 1989). That distinction — knowing a stereotype is not the same as endorsing it, and activation is not the same as application — is the conceptual core of implicit bias, and it is why the phenomenon is studied as a problem of automatic processes rather than of hidden conviction.
Implicit bias must be separated from adjacent notions. It is not the same as explicit prejudice, which is consciously held and directly reported; the two are distinct constructs that often correlate only weakly. It is not cognitive dissonance, the discomfort of holding conflicting beliefs, though the gap between implicit and explicit attitudes can generate it. And it is not a fixed trait: a person's implicit measure varies from context to context and from testing to testing, a volatility that turns out to be central to what the construct is.
Measuring Implicit Bias
Because implicit bias is by definition not accessible to self-report, it must be measured indirectly. The dominant instrument is the Implicit Association Test (IAT), which infers the strength of an association from how quickly a person can sort items when two concepts share a response key (Greenwald, McGhee, & Schwartz, 1998). In a race IAT, for instance, a participant sorts faces and words as fast as they can; in one block, Black and bad share a key while White and good share the other, and in another block the pairings reverse. If sorting is faster when White pairs with good than when Black does, the difference in reaction time is taken to index a stronger automatic association between the White category and positive attributes.
The raw latency difference is converted into a standardized effect size, the D score, by dividing the difference in mean latency between the two critical blocks by the pooled standard deviation of the individual's latencies. Scoring each person against their own variability makes the measure robust to general processing speed, so that a slow and a fast responder with the same underlying association receive similar scores. Conventionally, D values near 0.15, 0.35, and 0.65 are described as slight, moderate, and strong associations, though these labels are descriptive rather than diagnostic.
The first demonstration lets a reader run a simplified IAT and watch the compatible and incompatible blocks produce the latency difference from which a D score is computed.
Dividing the block difference by each person’s own variability makes the score comparable across fast and slow responders. The labels (slight, moderate, strong) are descriptive conventions, not diagnoses.
The IAT is not the only indirect measure. Evaluative and semantic priming tasks measure how a briefly presented category speeds or slows the evaluation of a following word; the affect misattribution procedure uses a category prime to bias judgments of neutral images; and sequential priming variants such as the weapon-identification task measure how a face category shifts the misidentification of ambiguous objects. These measures share the logic of inferring an association from its automatic downstream effect, but they correlate only modestly with one another — a reminder that “implicit bias” names a family of task-dependent measurements, not a single quantity read off a mental dial.
| Measure | What the participant does | Association inferred from |
|---|---|---|
| Implicit Association Test | Sorts items when two concepts share a response key | Latency difference between compatible and incompatible blocks (the D score) |
| Evaluative priming | Judges a target word as good or bad after a category prime | Speed-up or slow-down caused by a congruent versus incongruent prime |
| Affect misattribution procedure | Rates a neutral image (a pictograph) after a category prime | Shift in the pleasantness rating attributed to the neutral image |
| Weapon-identification task | Identifies an ambiguous object as a weapon or tool after a face prime | Bias in errors and speed toward misidentifying objects as weapons |
Dual-Process Models
Why should an indirect measure diverge from what a person sincerely reports? The leading answer is that implicit and explicit measures tap distinct kinds of processing, an instance of the broader dual-process theory of mind. The most developed account is the Associative-Propositional Evaluation (APE) model, which distinguishes associative processes — the automatic activation of a mental link based on similarity and past pairing — from propositional processes, which assess whether an activated thought is valid and consistent with other beliefs (Gawronski & Bodenhausen, 2006).
On the APE account, an implicit measure like the IAT primarily reflects the associative output: whatever link is activated, regardless of whether the person endorses it. An explicit measure reflects the propositional output: what the person is willing to affirm after evaluating that activation for truth and consistency. The two dissociate whenever a person's propositional reasoning rejects an association their environment has installed — the low-prejudice individual of Devine's account, who has the association available but refuses to endorse it. The model also predicts when the two will converge: as propositional processes come to accept an association, or as associative structure shifts to match a belief, implicit and explicit measures move back into alignment.
The Associative-Propositional Evaluation Pathway
Note. A single social input feeds both an associative and a propositional process; the two measures dissociate when propositional validation rejects an automatically activated association (Gawronski & Bodenhausen, 2006).
The second demonstration renders the APE distinction, letting a reader vary associative strength and propositional endorsement independently and see how each combination yields a characteristic pattern of implicit and explicit responses.
The implicit measure tracks associative activation; the explicit measure tracks propositional endorsement. Here they show a strong dissociation: an available association the person does not endorse (or vice versa).
This framing matters for interpretation. If an implicit measure reflects associative activation rather than endorsed belief, then a high score does not license the inference that a person “really” holds a prejudice they are concealing. It licenses only the weaker claim that an association is available and can, under the right conditions, leak into behavior — which is precisely why the question of predictive validity became so contested.
Predictive Validity
The scientific stakes of implicit bias rest on a single question: does an implicit measure predict how a person actually behaves? Early enthusiasm rested on a meta-analysis reporting that IAT scores predicted discriminatory behavior, and did so better than explicit measures in the socially sensitive domains where people are most motivated to misreport (Greenwald, Poehlman, Uhlmann, & Banaji, 2009). That result made the IAT appear to be a uniquely powerful window onto behavior that self-report could not reach.
A sharp rebuttal followed. A competing meta-analysis focused on studies predicting ethnic and racial discrimination and found the IAT-behavior correlation to be small and, the authors argued, too weak and too variable to support the individual-level claims being made for it in policy and law (Oswald, Mitchell, Blanton, Jaccard, & Tetlock, 2013). The disagreement was partly about which studies to include and partly about what a small correlation means: a modest average effect can still matter aggregated across many decisions, or can be near-useless for predicting a given individual's next act.
The most comprehensive synthesis to date, pooling hundreds of studies, put the average correlation between IAT scores and intergroup behavior at roughly r = 0.10 and confirmed that changes in implicit measures did not reliably translate into changes in behavior (Kurdi et al., 2019). The field now broadly accepts that implicit measures are weak predictors of any individual's behavior, while continuing to debate whether that weakness reflects a genuinely small effect, the psychometric limits of the measures, or a mismatch between stable-trait assumptions and a genuinely unstable phenomenon.
The third demonstration makes the levels-of-analysis point concrete: it shows how a correlation too weak to predict one person's behavior can nonetheless yield a strong, reliable relationship once scores and outcomes are aggregated across many people.
The same data yield a near-useless individual correlation and a strong group correlation once people are averaged into regions. Aggregation cancels the momentary noise in individual scores and leaves the stable structural signal.
Malleability and Reduction
If implicit bias contributes to discrimination, can it be reduced? A large collaborative study tested seventeen interventions against a control in a single competition and found that several could shift the race IAT immediately — the most effective invoked vivid counter-stereotypic exemplars or intentional strategies — but the effects were small and their durability untested (Lai et al., 2014). A follow-up delivered the decisive blow to optimism: every intervention that worked in the moment had its effect vanish within hours to days, leaving implicit measures back at baseline.
A meta-analysis of procedures designed to change implicit measures reached the same conclusion at scale: implicit measures can be moved, but the changes are weak, and — critically — changes in implicit measures were not accompanied by changes in explicit measures or behavior (Forscher et al., 2019). The practical implication is uncomfortable for the diversity-training industry that implicit-bias research helped inspire: there is little evidence that a brief intervention retraining an association produces any lasting change in how people act.
This does not mean bias is immovable. A distinct tradition treats prejudice reduction as an effortful, sustained habit change rather than a quick associative retraining — teaching people to recognize the situations that trigger a stereotype and to deploy deliberate strategies over weeks, an approach with more durable results than one-shot manipulations (Devine, 1989). The lesson of the malleability literature is that the target may be wrong: if implicit bias is not a stable individual trait, retraining individuals is the wrong lever.
The Bias of Crowds
The instability that frustrates intervention research motivated a reconceptualization. The bias of crowds model proposes that an implicit measure is best understood not as a stable individual disposition but as a reading of the situation a person is in at the moment of testing — a momentary sample of the concepts the environment has made accessible (Payne, Vuletich, & Lundberg, 2017). On this account, the notorious low test-retest reliability of the IAT is not a flaw to be engineered away but a signal: implicit bias fluctuates because the situations that cue it fluctuate.
The model resolves a paradox. Individual IAT scores are unstable and weakly predictive, yet aggregated implicit bias — the average IAT score of a region or institution — is highly stable over time and predicts consequential aggregate outcomes, from disparities in police violence to gaps in health and education, better than aggregated explicit attitudes do. The resolution is that aggregation averages out the momentary noise in individual scores and leaves the stable, structural signal: implicit bias, on this view, lives in situations and systems, and an individual's score is a sample of the biased environment they inhabit rather than a fixed fact about them.
Worked Example
Consider scoring a single participant's race IAT. In the compatible block — White/good and Black/bad sharing keys — the participant's mean sorting latency is 620 milliseconds. In the incompatible block, with the pairings reversed, the mean latency rises to 780 milliseconds. Across all of the participant's trials in the two blocks, the pooled standard deviation of the latencies is 250 milliseconds.
The D score is the difference in block means divided by that pooled standard deviation: (780 − 620) ÷ 250 = 160 ÷ 250 = 0.64. The positive sign indicates faster sorting in the compatible block — the pattern conventionally read as a stronger automatic association between White and positive attributes — and the magnitude, close to 0.65, falls at the boundary conventionally labeled a “strong” association.
The interpretive caution is the whole point of the predictive-validity debate. With an average IAT-behavior correlation near r = 0.10, this D of 0.64 explains on the order of one percent of the variance in any specific behavior it might be used to forecast (Kurdi et al., 2019). The score is a reliable-enough description of this person's sorting speed on this occasion; it is a poor basis for predicting what they will do next, and — under the bias-of-crowds reading — it may say as much about the testing situation as about the person. The third demonstration shows how scores like this one, individually weak, combine into a strong aggregate signal.
Discussion
Implicit bias is a genuine phenomenon with a fragile interpretation. That automatic associations exist, can be measured, and can influence judgment under some conditions is not seriously disputed; the theory of implicit social cognition that predicted them has been amply confirmed (Greenwald & Banaji, 1995). What two decades of scrutiny have overturned is the strong version that entered public discourse — that the IAT reveals a stable hidden prejudice, that this prejudice reliably drives an individual's discriminatory acts, and that brief training can correct it. Each of those claims has failed a serious empirical test (Oswald et al., 2013; Forscher et al., 2019).
The most productive current reading holds the phenomenon and the caution together. Implicit bias is real but weak at the individual level, unstable across situations, and resistant to durable change — properties that make it a poor target for individual retraining but a good diagnostic of biased environments when aggregated (Payne et al., 2017). This reframing shifts the practical question from “how do we fix biased people?” to “how do we change the situations and structures that make biased responses accessible?” — a question about institutions rather than individuals, and one the aggregate data are far better equipped to inform.
Current Directions
Three questions define the active front. The first is measurement and stability: because the bias-of-crowds model reinterprets the IAT's low reliability as substantive rather than as error, researchers are building measures and designs that deliberately capture situational variation — testing the same people across contexts and time — rather than treating a single score as a trait estimate (Payne et al., 2017). The second is what a small effect is good for: the meta-analytic consensus that individual-level prediction is weak has redirected effort toward aggregate and structural outcomes, where implicit measures retain predictive value, and toward clarifying when a small average effect is nonetheless practically consequential (Kurdi et al., 2019).
The third is intervention design. With brief associative retraining shown to be ineffective, work has moved toward two alternatives: durable, effortful habit-breaking approaches that treat bias reduction as sustained self-regulation, and structural changes that remove the discretion through which bias operates — blind review, standardized criteria, changed defaults — sidestepping the individual mind altogether (Forscher et al., 2019). The common thread is a retreat from the assumption that measuring and retraining individual associations is the route to less discriminatory behavior.
Common Misconceptions
- A high IAT score reveals a person's hidden true prejudice.
- An implicit measure reflects the automatic activation of an association, not an endorsed belief. Dual-process models distinguish the associative link that the IAT captures from the propositional judgment a person actually affirms, so a score licenses only the claim that an association is available, not that the person secretly holds a prejudice (Gawronski & Bodenhausen, 2006).
- Implicit bias strongly predicts how a person will discriminate.
- The most comprehensive meta-analysis puts the average correlation between IAT scores and intergroup behavior near r = 0.10, far too weak to forecast any individual's actions reliably; the measure's value, where it has any, is at the aggregate rather than the individual level (Kurdi et al., 2019).
- A short training session can eliminate implicit bias.
- Brief interventions can shift implicit measures for hours but the effect decays to baseline, and changes in implicit measures do not carry over to explicit measures or behavior — evidence that one-shot retraining does not produce lasting change (Forscher et al., 2019).
Glossary
- Affect misattribution procedure (AMP).
- An indirect measure in which a briefly shown group prime shifts how pleasant a neutral target is judged, and the misattributed affect indexes an automatic evaluation.
- Associative process.
- The automatic activation of a mental link based on similarity and past co-occurrence; in the APE model, the output an implicit measure primarily reflects.
- Attitude.
- A summary evaluation of an object along a good-bad dimension; implicit and explicit measures are taken to tap the same attitude through different processes.
- Automatic process.
- A mental operation that is fast, unintentional, and difficult to control; implicit bias is studied as one such process.
- Bias of crowds.
- The proposal that an implicit measure reflects the situation a person is tested in rather than a stable trait, so that bias is best studied at the aggregate, structural level.
- D score.
- The standardized IAT effect size: the difference in mean latency between the two critical blocks divided by the pooled standard deviation of the individual's latencies.
- Evaluative priming.
- A reaction-time measure in which a group prime speeds or slows the classification of a following word as good or bad, indexing an automatic evaluation.
- Explicit attitude.
- An evaluation a person is aware of and reports directly; often correlates only weakly with the corresponding implicit measure.
- Implicit Association Test (IAT).
- A reaction-time task that infers the strength of an association from how quickly a person sorts items when two concepts share a response key.
- Implicit bias.
- An automatic evaluation or stereotype that can operate without conscious awareness or intention and can diverge from consciously held beliefs.
- Implicit social cognition.
- The theory that attitudes and stereotypes influence judgment through traces of past experience that introspection cannot accurately report.
- Predictive validity.
- The degree to which scores on a measure forecast a criterion such as discriminatory behavior; for the IAT it is statistically reliable but small.
- Propositional process.
- The assessment of whether an activated thought is valid and consistent with other beliefs; in the APE model, the output an explicit measure reflects.
- Stereotype.
- A cognitive association between a social group and an attribute; when activated automatically it can bias judgment independently of an endorsed belief.
- Weapon-identification task.
- A sequential-priming measure in which a face category shifts the speed and accuracy of identifying an ambiguous object as a weapon or a tool.
Key Researchers
Mahzarin R. Banaji (b. 1956). Richard Clarke Cabot Professor of Social Ethics at Harvard University; co-developed implicit social cognition and the IAT and co-authored the leading synthesis of implicit-behavior relations. Faculty Page - ORCID - Google Scholar - Wikipedia
Patricia G. Devine (b. 1957). Professor of psychology at the University of Wisconsin-Madison; established the automatic/controlled distinction in stereotyping and developed the prejudice-habit-breaking approach to durable bias reduction. Faculty Page - ORCID - Google Scholar - Wikipedia
Bertram Gawronski (b. 1971). Professor of psychology at the University of Texas at Austin; developed the Associative-Propositional Evaluation model of implicit and explicit attitudes. Faculty Page - ORCID - Google Scholar - Wikipedia
Anthony G. Greenwald (b. 1939). Professor of psychology at the University of Washington; introduced the Implicit Association Test and, with Banaji, the theory of implicit social cognition. Faculty Page - ORCID - Google Scholar - Wikipedia
Brian A. Nosek (b. 1973). Professor of psychology at the University of Virginia and co-founder of Project Implicit; led large-scale studies of implicit measures and the durability of bias-reduction interventions. Faculty Page - ORCID - Google Scholar - Wikipedia
B. Keith Payne (living). Cary C. Boshamer Distinguished Professor at the University of North Carolina at Chapel Hill; developed the weapon-identification task and the bias-of-crowds model of implicit bias. Faculty Page - Google Scholar
Frequently Asked Questions
What is implicit bias?
Implicit bias is an automatic evaluation or stereotype that can operate without conscious awareness or intention and can diverge from a person's consciously held beliefs; it is inferred indirectly rather than from self-report (Greenwald & Banaji, 1995).
How is implicit bias measured?
Most often by the Implicit Association Test, which infers association strength from how quickly a person sorts items when two concepts share a response key; priming tasks and the affect misattribution procedure are alternatives (Greenwald et al., 1998).
What is a *D* score on the IAT?
The D score is the standardized IAT effect: the difference in mean sorting latency between the two critical blocks divided by the pooled standard deviation of the person's latencies, which makes the score comparable across fast and slow responders.
Why do implicit and explicit measures often disagree?
Dual-process models such as the APE model hold that implicit measures reflect automatic associative activation while explicit measures reflect propositional judgments about what is true, so the two dissociate when a person does not endorse an association they have (Gawronski & Bodenhausen, 2006).
Does implicit bias predict discrimination?
Only weakly at the individual level; the most comprehensive meta-analysis puts the average IAT-behavior correlation near r = 0.10, though aggregated implicit bias predicts group-level outcomes more strongly (Kurdi et al., 2019).
Can implicit bias be reduced?
Brief interventions shift implicit measures temporarily but the effects decay and do not change behavior; durable change appears to require sustained habit-breaking effort or structural changes rather than one-shot training (Forscher et al., 2019).
Is implicit bias a stable personality trait?
Probably not; individual scores are unstable across situations and time, which the bias-of-crowds model interprets as evidence that implicit bias reflects the situation rather than a fixed disposition (Payne et al., 2017).
Does a high IAT score mean I am secretly prejudiced?
No; a high score indicates that an automatic association is available, not that a person endorses a prejudice, and it is a weak predictor of how that person will actually behave (Oswald et al., 2013).
References
Devine, P. G. (1989). Stereotypes and prejudice: Their automatic and controlled components. Journal of Personality and Social Psychology, 56(1), 5-18. https://doi.org/10.1037/0022-3514.56.1.5
Forscher, P. S., Lai, C. K., Axt, J. R., Ebersole, C. R., Herman, M., Devine, P. G., & Nosek, B. A. (2019). A meta-analysis of procedures to change implicit measures. Journal of Personality and Social Psychology, 117(3), 522-559. https://doi.org/10.1037/pspa0000160
Gawronski, B., & Bodenhausen, G. V. (2006). Associative and propositional processes in evaluation: An integrative review of implicit and explicit attitude change. Psychological Bulletin, 132(5), 692-731. https://doi.org/10.1037/0033-2909.132.5.692
Greenwald, A. G., & Banaji, M. R. (1995). Implicit social cognition: Attitudes, self-esteem, and stereotypes. Psychological Review, 102(1), 4-27. https://doi.org/10.1037/0033-295X.102.1.4
Greenwald, A. G., McGhee, D. E., & Schwartz, J. L. K. (1998). Measuring individual differences in implicit cognition: The implicit association test. Journal of Personality and Social Psychology, 74(6), 1464-1480. https://doi.org/10.1037/0022-3514.74.6.1464
Greenwald, A. G., Poehlman, T. A., Uhlmann, E. L., & Banaji, M. R. (2009). Understanding and using the Implicit Association Test: III. Meta-analysis of predictive validity. Journal of Personality and Social Psychology, 97(1), 17-41. https://doi.org/10.1037/a0015575
Kurdi, B., Seitchik, A. E., Axt, J. R., Carroll, T. J., Karapetyan, A., Kaushik, N., ... Banaji, M. R. (2019). Relationship between the Implicit Association Test and intergroup behavior: A meta-analysis. American Psychologist, 74(5), 569-586. https://doi.org/10.1037/amp0000364
Lai, C. K., Marini, M., Lehr, S. A., Cerruti, C., Shin, J. L., Joy-Gaba, J. A., ... Nosek, B. A. (2014). Reducing implicit racial preferences: I. A comparative investigation of 17 interventions. Journal of Experimental Psychology: General, 143(4), 1765-1785. https://doi.org/10.1037/a0036260
Oswald, F. L., Mitchell, G., Blanton, H., Jaccard, J., & Tetlock, P. E. (2013). Predicting ethnic and racial discrimination: A meta-analysis of IAT criterion studies. Journal of Personality and Social Psychology, 105(2), 171-192. https://doi.org/10.1037/a0032734
Payne, B. K., Vuletich, H. A., & Lundberg, K. B. (2017). The bias of crowds: How implicit bias bridges personal and systemic prejudice. Psychological Inquiry, 28(4), 233-248. https://doi.org/10.1080/1047840X.2017.1335568