Abstract
Psychological generalization is a form of learning: once a response is attached to one stimulus, it transfers, in graded fashion, to other stimuli that resemble it. Plotted against similarity, that transfer traces a generalization gradient, the orderly falling-off of responding as a test stimulus departs from the trained one. The same rule works in reverse when discrimination training sharpens the gradient and, under opposing excitatory and inhibitory gradients, shifts its peak away from the non-reinforced value, while Shepard argued that the gradient obeys a near-universal exponential law of distance in an internal psychological space. Generalization is adaptive when it lets past learning meet novel cases, and maladaptive when conditioned fear spreads too widely, a mechanism now central to models of anxiety. Three interactive demonstrations let the reader shape a gradient, produce peak shift, and over-generalize fear.
Keywords: generalization, generalization gradient, discrimination, peak shift, fear generalization
Generalization is the process by which learning about one stimulus comes to govern behavior toward others that resemble it. It is the complement of discrimination: where discrimination is the capacity to respond differently to different stimuli, generalization is the tendency to respond similarly to similar ones, and every act of learning strikes some balance between the two. The phenomenon was among the first that Pavlov recorded, and it has remained central because it answers a basic question about any learned response—how far does it reach? A dog conditioned to salivate to a 1,000 Hz tone will also salivate, though less, to 900 Hz and to 1,100 Hz (Pavlov, 1927). That orderly spread, and the conditions that widen or narrow it, are the subject of this article.
- Generalization is the transfer of a learned response to stimuli that resemble the trained one; it is the mirror image of discrimination.
- Its signature measure is the generalization gradient: responding declines in orderly fashion as a test stimulus grows less similar to the training stimulus.
- Discrimination training sharpens the gradient, and opposing excitatory and inhibitory gradients can shift its peak away from the non-reinforced value—the peak-shift effect.
- Shepard proposed a universal law: generalization falls off exponentially with distance in an internal psychological space, across species and modalities.
- Over-generalization of conditioned fear to safe but similar stimuli is a core mechanism in anxiety disorders and a target of exposure therapy.
What Generalization Is
In its technical sense, generalization is the graded transfer of a learned response across a dimension of similarity. Three features define it. First, it is graded rather than all-or-none: responding does not simply spread to every other stimulus, but declines smoothly as similarity to the trained stimulus decreases. Second, it operates over a dimension—wavelength, tone frequency, size, spatial position, or a more abstract semantic space—along which stimuli can be ordered by resemblance. Third, it is learned in one place and expressed in another: nothing was ever explicitly trained about the test stimuli, so their control over behavior reveals how the original learning was represented.
This makes generalization more than a curiosity of the conditioning laboratory. It is the empirical face of a deep problem: no two situations are ever identical, so any useful learning must apply to cases never encountered during training. Generalization is how a learned response solves that problem, and the shape of the generalization gradient is a readout of the similarity structure the learner imposes on the world (Shepard, 1987). The same construct appears in Pavlovian conditioning, operant behavior, perceptual learning, and category learning; a review spanning a century of work found the gradient to be one of the most robust and general regularities in the study of behavior (Ghirlanda & Enquist, 2003).
Types of Psychological Generalization
MeSH indexes Psychological Generalization under learning and divides it into two direct subtypes, which separate the two things that can vary in a learned relationship: the stimulus that triggers a response, and the response itself. The distinction is orthogonal—an episode of learning can show either, both, or neither—and the MeSH tree is an indexing classification for the literature, not a theoretical claim that these are the only forms generalization takes.
| Subtype | In brief |
|---|---|
| Response Generalization | Transfer to responses that resemble the trained one: a reinforced action spreads to variants of that action, so learning to press a lever hard also raises the rate of moderate presses. The generalized quantity is behavior, not stimulus. |
| Stimulus Generalization | Transfer to stimuli that resemble the trained one: a response conditioned to one cue is elicited by physically similar cues, tracing the classic generalization gradient. This is the more heavily studied of the two subtypes. |
Most of the experimental literature, and most of this article, concerns stimulus generalization, because a stimulus continuum is easy to define and to test point by point. Response generalization is the same idea applied to the output side and is central to how operant conditioning produces flexible, variable behavior rather than a single fixed act (Skinner, 1938).
The Generalization Gradient
The generalization gradient is the function relating response strength to the similarity between a test stimulus and the training stimulus. Its canonical measurement is the post-discrimination gradient of Guttman and Kalish (1956): pigeons were reinforced for pecking a key illuminated at a single wavelength, then tested in extinction across a range of wavelengths. Responding peaked at the trained wavelength and fell away symmetrically on either side, producing the orderly, roughly bell-shaped curve that became the standard picture of generalization. Crucially, the birds had never been reinforced at the test wavelengths; the graded responding was pure transfer.
The gradient's shape is theoretically loaded. A flat gradient means the learner treats a wide range of stimuli as equivalent; a steep gradient means it treats even near neighbors as different. Shepard (1987) analyzed gradients across species, sensory modalities, and stimulus types and argued that when distance is measured in the appropriate internal psychological space—not physical units—the gradient is very generally an exponential decay function of that distance. This universal law of generalization treats the gradient not as a failure of discrimination but as a rational inference: given uncertainty about the size of the region in which a consequence applies, exponential falloff is the optimal way to extend a response to new stimuli.
Figure 1
The Generalization Gradient: Responding Falls Off With Distance From the Trained Stimulus
The first interactive demonstration sets the steepness of an exponential gradient and reads off the predicted response strength at each test stimulus, reproducing the Worked Example below.
Demo 1
Shape a Generalization Gradient
The trained value sits at the center of the dimension, where responding is maximal (Rmax = 100). Drag the slider to change the breadth parameter lambda and watch the gradient sharpen or spread. The table reads off the predicted response at fixed distances, reproducing the Worked Example.
| Distance d (nm) | 0 | 20 | 40 | 60 |
|---|---|---|---|---|
| Response R | 100.0 | 36.8 | 13.5 | 5.0 |
Discrimination and Peak Shift
Generalization and discrimination are two ends of one adjustable process. Discrimination training—reinforcing responses to one stimulus (S+) while withholding reinforcement for another (S-)—systematically narrows the generalization gradient around S+, because the animal learns that the reinforced region is smaller than a single training stimulus would suggest (Pavlov, 1927; Skinner, 1938). The associative account of this narrowing was formalized by Rescorla and Wagner (1972): reinforced and non-reinforced trials build competing associative strengths whose summation across the stimulus dimension reshapes the observed gradient.
The most striking product of discrimination training is the peak-shift effect. After training in which S+ and S- lie close together on a dimension, the peak of responding does not sit at S+ but shifts away from S-, to a value on the far side of S+ that was never reinforced. The standard explanation superimposes an excitatory gradient centered on S+ and an inhibitory gradient centered on S-; where the two are subtracted, the net gradient peaks beyond S+, displaced in the direction opposite the non-reinforced stimulus. Peak shift is important precisely because it shows that generalization is not a passive smearing of a response: the learner extracts a relational rule—respond to the stimulus that is more extreme in the reinforced direction—rather than merely responding most to the exact stimulus it was trained on.
Demo 2
Peak Shift From Discrimination Training
S+ is reinforced at 550 nm; S- is non-reinforced and sits below it. The dashed curve is the raw excitatory gradient; the solid curve is what is observed after subtracting the inhibitory gradient around S-. Move S- closer to S+ and the observed peak shifts farther past S+, to a wavelength that was never itself reinforced.
Generalization of Fear
Nowhere does generalization matter more than in fear. When a stimulus is paired with an aversive event it becomes a conditioned threat, and the fear it evokes generalizes to stimuli that resemble it—a spread that is adaptive when a genuine danger has variable forms, but pathological when it floods safe situations. The founding demonstration is Watson and Rayner's (1920) conditioning of the infant known as Little Albert, whose learned fear of a white rat generalized to a rabbit, a dog, and a fur coat: physically similar stimuli that had never themselves been paired with the aversive noise.
Contemporary work treats over-generalization of conditioned fear as a core mechanism of anxiety disorders. Dunsmoor and Paz (2015) reviewed evidence that anxious individuals show broader fear-generalization gradients—their fear falls off more slowly with distance from the conditioned threat—so that stimuli a healthy learner would treat as safe continue to evoke fear. At the neural level, Asok, Kandel, and Rayman (2019) describe how the balance among the amygdala, hippocampus, and prefrontal cortex sets the breadth of the fear gradient, with the hippocampus supporting the pattern separation that keeps similar-but-safe stimuli distinct from the threat. This reframing has direct clinical consequences: because a fear is not erased but inhibited by new learning, it can return when context changes (Bouton, 2004; Bouton, Maren, & McNally, 2021), and effective exposure therapy is best understood as building inhibitory learning that competes with an over-broad fear gradient rather than as deleting the original association (Craske, Hermans, & Vervliet, 2018).
Demo 3
Over-generalization of Fear
The center bar is the conditioned threat (CS+); the other eight stimuli are progressively less similar and are objectively safe. Widen the fear gradient and watch fear spill across the dashed avoidance threshold onto safe stimuli. Breadth, not peak height, is what turns adaptive caution into over-generalization.
Theories of Generalization
Three broad accounts of why generalization takes the form it does have organized the field, and they are complementary rather than exclusive. Table 2 summarizes them.
| Account | Core claim | Explains the gradient as | Key source |
|---|---|---|---|
| Stimulus-element / associative | Similar stimuli share sensory elements that carry associative strength | Summation of excitatory and inhibitory strength across shared elements | Rescorla & Wagner, 1972 |
| Universal law (similarity) | Generalization reflects distance in an internal psychological space | Exponential decay with psychological distance, optimal under uncertainty | Shepard, 1987 |
| Ecological / comparative | Gradients are shaped by the statistics of natural stimulus variation | An adaptive response calibrated across species and tasks | Ghirlanda & Enquist, 2003 |
The associative account, rooted in the tradition of classical conditioning, treats the gradient as the summed output of learning about the sensory elements that stimuli share; discrimination and peak shift follow directly from the competition between excitatory and inhibitory strengths (Rescorla & Wagner, 1972). Shepard's universal law abstracts away from the sensory details, locating the regularity in the geometry of an internal similarity space and arguing that exponential falloff is the rational policy for a learner uncertain about how far a consequence extends (Shepard, 1987). The ecological view, drawing on a century of comparative data, stresses that the breadth an animal actually adopts is tuned to the natural variability of the stimuli it must respond to (Ghirlanda & Enquist, 2003). Together they frame generalization as neither error nor accident but a calibrated inference about similarity.
Worked Example
The first demonstration computes an exponential generalization gradient exactly, following the form Shepard proposed. Let response strength at a test stimulus be R = Rmax × exp(−d / λ), where d is the distance of the test stimulus from the trained value, λ (lambda) is a breadth parameter setting how slowly the gradient decays, and Rmax is the response strength at the trained value. A large λ makes a broad, shallow gradient; a small λ makes a steep, narrow one.
Take a pigeon trained at 550 nm with Rmax = 100 pecks and a breadth of λ = 20 nm. At the trained value, d = 0, so R = 100 × exp(0) = 100. At a test stimulus 20 nm away (d = 20), R = 100 × exp(−20/20) = 100 × exp(−1) = 100 × 0.368 = 36.8. At 40 nm away, R = 100 × exp(−2) = 100 × 0.135 = 13.5, and at 60 nm away, R = 100 × exp(−3) = 5.0. Responding thus falls to about a third of its peak for every 20 nm of distance—the constant proportional decay that is the signature of an exponential gradient.
Now compare two learners tested at the same distance of d = 30 nm. A sharp discriminator with λ = 15 responds R = 100 × exp(−30/15) = 100 × exp(−2) = 13.5. A broad generalizer with λ = 45 responds R = 100 × exp(−30/45) = 100 × exp(−0.667) = 51.3. The same physical difference produces nearly four times the transfer in the broad learner—a quantitative statement of what it means for an anxious individual to over-generalize fear: not that the peak fear is higher, but that the gradient is shallower, so distant and objectively safe stimuli still command a substantial response (Dunsmoor & Paz, 2015).
Discussion
Treating generalization as a single construct clarifies a body of findings that are otherwise scattered across the study of conditioning, perception, and psychopathology. The generalization gradient is the common measurement; its steepness is the common variable; and discrimination, peak shift, and fear over-generalization are all statements about what widens or narrows it. The associative, universal-law, and ecological accounts are not rivals so much as descriptions at different levels—mechanism, computation, and function—of why a learned response reaches as far as it does.
The construct also reorganizes how a learned response's scope is understood. The intuitive view treats a conditioned response as attached to one stimulus, with any spread to others as noise. The generalization literature shows the opposite: spread is the rule, orderly, lawful, and often optimal, and it is discrimination that must be trained to contain it (Guttman & Kalish, 1956). This has a direct practical payoff in the clinic, where the goal of exposure therapy is not to erase a fear but to build competing inhibitory learning that narrows an over-broad gradient, an effort complicated by the context-dependence of what is learned (Bouton, 2004; Craske et al., 2018). It also warns against a symmetric error: a fear that appears contained in the therapy room may generalize poorly to the settings where it was acquired, so that apparent success is bounded by the very similarity structure the gradient describes.
Current Directions
The most active current work on generalization is at the intersection of computation and neuroscience. The universal-law tradition has been extended by treating generalization as Bayesian inference about the extent of a consequence, which recovers Shepard's exponential gradient as an optimal policy and predicts how the gradient should change with the learner's uncertainty (Tenenbaum & Griffiths, 2001). In parallel, the fear-generalization literature has become a translational model for anxiety: work mapping the amygdala–hippocampus–prefrontal circuit has begun to specify how pattern separation in the hippocampus keeps similar-but-safe stimuli distinct from a conditioned threat, and how a failure of that separation broadens the fear gradient (Asok et al., 2019). A convergent clinical line reframes exposure therapy in terms of inhibitory learning that must itself generalize across contexts to be durable, turning the context-dependence of extinction from a nuisance into a design principle for treatment (Craske et al., 2018; Bouton et al., 2021). Open questions concern how generalization over abstract and conceptual dimensions—category membership, meaning—relates to the perceptual gradients studied classically, and whether a single similarity space can span both.
Common Misconceptions
- Generalization is just a failure to discriminate.
- Generalization is an orderly, often optimal inference, not an error. Shepard showed that exponential falloff with psychological distance is the rational way to extend a response under uncertainty about how far a consequence applies (Shepard, 1987); discrimination is the process that narrows this default, not a competence that generalization lacks (Guttman & Kalish, 1956).
- Responding always peaks at the exact stimulus that was trained.
- After discrimination training with a nearby non-reinforced stimulus, the peak of responding shifts away from the non-reinforced value to a stimulus that was never reinforced—the peak-shift effect—showing that learners extract a relational rule rather than a fixed stimulus–response bond (Rescorla & Wagner, 1972).
- A high level of fear is what makes anxiety pathological.
- What distinguishes clinical fear is often the breadth of the generalization gradient, not its height: fear that falls off too slowly with distance from a genuine threat spreads to safe stimuli, so a shallow gradient—not a taller peak—is the mark of over-generalization (Dunsmoor & Paz, 2015).
Glossary
- Conditioned stimulus.
- A stimulus that, through pairing with a biologically significant event, comes to elicit a learned response; it is the stimulus from which conditioned responding generalizes.
- Discrimination.
- The capacity to respond differently to different stimuli; the complement of generalization, trained by reinforcing one stimulus while withholding reinforcement for another.
- Excitatory gradient.
- The generalization gradient of response strength centered on a reinforced stimulus (S+), reflecting learned tendencies to respond.
- Fear generalization.
- The spread of a conditioned fear response to stimuli that resemble the conditioned threat; a broad gradient extends fear to safe stimuli and figures in anxiety disorders.
- Generalization gradient.
- The function relating response strength to the similarity between a test stimulus and the training stimulus; its steepness indexes how narrowly the response is tuned.
- Generalization.
- The graded transfer of a learned response to stimuli, or responses, that resemble the trained one along a dimension of similarity.
- Inhibitory gradient.
- The gradient of learned suppression centered on a non-reinforced stimulus (S-); its subtraction from the excitatory gradient produces peak shift.
- Pattern separation.
- A hippocampal process that renders similar inputs as distinct representations, keeping similar-but-safe stimuli separable from a conditioned threat and thereby limiting fear generalization.
- Peak shift.
- The displacement of the peak of responding away from the reinforced stimulus, to the side opposite a nearby non-reinforced stimulus, following discrimination training.
- Psychological space.
- The internal representation of similarity in which, per Shepard, generalization declines exponentially with distance, regardless of the physical units of the stimulus.
- Response generalization.
- The transfer of learning to responses that resemble the trained response, so reinforcing one action raises the probability of similar actions.
- Stimulus generalization.
- The transfer of a conditioned response to stimuli physically similar to the training stimulus, tracing the classic generalization gradient.
- Summation.
- The combination of excitatory and inhibitory associative strengths across shared stimulus elements, whose net value the associative account identifies with observed responding.
- Universal law of generalization.
- Shepard's proposal that generalization is an exponential-decay function of distance in psychological space, holding across species, modalities, and stimulus types.
Key Researchers
Mark E. Bouton (b. 1953). Professor of Psychological Science at the University of Vermont; showed that what is learned is bound to its context, so an extinguished response renews when the context changes—directly constraining how far and how durably learning generalizes. ORCID - Faculty Page - Google Scholar
Michelle G. Craske (b. 1959). Professor of Psychology at the University of California, Los Angeles; reframed clinical fear as impaired inhibitory learning and over-broad generalization, developing inhibitory-learning approaches to exposure therapy. ORCID - Wikipedia - Google Scholar
Joseph E. Dunsmoor. Associate Professor of Neuroscience at the University of Texas at Austin; maps the behavioral and neural mechanisms of fear generalization and its bearing on anxiety, showing that anxious individuals adopt broader fear gradients. Faculty Page - Google Scholar
Ivan P. Pavlov (1849-1936). Physiologist at the Imperial Military Medical Academy, St. Petersburg; first documented stimulus generalization, showing that a conditioned salivary response transferred in graded fashion to tones near the trained one and that discrimination training sharpened the gradient. Wikipedia - Wikidata
Roger N. Shepard (1929-2022). Professor Emeritus of Psychology at Stanford University; proposed the universal law of generalization, holding that response probability declines exponentially with distance in an internal psychological space across species and modalities. Wikipedia - Faculty Page
Burrhus Frederic Skinner (1904-1990). Behaviorist at Harvard University; extended generalization to operant behavior, showing that a reinforced response is emitted to physically similar stimuli along a gradient and that differential reinforcement narrows it. Wikipedia - Wikidata
Frequently Asked Questions
What is generalization in psychology?
Generalization is the graded transfer of a learned response to stimuli, or responses, that resemble the trained one; a response conditioned to one stimulus is elicited, though more weakly, by physically similar stimuli (Pavlov, 1927).
What is a generalization gradient?
It is the function relating response strength to the similarity between a test stimulus and the training stimulus, typically peaking at the trained value and declining on either side, as first mapped systematically by Guttman and Kalish (1956).
How are generalization and discrimination related?
They are complementary: generalization is the tendency to respond similarly to similar stimuli, while discrimination is the trained capacity to respond differently to them; discrimination training narrows the generalization gradient (Guttman & Kalish, 1956).
What is the peak-shift effect?
After discrimination training in which a non-reinforced stimulus lies near the reinforced one, the peak of responding shifts away from the non-reinforced value to a stimulus that was never reinforced, reflecting the subtraction of an inhibitory gradient from an excitatory one (Rescorla & Wagner, 1972).
What is Shepard's universal law of generalization?
It is the proposal that generalization declines exponentially with distance in an internal psychological space, a regularity that holds across species, sensory modalities, and stimulus types and can be derived as an optimal inference under uncertainty (Shepard, 1987).
What is the difference between stimulus and response generalization?
Stimulus generalization is transfer to similar stimuli, whereas response generalization is transfer to similar responses; reinforcing one action raises the probability of related actions (Skinner, 1938).
How does fear generalization relate to anxiety?
Anxiety disorders are associated with broader fear-generalization gradients, so conditioned fear falls off slowly with distance from a genuine threat and spreads to safe but similar stimuli (Dunsmoor & Paz, 2015).
Why can a fear return after successful therapy?
Extinction inhibits rather than erases a conditioned association, so a fear can renew when the context changes; durable treatment depends on building inhibitory learning that itself generalizes across contexts (Bouton, 2004; Craske et al., 2018).
References
Asok, A., Kandel, E. R., & Rayman, J. B. (2019). The neurobiology of fear generalization. Frontiers in Behavioral Neuroscience, 12, 329. https://doi.org/10.3389/fnbeh.2018.00329
Bouton, M. E. (2004). Context and behavioral processes in extinction. Learning & Memory, 11(5), 485-494. https://doi.org/10.1101/lm.78804
Bouton, M. E., Maren, S., & McNally, G. P. (2021). Behavioral and neurobiological mechanisms of Pavlovian and instrumental extinction learning. Physiological Reviews, 101(2), 611-681. https://doi.org/10.1152/physrev.00016.2020
Craske, M. G., Hermans, D., & Vervliet, B. (2018). State-of-the-art and future directions for extinction as a translational model for fear and anxiety. Philosophical Transactions of the Royal Society B, 373(1742), 20170025. https://doi.org/10.1098/rstb.2017.0025
Dunsmoor, J. E., & Paz, R. (2015). Fear generalization and anxiety: Behavioral and neural mechanisms. Biological Psychiatry, 78(5), 336-343. https://doi.org/10.1016/j.biopsych.2015.04.010
Ghirlanda, S., & Enquist, M. (2003). A century of generalization. Animal Behaviour, 66(1), 15-36. https://doi.org/10.1006/anbe.2003.2174
Guttman, N., & Kalish, H. I. (1956). Discriminability and stimulus generalization. Journal of Experimental Psychology, 51(1), 79-88. https://doi.org/10.1037/h0046219
Pavlov, I. P. (1927). Conditioned reflexes: An investigation of the physiological activity of the cerebral cortex (G. V. Anrep, Trans.). Oxford University Press.
Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), Classical conditioning II: Current research and theory (pp. 64-99). Appleton-Century-Crofts.
Shepard, R. N. (1987). Toward a universal law of generalization for psychological science. Science, 237(4820), 1317-1323. https://doi.org/10.1126/science.3629243
Skinner, B. F. (1938). The behavior of organisms: An experimental analysis. Appleton-Century.
Tenenbaum, J. B., & Griffiths, T. L. (2001). Generalization, similarity, and Bayesian inference. Behavioral and Brain Sciences, 24(4), 629-640. https://doi.org/10.1017/S0140525X01000061
Watson, J. B., & Rayner, R. (1920). Conditioned emotional reactions. Journal of Experimental Psychology, 3(1), 1-14. https://doi.org/10.1037/h0069608