Abstract
Discrimination learning, a type of learning, is the process by which an organism comes to respond differently to two or more stimuli, typically because responding to one is reinforced while responding to another is not. It is a foundational paradigm of experimental psychology and the operation by which stimulus control is established. Early work divided over whether discrimination is acquired gradually through the summation of excitatory and inhibitory tendencies or suddenly through the testing of hypotheses. The mature theory treats a discrimination as the difference between a gradient of approach centred on the reinforced stimulus and a gradient of avoidance centred on the unreinforced one — a subtraction that predicts the counterintuitive peak shift — with later accounts adding selective attention to the relevant stimulus dimension. Three interactive demonstrations develop the generalization gradient, errorless training, and attentional weighting.
Keywords: discrimination learning, stimulus control, generalization gradient, peak shift, selective attention
Discrimination learning is learning that is manifested in the ability to respond differentially to different stimuli, and it is the experimental operation through which a stimulus comes to control a response (Spence, 1936). The procedure is simple in outline: one stimulus, conventionally written S+ or CS+, is paired with reinforcement, and a second, written S− or CS−, is not, so that over trials the organism comes to respond to the first and withhold responding to the second. What makes the paradigm central rather than merely useful is that it isolates the question of stimulus control — what feature of the environment a response has come under the government of — from the separate question of how a response is strengthened. A pigeon that pecks a key lit at 550 nanometres but not one lit at 560 has done more than acquire a peck; it has come to treat two points on a continuous physical dimension as different occasions for behaviour, and the shape of that difference is measurable. The earliest theories disagreed sharply about the underlying process. The continuity tradition held that discrimination is built gradually, trial by trial, as reinforcement accrues to the reinforced stimulus and extinction to the other (Spence, 1936), whereas the noncontinuity tradition held that animals entertain and reject hypotheses, so that learning appears suddenly when the correct hypothesis is adopted (Krechevsky, 1932). The sections below trace this debate, develop the generalization gradient and the stimulus control it reveals, derive the peak shift from Spence's summation theory, present errorless training and the attentional theories that followed, and close with the models that treat a discrimination as a shift of attention across stimulus dimensions.
- Discrimination learning is the acquisition of differential responding to two or more stimuli, and it is the operation that brings a response under the control of a specific stimulus.
- Early theory split between continuity accounts, in which discrimination accrues gradually, and noncontinuity accounts, in which animals test discrete hypotheses and learn abruptly.
- A generalization gradient plots responding against stimulus values around a trained stimulus; its width measures how tightly a response is controlled by that stimulus.
- Spence's summation theory models a discrimination as an excitatory gradient around S+ minus an inhibitory gradient around S−, and this subtraction predicts the peak shift: the strongest response moves away from S+ to the side opposite S−.
- Errorless training can establish a discrimination with almost no responses to S−, and doing so removes the peak shift and other by-products, showing that those by-products depend on the errors ordinary training permits.
The Discrimination Learning Task
A discrimination task presents at least two stimuli and arranges different consequences for responding to each. In the operant form a reinforced stimulus S+ signals that responding will be reinforced and an unreinforced stimulus S− signals that it will not, and differential responding is taken as the index of learning; in the classical form a CS+ is paired with an outcome and a CS− is presented without it. Because the two stimuli can be placed at chosen points on a physical continuum — two wavelengths, two line orientations, two tone frequencies — the paradigm converts a question about knowledge into a question about a measurable function, which is why it became the workbench on which theories of learning were tested (Spence, 1936). MeSH classifies discrimination learning as a form of learning, a descriptor exactly matching the concept treated here (D004193), and its long experimental history reflects that generality: the same procedure has been run with rats in mazes, pigeons at keys, monkeys at test trays, and humans in categorization studies.
The founding controversy concerned whether the learning is continuous or not. Spence's continuity theory held that every reinforced trial adds a small increment of excitatory tendency to the stimulus present and every unreinforced trial subtracts one, so that a discrimination is the slow separation of two response tendencies and any apparent suddenness is an artefact of when the tendencies cross a threshold (Spence, 1936). Against this, Krechevsky reported that rats in an insoluble discrimination showed systematic patterns of choice — position preferences, alternation — rather than the random responding a gradual account seemed to predict, and argued that the animal advances and abandons hypotheses until one is confirmed, so that learning is discontinuous (Krechevsky, 1932). The dispute was never settled by knockout; both processes turned out to operate, and the modern view retains the continuity theory's incremental machinery while granting that selective attention can make some dimensions effectively invisible until the animal attends to them. Evidence that discrimination involves more than the strengthening of a single response came from Harlow's demonstration that monkeys given hundreds of distinct two-object discriminations improve at learning itself, solving each new problem faster until they succeed in one trial — a learning set, or learning to learn, that no account of a single accruing association can explain (Harlow, 1949).
Generalization Gradients and Stimulus Control
Once a response is established to S+, testing the organism with a range of stimuli along the same dimension reveals how tightly the response is bound to the trained value. The function relating response strength to stimulus value is the generalization gradient, and its shape is the primary datum of stimulus control. Guttman and Kalish trained pigeons to peck a key of a single wavelength and then, in extinction, presented wavelengths spanning the visible spectrum; responding fell off in an orderly, roughly bell-shaped gradient centred on the training value, with more responses the closer a test wavelength lay to it (Guttman & Kalish, 1956). The gradient's existence shows that reinforcement at one value spreads its effect to neighbouring values, and its width indexes the precision of control: a narrow gradient means the response is sharply tuned to S+, a broad one that it generalizes widely. The formal tradition treats this falloff as a decreasing function of distance in an underlying psychological space, so that generalization between two stimuli reflects their similarity rather than their raw physical separation (Shepard, 1957).
A single physical stimulus is in fact a bundle of features, and discrimination training determines which of them acquires control. Reynolds trained two pigeons on a compound stimulus — a white triangle on a coloured background — reinforcing pecks to it and not to a different compound, and then tested the elements separately; one bird's responding proved to be controlled almost entirely by the colour and the other's by the triangle, although both had been reinforced in the presence of the whole compound (Reynolds, 1961). The result is a demonstration of selective attention in discrimination: reinforcement in the presence of a compound does not guarantee that every element gains control, and which element wins can differ between individuals trained identically. Stimulus control is therefore not read off the physical stimulus but is the joint product of the reinforcement contingency and the dimensions the organism attends to.
Peak Shift and the Summation Theory
The most theoretically productive phenomenon in the field is the peak shift, and it falls directly out of Spence's account of how a discrimination is composed. Spence proposed that reinforcing S+ builds an excitatory generalization gradient centred on S+, that not reinforcing S− builds an inhibitory gradient centred on S−, and that the net tendency to respond at any value is the excitatory gradient minus the inhibitory one (Spence, 1937). Because the inhibitory gradient centred on S− subtracts more from the near side of S+ than from the far side, the maximum of the net function is pulled away from S+, to the side opposite S−. The prediction is counterintuitive: after training to respond to S+ and not to S−, the organism responds most strongly not to S+ itself but to a value displaced beyond it, away from S−. Hanson confirmed this directly, training pigeons with S+ at 550 nanometres and S− at 555, 560, 570, or 590 nanometres and finding that the peak of the post-discrimination gradient shifted to wavelengths shorter than 550 — away from S− — with the size of the shift depending on how close S− had been (Hanson, 1959). The effect is not a curiosity of pigeon vision; it is a signature of the subtraction Spence described, and it has been reported across a range of species and stimulus dimensions when an intradimensional discrimination is trained. The same summation has a companion effect on the rate of responding: introducing an unreinforced S− also raises the rate of responding to S+ above its pre-discrimination level — behavioral contrast — the response-rate face of the summation that produces peak shift (Reynolds, 1961b). Figure 1 shows the composition, and the demonstration below lets the two gradients and their difference be manipulated so that the shifted peak emerges from the subtraction rather than being asserted.
Figure 1
Peak Shift as the Difference of an Excitatory and an Inhibitory Gradient
Model It
Where the Response Peaks After Discrimination Training
Reinforcing S+ builds an excitatory gradient; not reinforcing S- builds an inhibitory one. The net tendency to respond is their difference, and its peak sits on the far side of S+ from S-, the peak shift. Move the two sliders and watch where the peak lands: with equal-width gradients the shift is largest at an intermediate separation, and it shrinks both when S- is far (inhibition barely reaches the peak) and when S- is very close (inhibition subtracts almost symmetrically about S+).
Errorless Discrimination
Ordinary discrimination training allows the organism to respond to S− many times before the discrimination is complete, and Terrace showed that those errors are not inevitable. By introducing S− early, faintly, and briefly — dim and short at first, then gradually brightened and lengthened as training proceeded, a technique of fading — he trained pigeons to discriminate red from green and later a vertical from a horizontal line with almost no responses ever made to S− (Terrace, 1963). The finding mattered for more than efficiency. Terrace observed that discriminations learned without errors lacked properties that normally accompany discrimination learning: the birds trained errorlessly showed little or no peak shift and none of the emotional by-products, such as the burst of responding to S+ that follows exposure to S−, that errorful training produces. Because those phenomena were absent exactly when errors were absent, they could be attributed to the experience of non-reinforcement in the presence of S− rather than to the discrimination itself. Errorless learning thus dissociated the core achievement — differential responding — from its usual accompaniments, and it did so by manipulating the one variable, responding to S−, that the summation theory makes responsible for the inhibitory gradient. The demonstration below contrasts the accumulation of errors under abrupt and faded introductions of S−.
Compare
Errors With and Without Fading
A discrimination can be trained by presenting S- abruptly at full strength, or by fading it in, faint and brief at first and stronger over blocks. Lengthen the fade and the learner makes almost no responses to S-, because S- is never salient while the tendency to respond is still high.
Errors, abrupt S-
Errors, faded S-
Attention, Associability, and Similarity
That identically trained animals can come under the control of different elements (Reynolds, 1961) pushed learning theory toward an explicit role for attention, and two formal traditions supplied one. Mackintosh proposed that the associability of a stimulus — the rate at which it gains associative strength — is not fixed but varies with how good a predictor the stimulus has been: cues that reliably predict reinforcement command attention and are learned about quickly, while cues that predict nothing lose associability and are increasingly ignored (Mackintosh, 1975). On this account a discrimination is learned in part by learning where to look, and the difficulty of a discrimination reflects how readily the relevant dimension can win attention from competing ones. The idea also refines the associative machinery inherited from continuity theory: the Rescorla–Wagner model had already formalized how S+ acquires excitatory strength and S− acquires inhibitory strength through competition among cues for a limited pool of associative strength on each trial, explaining why a discrimination depends on the relative validity of the stimuli rather than their individual pairings (Rescorla & Wagner, 1972). Mackintosh's contribution was to let attention to a cue itself change with the cue's predictiveness.
The attentional idea reappears in the modeling of human categorization, where selective attention to a stimulus dimension is treated as a stretching of psychological space along that dimension. Building on the principle that generalization declines with distance in a similarity space (Shepard, 1957), Nosofsky's generalized context model represents each stimulus as a point in a multidimensional space and lets the learner place attention weights on the dimensions; increasing the weight on a dimension expands distances along it, so that stimuli differing on an attended dimension become more discriminable while those differing only on an ignored dimension remain confusable (Nosofsky, 1986). A discrimination that is hard because the relevant difference is small can thus be made easy by attending to — and so perceptually magnifying — the dimension that carries it. Table 1 sets the principal theoretical positions side by side, and the demonstration lets attention be weighted across two dimensions so that its effect on discriminability can be seen directly.
Table 1
Principal Theoretical Accounts of Discrimination Learning
| Account | Core claim | Signature evidence |
|---|---|---|
| Continuity (Spence) | Discrimination accrues gradually as excitation adds to S+ and inhibition to S− | Smooth generalization gradients; peak shift from summation |
| Noncontinuity (Krechevsky) | Animals test and discard hypotheses; learning is abrupt | Systematic pre-solution choice patterns |
| Summation (Spence) | Net response = excitatory gradient around S+ minus inhibitory gradient around S− | Peak shift away from S− (Hanson) |
| Cue competition (Rescorla–Wagner) | Cues compete for limited associative strength; relative validity governs learning | Discrimination depends on relative, not absolute, predictiveness |
| Attention (Mackintosh) | Associability rises for predictive cues and falls for non-predictive ones | Different elements control identically trained animals |
| Similarity space (Nosofsky) | Attention weights stretch a psychological space, raising discriminability on attended dimensions | Attention-weighted fits to human categorization |
Note. The accounts are layered rather than strictly rival: the summation theory elaborates continuity, and the attentional models add a mechanism for which dimension the incremental learning acts on (Mackintosh, 1975; Nosofsky, 1986).
Try It
Attend the Dimension That Carries the Difference
Two stimuli sit in a psychological space, far apart on the dimension that predicts reinforcement and close on one that does not. Attention weights the dimensions: sliding attention onto the relevant dimension stretches the space along it, pulling the pair apart and making them easy to tell apart.
Weighted distance
Discrimination accuracy
Worked Example
The peak shift can be derived exactly from Spence's summation theory, and doing so shows that the shifted maximum is a consequence of subtraction rather than an extra assumption. Model the excitatory gradient as a Gaussian centred on S+ and the inhibitory gradient as a smaller Gaussian centred on S−, and let the net tendency to respond at a wavelength be their difference where it is positive. Take S+ at 550 nanometres and S− at 560 nanometres, ten nanometres to the long-wavelength side. Give the excitatory gradient a height of 1.00 and the inhibitory gradient a height of 0.50, and let both have the same spread, with a standard deviation of 25 nanometres, so that the only asymmetry in the problem is the ten-nanometre gap between the two centres. Before any discrimination training the excitatory gradient alone governs responding and its maximum sits exactly at S+, 550 nanometres. After training, the inhibitory gradient centred on S− is subtracted. At 550 nanometres the excitatory gradient is at its full height of 1.000 while the inhibitory gradient contributes 0.462, giving a net of 0.538; at 545 nanometres the excitatory gradient has fallen only slightly to 0.980 while the inhibitory gradient has fallen further to 0.418, giving a net of 0.563; at 543 nanometres the net reaches its maximum. Sweeping the net function across the spectrum in tenth-nanometre steps locates the peak at 543.0 nanometres, a shift of 7.0 nanometres away from S+ in the direction opposite S−. The organism trained to respond to 550 and not to 560 therefore responds most strongly to 543, a wavelength it was never reinforced at, because that is where the excess of excitation over inhibition is greatest. The size of the shift depends on the separation between S+ and S−: when S− lies far away the inhibitory gradient scarcely reaches the excitatory peak and the shift is small, and bringing S− nearer enlarges the near-flank subtraction until, with two equal-width gradients, the shift reaches a maximum at an intermediate separation and then contracts again once S− is so close that inhibition subtracts almost symmetrically about S+. Hanson found the shift growing as S− approached S+ across the separations he tested, an ordering the equal-width model reproduces only over the wider separations and only if the inhibitory gradient is allowed to be the steeper of the two (Hanson, 1959; Spence, 1937).
Discussion
Discrimination learning earned its central place because it made stimulus control measurable and therefore theorizable. The generalization gradient turned a question about what an animal knows into the shape of a function, and once the function existed, its width, its centre, and its movement under training became data that a theory had to reproduce (Guttman & Kalish, 1956). Spence's summation theory met that demand with unusual economy: from the single idea that a discrimination is an excitatory gradient minus an inhibitory one, it derived the peak shift, a phenomenon no one would have proposed from intuition and that ordinary reinforcement accounts do not predict (Spence, 1937; Hanson, 1959). The theory's reach was then tested at its own joints. Errorless training showed that the peak shift and its emotional accompaniments are contingent on responding to S−, exactly the term the theory assigns to the inhibitory gradient, so that removing the errors removes the by-products and confirms the mechanism by its absence (Terrace, 1963). The later history is a steady incorporation of attention. That identically trained animals come under the control of different features (Reynolds, 1961) forced learning theory to specify not only how strength accrues but to what it accrues, and the attentional theories answered by letting a cue's associability track its predictiveness (Mackintosh, 1975) and by treating attention as a deformation of a similarity space in which categorization occurs (Nosofsky, 1986). What survives across the century of work is a layered picture rather than a winner. The incremental machinery of continuity theory remains the substrate; the summation of opposed gradients explains the geometry of stimulus control; and selective attention determines which dimension that machinery operates on. A paradigm introduced to adjudicate whether animals learn gradually or by insight (Spence, 1936; Krechevsky, 1932) became the setting in which the modern understanding of stimulus control, and much of the formal modeling of human categorization, was worked out.
Glossary
- Associability.
- The rate at which a stimulus gains associative strength; in Mackintosh's theory it rises for cues that predict reinforcement and falls for cues that do not, making it a variable of attention.
- Continuity theory.
- The view that a discrimination is acquired gradually, each reinforced trial adding a small increment of excitatory tendency and each unreinforced trial adding inhibition, with no sudden insight.
- Discrimination learning.
- Learning manifested in the ability to respond differentially to two or more stimuli, typically because responding to one is reinforced while responding to another is not.
- Errorless discrimination.
- A discrimination trained by introducing S− early, faint, and brief and then strengthening it gradually, so that the organism makes almost no responses to S− during acquisition.
- Excitatory gradient.
- The generalization gradient of the tendency to respond, centred on the reinforced stimulus S+ and declining with distance from it.
- Fading.
- The training technique of presenting S− at low intensity or short duration at first and increasing it across trials, the method by which errorless discrimination is produced.
- Generalization gradient.
- The function relating response strength to stimulus value around a trained stimulus; its width indexes how tightly the response is controlled by that stimulus.
- Inhibitory gradient.
- The generalization gradient of the tendency to withhold responding, centred on the unreinforced stimulus S− and subtracted from the excitatory gradient in summation theory.
- Intradimensional discrimination.
- A discrimination in which S+ and S− differ along a single physical dimension, such as two wavelengths, the arrangement under which the peak shift is obtained.
- Learning set.
- An acquired improvement in the ability to learn problems of a given class, such that new discriminations are solved ever faster; Harlow's learning to learn.
- Noncontinuity theory.
- The view that animals solve a discrimination by testing and discarding hypotheses, so that learning appears abruptly when the correct hypothesis is adopted.
- Peak shift.
- The displacement of the maximum of the post-discrimination gradient away from S+, to the side opposite S−, predicted by subtracting the inhibitory gradient from the excitatory one.
- Selective attention.
- The differential weighting of stimulus dimensions, so that reinforcement in the presence of a compound gives control to the attended element rather than to all of them equally.
- Stimulus control.
- The degree to which the presence or value of a stimulus governs the occurrence of a response, measured by the sharpness of the generalization gradient.
- Summation theory.
- Spence's account in which the net tendency to respond at a stimulus value is the excitatory gradient around S+ minus the inhibitory gradient around S−.
Key Researchers
Harry F. Harlow (1905-1981). Psychologist at the University of Wisconsin-Madison who demonstrated the learning set, showing that discrimination is accompanied by an acquired improvement in the capacity to learn. Wikipedia
Nicholas J. Mackintosh (1935-2015). Experimental psychologist at the University of Cambridge who developed the attentional theory of associability, in which a cue's learning rate tracks how well it predicts reinforcement. Wikipedia
Kenneth W. Spence (1907-1967). Learning theorist at the University of Iowa who formulated the continuity theory of discrimination and the summation account from which the peak shift is derived. Wikipedia
Herbert S. Terrace (b. 1936). Psychologist at Columbia University who introduced errorless discrimination learning and showed that the peak shift depends on responses made to S−. Faculty Page - Wikipedia
Frequently Asked Questions
What is discrimination learning?
Discrimination learning is learning that is manifested in the ability to respond differentially to two or more stimuli, usually established by reinforcing responses to one stimulus, S+, and not to another, S−, so that responding comes under the control of the reinforced stimulus (Spence, 1936).
What is the difference between continuity and noncontinuity theories?
Continuity theory holds that a discrimination accrues gradually as reinforcement adds excitation to S+ and non-reinforcement adds inhibition to S− (Spence, 1936), whereas noncontinuity theory holds that animals test and discard hypotheses and learn abruptly when the correct one is found (Krechevsky, 1932).
What is a generalization gradient?
A generalization gradient is the function relating response strength to stimulus value around a trained stimulus; pigeons trained at one wavelength respond in an orderly, bell-shaped gradient centred on it, and the gradient's width measures how tightly the response is controlled by that stimulus (Guttman & Kalish, 1956).
What is the peak shift?
The peak shift is the finding that after training to respond to S+ and not to a nearby S−, an organism responds most strongly not to S+ but to a value displaced beyond it, away from S−, as Hanson showed with pigeons trained on wavelength (Hanson, 1959).
How does Spence's summation theory explain the peak shift?
Summation theory models the net tendency to respond as an excitatory gradient centred on S+ minus an inhibitory gradient centred on S−; because inhibition subtracts more from the S− side, the maximum of the difference is pulled to the far side of S+, producing the peak shift (Spence, 1937).
What is errorless discrimination learning?
Errorless discrimination learning is a technique in which S− is introduced faint and brief and then gradually strengthened, so that the organism makes almost no responses to S−; Terrace found that discriminations trained this way lack the peak shift and its emotional by-products (Terrace, 1963).
What role does attention play in discrimination learning?
Reinforcement in the presence of a compound stimulus gives control to the attended element rather than to all elements equally, so identically trained animals can come under the control of different features, and Mackintosh's theory lets a cue's associability rise or fall with how well it predicts reinforcement (Mackintosh, 1975).
How is discrimination learning modeled in human categorization?
Nosofsky's generalized context model represents stimuli as points in a similarity space and lets learners place attention weights on dimensions, so that attending to a dimension expands distances along it and makes stimuli that differ on it more discriminable (Nosofsky, 1986).
References
Guttman, N., & Kalish, H. I. (1956). Discriminability and stimulus generalization. Journal of Experimental Psychology, 51(1), 79-88. https://doi.org/10.1037/h0046219
Hanson, H. M. (1959). Effects of discrimination training on stimulus generalization. Journal of Experimental Psychology, 58(5), 321-334. https://doi.org/10.1037/h0042606
Harlow, H. F. (1949). The formation of learning sets. Psychological Review, 56(1), 51-65. https://doi.org/10.1037/h0062474
Krechevsky, I. (1932). "Hypotheses" in rats. Psychological Review, 39(6), 516-532. https://doi.org/10.1037/h0073500
Mackintosh, N. J. (1975). A theory of attention: Variations in the associability of stimuli with reinforcement. Psychological Review, 82(4), 276-298. https://doi.org/10.1037/h0076778
Nosofsky, R. M. (1986). Attention, similarity, and the identification-categorization relationship. Journal of Experimental Psychology: General, 115(1), 39-57. https://doi.org/10.1037/0096-3445.115.1.39
Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), Classical conditioning II: Current research and theory (pp. 64-99). Appleton-Century-Crofts.
Reynolds, G. S. (1961). Attention in the pigeon. Journal of the Experimental Analysis of Behavior, 4(3), 203-208. https://doi.org/10.1901/jeab.1961.4-203
Reynolds, G. S. (1961). Behavioral contrast. Journal of the Experimental Analysis of Behavior, 4(1), 57-71. https://doi.org/10.1901/jeab.1961.4-57
Shepard, R. N. (1957). Stimulus and response generalization: A stochastic model relating generalization to distance in psychological space. Psychometrika, 22(4), 325-345. https://doi.org/10.1007/BF02288967
Spence, K. W. (1936). The nature of discrimination learning in animals. Psychological Review, 43(5), 427-449. https://doi.org/10.1037/h0056975
Spence, K. W. (1937). The differential response in animals to stimuli varying within a single dimension. Psychological Review, 44(5), 430-444. https://doi.org/10.1037/h0062885
Terrace, H. S. (1963). Discrimination learning with and without "errors." Journal of the Experimental Analysis of Behavior, 6(1), 1-27. https://doi.org/10.1901/jeab.1963.6-1