Abstract

A reinforcement schedule is a form of reinforcement — the rule that specifies which responses, and how many or how timed, produce a reinforcer within operant conditioning. The organizing insight of the field was that the schedule, not merely the reinforcer, controls behavior: each rule generates its own orderly, reproducible pattern of responding, visible in the cumulative record. This article traces the schedule from Skinner's cumulative recorder through Ferster and Skinner's catalogue of the four basic schedules — fixed and variable, ratio and interval — and their signature patterns, to the partial-reinforcement extinction effect, Herrnstein's matching law and Baum's generalized form that made concurrent schedules a quantitative science of choice, and Nevin's behavioral momentum. Three demonstrations let the reader generate the cumulative-record patterns, allocate behavior between concurrent schedules under the matching law, and compare extinction after continuous and partial reinforcement.

Keywords: reinforcement schedule, operant conditioning, matching law, cumulative record, behavioral momentum

The move that founded the experimental analysis of schedules was to stop asking whether a reinforcer works and start asking how its arrangement shapes the stream of behavior. A rat pressing a lever or a pigeon pecking a key rarely earns a reinforcer for every response; in nature and in the laboratory alike, reinforcement is intermittent, contingent on some number of responses or on the passage of time. The reinforcement schedule is the precise statement of that contingency, and the central empirical fact of the field is that each schedule imposes a characteristic rate and temporal pattern of responding so stable that a trained eye can read the schedule from the record (Ferster & Skinner, 1957; Staddon & Cerutti, 2003).

Key Takeaways
  • A reinforcement schedule is the rule relating responses to reinforcers; the schedule, not the reinforcer alone, determines the rate and pattern of behavior.
  • The four basic schedules cross two dimensions: whether reinforcement depends on a number of responses (ratio) or on time (interval), and whether the requirement is fixed or variable.
  • Each schedule produces a signature cumulative-record pattern: the fixed-interval scallop, the fixed-ratio break-and-run, and the high steady rates of the variable schedules.
  • Intermittently reinforced responses are more resistant to extinction than continuously reinforced ones — the partial-reinforcement extinction effect.
  • On concurrent schedules, relative response rates match relative reinforcement rates — Herrnstein's matching law, generalized by Baum to describe bias and undermatching.

What a Reinforcement Schedule Is

A reinforcement schedule is a rule that determines which occurrences of a response are followed by a reinforcer. The simplest rule, continuous reinforcement, reinforces every response; every other rule is a schedule of intermittent or partial reinforcement, delivering the reinforcer only on some responses (Skinner, 1938). The importance of the concept is that behavior is exquisitely sensitive to the rule. Under continuous reinforcement a response is acquired quickly and abandoned quickly when reinforcement stops; under intermittent reinforcement the same response is maintained at a steady rate and persists long after reinforcement ceases. The schedule is therefore not a peripheral detail of an experiment but the primary independent variable of operant research.

Skinner's methodological contribution was the instrument that made schedules visible. The cumulative recorder draws a line that steps up by a fixed amount with each response and never descends, so that the slope of the line is the momentary response rate and the whole record is a running total across the session. On such a record a fast, steady rate is a steep straight line, a pause is a horizontal segment, and an accelerating rate is an upward curve. Because the four basic schedules each impose a distinct rate-over-time profile, they draw four distinct curves, and the cumulative record became the signature by which a schedule is recognized (Ferster & Skinner, 1957).

The cumulative record

Each schedule carves its own signature into the cumulative record, where the line steps up with every response and never falls, so its slope is the response rate. Choose a schedule and read its pattern; the gold ticks mark reinforcers.

resp.time →
Fixed interval (FI). The scallop: a pause after each reinforcer, then an accelerating rate as the interval's end nears.

The curves are an illustrative model of the four canonical patterns (Ferster & Skinner, 1957), computed locally and not stored; real records vary across organisms and sessions.

MeSH files the descriptor for reinforcement schedule directly under reinforcement, reflecting the indexing decision to treat the schedule as a subordinate kind of the reinforcement operation. The relation is genuinely taxonomic here: a schedule is a way of arranging reinforcement, so a reinforcement schedule is a form of reinforcement, and the two are studied within the same operant framework rather than standing in opposition.

The Four Basic Schedules

The classification that organizes the field crosses two binary dimensions. The first is what the reinforcer is contingent on: a number of responses (a ratio schedule) or the first response after an interval of time has elapsed (an interval schedule). The second is whether that requirement is held constant (fixed) or varied around an average (variable). Crossing the two yields the four basic schedules — fixed ratio, variable ratio, fixed interval, and variable interval — whose systematic study filled Ferster and Skinner's Schedules of Reinforcement, the volume that defined the subject (Ferster & Skinner, 1957).

The distinction between ratio and interval schedules is the more consequential of the two, because it determines how response rate relates to reinforcement rate. On a ratio schedule the reinforcer depends only on the count of responses, so responding faster earns reinforcers faster, and ratio schedules therefore sustain very high rates. On an interval schedule the reinforcer becomes available with the passage of time and only the first response after it becomes available is reinforced, so responding faster earns very little extra, and interval schedules sustain moderate rates. The fixed-versus-variable distinction governs the temporal pattern rather than the overall rate: fixed requirements permit the organism to predict the next reinforcer and pause after each one, whereas variable requirements make the next reinforcer unpredictable and eliminate the pause (Staddon & Cerutti, 2003).

Table 1

The Four Basic Schedules and Their Characteristic Response Patterns

Schedule Reinforcer delivered Characteristic pattern
Fixed ratio (FR) After a fixed number of responses. Break-and-run: a post-reinforcement pause followed by a rapid, steady burst to the next reinforcer.
Variable ratio (VR) After a varying number of responses, averaging some value. Very high, steady rate with little or no pausing; the schedule underlying gambling.
Fixed interval (FI) For the first response after a fixed time has elapsed. Scallop: a pause after reinforcement, then an accelerating rate as the interval's end approaches.
Variable interval (VI) For the first response after a varying time, averaging some value. Moderate, steady rate; the stable baseline favored for studying other variables.

Signature Response Patterns

The characteristic patterns are not incidental; they are the primary data of schedule research and the reason the cumulative record matters. The fixed-interval scallop is the clearest case. Because reinforcement is available only after a fixed time, and because a response early in the interval is never reinforced, the organism comes to pause after each reinforcer and then accelerate its responding as the moment of availability nears, tracing a repeating scalloped curve. The pause is not a failure of motivation but an adaptation to the temporal contingency: responding early cannot pay, and the animal's behavior comes to reflect the passage of time (Ferster & Skinner, 1957).

The fixed-ratio pattern is the break-and-run: after each reinforcer the organism pauses — the post-reinforcement pause, longer for larger ratios — and then completes the response requirement in a rapid, near-constant burst. The variable schedules erase the pause. Because the next reinforcer might arrive at any moment, there is no predictable safe period in which to rest, and both variable-ratio and variable-interval schedules generate steady responding without the post-reinforcement pause. Variable ratio produces the highest sustained rates of any simple schedule, which is why it is the contingency behind gambling; variable interval produces a moderate, remarkably stable rate that makes it the preferred baseline for measuring the effect of other variables such as a drug, a punisher, or a second concurrent schedule (Staddon & Cerutti, 2003).

Figure 1

Idealized Cumulative Records of the Four Basic Schedules

Four cumulative records, one per basic schedule, each with responses on the vertical axis and time on the horizontal axis Four small line graphs. Variable ratio is a single steep, straight line, the highest rate. Fixed ratio is a break-and-run staircase: flat pauses after each reinforcer followed by rapid straight bursts. Fixed interval is a repeating scallop: a pause after each reinforcer that accelerates into the next. Variable interval is a moderate, straight line, less steep than variable ratio. In every panel the line only rises or stays level, never descends, because the record is cumulative. Variable ratio (VR) Fixed ratio (FR) Fixed interval (FI) Variable interval (VI) responses responses time → time → time → time →
Note. The signature cumulative-record patterns. Ratio schedules (top) sustain higher rates than interval schedules (bottom); fixed schedules (left of each pair) produce a post-reinforcement pause — the fixed-ratio break-and-run and the fixed-interval scallop — while variable schedules produce steady responding without pausing. The record only rises, since responses accumulate (Ferster & Skinner, 1957; Staddon & Cerutti, 2003).

The Partial-Reinforcement Extinction Effect

The most counterintuitive consequence of intermittent reinforcement appears not during reinforcement but during its withdrawal. Common sense suggests that a response reinforced every time should be the best learned and so the most persistent; the opposite is true. A response maintained on an intermittent schedule persists far longer in extinction than one that was continuously reinforced — the partial-reinforcement extinction effect. Humphreys gave the effect its classic demonstration, showing that conditioned responses acquired under a random alternation of reinforced and unreinforced trials extinguished much more slowly than those acquired under consistent reinforcement (Humphreys, 1939).

The partial-reinforcement extinction effect

When reinforcement stops, a continuously reinforced response fades fast, while an intermittently reinforced one persists. Thin the training schedule — raise the responses-per-reinforcer ratio — and watch the partial-reinforcement curve stretch out its resistance to extinction.

% resp.extinction trials →continuouspartial
Continuous reinforcement falls below 10% of baseline in about 14 trials. The partial schedule of one reinforcer per 5 responses takes about 47 trials — 3.4× as persistent.

An illustrative exponential-decay model of extinction (Humphreys, 1939), computed locally and not stored; it captures the direction and ordering of the effect, not the exact rates of any one experiment.

The effect is a discrimination problem. Extinction is learned only when the organism can tell that the contingency has changed, and a run of unreinforced responses is a much larger, more novel departure for an animal accustomed to a reinforcer on every response than for one accustomed to long unreinforced stretches. Intermittent reinforcement thus makes the onset of extinction hard to detect, and the response continues. The practical importance is large: behaviors inadvertently maintained on a thin intermittent schedule — a child's tantrum reinforced only occasionally, a gambler's play — are precisely the ones most resistant to elimination, a direct corollary of the schedule's grip on persistence (Staddon & Cerutti, 2003).

The Matching Law

Schedules acquire their deepest theoretical interest when two are available at once. On a concurrent schedule an organism chooses between two responses, each on its own schedule of reinforcement, and the question becomes how behavior is allocated between them. Herrnstein's answer, from pigeons choosing between two variable-interval keys, is the matching law: the proportion of responses directed to an alternative matches the proportion of reinforcers obtained from it. Relative behavior tracks relative reinforcement, so that a key delivering three-quarters of the reinforcers receives about three-quarters of the responses (Herrnstein, 1961).

Herrnstein went on to argue that matching is not confined to concurrent schedules but is a general principle: the rate of any single response is proportional to the reinforcement it earns relative to all reinforcement available in the situation, including reinforcement from extraneous, unmeasured sources. On this view a single schedule is a concurrent schedule in which the alternative is everything else the organism could do, and the absolute response rate is set by the fraction of total context reinforcement the measured response commands (Herrnstein, 1970). Baum then generalized the law to the form now standard, expressing the ratio of behaviors as a power function of the ratio of reinforcements. Two free parameters capture the systematic deviations from strict matching: bias, a constant preference for one alternative independent of reinforcement, and sensitivity, which is typically less than one — the pervasive finding of undermatching, in which behavior is allocated less extremely than reinforcement (Baum, 1974).

The matching law

Two keys, each on its own variable-interval schedule. The matching law says the share of responses to a key equals its share of the reinforcers. Set the two reinforcement rates and the sensitivity a; with a = 1 you get strict matching, and a < 1 gives the usual undermatching.

strict 75%A: 707B: 293
Reinforcement ratio A:B = 3.00. Predicted response share on key A = 70.7% (707 of 1000 responses), versus 75.0% under strict matching. Undermatching pulls allocation toward 50/50.

Matching states the outcome but not the mechanism, and why an organism should match became the law's central theoretical question. The leading answer is melioration: rather than computing an optimum, the organism continuously shifts behavior toward whichever alternative currently yields the higher local rate of reinforcement, and this moment-to-moment reallocation settles at the matching relation. Melioration is distinct from global maximization of total reinforcement, and the two diverge most sharply on concurrent ratio schedules, where maximization predicts exclusive choice of the richer alternative. Distinguishing matching, melioration, and maximization &mdash; whether behavior tracks local rates, probabilities, or a global optimum &mdash; remains the live question behind the modern re-examinations of concurrent schedules (Vaughan, 1981).

Behavioral Momentum

The matching law describes how reinforcement allocates behavior; behavioral momentum theory describes how reinforcement makes behavior resist change. Nevin drew a deliberate analogy to Newtonian momentum. The rate of a response in a given stimulus context is like an object's velocity, and its resistance to disruption &mdash; by extinction, by free reinforcers, by satiation &mdash; is like the object's mass. Crucially, the two are separable: response rate is governed by the response&ndash;reinforcer contingency, whereas resistance to change is governed by the total rate of reinforcement in the stimulus context, including reinforcers not contingent on the response at all (Nevin, 1992).

The consequence is that a richer reinforcement context builds a more persistent behavior even when it does not build a faster one. A response in a context signalling a high rate of reinforcement will be more resistant to extinction than the same response, at the same baseline rate, in a leaner context. Behavioral momentum has become the dominant account of response persistence and relapse, and its predictions &mdash; and their limits &mdash; remain under active test, particularly where reinforcement rate affects the resurgence of an extinguished response in ways the original theory does not fully capture (Craig & Shahan, 2016).

Worked Example

The matching law can be made concrete with a concurrent variable-interval schedule. Suppose two keys are arranged so that key A delivers reinforcers at 30 per hour and key B at 10 per hour, for a total of 40 reinforcers per hour. Strict matching predicts that the proportion of responses on key A equals the proportion of reinforcers from key A: BA / (BA + BB) = 30 / 40 = 0.75. Of 1,000 responses in the session, the law predicts 750 on key A and 250 on key B, a 3-to-1 response ratio matching the 3-to-1 reinforcement ratio.

Real animals usually deviate, and the generalized matching law quantifies the deviation. Writing it as a ratio, BA / BB = b (rA / rB)a, strict matching is the special case a = 1, b = 1. Undermatching, the common finding, is a < 1. With no bias (b = 1) and a typical sensitivity a = 0.8, the reinforcement ratio rA / rB = 30 / 10 = 3 gives a behavior ratio of BA / BB = 30.8 = 2.41. The predicted proportion on key A is then 2.41 / (2.41 + 1) = 0.707. Undermatching thus pulls allocation from the strict-matching 0.750 toward indifference at 0.500, yielding 0.707 &mdash; the animal still prefers the richer key, but less sharply than the reinforcement ratio alone would dictate (Herrnstein, 1961; Baum, 1974).

Current Directions

Contemporary schedule research is largely a quantitative debate about choice and persistence. One active front tests the reach of behavioral momentum theory. Craig and Shahan found that the rate of reinforcement, which the theory predicts should govern the persistence and resurgence of a response, does not affect resurgence in the way the theory requires, a result that has pushed the field toward revised quantitative models of how reinforcement history controls relapse (Craig & Shahan, 2016).

A second front returns to the matching law with modern data and analysis. Concurrent ratio schedules, long thought to yield exclusive preference for the richer alternative, have been re-examined to ask whether animals rate-match, probability-match, or optimize, with evidence that behavior on concurrent ratios is more graded and more dynamic than the classic account assumed (Baum et al., 2022). Related work analyzes the moment-to-moment dynamics of choice on concurrent ratio schedules, tracking how allocation develops within a session rather than only at steady state, and so moving the study of schedules from static equilibrium toward the temporal structure of choice itself (Bellow & Lattal, 2023).

Key Researchers

B. F. Skinner (1904-1990). Harvard University; he introduced the cumulative recorder and the systematic study of how patterns of reinforcement, rather than single rewards, control the rate and temporal structure of responding. Wikipedia

Charles B. Ferster (1922-1981). With Skinner he co-authored Schedules of Reinforcement (1957), the exhaustive catalogue of the response patterns generated by the fixed and variable, ratio and interval schedules that defined the field for a generation. Wikipedia

Richard J. Herrnstein (1930-1994). Harvard University; he discovered the matching law, turning the study of concurrent schedules from a descriptive catalogue into a quantitative science of choice. Wikipedia

John Anthony Nevin (1933-2018). University of New Hampshire; he developed behavioral momentum theory, showing that the reinforcement rate in a context determines how resistant responding is to disruption, a dynamic property beyond steady-state rate. Wikipedia

William M. Baum. University of California, Davis; he generalized the matching law to the power form that quantifies bias and undermatching, and continues to develop molar accounts of how schedules allocate behavior over time. ORCID

Timothy A. Shahan. Utah State University; he tests and extends behavioral momentum theory experimentally, showing where reinforcement-rate effects on persistence and resurgence diverge from the theory's predictions. ORCID

Kennon A. Lattal. West Virginia University; a leading contemporary analyst of operant schedules whose work on choice dynamics in concurrent ratio schedules extends the quantitative study of reinforcement into the moment-to-moment structure of responding. Faculty page &middot; Google Scholar

Commonly Confused With

Reinforcement
Reinforcement is the process by which a consequence strengthens the behavior it follows; a reinforcement schedule is the rule that specifies when and how often that consequence is delivered. The reinforcer answers what strengthens the behavior; the schedule answers on what contingency it is arranged. The same reinforcer produces entirely different rates and patterns of responding &mdash; and entirely different resistance to extinction &mdash; depending on the schedule that delivers it, which is precisely why the schedule, not the reinforcer alone, is the primary variable of operant research.

Discussion

The reinforcement schedule is the concept through which the study of learning became quantitative. Thorndike's law of effect established that consequences select behavior, but it treated the reinforcer as a discrete event stamping in a response (Thorndike, 1927). Skinner's reframing &mdash; that the arrangement of reinforcement, extended over time, is what controls the rate and pattern of behavior &mdash; shifted the unit of analysis from the single reinforced response to the schedule, and the cumulative records of the four basic schedules gave that shift its empirical spine (Ferster & Skinner, 1957). The orderliness of the patterns, reproducible across species and across decades, is among the most robust findings in psychology.

The concept's later history is a progressive mathematization. The partial-reinforcement extinction effect showed that a schedule determines not only how an organism responds while reinforced but how tenaciously it persists when reinforcement stops (Humphreys, 1939). The matching law and its generalization made choice between schedules predictable to a parameter (Herrnstein, 1961; Herrnstein, 1970; Baum, 1974), and behavioral momentum added a separate dimension of persistence governed by context reinforcement (Nevin, 1992). Current work continues to refine these quantitative accounts of choice and persistence rather than to overturn them, a sign of a mature theory still generating testable disagreement (Craig & Shahan, 2016; Baum et al., 2022; Bellow & Lattal, 2023; Staddon & Cerutti, 2003).

Glossary

Behavioral momentum.
The resistance of a response to disruption, treated by analogy to physical momentum and governed by the rate of reinforcement in the stimulus context.
Concurrent schedule.
An arrangement in which two or more schedules operate at once on different responses, so the organism allocates behavior between them; the setting for the matching law.
Continuous reinforcement.
A schedule that reinforces every occurrence of the response; produces fast acquisition and fast extinction.
Cumulative record.
A running total of responses over time whose slope is the response rate; the instrument by which schedule patterns are read.
Fixed-interval schedule.
Reinforces the first response after a fixed time has elapsed; produces the scallop pattern of a pause followed by acceleration.
Fixed-ratio schedule.
Reinforces after a fixed number of responses; produces break-and-run, a post-reinforcement pause followed by a rapid burst.
Matching law.
Herrnstein's principle that the proportion of responses on an alternative matches the proportion of reinforcers obtained from it.
Melioration.
The proposed mechanism of matching: the organism continuously shifts behavior toward whichever alternative currently yields the higher local rate of reinforcement, settling at the matching relation rather than at a computed global optimum.
Operant conditioning.
Learning in which the consequences of a response alter its future probability; the framework within which schedules are defined.
Partial-reinforcement extinction effect.
The finding that intermittently reinforced responses persist longer in extinction than continuously reinforced ones.
Post-reinforcement pause.
The break in responding immediately after a reinforcer on fixed schedules, longer for larger fixed-ratio requirements.
Ratio schedule.
A schedule on which reinforcement depends on a number of responses; sustains high rates because faster responding earns reinforcers faster.
Reinforcement.
The strengthening of behavior by its consequences; the operation the schedule arranges, and the MeSH category under which schedules are filed.
Undermatching.
The common deviation in which behavior is allocated less extremely than reinforcement, captured by a sensitivity parameter below one in the generalized matching law.
Variable-interval schedule.
Reinforces the first response after a varying average time; produces a moderate, stable rate favored as a baseline.
Variable-ratio schedule.
Reinforces after a varying average number of responses; produces the highest steady rates and underlies gambling.

Frequently Asked Questions

What is a reinforcement schedule?
A reinforcement schedule is the rule specifying which responses are followed by a reinforcer &mdash; for example, every response, every fifth response, or the first response after a minute. The schedule, not the reinforcer alone, determines the rate and pattern of behavior (Ferster & Skinner, 1957).

What are the four basic schedules of reinforcement?
Fixed ratio, variable ratio, fixed interval, and variable interval. They cross two dimensions: whether reinforcement depends on a number of responses (ratio) or on elapsed time (interval), and whether the requirement is fixed or varies around an average (Ferster & Skinner, 1957).

Why do ratio schedules produce higher response rates than interval schedules?
On a ratio schedule the reinforcer depends only on the number of responses, so responding faster earns reinforcers faster. On an interval schedule the reinforcer depends on elapsed time, so responding faster earns little extra, and rates stay moderate (Staddon & Cerutti, 2003).

What is the partial-reinforcement extinction effect?
It is the finding that a response reinforced only intermittently persists longer when reinforcement stops than a response that was reinforced every time, because the change to extinction is harder to discriminate (Humphreys, 1939).

What is the matching law?
The matching law states that on concurrent schedules the proportion of responses directed to an alternative matches the proportion of reinforcers obtained from it, so relative behavior tracks relative reinforcement (Herrnstein, 1961).

What is undermatching?
Undermatching is the common deviation in which behavior is allocated less extremely than reinforcement &mdash; a preference weaker than strict matching predicts &mdash; captured by a sensitivity parameter below one in Baum's generalized matching law (Baum, 1974).

What is behavioral momentum?
Behavioral momentum is the resistance of a response to disruption, treated by analogy to physical momentum. Its magnitude is governed by the total rate of reinforcement in the stimulus context, separately from the response rate itself (Nevin, 1992).

Why is variable-ratio reinforcement linked to gambling?
Variable-ratio schedules reinforce after an unpredictable number of responses and so generate very high, persistent rates of responding with no pausing, the pattern seen in slot-machine play and other forms of gambling (Ferster & Skinner, 1957).

References

Baum, W. M. (1974). On two types of deviation from the matching law: Bias and undermatching. Journal of the Experimental Analysis of Behavior, 22(1), 231-242. https://doi.org/10.1901/jeab.1974.22-231

Baum, W. M., Aparicio, C. F., & Alonso-Alvarez, B. (2022). Rate matching, probability matching, and optimization in concurrent ratio schedules. Journal of the Experimental Analysis of Behavior, 118(1), 96-131. https://doi.org/10.1002/jeab.771

Bellow, R. R., & Lattal, K. A. (2023). Choice dynamics in concurrent ratio schedules of reinforcement. Journal of the Experimental Analysis of Behavior, 119(3), 337-355. https://doi.org/10.1002/jeab.828

Craig, A. R., & Shahan, T. A. (2016). Behavioral momentum theory fails to account for the effects of reinforcement rate on resurgence. Journal of the Experimental Analysis of Behavior, 105(3), 375-392. https://doi.org/10.1002/jeab.207

Ferster, C. B., & Skinner, B. F. (1957). Schedules of reinforcement. Appleton-Century-Crofts. ISBN 9780133139327.

Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. Journal of the Experimental Analysis of Behavior, 4(3), 267-272. https://doi.org/10.1901/jeab.1961.4-267

Herrnstein, R. J. (1970). On the law of effect. Journal of the Experimental Analysis of Behavior, 13(2), 243-266. https://doi.org/10.1901/jeab.1970.13-243

Humphreys, L. G. (1939). The effect of random alternation of reinforcement on the acquisition and extinction of conditioned eyelid reactions. Journal of Experimental Psychology, 25(2), 141-158. https://doi.org/10.1037/h0058138

Nevin, J. A. (1992). An integrative model for the study of behavioral momentum. Journal of the Experimental Analysis of Behavior, 57(3), 301-316. https://doi.org/10.1901/jeab.1992.57-301

Skinner, B. F. (1938). The behavior of organisms: An experimental analysis. Appleton-Century.

Staddon, J. E. R., & Cerutti, D. T. (2003). Operant conditioning. Annual Review of Psychology, 54, 115-144. https://doi.org/10.1146/annurev.psych.54.101601.145124

Thorndike, E. L. (1927). The law of effect. The American Journal of Psychology, 39(1/4), 212-222. https://doi.org/10.2307/1415413

Vaughan, W. (1981). Melioration, matching, and maximization. Journal of the Experimental Analysis of Behavior, 36(2), 141-149. https://doi.org/10.1901/jeab.1981.36-141