Abstract

Probability learning is a form of learning in which a person or animal, choosing repeatedly among options whose rewards occur with fixed but unknown probabilities, gradually adjusts its choices to reflect those probabilities. Its signature finding is probability matching: rather than always selecting the more frequently rewarded option, the strategy that maximizes payoff, learners tend to choose each option in rough proportion to how often it is reinforced, a suboptimal pattern first quantified in the mid-twentieth century. Whether matching is a genuine strategy or an artifact of limited trials, weak incentives, and a search for pattern remains debated. Modern work recasts the problem as reinforcement learning under uncertainty, in which the brain tracks the volatility of the environment and tunes its learning rate accordingly. This article surveys the paradigm, the matching phenomenon, its interpretations, and its neural basis.

Keywords: probability learning, probability matching, reinforcement, decision making, learning rate

Probability learning is studied with a deceptively simple task: on each of many trials the learner predicts which of two events will occur, or chooses between two options, and one option pays off more often than the other according to a probability the learner is never told (Estes, 1950). Over trials, choices come to track the reinforcement probabilities, and the striking regularity is that they track them too closely, settling near a proportion that equals the reward probability rather than the higher accuracy that always choosing the better option would secure (Humphreys, 1939). The task is a bridge between simple conditioning and full decision making under uncertainty, and it exposes how people extract, and sometimes over-read, the statistical structure of their environment.

Key Takeaways
  • Probability learning is the acquisition, through repeated feedback, of choices that reflect the reward probabilities of options that are never explicitly stated.
  • Its hallmark is probability matching: choosing each option in proportion to its reward rate, which yields lower accuracy than consistently choosing the better option (maximizing).
  • Matching is robust across humans and animals but weakens with more trials, stronger incentives, and clearer feedback, fuelling debate over whether it is a strategy or an artifact.
  • Statistical-learning theory modelled matching as the asymptote of trial-by-trial conditioning; economics treats it as a departure from rational choice.
  • Modern accounts frame it as reinforcement learning in which the brain estimates environmental volatility and adjusts how fast it updates beliefs.

What Probability Learning Is

Probability learning is the process by which repeated experience of probabilistic outcomes shapes choice. In the canonical two-choice prediction task, a light flashes on the left or the right on each trial; the learner guesses which side before it appears, and the left light comes on with some fixed probability, say 0.75, while the right comes on the rest of the time. No trial is fully predictable, and no instruction reveals the odds. What the learner has to go on is only the running history of outcomes, and from that history a stable pattern of guessing emerges (Estes, 1950).

The construct sits between two neighbours. It is more than conditioning, because the outcomes are only partially predictable and the learner must integrate over a long, noisy history rather than form a single stimulus-response bond. It is less than deliberate decision making, because there is no described gamble to reason about, only experience to be accumulated. This intermediate status is exactly why probability learning became a testing ground: it asks what statistical regularity a mind will extract when it is given nothing but a stream of reinforced and unreinforced choices (Edwards, 1961).

The Probability-Matching Phenomenon

The central empirical fact is probability matching. When the better option is reinforced on a fraction p of trials, learners come to choose it on approximately a fraction p of trials rather than on every trial. If a light appears on the left 75 percent of the time, people end up guessing “left” about 75 percent of the time, leaving a quarter of their guesses on the side that is right only a quarter of the time (Grant, Hake, & Hornseth, 1951). The pattern is not confined to discrete prediction: in free-operant choice between concurrently reinforced alternatives, the relative rate of responding tracks the relative rate of reinforcement, a quantitative regularity Herrnstein formalized as the matching law, of which probability matching is the discrete-trial counterpart (Herrnstein, 1961).

Matching versus maximizing

A matcher chooses each option in proportion to how often it pays off; a maximizer always picks the better option. Set the reward probability of the better option and compare the two accuracies. The matcher never wins, and the gap is widest in the middle range.

63%matching75%maximizingexpected accuracy per trial
At a reward probability of 75%, a matcher is correct 62.5% of the time and a maximizer 75% — a sacrifice of 12.5 percentage points on every trial.

Matching is puzzling because it is demonstrably suboptimal. The strategy that maximizes reward, always choosing the more frequently reinforced option, guarantees an accuracy equal to p; matching yields an expected accuracy of only p² + (1 − p)², which is always lower whenever the two options differ. For a 75/25 split the maximizer is correct 75 percent of the time while the matcher is correct only 62.5 percent, a gap of 12.5 percentage points thrown away on every hundred trials (Vulkan, 2000). That a robust, cross-species behaviour should leave reward on the table is the tension the rest of the literature tries to resolve.

Figure 1

Accuracy of Matching Versus Maximizing as the Reward Probability Varies

Maximizing always yields at least as much accuracy as probability matching Two curves are plotted against the reward probability of the better option, from one-half to one, on the horizontal axis, with expected accuracy on the vertical axis. The maximizing line rises straight from one-half to one. The matching curve, equal to p squared plus one minus p squared, starts at one-half, dips slightly below the diagonal in the middle range, and rejoins it at both ends; the vertical gap between the two, widest near a reward probability of three-quarters, is the accuracy sacrificed by matching. Expected accuracy Reward probability of better option (p) 0.5 0.75 1.0 0.5 0.75 1.0 Maximizing (p) Matching (p²+(1−p)²)
Note. Maximizing accuracy equals p (blue). Matching accuracy equals p² + (1 − p)² (gold) and never exceeds it; the two coincide only at p = 0.5 and p = 1. At p = 0.75 the maximizer scores 0.75 and the matcher 0.625, the 12.5-percentage-point gap marked by the dashed line.

Interpreting Matching: Strategy or Artifact

Two broad readings of matching have competed for decades. The first, from statistical-learning theory, treats matching not as a choice but as a by-product of learning dynamics. In Estes's stimulus-sampling model, each trial conditions a sample of stimulus elements toward whichever response was reinforced; at equilibrium the proportion of elements favouring a response equals the probability it was reinforced, so choice probability converges on the reinforcement probability automatically. Matching, on this view, is simply where trial-by-trial conditioning settles, not a policy the learner selects (Estes, 1950).

The learning curve settles on the odds

A simple reinforcement learner updates its choice probability toward whichever option was just rewarded. Over trials it drifts from an unbiased start toward the reward probability itself — matching as the equilibrium of trial-by-trial conditioning, not a chosen policy.

0.00.51.0050100150200trialp = 0.75
With the better option rewarded 75% of the time, the learner’s choice probability settles near 75% after 200 trials — tracking the odds rather than committing to the better option.

The second reading, from the study of judgment and decision making, treats matching as a departure from rational choice that demands explanation, and asks how real it is. Careful experiments show that matching is fragile: it recedes when participants receive many trials, meaningful financial incentives, and clear feedback, and a substantial share of people maximize under favourable conditions (Shanks, Tunney, & McCarthy, 2002). This has led to the view that apparent matching often reflects an active but mistaken search for predictive pattern in what is actually a random sequence, together with the limited data of short experiments, rather than a fixed preference for matching itself (Vulkan, 2000).

Table 1. Competing accounts of probability matching and what each implies.
Account Matching is... Key implication
Statistical learning (Estes) The equilibrium of trial-by-trial conditioning. Matching is expected, not irrational; it is where the learning process settles.
Behavioural economics (Vulkan) A violation of expected-value maximization. Matching is a puzzle to be explained by cognitive limits or misperceived randomness.
Artifact view (Shanks et al.) A transient of too few trials and weak incentives. Given enough trials, incentives, and feedback, most learners shift toward maximizing.

Learning Rate and a Changing World

The classic paradigm holds the reward probabilities fixed, but real environments drift, and how quickly a learner should update depends on how fast the world is changing. In a stable world a single surprising outcome is noise and should barely move one's estimate; in a volatile world the same surprise may signal that the odds have shifted and should be weighted heavily. The efficient learner therefore sets a high learning rate when the environment is volatile and a low one when it is stable (Behrens, Woolrich, Walton, & Rushworth, 2007).

How volatility sets the learning rate

The efficient learner tunes how much the latest outcome counts to how fast the world is changing. Drag the environment from stable to volatile and watch the warranted learning rate rise, and with it how far a single surprising outcome shifts the current belief.

belief update after one surprising outcome01prior 0.300.57
At 50% volatility the warranted learning rate is about 0.39: one surprising outcome moves a belief held at 0.30 to 0.57, and beliefs track a genuine change in roughly 2.6 trials.

Human learners approximate this ideal. Behrens and colleagues showed that people raise their learning rate in blocks where reward probabilities change frequently and lower it in stable blocks, behaving as if they estimate volatility itself, and that activity in the anterior cingulate cortex tracks that estimate (Behrens, Woolrich, Walton, & Rushworth, 2007). This reframes probability learning as an inference problem: the brain is not merely counting outcomes but estimating how much recent outcomes should count. The neural substrate for the underlying error signal, the mismatch between expected and received reward, is the phasic activity of midbrain dopamine neurons, which encodes reward prediction error and so provides the teaching signal that drives probabilistic learning (Lerner, Holloway, & Seiler, 2021).

Worked Example

Consider the two-choice task with the left option rewarded on 75 percent of trials and the right on 25 percent. A learner who matches will, at asymptote, choose left on about 75 of every 100 trials and right on about 25. How often is that learner correct? Left is correct 75 percent of the time and is chosen 75 percent of the time; right is correct 25 percent of the time and is chosen 25 percent of the time. The expected accuracy is therefore 0.75 × 0.75 + 0.25 × 0.25 = 0.5625 + 0.0625 = 0.625, or 62.5 percent.

Now compare the maximizer, who always chooses left. That learner is correct exactly whenever left is the rewarded side, which is 75 percent of trials, for an accuracy of 75 percent. The matcher sacrifices 75 − 62.5 = 12.5 percentage points, twelve or thirteen extra errors in every hundred trials, purely by spreading guesses in proportion to the odds instead of committing to the better option. The arithmetic makes the puzzle concrete: matching is not a small rounding away from optimal but a systematic and sizable cost, which is why its persistence across people and species has demanded an explanation (Vulkan, 2000).

Current Directions

Contemporary research has largely dissolved the old strategy-versus-artifact dichotomy by embedding probability learning inside computational models of reinforcement learning, where matching, maximizing, and everything between fall out of the settings of a few parameters such as learning rate, choice stochasticity, and perseveration. Recent studies decompose sequential choice into these components and show that much of what looked like a preference for matching is better described as biased probability representation, sensitivity to sequential structure, and a tendency to repeat recent choices (Baumann, Schlegelmilch, & von Helversen, 2025). Foraging-style tasks, in which rewards persist and accumulate at unchosen options, further show that the balance between matching and maximizing shifts systematically with the payoff structure, so that the “irrationality” of matching is partly a rational response to a richer environment than the classic task assumed (Ellerby & Tunney, 2019).

The neuroscience has moved in parallel. The dopaminergic reward-prediction-error signal, once treated as a single scalar teaching signal, is now understood to be regionally and temporally diverse, carrying information about the value, timing, and even the identity of outcomes, which opens the question of how these richer signals support probabilistic learning in changing environments (Lerner, Holloway, & Seiler, 2021). The open problems that remain are how volatility estimation, dopaminergic error signalling, and the several biases now identified in choice combine to produce the deceptively simple curve of a learner homing in on the odds.

Key Researchers

William K. Estes (1919-2011). Indiana University, Rockefeller University, and Harvard University; he founded stimulus-sampling theory and the statistical-learning tradition, giving probability matching its first formal account as the equilibrium of trial-by-trial conditioning. Wikipedia

Ward Edwards (1927-2005). University of Michigan and University of Southern California; the founder of behavioral decision theory, his 1000-trial studies established the empirical regularities of human sequential prediction and linked probability learning to judgment under uncertainty. Wikipedia

David R. Shanks. University College London; he re-examined the matching phenomenon experimentally, showing that with adequate trials, incentives, and feedback learners shift toward maximizing, sharpening the debate over whether matching is a strategy or an artifact. ORCID

Timothy E. J. Behrens. University of Oxford; he showed with functional imaging and Bayesian modelling that humans track the volatility of a probabilistic environment and adjust their learning rate accordingly, giving probability learning a neural account. ORCID

Richard J. Tunney. Aston University; he co-authored the influential re-examination of probability matching and later foraging-task studies quantifying how reward persistence and accumulation shift choice between matching and maximizing. ORCID

Discussion

Probability learning has proved durable as a research topic because its simplest result is also its most stubborn. A learner given nothing but a stream of partially predictable outcomes reliably comes to choose in proportion to the odds, and reliably leaves reward on the table by doing so. For half a century that fact was read in two incompatible ways, as the natural resting point of a conditioning process (Estes, 1950) or as an irrational lapse from maximization (Vulkan, 2000), with careful experiments showing that the behaviour is real but fragile, receding under stronger incentives and longer training (Shanks, Tunney, & McCarthy, 2002).

The reinforcement-learning synthesis has largely reconciled these positions by dissolving the question they shared. Matching and maximizing are not rival policies a learner selects between but points on a continuum generated by the same updating machinery under different parameters, and the parameters themselves are adaptive: a learner who raises the learning rate in a volatile world and lowers it in a stable one is behaving well, not badly (Behrens, Woolrich, Walton, & Rushworth, 2007). What once looked like a single anomaly is now a window onto how the brain estimates the statistics of its environment and decides how much the latest outcome should change its mind, a computation that connects probability learning to decision making, to reward-based learning, and to the dopaminergic signalling that implements it (Lerner, Holloway, & Seiler, 2021).

Glossary

Asymptote.
The stable level a learning curve approaches after many trials; in probability learning, the choice proportion at which behaviour settles.
Conditioning.
The formation of a link between events or between a response and its reinforcement; the simpler process probability learning extends to partially predictable outcomes.
Decision making.
The selection of an action from alternatives on the basis of their expected outcomes, of which probability learning is an experience-based case.
Dopamine.
A midbrain neurotransmitter whose phasic activity encodes reward prediction error, supplying the teaching signal that drives probabilistic learning.
Expected value.
The average payoff of an option weighted by its probability; the benchmark against which matching is judged suboptimal.
Learning rate.
The weight a learner gives to the most recent outcome when updating a belief; high rates track change quickly, low rates smooth over noise.
Matching law.
Herrnstein's regularity that in free-operant choice the relative rate of responding equals the relative rate of reinforcement; probability matching is its discrete-trial counterpart.
Maximizing.
The reward-optimal policy of always choosing the option with the higher reinforcement probability, yielding accuracy equal to that probability.
Prediction task.
The canonical probability-learning procedure in which the learner guesses which of two events will occur on each trial before it happens.
Probability learning.
The acquisition, through repeated feedback, of choices that reflect the reward probabilities of options whose odds are never stated.
Probability matching.
Choosing each option in proportion to how often it is reinforced, a suboptimal pattern that scores below maximizing.
Reinforcement learning.
A computational framework in which an agent learns to choose actions that maximize cumulative reward by updating value estimates from prediction errors; the modern synthesis for probability learning.
Reinforcement.
The delivery of a rewarding outcome that follows a choice and raises the probability of repeating it.
Reward prediction error.
The difference between the reward expected and the reward received; a teaching signal encoded by midbrain dopamine neurons that drives learning.
Stimulus sampling.
Estes's theory that each trial conditions a random sample of stimulus elements, whose proportions at equilibrium produce probability matching.
Volatility.
The rate at which the reward probabilities of an environment change; higher volatility warrants a higher learning rate.

Frequently Asked Questions

What is probability learning?
Probability learning is the process by which repeated experience of probabilistic outcomes leads a person or animal to choose in a way that reflects the underlying reward probabilities, even though those probabilities are never stated (Estes, 1950).

What is probability matching?
Probability matching is the finding that learners choose each option roughly in proportion to how often it is reinforced; if one option pays off 75 percent of the time, they choose it about 75 percent of the time (Grant, Hake, & Hornseth, 1951).

Why is probability matching considered suboptimal?
Because the strategy that maximizes reward is to always pick the more frequently reinforced option, giving accuracy equal to its reward probability p, whereas matching yields only p² + (1 − p)², which is lower whenever the options differ (Vulkan, 2000).

How much accuracy does matching cost?
For a 75/25 split, a maximizer is correct 75 percent of the time and a matcher only 62.5 percent, a loss of 12.5 percentage points, roughly twelve or thirteen extra errors per hundred trials (Vulkan, 2000).

Do people always match?
No. Matching is robust but fragile; with many trials, real financial incentives, and clear feedback a substantial proportion of people shift toward maximizing (Shanks, Tunney, & McCarthy, 2002).

How did Estes explain matching?
In stimulus-sampling theory each trial conditions a sample of stimulus elements toward the reinforced response, and at equilibrium the proportion favouring a response equals its reinforcement probability, so matching is simply where conditioning settles (Estes, 1950).

What does volatility have to do with learning?
When an environment changes often, recent outcomes are more informative, so the efficient learner raises its learning rate in volatile conditions and lowers it in stable ones; people approximate this and the anterior cingulate tracks estimated volatility (Behrens, Woolrich, Walton, & Rushworth, 2007).

How is probability learning related to the brain's reward system?
The mismatch between expected and received reward, the reward prediction error, is encoded by phasic dopamine activity, providing the teaching signal that updates choice probabilities during probabilistic learning (Lerner, Holloway, & Seiler, 2021).

References

Baumann, C., Schlegelmilch, R., & von Helversen, B. (2025). Beyond risk preferences in sequential decision-making: How probability representation, sequential structure and choice perseverance bias optimal search. Cognition, 254, 106001. https://doi.org/10.1016/j.cognition.2024.106001

Behrens, T. E. J., Woolrich, M. W., Walton, M. E., & Rushworth, M. F. S. (2007). Learning the value of information in an uncertain world. Nature Neuroscience, 10(9), 1214-1221. https://doi.org/10.1038/nn1954

Edwards, W. (1961). Probability learning in 1000 trials. Journal of Experimental Psychology, 62(4), 385-394. https://doi.org/10.1037/h0041970

Ellerby, Z., & Tunney, R. J. (2019). Probability matching on a simple simulated foraging task: The effects of reward persistence and accumulation on choice behavior. Advances in Cognitive Psychology, 15(2), 111-126. https://doi.org/10.5709/acp-0261-2

Estes, W. K. (1950). Toward a statistical theory of learning. Psychological Review, 57(2), 94-107. https://doi.org/10.1037/h0058559

Grant, D. A., Hake, H. W., & Hornseth, J. P. (1951). Acquisition and extinction of a verbal conditioned response with differing percentages of reinforcement. Journal of Experimental Psychology, 42(1), 1-5. https://doi.org/10.1037/h0054051

Herrnstein, R. J. (1961). Relative and absolute strength of response as a function of frequency of reinforcement. Journal of the Experimental Analysis of Behavior, 4(3), 267-272. https://doi.org/10.1901/jeab.1961.4-267

Humphreys, L. G. (1939). Acquisition and extinction of verbal expectations in a situation analogous to conditioning. Journal of Experimental Psychology, 25(3), 294-301. https://doi.org/10.1037/h0053555

Lerner, T. N., Holloway, A. L., & Seiler, J. L. (2021). Dopamine, updated: Reward prediction error and beyond. Current Opinion in Neurobiology, 67, 123-130. https://doi.org/10.1016/j.conb.2020.10.012

Shanks, D. R., Tunney, R. J., & McCarthy, J. D. (2002). A re-examination of probability matching and rational choice. Journal of Behavioral Decision Making, 15(3), 233-250. https://doi.org/10.1002/bdm.413

Vulkan, N. (2000). An economist's perspective on probability matching. Journal of Economic Surveys, 14(1), 101-118. https://doi.org/10.1111/1467-6419.00106