Abstract

Association learning is a form of association: the process by which an organism comes to link events, cues, and responses so that one predicts or evokes another. This article covers the two experimental traditions that established it — Pavlovian conditioning, in which a neutral stimulus paired with a biologically significant one acquires a conditioned response, and instrumental conditioning, in which responses followed by satisfying consequences are strengthened. It develops the Rescorla-Wagner model, which recast learning as driven by prediction error rather than mere pairing and so explains blocking; the attentional theories of cue competition that extend it; statistical learning, in which regularities are extracted from mere exposure; and extinction, the context-dependent inhibitory learning that suppresses an association without erasing it. A worked example derives the blocking effect trial by trial.

Keywords: association learning, prediction error, classical conditioning, blocking, statistical learning

Association learning is the mechanism by which experience of the world's regularities reshapes behaviour: when two events reliably co-occur, or when an action reliably produces an outcome, the organism comes to treat the one as a sign of the other. It is the experimental, measurable core of the older philosophical association of ideas, and for a century it has been the central phenomenon of the science of learning. This article follows it from Pavlov and Thorndike, through the prediction-error models that made it quantitative, to the attentional, statistical, and extinction phenomena that define its modern form.

Key Takeaways
  • Association learning is the acquisition of a link between two stimuli, or between a response and its outcome, inferred from a measured change in behaviour when the events are paired.
  • It was put on an experimental footing by Pavlov's conditioned reflexes and Thorndike's law of effect, which established the two forms: stimulus-stimulus and response-outcome.
  • The Rescorla-Wagner model recast learning as error-driven — strength grows in proportion to the gap between the outcome and the outcome the present cues already predict — which explains the blocking effect.
  • Attentional theories add that a cue's associability also changes: Mackintosh ties it to predictiveness, Pearce and Hall to uncertainty, and human data show both operate.
  • Statistical learning extracts co-occurrence regularities from mere exposure, and extinction suppresses an association through new inhibitory learning without erasing the original.

What Association Learning Is

Association learning is the process by which an organism comes to link events in the world — a cue with an outcome, or an action with its consequence — so that experience of their relationship changes what it does. It is inferred, never observed directly: the association is a theoretical link whose existence is read off a measured change in behaviour when the paired events recur. MeSH files it beneath association, and the two words are often used interchangeably, but association learning names specifically the experimental, quantifiable acquisition of such links, as distinct from the broader philosophical principle that the mind is organised by connection.

The construct is defined by what it relates rather than by any single mechanism. On a functional view, learning is simply a change in an organism's behaviour that is due to regularities in its environment — a definition that identifies the phenomenon without committing to how it is realised, and so lets association learning, propositional reasoning, and statistical extraction all count as instances to be told apart empirically (De Houwer et al., 2013). That neutrality matters, because the century-long argument this article traces is precisely about which mechanism does the work: a bond strengthened by contiguity, an error-correcting rule, a shift of attention, or an inference about relations among events.

Pavlovian and Instrumental Learning

Two experimental traditions established association learning as a laboratory science, each isolating a different kind of link. Pavlov's programme showed that a neutral stimulus, repeatedly presented just before a biologically significant one, comes to evoke a conditioned response: the association between the conditioned stimulus (CS) and the unconditioned stimulus (US) is inferred from the response the CS acquires (Pavlov, 1927). Thorndike, working with animals in puzzle boxes, established the complementary instrumental case: responses followed by a satisfying consequence become more strongly connected to the situation in which they occurred, the principle he named the law of effect (Thorndike, 1911). Between them, Pavlov and Thorndike converted the association of ideas into two measurable forms — stimulus-stimulus and response-outcome.

Pavlovian conditioning is not the arbitrary substitution of one reflex for another. On a functional-behavioural analysis, the conditioned response is an adaptation that prepares the organism for the biologically important event the CS predicts, so its form reflects what the animal must do about the US rather than simply copying the reflex the US evokes (Domjan, 2005). This functional reading anticipates the central modern claim, developed below, that conditioning is the learning of predictive relations among events — the acquisition of information about what signals what — and not the mechanical stamping-in of a bond by mere temporal pairing.

Prediction Error and the Rescorla-Wagner Model

The question these findings forced was quantitative: what determines how much an association grows on a given trial? The answer that reorganised the field is the Rescorla-Wagner model. It holds that associative strength changes in proportion to the prediction error — the discrepancy between the outcome that occurs and the outcome already predicted by all the cues present on that trial: ΔV = αβ(λ − ΣV), where λ is the outcome's asymptotic strength and ΣV the summed strength of the cues present (Rescorla & Wagner, 1972). Learning is therefore fast when the outcome is surprising and slows to nothing as the cues come to predict it, producing the negatively accelerated acquisition curve.

The model's signature success is blocking, the effect Kamin reported when he found that a cue already trained to predict a shock prevents an added, redundant cue from being learned about at all (Kamin, 1969). If one cue already predicts the outcome, there is almost no error left for the redundant cue added alongside it, so it is barely learned — direct evidence that pairing alone is insufficient and that association tracks predictive information. Rescorla later drew the moral explicitly, arguing that conditioning is the learning of relations among events and the acquisition of information about contingency, not the transfer of a reflex by contiguity (Rescorla, 1988). The same error-correcting logic generalises: temporal-difference learning extends the Rescorla-Wagner rule from single trials to sequences of predictions and is the computational foundation of modern reinforcement learning (Sutton & Barto, 2018). The demonstration below drives the rule directly, showing how a pretrained cue steals the available error and blocks a second.

Blocking: a pretrained cue steals the error

Two cues, A and B, are paired with the same outcome. Set how many trials cue A is trained alone first, then watch the compound AB phase. Because learning corrects the error against the combined prediction of both cues, a well-trained A leaves almost no error for B to absorb, and B’s association stalls near zero however often it is paired with the outcome.

λ = 100B added1000cue Acue BTrial
final V(A) = 97.1
final V(B) = 2.9
20 trials total

With no pretraining the two cues share the outcome and both learn. With A already trained, the compound is fully predicted from the first compound trial, the error term is near zero, and B is blocked — direct evidence that association learning tracks predictive information, not mere co-occurrence.

Attention and Cue Competition

What the Rescorla-Wagner model holds fixed is the learning rate α — the cue's associability, how readily it enters new associations — and two classic theories set it free from opposite directions. Mackintosh proposed that a cue's associability rises when it is a good predictor of the outcome, so animals come to learn faster about informative cues and more slowly about uninformative ones (Mackintosh, 1975). Pearce and Hall proposed the reverse: a cue commands more processing when its outcome is uncertain, so associability is high while the consequences of a cue are not yet well predicted and falls once they are (Pearce & Hall, 1980). The two rules make opposite predictions for a cue that comes to predict its outcome reliably.

The integrative review of human studies concludes that both principles operate, on different timescales, and that attention and associative learning are reciprocally coupled rather than one being reducible to the other: predictiveness governs a slow, learned bias in what is attended, while surprise drives a faster, trial-by-trial modulation (Le Pelley et al., 2016). The demonstration below lets the reader switch between the Mackintosh and Pearce-Hall rules, and between a reliable and an unreliable outcome, and watch a single cue's associability diverge.

Attention: which cues become learnable

A cue’s associability — how readily it enters new associations — is not fixed. Switch between the two classic attentional rules and between a reliable and an unreliable outcome, and watch the associability of one cue evolve. The two theories make opposite predictions for a cue that comes to predict its outcome well.

1.00αTrial
final associability α = 0.05
predictiveness drives attention

For a reliable cue the two rules diverge sharply: Mackintosh raises its associability, because it is the best available predictor, while Pearce-Hall lowers it, because a well-predicted outcome is no longer surprising. Under partial reinforcement, where uncertainty persists, Pearce-Hall keeps associability high. Human data show both effects operate, which is why the integrative accounts combine the two.

Table 1. The major formal models of associative learning and what each holds to change with experience.
Model What changes with learning Signature phenomenon
Rescorla-Wagner (1972) Each cue's associative strength, updated by the shared prediction error summed over all cues present. Blocking: a pretrained cue leaves no error for a redundant added cue to absorb.
Mackintosh (1975) The cue's associability, which rises when the cue is the best available predictor of the outcome. Learned relevance: a reliable predictor is learned about faster on later problems.
Pearce-Hall (1980) The cue's associability, which stays high while its outcome remains uncertain and falls once it is well predicted. Surprise-driven attention: uncertain cues command more processing.
Propositional account Consciously held beliefs about the relations between events, formed and tested like hypotheses. Sensitivity to instructions and inferred structure that bare contiguity cannot explain.

Statistical Learning

Association learning need not involve reinforcement or even intention. In the founding demonstration of statistical learning, 8-month-old infants heard two minutes of a continuous stream of nonsense syllables in which the only cue to the word boundaries was statistical — the transitional probability from one syllable to the next was high within a word and low across a boundary — and afterwards distinguished the stream's words from equally familiar non-words, showing they had segmented it from the co-occurrence statistics alone (Saffran, Aslin, & Newport, 1996). The mechanism is a form of association learning by mere exposure: it extracts co-occurrence regularities automatically, without feedback and without explicit effort, and operates across modalities and species (Aslin, 2017).

Later work has complicated the picture in a productive way. A critical review argues that statistical learning is not a single, unitary ability but a family of modality- and stimulus-specific computations that share a name rather than a mechanism, which reframes how the phenomenon should be measured and what generalises across its many demonstrations (Frost, Armstrong, & Christiansen, 2019). The demonstration below makes the segmentation concrete: a threshold on the transitional probability cuts a continuous syllable stream into candidate words.

Statistical learning: finding words in a stream

A continuous stream of syllables carries no pauses, only statistics. Within a word each syllable always follows the last, so the transitional probability is 1.00; across a word boundary the probability drops to about 0.33, because each word can be followed by several others. Slide the threshold: cut the stream wherever the transitional probability falls below it, and the words fall out.

golatu|tibudo|daropi
|pabiku|golatu|daropi
|tibudo|pabiku|golatu
chunks recovered = 9
correct boundaries = 8/8
false cuts = 0
words recovered exactly

Any threshold between 0.33 and 1.00 cuts every boundary and none of the within-word transitions, recovering the four words exactly — no meaning, no reinforcement, only the co-occurrence statistics. This is how 8-month-old infants, and adults hearing an unfamiliar language, find candidate words after only minutes of exposure.

Extinction

An association once acquired can be weakened by presenting the cue without its outcome, but this extinction does not erase the original learning. The response returns with a change of context (renewal), with the passage of time (spontaneous recovery), or after a reminder of the outcome (reinstatement) — evidence that extinction is new, context-dependent inhibitory learning laid over the original association rather than its deletion (Bouton, 2004). The learner ends up holding two competing associations, one excitatory and one inhibitory, and which one governs behaviour depends on the context that retrieves it.

This has direct clinical force. Exposure therapy for fear and anxiety is extinction learning, and understanding it as inhibitory rather than erasive explains why treated fears relapse and reframes therapy as the task of making the new, safe learning win the retrieval competition across contexts (Craske, Hermans, & Vervliet, 2018). The behavioural and neurobiological mechanisms of extinction — the context-dependence, the recovery phenomena, and the circuits that suppress a learned association without deleting it — have since been synthesised across Pavlovian and instrumental learning (Bouton, Maren, & McNally, 2021).

The core phenomena that any account of association learning must explain are summarised in Table 1; each is a place where mere contiguity fails and a predictive, error-driven mechanism succeeds.

Table 1. Core phenomena of association learning and what each reveals.
Phenomenon What happens What it reveals
Acquisition Associative strength grows trial by trial toward an asymptote as the prediction error shrinks. Learning is graded and driven by the discrepancy between the outcome and the current expectation.
Blocking A cue added to another that already predicts the outcome gains almost no strength despite repeated pairing. Cues compete; only an unpredicted outcome leaves error for a new cue to absorb.
Contingency A cue paired with the outcome as often as the background is is not learned about at all. Conditioning tracks the predictive relation between cue and outcome, not the raw pairing count.
Extinction Presenting the cue without the outcome suppresses the response while leaving the original link intact. Extinction is new, context-dependent inhibitory learning rather than the deletion of the association.
Renewal An extinguished response reappears when the cue is met in a context other than the one where extinction occurred. The original association survives extinction and is gated by the retrieval context.

Worked Example

Blocking can be worked out by hand from the Rescorla-Wagner rule. Take two cues, A and B, each capable of predicting an outcome of asymptotic strength λ = 100, with combined learning rate αβ = 0.30. In a first phase cue A is trained alone; because only A is present, its strength follows the single-cue rule Vn = 100(1 − 0.7n), reaching V1 = 30, V2 = 51, V3 = 65.7, and after eight trials VA = 100(1 − 0.78) = 94.2.

Now the compound AB is trained. On every compound trial the error is computed against the summed prediction of both cues, ΔV = 0.30 × (100 − (VA + VB)), and that same increment is added to each. On the first compound trial the error is 100 − 94.2 = 5.8, so each cue gains 0.30 × 5.8 = 1.7: A rises to 95.9 and B to just 1.7. Because both cues always receive the identical increment, their difference stays fixed at VA − VB = 94.2 while their sum climbs to the ceiling of 100. After the twelve compound trials the sum is essentially 100, which splits as VA = 97.1 and VB = 2.9: cue B is blocked, having learned almost nothing despite twelve pairings with the outcome.

The contrast is the control condition. With no pretraining, A and B enter the compound phase together from zero, share every increment symmetrically, and each reaches V = 50 — so B learns fully. The only difference between B ending at 2.9 and B ending at 50 is whether another cue already predicted the outcome, which is exactly the point: association learning tracks predictive information, not the number of pairings. Figure 1 shows the two outcomes side by side.

Figure 1

Final associative strength of a redundant cue B after twelve compound AB trials (αβ = 0.30, λ = 100). When cue A is pretrained, it steals the prediction error and B is blocked; with no pretraining, the two cues share the outcome equally.

Blocking of a redundant cue A grouped bar chart. With no pretraining, cue A and cue B each reach associative strength fifty. With cue A pretrained for eight trials, cue A reaches ninety-seven and cue B reaches only about three, showing that B is blocked. Final strength V 0 50 100 λ = 100 A = 50 B = 50 No pretraining A = 97 B = 2.9 A pretrained (8 trials)
Note. Strengths follow ΔV = 0.30(λ − ΣV) with λ = 100. Pretraining drives VA to 94.2 before the compound; the fixed difference VA − VB = 94.2 is preserved as the sum approaches 100, giving VA = 97.1 and VB = 2.9. The blocked cue B learns almost nothing despite twelve pairings.

Discussion

Association learning has proved one of the most durable units of explanation in psychology, but its modern form is far from the simple bond it began as. Pavlov and Thorndike established that stimulus-stimulus and response-outcome links can be measured (Pavlov, 1927); (Thorndike, 1911); the Rescorla-Wagner model then turned the bare idea that associations form by contiguity into a precise error-correcting rule that predicts phenomena — blocking above all — which contiguity alone cannot (Rescorla & Wagner, 1972). The attentional theories complicate the picture further, making learning depend on the informativeness of cues and on recent surprise rather than on pairing per se (Mackintosh, 1975); (Pearce & Hall, 1980).

The tradition's central tension is whether association is the whole story. A strong challenge holds that much of what looks like automatic association in humans is better described as propositional reasoning about the relations between events — the learner forming and testing beliefs, not just accumulating bonds (Rescorla, 1988). Whether human learning is fundamentally associative or inferential remains genuinely open, and the evidence is mixed enough that the field has not settled it (Shanks, 2010). What is not in doubt is that the associative link — weighted, error-driven, and competing with its neighbours for the available prediction error — remains indispensable, whatever higher-level process may sit above it.

Current Directions

The most active bridge is to neuroscience. The Rescorla-Wagner error term found a physical realisation in the reward prediction error signalled by midbrain dopamine neurons, whose phasic firing tracks the discrepancy between received and predicted reward almost exactly as the theory requires (Schultz, Dayan, & Montague, 1997). The contemporary debate concerns how literally that identification should be taken and what the dopamine signal is really computing, tying association learning theory to reinforcement learning and to the neural code for value (Gershman & Uchida, 2019). The reinforcement-learning framework itself, built on the temporal-difference generalisation of the error rule, now supplies the common computational language for both the behavioural and the neural work (Sutton & Barto, 2018).

A second front concerns the scope and unity of the phenomenon. The reframing of statistical learning as a family of modality-specific computations rather than a single ability is reshaping how learning by mere exposure is studied and what should be expected to transfer across tasks (Frost, Armstrong, & Christiansen, 2019). Running through both fronts is the older, unresolved question of whether an associative account can be maintained against a propositional one, now pursued with the tools of both cognitive experiment and computational modelling (De Houwer et al., 2013).

Common Misconceptions

Associations form automatically whenever two things occur together.
Mere temporal pairing is neither necessary nor sufficient. The Rescorla-Wagner model and the blocking effect show that an association grows only to the extent that the outcome is surprising; a cue whose outcome is already predicted is barely learned however often it is paired (Rescorla & Wagner, 1972); (Rescorla, 1988).
Extinction erases a learned association.
Presenting a cue without its outcome suppresses the response but leaves the original association intact: it returns with a change of context, the passage of time, or a reminder. Extinction is new inhibitory learning layered over the old link, not its deletion (Bouton, 2004); (Bouton et al., 2021).
Association learning always needs reinforcement.
Statistical learning shows that co-occurrence regularities are extracted from mere exposure, with no reward, feedback, or intention — infants segment a speech stream from transitional probabilities alone (Saffran et al., 1996); (Aslin, 2017).

Glossary

Associability (α).
The readiness of a cue to enter new associations; the learning-rate parameter that attentional theories allow to change with a cue's predictiveness or the uncertainty of its outcome.
Association learning.
The acquisition of a link between two stimuli, or between a response and its outcome, inferred from a measured change in behaviour when the events are paired.
Blocking.
The finding that a cue already predicting an outcome prevents a second, redundant cue paired with it from being learned; a key demonstration that association tracks prediction error, not pairing.
Conditioned response (CR).
The learned response evoked by a conditioned stimulus after it has been paired with an unconditioned stimulus.
Conditioned stimulus (CS).
An initially neutral stimulus that, through pairing with a biologically significant one, comes to evoke a learned response.
Contingency.
The predictive relation between a cue and an outcome — how much the cue changes the probability of the outcome; on Rescorla's view, the true content of what conditioning learns.
Extinction.
The reduction of a conditioned response when the cue is presented without its outcome; new inhibitory learning that suppresses the original association rather than erasing it.
Instrumental (operant) conditioning.
Association learning in which a response is strengthened or weakened by its consequences, following Thorndike's law of effect.
Law of effect.
Thorndike's principle that responses followed by a satisfying consequence become more strongly associated with the situation in which they occur; the basis of instrumental association.
Pavlovian (classical) conditioning.
Association learning in which a neutral stimulus paired with a biologically significant one comes to evoke a conditioned response.
Prediction error.
The discrepancy between the outcome that occurs and the outcome the present cues already predict; the driving term of the Rescorla-Wagner model, so learning is proportional to surprise.
Rescorla-Wagner model.
The theory that associative strength changes in proportion to the prediction error, with all cues present sharing the available error; it explains the acquisition curve and blocking.
Reward prediction error.
A prediction error about reward, signalled by midbrain dopamine neurons; the candidate neural realisation of the Rescorla-Wagner error term.
Statistical learning.
The extraction of co-occurrence regularities from mere exposure, without reinforcement or explicit intent; demonstrated by infants segmenting a speech stream from transitional probabilities.
Transitional probability.
The probability that one item follows another; high within a learned unit (such as a word) and low across its boundary, the statistic that supports segmentation.
Unconditioned stimulus (US).
A stimulus that evokes a response without prior learning; the biologically significant event whose association a conditioned stimulus comes to signal.

Key Researchers

Samuel J. Gershman. Links association learning theory to reinforcement learning and to the dopaminergic reward-prediction-error signal, giving the Rescorla-Wagner error term a neural interpretation. ORCID - Google Scholar - Faculty page

Jan De Houwer. Leads the functional and propositional analysis of learning, arguing that much human association learning reflects reasoning about relations among events rather than the automatic formation of bonds. ORCID - Google Scholar

Stephen Maren. Studies the neurobiology of Pavlovian fear conditioning and extinction, and the brain circuits that suppress a learned association without deleting it. ORCID - Google Scholar - Faculty page

Ivan Pavlov (1849-1936). Founded the experimental study of classical conditioning, demonstrating that a neutral stimulus paired with a biologically significant one comes to evoke a learned, conditioned response. Wikipedia - Wikidata

Robert A. Rescorla (1940-2020). Co-authored the Rescorla-Wagner prediction-error model and reframed conditioning as the learning of predictive relations among events rather than the stamping-in of a reflex by contiguity. Wikipedia - Wikidata

Wolfram Schultz. Discovered that midbrain dopamine neurons signal a reward prediction error, giving the associative error term of the Rescorla-Wagner model a concrete neural substrate. ORCID - Google Scholar - Faculty page

David R. Shanks. Studies human associative and contingency learning and the boundary between associative and inferential accounts of how people learn about relations among events. ORCID - Google Scholar - Faculty page

Edward L. Thorndike (1874-1949). Formulated the law of effect from the puzzle-box experiments, establishing instrumental association as the complement to Pavlovian conditioning. Wikipedia - Wikidata

Frequently Asked Questions

What is association learning? Association learning is the process by which an organism comes to link events, so that a cue predicts an outcome or an action is tied to its consequence. It is inferred from a measured change in behaviour when the events are paired, and it covers Pavlovian conditioning, instrumental conditioning, and learning by mere exposure (De Houwer et al., 2013).

What is the difference between Pavlovian and instrumental conditioning? Pavlovian conditioning links two stimuli: a neutral cue paired with a biologically significant one comes to evoke a conditioned response. Instrumental conditioning links a response to its outcome: an action followed by a satisfying consequence is strengthened, following Thorndike's law of effect (Pavlov, 1927); (Thorndike, 1911).

What does the Rescorla-Wagner model say? It says associative strength changes in proportion to the prediction error, the gap between the outcome that occurs and the outcome the cues present already predict. Learning is fast when the outcome is surprising and slows as it becomes predicted, and all cues present share the available error, which explains blocking (Rescorla & Wagner, 1972).

What is blocking, and why does it matter? Blocking is the finding that a cue already predicting an outcome prevents a second, redundant cue paired with it from being learned. It matters because it shows that mere pairing does not build an association; what matters is whether the cue adds predictive information, which is exactly what a prediction-error model captures (Rescorla & Wagner, 1972); (Rescorla, 1988).

How does attention affect what is learned? A cue's associability is not fixed. Mackintosh held that it rises for cues that predict well; Pearce and Hall held that it rises for cues whose outcomes are uncertain. Human studies indicate both operate, on different timescales, so attention and association learning are reciprocally coupled (Mackintosh, 1975); (Le Pelley et al., 2016).

What is statistical learning? Statistical learning is the extraction of co-occurrence regularities from mere exposure, without reward or intention. In the founding study, 8-month-old infants segmented a continuous speech stream into words using only the transitional probabilities between syllables (Saffran et al., 1996); (Aslin, 2017).

Does extinction erase a learned association? No. Presenting the cue without the outcome reduces the response, but the association returns with a change of context, the passage of time, or a reminder of the outcome. Extinction is new, context-dependent inhibitory learning laid over the original link rather than its deletion (Bouton, 2004); (Bouton et al., 2021).

Is human learning really associative, or is it reasoning? This is genuinely unsettled. Prediction-error and attentional models already move beyond simple pairing, and one influential view holds that much human learning is propositional, forming and testing beliefs about relations among events rather than accreting bonds. The evidence supports parts of both accounts (De Houwer et al., 2013); (Shanks, 2010).

References

Aslin, R. N. (2017). Statistical learning: A powerful mechanism that operates by mere exposure. WIREs Cognitive Science, 8(1-2), e1373. https://doi.org/10.1002/wcs.1373

Bouton, M. E. (2004). Context and behavioral processes in extinction. Learning & Memory, 11(5), 485-494. https://doi.org/10.1101/lm.78804

Bouton, M. E., Maren, S., & McNally, G. P. (2021). Behavioral and neurobiological mechanisms of Pavlovian and instrumental extinction learning. Physiological Reviews, 101(2), 611-681. https://doi.org/10.1152/physrev.00016.2020

Craske, M. G., Hermans, D., & Vervliet, B. (2018). State-of-the-art and future directions for extinction as a translational model for fear and anxiety. Philosophical Transactions of the Royal Society B, 373(1742), 20170025. https://doi.org/10.1098/rstb.2017.0025

De Houwer, J., Barnes-Holmes, D., & Moors, A. (2013). What is learning? On the nature and merits of a functional definition of learning. Psychonomic Bulletin & Review, 20(4), 631-642. https://doi.org/10.3758/s13423-013-0386-3

Domjan, M. (2005). Pavlovian conditioning: A functional perspective. Annual Review of Psychology, 56, 179-206. https://doi.org/10.1146/annurev.psych.55.090902.141409

Frost, R., Armstrong, B. C., & Christiansen, M. H. (2019). Statistical learning research: A critical review and possible new directions. Psychological Bulletin, 145(12), 1128-1153. https://doi.org/10.1037/bul0000210

Gershman, S. J., & Uchida, N. (2019). Believing in dopamine. Nature Reviews Neuroscience, 20(11), 703-714. https://doi.org/10.1038/s41583-019-0220-7

Kamin, L. J. (1969). Predictability, surprise, attention, and conditioning. In B. A. Campbell & R. M. Church (Eds.), Punishment and aversive behavior (pp. 279-296). Appleton-Century-Crofts.

Le Pelley, M. E., Mitchell, C. J., Beesley, T., George, D. N., & Wills, A. J. (2016). Attention and associative learning in humans: An integrative review. Psychological Bulletin, 142(10), 1111-1140. https://doi.org/10.1037/bul0000064

Mackintosh, N. J. (1975). A theory of attention: Variations in the associability of stimuli with reinforcement. Psychological Review, 82(4), 276-298. https://doi.org/10.1037/h0076778

Pavlov, I. P. (1927). Conditioned reflexes: An investigation of the physiological activity of the cerebral cortex (G. V. Anrep, Trans.). Oxford University Press.

Pearce, J. M., & Hall, G. (1980). A model for Pavlovian learning: Variations in the effectiveness of conditioned but not of unconditioned stimuli. Psychological Review, 87(6), 532-552. https://doi.org/10.1037/0033-295X.87.6.532

Rescorla, R. A. (1988). Pavlovian conditioning: It's not what you think it is. American Psychologist, 43(3), 151-160. https://doi.org/10.1037/0003-066X.43.3.151

Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), Classical conditioning II: Current research and theory (pp. 64-99). Appleton-Century-Crofts.

Saffran, J. R., Aslin, R. N., & Newport, E. L. (1996). Statistical learning by 8-month-old infants. Science, 274(5294), 1926-1928. https://doi.org/10.1126/science.274.5294.1926

Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. Science, 275(5306), 1593-1599. https://doi.org/10.1126/science.275.5306.1593

Shanks, D. R. (2010). Learning: From association to cognition. Annual Review of Psychology, 61, 273-301. https://doi.org/10.1146/annurev.psych.093008.100519

Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.

Thorndike, E. L. (1911). Animal intelligence: Experimental studies. Macmillan.