Abstract
Association, which MeSH classifies under learning, is the principle that ideas, events, and mental states become linked in the mind so that one tends to evoke another. This article traces the concept from the British empiricists' association of ideas and the classical laws of association — contiguity, similarity, and contrast — through the experimental study of associative learning in Pavlov and Thorndike, to the Rescorla-Wagner model, which recast association as error-driven learning proportional to the discrepancy between an outcome and the outcome its cues already predict. It examines the attentional theories that extend that model, the reframing of conditioning as the learning of predictive relations among events rather than mere temporal pairing, and the spreading-activation account of semantic memory, in which associative links between concept nodes explain priming and the organisation of knowledge.
Keywords: association, laws of association, associative learning, prediction error, classical conditioning
Association is the oldest idea in the science of mind: the observation that one thought calls up another, and that experiences occurring together come to be linked so that later one revives the rest. It began as a philosopher's principle for how ideas cohere, became the central mechanism of a laboratory science of learning, and survives today as the organising assumption of computational models of memory and reward. This article follows that arc as a single thread — from the association of ideas, through the laws that were said to govern it, to the error-driven learning rules and semantic networks that give the principle its modern, quantitative form.
- Association is the linking of ideas, events, or stimuli so that one tends to evoke another; it is the common mechanism behind the association of ideas, associative learning, and the organisation of semantic memory.
- The classical laws of association — contiguity, similarity (resemblance), and contrast, with cause and effect added by Hume — describe the conditions under which two mental items become linked.
- Associative learning was put on an experimental footing by Pavlov's conditioned reflexes and Thorndike's law of effect, establishing that associations between stimuli and between actions and outcomes can be measured.
- The Rescorla-Wagner model recast association as error-driven: strength grows in proportion to the prediction error, the gap between the outcome and the outcome the present cues already predict, which explains blocking and the acquisition curve.
- In memory, spreading activation across a network of associatively linked concept nodes explains semantic priming and why related ideas come to mind together.
What Association Is
Association is the process by which two mental contents — ideas, images, stimuli, responses, or events — become linked, so that the later occurrence of one tends to bring the other to mind or to elicit the response the other controls. MeSH defines the descriptor tersely, but the psychological construct is broader than any single definition: it names at once a relation between mental items, the process that forms that relation, and the principle that the mind is organised by such links. The unifying claim of the associationist tradition is that complex mental life is built from simple connections formed by experience, rather than from innate structure imposed on it.
Three senses of the word recur across this article and are worth separating at the outset. The association of ideas is the classical, introspective sense: the tendency of one thought to summon another, which the British empiricists took to be the basic law of mental life. Associative learning is the experimental sense: the measurable acquisition of a link between a stimulus and an outcome, or between an action and its consequence, studied in conditioning experiments. And associative memory is the representational sense: the storage of knowledge as a network of linked concepts, in which retrieving one activates its neighbours. The sections that follow treat each in turn, but the same underlying idea runs through all three — that contiguity in experience builds connection in the mind.
Types of Association
In the Medical Subject Headings vocabulary, the descriptor Association (D001244) is placed beneath learning and carries one narrower descriptor of its own, shown in Table 1. MeSH is an indexing classification built for retrieving the biomedical literature, not a theory of the mind, so this single subtype captures only the learning-theoretic sense of association and leaves aside the association of ideas and the semantic-network sense treated later. The distinctions a psychologist actually draws — by content (idea-to-idea, stimulus-to-stimulus, or stimulus-to-response links), by the law that forms them (contiguity, similarity, contrast), or by whether they are excitatory or inhibitory — cut across the MeSH tree rather than nesting within it, and are orthogonal to the indexing hierarchy shown here.
| Subtype (MeSH descriptor) | What it denotes | Site status |
|---|---|---|
| Association Learning | The formation of a learned link between two stimuli, or between a stimulus and a response, through their paired occurrence; the learning-theoretic core of the sections below. | No dedicated page yet. |
The Laws of Association
The idea that thought is governed by laws of association is ancient — Aristotle noted that recollection moves from one item to another by similarity, contrast, or contiguity — but its canonical modern statement is Hume's. In A Treatise of Human Nature he proposed that ideas are connected by three principles: resemblance, contiguity in time or place, and cause and effect, which together act as a gentle, associative force drawing one idea after another (Hume, 1739). The later empiricists elaborated the list, but contiguity — the linking of items experienced close together in time or space — became the primary law, the one from which the others were often derived.
The associationists treated these laws as the mental analogue of the laws of physics: a small set of principles from which the whole of thought could be reconstructed. William James, while accepting association as fundamental, made a decisive move in reinterpreting it. Rather than a law relating ideas directly, he argued, association is a consequence of the brain's habits of conduction: when two processes have occurred together, the later excitation of one tends to propagate to the other, so the association of ideas is really the association of neural processes (James, 1890). This shift — from a law of ideas to a law of the nervous system — anticipated both the conditioning tradition and the neural-network models that would follow. The first experimental test of the laws came from Ebbinghaus, who used lists of nonsense syllables to measure how associations between items are built up by repetition and lost over time, turning association from a philosophical principle into a quantity that could be plotted (Ebbinghaus, 1913).
From Ideas to Behavior: Associative Learning
The twentieth century moved association out of the study and into the laboratory by asking not how ideas connect but how behaviour changes when events are paired. Pavlov's programme showed that a neutral stimulus, repeatedly presented just before a biologically significant one, comes to evoke a conditioned response: the association between the conditioned stimulus (CS) and the unconditioned stimulus (US) is inferred from the response the CS acquires (Pavlov, 1927). Thorndike, working with animals in puzzle boxes, established the complementary case of instrumental association: responses followed by a satisfying consequence become more strongly connected to the situation in which they occurred, the principle he named the law of effect (Thorndike, 1911). Between them, Pavlov and Thorndike converted the association of ideas into two measurable forms of association in behaviour — stimulus-stimulus and response-outcome.
Two features of associative learning show that it is more than the strengthening of a bond by repetition. First, an association once acquired can be extinguished by presenting the cue without the outcome, but extinction does not erase it: the response returns with a change of context (renewal), with the passage of time (spontaneous recovery), or after a reminder of the outcome (reinstatement), which shows that extinction is new, context-dependent inhibitory learning laid over the original association rather than its deletion (Bouton, 2004). Second, conditioning is functional: rather than an arbitrary substitution of one reflex for another, it prepares the organism for the biologically important event the CS predicts, so the form of the conditioned response reflects what the animal must do about the US (Domjan, 2005). The demonstration below shows the two phases of a conditioned association — its acquisition over reinforced trials and its extinction when the outcome is withdrawn — and the residual strength that extinction leaves behind.
Reinforced trials build an association; non-reinforced trials weaken it. Set how many conditioning trials the cue receives and watch associative strength climb toward the asymptote, then decay once the outcome is withheld. The same prediction-error rule drives both the rise and the fall.
Acquisition is negatively accelerated — most learning happens on the first few trials, when prediction error is large — and extinction mirrors it, decaying fastest at first. The behavioural response falls to near zero, but the model captures only the surface: renewal and spontaneous recovery show the original association is inhibited, not erased.
Models of Associative Strength
The question these findings forced was quantitative: what determines how much an association grows on a given trial? The answer that reorganised the field is the Rescorla-Wagner model. It holds that associative strength changes in proportion to the prediction error — the discrepancy between the outcome that occurs and the outcome already predicted by all the cues present on that trial: ΔV = αβ(λ − ΣV), where λ is the outcome's asymptotic strength and ΣV the summed strength of the cues present (Rescorla & Wagner, 1972). Learning is therefore fast when the outcome is surprising and slows to nothing as the cues come to predict it, producing the negatively accelerated acquisition curve. The model's signature success is blocking: if one cue already predicts the outcome, there is little error left for a second, redundant cue added alongside it, so the redundant cue is barely learned — direct evidence that pairing alone is insufficient and that association tracks predictive information. Rescorla later drew the moral explicitly, arguing that conditioning is the learning of relations among events and the acquisition of information about contingency, not the stamping-in of a reflex by contiguity (Rescorla, 1988).
What the Rescorla-Wagner model leaves out is a changing role for attention, and two classic theories supply it from opposite directions. Mackintosh proposed that a cue's associability rises when it is a good predictor of the outcome, so animals learn faster about informative cues (Mackintosh, 1975); Pearce and Hall proposed the reverse, that a cue commands more processing when its outcome is uncertain, so attention is paid to cues whose consequences are not yet well predicted (Pearce & Hall, 1980). The integrative review of human studies concludes that both principles operate, on different timescales, and that attention and associative learning are reciprocally coupled rather than one being reducible to the other (Le Pelley et al., 2016). The demonstration below lets the reader drive the Rescorla-Wagner rule directly, varying the learning rate and the outcome's strength and watching how the prediction error shapes the acquisition curve.
Associative strength grows by a fraction of the current prediction error on each trial. Set the learning rate and step through the trials to trace the acquisition curve. Turn on blocking to add a cue that the outcome is already predicted by another cue — and watch how little the new association can grow.
Each trial closes a fixed fraction of the remaining gap to the asymptote, so the increments shrink and the curve flattens — most learning happens early, when prediction error is large. Under blocking the pretrained cue leaves almost no error to teach the new cue, which stalls far below the asymptote however many trials pass.
Associations in Memory: Semantic Networks
The associationist principle re-entered cognitive psychology in the form of a theory of memory. In the spreading-activation account, semantic memory is a network of concept nodes joined by associative links whose strength reflects how closely related the concepts are; retrieving a concept activates its node, and that activation spreads outward along the links to neighbouring concepts, decaying with distance (Collins & Loftus, 1975). The model explains semantic priming — that reading NURSE speeds recognition of DOCTOR — as the pre-activation of an associate through the link between them, and it explains the graded structure of the effect: closely linked concepts prime each other strongly, distant ones weakly. Association here is not a bond between reflexes but a weighted edge in a graph of meaning.
The tradition continues in contemporary computational models. Current reviews of semantic memory contrast associative-network models, in which links are laid down between discrete concepts, with distributional models, which learn a concept's meaning from the statistics of the contexts in which its word appears (Kumar, 2021). The distributional, vector-space models are the direct computational descendants of associationism: they build meaning from co-occurrence — a statistical form of contiguity — and represent it as position in a high-dimensional space rather than as explicit links, a difference whose cognitive interpretation is still debated (Guenther et al., 2019). The demonstration below makes the spreading-activation idea concrete: activating one concept in a small associative network sends graded activation to its neighbours, strongest along the most direct, heavily weighted links.
Concepts in memory are joined by associative links. Click a concept to activate it and watch activation spread outward along the links, weakening with each step. This is why reading doctor speeds recognition of nurse: the neighbours are primed before you ever see them.
Activation falls by half at each step from the source, so directly linked concepts are primed strongly, concepts two links away only weakly, and unconnected concepts not at all. The graph distance between two concepts is the model’s measure of how closely they are associated.
Worked Example
The Rescorla-Wagner model makes the growth of an association fully quantitative, so a single cue's learning curve can be worked out by hand. Take one CS paired with a US, with combined learning rate αβ = 0.30 and asymptote λ = 1, starting from no association (V0 = 0). On each trial the strength gains a fixed fraction of the remaining prediction error, ΔV = 0.30 × (1 − V), which has the closed form Vn = 1 − 0.7n.
On trial 1 the prediction error is the full λ − V0 = 1, so ΔV = 0.30 and V rises to 0.30. On trial 2 the error has shrunk to 1 − 0.30 = 0.70, so ΔV = 0.30 × 0.70 = 0.21 and V reaches 0.51. Continuing, V3 = 0.657, V4 = 0.760, V5 = 0.832, and by trial 10 the strength is V10 = 1 − 0.710 = 0.972 — almost the whole way to the ceiling. The curve reaches half its asymptote when 0.7n = 0.5, that is at n = ln 0.5 / ln 0.7 = 1.94 trials, so most of the learning is done in the first two trials and each later trial adds less. This is the negatively accelerated acquisition curve that the associative-learning literature reports, produced by nothing more than the assumption that learning is proportional to surprise. Figure 1 plots the curve.
Rescorla-Wagner acquisition of a single conditioned association with learning rate αβ = 0.30 and asymptote λ = 1: associative strength climbs toward the ceiling as the prediction error shrinks.
Discussion
Few ideas in psychology have shown the durability of association. It began as the empiricists' single principle of mental cohesion (Hume, 1739), was reinterpreted by James as a property of neural conduction (James, 1890), and was made an experimental science by Pavlov and Thorndike (Pavlov, 1927); (Thorndike, 1911). Its modern form is quantitative and mechanistic: the Rescorla-Wagner model turned the bare idea that associations form by contiguity into a precise error-correcting rule that predicts phenomena — blocking above all — that contiguity alone cannot (Rescorla & Wagner, 1972), and the spreading-activation account carried the principle into the study of memory and meaning (Collins & Loftus, 1975).
The tradition's central tension is whether association is the whole story. The prediction-error and attentional models already complicate the simple picture, making learning depend on surprise and on the informativeness of cues rather than on mere pairing (Rescorla, 1988); (Le Pelley et al., 2016). A stronger challenge holds that much of what looks like automatic association in humans is better described as propositional reasoning about the relations between events — the learner forming and testing beliefs, not just accumulating bonds (De Houwer et al., 2013). Whether human learning is fundamentally associative or inferential remains genuinely open, and the evidence is mixed enough that the field has not settled it (Shanks, 2010). What is not in doubt is that the associative link — weighted, error-driven, and spreading — remains one of the most productive units of explanation the science of mind possesses.
Current Directions
The most active bridge is to neuroscience. The Rescorla-Wagner error term found a physical realisation in the reward prediction error signalled by midbrain dopamine neurons (Schultz, Dayan, & Montague, 1997), and the contemporary debate concerns exactly how literally that identification should be taken and what the dopamine signal is really computing, tying associative learning theory to reinforcement learning and to the neural code for value (Gershman & Uchida, 2019). In parallel, the behavioural and neurobiological mechanisms of extinction — the context-dependence, the recovery phenomena, the circuits that suppress a learned association without deleting it — have been synthesised across Pavlovian and instrumental learning, giving the clinical work on exposure therapy a mechanistic foundation (Bouton et al., 2021).
A second front is computational and concerns meaning. Vector-space (distributional) models now learn rich semantic representations from the statistics of language, and the open question is how far these co-occurrence-based systems capture human semantic memory and where they depart from it — whether meaning is better modelled as explicit associative links or as position in a learned space (Guenther et al., 2019); (Kumar, 2021). Running through both fronts is the older, unresolved question of whether the associative account can be maintained against a propositional one, now pursued with the tools of both cognitive experiment and computational modelling (De Houwer et al., 2013); (Shanks, 2010).
Common Misconceptions
- Associations form automatically whenever two things occur together.
- Mere temporal pairing is neither necessary nor sufficient. The Rescorla-Wagner model and the blocking effect show that an association grows only to the extent that the outcome is surprising; a cue whose outcome is already predicted is barely learned (Rescorla & Wagner, 1972); (Rescorla, 1988).
- Extinction erases a learned association.
- Presenting a cue without its outcome suppresses the response but leaves the original association intact: it returns with a change of context, the passage of time, or a reminder. Extinction is new inhibitory learning layered over the old link, not its deletion (Bouton, 2004); (Bouton et al., 2021).
- Association is a discredited relic of behaviorism.
- The associative principle is central to modern cognitive and computational psychology: spreading activation models semantic memory and priming, and prediction-error learning underpins contemporary accounts of reward and of distributional meaning (Collins & Loftus, 1975); (Gershman & Uchida, 2019).
Glossary
- Association of ideas.
- The classical, introspective principle that one idea tends to summon another with which it has been connected in experience; the founding notion of the associationist tradition.
- Associative learning.
- The acquisition of a link between two stimuli, or between a response and its outcome, inferred from a measured change in behaviour when the events are paired.
- Associative strength (V).
- The learned magnitude of the link between a cue and an outcome; the quantity the Rescorla-Wagner model updates on each trial toward an asymptote.
- Blocking.
- The finding that a cue already predicting an outcome prevents a second, redundant cue paired with it from being learned; a key demonstration that association tracks prediction error, not pairing.
- Conditioned stimulus (CS).
- An initially neutral stimulus that, through pairing with a biologically significant one, comes to evoke a learned response.
- Contiguity.
- Closeness in time or place between two events; the primary classical law of association, holding that items experienced together become linked.
- Extinction.
- The reduction of a conditioned response when the cue is presented without its outcome; new inhibitory learning that suppresses the original association rather than erasing it.
- Law of effect.
- Thorndike's principle that responses followed by a satisfying consequence become more strongly associated with the situation in which they occur; the basis of instrumental association.
- Prediction error.
- The discrepancy between the outcome that occurs and the outcome the present cues already predict; the driving term of the Rescorla-Wagner model, so learning is proportional to surprise.
- Rescorla-Wagner model.
- The theory that associative strength changes in proportion to the prediction error, with all cues present sharing the available error; it explains the acquisition curve and blocking.
- Resemblance (similarity).
- A classical law of association by which ideas that are alike tend to evoke one another; one of Hume's three connecting principles.
- Reward prediction error.
- A prediction error about reward, signalled by midbrain dopamine neurons; the candidate neural realisation of the Rescorla-Wagner error term.
- Semantic priming.
- The speeding of recognition of a word by a preceding associated word; explained by spreading activation between linked concept nodes.
- Spreading activation.
- The process by which activating one concept node in semantic memory sends activation to associatively linked nodes, decaying with distance; the associationist account of memory retrieval.
- Unconditioned stimulus (US).
- A stimulus that evokes a response without prior learning; the biologically significant event whose association a conditioned stimulus comes to signal.
Key Researchers
Hermann Ebbinghaus (1850-1909). Founded the experimental study of association and memory, using nonsense syllables to measure how associative bonds are formed by repetition and lost through forgetting. Wikipedia - Wikidata
Samuel J. Gershman. Links associative learning theory to reinforcement learning and to the dopaminergic reward-prediction-error signal, giving the Rescorla-Wagner error term a neural interpretation. ORCID - Google Scholar - Faculty page
Jan De Houwer. Leads the functional and propositional analysis of learning, arguing that much human associative learning reflects reasoning about relations among events rather than the automatic formation of bonds. ORCID - Google Scholar
David Hume (1711-1776). Gave the association of ideas its canonical formulation, naming resemblance, contiguity, and cause and effect as the three principles that connect ideas in the mind. Wikipedia
Ivan Pavlov (1849-1936). Founded the experimental study of classical conditioning, demonstrating that a neutral stimulus paired with a biologically significant one comes to evoke a learned, conditioned response. Wikipedia - Wikidata
Robert A. Rescorla (1940-2020). Co-authored the Rescorla-Wagner prediction-error model and reframed conditioning as the learning of predictive relations among events rather than the stamping-in of a reflex by contiguity. Wikipedia - Wikidata
David R. Shanks. Studies human associative and contingency learning and the boundary between associative and inferential accounts of how people learn about relations among events. ORCID - Google Scholar - Faculty page
Edward L. Thorndike (1874-1949). Formulated the law of effect from the puzzle-box experiments, establishing instrumental association as the complement to Pavlovian conditioning. Wikipedia
Frequently Asked Questions
What is association in psychology? Association is the linking of two mental contents (ideas, stimuli, responses, or events) so that the later occurrence of one tends to evoke the other. It names at once a relation, the process that forms it, and the principle that the mind is organised by such links, and it underlies the association of ideas, associative learning, and the structure of semantic memory (Hume, 1739).
What are the laws of association? The classical laws are contiguity (items experienced close in time or place become linked), similarity or resemblance (like ideas evoke one another), and contrast. Hume's canonical version named resemblance, contiguity, and cause and effect as the three connecting principles, with contiguity usually treated as primary (Hume, 1739).
How is associative learning different from the association of ideas? The association of ideas is the introspective, philosophical sense: one thought summoning another. Associative learning is the experimental sense, a measurable link between a stimulus and an outcome, or an action and its consequence, studied through conditioning by Pavlov and Thorndike (Pavlov, 1927); (Thorndike, 1911).
What does the Rescorla-Wagner model say? It says associative strength changes in proportion to the prediction error, the gap between the outcome that occurs and the outcome the cues present already predict. Learning is fast when the outcome is surprising and slows as it becomes predicted, and all cues present share the available error, which explains blocking (Rescorla & Wagner, 1972).
What is blocking, and why does it matter? Blocking is the finding that a cue already predicting an outcome prevents a second, redundant cue paired with it from being learned. It matters because it shows that mere pairing does not build an association; what matters is whether the cue adds predictive information, which is exactly what a prediction-error model captures (Rescorla & Wagner, 1972); (Rescorla, 1988).
Does extinction erase a learned association? No. Presenting the cue without the outcome reduces the response, but the association returns with a change of context, the passage of time, or a reminder of the outcome. Extinction is new, context-dependent inhibitory learning laid over the original link rather than its deletion (Bouton, 2004); (Bouton et al., 2021).
How does association explain memory? In the spreading-activation account, semantic memory is a network of concept nodes joined by weighted associative links. Retrieving a concept activates its node, and activation spreads to neighbours, decaying with distance, which explains semantic priming and why related ideas come to mind together (Collins & Loftus, 1975).
Is human learning really associative, or is it reasoning? This is genuinely unsettled. Prediction-error and attentional models already move beyond simple pairing, and one influential view holds that much human learning is propositional (forming and testing beliefs about relations among events) rather than the automatic accretion of bonds. The evidence supports parts of both accounts (De Houwer et al., 2013); (Shanks, 2010).
References
Bouton, M. E. (2004). Context and behavioral processes in extinction. Learning & Memory, 11(5), 485-494. https://doi.org/10.1101/lm.78804
Bouton, M. E., Maren, S., & McNally, G. P. (2021). Behavioral and neurobiological mechanisms of Pavlovian and instrumental extinction learning. Physiological Reviews, 101(2), 611-681. https://doi.org/10.1152/physrev.00016.2020
Collins, A. M., & Loftus, E. F. (1975). A spreading-activation theory of semantic processing. Psychological Review, 82(6), 407-428. https://doi.org/10.1037/0033-295X.82.6.407
De Houwer, J., Barnes-Holmes, D., & Moors, A. (2013). What is learning? On the nature and merits of a functional definition of learning. Psychonomic Bulletin & Review, 20(4), 631-642. https://doi.org/10.3758/s13423-013-0386-3
Domjan, M. (2005). Pavlovian conditioning: A functional perspective. Annual Review of Psychology, 56, 179-206. https://doi.org/10.1146/annurev.psych.55.090902.141409
Ebbinghaus, H. (1913). Memory: A contribution to experimental psychology (H. A. Ruger & C. E. Bussenius, Trans.). Teachers College, Columbia University. (Original work published 1885)
Gershman, S. J., & Uchida, N. (2019). Believing in dopamine. Nature Reviews Neuroscience, 20(11), 703-714. https://doi.org/10.1038/s41583-019-0220-7
Guenther, F., Rinaldi, L., & Marelli, M. (2019). Vector-space models of semantic representation from a cognitive perspective: A discussion of common misconceptions. Perspectives on Psychological Science, 14(6), 1006-1033. https://doi.org/10.1177/1745691619861372
Hume, D. (1739). A treatise of human nature. John Noon.
James, W. (1890). The principles of psychology (Vol. 1). Henry Holt and Company.
Kumar, A. A. (2021). Semantic memory: A review of methods, models, and current challenges. Psychonomic Bulletin & Review, 28(1), 40-80. https://doi.org/10.3758/s13423-020-01792-x
Le Pelley, M. E., Mitchell, C. J., Beesley, T., George, D. N., & Wills, A. J. (2016). Attention and associative learning in humans: An integrative review. Psychological Bulletin, 142(10), 1111-1140. https://doi.org/10.1037/bul0000064
Mackintosh, N. J. (1975). A theory of attention: Variations in the associability of stimuli with reinforcement. Psychological Review, 82(4), 276-298. https://doi.org/10.1037/h0076778
Pavlov, I. P. (1927). Conditioned reflexes: An investigation of the physiological activity of the cerebral cortex (G. V. Anrep, Trans.). Oxford University Press.
Pearce, J. M., & Hall, G. (1980). A model for Pavlovian learning: Variations in the effectiveness of conditioned but not of unconditioned stimuli. Psychological Review, 87(6), 532-552. https://doi.org/10.1037/0033-295X.87.6.532
Rescorla, R. A. (1988). Pavlovian conditioning: It's not what you think it is. American Psychologist, 43(3), 151-160. https://doi.org/10.1037/0003-066X.43.3.151
Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), Classical conditioning II: Current research and theory (pp. 64-99). Appleton-Century-Crofts.
Schultz, W., Dayan, P., & Montague, P. R. (1997). A neural substrate of prediction and reward. Science, 275(5306), 1593-1599. https://doi.org/10.1126/science.275.5306.1593
Shanks, D. R. (2010). Learning: From association to cognition. Annual Review of Psychology, 61, 273-301. https://doi.org/10.1146/annurev.psych.093008.100519
Thorndike, E. L. (1911). Animal intelligence: Experimental studies. Macmillan.