Abstract

Reversal learning is a type of learning: a procedure in which the reward contingencies of two discriminated options are swapped, so the correct stimulus becomes incorrect and vice versa. Because perception is unchanged and only the outcome mapping is reversed, performance after the switch isolates an organism's capacity to update learned value and abandon a previously reinforced response — a behavioral index of cognitive flexibility. Its signature measure is the perseverative error: continued choice of the formerly rewarded option after the reversal. Lesion, imaging, and computational work converge on the orbitofrontal cortex, the amygdala, and their dopaminergic and serotonergic modulation as the circuitry supporting flexible reversal. This article defines the task, distinguishes it from attentional set-shifting, traces its neural basis, and explains how reinforcement-learning models account for perseveration, with three interactive demonstrations.

Keywords: reversal learning, cognitive flexibility, perseveration, orbitofrontal cortex, reinforcement learning

Reversal learning is one of the oldest and most widely used probes of behavioral flexibility in comparative and clinical psychology. An organism first learns a simple discrimination — choose stimulus A rather than stimulus B, because A is reinforced and B is not — and, once that discrimination is stable, the experimenter reverses the contingency without warning: now B is reinforced and A is not. Nothing about the stimuli, the responses, or the reward itself has changed; only the association between a specific choice and its outcome has been inverted. Measuring how quickly behavior tracks that inversion, and how many perseverative responses to the old choice occur before it does, gives a clean behavioral read-out of how an organism updates learned value and suppresses a previously adaptive response (Izquierdo et al., 2017; Clark, Cools & Robbins, 2004).

Key Takeaways
  • Reversal learning swaps the reward contingencies of an already-learned discrimination, so post-switch performance isolates the updating of learned value rather than new perceptual learning.
  • The diagnostic measure is the perseverative error: continued choice of the formerly rewarded option after the reversal.
  • It differs from attentional set-shifting: reversal changes which stimulus within a dimension is rewarded, whereas set-shifting changes which stimulus dimension is relevant.
  • The orbitofrontal cortex, amygdala, and ventral striatum, modulated by dopamine and serotonin, form the core circuitry; orbitofrontal damage produces marked perseveration.
  • Reinforcement-learning models capture perseveration through the learning rate and the choice of what is updated, and reversal tasks are now standard assays of flexibility in psychiatric and computational research.

What Reversal Learning Is

A reversal-learning task has two phases built on the same pair of options. In acquisition, the organism learns a discrimination: one stimulus (or spatial location, or action) is consistently reinforced and the other is not, and choice comes to favor the rewarded option until it reaches a performance criterion. In reversal, the contingencies are exchanged — the previously unrewarded option now delivers reward, and the previously rewarded option no longer does — and the organism must detect the change and shift its choices accordingly (Izquierdo et al., 2017). The perceptual discrimination itself is never in question; the two stimuli remain equally discriminable throughout. What changes is only the mapping from choice to outcome, which is precisely why the task isolates the updating of learned reward value from every other component of learning.

The behavioral signature of reversal is the pattern of errors immediately after the switch. A perseverative error is a choice of the formerly correct option once it has stopped being reinforced; a run of such errors reflects the pull of the previously learned value before it has been overwritten. Once the organism stops perseverating, a briefer phase of regressive errors — occasional lapses back to the old choice — may persist while the new contingency is consolidated. Counting errors to criterion, and separating the perseverative from the regressive component, turns reversal learning into a quantitative assay of how readily a learned response can be abandoned when it no longer pays (Clark, Cools & Robbins, 2004; Jones & Mishkin, 1972).

Because the same procedure works with visual stimuli, spatial locations, or arbitrary actions, and with species from fish to primates, reversal learning has long served as a common currency for comparing flexibility across tasks and organisms. It is closely tied to the broader construct of cognitive flexibility, and it is distinct from — though often studied alongside — attentional set-shifting, a distinction developed in the next section (Kehagia, Murray & Robbins, 2010).

Reversals are frequently administered in a series: once behavior tracks one reversal, the contingencies are reversed again, and again. Across such a series many animals develop a reversal-learning set — errors to criterion fall with successive reversals as the organism comes to anticipate that contingencies change — and the rate of that improvement was an early comparative tool for ranking learning across the phylogenetic scale (Bitterman, 1965). Serial reversal later became the standard rodent assay on which the orbitofrontal specialization for reversal was established (Boulougouris, Dalley & Robbins, 2007).

Serial reversal: forming a reversal set

When reversals are repeated, many animals get progressively better at them: errors to criterion fall from one reversal to the next as the organism learns that the contingency will keep flipping. The steepness of that decline — the rate at which a reversal set forms — was an early yardstick for comparing flexibility across species. Set how fast the set forms and watch errors per reversal collapse toward an asymptote.

240errors to criterion241172123947556574849310reversal number
By the tenth reversal errors have fallen from 24 to 3. Over the whole series the learner makes 90 errors instead of the 240 a no-set learner would — a saving of 150 errors from forming the reversal set. Faster set formation bends the curve down sooner; a flat curve means each reversal is met as if it were the first.

Reversal Versus Set-Shifting

Reversal learning is frequently confused with attentional set-shifting, but the two probe different levels of flexibility, and a landmark dissociation established that they depend on different regions of prefrontal cortex. In a reversal, the relevant dimension stays the same — if color was the rewarded dimension, it remains so — and only the specific value within that dimension that is rewarded is swapped. In an extradimensional shift, by contrast, the rewarded value stays within the old dimension no longer; the organism must stop attending to color altogether and start attending to a previously irrelevant dimension such as shape (Dias, Robbins & Roberts, 1996).

Dias, Robbins, and Roberts showed that these two forms of flexibility are neurally separable in the marmoset: lesions of the orbitofrontal cortex selectively impaired reversal learning — an affective shift in which the value attached to a specific stimulus must be updated — while sparing extradimensional set-shifting, whereas lateral prefrontal lesions produced the opposite pattern, impairing the attentional shift while sparing reversal (Dias, Robbins & Roberts, 1996). This double dissociation is one of the foundational results of the field: it demonstrates that suppressing a specific learned stimulus-reward association and reallocating attention across dimensions are distinct operations with distinct neural substrates, and it grounds the use of reversal learning as a targeted probe of orbitofrontal, rather than general prefrontal, function (Clark, Cools & Robbins, 2004).

Discrimination reversal: counting perseverative errors

The learner arrives at the reversal already preferring option A (VA = 0.8, VB = 0.2). Now A pays nothing and B pays reward, and the learner always picks the higher-valued option. Each choice updates that option’s value by the learning rate α. Watch how many times it keeps choosing the old option A — the perseverative errors — before B’s value overtakes A’s.

1.00.50.0trial after reversalV(A) oldV(B) new
At α = 0.40 the learner commits 3 perseverative errors before switching to the newly rewarded option (red markers = chose old A, green = chose new B). A higher learning rate overwrites the old value faster and shortens perseveration; a lower one prolongs it, even though every non-reward is registered.

The distinction matters clinically because different disorders load on different components. Tasks that combine both, such as the intra-/extradimensional set-shift battery, can localize a flexibility deficit to the reversal stage or the shifting stage, and this resolution is what makes reversal learning a more specific measure than a global label of perseveration or rigidity would allow (Kehagia, Murray & Robbins, 2010).

Table 1

Reversal Learning Contrasted With Extradimensional Set-Shifting

Feature Reversal learning Extradimensional set-shifting
What changes Which value within a dimension is rewarded Which stimulus dimension is relevant
Locus of attention Stays on the same dimension Moves to a previously irrelevant dimension
Type of shift Affective (stimulus-reward value) Attentional
Critical region Orbitofrontal cortex Lateral prefrontal cortex
Marmoset lesion effect Impaired by orbitofrontal lesion; shifting spared Impaired by lateral prefrontal lesion; reversal spared

Note. The double dissociation of the two shifts by region was established by lesion in the marmoset (Dias, Robbins & Roberts, 1996).

The Neural Basis of Reversal

The circuitry of reversal learning centers on the orbitofrontal cortex, working with the amygdala and the ventral striatum. Early primate work implicated the orbital and inferior frontal cortex and the limbic structures beneath it: lesions there disrupted the ability to relearn stimulus-reinforcement associations when they were reversed, without abolishing the original discrimination (Jones & Mishkin, 1972). Rodent studies later refined the anatomy, showing that orbitofrontal lesions impair serial spatial reversal while lesions of adjacent medial prefrontal regions such as infralimbic and prelimbic cortex have distinct or lesser effects, establishing an orbitofrontal specialization for reversal rather than a diffuse prefrontal one (Boulougouris, Dalley & Robbins, 2007).

The functional interpretation of the orbitofrontal contribution has shifted over time. The classical account cast the region as an inhibitory controller that suppresses the prepotent, previously rewarded response. A more recent and better-supported view holds that the orbitofrontal cortex represents expected outcomes and the current task state — what reward each choice currently predicts — so that a reversal deficit reflects a failure to update outcome expectancies rather than a failure of response inhibition per se (Schoenbaum et al., 2009; Rudebeck & Murray, 2014). On this account the orbitofrontal cortex is less a brake than a predictive map that must be revised when outcomes change, and reversal learning is the behavioral test of whether that revision has occurred.

Figure 1

Perseverative Errors Across a Contingency Reversal

Accuracy rises during acquisition, drops at the reversal point, then recovers A line rises from chance to high accuracy during acquisition. A dashed vertical line marks the reversal. Accuracy falls below chance immediately after the reversal, reflecting perseverative choice of the old option, then climbs back to high accuracy. 100% 50% 0% acquisition reversal reversal point perseverative errors
Note. During acquisition accuracy climbs above chance (the dashed horizontal line). At the reversal (gold line) the previously correct choice becomes incorrect, so accuracy falls sharply as the organism perseverates on the old option (red), then recovers as the new contingency is learned (green). Original schematic.

Neuromodulation tunes this circuit. Serotonin depletion in the orbitofrontal cortex selectively increases perseverative responding, dissociating the chemical control of reversal from that of set-shifting, and dopamine sets the rate at which value predictions are updated from reward and its omission (Boulougouris, Dalley & Robbins, 2007; Cools et al., 2002). Human functional imaging localizes the same operations: event-related fMRI during probabilistic reversal reveals ventral prefrontal and striatal activity time-locked to the reversal itself, and dopaminergic manipulations bidirectionally modulate the flexibility of switching (Cools et al., 2002; Clark, Cools & Robbins, 2004).

Computational Accounts

Reinforcement-learning models give reversal learning a precise mechanistic vocabulary. In the simplest delta-rule model, each option carries an expected value that is nudged toward the observed reward after every choice, by an amount set by the learning rate. Perseveration after a reversal is then a direct consequence of how strongly the old value was learned and how quickly it decays: a low learning rate makes value estimates stable and hard to overturn, producing many perseverative errors, while a high learning rate tracks the switch quickly but at the cost of being buffeted by every chance non-reward (Costa et al., 2015).

Probabilistic reversal: telling a switch from noise

Now the correct option is rewarded only most of the time, so a stray non-reward no longer means the world has changed. Option A is correct for the first 30 trials; then the contingency reverses to B. The curve is the fraction of a cohort of 240 value-tracking learners (α = 0.35) choosing correctly on each trial. Set the reward probability p and watch how cleanly the group tracks the switch.

100%50%0%reversaltrial
At p = 0.80 the cohort averages 7.0 errors and settles at about 99% correct, climbing back above 75% roughly 8 trials after the switch. As p approaches chance the reward signal is buried in noise, so the group tracks the reversal slowly and plateaus far below ceiling; a cleaner contingency is overturned fast and completely.

This framing reframes reversal as a problem of statistical inference under uncertainty. A Bayesian account treats the reversal as a hidden change in the state of the world that the organism must infer from a run of unexpected outcomes, and this perspective explains why reversal is faster when the environment is known to be volatile: an agent that expects change interprets a few surprising outcomes as evidence of a switch rather than as noise (Costa et al., 2015). Single-neuron recordings support the inferential view: prefrontal population activity predicts the animal's internal state switch around a reversal before behavior fully changes, consistent with the orbitofrontal cortex signaling the current task state rather than merely inhibiting a response (Bartolo & Averbeck, 2020). Circuit-level manipulation reinforces the point: dissociable orbitofrontal contributions govern distinct reinforcement-learning processes such as learning from reward versus its omission (Groman et al., 2019).

Worked Example

Consider a delta-rule learner choosing between options A and B. Each option has an expected value V updated after every choice by V ← V + α(r − V), where r is the reward received (1 or 0) and α is the learning rate. Suppose acquisition has left the learner confident in A: VA = 0.8 and VB = 0.2. The contingencies now reverse — A yields r = 0, B yields r = 1 — and the learner chooses greedily, always taking the option with the higher current value.

With a learning rate of α = 0.4, each choice of A (now unrewarded) multiplies its value by (1 − α) = 0.6, since VA ← VA + 0.4(0 − VA) = 0.6 VA. So VA falls 0.8 → 0.48 → 0.288 → 0.173 across three trials, and only on the third update does it drop below VB = 0.2. The learner therefore commits three perseverative errors — three continued choices of the old option A — before switching to B. Analytically, perseveration lasts while 0.8 × 0.6n > 0.2, i.e. 0.6n > 0.25, which fails first at n = ⌈log 0.25 / log 0.6⌉ = ⌈2.71⌉ = 3. On the fourth trial the learner finally chooses B and is rewarded, so VB ← 0.2 + 0.4(1 − 0.2) = 0.52, cementing the new preference.

The learning rate controls the whole picture. Raising α to 0.7 shrinks perseveration to two errors (0.8 → 0.24 → 0.072, dropping below 0.2 on the second trial), because value is overwritten faster; lowering it would extend perseveration further. This is exactly why reversal tasks are used to estimate an individual's learning rate and its asymmetry between reward and non-reward: the number of perseverative errors is a behavioral window onto how the underlying value updates are parameterized (Costa et al., 2015; Groman et al., 2019).

Discussion

Reversal learning has earned its long tenure as a workhorse task because it isolates a single, well-defined operation — overturning a learned stimulus-reward association — while holding perception, motor demands, and reward magnitude constant. That specificity is what allowed the field to move from a global notion of behavioral rigidity to a componential one, separating the affective updating supported by the orbitofrontal cortex from the attentional shifting supported by lateral prefrontal cortex (Dias, Robbins & Roberts, 1996; Clark, Cools & Robbins, 2004). The convergence of lesion, pharmacological, imaging, and single-unit evidence on the orbitofrontal-amygdala-striatal circuit, and on its dopaminergic and serotonergic modulation, makes reversal one of the better-understood assays of flexible behavior at the level of mechanism (Izquierdo et al., 2017; Schoenbaum et al., 2009).

The interpretive shift from inhibition to state-representation has broadened the task's reach. If the orbitofrontal cortex encodes what each choice currently predicts, then reversal deficits index a failure to update an internal model, which connects the paradigm to computational psychiatry: abnormally high or low learning rates, or a distorted balance between learning from reward and from its absence, characterize conditions marked by inflexible or unstable choice, and reversal tasks provide a compact way to estimate those parameters (Costa et al., 2015; Bartolo & Averbeck, 2020). The task thus sits at the intersection of comparative psychology, systems neuroscience, and clinical modeling — a simple procedure whose measure, the perseverative error, remains as informative today as when it was first counted.

Current Directions

The most active current work treats reversal learning through the lens of reinforcement-learning computation and circuit dissection. Optogenetic and lesion studies in rodents now separate the orbitofrontal cortex's contribution to distinct learning processes — updating from received reward versus from omitted reward — showing that flexible reversal is not a single function but a set of dissociable value-updating operations distributed across orbitofrontal circuits (Groman et al., 2019). A parallel anatomical refinement distinguishes medial from lateral orbitofrontal cortex, whose roles in serial reversal are dissociable and, in some manipulations, paradoxically opposed, cautioning against treating the orbitofrontal cortex as a uniform structure (Hervig et al., 2020).

A second front concerns the internal representation of task state. Simultaneous recordings show that prefrontal populations predict an animal's covert switch of strategy around a reversal, favoring accounts in which the cortex infers a hidden change of state rather than simply inhibiting the old response (Bartolo & Averbeck, 2020). As probabilistic reversal tasks are increasingly used as biomarkers in computational psychiatry, their psychometric properties have come under scrutiny: recent work establishes that the behavioral and model-derived readouts of a probabilistic reversal task are sufficiently reliable to support individual-differences and clinical research, a prerequisite for using learning-rate estimates as stable traits (Waltmann, Schlagenhauf & Deserno, 2022).

Common Misconceptions

Reversal learning and set-shifting are the same test of flexibility.
They are dissociable. Reversal swaps which value within a dimension is rewarded and depends on the orbitofrontal cortex; extradimensional set-shifting changes which dimension is relevant and depends on lateral prefrontal cortex. The two were doubly dissociated by lesion in the marmoset (Dias, Robbins & Roberts, 1996).
The orbitofrontal cortex reverses behavior by inhibiting the old response.
The better-supported account is that the orbitofrontal cortex represents expected outcomes and the current task state, so a reversal deficit reflects a failure to update outcome expectancies rather than a failure of response inhibition (Schoenbaum et al., 2009; Rudebeck & Murray, 2014).
Perseverative errors simply mean the organism did not notice the change.
Perseveration is a lawful consequence of value updating, not mere inattention. A delta-rule learner with a low learning rate makes many perseverative errors even while registering each non-reward, because the strongly learned value decays slowly (Costa et al., 2015).

Glossary

Acquisition.
The first phase of a reversal task, in which the initial discrimination is learned to a performance criterion before the contingencies are reversed.
Amygdala.
A limbic structure that encodes stimulus-reward associations and, with the orbitofrontal cortex, supports the updating of value during reversal.
Cognitive flexibility.
The capacity to adjust behavior to changing contingencies; reversal learning provides one of its most direct behavioral measures.
Delta rule.
A reinforcement-learning update in which an option's expected value is shifted toward the observed reward by a fraction set by the learning rate.
Discrimination learning.
Learning to respond differently to two stimuli because they carry different consequences; the substrate a reversal later inverts.
Extradimensional shift.
A change in which stimulus dimension is relevant to reward; distinct from reversal and dependent on lateral rather than orbital prefrontal cortex.
Learning rate.
The parameter (α) governing how much each outcome updates an option's expected value; it controls how many perseverative errors a reversal produces.
Orbitofrontal cortex.
The ventral prefrontal region whose damage produces marked reversal deficits; on current accounts it represents expected outcomes and the current task state.
Perseverative error.
A choice of the formerly rewarded option after the contingencies have reversed; the diagnostic measure of a reversal-learning deficit.
Probabilistic reversal.
A reversal task in which reward follows the correct option only most of the time, so the learner must distinguish a genuine switch from ordinary reward noise.
Regressive error.
An occasional lapse back to the old choice after the switch has largely been learned, distinct from the initial run of perseverative errors.
Reinforcement.
The strengthening of a choice by its outcome; reversal learning manipulates which choice is reinforced while holding everything else fixed.
Reversal learning.
A procedure that swaps the reward contingencies of a learned discrimination to measure how readily an organism updates value and abandons a previously reinforced response.
Reversal set.
The improvement in reversal performance across a series of successive reversals, in which errors to criterion fall as the organism learns that contingencies reverse.
Serotonin.
A neuromodulator whose depletion in the orbitofrontal cortex selectively increases perseverative responding during reversal.
Set-shifting.
Shifting attention from one stimulus dimension to another; a form of flexibility dissociable from reversal and dependent on lateral prefrontal cortex.

Key Researchers

Roshan Cools. Donders Institute, Radboud University Nijmegen; she established the dopaminergic control of probabilistic reversal learning in humans, using event-related fMRI to define the ventral prefrontal and striatal mechanisms and showing that dopamine bidirectionally modulates flexible switching. ORCID

Alicia Izquierdo. University of California, Los Angeles; she authored the definitive modern synthesis of the neural basis of reversal learning, integrating rodent, primate, and human evidence on orbitofrontal-amygdala-striatal circuits and their neuromodulation. ORCID

Nicholas J. Mackintosh (1935-2015). University of Cambridge; a foundational figure in the comparative study of discrimination learning and attention, whose analyses of reversal, learning sets, and selective attention frame how reversal performance indexes flexibility across species. Wikipedia

Angela C. Roberts. University of Cambridge; co-author of the founding dissociation of affective and attentional shifts in prefrontal cortex, her marmoset lesion and imaging work localizes reversal learning to the orbitofrontal cortex and its serotonergic control. ORCID

Trevor W. Robbins. University of Cambridge; with Roberts and colleagues he defined the modern neuropsychology of reversal learning, dissociating affective from attentional shifts and mapping serial reversal onto orbitofrontal cortex and its monoaminergic modulation. ORCID - Wikipedia

Geoffrey Schoenbaum. National Institute on Drug Abuse, NIH; he reframed the orbitofrontal cortex's role in reversal, arguing that it signals expected outcomes and the current task state rather than simply inhibiting responses, so that reversal deficits reflect a failure to update outcome expectancies. ORCID

Frequently Asked Questions

What is reversal learning?
Reversal learning is a procedure in which the reward contingencies of a learned discrimination are swapped, so the previously correct option becomes incorrect and vice versa. Because only the choice-outcome mapping changes, performance after the switch measures how readily an organism updates learned value (Izquierdo et al., 2017).

What is a perseverative error?
A perseverative error is a continued choice of the formerly rewarded option after the contingencies have reversed. A run of such errors reflects the pull of the previously learned value before it has been overwritten, and counting them is the standard measure of a reversal deficit (Clark, Cools & Robbins, 2004).

How is reversal learning different from set-shifting?
Reversal changes which value within a dimension is rewarded, while extradimensional set-shifting changes which dimension is relevant. A marmoset lesion study doubly dissociated them: orbitofrontal damage impaired reversal, and lateral prefrontal damage impaired shifting (Dias, Robbins & Roberts, 1996).

Which brain region is most important for reversal learning?
The orbitofrontal cortex, working with the amygdala and ventral striatum. Lesions there impair the relearning of reversed stimulus-reward associations without abolishing the original discrimination (Jones & Mishkin, 1972).

Does the orbitofrontal cortex reverse behavior by inhibiting the old response?
Current evidence favors a different view: the orbitofrontal cortex represents expected outcomes and the current task state, so a reversal deficit reflects a failure to update outcome expectancies rather than a failure of response inhibition (Schoenbaum et al., 2009).

How do reinforcement-learning models explain perseveration?
In a delta-rule model, each option's value is updated toward the observed reward by an amount set by the learning rate. A low learning rate makes value stable and slow to overturn, producing many perseverative errors even when each non-reward is registered (Costa et al., 2015).

What is a probabilistic reversal task?
It is a reversal task in which the correct option is rewarded only most of the time, so the learner must distinguish a genuine contingency switch from ordinary reward noise. Such tasks are widely used in human imaging and computational psychiatry (Cools et al., 2002).

Why is reversal learning used to study psychiatric disorders?
Because it yields compact, model-based estimates of value updating — learning rates and their reward/non-reward asymmetry — whose reliability supports individual-differences research, letting inflexible or unstable choice be quantified in a single task (Waltmann, Schlagenhauf & Deserno, 2022).

References

Bartolo, R., & Averbeck, B. B. (2020). Prefrontal cortex predicts state switches during reversal learning. Neuron, 106(6), 1044-1054. https://doi.org/10.1016/j.neuron.2020.03.024

Bitterman, M. E. (1965). Phyletic differences in learning. American Psychologist, 20(6), 396-410. https://doi.org/10.1037/h0022328

Boulougouris, V., Dalley, J. W., & Robbins, T. W. (2007). Effects of orbitofrontal, infralimbic and prelimbic cortical lesions on serial spatial reversal learning in the rat. Behavioural Brain Research, 179(2), 219-228. https://doi.org/10.1016/j.bbr.2007.02.005

Clark, L., Cools, R., & Robbins, T. W. (2004). The neuropsychology of ventral prefrontal cortex: Decision-making and reversal learning. Brain and Cognition, 55(1), 41-53. https://doi.org/10.1016/s0278-2626(03)00284-7

Cools, R., Clark, L., Owen, A. M., & Robbins, T. W. (2002). Defining the neural mechanisms of probabilistic reversal learning using event-related functional magnetic resonance imaging. Journal of Neuroscience, 22(11), 4563-4567. https://doi.org/10.1523/jneurosci.22-11-04563.2002

Costa, V. D., Tran, V. L., Turchi, J., & Averbeck, B. B. (2015). Reversal learning and dopamine: A Bayesian perspective. Journal of Neuroscience, 35(6), 2407-2416. https://doi.org/10.1523/jneurosci.1989-14.2015

Dias, R., Robbins, T. W., & Roberts, A. C. (1996). Dissociation in prefrontal cortex of affective and attentional shifts. Nature, 380(6569), 69-72. https://doi.org/10.1038/380069a0

Groman, S. M., Keistler, C., Keip, A. J., Hammarlund, E., DiLeone, R. J., Pittenger, C., Lee, D., & Taylor, J. R. (2019). Orbitofrontal circuits control multiple reinforcement-learning processes. Neuron, 103(4), 734-746. https://doi.org/10.1016/j.neuron.2019.05.042

Hervig, M. E., Fiddian, L., Piilgaard, L., Božić, T., Blanco-Pozo, M., Knudsen, C., Olesen, S. F., Als&ioml;ö, J., & Robbins, T. W. (2020). Dissociable and paradoxical roles of rat medial and lateral orbitofrontal cortex in visual serial reversal learning. Cerebral Cortex, 30(3), 1016-1029. https://doi.org/10.1093/cercor/bhz144

Izquierdo, A., Brigman, J. L., Radke, A. K., Rudebeck, P. H., & Holmes, A. (2017). The neural basis of reversal learning: An updated perspective. Neuroscience, 345, 12-26. https://doi.org/10.1016/j.neuroscience.2016.03.021

Jones, B., & Mishkin, M. (1972). Limbic lesions and the problem of stimulus-reinforcement associations. Experimental Neurology, 36(2), 362-377. https://doi.org/10.1016/0014-4886(72)90030-1

Kehagia, A. A., Murray, G. K., & Robbins, T. W. (2010). Learning and cognitive flexibility: Frontostriatal function and monoaminergic modulation. Current Opinion in Neurobiology, 20(2), 199-204. https://doi.org/10.1016/j.conb.2010.01.007

Rudebeck, P. H., & Murray, E. A. (2014). The orbitofrontal oracle: Cortical mechanisms for the prediction and evaluation of specific behavioral outcomes. Neuron, 84(6), 1143-1156. https://doi.org/10.1016/j.neuron.2014.10.049

Schoenbaum, G., Roesch, M. R., Stalnaker, T. A., & Takahashi, Y. K. (2009). A new perspective on the role of the orbitofrontal cortex in adaptive behaviour. Nature Reviews Neuroscience, 10(12), 885-892. https://doi.org/10.1038/nrn2753

Waltmann, M., Schlagenhauf, F., & Deserno, L. (2022). Sufficient reliability of the behavioral and computational readouts of a probabilistic reversal learning task. Behavior Research Methods, 54(6), 2993-3014. https://doi.org/10.3758/s13428-021-01739-7