Reversal Learning

Task-family page: reversal learning — the canonical non-verbal assay of behavioural flexibility when reward contingencies change. Companion tasks: intradimensional-extradimensional-shift (set-shifting), two-armed-bandit-task (probabilistic RL), affective-bias-test (value/memory bias). Hubs: nonverbal-cognitive-tasks (cross-condition) and autism-nonverbal-cognitive-tasks (ASD translation). Platforms: automated-cognitive-testing-devices, campden-instruments.

Reversal learning paradigms are among the most widely used tests of cognitive flexibility, and have been used as cross-species assays for altered cognitive processes in a host of neuropsychiatric conditions.1 The logic is simple: a subject first acquires a stimulus–outcome (or response–outcome) contingency, and then — without warning — the contingency reverses; the question is how quickly and cleanly the subject adapts.

What the task measures

The construct validity of reversal learning has been revised in recent years: the older framing of the task as primarily a measure of response inhibition is no longer accepted. The current view is that all reversal paradigms, irrespective of exact design, test a combination of (i) learning from rewards and reward omission, (ii) estimating the likelihood/prior probability that a reversal can occur, and (iii) rule use (a “learning set”) working alongside reward learning — so intact performance requires dynamic reward representations and an expectation of change, not merely motor inhibition.1 Across species the task exists in deterministic and probabilistic forms; human versions are almost always probabilistic, while rodent and monkey designs are often deterministic.1

Variants

VariantDesignTypical implementationKey observations
Deterministic discrimination reversal100%/0% feedback; stage advances to criterionTouchscreen pairwise visual discrimination (mouse); spatial/visual mazes (rat); lever/nosepoke tasksStandard in rodents and monkeys; models genetic disorders with stage-by-stage criteria (e.g., >80% accuracy) 2
Probabilistic reversal learning (PRL)Correct choice reinforced stochastically (e.g., 80:20); reversal signalled by the reward pattern, not a ruleHuman screen tasks (almost always probabilistic 1); rat operant PRL 3; mouse operant PRL 4Can’t be performed errorlessly; measures feedback integration, win-stay/lose-shift, and switch dynamics 3
Serial reversalsMultiple successive contingency flips; tracks learning-set formation and cumulative perseverationMouse touchscreen (e.g., 4 serial reversals) 5; rat serial spatial reversals 1Repeated reversal reveals savings and/or progressive perseveration; useful for detecting load-dependent deficits 5
Reversal stages inside set-shifting batteriesSDR / CDR / IDR / EDR stages embedded in the CANTAB-style IED progressionHuman CANTAB, MonkeyCantab, rodent 7-stage digging/touchscreen protocolsPure within-dimension contingency reversals; see intradimensional-extradimensional-shift and parametric-ied-stimulus-design

Outcome measures

  • Trials / errors to criterion — the headline metric in deterministic designs.2
  • Perseverative errors — responses continuing the previously rewarded rule despite the reversal; regressive errors — returning to the old (now incorrect) response after the initial shift to the new correct target. Regressive errors specifically index the ability to maintain the new choice.67
  • Win-stay / lose-shift — trial-level sensitivities to positive vs negative feedback; dissociate reward sensitivity from flexibility proper.3
  • Computational parameters — delta-rule or Q-learning fits yield learning rate, reinforcement sensitivity, and choice stochasticity; hierarchical Gaussian filter models decompose updating of volatile contingencies.73
  • Correction trials — errors can trigger repeats of the trial (same stimulus, same location) until the correct response is made, preventing positional response strategies.2

Human evidence — autism spectrum disorder

  • PRL reveals a reversal-specific deficit. Autistic adults (41 vs 37 matched controls) performed normally during acquisition but at reversal made more regressive errors — reverting to the previously preferred response — and regressive-error counts correlated with independently rated restricted and repetitive behaviours, implicating frontostriatal flexibility circuitry.8
  • PRL is outcome-measure ready. In a within-subjects trial of a group intervention, the PRL task was highly feasible, showed test–retest reproducibility and detected change — a rare example of a non-verbal flexibility task validated as a treatment endpoint.6
  • But flexible updating is not uniformly impaired. In a developmentally adapted volatile-reward task (stable 75:25 vs volatile 80:20↔20:80 conditions), autistic children increased their learning rate in the volatile environment just like typically developing children and adults — no group difference, no interaction — constraining predictive-coding accounts of autism.9
  • Development matters. In typical development, children make more overall and regressive errors but fewer perseverative errors than adolescents, and everyday rigidity (parent ratings) relates to less explorative choice during PRL — a framing that applies directly to paediatric studies.7
  • Psychometrics matter. For use as an individual-difference or clinical measure, PRL indices reach good-to-excellent reliability only with specific modelling choices: behavioural indices via mixed-effects models pooling sessions, computational indices via hierarchical estimation with empirical priors.10

Human evidence — mood and anxiety (brief)

Depression has long been linked to abnormal reactions to positive and negative feedback, and reversal tasks are a natural probe:3

  • Youths with bipolar disorder showed impaired probabilistic reversal learning; MDD showed a trend (p = 0.07); anxiety and severe mood dysregulation did not — i.e., deficits were not uniform across paediatric mood diagnoses.11
  • In unmedicated major depression, a task separating reward- from punishment-based reversals found impaired reward (but not punishment) reversal together with attenuated anteroventral striatal response to unexpected reward — linking reversal deficits to anhedonia circuitry.12

Neural basis and neurochemistry

  • Cortical. Human neuroimaging shows increased OFC and medial PFC activity during reversals; in rodents, excitotoxic lesions and glutamatergic manipulations of OFC impair visual reversal learning.1 The OFC picture is still evolving: fiber-sparing excitotoxic lesions limited to macaque OFC areas 11/13/14 left reversal intact, contrasting with older aspiration/marmoset lesion results — evidence for finer-grained, subregion-specific roles.1
  • Striatal. Ventral (anteroventral) striatal responses to unexpected reward are attenuated in depression during reward-based reversal.12 Cortical–striatal–amygdala circuits and dopamine, serotonin and glutamate systems are all implicated across the species literature.1
  • Serotonin (5-HT). In the rat PRL task, 5-HT manipulations shift feedback sensitivity in a dose- and time-dependent way: a single low-dose citalopram increased lose-shift behaviour and reduced reversals completed; higher doses and repeated citalopram had the opposite effect (more reversals, more win-stay), and global 5-HT depletion mirrored the repeated-treatment profile in reverse. The interpretation: elevated 5-HT → less negative-feedback sensitivity, more positive-feedback sensitivity — paralleling human 5-HT findings and hypotheses about depression.3 Note that systemic depletion also affects discrimination/reward learning more broadly, so specificity depends on the manipulation.1

Rodent protocols and model findings

The standard touchscreen protocol (Bussey-Saksida chambers; see automated-cognitive-testing-devices, campden-instruments) runs habituation plus five pre-training stages (initial touch → must touch → must initiate → punish incorrect), then a pairwise visual discrimination (>80% criterion), then the reversal, with correction trials for errors and a strawberry-milk reward.2 Strain matters: BALB/c mice, classically poor learners on this platform, complete discrimination and reversal in fewer sessions than C57BL/6 under an adapted protocol — and dissociations between “sessions to criterion” and accuracy illustrate why both learning speed and reversal-specific errors should be reported.2

Model findings form a translational ladder:

  • Fmr1 knockout mice (fragile X syndrome) — the expected executive deficit appears only under high cognitive load: more learning-type errors during a second reversal when salient distractors conflict with the task set, plus more reward-collection attempts during timeouts after first-reversal errors.5 A caution against null results under low-load conditions.
  • BTBR mice (idiopathic autism model) — impaired operant probabilistic reversal learning in both sexes, extending the model’s behavioural inflexibility to the feedback-integration version of the task.4
  • Shank3 rats (Phelan–McDermid syndrome) — the most prominent adult cognitive phenotype was touchscreen visual discrimination and reversal impairment, with rapid, error-prone responding, against largely normal social behaviour.13
  • OFC lesion literature — see above; the classical rodent finding is lesion- and dimension-specific reversal impairment.1

Non-human primates

  • Deterministic or probabilistic reversals are standard in macaque testing; the classic lesion dissociations (OFC vs lateral PFC contributions) come from this literature, with the newer subregion-restricted lesion results refining the OFC story.1 MonkeyCantab embeds four reversal stages (SDR, CDR, IDR, EDR) inside the IED sequence — see intradimensional-extradimensional-shift and lafayette-instrument.
  • VPA-exposed marmosets (autism model): adults persisted with previously acquired strategies on reversal learning — specifically for strategies that had taken many days to acquire — and childhood gaze deficits predicted this adult inflexibility.14

Cross-species summary

HumanRodentMonkey
Deterministic reversalRare (research batteries)Standard (touchscreen/maze) 2Standard; OFC lesion literature 1
Probabilistic reversalThe default human version 1Rat PRL (5-HT pharmacology) 3; mouse PRL (BTBR, ASD models) 4Available; less standard
Serial reversalsLearning-set studies 1Mouse (Fmr1) 5; rats 1Reference (Harlow learning set) 1
ASD-relevant findingsRegressive errors ↔ repetitive behaviours 8; outcome-feasibility 6; intact volatile updating 9Reversal deficits in Shank3 rat 13; BTBR 4; load-dependent Fmr1 5VPA marmoset perseveration 14
Clinical outcome usePRL feasible/reproducible/sensitive to change 6; reliability dependent on modelling 10

Relation to other task pages

  • intradimensional-extradimensional-shift — reversal learning reverses the contingency within a given dimension; the extradimensional shift requires abandoning the dimension altogether (attentional set-shifting). The CANTAB IED sequence interleaves both (four reversal stages among nine), and the classic dissociation is reversal ↔ OFC/5-HT circuits vs ED shift ↔ medial PFC/catecholamine circuits (see the IED page for the rodent pharmacology). In ASD research both variants are used and both show heterogeneous, small-to-moderate effects.18
  • stop-signal-task — action cancellation (withholding an already-prepared response) is a distinct facet of “inhibition” from contingency reversal (switching the correct response after feedback); the revised construct-validity view above — that reversal is not primarily response inhibition1 — is why the two task families should be annotated separately.
  • two-armed-bandit-task — PRL is effectively a two-choice bandit that occasionally reverses; both are modelled with delta-rule/Q-learning machinery (learning rate, learning noise), but bandit tasks foreground explore/exploit trade-offs while PRL foregrounds feedback sensitivity and switch/stay dynamics.73 The bandit page covers the computational microstructure in depth.
  • affective-bias-test / cognitive-judgement-bias — “hot” affective-bias assays of how mood distorts value/interpretation; reversal learning is closer to “cold” flexibility, but the mood-disorder findings above (reward-vs-punishment reversal asymmetry) are where the hot/cold distinction blurs.123
  • delayed-match-to-sample / spatial-working-memory / delayed-response-tasks-working-memory — memory-load tasks; orthogonal to contingency change, though both are non-verbal touchscreen staples of the same cross-species battery families.

Open questions

  • Feedback sensitivity vs flexibility: reversal “deficits” can reflect altered reward/punishment sensitivity (e.g., the 5-HT and MDD reward-reversal asymmetries) rather than inflexibility per se — computational decomposition is needed to separate them.123
  • Reliability and modelling: individual-level use requires hierarchical, session-pooled estimation; naive summary statistics can undercut reliability.10
  • Developmental trajectories: child/adolescent differences in error types indicate reversal performance is not a fixed trait across development.7
  • Load- and salience-dependence: deficits may be latent until cognitive load or sensory conflict is raised — null findings require load manipulations.5
  • Cross-species design harmonisation: deterministic (animal-default) and probabilistic (human-default) versions are not interchangeable; deterministic PRL-style tasks in humans remain comparatively rare.1

intradimensional-extradimensional-shift | two-armed-bandit-task | stop-signal-task | delayed-match-to-sample | spatial-working-memory | affective-bias-test | cognitive-judgement-bias | nonverbal-cognitive-tasks | autism-nonverbal-cognitive-tasks | cantab | automated-cognitive-testing-devices | preclinical-drug-screening

References

Footnotes

  1. raw/papers/izquierdo-2017-reversal-learning-neural-basis.md 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19

  2. raw/papers/turner-2017-mouse-visual-reversal-protocol.md 2 3 4 5 6

  3. raw/papers/bari-2010-serotonin-rat-prl.md 2 3 4 5 6 7 8 9 10

  4. raw/papers/alvarez-2023-btbr-prl.md 2 3 4

  5. raw/papers/dickson-2013-fmr1-serial-reversal.md 2 3 4 5 6

  6. raw/papers/schmitt-2021-prl-outcome-measure.md 2 3 4

  7. raw/papers/weiss-2020-prl-development.md 2 3 4 5

  8. raw/papers/dcruz-2013-prl-autism.md 2 3

  9. raw/papers/manning-2017-volatile-learning-asd.md 2

  10. raw/papers/waltmann-2022-prl-reliability.md 2 3

  11. raw/papers/dickstein-2010-prl-youth-mood.md

  12. raw/papers/robinson-2012-mdd-reward-reversal.md 2 3 4

  13. raw/papers/pearson-2026-shank3-rat-touchscreen.md 2

  14. raw/papers/nakagami-2022-marmoset-vpa-attention.md 2