Towards an Integrative Neuroscience of Metacognition (capture)

Authors: Stephen M. Fleming — Department of Experimental Psychology and Institute of Cognitive Neuroscience, University College London; Max Planck UCL Centre for Computational Psychiatry and Ageing Research; Program for Brain, Mind and Consciousness, Canadian Institute for Advanced Research (CIFAR). Citation: Fleming SM (2026). Towards an integrative neuroscience of metacognition. Nature Reviews Neuroscience. Published online 16 September 2026. doi:10.1038/s41583-026-01081-x. Source: https://www.nature.com/articles/s41583-026-01081-x · Captured 2026-09-29. Capture note: Full text captured from a user-provided copy of the published PDF (subscription article, 18 pp). Page furniture removed; reference list reduced to the annotated key references (editorial notes retained as printed).

Abstract

Metacognition enables organisms to self-evaluate their own cognitive processes. In humans, this capacity underpins adaptive learning, social coordination and the formation of self-beliefs about skill and ability. A core motif that underpins multiple forms of metacognition is the ability to form estimates of performance in a self-directed frame of reference, expressed as propositional confidence about one’s own behaviour and mental processes. Convergent behavioural, computational and neural evidence obtained across rodents, non-human primates and humans indicates that these self-evaluations arise from structured transformations of uncertainty. Dynamic evidence accumulation processes provide candidate building blocks for local confidence, whereas multimodal prefrontal brain areas integrate these signals to support abstract, global self-beliefs. This architecture supports metacognitive control at multiple timescales, from rapid error correction to strategic task arbitration and social communication. Dysfunctional metacognition can lead to disorders of reality monitoring and insight, in which erroneous mental representations are endorsed with aberrant confidence. By integrating computational, cross-species and clinical perspectives, metacognitive neuroscience offers a bridge from neural circuit mechanisms to subjective experience.

Introduction

Metacognition is a broad, umbrella term for the processes that allow us to reflect on and self-evaluate mental and physical function and use these insights to control behaviour1. When lining up to take a penalty kick in football, you might estimate your chances of precisely placing the ball in the top-right hand corner of the goal. If you realize that you are not up to the task, you could hand over the task to a more confident teammate. Or, when driving down a foggy road at night you might deliberately decide to slow down, because you realize that the visual information that you have available is not sufficient to efficiently stop at higher speeds. These examples illustrate the intimate connections between metacognitive self-evaluations and adaptive control, guiding how we seek information and advice, how we regulate our learning and collaborate with others, and how or whether we offload to external aids. Metacognition research has its origins in cognitive psychology, particularly within educational contexts in which researchers have been interested in how children develop beliefs about their own knowledge2. Historically, metacognition was considered a uniquely human faculty. More recently, it has become possible to develop task-based measures and computational metrics of metacognition that can be deployed across both humans and non-human animals3,4 — metrics that are now proving useful for probing metacognition in artificial intelligence systems (Box 1). Previous review articles have focused on psychological5–7 and computational8,9 aspects of metacognition; here, I focus on its neural basis, with the aim of integrating findings from systems and cognitive neuroscience. In this Review Article, the term metacognition is used to refer primarily to metacognitive monitoring, as this aspect has received the most attention from neuroscientists. Metacognitive monitoring refers to the ongoing assessment of the reliability or quality of one’s own mental processes — such as perception, memory, reasoning or decision-making — and typically involves the generation of second-order representations of confidence about a first-order cognitive process. Confidence formation is therefore a core variable of interest for metacognition research. Metacognitive monitoring can be harnessed for adaptive metacognitive control (Fig. 1a), in which metacognitive estimates are used either for regulating behaviour and/or for communicating metacognitive assessments to others10. I turn to these functional roles of metacognition later in the Review. From a functional standpoint, introspection and metacognition are similar psychological concepts. However, because the term introspection tends to presume a connection with consciousness, I use the more neutral term metacognition here, and instead consider in Box 2 how metacognitive mechanisms may themselves be important for developing functional explanations of consciousness. I begin by articulating the quantitative approaches that can be used to define and measure metacognition in the laboratory (Fig. 1b). I then aim to synthesize results obtained in rodents, monkeys and humans that illuminate the neural basis of canonical metacognitive judgements and showcase how a network of brain regions, centred on hubs in the prefrontal cortex (PFC), build up beliefs about self-performance via interactions with domain-specific and modality-specific areas. I ask how metacognitive processes that are expressed on a local, moment-tomoment timescale shape the formation of global self-beliefs about skills and abilities over longer timescales (Fig. 1c). Finally, I ask how distortions in metacognitive processes may underpin a variety of mental health symptoms in which there is a disconnect between subjective reality and objective performance.

Multiple forms of metacognition are underpinned by the ability to form estimates of uncertainty in a self-directed frame of reference, with these estimates expressed as propositional confidence about our own behaviour and mental processes7. This capacity allows us to estimate feelings of surety in our own (hypothetical) decisions or actions, resulting in covert propositions such as “I think I will remember this word.” Propositional confidence can be distinguished computationally from other forms of world-directed or distributional uncertainty11,12. The self-directed nature of metacognition makes it challenging to design tasks that truly tap into metacognitive processes, as not all behaviour that seems metacognitive is metacognitive in nature13. For instance, selectively avoiding difficult choices may be driven by properties of the stimulus or situation, rather than representing to oneself that the decision is difficult. Tasks designed to measure metacognition require two components: a first-order behaviour or mental state to be monitored and a way of eliciting a second-order evaluation of the first-order behaviour or mental state (Fig. 1b). In humans, these judgements are usually explicit: subjects are instructed to press a button or select a point on a scale to indicate their confidence. In other cases, such as confidence forced-choice paradigms, they are required to select the decision that they feel most confident about from a pair of decisions14. In animal metacognition research, confidence estimates are elicited via so-called implicit measures of confidence, described further below. When we are comparing the literature on metacognition across species, it is important to keep in mind that behavioural and neural markers of metacognition in animals may be tapping into implicit forms of metacognition: whether there is a cross-species analogue of explicit metacognition remains unclear15. A popular approach to assess metacognition in non-human animals harnesses opt-out paradigms, in which the animal is initially trained to produce a first-order behaviour (for instance, a perceptual discrimination judgement) and then trained to select an additional opt-out response when uncertain about its primary decision4. Critics of opt-out paradigms highlight ambiguity as to whether they are being solved in a world-directed or self-directed manner, because an effective opt-out decision can be made by tracking stimulus strength, rather than performance conditional on a chosen action15. Indeed, on opt-out trials, it may be the case that a primary decision was never formed. However, if the opportunity to opt-out is offered randomly and intermittently and is unknown to the decision-maker when viewing the stimulus, provisional formation of a primary decision is more likely16. An alternative approach has used variants on post-(or peri-)decision wagering — in which subjects make incentivised wagers on their performance, elicited as a second-stage behaviour (such as a button press) after each trial of the first-order task. The typical setup of a post-decision wagering experiment involves training the animal or human to understand that a high confidence response is risky, as it will only be highly rewarded when correct, whereas a low confidence response is guaranteed to receive a smaller reward irrespective of response correctness. This creates incentives to engage in metacognitive monitoring to wager effectively. In a variant of the post-decision wagering paradigm developed for use with rodents3 and human infants17, subjects are trained to wait for an uncertain reward that is more likely to be delivered when they are correct, and withheld when incorrect. At any point in this waiting period, they can give up to start a new trial. The optimal strategy here is to wait longer when confidence in being correct (and therefore the subjective probability of a reward

Box 1 | From biological to artificial metacognition Given the importance of metacognition for monitoring and evaluating knowledge and performance it has become increasingly pressing to understand the metacognitive capacities of artificial intelligence (AI) systems187. The incorporation of effective metacognition into generative AI systems such as large language models has been historically difficult to achieve, with notable disconnections between confidence and accuracy in safety-critical domains such as medical decision-making188. This can lead to a ‘calibration gap’, whereby humans assume artificial systems are more accurate than they really are189. There is, however, growing evidence that artificial neural networks can form and update calibrated confidence estimates87, estimate their own knowledge states190, identify their own behavioural propensities191,192 and potentially have access to their own internal workings193,194 — albeit with profiles that may diverge substantially from both normative standards and human judgement195. Artificial neural networks also provide a useful testbed for exploring the functional profile of different metacognitive architectures. For instance, one study found that idiosyncratic biases in confidence emerge within a convolutional neural network trained

to discriminate handwritten digits — indicating that such biases may not be unique limitations of human metacognition but rather a core feature of how high-dimensional evidence spaces are mapped to confidence87. The mechanisms supporting such ‘introspection’ in artificial systems are likely to be very different from architectures supporting human metacognition196 and will depend on the tasks the systems have been optimized to solve. If a main function of introspective systems is to allow performance prediction and error monitoring, then a low-dimensional metacognitive signal may suffice. But if there is a functional need to communicate the internal states of the system to others, then a more granular introspective system may be required. In either case, the development of such metacognitive capacities in AI systems will likely be important in managing both self and social regulation in human–AI interactions187. In turn, as the first-order capabilities of AI become more sophisticated, the demands placed on human metacognition will increase, as people will be required to know when and how to efficiently offload cognitive work to AI and monitor its output197.

being available) increases18. Figure 1b illustrates the various approaches that have been used to measure metacognition. Once first-order and second-order behaviours have been elicited over multiple trials or episodes of a task, these data can be analysed to quantify a metacognitive profile. The field has made use of two broad classes of summary statistic19. The first is metacognitive bias — the level of a metacognitive assessment (for instance, high or low confidence) relative to performance (this is sometimes referred to as calibration or simply confidence level). The second is referred to as metacognitive sensitivity — the extent to which metacognitive judgements correlate with behavioural performance from trial to trial. Estimates of metacognitive sensitivity are easily confounded with changes in first-order performance or metacognitive bias19. When sensitivity is corrected for these confounds, it is sometimes referred to as metacognitive efficiency. A popular approach to distinguishing bias and sensitivity in metacognitive judgements is to apply the meta-d′ model20, which grew out of established signal detection theory (SDT) approaches within cognitive psychology. SDT provides a simple framework for measuring how humans and animals make decisions under uncertainty across a range of tasks. A classical SDT analysis returns a sensitivity parameter, d′, which reflects a participant’s performance on the first-order perception, memory or decision task. Meta-d′ reflects the first-order sensitivity parameter that best explains an agent’s confidence ratings under a lossless SDT model. Meta-d′ is a descriptive statistic and admits of various computational architectures underpinning better or worse metacognition.

selective impairment in the accuracy of prospective metacognitive judgements of memory compared to an amnesic control group, despite recognition memory performance being similar. This research identified the human PFC as a key neural substrate for metacognition22 and underpinned subsequent exploration of the neural correlates of metacognitive sensitivity using performance-corrected tools and metrics. Early investigations of individual differences in metacognitive sensitivity also added weight to the proposal that the PFC supports metacognitive capacity, identifying links with its structure23–25, function26,27 and connectivity28–30. This work was complemented by studies of the lifespan trajectory of metacognition, showing that metacognitive sensitivity matures through adolescence25,31, plateaux in adulthood32 and can remain robust into older age despite other aspects of task performance deteriorating33,34. A second wave of studies beginning in the 2000s used functional brain imaging to chart the neural correlates of metacognitive judgements. A first approach, in humans, compared the requirement for a metacognitive judgement to a control condition, highlighting selective increases in activation within the medial frontal cortex (specifically, the dorsal anterior cingulate cortex) and lateral frontopolar cortex (FPC)26,35,36. A second approach sought to relate changes in neural activity to changes in the level of a metacognitive judgement — for instance, correlating the degree of confidence with increases or decreases in the BOLD signal or neural firing rate. Here again, studies identified neural correlates of confidence in the frontal and parietal cortex of monkeys16,37,38, humans26,39,40 and rodents41. In rodents, lesions to the orbitofrontal cortex (OFC) and ACC have been shown to reduce metacognitive sensitivity, while leaving first-order performance intact42,43. In both humans and non-human primates, lesions or inactivation of the FPC have been linked to selective metacognitive deficits38,44–47 (Fig. 2), consistent with its position at the apex of a neurocognitive hierarchy48. The earliest work in this second wave predominantly focused on metacognitive judgements about percepts and memories that have objectively correct or incorrect answers — truth-conditions that can be

Neural basis of metacognition

Neural correlates of metacognition: a brief history A first wave of studies starting in the 1980s identified putative neural correlates of metacognition through observational studies relating metacognitive sensitivity to features of brain structure and function. Shimamura and Squire21 found that people with Korsakoff’s syndrome (a chronic neurological disorder associated with alcohol misuse) had

a

b PDW task Stimulus

Control

Monitoring

Meta-level (second order) processes

Response

Peri-DW task

Opt-out paradigm

Peri-DW

Trial 1

Stimulus

Response

Trial 2 Stimulus

Response

Stimulus

Stimulus Respond

Aspects of self-performance evaluated

Confidence forced choice

Prospective confidence task

Peri-DW

Peri-DW

Global

Confidence forced-choice task

Opt-out

Object-level (first order) processes

c

Confidence estimate or PDW

Stimulus

Judgements of learning or feelings of knowing

Response

Peri-DW

Fluctuations in metacognitive estimates

Cognitive faculties

Examples “I have a good memory.” “I should do well at physics.”

Task performance “I did ok on my history exam.” Decisions

“I made a mistake on this question.”

Local

Fig. 1 | Tasks and frameworks for characterizing metacognition. a, The classic Nelson and Narens model of metacognition, incorporating monitoring and control10. Meta-level (second-order) processes are proposed to monitor and control first-order (object-level) mental processes. b, A range of different paradigms have been developed for the measurement of metacognition across species. In retrospective confidence/post-decision wagering (PDW) tasks a confidence estimate or post-decision wager is elicited following a first-order task response. In a confidence forced-choice procedure the participant is asked to choose which of two trials they performed better on. In a peri-decision wagering (peri-DW) task the first-order response and the confidence estimate are elicited simultaneously. In an opt-out paradigm, participants can either choose to respond or opt-out conditional on their assessment of the difficulty of the trial.

Other tasks measure confidence about future decisions or actions (prospective confidence), using measures that include judgements of learning or feelings of knowing. The red and blue shading roughly corresponds to the stages of the task behaviour that map onto first-order and second-order components of the model in panel a. c, A hierarchical framework for metacognition posits that reciprocal interactions between local and global self-evaluations unfold over different timescales101. Estimates of self-performance shown on the left span from local, moment-to-moment evaluations of individual decisions and actions up to more global, long-run evaluations of cognitive faculties. The traces in the centre illustrate the fluctuations in metacognitive estimates at different levels over time, with reciprocal and interactive relationships between levels. Right: examples of the ways in which these estimates might be expressed are shown.

verified by the experimenter. However, in daily life we often self-evaluate decisions without an obvious correct answer — so-called value-based or subjective decisions. By asking for explicit judgements of confidence in subjective decisions (such as which of two snack items people preferred to eat), it was possible to show that people could recognize suboptimal subjective choices29. Notably, confidence in value-based decision-making exhibited similar neural correlates to those found in studies of perceptual and mnemonic metacognition: whereas the ventromedial PFC carried multiplexed confidence and value signals, the FPC specifically tracked both confidence and individual differences in metacognitive sensitivity29,49. When interpreting studies of the neural correlates of confidence, performance confounds must be taken into account. Fluctuations in confidence are often associated with changes in both stimulus strength and choice accuracy, rendering it ambiguous as to whether the observed neural response is covarying with factors affecting first-order performance, metacognition, or a mixture of the two. The advantage of computational frameworks is that they make explicit the latent processes involved, allowing us to discover the building blocks of a

canonical metacognitive judgement and dissociate them from factors driving first-order task performance50. Below, I consider how findings on the neural basis of metacognition relate to hypotheses derived from different computational accounts.

Neurocomputational approaches to metacognition The simplest models of metacognition, sometimes referred to as direct access or first-order models, assume that the subject has perfect access to the evidence or process underpinning their first-order behaviour. According to these accounts, there should be minimal dissociations between performance and metacognition, and the evidence used for both is computed in the same (world-directed) frame of reference. Indeed, such accounts are not really metacognitive, in that apparently metacognitive behaviour is being explained by a first-order (non-metacognitive) process. However, first-order accounts struggle to accommodate cases in which the psychological or neural characteristics of metacognition diverge from those supporting first-order performance51. In such cases, additional computations supporting metacognition need to be elucidated.

Two broad cross-cutting distinctions help organize these more advanced computational approaches to confidence and metacognition. First, we can divide models into static and dynamic, based on whether or not they seek to model the latencies of task performance and metacognitive judgements. Second, in both static and dynamic models we can identify serial, parallel and hierarchical architectures relating performance to metacognition8,51. Static models of metacognition. Static models of metacognition seek to characterize quantitative relationships between performance and self-evaluation while remaining agnostic as to the dynamics of these relationships over time (either within a trial, or over multiple trials). Static models are useful because they generate statistical signatures of confidence formation — data patterns that are indicative of metacognition in a way that cannot be easily explained by other underlying decision processes. A useful statistical signature of confidence that can be straightforwardly derived from static models is the so-called ‘folded-X’ pattern18. When plotted against objective measures of stimulus strength, propositional confidence should increase with evidence strength for correct trials, and decrease with evidence strength for error trials (Fig. 3a). A folded-X pattern has been observed in human and animal experiments designed to assess confidence18 and used as a non-verbal indicator that physiological or neural signals are confidence-related52.

For example, in a seminal study, the activity of neurons in the rodent OFC was shown to exhibit a folded-X pattern in an odour discrimination task41 (Fig. 3b). This activity pattern predicted post-decision wagering behaviours and generalized across both auditory and olfactory decisions53. However, recent findings suggest that whether or not a folded-X pattern is observed is context-dependent: manipulations of evidence strength that allow the observer to infer task difficulty, such as stimulus reliability or visibility, can lead to increases in confidence for both correct and error trials54,55. Other static models have sought to isolate the noise or imprecision in metacognitive judgements from aspects of first-order task performance. In the ‘confidence as a noisy decision-reliability estimate’ or CASANDRE model, metacognition is proposed to be limited by the extent to which individuals have meta-uncertainty about the world-directed uncertainties affecting their task performance56. Meta-uncertainty can be estimated from fits to psychometric functions that reveal how stimulus strength differentially influences performance and confidence. Fitted meta-uncertainty is partly domain-general, showing cross-task correlations that are reduced as the distance between tasks in task space increases56,57, and has been estimated in both humans and monkeys58. CASANDRE is a hierarchical, static model, in that a ceiling on metacognitive performance is provided by the degree of meta-uncertainty. By contrast, parallel architectures such as the ‘confidence-noise confidence-boost’

Box 2 | Perceptual metacognition and consciousness Consciousness refers to subjective experience. Although metacognition and consciousness are distinct concepts, there is growing interest in the role of metacognitive processes in computational explanations of human consciousness. This approach has its origins in higher-order theories of consciousness, in which a central aspect of becoming conscious of a mental state is the creation of a metacognitive representation of that mental state198. Metacognitive sensitivity has been shown to be ‘consciousness selective’, in that people are typically better able to monitor their performance when they are also conscious of the relevant information185. One proposal that accounts for this linkage holds that percepts become conscious when they are deemed to be strong or reliable enough to impinge upon cognition117,199. A core metacognitive operation that is relevant to this argument is known as ‘reality monitoring’ — determining which aspects of internal neural activity track an individual’s current environment and which are engaged in internal aspects of planning, simulation and so forth. Studies have shown that mnemonic reality monitoring — figuring out whether something had previously been experienced or just imagined — recruits similar regions of the frontopolar cortex as explicit metacognition200. Recent work has developed paradigms for studying perceptual reality monitoring by asking subjects to imagine low-level visual stimuli embedded in dynamic noise, and then occasionally ‘fading in’ a real stimulus. When imagination is congruent with external stimulation and when it is particularly vivid, subjects are more likely to attribute their experience to reality — suggesting that a central driver of reality monitoring is the strength and precision of ongoing perceptual processing201. Using functional magnetic resonance imaging, it was found that anterior insula activity

both correlated with reality judgements and tracked fluctuations in the strength of mid-level visual cortical activity, furnishing a potential neural substrate for the higher-order monitoring of sensory strength202. There is natural overlap here with work on neural ‘read-outs’ of sensory uncertainty. The metacognitive computation involved in reality monitoring is likely implicit — we are not typically explicitly aware that we are engaged in this process, except in special circumstances (for instance, when trying to figure out whether a seemingly illusory experience is real). Accordingly, people have limited insight into the trial-to-trial confusions generated in perceptual reality monitoring experiments203. Instead, gaining insight into one’s hallucinatory experiences may depend on additional layers of explicit metacognition204. Within this framework, aphantasia — a lack of conscious imagery, despite intact performance on a range of tasks — may reflect a change in metacognition. The idea is that people can still internally activate the relevant perceptual representations to solve a task, but higher-order metacognitive representations (correctly) discard them as internally generated, and they are not consciously experienced154. Indeed, people with aphantasia have been suggested to exhibit a dysconnection syndrome, in which regions of the brain implicated in perceptual metacognition (such as the anterior prefrontal cortex) are decoupled from sensory areas205, potentially underpinning this change in conscious experience. It is challenging to develop animal models of distortions in reality monitoring. However, one promising approach combines measures of metacognitive confidence with perceptual tasks, in order to identify confident but false percepts206.

a

Task performance

Metacognitive sensitivity 1.0

Accuracy

OFC

Accuracy

1.0

0.5

0.7 0.4

0

20

40

60

–1.5

Constrast between stimuli (odour-mixture ratio) (%)

0.0

Control

∆recognition (∆d’ (POST–PRE))

Lateral FPC

Metacognitive sensitivity

0.4

0.4

n.s. 0.0

–0.4

–0.8

0.0

–0.4

–0.8

Muscimol

c

50

Chance performance level

40 30 20 10 0

Healthy controls

Metacognitive efficiency (meta-d’/d’)

Performance (% correct)

60

Saline

Metacognitive efficiency

Task performance 80 70

Muscimol

Task performance

∆confidence (∆meta-d’ (POST–PRE))

b

1.5

Confidence (Z-scored waiting time)

1.4 1.2

Optimal metacognitive efficiency ** *

1.0 0.8 0.6 0.4 0.2 0.0

aPFC lesions

Temporal lobe lesions

Fig. 2 | Convergent evidence for a role of prefrontal cortex in metacognition across species. a, In rodents completing a post-decision wagering task in which confidence is indicated by the time the animals are willing to wait for reward delivery, reversible inactivation of the orbitofrontal cortex (OFC) via infusion of the GABA-A receptor agonist muscimol leads to a deficit in metacognitive sensitivity (the relationship between confidence and accuracy) (right) without changing task performance (left)42. b, In monkeys engaged in a post-decision wagering task, reversible inactivation of the lateral frontopolar cortex (FPC) via injection of muscimol leads to a deficit in metacognitive sensitivity (right) but not recognition memory task performance (left)38. PRE and POST refer to before and after the injection, respectively; meta-d′ is a measure of metacognitive sensitivity extracted from a signal detection theory model. *P < 0.05; n.s., not significant. c, In humans, lesions to the anterior prefrontal cortex (aPFC) are associated with deficits in metacognitive efficiency for perceptual judgements (right) despite intact perceptual task performance (left)44. **P < 0.01; *P < 0.05. Panel a adapted with permission from ref. 42, Elsevier. Panel b adapted with permission from ref. 38, Elsevier. Panel c adapted with permission from ref. 44, CC BY 4.0 (https://creativecommons. org/licenses/by/4.0/).

model allow for independent sources of evidence to contribute to a confidence judgement, such that metacognition may be ‘boosted’ relative to performance14,51. Combining these static models with neural measures offers powerful opportunities for future work on the neural basis of metacognition58. Dynamic models of metacognition. A limitation of static models of metacognition is that they do not make predictions about the timescales of performance and self-evaluation, restricting the extent to which they can be related to the dynamics of brain activity. Just as static models have often been grounded in SDT, most dynamic models of metacognition have built on evidence accumulation (EA) models. EA models propose that the decision-maker accumulates noisy evidence in favour of different choice options until a threshold is reached and a decision is made59. EA models were initially developed to account for first-order perceptual and mnemonic decision-making but can be extended to accommodate confidence by assuming that metacognitive evaluation rests on extracting the degree of evidence supporting the first-order decision60. From this shared starting point, models diverge as to their core assumptions about the mechanisms for generating confidence estimates (Fig. 3c). Mirroring the divide between hierarchical and parallel static models, distinct model variants assume either serial or parallel EA for confidence. Serial models propose that, after a decision, additional post-decisional accumulation continues, capturing time-dependent fluctuations in subsequent confidence level61–63. By contrast, parallel models allow confidence to be continuously ‘read out’ from the evidence available during the pre-decision phase16. There is evidence for behavioural and neural correlates of both of these processes (Fig. 3d). In support of parallel models, the dynamics of EA in the pre-decision phase have been shown to predict confidence levels. In monkeys, the rate of EA in lateral intraparietal cortex firing rates has been linked to confidence measured using both opt-out choices16 and peri-decision wagering64. In humans, research using electro­­-encephalography (EEG) has found that the centroparietal positivity, an event-related potential component linked to EA in perceptual decision-making, predicts confidence level measured via explicit ratings65, including those measured using psychophysical techniques to dissociate confidence from performance, reaction time and stimulus strength66. There is also evidence for a parallel encoding of confidence. In a recent study in monkeys, peri-decision wagering was used to simultaneously measure choices and confidence in a motion discrimination task64. The authors found that the monkeys’ behaviour was well-fit by an EA model in which choice and confidence are evaluated concurrently, but not by a serial choice-then-confidence model. Single-trial decoding of neural activity in the lateral intraparietal cortex revealed that choice and confidence-predictive signals emerged simultaneously at the start of the trial, consistent with the parallel evolution of task performance and metacognition, and with previous research proposing that decisions and confidence estimates are encoded in parallel cortico-thalamic loops67. In an experiment using human magnetoencephalography (MEG), a similar decoding profile was observed, with both choice and confidence signals emerging in parallel at stimulus onset, and occupying different sensor subspaces (implying distinct neuronal populations) across the frontal and parietal cortex68. These parallel signatures of EA for choice and confidence may explain the capacity to dynamically develop metacognitive estimates ‘online’ for use in the rapid correction of errors69–71 and the control of decision policies72,73. Finally, there is evidence for post-decisional processing contributing to confidence formation74, consistent with a serial architecture.

Much of this work has been influenced by early discoveries about the neural basis of error monitoring. A canonical finding is that posterior medial frontal cortex (pMFC) neurons activate when errors are made in both macaques75 and humans76, generating a classical error-related negativity at the scalp surface74,77. A later event-related potential component, the Pe, has been linked to explicit metacognitive reports of error78,79 and is thought to be generated by the insula cortex80. More recently, work in humans using EEG has revealed that decision errors trigger a second phase of EA (linked to a similar centroparietal positivity potential associated with pre-decisional EA) that can explain features of confidence, error monitoring and changes of mind71,74,81,82. Similarly, using functional magnetic resonance imaging (fMRI), regions commonly implicated in error monitoring such as the pMFC have been shown to harbour graded signatures of evidence for or against a previous choice83. Some of the debate over parallel and serial architectures for confidence reports may be due to the timing of the metacognitive judgement in different tasks. When metacognitive judgements are collected retrospectively, the importance of a post-decisional phase of EA is emphasized37. By contrast, when metacognitive judgements are elicited concurrently with a choice (as in peri-decision wagering paradigms), emphasis is placed on parallel signatures of confidence. As has previously been noted, there is no reason that post-decisional contributions to confidence judgements should be mutually exclusive of a provisional confidence estimate emerging during the decision process8. In practice, in naturalistic settings, both pre- and post-decisional dynamics are likely to contribute to the continuous evolution of the metacognitive state of the system. Sources of information for metacognitive computation. What information informs these neural signatures of propositional confidence? In the previous section we explored static and dynamic models that characterize how latent decision variables (such as EA or meta-uncertainty) influence confidence formation. But in isolation, these models are largely silent on the kinds of sensory and mnemonic representations that inform these computations. One prominent proposal is that, within a given task, world-directed uncertainty needs to be extracted by brain areas downstream of sensory regions to enable metacognitive behaviours12,56. World-directed uncertainty refers to external quantities, such as sensory uncertainty about line orientation, or cognitive uncertainty about whether Manchester United will win the English Premier League. These uncertainties are not themselves metacognitive — they are about aspects of the world, rather than the self — but they provide crucial inputs to metacognitive judgements (Fig. 4a). For instance, when making a judgement about line orientation, it is proposed that a summary of the implicit or ‘distributional’ uncertainty can be extracted from the activity of neural populations in visual cortex tracking line orientation84. Similarly, in an auditory task, the relevant uncertainties are suggested to be implicitly encoded by populations of auditory cortex neurons85. Recent work has sought to understand how the neural encoding of world-directed sensory, mnemonic and motor uncertainty is connected to the downstream neural mechanisms underlying self-evaluative judgement. To this end, a machine learning approach has been used to decode trial-by-trial uncertainty in the representation of the orientations of a visual grating from fMRI data collected from the human visual cortex. Reported confidence was negatively correlated with the decoder’s readout of uncertainty, even for identical stimuli86. In turn, the decoder’s readout of sensory uncertainty was correlated with

fMRI signals in the PFC, suggesting that domain-specific uncertainty estimates in sensory regions inform downstream representations of propositional confidence in prefrontal regions. In another recent study, monkeys were trained on an orientation discrimination task and reported their confidence using a peri-decision wagering procedure58. Although both linear and nonlinear decoders applied to neuronal activity in the primary visual cortex predicted first-order choices, nonlinear decoders consistently outperformed linear decoders in predicting the monkey’s confidence estimates, suggesting that confidence arises from an additional transformation of the sensory signals that inform perceptual decisions87. Similarly, another fMRI study showed that a proxy for sensory uncertainty (motion coherence, which was dissociated from choice difficulty) was related to activity in the extrastriate visual and parietal cortex, whereas activity related to propositional confidence was observed within the perigenual anterior cingulate cortex (pgACC)88. These results are consistent with the idea that world-directed visual uncertainty (tracked in visual and parietal cortex) is converted to a self-directed format in downstream PFC. It remains unclear whether similar mechanisms apply beyond sensory and motor representations, contributing, for example, to confidence in being able to remember something. A recent study in macaque monkeys recorded from neural populations in the lateral PFC during a spatial working memory task. Population activity in this region during the delay period in which the animals were required to remember the locations of the stimuli jointly encoded both the remembered locations and their associated uncertainty, with a decoded measure of working memory strength predicting both recall accuracy and opt-out decisions on a trial-by-trial basis. Notably, a distinct metacognitive signal — reflecting the decision to opt-out — emerged approximately 200 ms after the working memory representation in a different subspace of the same PFC population, consistent with the idea that the metacognitive judgement is informed by a readout of mnemonic uncertainty89. fMRI evidence obtained in humans also supports the proposal that population-level representations of uncertainty are maintained in the brain areas that track mnemonic content during visual working memory tasks90. Other forms of world-directed uncertainty encompass our own bodies. For instance, uncertainty about the extent to which our muscles and joints are able to accomplish particular actions, such as rapidly pointing to a target91. Compared to sensory or mnemonic uncertainty, less is known about how motor and action uncertainty is tracked and read out for use in online metacognitive evaluation. However, in tasks requiring skilled action, this is likely to be a key contributor to confidence formation92,93. In humans, motor kinematics have been shown to impact upon error-related brain activity94 and aspects of action production including speech prosody, response time and movement time have all been linked to confidence95–97. Another influence of action on metacognitive evaluation concerns memory for one’s own decisions and actions. Because the evidence contributing to confidence estimates may be subject to further processing (either parallel or serial, and potentially in downstream brain regions), the decision or action that was actually taken often needs to be stored for later use in self-evaluation51. Consistent with this proposal, in human neuroimaging studies, activity in the frontopolar and insula cortex, and enhanced functional connectivity between motor and frontal cortex, was increased when self-action information was available to inform metacognitive judgements98,99. One implication of this computational constraint is that metacognition can be improved by making initial choices explicit, a prediction borne out by experiment98–100.

b

15

Probability density

Signal Noise

15

Firing rate (spikes s–1)

Weak evidence

Firing rate (spikes s–1)

a

10 5 0 0.0

0.5

1.0

Binaural contrast

c

10 5 0 50

53

58

95

Strongest odour (%)

Correct choice port

OFC

Strong evidence

Trial start port

Probability density

Incorrect choice port

Correct Reward

Odour ‘puff’ c Decision variable

1.0

Correct

Confidence

Error

Leaving

Folded X-pattern

0.5 Weak

Evidence

Error

Strong

No reward

Stimulus

Decision

Time investment

Outcome

High confidence Pre-decisional models

Post-decisional models

Confidence variable

c

Time

Time

60

3

µV m–2

CPP amplitude (µV)

4

Time

2 1

40

Prediction accuracy

d

Parallel models

Decision variable

Decision variable

Choice

Low confidence

20

0

0 0

200

400

600

800

Time from stimulus onset (ms)

–0.4

Error

Taken together, these studies support a computational scheme in which effective metacognition involves tracking world-directed uncertainty in a self-directed format to enable adaptive control of behaviour7,51,74.

Time (s)

0.4

0.8

Choice

Wager

–0.4 –0.2

0.0

0.9 0.8 0.7 0.6 0.5 0.0

0.2

0.4

0.6

0.2

Time from motion and saccade onset (s)

Timescales of metacognition. Metacognitive processes operate dynamically, over multiple interacting timescales: local metacognitive processes track moment-to-moment performance, whereas global metacognition encompasses self-estimates of skill or ability that are maintained

Fig. 3 | Overview of static and dynamic models of metacognition. a, A static signal detection theory-based framework for confidence judgements. Upper panels: according to classical signal detection theory, a decision variable is obtained from either a signal (solid line) or noise (dotted line) distribution and compared with a criterion (c) to make a decision. In the example shown, the correct decision is obtained for above-criterion responses (blue-shaded areas). When evidence is weak (that is, when there is more overlap between the signal and noise distributions), the chances of making an error (the red shaded areas) are greater than when the evidence is strong. Confidence reflects the distance of the decision variable from the criterion. Lower panel: when averaging over trials, this produces a characteristic ‘folded-X’ pattern: as evidence strength increases, confidence decreases for errors and increases for correct decisions. b, This folded-X signature has been observed in neuronal firing in the rodent orbitofrontal cortex (OFC) for auditory and olfactory decision tasks41,53. Both tasks involved an initial perceptual decision followed by an unpredictable time interval for reward delivery, during which firing in the OFC was measured. The amount of time the animal was willing to wait for a reward before moving on to start the next trial is a proxy of confidence. c, Schematic representations of the ways in which dynamic models of decision-making accommodate variability in metacognitive judgements. Left: predecisional accounts propose that the dynamics of evidence accumulation before a decision are monitored to inform confidence. Middle: post-decisional accounts propose that confidence is informed by the continued accumulation of evidence after a choice has been made. Right: parallel accounts propose that decisions and confidence estimates are formed by accumulating evidence in parallel neuronal populations. d, Empirical evidence in humans and monkeys supports a role for all three dynamic signatures in metacognitive evaluation. Left: electroencephalogram data from humans demonstrating that the slope of the centroparietal positivity

(CPP) correlates with perceptual decision confidence before a decision being made, consistent with pre-decisional accounts of confidence formation66, are shown. Black points below the plot denote the centre of 200-ms time windows in which a linear effect of confidence on CPP slope reached significance (P < 0.05) after correcting for multiple comparisons across time. Middle: human electroencephalogram data demonstrating that a second-stage CPP is triggered when errors are made, consistent with post-decisional accounts of confidence formation, are shown. Deeper colours indicate faster error detection responses81. The vertical dashed lines represent median response times (RTs). The grey markers indicate timepoints when linear regression of RT on signal amplitude reached significance (P < 0.05); black markers indicate centre of 150-ms time windows in which regression of RT on signal slope reached significance (P < 0.05; one-tailed predicting steeper slope for faster RTs). Right: prediction accuracy for decoders fit to monkey lateral intraparietal cortex neuronal population data predicting choice (purple) and peri-decision wagers (yellow) as a function of time and aligned to motion onset and saccade onset64 is shown. The parallel onset and evolution of choice and wager (confidence) decoders is consistent with parallel accounts of confidence formation. The coloured bars at the top indicate when the accuracy for the corresponding decoder was significantly greater than chance (one-tailed Wilcoxon rank-sum test with Šidák’s correction). The black bar indicates when prediction accuracy was significantly different for choice versus wager (Wilcoxon signed-rank test with Šidák correction). Part a adapted with permission from ref. 52, CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). Part b adapted with permission from ref. 53, Elsevier. The left panel of part d reprinted with permission from ref. 66, Sage. The middle panel of part d reprinted with permission from ref. 81, CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). The right panel of part d reprinted from ref. 64, Springer Nature Limited.

over days, weeks or even years101. There has been increasing interest in understanding the interplay between these timescales of metacognition. Initial findings in humans documented ‘confidence leak’ between trials, an autoregressive phenomenon in which confidence on previous trials predicts confidence on subsequent trials, despite the trials being independent102. These observations hinted at slower dynamics to metacognition that may be obscured by classical single-trial analyses102. Laboratory studies have found that fluctuations in local confidence from trial to trial significantly predict global confidence estimates103,104, with computational modelling indicating that people learn both from external feedback and local confidence to construct a self-model of their skills and abilities across a range of task domains105. Combining a local-global confidence integration task with fMRI revealed a posterior-to-anterior gradient within the PFC: pMFC (dorsal anterior cingulate cortex) activity was sensitive to local (but not global) confidence whereas more anterior ventromedial PFC was modulated by both local confidence and by the level of global belief in performance across blocks106. In a Bayesian framework, global confidence may reflect a higher-level prior on the precision of perceptual and cognitive representations85,107. Oscillatory activity in the form of neuronal oscillations at alpha and beta frequencies has been linked to priors on confidence in both auditory108 and visual109 perceptual tasks in humans. However, compared to ‘bottom-up’ influences on confidence formation, less is known about the neural basis of these contextual influences on metacognitive judgements110,111. One idea is that metacognitive priors are inherited from our sociocultural milieu and possibly only loosely connected to the task at hand112.

(Fig. 4). The sensory and motor cortices contain rich information about task-specific uncertainties, which are tracked and transformed into a self-directed frame of reference in downstream circuits supported by the parietal and prefrontal association cortices. Dynamic models of confidence formation suggest that the brain extracts, in real time, estimates of decision reliability that are informed by the unfolding perception–action cycle — a computation that is linked to EA signals in the parietal cortex. In turn, self-directed (propositional) confidence is represented by multimodal areas of the PFC (insula and medial PFC, together with their cortico-thalamic and cortico-striatal loops) and further broadcast to the anterior PFC for use in self-regulation and public communication. These integrated, cross-domain representations of confidence are likely to be informed by broader cues and priors that (in humans, and possibly other animals) are supported by a ‘self-model’ that shares neural machinery with social cognition40,113. Finally, a posterior-to-anterior gradient within the human PFC may support the more stable, long-run metacognitive estimates that underpin global self-beliefs about skill and ability. There is evidence of notable commonalities in this architecture across species. First, both human and monkey metacognitive judgements in perceptual tasks are sensitive to fluctuations in uncertainty estimated from neural populations in the early visual cortex58,86. Second, in studies of perceptual decision-making, confidence is intimately linked to the properties of neural EA pre- and post-decisionally in both humans and monkeys37,64–66,74. Third, it is striking that medial aspects of agranular PFC carry domain-general confidence or error signals in mice, monkeys and humans53,75,88,114. Finally, there is evidence for a causal contribution of anterior/lateral aspects of the granular PFC to metacognitive judgements in both humans and monkeys38,44,47. There are, however, salient differences between species (Fig. 4b). Findings of PFC involvement in metacognition seem more prevalent

An integrated model of metacognition We can now tentatively advance a working hypothesis for the neural basis of metacognition in simple cognitive and perceptual tasks

a

Self-model

CW

CCW

Communication

CW

Possible line orientations

b

Communication or control

Local confidence

Probability

Probability

Sensory uncertainty

Confidence-based behaviour

CCW

Human pre-SMA PCC/Prec. dmPFC

IPS

dACC

dIPFC

vmPFC

V1

pgACC

Insula

FPC aIPFC

Sensory uncertainty Local confidence Communication or control Self-model Monkey

Global confidence

Rodent LIP

ACC

dIPFC V1

OFC

FPC

Fig. 4 | Towards an integrative neuroscience of metacognition. a, A graphical illustration of the components of a perceptual metacognitive judgement. A belief about possible world states (here, the curve represents a probability distribution over possible line orientations) generates sensory uncertainty, which is converted into (local) propositional confidence with regard to a categorical decision of whether the stimulus is tilted clockwise (CW) or anticlockwise (CCW). A propositional confidence estimate is then readout for communication or metacognitive control (confidence-based behaviours). Background beliefs about a range of factors influencing self-performance are furnished by a self-model and modulate the formation and usage of confidence. b, An illustration showing how the different computational components of metacognition map onto the functional anatomy of humans, monkeys and rodents. Not shown are cortico-basal ganglia and cortico-thalamic loops or

macaque medial frontal areas linked to error monitoring. Regions are colour-coded by their association with different computational signatures in the literature so far; where more than one signature has been associated with a particular region, a colour gradient is used. ACC, anterior cingulate cortex; OFC, orbitofrontal cortex; dlPFC, dorsolateral prefrontal cortex; FPC, frontopolar cortex; LIP, lateral intraparietal area; IPS, intraparietal sulcus; alPFC, anterolateral prefrontal cortex; PCC, posterior cingulate cortex; Prec., precuneus; pre-SMA, pre-supplementary motor area; dACC, dorsal anterior cingulate cortex; dmPFC, dorsomedial prefrontal cortex; vmPFC, ventromedial prefrontal cortex; pgACC, perigenual anterior cingulate cortex. Part a adapted with permission from ref. 7, CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/).

in human studies than non-human animal studies. There are several potential reasons for this. First, there may be a sampling bias in the types of neural recordings used in animal studies of confidence and metacognition. In support of this, when fMRI was used to obtain whole-brain coverage in non-human primate research, the causal importance of the anterior PFC for metacognition in macaque monkeys was uncovered38,46. Second, reliance on post-decisional, explicit reports in human studies may emphasize the recruitment of prefrontal areas sensitive to self-action and conscious self-reflection. This highlights the importance of developing tasks to compare the neural basis of both implicit and explicit metacognitive measures in human studies, to enable more appropriate comparisons to be made with neural signatures obtained in other species. Finally, it is possible that, even when task features and metacognitive measures are matched, different species may have adopted different algorithmic solutions to metacognitive performance. This possibility was recently supported by a study that compared the neural basis of metacognition across humans and monkeys performing the same prospective confidence task115. The task involved comparing an estimate of (future) self-performance in a motion discrimination task against a fluctuating probability of receiving an external reward. Using fMRI, it was shown that the monkey lateral PFC coded self-performance and external reward probability in different subregions. In humans, however, these two latent quantities were jointly tracked in the lateral anterior PFC (Brodmann area 47 (ref. 116)). By comparing the same task across species, insights into the aspects of metacognitive machinery that are conserved can be exposed, potentially providing information that can guide novel interventions to remedy metacognitive dysfunction.

Functional roles of metacognition Once metacognitive estimates of self-performance are formed, they can be used in the guidance of future behaviour. As noted above, when compared to the increasingly rich picture of metacognitive monitoring that has been provided by neuroscience research, less work has focused on the neural basis of metacognitive control. In this section I explore what is known, and describe opportunities for future work. Like metacognitive monitoring, metacognitive control processes operate over both shorter (local) and longer (global) timescales. Cross-cutting these different timescales are distinct domains of control, which I organize into self and social regulation (Fig. 5).

Self regulation On a local timescale, metacognition plays an epistemic role in rapid error correction, adaptive regulation of evidence gathering72,117, selection of the next step in sequential decision-making118 and the control of information seeking119–121. In studies in humans, for example, the current level of confidence during EA has been proposed to regulate how much evidence is required to commit to a decision72,122. At slightly longer timescales, fMRI correlates of perceptual confidence in human medial PFC (vmPFC and ACC) were found to predict trial-by-trial variation in curiosity and information seeking121. Confidence can also act as a proxy feedback signal when none is available, shaping learning and updating of beliefs123,124. In macaque monkeys, for example, dopamine neurons were found to encode teaching signals (reward prediction errors) that scaled in proportion to confidence125. It remains to be determined how the machinery of metacognitive monitoring described in earlier sections furnishes the relevant confidence signals for use in learning. The formation of more global metacognitive estimates is likely to be crucial to allow the strategic arbitration between tasks, or strategies

within a task. For instance, if I am deciding whether to invest my time in becoming a musician or a scientist, I might reflect on my past performance in each endeavour to estimate my likely chances of future success. Neurons in monkey pMFC were found to integrate information about previous trials to update long-run performance beliefs and drive decisions about whether to switch strategy126. Such arbitration may be supported in part by a ‘common currency’ for confidence — metacognitive signals that transcend particular tasks, allowing them to be compared in a common frame of reference53,114,127. Multiplexed domain-general and domain-specific metacognitive signals could then be readout by downstream regions to support task-level arbitration128,129 and strategy selection130. Evidence for this view comes from studies that have examined the neural correlates of confidence across multiple tasks, finding common currency representations that generalize over task in the medial PFC in both humans114 and rodents53. Other work has implicated the FPC in ‘tagging’ different first-order behaviours with degrees of confidence or uncertainty to allow efficient arbitration129,130. Over more extended timescales, stable beliefs about our skills and abilities may shape our motivation to engage in future endeavours, and impact on feelings of self-worth and self-esteem101,131.

Social regulation Metacognitive capacities are central to adaptive social behaviour, allowing both the public sharing of our own confidence and adaptive deferral to others, including artificial intelligence (Box 1). In one study it was found that the sharing of metacognitive estimates between two observers allowed them to perform a psychophysical task together better than either of them could have achieved alone132. If we wish to strategically influence group decision-making, it may be advantageous to overstate (or understate) our confidence133,134. In a study of the neural basis of these strategic adjustments of ‘public’ metacognitive statements using fMRI in humans, it was found that, whereas activity in the medial PFC (specifically, pgACC) covaried with private estimates of confidence in a random dot motion decision, the FPC additionally carried information about the extent to which one’s confidence level should be adjusted when engaging in public communication135. The recruitment of these regions may be sequenced in time: a study using EEG-informed fMRI in humans was able to separate early neural activations correlating with confidence in pgACC from later activations in FPC recruited at the time of metacognitive communication136. In turn, a nascent literature has developed to understand how we develop beliefs about the skills and capacities of others: a form of other-directed metacognitive evaluation7. Studies in humans have highlighted a role for networks involved in both theory of mind137,138 and prospective self-simulation (simulating how we would act in similar situations)139 in this process (Fig. 5b).

Metacognition as a clinical target

Altered metacognition in brain disorders Dysfunctional metacognition has been hypothesized to underpin symptoms in a range of psychiatric and neurological disorders, and has also increasingly become a target for novel therapeutic interventions. Disturbances in metacognition have been linked to psychiatric diagnoses of anxiety, depression, obsessive–compulsive disorder and psychosis140. For instance, obsessive–compulsive disorder has classically been considered a ‘disorder of doubt’ in which people lack trust in the veracity of their own perceptual, cognitive and mnemonic processes141. By contrast, people with schizophrenia have been characterized as having poor domain-general metacognitive sensitivity142,

a

Global

Metacognitive control • Motivation • Self worth • Help-seeking

Cognitive faculties

Task performance

• Arbitration • Forecasting • Offloading

Decisions

Local

• Learning • Information seeking • Communication and influence

• Evidence accumulation

• Self-regulation • Social regulation

b

Metacognition

Mentalizing (neurosynth)

vmPFC

c Interventions

Computational models of metacognitive behaviours Symptoms and lived experiences

Algorithms for subjective experience and supporting neural mechanisms

Novel genetic and molecular targets Identification of neural signatures across spatial scales

and often manifest impaired clinical and cognitive insight into the abnormal nature of their experiences143. More recently, transdiagnostic methods have been applied in large general population samples to relate metacognitive profiles to variation in symptom dimensions144. One way of organizing these findings is to posit a phenotypic space spanning changes in metacognitive bias and sensitivity across either broad or narrow task domains. For instance, anxiety-depression scores have been associated with lower local confidence (a change in metacognitive bias but not sensitivity145),

Fig. 5 | New frontiers in metacognition research. a, Extending the hierarchical framework for metacognitive monitoring101 to incorporate different timescales of metacognitive control. Aspects of control acting on shorter timescales (for instance, on an individual decision) are situated at the local level, whereas aspects acting over longer timescales (for instance, affecting multiple trials) are situated at more global levels. These timescales cross-cut both self-regulatory and social functions of metacognition. b, Meta-analyses indicate partial overlap between brain regions engaged in self-directed metacognition and those involved in mentalising (thinking about the mental states of others), particularly within the ventromedial prefrontal cortex40,184. The colour scales indicate the strength of associations between activations and metacognition/mentalising. Note that although both maps are corrected for multiple comparisons across the wholebrain volume, the numerical values and thresholds are not comparable, as they are obtained via different meta-analytic methods (activation likelihood estimation for metacognition, cluster-level family-wise error rate corrected at P < 0.05; multilevel kernel density analysis for mentalising, false discovery rate corrected at P < 0.01). Overlap was computed from voxels that surpassed these thresholds in both maps. c, Integrative approaches in metacognitive neuroscience show promise for tackling disorders of subjective experience. Metacognitive measures are consciousness selective and track distortions in subjective reality185 (Box 2). These measures can be used in conjunction with computational models to identify algorithms (for instance, for confidence formation or reality monitoring) that can be probed in animal models using non-verbal assays. Neural mechanisms supporting these algorithms can be validated across species by harnessing analysis tools such as representational similarity analysis to probe the structure of neural representations across techniques and spatial scale186. Knowledge of the implementation of these algorithms in animal models guides novel insights into their genetic and molecular basis, potentially unlocking new interventions to tackle subjective symptoms in humans. Interventions target neurobiological mechanisms to alter parameters of computational models of metacognition. vmPFC, ventromedial PFC. Part b adapted with permission from ref. 40, CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/).

whereas a symptom dimension tracking compulsive behaviours and intrusive thoughts is associated with both overconfidence at the local, trial level, and underconfidence (doubt) at more global levels146,147 — findings which generalize across distinct task domains148,149. When studying links between metacognition and mental health, care is needed to ensure that conclusions are not unduly affected by idiosyncratic biases affecting self-report metrics of both confidence and symptoms150,151. A common feature of a range of neurological disorders is lack of awareness of cognitive problems, known as anosognosia152 and often associated with impairments in prefrontal function. For example, around 75% of people with frontotemporal dementia show impaired awareness of illness153. Disparate forms of anosognosia can be considered failures to update a global metacognitive model of personal-level beliefs about one’s own abilities152,154. Objective measures of metacognitive sensitivity have been linked to levels of tauopathy (a major proteinopathy in Alzheimer disease) in older adults155, and to changes in prefrontal function in people with substance-use disorder156. These impairments in metacognition have functional consequences, such as a reluctance to seek help, take medication, or take steps to compensate for cognitive changes157. However, despite its clinical importance, metacognitive capacity is not yet captured by standard neuropsychological assessments. Functional disorders — cases in which people report subjective symptoms in the absence of deficit — perhaps reflect a pure expression of a change in metacognition, without any underlying measurable alteration in cognition or performance. However, few studies have used objective measures of metacognition in functional disorders158.

An important open question concerns the causal relationships between metacognitive disturbances and clinical symptoms. If metacognition is a root cause of mental ill health, then improving or modifying it may become a viable treatment target. Alternatively, if it is primarily an accompanying symptom, it may instead serve as a novel marker of underlying dysfunction. These questions could be addressed initially with longitudinal data and then with intervention studies.

previous findings indicating that unexpected physiological arousal compromises the quality of metacognitive judgements171 and suggesting that interoceptive uncertainty informs metacognitive evaluations

Interventions

Adaptive control

Metacognitive sensitivity

The prevalence of distortions in metacognition in clinical conditions, and the importance of effective metacognition in educational and social settings, has prompted growing interest in the potential for interventions. Initial work explored whether behavioural training may be effective for increasing metacognitive efficiency, with some evidence that people’s confidence reports could be shaped via feedback, although results were variable across studies159–162. Neurocomputational modelling is important to understand the locus of action of these effects. In one recent study, the fitted value of meta-uncertainty within the CASANDRE model was lower for (well-trained) monkeys than for human participants on their first visit to the lab58. However, humans were able to approach similar levels of metacognitive competence after further training with an additional 2,200 trials of the task. In clinically-focused studies, systematic protocols were developed for metacognitive training that seeks to provide a structured education about common cognitive biases and ‘plant the seeds of doubt’ in people who are prone to overconfidence and/or lack insight162. Metacognitive training has been shown to have consistent effects in reducing delusions and positive symptoms of psychosis, leading to it being recommended in national treatment guidelines in Australia, New Zealand and Germany. How such training may be altering the neural and computational components of confidence formation described above remains to be established. Other work has shown potential benefits of meditation163, neurofeedback164 and brain stimulation for modulating metacognition. Multi-region transcranial magnetic stimulation (TMS) has proven particularly illuminating in probing the causal contribution of particular brain areas. For instance, one study showed that although dlPFC stimulation reduced metacognitive bias (confidence level), FPC stimulation improved metacognitive efficiency47 (see also refs. 165,166). Another study revealed a double dissociation between the effects of TMS on decision-making and metacognition: neurostimulation aimed at boosting back projections from cortical regions V5/MT to V1/V2 enhanced motion sensitivity without impacting metacognition, whereas boosting intraparietal sulcus-to-V1/V2 back projections increased metacognitive efficiency without impacting motion sensitivity167. In a meta-analysis of 25 TMS and transcranial electrical stimulation studies of metacognition in humans, the most robust finding was a contribution of anterior and lateral portions of PFC to metacognition in both perceptual decision-making and memory tasks168. As for much exploratory work in the cognitive neurosciences, what is now needed is the combination of these promising methods with computational models that can develop precise hypotheses about the processes being targeted by particular types of stimulation. Finally, a small number of studies have explored pharmacological effects on metacognition and confidence. Blocking noradrenaline function improved metacognitive efficiency in a perceptual task169, whereas enhancing catecholaminergic neuromodulation changed perceptual sensitivity, but not metacognition170. There is also nascent work on the role of bodily states in multimodal metacognitive judgements, with

The ability to adjust thoughts,

How effectively confidence distinguishes

decisions or actions as circumstances

between correct and incorrect

change. Facilitated by metacognitive

judgements. Estimates are often

monitoring of current performance levels,

confounded with changes in first-order

which allows behaviour to be modified to

performance and metacognitive bias.

Glossary

pursue one’s goals.

Bayesian framework

Meta-uncertainty Imprecision in an observer’s own

An approach to perception,

estimate of the reliability of their

cognition and metacognition

perceptual or cognitive system.

in which the brain is treated as

Because this second-order estimate

representing beliefs as probability

is itself noisy, it imposes a ceiling on

distributions, and updating them by

metacognitive performance even when

combining prior expectations with

first-order sensitivity is intact. A central

incoming evidence weighted by

parameter of the CASANDRE model.

its reliability. Within this framework confidence is expressed as a probability,

Neurofeedback

and expectations about reliability as

A training technique in which brain activity

priors.

is measured in real time and displayed

BOLD signal

back to the participant, who is rewarded for producing a target activity pattern.

The blood-oxygen-level-dependent

It offers a route to directly modulating

signal measured by functional magnetic

specific neural signals, including those

resonance imaging. It reflects changes in

associated with confidence.

blood oxygenation and flow associated with nearby neural activity, providing an

Propositional confidence

indirect measure of brain activity.

Confidence in one’s own (possibly hypo­

Machine learning

thetical) decisions or actions, expressed as covert propositions such as “I think I will

A set of computational

remember this word.” Distinguished from

methods that identify patterns

uncertainty about states of the world.

in data and use them to make predictions or classifications. In

Psychometric functions

neuroscience they are typically

Curves describing how a behavioural

used as decoders: a classifier is

measure changes with the physical

trained to readout a variable — such

strength of a stimulus. Fitting

as a stimulus feature or its associated

them separately for choices

uncertainty — from patterns of brain

and for confidence ratings can

activity.

dissociate first-order sensitivity from

Metacognitive bias

metacognitive parameters.

A difference in subjective

Theory of mind

confidence despite objective

The capacity to attribute mental

task performance remaining

states such as beliefs, intentions and

constant.

knowledge to other agents, and to

Metacognitive efficiency

use those attributions to predict or explain their behaviour (also known as

The level of metacognitive sensitivity

mentalising). Partially overlapping brain

corrected for differences in first-order

networks support theory of mind and

performance.

self-directed metacognition.

across domains131. New experimental paradigms now allow for the direct investigation and modelling of metacognition about interoceptive states such as respiration and heart rate172,173. Given the potential clinical importance of metacognitive interventions, this is an area ripe for further investigation.

Future directions

Innovating new measures, metrics and models Advances in measurement have created new opportunities for charting metacognition across populations and task domains. However, the development of these measures and computational models has not been matched by equivalent investment in psychometric evaluations of associated model parameters174. In addition, popular frameworks, such as SDT-based and EA-based models, have led to a focus on retrospective confidence estimates in two-alternative forced-choice tasks, precluding the evolution of more naturalistic measures175,176. One approach to rectifying this will be to develop more flexible ways of relating task performance to metacognition, such as relative psychometric function analysis177, or computational approaches that cast metacognition as a problem of optimizing mutual information between behaviour and reflective judgements178. The idea that decision confidence should be framed as a probability of a first-order behaviour being objectively accurate or correct has been challenged on both empirical and theoretical grounds. Theoretically, it is often hard to define accuracy for subjective decisions, such as value-based choices or aesthetic judgements, and yet empirical studies show that people hold a meaningful sense of confidence in these choices29. Instead, more recent frameworks have suggested that the ground truth for metacognition should not be objective correctness but self-consistency — the probability of making the same choice across multiple presentations of the same decision problem56,179,180. Key to the self-consistency approach is that metacognitive estimates arise from an overall evaluation of the quality of the perception–action cycle, rather than a direct evaluation of the sensory input8. There is also the question of meta-metacognition: behavioural evidence shows that humans can engage in third or even fourth-order self-evaluations, tracking the reliability of their own confidence judgements181. It would be of interest to probe whether the same confidence formation motifs documented here could be applied to metacognitive judgements themselves, providing a powerful engine for recursive self-modelling182.

Integrative approaches Metacognitive neuroscience crosses species (from rodents to humans) and levels of analysis (from single neurons to behaviour), while retaining a strong connection to human subjective experience (Box 2). This blend positions metacognition research as potentially transformative for the discovery of conserved circuit-level motifs that underpin disturbances in conscious experience, such as those generating a sense of unreality in psychosis, or persistent underconfidence in anxiety disorders. However, achieving this goal will require greater integration between human and animal laboratories to develop cross-species behavioural assays and computational models to bridge the gap between neural circuits and experience (Fig. 5c). It also remains unknown which aspects of metacognition are shared across species. Rodents lack the granular lateral part of the PFC that is thought to be essential for some metacognitive functions in primates, particularly perhaps those supporting flexible task arbitration and control. However, both rodents and primates possess agranular frontal cortex and insula, which may

be important for implicit aspects of confidence and reality monitoring. Fully addressing this question requires consideration of how to integrate neural data across species, in which identifying homologues between circuits and brain regions is fraught. A goal for future research is to identify common building blocks of metacognition, while being alive to the potential that certain aspects, such as explicit self-knowledge, may be uniquely human. It will also be of interest to expand the study of the neural basis of metacognition beyond rodents and primates in order to probe minimal circuits that may support forms of metacognitive computation — for instance, the capacity to opt-out of difficult choices displayed by honeybees183. This endeavour will benefit from adopting performance-controlled assays of metacognition that have been developed in work on mammals. Even if some elements of metacognition may be restricted to primates or even humans, the tools provided by metacognition research offer a bridge from circuit-level analysis in animal models to subjective symptoms in humans and harbour considerable potential for an integrative neuroscience of psychopathology.

Conclusion Metacognition is no longer restricted to philosophical or psychological enquiry. Across species, convergent behavioural, computational and neural evidence indicates that self-evaluation arises from structured transformations of uncertainty into a format that can be used for adaptive control. Dynamic EA processes provide candidate building blocks for local confidence, whereas multimodal prefrontal circuits integrate these signals into more abstract, global self-beliefs. This architecture supports adaptive control at multiple timescales, from rapid error correction to strategic task arbitration and social communication. Important open questions remain. We lack precise psychometric characterization of many metacognitive metrics, and cross-species tasks are still only partially aligned. The field must also clarify which aspects of metacognition are evolutionarily conserved and which depend on uniquely human forms of self-modelling. Nonetheless, the integrative framework emerging from recent work suggests that metacognition research offers a tractable computational handle on subjective experience. By linking circuit-level mechanisms to distortions in confidence and self-belief, metacognitive neuroscience offers a principled route towards understanding — and potentially remediating — disorders of subjective reality.

Annotated key references (editorial notes as printed)

(Editorial annotations from the published reference list; each note introduces a key reference cited in the main text.)

  • This paper introduced metacognition as a distinct psychological construct, setting the agenda for research on its cognitive and neural basis. Flavell’s distinction between metacognitive knowledge and experience is a forerunner of global and local aspects of metacognition. (cognitive–developmental inquiry. Am. Psychol. 34, 906–911 (1979).)
  • This review sets out a normative computational framework for behavioural signatures of confidence, allowing metacognition to be quantified and studied using both explicit and implicit measures in humans and non-human animals. (humans and animals. Philos. Trans. R. Soc. B 367, 1322–1337 (2012).)
  • This foundational paper formulates a distinction between meta and object levels of cognition and articulates relationships between metacognitive monitoring and control — a framework that continues to organize the field. (Psychol. Learn. Motiv. Adv. Res. Theory 26, 125–173 (1990).)
  • This computational perspective distinguishes world-directed uncertainty from self-directed confidence, arguing that the two serve different behavioural goals. (quantities for different goals. Nat. Neurosci. 19, 366–374 (2016).)
  • This paper showed that monkeys will decline a difficult perceptual decision in favour of a guaranteed smaller reward, and that firing rates linked to EA in the lateral intraparietal cortex also predict the opt-out decision. (neurons in the parietal cortex. Science 324, 759–764 (2009).)
  • This tutorial review distinguishes between metacognitive bias, sensitivity and efficiency, and sets out why interpreting findings in the neuroscience of metacognition requires measures of sensitivity to be corrected for first-order performance. ((2014).)
  • This paper introduced meta-d’, a signal detection theoretic measure that expresses metacognitive sensitivity on the same scale as first-order sensitivity and which has become a standard tool for the field. ((2012).)
  • This paper pioneered single-neuron correlates of decision confidence, showing that firing in the rodent orbitofrontal cortex during odour categorisation follows the folded-X pattern predicted by statistical models. (behavioural impact of decision confidence. Nature 455, 227–231 (2008).)
  • This paper provided causal evidence that the rodent orbitofrontal cortex is required for metacognition, with inactivation abolishing confidence-dependent waiting for reward while leaving decision accuracy intact. (confidence. Neuron. 84, 190–201 (2014).)
  • This paper highlighted how different prefrontal subregions support different computational components of visual metacognition, with dorsolateral PFC TMS altering confidence level and anterior PFC TMS altering metacognitive efficiency. (visual metacognition. J. Neurosci. 38, 5078–5087 (2018).)
  • This paper introduced a Bayesian framework for metacognition, bringing together the literatures on error-monitoring and confidence and providing a taxonomy for first-order, serial, parallel and hierarchical architectures that can be used to organize computational accounts of confidence. (framework for metacognitive computation. Psychol. Rev. 124, 91–114 (2017).)
  • This paper showed that confidence signals in the rodent orbitofrontal cortex generalize across sensory modalities and behavioural readouts, providing evidence for a common currency. (representation of confidence in orbitofrontal cortex. Cell 182, 112–126.e18 (2020).)
  • This paper introduced the CASANDRE model, in which confidence reflects a noisy estimate of decision reliability and meta-uncertainty (uncertainty about one’s own cognitive and perceptual performance) acts as a limit on metacognitive performance. (reliability estimate. Nat. Hum. Behav. 7, 142–154 (2023).)
  • This paper used peri-decision wagering in monkeys to show that choice and confidence are deliberated concurrently rather than sequentially, with decodable signals for both emerging together in the lateral intraparietal cortex. (and confidence judgment. Nat. Neurosci. 29, 159–17 (2025).)
  • This integrative review unifies distinct literatures on performance monitoring, relating error monitoring to confidence formation through the lens of post-decisional processing. (post-decisional performance monitoring: An integrative review. eLife 10, e67556 (2021).)
  • This paper decoded trial-by-trial sensory uncertainty from human visual cortex and showed that reported confidence tracks this readout in PFC, linking variability in sensory uncertainty to the formation of propositional confidence. ((2022).)
  • This review proposes that metacognition operates over both local and global timescales, and harnesses this distinction to organise findings on metacognitive disturbance in a range of psychiatric conditions. (metacognition shape mental health. Biol. Psychiatry 90, 436–446 (2021).)
  • This paper compared humans and monkeys performing the same prospective confidence task, revealing that the two species track self-performance using different neural architectures within the PFC. ((2026).)
  • This paper provides a systematic assessment of the psychometric properties of current popular measures and computational models of metacognition. (metacognition. Nat. Commun. 16, 701 (2025).)
  • This review marshals evidence that metacognitive sensitivity is consciousness-selective, integrating theoretical perspectives with empirical data. (e1628 (2023).)
  • This review compares how humans and large language models represent and communicate uncertainty and sets out why the calibration of AI confidence matters for human–artificial intelligence collaboration. (and large language models. Curr. Dir. Psychol. Sci. 35, 131–139 (2025).)