Cortical Microstructural Variations Correlate with Individual

Differences in Gamified Exploration–Exploitation Behaviours

1 Ashley Tyrer, Niia Nikolova¹, Magda Dubois², Leah Banellis¹, Melina Vejlø¹, Tobias U.* 2,3† 1,4† Hauser, Micah Allen

Hauser 2,3† , Micah Allen 1,4†

1 Center of Functionally Integrative Neuroscience, Aarhus University, Denmark 2 Max Planck UCL Centre for Computational Psychiatry and Ageing Research, University College London, United Kingdom 3 Department for Psychiatry and Psychotherapy, University of Tübingen, Germany 4 Cambridge Psychiatry, University of Cambridge, United Kingdom

Abstract

The exploration-exploitation trade-off is ubiquitous in our everyday lives, and individuals display considerable variability in their preferred decision-making strategies. Most previous work pertaining to neural signatures of exploration is restricted to functional pathways. However, the specific contributions of cortical microarchitectures to high-level cognitive processes such as decision-making are as yet unknown. Here, we investigated the neuroanatomical foundations of inter-individual variability in decision-making strategies. To this end, 122 healthy participants completed a gamified multi-armed bandit paradigm aimed at teasing apart distinct exploration-exploitation decision strategies. We also collected whole-brain quantitative MRI maps indexing microstructural features of cortical myelination and iron content. Through computational modelling, we disentangled individual-specific exploration strategies, including value-free random exploration. Whole-brain regression analyses identified significant associations between value-free exploration and increased cortical myelination in right frontal brain areas with reported links to impulsivity. By elucidating the brain microstructural correlates of distinct exploration-exploitation strategies, we aimed to further our understanding of why individuals differ in their decision-making capabilities, and how decision-making may become aberrant in mental health conditions.


bioRxiv preprint doi: https://doi.org/10.1101/2025.10.08.681181; this version posted July 9, 2026. The copyright holder for this preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made available under a CC-BY-ND 4.0 International license.

Introduction

It has been long established that individuals demonstrate considerable variability in the ways they make decisions and process information pertaining to rewards (Gershman, 2018; Schulz & Gershman, 2019). However, the neurobiological basis of these differences is not currently well characterised. A plausible explanation may be that inter-individual differences in cortical microstructures are driving this behavioural variability (Ziegler et al., 2019). Cortical microstructures such as myeloarchitecture play a vital role in the efficiency of neural signalling, by improving conduction velocity and reducing signal-to-noise ratio (Nave, 2010; Purves et al., 2001). By investigating how individual-specific differences in the brain’s structural features relate to decision-making, we can gain insight into the latent causes of such wide-ranging behavioural differences. Through this approach, we may reveal why aberrant or suboptimal decision-making strategies are often observed in individuals suffering from mental health conditions such as ADHD or depression (Must et al., 2013; Schulze et al., 2021).

Inter-individual differences in learning and decision-making become apparent when examining decision-making dilemmas such as the exploration-exploitation trade-off. In such situations, one must choose either the option with the greatest known value (exploitation), or an alternative, lesser-known option which may potentially yield yet greater rewards but at the risk of disappointing outcomes (exploration). The methods by which agents select different decision-making strategies is highly idiosyncratic and context-specific, and individuals vary greatly in their preferred decision strategies (Frank et al., 2009; Somerville et al., 2017; von Helversen et al., 2018).

Recent advances have led to the development of computational models that allow for the estimation of subject-specific parameters to describe key signatures of these exploration-exploitation behaviours (Gershman, 2018; Schwartenbeck et al., 2019). Previous work has identified the employment of exploration strategies that consider all available choices to be equally likely by disregarding all existing knowledge of the decision space, including any uncertainty and expectation estimates, thus reducing computational complexity (Daw et al., 2006; Wilson et al., 2014). This strategy is referred to as value-free random exploration. A recent study by Dubois et al. (2021) found clear evidence that humans employ a combination of complex resource-heavy approaches and simpler heuristic exploration strategies such as value-free random exploration, as well as uncertainty-led exploration and novelty exploration, to estimate an optimal decision method (Dubois et al., 2021). Recent work has also highlighted the neurochemical basis for decision-making, demonstrating that noradrenergic and dopaminergic activity play crucial roles in modulating exploration-exploitation strategy selection (Chakroun et al., 2019; Cremer et al., 2023; Dubois & Hauser, 2022). However, these neuromodulatory influences are expressed through distributed cortical circuits, suggesting that variability in the structural properties of such circuits controls how neurochemical signals are integrated and translated into behaviour. These findings highlight the potential for individual-specific neurobiological foundations of decision-making, yet further research is necessary to confirm this link between biology and behaviour.

Numerous previous studies have identified specific brain regions and functional connectivity pathways underlying key decision variables. In particular, the ventromedial prefrontal cortex (vmPFC) has been repeatedly implicated in choice probability, valuation, and subjective utility of available choices (Cockburn et al., 2022; Daw et al., 2006). The frontopolar cortex has also been highlighted in many studies as a locus of exploratory behaviours and choice switching (Boorman et al., 2009; Daw et al., 2006), and the posterior cingulate cortex (PCC) has been implicated in tracking value computations across diverse decision-making contexts (Clithero & Rangel, 2014). In addition to functional pathways, brain structural properties such as cortical thickness have been recently shown to contribute to decision-making processes (Filmer et al., 2023; Smid et al., 2023; Sunderaraman et al., 2022) and are subject to age-related changes (Liu et al., 2025). The behavioural variability we observe in decision-making contexts may therefore result from inter-individual differences in brain structure, which can arise from developmental or genetic sources, and in turn influence function (Frank et al., 2009). Examining such cortical differences through microstructural imaging can potentially reveal the underlying neuroanatomical origins of these observed differences in behaviour and brain anatomical pathways. More specifically, exploration-exploitation behaviours rely on the balance between flexible, stochastic policy control, often associated with prefrontal networks, and more stable, value-based action selection frequently attributed to parietal and sensorimotor regions (Daw et al., 2006; Klein-Flügge & Bestmann, 2012). Such processes are likely to rely heavily on the efficiency of neural signalling and underlying structural connectivity.

Quantitative MRI is a valuable method for the quantification and investigation of in vivo brain microstructures, providing a non-invasive means of assessing neuroanatomical features relevant to cognitive function (Callaghan et al., 2014; Weiskopf et al., 2015; Ziegler et al., 2019). These indices extend beyond conventional morphometric measures such as cortical volume or thickness to index variation in quantitative, non-arbitrary units with direct histological correlates. For example, previous work has exploited these recent advances in quantitative neuroimaging to correlate model parameters of respiratory interoception with in vivo indicators of cortical histology, revealing multiple neuroanatomical contributions to interoceptive sensitivity, precision, and metacognition (Nikolova et al., 2025). In particular, the R2* quantitative contrast, which indexes the iron content of cortical tissues, has been repeatedly associated with noradrenergic and dopaminergic activity due to the accumulation of the neuromelanin-iron complex in noradrenergic neurons in the locus coeruleus, and dopaminergic neurons in the substantia nigra (Riley et al., 2023; Zucca et al., 2017). Noradrenaline is broadly recognised as a key contributor to higher-level cognitive processes including decision-making and uncertainty processing (Hauser et al., 2017; Lawson et al., 2021; Yu & Dayan, 2005). Therefore, investigating this relationship between structural brain components and neurotransmitters with putative roles in decision-making may advance our understanding of why decision-making behaviours differ so greatly between individuals.

In this study, we aimed to examine the neuroanatomical foundations of inter-individual differences in exploration-exploitation behaviours. To this end, 122 participants completed a gamified exploration task, constructed as a multi-armed bandit, which enabled the quantification of subject-level variability in exploration-exploitation behaviours (Dubois et al., 2021; Dubois and Hauser, 2022). We applied a whole-brain, cluster-corrected multiple linear regression analysis on three distinct neuroanatomical maps generated through quantitative MRI, relating cortical architectures with task parameter estimates characterising exploration-exploitation behaviours. Through associating individual-specific model parameter estimates denoting behavioural heuristics with cortical microstructures, we sought to elucidate the neurobiological bases of decision-making processes, and how they might drive inter-individual differences in learning.


bioRxiv preprint doi: https://doi.org/10.1101/2025.10.08.681181; this version posted July 9, 2026. The copyright holder for this preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made available under a CC-BY-ND 4.0 International license.

Materials and Methods

Participants

A total of 566 (360 females, 205 males, 1 other) participants (median age = 24, age range = 18 - 56) participated in the Visceral Mind Project, a large-scale neuroimaging project at the Center of Functionally Integrative Neuroscience, Aarhus University. Participants were recruited through local advertisements such as social media, posters, and flyers, and also via the SONA online participation pool system. Participants were required to have normal or corrected-to-normal vision and fluency in Danish or English. Exclusion criteria required that participants were not taking any medications excluding contraceptives or over-the-counter antihistamines. In addition, participants were required to be compatible with standard MRI scanning requirements (i.e., not pregnant/breastfeeding, no metal implants, no claustrophobia, etc.). Participants from this dataset completed multiple behavioural tasks, physiological recordings, and MRI scanning, in addition to psychiatric and lifestyle inventories, across three separate visits scheduled on different days. The MRI and behavioural data reported in this study were collected on different days. The study was granted ethical approval by the local Region Midtjylland Ethics Committee and was conducted in accordance with the Declaration of Helsinki (2013). All participants provided written informed consent and were remunerated for their participation.


[Image: Im0]

Figure 1 Participant inclusion and demographics. (A) Participant inclusion diagram, detailing data quality controls applied to both the behavioural and neuroimaging data. (B) Distributions of age and gender of the 122 participants included in the behavioural analyses. ( C) Maggie’s Farm (MF) task design. Participants could choose from three possible bandits (trees) to obtain samples (apples) of varying value (apple size), to maximise their reward. At the start of each trial, participants were shown initial samples in the crate at the bottom of the screen. Top right: trial horizon and number of remaining draws were indicated by the number of empty crates. Our analyses primarily focused on the initial draw for both horizon conditions.


Exploration-Exploitation Paradigm – ‘Maggie’s Farm’

‘Maggie’s Farm’ is a multi-armed bandit paradigm designed to examine human exploration-exploitation strategies in decision-making (Dubois et al., 2021; Dubois & Hauser, 2022) ( Figure 1C). On each trial, participants were asked to choose between three “bandits” (represented by trees bearing apples), each associated with a given reward distribution (apple size), to maximise their total reward. Each bandit’s reward (apple size) was drawn from a normal distribution with a fixed sampling variance. Bandits displayed either high prior reward information (i.e., three initial samples), or limited prior information (one initial sample), while one novel bandit provided no prior reward information. Bandits also had either a standard or a low reward mean. At the start of each trial, participants were shown initial reward samples displayed on a wooden crate at the bottom of the screen. They then selected a bandit to sample from. Participants were instructed to collect the biggest apples before the end of each trial (sunset). Participants could perform either one draw (short horizon) or six draws (long horizon) from the selected bandit, depending on the trial condition. By analysing participants’ choices, the task distinguishes between different exploration strategies, such as selecting novel options to reduce uncertainty or choosing randomly without regard for expected value. See Supplementary Information for detailed task instructions, and Dubois et al. (2021) for further information regarding details of the paradigm.

Statistical Analyses of Behaviour

To first assess normality in participants’ choice behaviour, behavioural heuristics, and model parameter estimates, we conducted Shapiro-Wilk normality tests. Due to non-normal distributions in the behavioural data, we proceeded with non-parametric statistical analyses for all comparisons and correlations. To examine whether trial horizon manipulation influenced exploratory behaviours (and as positive controls to demonstrate consistency with the findings of Dubois et al. 2021; Dubois and Hauser 2022), we compared the choice frequencies of different bandit values in the long versus short horizon trials using two-sided Wilcoxon signed-rank two-tailed tests. To identify potential correlations between behavioural metrics and model parameters, we conducted partial Spearman’s rank correlation analyses correcting for age, gender, and body mass index (BMI) and correcting for multiple comparisons using false discovery rate (FDR) correction (see Supplementary Figure 1 for partial correlation matrix). To investigate the association between participants’ choice behaviours and model heuristics, and mental health factors extracted from psychiatric surveys, we applied partial Spearman’s rank correlation analyses, correcting for age, gender, and BMI and correcting for multiple comparisons using FDR correction (see Supplementary Information, Supplementary Figures 2 and 3 for results, factor structure, and further details).

Computational Modelling of Behaviour

The computational model used here employs Thompson sampling (capturing uncertainty-driven value-based random exploration), with the addition of both value-free random exploration and novelty exploration following the framework described in Dubois et al. (see Dubois et al. 2021; Dubois and Hauser 2022, for further model details).

This previous validation work examining the same task in a different sample (Dubois et al., 2021) compared multiple computational models, consisting of three base models: the UCB model which includes the UCB algorithm and a softmax choice function; the Thompson model which employs Thompson sampling; and a hybrid model which combined the UCB model and the Thompson model. These computationally-demanding models were compared both with and without the addition of simpler heuristic exploration strategies, i.e., value-free random exploration (ε-greedy), and novelty exploration (novelty bonus η). Here, we employ the winning model from this model comparison, i.e., the model with the greatest held-out data likelihood (%): the Thompson model with the addition of ε-greedy and novelty bonus η.

To model value-free random exploration in line with these previous analyses, we added an ε-greedy component to the decision rule, ensuring that every ε% of the time, an option alternative to the one predicted is chosen. Also in line with previous work, we added a novelty bonus η to the computation of the value of novel bandits in order to model novelty exploration.

The value of each bandit, 𝑖, is defined as follows:

2 A sample 𝑥 ~ 𝑁(𝑉 σ) is taken from each bandit. The probability of choosing bandit 𝑖 𝑖,𝑡 𝑖,𝑡 𝑖,𝑡 depends on the probability that all pairwise differences between the sample from bandit 𝑖 and the other bandits 𝑗≠𝑖 were greater or equal to 0 (see Speekenbrink & Konstantinidis, (2015) for the probability of maximum utility choice rule). In Maggie’s Farm, two pairwise

x_{i,t}\!N\!\left(V_{i,t}\sigma_{i,t}^{2}\right) $$ differences scores (contained in the two-dimensional vector 𝑢) were computed for each bandit due to the presence of three bandits at a time. The probability of choosing bandit 𝑖 is described by:

P(c_{t}=,i)=,P(\forall j:,x{{i,t}}>,,;,x{{j,t}})^{*},,leftleft(1,-,,c{{v f}}\varepsilon)+,,c{_{v f}}\frac{\varepsilon}{3}

P\left(c_{t}=\ x{}{i}\right)=\iint{0}

\Phi

M _ {i, t} = A _ {i} \left(V _ {1, t} V _ {2, t} V _ {3, t}\right) \text {a n d c o v a r i a n c e m a t r i x}: C _ {i, t} = A _ {i} \left(\sigma_ {1, t} 0 0 0 \sigma_ {2, t} 0 0 0 \sigma_ {3, t}\right) A _ {i} ^ {T}

A_{i}

A _ {1} = (1 - 1 0 1 0 - 1)

We used the maximum a posteriori (MAP) probability estimate for fitting parameter values, employing the fmincon optimisation function in MATLAB, in line with previous work (see Dubois et al. 2021; Dubois and Hauser 2022, for further details regarding parameter estimation). ## Multi-Parameter Brain Mapping The multi-parameter pipeline employed here follows that described in Nikolova et al. (2025). We utilised a well-established qMRI protocol (Weiskopf et al., 2013, 2015) to map percent saturation resulting from magnetisation transfer (MT), longitudinal relaxation rate (R1) and effective transverse relaxation rate (R2*), followed by voxel-based quantification to examine how subject-specific parameter estimates of exploration-exploitation behaviours correlate with different characteristics of brain microstructure. ## Data Acquisition Neuroimaging data were acquired using a 3T MR system (Magnetom Prisma, Siemens healthcare, Erlangen, Germany), using a standard 32-channel radiofrequency (RF) head coil and a body coil. We obtained high-resolution whole brain T1-weighted anatomical images (0.8 mm³ isotropic) using an MP-RAGE sequence (repetition time = 2.2 s, echo time = 2.51 ms, matrix size = 256 x 256 x 192 voxels, flip angle = 8°, AP acquisition direction).

(0.8~\mathsf{m m}

21

6

220

30

115

65

5

## Map Creation All qMRI images were pre-processed using the hMRI toolbox v. 0.5.0 (January 2023) (Tabelow et al., 2019) and SMP12 (version 12.r7771, Wellcome Trust Centre for Neuroimaging, http://www.fil.ion.ucl.ac.uk/spm/), to correct the raw qMRI images for spatial transmit, receive field inhomogeneities and obtain quantitative MT, PD, R1 and R2* estimate maps. This correction was executed using empirical data, i.e., the RF sensitivity map and B1 maps collected from each participant. Excluding enabling imperfect spoiling correction, the hMRI toolbox was configured using default settings. All images were reoriented to MNI standard space prior to map creation. This processing produced four maps modelling different aspects of tissue microstructure: an MT map sensitive to myeloarchitectural integrity (Helms et al., 2008), a PD map representing tissue water content, an R1 map reflecting myelination, iron concentration and water content (primarily driven by myelination) (Lutti et al., 2014), and an R2* map sensitive to tissue iron concentration (Langkammer et al., 2010).

\mathsf{R}2

\mathsf{R}2

MT saturation and R1 both reflect qualities of myelin, however they index different underlying biophysical properties (Filo et al., 2019; Weiskopf et al., 2021). R1 is predominantly shaped by tissue composition, water content, and iron, whereas MT saturation indexes macromolecular content through magnetisation transfer. As a result, subtle variations in microstructural myelin components may exert a greater effect on one map type compared with the other. However, the in-depth understanding of the neuroarchitecture-to-MPM contrast association remains an active research area, and some biophysical constructs contribute to both map types. The unified segmentation approach (Ashburner & Friston, 2005) was used to segment MT saturation maps into grey matter (GM), white matter (WM) and cerebrospinal fluid (CSF) probability maps. We used tissue probability maps based on multi-parametric maps developed by Lorio et al. (2016), without bias field correction given that MT maps do not show significant bias field modulation. We then used the GM and WM probability maps to perform inter-subject registration using Diffeomorphic Image Registration (DARTEL), a nonlinear diffeomorphic algorithm (Ashburner, 2007). The MT, PD, R1 and R2* maps were then normalised to MNI space (at isotropic 1 mm resolution) using the resulting DARTEL template and participant-specific deformation fields. The nonlinear registration of the quantitative maps was based on the MT maps due to their high contrast in subcortical structures, and a WM-GM contrast in the cortex similar to T1 weighted images (Helms et al., 2009). Finally, tissue-weighted smoothing was applied using a kernel of 8 mm full width at half maximum (FWHM) using the voxel-based quantification (VBQ) approach (Draganski et al., 2011). Importantly, in contrast to voxel-based morphometry analysis, this VBQ smoothing approach aims to minimise partial volume effects and optimally preserves the quantitative values of the original qMRI images by not modulating the parameter maps to account for volume changes. The resulting GM segments for each map were used for all statistical analyses. ## MRI Quality Control and Participant Exclusions Multi-parameter mapping contrast images were acquired for 503 total participants. Following inspection, several participants were removed from all analyses for reasons related to either MRI or behavioural data. Three participants were excluded immediately following MR data collection for medical reasons (one cerebral palsy and two other suspected brain abnormalities). Three participants were excluded due to errors in the scanning sequences. As MPM data is known to be particularly sensitive to motion artefacts, we conducted a thorough quality control (QC) assessment to identify high-motion images. Visual QC HTML reports were created of each participant using the hMRI-vQC toolbox (Sherif et al., 2022), and all reports were visually inspected and labelled by two researchers. Doubtful cases were discussed, and a further 57 participants were excluded due to excessive motion affecting the tissue-class segmentation. Data from the remaining 443 participants was used in the spatial analysis and template creation using DARTEL. A total of 196 participated in the Maggie’s Farm task, 122 of which had complete datasets for the task (median age: 24, age range: 18-52, 67 females, 54 males, 1 other). Participants with missing or incomplete behavioural data were excluded from these analyses. The overlap between the behavioural paradigm and MPM data left 111 participants for VBQ analyses. Following additional QC using the CAT12 toolbox in SPM12, a further four participants were excluded from VBQ analyses for MT maps (MT: final N = 107), four for R1 maps (R1: final N = 107), and three for R2* maps (R2*: final N = 108) (**Figure 1**). ## Voxel Based Quantification Analysis Grey and white matter masks were generated based on our samples, by averaging the smoothed, modulated GM and WM segment images, and thresholding the result at p > .2. Inter-subject variation in MT, R1 and R2* GM maps were modelled in separate multiple linear regression analyses. A total of 16 regressors of interest were used in the VBQ analysis, consisting of behavioural heuristics and subject-specific model parameter estimates generated through Thompson sampling modelling (see **Table 2** for a full list of regressors included). Additionally, we included age, gender, body mass index (BMI) and total intracranial volume (TIV) as nuisance covariates in all analyses, following recommended procedures for computational neuroanatomy (Ridgway et al., 2008). VBQ data is often subject to non-stationarity, therefore we applied Threshold-Free Cluster Enhancement (TFCE) correction to our contrasts in combination with a GM mask, which provides non-parametric statistics and is robust to such non-stationarity in the data. We analysed whole-brain maps of each positive and negative t-contrast using a TFCE-corrected FWE-cluster p-value with p < .05 inclusion threshold (Hupé, 2015; Ridgway et al., 2008). All statistical analyses were conducted in SPM12, and we utilised the JuBrain Anatomy Toolbox v. 3.0 (Eickhoff et al., 2005) to determine anatomical labels and regional percentages. --- ## Results ## Exploration-Exploitation Behaviours are Modulated by Trial Horizon To evaluate participants’ chosen exploration-exploitation strategies, we first analysed participants’ choice behaviour in the multi-armed bandit task, Maggie’s Farm, and estimated subject-specific behavioural heuristics through Thompson sampling modelling (**Figure 1C**). Participants chose to sample less from the high-value bandit in long horizon trials compared with short horizon trials (two-sided Wilcoxon signed-rank two-tailed test: V = 336.5, p < .001), demonstrating that participants were willing to forego selecting the bandit with the greatest expected reward outcome to potentially learn about alternative bandits (**Figure 2A**). Consequently, participants chose to sample more from the novel bandit (V = 514.0, p < .001) and from the low-value bandit (V = 1076.5, p < .001) in long horizon trials versus short horizon trials (**Figure 2A**). This demonstrates that participants chose to engage in more exploratory behaviours in the long horizon trials, where exploration is beneficial, compared with short horizon trials, where exploration is less beneficial. This behaviour is also consistent with that presented in previous studies examining choice behaviours in the same task (Dubois et al., 2021; Dubois & Hauser, 2022), which found that manipulating trial horizon drove exploratory behaviours.

(V=514.0,,p<<.001)

\sigma_{0}

p!<!.01

--- ## Participants utilise value-free exploration when exploration may be beneficial To formally quantify how participants employ different exploration strategies, we examined differences in behavioural model parameter estimates between long and short horizon trials. The ε-greedy parameter indexes value-free random exploration, in that ε% of the time each available bandit has an equal probability of being selected by the participant. Participants had higher values of ε-greedy, i.e., greater value-free random exploration, in the long horizon versus short horizon trials (two-sided Wilcoxon signed-rank two-tailed test: V = 1167.0, p < .001) ( **Figure 2C**), suggesting that participants employed value-free explorative strategies in trials where exploration was beneficial. Participants also had significantly greater novelty bonuses, i.e., they showed greater novelty exploration in long horizon trials versus short horizon (V = 643.0, p < .001) (Figure 2C). Additionally, participants exhibited higher prior variances, i.e., uncertainty-driven value-based random exploration, or greater uncertainty about a bandit’s mean before seeing any samples, in long horizon trials versus short horizon trials (V = 2624.0, p = 0.00398) (**Figure 2C**), further bolstering the suggestion that exploration strategies were promoted during long horizon trials, which allowed the participant to subsequently exploit the information gathered to obtain greater rewards overall. These findings also remain consistent with those in Dubois et al. (2021, 2022), further demonstrating that long horizon trials engendered robust value-free exploratory behaviours.

(V=643.0,,p<<.001)

## Value-free random exploration leads to greater rewards in the long run To establish whether this choice behaviour yielded greater rewards in the long run, we examined reward earned in the initial samples of short horizon trials, initial samples in long horizon trials, and average reward earned across all six samples in the long horizon. As expected, participants earned fewer rewards in the first sample of long horizon trials versus short horizon trials due to the horizon-specific choice behaviour described above (FDR-corrected two-sided Wilcoxon signed-rank two-tailed test: V = 1015.5, p < .001) ( **Figure 2D**). However, participants earned higher rewards on average across all samples in the long horizon compared to the short horizon initial sample (V = 10.0, p < .001), suggesting that participants initially selected less optimal bandits to gather information about lesser-known bandits, then applied this newly acquired information to select more optimal bandits in the future and maximise their rewards. --- ## Voxel-Based Quantification reveals neuroanatomical correlates of value-free exploration We next investigated how these behavioural metrics and model parameter estimates of value-free random exploration in our exploration-exploitation paradigm relates to brain microstructural indices. To achieve this, we employed whole-brain multiple linear regression with TFCE correction against participants’ ε-greedy parameter estimates in short horizon trials and long horizon trials separately, while adjusting for age, gender, BMI, and total intracranial volume (TIV). We found that R1 map values in the right superior frontal gyrus (TFCE corrected: k = 15138, pFWE<sub>corr</sub> = .013, peak voxel coordinates: x = 8.0, y = -11.2, z = 77.6) and right middle frontal gyrus (TFCE corrected: k = 3467, pFWE<sub>corr</sub> = .021, peak voxel coordinates: x = 36.0, y = 12.8, z = 55.2) were significantly positively correlated with ε-greedy parameter estimates specifically in long horizon trials (**Figure 3A-C**). In contrast, we found no significant correlations between R1 map values and ε-greedy values in short horizon trials. This finding specifically in long horizon but not short horizon trials therefore indicates a significant positive association between value-free random exploration in trials where exploration is more beneficial, and myelination in right frontal gyri, which have been previously linked to the control of impulsive behaviours (Hu et al., 2016).

k!=!15138

p F\mathsf{M M E}_{\mathsf{c o r r}}=.013

k=3467

x=36.0,,y=

k=5840

k=3883,D F W E_{c O r r}=0.36

x=24.8,y=-46.4,2=47.21

undefined

24.9\pm5.0

23.1\pm2.9

--- **Table 2: Behavioural model parameter glossary** | Parameter | Definition | | --- | --- | | Xi SH | Epsilon greedy parameter for short horizon trials | | Xi LH | Epsilon greedy parameter for long horizon trials | | ηSH | Novelty bonus for short horizon trials | | ηLH | Novelty bonus for long horizon trials | | $\sigma_{0}$SH | Prior variance for short horizon trials | | $\sigma_{0}$LH | Prior variance for long horizon trials | | Q0 | Expected value before seeing any samples | | Score SH | Score in the short horizon (i.e., the size of the single apple picked) | | Score all LH trials | Total score in the long horizon (i.e., the sum of all the apples picked) | | Score first LH trial | Score on the first trial in the long horizon (i.e., the size of the first apple picked) | | Behaviour EV SH | Expected value of chosen bandit (i.e., exploitation), in the short horizon | | Behaviour EV LH | Expected value of chosen bandit (i.e., exploitation), in the long horizon | | Behaviour IS SH | Number of samples that were displayed for the bandit that they chose (i.e., information seeking), in the short horizon | | Behaviour IS LH | Number of samples that were displayed for the bandit that they chose (i.e., information seeking), in the long horizon | | Behaviour Consist SH | Consistency measure, i.e., whether participants made the same choice when observing the same data, in the short horizon | | Behaviour Consist LH | Consistency measure, i.e., whether participants made the same choice when observing the same data, in the long horizon |

\sigma_{0},H H

{\mathfrak{Q}}_{0}

undefined

{\bf Q}_{0},

{\mathfrak{Q}}_{0}

{\mathfrak{Q}}_{0}

--- (negative loading) (**Supplementary Figure 3**). Together, these findings indicate that participants who experience greater social anxiety and reduced confidence in oneself may place inflated value expectations upon unknown options. ## Supplementary Information 2: Detailed Task Instructions Prior to starting the task, participants underwent a short task tutorial with the following instructions, alongside visual examples of bandits (trees bearing apples) and rewards (apples): “Apples come in different shades and sizes. You need to help us collect the BIGGEST ones before sunset. You can only pick apples until sunset. Sometimes you will start at noon and will pick six apples, sometimes you will start in the afternoon and will pick only one apple. On each day you will collect apples from new trees. Some of the trees are better than others. To help you, some apples were already picked up before you arrived. These are the three different types of trees. Each tree produces an apple with a corresponding colour, but it’s their size that matters! For example, if we look at the apples produced, the yellow tree seems to be better on this day because it produces bigger apples ON AVERAGE. This is a big apple from the yellow tree. This is a medium apple from the red tree. This is a small apple from the yellow tree. This is a small apple from the blue tree. And finally this is a big apple from the blue tree. You will press the key 1, 2, or 3 to select the tree you wish to pick an apple from. If the sun is here, it is still early and you will be able to pick six apples. If the sun is almost set, you will only be able to pick one apple. At the end of each day, you will see the amount of juice extracted from the apples you collected. The juice will be poured in a small glass if one apple was picked, and in a big glass if six apples were picked.” --- ## Supplementary Figure 2 **SF2** Multi-level factor solution derived from the responses in eight mental health surveys of [Image: Im0] 566 participants in the Visceral Mind Project (Banellis et al. 2025). Outer circle denotes the two higher-level factors: affective symptoms and ADHD & somatic. Inner circle denotes the lower-level factors: social anxiety, self-confidence, sleep, depression, stress, autism, negative thoughts, restlessness, somatic symptoms, and impulsivity. The eleventh lower-level factor, inattentiveness, was not encompassed within either of the two higher-level factors and was not examined here. Numbers denote factor loadings for each lower-level factor. For more details regarding the exploratory factor analysis (EFA) methods, see Banellis et al. (2025). --- bioRxiv preprint doi: https://doi.org/10.1101/2025.10.08.681181; this version posted July 9, 2026. The copyright holder for this preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made available under a CC-BY-ND 4.0 International license. ## Supplementary Figure 3 **SF3** Left: The exploratory factor analysis (EFA)-derived ‘Affective Symptoms’ higher-order factor was significantly correlated with prior mean Q₀, i.e., participants’ prior beliefs about a bandit's mean value before seeing any samples (r(107) = 0.32, punc < .001, pcor = .00667; FDR-corrected partial correlations correcting for age, gender, and BMI). Right: There were no significant correlations between either of the EFA higher-order factors and participants’ behaviour (i.e., the frequency of picking the low-value, novel, or high-value bandits). Bottom: FDR-corrected post-hoc partial correlations correcting for age, gender, and BMI revealed that the association between the affective symptoms higher-order factor and prior mean Q₀ is driven by significant positive correlation with the social anxiety factor (r(107) = 0.23, p<sub>unc</sub> = .0105, p<sub>cor</sub> = .0315) and significant negative correlation with the self confidence factor (r(107) = -0.28, p<sub>unc</sub> = .00248, p<sub>cor</sub> = .0149). For significant uncorrected correlations, p values are printed on the heatmaps, and p values for significant corrected correlations are indicated in bold with asterisks (*p < .05, **p < .01, ***p < .001).

(\mathtt{r}(\mathtt{107}),=,\mathtt{0.32},;\rho_{u n e},<,\mathtt{001},;\rho_{c o r},=,\mathtt{.00667};

{\mathfrak{Q}}_{0}

(1117)=0.23,,\rho_{u n c}=

p_{c o r}=.0315)

p_{u n c}=.00248

=-0.28

p_{c o r}=.0149,

(^{\star}p<0.05,^{\star}p<0.1,

--- **Supplementary Figure 4** | R1 Long horizon trials | $\varepsilon$-greedy | Picked low-value bandit | | --- | --- | --- | | R1 Long horizon trials | 2820 2880 2940 3000 3060 | 0 20 40 60 80 | | MT Short horizon trials | 1800 2000 2200 2400 2600 | 0 300 600 900 1200 | **SF4** Comparison of main VBQ findings examining associations between cortical microstructures and ε-greedy parameter estimates (left), versus the simpler behavioural metric of bandit selection frequency (right). --- ## Supplementary Figure 5 | R1Long horizon trials | Original Model | Reduced Model | | --- | --- | --- | | R1Long horizon trials | 2820 2880 2940 3000 3060 | 0 200 400 600 800 | | MTShort horizon trials | 1800 2000 2200 2400 2600 | 0 300 600 900 1200 | **SF5** Comparison of main VBQ findings for the original model comprising 16 regressors of interest and four nuisance covariates (age, gender, BMI, and TIV), and a reduced model comprising ε-greedy for short and long horizon trials only, plus the same four nuisance covariates. --- |bioRxiv preprint|doi: [https://doi.org/10.1101/2025.10.08.681181|](https://doi.org/10.1101/2025.10.08.681181|); this version posted July 9, 2026.|The copyright holder for this preprint| |---|---|---|---| |(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is||CC-BY-ND 4.0 International license|.| made available under a ## Supplementary Figure 6 **SF6** Meta-analytic analyses for right SFG and right PostCG seed regions identified through VBQ analyses, displaying areas of significant coactivation. SFG = superior frontal gyrus; PostCG = postcentral gyrus. --- bioRxiv preprint doi: https://doi.org/10.1101/2025.10.08.681181; this version posted July 9, 2026. The copyright holder for this preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made available under a CC-BY-ND 4.0 International license. ## Supplementary Tables ## Supplementary Table 1 | R1 Whole Brain Clusters for Raw Behavioural Data | | | | | | | | | --- | --- | --- | --- | --- | --- | --- | --- | | Region | k | p(FWE-corr) | p(uncorr.) | TFCE | x | y | z | | Left lateral ventricle | 1 | 1.00 | 0.049 | 132.53 | -24.0 | -52.8 | 13.6 | | Left lateral ventricle | 1 | 1.00 | 0.049 | 132.04 | -23.2 | -51.2 | 12.8 | | Right lateral ventricle | 19 | 1.00 | 0.043 | 123.99 | 34.4 | -42.4 | 3.2 | | Right lateral ventricle | 82 | 1.00 | 0.030 | 81.92 | 23.2 | -30.4 | 24 | **Supplementary Table 1:** Summary of whole brain VBQ results: positive correlation between R1 map values and low-value bandit selection frequency in long horizon trials. Clusters are ordered by FWE-corrected p value. --- ## Supplementary Table 2 | MT Whole Brain Clusters for Raw Behavioural Data | | | | | | | | | --- | --- | --- | --- | --- | --- | --- | --- | | Region | k | p(FWE-corr) | p(uncorr.) | TFCE | x | y | z | | Left putamen | 3949 | 0.200 | 0.015 | 1207.01 | -33.6 | -6.4 | 2.4 | | Left pallidum | 657 | 0.208 | 0.018 | 1177.95 | -18.4 | -10.4 | -2.4 | | Right basal forebrain | 3433 | 0.211 | 0.016 | 1167.07 | 20.8 | 2.4 | -15.2 | | Right frontal operculum | 1771 | 0.241 | 0.011 | 1073.67 | 40.8 | 16.8 | 10.4 | | Right posterior orbital gyrus | 465 | 0.246 | 0.020 | 1057.30 | 24.0 | 16.0 | -14.4 | | Left parahippocampal gyrus | 3517 | 0.268 | 0.013 | 997.49 | -12.8 | -7.2 | -29.6 | | Left ventral DC | 549 | 0.295 | 0.019 | 932.45 | -12.8 | -20.0 | -18.4 | | Left parahippocampal gyrus | 1 | 0.313 | 0.019 | 890.73 | -20.0 | 30.4 | -28.8 | | Left parahippocampal gyrus | 10 | 0.316 | 0.019 | 884.47 | -24.0 | -26.4 | -29.6 | | Right OrIFG | 1149 | 0.335 | 0.017 | 846.34 | 48.0 | 29.6 | -10.4 | **Supplementary Table 2:** Summary of top ten whole brain VBQ results: negative correlation between MT Saturation and low-value bandit selection frequency in short horizon trials. Clusters are ordered by FWE-corrected p value. DC = diencephalon; OrIFG = orbital part of inferior frontal gyrus. --- bioRxiv preprint doi: https://doi.org/10.1101/2025.10.08.681181; this version posted July 9, 2026. The copyright holder for this preprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made available under a CC-BY-ND 4.0 International license. **Supplementary Table 3** | Name | z-score | Posterior prob. | Func. conn.(r) | Meta-analytic coact.(r) | | --- | --- | --- | --- | --- | | foot | 8.35 | 0.95 | 0.48 | 0.19 | | muscle | 6.02 | 0.92 | 0.23 | 0.13 | | suppressed | 5.95 | 0.92 | 0.01 | 0 | | substantia | 5.13 | 0.92 | -0.05 | -0.01 | | motor premotor | 4.82 | 0.91 | 0.28 | 0.12 | | distractor | 4.75 | 0.91 | -0.01 | -0.01 | | electrical | 5.12 | 0.9 | 0.19 | 0.16 | | cortex dorsal | 4.58 | 0.9 | 0.02 | 0.01 | | midbrain | 4.9 | 0.88 | -0.06 | 0 | | coordination | 4.55 | 0.88 | 0.27 | 0.22 | **Supplementary Table 3:** Top ten meta-analysis map associations for the right SFG, ordered by posterior probability. ## Supplementary Table 4 | Name | z-score | Posterior prob. | Func. conn.(r) | Meta-analytic coact.(r) | | --- | --- | --- | --- | --- | | postcentral gyrus | 6.54 | 0.89 | 0.14 | 0.2 | | pitch | 5.57 | 0.89 | 0.19 | 0.26 | | speech production | 5.4 | 0.89 | 0.35 | 0.44 | | taste | 5.04 | 0.89 | 0.02 | 0.09 | | vocal | 4.78 | 0.89 | 0.24 | 0.31 | | postcentral | 5.49 | 0.85 | 0.1 | 0.26 | | multisensory | 4.38 | 0.85 | 0.07 | 0.13 | **Supplementary Table 4:** Meta-analysis map associations for the right postcentral gyrus, ordered by posterior probability.