Non-Verbal Cognitive Tasks for Autism Spectrum Disorder
ASD deep-dive companion to nonverbal-cognitive-tasks: which language-free cognitive tasks best distinguish autism in humans, and how the same task families translate to mouse and monkey models — in particular SHANK3 models of Phelan–McDermid syndrome (PMS). Batteries covered here: cantab (CANTAB-ASD), the NIH Toolbox Cognition Battery, and the ABC-CT biomarker battery. Animal test platforms: automated-cognitive-testing-devices (Monkey CANTAB, Bussey-Saksida, CageLab and neighbours).
Overview
ASD is defined behaviourally — social-communication differences plus restricted/repetitive behaviour — and it has no single cognitive signature. Non-verbal (language-free) tasks are used for three reasons: they remove the language confound (many autistic people are minimally verbal, and toddlers cannot be tested verbally); the same task logic can be implemented in rodents and non-human primates (NHPs); and they yield objective, quantitative endpoints for intervention trials. The best-validated group-level signals are social-attention eye-tracking and modest, domain-general executive-function differences.123
Non-verbal task families with ASD evidence and cross-species reach:
| Domain | Best human tasks | Cross-species analogs |
|---|---|---|
| Visuospatial working memory | [[spatial-working-memory | CANTAB SWM]] |
| Cognitive flexibility (set-shifting, reversal) | [[intradimensional-extradimensional-shift | CANTAB IED]]; [[reversal-learning |
| Associative / episodic memory | CANTAB PAL; NIH-TCB Picture Sequence Memory | paired-association learning in macaques 4 |
| Social attention & gaze | eye-tracking batteries (ABC-CT OMI; social visual engagement) | eye contact in SHANK3 monkeys 5; gaze in VPA marmosets 6 — no direct rodent analog |
| Inhibitory control, attention, speed | NIH-TCB Flanker, Pattern Comparison | not yet standard in ASD models — rodent test harmonisation is still being called for 7 |
Why ASD resists a simple cognitive signature
Three lines of evidence set the ceiling on “tasks that distinguish ASD”:
- Meta-analytic EF differences are real but moderate and unspecific. Demetriou et al. (2018) pooled 235 studies (14,081 participants) and found g = 0.48 overall, similar across EF subdomains, with no support for fractionation into distinct subdomain deficits — and “only a small number of EF measures achieved clinical sensitivity”.1
- Diagnostic specificity is the weak point. Kofler et al. (2024) conclude that most studies fail to account for ASD–ADHD co-occurrence, that traditional neuropsychological tests and rating scales have poor specificity and construct validity, and that the most parsimonious reading of the field is that children with ADHD and/or ASD perform moderately worse than neurotypical children on a broad range of tests, with unique-vs-shared subprofiles still largely unknown.2
- Group effects have shrunk over time. Across 11 meta-analyses of autism–control comparisons, effect sizes for 7 distinct differences decreased over publication year, 5 of them significantly — consistent with broadened (more heterogeneous) diagnostic boundaries rather than simply better measurement.3
Human task evidence, by domain
Visuospatial working memory
The classic finding: high-functioning individuals with autism made more errors on the CANTAB SWM self-ordered search and were less likely to use a consistent search strategy, with deficits emerging when memory load exceeded a limited capacity.8 Cambridge Cognition highlights SWM as “highly sensitive to cognitive dysfunction in autism” in its ASD battery.9 Meta-analytic summaries put autistic WM differences at moderate size (school-age children larger than adolescents), with visuospatial tasks yielding larger effects than verbal ones.2
Cognitive flexibility: set-shifting and reversal
- An early landmark, Ozonoff et al. (2004), found deficits in planning efficiency (Stockings of Cambridge) and extradimensional set-shifting (IED) in 79 autistic vs 70 matched controls across ages 6–47.10
- The 2024 meta-analysis of cognitive flexibility (59 studies; 2,122 autistic participants without intellectual disability) found a small-to-moderate overall effect, with perseverative errors producing the largest effects, and substantial heterogeneity across task/age/outcome choices.11
- Probabilistic reversal learning (PRL; reversal-learning) is the sharpest human probe: ASD participants performed normally during acquisition but at reversal made more regressive errors — reverting to the previously preferred response — and regressive-error counts correlated with independently rated restricted/repetitive behaviours, implicating frontostriatal flexibility circuitry.12 PRL has also passed feasibility, test–retest, and intervention-sensitivity checks as a cognitive-flexibility outcome measure in a within-subjects ASD trial.13
- Nuance: the NIH-Toolbox DCCS (a set-shifting task) did not differentiate autistic children after age/IQ control,14 while in adolescents/young adults the flexibility composite did.15 Task demand and age matter.
Inhibitory control, attention, and processing speed
In autistic children aged 3–17 the NIH-TCB Flanker (inhibitory control/attention) and Pattern Comparison (processing speed) showed small but significant diagnosis-related differences after controlling age and IQ; picture-sequence memory and DCCS did not.14 In an IQ-matched adolescent/young-adult cohort (12–22), ASD showed poorer inhibitory control, cognitive flexibility, episodic memory, and processing speed — all Fluid composite components — while vocabulary and word reading (Crystallized) were spared: a fluid-over-crystallized profile that persisted across ages, with latent-profile analysis splitting autistic participants into a broad-impairment group and a crystallized-stronger group.15
Construct note: these measures index interference control and prepotent-response inhibition — meta-analysis of 41 inhibition studies found a moderate ASD effect on prepotent response inhibition (ES 0.55, moderated by age) and a smaller effect on interference control (ES 0.31), with substantial heterogeneity,16 while the narrower action-cancellation construct measured by the stop-signal task (stop-signal-task) shows little SSRT difference in ASD.17
Social attention, face processing, and emotion
This is where non-verbal tasks achieve their strongest discrimination:
- ABC-CT eye-tracking battery (280 autistic, 119 TD children, 6–11; Activity Monitoring, Social Interactive, Static Scenes, Biological Motion, Pupillary Light Reflex). Gaze-to-faces carried the signal: Activity-Monitoring %Face d = −1.04; the aggregated Oculomotor Index (OMI) d = −0.79; six-week stability (ICC) mostly moderate-to-high. The consortium explicitly framed these as group-asymmetry markers for mechanistic trials, “not … individual diagnostic precision”.18
- Social visual engagement vs expert diagnosis (475 children, 16–30 months): a single eye-tracking index of preferential social looking reached 71.0% sensitivity / 80.7% specificity, rising to 78.0%/85.4% in children with certain diagnoses — the strongest individual-level discriminative performance of any non-verbal battery to date.19
- GeoPref Test (1,863 toddlers; 1-minute social-vs-geometric movie): 98% specificity but only 17% sensitivity (33% with saccades) at a 69% fixation threshold — i.e. a subtype marker of an ASD segment, highly heritable (twin concordance).20
- Meta-analysis 2025 (57 studies, ≤36 months): pooled social-fixation difference g ≈ 0.65; machine-learning pipelines reached sensitivity up to 89% and specificity up to 86%; heterogeneity and protocol standardisation remain the blockers to clinical translation.21
- The Selective Social Attention task (ABC-CT-adjacent) replicated toddler effects in 4–12-year-olds: reduced face looking during socially engaging conditions, with %Face correlating with SRS symptom scores.22
- Emotion recognition (CANTAB ERT is inside the vendor’s ASD battery): meta-analysis shows non-selective impairment across all basic emotions, relatively specific to ASD vs other clinical groups but not to emotion-specific or face-specific processing — a softer discriminator than social attention.239
Associative learning and episodic memory
CANTAB PAL (paired associates learning) is a member of the vendor’s ASD battery (episodic-memory endpoint).9 On the NIH-TCB, picture-sequence (episodic) memory differentiated IQ-matched autistic adolescents,15 but not children in the feasibility study.14 The strongest cross-species motif here is the paired-association learning deficit in SHANK3-monkey models (see below),4 which maps onto PAL-like tasks in all three species.
Batteries
CANTAB (and the CANTAB-ASD battery)
The vendor’s ASD battery packages: Emotion Recognition Task (ERT), Match to Sample Visual Search (MTS), Multitasking Test (MTT), One Touch Stockings of Cambridge (OTS), Paired Associates Learning (PAL), Reaction Time (RTI), Spatial Working Memory (SWM) — covering executive function, episodic memory, planning, and processing speed, with the claim that the tests discriminate ASD from comorbid ADHD and depression and are sensitive across severity and age.9 Underlying literature: OTS + IED deficits (Ozonoff 2004),10 SWM deficits load-dependent (Steele 2007),8 and ASD–ADHD differentiation by EF profile — interestingly, in one carefully matched study ADHD (not ASD) was the group with reduced CANTAB SWM, while ASD was distinguished by WISC perceptual-reasoning relative weakness, illustrating that EF tasks alone rarely separate diagnoses cleanly.24 The CANTAB lineage is the most cross-species-extended test family in existence: human CANTAB → Monkey CANTAB IntelliStation → Bussey-Saksida rodent touchscreens (see automated-cognitive-testing-devices, lafayette-instrument, campden-instruments).
NIH Toolbox Cognition Battery (NIHTB-CB)
A standardised iPad battery; Fluid composite = DCCS, Flanker, Picture Sequence Memory, List Sorting, Pattern Comparison; Crystallized composite = Picture Vocabulary + Oral Reading Recognition (years 7+),25 plus early-childhood versions for ages 3–6.25 In ASD:
- Feasibility (116 autistic vs 80 TD, ages 3–17): 57% of autistic children completed all four tested tasks vs 88% of TD; those with IQ ≤ 70 completed on average 2.4 of 4 tasks, with task-completion attrition highest for processing-speed and memory tasks; performance was strongly IQ-dependent; only Flanker and processing speed carried diagnosis effects.14
- Profile (IQ-matched 12–22-year-olds): fluid-over-crystallized impairment with LPA-derived subgroups;15 both papers also note that mental-age-based administration (from the ID literature) improves feasibility in intellectual disability,14 and the battery has been formally validated in ID populations.14
ABC-CT (Autism Biomarkers Consortium for Clinical Trials)
A multi-site effort to qualify objective biomarkers for ASD trials: an eye-tracking battery (OMI; above),18 an EEG battery analysed with the same psychometric rigour, and video-tracking — deliberately built as trial-readiness instruments rather than diagnostic tests.26
Mouse models: Shank3 and beyond
The standard behavioural battery
The field’s canonical framing tests the three diagnostic domains — social interaction, communication, and repetitive behaviour — with an expanding cognitive add-on layer.27 A 2024 scoping review of rodent ASD models (genetic: Shank3, Fmr1, 16p11.2, Cntnap2, Mecp2, BTBR; environmental: VPA, poly(I:C)/MIA, propionic acid) calls for a comprehensive scoring system so models can be compared by how many neurobehavioral domains they capture rather than single assays.7 Cross-model comparison for SHANK3 specifically shows phenotypes that track construct and residual-isoform differences, mirroring clinical heterogeneity.28
Cognitive tasks in Shank3 and other ASD models
- Shank3B mice (+/−, PDZ-domain deletion): slower to reach criterion in the touchscreen pairwise visual discrimination task (with trends toward more errors), while open-field activity and three-chamber sociability were normal — i.e. an early, cognitive phenotype in a model without gross social/activity confounds.29 (The task was explicitly chosen for its similarity to CANTAB/NIH Toolbox assays — the translation argument in miniature.)
- Shank3 rats (Long-Evans; WT/HET/KO longitudinal): the most prominent adult deficit was touchscreen visual discrimination and reversal learning — rapid, error-prone responding — against a background of mostly normal social behaviour; female knockouts showed reduced distress ultrasonic vocalisations early; subtle juvenile memory/coordination abnormalities preceded the adult executive dysfunction.30
- BTBR mice (idiopathic model): impairments in operant probabilistic reversal learning in both sexes — connecting the mouse assay to the human PRL findings.31
- A Shank3 transgenic line showed pronounced repetitive/stereotypic behaviour and anxiety-related phenotypes with only mild social change and slightly impaired cognitive flexibility across development.32
- Platforms for these tasks: Bussey-Saksida touchscreen chambers (Campden/Lafayette) and open-source alternatives; rodent touchscreen batteries were designed explicitly for cross-species task homology (see automated-cognitive-testing-devices).3334
Monkey models: SHANK3 and other NHP models
SHANK3 macaques
- Founder generation (CRISPR-Cas9, germline-transmissible; Zhou et al. 2019, Nature): sleep disturbances, motor deficits, increased repetitive behaviours, and social and learning impairments, with altered local/global functional connectivity.35
- A single SHANK3-mutant cynomolgus (SHANK3M3) followed longitudinally against matched controls: delayed growth, increased stereotyped behaviour, reduced exploration, anxiety-related crook-tail posture, reduced social initiation, and reduced eye contact; chronic fluoxetine (2.5 mg/kg/day) improved social interaction and partially normalised brain glucose metabolism — the first drug-response evidence in a SHANK3 primate.5
- F1 generation cohort (Jiang et al. 2026, Neuron; larger, non-mosaic): sleep disturbances, diminished exploration, atypical social interactions, stereotypical behaviours, altered functional connectivity, and a markedly diminished auditory ERP response. Cognitively, the F1 monkeys showed no major working-memory deficit but impaired paired-association learning — the cleanest evidence that the SHANK3 haploinsufficiency signature in primates is selective, and that explicit memory tasks (PAL-like) are more sensitive than WM tasks. A multi-task behavioural array was developed to score autism-related phenotypes systematically, and revealed pronounced inter-individual heterogeneity.4
Other NHP models and the monkey-task landscape
- MECP2-overexpressing macaques: autism-like behaviours (repetitive locomotion, increased stress responses, reduced social interaction) with germline transmission — the proof-of-principle for transgenic NHP models of syndromic autism.36
- Maternal VPA macaques: neurogenesis defects and autism-like behaviours in offspring.37
- VPA marmosets: reduced early synaptogenesis (then increased at juvenile ages, as in human tissue), transiently increased plasticity in infancy associated with altered vocalisations, and gene-expression changes resembling human ASD;38 a separate cohort showed reduced gaze to conspecifics in childhood that predicted adult social deficits and perseverative reversal-learning behaviour — the strongest longitudinal childhood-marker→adult-phenotype chain in any animal model.6
- Task availability is the bottleneck: a review of macaque screening tools concluded that limited phenotype-screening batteries constrain the translational utility of NHP models and called for specialised tests and rodent–macaque collaboration.39 New instruments continue to be developed, e.g. a scalable touchscreen delayed non-match-to-position task validated in 12 marmosets and 71 humans that parametrically manipulates delay and spatial interference and works identically in both species.40
Cross-species task map
| Task family | Human evidence (ASD) | Mouse | Monkey |
|---|---|---|---|
| Visuospatial WM (self-ordered / DNMTP) | SWM errors ↑, strategy ↓ 8 | (no strong Shank3 WM phenotype reported) | DNMTP validated marmoset↔human 40; WM intact in SHANK3 F1 macaques 4 |
| Discrimination & reversal / PRL | regressive errors ↔ RRB 12; PRL outcome-ready 13 | discrimination & reversal deficits in Shank3 rats 30; PRL deficits in BTBR 31 | reversal perseveration in VPA marmosets 6 |
| Paired-association / associative memory | PAL in CANTAB-ASD battery 9 | PAL tasks available on rodent touchscreen platforms | impaired in SHANK3 F1 macaques 4 |
| Social attention / gaze | OMI d≈−0.8; 71/81% sens/spec 1819 | no direct analog (sociability assays as proxies 27) | eye contact ↓ (SHANK3) 5; gaze ↓ predicts inflexibility (marmoset) 6 |
| Repetitive behaviour | PRL/EF correlates; RRB scales 12 | grooming/stereotypies (reported across Shank3 lines 3228) | stereotypies in SHANK3 founders 35 and MECP2 monkeys 36 |
| Emotion recognition | ERT; non-selective impairment 23 | absent | — |
Translation strategy and platforms
- The identical-test strategy is the field’s ideal: the first direct demonstration ran the same object-location paired-associates task in Dlg2 knockout mice and in humans carrying DLG2 CNV deletions, showing genotype-linked PAL impairment in both species.33 (DLG2 is a postsynaptic-density scaffold gene — the same protein family logic as SHANK3, making this the template for SHANK3 co-clinical studies.)
- Co-clinical trials: computer-automated cognitive batteries (CANTAB-family) are argued to be the practical bridge for mouse→human trials.34 A 2025 systematic review of actual human–rodent touchscreen co-clinical studies found only 6 eligible studies; behavioural flexibility and visuospatial cognition showed the best comparability, with methodological diversity the main gap.41
- For implementing these tasks in monkeys and mice at scale, see the platform landscape in automated-cognitive-testing-devices — including Monkey CANTAB, Bussey-Saksida chambers and the self-adapting in-cage systems (CageLab and neighbours).
Recent literature (2023–2026 highlights)
- Jiang et al. 2026 (Neuron) — F1 SHANK3 macaques: selective cognitive signature (PAL-impaired, WM-spared); multi-task phenotype array.4
- Pearson et al. 2026 (bioRxiv, NIMH) — Shank3 Long-Evans rats: late-emerging, touchscreen-defined executive dysfunction as the leading adult phenotype.30
- Vanderlip et al. 2026 (Neurobiol Learn Mem) — cross-species delayed non-match-to-position (marmosets + humans) with parametric delay/interference.40
- Tecar et al. 2025 (J Clin Med) — eye-tracking early-screening meta-analysis: 57 studies, g ≈ 0.65, ML pipelines up to 89%/86%.21
- Martins et al. 2025 (J Neurochem) — systematic review of human–rodent co-clinical touchscreen trials; flexibility + visuospatial most comparable; field thin (6 studies).41
- Lage et al. 2024 (Neurosci Biobehav Rev) — cognitive-flexibility meta-analysis; perseverative errors largest effect.11
- Kofler et al. 2024 (Nat Rev Psychol) — ADHD/ASD EF review; specificity and co-occurrence as the core methodological problems.2
- Ongoing: ABC-CT-derived measures are the current template for trial-ready non-verbal endpoints in ASD,1826 and SHANK3-macaque platforms are moving toward drug testing.54
Open questions
- Which task, for what purpose? Screening/subtype detection (eye-tracking GeoPref-type, high specificity/low sensitivity) and mechanistic trial endpoints (OMI, PRL, PAL) are different products with different psychometrics.1820
- Human–animal discordance: humans show visuospatial WM weaknesses; SHANK3 macaques show intact WM but PAL deficits.84 Is that model-specific, task-specific (self-ordered search vs DNMTP), or a species difference in what the gene does?
- Specificity vs transdiagnostic value: EF and flexibility differences overlap heavily with ADHD; the ASD-unique cognitive core remains contested.2
- Standardisation: from mouse batteries (call for a comprehensive scoring system7) to eye-tracking protocols21 and co-clinical task harmonisation,41 the field lacks shared standards.
- Predictive validity: PRL13 and eye-tracking18 are the first non-verbal measures with published trial-feasibility data; none has yet been shown to track a successful disease-modifying intervention.
Related pages
nonverbal-cognitive-tasks · cantab · automated-cognitive-testing-devices · preclinical-drug-screening — task pages: intradimensional-extradimensional-shift, reversal-learning, stop-signal-task, delayed-match-to-sample, spatial-working-memory, emotion-recognition-task, one-touch-stockings-cambridge, delayed-response-tasks-working-memory — entities: lafayette-instrument, campden-instruments
References
Footnotes
-
raw/papers/jiang-2026-f1-shank3-macaques-neuron.md ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8
-
raw/papers/nakagami-2022-marmoset-vpa-attention.md ↩ ↩2 ↩3 ↩4
-
raw/papers/steele-2007-spatial-working-memory-autism.md ↩ ↩2 ↩3 ↩4
-
raw/papers/jones-2021-nih-toolbox-feasibility-autism.md ↩ ↩2 ↩3 ↩4 ↩5 ↩6
-
raw/papers/shic-2022-abcct-eye-tracking-battery.md ↩ ↩2 ↩3 ↩4 ↩5 ↩6
-
McPartland JC et al. (2020). The Autism Biomarkers Consortium for Clinical Trials (ABC-CT): Scientific Context, Study Design, and Progress Toward Biomarker Qualification. Front Integr Neurosci 14:16. https://doi.org/10.3389/fnint.2020.00016 ↩ ↩2
-
raw/papers/nithianantharajah-2015-identical-touchscreen.md ↩ ↩2
-
raw/papers/martins-2025-touchscreen-coclinical-review.md ↩ ↩2 ↩3