OmniScientist — Project Page

URL: https://omni-scientist.github.io Source code: https://github.com/Omni-Scientist/OmniScientist Contact: libobo@nus.edu.sg, haofei7419@gmail.com (corresponding: Hao Fei)

Key Claims

  • One engine, many domains. OmniScientist is an omni-modal, omni-discipline AI scientist that takes raw heterogeneous evidence (images, signals, audio, video, 3-D structures, trajectories, tables, formulae, graphs) and produces a finished scientific study end-to-end.
  • Pipeline: Perceive → Hypothesis → Experiment → Report → Feedback. A perception layer plus 3 autonomous agents (ideation, experiment, writeup) operate within a deterministic pipeline.
  • Rigour checks: Idea checks (novelty screening), rigour checks (statistical validity, execution provenance), and claim checks (numerical traceability) are enforced in code.
  • 36 real-data cases spanning 5 discipline families, 4 evidence families, and 20+ disciplines (pathology, seismology, astronomy, cardiology, bioacoustics, materials science, remote sensing, plant science, radiology, etc.).

Disciplines Covered (from hero tiles)

Tropical cyclones, marine bioacoustics, symbolic physics, radiology, plant phenotyping, bioacoustics, pathology, cardiology, movement ecology, materials science, plant pathology, hyperspectral sensing, seismology, remote sensing, transportation, LiDAR perception, cell biology, regulatory genomics, machine diagnostics, mechanical CAD, mineralogy, astronomy, atmospheric science. Also: gravitational waves, sleep medicine, digital rock physics, medical imaging, ethology, autonomous driving, marine ecology.

Representative Findings (with metrics)

  1. Seismology: 21.7% of noise-labelled STEAD traces carry coherent polarized transients (163 of 750). Detector calibrated on IAAFT surrogates at 1% false-alarm rate; amplitude-only detector flags only 2.0% — polarization and cross-channel coincidence carry the finding.
  2. Cardiology (audio): Metadata-only AUC drops from 0.60 to 0.35 under leave-one-cohort-out — the shortcut flips sign across sites (drop 0.248, p < 0.0001). The murmur-band acoustic feature barely moves (0.701 → 0.656, not significant).
  3. Astronomy (imaging): 83.8% agreement on galaxy morphology across SDSS and DECaLS surveys (chance = 35.0%, Cohen’s κ = 0.75, CI 0.64–0.85, permutation p = 0.0002, 105 galaxies).
  4. Materials informatics (table): Random-split vs leave-one-family-out error gap 3.1–7.0× (cuprates RMSE 12.1 → 51.6 K, iron-based 6.6 → 20.5 K).
  5. Remote sensing (imaging): Orientation-only classifier collapses from 62.0% to 50.7% when patches are rotated (forest recall 0.68 → 0.27).

Backbone Benchmark (cross-family judge panel)

Scores 0–10 across 7 dimensions (Novelty, Soundness, Clarity, Significance, Reproducibility, MM Grounding, Factual). Composite = mean.

BackboneNoveltySound.ClaritySignif.Reprod.MM ground.FactualComposite
GLM-5.26.27.16.86.45.96.67.56.6
Sonnet 56.37.07.06.36.15.17.86.5
Kimi K2.76.27.26.76.25.55.88.06.5
GPT-5.65.26.36.35.05.24.27.75.7
Qwen3.5-122B4.75.56.24.84.84.86.55.3
Qwen3.5-27B4.95.55.94.94.64.96.35.3
Gemma-4-31B4.75.15.64.54.44.76.55.1
Gemma-4-26B4.44.45.04.03.73.85.14.3
Qwen3.5-9B4.04.14.83.73.73.94.84.1

GLM-5.2 achieves the highest composite score (6.6), with the best multimodal grounding (6.6) and soundness (7.1).

Five Papers Produced End-to-End

PaperDisciplineEvidenceScore
Cramér-Rao scaling of exponent precision in monomial Feynman lawsPhysicsFormula7.1
Sequential versus bursty leaf initiation in 3-D plant scansPlant science3-D scan7.2
A continuum-removed index for residue and tillage orderingRemote sensingHyperspectral6.3
Transient-impulsivity features across machine typesMachineryAudio7.1
Coherent polarized signals in noise-labelled STEAD tracesSeismologySignal6.9

Interactive Demos

Three complete runs with recorded traces (evidence, code, checks, feedback loops):

  1. Seismology — 80 steps, 9 looked at, 33 papers read, 36 code runs, 3 sent back
  2. Astronomy — 42 steps, 12 looked at, 24 papers read, 11 code runs, 3 sent back
  3. Bioacoustics — 85 steps, 7 looked at, 36 papers read, 25 code runs, 6 sent back

Installation

Supports Mac/Windows/Linux desktop, CLI, and agent skill (Claude Code, Codex, Hermes Agent). One-command install via setup document.

Key Distinguishing Feature

In paired comparisons against a “blind” variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments. This demonstrates that lifecycle-wide perception (working from raw evidence, not precomputed summaries) is essential for evidence-grounded scientific discovery.