Prior art: agents and benchmarks for science

Drafted registry · not yet discussed

How to read the last column

  • A benchmark score is a measurement of the tasks, not of “doing science” (Flake and Fried 2020; Messick 1995)
  • For each task set, inspect who supplies the question, data, tools, and success criterion; the boundaries differ.
  • Read the last column as an inventory of open problems, not as criticism

The group’s own system

  • GAIA: thirteen role-specialized subagents under one orchestrator; read-only auditor; lab notebook; provenance keeper that discloses AI use; human gates
  • Beside it: an evaluation vault of golden items with planted flaws and clean controls
  • This book does not establish performance of the group’s runtime; meeting 13 specifies a mode-based adapter and evaluation.
  • Implementation and run status must be checked against actual artifacts before claims about its capabilities.
  • Roster private and under review; public summary in the design notes

Systems that predate language models

  • Language-model systems have important predecessors.
  • One of these was built the way this seminar proposes: from a historian’s reconstruction of a scientist’s notebooks
  • Simon’s account of discovery is relevant philosophical background, not a demonstrated common commitment of every system (Simon 1973).
Project What it is What it covered What its results do not establish
BACON — Langley (1981) Heuristic search for invariants in tables of numeric data Rediscovered Kepler’s third law, Ohm’s law, and others from the original data Anything about choosing the variables or the data; both were supplied
KEKADA — Kulkarni and Simon (1988) A model of Krebs’s discovery of the urea cycle, built from Holmes’s notebook-based reconstruction; surprise and confirmation are explicit control signals Reproduces the sequence of experiments Krebs ran, including the exploratory phase before any hypothesis That the strategy transfers to another scientist or problem. It is one episode, modelled after the fact
Robot Scientist Adam — King et al. (2009) A closed loop: hypothesis generation, experiment selection, robotic execution, and analysis, in yeast functional genomics Reported autonomous functional-genomics discoveries with follow-up checks That the loop generalizes beyond a domain with an enumerable hypothesis space and a fully automated assay
Eureqa — Schmidt and Lipson (2009) Symbolic regression over experimental time series Recovered Lagrangians and Hamiltonians for pendulums and oscillators Which variables to measure; the search space is fixed by the input columns
Literature-based discovery — Swanson (1986) Connecting two literatures that never cite each other, first done by reading and later by program The fish-oil and Raynaud’s syndrome link, later confirmed clinically How often the method produces false connections. The successes are what got reported
  • Arguments rather than systems: Kitano (2021) states the closed-loop design explicitly; Wang et al. (2023) is the standard review, an index rather than evidence

Benchmarks and evaluations

Project What it is What it covers What its scores do not establish
ScienceAgentBench — Chen et al. (2025) 102 executable tasks drawn from 44 peer-reviewed papers, validated by subject-matter experts Data-driven analysis: given a question and data, produce working code and the right output Question formation, novelty, or discovery. The task supplies what earlier meetings identified as the difficult and interesting choices
DiscoveryBench (ICLR 2025; preprint arXiv:2407.01725) 264 tasks derived from published papers, each a multi-step search for a stated relationship in supplied data, plus 903 synthetic tasks Data-driven discovery as a workflow: the system must find the relationship, not only run an analysis. The best system reported scored 25% Question formation: the goal is stated in words for every task, and the data are supplied. Transfer beyond the six domains covered is untested
MLE-bench (ICLR 2025; preprint arXiv:2410.07095) 75 Kaggle competitions run as machine-learning engineering tasks Engineering competence: building a model that scores well against a defined metric Any scientific claim. A Kaggle task arrives with the target, the metric, and the guarantee that a solution exists
Gravity-Bench-v1 (ICML 2025; preprint arXiv:2501.18411) Simulated gravitational systems where the agent must plan its own observations Closer to inquiry than most: the agent decides what to measure, under a budget, including out-of-distribution physics Performance with real instruments, real noise, real calibration, or a real literature to be wrong about
LLM-SRBench (ICML 2025; preprint arXiv:2504.10415) Equation-discovery tasks built to resist recall of known formulas Whether a system can find a governing relation in data rather than retrieve one it has seen Inquiry beyond the symbolic-regression step. The data, the variables, and the goal are given
FML-bench (arXiv:2510.10472) Machine-learning research tasks scored on multiple axes Broader research behavior than single-task benchmarks Transfer outside machine learning, which is the field these agents were largely trained on
Human study of LLM ideation — Si et al. (2025) A blinded comparison of research ideas written by over a hundred NLP researchers and by a language-model agent, judged by expert reviewers Novelty, excitement, and feasibility of the ideas themselves; the model’s ideas rated higher on novelty and lower on feasibility That a novel-rated idea leads to good work. Execution is a separate study, and the field is NLP
NovBench — Wu et al. (2026) Evaluation of model judgments of paper novelty Whether a model’s novelty assessment agrees with expert assessment on NLP papers That the judgments transfer to Earth science. Whether they do is itself something to check. See meeting 10

Surveys and registries

  • Two surveys cover the space more completely; start there when extending the tables

  • A Survey of AI Scientists (arXiv:2510.23045). Tie and colleagues, 2025. Broad coverage of automated-scientist systems and their components.

  • Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap (arXiv:2608.05179). Ding and colleagues, 2026. Screens a large candidate set down to a coded sample of runnable systems, and organizes them around verification, the same concern as the last column of the tables above.

  • A survey is a secondary source; where a claim matters, read the primary paper

Questions for the next evidence review

  • This registry is a starting sample, not a systematic absence search. Record a query, date, source set, and limits before claiming a capability has no precedent.
  • For each benchmark, inspect coverage of question choice, measurement response, trace inference, anomaly follow-up, justified limits, field constraints, and scientific gains.
  • Separate task execution, process adequacy, and demonstrated advance. A supplied answer target can measure a useful capability while leaving other questions open.
  • The historical evaluation proposes three suites and explicit evidence boundaries. Its difference from existing evaluations must be documented before any priority claim.
  • Existing geoscience resources include source-inversion exercises (Mai et al. 2016) and seismological ML data/model infrastructure (Woollam et al. 2022). Reuse suitable components rather than claiming a blank field.
  • Assess collective topic coverage separately from individual productivity. Hao et al. (2026) is a scientific-publication study; Messeri and Crockett (2024) is a conceptual analysis, not a second causal demonstration of the same effect.

Adding an entry

  • Corrections welcome, especially from anyone who has run one of these systems
  • Accuracy in the last column matters more than coverage; overstating a limitation is as much a problem as omitting it · how to contribute