Prior art: agents and benchmarks for science
Drafted registry · not yet discussed
- Systems that automate parts of research exist, and so do benchmarks that score them
- Before proposing an architecture or a test, find out whether it has been built, and what happened
- Entries carry source records; verification of identity is distinct from checking every interpretation or current system capability.
How to read the last column
- A benchmark score is a measurement of the tasks, not of “doing science” (Flake and Fried 2020; Messick 1995)
- For each task set, inspect who supplies the question, data, tools, and success criterion; the boundaries differ.
- Read the last column as an inventory of open problems, not as criticism
The group’s own system
- GAIA: thirteen role-specialized subagents under one orchestrator; read-only auditor; lab notebook; provenance keeper that discloses AI use; human gates
- Beside it: an evaluation vault of golden items with planted flaws and clean controls
- This book does not establish performance of the group’s runtime; meeting 13 specifies a mode-based adapter and evaluation.
- Implementation and run status must be checked against actual artifacts before claims about its capabilities.
- Roster private and under review; public summary in the design notes
Systems that predate language models
- Language-model systems have important predecessors.
- One of these was built the way this seminar proposes: from a historian’s reconstruction of a scientist’s notebooks
- Simon’s account of discovery is relevant philosophical background, not a demonstrated common commitment of every system (Simon 1973).
| Project | What it is | What it covered | What its results do not establish |
|---|---|---|---|
| BACON — Langley (1981) | Heuristic search for invariants in tables of numeric data | Rediscovered Kepler’s third law, Ohm’s law, and others from the original data | Anything about choosing the variables or the data; both were supplied |
| KEKADA — Kulkarni and Simon (1988) | A model of Krebs’s discovery of the urea cycle, built from Holmes’s notebook-based reconstruction; surprise and confirmation are explicit control signals | Reproduces the sequence of experiments Krebs ran, including the exploratory phase before any hypothesis | That the strategy transfers to another scientist or problem. It is one episode, modelled after the fact |
| Robot Scientist Adam — King et al. (2009) | A closed loop: hypothesis generation, experiment selection, robotic execution, and analysis, in yeast functional genomics | Reported autonomous functional-genomics discoveries with follow-up checks | That the loop generalizes beyond a domain with an enumerable hypothesis space and a fully automated assay |
| Eureqa — Schmidt and Lipson (2009) | Symbolic regression over experimental time series | Recovered Lagrangians and Hamiltonians for pendulums and oscillators | Which variables to measure; the search space is fixed by the input columns |
| Literature-based discovery — Swanson (1986) | Connecting two literatures that never cite each other, first done by reading and later by program | The fish-oil and Raynaud’s syndrome link, later confirmed clinically | How often the method produces false connections. The successes are what got reported |
Benchmarks and evaluations
| Project | What it is | What it covers | What its scores do not establish |
|---|---|---|---|
| ScienceAgentBench — Chen et al. (2025) | 102 executable tasks drawn from 44 peer-reviewed papers, validated by subject-matter experts | Data-driven analysis: given a question and data, produce working code and the right output | Question formation, novelty, or discovery. The task supplies what earlier meetings identified as the difficult and interesting choices |
| DiscoveryBench (ICLR 2025; preprint arXiv:2407.01725) | 264 tasks derived from published papers, each a multi-step search for a stated relationship in supplied data, plus 903 synthetic tasks | Data-driven discovery as a workflow: the system must find the relationship, not only run an analysis. The best system reported scored 25% | Question formation: the goal is stated in words for every task, and the data are supplied. Transfer beyond the six domains covered is untested |
| MLE-bench (ICLR 2025; preprint arXiv:2410.07095) | 75 Kaggle competitions run as machine-learning engineering tasks | Engineering competence: building a model that scores well against a defined metric | Any scientific claim. A Kaggle task arrives with the target, the metric, and the guarantee that a solution exists |
| Gravity-Bench-v1 (ICML 2025; preprint arXiv:2501.18411) | Simulated gravitational systems where the agent must plan its own observations | Closer to inquiry than most: the agent decides what to measure, under a budget, including out-of-distribution physics | Performance with real instruments, real noise, real calibration, or a real literature to be wrong about |
| LLM-SRBench (ICML 2025; preprint arXiv:2504.10415) | Equation-discovery tasks built to resist recall of known formulas | Whether a system can find a governing relation in data rather than retrieve one it has seen | Inquiry beyond the symbolic-regression step. The data, the variables, and the goal are given |
| FML-bench (arXiv:2510.10472) | Machine-learning research tasks scored on multiple axes | Broader research behavior than single-task benchmarks | Transfer outside machine learning, which is the field these agents were largely trained on |
| Human study of LLM ideation — Si et al. (2025) | A blinded comparison of research ideas written by over a hundred NLP researchers and by a language-model agent, judged by expert reviewers | Novelty, excitement, and feasibility of the ideas themselves; the model’s ideas rated higher on novelty and lower on feasibility | That a novel-rated idea leads to good work. Execution is a separate study, and the field is NLP |
| NovBench — Wu et al. (2026) | Evaluation of model judgments of paper novelty | Whether a model’s novelty assessment agrees with expert assessment on NLP papers | That the judgments transfer to Earth science. Whether they do is itself something to check. See meeting 10 |
Surveys and registries
Two surveys cover the space more completely; start there when extending the tables
A Survey of AI Scientists (arXiv:2510.23045). Tie and colleagues, 2025. Broad coverage of automated-scientist systems and their components.
Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap (arXiv:2608.05179). Ding and colleagues, 2026. Screens a large candidate set down to a coded sample of runnable systems, and organizes them around verification, the same concern as the last column of the tables above.
A survey is a secondary source; where a claim matters, read the primary paper
Questions for the next evidence review
- This registry is a starting sample, not a systematic absence search. Record a query, date, source set, and limits before claiming a capability has no precedent.
- For each benchmark, inspect coverage of question choice, measurement response, trace inference, anomaly follow-up, justified limits, field constraints, and scientific gains.
- Separate task execution, process adequacy, and demonstrated advance. A supplied answer target can measure a useful capability while leaving other questions open.
- The historical evaluation proposes three suites and explicit evidence boundaries. Its difference from existing evaluations must be documented before any priority claim.
- Existing geoscience resources include source-inversion exercises (Mai et al. 2016) and seismological ML data/model infrastructure (Woollam et al. 2022). Reuse suitable components rather than claiming a blank field.
- Assess collective topic coverage separately from individual productivity. Hao et al. (2026) is a scientific-publication study; Messeri and Crockett (2024) is a conceptual analysis, not a second causal demonstration of the same effect.
Adding an entry
- Corrections welcome, especially from anyone who has run one of these systems
- Accuracy in the last column matters more than coverage; overstating a limitation is as much a problem as omitting it · how to contribute