What architecture follows from practice, and what would show it worked?
Meeting 13 of 13
Drafted v0.4 · not yet held · record discussion
Central question
Do agents organized around human-defined epistemic modes produce greater warranted scientific gains on documented past work than comparable planners?
Anchor and companion
- Everyone reads the selected anchor sections and the companion abstract/overview. The rotating reader presents the companion in depth.
- Pairing: Implemented laboratory agent / concrete agent benchmark. Full records and access notes: bibliography.
Anchor
Boiko, D. A., MacKnight, R., Kline, B., & Gomes, G. (2023). Autonomous chemical research with large language models. Nature, 624, 570–578. https://doi.org/10.1038/s41586-023-06792-0
Read the system architecture and one experimental demonstration. Identify which steps the system performed and which a person specified for it.
Companion
Chen, Z., Chen, S., Ning, Y., Zhang, Q., Wang, B., Yu, B., Li, Y., Liao, Z., Wei, C., Lu, Z., Dey, V., Xue, M., Baker, F. N., Burns, B., Adu-Ampratwum, D., Huang, X., Ning, X., Gao, S., Su, Y., & Sun, H. (2025). ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery. In International Conference on Learning Representations. https://proceedings.iclr.cc/paper_files/paper/2025/hash/f12b4df26344f3be803c06b555252efe-Abstract-Conference.html
Read the task construction, evaluation, and stated limitations. The full accepted conference paper and proceedings are linked.
Why these readings belong together
Boiko and colleagues demonstrate a tool-using system that acts in a laboratory. ScienceAgentBench specifies executable tasks drawn from published studies and scores them. Between them they cover what an agent is built to do and what a score of its behavior can mean. A demonstration does not establish autonomy across every scientific mode, and success on a supplied analysis task does not by itself evaluate question choice, unexpected discovery, or field-level scientific value. (Boiko et al. 2023; Chen et al. 2025)
Prepare and discuss
Read the selected anchor sections and the companion abstract or overview. Bring one source passage or artifact relevant to the case; the rotating reader presents the companion in depth.
- Who specified the objective, supplied the tools, and checked the evidence?
- Which construct does each proposed metric actually observe, and where could a correct final answer conceal invalid reasoning, leakage, or an unreproducible workflow?
- How might a shared automated workflow narrow the questions this group asks, and would our evaluation detect that?
Case exercise: scientific gain and epistemic mode
Complete one chain from a documented human case to a mode specification, procedure, claim record, and historical evaluation packet. Define a meaningful gain, a fair baseline, controls, information boundaries, and an assessment that could overturn the proposed benefit.
- Profiles to examine: estimation, exploration, hypothesis, methods.
- Record source kind and unknown chronology; distinguish documented practice, philosophical argument, association, and our proposed agent rule.
- Ask what was gained, what warrants it, which action helped, and what a comparable agent would need to demonstrate.
The agent
- After the meeting, say what these papers change in the unit skill, agents/skills/unit-of-inquiry/SKILL.md: a rule at a node, a new composition of units, or a planted flaw. Meeting 1’s page shows the form.
- Modes exercised here, whose “agent actions” sections the rule would enter: estimation, exploration, hypothesis, methods.
- Written from the texts after the papers are read in full, as for meeting 1; nothing here yet.
Optional extensions
- Simon (1973): Does Scientific Discovery Have a Logic?. Philosophy of Science (1973). Simon argues, against Popper, that discovery is problem solving with describable heuristics. The computer discovery programs listed in meeting 13 were built on this argument.
- Messick (1995): Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. Publisher/DOI page; full text may require institutional access.
- Langley (1981): Data-Driven Discovery of Physical Laws. Cognitive Science (1981). BACON recovered Kepler’s third law and Ohm’s law from tables of numbers by searching for invariants. A historical example of invariant search with supplied variables and data; inspect its representation limits.
- Kulkarni and Simon (1988): The Processes of Scientific Discovery: The Strategy of Experimentation. Cognitive Science (1988). KEKADA reproduces Krebs’s path to the urea cycle, built from Holmes’s notebook-based reconstruction; surprise is one of its control signals. The one prior system built the way this seminar proposes to build: from documented human practice.
- King et al. (2009): The Automation of Science. Science (2009). The Robot Scientist Adam: a closed generate-and-test loop with a real laboratory and reported functional-genomics discoveries. Its enumerable task and assay constraints matter to comparison.
- Schmidt and Lipson (2009): Distilling Free-Form Natural Laws from Experimental Data. Science (2009). Symbolic regression recovers conservation laws from motion-capture data. What “law discovery” means when the variables have already been chosen.
- Wang et al. (2023): Scientific discovery in the age of artificial intelligence. Nature (2023). The standard review. Use it as an index; it is not evidence about any one system.
- Kitano (2021): Nobel Turing Challenge: creating the engine for scientific discovery. npj Systems Biology and Applications (2021). The explicit proposal for a closed-loop AI scientist and what it would have to do. The clearest statement of the design this seminar is questioning.
- Gottweis et al. (2026): Accelerating scientific discovery with Co-Scientist. Nature (2026). The multi-agent hypothesis tournament, now peer reviewed. Compare the validated cases with the number of proposals generated.
- Szymanski et al. (2023): An autonomous laboratory for the accelerated synthesis of inorganic materials. Background for the meeting; compare its population, evidence, and scope.
- Leeman et al. (2024): Challenges in High-Throughput Inorganic Materials Prediction and Autonomous Synthesis. Background for the meeting; compare its population, evidence, and scope.
- Kapoor and Narayanan (2023): Leakage and the reproducibility crisis in machine-learning-based science. Patterns (2023). A taxonomy of leakage with cases across fields. Read before reusing any benchmark.
- Rainforth et al. (2024): Modern Bayesian Experimental Design. Statistical Science (2024). Expected information gain as the rule for choosing the next experiment. Useful when the model and observation spaces are specified; changing representations raises additional design questions.
- Lehman and Stanley (2011): Abandoning Objectives: Evolution Through the Search for Novelty Alone. Evolutionary Computation (2011). Searching for behavioral novelty instead of for an objective solved problems that objective search could not, because the objective was deceptive. The machine-side argument that a gate can narrow the search, stated so it can be tested.
- Gil et al. (2016): Toward the Geoscience Paper of the Future: Best practices for documenting and sharing research from data to software to provenance. Earth and Space Science (2016). The geoscience research workflow specified end to end, with the documentation, software, data, and provenance a paper should carry to be reproducible. The nearest thing in the literature to a general account of how a geoscience result is produced.
- Gil et al. (2018): Intelligent systems for geosciences: an essential research agenda. Communications of the ACM (2018). A research agenda for intelligent systems in the geosciences, written by geoscientists and computer scientists together.
- Messeri and Crockett (2024): Artificial intelligence and illusions of understanding in scientific research. Conceptual analysis of AI and scientific understanding; a prompt for evaluating what a fluent output conceals.
- Flake and Fried (2020): Measurement Schmeasurement: Questionable Measurement Practices and How to Avoid Them. Measurement validity: a repeatable benchmark score can measure the wrong construct.
- Si et al. (2025): The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas. Executed research outcomes in NLP, separate from the earlier idea-rating study.
- Woollam et al. (2022): SeisBench—A Toolbox for Machine Learning in Seismology. Seismological ML data/model infrastructure that can support capability evaluation.
- Dekoninck et al. (2024): Evading Data Contamination Detection for Language Models is (too) Easy. Limits of contamination detection; negative probes do not establish no prior exposure.
Record after the meeting
Sign and date the notes; preserve disagreements and what changed your assessment. Keep confidential examples in private notes.