What architecture follows from practice, and what would show it worked?

Meeting 13 of 13

Drafted v0.4 · not yet held · record discussion

Central question

Do agents organized around human-defined epistemic modes produce greater warranted scientific gains on documented past work than comparable planners?

Anchor and companion

  • Everyone reads the selected anchor sections and the companion abstract/overview. The rotating reader presents the companion in depth.
  • Pairing: Implemented laboratory agent / concrete agent benchmark. Full records and access notes: bibliography.

Anchor

Boiko, D. A., MacKnight, R., Kline, B., & Gomes, G. (2023). Autonomous chemical research with large language models. Nature, 624, 570–578. https://doi.org/10.1038/s41586-023-06792-0

Read the system architecture and one experimental demonstration. Identify which steps the system performed and which a person specified for it.

Companion

Chen, Z., Chen, S., Ning, Y., Zhang, Q., Wang, B., Yu, B., Li, Y., Liao, Z., Wei, C., Lu, Z., Dey, V., Xue, M., Baker, F. N., Burns, B., Adu-Ampratwum, D., Huang, X., Ning, X., Gao, S., Su, Y., & Sun, H. (2025). ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery. In International Conference on Learning Representations. https://proceedings.iclr.cc/paper_files/paper/2025/hash/f12b4df26344f3be803c06b555252efe-Abstract-Conference.html

Read the task construction, evaluation, and stated limitations. The full accepted conference paper and proceedings are linked.

Why these readings belong together

Boiko and colleagues demonstrate a tool-using system that acts in a laboratory. ScienceAgentBench specifies executable tasks drawn from published studies and scores them. Between them they cover what an agent is built to do and what a score of its behavior can mean. A demonstration does not establish autonomy across every scientific mode, and success on a supplied analysis task does not by itself evaluate question choice, unexpected discovery, or field-level scientific value. (Boiko et al. 2023; Chen et al. 2025)

Prepare and discuss

Read the selected anchor sections and the companion abstract or overview. Bring one source passage or artifact relevant to the case; the rotating reader presents the companion in depth.

  1. Who specified the objective, supplied the tools, and checked the evidence?
  2. Which construct does each proposed metric actually observe, and where could a correct final answer conceal invalid reasoning, leakage, or an unreproducible workflow?
  3. How might a shared automated workflow narrow the questions this group asks, and would our evaluation detect that?

Case exercise: scientific gain and epistemic mode

Complete one chain from a documented human case to a mode specification, procedure, claim record, and historical evaluation packet. Define a meaningful gain, a fair baseline, controls, information boundaries, and an assessment that could overturn the proposed benefit.

  • Profiles to examine: estimation, exploration, hypothesis, methods.
  • Record source kind and unknown chronology; distinguish documented practice, philosophical argument, association, and our proposed agent rule.
  • Ask what was gained, what warrants it, which action helped, and what a comparable agent would need to demonstrate.

Shared output

One episode entry and three sentences: what changed scientifically; what evidence warrants that account; what agent action or evaluation follows. Longer worksheets are optional.

The agent

  • After the meeting, say what these papers change in the unit skill, agents/skills/unit-of-inquiry/SKILL.md: a rule at a node, a new composition of units, or a planted flaw. Meeting 1’s page shows the form.
  • Modes exercised here, whose “agent actions” sections the rule would enter: estimation, exploration, hypothesis, methods.
  • Written from the texts after the papers are read in full, as for meeting 1; nothing here yet.

Optional extensions

Record after the meeting

Sign and date the notes; preserve disagreements and what changed your assessment. Keep confidential examples in private notes.