From human-defined modes to scientific agents
Drafted v0.4 design · exported instructions available · scientific performance untested
Objective and evidence chain
- Primary objective: produce meaningful warranted gains in knowledge, understanding, or research capability under a stated budget.
- Human scientific record → mode specification → executable procedure → historical evaluation → assessed scientific gain.
- Novelty characterizes what changes; correctness/warrant supports its scope; impact records later consequences. An agent can estimate expected advance for planning, but those estimates require subsequent assessment.
- History-informed discovery systems are precedents (Langley 1981; Kulkarni and Simon 1988). Their existence motivates this project; their task boundaries do not establish general agent competence.
Worked design example
- Illustrative question: does seasonal seismic velocity variation help constrain groundwater effects? This is a teaching scenario, not a reported result from a private project.
- Observe the series and coverage; estimate velocity change with uncertainty and processing assumptions.
- Explore variability; follow a jump with a measurement-chain check before attributing a cause.
- Generate rivals such as physical changes and changes in the noise/measurement chain; identify observables with different consequences under those rivals.
- Transfer hydrological information only after checking correspondence of target quantities, scales, and assumptions.
- Return a bounded result, a meaningful limitation, or a proposed next observation. A causal claim needs more than a fitted seasonal association.
- This path is illustrative; the planner may enter a different profile when the evidence requires it.
What the readings support
- Si’s ideation study and execution study are distinct (Si, Yang, et al. 2025; Si, Hashimoto, et al. 2025). In the execution study, 43 researchers implemented assigned NLP ideas; model-generated ideas’ ratings fell more from ideation to execution on the reported measures. This does not show that all scientific hypothesis generation is poor.
- Oreskes et al. (1994) permits partial confirmation while rejecting complete verification/validation of natural-system models. State the purpose and supported scope.
- Steinle (1997) and Karaca (2013a) motivate exploratory variation; Karaca (2013b) changes the reported chronology of Bjorken’s hypothesis. Existence of a hypothesis does not show how it guided an action.
- Leeman et al. (2024) offers a critical materials reanalysis; attribute its conclusions and compare the original evidence. It does not establish a single failure mechanism for all scientific agents.
- Shi and Evans (2023) and Sourati and Evans (2023) motivate testing cross-field connections. Neither makes a broad portfolio or an unusual reference combination sufficient for advance.
Implementation state and integration
| Component | State here | Evidence needed next |
|---|---|---|
| Versioned mode specifications | Drafted and exported from source. | Human review on contrasting cases; revisions recorded. |
| Claim/episode record | Templates and public worksheets. | Capture fidelity checked against actual artifacts. |
| Historical evaluation | Three reading-based case specifications. | Permitted source packets, independent references, and frozen splits. |
| Scientific-agent runtime | Not implemented or run in this repository. | Connect an existing runner, execute a case, and inspect artifacts. |
| Scientific benefit | Untested. | Comparable policies on independent held-out cases; prospective evidence later. |
- The group’s existing agent family can consume these specifications without a new persona per mode. This page does not assert that its notebook, auditor, or harness has run successfully.
- Adapter responsibilities: pass source version and mode state to the planner; let the study designer act without inventing a hypothesis; preserve action/claim records; give the assessor claim-appropriate checks; expose mode and case axes in evaluation.
- First integration check: one nonconfidential toy task must emit a faithful action and claim record. Then complete an evidence-linked historical packet. A toy run is an engineering check, not evidence of discovery skill.
- Correlated prompting roles are not independent expertise. A separate prompt or agent does not automatically provide independent evidence.
Evaluation that could change the design
- Compare competent generic, hypothesis-testing, and mode-selecting planners with comparable model, tools, guidance, evidence, budgets, and assistance.
- Test supplied-mode competence separately from selection/transitions; remove mode organization while preserving substantive guidance in an ablation.
- Primary outcome is case-level warranted gain the protocol; report scientific magnitude, unsupported claims, cost, novelty, and later impact separately.
- Historical reference work is not a required path or a clean modern human-speed baseline. A remembered answer is not a newly demonstrated scientific gain.
- Preserve failures, unknowns, and unanticipated valid gains; evaluate breadth at project, competence, and portfolio scales.
- A broader literature study can describe reported practices and uptake. It cannot substitute for testing the procedures.