From human-defined modes to scientific agents

Drafted v0.4 design · exported instructions available · scientific performance untested

Objective and evidence chain

  • Primary objective: produce meaningful warranted gains in knowledge, understanding, or research capability under a stated budget.
  • Human scientific record → mode specification → executable procedure → historical evaluation → assessed scientific gain.
  • Novelty characterizes what changes; correctness/warrant supports its scope; impact records later consequences. An agent can estimate expected advance for planning, but those estimates require subsequent assessment.
  • History-informed discovery systems are precedents (Langley 1981; Kulkarni and Simon 1988). Their existence motivates this project; their task boundaries do not establish general agent competence.

One shared state, several procedures

  • State: question and its origin/revisions, evidence and source kind, measurement assumptions, candidate explanations, claim dependencies, mode IDs/versions, commitments, remaining budget.
  • Each procedure changes actions and representations; mode names attached to otherwise identical behavior are insufficient.
  • Humans edit the authoritative specifications in modes/. The exporter creates agent instructions from them with source hashes; it does not translate prose into a proven reasoning algorithm.
  • Agents can suggest a new mode or revision, or mark work outside the vocabulary. They cannot silently alter the frozen rubric or redefine success during evaluation.
Procedure Observable behavior Scientific assessment
Estimate with checks Separate measured data from latent target; inspect identifiability, response, regularization, and uncertainty. Useful estimate or consequential identified limit; misleading precision is a failure.
Follow an anomaly Generate artifact/physical possibilities; select an informative initial reliability check. Credible candidate and justified next observation; unique events may merit investigation.
Generate and assess rivals Produce substantive alternatives, then separately choose differing evidential consequences. Resolve a meaningful contrast or establish why current evidence cannot.
Reframe a question Revise the scientific target when uncertainty, representation, or resources demand it. A tractable question whose answer would matter, rather than an unrelated easy task.
Transfer a method/concept Map variables, units, scales, assumptions, failure conditions, and a target-domain check. Supported integration that contributes to gain rather than borrowed terminology.
  • The remaining profiles cover description, theory, simulation, instruments, methods, synthesis, and robustness; they can share actions while owing different evidence.
  • Require claim-appropriate evidence for a supported status, not a hypothesis before every action. Conjectures and provisional outputs are permitted.
  • Separate action restrictions, commitments before evidence access, and claim warrant in the record. Evaluate their effects rather than assuming a universal gate.

Worked design example

  • Illustrative question: does seasonal seismic velocity variation help constrain groundwater effects? This is a teaching scenario, not a reported result from a private project.
  • Observe the series and coverage; estimate velocity change with uncertainty and processing assumptions.
  • Explore variability; follow a jump with a measurement-chain check before attributing a cause.
  • Generate rivals such as physical changes and changes in the noise/measurement chain; identify observables with different consequences under those rivals.
  • Transfer hydrological information only after checking correspondence of target quantities, scales, and assumptions.
  • Return a bounded result, a meaningful limitation, or a proposed next observation. A causal claim needs more than a fitted seasonal association.
  • This path is illustrative; the planner may enter a different profile when the evidence requires it.

What the readings support

  • Si’s ideation study and execution study are distinct (Si, Yang, et al. 2025; Si, Hashimoto, et al. 2025). In the execution study, 43 researchers implemented assigned NLP ideas; model-generated ideas’ ratings fell more from ideation to execution on the reported measures. This does not show that all scientific hypothesis generation is poor.
  • Oreskes et al. (1994) permits partial confirmation while rejecting complete verification/validation of natural-system models. State the purpose and supported scope.
  • Steinle (1997) and Karaca (2013a) motivate exploratory variation; Karaca (2013b) changes the reported chronology of Bjorken’s hypothesis. Existence of a hypothesis does not show how it guided an action.
  • Leeman et al. (2024) offers a critical materials reanalysis; attribute its conclusions and compare the original evidence. It does not establish a single failure mechanism for all scientific agents.
  • Shi and Evans (2023) and Sourati and Evans (2023) motivate testing cross-field connections. Neither makes a broad portfolio or an unusual reference combination sufficient for advance.

Implementation state and integration

Component State here Evidence needed next
Versioned mode specifications Drafted and exported from source. Human review on contrasting cases; revisions recorded.
Claim/episode record Templates and public worksheets. Capture fidelity checked against actual artifacts.
Historical evaluation Three reading-based case specifications. Permitted source packets, independent references, and frozen splits.
Scientific-agent runtime Not implemented or run in this repository. Connect an existing runner, execute a case, and inspect artifacts.
Scientific benefit Untested. Comparable policies on independent held-out cases; prospective evidence later.
  • The group’s existing agent family can consume these specifications without a new persona per mode. This page does not assert that its notebook, auditor, or harness has run successfully.
  • Adapter responsibilities: pass source version and mode state to the planner; let the study designer act without inventing a hypothesis; preserve action/claim records; give the assessor claim-appropriate checks; expose mode and case axes in evaluation.
  • First integration check: one nonconfidential toy task must emit a faithful action and claim record. Then complete an evidence-linked historical packet. A toy run is an engineering check, not evidence of discovery skill.
  • Correlated prompting roles are not independent expertise. A separate prompt or agent does not automatically provide independent evidence.

Evaluation that could change the design

  • Compare competent generic, hypothesis-testing, and mode-selecting planners with comparable model, tools, guidance, evidence, budgets, and assistance.
  • Test supplied-mode competence separately from selection/transitions; remove mode organization while preserving substantive guidance in an ablation.
  • Primary outcome is case-level warranted gain the protocol; report scientific magnitude, unsupported claims, cost, novelty, and later impact separately.
  • Historical reference work is not a required path or a clean modern human-speed baseline. A remembered answer is not a newly demonstrated scientific gain.
  • Preserve failures, unknowns, and unanticipated valid gains; evaluate breadth at project, competence, and portfolio scales.
  • A broader literature study can describe reported practices and uptake. It cannot substitute for testing the procedures.