Methods and computation

Drafted human-editable specification · v0.4.0 · no agent-performance claim

Epistemic purpose

  • Expand what can be inferred or computed with an inspectable method.

When this mode helps

  • The contribution is an algorithm, representation, or computational procedure; using an established method for a new estimate also invokes estimation.

Agent actions and representations

  • Specify inputs and assumptions; implement or adapt a method; test known cases and failure conditions; compare baselines; preserve reproducible artifacts and costs.
  • Inputs: the question, available evidence and provenance, constraints, prior claims, and remaining budget.
  • Outputs: inspectable artifacts, updated question/evidence state, and a claim record when a substantive claim is made.

Evidence and claim scope

  • Demonstrate recovery or performance on cases appropriate to the claim, separate development from evaluation, and account for relevant uncertainty and costs. Adoption alone is not correctness.
  • Distinguish a warranted decision at the time from a claim’s later assessed adequacy; append follow-up without rewriting the original record.

Transitions and stopping

  • Enter estimation or observation when applying the method; use robustness for transfer; reframe the problem if the method solves the wrong target.
  • Neighbouring profiles: estimation, observation, robustness, exploration.

Characteristic failure

  • Leakage produces apparent performance; synthetic recovery is presented as real-world validity; a new implementation adds no useful capability.

Human evidence and borrowing

  • Shapiro and Campillo (2004) and Shapiro et al. (2005) are geoscience method/application cases. Kapoor and Narayanan (2023) discusses leakage in science; Flake and Fried (2020) motivates checking what a score measures.
  • Cross-field comparison: Langley (1981) and Kulkarni and Simon (1988) provide historical discovery-system comparisons; their restricted tasks do not establish general scientific autonomy.
  • Science of process/impact: Chen et al. (2025) evaluates executable scientific tasks and costs. Task completion is evidence of capability, requiring further assessment to establish advance.
  • The actions and transitions above are design hypotheses derived from these sources, not established optimal policies. Missing process chronology remains unknown.

Evaluation against past work

  • Show a relevant capability gain against the prior method at comparable cost or explain the tradeoff. Include a data-leakage trap and an out-of-domain failure.
  • Use the historical evaluation to choose a source record, preserve evidence boundaries, and assess a meaningful gain. A synthetic control tests a constructed case, not a historical discovery.
  • Assessment definitions · Agent implementation