Methods and computation
Drafted human-editable specification · v0.4.0 · no agent-performance claim
Epistemic purpose
- Expand what can be inferred or computed with an inspectable method.
When this mode helps
- The contribution is an algorithm, representation, or computational procedure; using an established method for a new estimate also invokes estimation.
Agent actions and representations
- Specify inputs and assumptions; implement or adapt a method; test known cases and failure conditions; compare baselines; preserve reproducible artifacts and costs.
- Inputs: the question, available evidence and provenance, constraints, prior claims, and remaining budget.
- Outputs: inspectable artifacts, updated question/evidence state, and a claim record when a substantive claim is made.
Evidence and claim scope
- Demonstrate recovery or performance on cases appropriate to the claim, separate development from evaluation, and account for relevant uncertainty and costs. Adoption alone is not correctness.
- Distinguish a warranted decision at the time from a claim’s later assessed adequacy; append follow-up without rewriting the original record.
Transitions and stopping
- Enter estimation or observation when applying the method; use robustness for transfer; reframe the problem if the method solves the wrong target.
- Neighbouring profiles: estimation, observation, robustness, exploration.
Characteristic failure
- Leakage produces apparent performance; synthetic recovery is presented as real-world validity; a new implementation adds no useful capability.
Human evidence and borrowing
- Shapiro and Campillo (2004) and Shapiro et al. (2005) are geoscience method/application cases. Kapoor and Narayanan (2023) discusses leakage in science; Flake and Fried (2020) motivates checking what a score measures.
- Cross-field comparison: Langley (1981) and Kulkarni and Simon (1988) provide historical discovery-system comparisons; their restricted tasks do not establish general scientific autonomy.
- Science of process/impact: Chen et al. (2025) evaluates executable scientific tasks and costs. Task completion is evidence of capability, requiring further assessment to establish advance.
- The actions and transitions above are design hypotheses derived from these sources, not established optimal policies. Missing process chronology remains unknown.
Evaluation against past work
- Show a relevant capability gain against the prior method at comparable cost or explain the tradeoff. Include a data-leakage trap and an out-of-domain failure.
- Use the historical evaluation to choose a source record, preserve evidence boundaries, and assess a meaningful gain. A synthetic control tests a constructed case, not a historical discovery.
- Assessment definitions · Agent implementation