Replication and robustness assessment
Drafted human-editable specification · v0.4.0 · no agent-performance claim
Epistemic purpose
- Establish what survives repetition, perturbation, or transfer, and correct consequential errors.
When this mode helps
- The question concerns reliability, domain of applicability, or a challenged result. Informative failures can advance knowledge.
Agent actions and representations
- Reproduce analyses; change relevant assumptions; use independent observations; test transfer; localize disagreement; update claim scope and preserve earlier assessments.
- Inputs: the question, available evidence and provenance, constraints, prior claims, and remaining budget.
- Outputs: inspectable artifacts, updated question/evidence state, and a claim record when a substantive claim is made.
Evidence and claim scope
- State what was repeated and what was independent, with sensitivity, uncertainty, and conditions of failure. Reproducible code alone does not establish empirical adequacy.
- Distinguish a warranted decision at the time from a claim’s later assessed adequacy; append follow-up without rewriting the original record.
Transitions and stopping
- Return to the originating mode with revised scope; obtain better observations when disagreement is unresolvable; reframe after a consequential correction.
- Neighbouring profiles: instruments, methods, estimation, hypothesis.
Characteristic failure
- Treat agreement as proof; repeat a shared bias; call every difference a refutation; dismiss replication as unoriginal.
Human evidence and borrowing
- Lindsey et al. (2020) supplies measurement assessment; Beven and Freer (2001) motivates sensitivity to equifinality. Nosek et al. (2018) separates planned and unplanned analyses; it does not ban exploratory actions.
- Cross-field comparison: Leeman et al. (2024) supplies a materials-science reanalysis, whose conclusions should be attributed and compared with original evidence; it is not a blanket verdict on laboratory agents.
- Science of process/impact: Petersen et al. (2025) challenges interpretation of an impact indicator. Flake and Fried (2020) explains why measurement validity needs attention beyond a repeatable score.
- The actions and transitions above are design hypotheses derived from these sources, not established optimal policies. Missing process chronology remains unknown.
Evaluation against past work
- Demonstrate a corrected conclusion, narrower domain, or meaningful independent support. Include shared systematic errors and a legitimate failure to transfer.
- Use the historical evaluation to choose a source record, preserve evidence boundaries, and assess a meaningful gain. A synthetic control tests a constructed case, not a historical discovery.
- Assessment definitions · Agent implementation