Hypothesis assessment

Drafted human-editable specification · v0.4.1 · no agent-performance claim

Epistemic purpose

  • Discriminate among explanations or stated expectations using informative evidence.

When this mode helps

  • The scientific question concerns alternatives that can have different evidential consequences. Rival generation and discrimination are separate actions.

Agent actions and representations

  • Generate substantive rivals; identify auxiliary assumptions; derive differing predictions; choose a discriminating observation; compare results; revise the rival set when needed.
  • Inputs: the question, available evidence and provenance, constraints, prior claims, and remaining budget.
  • Outputs: inspectable artifacts, updated question/evidence state, and a claim record when a substantive claim is made.

Evidence and claim scope

  • State what each outcome supports and what remains unresolved. Evidence need not exclude every imaginable rival to support a bounded claim; shared assumptions and untested alternatives limit its strength.
  • Distinguish a warranted decision at the time from a claim’s later assessed adequacy; append follow-up without rewriting the original record.

Transitions and stopping

  • Reframe or return to exploration when all rivals accommodate the available evidence; acquire a new observable; use estimation to quantify the predicted contrast.
  • Neighbouring profiles: exploration, theory, estimation, simulation, synthesis.

Characteristic failure

  • Verbal variants masquerade as rivals; all alternatives pass the proposed test; an auxiliary assumption absorbs every failure; a causal claim outruns an association.
  • The absorbing auxiliary is the Duhem problem (Duhem 1954; Quine 1951): a failed prediction refutes the conjunction of hypothesis, auxiliaries, and conditions, and logic does not say which to drop. This mode runs the whole loop of the unit graph; record at node X which conjunct was revised and why.

Human evidence and borrowing

  • Platt (1964) offers a methodological argument; Cleland (2001) and Cleland (2002) distinguish historical and experimental reasoning. Sykes (1967) is the primary tectonic test case, not a controlled evaluation of agents.
  • Cross-field comparison: King et al. (2009) illustrates executable hypothesis testing in yeast with assay constraints; borrowing requires identifying which geoscience observations can play the comparable discriminating role.
  • Science of process/impact: Si et al. (2025) distinguishes idea ratings from executed research outcomes in NLP. Its results motivate testing rival quality through consequences rather than fluent phrasing.
  • The actions and transitions above are design hypotheses derived from these sources, not established optimal policies. Missing process chronology remains unknown.

Evaluation against past work

  • Resolve a scientifically important contrast or demonstrate why current evidence cannot. Include distinguishable rivals, equifinal rivals, and a missing-data case; score gain and unsupported exclusions.
  • Use the historical evaluation to choose a source record, preserve evidence boundaries, and assess a meaningful gain. A synthetic control tests a constructed case, not a historical discovery.
  • Assessment definitions · Agent implementation