Hypothesis testing when several models fit

Meeting 9 of 13

Drafted v0.4 · not yet held · record discussion

Central question

What does a successful model test establish, and what should we do when incompatible models fit the same observations?

Anchor and companion

  • Everyone reads the selected anchor sections and the companion abstract/overview. The rotating reader presents the companion in depth.
  • Pairing: Limits of validation / practical model comparison. Full records and access notes: bibliography.

Anchor

Oreskes, N., Shrader-Frechette, K., & Belitz, K. (1994). Verification, Validation, and Confirmation of Numerical Models in the Earth Sciences. Science, 263(5147), 641–646. https://doi.org/10.1126/science.263.5147.641

Read Oreskes and colleagues in full.

Companion

Beven, K., & Freer, J. (2001). Equifinality, data assimilation, and uncertainty estimation in mechanistic modelling of complex environmental systems using the GLUE methodology. Journal of Hydrology, 249(1–4), 11–29. https://doi.org/10.1016/S0022-1694(01)00421-8

Read Beven and Freer’s equifinality argument and environmental-model example; inspect how acceptable models are retained.

Why these readings belong together

The first paper challenges unrestricted truth claims about numerical models of open systems. The second develops an approach to multiple acceptable parameterizations and structures. This is not a claim that software cannot be tested or that predictive skill cannot be assessed. (Oreskes et al. 1994; Beven and Freer 2001)

Prepare and discuss

Read the selected anchor sections and the companion abstract or overview. Bring one source passage or artifact relevant to the case; the rotating reader presents the companion in depth.

  1. What has been tested: code, numerical accuracy, a parameterization, or a causal explanation?
  2. Which independent observable distinguishes the acceptable models?
  3. How do likelihood choices, thresholds, and model inadequacy enter the uncertainty claim?

Case exercise: scientific gain and epistemic mode

Distinguish an estimated quantity, a prediction, and an explanation in the model case. Compare assumptions and equifinal alternatives; state the meaningful gain and a supported limit rather than claiming complete validation.

  • Profiles to examine: estimation, simulation, hypothesis.
  • Record source kind and unknown chronology; distinguish documented practice, philosophical argument, association, and our proposed agent rule.
  • Ask what was gained, what warrants it, which action helped, and what a comparable agent would need to demonstrate.

Shared output

One episode entry and three sentences: what changed scientifically; what evidence warrants that account; what agent action or evaluation follows. Longer worksheets are optional.

The agent

  • After the meeting, say what these papers change in the unit skill, agents/skills/unit-of-inquiry/SKILL.md: a rule at a node, a new composition of units, or a planted flaw. Meeting 1’s page shows the form.
  • Modes exercised here, whose “agent actions” sections the rule would enter: estimation, simulation, hypothesis.
  • Written from the texts after the papers are read in full, as for meeting 1; nothing here yet.

Optional extensions

Record after the meeting

Sign and date the notes; preserve disagreements and what changed your assessment. Keep confidential examples in private notes.