Hypothesis testing when several models fit

Meeting 9 | Tuesday, November 3, 2026 | 1:00-2:00 PM Pacific

Central question

What does a successful model test establish, and what should we do when incompatible models fit the same observations?

The paper pair

Pairing type: Limits of validation / practical model comparison.

Paper A: Verification, Validation, and Confirmation of Numerical Models in the Earth Sciences

  1. Science (1994). Paper / publisher record

Read Oreskes and colleagues in full.

Access: Publisher/DOI page; full text may require institutional access.

Paper B: Equifinality, data assimilation, and uncertainty estimation in mechanistic modelling of complex environmental systems using the GLUE methodology

  1. Journal of Hydrology (2001). Paper / publisher record

Read Beven and Freer’s equifinality argument and environmental-model example; inspect how acceptable models are retained.

Access: Publisher/DOI page; full text may require institutional access.

Why these papers belong together

The first paper challenges unrestricted truth claims about numerical models of open systems. The second develops an approach to multiple acceptable parameterizations and structures. This is not a claim that software cannot be tested or that predictive skill cannot be assessed. (Oreskes et al. 1994; Beven and Freer 2001)

Prepare before the meeting

Bring two plausible explanations or model configurations for one mature research problem. List shared and differing predictions.

Discussion questions

  1. What has been tested: code, numerical accuracy, a parameterization, or a causal explanation?
  2. Which independent observable distinguishes the acceptable models?
  3. How do likelihood choices, thresholds, and model inadequacy enter the uncertainty claim?

One-hour meeting

Time Activity
0-10 min Independent first judgments; surface disagreements.
10-25 min Compare the papers: claim, evidence, assumptions, and limits.
25-45 min Work through the case exercise below.
45-55 min Translate the discussion into agent requirements and tests.
55-60 min Record an output and one unresolved disagreement.

Case exercise

Compare an exploration-first workflow with explicit competing models. Define success before seeing their outputs. Use a held-out observable or regime rather than reward an increasingly flexible fit to the same data. Treat GLUE weighting choices explicitly rather than as uniquely determined Bayesian probabilities.

Agent-design or evaluation output

Inquiry-mode routing policy and model-test card: claim, alternatives, measurement assumptions, acceptance criterion, held-out evidence, and conditions for model revision.

Record after the meeting

Record the evidence for your main claim, what changed your mind, what remains unresolved, and one change to the agent or its evaluation. Keep confidential examples in private group notes rather than committing them to this public book.