The unit of inquiry: reasoning, hypotheses, tests

Drafted module 0 · read before meeting 1 · the verbs every mode page uses · not yet discussed

0. Three things people mean by “the scientific method”

The textbook method The AI discovery loop The earth sciences’ own account
Where it is stated Science texts “from grade school through college,” which Cleland says “invariably provide one (or a combination) of two accounts, scientific inductivism or falsificationism” (Cleland 2001, 987); the school version taken apart by Windschitl et al. (2008) “Closed-the-loop of hypothesis generation, experimental planning, and execution” (Kitano 2021, 1); a system “mirroring the reasoning process underpinning the scientific method” that, “given a research goal specified in natural language,” proposes hypotheses and protocols (Gottweis et al. 2026); “active learning closes the loop by proposing improved follow-up recipes” (Szymanski et al. 2023, 86); a system that “autonomously designs, plans and performs complex experiments” (Boiko et al. 2023, 570); Adam, the first closed loop (King et al. 2009) Chamberlin’s multiple working hypotheses, written by a geologist against the “ruling theory” (Chamberlin 1965); Gilbert on where a geological hypothesis comes from (Gilbert 1886, 1896); Cleland on historical science (Cleland 2001, 2002); Frodeman on geology as interpretive and historical (Frodeman 1995)
The steps Observe, ask, hypothesize, predict, test, conclude; one hypothesis at a time; confirmation accumulates or a failure refutes Goal in; generate hypotheses or candidates; choose the next experiment by a score; execute; analyze; repeat until the goal is met Observe puzzling traces; formulate incompatible rivals; hunt for the trace that discriminates; retain the best explanation; expect it to be deposed
Where it starts (node Q) An observation, then a question A goal supplied from outside: six tasks (Boiko et al. 2023), a target list (Szymanski et al. 2023), a research goal (Gottweis et al. 2026) Puzzling traces of a past event (Cleland 2002, 480)
Hypotheses at G One Many, ranked by a score; rarely incompatible rivals Several, incompatible (Chamberlin 1965; Cleland 2002, 483–84)
Conditions at C Set Set, by a robot or a computation Found, by fieldwork (Cleland 2002, 484)
The control series at V Absent from the diagram; present in the laboratory Mostly absent: a positive-only verifier scored itself (Leeman et al. 2024) Absent by necessity; replaced by independent traces (Cleland 2002, 491)
Decision at X or K Refute or confirm A score threshold Best explanation within the stated set (Cleland 2002, 481–82)
What it leaves out Auxiliaries, the series, rival generation, where questions come from (Cleland 2001, 987–88; Windschitl et al. 2008) Choosing the question; enumerating rivals; abstaining (Leeman et al. 2024) Manipulation; hence the method of difference
  • The textbook version is Platt’s loop with one branch: he called that “The Great Man With a Single Hypothesis” (Platt 1964, 350). Chamberlin’s remedy is what makes it strong.
  • The AI loop is Platt’s loop with node G automated and node Q supplied. Kitano’s own example iterates “to verify or reject h1” against “h2, h3, and h4” (Kitano 2021, 4), so the design has room for rivals; the published systems mostly optimize a single target.
  • The earth sciences run both patterns. Where nature repeats herself, a catalog of events supports regularities the way an experimental program would, minus the power to set the condition (Cleland 2002, 485); where an event is unique, the trace pattern rules. That is why one word will not do, and why the ledger records a node letter instead.

Where inquiry starts

  • The readings give six entry states for node Q. They are not rival theories; they are what different modes begin with, and the ledger’s “question origin” column records which.
Entry state Who says so Example in the reading
A surprising observation or anomaly Peirce, the surprising fact (Peirce 1992); Cleland’s prototype, “puzzling traces” (Cleland 2002, 480); Platt’s own description of practice, where the logical tree on the blackboard begins with “the hot new result just up from the laboratory” (Platt 1964, 348); Klahr and Simon: “a phenomenon led to a hypothesis, rather than a hypothesis to experimental phenomena,” and this “is not a singular case in scientific history but a frequent occurrence” (Klahr and Simon 1999, 537) Becquerel’s plate; Cech’s catalytic products; the iridium layer (Cleland 2002, 482, 486); the Curies’ pitchblende, more radioactive than pure uranium (Klahr and Simon 1999, 537)
An unsolved problem or question Popper, science begins with problems (Popper 2005); Platt’s prescription, “every problem in science” (Platt 1964, 347); Gilbert’s topographic problem (Gilbert 1896) Coon Butte; the origin of the dinosaurs’ extinction before 1980 (Cleland 2001, 988)
A hypothesis or a theory’s prediction Dicke predicting the background radiation (Cleland 2001, 990); the Viking designers (Cleland 2002, 479); Vine and Matthews (Vine and Matthews 1963) Platt’s steps 1 and 2 presuppose this state
A body of data without a theory The hypothesis the BACON programs were built to test: “especially in new fields of science, where theory is poorly developed or absent, observation typically precedes hypothesis construction,” and “the first step in progress is to find patterns (laws) in data” (Klahr and Simon 1999, 537; Langley et al. 1987) Kepler’s third law; Balmer’s formula, found by a geometry teacher “thoroughly innocent of the physics”; Planck’s revision of Wien’s law in one evening, “pure numerology” (Klahr and Simon 1999, 534–36)
A new capability Cleland: finding attenuated traces “may require advances in technology,” antennas and the cyclotron (Cleland 2001, 990); Hacking’s playing with equipment, via Cleland (2002), p. 486; Klahr and Simon: new representations come from “new ways of seeing experiments triggered by new instruments of observation” (Klahr and Simon 1999, 539) Dark-fibre sensing, meeting 5 (Lindsey et al. 2019); measuring the volumes of gases, which turned phlogiston into oxygen; Muller’s x-rays, which raised the mutation rate of breeding experiments (Klahr and Simon 1999, 539)
A goal supplied by someone else Every AI system in the prior-art chapter; funders, in Platt’s closing paragraph (Platt 1964, 352) Six chemistry tasks; fifty-seven target compounds (Boiko et al. 2023; Szymanski et al. 2023)
  • Platt’s description and his prescription start in different places: the blackboards begin with a result, the four steps begin with hypotheses already devised. The gap between them is the exploratory phase Cleland describes and Platt does not (Cleland 2002, 486).
  • The data-without-theory state differs from the anomaly state in what it lacks. An anomaly needs an expectation to violate; a table of numbers with no theory has none, and the work is pattern-finding with feedback from each misfit. It enters the graph before G, and what it retains at K is a regularity, not a mechanism: Planck’s formula is still accepted, while the physical explanation he spent the following months on “bears little relation to today’s quantum mechanics” (Klahr and Simon 1999, 534).

1. Three cognitive tasks

  • Peirce’s framing makes the three symmetric: each rearranges a rule, a case, and a result (Peirce 1992, originally 1878).
Task Form What it guarantees Geoscience example
Deduction Rule + case, therefore result The result is certain if the premises are true; nothing new is added If ridge-ridge transforms exist (rule) and this offset is one (case), first motions on the offset segment are opposite to the ridge-offset sense (result) (Wilson 1965)
Induction Case + result, therefore rule The rule is probable, never certain; the next case can break it Twenty catalogs (cases) each give b near 1 (results), therefore b near 1 in general (rule)
Abduction Rule + result, therefore case The case is the best available explanation, and only that Peat-mud couplets with sand sheets (result) and what subduction earthquakes do to coasts (rule), therefore a great earthquake (case) (Atwater 1987)
  • Deduction preserves truth and adds no content. Induction and abduction add content and can be wrong.
  • Aristotle named the first two. Peirce named the third and made it the logic of discovery; Hanson built the case that observation itself is shaped by it (Hanson 1958).

2. Inference, hypothesis, prediction, explanation

  • Inference: any reasoned move from premises or evidence to a conclusion. The umbrella term; the three tasks above are its kinds.
  • Hypothesis: a candidate claim, not yet established, with consequences that could be checked. How the word acquired that meaning: Glass and Hall (2008). The working hypothesis against the ruling theory: Chamberlin (1965), first printed in 1890. Where a geological hypothesis comes from, by analogy and then enumeration: Gilbert (1886); Gilbert (1896).
  • Prediction versus explanation: a prediction is a consequence derived before the observation; an explanation is a hypothesis fitted to an observation already in hand. Whether the order matters for how much the evidence counts is a live dispute (Musgrave 1974; Douglas and Magnus 2013).

3. The unit: one pass through the schema

\[H + A + C \Rightarrow E\]

  • \(H\): the hypothesis under test.

  • \(A\): the auxiliary assumptions, such as the theory of the instrument, the background laws, the dating method.

  • \(C\): the conditions, what the system was in, either set by the investigator or found in nature.

  • \(E\): the expected observation.

  • The schema is in the assigned reading. Cleland writes it as a test implication inferred from \(H\): if condition \(C\) is brought about, event \(E\) will occur (Cleland 2001, 987). The auxiliaries \(A\) enter one page later, as the “enormous number of auxiliary assumptions about equipment and background conditions” any real test carries (Cleland 2001, 988).

  • The verbs are defined on the schema. Each has an owner who fixed its meaning.

Verb On the schema Who fixed the meaning
Test Arrange or find \(C\), derive \(E\), observe Hempel (1945)
Confirm Observe \(E\); the conjunction survives; strength depends on how unlikely \(E\) was without \(H\) Hempel (1945); the Bayesian form is \(P(E \mid H) / P(E)\)
Corroborate Survive a test that could have failed, without calling that confirmation Popper (2005), first published 1934
Falsify Observe \(\neg E\); therefore \(\neg(H + A + C)\); logic does not say which conjunct to drop Popper (2005); Duhem (1954); Quine (1951); stated for geologists in Cleland (2001), p. 988
Control Hold \(C\) and vary \(A\) after a failure, against false negatives; repeat after a success, against false positives; remove \(C\) to see whether it was needed Mill’s method of difference (Mill 2011); Cleland (2001), p. 988; Cleland (2002), pp. 477–478
Exclude Falsify one member of a stated set of rivals Platt (1964), after Bacon’s instances of the fingerpost (Bacon 2004)
Infer to the best explanation Retain the \(H\) that makes \(E\) least surprising among the rivals; a retention move, not an exclusion Harman (1965); Lipton (2004)
Severely test A test the \(H\) would probably have failed if it were false Mayo (1996); Mayo and Spanos (2006)
  • The Bayesian line ties them together: \(P(H \mid E) \propto P(E \mid H)\,P(H)\). Cleland suspects her differences between historical and experimental science can be rewritten this way (Cleland 2002, 495); Tarantola (2006) is the geophysicist’s version.

4. The thinkers, and the node each owns

  • Dates are the original publications. Citations point at the editions in the bibliography.
Year Thinker Contribution Node on the graph
1620 Bacon (Bacon 2004) Eliminative induction; the instance of the fingerpost that decides between roads Design, Exclude
1840 Whewell (Whewell 2014) Hypotheses come first; consilience when independent classes of facts agree Generate, Retain
1843 Mill (Mill 2011) Methods of agreement and difference: what controlling \(C\) means Conditions
1878 Peirce (Peirce 1992) Abduction as the third task Generate
1886 Gilbert (Gilbert 1886) Hypotheses by analogy, in geology Generate
1890 Chamberlin (Chamberlin 1965) Hold a family of hypotheses at once Generate
1906 Duhem (Duhem 1954) No crucial experiment in physics; refutation hits the conjunction Exclude, Protect
1934 Popper (Popper 2005) Falsification; the asymmetry between refuting and confirming Exclude
1945 Hempel (Hempel 1945) The logic of confirmation and its paradoxes Retain
1951 Quine (Quine 1951) Duhem’s holism generalized to all belief Protect
1958 Hanson (Hanson 1958) Discovery has a logic; observation is theory-laden Generate, Observe
1964 Platt (Platt 1964) Strong inference The loop
1965 Harman (Harman 1965) Inference to the best explanation, named Retain
1970 Lakatos (Lakatos 1970) Protective belt; research programmes Protect
1988 Klahr and Dunbar, via Klahr and Simon (1999), p. 538 Discovery as search in two spaces, hypotheses and experiments; theorists and experimenters Generate, Design
1996 Mayo (Mayo 1996) Severity as the measure of a test Design
1999 Klahr and Simon (Klahr and Simon 1999) Discovery as heuristic search in several spaces; surprise as a heuristic for the next step; the starting state as knowledge plus expectation Q, the surprise edge from M
2001 Cleland (Cleland 2001) Historical versus experimental: the asymmetry of overdetermination Conditions, Retain

5. The graph

  • One unit as a workflow. Ten working nodes and one escape. Every thinker above sits on one of them. The meeting 1 page places Platt and Cleland on it.
%%{init: {"theme": "base", "themeVariables": {"fontSize": "17px", "fontFamily": "Source Serif 4, Georgia, serif", "primaryColor": "#e7f0f2", "primaryBorderColor": "#1c6a7f", "primaryTextColor": "#153742", "lineColor": "#49616a", "secondaryColor": "#ffffff", "tertiaryColor": "#f5f4f2", "edgeLabelBackground": "#f5f4f2"}, "flowchart": {"htmlLabels": true, "nodeSpacing": 26, "rankSpacing": 40, "curve": "basis", "padding": 10}}}%%
flowchart TD
    Q["Q. Question or anomaly"] --> G["G. Generate rivals H1 … Hn<br/><b>abduction</b><br/><small>Peirce · Gilbert · Chamberlin</small>"]
    G --> D["D. Derive E_i from H_i + A + C<br/><b>deduction</b><br/><small>Hempel · Cleland</small>"]
    D --> S["S. Design: choose the C and the observable<br/>where the E_i differ<br/><b>severe test</b><br/><small>Bacon's fingerpost · Mill · Mayo</small>"]
    S --> C["C. Set the conditions (experiment)<br/>or find them (trace)"]
    C --> O["O. Observe<br/><small>theory-laden · Hanson</small>"]
    O --> M{"M. Match E_i?"}
    M --> V["V. <b>Control</b> the series: hold C and vary A;<br/>repeat after a success; remove C<br/><small>Mill · Cleland</small>"]
    M -. "surprise: no E_i matches<br/>find its scope, then its mechanism<br/><small>Klahr · Simon</small>" .-> Q
    V -- "run again" --> C
    V -- "not E_i survives the series" --> X["X. Exclude H_i<br/><b>falsify</b> · eliminative <b>induction</b><br/><small>Popper · Platt</small>"]
    V -- "E_i survives the series" --> K["K. Retain H_i<br/><b>induction</b>: confirm, corroborate<br/>or <b>best explanation</b><br/><small>Hempel · Popper · Harman</small>"]
    X -. "or protect H_i" .-> P["P. Revise A or C instead<br/><small>Duhem · Quine · Lakatos</small>"]
    P -.-> D
    X --> R["R. Refine the survivors<br/>into subhypotheses"]
    K --> R
    R --> G

    classDef step fill:#e7f0f2,stroke:#1c6a7f,stroke-width:1.5px,color:#153742;
    classDef branch fill:#ffffff,stroke:#1c6a7f,stroke-width:2px,color:#153742;
    classDef escape fill:#f6e9e2,stroke:#a1512b,stroke-width:1.5px,color:#5a2e17;
    class Q,G,D,S,C,O,V,X,K,R step;
    class M branch;
    class P escape;
Figure 1: The unit of inquiry as a workflow. Solid arrows are one pass; V is the control series Cleland says is the real unit of experimental work, and its loop back to C is where most laboratory time goes; the dashed arrow from X is the Duhem escape, revising an auxiliary or a condition instead of the hypothesis; the dashed arrow from M is surprise, an outcome no rival predicted, which sends the search back to establish the new phenomenon and reframe the question. Node letters are used in the tables below and in every meeting’s ledger row.
  • The surprise edge (graph v0.3). A mismatch with one \(E_i\) goes to V and then to X or P. An outcome that matches none of the \(E_i\) is a different event: the rival set was wrong, not one rival. Klahr and Simon document the move in history, in models, and in the laboratory: “in the face of surprise, scientists frequently divert the path of exploration to ascertain the scope and import of the surprising phenomenon and to determine its mechanism” (Klahr and Simon 1999, 537). Krebs’s ornithine yield and Faraday’s transient current both went this way, and KEKADA uses surprise as the heuristic that chooses its next experiment (Klahr and Simon 1999, 535, 537). The edge re-enters at Q as the anomaly entry state, which first establishes that the surprise is a phenomenon.
  • The graph as search in several spaces. Klahr and Simon’s spaces sit on the nodes as follows (Klahr and Simon 1999, 538–39).
Search space Node What is chosen there
Hypotheses G, R The rivals and their refinements; ordered by plausibility unless a rule says otherwise
Experimental paradigms S Which factors vary and which are held, the class of experiment
Experiments S, C The settings within the paradigm
Data representations O, and before G The features in which the phenomenon is described; where Monod’s activation became inhibition and Faraday’s currents became lines of force (Klahr and Simon 1999, 533–34, 539)
Instruments The capability entry state; O What can be observed at all; the gas measurements behind the oxygen theory (Klahr and Simon 1999, 539)
Strategies The method choice, section 6 Which composition of units is run; Klahr and Dunbar’s theorists search G first, their experimenters S first (Klahr and Simon 1999, 538)

6. Composing units into a mode

  • The graph above is one unit. A mode of inquiry is a way of composing units, and the compositions are read off the papers, not invented. Two are defined by the meeting 1 readings; a third follows from them; later meetings add their own, with provenance, in the log below.
Composition Structure Forced by
Parallel, shared test \(n\) units share one S, C, O, and M: one crucial experiment whose outcomes differ across the rivals; every outcome excludes at least one; survivors go to R and back to G. Platt’s strong inference Platt (1964), p. 347
Trace series C found; O repeated over independent traces stands in for V over varied A; K decided by best explanation within the stated rival set Cleland (2002), pp. 484, 487–491
Nature’s repetition C found but repeated by nature; a regularity is fitted across the repetitions; no power to set or remove C Cleland (2002), p. 485
  • Which composition applies is decided by two state variables: whether C is set or found, and whether the event is unique or repeated. That decision is what a mode page’s “when this mode helps” section records.
  • The meeting 1 page draws the first composition with the unit’s own letters. The seminar’s longer aim is the fourth row and onward: a geoscientific workflow composed from units, one composition per meeting, until the ledger’s episodes can each be written as a path.

7. AI agents: how we build one from the graph

  • The seminar’s agents are not personas. An agent here is a controller that runs the graph, a set of rubrics it is held to, and references that explain the rubrics to the people who revise them. The three are kept in separate files and the controller cites the rubrics by version (agents/README.md).

Three layers

Layer What it holds Who edits it What loads into the agent
Rubrics The graph on this page; the mode pages, which say what evidence each kind of claim owes; the claim record it must emit The group, one meeting at a time, with the paper that forced each change On demand, at the node that needs it; the mode pages are exported to agents/mode-instructions.json
Controller A skill: one line per node, do, produce, move on, do not; a decision table that picks the method from the situation; the outputs Derived from the rubrics after each meeting In full, when the skill fires: agents/skills/unit-of-inquiry/SKILL.md
References The reasons and sources behind every rule; the methods with their guards; a worked example; the evaluator’s planted-flaw items; the change log With the rubrics Not by the acting agent; the evaluator loads its items, people load the rest
  • The rule that keeps it lean: one fact lives in one place, and the others cite it by version. A controller that grows past about a thousand words is split.
  • The evaluator’s items are never in the acting agent’s context. An agent that holds its own grading rubric scores itself, which is what the A-Lab reanalysis found (Leeman et al. 2024).

Defining a scientific method

  • A method is a composition of the unit (section 6). To define one, the group writes one row: how the composition runs the graph, its guard, and the paper that defined it. That row is the rubric.
  • The controller then gets two things from the row: an entry in its method table, and the state test that selects it, which uses the same variables for every method: can the conditions be set, does nature repeat the event, is the phenomenon stable.
  • The glossary supplies the second half of the rubric. A method says how the nodes are composed; a mode page says what a claim at K owes given the kind of work. Strong inference in the hypothesis mode owes a discriminating outcome; nature’s repetition in the regularities mode owes out-of-sample fit and a stated range.
  • Four methods are defined so far, from the meeting 1 readings. Each later meeting defines the one its papers document, in the same row form, with provenance in the graph’s log.

Using a scientific method

  • A run starts from a trigger, one of the six entry states of section 0, and a stated situation. The controller enters the graph at the node the trigger implies, chooses the method from the state test unless a person names one, and moves node by node.
  • At each node it produces the artifact the node owes and refuses the moves the node forbids. At K it writes a claim record with a status. At any node after M it may abstain with the observation that would decide.
  • The run’s record is the node-labelled action log, the method declaration with the state that chose it, and the claim records. The agent’s own rationale is not the record.
  • Terminal states are four: a warranted claim, a principled abstention, a reframed question, an artifact verdict. Two of the four are not discoveries, and the evaluation scores them as correct where the evidence warranted them.

What the group keeps for itself

  • The rubrics. Nothing in the controller is edited except by shortening a rubric the group has changed.
  • The verification contract: for every delegated action, the check a non-specialist can run to see that it was done wrong. It is written into the claim record’s evidence fields, not into the agent’s confidence.
  • The separation of powers: the controller that acts, the evaluator that scores, and the notebook that records are different components. The group’s own agent family already has that shape; the seminar changes what its designer can plan and what its auditor must ask, not its structure (meeting 13).

Where the seminar’s outputs enter the agent

Seminar output Becomes Where
A revised node rule One line in the controller; the full rule with its source in the references After the meeting that changed it
A new composition A row in section 6; a method entry and a state test in the controller The meeting whose papers document it
A mode page’s evidence column What a claim record must show at K for that mode Exported automatically
A ledger episode A historical item the evaluator can score, on process at the date Historical evaluation
A failure mode from the reading A planted-flaw item for the evaluator, never a line the agent sees references/evaluation.md

8. Using the graph

  • Meeting 1: place one recent paper of your own on the graph; mark the nodes it performed, the nodes it skipped, and whether \(C\) was set or found. Then place Platt and Cleland.
  • Every later meeting: the ledger’s “decisive move” names a node letter. The mode specifications say which nodes each mode exercises; hypothesis assessment is the mode that runs the whole loop.
  • Discussion question 3 of meeting 1, whether the papers describe how ideas are generated or how claims are warranted, becomes a pointing exercise: G against K.

The graph is a living object

  • It grows as the meetings read new papers. A node or an edge is added when a paper documents a move the graph cannot place; nothing is added for a move it already places.
  • Every change is logged here with the meeting and the paper that forced it, so the graph’s version can be cited beside a ledger row.
Version Date Change Forced by
0.1 September 16, 2026 Nodes Q, G, D, S, C, O, M, X, K, P, R Platt (1964); Cleland (2001)
0.2 September 16, 2026 Node V, the control series, between M and X or K, with its loop back to C Cleland (2002), pp. 477–478, on rereading the full text
0.3 September 23, 2026 The surprise edge, M to Q, for an outcome no rival predicted; a sixth entry state, data without a theory; the starting state and the search spaces placed on the nodes Klahr and Simon (1999), pp. 532–540, read in full for the meeting 1 follow-up