%%{init: {"theme": "base", "themeVariables": {"fontSize": "17px", "fontFamily": "Source Serif 4, Georgia, serif", "primaryColor": "#e7f0f2", "primaryBorderColor": "#1c6a7f", "primaryTextColor": "#153742", "lineColor": "#49616a", "secondaryColor": "#ffffff", "tertiaryColor": "#f5f4f2", "edgeLabelBackground": "#f5f4f2"}, "flowchart": {"htmlLabels": true, "nodeSpacing": 22, "rankSpacing": 38, "curve": "basis", "padding": 10}}}%%
flowchart TD
Q["Q. Problem"] --> G["G. Generate rivals H_1, H_2, H_3<br/><b>abduction</b><br/><small>Chamberlin · Platt step 1</small>"]
G --> D1["D. Derive E_1<br/>from H_1 + A + C<br/><b>deduction</b>"]
G --> D2["D. Derive E_2<br/>from H_2 + A + C<br/><b>deduction</b>"]
G --> D3["D. Derive E_3<br/>from H_3 + A + C<br/><b>deduction</b>"]
D1 & D2 & D3 --> S["S. One crucial experiment: choose C and the observable<br/>where E_1, E_2, E_3 differ<br/><b>severe test</b><br/><small>Bacon's fingerpost · Platt step 2</small>"]
S --> C["C. Set the conditions"]
C --> O["O. Observe<br/><small>Platt step 3, 'a clean result'</small>"]
O -.-> V["V. Control the series<br/><small>not described by Platt</small>"]
V -.-> M
O --> M{"M. Which E_i?"}
M -- "E_1" --> R1["X. Exclude H_2, H_3 · K. Retain H_1<br/><b>falsify</b> · eliminative <b>induction</b>"]
M -- "E_2" --> R2["X. Exclude H_1, H_3 · K. Retain H_2"]
M -- "E_3" --> R3["X. Exclude H_1, H_2 · K. Retain H_3"]
M -- "no clean separation" --> NC["No exclusion: redesign S<br/><small>Platt, p. 347: 'insecure and must be rechecked'</small>"]
NC --> S
R1 & R2 & R3 --> R["R. Refine the survivor into subhypotheses<br/><small>Platt step 1'</small>"]
R --> G
classDef step fill:#e7f0f2,stroke:#1c6a7f,stroke-width:1.5px,color:#153742;
classDef unit fill:#ffffff,stroke:#1c6a7f,stroke-width:1.5px,color:#153742;
classDef branch fill:#ffffff,stroke:#1c6a7f,stroke-width:2px,color:#153742;
classDef escape fill:#f6e9e2,stroke:#a1512b,stroke-width:1.5px,color:#5a2e17;
classDef ghost fill:#f5f4f2,stroke:#49616a,stroke-width:1px,stroke-dasharray:4 3,color:#49616a;
class Q,G,S,C,O,R step;
class D1,D2,D3,R1,R2,R3 unit;
class M branch;
class NC escape;
class V ghost;
Is there a scientific method?
Meeting 1 of 13
Discussed v0.4 · held September 15, 2026 · record discussion
Central question
Is hypothesis-testing a universal account of inquiry, a standard for justification, or one useful mode among several?
Anchor and companion
- Everyone reads the selected anchor sections and the companion abstract/overview. The rotating reader presents the companion in depth.
- Pairing: Methodological ideal / evidential counterpoint. Full records and access notes: bibliography.
Anchor
Platt, J. R. (1964). Strong Inference. Science, 146(3642), 347–353. https://doi.org/10.1126/science.146.3642.347
Read Platt in full. Identify the role of alternative explanations and discriminating observations.
Companion
Cleland, C. E. (2001). Historical science, experimental science, and the scientific method. Geology, 29(11), 987–990. https://doi.org/10.1130/0091-7613(2001)029%3C0987:HSESAT%3E2.0.CO;2
Read Cleland’s Geology paper in full. It makes the argument in four pages, for geologists.
Why these readings belong together
Platt offers a prescriptive strategy; Cleland analyzes differences between experimental interventions and reconstructing past events. Neither paper directly records this group’s daily research. Test both against an actual project rather than declaring one a complete scientific method. (Platt 1964; Cleland 2001)
Where the two papers sit on the unit graph
- Read the unit of inquiry first. It defines deduction, induction, abduction, hypothesis, and the verbs test, confirm, exclude, and corroborate, and draws one unit of hypothesis testing as a workflow of nine nodes, Q to R.
- Both papers are overlays on that graph. Neither adds a node.
Platt
- Strong inference is the parallel composition of unit page, section 6: \(n\) units with their own D share one S, C, O, and M; the figure below draws it with the unit’s letters (Platt 1964, 347).
- What makes it strong is one property of node S: every outcome at M excludes at least one rival. The tree branches and prunes on each pass.
- Platt’s step 3, “carrying out the experiment so as to get a clean result” (Platt 1964, 347), is where node V hides. He never describes the control series; Cleland says it is most of the work (Cleland 2002, 477–78).
- Its parts are older than Platt: Bacon’s eliminative induction at S and X, Chamberlin’s family at G, Popper’s asymmetry at X (Platt 1964, 349–50).
- Platt supplies nothing at G beyond “invent” and sends the reader to Pólya (Platt 1964, 347). The generative step is outside the method.
- Fifty years on: whether strong inference described practice or inspired it (Davis 2006), and how it is cited now (Fudge 2014).
- Platt says the scheme is Baconian: the conditional inductive tree, the instances of the fingerpost, induction by rejections and exclusions (Platt 1964, 349–50; Bacon 2004); Chamberlin’s multiple working hypotheses round it out (Platt 1964, 350; Chamberlin 1965); Popper supplies the asymmetry at X (Platt 1964, 350).
- The exclusion step contradicts \(H_i + A\), not \(H_i\) alone. Platt does not draw the distinction; Cleland makes it the centre of her case, with Neptune as the example (Cleland 2001, 988).
Cleland
- Her thesis in one sentence: two patterns of evidential reasoning, “from causes (test conditions) to effects, with the concomitant worries about ruling out false positives and false negatives, and from effects (traces) to causes, with the concomitant worries about ruling out alternative explanations” (Cleland 2002, 484–85). The first runs D, S, C, O, M, V; the second runs O, G, K.
| Node | Objection | Where she says it |
|---|---|---|
| C | Controlled experiments are impossible for past events: the time frame is too long, the conditions too complex. A smoking gun already exists; it is uncovered by fieldwork, not produced by varying conditions | Cleland (2001), p. 988; Cleland (2002), p. 484 |
| V | Little in historical research resembles controlling for false positives and negatives. Generating rivals is not entertaining auxiliaries: an auxiliary is independent of \(H\) and can be dropped to save it; rivals are incompatible with \(H\) | Cleland (2002), pp. 483–484 |
| X | Refutation hits the conjunction; in practice \(A\) is revised first. Astronomers kept Newton and found Neptune | Cleland (2001), p. 988 |
| K | The smoking gun is a retention move, inference to the best explanation within a stated set of rivals; it can be deposed later, and it need not touch the rivals: shocked quartz is “merely irrelevant” to the contagion hypothesis | Cleland (2002), pp. 481–483 |
| before G | An exploratory phase in which the phenomenon is not yet stable and no \(H\) is worth deriving from; explore until it repeats, then test. Becquerel, Faraday, Cech; Viking as the case where the controls were fixed before any data | Cleland (2002), pp. 479–480, 486 |
- Her positive claim, the asymmetry of overdetermination, is why the found-\(C\) path works: a past event leaves many independent traces, so O can be repeated on new traces without control (Cleland 2001, 989–90; 2002, 487–90). Both papers allow that it may be probabilistic rather than strict (Cleland 2001, 989; 2002, 490–91).
- The two patterns are not tied to fields. Experimentalists reason like historians when a test fails or a phenomenon surprises them; historians reason like experimentalists when nature repeats herself, as with Cepheid variables, though without the power to set or remove the test condition (Cleland 2002, 485–86). That last case is the statistical mode of this seminar: an earthquake catalog is nature’s repeated experiment.
- A smoking gun can be predicted before it is found (Dicke’s background radiation) or stumbled on (the Alvarez iridium) (Cleland 2001, 990). Prediction versus accommodation cuts across the historical and experimental divide.
- Cleland’s answer to why the earth sciences do not recognize themselves in the textbook: their hypotheses are about particular past events, not regularities (Cleland 2002, 480); the evidence is traces that already exist; and the taught account and the practiced one differ. Three geologists’ textbook says “hypotheses cannot be proved, only disproved,” while the field’s practice is a hunt for positive evidence (Cleland 2001, 988; 2002, 483). Philosophers made it worse by describing historical science as narrative, which fits archaeology and not geology or astronomy (Cleland 2002, 475). Computer simulation does not change the historical character of the hypotheses (Cleland 2001, 987, 989).
Read against the texts, September 16, 2026
- Basis: all three papers read in full. Cleland 2002 from the publisher’s PDF at Cambridge Core; Platt and Cleland 2001 from copies posted on university course pages, not stored in this repository. Page numbers are the journals’ own. What follows is what the notes above and the sealed analysis had not recorded.
In Platt, and not in the notes
- The yardstick (p. 351): surveys, taxonomy, instrument design, systematic measurement, and theoretical computation “have their proper and honored place, provided they are parts of a chain of precise induction.” Platt subordinates every other mode to the loop. The assigned paper is the textual opponent of the glossary’s pluralism.
- The two boxes (pp. 351–352): “The logical box is coarse but strong. The mathematical box is fine-grained but flimsy.” Five-decimal agreement excludes; one- or two-decimal fit “may be a trap”; “we substitute correlations for causal studies.” Bohr and Schrödinger predict the same Rydberg constant, so agreement alone excludes nothing. Platt’s position on the statistical and estimation modes, and on data-driven discovery, is here, and it is the equifinality problem of meeting 9.
- Pasteur (p. 351): moving to a new biological problem every two or three years and beating experts who “knew a hundred times as much,” which Platt attributes to method, not knowledge. The assigned paper’s own breadth claim, for meeting 11.
- When the method pays (p. 349): in “high-information” fields, where the hypothesis space is large or observations are expensive, and where leadership taught it. Platt’s own contingency claim about method and field speed, testable in the corpus study.
- The loop in hardware (p. 349): coincidence circuits that “run through a complete logical tree in a microsecond.” The first automated excluder in the paper.
- The pathologies (p. 350): the Frozen Method, the Eternal Surveyor, the Never Finished, the Great Man With a Single Hypothesis, the Little Club of Dependents, the Vendetta, the All-Encompassing Theory. The glossary uses one of the seven; the list is a ready-made failure taxonomy.
- Platt’s evidence: the published form of papers (Lederberg’s “subject to denial” list, Monod and Jacob’s “logical density,” a 1964 paper’s “our conclusions might be invalid if”), notebooks (Faraday, Fermi), and recollected meetings (Szilard at Boulder, 1958), pp. 348, 351–352. Method read off narratives: standing question 5 applied to the paper itself.
- Footnote 14: Platt cites Kuhn 1962 and says the modified view “does not invalidate any of these conclusions.” He knew the objection.
In Cleland 2001, and not in the notes
- The schema is hers (p. 987): a test implication, if \(C\) then \(E\), inferred from \(H\); auxiliaries on p. 988. The unit page’s formula comes from the assigned reading.
- Two textbook accounts, inductivism and falsificationism, both “deeply flawed, both logically and as accounts of the actual practices of scientists” (pp. 987–988). The paper answers the question of what the original thinkers proposed, in one page.
- Experimentalists control, they do not falsify (pp. 988, 990): after a failure they vary auxiliaries against false negatives; after a success they vary them again, and remove \(C\), against false positives. “They are not trying to disprove their hypotheses or to save them from falsification.”
- A smoking gun can be predicted (Dicke) or stumbled on (Alvarez), p. 990.
- Laboratory work in historical science sharpens traces or tests an auxiliary (Miller-Urey), p. 989. Computer models determine consequences under stated conditions and cannot say which conditions obtained: snowball Earth, p. 989. Meeting 9.
- The probabilistic loosening of overdetermination is already in the short paper (p. 989), not only in the long one.
In Cleland 2002, and not in the notes
- The unit is the series, not the test (pp. 477–478): “The historical tendency of philosophers to take the isolated experiment as the descriptive unit of experimental research has obscured this character.” Each experiment is designed in light of the last; without such a series “most researchers are reluctant to submit their work.” This is the finding that added node V to the unit graph.
- Rivals are not auxiliaries (pp. 483–484): an auxiliary is independent of \(H\) and can be dropped to save it; rivals are incompatible with \(H\). This is why G and P are different nodes.
- The one-sentence thesis (pp. 484–485): causes to effects with false-positive and false-negative worries; effects to causes with rival-explanation worries.
- Nature repeats herself (p. 485): Cepheid variables give a body of evidence “resembling what an experimental program would provide,” without the power to set or remove the test condition. The statistical mode, in the assigned philosophy.
- Experimental hypotheses “postulate regularities among event-types,” which “may be statistical” and “may or may not be causal” (p. 476, note 2). Cleland places functional regularities inside experimental science’s own objects.
- Two criteria for a good historical explanation, unification (Kitcher) and a causal mechanism (Salmon); acceptance of drift waited for the mechanism (p. 481). Meeting 4.
- Independent smoking guns (p. 491): microspherules, fullerenes, soot; “had iridium been absent” the case would stand. Convergence of independent traces, which is the evidence column of the synthesis mode.
- Lakatos: a “superficial resemblance”; he ignored protection against misleading confirmations (p. 478 and note 4).
- Her idealization (p. 476): “simple, idealized experiments,” setting aside Franklin’s complications with instruments and unrepeatable experiments. The limit of her account for meeting 5.
Corrections made to these pages from the reading
- The Bacon citation moved from p. 349 to pp. 349–350; “proper rejections and exclusions” is on 350.
- The “conditions found, not set” objection now cites Cleland 2002, p. 484, where she says it, with 2001, p. 988, for the impossibility of controlled experiments.
- The unit graph gained node V, the control series, and the verb table gained “control.”
What to settle in the room
- Are Platt’s exclusion at X and Cleland’s smoking gun at K one Bayesian move under two constraints, or two different moves? Cleland’s last sentence says one (Cleland 2002, 495); the two papers read as two.
- Both papers put the hard part before the loop: Platt at G, Cleland in the exploratory phase before it.
Held September 15, 2026
- Source: the group’s shared meeting notes for September 15. No slides are recorded for this meeting. Beyond the anchor and companion, the discussion drew on two readings listed elsewhere on this site: the science-of-science review Fortunato et al. (2018) and Simon’s argument that discovery has a logic, Simon (1973), an optional reading for meeting 13. The group has since identified Klahr and Simon’s review, Klahr and Simon (1999), as the source of the discussion of prior knowledge; it was read in full on September 23, 2026, from a copy supplied by the convenor and not stored here. Simon (1973) has not been read for this page.
What was discussed
| Reading | Role on this page | What the room took from it |
|---|---|---|
| Platt (1964) | Anchor | Strong inference (the notes say “strong hypothesis”); falsification between rivals as the method |
| Cleland (2001) | Companion | Assume the conditions, derive the predicted event, test the hypothesis by whether it occurs; the smoking gun as the evidence that discriminates between hypotheses |
| Fortunato et al. (2018) | Background, cited in the glossary | Science grows and changes, so a data-driven science of science can show where opportunities for advance remain; small teams against large teams; how scientists choose experiments |
| Simon (1973) | Optional reading, meeting 13; not read for this page | Where new ideas come from; discovery as search; the British Museum algorithm against heuristic search |
| Klahr and Simon (1999) | Identified after the meeting; read in full | Starting knowledge and the path of search; discovery as problem solving in several spaces |
What the room said
- The scientific method, as the room stated it. A hypothesis H is stated, conditions C are assumed, and a testable implication is derived: if C is brought about, event E will occur. H is accepted or rejected according to whether E occurs.
- Three kinds of cognitive work, induction, deduction, and abduction, and a fourth element, the conditions, which are “never really fully represented and reproducible.”
- Prior knowledge sets the range of hypotheses. With strong prior knowledge the range is narrow and the work is a hunt for the smoking gun that shows the expectation is right; with little prior knowledge the range is broader.
- Two textbook accounts. Inductivism: repeated successes confirm H. Falsificationism: H cannot be confirmed, only refuted.
- Less difference between experimental and historical science than Cleland draws.
- The smoking gun cuts both ways. It discriminates between hypotheses, and it biases the work toward confirming one.
- Falsification is not done as rigorously as Chamberlin and Platt describe, because conditions vary. Falsifying a hypothesis needs conditions that are known, met, and repeated, which is very hard in nature; a failed prediction can be a misleading disconfirmation.
- Models of the past are theory. The notes quote Cleland: “modeling past events is theoretical work, and while it may yield predictions, these predictions are only as secure as the assumptions upon which the model is based.”
- Science of science. Small teams bring in older, less popular ideas and open new directions, and last longest when they keep a stable core; large teams refine recent, popular ideas and reach high but often short-lived impact. Scientists’ actual strategy for choosing experiments mines the hubs they already care about efficiently but does not reveal the whole knowledge network; both high-impact targets and unexplored ones need attention. There are optimal team sizes, and smaller teams try to disrupt. The design question the room drew: how to build a small, stable core of agents.
- Simon: discovery is search. If some ways of hunting for hypotheses are systematically better than others, discovery has a logic. It does not need to solve the problem of induction: discovery recodes the data in hand into a more compact pattern, and whether the pattern fits those data can be checked with certainty; whether it will hold on future data belongs to testing. Brute-force enumeration, the British Museum algorithm, against heuristic search, which uses structure in the data to prune the space. Generating hypotheses is a search process, some search processes are demonstrably better than others, and studying why gives a logic of discovery.
- What the group needs to quantify next: disruptive ideas, scientific consensus, novelty, and what counts as an advance, a big leap against incremental work. Meeting 10 took up novelty and uptake the following week, on September 22.
Read against the texts
- The room’s definition is Cleland’s schema without the auxiliaries. Cleland writes the test implication as “if C then E,” inferred from H (Cleland 2001, 987), and adds the auxiliary assumptions a page later (Cleland 2001, 988). The room’s remark that conditions are never fully represented is the same point from the other side: what is not written into C sits in A, and a failed prediction cannot say which conjunct failed (Duhem 1954). The unit graph writes the schema as \(H + A + C \Rightarrow E\) for that reason.
- On falsification, the room agrees with Cleland against Platt. Cleland’s own claim is that experimentalists “are not trying to disprove their hypotheses”: after a failed prediction they vary the auxiliaries to guard against a false negative, the room’s misleading disconfirmation (Cleland 2001, 988). What the room calls unrigorous falsification is, in her account, the control series, node V on the graph (Cleland 2002, 477–78). Where conditions can only be found, V is unavailable, and the found-C methods replace it with independent traces.
- “Less difference than Cleland draws” is partly hers. She allows that experimentalists reason like historians when a test fails or a result surprises them, and that historians reason like experimentalists when nature repeats herself (Cleland 2002, 485–86). What she maintains is the asymmetry of overdetermination: a past event leaves many independent traces (Cleland 2001, 989–90). The notes do not record the room disputing that. Whether the difference is one of kind or of degree stays open.
- The smoking gun and confirmation bias. Cleland’s defense is that a smoking gun is judged within a stated set of rivals, and the survivor is retained as the best explanation, not proved (Cleland 2002, 481–82). The bias the room named is the risk when the set is too small, which is Chamberlin’s warning against the ruling theory (Chamberlin 1965). On the graph, that is why G requires at least two incompatible rivals and why K records “best explanation within the stated set.”
- Prior knowledge narrows G in both directions. It supplies hypotheses by analogy (Gilbert 1886), and it hides the rivals nobody thought to list. The notes’ contrast, narrow range with prior knowledge, broad range without, is a claim about the size of the rival set at G that the ledger can record case by case.
- Klahr and Simon on starting knowledge. The review does not say that prior knowledge predetermines the path. It documents two narrowing effects: familiarity with a domain biases which hypotheses seem plausible, and plausible ones are tested first (Klahr and Simon 1999, 538); and recognition can transfer negatively, as Einstellung and functional fixedness, which is how Monod and Jacob’s belief in activation delayed inhibitory control (Klahr and Simon 1999, 532–34). It also documents two opening effects: a surprise needs an expectation to violate, and analogy draws on everything stored in memory (Klahr and Simon 1999, 533, 537). Its cases locate the narrowing in knowledge held from one domain and in one representation, not in the amount known: Balmer was a geometry teacher, and the mathematicians who refit Planck’s curve in minutes did not recognize Planck’s law (Klahr and Simon 1999, 534). And the path is steered by outcomes as it goes: models given only Krebs’s starting knowledge found his reaction path, with surprise choosing each next experiment (Klahr and Simon 1999, 535). The full account, and what it changes on the graph, is on the unit page.
- The quotation is not located. Its argument matches Cleland on computer models, which determine consequences under stated conditions and cannot say which conditions obtained (Cleland 2001, 989); the wording has not been checked against either Cleland paper.
- The science-of-science points trace to three sources, one outside this site. Small teams disrupt and large teams develop: Wu et al. (2019), whose team-size effect Petersen and colleagues find small once reference-list growth is corrected, and reversing above eight authors (Petersen et al. 2025). The experiment-choice finding: Rzhetsky et al. (2015), an optional reading for meeting 6. That small teams last longest with a stable core is not in this site’s bibliography and has not been traced to a source. “Optimal team sizes” is not traced to a reading either. The design question, a small stable core of agents, leans on the contested part.
- Simon answers discussion question 3. Klahr and Simon state the position and attribute it to Simon (1973): much empirical work is done “in the context of discovery rather than the context of verification,” and “theories cannot be tested until they have been created” (Klahr and Simon 1999, 529). On the graph that is G against K, and the room took this side: generation is search and can be done better or worse. The review gives the search argument in its own terms: with \(b\) branches at each of \(m\) moves, exhaustive search is beyond human capacity, so “effective problem solving depends in large part on processes that constrain search judiciously” (Klahr and Simon 1999, 532). The notes’ “British Museum algorithm” and “recoding the data into a more compact pattern” do not appear in the review; they are taken to come from Simon (1973) and remain unchecked against it. The nearest statement in the review defines the goal of hypothesis search as a hypothesis that accounts for the data “in a more concise or universal form” (Klahr and Simon 1999, 538).
The three questions, as the meeting left them
- Can an observational study discriminate hypotheses without controlling the system? Yes, by the smoking gun, with conditions assumed rather than set, and with the risk of confirmation bias when the rival set is small.
- Which research actions are not well described as a test? Pattern-finding in Simon’s sense: recoding data into a compact pattern is discovery, and whether it will hold is a later test. Exploration is the other. Klahr and Simon name both: exploratory experiments “guided by no specific hypothesis to be tested, and no clear control condition,” whose goal is “to permit phenomena to appear” (Klahr and Simon 1999, 526), and data-driven discovery, now the unit page’s sixth entry state.
- Generation or warrant? Both, following Simon. The notes record no disagreement.
Open after the meeting
- Whether experimental and historical science differ in kind or in degree.
- Where the stable-core finding comes from, and whether it holds for agents.
- Whether reading widely narrows the search. Klahr and Simon’s cases suggest the narrowing comes from reading within one domain and one representation; an agent that has read everything carries the strongest plausibility prior of all, which is the concern meeting 10 raised through Machado.
- Whether the case exercise was done, each person’s own paper placed on the unit graph; the notes do not say.
- Quantifying disruption, consensus, novelty, and advance, taken up in meeting 10.
Prepare and discuss
Read the selected anchor sections and the companion abstract or overview. Bring one source passage or artifact relevant to the case; the rotating reader presents the companion in depth.
- Can an observational study discriminate hypotheses without controlling the system?
- Which research actions are not well described as a test?
- Are we describing how ideas are generated or how claims are warranted?
Case exercise: scientific gain and epistemic mode
Compare Platt and Cleland on one documented case. State a meaningful before/after scientific gain, then identify which proposed mode definition succeeds or fails on the case. Distinguish a normative argument from evidence of what scientists did.
- First, ten minutes: place one recent paper of your own on the unit graph. Mark the nodes it performed, the nodes it skipped, and whether \(C\) was set or found. Then place Platt and Cleland.
- Profiles to examine: hypothesis, synthesis, exploration.
- Record source kind and unknown chronology; distinguish documented practice, philosophical argument, association, and our proposed agent rule.
- Ask what was gained, what warrants it, which action helped, and what a comparable agent would need to demonstrate.
The agent
- The unit skill, agents/skills/unit-of-inquiry/SKILL.md, runs a research agent from a trigger event along the graph, following a method people have defended. Meeting 1 wrote its first version: the node procedures and the strong-inference and trace-series methods come from these two papers. Every later meeting adds to it. The rules below are the ones these two papers force.
- From Platt: no test without at least two incompatible rivals at G, or a logged reason why none could be formulated; before any experiment, the agent answers Platt’s Question, which hypothesis this outcome would disprove, and an experiment with no answer is labelled exploration (Platt 1964, 350, 352). The parallel composition, one crucial experiment shared by \(n\) units, is the skill’s first composition (Platt 1964, 347).
- From Cleland: the agent records whether C was set or found; an experimental claim cannot leave V after one run; controls fixed before data may be revised after data with the reason logged, the Viking rule; rejecting an auxiliary is node P, not X, and is logged as such; a best explanation is retained with that status, never as an exclusion; “not decidable, and this would decide it” is a legal terminal state (Cleland 2001, 988; 2002, 477–80, 483–84, 492–94).
- Planted flaws that test the rules: a forced single hypothesis; a test all rivals pass; a claim leaving V after one run; an exclusion that silently rejected A; a smoking gun reported as an exclusion; abstention on an answerable case.
- Open: whether exclusion and best explanation are one rule with two thresholds or two rules. The room decides; the skill file records the answer.
- Proposed from the September 15 discussion, not yet adopted:
- At G: log how each rival was generated (analogy, known cause, a pattern in the data), so that generation strategies can be compared, as Simon’s argument requires (Klahr and Simon 1999, 529, 532). The evaluator adds a brute-force baseline, enumeration without heuristics, as the control.
- At G: when prior knowledge leaves one rival, log what excluded the others before any test. A single rival with no logged reason is the ruling theory (Chamberlin 1965).
- Adopted into the skill (0.2.2) and the graph (v0.3) from Klahr and Simon (1999): record the starting expectation and the plausibility order of the first rivals; keep one implausible rival until a test could favour it; and send an outcome no rival predicted along the new surprise edge from M, not to X.
- At M and V: when conditions were assumed rather than set, a failed prediction goes to P with the assumption named, not to X. This is the room’s misleading disconfirmation, and the existing planted flaw, an exclusion that silently rejected A, tests it.
- For meeting 13: the architecture question the room raised, a small stable core of agents, with the caveat that the team-size evidence it rests on is contested (Petersen et al. 2025).
Optional extensions
On reasoning and testing. The sources behind the unit of inquiry, listed there by node. Deduction, induction, and abduction: Peirce (1992); Hanson (1958); Okasha (2016) as the one-hour introduction. The hypothesis and where it comes from: Gilbert (1886); Whewell (2014); Mill (2011). Confirmation, falsification, and the Duhem problem: Hempel (1945); Popper (2005); Duhem (1954); Quine (1951); Lakatos (1970). Eliminative induction and severity: Bacon (2004); Mayo (1996); Mayo and Spanos (2006). Inference to the best explanation: Harman (1965); Lipton (2004). Prediction against accommodation: Musgrave (1974); Douglas and Magnus (2013). Strong inference fifty years on: Davis (2006); Fudge (2014). The school version of the method and what it leaves out, from the University of Washington’s College of Education: Windschitl et al. (2008).
- Chamberlin (1965): The Method of Multiple Working Hypotheses. Science (1965). Written by a geologist, and the source Platt built on. Chamberlin’s target is the single ruling hypothesis; his remedy is to hold a family of them at once. Decide whether Platt’s loop and Chamberlin’s family are the same practice.
- Holmes (1987): Scientific Writing and Scientific Discovery. Isis (1987). Holmes compared laboratory notebooks with the published papers written from them and describes what the paper leaves out. The basis for standing question 5.
- Gilbert (1896): The Origin of Hypotheses, Illustrated by the Discussion of a Topographic Problem. Science (1896). Grove Karl Gilbert on where a geological hypothesis comes from, worked through his own attempt to explain the Coon Butte crater: by analogy, by enumeration of the possible causes, and by the test that separates them. The geological ancestor of Platt’s method, and an account of the generative step Platt hands to Pólya.
- Johnson (1933): Role of Analysis in Scientific Investigation. Geological Society of America Bulletin (1933). A geologist’s account of analysis as the method of geological investigation, with the failure modes it guards against.
- Frodeman (1995): Geological reasoning: Geology as an interpretive and historical science. Geological Society of America Bulletin (1995). Geology as an interpretive and historical science, against the physics model of method. The geological statement of Cleland’s thesis, seven years earlier.
- Kleinhans et al. (2005): Terra Incognita: Explanation and Reduction in Earth Science. International Studies in the Philosophy of Science (2005). What explanation and reduction mean in earth science, where the systems are open and the histories unique.
- Elliott and Brook (2007): Revisiting Chamberlin: Multiple Working Hypotheses for the 21st Century. BioScience (2007). Chamberlin’s multiple working hypotheses reconsidered for present-day practice, with the model-selection tools that now implement them.
- Cleland (2002): Methodological and Epistemic Differences between Historical Science and Experimental Science. Longer development of experimental/historical differences; optional elaboration of the companion.
Record after the meeting
Sign and date the notes; preserve disagreements and what changed your assessment. Keep confidential examples in private notes.