Evaluating scientific advance against past work
Drafted case specifications and protocol · no agent runs or historical scores reported
What the comparison establishes
- Primary question: do human-defined modes improve warranted scientific gains under comparable information and resource constraints?
- Assess a result against past science and compare agent policies separately. Modern tools and later knowledge prevent a simple speed comparison with the original scientists.
- Reproduction demonstrates task competence; contemporary novelty requires a current prior-art comparison. Historical uptake belongs to the original contribution, not automatically to an agent’s alternative proposal.
- Outcome rubric · Mode definitions.
Three task suites
| Suite | Input and action | Scientific assessment |
|---|---|---|
| Decision reconstruction | Evidence available at a dated decision point; propose a question, representation, or next action before later artifacts are revealed. | Was the action justified then, and could it resolve a consequential uncertainty? A different adequate path is acceptable. |
| Reproduction/reassessment | Past claim, data, and available methods; reproduce, qualify, or challenge with inspectable outputs. | Does the analysis recover or correct the result and support its asserted scope? |
| Extension/transfer | Established result plus held-out period, region, observable, or related problem. | Is there a warranted gain beyond the supplied result under checked assumptions? |
- Source categories are a separate axis: completed local projects, external historical cases, and prospective/newly held-out cases.
- Inventory four to six completed local projects as candidates; suitability and access have not been established. Keep private artifacts private and add external cases to test beyond local habits.
- Authors can explain records; include another qualified assessor where feasible. Freeze related projects together rather than splitting closely related episodes across development and evaluation.
Three initial case specifications
- These specify reading-based development exercises now. They are not executable benchmark packets or certified reconstructions of private thought.
- Before held-out use: obtain permitted artifacts, verify source passages and dates, freeze an information packet and a separate assessor packet, and confirm an independent scoring reference.
Coastal earthquake reconstruction
- Sources: Atwater (1987); Nelson et al. (1996); later dating/context in Satake et al. (1996) and Yamaguchi et al. (1997). Primary reports; action chronology beyond those records is unknown.
- Profiles: synthesis, observation, hypothesis assessment.
- Starting packet: an instructor-selected set of coastal observations from the earlier report, without later regional/date conclusions. Explicitly list what the selection omits.
- Task: construct rival histories and propose an observation that distinguishes them; state the supported spatial and temporal scope.
- Meaningful gain: a better-constrained historical account or an informative next trace that separates substantive alternatives. A famous date recalled without support is not the gain.
- Assessor packet: published interpretations, later independent traces, limitations, and alternative adequate explanations of why a proposed trace is informative.
- Controls: several traces with a shared dating assumption; insufficient spatial coverage; a later decisive trace revealed in a second stage.
- Assessment: evidence dependencies, relevance of discriminating traces, scope of conclusion, and how the conclusion changes on reveal. Do not penalize warranted earlier uncertainty because later evidence resolves it.
- Exposure: famous historical outcome; prior knowledge possible even when recall probes fail. Suitable for discussion and diagnostic process assessment with that limitation.
Concept and tectonic test
- Sources: Wilson (1965) and Sykes (1967). Separate conceptual proposal from subsequent observational testing.
- Profiles: theory, hypothesis assessment, estimation.
- Starting packet: dated geometrical/conceptual material and assumptions, withholding the later test result; representations must be checked against the primary sources.
- Task: derive an observable contrast between the proposed account and a stated alternative, then assess the supplied observations in a second stage.
- Meaningful gain: a valid distinguishable consequence, followed by a warranted assessment rather than a retelling of the accepted theory.
- Assessor packet: derivation/geometry checks, published observational interpretation, uncertainty, and adequate alternative test designs.
- Controls: an inconsistent geometry, an observation that both accounts accommodate, and missing information needed to interpret an apparent contrast.
- Assessment: validity under assumptions, discriminating value, uncertainty, and claim revision. Match of prose to the historical narrative is not a score.
- Exposure: canonical case; alternative scientifically checked geometries can test reasoning but are labelled constructed variants.
Sensing capability and trustworthy measurement
- Sources: Lindsey et al. (2019) and Lindsey et al. (2020). Treat the instrument/application and calibration discussions as different evidential claims.
- Profiles: instruments, estimation, robustness.
- Starting packet: a specified measurement target, selected response/processing information and observations, with calibration evidence staged separately. Public data availability must be confirmed before computational scoring.
- Task: identify what is observed versus inferred; propose response checks; revise the capability claim after calibration evidence.
- Meaningful gain: a supported measurement capability or a consequential correction to its range, not merely a convincing signal plot.
- Assessor packet: independently checked response relations and published scope, uncertainty, and data provenance.
- Controls: a response change mimicking a physical feature; shared bias in two comparison chains; a valid capability extension.
- Assessment: recovery in physical units where available, uncertainty, localization of error, and one scientifically relevant new question enabled.
- Exposure: source familiarity remains possible. A constructed artifact is a diagnostic control, not a new finding about the published work.
From a specification to a frozen packet
- Identify exact artifacts, rights/access, version and dates; separate public and private files.
- Create a starting-information manifest and later-evidence manifest; record unavailable/unknown chronology.
- Choose mode IDs/versions, task family, baseline, budget, meaningful gain, and evidence requirements.
- Have two readers try the task and review the scoring reference. Resolve ambiguity on development cases; retain disagreement and alternative adequate answers.
- Freeze project-level splits, model/tool versions, assistance, stopping rules, and assessment protocol; record all runs and failures.
- Score case-level gains using claim records; keep novelty, adequacy/warrant, impact, and cost visible.
Contamination and evidence access
- Record evidence of prior knowledge detected / no evidence detected by these probes / exposure unresolved. A negative recall probe does not certify absence of training exposure (Dekoninck et al. 2024).
- Record exact model version, documented training bounds if available, retrieval/tool access, artifact history, and probes. Run probes outside evaluation contexts.
- A post-cutoff paper may have an earlier preprint; private details may have appeared elsewhere. Process scores can benefit from knowing the outcome too.
- Use sensitivity analyses across exposure categories. New independently held-out artifacts and prospective outcomes support stronger claims than canonical-history replay alone.
Reuse and extension
- Mai et al. (2016) supplies source-inversion benchmark exercises; Woollam et al. (2022) supplies seismological ML infrastructure. Their capabilities can support evaluation without defining all scientific advance.
- Chen et al. (2025) supplies a model for expert-validated tasks extracted from papers; our mode/gain assessment extends the task question, without a claim of priority.
- Next coverage: an empirical-law/forecast case and a cross-field transfer. Select dated primary records and independent references before treating these as benchmark results.
- Corpus study: population patterns require a sampling design; an initial bounded comparison sample does not.