Evaluating scientific advance against past work

Drafted case specifications and protocol · no agent runs or historical scores reported

What the comparison establishes

  • Primary question: do human-defined modes improve warranted scientific gains under comparable information and resource constraints?
  • Assess a result against past science and compare agent policies separately. Modern tools and later knowledge prevent a simple speed comparison with the original scientists.
  • Reproduction demonstrates task competence; contemporary novelty requires a current prior-art comparison. Historical uptake belongs to the original contribution, not automatically to an agent’s alternative proposal.
  • Outcome rubric · Mode definitions.

Three task suites

Suite Input and action Scientific assessment
Decision reconstruction Evidence available at a dated decision point; propose a question, representation, or next action before later artifacts are revealed. Was the action justified then, and could it resolve a consequential uncertainty? A different adequate path is acceptable.
Reproduction/reassessment Past claim, data, and available methods; reproduce, qualify, or challenge with inspectable outputs. Does the analysis recover or correct the result and support its asserted scope?
Extension/transfer Established result plus held-out period, region, observable, or related problem. Is there a warranted gain beyond the supplied result under checked assumptions?
  • Source categories are a separate axis: completed local projects, external historical cases, and prospective/newly held-out cases.
  • Inventory four to six completed local projects as candidates; suitability and access have not been established. Keep private artifacts private and add external cases to test beyond local habits.
  • Authors can explain records; include another qualified assessor where feasible. Freeze related projects together rather than splitting closely related episodes across development and evaluation.

Three initial case specifications

  • These specify reading-based development exercises now. They are not executable benchmark packets or certified reconstructions of private thought.
  • Before held-out use: obtain permitted artifacts, verify source passages and dates, freeze an information packet and a separate assessor packet, and confirm an independent scoring reference.

Coastal earthquake reconstruction

  • Sources: Atwater (1987); Nelson et al. (1996); later dating/context in Satake et al. (1996) and Yamaguchi et al. (1997). Primary reports; action chronology beyond those records is unknown.
  • Profiles: synthesis, observation, hypothesis assessment.
  • Starting packet: an instructor-selected set of coastal observations from the earlier report, without later regional/date conclusions. Explicitly list what the selection omits.
  • Task: construct rival histories and propose an observation that distinguishes them; state the supported spatial and temporal scope.
  • Meaningful gain: a better-constrained historical account or an informative next trace that separates substantive alternatives. A famous date recalled without support is not the gain.
  • Assessor packet: published interpretations, later independent traces, limitations, and alternative adequate explanations of why a proposed trace is informative.
  • Controls: several traces with a shared dating assumption; insufficient spatial coverage; a later decisive trace revealed in a second stage.
  • Assessment: evidence dependencies, relevance of discriminating traces, scope of conclusion, and how the conclusion changes on reveal. Do not penalize warranted earlier uncertainty because later evidence resolves it.
  • Exposure: famous historical outcome; prior knowledge possible even when recall probes fail. Suitable for discussion and diagnostic process assessment with that limitation.

Concept and tectonic test

  • Sources: Wilson (1965) and Sykes (1967). Separate conceptual proposal from subsequent observational testing.
  • Profiles: theory, hypothesis assessment, estimation.
  • Starting packet: dated geometrical/conceptual material and assumptions, withholding the later test result; representations must be checked against the primary sources.
  • Task: derive an observable contrast between the proposed account and a stated alternative, then assess the supplied observations in a second stage.
  • Meaningful gain: a valid distinguishable consequence, followed by a warranted assessment rather than a retelling of the accepted theory.
  • Assessor packet: derivation/geometry checks, published observational interpretation, uncertainty, and adequate alternative test designs.
  • Controls: an inconsistent geometry, an observation that both accounts accommodate, and missing information needed to interpret an apparent contrast.
  • Assessment: validity under assumptions, discriminating value, uncertainty, and claim revision. Match of prose to the historical narrative is not a score.
  • Exposure: canonical case; alternative scientifically checked geometries can test reasoning but are labelled constructed variants.

Sensing capability and trustworthy measurement

  • Sources: Lindsey et al. (2019) and Lindsey et al. (2020). Treat the instrument/application and calibration discussions as different evidential claims.
  • Profiles: instruments, estimation, robustness.
  • Starting packet: a specified measurement target, selected response/processing information and observations, with calibration evidence staged separately. Public data availability must be confirmed before computational scoring.
  • Task: identify what is observed versus inferred; propose response checks; revise the capability claim after calibration evidence.
  • Meaningful gain: a supported measurement capability or a consequential correction to its range, not merely a convincing signal plot.
  • Assessor packet: independently checked response relations and published scope, uncertainty, and data provenance.
  • Controls: a response change mimicking a physical feature; shared bias in two comparison chains; a valid capability extension.
  • Assessment: recovery in physical units where available, uncertainty, localization of error, and one scientifically relevant new question enabled.
  • Exposure: source familiarity remains possible. A constructed artifact is a diagnostic control, not a new finding about the published work.

From a specification to a frozen packet

  1. Identify exact artifacts, rights/access, version and dates; separate public and private files.
  2. Create a starting-information manifest and later-evidence manifest; record unavailable/unknown chronology.
  3. Choose mode IDs/versions, task family, baseline, budget, meaningful gain, and evidence requirements.
  4. Have two readers try the task and review the scoring reference. Resolve ambiguity on development cases; retain disagreement and alternative adequate answers.
  5. Freeze project-level splits, model/tool versions, assistance, stopping rules, and assessment protocol; record all runs and failures.
  6. Score case-level gains using claim records; keep novelty, adequacy/warrant, impact, and cost visible.

Contamination and evidence access

  • Record evidence of prior knowledge detected / no evidence detected by these probes / exposure unresolved. A negative recall probe does not certify absence of training exposure (Dekoninck et al. 2024).
  • Record exact model version, documented training bounds if available, retrieval/tool access, artifact history, and probes. Run probes outside evaluation contexts.
  • A post-cutoff paper may have an earlier preprint; private details may have appeared elsewhere. Process scores can benefit from knowing the outcome too.
  • Use sensitivity analyses across exposure categories. New independently held-out artifacts and prospective outcomes support stronger claims than canonical-history replay alone.

Reuse and extension

  • Mai et al. (2016) supplies source-inversion benchmark exercises; Woollam et al. (2022) supplies seismological ML infrastructure. Their capabilities can support evaluation without defining all scientific advance.
  • Chen et al. (2025) supplies a model for expert-validated tasks extracted from papers; our mode/gain assessment extends the task question, without a claim of priority.
  • Next coverage: an empirical-law/forecast case and a cross-field transfer. Select dated primary records and independent references before treating these as benchmark results.
  • Corpus study: population patterns require a sampling design; an initial bounded comparison sample does not.