Scientific advance: records and evaluation

Drafted working instruments · assessors and the group must test them

Inquiry-mode ledger

  • One row per episode, signed and dated; five minutes at the end of a meeting. Live table.
  • The central question is what changed scientifically and what justified the next action.
  • Multiple modes and unknown chronology are allowed. An imagined alternative policy is a counterfactual judgment, not an observed treatment effect.
Field Record
Episode/source Identifier, source location, record type, contributor/date.
Question before → after Include unchanged/unknown. Origin and reframing are different variables.
Action → result What was done; artifact or observation; what changed.
Modes IDs and versions, multiple if appropriate; unknown/outside vocabulary permitted.
Gain and warrant Before/after scientific gain; evidence and supported scope; what remains provisional.
Next action/agent lesson One useful action, alternative, or testable capability.
  • Expanded process record: information available at the time, alternatives, assumptions, constraints, discarded paths, commitment history, disciplinary resources, and missing steps.
  • Keep raw artifacts, contemporaneous statements, and retrospective interpretations distinct (Holmes 1987). Never infer complete chronology from a paper’s rhetorical order.

Claim record

  • For substantive findings, conjectures, or capability claims; ordinary tool calls need an action log, not a full novelty search.
  • One template is shared by the notebook, agent output, and assessment. Preserve the original record and append dated revisions.
Field Required content
Identity Claim ID, case/project, date, text, scope; mode IDs and source versions.
Scientific gain Before/after change and the case-level gain it supports; why that change matters.
Novelty Reference date, search limits, closest precedent, new component, uncertainty.
Warrant/adequacy Evidence required/supplied, assumptions, checks and results, unresolved or contradicted components.
Status Conjectural, provisional, supported within scope, contradicted, or unresolved; assessor/date/reason.
Impact Observed use with sources/dates; separately, any prediction with target, horizon, and later check.
Dependencies Data/code/source artifacts and other claims on which this one depends.
Follow-up Corroborated, challenged, narrowed, overturned, or no informative follow-up located; search date and evidence.
  • Correct procedure and warranted assertion at a date do not guarantee eventual truth. Continued citation and absence of contradiction do not establish correctness.
  • Count case-level gains, not claim cards: several claims may support one advance.
  • Machine-readable template; blank values are prompts for completion, not observed results.

Novelty Dossier v1

Field Record
Atomic claim Observation, method, concept, explanation, connection, or question.
Reference frame Community, corpus, languages, publication cutoff, search date.
Closest precedent Specific work and overlapping passage/result.
New component What apparently goes beyond the precedent and in which respect.
Search for equivalents Earlier terminology, adjacent fields, and searches likely to uncover a precedent.
Limits/judgment Missing sources; exact/close/component precedent, apparently new component, or unresolved.
  • “No precedent located within this search” is a bounded finding. A historical replay is not contemporary novelty.
  • Human judgments are comparators with disagreement (Boudreau et al. 2016). Citation/semantic novelty does not establish correctness or significance.
  • Candidate checks: paraphrase of a known result; established method with new data; concealed precedent. Report false and missed novelty, and search coverage.

Scientific Advance Profile v1

  • Working definition: a meaningful warranted improvement in knowledge, understanding, or research capability. Compare knowledge and understanding accounts (Bird 2007; Dellsén 2016).
  • State the frame: supplied task/project, historical state of a field, or contemporary knowledge.
Field Question
Starting state What was already known, possible, and available to the investigator?
Gain What can now be known, explained, measured, predicted, or investigated?
Magnitude/significance Quantitative change where meaningful; why it matters for this scientific question.
Evidence What warrants the gain, at what uncertainty and scope? What would change the assessment?
Contribution role Corrective, enabling, integrative, cumulative, transformative; overlap is possible.
Dependencies Instruments, datasets, methods, people, standards, and prior results.
Consequences Actual downstream use and time window; distinguish prospective promise.
  • Primary case outcome: achieved/not demonstrated/unassessable meaningful warranted gain, judged against criteria fixed before evaluation. Unassessable and incomplete cases do not count as successes and stay visible in the full denominator.
  • Report success fraction by case family/mode, magnitude in scientific units, uncertainty, and all failures. An aggregate is conditional on case selection, not a universal measure of science.
  • Prespecify meaningful gains and a procedure for assessing unanticipated valid gains; judges must provide evidence and reasons for any such award.
  • Novelty, warrant, impact, and cost remain visible beside the primary outcome. Avoid a default weighted sum or product.
  • First assess evidence/gain with venue, citations, and agent identity hidden where feasible. Then reveal documented uptake to assess realized impact. Retain disagreements; calibrate on development cases before freezing the rubric.
  • Do not rank enabling between incremental and field-shaping: those describe role and reach. A correction or replication can produce a major gain.

Interdisciplinary transfer profile

  • Hypothesis: appropriate integration can improve scientific gains; fluent borrowing is insufficient (Wagner et al. 2011; Shi and Evans 2023).
  • Record source/target concepts, variables, units, scales, assumptions, mathematical or mechanistic correspondence, adaptation, failure conditions, and a check on the target.
  • Record competence as held, borrowed from an expert, checked through tools/evidence, or missing. Assess depth through demonstrated understanding, not reference concentration.
  • Compare access to the same cross-field material with/without an explicit transfer procedure. Include beneficial, irrelevant, and invalid transfers.
  • Assess project gains, individual/team competence, and portfolio diversity separately; effects in one do not establish the others (Kitcher 1990; Leahey et al. 2017; Hao et al. 2026).

Agent evaluation protocol v1

  • Follow the historical evaluation specification. Past papers supply scientific reference work, not a compulsory sequence of actions.
  • Three policies: competent generic planner with the same guidance; hypothesis-testing planner allowed inspection/calibration; planner selecting human-defined modes. A rigid action-gate rule is an optional ablation.
  • Hold model/version, tools, evidence, budgets, and human assistance comparable; separate mode competence with a supplied mode from mode selection.
  • Freeze task packets, scoring, information access, costs, stopping rules, and policy versions before evaluation. Record commitments relative to access, not only collection dates (Nosek et al. 2018).
  • Reserve entire projects/case families to avoid leakage between related development and evaluation episodes. Repeat runs, but do not count repeats as independent scientific cases.
  • Assess artifacts and actions, not generated accounts of hidden reasoning. Human assessors should be blind to policy where feasible.
Control What it reveals
Positive and negative cases Failure to recover a real gain or invention of a gain unsupported by the evidence.
Measurement artifact Mistaking response/processing for a physical effect.
Missing evidence Appropriate limitation or next measurement, versus unjustified certainty or generic abstention.
Independent period/region Transfer beyond the evidence used to formulate the claim.
Valuable correction Whether a novelty-focused agent misses a consequential correction.
Same guidance without mode structure Whether improvement comes from mode organization rather than more instructions.
  • Choose independent cases and outcome thresholds first; use development variability/cost to set repeated-run budgets. No fixed run count guarantees broad evidence.
  • The initial reference sample permits bounded comparisons. A population claim needs the corpus study.