Scientific advance and modes of inquiry

Drafted working definitions · the group tests and revises them

Purpose and use

  • Scientific advance is the primary outcome: a warranted gain in knowledge, understanding, or research capability relative to a stated starting point.
  • Human-defined epistemic modes organize how an agent pursues that gain; novelty, correctness/warrant, and impact help assess it.
  • These overlapping profiles are proposals for geosciences. They draw on philosophy, history, and empirical studies of scientific process and impact (Hacking 1992; Fortunato et al. 2018).
  • Hacking’s discussion of styles supplies historical/philosophical context, not validation of this vocabulary. Our modes are not the science-policy distinction called Mode 1/Mode 2.
  • The reasoning vocabulary the profiles assume, deduction, induction, abduction, and the verbs test, confirm, exclude, and corroborate, is defined once on the unit of inquiry, with the graph every episode is placed on.
  • Keep purpose, actions, and context linked: a mode states the epistemic purpose and evidence standards; actions implement it; context records contributions, users, infrastructure, and constraints.
  • Multiple modes may describe one episode. Do not force a decisive mode for a whole paper; allow unknown, insufficient record, and work outside the current vocabulary.

Working mode specifications

  • Each linked page is the authoritative version for teaching and agent-instruction export. Its actions and transitions are design hypotheses.
  • Eleven initial profiles: the previous use/co-production row becomes context; estimation and statistical practice become explicit. The count is provisional.
Profile Scientific gain it can support
Observation and description A credible phenomenon, distribution, baseline, or consequential coverage limit.
Exploratory inquiry A developed phenomenon, representation, question, or credible candidate for follow-up.
Hypothesis assessment Discrimination among alternatives; rival generation is a separate action from testing.
Theory and mechanistic modeling A concept, explanation, or valid derivation that changes what can be understood or tested.
Simulation and prediction Assessed forecasts or useful conditional model consequences.
Instruments and observing systems A reliable new measurement capability.
Methods and computation A new inference or computational capability.
Synthesis and reconstruction A constrained history or supported connection across evidence.
Replication and robustness assessment A corrected result, supported range, or meaningful independent corroboration.
Inferential estimation A latent quantity with uncertainty, resolution, and identified limits.
Statistical modeling and empirical regularities An informative relation or distribution with a warranted range of application.

Coding an episode

  • Record the question before/after, actions and artifacts, modes, evidence basis, and next decision in the ledger.
  • Keep contribution types separate: new data, estimate, instrument, method, concept, synthesis, or replication; enabling/corrective/integrative are roles, not importance ranks.
  • Record understanding/use aims, co-production, infrastructure, constraints, question origin, and later reframing. Usefulness and legitimacy require their own evidence (Stokes 1997; Cash et al. 2003).
  • A tomography episode estimates a structure using inversion and sensitivity analysis; a later episode uses it to assess tectonic explanations. The resulting image is inferred, not a direct observation.
  • A derivation can follow observations and be mathematically valid before any new measurement. Empirical adequacy is a further claim.
  • A fitted seismic relation can be descriptive, predictive, and mechanistic in different respects. Breiman’s cultures and Shmueli’s objectives are different lenses (Breiman 2001; Shmueli 2010).

Terms of assessment

Scientific advance

  • A meaningful, evidenced before/after improvement in what can be known, explained, predicted, measured, or investigated; knowledge and understanding accounts are open to discussion (Bird 2007; Dellsén 2016).
  • State the frame: task/project, historical knowledge at a date, or contemporary field. Reproducing an old discovery can demonstrate task competence without a new field advance.
  • Primary evaluation: the fraction of held-out cases achieving a prespecified meaningful warranted gain, with magnitude, failures, uncertainty, and case-family results. The rubric specifies its limits.
  • Neither a fixed novelty × correctness × impact formula nor contribution count defines advance.

Idea

  • The unit that novelty and uptake are assigned to. The readings use four different units, and a novelty score is only comparable with another computed on the same unit.
Unit Definition in the source Source
A research idea A sufficiently specified research project that proposes a motivated method and a practical way to test it, not a short concept or hypothesis alone Si et al. (2025), as summarized in the meeting 10 notes
A link between two concepts Two concepts from a topic model put together in one document for the first time in the corpus Hofstra et al. (2020)
A combination of prior work A pair of journals co-cited in one reference list Uzzi et al. (2013); Wang et al. (2017)
A contribution or claim An observation, method, concept, explanation, connection, or question, stated atomically Novelty Dossier
  • On the unit graph, an idea in Si’s sense is a pass through Q, G, D, and S without C, O, or M: a question, a hypothesis, and a designed test, not yet run. A link in Hofstra’s sense is part of what happens at G, generation by joining two concepts.

Novelty

  • Newness of a specific contribution relative to a stated corpus and date (Uzzi et al. 2013; Fontana et al. 2020).
  • Record closest precedent, new component, search coverage, and uncertainty; semantic or citation distance is a proxy (Shibayama et al. 2021).
  • Say “no precedent located within this search.” Distinguish discovery data from later evaluation data. Novelty Dossier.
  • The same word names different constructs. Meeting 10 read three of them side by side:
Construct What is new How it is judged Source
Atypical combination A pair of cited journals that co-occurs less often than in a randomized citation network; atypical pairs are those in the 10th percentile of z-scores Computed; no human in the loop Uzzi et al. (2013)
First-ever combination A journal pair cited together for the first time, weighted by how difficult the pairing is Computed Wang et al. (2017)
New conceptual link A concept pair never linked before in the corpus; counted per thesis Computed from a topic model Hofstra et al. (2020)
Semantic distance Distance between cited works, or between concepts, in an embedding space Computed; validated against authors’ own ratings in one field Shibayama et al. (2021)
Rated difference from prior art “Whether the idea is creative and different from existing works on the topic, and brings fresh insights,” with all work online before a stated cutoff treated as prior 1 to 10 expert rating against written anchors Si et al. (2025)
Originality as a review criterion One of four dimensions a funder’s panel scores Expert panel; compared with a language model’s ranking Machado (2026)
  • Two rules follow for this group’s records. State the construct with the score: “novel” without its unit and its judge is not a finding. And state the cutoff date: Si’s reviewers used July 2024; a computed measure uses the corpus it was run on.
  • Si’s rubric bundles two things, difference from prior work and “fresh insights.” An idea can be unprecedented and useless. The Dossier keeps them apart: difference goes in “new component,” usefulness in the advance profile.

Proximal and distal novelty

  • Proximal novelty links two concepts inside one cluster of a field’s vocabulary. Distal novelty bridges two clusters. Distance is the embedding distance between the linked concepts, averaged per document (Hofstra et al. 2020).
  • Distal novelty is the continuous form of what the combination measures call atypicality and what the breadth literature calls interdisciplinarity. Fontana’s point applies: a distance score may be measuring interdisciplinarity rather than newness (Fontana et al. 2020).
  • In Hofstra’s data, uptake falls steeply with distance: distal links are “difficult to integrate into localized conversations within prevailing fields” (Hofstra et al. 2020). Distal novelty is the kind least likely to travel, not the kind least likely to matter.

Correctness and evidential warrant

  • Assess mathematical/empirical adequacy separately from what the available evidence justified at the time.
  • The mode specifies relevant evidence: validity of a derivation, calibrated uncertainty, independent observations, discriminating consequences, or sensitivity to assumptions.
  • A completed checklist cannot guarantee truth. Record supported scope, unknowns, contradictions, and dated later reassessment (Oreskes et al. 1994; Flake and Fried 2020).
  • No contradiction found, continued citation, and reproducible code are not sufficient evidence of correctness.

Impact

  • Observed downstream use or consequence: later explanations, measurements, datasets, methods, or decisions.
  • Field/time-qualified citations measure some uptake and attention; record follow-up windows and missed forms of use (Bornmann and Daniel 2008; Waltman 2016; Petersen et al. 2025).
  • Keep predicted impact dated and separate from realized impact; an overlooked gain can still be an advance (Wang et al. 2017).

Uptake

  • Uptake is impact counted on the unit of novelty: the later reuse of one specific new element. Hofstra counts, for each new concept link a thesis introduced, how many later theses use that link (Hofstra et al. 2020). Citations per paper, Uzzi’s outcome, count attention to the whole paper, not reuse of its new part (Uzzi et al. 2013).
  • Impactful novelty is novelty that is taken up. Hofstra models uptake only where novelty is greater than zero, so “not adopted” is not confused with “nothing to adopt” (Hofstra et al. 2020).
  • Innovation = novelty × uptake, kept as two variables. The split lets two separate questions be asked: who creates new links, and whose links spread. In Hofstra’s data they have opposite answers. Scholars with fewer demographically similar peers introduce more new links, and their links are reused less (Hofstra et al. 2020).
  • Before- and after-the-fact estimates of uptake are different quantities. Si’s “excitement” rating is a reviewer’s estimate before the work is done of “potential impact and influence if developed fully” (Si et al. 2025). Hofstra’s uptake is observed afterwards. The advance profile keeps the two in separate fields.

Breadth, interdisciplinarity, and depth

  • Breadth is a candidate contributor to advance: test when integrating another field’s concepts, methods, or observations helps (Wagner et al. 2011; Shi and Evans 2023).
  • Distinguish project integration, competent individual/team reach, and diversity across a portfolio. Reference diversity alone demonstrates none of these.
  • Depth means demonstrated competence and understanding. Reference concentration and topic specialization are proxies for concentration, not its definition (Teodoridis et al. 2019; Rassenfosse et al. 2022).
  • Empirical studies differ in populations and outcomes: citation combinations (Uzzi et al. 2013), career costs/uptake (Leahey et al. 2017), and scientific AI/topic coverage (Hao et al. 2026). A writing experiment (Doshi and Hauser 2024) is a different setting. No combined universal breadth recipe follows.
  • Breadth of inputs and distance of a link are measured differently. Reference diversity measures what was read, over the whole reference list (Stirling 2007; Wagner et al. 2011). Distal novelty measures how far apart two joined concepts are, one link at a time (Hofstra et al. 2020). A paper can read broadly and join nothing distant.

Field-shaping

  • Retrospective change in a field’s questions, capabilities, practices, or explanatory resources.
  • Citation disruption, canonical status, and delayed recognition observe different aspects (Wu et al. 2019; Petersen et al. 2025; Chu and Evans 2021; Ke et al. 2015).
  • Enabling is a contribution role, not a middle rank between incremental and field-shaping.
  • Field-shaping differs from uptake in what is counted. Uptake counts reuse of a new element. Field-shaping asks whether the later literature reorganized around the work: whether its citers stop citing what it built on (Funk and Owen-Smith 2017; Wu et al. 2019), whether the field’s canon turns over (Chu and Evans 2021), or whether recognition arrives late and from another field (Ke et al. 2015). A widely reused link can leave a field’s structure unchanged; a paper cited little for a decade can reshape one.
  • Uzzi’s high-impact profile, a conventional base with an intrusion of atypical combinations (Uzzi et al. 2013), is a claim about citation impact, not about field-shaping. Meeting 10 recorded it as “a strong base and a niche where you add the novelty.”

Action, commitment, and claim status

Distinction Record and evaluate
Action restriction Does a rule require a hypothesis before every next action? Treat this as a policy to test, not a premise.
Prospective commitment What was fixed, when, before access to which evaluation evidence? Record amendments, blinding, unknown, and not applicable (Nosek et al. 2018).
Claim warrant Does the asserted scope/status match the evidence supplied for this kind of claim? Allow conjectures and provisional findings.
  • A frozen forecast, blinded test, and registered analysis plan are related but distinct. Data can already exist before a legitimate prospective commitment about access or analysis.
  • Ioannidis models false-positive risks under assumptions; it does not prescribe a universal hypothesis-before-action rule (Ioannidis 2005).
  • A turning point may be identifiable, but published narrative is not complete chronology. Separate raw artifacts, contemporaneous statements, and retrospective interpretations (Holmes 1987).
  • Revise a definition by recording the case it fails on and the changed version. An agent can suggest revisions; it cannot silently redefine its evaluation criteria.