Scientific practice and impact: a study design

Drafted optional population study · no corpus results reported

Two scales of measurement

  • The agent evaluation needs a bounded reference sample, explicit prior-art searches, and direct case-level assessment of scientific gains.
  • A larger corpus is needed for claims about population mode shares or associations with novelty/impact. It is not a prerequisite for every quantitative comparison.
  • Literature distributions supply context, not ceilings on which unfamiliar modes an agent should pursue.

Questions

  1. Which epistemic work is presented in a defined geoscience population, and how does it vary over time?
  2. How do measured novelty and later uptake differ across overlapping mode labels and contribution types?
  3. When does substantive interdisciplinary integration accompany a demonstrated scientific gain on a closely read subset?
  4. Where dated process records exist, how do they differ from the publication’s presented account?

Population and records

  • Define the target population and time range before sampling. Begin with seismology; screen subject matter within BSSA, SRL, GJI, JGR Solid Earth, and broader journals rather than treating every article as seismology.
  • Record source/version, query, inclusion rules, topic-screen sensitivity, missing abstracts/references, and access restrictions.
  • Add a second physical-science setting after the local pilot. Bioengineering/biology process comparisons can inform modes without substituting for physical-science coverage.
  • Unit for publication analysis: paper or identified claim, explicitly stated. Unit for process analysis: a documented episode. Neither reveals unrecorded actions automatically.
  • Older feasibility counts in the v0.3 design were exploratory snapshots, not a frozen corpus definition; obtain fresh counts when specifying the sample.

Coding

  • Multiple mode IDs, actions, and context/contribution attributes; unknown, insufficient record, and outside vocabulary permitted.
  • Record direct observation versus inferential estimation, statistical objectives, question origin/reframing, evidence basis, assumptions, and shared infrastructure.
  • Hypothesis timing is “as presented” unless dated evidence exists. Commitments record what was fixed before access to which evidence, with unknown/not-applicable states.
  • Keep abstracts-only and full-text coding separate. Mask prestige/citation information where feasible, acknowledging that full texts may reveal identity and era.
  • Two readers annotate a development subset independently; report disagreements and revise rules. Compare with a simpler codebook for coverage, usefulness, and agreement.

Sampling and validation

  • Pilot a bounded sample before promising a full corpus. Include a probability-based component for population estimation and enriched rare-mode cases for diagnostic evaluation; report enrichment separately.
  • Reserve independent held-out papers/projects, stratified for period, mode coverage, and access. A fixed 300-paper holdout does not guarantee precision for rare labels.
  • Validate per-label precision/recall, calibration, coverage, and human agreement. Carry classification and sampling uncertainty into reported estimates; low agreement is a measurement problem to investigate (Flake and Fried 2020).
  • Retain holdout isolation after every codebook/classifier change. Material used to tune a reader is development data, not an independent evaluation set.

Measures and interpretations

Construct Measurement and limit
Scientific advance Close-reading subset: evidenced before/after gain using the profile. No automatic field-wide advance score.
Novelty Claim-level closest precedent at a date; citation/semantic measures as proxies, with reference-corpus and terminology limits (Fontana et al. 2020; Shibayama et al. 2021).
Correctness/warrant Claim-specific expert assessment with explicit evidence and unknowns. Later citations, corrections, and contradictions locate reassessment; none is assumed a validated truth proxy.
Impact Comparable follow-up windows, field/time-qualified citations, documented data/method use, and missing uptake (Waltman 2016; Bornmann and Daniel 2008). Separate predicted from observed impact.
Integration Cited-field diversity as a proxy; direct coding of which concept/method transferred, assumptions, and contribution to the result on a subset (Wagner et al. 2011; Shi and Evans 2023).
Depth/competence Demonstrated understanding and checks on a subset. Reference concentration measures concentration; it is not an independent inverse measure of breadth or a measure of competence.
Field-shaping Citation disruption, canonical uptake, and delayed recognition separately, with limitations (Wu et al. 2019; Petersen et al. 2025; Chu and Evans 2021; Ke et al. 2015).

Analysis and decision points

  • Describe shares and distributions with uncertainty by field/period and accessible record type. Multi-label shares may sum above 100%.
  • State a causal question and assumptions before causal language. Field/year adjustment alone does not establish effects; distinguish confounders, mediators, selection, and unequal follow-up.
  • Include ordinary work and failed paths where records exist; published-paper samples miss much unsuccessful inquiry.
  • Freeze the analysis after development and before evaluating held-out outcomes; report later exploration as such (Tukey 1962; Nosek et al. 2018).
  • Deliver a reproducible sample definition, versioned codebook, validation report, and uncertainty-qualified profiles. Populate mode pages with results only after the study runs.
  • Progress depends on record quality and measurement validity, not predetermined meeting deadlines. The reading club remains useful if the larger study is not pursued.