Scientific practice and impact: a study design
Drafted optional population study · no corpus results reported
Two scales of measurement
- The agent evaluation needs a bounded reference sample, explicit prior-art searches, and direct case-level assessment of scientific gains.
- A larger corpus is needed for claims about population mode shares or associations with novelty/impact. It is not a prerequisite for every quantitative comparison.
- Literature distributions supply context, not ceilings on which unfamiliar modes an agent should pursue.
Questions
- Which epistemic work is presented in a defined geoscience population, and how does it vary over time?
- How do measured novelty and later uptake differ across overlapping mode labels and contribution types?
- When does substantive interdisciplinary integration accompany a demonstrated scientific gain on a closely read subset?
- Where dated process records exist, how do they differ from the publication’s presented account?
- Step 0: a documented prior-art search in Novelty Dossier form. Do not claim these questions have never been measured.
- Starting references span strategies, combinations, interdisciplinary work, and measurement limits (Foster et al. 2015; Uzzi et al. 2013; Shi and Evans 2023; Leahey et al. 2017; Fortunato et al. 2018; Petersen et al. 2025). They study different populations and outcomes.
Population and records
- Define the target population and time range before sampling. Begin with seismology; screen subject matter within BSSA, SRL, GJI, JGR Solid Earth, and broader journals rather than treating every article as seismology.
- Record source/version, query, inclusion rules, topic-screen sensitivity, missing abstracts/references, and access restrictions.
- Add a second physical-science setting after the local pilot. Bioengineering/biology process comparisons can inform modes without substituting for physical-science coverage.
- Unit for publication analysis: paper or identified claim, explicitly stated. Unit for process analysis: a documented episode. Neither reveals unrecorded actions automatically.
- Older feasibility counts in the v0.3 design were exploratory snapshots, not a frozen corpus definition; obtain fresh counts when specifying the sample.
Coding
- Multiple mode IDs, actions, and context/contribution attributes; unknown, insufficient record, and outside vocabulary permitted.
- Record direct observation versus inferential estimation, statistical objectives, question origin/reframing, evidence basis, assumptions, and shared infrastructure.
- Hypothesis timing is “as presented” unless dated evidence exists. Commitments record what was fixed before access to which evidence, with unknown/not-applicable states.
- Keep abstracts-only and full-text coding separate. Mask prestige/citation information where feasible, acknowledging that full texts may reveal identity and era.
- Two readers annotate a development subset independently; report disagreements and revise rules. Compare with a simpler codebook for coverage, usefulness, and agreement.
Sampling and validation
- Pilot a bounded sample before promising a full corpus. Include a probability-based component for population estimation and enriched rare-mode cases for diagnostic evaluation; report enrichment separately.
- Reserve independent held-out papers/projects, stratified for period, mode coverage, and access. A fixed 300-paper holdout does not guarantee precision for rare labels.
- Validate per-label precision/recall, calibration, coverage, and human agreement. Carry classification and sampling uncertainty into reported estimates; low agreement is a measurement problem to investigate (Flake and Fried 2020).
- Retain holdout isolation after every codebook/classifier change. Material used to tune a reader is development data, not an independent evaluation set.
Measures and interpretations
| Construct | Measurement and limit |
|---|---|
| Scientific advance | Close-reading subset: evidenced before/after gain using the profile. No automatic field-wide advance score. |
| Novelty | Claim-level closest precedent at a date; citation/semantic measures as proxies, with reference-corpus and terminology limits (Fontana et al. 2020; Shibayama et al. 2021). |
| Correctness/warrant | Claim-specific expert assessment with explicit evidence and unknowns. Later citations, corrections, and contradictions locate reassessment; none is assumed a validated truth proxy. |
| Impact | Comparable follow-up windows, field/time-qualified citations, documented data/method use, and missing uptake (Waltman 2016; Bornmann and Daniel 2008). Separate predicted from observed impact. |
| Integration | Cited-field diversity as a proxy; direct coding of which concept/method transferred, assumptions, and contribution to the result on a subset (Wagner et al. 2011; Shi and Evans 2023). |
| Depth/competence | Demonstrated understanding and checks on a subset. Reference concentration measures concentration; it is not an independent inverse measure of breadth or a measure of competence. |
| Field-shaping | Citation disruption, canonical uptake, and delayed recognition separately, with limitations (Wu et al. 2019; Petersen et al. 2025; Chu and Evans 2021; Ke et al. 2015). |
Analysis and decision points
- Describe shares and distributions with uncertainty by field/period and accessible record type. Multi-label shares may sum above 100%.
- State a causal question and assumptions before causal language. Field/year adjustment alone does not establish effects; distinguish confounders, mediators, selection, and unequal follow-up.
- Include ordinary work and failed paths where records exist; published-paper samples miss much unsuccessful inquiry.
- Freeze the analysis after development and before evaluating held-out outcomes; report later exploration as such (Tukey 1962; Nosek et al. 2018).
- Deliver a reproducible sample definition, versioned codebook, validation report, and uncertainty-qualified profiles. Populate mode pages with results only after the study runs.
- Progress depends on record quality and measurement validity, not predetermined meeting deadlines. The reading club remains useful if the larger study is not pursued.