Scientific novelty: new to whom, compared with what?
Meeting 10 | Tuesday, November 10, 2026 | 1:00-2:00 PM Pacific
Central question
How do people and agents establish what is genuinely new without equating novelty with unfamiliar wording or unusual citation combinations?
The paper pair
Pairing type: Bibliometric indicator / construct-validity challenge.
Paper A: Atypical Combinations and Scientific Impact
- Science (2013). Paper / publisher record
Read Uzzi’s main article and identify its operational definition.
Access: Publisher/DOI page; full text may require institutional access.
Paper B: New and atypical combinations: An assessment of novelty and interdisciplinarity
- Research Policy (2020). Paper / publisher record
Read Fontana’s introduction, validation results, and conclusion; a discussant inspects the physics-paper tests.
Access: Publisher/DOI page; full text may require institutional access.
Why these papers belong together
Uzzi associates citation impact with conventional foundations plus some atypical combinations. Fontana examines whether combination indicators identify novelty rather than interdisciplinarity. The pair supports testing a proxy, not replacing scientific judgment with a distance score. (Uzzi et al. 2013; Fontana et al. 2020)
Prepare before the meeting
Bring a 150-word contribution claim and two closest precedents. A facilitator prepares anonymous human and agent assessments in advance.
Discussion questions
- What community, corpus, and time define the relevant prior knowledge?
- Can a conventional method yield a genuinely new observation?
- What should an agent conclude when a search finds no precedent?
One-hour meeting
| Time | Activity |
|---|---|
| 0-10 min | Independent first judgments; surface disagreements. |
| 10-25 min | Compare the papers: claim, evidence, assumptions, and limits. |
| 25-45 min | Work through the case exercise below. |
| 45-55 min | Translate the discussion into agent requirements and tests. |
| 55-60 min | Record an output and one unresolved disagreement. |
Case exercise
Compare a paraphrase of a known result, a new observation with a conventional method, and a conceptual reinterpretation. Audit assessments against the same source packet; then reveal an overlooked precedent and require revision. Distinguish exact, close, and component precedents.
Agent-design or evaluation output
Novelty Dossier v1: atomic claims, novelty types, closest prior art, attempted disconfirmation, search coverage, apparently new component, uncertainty, and separate correctness/value judgments. Track missed novelty, false novelty, prior-art recall, claim localization, calibration, and paraphrase robustness.
Preserve distinctions and avoid a false historical test
Separate empirical, methodological/instrumental, conceptual, explanatory/theoretical, recombinant, and question novelty. A novelty judgment is relative to a specified corpus and time. It is not a correctness or importance judgment.
Historical replay is a discussion exercise, not automatically a leakage-free benchmark. Restricting retrieved papers does not remove later discoveries from a model’s training. Stronger evaluation can use prospective held-out work or controlled cases, with explicit limits on what those tasks represent.
Optional extensions
- Boudreau et al. (2016): Looking Across and Looking Beyond the Knowledge Frontier: Intellectual Distance, Novelty, and Resource Allocation in Science. Publisher/DOI page; full text may require institutional access.
- Wu et al. (2026): NovBench: Evaluating Large Language Models on Academic Paper Novelty Assessment. Open ACL Anthology proceedings with PDF and BibTeX.
- Shibayama et al. (2021): Measuring novelty in science with word embedding. Open access; consult the linked 2026 correction to Table 4.
- Machado (2026): Generative AI bias against scientific novelty: a cautionary tale from a small-sample evaluation of research proposals. Publisher article; the experiment concerns one funder, year, and model configuration.
For Shibayama and colleagues, also consult the publisher’s correction to Table 4. NovBench evaluates NLP-paper novelty assessments; transfer to Earth science must itself be tested. Human review and one-model proposal-selection studies are comparators, not infallible ground truth or proof of a universal AI bias.
Record after the meeting
Record the evidence for your main claim, what changed your mind, what remains unresolved, and one change to the agent or its evaluation. Keep confidential examples in private group notes rather than committing them to this public book.