What has been measured: novelty, impact, breadth, advance
Drafted registry of metrics and findings · read against the texts where an open copy exists · master list begun at meeting 10, September 22, 2026
- The seminar’s outcome is scientific advance, assessed through novelty, correctness, and impact (glossary). This page records what the measurement literature has found useful, what it found, and what it cannot see, so the club reads the papers with the constructs already in view.
- Most of this literature predates language models. Its metrics are built from citation graphs, journal categories, and keyword co-occurrence. Section 7 says which of them embeddings and computational reading can revisit, and which they cannot fix.
- The club does not take these findings on trust. Section 6 says how they are tested against the group’s own judgments, meeting by meeting, using tools/indicators.py.
- Basis column: full means the full text was read from an open copy on September 17, 2026; abstract means the abstract from a registry; record means bibliographic record only, and the finding column reports the paper’s stated claim as commonly cited, unverified here.
0. Master list by construct
- One row per metric, grouped by the construct it is used for. The definitions of the constructs are in the glossary. Findings, data, and limitations are in sections 1 to 4, by paper. Started at meeting 10, September 22, 2026. Add a row when a reading introduces a metric, with the unit it counts on.
- “Unit” is what receives the score. Two metrics with different units are not two measurements of one thing.
| Construct | Metric | Unit | Computation or scale | Source |
|---|---|---|---|---|
| Novelty | Atypicality: tail z-score of co-cited journal pairs | Reference list | z-score of each pair’s frequency against a randomized citation network; the 10th-percentile pair marks the atypical tail | Uzzi et al. (2013) |
| Novelty | Conventionality: median z-score of the same pairs | Reference list | As above; high median means a conventional base | Uzzi et al. (2013) |
| Novelty | First-ever journal pairs, difficulty-weighted | Reference list | Count of pairs never co-cited before, weighted by how distant the two journals are | Wang et al. (2017) |
| Novelty | Semantic distance between cited works | Reference list | Cosine distance between embeddings of the cited works’ titles, abstracts, or keywords | Shibayama et al. (2021) |
| Novelty | New concept links | Document (thesis) | Count of concept pairs from a topic model that never co-occurred before in the corpus | Hofstra et al. (2020) |
| Novelty | Surprise of content and context combinations | Paper | Improbability of the paper’s keyword and cited-journal combinations under a hypergraph model fitted to earlier papers | Shi and Evans (2023) |
| Novelty | Research strategy: tradition versus innovation | Paper | Whether the chemical relationships a paper states are new to the network or repeat existing ones | Foster et al. (2015) |
| Novelty | Expert novelty rating | Research idea | 1 to 10 against written anchors (5: not enough for a paper; 6: probably enough); prior art is everything online before a stated cutoff; a low score must name the similar work | Si, Yang, et al. (2025) |
| Novelty | Originality as a review criterion | Proposal | Panel score; compared with a language model’s pairwise ranking | Machado (2026) |
| Novelty | Model novelty assessment against expert reviews | Paper | Agreement of a model’s novelty judgment with reviewers’ written assessments | Wu et al. (2026) |
| Novelty (validation) | Construct validity of combination indicators | Paper | Indicators tested against expert judgment and against interdisciplinarity | Fontana et al. (2020) |
| Distance | Proximal against distal novelty | Concept link | Embedding distance between the two linked concepts, averaged per document | Hofstra et al. (2020) |
| Distance | Intellectual distance from the evaluator | Proposal and evaluator | Overlap between a proposal’s topic and the evaluator’s own expertise | Boudreau et al. (2016) |
| Breadth | Variety, balance, disparity | Reference list | Number of categories cited, evenness across them, and how dissimilar they are | Stirling (2007); Yegros-Yegros et al. (2015); Wang et al. (2015) |
| Breadth | Integration score | Reference list | Rao-Stirling-type diversity over Web of Science categories | Porter and Rafols (2009) |
| Breadth | Interdisciplinarity weighted by disciplinary similarity | Scientist’s output | Share of output outside one discipline, discounted by similarity | Leahey et al. (2017) |
| Depth | Specialization | Scientist’s output | Concentration of MeSH terms over a career | Rassenfosse et al. (2022) |
| Depth | Exploration against exploitation | Scientist’s output | Topic diversity of outputs from learned representations, before and after a hot streak | Liu et al. (2021) |
| Uptake | Uptake per new link | Concept link | Number of later documents that reuse a link; modelled only where novelty is greater than zero | Hofstra et al. (2020) |
| Uptake (predicted) | Excitement rating | Research idea | 1 to 10: potential impact and influence if developed fully | Si, Yang, et al. (2025) |
| Impact | Citations; top-percentile membership | Paper | Count, or membership of the top 5% (Uzzi) or top 1% (Wang) within field and year | Uzzi et al. (2013); Wang et al. (2017) |
| Impact | Relative citations | Paper | Citations divided by the field-and-year mean | Radicchi et al. (2008) |
| Impact | Fitness, immediacy, longevity | Paper | Parameters of a fitted model of the citation history | Wang et al. (2013) |
| Impact | Q parameter | Scientist | Constant factor in the impact of a scientist’s papers once productivity is removed | Sinatra et al. (2016) |
| Field-shaping | CD disruption index | Paper, patent, software | Whether citers of the focal work also cite its references; −1 consolidating to +1 disruptive | Funk and Owen-Smith (2017); Wu et al. (2019); Park et al. (2023); critiques Petersen et al. (2025); Leibel and Bornmann (2024) |
| Field-shaping | Canon ossification | Field and year | Turnover of the most-cited list against field size | Chu and Evans (2021) |
| Field-shaping | Beauty coefficient | Paper | Area between the citation curve and a line from publication to the citation peak | Ke et al. (2015) |
| Field-shaping | Outsider entry | Field | Output of non-collaborators after a star’s death | Azoulay et al. (2019) |
| Feasibility and quality | Feasibility; expected effectiveness; overall | Research idea | 1 to 10 each; feasibility for a PhD student in one to two months; effectiveness against baselines; overall as likely acceptance at a leading venue | Si, Yang, et al. (2025) |
| Feasibility and quality | Scientific merit; team capacity; feasibility | Proposal | Funder’s panel dimensions, alongside originality | Machado (2026) |
| Advance (field level) | Research productivity | Sector | Idea output per effective researcher | Bloom et al. (2020) |
- Three rows are expert rubrics, not computations: Si’s five dimensions, Machado’s four, and the Novelty Dossier. The meeting 10 discussion called them the subjective and objective rubrics and asked for both. A computed score says where to look. A rubric records what an expert found when they looked.
1. Novelty
| Paper | Metric | Data | Main finding | Stated limitation | Basis |
|---|---|---|---|---|---|
| Uzzi et al. (2013) | Atypicality of journal pairs in a reference list against a shuffled null; conventionality as the median pair z-score | 17.9 million papers, all fields, Web of Science | Highest impact comes from highly conventional combinations with an intrusion of atypical ones; such papers twice as likely to be in the top 5%; teams 37.7% more likely than solo authors to insert atypical pairs | Novelty is a property of the reference list, not of the claim; a journal pair is a coarse unit | abstract |
| Wang et al. (2017) | First-ever journal pairs in a reference list, weighted by how difficult the pairing is | One million articles, all fields, WoS, long citation windows | Highly novel papers have over three times the probability of being top 1% cited given a long window; higher variance; published in lower-impact-factor journals; less cited in short windows | Short-window indicators and journal impact factors are biased against novelty; the finding is about journal-pair novelty only | abstract |
| Fontana et al. (2020) | Tests whether combination indicators (Uzzi’s) measure novelty or interdisciplinarity | Physics papers | Combination indicators partly capture interdisciplinarity rather than novelty; validation against expert judgment is weak | The paper’s own point: construct validity of the indicators | record |
| Shibayama et al. (2021) | Semantic distance between a paper’s cited references, from word embeddings of their titles, abstracts, or keywords | Biomedicine, WoS; validation on 321 survey respondents rating one of their own 2001 to 2006 papers; 2,000 articles from 2010 for citation prediction | The embedding measure correlates with self-reported novelty and predicts ten-year citations; keyword-based variants correlate best | Validation is one field and one survey; most prior novelty measures were never validated at all; computation needs the whole database | full |
| Boudreau et al. (2016) | Expert scores of research proposals as a function of intellectual distance and novelty | 2,130 randomized evaluator-proposal pairs at one university | Evaluators score proposals lower when closer to their own expertise and when highly novel | One institution, one competition | abstract |
| Foster et al. (2015) | Research strategies coded on chemical-relationship networks from abstracts: tradition versus innovation, consolidate versus bridge | Millions of MEDLINE abstracts | The mix of strategies is stable over time; innovation is rare; an innovative paper is more likely to be high impact but the reward does not compensate for the risk of not publishing; prizewinners gamble | Biomedicine and chemistry only; novelty defined on chemical entities | abstract |
| Hofstra et al. (2020) | New links between concepts in dissertation abstracts, from word embeddings; uptake of each new link in later documents | 1.2 million US doctoral theses 1977 to 2015, ProQuest, linked to WoS careers | Underrepresented groups introduce more novel links; their links are taken up less and lead to faculty positions less often; uptake falls steeply with the embedding distance of a link, and gender minorities introduce slightly more distal links, which explains part but not all of the discount | Abstracts only, by necessity; novelty of a thesis, not of a claim | full |
| Shi and Evans (2023) | Surprise of a paper’s combination of contents (keywords) and contexts (cited journals) under a hypergraph embedding model that predicts realized combinations | MEDLINE 19.9 million papers 1865 to 2009; 541,448 physics papers 1893 to 2013; 6.5 million patents 1976 to 2015 | Surprise predicts top-10% citation impact: most surprising context combinations four times more likely to be hits in biomedicine, content two times, both about five; surprising advances come disproportionately from outsiders publishing to a distant field; Nobel papers have high content novelty and low context novelty | Contents as keywords and contexts as cited journals are coarse; association with citations, not a causal test of advance | full |
| Machado (2026) | GPT-4o funding decisions from pairwise comparisons against expert scores | 139 monodisciplinary proposals to the Austrian Science Fund, 2021 | The model matched human funding decisions in 60% of cases; novelty was the only review dimension that predicted disagreement, and proposals humans rated more novel were the ones the model declined | One funder, one year, one model; mechanism not identified | full |
| Wu et al. (2026) | Quality of language-model novelty assessments against expert-written reviews | 1,684 paper-review pairs from one NLP conference | Current models show limited understanding of novelty; fine-tuned models fail at instruction following | Accepted NLP papers only; selection bias acknowledged | full |
| Si, Yang, et al. (2025); Si, Hashimoto, et al. (2025) | Blind expert ratings of research ideas from researchers and a model; then ratings after execution | Over 100 NLP researchers; 43 executed ideas | Model ideas rated more novel and less feasible; after execution the model ideas’ scores fell more on every measure | NLP only; ideas, not discoveries | record (ICLR); full (execution preprint, read September 16) |
- What the literature found useful: two measures survive validation, atypical or first-ever reference pairs with a long citation window, and embedding distance between references or concepts. Both measure novelty of inputs or of vocabulary, not of the claim. That gap is what the Novelty Dossier’s “closest precedent” field records by hand.
2. Impact and its normalization
| Paper | Metric | Data | Main finding | Stated limitation | Basis |
|---|---|---|---|---|---|
| Radicchi et al. (2008) | Relative citations c divided by the field-and-year mean c_0 | WoS, several disciplines | Citation distributions across fields collapse onto one curve when rescaled by c_0; rankings by c_0-normalized counts are unbiased across fields where raw counts are not | Individual-level biases not addressed; the universality was later qualified in the literature | full |
| Wang et al. (2013) | A mechanistic model of a paper’s citation history with three parameters: fitness, immediacy, longevity; predicts ultimate impact | 463,348 Physical Review papers 1893 to 2010; Cell and NEJM 1996 to 2006 | Citation histories collapse onto one universal curve; the fitness parameter predicts long-term impact where impact factor and short-term counts do not | The logistic form cannot capture asymmetric curves; physics-heavy validation | full |
| Ke et al. (2015) | The beauty coefficient B, parameter-free: the area between a straight line from publication to the citation peak and the actual curve | 22 million WoS papers, all disciplines, over a century; 384,649 APS papers | Delayed recognition is not exceptional but a continuous spectrum; top sleeping beauties come disproportionately from physics, chemistry, mathematics and often awaken in a different discipline; short-term metrics miss them | B depends on citation windows long enough to contain the awakening | full |
| Bornmann and Daniel (2008) | Review of about thirty studies of why scientists cite | Studies from the 1960s to 2005 | Citing is not motivated only by intellectual influence; but the other motives are not so random as to destroy citations as an impact measure; the studies themselves are of poor and unreplicable design | The review’s own verdict on its evidence base | abstract |
| Waltman (2016) | Review of citation impact indicators: databases, selection of publications, field normalization, counting methods, journal indicators | Literature review | Field normalization and counting method change rankings; database coverage differs by field; recommendations for further research | A review, not a measurement | full |
| Sinatra et al. (2016) | Position of a scientist’s highest-impact paper in the sequence of their papers; the Q parameter | 2,887 physicists and scientists in several fields | Controlling for productivity, the highest-impact paper falls at a random position in the career; Q is constant over a career and captures the scientist’s contribution beyond luck | Impact is citations within ten years; position in a sequence, not calendar time | abstract |
| Petersen et al. (2025) | Re-analysis of the CD disruption index | SciSciNet, 7.8 million articles 1995 to 2015; PNAS 2011 to 2015 as a quasi-experiment on reference-list length | CD falls over time because reference lists lengthen, a structural effect unrelated to innovation; corrected, disruptiveness rose 2005 to 2015 and the team-size effect is small and reverses above eight authors | Omitted-variable bias in any CD covariate analysis | full |
- What survives: field-and-year normalization is necessary and standard; long windows matter; short-term counts and journal impact factors penalize novelty (Wang 2017) and miss delayed recognition (Ke 2015). Citations measure attention and use together and cannot separate them (Bornmann).
3. Field-shaping and disruption
| Paper | Metric | Data | Main finding | Stated limitation | Basis |
|---|---|---|---|---|---|
| Funk and Owen-Smith (2017) | The CD index: whether works citing a focal work also cite its references (consolidating) or not (destabilizing) | US patents; university commercialization | Federal funding pushes toward destabilizing inventions, commercial ties toward consolidating ones | Patents, not papers | abstract |
| Wu et al. (2019) | CD applied to papers, patents, and software; team size | Tens of millions of papers, patents, and software products, 1954 to 2014 | Small teams disrupt, large teams develop, across all three record types | The claim later contested on measurement grounds (Petersen 2025) | record |
| Park et al. (2023) | CD over time | 45 million papers and 3.9 million patents, 1945 to 2010 | Papers and patents have become less disruptive over time in every field | Contested: reference-list inflation and data anomalies reproduce the trend (Petersen 2025) | record |
| Leibel and Bornmann (2024) | Review of the disruption index and its variants | Literature review | Convergent validity is inconclusive; DI_5 variants validate better than DI_1; the indices carry inconsistency, time-sensitive biases, and data-induced biases; not ready for evaluation practice | The review’s own conclusion | abstract |
| Chu and Evans (2021) | Field size against canon turnover, new-paper entry into the most-cited list, and disruption | WoS, by field and year | When a field publishes many papers per year, citations flow to already-cited papers, the most-cited list ossifies, new papers rarely become highly cited, and new papers are less disruptive | Association across fields; the theory of cognitive overload is argued, not tested directly | abstract |
| Azoulay et al. (2019) | Entry of outsiders into a field after a star scientist’s premature death | Life sciences, WoS | Collaborators’ output falls; non-collaborators’ output rises, draws on a different corpus, and is more often highly cited | Life sciences only; a specific natural experiment | abstract |
- What survives: the construct “disruption” is real enough to argue about and its index is not yet trustworthy. The seminar’s own numbers (section 4) show why: the paper that tested a concept scores as “developing” and the concept as “disruptive,” which is the ledger’s distinction between G and X, not a ranking of worth.
4. Breadth, depth, and the conditions of advance
| Paper | Metric | Data | Main finding | Stated limitation | Basis |
|---|---|---|---|---|---|
| Stirling (2007) | Diversity as variety, balance, and disparity, with a general heuristic | Framework | Three necessary properties of diversity; the heuristic exposes weightings | A framework; the categories it is applied to are conventions | abstract |
| Porter and Rafols (2009) | Integration score over WoS categories in references | Six fields, 1975 to 2005 | Interdisciplinarity increased modestly over thirty years | Depends on WoS category conventions | record |
| Wagner et al. (2011) | Review of approaches to measuring interdisciplinary research | Literature review | Diverse references show what was read, not what was integrated | The review’s own point | record |
| Yegros-Yegros et al. (2015) | Variety, balance, disparity of cited WoS categories, Tobit regression on citations | Articles from 2005 in four WoS categories: cell biology, electrical engineering, food science, atomic and molecular physics | Variety raises impact; balance and disparity lower it; all three are inverted-U in citations | Four categories, one year; publication level rather than group level; citations miss new avenues that are lowly cited | full |
| Wang et al. (2015) | The same three dimensions by factor analysis; Poisson models with journal fixed effects; three-year and thirteen-year citations | All WoS articles of 2001 | Long-term citations rise at an increasing rate with variety, fall with balance, rise at a decreasing rate with disparity; variety and disparity lower short-term citations | Reference-based measures cannot capture integration in the research process | full |
| Leahey et al. (2017) | Interdisciplinarity of a scientist’s output weighted by disciplinary similarity; productivity and citations | About 900 research-center scientists, 32,000 articles | Interdisciplinary scientists publish fewer papers and are cited more; a production penalty and a reception benefit | Research-center scientists; US | abstract |
| Teodoridis et al. (2019) | Specialist versus generalist mathematicians before and after the Soviet collapse changed the pace of a subfield | 6,358 mathematicians, 1980 to 2000 | Generalists performed best when the pace of change was slower; specialists gained the advantage when the pace increased, and became more sought-after collaborators | Theoretical mathematics; one natural experiment | full |
| Teodoridis (2018) | Team composition after the Kinect made motion sensing cheap | Publications using motion sensing | Cheap technology substituted for area specialists and brought in generalists and outside-area specialists | One technology, one field | abstract |
| Rassenfosse et al. (2022) | Within-researcher specialization from MeSH-term concentration, panel regression on citations | Almost 30,000 established biomedical researchers, PubMed linked to WoS | 25% more citations per standard deviation of specialization; up to 75% early in a career; large changes of direction also raise impact | Conditional on long-term success; biomedicine only | full |
| Jones (2009) | Age at first invention, specialization, teamwork over time | US inventor microdata | All three rise over time and are higher in deeper fields; a “burden of knowledge” model | Inventors, not scientists | abstract |
| Wuchty et al. (2007) | Team versus solo authorship and citations | 19.9 million papers over five decades; 2.1 million patents | Teams dominate and produce the most-cited work in nearly all fields; a team paper is 6.3 times more likely to reach 1,000 citations | Team size confounded with resources and field | abstract; number from the Fortunato review |
| Azoulay et al. (2011) | HHMI investigators against matched NIH grantees, propensity weighting and difference-in-differences | Academic life sciences | HHMI investigators produce high-impact papers at a much higher rate and change direction toward novel lines | Observational; no exogenous assignment, as the authors say | abstract |
| Liu et al. (2021) | Topic diversity of a career’s outputs from deep-learning representations, before and after a hot streak | 20,040 scientists; artists and film directors | Hot streaks begin at a transition from exploration of diverse topics to exploitation of one | Observational; omitted variables | full |
| Bloom et al. (2020) | Research productivity as idea output per effective researcher | Semiconductors, agriculture, medicine, firm data | Research effort rises while research productivity falls sharply; Moore’s law now needs eighteen times the researchers of the early 1970s | Measurement of R&D inputs over time; composition effects, discussed by the authors | full |
| Hao et al. (2026) | AI-tool use in papers against output, citations, and topic coverage | 41 million papers | AI users publish more and are cited more; the fields they work in narrow | Tools, not agents; association | record |
| Doshi and Hauser (2024) | Story creativity with and without AI ideas | Online experiment with short stories | Individual creativity rises; the stories become more similar to one another | Creative writing, not science | abstract |
| Sourati and Evans (2023) | Prediction of future discoveries by models that include the distribution of human expertise | Materials and therapies literatures | Human-aware models predict discoveries up to 400% better and can be turned to generate hypotheses humans are unlikely to reach | Prediction of what will be discovered, not whether it is right | abstract |
| Tshitoyan et al. (2019) | Word embeddings trained on materials abstracts up to a cutoff, tested on later discoveries | 3.3 million materials-science abstracts | Embeddings trained before a cutoff pick out materials later reported as thermoelectrics | One property, one field | record |
- What survives on breadth: the three dimensions of diversity behave differently and nonlinearly, so “more breadth” is not a variable; the reception benefit and the production and funding costs are separate outcomes; and the direction of the specialist advantage depends on the pace of the field, in the one natural experiment that could test it. None of it was measured in the earth sciences.
- Correction recorded September 17, 2026: earlier seminar documents stated the Teodoridis result backwards. The paper’s finding is that specialists gained when the pace of change increased.
5. Where the metrics attach on the unit graph
| Construct | Node | What the metric sees | What it misses |
|---|---|---|---|
| Novelty | G, and the claim at K | The inputs (references, keywords) or the vocabulary of the claim, at a date | Whether the claim itself had a precedent; the Dossier’s job |
| Correctness | V, X, K | Nothing directly; survival and contradiction are proxies | Whether the evidence owed was supplied; the mode page’s job |
| Impact | After R | Uptake by citation, normalized by field and year, over a long window | Use without citation: data, methods, standards, decisions |
| Field-shaping | After R, across many units | Whether citers also cite the references (CD); canon turnover; delayed awakening | Confounded by reference-list length and field size |
| Breadth | G and D | Diversity of cited categories; author topic concentration | Integration in the work; borrowed competence and its check |
6. How the club uses them as it proceeds
- Meetings 1 to 9: judgments first, numbers sealed. For each anchor and companion paper, the ledger row records the group’s judgment of novelty at the date (closest precedent), warrant, and uptake, before anyone looks at a number. The indicator tool has already computed citations, field percentile, CD_5, and the beauty coefficient for the assigned papers; the file stays in the private repository until meeting 10.
- Meeting 10, novelty: unseal the numbers. Compare the group’s closest-precedent judgments with reference-pair atypicality and with an embedding distance computed over the assigned papers’ reference lists, the Shibayama measure redone with a current embedding model. Record where the metric and the room disagree and why.
- Meeting 11, breadth: compute variety, balance, and disparity of cited categories for one paper of each participant, alongside their own held, borrowed, trusted, absent inventory from the case exercise. The comparison is the paper’s point: reference diversity against integration in the work.
- Meeting 12, advance: compare the group’s blinded advance ratings of the nine practice episodes with CD_5, percentile, and B. The Wilson and Sykes pair is the planned disagreement: the index files the test as development and the concept as disruption.
- Meeting 13 and after: the metrics that survived the three comparisons enter the historical evaluation as scores, with the disagreements as the construct-validity record. The corpus study inherits only those.
- First numbers, computed September 17, 2026 from OpenAlex, which is not the Web of Science data the original papers used. CD_5 here omits the n_k term; the full index needs each reference’s citers and is computed on request.
| Paper | Year | Citations | Field percentile | CD_5 without n_k |
|---|---|---|---|---|
| Platt (1964) | 1964 | 2886 | 99 to 100 | +0.84 |
| Cleland (2001) | 2001 | 293 | 99 to 100 | +0.75 |
| Wilson (1965) | 1965 | 1365 | 99 to 100 | +0.37 |
| Sykes (1967) | 1967 | 708 | 96 to 100 | −0.78 |
| Atwater (1987) | 1987 | 687 | 99 to 100 | −0.49 |
| Lindsey et al. (2019) | 2019 | 681 | 89 to 100 | −0.33 |
- The beauty coefficient is not reported: OpenAlex’s yearly counts start in 2012, so a 1964 paper has no dormancy period in the data.
7. Revisiting the pre-LLM studies with embeddings and computational reading
- Nearly every metric above was built on what a database could count: citations, journal categories, keyword co-occurrence. Three things have changed: documents can be embedded, so distance is continuous and does not need category conventions; a model can read a citing sentence and say what the citation does; and a model can read a full text and say what was used. Each fixes one known limitation and introduces one of its own.
| Metric | Pre-LLM operationalization and its known problem | Revisit | What it would fix | What it cannot fix |
|---|---|---|---|---|
| Novelty of inputs | Journal-pair atypicality (Uzzi; Wang 2017); measures interdisciplinarity as much as novelty (Fontana) | Embedding distance between the cited works’ content (Shibayama’s design with current models); separate content from context as Shi and Evans do | Category conventions; coarse journal units | Novelty of the claim; leakage when the embedding model has read the later literature |
| Novelty of the claim | Not measured at scale; expert judgment disagrees with itself (Boudreau) | A reading agent extracts the claim and searches a corpus cut at the date for its closest precedent, the Novelty Dossier automated | Scale | Model conservatism against novelty (Machado; NovBench); the corpus cut does not remove the outcome from the model’s training |
| Disruption | CD on citation counts; confounded by reference-list length (Petersen) and of uncertain validity (Leibel) | Classify each citation’s function by reading the citing sentence: builds on, uses, tests, supersedes, mentions; compute disruption on the “supersedes” subset | The construct problem: whether citers ignore the references because the paper replaced them or because reference lists are short | Coverage of full text; the classifier’s own validity, which needs human labels |
| Impact beyond citations | Citation counts only; data, method, and standard uptake invisible (Bornmann; Waltman) | Read full texts for use of a method, dataset, instrument, or catalog without citation; link to data repositories and software | Uptake that citations miss, which is most of instrument and infrastructure impact | Full-text access; attribution when use is unacknowledged |
| Interdisciplinarity | Variety, balance, disparity over WoS categories in references (Yegros; Wang 2015) | The same three dimensions over topic embeddings of the cited works, no categories; plus disparity between the paper’s own content and its references | Category conventions; disciplines that categories split or merge | Integration in the work, still; needs the borrowed-competence inventory |
| Depth and specialization | MeSH or topic concentration of an author’s outputs (Rassenfosse; Teodoridis) | Concentration of an author’s outputs in embedding space over time (Liu 2021 did this for hot streaks) | Field-specific vocabularies; comparability across fields | Competence, which is not concentration |
| Field-shaping | Disruption; canon turnover (Chu); sleeping beauties (Ke) | Emergence of a new region in topic-embedding space whose papers cite the focal work; awakening detected in a different field by embedding, as Ke found by categories | Cross-field awakenings that categories hide | The retrospective nature of the construct |
| Correctness | Not measured; survival as a proxy | Read later papers for replication, contradiction, or correction of a specific claim | The proxy’s blindness to silent failure | Ground truth, which needs the mode rubric applied by people |
| Question choice and breadth of a field | Topic coverage of a field over time (Hao) | Topic-embedding diversity of a field’s questions, computed prospectively as the group’s own work proceeds | Measuring the portfolio effect the agent evaluation needs | Attribution to the tool rather than to the field’s trend |
- Two cautions run through the column. Historical replay leaks: a model trained after the outcome cannot be a blind judge of novelty before it, which is why the historical evaluation scores process rather than outcome on known episodes. And the models themselves are measured to be conservative about novelty (Machado 2026) and weak at judging it (NovBench 2026), so an embedding-based novelty score is a second instrument to validate, not a replacement for the first.
- The seismology corpus is the test bed: the assigned papers first, then the corpus study sample, each metric run in its original form and its embedding form, disagreements inspected by hand.
8. What this page does not claim
- That any metric here measures advance. Each measures a component or a correlate, in one or a few fields, mostly biomedicine and physics, with citations as the outcome.
- That the findings transfer to the earth sciences. None of the studies above measured them; the corpus study would.
- That the numbers in section 6 are comparable with the published ones. They come from a different database.