Bibliometric analyses rely on accurate citation counts, yet bibliographic databases routinely contain variant representations of the same cited reference, differing in journal abbreviation style, author name format, punctuation, or metadata completeness, that fragment citation links and distort standard indicators such as the h-index and journal impact metrics. We propose an unsupervised, multi-phase reference matching algorithm designed to consolidate these variants without requiring training data or external authority files beyond the ISO 4 List of Title Word Abbreviations (LTWA). The pipeline operates in seven phases: (i) format detection and string normalisation, which parses heterogeneous reference styles and standardises author names, titles, and pagination; (ii) ISO 4 journal-name normalisation, which maps both abbreviated and full journal names to a canonical short form using the LTWA; (iii) exact matching on DOI identifiers and normalised reference strings; (iv) blocking by first-author surname and publication year; (v) within-block fuzzy matching that combines Jaro-Winkler similarity with agglomerative hierarchical clustering to group near-duplicate references; (vi) post-processing metadata reconciliation, which merges complementary fields across matched records; and (vii) canonical representative selection, which elects the most informative variant as the group representative. Evaluation on a synthetic benchmark of 1.064 source articles under 17 controlled perturbation scenarios, yields precision, recall, and F1 scores above 0.95 in 15 of 17 scenarios, with the two lowest-scoring scenarios still achieving F1 ≥ 0.78. Validation on two real-world Scopus datasets demonstrates that the algorithm reduces unique cited-reference counts by 4.6−11.1 %, with corresponding increases in concentration-based bibliometric indicators. The algorithm is implemented in the open-source bibliometrix R package, which has been widely adopted in the scientometric community.

A multi-phase reference matching algorithm for bibliometric analysis: design, implementation, and evaluation / Aria, M., D'Aniello, L., Spano, M.. - In: SCIENTOMETRICS. - ISSN 0138-9130. - (2026), pp. 1-36. [10.1007/s11192-026-05763-2]

A multi-phase reference matching algorithm for bibliometric analysis: design, implementation, and evaluation

Aria, Massimo
;
D'Aniello, Luca;Spano, Maria
2026

Abstract

Bibliometric analyses rely on accurate citation counts, yet bibliographic databases routinely contain variant representations of the same cited reference, differing in journal abbreviation style, author name format, punctuation, or metadata completeness, that fragment citation links and distort standard indicators such as the h-index and journal impact metrics. We propose an unsupervised, multi-phase reference matching algorithm designed to consolidate these variants without requiring training data or external authority files beyond the ISO 4 List of Title Word Abbreviations (LTWA). The pipeline operates in seven phases: (i) format detection and string normalisation, which parses heterogeneous reference styles and standardises author names, titles, and pagination; (ii) ISO 4 journal-name normalisation, which maps both abbreviated and full journal names to a canonical short form using the LTWA; (iii) exact matching on DOI identifiers and normalised reference strings; (iv) blocking by first-author surname and publication year; (v) within-block fuzzy matching that combines Jaro-Winkler similarity with agglomerative hierarchical clustering to group near-duplicate references; (vi) post-processing metadata reconciliation, which merges complementary fields across matched records; and (vii) canonical representative selection, which elects the most informative variant as the group representative. Evaluation on a synthetic benchmark of 1.064 source articles under 17 controlled perturbation scenarios, yields precision, recall, and F1 scores above 0.95 in 15 of 17 scenarios, with the two lowest-scoring scenarios still achieving F1 ≥ 0.78. Validation on two real-world Scopus datasets demonstrates that the algorithm reduces unique cited-reference counts by 4.6−11.1 %, with corresponding increases in concentration-based bibliometric indicators. The algorithm is implemented in the open-source bibliometrix R package, which has been widely adopted in the scientometric community.
2026
A multi-phase reference matching algorithm for bibliometric analysis: design, implementation, and evaluation / Aria, M., D'Aniello, L., Spano, M.. - In: SCIENTOMETRICS. - ISSN 0138-9130. - (2026), pp. 1-36. [10.1007/s11192-026-05763-2]
File in questo prodotto:
File Dimensione Formato  
s11192-026-05763-2.pdf

accesso aperto

Tipologia: Versione Editoriale (PDF)
Licenza: Creative commons
Dimensione 2.21 MB
Formato Adobe PDF
2.21 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11588/1061196
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact