Author name normalization is a crucial preprocessing step for bibliometric analysis, research evaluation, and the study of scientific collaboration networks, especially when data originate from heterogeneous and weakly standardized institutional repositories. This paper presents a Python-based, conservative pipeline for the normalization of author strings in bibliographic records extracted from IRIS/CRIS repositories, with a case study focused on the Department of Economics at the University of Modena and Reggio Emilia. The proposed approach addresses recurrent sources of syntactic heterogeneity in author fields, including inconsistent separators, capitalization variants, initials, orthographic variants, onomastic particles, compound surnames, and given-name/surname inversions. The pipeline combines explainable rule-based preprocessing and parsing with a corpus-adaptive expansion of initials based on internal frequency evidence and a final majority-based canonicalization of inverted full-name pairs applied after frequency stabilization. Unlike full author-name disambiguation systems, the method does not aim to resolve author identity through external or relational metadata such as ORCID identifiers, affiliations, co-author networks, venues, titles, or topics. Rather, it provides an auditable preprocessing layer that consolidates surface variants only when the available corpus evidence supports a reliable decision, while explicitly retaining unresolved cases when evidence is insufficient or ambiguity remains. Applied to 10,271 publication records and 34,746 author–publication pairs, the pipeline reduces syntactic heterogeneity and produces traceable outputs for downstream bibliometric analyses. The revised evaluation complements internal outcome counts with manual validation and sensitivity analysis of the majority threshold, thereby clarifying the precision–coverage trade-off of the proposed approach. The contribution lies in providing a transparent, reproducible, and adaptable normalization workflow for institutional repository data, designed to improve metadata quality before subsequent bibliometric or science-mapping analyses.
A hybrid rule-based and data-driven pipeline for author name normalization in IRIS institutional repositories
Pasquale Pavone
;
2026-01-01
Abstract
Author name normalization is a crucial preprocessing step for bibliometric analysis, research evaluation, and the study of scientific collaboration networks, especially when data originate from heterogeneous and weakly standardized institutional repositories. This paper presents a Python-based, conservative pipeline for the normalization of author strings in bibliographic records extracted from IRIS/CRIS repositories, with a case study focused on the Department of Economics at the University of Modena and Reggio Emilia. The proposed approach addresses recurrent sources of syntactic heterogeneity in author fields, including inconsistent separators, capitalization variants, initials, orthographic variants, onomastic particles, compound surnames, and given-name/surname inversions. The pipeline combines explainable rule-based preprocessing and parsing with a corpus-adaptive expansion of initials based on internal frequency evidence and a final majority-based canonicalization of inverted full-name pairs applied after frequency stabilization. Unlike full author-name disambiguation systems, the method does not aim to resolve author identity through external or relational metadata such as ORCID identifiers, affiliations, co-author networks, venues, titles, or topics. Rather, it provides an auditable preprocessing layer that consolidates surface variants only when the available corpus evidence supports a reliable decision, while explicitly retaining unresolved cases when evidence is insufficient or ambiguity remains. Applied to 10,271 publication records and 34,746 author–publication pairs, the pipeline reduces syntactic heterogeneity and produces traceable outputs for downstream bibliometric analyses. The revised evaluation complements internal outcome counts with manual validation and sensitivity analysis of the majority threshold, thereby clarifying the precision–coverage trade-off of the proposed approach. The contribution lies in providing a transparent, reproducible, and adaptable normalization workflow for institutional repository data, designed to improve metadata quality before subsequent bibliometric or science-mapping analyses.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
