TY - GEN
T1 - Comparative Performance Analysis of String Matching Algorithms and Data Matching Frameworks Using Python Libraries on Academic Datasets
AU - Waradana, Muhammad Ridho
AU - Rakhmanda, Venia Anisya
AU - Bhamakerti, Ganendra Aby
AU - Rosyidan, Fikri Yoma
AU - Rakhmawati, Nur Aini
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - This study conducts a comparative analysis of diverse string matching algorithms (Jaro-Winkler, Edit Distance, Jaccard, Hamming) and data matching frameworks (Splink, RLTK, etc.) implemented using Python libraries on degraded academic datasets. Methodologically, performance was quantified via precision, recall, and F1-score across varying degrees of data corruption. Findings indicate that no single method achieves universal optimality. Jaro-Winkler and the Splink framework demonstrate superior precision, minimizing false positives. Conversely, Edit Distance and the RLTK framework yield maximum recall, capturing a broader range of potential matches but at a higher risk of false positives. The implication is that reliable entity resolution necessitates a tiered, adaptive pipeline. This integrated approach balances the trade-offs between precision and recall, supporting consistent, cost-effective, and reliable entity resolution across heterogeneous academic records.
AB - This study conducts a comparative analysis of diverse string matching algorithms (Jaro-Winkler, Edit Distance, Jaccard, Hamming) and data matching frameworks (Splink, RLTK, etc.) implemented using Python libraries on degraded academic datasets. Methodologically, performance was quantified via precision, recall, and F1-score across varying degrees of data corruption. Findings indicate that no single method achieves universal optimality. Jaro-Winkler and the Splink framework demonstrate superior precision, minimizing false positives. Conversely, Edit Distance and the RLTK framework yield maximum recall, capturing a broader range of potential matches but at a higher risk of false positives. The implication is that reliable entity resolution necessitates a tiered, adaptive pipeline. This integrated approach balances the trade-offs between precision and recall, supporting consistent, cost-effective, and reliable entity resolution across heterogeneous academic records.
KW - Academic Datasets
KW - Data Matching
KW - Python Libraries
KW - Similarity
KW - String Matching
UR - https://www.scopus.com/pages/publications/105036311992
U2 - 10.1109/3ict68299.2025.11442241
DO - 10.1109/3ict68299.2025.11442241
M3 - Conference contribution
AN - SCOPUS:105036311992
T3 - 2025 International Conference on Innovation and Intelligence for Informatics, Computing,and Technologies, 3ICT 2025
BT - 2025 International Conference on Innovation and Intelligence for Informatics, Computing,and Technologies, 3ICT 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 International Conference on Innovation and Intelligence for Informatics, Computing,and Technologies, 3ICT 2025
Y2 - 17 November 2025 through 19 November 2025
ER -