In our primary work (Gagliardelli et al., 2024), we presented Generalized Supervised Meta-blocking, a novel approach that is based on the task of associating every pair of candidates with a probabilistic estimation of their matching likelihood. It formulates Meta-blocking as a probabilistic binary classification task, using a wide variety of weighting schemes as features. The resulting probability estimations can be combined with any pruning algorithm that improves blocking by retaining the best candidate comparisons, drastically reducing false positives without any significant impact on true positives. They can also be used by state-of-the-art progressive Entity Resolution methods to identify the most promising candidates as early as possible. We experimentally identified the best pruning algorithms, their optimal sets of features, and the minimum possible size of the training set. Our experiments demonstrate that the resulting approaches achieve excellent performance in several established benchmark datasets. This work introduces a companion reproducible paper of our previous work (Gagliardelli et al., 2024) to describe how to reproduce the entire experimental study and discuss how to extend it with new features, classification algorithms, and datasets. The presented reproducibility methodology is based on a Docker image and the NVIDIA CUDA Toolkit 12.0, which are evaluated on several Linux-based platforms equipped with at least 128 GB RAM, a GPU and at least 300 GB of free disk space. The reproducibility methodology leads to a weak reproducibility, with minor deviations from the experimental results reported in our primary work (Gagliardelli et al., 2024).

Reproducible experiments on a Generalized Approach to Supervised Meta-blocking for Scalable Entity Resolution

Gagliardelli, Luca
;
2026-01-01

Abstract

In our primary work (Gagliardelli et al., 2024), we presented Generalized Supervised Meta-blocking, a novel approach that is based on the task of associating every pair of candidates with a probabilistic estimation of their matching likelihood. It formulates Meta-blocking as a probabilistic binary classification task, using a wide variety of weighting schemes as features. The resulting probability estimations can be combined with any pruning algorithm that improves blocking by retaining the best candidate comparisons, drastically reducing false positives without any significant impact on true positives. They can also be used by state-of-the-art progressive Entity Resolution methods to identify the most promising candidates as early as possible. We experimentally identified the best pruning algorithms, their optimal sets of features, and the minimum possible size of the training set. Our experiments demonstrate that the resulting approaches achieve excellent performance in several established benchmark datasets. This work introduces a companion reproducible paper of our previous work (Gagliardelli et al., 2024) to describe how to reproduce the entire experimental study and discuss how to extend it with new features, classification algorithms, and datasets. The presented reproducibility methodology is based on a Docker image and the NVIDIA CUDA Toolkit 12.0, which are evaluated on several Linux-based platforms equipped with at least 128 GB RAM, a GPU and at least 300 GB of free disk space. The reproducibility methodology leads to a weak reproducibility, with minor deviations from the experimental results reported in our primary work (Gagliardelli et al., 2024).
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11389/94215
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact