Salvatore Bufi

dblp:356/7978 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0001-6979-1715ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Information retrieval · 43% Knowledge graphs · 38% Data integration and cleaning · 19%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › evaluation
benchmark
0.912025
DataRec: A Python Library for Standardized and Reproducible Data Management in Recommender Systems · SIGIR 2025
Data integration and cleaning
data preprocessing
0.912025
DataRec: A Python Library for Standardized and Reproducible Data Management in Recommender Systems · SIGIR 2025
Knowledge graphs › link prediction
inductive link prediction
0.912025
Type-Less yet Type-Aware Inductive Link Prediction with Pretrained Language Models · EMNLP 2025
Knowledge graphs
link prediction
0.912025
Type-Less yet Type-Aware Inductive Link Prediction with Pretrained Language Models · EMNLP 2025
Information retrieval › evaluation › evaluation methodology
reproducibility
0.912025
DataRec: A Python Library for Standardized and Reproducible Data Management in Recommender Systems · SIGIR 2025
Information retrieval
evaluation
0.312025
DataRec: A Python Library for Standardized and Reproducible Data Management in Recommender Systems · SIGIR 2025

Methods — techniques the papers use, named apart from their topics

pre-trained language model · 1.7
YearPublicationVenuePosition
2025 Type-Less yet Type-Aware Inductive Link Prediction with Pretrained Language Models
abstract
Alessandro De Bellis, Salvatore Bufi, Giovanni Servedio, Vito Walter Anelli, Tommaso Di Noia, Eugenio Di Sciascio. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Alessandro De Bellis, Salvatore Bufi, Giovanni Servedio, Vito Walter Anelli, Tommaso Di Noia, Eugenio Di Sciascio
EMNLP2
2025 DataRec: A Python Library for Standardized and Reproducible Data Management in Recommender Systems
abstract
Recommender systems have demonstrated a significant impact across diverse domains, yet ensuring the reproducibility of experimental findings remains a persistent challenge.A primary obstacle lies in the fragmented and often opaque data management strategies employed during the preprocessing stage, where decisions about dataset selection, filtering, and splitting can substantially influence outcomes.To address these limitations, we introduce DataRec, an open-source Python-based library specifically designed to unify and streamline data handling in recommender system research.By providing reproducible routines for dataset preparation, data versioning, and seamless integration with other frameworks, DataRec promotes methodological standardization, interoperability, and comparability across different experimental setups.Our design is informed by an in-depth review of 55 stateof-the-art recommendation studies, ensuring that DataRec adopts best practices while addressing common pitfalls in data management.Ultimately, our contribution facilitates fair benchmarking, enhances reproducibility, and fosters greater trust in experimental results within the broader recommender systems community.The DataRec library, documentation, and examples are freely available at https://github.com/sisinflab/DataRec.
Alberto Carlo Maria Mancino, Salvatore Bufi, Angela Di Fazio, Antonio Ferrara 0001, Daniele Malitesta, Claudio Pomo, Tommaso Di Noia
SIGIR2
2025 Legal but Unfair: Auditing the Impact of Data Minimization on Fairness and Accuracy Trade-off in Recommender Systems
abstract
Data minimization, required by recent data privacy regulations, is crucial for user privacy, but its impact on recommender systems remains largely unclear.The core problem lies in the fact that reducing or altering the training data of these systems can drastically affect their performance.While previous research has explored how data minimization affects recommendation accuracy, a critical gap remains: How does data minimization impact consumers' and providers' fairness?This study addresses this gap by systematically examining how data minimization influences multiple objectives in recommender systems, i.e., the trade-offs between accuracy, user fairness, and provider fairness.Our investigation includes (i) an analysis of how the data minimization strategies affect RS performance across these objectives, (ii) an assessment of data minimization techniques to determine those that can balance better the trade-off among the considered objectives, and (iii) an evaluation of the robustness of different recommendation models under diverse minimization strategies to identify those that best maintain performance.The findings reveal that data minimization can sometimes undermine provider fairness, albeit enhancing groupbased consumer fairness to the detriment of accuracy.Additionally, different strategies can offer diverse trade-offs for the assessed objectives.The source code supporting this study is available at https://github.com/salvatore-bufi/DataMinimizationFairness.
Salvatore Bufi, Vincenzo Paparella, Vito Walter Anelli, Tommaso Di Noia
UMAP1
2023 KGTORe: Tailored Recommendations through Knowledge-aware GNN Models
abstract
Knowledge graphs (KG) have been proven to be a powerful source of side information to enhance the performance of recommendation algorithms. Their graph-based structure paves the way for the adoption of graph-aware learning models such as Graph Neural Networks (GNNs). In this respect, state-of-the-art models achieve good performance and interpretability via user-level combinations of intents leading users to their choices. Unfortunately, such results often come from and end-to-end learnings that considers a combination of the whole set of features contained in the KG without any analysis of the user decisions. In this paper, we introduce KGTORe, a GNN-based model that exploits KG to learn latent representations for the semantic features, and consequently, interpret the user decisions as a personal distillation of the item feature representations. Differently from previous models, KGTORe does not need to process the whole KG at training time but relies on a selection of the most discriminative features for the users, thus resulting in improved performance and personalization. Experimental results on three well-known datasets show that KGTORe achieves remarkable accuracy performance and several ablation studies demonstrate the effectiveness of its components. The implementation of KGTORe is available at: https://github.com/sisinflab/KGTORe.
Alberto Carlo Maria Mancino, Antonio Ferrara 0001, Salvatore Bufi, Daniele Malitesta, Tommaso Di Noia, Eugenio Di Sciascio
RecSys3