Alejandro Bellogín

dblp:23/1049 · also Alejandro Bellogín Kouki · DBLP profile ↗
← Back
65ranked-venue papers in the field
20as first author
23since 2021 · last 2026
0000-0001-6368-2510ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 52 (16 first)Data Mining & Knowledge Discovery · 6 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 4 (2 first)Database Systems & Data Management · 2 (1 first)Other / Interdisciplinary · 1
YearPublicationVenuePosition
2026 SynIRgy: Synthetic Data and Simulation Synergy for Information Retrieval
Manel Slokom, Alejandro Bellogín, Andrea Barraza-Urbina
ECIR (3)2
2026 ACE: Semantically-Grounded Graph Alignment via Affective Contrastive Learning
abstract
Graph Contrastive Learning (GCL) methods for recommendation learn representations by propagating signals over user-item interaction graphs. However, modeling these graphs with homogeneous edges can lead to semantic-structural misalignment, where information is exchanged between structurally adjacent but semantically dissimilar items, adversely affecting retrieval quality. Existing solutions typically rely on auxiliary encoders or additional supervision, increasing model complexity and training cost. We propose ACE (Affective Contrastive Embeddings), a framework that improves representation alignment by incorporating affective semantics into contrastive learning. ACE encodes the affective dimensions of valence and arousal as a topological prior, encouraging consistency between learned embeddings and an affective semantic space distilled from large language models. To operationalize this alignment, we introduce a Semantically Weighted Noise Contrastive Estimation (SW-NCE) loss that modulates contrastive gradients according to users' affective preferences. Experiments on Amazon, Last.fm, and SiTunes demonstrate that ACE consistently improves top-K retrieval performance over 11 baselines while reducing computational overhead. These results indicate that affective geometric alignment is an effective and efficient mechanism for enhancing graph-based retrieval models.
Potito Aghilar, Sabino Roccotelli, Vito Walter Anelli, Alejandro Bellogín, Michelantonio Trizio, Tommaso Di Noia
SIGIR4
2025 Improving Novelty and Diversity of Nearest-Neighbors Recommendation by Exploiting Dissimilarities
Pablo Sánchez 0001, Javier Sanz-Cruzado, Alejandro Bellogín
ECIR (4)3
2025 The Impact of Mainstream-Driven Algorithms on Recommendations for Children
Robin Ungruh, Alejandro Bellogín, Maria Soledad Pera
ECIR (3)2
2025 Context Trails: A Dataset to Study Contextual and Route Recommendation
abstract
Recommender systems in the tourism domain are gaining increasing attention, yet the development of diverse recommendation tasks remains limited, largely due to the scarcity of public datasets.This paper introduces Context Trails, a novel dataset addressing this gap.Context Trails distinguishes itself by including not only user interactions with touristic venues, but also the itineraries (trails or routes) followed by users.Furthermore, it enriches existing item features (e.g., category, coordinates) with contextual attributes related to the interaction moment (e.g., weather) and the venue itself (e.g., opening hours).Beyond a detailed description of the dataset's characteristics, we evaluate the performance of several baseline algorithms across three distinct recommendation tasks: classical recommendation, route recommendation, and contextual recommendation.We believe this dataset will foster further research and development of advanced recommender systems within the tourism domain.
Pablo Sánchez 0001, Alejandro Bellogín, Jose L. Jorro-Aragoneses
RecSys2
2025 Impacts of Mainstream-Driven Algorithms on Recommendations for Children Across Domains: A Reproducibility Study
abstract
Children are often exposed to items curated by recommendation algorithms. Yet, research seldom considers children as a user group, and when it does, it is anchored on datasets where children are underrepresented, risking overlooking their interests, favoring those of the majority, i.e., mainstream users. Recently, Ungruh et al. demonstrated that children's consumption patterns and preferences differ from those of mainstream users, resulting in inconsistent recommendation algorithm performance and behavior for this user group. These findings, however, are based on two datasets with a limited child user sample. We reproduce and replicate this study on a wider range of datasets in the movie, music, and book domains, uncovering interaction patterns and aspects of child-recommender interactions consistent across domains, as well as those specific to some user samples in the data. We also extend insights from the original study with popularity bias metrics, given the interpretation of results from the original study. With this reproduction and extension, we uncover consumption patterns and differences between age groups stemming from intrinsic differences between children and others, and those unique to specific datasets or domains.
Robin Ungruh, Alejandro Bellogín, Dominik Kowald, Maria Soledad Pera
RecSys2
2025 From Previous Plays to Long-Term Tastes: Exploring the Long-term Reliability of Recommender Systems Simulations for Children
Robin Ungruh, Alejandro Bellogín, Maria Soledad Pera
RecSys2
2025 International Workshop on Algorithmic Bias in Search and Recommendation (BIAS 2025)
abstract
Designing search and recommendation models that are both efficient and effective has long been a central objective for both industry professionals and academic researchers. Yet, growing evidence highlights how models trained on historical data can reinforce pre-existing biases, potentially leading to harmful outcomes. Addressing these challenges by defining, evaluating, and mitigating bias across development workflows is a crucial step toward the responsible deployment of search and recommendation models in practice. The BIAS 2025 workshop seeks to gather innovative research and foster a shared space for dialogue among researchers and practitioners committed to advancing this fundamental direction. Workshop website: https://biasinrecsys.github.io/sigir2025/.
Alejandro Bellogín, Ludovico Boratto, Styliani Kleanthous, Elisabeth Lex, Francesca Maridina Malloci, Mirko Marras
SIGIR1
2025 From Monolith to Mosaic: Uncovering Behavioral Differences for Choice Models in Recommender Systems Simulations
abstract
Simulation is widely used in recommender systems research to study algorithm behavior and its impact on users. A common strategy involves adopting a universal choice model to represent users, assuming all follow the same consumption patterns. This one-size-fits-all approach overlooks the diversity in user preferences and decision-making patterns. In this work, we scrutinize whether this universal view fails to account for unique user behavior, thus harming realism and reliability of simulation outcomes. We conduct multiple simulations with various recommendation algorithms and choice models in the movie domain, comparing outcomes to users' organic consumption patterns. Further, we evaluate whether a holistic model that captures users' differences in behavior would better reflect a wide user base. Our findings highlight the limitations of using a naive, universal choice model and emphasize the need for more nuanced, user-specific approaches to make contributions from simulation studies more reflective of real-world effects.
Robin Ungruh, Alejandro Bellogín, Maria Soledad Pera
SIGIR2
2025 The role of recommendation algorithms in the formation of disinformation networks
abstract
Disinformation on social networks, especially those that share media content, remains a critical issue with far-reaching societal implications. Although extensive research has addressed the prevalence and mitigation of false information, the specific impact of recommendation algorithms on the creation and consolidation of disinformation networks has not been thoroughly examined. In this work, we bridge this gap by simulating how various recommendation techniques — ranging from basic yet foundational approaches such as popularity-based and content-based methods — shape network dynamics and facilitate disinformation spread. These classical algorithms are essential building blocks of modern hybrid and task-specific recommender systems; understanding their effects is thus crucial for assessing systemic risks. Using a dataset comprising tweets from 275 disinformation agents and 275 legitimate journalism agents, we conduct a realistic simulation grounded in probabilistic click models of user behavior and real-world social media data. Our findings reveal that certain recommendation approaches can significantly reinforce the cohesion and visibility of disinformation networks, thereby amplifying their reach. These results underscore the necessity for algorithmic accountability and the design of ethically responsible recommender systems to maintain information integrity on social platforms. • Dense and cohesive network structures facilitate rapid disinformation spread. • Higher accuracy algorithms correlate with more cohesive disinformation networks. • DeepRank and Word2Vec contribute to formation and consolidation of disinformation networks. • New metrics reveal disinformation amplification in legitimate networks by certain algorithms. • Promoting diversity in recommenders can reduce the formation of disinformation networks.
Pau Muñoz, Raúl Barba-Rojas, Fernando Díez, Alejandro Bellogín
Inf. Process. Manag.4
2025 Smart Imputation, Better Recommendations: Improving Traditional Point-of-Interest Recommendation through Data Augmentation
abstract
Data sparsity is a persistent challenge in recommender systems, especially in specific domains like Point-of-Interest (POI) recommendation, where it significantly impacts model performance. While classical recommender systems have used various imputation and data augmentation mechanisms to address data sparsity, these methods have not been extensively explored in the POI recommendation domain. In this work, we propose a generic imputation framework to study the use of data augmentation techniques to generate synthetic check-ins and analyze their effects on the POI recommendation scenario. Our main goal is to enhance the performance of various traditional recommenders by increasing the training set interactions, considering specific characteristics of the domain, such as geographical information. We apply these techniques in six different cities from a global Foursquare check-in dataset, as well as in two additional cities from the Gowalla dataset, and a separate dataset from Yelp, ensuring a comprehensive evaluation across multiple data sources. Our imputation approach evidences improvements for most models. In several cases, these improvements exceeded 100% for ranking accuracy, measured in terms of nDCG, without considerably compromising novelty or diversity. Data and code are released at https://github.com/pablosanchezp/ImputationForPOIRecsys .
Pablo Sánchez 0001, Alejandro Bellogín
ACM Trans. Intell. Syst. Technol.2
2024 Recommendation Fairness in eParticipation: Listening to Minority, Vulnerable and NIMBY Citizens
Marina Alonso-Cortés, Iván Cantador, Alejandro Bellogín
ECIR (4)3
2024 International Workshop on Algorithmic Bias in Search and Recommendation (BIAS)
abstract
Creating efficient and effective search and recommendation algorithms has been the main objective of industry practitioners and academic researchers over the years. However, recent research has shown how these algorithms trained on historical data lead to models that might exacerbate existing biases and generate potentially negative outcomes. Defining, assessing, and mitigating these biases throughout experimental pipelines is a primary step for devising search and recommendation algorithms that can be responsibly deployed in real-world applications. This workshop aims to collect novel contributions in this field and offer a common ground for interested researchers and practitioners. More information about the workshop is available at https://biasinrecsys.github.io/sigir2024/
Alejandro Bellogín, Ludovico Boratto, Styliani Kleanthous, Elisabeth Lex, Francesca Maridina Malloci, Mirko Marras
SIGIR1
2024 Correction to: Bias characterization, assessment, and mitigation in location-based recommender systems
Pablo Sánchez 0001, Alejandro Bellogín, Ludovico Boratto
Data Min. Knowl. Discov.2
2023 Challenging the Myth of Graph Collaborative Filtering: a Reasoned and Reproducibility-driven Analysis
abstract
The success of graph neural network-based models (GNNs) has significantly advanced recommender systems by effectively modeling users and items as a bipartite, undirected graph. However, many original graph-based works often adopt results from baseline papers without verifying their validity for the specific configuration under analysis. Our work addresses this issue by focusing on the replicability of results. We present a code that successfully replicates results from six popular and recent graph recommendation models (NGCF, DGCF, LightGCN, SGL, UltraGCN, and GFCF) on three common benchmark datasets (Gowalla, Yelp 2018, and Amazon Book). Additionally, we compare these graph models with traditional collaborative filtering models that historically performed well in offline evaluations. Furthermore, we extend our study to two new datasets (Allrecipes and BookCrossing) that lack established setups in existing literature. As the performance on these datasets differs from the previous benchmarks, we analyze the impact of specific dataset characteristics on recommendation accuracy. By investigating the information flow from users’ neighborhoods, we aim to identify which models are influenced by intrinsic features in the dataset structure. The code to reproduce our experiments is available at: https://github.com/sisinflab/Graph-RSs-Reproducibility.
Vito Walter Anelli, Daniele Malitesta, Claudio Pomo, Alejandro Bellogín, Eugenio Di Sciascio, Tommaso Di Noia
RecSys4
2023 Bias characterization, assessment, and mitigation in location-based recommender systems
abstract
Abstract Location-Based Social Networks stimulated the rise of services such as Location-based Recommender Systems. These systems suggest to users points of interest (or venues) to visit when they arrive in a specific city or region. These recommendations impact various stakeholders in society, like the users who receive the recommendations and venue owners. Hence, if a recommender generates biased or polarized results, this affects in tangible ways both the experience of the users and the providers’ activities. In this paper, we focus on four forms of polarization, namely venue popularity, category popularity, venue exposure, and geographical distance. We characterize them on different families of recommendation algorithms when using a realistic (temporal-aware) offline evaluation methodology while assessing their existence. Besides, we propose two automatic approaches to mitigate those biases. Experimental results on real-world data show that these approaches are able to jointly improve the recommendation effectiveness, while alleviating these multiple polarizations.
Pablo Sánchez 0001, Alejandro Bellogín, Ludovico Boratto
Data Min. Knowl. Discov.2
2023 A unifying and general account of fairness measurement in recommender systems
abstract
Fairness is fundamental to all information access systems, including recommender systems. However, the landscape of fairness definition and measurement is quite scattered with many competing definitions that are partial and often incompatible. There is much work focusing on specific – and different – notions of fairness and there exist dozens of metrics of fairness in the literature, many of them redundant and most of them incompatible. In contrast, to our knowledge, there is no formal framework that covers all possible variants of fairness and allows developers to choose the most appropriate variant depending on the particular scenario. In this paper, we aim to define a general, flexible, and parameterizable framework that covers a whole range of fairness evaluation possibilities. Instead of modeling the metrics based on an abstract definition of fairness, the distinctive feature of this study compared to the current state of the art is that we start from the metrics applied in the literature to obtain a unified model by generalization. The framework is grounded on a general work hypothesis: interpreting the space of users and items as a probabilistic sample space, two fundamental measures in information theory (Kullback–Leibler Divergence and Mutual Information) can capture the majority of possible scenarios for measuring fairness on recommender system outputs. In addition, earlier research on fairness in recommender systems could be viewed as single-sided, trying to optimize some form of equity across either user groups or provider/procurer groups, without considering the user/item space in conjunction, thereby overlooking/disregarding the interplay between user and item groups. Instead, our framework includes the notion of statistical independence between user and item groups. We finally validate our approach experimentally on both synthetic and real data according to a wide range of state-of-the-art recommendation algorithms and real-world data sets, showing that with our framework we can measure fairness in a general, uniform, and meaningful way.
Enrique Amigó, Yashar Deldjoo, Stefano Mizzaro, Alejandro Bellogín
Inf. Process. Manag.4
2022 A reproducible POI recommendation framework: Works mapping and benchmark evaluation
abstract
This work is a companion reproducibility paper that presents a framework to reproduce our previous experiments and results reported in Werneck et al. (2021). In that previous paper, we introduced a systematic mapping process of points-of-interest (POI) recommendation methods and provided a uniform evaluation methodology based on metrics covering different aspects besides accuracy. Due to the lack of reproducible and extensible benchmarks, our work introduces a reproducibility framework for POI methods based on a collection of Python software libraries and a Docker image. Our proposal is composed of: (1) a package to perform a protocol that reproduces our systematic mapping process Werneck et al. (2021), containing all collected data, insightful views on current advances and opened challenges; and (2) an extensible benchmark to perform a protocol to reproduce experimental evaluations on POI recommendation, considering different datasets, metrics, and the strongest baselines in the literature. This work also demonstrates all processes required to instantiate its framework. Moreover, our work can be considered at least weakly reproducible, since we were able to reproduce the results of the previous paper, leading us to the same conclusions.
Heitor Werneck, Nícollas Silva, Adriano C. M. Pereira, Matheus Carvalho Viana, Alejandro Bellogín, Jorge Martinez-Gil, Fernando Mourão, Leonardo Rocha 0001
Inf. Syst.5
2021 V-Elliot: Design, Evaluate and Tune Visual Recommender Systems
abstract
The paper introduces Visual-Elliot (V-Elliot), a reproducibility framework for Visual Recommendation systems (VRSs) based on Elliot. framework provides the widest set of VRSs compared to other recommendation frameworks in the literature (i.e., 6 state-of-the-art models which have been commonly employed as baselines in recent works). The framework pipeline spans from the dataset preprocessing and item visual features loading to easily train and test complex combinations of visual models and evaluation settings. V-Elliot provides an extended set of features to ease the design, testing, and integration of novel VRSs into V-Elliot. The framework exploits of dataset filtering/splitting functions, 40 evaluation metrics, five hyper-parameter optimization methods, more than 50 recommendation algorithms, and two statistical hypothesis tests. The files of this demonstration are available at: github.com/sisinflab/elliot.
Vito Walter Anelli, Alejandro Bellogín, Antonio Ferrara 0001, Daniele Malitesta, Felice Antonio Merra, Claudio Pomo, Francesco M. Donini, Tommaso Di Noia
RecSys2
2021 Reenvisioning the comparison between Neural Collaborative Filtering and Matrix Factorization
abstract
Collaborative filtering models based on matrix factorization and learned similarities using Artificial Neural Networks (ANNs) have gained significant attention in recent years. This is, in part, because ANNs have demonstrated very good results in a wide variety of recommendation tasks. However, the introduction of ANNs within the recommendation ecosystem has been recently questioned, raising several comparisons in terms of efficiency and effectiveness. One aspect most of these comparisons have in common is their focus on accuracy, neglecting other evaluation dimensions important for the recommendation, such as novelty, diversity, or accounting for biases. In this work, we replicate experiments from three different papers that compare Neural Collaborative Filtering (NCF) and Matrix Factorization (MF), to extend the analysis to other evaluation dimensions. First, our contribution shows that the experiments under analysis are entirely reproducible, and we extend the study including other accuracy metrics and two statistical hypothesis tests. Second, we investigated the Diversity and Novelty of the recommendations, showing that MF provides a better accuracy also on the long tail, although NCF provides a better item coverage and more diversified recommendation lists. Lastly, we discuss the bias effect generated by the tested methods. They show a relatively small bias, but other recommendation baselines, with competitive accuracy performance, consistently show to be less affected by this issue. This is the first work, to the best of our knowledge, where several complementary evaluation dimensions have been explored for an array of state-of-the-art algorithms covering recent adaptations of ANNs and MF. Hence, we aim to show the potential these techniques may have on beyond-accuracy evaluation while analyzing the effect on reproducibility these complementary dimensions may spark. The code to reproduce the experiments is publicly available on GitHub at https://tny.sh/Reenvisioning.
Vito Walter Anelli, Alejandro Bellogín, Tommaso Di Noia, Claudio Pomo
RecSys2
2021 Elliot: A Comprehensive and Rigorous Framework for Reproducible Recommender Systems Evaluation
abstract
Recommender Systems have shown to be an effective way to alleviate the over-choice problem and provide accurate and tailored recommendations. However, the impressive number of proposed recommendation algorithms, splitting strategies, evaluation protocols, metrics, and tasks, has made rigorous experimental evaluation particularly challenging. Puzzled and frustrated by the continuous recreation of appropriate evaluation benchmarks, experimental pipelines, hyperparameter optimization, and evaluation procedures, we have developed an exhaustive framework to address such needs. Elliot is a comprehensive recommendation framework that aims to run and reproduce an entire experimental pipeline by processing a simple configuration file. The framework loads, filters, and splits the data considering a vast set of strategies (13 splitting methods and 8 filtering approaches, from temporal training-test splitting to nested K-folds Cross-Validation). Elliot(https://github.com/sisinflab/elliot) optimizes hyperparameters (51 strategies) for several recommendation algorithms (50), selects the best models, compares them with the baselines providing intra-model statistics, computes metrics (36) spanning from accuracy to beyond-accuracy, bias, and fairness, and conducts statistical analysis (Wilcoxon and Paired t-test).
Vito Walter Anelli, Alejandro Bellogín, Antonio Ferrara 0001, Daniele Malitesta, Felice Antonio Merra, Claudio Pomo, Francesco M. Donini, Tommaso Di Noia
SIGIR2
2021 Explaining recommender systems fairness and accuracy through the lens of data characteristics
abstract
The impact of data characteristics on the performance of classical recommender systems has been recently investigated and produced fruitful results about the relationship they have with recommendation accuracy. This work provides a systematic study on the impact of broadly chosen data characteristics (DCs) of recommender systems. This is applied to the accuracy and fairness of several variations of CF recommendation models. We focus on a suite of DCs that capture properties about the structure of the user–item interaction matrix, the rating frequency, item properties, or the distribution of rating values. Experimental validation of the proposed system involved large-scale experiments by performing 23,400 recommendation simulations on three real-world datasets in the movie (ML-100K and ML-1M) and book domains (BookCrossing). The validation results show that the investigated DCs in some cases can have up to 90% of explanatory power – on several variations of classical CF algorithms –, while they can explain – in the best case – about 40% of fairness results (measured according to user gender and age sensitive attributes). Therefore, this work evidences that it is more difficult to explain variations in performance when dealing with fairness dimension than accuracy.
Yashar Deldjoo, Alejandro Bellogín, Tommaso Di Noia
Inf. Process. Manag.2
2021 On the effects of aggregation strategies for different groups of users in venue recommendation
abstract
Artículos en revistas
Pablo Sánchez 0001, Alejandro Bellogín
Inf. Process. Manag.2
2020 Time and sequence awareness in similarity metrics for recommendation
Pablo Sánchez 0001, Alejandro Bellogín
Inf. Process. Manag.2
2020 Assessing ranking metrics in top-N recommendation
Daniel Valcarce, Alejandro Bellogín, Javier Parapar, Pablo Castells
Inf. Retr. J.2
2020 Exploiting recommendation confidence in decision-aware recommender systems
Rus M. Mesas, Alejandro Bellogín
J. Intell. Inf. Syst.2
2019 Attribute-based evaluation for recommender systems: incorporating user and item attributes in evaluation metrics
abstract
Research in Recommender Systems evaluation remains critical to study the efficiency of developed algorithms. Even if different aspects have been addressed and some of its shortcomings - such as biases, robustness, or cold start - have been analyzed and solutions or guidelines have been proposed, there are still some gaps that need to be further investigated. At the same time, the increasing amount of data collected by most recommender systems allows to gather valuable information from users and items which is being neglected by classical offline evaluation metrics. In this work, we integrate such information into the evaluation process in two complementary ways: on the one hand, we aggregate any evaluation metric according to the groups defined by the user attributes, and, on the other hand, we exploit item attributes to consider some recommended items as surrogates of those interacted by the user, with a proper penalization. Our results evidence that this novel evaluation methodology allows to capture different nuances of the algorithms performance, inherent biases in the data, and even fairness of the recommendations.
Pablo Sánchez 0001, Alejandro Bellogín
RecSys2
2019 Building user profiles based on sequences for content and collaborative filtering
Pablo Sánchez 0001, Alejandro Bellogín
Inf. Process. Manag.2
2018 Time-Aware Novelty Metrics for Recommender Systems
Pablo Sánchez 0001, Alejandro Bellogín
ECIR2
2018 Measuring anti-relevance: a study on when recommendation algorithms produce bad suggestions
abstract
Typically, performance of recommender systems has been measured focusing on the amount of relevant items recommended to the users. However, this perspective provides an incomplete view of an algorithm's quality, since it neglects the amount of negative recommendations by equating the unknown and negatively interacted items when computing ranking-based evaluation metrics. In this paper, we propose an evaluation framework where anti-relevance is seamlessly introduced in several ranking-based metrics; in this way, we obtain a different perspective on how recommenders behave and the type of suggestions they make. Based on our results, we observe that non-personalized approaches tend to return less bad recommendations than personalized ones, however the amount of unknown recommendations is also larger, which explains why the latter tend to suggest more relevant items. Our metrics based on anti-relevance also show the potential to discriminate between algorithms whose performance is very similar in terms of relevance.
Pablo Sánchez 0001, Alejandro Bellogín
RecSys2
2018 Towards an open, collaborative REST API for recommender systems
abstract
Recommender Systems aim to suggest relevant items to users, however, for this they need to properly obtain/serve different types of data from/to the users of such systems. In this work, we propose and show an example implementation for a common REST API focused on Recommender Systems. This API meets the most typical requirements faced by Recommender Systems practitioners while, at the same time, is open and flexible to be extended, based on the feedback from the community. We also present a Web client that demonstrates the functionalities of the proposed API.
Alejandro Bellogín
RecSys2
2018 On the robustness and discriminative power of information retrieval metrics for top-N recommendation
abstract
The evaluation of Recommender Systems is still an open issue in the field. Despite its limitations, offline evaluation usually constitutes the first step in assessing recommendation methods due to its reduced costs and high reproducibility. Selecting the appropriate metric is a critical and ranking accuracy usually attracts the most attention nowadays. In this paper, we aim to shed light on the advantages of different ranking metrics which were previously used in Information Retrieval and are now used for assessing top-N recommenders. We propose methodologies for comparing the robustness and the discriminative power of different metrics. On the one hand, we study cut-offs and we find that deeper cut-offs offer greater robustness and discriminative power. On the other hand, we find that precision offers high robustness and Normalised Discounted Cumulative Gain provides the best discriminative power.
Daniel Valcarce, Alejandro Bellogín, Javier Parapar, Pablo Castells
RecSys2
2017 Evaluating Decision-Aware Recommender Systems
abstract
The main goal of a Recommender System is to suggest relevant items to users, although other utility dimensions - such as diversity, novelty, confidence, possibility of providing explanations - are often considered. In this work, in order to increase the amount of relevant items presented to the user, we analyse how the system could measure the confidence on its own recommendations, so it has the capability of taking decisions about whether an item should be recommended or not. A direct consequence of this design is that the number of suggested items decreases, impacting in some of the beyond-accuracy dimensions (especially, coverage). We present an evaluation of different decision-aware techniques that can be applied to some families of recommender systems, and explore evaluation metrics that allow to combine more than one evaluation dimension. Empiric results show that large precision improvements are obtained when using these approaches at the expense of user and item coverage.
Rus M. Mesas, Alejandro Bellogín
RecSys2
2017 Statistical biases in Information Retrieval metrics for recommender systems
Alejandro Bellogín, Pablo Castells, Iván Cantador
Inf. Retr. J.1
2017 Collaborative filtering based on subsequence matching: A new approach
Alejandro Bellogín, Pablo Sánchez 0001
Inf. Sci.1
2016 The strange case of reproducibility versus representativeness in contextual suggestion test collections
abstract
The most common approach to measuring the effectiveness of Information Retrieval systems is by using test collections. The Contextual Suggestion (CS) TREC track provides an evaluation framework for systems that recommend items to users given their geographical context. The specific nature of this track allows the participating teams to identify candidate documents either from the Open Web or from the ClueWeb12 collection, a static version of the web. In the judging pool, the documents from the Open Web and ClueWeb12 collection are distinguished. Hence, each system submission should be based only on one resource, either Open Web (identified by URLs) or ClueWeb12 (identified by ids). To achieve reproducibility, ranking web pages from ClueWeb12 should be the preferred method for scientific evaluation of CS systems, but it has been found that the systems that build their suggestion algorithms on top of input taken from the Open Web achieve consistently a higher effectiveness. Because most of the systems take a rather similar approach to making CSs, this raises the question whether systems built by researchers on top of ClueWeb12 are still representative of those that would work directly on industry-strength web search engines. Do we need to sacrifice reproducibility for the sake of representativeness? We study the difference in effectiveness between Open Web systems and ClueWeb12 systems through analyzing the relevance assessments of documents identified from both the Open Web and ClueWeb12. Then, we identify documents that overlap between the relevance assessments of the Open Web and ClueWeb12, observing a dependency between relevance assessments and the source of the document being taken from the Open Web or from ClueWeb12. After that, we identify documents from the relevance assessments of the Open Web which exist in the ClueWeb12 collection but do not exist in the ClueWeb12 relevance assessments. We use these documents to expand the ClueWeb12 relevance assessments. Our main findings are twofold. First, our empirical analysis of the relevance assessments of 2 years of CS track shows that Open Web documents receive better ratings than ClueWeb12 documents, especially if we look at the documents in the overlap. Second, our approach for selecting candidate documents from ClueWeb12 collection based on information obtained from the Open Web makes an improvement step towards partially bridging the gap in effectiveness between Open Web and ClueWeb12 systems, while at the same time we achieve reproducible results on well-known representative sample of the web.
Thaer Samar, Alejandro Bellogín, Arjen P. de Vries
Inf. Retr. J.2
2016 A Framework for Dataset Benchmarking and Its Application to a New Movie Rating Dataset
abstract
Rating datasets are of paramount importance in recommender systems research. They serve as input for recommendation algorithms, as simulation data, or for evaluation purposes. In the past, public accessible rating datasets were not abundantly available, leaving researchers no choice but to work with old and static datasets like MovieLens and Netflix. More recently, however, emerging trends as social media and smartphones are found to provide rich data sources which can be turned into valuable research datasets. While dataset availability is growing, a structured way for introducing and comparing new datasets is currently still lacking. In this work, we propose a five-step framework to introduce and benchmark new datasets in the recommender systems domain. We illustrate our framework on a new movie rating dataset—called MovieTweetings—collected from Twitter. Following our framework, we detail the origin of the dataset, provide basic descriptive statistics, investigate external validity, report the results of a number of reproducible benchmarks, and conclude by discussing some interesting advantages and appropriate research use cases.
Simon Dooms, Alejandro Bellogín, Toon De Pessemier, Luc Martens
ACM Trans. Intell. Syst. Technol.2
2015 Replicable Evaluation of Recommender Systems
Alan Said, Alejandro Bellogín
RecSys2
2014 Challenges on Combining Open Web and Dataset Evaluation Results: The Case of the Contextual Suggestion Track
Alejandro Bellogín, Thaer Samar, Arjen P. de Vries, Alan Said
ECIR1
2014 Effects of Position Bias on Click-Based Recommender Evaluation
Katja Hofmann, Anne Schuth, Alejandro Bellogín, Maarten de Rijke
ECIR3
2014 REDD 2014 - international workshop on recommender systems evaluation: dimensions and design
abstract
Evaluation is a cardinal issue in recommender systems; as in any technical discipline, it highlights to a large extent the problems that need to be solved by the field and, hence, leads the way for algorithmic research and development in the community. Yet, in the field of recommender systems, there still exists considerable disparity in evaluation methods, metrics and experimental designs, as well as a significant mismatch between evaluation methods in the lab and what constitutes an effective recommendation for real users and businesses. Even after the relevant quality dimensions have been defined, a clear evaluation protocol should be specified in detail and agreed upon, allowing for the comparison of results and experiments conducted by different authors. This would enable any contribution to the same problem to be incremental and add up on top of previous work, rather than grow sideways. The REDD 2014 workshop seeks to provide an informal forum to tackle such issues and to move towards better understood and shared evaluation methodologies, allowing one to leverage the efforts and the workforce of the academic community towards meaningful and relevant directions in real-world developments.
Panagiotis Adamopoulos, Alejandro Bellogín, Pablo Castells, Paolo Cremonesi, Harald Steck
RecSys2
2014 Implicit vs. explicit trust in social matrix factorization
abstract
Incorporating social trust in Matrix Factorization (MF) methods demonstrably improves accuracy of rating prediction. Such approaches mainly use the trust scores explicitly expressed by users. However, it is often challenging to have users provide explicit trust scores of each other. There exist quite a few works, which propose Trust Metrics (TM) to compute and predict trust scores between users based on their interactions. In this paper, we first evaluate several TMs to find out which one can best predict trust scores compared to the actual trust scores explicitly expressed by users. And, second, we propose to incorporate these trust scores inferred from the candidate TMs into social matrix factorization (MF). We investigate if incorporating the implicit trust scores in MF can make rating prediction as accurate as the MF on explicit trust scores. The reported results support the idea of employing implicit trust into MF whenever explicit trust is not available, since the performance of both models is similar.
Soude Fazeli, Babak Loni, Alejandro Bellogín, Hendrik Drachsler, Peter B. Sloep
RecSys3
2014 Comparative recommender system evaluation: benchmarking recommendation frameworks
abstract
Recommender systems research is often based on comparisons of predictive accuracy: the better the evaluation scores, the better the recommender. However, it is difficult to compare results from different recommender systems due to the many options in design and implementation of an evaluation strategy. Additionally, algorithmic implementations can diverge from the standard formulation due to manual tuning and modifications that work better in some situations.
Alan Said, Alejandro Bellogín
RecSys2
2014 Rival: a toolkit to foster reproducibility in recommender system evaluation
abstract
Currently, it is difficult to put in context and compare the results from a given evaluation of a recommender system, mainly because too many alternatives exist when designing and implementing an evaluation strategy. Furthermore, the actual implementation of a recommendation algorithm sometimes diverges considerably from the well-known ideal formulation due to manual tuning and modifications observed to work better in some situations. RiVal - a recommender system evaluation toolkit - allows for complete control of the different evaluation dimensions that take place in any experimental evaluation of a recommender system: data splitting, definition of evaluation strategies, and computation of evaluation metrics. In this demo we present some of the functionality of RiVal and show step-by-step how RiVal can be used to evaluate the results from any recommendation framework and make sure that the results are comparable and reproducible.
Alan Said, Alejandro Bellogín
RecSys2
2014 Neighbor Selection and Weighting in User-Based Collaborative Filtering: A Performance Prediction Approach
abstract
User-based collaborative filtering systems suggest interesting items to a user relying on similar-minded people called neighbors. The selection and weighting of these neighbors characterize the different recommendation approaches. While standard strategies perform a neighbor selection based on user similarities, trust-aware recommendation algorithms rely on other aspects indicative of user trust and reliability. In this article we restate the trust-aware recommendation problem, generalizing it in terms of performance prediction techniques, whose goal is to predict the performance of an information retrieval system in response to a particular query. We investigate how to adopt the preceding generalization to define a unified framework where we conduct an objective analysis of the effectiveness (predictive power) of neighbor scoring functions. The proposed framework enables discriminating whether recommendation performance improvements are caused by the used neighbor scoring functions or by the ways these functions are used in the recommendation computation. We evaluated our approach with several state-of-the-art and novel neighbor scoring functions on three publicly available datasets. By empirically comparing four neighbor quality metrics and thirteen performance predictors, we found strong predictive power for some of the predictors with respect to certain metrics. This result was then validated by checking the final performance of recommendation strategies where predictors are used for selecting and/or weighting user neighbors. As a result, we have found that, by measuring the predictive power of neighbor performance predictors, we are able to anticipate which predictors are going to perform better in neighbor-scoring-powered versions of a user-based collaborative filtering algorithm.
Alejandro Bellogín, Pablo Castells, Iván Cantador
ACM Trans. Web1
2013 Document Difficulty Framework for Semi-automatic Text Classification
Miguel Martinez-Alvarez, Alejandro Bellogín, Thomas Roelleke
DaWaK2
2013 Artist Popularity: Do Web and Social Music Services Agree?
Alejandro Bellogín, Arjen P. de Vries, Jiyin He
ICWSM1
2013 Workshop on reproducibility and replication in recommender systems evaluation: RepSys
abstract
Experiment replication and reproduction are key requirements for empirical research methodology, and an important open issue in the field of Recommender Systems. When an experiment is repeated by a different researcher and exactly the same result is obtained, we can say the experiment has been replicated. When the results are not exactly the same but the conclusions are compatible with the prior ones, we have a reproduction of the experiment. Reproducibility and replication involve recommendation algorithm implementations, experimental protocols, and evaluation metrics. While the problem of reproducibility and replication has been recognized in the Recommender Systems community, the need for a clear solution remains largely unmet, which motivates the present workshop.
Alejandro Bellogín, Pablo Castells, Alan Said, Domonkos Tikk
RecSys1
2013 Probabilistic collaborative filtering with negative cross entropy
abstract
Relevance-Based Language Models are an effective IR approach which explicitly introduces the concept of relevance in the statistical Language Modelling framework of Information Retrieval. These models have shown to achieve state-of-the-art retrieval performance in the pseudo relevance feedback task. In this paper we propose a novel adaptation of this language modeling approach to rating-based Collaborative Filtering. In a memory-based approach, we apply the model to the formation of user neighbourhoods, and the generation of recommendations based on such neighbourhoods. We report experimental results where our method outperforms other standard memory-based algorithms in terms of ranking precision.
Alejandro Bellogín, Javier Parapar, Pablo Castells
RecSys1
2013 Relevance-based language modelling for recommender systems
Javier Parapar, Alejandro Bellogín, Pablo Castells, Álvaro Barreiro
Inf. Process. Manag.2
2013 Bridging memory-based collaborative filtering and text retrieval
Alejandro Bellogín, Jun Wang 0012, Pablo Castells
Inf. Retr.1
2013 A comparative study of heterogeneous item recommendations in social systems
Alejandro Bellogín, Iván Cantador, Pablo Castells
Inf. Sci.1
2013 An empirical comparison of social, collaborative filtering, and hybrid recommenders
abstract
In the Social Web, a number of diverse recommendation approaches have been proposed to exploit the user generated contents available in the Web, such as rating, tagging, and social networking information. In general, these approaches naturally require the availability of a wide amount of these user preferences. This may represent an important limitation for real applications, and may be somewhat unnoticed in studies focusing on overall precision, in which a failure to produce recommendations gets blurred when averaging the obtained results or, even worse, is just not accounted for, as users with no recommendations are typically excluded from the performance calculations. In this article, we propose a coverage metric that uncovers and compensates for the incompleteness of performance evaluations based only on precision. We use this metric together with precision metrics in an empirical comparison of several social, collaborative filtering, and hybrid recommenders. The obtained results show that a better balance between precision and coverage can be achieved by combining social-based filtering (high accuracy, low coverage) and collaborative filtering (low accuracy, high coverage) recommendation techniques. We thus explore several hybrid recommendation approaches to balance this trade-off. In particular, we compare, on the one hand, techniques integrating collaborative and social information into a single model, and on the other, linear combinations of recommenders. For the last approach, we also propose a novel strategy to dynamically adjust the weight of each recommender on a user-basis, utilizing graph measures as indicators of the target user's connectedness and relevance in a social network.
Alejandro Bellogín, Iván Cantador, Fernando Díez, Pablo Castells, Enrique Chavarriaga
ACM Trans. Intell. Syst. Technol.1
2012 Time feature selection for identifying active household members
abstract
Popular online rental services such as Netflix and MoviePilot often manage household accounts. A household account is usually shared by various users who live in the same house, but in general does not provide a mechanism by which current active users are identified, and thus leads to considerable difficulties for making effective personalized recommendations. The identification of the active household members, defined as the discrimination of the users from a given household who are interacting with a system (e.g. an on-demand video service), is thus an interesting challenge for the recommender systems research community. In this paper, we formulate the above task as a classification problem, and address it by means of global and local feature selection methods and classifiers that only exploit time features from past item consumption records. The results obtained from a series of experiments on a real dataset show that some of the proposed methods are able to select relevant time features, which allow simple classifiers to accurately identify active members of household accounts.
Pedro G. Campos, Alejandro Bellogín, Fernando Díez, Iván Cantador
CIKM2
2012 Using graph partitioning techniques for neighbour selection in user-based collaborative filtering
abstract
Spectral clustering techniques have become one of the most popular clustering algorithms, mainly because of their simplicity and effectiveness. In this work, we make use of one of these techniques, Normalised Cut, in order to derive a cluster-based collaborative filtering algorithm which outperforms other standard techniques in the state-of-the-art in terms of ranking precision. We frame this technique as a method for neighbour selection, and we show its effectiveness when compared with other cluster-based methods. Furthermore, the performance of our method could be improved if standard similarity metrics -- such as Pearson's correlation -- are also used when predicting the user's preferences.
Alejandro Bellogín, Javier Parapar
RecSys1
2011 Structured collaborative filtering
abstract
In a general collaborative filtering (CF) setting, a user profile contains a set of previously rated items and is used to represent the user's interest. Unfortunately, most CF approaches ignore the underlying structure of user profiles. In this paper, we argue that a certain class of interest is best represented jointly by several items, drawing an analogy to "phrases" in text retrieval, which are not equivalent to the separate meaning of their words. At an alternative stance, we also consider the situation where, analogously to word synonyms, two items might be substitutable when representing a class of interest. We propose an approach integrating these two notions as opposing poles on a continuum spectrum. Upon this, we model the underlying structure in user profiles, drawing an analogy with text retrieval. The approach gives rise to a novel structured Vector Space Model for CF. We show that item-based CF approaches are a special case of the proposed method.
Alejandro Bellogín, Jun Wang 0012, Pablo Castells
CIKM1
2011 Text Retrieval Methods for Item Ranking in Collaborative Filtering
Alejandro Bellogín, Jun Wang 0012, Pablo Castells
ECIR1
2011 Predicting performance in recommender systems
abstract
Performance prediction has gained growing attention in the Information Retrieval field since the late nineties and has become an established research topic in the field. Our work restates the problem in the area of Recommender Systems, where it has barely been researched so far, despite being an appealing problem, as it enables an array of strategies for deciding when to deliver or hold back recommendations based on their foreseen accuracy. We investigate the adaptation and definition of different performance predictors based on the available user and item features. The properties of the predictor are empirically studied by checking the correlation of the predictor output with a performance measure. Then, we propose to introduce the performance predictor in a recommender system to produce a dynamic strategy. Depending on how the predictor is introduced we analyze two different problems: dynamic neighbor weighting in collaborative filtering and dynamic weighting of ensemble recommenders.
Alejandro Bellogín
RecSys1
2011 Precision-oriented evaluation of recommender systems: an algorithmic comparison
abstract
There is considerable methodological divergence in the way precision-oriented metrics are being applied in the Recommender Systems field, and as a consequence, the results reported in different studies are difficult to put in context and compare. We aim to identify the involved methodological design alternatives, and their effect on the resulting measurements, with a view to assessing their suitability, advantages, and potential shortcomings. We compare five experimental methodologies, broadly covering the variants reported in the literature. In our experiments with three state-of-the-art recommenders, four of the evaluation methodologies are consistent with each other and differ from error metrics, in terms of the comparative recommenders' performance measurements. The other procedure aligns with RMSE, but shows a heavy bias towards known relevant items, considerably overestimating performance.
Alejandro Bellogín, Pablo Castells, Iván Cantador
RecSys1
2011 Self-adjusting hybrid recommenders based on social network analysis
abstract
Ensemble recommender systems successfully enhance recom-mendation accuracy by exploiting different sources of user prefe-rences, such as ratings and social contacts. In linear ensembles, the optimal weight of each recommender strategy is commonly tuned empirically, with limited guarantee that such weights are optimal afterwards. We propose a self-adjusting hybrid recommendation approach that alleviates the social cold start situation by weighting the recommender combination dynamically at recommendation time, based on social network analysis algorithms. We show empirical results where our approach outperforms the best static combination for different hybrid recommenders.
Alejandro Bellogín, Pablo Castells, Iván Cantador
SIGIR1
2011 An Enhanced Semantic Layer for Hybrid Recommender Systems: Application to News Recommendation
abstract
Recommender systems have achieved success in a variety of domains, as a means to help users in information overload scenarios by proactively finding items or services on their behalf, taking into account or predicting their tastes, priorities, or goals. Challenging issues in their research agenda include the sparsity of user preference data and the lack of flexibility to incorporate contextual factors in the recommendation methods. To a significant extent, these issues can be related to a limited description and exploitation of the semantics underlying both user and item representations. The authors propose a three-fold knowledge representation, in which an explicit, semantic-rich domain knowledge space is incorporated between user and item spaces. The enhanced semantics support the development of contextualisation capabilities and enable performance improvements in recommendation methods. As a proof of concept and evaluation testbed, the approach is evaluated through its implementation in a news recommender system, in which it is tested with real users. In such scenario, semantic knowledge bases and item annotations are automatically produced from public sources.
Iván Cantador, Pablo Castells, Alejandro Bellogín
Int. J. Semantic Web Inf. Syst.3
2010 A Performance Prediction Approach to Enhance Collaborative Filtering Performance
Alejandro Bellogín, Pablo Castells
ECIR1
2010 Content-based recommendation in social tagging systems
abstract
We present and evaluate various content-based recommendation models that make use of user and item profiles defined in terms of weighted lists of social tags. The studied approaches are adaptations of the Vector Space and Okapi BM25 information retrieval models. We empirically compare the recommenders using two datasets obtained from Delicious and Last.fm social systems, in order to analyse the performance of the approaches in scenarios with different domains and tagging behaviours.
Iván Cantador, Alejandro Bellogín, David Vallet
RecSys2
2009 Predicting Neighbor Goodness in Collaborative Filtering
Alejandro Bellogín, Pablo Castells
FQAS1
2008 Ontology-Based Personalised and Context-Aware Recommendations of News Items
abstract
News@hand is a news recommender system that makes use of semantic technologies to provide several on-line news recommendation services. News contents and user preferences are described in terms of concepts appearing in a set of domain ontologies. Based on the similarities between item descriptions and user profiles, and the se-mantic relations between concepts, content-based and collaborative recommendation models are supported by the system. In this paper, we evaluate a model that personalises the order in which news articles are shown to the user according to his long-term interest profile, and other model that reorders the news items lists taking into account the current semantic context of interest of the user. The combination of those models is investigated showing significant improvements on the experimental tasks performed.
Iván Cantador, Alejandro Bellogín, Pablo Castells
Web Intelligence2