VLDB 2026 Research / reviewers in the wild / expert
Pablo Castells
dblp:c/PabloCastells
· DBLP profile ↗
72ranked-venue papers in the field
4as first author
16since 2021 · last 2025
0000-0003-0668-6317ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 60 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 5Database Systems & Data Management · 3 (1 first)Data Mining & Knowledge Discovery · 3 (1 first)Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Toward Holistic Evaluation of Recommender Systems Powered by Generative ModelsabstractRecommender systems powered by generative models (Gen-RecSys) extend beyond classical item-ranking by producing open-ended content, which simultaneously unlocks richer user experiences and introduces new risks. On one hand, these systems can enhance personalization and appeal through dynamic explanations and multi-turn dialogues. On the other hand, they might venture into unknown territory-hallucinating nonexistent items, amplifying bias, or leaking private information. Traditional accuracy metrics cannot fully capture these challenges, as they fail to measure factual correctness, content safety, or alignment with user intent. Yashar Deldjoo, Nikhil Mehta 0002, Maheswaran Sathiamoorthy, Shuai Zhang 0007, Pablo Castells, Julian J. McAuley |
SIGIR | 5 |
| 2025 | GENNEXT: The Next Generation of IR and Recommender Systems with Language Agents, Generative Models, and Conversational AIabstractWe present GENNEXT, a workshop dedicated to exploring the integration of language agents, generative models, and conversational AI within information retrieval (IR) and recommender systems (RS). Building on the success of our recent RecSys'24 workshop, GENNEXT aims to advance discussions on the applications of language agents powered by Large Language Models (LLMs). The workshop will focus on enhancing interactivity between users and systems through multi-turn dialogues, improving creative content generation, advancing personalization, and enabling multifaceted, context-aware decision-making. For example, a language agent could respond to a query like ''Suggest an eco-friendly food tour for a weekend in my city'' by using a recommendation API to identify eateries specializing in sustainable or organic cuisine and a pollution API to ensure the selected routes have low air pollution levels. Yashar Deldjoo, Scott Sanner, Enrico Palumbo, Hugues Bouchard, Shuai Zhang 0007, Pablo Castells, Julian J. McAuley |
SIGIR | 6 |
| 2025 | Impression-Aware Recommender SystemsabstractNovel data sources bring new opportunities to improve the quality of recommender systems and serve as a catalyst for the creation of new paradigms on personalized recommendations. Impressions are a novel data source containing the items shown to users on their screens. Past research focused on providing personalized recommendations using interactions and occasionally using impressions when such a data source was available. Interest in impressions has increased due to their potential to provide more accurate recommendations. Despite this increased interest, research in recommender systems using impressions is still dispersed. Many works have distinct interpretations of impressions and use impressions in recommender systems in numerous different manners. To unify those interpretations into a single framework, we present a systematic literature review on recommender systems using impressions, focusing on three fundamental perspectives: recommendation models , datasets , and evaluation methodologies . We define a theoretical framework to delimit recommender systems using impressions and a novel paradigm for personalized recommendations, called impression-aware recommender systems. We propose a classification system for recommenders in this paradigm, which we use to categorize the recommendation models, datasets, and evaluation methodologies used in past research. Last, we identify open questions and future directions, highlighting missing aspects in the reviewed literature. Fernando Benjamín Pérez Maurera, Maurizio Ferrari Dacrema, Pablo Castells, Paolo Cremonesi |
Trans. Recomm. Syst. | 3 |
| 2025 | Introduction to the Special Issue on Trustworthy Recommender SystemsabstractThis editorial introduces the Special Issue on Trustworthy Recommender Systems , hosted by the ACM Transactions on Recommender Systems in 2024. We provide an overview on the multifaceted aspects of trustworthiness and point to recent regulations that underline the importance of the topic, also beyond technical perspectives. Subsequently, we present the nine articles constituting the special issue: one survey that reviews over 400 papers, categorizing them according to five trustworthiness dimensions, and eight research articles . We categorize and introduce the latter according to the major trustworthiness dimensions they address, specifically into privacy/security , transparency/explainability , and bias/fairness . We provide a summary of their main contributions and end with a brief personal statement about envisioned challenges ahead. Markus Schedl, Yashar Deldjoo, Pablo Castells, Emine Yilmaz |
Trans. Recomm. Syst. | 3 |
| 2024 | The 1st International Workshop on Risks, Opportunities, and Evaluation of Generative Models in Recommendation (ROEGEN)abstractWe present an overview of a workshop focused on the exploration of generative models within recommender systems (RS). It highlights the dual nature of these technologies: on the one hand, they offer groundbreaking opportunities for enhancing RS through improved personalization, innovative content creation, and interactive user experiences; on the other hand, they introduce a range of challenges, including bias, misinformation, privacy concerns, and environmental impact. Yashar Deldjoo, Julian J. McAuley, Scott Sanner, Pablo Castells, Shuai Zhang 0007, Enrico Palumbo |
RecSys | 4 |
| 2024 | Temporal Conformity-aware Hawkes Graph Network for RecommendationsabstractMany existing recommender systems (RSs) assume user behavior is governed solely by their interests. However, the peer effect often influences individual decision-making, which leads to conformity behavior. Conventional solutions that eliminate indiscriminately such bias may cause RSs to neglect valuable information and depersonalize the recommendation results. Also, conformity can transform into user interest, e.g., discovering new tastes after a glance at popular music. By better representing different forms of conformity influence, we can do a better job at interest mining and debiasing. In certain extreme circumstances, the herd effect may be exacerbated by user anxiety with uncertainty (e.g., panic buying during the COVID-19 pandemic). RSs may thus fail to respond in time due to sudden and dramatic changes. Moreover, many existing studies potentially conflate conformity bias with popularity bias and lump together various factors responsible for differences in popularity. In this paper, we identify two distinct types of conformity behavior: informational conformity and normative conformity. To address this, we introduce the TCHN model, which utilizes attentional Hawkes processes to disentangle user self-interest and conformity in a personalized manner. Our approach incorporates temporal graph attention networks to capture users' stable and volatile dynamics. We conduct experiments on three real-world datasets, which uncover diverse levels of conformity among users. The results show that TCHN excels in recommendation accuracy, diversity, and fairness across various user groups. Chenglong Ma 0001, Yongli Ren, Pablo Castells, Mark Sanderson |
WWW | 3 |
| 2023 | Workshop on Learning and Evaluating Recommendations with Impressions (LERI)abstractRecommender systems typically rely on past user interactions as the primary source of information for making predictions. However, although highly informative, past user interactions are strongly biased. Impressions, on the other hand, are a new source of information that indicate the items displayed on screen when the user interacted (or not) with them, and have the potential to impact the field of recommender systems in several ways. Early research on impressions was constrained by the limited availability of public datasets, but this is rapidly changing and, as a consequence, interest in impressions has increased. Impressions present new research questions and opportunities, but also bring new challenges. Several works propose to use impressions as part of recommender models in various ways and discuss their information content. Others explore their potential in off-policy-estimation and reinforcement learning. Overall, the interest of the community is growing, but efforts in this direction remain disconnected. Therefore, we believe that a workshop would be useful in bringing the community together. Maurizio Ferrari Dacrema, Pablo Castells, Justin Basilico, Paolo Cremonesi |
RecSys | 2 |
| 2022 | Debiased Balanced Interleaving at Amazon SearchabstractInterleaving is an online evaluation technique that has shown to be orders of magnitude more sensitive than traditional A/B tests. It presents users with a single merged result of the compared rankings and then attributes user actions back to the evaluated rankers. Different interleaving methods in the literature have their advantages and limitations with respect to unbiasedness, sensitivity, preservation of user experience, and implementation and computation complexity. We propose a new interleaving method that utilizes a counterfactual evaluation framework for credit attribution while sticking to the simple ranking merge policy of balanced interleaving, and formally derive an unbiased estimator for comparing rankers with theoretical guarantees. We then confirm the effectiveness of our method with both synthetic and real experiments. We also discuss practical considerations of bringing different interleaving methods from the literature into a large-scale experiment, and show that our method achieves a favorable tradeoff in implementation and computation complexity while preserving statistical power and reliability. We have successfully implemented our method and produced consistent conclusions at the scale of billions of search queries. We report 10 online experiments that apply our method to e-commerce search, and observe a 60x sensitivity gain over A/B tests. We also find high correlations between our proposed estimator and corresponding A/B metrics, which helps interpret interleaving results in the magnitude of A/B measurements. Nan Bi, Pablo Castells, Daniel Gilbert, Slava Galperin, Patrick Tardif, Sachin Ahuja |
CIKM | 2 |
| 2022 | Addressing Cold Start in Product Search via Empirical BayesabstractCold start is a challenge in product search. Profuse literature addresses related problems such as bias and diversity in search, and cold start is a classic topic in recommender systems research. While search cold start might be seen conceptually as a particular case in such areas, we find that available solutions fail to specifically and practically solve the cold-start problem in product search. The problem is complex as exposing new products may come at the expense of primary business metrics (e.g. revenue), and involves a complex balance between customer satisfaction, seller satisfaction, business performance, short-term gains and long-term value. Cuize Han, Pablo Castells, Parth Gupta, Vamsi Salaka |
CIKM | 2 |
| 2022 | NEST: Simulating Pandemic-like Events for Collaborative Filtering by Modeling User Needs EvolutionabstractWe outline a simulation-based study of the effect rapid population-scale concept drifts have on Collaborative Filtering (CF) models. We create a framework for analyzing the effects of macro-trends in population dynamics on the behavior of such models. Our framework characterizes population-scale concept drifts in item preferences and provides a lens to understand the influence events, such as a pandemic, have on CF models. Our experimental results show the initial impact on CF performance at the initial stage of such events, followed by an aggravated population herding effect during the event. The herding introduces a popularity bias that may benefit affected users, but which comes at the expense of a normal user experience. We propose an adaptive ensemble method that can effectively apply optimal algorithms to cope with the change brought about by different stages of the event. Chenglong Ma 0001, Yongli Ren, Pablo Castells, Mark Sanderson |
CIKM | 3 |
| 2022 | Evaluation of Herd Behavior Caused by Population-scale Concept Drift in Collaborative FilteringabstractConcept drift in stream data has been well studied in machine learning applications. In the field of recommender systems, this issue is also widely observed, as known as temporal dynamics in user behavior. Furthermore, in the context of COVID-19 pandemic related contingencies, people shift their behavior patterns extremely and tend to imitate others' opinions. The changes in user behavior may not be always rational. Thus, irrational behavior may impair the knowledge learned by the algorithm. It can cause herd effects and aggravate the popularity bias in recommender systems due to the irrational behavior of users. However, related research usually pays attention to the concept drift of individuals and overlooks the synergistic effect among users in the same social group. We conduct a study on user behavior to detect the collaborative concept drifts among users. Also, we empirically study the increase of experience of individuals can weaken herding effects. Our results suggest the CF models are highly impacted by the herd behavior and our findings could provide useful implications for the design of future recommender algorithms. Chenglong Ma 0001, Yongli Ren, Pablo Castells, Mark Sanderson |
SIGIR | 3 |
| 2022 | RELISON: A Framework for Link Recommendation in Social NetworksabstractLink recommendation is an important and compelling problem at the intersection of recommender systems and online social networks. Given a user, link recommenders identify people in the platform the user might be interested in interacting with. We present RELISON, an extensible framework for running link recommendation experiments. The library provides a wide range of algorithms, along with tools for evaluating the produced recommendations. RELISON includes algorithms and metrics that consider the potential effect of recommendations on the properties of online social networks. For this reason, the library also implements network structure analysis metrics, community detection algorithms, and network diffusion simulation functionalities. The library code and documentation is available at https://github.com/ir-uam/RELISON. Javier Sanz-Cruzado, Pablo Castells |
SIGIR | 2 |
| 2022 | Human Preferences as Dueling BanditsabstractThe dramatic improvements in core information retrieval tasks engendered by neural rankers create a need for novel evaluation methods. If every ranker returns highly relevant items in the top ranks, it becomes difficult to recognize meaningful differences between them and to build reusable test collections. Several recent papers explore pairwise preference judgments as an alternative to traditional graded relevance assessments. Rather than viewing items one at a time, assessors view items side-by-side and indicate the one that provides the better response to a query, allowing fine-grained distinctions. If we employ preference judgments to identify the probably best items for each query, we can measure rankers by their ability to place these items as high as possible. We frame the problem of finding best items as a dueling bandits problem. While many papers explore dueling bandits for online ranker evaluation via interleaving, they have not been considered as a framework for offline evaluation via human preference judgments. We review the literature for possible solutions. For human preference judgments, any usable algorithm must tolerate ties, since two items may appear nearly equal to assessors, and it must minimize the number of judgments required for any specific pair, since each such comparison requires an independent assessor. Since the theoretical guarantees provided by most algorithms depend on assumptions that are not satisfied by human preference judgments, we simulate selected algorithms on representative test cases to provide insight into their practical utility. Based on these simulations, one algorithm stands out for its potential. Our simulations suggest modifications to further improve its performance. Using the modified algorithm, we collect over 10,000 preference judgments for pools derived from submissions to the TREC 2021 Deep Learning Track, confirming its suitability. We test the idea of best-item evaluation and suggest ideas for further theoretical and practical progress. Xinyi Yan, Chengxi Luo, Charles L. A. Clarke, Nick Craswell, Ellen M. Voorhees, Pablo Castells |
SIGIR | 6 |
| 2021 | SimuRec: Workshop on Synthetic Data and Simulation Methods for Recommender Systems ResearchabstractThere is significant interest lately in using synthetic data and simulation infrastructures for various types of recommender systems research. However, there are not currently any clear best practices around how best to apply these methods. We proposed a workshop to bring together researchers and practitioners interested in simulating recommender systems and their data to discuss the state of the art of such research and the pressing open methodological questions. The workshop resulted in a report authored by the participants that documents currently-known best practices on which the group has consensus and lays out an agenda for further research over the next 3–5 years to fill in places where we currently lack the information needed to make methodological recommendations. Michael D. Ekstrand, Allison Chaney, Pablo Castells, Robin D. Burke, David Rohde, Manel Slokom |
RecSys | 3 |
| 2021 | Guest editorial: special issue on ECIR 2020
Joemon M. Jose, Emine Yilmaz, João Magalhães, Pablo Castells |
Inf. Retr. J. | 4 |
| 2021 | Popularity Bias in False-positive Metrics for Recommender Systems EvaluationabstractWe investigate the impact of popularity bias in false-positive metrics in the offline evaluation of recommender systems. Unlike their true-positive complements, false-positive metrics reward systems that minimize recommendations disliked by users. Our analysis is, to the best of our knowledge, the first to show that false-positive metrics tend to penalise popular items, the opposite behavior of true-positive metrics—causing a disagreement trend between both types of metrics in the presence of popularity biases. We present a theoretical analysis of the metrics that identifies the reason that the metrics disagree and determines rare situations where the metrics might agree—the key to the situation lies in the relationship between popularity and relevance distributions, in terms of their agreement and steepness —two fundamental concepts we formalize. We then examine three well-known datasets using multiple popular true- and false-positive metrics on 16 recommendation algorithms. Specific datasets are chosen to allow us to estimate both biased and unbiased metric values. The results of the empirical study confirm and illustrate our analytical findings. With the conditions of the disagreement of the two types of metrics established, we then determine under which circumstances true-positive or false-positive metrics should be used by researchers of offline evaluation in recommender systems. 1 Elisa Mena-Maldonado, Rocío Cañamares, Pablo Castells, Yongli Ren, Mark Sanderson |
ACM Trans. Inf. Syst. | 3 |
| 2020 | Axiomatic Analysis of Contact Recommendation Methods in Social Networks: An IR Perspective
Javier Sanz-Cruzado, Craig Macdonald, Iadh Ounis, Pablo Castells |
ECIR (1) | 4 |
| 2020 | On Target Item Sampling in Offline Recommender System EvaluationabstractTarget selection is a basic yet often implicit decision in the configuration of offline recommendation experiments. In this paper we research the impact of target sampling on the outcome of comparative recommender system evaluation. Specifically, we undertake a detailed analysis considering the informativeness and consistency of experiments across the target size axis. We find that comparative evaluation using reduced target sets contradicts in many cases the corresponding outcome using large targets, and we provide a principled explanation for these disagreements. We further seek to determine which among the contradicting results may be more reliable. Through comparison to unbiased evaluation, we find that minimum target sets incur in substantial distortion in pairwise system comparisons, while maximum sets may not be ideal either, and better options may lie in between the extremes. We further find means for informing the target size setting in the common case where unbiased evaluation is not possible, by an assessment of the discriminative power of evaluation, that remarkably aligns with the agreement with unbiased evaluation. Rocío Cañamares, Pablo Castells |
RecSys | 2 |
| 2020 | Agreement and Disagreement between True and False-Positive Metrics in Recommender Systems EvaluationabstractFalse-positive metrics can capture an important side of recommendation quality, focusing on the impact of suggestions that are disliked by users, as a complement of common metrics that only measure the amount of successful recommendations. In this paper we research the extent to which false-positive metrics agree or disagree with true-positive metrics in the offline evaluation of recommender systems. We discover a surprising degree of systematic disagreement that was occasionally noted but not explained in the literature by previous authors. We find an explanation for the discrepancy be-tween the metrics in the effect of popularity biases, which impact false and true-positive metrics in very different ways: instead of rewarding the recommendation of popular items, as with true-positive, false-positive metrics penalize the popular. We determine precise conditions and cases in the general trends, with a formal explanation for our findings, which we confirm and illustrate empirically in experiments with different datasets. Elisa Mena-Maldonado, Rocío Cañamares, Pablo Castells, Yongli Ren, Mark Sanderson |
SIGIR | 3 |
| 2020 | Effective contact recommendation in social networks by adaptation of information retrieval models
Javier Sanz-Cruzado, Pablo Castells, Craig Macdonald, Iadh Ounis |
Inf. Process. Manag. | 2 |
| 2020 | Offline evaluation options for recommender systems
Rocío Cañamares, Pablo Castells, Alistair Moffat |
Inf. Retr. J. | 2 |
| 2020 | Assessing ranking metrics in top-N recommendation
Daniel Valcarce, Alejandro Bellogín, Javier Parapar, Pablo Castells |
Inf. Retr. J. | 4 |
| 2019 | Information Retrieval Models for Contact Recommendation in Social Networks
Javier Sanz-Cruzado, Pablo Castells |
ECIR (1) | 2 |
| 2019 | Multi-armed recommender system bandit ensemblesabstractIt has long been found that well-configured recommender system ensembles can achieve better effectiveness than the combined systems separately. Sophisticated approaches have been developed to automatically optimize the ensembles' configuration to maximize their performance gains. However most work in this area has targeted simplified scenarios where algorithms are tested and compared on a single non-interactive run. In this paper we consider a more realistic perspective bearing in mind the cyclic nature of the recommendation task, where a large part of the system's input is collected from the reaction of users to the recommendations they are delivered. The cyclic process provides the opportunity for ensembles to observe and learn about the effectiveness of the combined algorithms, and improve the ensemble configuration progressively. Rocío Cañamares, Marcos Redondo, Pablo Castells |
RecSys | 3 |
| 2019 | A simple multi-armed nearest-neighbor bandit for interactive recommendationabstractThe cyclic nature of the recommendation task is being increasingly taken into account in recommender systems research. In this line, framing interactive recommendation as a genuine reinforcement learning problem, multi-armed bandit approaches have been increasingly considered as a means to cope with the dual exploitation/exploration goal of recommendation. In this paper we develop a simple multi-armed bandit elaboration of neighbor-based collaborative filtering. The approach can be seen as a variant of the nearest-neighbors scheme, but endowed with a controlled stochastic exploration capability of the users' neighborhood, by a parameter-free application of Thompson sampling. Our approach is based on a formal development and a reasonably simple design, whereby it aims to be easy to reproduce and further elaborate upon. We report experiments using datasets from different domains showing that neighbor-based bandits indeed achieve recommendation accuracy enhancements in the mid to long run. Javier Sanz-Cruzado, Pablo Castells, Esther López |
RecSys | 2 |
| 2018 | Enhancing structural diversity in social networks by recommending weak tiesabstractContact recommendation has become a common functionality in online social platforms, and an established research topic in the social networks and recommender systems fields. Predicting and recommending links has been mainly addressed to date as an accuracy-targeting problem. In this paper we put forward a different perspective, considering that correctly predicted links may not be all equally valuable. Contact recommendation brings an opportunity to drive the structural evolution of a social network towards desirable properties of the network as a whole, beyond the sum of the isolated gains for the individual users to whom recommendations are delivered -global properties that we may want to assess and promote as explicit recommendation targets. Javier Sanz-Cruzado, Pablo Castells |
RecSys | 2 |
| 2018 | On the robustness and discriminative power of information retrieval metrics for top-N recommendationabstractThe evaluation of Recommender Systems is still an open issue in the field. Despite its limitations, offline evaluation usually constitutes the first step in assessing recommendation methods due to its reduced costs and high reproducibility. Selecting the appropriate metric is a critical and ranking accuracy usually attracts the most attention nowadays. In this paper, we aim to shed light on the advantages of different ranking metrics which were previously used in Information Retrieval and are now used for assessing top-N recommenders. We propose methodologies for comparing the robustness and the discriminative power of different metrics. On the one hand, we study cut-offs and we find that deeper cut-offs offer greater robustness and discriminative power. On the other hand, we find that precision offers high robustness and Normalised Discounted Cumulative Gain provides the best discriminative power. Daniel Valcarce, Alejandro Bellogín, Javier Parapar, Pablo Castells |
RecSys | 4 |
| 2018 | Should I Follow the Crowd?: A Probabilistic Analysis of the Effectiveness of Popularity in Recommender SystemsabstractThe use of IR methodology in the evaluation of recommender systems has become common practice in recent years. IR metrics have been found however to be strongly biased towards rewarding algorithms that recommend popular items "the same bias that state of the art recommendation algorithms display. Recent research has confirmed and measured such biases, and proposed methods to avoid them. The fundamental question remains open though whether popularity is really a bias we should avoid or not; whether it could be a useful and reliable signal in recommendation, or it may be unfairly rewarded by the experimental biases. We address this question at a formal level by identifying and modeling the conditions that can determine the answer, in terms of dependencies between key random variables, involving item rating, discovery and relevance. We find conditions that guarantee popularity to be effective or quite the opposite, and for the measured metric values to reflect a true effectiveness, or qualitatively deviate from it. We exemplify and confirm the theoretical findings with empirical results. We build a crowdsourced dataset devoid of the usual biases displayed by common publicly available data, in which we illustrate contradictions between the accuracy that would be measured in a common biased offline experimental setting, and the actual accuracy that can be measured with unbiased observations. Rocío Cañamares, Pablo Castells |
SIGIR | 2 |
| 2018 | From the PRP to the Low Prior Discovery Recall Principle for Recommender SystemsabstractWe revisit the Probability Ranking Principle in the context of recommender systems. We find a key difference in the retrieval protocol with respect to query-based search, that leads to the identification of different optimal ranking principles for discovery-oriented recommendation. Based on this finding, we revise the effectiveness of common non-personalized ranking functions in respect to the new principles. We run an experiment confirming and illustrating our theoretical analysis, and providing further observations and hints for reflection and future research. Rocío Cañamares, Pablo Castells |
SIGIR | 2 |
| 2017 | A Probabilistic Reformulation of Memory-Based Collaborative Filtering: Implications on Popularity BiasesabstractWe develop a probabilistic formulation giving rise to a formal version of heuristic k nearest-neighbor (kNN) collaborative filtering. Different independence assumptions in our scheme lead to user-based, item-based, normalized and non-normalized variants that match in structure the traditional formulations, while showing equivalent empirical effectiveness. The probabilistic formulation provides a principled explanation why kNN is an effective recommendation strategy, and identifies a key condition for this to be the case. Moreover, a natural explanation arises for the bias of kNN towards recommending popular items. Thereupon the kNN variants are shown to fall into two groups with similar trends in behavior, corresponding to two different notions of item popularity. We show experiments where the comparative performance of the two groups of algorithms changes substantially, which suggests that the performance measurements and comparison may heavily depend on statistical properties of the input data sample. Rocío Cañamares, Pablo Castells |
SIGIR | 2 |
| 2017 | Statistical biases in Information Retrieval metrics for recommender systems
Alejandro Bellogín, Pablo Castells, Iván Cantador |
Inf. Retr. J. | 2 |
| 2014 | REDD 2014 - international workshop on recommender systems evaluation: dimensions and designabstractEvaluation is a cardinal issue in recommender systems; as in any technical discipline, it highlights to a large extent the problems that need to be solved by the field and, hence, leads the way for algorithmic research and development in the community. Yet, in the field of recommender systems, there still exists considerable disparity in evaluation methods, metrics and experimental designs, as well as a significant mismatch between evaluation methods in the lab and what constitutes an effective recommendation for real users and businesses. Even after the relevant quality dimensions have been defined, a clear evaluation protocol should be specified in detail and agreed upon, allowing for the comparison of results and experiments conducted by different authors. This would enable any contribution to the same problem to be incremental and add up on top of previous work, rather than grow sideways. The REDD 2014 workshop seeks to provide an informal forum to tackle such issues and to move towards better understood and shared evaluation methodologies, allowing one to leverage the efforts and the workforce of the academic community towards meaningful and relevant directions in real-world developments. Panagiotis Adamopoulos, Alejandro Bellogín, Pablo Castells, Paolo Cremonesi, Harald Steck |
RecSys | 3 |
| 2014 | Coverage, redundancy and size-awareness in genre diversity for recommender systemsabstractThere is increasing awareness in the Recommender Systems field that diversity is a key property that enhances the usefulness of recommendations. Genre information can serve as a means to measure and enhance the diversity of recommendations and is readily available in domains such as movies, music or books. In this work we propose a new Binomial framework for defining genre diversity in recommender systems that takes into account three key properties: genre coverage, genre redundancy and recommendation list size-awareness. We show that methods previously proposed for measuring and enhancing recommendation diversity - including those adapted from search result diversification - fail to address adequately these three properties. We also propose an efficient greedy optimization technique to optimize Binomial diversity. Experiments with the Netflix dataset show the properties of our framework and comparison with state of the art methods. Saul Vargas, Linas Baltrunas, Alexandros Karatzoglou, Pablo Castells |
RecSys | 4 |
| 2014 | Improving sales diversity by recommending users to itemsabstractSales diversity is considered a key feature of Recommender Systems from a business perspective. Sales diversity is also linked with the long-tail novelty of recommendations, a quality dimension from the user perspective. We explore the inversion of the recommendation task as a means to enhance sales diversity - and indirectly novelty - by selecting which users an item should be recommended to instead of the other way around. We address the inverted task by two approaches: a) inverting the rating matrix, and b) defining a probabilistic reformulation which isolates the popularity component of arbitrary recommendation algorithms. We find that the first approach gives rise to interesting reformulations of nearest-neighbor algorithms, which essentially introduce a new neighbor selection policy. The second approach, as well as the first, ultimately result in substantial sales diversity enhancements, and improved trade-offs with recommendation precision and novelty. Two experiments on movie and music recommendation datasets show the effectiveness of the resulting approach, even when compared to direct optimization approaches of the target metrics proposed in prior work. Saul Vargas, Pablo Castells |
RecSys | 2 |
| 2014 | Diversity and novelty in web search, recommender systems and data streamsabstractThis tutorial aims to provide a unifying account of current research on diversity and novelty in the domains of web search, recommender systems, and data stream processing. Rodrygo L. T. Santos, Pablo Castells, Ismail Sengör Altingövde, Fazli Can |
WSDM | 2 |
| 2014 | Introduction to the Special Issue on Diversity and Discovery in Recommender Systemsabstractintroduction Share on Introduction to the Special Issue on Diversity and Discovery in Recommender Systems Authors: Pablo Castells Universidad Autónoma de Madrid Universidad Autónoma de MadridView Profile , Jun Wang University College London University College LondonView Profile , Rubén Lara Telefónica Digital Telefónica DigitalView Profile , Dell Zhang Birkbeck, University of London Birkbeck, University of LondonView Profile Authors Info & Claims ACM Transactions on Intelligent Systems and TechnologyVolume 5Issue 4January 2015 Article No.: 52pp 1–3https://doi.org/10.1145/2668113Online:15 December 2014Publication History 4citation289DownloadsMetricsTotal Citations4Total Downloads289Last 12 Months4Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Pablo Castells, Jun Wang 0012, Rubén Lara, Dell Zhang |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2014 | Neighbor Selection and Weighting in User-Based Collaborative Filtering: A Performance Prediction ApproachabstractUser-based collaborative filtering systems suggest interesting items to a user relying on similar-minded people called neighbors. The selection and weighting of these neighbors characterize the different recommendation approaches. While standard strategies perform a neighbor selection based on user similarities, trust-aware recommendation algorithms rely on other aspects indicative of user trust and reliability. In this article we restate the trust-aware recommendation problem, generalizing it in terms of performance prediction techniques, whose goal is to predict the performance of an information retrieval system in response to a particular query. We investigate how to adopt the preceding generalization to define a unified framework where we conduct an objective analysis of the effectiveness (predictive power) of neighbor scoring functions. The proposed framework enables discriminating whether recommendation performance improvements are caused by the used neighbor scoring functions or by the ways these functions are used in the recommendation computation. We evaluated our approach with several state-of-the-art and novel neighbor scoring functions on three publicly available datasets. By empirically comparing four neighbor quality metrics and thirteen performance predictors, we found strong predictive power for some of the predictors with respect to certain metrics. This result was then validated by checking the final performance of recommendation strategies where predictors are used for selecting and/or weighting user neighbors. As a result, we have found that, by measuring the predictive power of neighbor performance predictors, we are able to anticipate which predictors are going to perform better in neighbor-scoring-powered versions of a user-based collaborative filtering algorithm. Alejandro Bellogín, Pablo Castells, Iván Cantador |
ACM Trans. Web | 2 |
| 2013 | Workshop on reproducibility and replication in recommender systems evaluation: RepSysabstractExperiment replication and reproduction are key requirements for empirical research methodology, and an important open issue in the field of Recommender Systems. When an experiment is repeated by a different researcher and exactly the same result is obtained, we can say the experiment has been replicated. When the results are not exactly the same but the conclusions are compatible with the prior ones, we have a reproduction of the experiment. Reproducibility and replication involve recommendation algorithm implementations, experimental protocols, and evaluation metrics. While the problem of reproducibility and replication has been recognized in the Recommender Systems community, the need for a clear solution remains largely unmet, which motivates the present workshop. Alejandro Bellogín, Pablo Castells, Alan Said, Domonkos Tikk |
RecSys | 2 |
| 2013 | Probabilistic collaborative filtering with negative cross entropyabstractRelevance-Based Language Models are an effective IR approach which explicitly introduces the concept of relevance in the statistical Language Modelling framework of Information Retrieval. These models have shown to achieve state-of-the-art retrieval performance in the pseudo relevance feedback task. In this paper we propose a novel adaptation of this language modeling approach to rating-based Collaborative Filtering. In a memory-based approach, we apply the model to the formation of user neighbourhoods, and the generation of recommendations based on such neighbourhoods. We report experimental results where our method outperforms other standard memory-based algorithms in terms of ranking precision. Alejandro Bellogín, Javier Parapar, Pablo Castells |
RecSys | 3 |
| 2013 | Workshop on benchmarking adaptive retrieval and recommender systems: BARS 2013abstractEvaluating adaptive and personalized information retrieval tech-niques is known to be a difficult endeavor. The rapid evolution of novel technologies in this scope raises additional challenges that further stress the need for new evaluation approaches and method-ologies. The BARS 2013 workshop seeks to provide a specific venue for work on novel, personalization-centric benchmarking approaches to evaluate adaptive retrieval and recommender systems. Pablo Castells, Frank Hopfgartner, Alan Said, Mounia Lalmas-Roelleke |
SIGIR | 1 |
| 2013 | Diversity and novelty in information retrievalabstractThis tutorial aims to provide a unifying account of current research on diversity and novelty in different IR domains, namely, in the context of search engines, recommender systems, and data streams. Rodrygo L. T. Santos, Pablo Castells, Ismail Sengör Altingövde, Fazli Can |
SIGIR | 2 |
| 2013 | Personalization and Recommendation in Information Access
Juan M. Fernández-Luna, Juan F. Huete, Pablo Castells |
Inf. Process. Manag. | 3 |
| 2013 | Relevance-based language modelling for recommender systems
Javier Parapar, Alejandro Bellogín, Pablo Castells, Álvaro Barreiro |
Inf. Process. Manag. | 3 |
| 2013 | Bridging memory-based collaborative filtering and text retrieval
Alejandro Bellogín, Jun Wang 0012, Pablo Castells |
Inf. Retr. | 3 |
| 2013 | A comparative study of heterogeneous item recommendations in social systems
Alejandro Bellogín, Iván Cantador, Pablo Castells |
Inf. Sci. | 3 |
| 2013 | An empirical comparison of social, collaborative filtering, and hybrid recommendersabstractIn the Social Web, a number of diverse recommendation approaches have been proposed to exploit the user generated contents available in the Web, such as rating, tagging, and social networking information. In general, these approaches naturally require the availability of a wide amount of these user preferences. This may represent an important limitation for real applications, and may be somewhat unnoticed in studies focusing on overall precision, in which a failure to produce recommendations gets blurred when averaging the obtained results or, even worse, is just not accounted for, as users with no recommendations are typically excluded from the performance calculations. In this article, we propose a coverage metric that uncovers and compensates for the incompleteness of performance evaluations based only on precision. We use this metric together with precision metrics in an empirical comparison of several social, collaborative filtering, and hybrid recommenders. The obtained results show that a better balance between precision and coverage can be achieved by combining social-based filtering (high accuracy, low coverage) and collaborative filtering (low accuracy, high coverage) recommendation techniques. We thus explore several hybrid recommendation approaches to balance this trade-off. In particular, we compare, on the one hand, techniques integrating collaborative and social information into a single model, and on the other, linear combinations of recommenders. For the last approach, we also propose a novel strategy to dynamically adjust the weight of each recommender on a user-basis, utilizing graph measures as indicators of the target user's connectedness and relevance in a social network. Alejandro Bellogín, Iván Cantador, Fernando Díez, Pablo Castells, Enrique Chavarriaga |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2012 | Workshop on recommendation utility evaluation: beyond RMSE - RUE 2012abstractMeasuring the error in rating prediction has been by far the dominant evaluation methodology in the Recommender Systems literature. Yet there seems to be a general consensus that this criterion alone is far from being enough to assess the practical effectiveness of a recommender system. Information Retrieval metrics have started to be used to evaluate item selection and ranking rather than rating prediction, but considerable divergence remains in the adoption of such metrics by different authors. On the other hand, recommendation utility includes other key dimensions and concerns beyond accuracy, such as novelty and diversity, user engagement, and business performance. While the need for further extension, formalization, clarification and standardization of evaluation methodologies is recognized in the community, this need is still unmet for a large extent. The RUE 2012 workshop sought to identify and better understand the current gaps in recommender system evaluation methodologies, help lay directions for progress in addressing them, and contribute to the consolidation and convergence of experimental methods and practice. Xavier Amatriain, Pablo Castells, Arjen P. de Vries, Christian Posse |
RecSys | 2 |
| 2012 | Personalized diversification of search resultsabstractSearch personalization and diversification are often seen as opposing alternatives to cope with query uncertainty, where, given an ambiguous query, it is either preferable to adapt the search result to a specific aspect that may interest the user (personalization) or to regard multiple aspects in order to maximize the probability that some query aspect is relevant to the user (diversification). In this work, we question this antagonistic view, and hypothesize that these two directions may in fact be effectively combined and enhance each other. We research the introduction of the user as an explicit random variable in state of the art diversification methods, thus developing a generalized framework for personalized diversification. In order to evaluate our hypothesis, we conduct an evaluation with real users using crowdsourcing services. The obtained results suggest that the combination of personalization and diversification achieves competitive performance, improving the base-line, plain personalization, and plain diversification approaches in terms of both diversity and accuracy measures. David Vallet, Pablo Castells |
SIGIR | 2 |
| 2012 | Explicit relevance models in intent-oriented information retrieval diversificationabstractThe intent-oriented search diversification methods developed in the field so far tend to build on generative views of the retrieval system to be diversified. Core algorithm components in particular redundancy assessment are expressed in terms of the probability to observe documents, rather than the probability that the documents be relevant. This has been sometimes described as a view considering the selection of a single document in the underlying task model. In this paper we propose an alternative formulation of aspect-based diversification algorithms which explicitly includes a formal relevance model. We develop means for the effective computation of the new formulation, and we test the resulting algorithm empirically. We report experiments on search and recommendation tasks showing competitive or better performance than the original diversification algorithms. The relevance-based formulation has further interesting properties, such as unifying two well-known state of the art algorithms into a single version. The relevance-based approach opens alternative possibilities for further formal connections and developments as natural extensions of the framework. We illustrate this by modeling tolerance to redundancy as an explicit configurable parameter, which can be set to better suit the characteristics of the IR task, or the evaluation metrics, as we illustrate empirically. Saul Vargas, Pablo Castells, David Vallet |
SIGIR | 2 |
| 2011 | Structured collaborative filteringabstractIn a general collaborative filtering (CF) setting, a user profile contains a set of previously rated items and is used to represent the user's interest. Unfortunately, most CF approaches ignore the underlying structure of user profiles. In this paper, we argue that a certain class of interest is best represented jointly by several items, drawing an analogy to "phrases" in text retrieval, which are not equivalent to the separate meaning of their words. At an alternative stance, we also consider the situation where, analogously to word synonyms, two items might be substitutable when representing a class of interest. We propose an approach integrating these two notions as opposing poles on a continuum spectrum. Upon this, we model the underlying structure in user profiles, drawing an analogy with text retrieval. The approach gives rise to a novel structured Vector Space Model for CF. We show that item-based CF approaches are a special case of the proposed method. Alejandro Bellogín, Jun Wang 0012, Pablo Castells |
CIKM | 3 |
| 2011 | Text Retrieval Methods for Item Ranking in Collaborative Filtering
Alejandro Bellogín, Jun Wang 0012, Pablo Castells |
ECIR | 3 |
| 2011 | Precision-oriented evaluation of recommender systems: an algorithmic comparisonabstractThere is considerable methodological divergence in the way precision-oriented metrics are being applied in the Recommender Systems field, and as a consequence, the results reported in different studies are difficult to put in context and compare. We aim to identify the involved methodological design alternatives, and their effect on the resulting measurements, with a view to assessing their suitability, advantages, and potential shortcomings. We compare five experimental methodologies, broadly covering the variants reported in the literature. In our experiments with three state-of-the-art recommenders, four of the evaluation methodologies are consistent with each other and differ from error metrics, in terms of the comparative recommenders' performance measurements. The other procedure aligns with RMSE, but shows a heavy bias towards known relevant items, considerably overestimating performance. Alejandro Bellogín, Pablo Castells, Iván Cantador |
RecSys | 2 |
| 2011 | Workshop on novelty and diversity in recommender systems - DiveRS 2011abstractNovelty and diversity have been identified as key dimensions of recommendation utility in real scenarios, and a fundamental research direction to keep making progress in the field. Yet recommendation novelty and diversity remain a largely open area for research. The DiveRS workshop gathered researchers and practitioners interested in the role of these dimensions in recommender systems. The workshop seeks to advance towards a better understanding of what novelty and diversity are, how they can improve the effectiveness of recommendation methods and the utility of their outputs. The workshop pursued the identification of open problems, relevant research directions, and opportunities for innovation in the recommendation business. Pablo Castells, Jun Wang 0012, Rubén Lara, Dell Zhang |
RecSys | 1 |
| 2011 | Rank and relevance in novelty and diversity metrics for recommender systems
Saul Vargas, Pablo Castells |
RecSys | 2 |
| 2011 | Self-adjusting hybrid recommenders based on social network analysisabstractEnsemble recommender systems successfully enhance recom-mendation accuracy by exploiting different sources of user prefe-rences, such as ratings and social contacts. In linear ensembles, the optimal weight of each recommender strategy is commonly tuned empirically, with limited guarantee that such weights are optimal afterwards. We propose a self-adjusting hybrid recommendation approach that alleviates the social cold start situation by weighting the recommender combination dynamically at recommendation time, based on social network analysis algorithms. We show empirical results where our approach outperforms the best static combination for different hybrid recommenders. Alejandro Bellogín, Pablo Castells, Iván Cantador |
SIGIR | 2 |
| 2011 | On diversifying and personalizing web searchabstractDiversification and personalization methods are common ap-proaches to deal with the one-size-fits-all paradigm of Web search engines. We performed a user study with 190 subjects where we analyzed the effects of diversification and personalization methods in a Web search engine. The obtained results suggest that our proposed combination of diversification and personalization factors may be a way to overcome the notion of intrusiveness in personalized approaches. David Vallet, Pablo Castells |
SIGIR | 2 |
| 2011 | Intent-oriented diversity in recommender systemsabstractDiversity as a relevant dimension of retrieval quality is receiving increasing attention in the Information Retrieval and Recommender Systems (RS) fields. The problem has nonetheless been approached under different views and formulations in IR and RS respectively, giving rise to different models, methodologies, and metrics, with little convergence between both fields. In this poster we explore the adaptation of diversity metrics, techniques, and principles from ad-hoc IR to the recommendation task, by introducing the notion of user profile aspect as an analogue of query intent. As a particular approach, user aspects are automatically extracted from latent item features. Empirical results support the proposed approach and provide further insights. Saul Vargas, Pablo Castells, David Vallet |
SIGIR | 2 |
| 2011 | An Enhanced Semantic Layer for Hybrid Recommender Systems: Application to News RecommendationabstractRecommender systems have achieved success in a variety of domains, as a means to help users in information overload scenarios by proactively finding items or services on their behalf, taking into account or predicting their tastes, priorities, or goals. Challenging issues in their research agenda include the sparsity of user preference data and the lack of flexibility to incorporate contextual factors in the recommendation methods. To a significant extent, these issues can be related to a limited description and exploitation of the semantics underlying both user and item representations. The authors propose a three-fold knowledge representation, in which an explicit, semantic-rich domain knowledge space is incorporated between user and item spaces. The enhanced semantics support the development of contextualisation capabilities and enable performance improvements in recommendation methods. As a proof of concept and evaluation testbed, the approach is evaluated through its implementation in a news recommender system, in which it is tested with real users. In such scenario, semantic knowledge bases and item annotations are automatically produced from public sources. Iván Cantador, Pablo Castells, Alejandro Bellogín |
Int. J. Semantic Web Inf. Syst. | 2 |
| 2011 | Effects of Usage-Based Feedback on Video Retrieval: A Simulation-Based StudyabstractWe present a model for exploiting community-based usage information for video retrieval, where implicit usage information from past users is exploited in order to provide enhanced assistance in video retrieval tasks, and alleviate the effects of the semantic gap problem. We propose a graph-based model for all types of implicit and explicit feedback, in which the relevant usage information is represented. Our model is designed to capture the complex interactions of a user with an interactive video retrieval system, including the representation of sequences of user-system interaction during a search session. Building upon this model, four recommendation strategies are defined and evaluated. An evaluation strategy is proposed based on simulated user actions, which enables the evaluation of our recommendation strategies over a usage information pool obtained from 24 users performing four different TRECVid tasks. Furthermore, the proposed simulation approach is used to simulate usage information pools with different characteristics, with which the recommendation approaches are further evaluated on a larger set of tasks, and their performance is studied with respect to the scalability and quality of the available implicit information. David Vallet, Frank Hopfgartner, Joemon M. Jose, Pablo Castells |
ACM Trans. Inf. Syst. | 4 |
| 2011 | Semantically enhanced Information Retrieval: An ontology-based approach
Miriam Fernández, Iván Cantador, Vanessa López, David Vallet, Pablo Castells, Enrico Motta |
J. Web Semant. | 5 |
| 2010 | A Performance Prediction Approach to Enhance Collaborative Filtering Performance
Alejandro Bellogín, Pablo Castells |
ECIR | 2 |
| 2010 | Workshop on the practical use of recommender systems algorithms & technologyabstractUser modeling, adaptation, and personalization techniques have hit the mainstream. The explosion of social network websites, on-line user-generated content platforms, and the tremendous growth in computational power of mobile devices are generating incredibly large amounts of user data, and an increasing desire of users to "personalize" (their desktop, e-mail, news site, phone). The potential value of personalization has become clear both as a commodity for the benefit or enjoyment of end-users, and as an enabler of new or better services -- a strategic opportunity to enhance and expand businesses. An exciting characteristic of recommender systems is that they draw the interest of industry and businesses while posing very interesting research and scientific challenges. Jérôme Picault, Dimitre Kostadinov, Pablo Castells, Alejandro Jaimes |
RecSys | 3 |
| 2010 | Inferring user intent in web search by exploiting social annotationsabstractThis is the author's version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record was published in Proceedings of the 33rd international ACM SIGIR conference on Research and development in information retrieval, http://dx.doi.org/10.1145/1835449.1835636 José M. Conde, David Vallet, Pablo Castells |
SIGIR | 3 |
| 2009 | Predicting Neighbor Goodness in Collaborative Filtering
Alejandro Bellogín, Pablo Castells |
FQAS | 2 |
| 2008 | Ontology-Based Personalised and Context-Aware Recommendations of News ItemsabstractNews@hand is a news recommender system that makes use of semantic technologies to provide several on-line news recommendation services. News contents and user preferences are described in terms of concepts appearing in a set of domain ontologies. Based on the similarities between item descriptions and user profiles, and the se-mantic relations between concepts, content-based and collaborative recommendation models are supported by the system. In this paper, we evaluate a model that personalises the order in which news articles are shown to the user according to his long-term interest profile, and other model that reorders the news items lists taking into account the current semantic context of interest of the user. The combination of those models is investigated showing significant improvements on the experimental tasks performed. Iván Cantador, Alejandro Bellogín, Pablo Castells |
Web Intelligence | 3 |
| 2007 | Automatising the learning of lexical patterns: An application to the enrichment of WordNet by extracting semantic relationships from Wikipedia
Maria Ruiz-Casado, Enrique Alfonseca, Pablo Castells |
Data Knowl. Eng. | 3 |
| 2007 | An Adaptation of the Vector-Space Model for Ontology-Based Information RetrievalabstractSemantic search has been one of the motivations of the semantic Web since it was envisioned. We propose a model for the exploitation of ontology-based knowledge bases to improve search over large document repositories. In our view of information retrieval on the semantic Web, a search engine returns documents rather than, or in addition to, exact values in response to user queries. For this purpose, our approach includes an ontology-based scheme for the semiautomatic annotation of documents and a retrieval system. The retrieval model is based on an adaptation of the classic vector-space model, including an annotation weighting algorithm, and a ranking algorithm. Semantic search is combined with conventional keyword-based retrieval to achieve tolerance to knowledge base incompleteness. Experiments are shown where our approach is tested on corpora of significant scale, showing clear improvements with respect to keyword-based search Pablo Castells, Miriam Fernández, David Vallet |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2006 | Probabilistic Score Normalization for Rank Aggregation
Miriam Fernández, David Vallet, Pablo Castells |
ECIR | 3 |
| 2006 | Multilayered Semantic Social Network Modeling by Ontology-Based User Profiles Clustering: Application to Collaborative Filtering
Iván Cantador, Pablo Castells |
EKAW | 2 |
| 2006 | Using historical data to enhance rank aggregationabstractRank aggregation is a pervading operation in IR technology. We hypothesize that the performance of score-based aggregation may be affected by artificial, usually meaningless deviations consis-tently occurring in the input score distributions, which distort the combined result when the individual biases differ from each other. We propose a score-based rank aggregation model where the source scores are normalized to a common distribution before being combined. Early experiments on available data from several TREC collections are shown to support our proposal. Miriam Fernández, David Vallet, Pablo Castells |
SIGIR | 3 |
| 2005 | An Ontology-Based Information Retrieval Model
David Vallet, Miriam Fernández, Pablo Castells |
ESWC | 3 |
| 2005 | Automatic Extraction of Semantic Relationships for WordNet by Means of Pattern Learning from Wikipedia
Maria Ruiz-Casado, Enrique Alfonseca, Pablo Castells |
NLDB | 3 |