Cedric Waterschoot

dblp:300/5293 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0003-4903-2604ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 GMAP 2026: 5th Workshop on Group Modeling, Adaptation and Personalization
abstract
Group Recommender Systems (GRSys) are designed to recommend items that address the needs of groups of people. Compared to individual users, groups are dynamic entities where interpersonal relationships, group dynamics, emotional contagion, etc., substantially affect the group’s needs. Nevertheless, these characteristics are often poorly defined or overlooked in system modeling. The fifth GMAP workshop brought together a community of scholars focused on group modeling, adaptation, and personalization. The event was dedicated to exploring the challenges and opportunities of supporting collective decision-making, fostering interdisciplinary dialogue, and forging new collaborations. The three presented papers covered a diverse range of topics, revisiting assumptions about similarity, task definitions, and fairness perception in particular scenarios within the realm of group modeling, adaptation, and personalization.
Francesco Barile, Amra Delic, Ladislav Peska, Cedric Waterschoot
UMAP4
2026 FIRE: A Modular Framework for Prototyping and Evaluating Interactive Recommender Explanations
abstract
Explanations in recommender systems are important for enhancing transparency, trust, and system understanding. However, as research interest in interactive and chat-based explanations grows, the absence of standardized, reusable interface components limits their systematic evaluation. We introduce the Framework for Interactive Recommender Explanations (FIRE) to bridge this gap. FIRE provides an open-source library of modular explanation styles with varying levels of interactivity — spanning from static lists and draggable charts to conversational agents. To test these modalities in online user studies, FIRE integrates a supporting serverless experiment engine equipped with dynamic routing, crowdsourcing integration, and complete interaction logging. Currently configured for group recommenders but built for broader applicability, FIRE helps to standardize the evaluation of the explanation layer, thereby driving reproducible experimental design.
Ulysse Maes, Cedric Waterschoot, Francesco Barile, Nava Tintarev
UMAP2
2026 Who is the Fairest of Them All? Using Large Language Models for Fairness Assessments
abstract
As algorithmic decision-making becomes more prevalent, ensuring these systems align with human perceptions of fairness becomes critical. While user studies usually can provide ground truth for fairness assessments, the emergence of Large Language Models (LLMs) as judges offers a potential path toward scalable and real-time evaluation. This study investigates whether LLMs can accurately replicate human fairness distributions using a dataset of human evaluations regarding group recommendation fairness. We evaluate the performance of base models as well as a fine-tuned model against human ground truth. Our findings reveal that base models often produce biased judgments that fail to capture the variety of human responses. Additionally, our results demonstrate that fine-tuning LLMs on human judgment data significantly improves their ability to mirror human fairness assessments on unseen data. This work highlights the limitations of off-the-shelf models and the promise of fine-tuning for automating algorithmic fairness evaluation. Our methodology can be extended beyond fairness to evaluate relevant subjective judgments such as user satisfaction to ground and adapt algorithmic outcomes.
Cedric Waterschoot, Nava Tintarev, Francesco Barile
UMAP1
2026 Critical reflections on user studies' evaluation methods for group recommender systems
abstract
Social choice-based aggregation strategies are often used in group recommender systems to aggregate individual preferences or recommendations. However, previous works evaluating group recommenders with user studies found that the diversity of the group members’ preferences impacts the effectiveness of the strategies. In this paper, we highlight and address the methodological limitations of those previous works. Specifically, the methodologies we introduce demonstrate the following three novelties: 1) We evaluated the strategies from the viewpoint of an “internal evaluator”; 2) We introduced a novel methodology for modeling a fictional but realistic group with specific preference profiles for the group members, defining scenarios with concrete users and items, that are still mapped to specific group configurations; 3) We evaluated the understanding of the participants, by measuring how well they can successfully apply the aggregation strategy to a new scenario. To do this we performed a randomized controlled trial (n=444) using a mixed design with two between-subject factors (the used aggregation strategy and the presented explanation type ), and a within-subject factor (the group configuration ). Our results, with friend groups, showed significant differences in the effectiveness of the aggregation strategies depending on the specific group configuration (i.e., depending on the internal diversity of group members’ preferences), with noticeable differences between evaluations of what is good for the group – external evaluation – and what is good for the participant – internal evaluation . We conclude with methodological implications for group recommender systems.
Francesco Barile, Pierre Hurlin, Cedric Waterschoot, Nava Tintarev
Int. J. Hum. Comput. Stud.3
2025 Consistent Explainers or Unreliable Narrators? Understanding LLM-generated Group Recommendations
abstract
Large Language Models (LLMs) are increasingly being implemented as joint decision-makers and explanation generators for Group Recommender Systems (GRS). In this paper, we evaluate these recommendations and explanations by comparing them to social choice-based aggregation strategies. Our results indicate that LLM-generated recommendations often resembled those produced by Additive Utilitarian (ADD) aggregation. However, the explanations typically referred to averaging ratings (resembling but not identical to ADD aggregation). Group structure, uniform or divergent, did not impact the recommendations. Furthermore, LLMs regularly claimed additional criteria such as user or item similarity, diversity, or used undefined popularity metrics or thresholds. Our findings have important implications for LLMs in the GRS pipeline as well as standard aggregation strategies. Additional criteria in explanations were dependent on the number of ratings in the group scenario, indicating potential inefficiency of standard aggregation methods at larger item set sizes. Additionally, inconsistent and ambiguous explanations undermine transparency and explainability, which are key motivations behind the use of LLMs for GRS.
Cedric Waterschoot, Nava Tintarev, Francesco Barile
RecSys1
2025 With Friends Like These, Who Needs Explanations? Evaluating User Understanding of Group Recommendations
abstract
Group Recommender Systems (GRS) employing social choice-based aggregation strategies have previously been explored in terms of perceived consensus, fairness, and satisfaction. At the same time, the impact of textual explanations has been examined, but the results suggest a low effectiveness of these explanations. However, user understanding remains fairly unexplored, even if it can contribute positively to transparent GRS. This is particularly interesting to study in more complex or potentially unfair scenarios when user preferences diverge, such as in a minority scenario (where group members have similar preferences, except for a single member in a minority position). In this paper, we analyzed the impact of different types of explanations on user understanding of group recommendations. We present a randomized controlled trial (n = 271) using two between-subject factors: (i) the aggregation strategy (additive, least misery, and approval voting), and (ii) the modality of explanation (no explanation, textual explanation, or multimodal explanation). We measured both subjective (self-perceived by the user) and objective understanding (performance on model simulation, counterfactuals and error detection). In line with recent findings on explanations for machine learning models, our results indicate that more detailed explanations, whether textual or multimodal, did not increase subjective or objective understanding. However, we did find a significant effect of aggregation strategies on both subjective and objective understanding. These results imply that when constructing GRS, practitioners need to consider that the choice of aggregation strategy can influence the understanding of users. Post-hoc analysis also suggests that there is value in analyzing performance on different tasks, rather than through a single aggregated metric of understanding.
Cedric Waterschoot, Raciel Yera, Nava Tintarev, Francesco Barile
UMAP1
2024 The Impact of Featuring Comments in Online Discussions
Cedric Waterschoot, Ernst van den Hemel, Antal van den Bosch
ASONAM (2)1
2022 Detecting Minority Arguments for Mutual Understanding: A Moderation Tool for the Online Climate Change Debate
abstract
Moderating user comments and promoting healthy understanding is a challenging task, especially in the context of polarized topics such as climate change. We propose a moderation tool to assist moderators in promoting mutual understanding in regard to this topic. The approach is twofold. First, we train classifiers to label incoming posts for the arguments they entail, with a specific focus on minority arguments. We apply active learning to further supplement the training data with rare arguments. Second, we dive deeper into singular arguments and extract the lexical patterns that distinguish each argument from the others. Our findings indicate that climate change arguments form clearly separable clusters in the embedding space. These classes are characterized by their own unique lexical patterns that provide a quick insight in an argument’s key concepts. Additionally, supplementing our training data was necessary for our classifiers to be able to adequately recognize rare arguments. We argue that this detailed rundown of each argument provides insight into where others are coming from. These computational approaches can be part of the toolkit for content moderators and researchers struggling with polarized topics.
Cedric Waterschoot, Ernst van den Hemel, Antal van den Bosch
COLING1
2021 Calculating Argument Diversity in Online Threads
abstract
We propose a method for estimating argument diversity and interactivity in online discussion threads. Using a case study on the subject of Black Pete ("Zwarte Piet") in the Netherlands, the approach for automatic detection of echo chambers is presented. Dynamic thread scoring calculates the status of the discussion on the thread level, while individual messages receive a contribution score reflecting the extent to which the post contributed to the overall interactivity in the thread. We obtain platform-specific results. Gab hosts only echo chambers, while the majority of Reddit threads are balanced in terms of perspectives. Twitter threads cover the whole spectrum of interactivity. While the results based on the case study mirror previous research, this calculation is only the first step towards better understanding and automatic detection of echo effects in online discussions.
Cedric Waterschoot, Antal van den Bosch, Ernst van den Hemel
LDK1