VLDB 2026 Research / reviewers in the wild / expert
Alan Medlar
dblp:46/9459
· DBLP profile ↗
20ranked-venue papers in the field
4as first author
14since 2021 · last 2026
0000-0002-5139-9483ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 20 (4 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Expanded Tag Genomes for Cross-Domain RecommendationabstractTag genome is widely used in recommender systems research to, for example, measure item similarity, make recommendations and generate recommendation explanations. Applying tag genome to problems in cross-domain recommendation, however, is complicated by the limited item overlap between cross-domain recommendation data sets and the available tag genomes. Furthermore, existing tag prediction models rely on content-based features that are not readily available in a majority of recommendation data sets. To address these issues, we generated tag genomes for both movies and books based on the Amazon data set, which is widely used in cross-domain recommendation research. These new tag genomes are over 200 × larger than the previous versions and can support comparative evaluation of tag-based and collaborative methods, facilitate the development of new cross-domain recommendation algorithms and provide a foundation for studying phenomena, such as serendipity and diversity, across multiple domains. Both data sets and the data generation pipeline are freely available at https://github.com/Bionic1251/Expanded-Tag-Genomes. Denis Kotkov, Alan Medlar, Dorota Glowacka, Martin Halvey |
CHIIR | 2 |
| 2025 | Paths and Recreation: Inclusive Recommendation of Physical Activities in your NeighbourhoodabstractWe present Paths and Recreation, an interactive recommender system that integrates weather data and users’ current locations with a municipal service map to recommend nearby points-of-interest related to physical activities, such as outdoor gyms and swimming pools. The system can also generate and recommend routes for walking and cycling through your neighbourhood. Our primary goal was to make the recommendation of physical activities as inclusive as possible. We, therefore, included health-related filters to help users select activities at an appropriate level of intensity, and accessibility filters to assist, for example, wheelchair users, those with reduced mobility and people pushing strollers. We conducted a user study (N=16) to compare our system to common mobile apps (i.e., weather app, maps and web browser). In the study, participants needed to identify appropriate nearby physical activities by taking various health and accessibility requirements into consideration. Participants highlighted our system’s ease of use and found the integration of weather and health conditions to be useful for decision-making. Denis Kotkov, Alan Medlar, Dorota Glowacka |
CHIIR | 2 |
| 2025 | Don't Get Ahead of Yourself: A Critical Study on Data Leakage in Offline Evaluation of Sequential Recommenders
Huy Hoang Le, Yang Liu 0254, Alan Medlar, Dorota Glowacka |
RecSys | 3 |
| 2025 | Disentangling User and Item Sequence Patterns in Sequential Recommendation Data SetsabstractSequential recommenders use the ordering of user-item interactions to perform next-item prediction. Several studies have attempted to estimate how much sequential information is available in data sets used for the offline evaluation of sequential recommenders by randomly shuffling users' interaction histories and breaking the sequential dependencies between interactions. However, random shuffling fails to distinguish between sequential patterns from user behaviour (users consuming items based on previous interactions) and item availability (when items enter the system and become available for user consumption).In this article, we analyse several widely used data sets in sequential recommendation studies using two shuffling techniques: random shuffling and constrained shuffling. While random shuffling reorders interactions arbitrarily, constrained shuffling does not allow user-item interactions to occur prior to the item's first appearance in the data set. Our experiments show that sequential information can either come exclusively from user behaviour patterns or item availability, or from a combination of the two. These findings have implications for understanding evaluation results in sequential recommendation and highlights why some data sets may be less appropriate for offline evaluation given how little sequential information comes from user behaviour. Yang Liu 0254, Alan Medlar, Dorota Glowacka |
RecSys | 3 |
| 2024 | The Dark Matter of Serendipity in Recommender SystemsabstractSerendipity has been recognized as a valuable property of recommender systems. While there is a lack of consensus on the precise definition of serendipity, it is often conceptualized in terms of the relevance, novelty and unexpectedness of recommendations. However, the common understanding and original meaning of serendipity is conceptually broader, requiring serendipitous encounters to be neither novel nor unexpected. Recent work has highlighted the various ways in which serendipity can manifest, leading to a more generalized definition of serendipity. In this paper, we conducted an observational study where we collected 2002 survey responses from 397 users of an online article recommender system. In our study, we found a significant proportion of serendipitous recommendations were missed by the conventional definitions used in the recommender systems research literature, exposing the “dark matter” of serendipity that has been overlooked in prior studies. Interestingly, users’ opinions of which articles should be considered serendipitous did not strongly align with any of the definitions investigated. Furthermore, despite several user behaviors being significantly associated with a majority of definitions of serendipity, the overall goodness of fit was very low. Our findings highlight the issues of evaluating serendipity in recommender systems and the challenge of reconciling serendipity with user expectations. Denis Kotkov, Alan Medlar, Triin Kask, Dorota Glowacka |
CHIIR | 2 |
| 2024 | On the Negative Perception of Cross-domain Recommendations and ExplanationsabstractRecommender systems typically operate within a single domain, for example, recommending books based on users' reading habits. If such data is unavailable, it may be possible to make cross-domain recommendations and recommend books based on user preferences from another domain, such as movies. However, despite considerable research on cross-domain recommendations, no studies have investigated their impact on users' behavioural intentions or system perceptions compared to single-domain recommendations. Similarly, while single-domain explanations have been shown to improve users' perceptions of recommendations, there are no comparable studies for the cross-domain case. Denis Kotkov, Alan Medlar, Yang Liu 0254, Dorota Glowacka |
SIGIR | 2 |
| 2023 | Rethinking Serendipity in Recommender SystemsabstractRecommender systems suggest items, such as movies or books, to users based on their interests. These systems often suggest items that users are either already familiar with or could easily have found on their own without additional assistance. To overcome these problems, recommender systems aim to suggest serendipitous items. While there is a lack of consensus in the recommender systems research community on the definition of serendipity, it is often conceptualized as a complex combination of relevance, novelty and unexpectedness. However, the common understanding and original meaning of serendipity is conceptually broader, requiring serendipitous encounters to be neither novel nor unexpected. Recent work in the social sciences has highlighted the various ways that serendipity can manifest, leading to a more generalized definition of serendipity. We argue that the study of serendipity in recommender systems would benefit from considering items that are serendipitous under this more general definition, giving us a deeper understanding of the item characteristics and behavioral impact of serendipitous recommendations. These findings will help us to better optimize recommender systems for serendipity. In this paper, we explore various definitions of serendipity and propose a novel formalization of what it means for recommendations to be serendipitous. Lastly, we present an experimental design for how serendipity can be measured in a deployed recommender system. Denis Kotkov, Alan Medlar, Dorota Glowacka |
CHIIR | 2 |
| 2023 | What We Evaluate When We Evaluate Recommender Systems: Understanding Recommender Systems' Performance using Item Response TheoryabstractCurrent practices in offline evaluation use rank-based metrics to measure the quality of top-n recommendation lists. This approach has practical benefits as it centres assessment on the output of the recommender system and, therefore, measures performance from the perspective of end-users. However, this methodology neglects how recommender systems more broadly model user preferences, which is not captured by only considering the top-n recommendations. In this article, we use item response theory (IRT), a family of latent variable models used in psychometric assessment, to gain a comprehensive understanding of offline evaluation. We use IRT to jointly estimate the latent abilities of 51 recommendation algorithms and the characteristics of 3 commonly used benchmark data sets. For all data sets, the latent abilities estimated by IRT suggest that higher scores from traditional rank-based metrics do not reflect improvements in modeling user preferences. Furthermore, we show that the top-n recommendations with the most discriminatory power are biased towards lower difficulty items, leaving much room for improvement. Lastly, we highlight the role of popularity in evaluation by investigating how user engagement and item popularity influence recommendation difficulty. Yang Liu 0254, Alan Medlar, Dorota Glowacka |
RecSys | 2 |
| 2023 | On the Consistency, Discriminative Power and Robustness of Sampled Metrics in Offline Top-N Recommender System EvaluationabstractNegative item sampling in offline top-n recommendation evaluation has become increasingly wide-spread, but remains controversial. While several studies have warned against using sampled evaluation metrics on the basis of being a poor approximation of the full ranking (i.e. using all negative items), others have highlighted their improved discriminative power and potential to make evaluation more robust. Unfortunately, empirical studies on negative item sampling are based on relatively few methods (between 3-12) and, therefore, lack the statistical power to assess the impact of negative item sampling in practice. Yang Liu 0254, Alan Medlar, Dorota Glowacka |
RecSys | 2 |
| 2022 | The Tag Genome Dataset for BooksabstractAttaching tags to items, such as books or movies, is found in many online systems. While a majority of these systems use binary tags, continuous item-tag relevance scores, such as those in tag genome, offer richer descriptions of item content. For example, tag genome for movies assigns the tag “gangster” to the movie “The Godfather (1972)” with a score of 0.93 on a scale of 0 to 1. Tag genome has received considerable attention in recommender systems research and has been used in a wide variety of studies, from investigating the effects of recommender systems on users to generating ideas for movies that appeal to certain user groups. Denis Kotkov, Alan Medlar, Alexandr V. Maslov, Umesh Raj Satyal, Mats Neovius, Dorota Glowacka |
CHIIR | 2 |
| 2022 | ROGUE: A System for Exploratory Search of GANsabstractImage retrieval from generative adversarial networks (GANs) is challenging for several reasons. First, there are no clear mappings between the GAN's latent space and useful semantic features, making it difficult for users to navigate. Second, the number of unique images that can be generated is exceptionally high, taxing the scaling properties of existing search algorithms. In this article, we present ROGUE, a system to support exploratory search of images generated from GANs. We demonstrate how to implement features that are commonly found in exploratory search interfaces, such as faceted search and relevance feedback, in the context of GAN search. We additionally use reinforcement learning to help users navigate the image space [8], trading off exploration (showing diverse images) and exploitation (showing images predicted to receive positive relevance feedback). Finally, we present a usability study where participants were situated in the role of a casting director who needs to explore actors' headshots for an upcoming movie. The system obtained an average SUS score of 72.8 and all participants reported being either satisfied or very satisfied with the images they identified with the system. The system is shown in this accompanying video: https://vimeo.com/680036160. Yang Liu 0254, Alan Medlar, Dorota Glowacka |
SIGIR | 2 |
| 2022 | Lexical ambiguity detection in professional discourseabstractProfessional discourse is the language used by specialists, such as lawyers, doctors and academics, to communicate the knowledge and assumptions associated with their respective fields. Professional discourse can be especially difficult for non-specialists to understand due to the lexical ambiguity of commonplace words that have a different or more specific meaning within a specialist domain. This phenomena also makes it harder for specialists to communicate with the general public because they are similarly unaware of the potential for misunderstandings. In this article, we present an approach for detecting domain terms with lexical ambiguity versus everyday English. We demonstrate the efficacy of our approach with three case studies in statistics, law and biomedicine. In all case studies, we identify domain terms with a [email protected] greater than 0.9, outperforming the best performing baseline by 18.1–91.7%. Most importantly, we show this ranking is broadly consistent with semantic differences. Our results highlight the difficulties that existing semantic difference methods have in the cross-domain setting, which rank non-domain terms highly due to noise or biases in the data. We additionally show that our approach generalizes to short phrases and investigate its data efficiency by varying the number of labeled examples. Yang Liu 0254, Alan Medlar, Dorota Glowacka |
Inf. Process. Manag. | 2 |
| 2021 | Query Suggestions as Summarization in Exploratory SearchabstractQuery suggestions have been shown to benefit users performing information retrieval tasks. In exploratory search, however, users may lack the necessary domain knowledge to assess the relevance of query suggestions with respect to their information needs. In this article, we investigate the use of alternative queries in exploratory search. Alternative queries are queries that would retrieve similar search results to those currently visible on-screen. They are independent of the original search query and can, therefore, be updated dynamically as users scroll through search results. In addition to being follow-on queries, alternative queries serve as keyword summaries of the current search results page to help users assess whether results are inline with their search intents. We investigated the use of alternative queries in scientific literature search and their impact on user behavior and perception. In a user study, participants inspected half as many documents per query when alternative queries were present, but were exposed to over 40% more search results overall. Despite using them extensively as follow-on queries, user feedback focused on the summarization properties offered by alternative queries; finding it reassuring that documents were relevant to their search goals. Alan Medlar, Dorota Glowacka |
CHIIR | 1 |
| 2021 | Exploratory Search of GANs with Contextual BanditsabstractInteractive image retrieval involves users searching a collection of images to satisfy their subjective information needs. However, even large image collections are finite and therefore may not be able to satisfy users. An alternate approach would be to explore a generative adversarial network (GAN) and model users' search intents directly in terms of the latent space used by the GAN to generate images. In this article, we present a simulation study exploring the performance of Gaussian Process bandits in the context of interactive GAN exploration. We used recent advances in interpretable GAN controls to investigate the scalability of different approaches in terms of image space dimensionality. While we present several experiments with promising results, none of the approaches tested scale sufficiently well to explore the entire GAN image space. Ivan Kropotov, Alan Medlar, Dorota Glowacka |
CIKM | 2 |
| 2019 | Holes in the Outline: Subject-dependent Abstract Quality and its Implications for Scientific Literature SearchabstractScientific literature search engines typically index abstracts instead of the full-text of publications. The expectation is that the abstract provides a comprehensive summary of the article, enumerating key points for the reader to assess whether their information needs could be satisfied by reading the full-text. Furthermore, from a practical standpoint, obtaining the full-text is more complicated due to licensing issues, in the case of commercial publishers, and resource limitations of public repositories and pre-print servers. Chien-Yu Huang, Arlene Casey, Dorota Glowacka, Alan Medlar |
CHIIR | 4 |
| 2019 | How Relevance Feedback is Framed Affects User Experience, but not BehaviourabstractRetrieval systems based on machine learning require both positive and negative examples to perform inference, which is usually obtained through relevance feedback. Unfortunately, explicit negative relevance feedback is thought to have poor user experience. Instead, systems typically rely on implicit negative feedback. In this study, we confirm that, in the case of binary relevance feedback, users prefer giving positive feedback (and implicit negative feedback) over negative feedback (and implicit positive feedback). These two feedback mechanisms are functionally equivalent, capturing the same information from the user, but differ in how they are framed. Despite users' preference for positive feedback, there were no significant differences in behaviour. As users were not shown how feedback influenced search results, we hypothesise that previously reported results could, at least in part, be due to cognitive biases related to user perception of negative feedback. Dhruv Tripathi, Alan Medlar, Dorota Glowacka |
CHIIR | 2 |
| 2018 | How Consistent is Relevance Feedback in Exploratory Search?abstractSearch activities involving knowledge acquisition, investigation and synthesis are collectively known as exploratory search. Exploratory search is challenging for users, who may be unable to formulate search queries, have ill-defined search goals or may even struggle to understand search results. To ameliorate these difficulties, reinforcement learning-based information retrieval systems were developed to provide adaptive support to users. Reinforcement learning is used to build a model of user intent based on relevance feedback provided by the user. But how reliable is relevance feedback in this context? To answer this question, we developed a novel permutation-based metric for scoring the consistency of relevance feedback. We used this metric to perform a retrospective analysis of interaction data from lookup and exploratory search experiments. Our analysis shows that for lookup search relevance judgments are highly consistent, supporting previous findings that relevance feedback improves retrieval performance. For exploratory search, however, the distribution of consistency scores shows considerable inconsistency. Alan Medlar, Dorota Glowacka |
CIKM | 1 |
| 2017 | Using Topic Models to Assess Document Relevance in Exploratory Search User StudiesabstractEvaluation is crucial in assessing the effectiveness of new information retrieval and human computer interaction techniques and systems. Relevance judgements are often performed by humans, which makes obtaining them expensive and time consuming. Consequently, relevance judgements are usually performed only on a subset of a given collection of data or experimental results with a focus on the top ranked documents. However, when assessing the performance of exploratory search systems, the diversity or subjective relevance of documents that the user was presented with over a search session are often of more importance than the relative ranking of top documents. In order to perform these types of assessment, all the documents in a given collection need to be judged for relevance. In this paper, we propose an approach based on topic modeling that can greatly accelerate document relevance judgment of an entire document collection with an expert assessor needing to mark only a small subset of documents from a given collection. Experimental results show a substantial overlap between relevance judgments compared to a human assessor. Alan Medlar, Dorota Glowacka |
CHIIR | 1 |
| 2016 | PULP: A System for Exploratory Search of Scientific LiteratureabstractDespite the growing importance of exploratory search, information retrieval (IR) systems tend to focus on lookup search. Lookup searches are well served by optimising the precision and recall of search results, however, for exploratory search this may be counterproductive if users are unable to formulate an appropriate search query. We present a system called PULP that supports exploratory search for scientific literature, though the system can be easily adapted to other types of literature. PULP uses reinforcement learning (RL) to avert the user from context traps resulting from poorly chosen search queries, trading off between exploration (presenting the user with diverse topics) and exploitation (moving towards more specific topics). Where other RL-based systems suffer from the "cold start" problem, requiring sufficient time to adjust to a user's information needs, PULP initially presents the user with an overview of the dataset using temporal topic models. Topic models are displayed in an interactive alluvial diagram, where topics are shown as ribbons that change thickness with a given topics relative prevalence over time. Interactive, exploratory search sessions can be initiated by selecting topics as a starting point. Alan Medlar, Kalle Ilves, Wray L. Buntine, Dorota Glowacka |
SIGIR | 1 |
| 2015 | Balancing Exploration and Exploitation: Empirical Parameterization of Exploratory Search SystemsabstractExploratory searches are where a user has insufficient knowledge to define exact search criteria or does not otherwise know what they are looking for. Reinforcement learning techniques have demonstrated great potential for supporting exploratory search in information retrieval systems as they allow the system to trade-off exploration (presenting the user with alternatives topics) and exploitation (moving toward more specific topics). Users of such systems, however, often feel that the system is not responsive to user needs. This problem is not an inherent feature of such systems, but is caused by the exploration rate parameter being inappropriately tuned for a given system, dataset or user. Kumaripaba Athukorala, Alan Medlar, Kalle Ilves, Dorota Glowacka |
CIKM | 2 |