EDBT 2026 Demo / reviewers in the wild / expert
Daan Odijk
dblp:37/9998
· DBLP profile ↗
17ranked-venue papers in the field
5as first author
5since 2021 · last 2025
0000-0003-0369-8857ORCID · corroborated
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 16 (5 first)Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RADio* - An Introduction to Measuring Normative Diversity in News RecommendationsabstractIn traditional recommender system literature, diversity is often seen as the opposite of similarity and typically defined as the distance between identified topics, categories, or word models. However, this is not expressive of the social science’s interpretation of diversity, which accounts for a news organization’s norms and values and which we here refer to as normative diversity. We introduce RADio, a versatile metrics framework to evaluate recommendations according to these normative goals. RADio introduces a rank-aware Jensen Shannon (JS) divergence. This combination accounts for (i) a user’s decreasing propensity to observe items further down a list and (ii) full distributional shifts as opposed to point estimates. We evaluate RADio’s ability to reflect five normative concepts in news recommendations on the Microsoft News Dataset and six (neural) recommendation algorithms, with the help of our metadata enrichment pipeline. We find that RADio provides insightful estimates that can potentially be used to inform news recommender system design. Sanne Vrijenhoek, Gabriel Bénédict, Mateo Gutierrez Granada, Daan Odijk |
Trans. Recomm. Syst. | 4 |
| 2023 | Intent-Satisfaction Modeling: From Music to Video StreamingabstractLogged behavioral data is a common resource for enhancing the user experience on streaming platforms. In music streaming, Mehrotra et al. have shown how complementing behavioral data with user intent can help predict and explain user satisfaction. Do their findings extend to video streaming? Compared to music streaming, video streaming platforms provide relatively shallow catalogs. Finding the right content demands more active and conscious commitment from users than in the music streaming setting. Video streaming platforms, in particular, could thus benefit from a better understanding of user intents and satisfaction level. We replicate Mehrotra et al.’s study from music to video streaming and extend their modeling framework on two fronts: (i) improved modeling accuracy (random forests), and (ii) interpretability (Bayesian models). Like the original study, we find that user intent affects behavior and satisfaction itself, even if to a lesser degree, based on data analysis and modeling. By proposing a grouping of intents into decisive and explorative categories we highlight a tension: decisive video streamers are not as keen to interact with the user interface as exploration-seeking ones. Meanwhile, music streamers explore by listening. In this study, we find that in video streaming, unsatisfied users provide the main signal: intent influences satisfaction levels together with behavioral data, depending on our decisive vs. explorative grouping. Gabriel Bénédict, Daan Odijk, Maarten de Rijke |
Trans. Recomm. Syst. | 2 |
| 2022 | RADio - Rank-Aware Divergence Metrics to Measure Normative Diversity in News RecommendationsabstractIn traditional recommender system literature, diversity is often seen as the opposite of similarity, and typically defined as the distance between identified topics, categories or word models. However, this is not expressive of the social science’s interpretation of diversity, which accounts for a news organization’s norms and values and which we here refer to as normative diversity. We introduce RADio, a versatile metrics framework to evaluate recommendations according to these normative goals. RADio introduces a rank-aware Jensen Shannon (JS) divergence. This combination accounts for (i) a user’s decreasing propensity to observe items further down a list and (ii) full distributional shifts as opposed to point estimates. We evaluate RADio’s ability to reflect five normative concepts in news recommendations on the Microsoft News Dataset and six (neural) recommendation algorithms, with the help of our metadata enrichment pipeline. We find that RADio provides insightful estimates that can potentially be used to inform news recommender system design. Sanne Vrijenhoek, Gabriel Bénédict, Mateo Gutierrez Granada, Daan Odijk, Maarten de Rijke |
RecSys | 4 |
| 2021 | Recommenders with a Mission: Assessing Diversity in News RecommendationsabstractNews recommenders help users to find relevant online content and have the potential to fulfill a crucial role in a democratic society, directing the scarce attention of citizens towards the information that is most important to them. Simultaneously, recent concerns about so-called filter bubbles, misinformation and selective exposure are symptomatic of the disruptive potential of these digital news recommenders. Recommender systems can make or break filter bubbles, and as such can be instrumental in creating either a more closed or a more open internet. Current approaches to evaluating recommender systems are often focused on measuring an increase in user clicks and short-term engagement, rather than measuring the user's longer term interest in diverse and important information. Sanne Vrijenhoek, Mesut Kaya, Nadia Metoui, Judith Möller, Daan Odijk, Natali Helberger |
CHIIR | 5 |
| 2021 | Recommendations at VideolandabstractShare on Recommendations at Videoland Authors: Mateo Gutierrez Granada RTL Nederland B.V., Netherlands RTL Nederland B.V., NetherlandsView Profile , Daan Odijk RTL Nederland B.V., Netherlands RTL Nederland B.V., NetherlandsView Profile Authors Info & Claims RecSys '21: Fifteenth ACM Conference on Recommender SystemsSeptember 2021 Pages 580–582https://doi.org/10.1145/3460231.3474617Online:13 September 2021Publication History 0citation278DownloadsMetricsTotal Citations0Total Downloads278Last 12 Months278Last 6 weeks11 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Mateo Gutierrez Granada, Daan Odijk |
RecSys | 2 |
| 2018 | The birth of collective memories: Analyzing emerging entities in text streamsabstractWe study how collective memories are formed online. We do so by tracking entities that emerge in public discourse, that is, in online text streams such as social media and news streams, before they are incorporated into Wikipedia, which, we argue, can be viewed as an online place for collective memory. By tracking how entities emerge in public discourse, that is, the temporal patterns between their first mention in online text streams and subsequent incorporation into collective memory, we gain insights into how the collective remembrance process happens online. Specifically, we analyze nearly 80,000 entities as they emerge in online text streams before they are incorporated into Wikipedia. The online text streams we use for our analysis comprise of social media and news streams, and span over 579 million documents in a time span of 18 months. We discover two main emergence patterns: entities that emerge in a “bursty” fashion, that is, that appear in public discourse without a precedent, blast into activity and transition into collective memory. Other entities display a “delayed” pattern, where they appear in public discourse, experience a period of inactivity, and then resurface before transitioning into our cultural collective memory. David Graus, Daan Odijk, Maarten de Rijke |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2017 | Online Learning to Rank for Recommender SystemsabstractBlendle is a New York Times backed startup that builds a platform where users can explore and support the world's best journalism. Users can read all content from 120 publications and only pay for what they read. Every morning, at Blendle, we have a huge cold-start problem when over 8.000 new articles from the latest editions of newspapers arrive in our system. At that moment, these articles are read by virtually no-one and we are tasked with sending out personalised newsletters to over 1 million users. We can thus not rely on collaborative filtering type of recommendations, nor can we use the popularity of the articles as clues for what our user might want to read. We overcome our cold-start problem by a mix of curation by our editorial team and an automated analysis of the content of these articles. We extract named entities, semantic links, authors, the language and plenty of stylometrics. For each of our users, we build a very fine grained profile based on the attributes of the articles that they read. The combination of enriched articles and user profiles is fed into our machine learning pipeline. We are currently experimenting with an online learning to rank setup, where each of our users is exposed to a slightly perturbed version of our ranking model. We observe the interactions of our users to infer in which direction we should be updating the model. Our editorial team gets up at around 5am every morning to read what was published over night. They are done reading and recommending their selection of articles around 8am, which is also the time we would ideally send out the newsletter so that our users, on their commute to work, can read our newsletter. These timing restrictions pose yet another challenge: our content analysis and machine learning pipeline needs to be really fast. We solve this by using a streaming infrastructure build on Kafka. In this infrastructure, an article is analysed and scored for relevance towards each of our users as soons as it arrives. This has the advantage that at 8am, when our editorial team is done reading, personalisation is much more lightweight. We use the precomputed relevance scores and balance them with diversity to arrive at a unique ranking for each of our users. In this talk, I will detail how we enrich articles in a streaming fashion and how we use online learning methods to learn a ranking model. I will also talk about how we deal with the time constraints of the problem we are trying to solve. Daan Odijk, Anne Schuth |
RecSys | 1 |
| 2016 | Balancing Relevance Criteria through Multi-Objective OptimizationabstractOffline evaluation of information retrieval systems typically focuses on a single effectiveness measure that models the utility for a typical user. Such a measure usually combines a behavior-based rank discount with a notion of document utility that captures the single relevance criterion of topicality. However, for individual users relevance criteria such as credibility, reputability or readability can strongly impact the utility. Also, for different information needs the utility can be a different mixture of these criteria. Because of the focus on single metrics, offline optimization of IR systems does not account for different preferences in balancing relevance criteria. Joost van Doorn, Daan Odijk, Diederik M. Roijers, Maarten de Rijke |
SIGIR | 2 |
| 2015 | Struggling and Success in Web SearchabstractWeb searchers sometimes struggle to find relevant information. Struggling leads to frustrating and dissatisfying search experiences, even if searchers ultimately meet their search objectives. Better understanding of search tasks where people struggle is important in improving search systems. We address this important issue using a mixed methods study using large-scale logs, crowd-sourced labeling, and predictive modeling. We analyze anonymized search logs from the Microsoft Bing Web search engine to characterize aspects of struggling searches and better explain the relationship between struggling and search success. To broaden our understanding of the struggling process beyond the behavioral signals in log data, we develop and utilize a crowd-sourced labeling methodology. We collect third-party judgments about why searchers appear to struggle and, if appropriate, where in the search task it became clear to the judges that searches would succeed (i.e., the pivotal query). We use our findings to propose ways in which systems can help searchers reduce struggling. Key components of such support are algorithms that accurately predict the nature of future actions and their anticipated impact on search outcomes. Our findings have implications for the design of search systems that help searchers struggle less and succeed more. Daan Odijk, Ryen W. White, Ahmed Awadallah 0001, Susan T. Dumais |
CIKM | 1 |
| 2015 | Supporting Exploration of Historical Perspectives Across Collections
Daan Odijk, Cristina Garbacea, Thomas Schoegje, Laura Hollink, Victor de Boer, Kees Ribbens, Jacco van Ossenbruggen |
TPDL | 1 |
| 2015 | Dynamic Query Modeling for Related Content FindingabstractWhile watching television, people increasingly consume additional content related to what they are watching. We consider the task of finding video content related to a live television broadcast for which we leverage the textual stream of subtitles associated with the broadcast. We model this task as a Markov decision process and propose a method that uses reinforcement learning to directly optimize the retrieval effectiveness of queries generated from the stream of subtitles. Our dynamic query modeling approach significantly outperforms state-of-the-art baselines for stationary query modeling and for text-based retrieval in a television setting. In particular we find that carefully weighting terms and decaying these weights based on recency significantly improves effectiveness. Moreover, our method is highly efficient and can be used in a live television setting, i.e., in near real time. Daan Odijk, Edgar Meij, Isaac Sijaranamual, Maarten de Rijke |
SIGIR | 1 |
| 2014 | Query-Dependent Contextualization of Streaming Data
Nikos Voskarides, Daan Odijk, Manos Tsagkias, Wouter Weerkamp, Maarten de Rijke |
ECIR | 2 |
| 2014 | Entity linking and retrieval for semantic searchabstractMore and more search engine users are expecting direct answers to their information needs, rather than links to documents. Semantic search and its recent applications enabled search engines to organize their wealth of information around entities. Entity linking and retrieval provide the building stones for organizing the web of entities. This tutorial aims to cover all facets of semantic search from a unified point of view and connect real-world applications with results from scientific publications. We provide a comprehensive overview of entity linking and retrieval in the context of semantic search and thoroughly explore techniques for query understanding, entity-based retrieval and ranking on unstructured text, structured knowledge repositories, and a mixture of these. We point out the connections between published approaches and applications, and provide hands-on examples on real-world use cases and datasets. Edgar Meij, Krisztian Balog, Daan Odijk |
WSDM | 3 |
| 2013 | Entity linking and retrievalabstractThis full-day tutorial presents a comprehensive introduction to entity linking and retrieval. Part I provides a detailed overview of entity linking: identifying and disambiguating entity occurrences in unstructured text. Part II focuses on entity retrieval, by first considering scenarios where explicit representations of entities are available, and then moving to a setting where evidence needs to be collected and aggregated from multiple documents or even collections, thereby combining techniques from both entity linking and entity retrieval. Part III concludes the tutorial with an overview and hands-on comparative analysis of applications and publicly available toolkits and web services. Edgar Meij, Krisztian Balog, Daan Odijk |
SIGIR | 3 |
| 2013 | ThemeStreams: visualizing the stream of themes discussed in politicsabstractThe political landscape is fluid. Discussions are always ongoing and new "hot topics" continue to appear in the headlines. But what made people start talking about that topic? And who started it? Because of the speed at which discussions sometimes take place this can be difficult to track down. We describe ThemeStreams: a demonstrator that maps political discussions to themes and influencers and illustrate how this mapping is used in an interactive visualization that shows us which themes are being discussed, and that helps us answer the question "Who put this issue on the map?" in streams of political data. Ork de Rooij, Daan Odijk, Maarten de Rijke |
SIGIR | 2 |
| 2012 | Semantic Document Selection - Historical Research on Collections That Span Multiple Centuries
Daan Odijk, Ork de Rooij, Maria-Hendrike Peetz, Toine Pieters, Maarten de Rijke, Stephen Snelders |
TPDL | 1 |
| 2011 | Instant Bag-of-Words served on a laptopabstractThis demo showcases our realtime implementation of concept classification using the Bag-of-Words method embedded within MediaTable, our interactive categorization tool for large multimedia collections. MediaTable allows the users to open images from disk or download these directly from the internet. Each image is then processed using the Bag-of-Words method, which computes classification scores for 20 distinct concepts classes on the fly. These are then seamlessly displayed in the interface. Jasper R. R. Uijlings, Ork de Rooij, Daan Odijk, Arnold W. M. Smeulders, Marcel Worring |
ICMR | 3 |