David Rohde

dblp:10/4983 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0002-0661-6266ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Inferential and Causal Principles for Better Understanding Interactive Systems
abstract
The practice of building high performance personalized interactive systems is founded on an artful combination of diverse machine learning methodologies including collaborative filtering, content based recommendation, contextual bandits, off policy estimation, click models, attribution and A/B testing. While combining these methods has been spectacularly successful, each methods finds justification using its own stylized protocols and there is little attention to the over-aching principles behind building reward optimizing recommender systems. This tutorial will focus on inferential principles including causal inference and relate these principles to current best practice in machine learning. The inferential principles covered will include Bayesian decision theory, coherence, the likelihood and conditionality principle as well as causal principles such as ignoreability, the do-calculus, Rubin Causal model, randomization and A/B testing. These foundations will then be used to investigate the differences between recommender systems best practices and approaches directly informed by inferential and causal principles, this section will both challenge best practice and the applicability of academic approaches. A lot of attention will be given to the pervasive problem of self confounding which will cover how engineering and machine learning best practice often results in production systems that unnecessarily suffer from confounding.
David Rohde
UMAP1
2025 Bayesian Off-Policy Evaluation and Learning for Large Action Spaces
abstract
In interactive systems, actions are often correlated, presenting an opportunity for more sample-efficient off-policy evaluation (OPE) and learning (OPL) in large action spaces. We introduce a unified Bayesian framework to capture these correlations through structured and informative priors. In this framework, we propose sDM, a generic Bayesian approach for OPE and OPL, grounded in both algorithmic and theoretical foundations. Notably, sDM leverages action correlations without compromising computational efficiency. Moreover, inspired by online Bayesian bandits, we introduce Bayesian metrics that assess the average performance of algorithms across multiple problem instances, deviating from the conventional worst-case assessments. We analyze sDM in OPE and OPL, highlighting the benefits of leveraging action correlations. Empirical evidence showcases the strong performance of sDM.
Imad Aouali, Victor-Emmanuel Brunel, David Rohde, Anna Korba
AISTATS3
2024 Why the Shooting in the Dark Method Dominates Recommender Systems Practice
abstract
The introduction of A/B Testing represented a great leap forward in recommender systems research. Like the randomized control trial for evaluating drug efficacy; A/B Testing has equipped recommender systems practitioners with a protocol for measuring performance as defined by actual business metrics and with minimal assumptions. While A/B testing provided a way to measure the performance of two or more candidate systems, it provides no guide for determining what policy we should test. The focus of this industry talk is to better understand, why the development of A/B testing was the last great leap forward in the development of reward optimizing recommender systems despite more than a decade of efforts in both industry and academia. The talk will survey: industry best practice, standard theories and tools including: collaborative filtering (MovieLens RecSys), contextual bandits, attribution, off-policy estimation, causal inference, click through rate models and will explain why we have converged on a fundamentally heuristic solution or guess and check type method. The talk will offer opinions about which of these theories are useful, and which are not and make a concrete proposal to make progress based on a non-standard use of deep learning tools.
David Rohde
RecSys1
2024 Unified PAC-Bayesian Study of Pessimism for Offline Policy Learning with Regularized Importance Sampling
abstract
Off-policy learning (OPL) often involves minimizing a risk estimator based on importance weighting to correct bias from the logging policy used to collect data. However, this method can produce an estimator with a high variance. A common solution is to regularize the importance weights and learn the policy by minimizing an estimator with penalties derived from generalization bounds specific to the estimator. This approach, known as pessimism, has gained recent attention but lacks a unified framework for analysis. To address this gap, we introduce a comprehensive PAC-Bayesian framework to examine pessimism with regularized importance weighting. We derive a tractable PAC-Bayesian generalization bound that universally applies to common importance weight regularizations, enabling their comparison within a single framework. Our empirical results challenge common understanding, demonstrating the effectiveness of standard IW regularization techniques.
Imad Aouali, Victor-Emmanuel Brunel, David Rohde, Anna Korba
UAI3
2023 Fast Offline Policy Optimization for Large Scale Recommendation
abstract
Personalised interactive systems such as recommender systems require selecting relevant items from massive catalogs dependent on context. Reward-driven offline optimisation of these systems can be achieved by a relaxation of the discrete problem resulting in policy learning or REINFORCE style learning algorithms. Unfortunately, this relaxation step requires computing a sum over the entire catalogue making the complexity of the evaluation of the gradient (and hence each stochastic gradient descent iterations) linear in the catalogue size. This calculation is untenable in many real world examples such as large catalogue recommender systems, severely limiting the usefulness of this method in practice. In this paper, we derive an approximation of these policy learning algorithms that scale logarithmically with the catalogue size. Our contribution is based upon combining three novel ideas: a new Monte Carlo estimate of the gradient of a policy, the self normalised importance sampling estimator and the use of fast maximum inner product search at training time. Extensive experiments show that our algorithm is an order of magnitude faster than naive approaches yet produces equally good policies.
Otmane Sakhi, David Rohde, Alexandre Gilotte
AAAI2
2023 Exponential Smoothing for Off-Policy Learning
abstract
Off-policy learning (OPL) aims at finding improved policies from logged bandit data, often by minimizing the inverse propensity scoring (IPS) estimator of the risk. In this work, we investigate a smooth regularization for IPS, for which we derive a two-sided PAC-Bayes generalization bound. The bound is tractable, scalable, interpretable and provides learning certificates. In particular, it is also valid for standard IPS without making the assumption that the importance weights are bounded. We demonstrate the relevance of our approach and its favorable performance through a set of learning tasks. Since our bound holds for standard IPS, we are able to provide insight into when regularizing IPS is useful. Namely, we identify cases where regularization might not be needed. This goes against the belief that, in practice, clipped IPS often enjoys favorable performance than standard IPS in OPL.
Imad Aouali, Victor-Emmanuel Brunel, David Rohde, Anna Korba
ICML3
2022 Reward Optimizing Recommendation using Deep Learning and Fast Maximum Inner Product Search
abstract
How can we build and optimize a recommender system that must rapidly fill slates (i.e. banners) of personalized recommendations? The combination of deep learning stacks with fast maximum inner product search (MIPS) algorithms have shown it is possible to deploy flexible models in production that can rapidly deliver personalized recommendations to users. Albeit promising, this methodology is unfortunately not sufficient to build a recommender system which maximizes the reward, e.g. the probability of click. Usually instead a proxy loss is optimized and A/B testing is used to test if the system actually improved performance. This tutorial takes participants through the necessary steps to model the reward and directly optimize the reward of recommendation engines built upon fast search algorithms to produce high-performance reward-optimizing recommender systems.
Imad Aouali, Amine Benhalloum, Martin Bompaire, Achraf Ait Sidi Hammou, Benjamin Heymann, David Rohde, Otmane Sakhi, Flavian Vasile, Maxime Vono
KDD7
2021 Bayesian Causal Inference for Real World Interactive Systems
abstract
Machine learning has allowed many systems that we interact with to improve performance and personalize. Recommender systems in particular are one of the largest users of machine learning in production environments that have improved performance of real-world systems. Learning in these interactive systems requires models that combine very diverse signals, including the logs of the interactive system (indicating if the intervention succeeded or failed) augmented with other data sources including: collaborative filtering, text, and image data. Bayesian inference is a compelling method to combine these diverse signals in a principled manner, but deployment of systems based on Bayesian principles remain challenging. The reward signal in the system logs is often uneven. Accurate estimation of reward is possible for exploiting actions, but often poor for other actions (exploration). Non-Bayesian methods such as inverse propensity score methods, the reinforce algorithm, and other heuristic-based approaches currently dominate practice. These commonly-used heuristics are often ineffective at leveraging diverse data. In contrast, Bayesian methods offer a principled, robust framework for learning from uneven signals and combining different types of information. Drawing upon the bandit and reinforcement learning community, in this workshop we will explore innovations in Bayesian inference for real world interactive systems, and consider advantages and limitations of the Bayesian approach.
Nicolas Chopin, Mike Gartrell, Dawen Liang, Alberto Lumbreras, David Rohde, Yixin Wang 0002
KDD5
2021 SimuRec: Workshop on Synthetic Data and Simulation Methods for Recommender Systems Research
abstract
There is significant interest lately in using synthetic data and simulation infrastructures for various types of recommender systems research. However, there are not currently any clear best practices around how best to apply these methods. We proposed a workshop to bring together researchers and practitioners interested in simulating recommender systems and their data to discuss the state of the art of such research and the pressing open methodological questions. The workshop resulted in a report authored by the participants that documents currently-known best practices on which the group has consensus and lays out an agenda for further research over the next 3–5 years to fill in places where we currently lack the information needed to make methodological recommendations.
Michael D. Ekstrand, Allison Chaney, Pablo Castells, Robin D. Burke, David Rohde, Manel Slokom
RecSys5
2020 Joint Policy-Value Learning for Recommendation
abstract
Conventional approaches to recommendation often do not explicitly take into account information on previously shown recommendations and their recorded responses. One reason is that, since we do not know the outcome of actions the system did not take, learning directly from such logs is not a straightforward task. Several methods for off-policy or counterfactual learning have been proposed in recent years, but their efficacy for the recommendation task remains understudied. Due to the limitations of offline datasets and the lack of access of most academic researchers to online experiments, this is a non-trivial task. Simulation environments can provide a reproducible solution to this problem.
Olivier Jeunen, David Rohde, Flavian Vasile, Martin Bompaire
KDD2
2020 BLOB: A Probabilistic Model for Recommendation that Combines Organic and Bandit Signals
abstract
A common task for recommender systems is to build a profile of the interests of a user from items in their browsing history and later to recommend items to the user from the same catalog. The users' behavior consists of two parts: the sequence of items that they viewed without intervention (the organic part) and the sequences of items recommended to them and their outcome (the bandit part).
Otmane Sakhi, Stephen Bonner, David Rohde, Flavian Vasile
KDD3
2020 Bayesian Value Based Recommendation: A modelling based alternative to proxy and counterfactual policy based recommendation
abstract
We develop the value based approach to Recommender systems. The value approach is a model based approach that allows forecasting of actual A/B test performance. It contrasts with the proxy based approach, which attempt to order the performance of different recommendation systems, but not forecast actual performance. It also contrasts with policy based approaches which also produce a performance forecast but use propensity scores to by-pass the requirement for a model. Value based approaches are a state of the art approach for combining organic and bandit signals that can utilise the three fundamental distances of recommendation. Their deployment requires sophisticated modelling and Bayesian computation. This tutorial develops the theory of value based recommendation and demonstrates the approach with examples in python notebooks.
David Rohde, Flavian Vasile, Otmane Sakhi
RecSys1
2020 A Gentle Introduction to Recommendation as Counterfactual Policy Learning
abstract
The objective of this tutorial is to give a structured overview of the conceptual frameworks behind current state-of-the-art recommender systems, explain their underlying assumptions, the resulting methods and their shortcomings, and to introduce an exciting new class of approaches that frames the task of recommendation as a counterfactual policy learning problem. The tutorial can be divided into two modules. In module 1, participants learn about current approaches for building real-world recommender systems that comprise mainly of two frameworks, namely: recommendation as optimal auto-completion of user behaviour and recommendation as reward modelling. In module 2, we present the framework of recommendation as a counterfactual policy learning problem and go over the theoretical guarantees that address the shortcomings of the previous frameworks. We then proceed to go over the associated algorithms and test them against classical methods in RecoGym, an open-source recommendation simulation environment.
Flavian Vasile, David Rohde, Olivier Jeunen, Amine Benhalloum
UMAP2
2016 Semiparametric Mean Field Variational Bayes: General Principles and Numerical Issues
abstract
We introduce the term semiparametric mean field variational Bayes to describe the relaxation of mean field variational Bayes in which some density functions in the product density restriction are pre-specified to be members of convenient parametric families. This notion has appeared in various guises in the mean field variational Bayes literature during its history and we endeavor to unify this important topic. We lay down a general framework and explain how previous relevant methodologies fall within this framework. A major contribution is elucidation of numerical issues that impact semiparametric mean field variational Bayes in practice.
David Rohde, Matthew P. Wand
J. Mach. Learn. Res.1
2012 Visualization of Predictive Distributions for Discrete Spatial-Temporal Log Cox Processes Approximated with MCMC
David Rohde, Jonathan Corcoran, Gentry White, Ruth Huang
IDEAL1
2004 Machine Learning for Matching Astronomy Catalogues
David Rohde, Michael Drinkwater, Marcus Gallagher, Tom Downs, Marianne Doyle
IDEAL1