EDBT 2026 Demo / reviewers in the wild / expert
Ian Osband
dblp:131/6683
· DBLP profile ↗
21ranked-venue papers
13as first author
5since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 13 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
18 papers |
Reinforcement learning · 76% Trustworthy machine learning · 11% Probabilistic and Bayesian machine learning · 6% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 100% |
Topics — the 30 heaviest of 31, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
exploration |
3.2 | 10 | 2020 | Hypermodels for Exploration · ICLR 2020 Deep Exploration via Randomized Value Functions · J. Mach. Learn. Res. 2019 Randomized Prior Functions for Deep Reinforcement Learning · NeurIPS 2018 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
1.6 | 3 | 2023 | Epistemic Neural Networks · NeurIPS 2023 The Neural Testbed: Evaluating Joint Predictions · NeurIPS 2022 Randomized Prior Functions for Deep Reinforcement Learning · NeurIPS 2018 |
Machine learning › Reinforcement learning
value-based reinforcement learning |
1.0 | 3 | 2019 | Deep Exploration via Randomized Value Functions · J. Mach. Learn. Res. 2019 The Uncertainty Bellman Equation and Exploration · ICML 2018 Deep Q-learning From Demonstrations · AAAI 2018 |
Machine learning › Reinforcement learning › exploration
randomized value functions |
0.6 | 2 | 2019 | Deep Exploration via Randomized Value Functions · J. Mach. Learn. Res. 2019 Generalization and Exploration via Randomized Value Functions · ICML 2016 |
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning |
0.5 | 3 | 2014 | Model-based Reinforcement Learning and the Eluder Dimension · NIPS 2014 Near-optimal Reinforcement Learning in Factored MDPs · NIPS 2014 (More) Efficient Reinforcement Learning via Posterior Sampling · NIPS 2013 |
Machine learning › Probabilistic and Bayesian machine learning › sampling
posterior sampling |
0.5 | 2 | 2017 | Why is Posterior Sampling Better than Optimism for Reinforcement Learning? · ICML 2017 (More) Efficient Reinforcement Learning via Posterior Sampling · NIPS 2013 |
Machine learning › Learning theory › online learning
regret bounds |
0.5 | 2 | 2017 | Minimax Regret Bounds for Reinforcement Learning · ICML 2017 (More) Efficient Reinforcement Learning via Posterior Sampling · NIPS 2013 |
Machine learning › Probabilistic and Bayesian machine learning
probabilistic inference |
0.4 | 1 | 2020 | Making Sense of Reinforcement Learning and Probabilistic Inference · ICLR 2020 |
Machine learning › Reinforcement learning › exploration
exploration-exploitation tradeoff |
0.4 | 2 | 2014 | Near-optimal Reinforcement Learning in Factored MDPs · NIPS 2014 (More) Efficient Reinforcement Learning via Posterior Sampling · NIPS 2013 |
Machine learning › Reinforcement learning
bellman equation |
0.3 | 1 | 2018 | The Uncertainty Bellman Equation and Exploration · ICML 2018 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
concurrent reinforcement learning |
0.3 | 1 | 2018 | Scalable Coordinated Exploration in Concurrent Reinforcement Learning · NeurIPS 2018 |
Machine learning › Reinforcement learning › exploration › multi-robot exploration
coordinated exploration |
0.3 | 1 | 2018 | Scalable Coordinated Exploration in Concurrent Reinforcement Learning · NeurIPS 2018 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.3 | 1 | 2018 | Deep Q-learning From Demonstrations · AAAI 2018 |
Robotics › Robot manipulation
learning from demonstration |
0.3 | 1 | 2018 | Deep Q-learning From Demonstrations · AAAI 2018 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.3 | 1 | 2018 | Scalable Coordinated Exploration in Concurrent Reinforcement Learning · NeurIPS 2018 |
Machine learning › Reinforcement learning › exploration
uncertainty-guided exploration |
0.3 | 1 | 2018 | The Uncertainty Bellman Equation and Exploration · ICML 2018 |
Machine learning › Reinforcement learning › regret minimization
bayesian regret |
0.3 | 1 | 2017 | Why is Posterior Sampling Better than Optimism for Reinforcement Learning? · ICML 2017 |
Machine learning › Reinforcement learning › markov decision process
finite-horizon MDP |
0.3 | 1 | 2017 | Why is Posterior Sampling Better than Optimism for Reinforcement Learning? · ICML 2017 |
Machine learning › Reinforcement learning
markov decision process |
0.3 | 1 | 2017 | Why is Posterior Sampling Better than Optimism for Reinforcement Learning? · ICML 2017 |
Machine learning › Reinforcement learning › dynamic programming › value iteration
optimistic value iteration |
0.3 | 1 | 2017 | Minimax Regret Bounds for Reinforcement Learning · ICML 2017 |
Machine learning › Reinforcement learning › dynamic programming
value iteration |
0.3 | 1 | 2017 | Minimax Regret Bounds for Reinforcement Learning · ICML 2017 |
Machine learning › Reinforcement learning › deep reinforcement learning
deep q-network |
0.2 | 1 | 2016 | Deep Exploration via Bootstrapped DQN · NIPS 2016 |
Machine learning › Reinforcement learning › value function approximation
randomized least-squares value iteration |
0.2 | 1 | 2016 | Generalization and Exploration via Randomized Value Functions · ICML 2016 |
Machine learning › Reinforcement learning
value function approximation |
0.2 | 1 | 2016 | Generalization and Exploration via Randomized Value Functions · ICML 2016 |
Machine learning › Reinforcement learning › exploration
efficient exploration |
0.2 | 1 | 2014 | Near-optimal Reinforcement Learning in Factored MDPs · NIPS 2014 |
Machine learning › Learning theory › online learning
eluder dimension |
0.2 | 1 | 2014 | Model-based Reinforcement Learning and the Eluder Dimension · NIPS 2014 |
Machine learning › Reinforcement learning › markov decision process
factored markov decision process |
0.2 | 1 | 2014 | Near-optimal Reinforcement Learning in Factored MDPs · NIPS 2014 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.2 | 1 | 2014 | Model-based Reinforcement Learning and the Eluder Dimension · NIPS 2014 |
Machine learning › Reinforcement learning › regret minimization
near-optimal regret bounds |
0.2 | 1 | 2014 | Near-optimal Reinforcement Learning in Factored MDPs · NIPS 2014 |
Performance modeling and evaluation
benchmarking |
0.1 | 1 | 2020 | Behaviour Suite for Reinforcement Learning · ICLR 2020 |
Methods — techniques the papers use, named apart from their topics
ensemble methods · 1.0joint prediction · 0.7randomized value functions · 0.6neural network data generating process · 0.6bayesian deep learning · 0.6reinforcement learning · 0.4probabilistic inference · 0.4hypernetwork · 0.4regret analysis · 0.4prioritized replay · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Epistemic Neural NetworksabstractIntelligence relies on an agent's knowledge of what it does not know.
This capability can be assessed based on the quality of joint predictions of labels across multiple inputs.
In principle, ensemble-based approaches can produce effective joint predictions, but the computational costs of large ensembles become prohibitive.
We introduce the epinet: an architecture that can supplement any conventional neural network, including large pretrained models, and can be trained with modest incremental computation to estimate uncertainty.
With an epinet, conventional neural networks outperform very large ensembles, consisting of hundreds or more particles, with orders of magnitude less computation.
The epinet does not fit the traditional framework of Bayesian neural networks.
To accommodate development of approaches beyond BNNs, such as the epinet, we introduce the epistemic neural network (ENN) as a general interface for models that produce joint predictions. Ian Osband, Zheng Wen 0002, Seyed Mohammad Asghari, Vikranth Reddy Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu, Benjamin Van Roy |
NeurIPS | 1 |
| 2023 | Approximate Thompson Sampling via Epistemic Neural NetworksabstractThompson sampling (TS) is a popular heuristic for action selection, but it requires sampling from a posterior distribution. Unfortunately, this can become computationally intractable in complex environments, such as those modeled using neural networks. Approximate posterior samples can produce effective actions, but only if they reasonably approximate joint predictive distributions of outputs across inputs. Notably, accuracy of marginal predictive distributions does not suffice. Epistemic neural networks (ENNs) are designed to produce accurate joint predictive distributions. We compare a range of ENNs through computational experiments that assess their performance in approximating TS across bandit and reinforcement learning environments. The results indicate that ENNs serve this purpose well and illustrate how the quality of joint predictive distributions drives performance. Further, we demonstrate that the epinet – a small additive network that estimates uncertainty – matches the performance of large ensembles at orders of magnitude lower computational cost. This enables effective application of TS with computation that scales gracefully to complex environments. Ian Osband, Zheng Wen 0002, Seyed Mohammad Asghari, Vikranth Reddy Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu, Benjamin Van Roy |
UAI | 1 |
| 2022 | The Neural Testbed: Evaluating Joint PredictionsabstractPredictive distributions quantify uncertainties ignored by point estimates. This paper introduces The Neural Testbed: an open source benchmark for controlled and principled evaluation of agents that generate such predictions. Crucially, the testbed assesses agents not only on the quality of their marginal predictions per input, but also on their joint predictions across many inputs. We evaluate a range of agents using a simple neural network data generating process.Our results indicate that some popular Bayesian deep learning agents do not fare well with joint predictions, even when they can produce accurate marginal predictions. We also show that the quality of joint predictions drives performance in downstream decision tasks. We find these results are robust across choice a wide range of generative models, and highlight the practical importance of joint predictions to the community. Ian Osband, Zheng Wen 0002, Seyed Mohammad Asghari, Vikranth Reddy Dwaracherla, Xiuyuan Lu, Morteza Ibrahimi, Dieterich Lawson, Botao Hao, Brendan O'Donoghue, Benjamin Van Roy |
NeurIPS | 1 |
| 2022 | Evaluating high-order predictive distributions in deep learningabstractMost work on supervised learning research has focused on marginal predictions. In decision problems, joint predictive distributions are essential for good performance. Previous work has developed methods for assessing low-order predictive distributions with inputs sampled i.i.d. from the testing distribution. With low-dimensional inputs, these methods distinguish agents that effectively estimate uncertainty from those that do not. We establish that the predictive distribution order required for such differentiation increases greatly with input dimension, rendering these methods impractical. To accommodate high-dimensional inputs, we introduce dyadic sampling, which focuses on predictive distributions associated with random pairs of inputs. We demonstrate that this approach efficiently distinguishes agents in high-dimensional examples involving simple logistic regression as well as complex synthetic and empirical data. Ian Osband, Zheng Wen 0002, Seyed Mohammad Asghari, Vikranth Reddy Dwaracherla, Xiuyuan Lu, Benjamin Van Roy |
UAI | 1 |
| 2021 | Matrix games with bandit feedbackabstractWe study a version of the classical zero-sum matrix game with unknown payoff matrix and bandit feedback, where the players only observe each others actions and a noisy payoff. This generalizes the usual matrix game, where the payoff matrix is known to the players. Despite numerous applications, this problem has received relatively little attention. Although adversarial bandit algorithms achieve low regret, they do not exploit the matrix structure and perform poorly relative to the new algorithms. The main contributions are regret analyses of variants of UCB and K-learning that hold for any opponent, e.g., even when the opponent adversarially plays the best-response to the learner’s mixed strategy. Along the way, we show that Thompson fails catastrophically in this setting and provide empirical comparison to existing algorithms. Brendan O'Donoghue, Tor Lattimore, Ian Osband |
UAI | 3 |
| 2020 | Hypermodels for Exploration
Vikranth Reddy Dwaracherla, Xiuyuan Lu, Morteza Ibrahimi, Ian Osband, Zheng Wen 0002, Benjamin Van Roy |
ICLR | 4 |
| 2020 | Making Sense of Reinforcement Learning and Probabilistic Inference
Brendan O'Donoghue, Ian Osband, Catalin Ionescu |
ICLR | 2 |
| 2020 | Behaviour Suite for Reinforcement Learning
Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, Eren Sezener, Andre Saraiva 0001, Katrina McKinney, Tor Lattimore, Csaba Szepesvári, Satinder Singh 0001, Benjamin Van Roy, Richard S. Sutton, David Silver 0001, Hado van Hasselt |
ICLR | 1 |
| 2019 | Deep Exploration via Randomized Value FunctionsabstractWe study the use of randomized value functions to guide deep exploration in reinforcement learning. This offers an elegant means for synthesizing statistically and computationally efficient exploration with common practical approaches to value function learning. We present several reinforcement learning algorithms that leverage randomized value functions and demonstrate their efficacy through computational studies. We also prove a regret bound that establishes statistical efficiency with a tabular representation. Ian Osband, Benjamin Van Roy, Daniel Russo 0001, Zheng Wen 0002 |
J. Mach. Learn. Res. | 1 |
| 2018 | Deep Q-learning From DemonstrationsabstractDeep reinforcement learning (RL) has achieved several high profile successes in difficult decision-making problems. However, these algorithms typically require a huge amount of data before they reach reasonable performance. In fact, their performance during learning can be extremely poor. This may be acceptable for a simulator, but it severely limits the applicability of deep RL to many real-world tasks, where the agent must learn in the real environment. In this paper we study a setting where the agent may access data from previous control of the system. We present an algorithm, Deep Q-learning from Demonstrations (DQfD), that leverages small sets of demonstration data to massively accelerate the learning process even from relatively small amounts of demonstration data and is able to automatically assess the necessary ratio of demonstration data while learning thanks to a prioritized replay mechanism. DQfD works by combining temporal difference updates with supervised classification of the demonstrator’s actions. We show that DQfD has better initial performance than Prioritized Dueling Double Deep Q-Networks (PDD DQN) as it starts with better scores on the first million steps on 41 of 42 games and on average it takes PDD DQN 83 million steps to catch up to DQfD’s performance. DQfD learns to out-perform the best demonstration given in 14 of 42 games. In addition, DQfD leverages human demonstrations to achieve state-of-the-art results for 11 games. Finally, we show that DQfD performs better than three related algorithms for incorporating demonstration data into DQN. Todd Hester, Matej Vecerík, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Daniel Horgan, John Quan, Andrew Sendonaris, Ian Osband, Gabriel Dulac-Arnold, John P. Agapiou, Joel Z. Leibo, Audrunas Gruslys |
AAAI | 10 |
| 2018 | Noisy Networks For Exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Matteo Hessel, Ian Osband, Alex Graves, Volodymyr Mnih, Rémi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, Shane Legg |
ICLR (Poster) | 6 |
| 2018 | The Uncertainty Bellman Equation and Exploration
Brendan O'Donoghue, Ian Osband, Rémi Munos, Volodymyr Mnih |
ICML | 2 |
| 2018 | Scalable Coordinated Exploration in Concurrent Reinforcement LearningabstractWe consider a team of reinforcement learning agents that concurrently operate in a common environment, and we develop an approach to efficient coordinated exploration that is suitable for problems of practical scale. Our approach builds on the seed sampling concept introduced in Dimakopoulou and Van Roy (2018) and on a randomized value function learning algorithm from Osband et al. (2016). We demonstrate that, for simple tabular contexts, the approach is competitive with those previously proposed in Dimakopoulou and Van Roy (2018) and with a higher-dimensional problem and a neural network value function representation, the approach learns quickly with far fewer agents than alternative exploration schemes. Maria Dimakopoulou, Ian Osband, Benjamin Van Roy |
NeurIPS | 2 |
| 2018 | Randomized Prior Functions for Deep Reinforcement LearningabstractDealing with uncertainty is essential for efficient reinforcement learning. There is a growing literature on uncertainty estimation for deep learning from fixed datasets, but many of the most popular approaches are poorly-suited to sequential decision problems. Other methods, such as bootstrap sampling, have no mechanism for uncertainty that does not come from the observed data. We highlight why this can be a crucial shortcoming and propose a simple remedy through addition of a randomized untrainable `prior' network to each ensemble member. We prove that this approach is efficient with linear representations, provide simple illustrations of its efficacy with nonlinear representations and show that this approach scales to large-scale problems far better than previous attempts. Ian Osband, John Aslanides, Albin Cassirer |
NeurIPS | 1 |
| 2017 | Minimax Regret Bounds for Reinforcement LearningabstractWe consider the problem of provably optimal exploration in reinforcement learning for finite horizon MDPs. We show that an optimistic modification to value iteration achieves a regret bound of $\tilde {O}( \sqrt{HSAT} + H^2S^2A+H\sqrt{T})$ where $H$ is the time horizon, $S$ the number of states, $A$ the number of actions and $T$ the number of time-steps. This result improves over the best previous known bound $\tilde {O}(HS \sqrt{AT})$ achieved by the UCRL2 algorithm. The key significance of our new results is that when $T\geq H^3S^3A$ and $SA\geq H$, it leads to a regret of $\tilde{O}(\sqrt{HSAT})$ that matches the established lower bound of $\Omega(\sqrt{HSAT})$ up to a logarithmic factor. Our analysis contain two key insights. We use careful application of concentration inequalities to the optimal value function as a whole, rather than to the transitions probabilities (to improve scaling in $S$), and we define Bernstein-based “exploration bonuses” that use the empirical variance of the estimated values at the next states (to improve scaling in $H$). Mohammad Gheshlaghi Azar, Ian Osband, Rémi Munos |
ICML | 2 |
| 2017 | Why is Posterior Sampling Better than Optimism for Reinforcement Learning?abstractComputational results demonstrate that posterior sampling for reinforcement learning (PSRL) dramatically outperforms existing algorithms driven by optimism, such as UCRL2. We provide insight into the extent of this performance boost and the phenomenon that drives it. We leverage this insight to establish an $\tilde{O}(H\sqrt{SAT})$ Bayesian regret bound for PSRL in finite-horizon episodic Markov decision processes. This improves upon the best previous Bayesian regret bound of $\tilde{O}(H S \sqrt{AT})$ for any reinforcement learning algorithm. Our theoretical results are supported by extensive empirical evaluation. Ian Osband, Benjamin Van Roy |
ICML | 1 |
| 2016 | Generalization and Exploration via Randomized Value FunctionsabstractWe propose randomized least-squares value iteration (RLSVI) – a new reinforcement learning algorithm designed to explore and generalize efficiently via linearly parameterized value functions. We explain why versions of least-squares value iteration that use Boltzmann or epsilon-greedy exploration can be highly inefficient, and we present computational results that demonstrate dramatic efficiency gains enjoyed by RLSVI. Further, we establish an upper bound on the expected regret of RLSVI that demonstrates near-optimality in a tabula rasa learning context. More broadly, our results suggest that randomized value functions offer a promising approach to tackling a critical challenge in reinforcement learning: synthesizing efficient exploration and effective generalization. Ian Osband, Benjamin Van Roy, Zheng Wen 0002 |
ICML | 1 |
| 2016 | Deep Exploration via Bootstrapped DQNabstractEfficient exploration remains a major challenge for reinforcement learning (RL). Common dithering strategies for exploration, such as epsilon-greedy, do not carry out temporally-extended (or deep) exploration; this can lead to exponentially larger data requirements. However, most algorithms for statistically efficient RL are not computationally tractable in complex environments. Randomized value functions offer a promising approach to efficient exploration with generalization, but existing algorithms are not compatible with nonlinearly parameterized value functions. As a first step towards addressing such contexts we develop bootstrapped DQN. We demonstrate that bootstrapped DQN can combine deep exploration with deep neural networks for exponentially faster learning than any dithering strategy. In the Arcade Learning Environment bootstrapped DQN substantially improves learning speed and cumulative performance across most games. Ian Osband, Charles Blundell, Alexander Pritzel, Benjamin Van Roy |
NIPS | 1 |
| 2014 | Near-optimal Reinforcement Learning in Factored MDPs
Ian Osband, Benjamin Van Roy |
NIPS | 1 |
| 2014 | Model-based Reinforcement Learning and the Eluder Dimension
Ian Osband, Benjamin Van Roy |
NIPS | 1 |
| 2013 | (More) Efficient Reinforcement Learning via Posterior SamplingabstractMost provably efficient learning algorithms introduce optimism about poorly-understood states and actions to encourage exploration. We study an alternative approach for efficient exploration, posterior sampling for reinforcement learning (PSRL). This algorithm proceeds in repeated episodes of known duration. At the start of each episode, PSRL updates a prior distribution over Markov decision processes and takes one sample from this posterior. PSRL then follows the policy that is optimal for this sample during the episode. The algorithm is conceptually simple, computationally efficient and allows an agent to encode prior knowledge in a natural way. We establish an $\tilde{O}(\tau S \sqrt{AT} )$ bound on the expected regret, where $T$ is time, $\tau$ is the episode length and $S$ and $A$ are the cardinalities of the state and action spaces. This bound is one of the first for an algorithm not based on optimism and close to the state of the art for any reinforcement learning algorithm. We show through simulation that PSRL significantly outperforms existing algorithms with similar regret bounds. Ian Osband, Daniel Russo 0001, Benjamin Van Roy |
NIPS | 1 |