EDBT 2026 Demo / reviewers in the wild / expert
Alessandro G. Bottero
dblp:336/1676
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 61% Trustworthy machine learning · 22% Probabilistic and Bayesian machine learning · 17% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › value-based reinforcement learning
distributional reinforcement learning |
0.8 | 1 | 2024 | Value-Distributional Model-Based Reinforcement Learning · J. Mach. Learn. Res. 2024 |
Machine learning › Trustworthy machine learning › uncertainty estimation › epistemic uncertainty
epistemic uncertainty quantification |
0.8 | 1 | 2024 | Value-Distributional Model-Based Reinforcement Learning · J. Mach. Learn. Res. 2024 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.8 | 1 | 2024 | Value-Distributional Model-Based Reinforcement Learning · J. Mach. Learn. Res. 2024 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.6 | 1 | 2022 | Information-Theoretic Safe Exploration with Gaussian Processes · NeurIPS 2022 |
Machine learning › Reinforcement learning › safe reinforcement learning
safe exploration |
0.6 | 1 | 2022 | Information-Theoretic Safe Exploration with Gaussian Processes · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
soft actor-critic · 0.8quantile regression · 0.8bellman operator · 0.8information-theoretic criterion · 0.6gaussian process posterior · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Value-Distributional Model-Based Reinforcement LearningabstractQuantifying uncertainty about a policy's long-term performance is important to solve sequential decision-making tasks. We study the problem from a model-based Bayesian reinforcement learning perspective, where the goal is to learn the posterior distribution over value functions induced by parameter (epistemic) uncertainty of the Markov decision process. Previous work restricts the analysis to a few moments of the distribution over values or imposes a particular distribution shape, e.g., Gaussians. Inspired by distributional reinforcement learning, we introduce a Bellman operator whose fixed-point is the value distribution function. Based on our theory, we propose Epistemic Quantile-Regression (EQR), a model-based algorithm that learns a value distribution function. We combine EQR with soft actor-critic (SAC) for policy optimization with an arbitrary differentiable objective function of the learned value distribution. Evaluation across several continuous-control tasks shows performance benefits with respect to both model-based and model-free algorithms. The code is available at https://github.com/boschresearch/dist-mbrl. Carlos E. Luis, Alessandro G. Bottero, Julia Vinogradska, Felix Berkenkamp, Jan Peters 0001 |
J. Mach. Learn. Res. | 2 |
| 2023 | Model-Based Uncertainty in Value FunctionsabstractWe consider the problem of quantifying uncertainty over expected cumulative rewards in model-based reinforcement learning. In particular, we focus on characterizing the variance over values induced by a distribution over MDPs. Previous work upper bounds the posterior variance over values by solving a so-called uncertainty Bellman equation, but the over-approximation may result in inefficient exploration. We propose a new uncertainty Bellman equation whose solution converges to the true posterior variance over values and explicitly characterizes the gap in previous work. Moreover, our uncertainty quantification technique is easily integrated into common exploration strategies and scales naturally beyond the tabular setting by using standard deep reinforcement learning architectures. Experiments in difficult exploration tasks, both in tabular and continuous control settings, show that our sharper uncertainty estimates improve sample-efficiency. Carlos E. Luis, Alessandro G. Bottero, Julia Vinogradska, Felix Berkenkamp, Jan Peters 0001 |
AISTATS | 2 |
| 2022 | Information-Theoretic Safe Exploration with Gaussian ProcessesabstractWe consider a sequential decision making task where we are not allowed to evaluate parameters that violate an a priori unknown (safety) constraint. A common approach is to place a Gaussian process prior on the unknown constraint and allow evaluations only in regions that are safe with high probability. Most current methods rely on a discretization of the domain and cannot be directly extended to the continuous case. Moreover, the way in which they exploit regularity assumptions about the constraint introduces an additional critical hyperparameter. In this paper, we propose an information-theoretic safe exploration criterion that directly exploits the GP posterior to identify the most informative safe parameters to evaluate. Our approach is naturally applicable to continuous domains and does not require additional hyperparameters. We theoretically analyze the method and show that we do not violate the safety constraint with high probability and that we explore by learning about the constraint up to arbitrary precision. Empirical evaluations demonstrate improved data-efficiency and scalability. Alessandro G. Bottero, Carlos E. Luis, Julia Vinogradska, Felix Berkenkamp, Jan Peters 0001 |
NeurIPS | 1 |