EDBT 2026 Demo / reviewers in the wild / expert
Matteo Turchetta
dblp:182/1966
· DBLP profile ↗
13ranked-venue papers
5as first author
7since 2021 · last 2023
0000-0001-5881-3096ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Reinforcement learning · 44% Optimization for machine learning · 27% Motion planning and robot control · 20% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 79% Environmental and earth informatics · 21% |
Topics — the 21 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › safe reinforcement learning
safe exploration |
2.3 | 5 | 2023 | GoSafeOpt: Scalable safe exploration for global optimization of dynamical systems · Artif. Intell. 2023 Near-Optimal Multi-Agent Learning for Safe Coverage Control · NeurIPS 2022 Safe Reinforcement Learning via Curriculum Induction · NeurIPS 2020 |
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization |
1.7 | 4 | 2021 | Safe and Efficient Model-free Adaptive Control via Bayesian Optimization · ICRA 2021 Mixed-Variable Bayesian Optimization · IJCAI 2020 Robust Model-free Reinforcement Learning with Multi-objective Bayesian Optimization · ICRA 2020 |
Robotics › Motion planning and robot control
robot control |
1.5 | 4 | 2023 | Safe and Efficient Model-free Adaptive Control via Bayesian Optimization · ICRA 2021 GoSafe: Globally Optimal Safe Robot Learning · ICRA 2021 Safe Model-based Reinforcement Learning with Stability Guarantees · NIPS 2017 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
safe bayesian optimization |
1.4 | 3 | 2021 | Safe and Efficient Model-free Adaptive Control via Bayesian Optimization · ICRA 2021 GoSafe: Globally Optimal Safe Robot Learning · ICRA 2021 Safe Exploration for Interactive Machine Learning · NeurIPS 2019 |
Machine learning › Reinforcement learning
safe reinforcement learning |
1.3 | 3 | 2022 | Near-Optimal Multi-Agent Learning for Safe Coverage Control · NeurIPS 2022 Safe Reinforcement Learning via Curriculum Induction · NeurIPS 2020 Safe Model-based Reinforcement Learning with Stability Guarantees · NIPS 2017 |
Machine learning › Reinforcement learning
policy learning |
0.6 | 1 | 2022 | Learning Long-Term Crop Management Strategies with CyclesGym · NeurIPS 2022 |
Robotics › Motion planning and robot control › robot control
adaptive control |
0.5 | 1 | 2021 | Safe and Efficient Model-free Adaptive Control via Bayesian Optimization · ICRA 2021 |
Machine learning › Reinforcement learning
reward learning |
0.5 | 1 | 2021 | Information Directed Reward Learning for Reinforcement Learning · NeurIPS 2021 |
Robotics › Motion planning and robot control › robot learning
safe robot learning |
0.5 | 1 | 2021 | GoSafe: Globally Optimal Safe Robot Learning · ICRA 2021 |
Machine learning › Reinforcement learning › robust reinforcement learning
model-free robust reinforcement learning |
0.4 | 1 | 2020 | Robust Model-free Reinforcement Learning with Multi-objective Bayesian Optimization · ICRA 2020 |
Machine learning › Optimization for machine learning
multi-objective optimization |
0.4 | 1 | 2020 | Robust Model-free Reinforcement Learning with Multi-objective Bayesian Optimization · ICRA 2020 |
Machine learning › Reinforcement learning
robust reinforcement learning |
0.4 | 1 | 2020 | Robust Model-free Reinforcement Learning with Multi-objective Bayesian Optimization · ICRA 2020 |
Machine learning › Reinforcement learning › safe reinforcement learning
model-based safe reinforcement learning |
0.3 | 1 | 2017 | Safe Model-based Reinforcement Learning with Stability Guarantees · NIPS 2017 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.2 | 1 | 2016 | Safe Exploration in Finite Markov Decision Processes with Gaussian Processes · NIPS 2016 |
GPUs and heterogeneous computing
GPU computing |
0.2 | 1 | 2023 | ChromaX: a fast and scalable breeding program simulator · Bioinform. 2023 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent learning |
0.2 | 1 | 2022 | Near-Optimal Multi-Agent Learning for Safe Coverage Control · NeurIPS 2022 |
Machine learning › Reinforcement learning
regret minimization |
0.2 | 1 | 2022 | Near-Optimal Multi-Agent Learning for Safe Coverage Control · NeurIPS 2022 |
Environmental and earth informatics
agriculture |
0.2 | 1 | 2022 | Learning Long-Term Crop Management Strategies with CyclesGym · NeurIPS 2022 |
Machine learning › Efficient and distributed learning
active learning |
0.1 | 1 | 2021 | Information Directed Reward Learning for Reinforcement Learning · NeurIPS 2021 |
Robotics › Motion planning and robot control › robot learning › data-driven control
model-free control |
0.1 | 1 | 2021 | Safe and Efficient Model-free Adaptive Control via Bayesian Optimization · ICRA 2021 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
0.1 | 1 | 2020 | Mixed-Variable Bayesian Optimization · IJCAI 2020 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 2.3bayesian optimization · 1.4monte carlo simulation · 1.3GPU acceleration · 1.3crop growth model · 1.1gaussian process · 0.9safe optimization · 0.7model-free learning · 0.7submodular coverage function · 0.6exploration-exploitation trade-off · 0.6sample-efficient optimization · 0.5GOOSE · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | GoSafeOpt: Scalable safe exploration for global optimization of dynamical systemsabstractLearning optimal control policies directly on physical systems is challenging. Even a single failure can lead to costly hardware damage. Most existing model-free learning methods that guarantee safety, i.e., no failures, during exploration are limited to local optima. This work proposes GoSafeOpt as the first provably safe and optimal algorithm that can safely discover globally optimal policies for systems with high-dimensional state space. We demonstrate the superiority of GoSafeOpt over competing model-free safe learning methods in simulation and hardware experiments on a robot arm. Bhavya Sukhija, Matteo Turchetta, David Lindner, Andreas Krause 0001, Sebastian Trimpe, Dominik Baumann |
Artif. Intell. | 2 |
| 2023 | ChromaX: a fast and scalable breeding program simulatorabstractSUMMARY: ChromaX is a Python library that enables the simulation of genetic recombination, genomic estimated breeding value calculations, and selection processes. By utilizing GPU processing, it can perform these simulations up to two orders of magnitude faster than existing tools with standard hardware. This offers breeders and scientists new opportunities to simulate genetic gain and optimize breeding schemes. AVAILABILITY AND IMPLEMENTATION: The documentation is available at https://chromax.readthedocs.io. The code is available at https://github.com/kora-labs/chromax. Omar G. Younis, Matteo Turchetta, Daniel Ariza Suarez, Steven Yates, Bruno Studer, Ioannis N. Athanasiadis, Andreas Krause 0001, Joachim M. Buhmann, Luca Corinzia |
Bioinform. | 2 |
| 2022 | Near-Optimal Multi-Agent Learning for Safe Coverage ControlabstractIn multi-agent coverage control problems, agents navigate their environment to reach locations that maximize the coverage of some density. In practice, the density is rarely known $\textit{a priori}$, further complicating the original NP-hard problem. Moreover, in many applications, agents cannot visit arbitrary locations due to $\textit{a priori}$ unknown safety constraints. In this paper, we aim to efficiently learn the density to approximately solve the coverage problem while preserving the agents' safety. We first propose a conditionally linear submodular coverage function that facilitates theoretical analysis. Utilizing this structure, we develop MacOpt, a novel algorithm that efficiently trades off the exploration-exploitation dilemma due to partial observability, and show that it achieves sublinear regret. Next, we extend results on single-agent safe exploration to our multi-agent setting and propose SafeMac for safe coverage and exploration. We analyze SafeMac and give first of its kind results: near optimal coverage in finite time while provably guaranteeing safety. We extensively evaluate our algorithms on synthetic and real problems, including a bio-diversity monitoring task under safety constraints, where SafeMac outperforms competing methods. Manish Prajapat, Matteo Turchetta, Melanie Nicole Zeilinger, Andreas Krause 0001 |
NeurIPS | 2 |
| 2022 | Learning Long-Term Crop Management Strategies with CyclesGymabstractTo improve the sustainability and resilience of modern food systems, designing improved crop management strategies is crucial. The increasing abundance of data on agricultural systems suggests that future strategies could benefit from adapting to environmental conditions, but how to design these adaptive policies poses a new frontier. A natural technique for learning policies in these kinds of sequential decision-making problems is reinforcement learning (RL). To obtain the large number of samples required to learn effective RL policies, existing work has used mechanistic crop growth models (CGMs) as simulators. These solutions focus on single-year, single-crop simulations for learning strategies for a single agricultural management practice. However, to learn sustainable long-term policies we must be able to train in multi-year environments, with multiple crops, and consider a wider array of management techniques. We introduce CYCLESGYM, an RL environment based on the multi-year, multi-crop CGM Cycles. CYCLESGYM allows for long-term planning in agroecosystems, provides modular state space and reward constructors and weather generators, and allows for complex actions. For RL researchers, this is a novel benchmark to investigate issues arising in real-world applications. For agronomists, we demonstrate the potential of RL as a powerful optimization tool for agricultural systems management in multi-year case studies on nitrogen (N) fertilization and crop planning scenarios. Matteo Turchetta, Luca Corinzia, Scott Sussex, Amanda Burton, Juan Herrera, Ioannis N. Athanasiadis, Joachim M. Buhmann, Andreas Krause 0001 |
NeurIPS | 1 |
| 2021 | GoSafe: Globally Optimal Safe Robot LearningabstractWhen learning policies for robotic systems from data, safety is a major concern, as violation of safety constraints may cause hardware damage. SafeOpt is an efficient Bayesian optimization (BO) algorithm that can learn policies while guaranteeing safety with high probability. However, its search space is limited to an initially given safe region. We extend this method by exploring outside the initial safe area while still guaranteeing safety with high probability. This is achieved by learning a set of initial conditions from which we can recover safely using a learned backup controller in case of a potential failure. We derive conditions for guaranteed convergence to the global optimum and validate GoSafe in hardware experiments. Dominik Baumann, Alonso Marco, Matteo Turchetta, Sebastian Trimpe |
ICRA | 3 |
| 2021 | Safe and Efficient Model-free Adaptive Control via Bayesian OptimizationabstractAdaptive control approaches yield high-performance controllers when a precise system model or suitable parametrizations of the controller are available. Existing data-driven approaches for adaptive control mostly augment standard model-based methods with additional information about uncertainties in the dynamics or about disturbances. In this work, we propose a purely data-driven, model-free approach for adaptive control. Tuning low-level controllers based solely on system data raises concerns on the underlying algorithm safety and computational performance. Thus, our approach builds on GOOSE, an algorithm for safe and sample-efficient Bayesian optimization. We introduce several computational and algorithmic modifications in GOOSE that enable its practical use on a rotational motion system. We numerically demonstrate for several types of disturbances that our approach is sample efficient, outperforms constrained Bayesian optimization in terms of safety, and achieves the performance optima computed by grid evaluation. We further demonstrate the proposed adaptive control approach experimentally on a rotational motion system. Christopher König, Matteo Turchetta, John Lygeros, Alisa Rupenyan, Andreas Krause 0001 |
ICRA | 2 |
| 2021 | Information Directed Reward Learning for Reinforcement LearningabstractFor many reinforcement learning (RL) applications, specifying a reward is difficult. In this paper, we consider an RL setting where the agent can obtain information about the reward only by querying an expert that can, for example, evaluate individual states or provide binary preferences over trajectories. From such expensive feedback, we aim to learn a model of the reward function that allows standard RL algorithms to achieve high expected return with as few expert queries as possible. For this purpose, we propose Information Directed Reward Learning (IDRL), which uses a Bayesian model of the reward function and selects queries that maximize the information gain about the difference in return between potentially optimal policies. In contrast to prior active reward learning methods designed for specific types of queries, IDRL naturally accommodates different query types. Moreover, by shifting the focus from reducing the reward approximation error to improving the policy induced by the reward model, it achieves similar or better performance with significantly fewer queries. We support our findings with extensive evaluations in multiple environments and with different types of queries. David Lindner, Matteo Turchetta, Sebastian Tschiatschek, Kamil Ciosek, Andreas Krause 0001 |
NeurIPS | 2 |
| 2020 | Robust Model-free Reinforcement Learning with Multi-objective Bayesian OptimizationabstractIn reinforcement learning (RL), an autonomous agent learns to perform complex tasks by maximizing an exogenous reward signal while interacting with its environment. In real world applications, test conditions may differ substantially from the training scenario and, therefore, focusing on pure reward maximization during training may lead to poor results at test time. In these cases, it is important to trade-off between performance and robustness while learning a policy. While several results exist for robust, model-based RL, the model-free case has not been widely investigated. In this paper, we cast the robust, model-free RL problem as a multi-objective optimization problem. To quantify the robustness of a policy, we use delay margin and gain margin, two robustness indicators that are common in control theory. We show how these metrics can be estimated from data in the model-free setting. We use multi-objective Bayesian optimization (MOBO) to solve efficiently this expensive-to-evaluate, multi-objective optimization problem. We show the benefits of our robust formulation both in sim-to-real and pure hardware experiments to balance a Furuta pendulum. Matteo Turchetta, Andreas Krause 0001, Sebastian Trimpe |
ICRA | 1 |
| 2020 | Mixed-Variable Bayesian OptimizationabstractThe optimization of expensive to evaluate, black-box, mixed-variable functions, i.e. functions that have continuous and discrete inputs, is a difficult and yet pervasive problem in science and engineering. In Bayesian optimization (BO), special cases of this problem that consider fully continuous or fully discrete domains have been widely studied. However, few methods exist for mixed-variable domains and none of them can handle discrete constraints that arise in many real-world applications. In this paper, we introduce MiVaBo, a novel BO algorithm for the efficient optimization of mixed-variable functions combining a linear surrogate model based on expressive feature representations with Thompson sampling. We propose an effective method to optimize its acquisition function, a challenging problem for mixed-variable domains, making MiVaBo the first BO method that can handle complex constraints over the discrete variables. Moreover, we provide the first convergence analysis of a mixed-variable BO algorithm. Finally, we show that MiVaBo is significantly more sample efficient than state-of-the-art mixed-variable BO algorithms on several hyperparameter tuning tasks, including the tuning of deep generative models. Erik A. Daxberger, Anastasia Makarova, Matteo Turchetta, Andreas Krause 0001 |
IJCAI | 3 |
| 2020 | Safe Reinforcement Learning via Curriculum InductionabstractIn safety-critical applications, autonomous agents may need to learn in an environment where mistakes can be very costly. In such settings, the agent needs to behave safely not only after but also while learning. To achieve this, existing safe reinforcement learning methods make an agent rely on priors that let it avoid dangerous situations during exploration with high probability, but both the probabilistic guarantees and the smoothness assumptions inherent in the priors are not viable in many scenarios of interest such as autonomous driving. This paper presents an alternative approach inspired by human teaching, where an agent learns under the supervision of an automatic instructor that saves the agent from violating constraints during learning. In this model, we introduce the monitor that neither needs to know how to do well at the task the agent is learning nor needs to know how the environment works. Instead, it has a library of reset controllers that it activates when the agent starts behaving dangerously, preventing it from doing damage. Crucially, the choices of which reset controller to apply in which situation affect the speed of agent learning. Based on observing agents' progress the teacher itself learns a policy for choosing the reset controllers, a curriculum, to optimize the agent's final policy reward. Our experiments use this framework in two environments to induce curricula for safe and efficient learning. Matteo Turchetta, Andrey Kolobov, Shital Shah, Andreas Krause 0001, Alekh Agarwal |
NeurIPS | 1 |
| 2019 | Safe Exploration for Interactive Machine LearningabstractIn interactive machine learning (IML), we iteratively make decisions and obtain noisy observations of an unknown function. While IML methods, e.g., Bayesian optimization and active learning, have been successful in applications, on real-world systems they must provably avoid unsafe decisions. To this end, safe IML algorithms must carefully learn about a priori unknown constraints without making unsafe decisions. Existing algorithms for this problem learn about the safety of all decisions to ensure convergence. This is sample-inefficient, as it explores decisions that are not relevant for the original IML objective. In this paper, we introduce a novel framework that renders any existing unsafe IML algorithm safe. Our method works as an add-on that takes suggested decisions as input and exploits regularity assumptions in terms of a Gaussian process prior in order to efficiently learn about their safety. As a result, we only explore the safe set when necessary for the IML problem. We apply our framework to safe Bayesian optimization and to safe exploration in deterministic Markov Decision Processes (MDP), which have been analyzed separately before. Our method outperforms other algorithms empirically. Matteo Turchetta, Felix Berkenkamp, Andreas Krause 0001 |
NeurIPS | 1 |
| 2017 | Safe Model-based Reinforcement Learning with Stability GuaranteesabstractReinforcement learning is a powerful paradigm for learning optimal policies from experimental data. However, to find optimal policies, most reinforcement learning algorithms explore all possible actions, which may be harmful for real-world systems. As a consequence, learning algorithms are rarely applied on safety-critical systems in the real world. In this paper, we present a learning algorithm that explicitly considers safety, defined in terms of stability guarantees. Specifically, we extend control-theoretic results on Lyapunov stability verification and show how to use statistical models of the dynamics to obtain high-performance control policies with provable stability certificates. Moreover, under additional regularity assumptions in terms of a Gaussian process prior, we prove that one can effectively and safely collect data in order to learn about the dynamics and thus both improve control performance and expand the safe region of the state space. In our experiments, we show how the resulting algorithm can safely optimize a neural network policy on a simulated inverted pendulum, without the pendulum ever falling down. Felix Berkenkamp, Matteo Turchetta, Angela P. Schoellig, Andreas Krause 0001 |
NIPS | 2 |
| 2016 | Safe Exploration in Finite Markov Decision Processes with Gaussian ProcessesabstractIn classical reinforcement learning agents accept arbitrary short term loss for long term gain when exploring their environment. This is infeasible for safety critical applications such as robotics, where even a single unsafe action may cause system failure or harm the environment. In this paper, we address the problem of safely exploring finite Markov decision processes (MDP). We define safety in terms of an a priori unknown safety constraint that depends on states and actions and satisfies certain regularity conditions expressed via a Gaussian process prior. We develop a novel algorithm, SAFEMDP, for this task and prove that it completely explores the safely reachable part of the MDP without violating the safety constraint. To achieve this, it cautiously explores safe states and actions in order to gain statistical confidence about the safety of unvisited state-action pairs from noisy observations collected while navigating the environment. Moreover, the algorithm explicitly considers reachability when exploring the MDP, ensuring that it does not get stuck in any state with no safe way out. We demonstrate our method on digital terrain models for the task of exploring an unknown map with a rover. Matteo Turchetta, Felix Berkenkamp, Andreas Krause 0001 |
NIPS | 1 |