EDBT 2026 Demo / reviewers in the wild / expert
Manu Orsini
dblp:267/2234
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 70% Optimization for machine learning · 30% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
imitation learning |
1.0 | 2 | 2021 | What Matters for Adversarial Imitation Learning? · NeurIPS 2021 Hyperparameter Selection for Imitation Learning · ICML 2021 |
Machine learning › Reinforcement learning
actor-critic methods |
0.5 | 1 | 2021 | What Matters for On-Policy Deep Actor-Critic Methods? A Large-Scale Study · ICLR 2021 |
Machine learning › Reinforcement learning › imitation learning › occupancy matching
adversarial imitation learning |
0.5 | 1 | 2021 | What Matters for Adversarial Imitation Learning? · NeurIPS 2021 |
Machine learning › Optimization for machine learning
hyperparameter optimization |
0.5 | 1 | 2021 | Hyperparameter Selection for Imitation Learning · ICML 2021 |
Machine learning › Optimization for machine learning › hyperparameter optimization
hyperparameter sensitivity |
0.5 | 1 | 2021 | What Matters for On-Policy Deep Actor-Critic Methods? A Large-Scale Study · ICLR 2021 |
Machine learning › Reinforcement learning
continuous control |
0.3 | 2 | 2021 | What Matters for Adversarial Imitation Learning? · NeurIPS 2021 Hyperparameter Selection for Imitation Learning · ICML 2021 |
Methods — techniques the papers use, named apart from their topics
empirical study · 1.0proxy reward functions · 0.5large-scale empirical study · 0.5adversarial training · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | What Matters for On-Policy Deep Actor-Critic Methods? A Large-Scale Study
Marcin Andrychowicz, Anton Raichuk, Piotr Stanczyk, Manu Orsini, Sertan Girgin, Raphaël Marinier, Léonard Hussenot, Matthieu Geist, Olivier Pietquin, Marcin Michalski, Sylvain Gelly, Olivier Bachem |
ICLR | 4 |
| 2021 | Hyperparameter Selection for Imitation LearningabstractWe address the issue of tuning hyperparameters (HPs) for imitation learning algorithms in the context of continuous-control, when the underlying reward function of the demonstrating expert cannot be observed at any time. The vast literature in imitation learning mostly considers this reward function to be available for HP selection, but this is not a realistic setting. Indeed, would this reward function be available, it could then directly be used for policy training and imitation would not be necessary. To tackle this mostly ignored problem, we propose a number of possible proxies to the external reward. We evaluate them in an extensive empirical study (more than 10’000 agents across 9 environments) and make practical recommendations for selecting HPs. Our results show that while imitation learning algorithms are sensitive to HP choices, it is often possible to select good enough HPs through a proxy to the reward function. Léonard Hussenot, Marcin Andrychowicz, Damien Vincent, Robert Dadashi, Anton Raichuk, Sabela Ramos, Nikola Momchev, Sertan Girgin, Raphaël Marinier, Lukasz Stafiniak, Manu Orsini, Olivier Bachem, Matthieu Geist, Olivier Pietquin |
ICML | 11 |
| 2021 | What Matters for Adversarial Imitation Learning?abstractAdversarial imitation learning has become a popular framework for imitation in continuous control. Over the years, several variations of its components were proposed to enhance the performance of the learned policies as well as the sample complexity of the algorithm. In practice, these choices are rarely tested all together in rigorous empirical studies.It is therefore difficult to discuss and understand what choices, among the high-level algorithmic options as well as low-level implementation details, matter. To tackle this issue, we implement more than 50 of these choices in a generic adversarial imitation learning frameworkand investigate their impacts in a large-scale study (>500k trained agents) with both synthetic and human-generated demonstrations. We analyze the key results and highlight the most surprising findings. Manu Orsini, Anton Raichuk, Léonard Hussenot, Damien Vincent, Robert Dadashi, Sertan Girgin, Matthieu Geist, Olivier Bachem, Olivier Pietquin, Marcin Andrychowicz |
NeurIPS | 1 |