VLDB 2026 Research / reviewers in the wild / expert
Victor Villin
dblp:304/3021
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Reinforcement learning · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › reinforcement learning environment
environment design |
0.8 | 1 | 2024 | Environment Design for Inverse Reinforcement Learning · ICML 2024 |
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
0.8 | 1 | 2024 | Environment Design for Inverse Reinforcement Learning · ICML 2024 |
Machine learning › Reinforcement learning › reward learning
reward learning from demonstrations |
0.8 | 1 | 2024 | Environment Design for Inverse Reinforcement Learning · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
adaptive environment selection · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Minimax Approach to Ad Hoc Teamwork
Victor Villin, Thomas Kleine Buening, Christos Dimitrakakis |
AAMAS | 1 |
| 2024 | Environment Design for Inverse Reinforcement LearningabstractLearning a reward function from demonstrations suffers from low sample-efficiency. Even with abundant data, current inverse reinforcement learning methods that focus on learning from a single environment can fail to handle slight changes in the environment dynamics. We tackle these challenges through adaptive environment design. In our framework, the learner repeatedly interacts with the expert, with the former selecting environments to identify the reward function as quickly as possible from the expert’s demonstrations in said environments. This results in improvements in both sample-efficiency and robustness, as we show experimentally, for both exact and approximate inference. Thomas Kleine Buening, Victor Villin, Christos Dimitrakakis |
ICML | 2 |