Victor Villin

dblp:304/3021 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › reinforcement learning environment
environment design
0.812024
Environment Design for Inverse Reinforcement Learning · ICML 2024
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning
0.812024
Environment Design for Inverse Reinforcement Learning · ICML 2024
Machine learning › Reinforcement learning › reward learning
reward learning from demonstrations
0.812024
Environment Design for Inverse Reinforcement Learning · ICML 2024

Methods — techniques the papers use, named apart from their topics

adaptive environment selection · 0.8
YearPublicationVenuePosition
2025 A Minimax Approach to Ad Hoc Teamwork
Victor Villin, Thomas Kleine Buening, Christos Dimitrakakis
AAMAS1
2024 Environment Design for Inverse Reinforcement Learning
abstract
Learning a reward function from demonstrations suffers from low sample-efficiency. Even with abundant data, current inverse reinforcement learning methods that focus on learning from a single environment can fail to handle slight changes in the environment dynamics. We tackle these challenges through adaptive environment design. In our framework, the learner repeatedly interacts with the expert, with the former selecting environments to identify the reward function as quickly as possible from the expert’s demonstrations in said environments. This results in improvements in both sample-efficiency and robustness, as we show experimentally, for both exact and approximate inference.
Thomas Kleine Buening, Victor Villin, Christos Dimitrakakis
ICML2