Erwan Lecarpentier

dblp:203/9218 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 59% Planning, search and constraint satisfaction · 32% Motion planning and robot control · 10%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › non-stationary reinforcement learning
continual reinforcement learning
0.512021
Lipschitz Lifelong Reinforcement Learning · AAAI 2021
Machine learning › Reinforcement learning
transfer learning in reinforcement learning
0.512021
Lipschitz Lifelong Reinforcement Learning · AAAI 2021
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
minimax search
0.412019
Non-Stationary Markov Decision Processes, a Worst-Case Approach using Model-Based Reinforcement Learning · NeurIPS 2019
Machine learning › Reinforcement learning
model-based reinforcement learning
0.412019
Non-Stationary Markov Decision Processes, a Worst-Case Approach using Model-Based Reinforcement Learning · NeurIPS 2019
Machine learning › Reinforcement learning › markov decision process
non-stationary markov decision process
0.412019
Non-Stationary Markov Decision Processes, a Worst-Case Approach using Model-Based Reinforcement Learning · NeurIPS 2019
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
robust planning
0.412019
Non-Stationary Markov Decision Processes, a Worst-Case Approach using Model-Based Reinforcement Learning · NeurIPS 2019
Robotics › Motion planning and robot control › robot control
open-loop control
0.312018
Open Loop Execution of Tree-Search Algorithms · IJCAI 2018
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
search-based planning
0.312018
Open Loop Execution of Tree-Search Algorithms · IJCAI 2018
Machine learning › Reinforcement learning › sample efficiency
PAC-MDP
0.112021
Lipschitz Lifelong Reinforcement Learning · AAAI 2021
Machine learning › Reinforcement learning
markov decision process
0.112019
Non-Stationary Markov Decision Processes, a Worst-Case Approach using Model-Based Reinforcement Learning · NeurIPS 2019

Methods — techniques the papers use, named apart from their topics

lipschitz continuity · 0.9value transfer · 0.5risk-averse tree search · 0.4minimax search · 0.4stochastic planning · 0.3monte carlo tree search · 0.3
YearPublicationVenuePosition
2022 LUCIE: an evaluation and selection method for stochastic problems
abstract
Selection in genetic algorithms is difficult for stochastic problems due to noise in the fitness space. Common methods to deal with this fitness noise include sampling multiple fitness values, which can be expensive. We propose LUCIE, the Lower Upper Confidence Intervals Elitism method, which selects individuals based on confidence. By focusing evaluation on separating promising individuals from others, we demonstrate that LUCIE can be effectively used as an elitism mechanism in genetic algorithms. We provide a theoretical analysis on the convergence of LUCIE and demonstrate its ability to select fit individuals across multiple types of noise on the OneMax and LeadingOnes problems. We also evaluate LUCIE as a selection method for neuroevolution on control policies with stochastic fitness values.
Erwan Lecarpentier, Paul Templier, Emmanuel Rachelson, Dennis Wilson
GECCO1
2021 Lipschitz Lifelong Reinforcement Learning
abstract
We consider the problem of knowledge transfer when an agent is facing a series of Reinforcement Learning (RL) tasks. We introduce a novel metric between Markov Decision Processes and establish that close MDPs have close optimal value functions. Formally, the optimal value functions are Lipschitz continuous with respect to the tasks space. These theoretical results lead us to a value-transfer method for Lifelong RL, which we use to build a PAC-MDP algorithm with improved convergence rate. Further, we show the method to experience no negative transfer with high probability. We illustrate the benefits of the method in Lifelong RL experiments.
Erwan Lecarpentier, David Abel, Kavosh Asadi, Yuu Jinnai, Emmanuel Rachelson, Michael L. Littman
AAAI1
2019 Non-Stationary Markov Decision Processes, a Worst-Case Approach using Model-Based Reinforcement Learning
abstract
This work tackles the problem of robust zero-shot planning in non-stationary stochastic environments. We study Markov Decision Processes (MDPs) evolving over time and consider Model-Based Reinforcement Learning algorithms in this setting. We make two hypotheses: 1) the environment evolves continuously with a bounded evolution rate; 2) a current model is known at each decision epoch but not its evolution. Our contribution can be presented in four points. 1) we define a specific class of MDPs that we call Non-Stationary MDPs (NSMDPs). We introduce the notion of regular evolution by making an hypothesis of Lipschitz-Continuity on the transition and reward functions w.r.t. time; 2) we consider a planning agent using the current model of the environment but unaware of its future evolution. This leads us to consider a worst-case method where the environment is seen as an adversarial agent; 3) following this approach, we propose the Risk-Averse Tree-Search (RATS) algorithm, a zero-shot Model-Based method similar to Minimax search; 4) we illustrate the benefits brought by RATS empirically and compare its performance with reference Model-Based algorithms.
Erwan Lecarpentier, Emmanuel Rachelson
NeurIPS1
2018 Open Loop Execution of Tree-Search Algorithms
abstract
In the context of tree-search stochastic planning algorithms where a generative model is available, we consider on-line planning algorithms building trees in order to recommend an action. We investigate the question of avoiding re-planning in subsequent decision steps by directly using sub-trees as action recommender. Firstly, we propose a method for open loop control via a new algorithm taking the decision of re-planning or not at each time step based on an analysis of the statistics of the sub-tree. Secondly, we show that the probability of selecting a suboptimal action at any depth of the tree can be upper bounded and converges towards zero. Moreover, this upper bound decays in a logarithmic way between subsequent depths. This leads to a distinction between node-wise optimality and state-wise optimality. Finally, we empirically demonstrate that our method achieves a compromise between loss of performance and computational gain.
Erwan Lecarpentier, Guillaume Infantes, Charles Lesire, Emmanuel Rachelson
IJCAI1