Edwin Hamel-De le Court

dblp:249/2268 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
abstract
Many reinforcement learning algorithms, particularly those that rely on return estimates for policy improvement, can suffer from poor sample efficiency and training instability due to high-variance return estimates. In this paper we leverage new results from off-policy evaluation; it has recently been shown that well-designed behaviour policies can be used to collect off-policy data for provably lower variance return estimates. This result is surprising as it means collecting data on-policy is not variance optimal. We extend this key insight to the online reinforcement learning setting, where both policy evaluation and improvement are interleaved to learn optimal policies. Off-policy RL has been well studied (e.g., IMPALA), with correct and truncated importance weighted samples for de-biasing and managing variance appropriately. Generally these approaches are concerned with reconciling data collected from multiple workers in parallel, while the policy is updated asynchronously, mismatch between the workers and policy is corrected in a mathematically sound way. Here we consider only one worker - the behaviour policy, which is used to collect data for policy improvement, with provably lower variance return estimates. In our experiments we extend two policy-gradient methods with this regime, demonstrating better sample efficiency and performance over a diverse set of environments.
Alexander W. Goodall, Edwin Hamel-De le Court, Francesco Belardinelli
AAAI2
2026 SC²: Safe Control via Shielding for CPCTL Specifications
abstract
In real-world scenarios, reinforcement learning (RL) agents must not only maximize reward but also behave safely, including during training. This has led to growing interest in Safe RL, where the objective is to learn an optimal policy among those satisfying given safety constraints. Most existing approaches focus on constraints expressed either as expected costs or as avoidance properties. However, safety in dynamical systems is often expressed using rich temporal languages, such as Probabilistic Computation Tree Logic (PCTL). In this paper, we address the Safe RL problem under constraints expressed in CPCTL, a fragment of PCTL that generalizes avoidance constraints and enables the specification of complex, nested behaviors. To this end, we leverage Shielding, a technique that restricts the agent’s actions during both training and deployment to enforce safety over an infinite horizon. We first introduce a general framework based on an augmentation method and provide its theoretical foundations. Building on this framework, we propose an algorithm that is provably safe at all times, including during training, while remaining optimal among all safe policies. Finally, we present an experimental evaluation demonstrating the effectiveness of our approach.
Edwin Hamel-De le Court, Gaspard Ohlmann, Francesco Belardinelli
KR1
2025 Probabilistic Shielding for Safe Reinforcement Learning
abstract
In real-life scenarios, a Reinforcement Learning (RL) agent aiming to maximize their reward, must often also behave in a safe manner, including at training time. Thus, much attention in recent years has been given to Safe RL, where an agent aims to learn an optimal policy among all policies that satisfy a given safety constraint. However, strict safety guarantees are often provided through approaches based on linear programming, and thus have limited scaling. In this paper we present a new, scalable method, which enjoys strict formal guarantees for Safe RL, in the case where the safety dynamics of the Markov Decision Process (MDP) are known, and safety is defined as an undiscounted probabilistic avoidance property. Our approach is based on state-augmentation of the MDP, and on the design of a shield that restricts the actions available to the agent. We show that our approach provides a strict formal safety guarantee that the agent stays safe at training and test time. Furthermore, we demonstrate that our approach is viable in practice through experimental evaluation.
Edwin Hamel-De le Court, Francesco Belardinelli, Alexander W. Goodall
AAAI1
2022 Two-Player Boundedness Counter Games
Emmanuel Filiot, Edwin Hamel-De le Court
CONCUR2
2022 Combination of roots and boolean operations: An application to state complexity
Pascal Caron, Edwin Hamel-De le Court, Jean-Gabriel Luque
Inf. Comput.2
2020 A Study of a Simple Class of Modifiers: Product Modifiers
Pascal Caron, Edwin Hamel-De le Court, Jean-Gabriel Luque
DLT2