EDBT 2026 Demo / reviewers in the wild / expert
Charlie Gauthier
dblp:400/8152
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Reinforcement learning · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning |
0.9 | 1 | 2025 | Safety Representations for Safer Policy Learning · ICLR 2025 |
Machine learning › Reinforcement learning › safe reinforcement learning
safe exploration |
0.9 | 1 | 2025 | Safety Representations for Safer Policy Learning · ICLR 2025 |
Machine learning › Reinforcement learning
safe reinforcement learning |
0.9 | 1 | 2025 | Safety Representations for Safer Policy Learning · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
state augmentation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Safety Representations for Safer Policy LearningabstractReinforcement learning algorithms typically necessitate extensive exploration of the state space to find optimal policies. However, in safety-critical applications, the risks associated with such exploration can lead to catastrophic consequences. Existing safe exploration methods attempt to mitigate this by imposing constraints, which often result in overly conservative behaviours and inefficient learning. Heavy penalties for early constraint violations can trap agents in local optima, deterring exploration of risky yet high-reward regions of the state space. To address this, we introduce a method that explicitly learns state-conditioned safety representations. By augmenting the state features with these safety representations, our approach naturally encourages safer exploration without being excessively cautious, resulting in more efficient and safer policy learning in safety-critical scenarios. Empirical evaluations across diverse environments show that our method significantly improves task performance while reducing constraint violations during training, underscoring its effectiveness in balancing exploration with safety. Kaustubh Mani, Vincent Mai, Charlie Gauthier, Annie S. Chen, Samer B. Nashed, Liam Paull |
ICLR | 3 |
| 2025 | Perpetua: Multi-Hypothesis Persistence Modeling for Semi-Static EnvironmentsabstractMany robotic systems require extended deployments in complex, dynamic environments. In such deployments, parts of the environment may change between subsequent robot observations. Most robotic mapping or environment modeling algorithms are incapable of representing dynamic features in a way that enables predicting their future state. Instead, they opt to filter certain state observations, either by removing them or some form of weighted averaging. This paper introduces Perpetua, a method for modeling the dynamics of semi-static features. Perpetua is able to: incorporate prior knowledge about the dynamics of the feature if it exists, track multiple hypotheses, and adapt over time to enable predicting of future feature states. Specifically, we chain together mixtures of "persistence" and "emergence" filters to model the probability that features will disappear or reappear in a formal Bayesian framework. The approach is an efficient, scalable, general, and robust method for estimating the states of features in an environment, both in the present as well as at arbitrary future times. Through experiments on simulated and real-world data, we find that Perpetua yields better accuracy than similar approaches while also being online adaptable and robust to missing observations. Miguel A. Saavedra-Ruiz, Samer B. Nashed, Charlie Gauthier, Liam Paull |
IROS | 3 |