Vincent Mai

dblp:229/0382 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0003-2823-504XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 85% Trustworthy machine learning · 15%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning
0.912025
Safety Representations for Safer Policy Learning · ICLR 2025
Machine learning › Reinforcement learning › safe reinforcement learning
safe exploration
0.912025
Safety Representations for Safer Policy Learning · ICLR 2025
Machine learning › Reinforcement learning
safe reinforcement learning
0.912025
Safety Representations for Safer Policy Learning · ICLR 2025
Machine learning › Reinforcement learning › sample efficiency
sample-efficient reinforcement learning
0.612022
Sample Efficient Deep Reinforcement Learning via Uncertainty Estimation · ICLR 2022
Machine learning › Trustworthy machine learning
uncertainty estimation
0.612022
Sample Efficient Deep Reinforcement Learning via Uncertainty Estimation · ICLR 2022
Machine learning › Reinforcement learning
deep reinforcement learning
0.212022
Sample Efficient Deep Reinforcement Learning via Uncertainty Estimation · ICLR 2022

Methods — techniques the papers use, named apart from their topics

state augmentation · 0.9uncertainty estimation · 0.6
YearPublicationVenuePosition
2025 Safety Representations for Safer Policy Learning
abstract
Reinforcement learning algorithms typically necessitate extensive exploration of the state space to find optimal policies. However, in safety-critical applications, the risks associated with such exploration can lead to catastrophic consequences. Existing safe exploration methods attempt to mitigate this by imposing constraints, which often result in overly conservative behaviours and inefficient learning. Heavy penalties for early constraint violations can trap agents in local optima, deterring exploration of risky yet high-reward regions of the state space. To address this, we introduce a method that explicitly learns state-conditioned safety representations. By augmenting the state features with these safety representations, our approach naturally encourages safer exploration without being excessively cautious, resulting in more efficient and safer policy learning in safety-critical scenarios. Empirical evaluations across diverse environments show that our method significantly improves task performance while reducing constraint violations during training, underscoring its effectiveness in balancing exploration with safety.
Kaustubh Mani, Vincent Mai, Charlie Gauthier, Annie S. Chen, Samer B. Nashed, Liam Paull
ICLR2
2024 Correction to: Multi-agent reinforcement learning for fast-timescale demand response of residential loads
Vincent Mai, Philippe Maisonneuve, Hadi Nekoei, Liam Paull, Antoine Lesage-Landry
Mach. Learn.1
2024 Multi-agent reinforcement learning for fast-timescale demand response of residential loads
Vincent Mai, Philippe Maisonneuve, Hadi Nekoei, Liam Paull, Antoine Lesage-Landry
Mach. Learn.1
2022 Sample Efficient Deep Reinforcement Learning via Uncertainty Estimation
Vincent Mai, Kaustubh Mani, Liam Paull
ICLR1