Prajwal Koirala

dblp:367/3704 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Reinforcement learning · 91% Generative modeling · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › safe reinforcement learning
constrained policy optimization
0.912025
Latent Safety-Constrained Policy Approach for Safe Offline Reinforcement Learning · ICLR 2025
Machine learning › Reinforcement learning
offline reinforcement learning
0.912025
Latent Safety-Constrained Policy Approach for Safe Offline Reinforcement Learning · ICLR 2025
Machine learning › Reinforcement learning › safe reinforcement learning
offline safe reinforcement learning
0.912025
Latent Safety-Constrained Policy Approach for Safe Offline Reinforcement Learning · ICLR 2025
Machine learning › Generative modeling › variational autoencoder
conditional variational autoencoder
0.312025
Latent Safety-Constrained Policy Approach for Safe Offline Reinforcement Learning · ICLR 2025

Methods — techniques the papers use, named apart from their topics

conditional variational autoencoder · 0.9advantage-weighted regression · 0.9
YearPublicationVenuePosition
2025 Latent Safety-Constrained Policy Approach for Safe Offline Reinforcement Learning
abstract
In safe offline reinforcement learning, the objective is to develop a policy that maximizes cumulative rewards while strictly adhering to safety constraints, utilizing only offline data. Traditional methods often face difficulties in balancing these constraints, leading to either diminished performance or increased safety risks. We address these issues with a novel approach that begins by learning a conservatively safe policy through the use of Conditional Variational Autoencoders, which model the latent safety constraints. Subsequently, we frame this as a Constrained Reward-Return Maximization problem, wherein the policy aims to optimize rewards while complying with the inferred latent safety constraints. This is achieved by training an encoder with a reward-Advantage Weighted Regression objective within the latent constraint space. Our methodology is supported by theoretical analysis, including bounds on policy performance and sample complexity. Extensive empirical evaluation on benchmark datasets, including challenging autonomous driving scenarios, demonstrates that our approach not only maintains safety compliance but also excels in cumulative reward optimization, surpassing existing methods. Additional visualizations provide further insights into the effectiveness and underlying mechanisms of our approach.
Prajwal Koirala, Zhanhong Jiang, Soumik Sarkar, Cody H. Fleming
ICLR1