VLDB 2026 Research / reviewers in the wild / expert
Mohit Prashant
dblp:343/4154
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 56% Trustworthy machine learning · 34% Generative modeling · 10% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
safe reinforcement learning |
1.9 | 2 | 2026 | CAPO: A Unified Policy Gradient Approach for Reward and Cost Optimization in Safe Reinforcement Learning (Student Abstract) · AAAI 2026 Guaranteeing Out-Of-Distribution Detection in Deep RL via Transition Estimation · AAAI 2025 |
Machine learning › Reinforcement learning › policy optimization
policy gradient |
1.0 | 1 | 2026 | CAPO: A Unified Policy Gradient Approach for Reward and Cost Optimization in Safe Reinforcement Learning (Student Abstract) · AAAI 2026 |
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection |
0.9 | 1 | 2025 | Guaranteeing Out-Of-Distribution Detection in Deep RL via Transition Estimation · AAAI 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 1 | 2025 | Guaranteeing Out-Of-Distribution Detection in Deep RL via Transition Estimation · AAAI 2025 |
Machine learning › Generative modeling › variational autoencoder
conditional variational autoencoder |
0.3 | 1 | 2025 | Guaranteeing Out-Of-Distribution Detection in Deep RL via Transition Estimation · AAAI 2025 |
Machine learning › Generative modeling
variational autoencoder |
0.3 | 1 | 2025 | Guaranteeing Out-Of-Distribution Detection in Deep RL via Transition Estimation · AAAI 2025 |
Methods — techniques the papers use, named apart from their topics
second-order taylor approximation · 1.0reward augmentation · 1.0cost shaping · 1.0reconstruction loss · 0.9conditional variational autoencoder · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CAPO: A Unified Policy Gradient Approach for Reward and Cost Optimization in Safe Reinforcement Learning (Student Abstract)abstractIn safe reinforcement learning (SRL), there exists an inherent conflict between maximizing reward and minimizing cost. We propose a novel approach that effectively resolve the conflict between maximizing reward and minimizing cost in joint optimization.When the cost exceeds the threshold, we perform cost-reducing updates. Otherwise, we compute policy gradients that maximize expected rewards, while using second-order Taylor approximation to evaluate whether these reward-maximizing gradients would violate the cost constraint. If constraint violation is detected, we adjust the gradient direction to maintain safety compliance; otherwise, we execute standard reward-increasing policy updates. This approach helps ensure that reward-seeking updates do not inadvertently increase costs, thereby reducing the likelihood of constraint violations. Empirical tests show our framework successfully manages reward-cost trade-offs through reward augmentation and cost shaping, improving both performance and safety without switching optimization strategies. Results demonstrate that concurrent treatment of both objectives in one policy gradient update is viable for improving safe reinforcement learning methods. Mohit Prashant, Arvind Easwaran |
AAAI | 2 |
| 2025 | Guaranteeing Out-Of-Distribution Detection in Deep RL via Transition EstimationabstractAn issue concerning the use of deep reinforcement learning (RL) agents is whether they can be trusted to perform reliably when deployed, as training environments may not reflect real-life environments. Anticipating instances outside their training scope, learning-enabled systems are often equipped with out-of-distribution (OOD) detectors that alert when a trained system encounters a state it does not recognize or in which it exhibits uncertainty. There exists limited work conducted on the problem of OOD detection within RL, with prior studies being unable to achieve a consensus on the definition of OOD execution within the context of RL. By framing our problem using a Markov Decision Process, we assume there is a transition distribution mapping each state-action pair to another state with some probability. Based on this, we consider the following definition of OOD execution within RL: A transition is OOD if its probability during real-life deployment differs from the transition distribution encountered during training. As such, we utilize conditional variational autoencoders (CVAE) to approximate the transition dynamics of the training environment and implement a conformity-based detector using reconstruction loss that is able to guarantee OOD detection with a pre-determined confidence level. We evaluate our detector by adapting existing benchmarks and compare it with existing OOD detection models for RL. Mohit Prashant, Arvind Easwaran, Michael Yuhas |
AAAI | 1 |
| 2025 | Improving Reinforcement Learning Sample-Efficiency Using Local ApproximationabstractIn this study, we derive Probably Approximately Correct (PAC) bounds on the asymptotic sample-complexity for RL within the infinite-horizon Markov Decision Process (MDP) setting that are sharper than those in existing literature. The aim of PAC learning is to converge to a near-optimal value function with guarantees on the ‘nearness’, i.e. ϵ, of the synthesized solution. With this, the premise of our study is twofold: firstly, the further two states are from each other, transition-wise, the less relevant the value of the first state is when learning the ϵ-optimal value of the second; secondly, the amount of ‘effort’, sample-complexity-wise, expended in learning the ϵ-optimal value of a state is independent of the number of samples required to learn the ϵ-optimal value of a second state that is a sufficient number of transitions away from the first. Inversely, states within each other’s vicinity have values that are dependent on each other and will require a similar number of samples to learn. By approximating the original MDP using smaller MDPs constructed using subsets of the original state-space, we are able to reduce the sample-complexity by a logarithmic factor to O(SA log A) timesteps, where S and A are the state and action space sizes. We are able to extend these results to an infinite-horizon, model-free setting by constructing a PAC-MDP algorithm with the aforementioned sample-complexity. Mohit Prashant, Arvind Easwaran |
ECAI | 1 |