EDBT 2026 Demo / reviewers in the wild / expert
Abhinav Verma 0001
dblp:01/1084-1
· DBLP profile ↗
10ranked-venue papers
3as first author
3since 2021 · last 2026
0000-0002-9820-8285ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Reinforcement learning · 65% Motion planning and robot control · 18% Planning, search and constraint satisfaction · 7% | |
| Software engineering, system software, and programming languages
4 papers |
Program synthesis and code generation · 71% Software maintenance and evolution · 29% | |
| Theoretical computer science
3 papers |
Automata and formal languages · 100% |
Topics — the 23 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
safe reinforcement learning |
2.1 | 3 | 2026 | Robust Adaptive Multi-Step Predictive Shielding (Student Abstract) · AAAI 2026 Compositional Policy Learning in Stochastic Control Systems with Formal Guarantees · NeurIPS 2023 Neurosymbolic Reinforcement Learning with Formally Verified Exploration · NeurIPS 2020 |
Machine learning › Reinforcement learning › safe reinforcement learning
verifiable reinforcement learning |
1.0 | 2 | 2023 | Compositional Policy Learning in Stochastic Control Systems with Formal Guarantees · NeurIPS 2023 Verifiable and Interpretable Reinforcement Learning through Program Synthesis · AAAI 2019 |
Robotics › Motion planning and robot control › robot control › safe control
control barrier functions |
1.0 | 1 | 2026 | Robust Adaptive Multi-Step Predictive Shielding (Student Abstract) · AAAI 2026 |
Machine learning › Reinforcement learning › safe reinforcement learning
shielding |
1.0 | 1 | 2026 | Robust Adaptive Multi-Step Predictive Shielding (Student Abstract) · AAAI 2026 |
Machine learning › Reinforcement learning › policy learning › policy parameterization
programmatic policy |
0.7 | 2 | 2019 | Verifiable and Interpretable Reinforcement Learning through Program Synthesis · AAAI 2019 Programmatically Interpretable Reinforcement Learning · ICML 2018 |
Machine learning › Reinforcement learning › off-policy reinforcement learning
experience replay |
0.7 | 1 | 2023 | Eventual Discounting Temporal Logic Counterfactual Experience Replay · ICML 2023 |
Robotics › Motion planning and robot control › robot control › learning control
neural network policy |
0.7 | 1 | 2023 | Compositional Policy Learning in Stochastic Control Systems with Formal Guarantees · NeurIPS 2023 |
Machine learning › Reinforcement learning
off-policy reinforcement learning |
0.7 | 1 | 2023 | Eventual Discounting Temporal Logic Counterfactual Experience Replay · ICML 2023 |
Robotics › Motion planning and robot control
robot control |
0.7 | 1 | 2023 | Compositional Policy Learning in Stochastic Control Systems with Formal Guarantees · NeurIPS 2023 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › heuristic search › best-first search
a* search |
0.4 | 1 | 2020 | Learning Differentiable Programs with Admissible Neural Heuristics · NeurIPS 2020 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
heuristic search |
0.4 | 1 | 2020 | Learning Differentiable Programs with Admissible Neural Heuristics · NeurIPS 2020 |
Machine learning › Reinforcement learning
continuous control |
0.4 | 1 | 2019 | Control Regularization for Reduced Variance Reinforcement Learning · ICML 2019 |
Machine learning › Reinforcement learning
model-free reinforcement learning |
0.4 | 1 | 2019 | Control Regularization for Reduced Variance Reinforcement Learning · ICML 2019 |
Machine learning › Reinforcement learning › policy search
programmatic reinforcement learning |
0.4 | 1 | 2019 | Imitation-Projected Programmatic Reinforcement Learning · NeurIPS 2019 |
Machine learning › Optimization for machine learning
variance reduction |
0.4 | 1 | 2019 | Control Regularization for Reduced Variance Reinforcement Learning · ICML 2019 |
Program synthesis and code generation › controller synthesis
policy synthesis |
0.4 | 1 | 2019 | Verifiable and Interpretable Reinforcement Learning through Program Synthesis · AAAI 2019 |
Automata and formal languages
finite automata |
0.4 | 1 | 2019 | Representing Formal Languages: A Comparison Between Finite Automata and Recurrent Neural Networks · ICLR (Poster) 2019 |
Automata and formal languages
formal language representation |
0.4 | 1 | 2019 | Representing Formal Languages: A Comparison Between Finite Automata and Recurrent Neural Networks · ICLR (Poster) 2019 |
Automata and formal languages
recurrent neural networks |
0.4 | 1 | 2019 | Representing Formal Languages: A Comparison Between Finite Automata and Recurrent Neural Networks · ICLR (Poster) 2019 |
Machine learning › Trustworthy machine learning › interpretability
explainable reinforcement learning |
0.3 | 1 | 2018 | Programmatically Interpretable Reinforcement Learning · ICML 2018 |
Machine learning › Reinforcement learning
policy learning |
0.3 | 1 | 2018 | Programmatically Interpretable Reinforcement Learning · ICML 2018 |
Software maintenance and evolution
code search |
0.3 | 1 | 2018 | Programmatically Interpretable Reinforcement Learning · ICML 2018 |
Machine learning › Trustworthy machine learning
interpretability |
0.1 | 1 | 2019 | Verifiable and Interpretable Reinforcement Learning through Program Synthesis · AAAI 2019 |
Methods — techniques the papers use, named apart from their topics
mirror descent · 1.6symbolic verification · 1.4reach-avoid supermartingale · 1.3compositional policy learning · 1.3SpectRL · 1.3learned dynamics model · 1.0control barrier functions · 1.0linear temporal logic · 0.7eventual discounting · 0.7counterfactual reasoning · 0.7symbolic policy verification · 0.4neurosymbolic policy learning · 0.4neural heuristics · 0.4iterative deepening depth-first search · 0.4a* search · 0.4recurrent neural network · 0.4program synthesis · 0.4imitation learning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Adaptive Multi-Step Predictive Shielding (Student Abstract)abstractEnsuring safety in deep reinforcement learning is challenging, as formal methods that provide strong guarantees often fail to scale to complex, high-dimensional systems. We introduce RAMPS, a scalable shielding framework that pairs a general-purpose, learned linear dynamics model with a robust, multi-step Control Barrier Function (CBF) for real-time safety interventions. Experiments show RAMPS significantly reduces safety violations in high-dimensional environments compared to state-of-the-art methods, without sacrificing task performance. Tanmay Ambadkar, Darshan Chudiwal, Greg Anderson 0003, Abhinav Verma 0001 |
AAAI | 4 |
| 2023 | Eventual Discounting Temporal Logic Counterfactual Experience ReplayabstractLinear temporal logic (LTL) offers a simplified way of specifying tasks for policy optimization that may otherwise be difficult to describe with scalar reward functions. However, the standard RL framework can be too myopic to find maximally LTL satisfying policies. This paper makes two contributions. First, we develop a new value-function based proxy, using a technique we call eventual discounting, under which one can find policies that satisfy the LTL specification with highest achievable probability. Second, we develop a new experience replay method for generating off-policy data from on-policy rollouts via counterfactual reasoning on different ways of satisfying the LTL specification. Our experiments, conducted in both discrete and continuous state-action spaces, confirm the effectiveness of our counterfactual experience replay approach. Cameron Voloshin, Abhinav Verma 0001, Yisong Yue |
ICML | 2 |
| 2023 | Compositional Policy Learning in Stochastic Control Systems with Formal GuaranteesabstractReinforcement learning has shown promising results in learning neural network policies for complicated control tasks. However, the lack of formal guarantees about the behavior of such policies remains an impediment to their deployment. We propose a novel method for learning a composition of neural network policies in stochastic environments, along with a formal certificate which guarantees that a specification over the policy's behavior is satisfied with the desired probability. Unlike prior work on verifiable RL, our approach leverages the compositional nature of logical specifications provided in SpectRL, to learn over graphs of probabilistic reach-avoid specifications. The formal guarantees are provided by learning neural network policies together with reach-avoid supermartingales (RASM) for the graph’s sub-tasks and then composing them into a global policy. We also derive a tighter lower bound compared to previous work on the probability of reach-avoidance implied by a RASM, which is required to find a compositional policy with an acceptable probabilistic threshold for complex tasks with multiple edge policies. We implement a prototype of our approach and evaluate it on a Stochastic Nine Rooms environment. Dorde Zikelic, Mathias Lechner, Abhinav Verma 0001, Krishnendu Chatterjee, Thomas A. Henzinger |
NeurIPS | 3 |
| 2020 | Neurosymbolic Reinforcement Learning with Formally Verified ExplorationabstractWe present REVEL, a partially neural reinforcement learning (RL) framework for provably safe exploration in continuous state and action spaces. A key challenge for provably safe deep RL is that repeatedly verifying neural networks within a learning loop is computationally infeasible. We address this challenge using two policy classes: a general, neurosymbolic class with approximate gradients and a more restricted class of symbolic policies that allows efficient verification. Our learning algorithm is a mirror descent over policies: in each iteration, it safely lifts a symbolic policy into the neurosymbolic space, performs safe gradient updates to the resulting policy, and projects the updated policy into the safe symbolic subset, all without requiring explicit verification of neural networks. Our empirical results show that REVEL enforces safe exploration in many scenarios in which Constrained Policy Optimization does not, and that it can discover policies that outperform those learned through prior approaches to verified exploration. Greg Anderson 0003, Abhinav Verma 0001, Isil Dillig, Swarat Chaudhuri |
NeurIPS | 2 |
| 2020 | Learning Differentiable Programs with Admissible Neural HeuristicsabstractWe study the problem of learning differentiable functions expressed as programs in a domain-specific language. Such programmatic models can offer benefits such as composability and interpretability; however, learning them requires optimizing over a combinatorial space of program "architectures". We frame this optimization problem as a search in a weighted graph whose paths encode top-down derivations of program syntax. Our key innovation is to view various classes of neural networks as continuous relaxations over the space of programs, which can then be used to complete any partial program. All the parameters of this relaxed program can be trained end-to-end, and the resulting training loss is an approximately admissible heuristic that can guide the combinatorial search. We instantiate our approach on top of the A* and Iterative Deepening Depth-First Search algorithms and use these algorithms to learn programmatic classifiers in three sequence classification tasks. Our experiments show that the algorithms outperform state-of-the-art methods for program learning, and that they discover programmatic classifiers that yield natural interpretations and achieve competitive accuracy. Ameesh Shah, Eric Zhan, Jennifer J. Sun, Abhinav Verma 0001, Yisong Yue, Swarat Chaudhuri |
NeurIPS | 4 |
| 2019 | Verifiable and Interpretable Reinforcement Learning through Program SynthesisabstractWe study the problem of generating interpretable and verifiable policies for Reinforcement Learning (RL). Unlike the popular Deep Reinforcement Learning (DRL) paradigm, in which the policy is represented by a neural network, the aim of this work is to find policies that can be represented in highlevel programming languages. Such programmatic policies have several benefits, including being more easily interpreted than neural networks, and being amenable to verification by scalable symbolic methods. The generation methods for programmatic policies also provide a mechanism for systematically using domain knowledge for guiding the policy search. The interpretability and verifiability of these policies provides the opportunity to deploy RL based solutions in safety critical environments. This thesis draws on, and extends, work from both the machine learning and formal methods communities. Abhinav Verma 0001 |
AAAI | 1 |
| 2019 | Representing Formal Languages: A Comparison Between Finite Automata and Recurrent Neural Networks
Joshua J. Michalenko, Ameesh Shah, Abhinav Verma 0001, Richard G. Baraniuk, Swarat Chaudhuri, Ankit B. Patel |
ICLR (Poster) | 3 |
| 2019 | Control Regularization for Reduced Variance Reinforcement LearningabstractDealing with high variance is a significant challenge in model-free reinforcement learning (RL). Existing methods are unreliable, exhibiting high variance in performance from run to run using different initializations/seeds. Focusing on problems arising in continuous control, we propose a functional regularization approach to augmenting model-free RL. In particular, we regularize the behavior of the deep policy to be similar to a policy prior, i.e., we regularize in function space. We show that functional regularization yields a bias-variance trade-off, and propose an adaptive tuning strategy to optimize this trade-off. When the policy prior has control-theoretic stability guarantees, we further show that this regularization approximately preserves those stability guarantees throughout learning. We validate our approach empirically on a range of settings, and demonstrate significantly reduced variance, guaranteed dynamic stability, and more efficient learning than deep RL alone. Richard Cheng, Abhinav Verma 0001, Gábor Orosz, Swarat Chaudhuri, Yisong Yue, Joel W. Burdick |
ICML | 2 |
| 2019 | Imitation-Projected Programmatic Reinforcement LearningabstractWe study the problem of programmatic reinforcement learning, in which policies are represented as short programs in a symbolic language. Programmatic policies can be more interpretable, generalizable, and amenable to formal verification than neural policies; however, designing rigorous learning approaches for such policies remains a challenge. Our approach to this challenge - a meta-algorithm called PROPEL - is based on three insights. First, we view our learning task as optimization in policy space, modulo the constraint that the desired policy has a programmatic representation, and solve this optimization problem using a form of mirror descent that takes a gradient step into the unconstrained policy space and then projects back onto the constrained space. Second, we view the unconstrained policy space as mixing neural and programmatic representations, which enables employing state-of-the-art deep policy gradient approaches. Third, we cast the projection step as program synthesis via imitation learning, and exploit contemporary combinatorial methods for this task. We present theoretical convergence results for PROPEL and empirically evaluate the approach in three continuous control domains. The experiments show that PROPEL can significantly outperform state-of-the-art approaches for learning programmatic policies. Abhinav Verma 0001, Hoang Minh Le 0002, Yisong Yue, Swarat Chaudhuri |
NeurIPS | 1 |
| 2018 | Programmatically Interpretable Reinforcement LearningabstractWe present a reinforcement learning framework, called Programmatically Interpretable Reinforcement Learning (PIRL), that is designed to generate interpretable and verifiable agent policies. Unlike the popular Deep Reinforcement Learning (DRL) paradigm, which represents policies by neural networks, PIRL represents policies using a high-level, domain-specific programming language. Such programmatic policies have the benefits of being more easily interpreted than neural networks, and being amenable to verification by symbolic methods. We propose a new method, called Neurally Directed Program Search (NDPS), for solving the challenging nonsmooth optimization problem of finding a programmatic policy with maximal reward. NDPS works by first learning a neural policy network using DRL, and then performing a local search over programmatic policies that seeks to minimize a distance from this neural “oracle”. We evaluate NDPS on the task of learning to drive a simulated car in the TORCS car-racing environment. We demonstrate that NDPS is able to discover human-readable policies that pass some significant performance bars. We also show that PIRL policies can have smoother trajectories, and can be more easily transferred to environments not encountered during training, than corresponding policies discovered by DRL. Abhinav Verma 0001, Vijayaraghavan Murali, Rishabh Singh, Pushmeet Kohli, Swarat Chaudhuri |
ICML | 1 |