Tanmay Ambadkar

dblp:332/6387 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2026
0000-0003-0544-6864ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 60% Knowledge representation and reasoning · 20% Motion planning and robot control · 20%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
safe reinforcement learning
2.022026
Robust Adaptive Multi-Step Predictive Shielding (Student Abstract) · AAAI 2026
Specification-Guided Reinforcement Learning · AAAI 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning
automated reasoning and model checking
1.012026
Specification-Guided Reinforcement Learning · AAAI 2026
Robotics › Motion planning and robot control › robot control › safe control
control barrier functions
1.012026
Robust Adaptive Multi-Step Predictive Shielding (Student Abstract) · AAAI 2026
Machine learning › Reinforcement learning › safe reinforcement learning
shielding
1.012026
Robust Adaptive Multi-Step Predictive Shielding (Student Abstract) · AAAI 2026

Methods — techniques the papers use, named apart from their topics

temporal logic · 1.0learned dynamics model · 1.0formal methods · 1.0control barrier functions · 1.0
YearPublicationVenuePosition
2026 Specification-Guided Reinforcement Learning
abstract
While Reinforcement Learning (RL) has demonstrated remarkable success in solving complex sequential decision-making problems, its application in real-world, safety-critical systems is hindered by its reliance on carefully engineered reward functions. Designing effective rewards is notoriously challenging and can lead to unintended or unsafe behaviors, a phenomenon known as reward hacking. Specification-guided RL has emerged as a principled alternative, leveraging formal methods to directly encode high-level objectives, safety requirements, and behavioral constraints. However, the practical utility of this approach is often limited by coarse or under-specified logical formulas and the computational challenge of enforcing safety at scale. This thesis addresses these limitations by developing a unified framework for the automated refinement, scalable enforcement, and flexible adaptation of formal specifications in RL.
Tanmay Ambadkar
AAAI1
2026 Robust Adaptive Multi-Step Predictive Shielding (Student Abstract)
abstract
Ensuring safety in deep reinforcement learning is challenging, as formal methods that provide strong guarantees often fail to scale to complex, high-dimensional systems. We introduce RAMPS, a scalable shielding framework that pairs a general-purpose, learned linear dynamics model with a robust, multi-step Control Barrier Function (CBF) for real-time safety interventions. Experiments show RAMPS significantly reduces safety violations in high-dimensional environments compared to state-of-the-art methods, without sacrificing task performance.
Tanmay Ambadkar, Darshan Chudiwal, Greg Anderson 0003, Abhinav Verma 0001
AAAI1