Praveen Paruchuri

dblp:22/1006 · DBLP profile ↗
← Back
22ranked-venue papers
1as first author
12since 2021 · last 2025
0000-0001-8071-5409ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 7 since 2021
YearPublicationVenuePosition
2025 M-SAT: Multi-State-Action Tokenisation in Decision Transformers for Multi-Discrete Actions
abstract
Effective decision-making in complex environments with multi-discrete action spaces poses significant challenges for agent architectures, particularly in image-based settings. While Decision Transformers have shown promise in various domains, their performance often suffers in environments where agents must handle multi-discrete actions. Existing enhancements to Decision Transformer architectures have yet to address this critical issue, limiting their ability to support agents in learning robust policies in these environments. To address this gap, we propose Multi-State Action Tokenisation (M-SAT), a novel approach designed to improve agent decision-making by tokenising actions at the individual action level and incorporating auxiliary state information. This disentanglement of actions improves both the performance of agents and the interpretability of individual actions within attention layers, fostering better visibility into agent decision processes. Importantly, M-SAT facilitates the development of more interpretable and transparent agents capable of making complex decisions in dynamic environments involving multi-discrete action spaces. We evaluate M-SAT on the challenging ViZDoom environments, focusing on scenarios with multi-discrete action spaces and image-based observations, such as Deadly Corridor, My Way Home and Death Match. Our approach demonstrates superior performance compared to baseline Decision Transformers, with no additional data or significant computational overheads. Furthermore, we observe that M-SAT does not require positional encoding to achieve high performance, with its removal occasionally leading to further improvements. These findings suggest that M-SAT enables more efficient and interpretable agent-based decision-making in multi-discrete action spaces.
Perusha Moodley, Dhillu Thambi, Mark Trovinger, Pramod Kaushik, Praveen Paruchuri, Xia Hong 0001, Benjamin Rosman
IJCNN5
2024 Preserving the Privacy of Reward Functions in MDPs through Deception
abstract
Preserving the privacy of preferences (or rewards) of a sequential decision-making agent when decisions are observable is crucial in many physical and cybersecurity domains. For instance, in wildlife monitoring, forest rangers must conduct surveillance without revealing animal locations to poachers. This paper addresses privacy preservation in planning over a sequence of actions in MDPs, where the reward function represents the preference structure to be protected. Observers can use Inverse RL (IRL) to learn these preferences, making this a challenging task. Current research on Differential Privacy (DP) in this setting fails to ensure a lower bound on the minimum expected reward and offers theoretical guarantees that are inadequate against IRL-based observers. To bridge this gap, we propose a novel approach rooted in the theory of deception. Deception includes two models: dissimulation (hiding the truth) and simulation (showing the wrong). As our first contribution, we theoretically demonstrate a significant privacy leak in the current dissimulation-based method. Our second contribution is a novel RL-based planning algorithm that uses simulation to effectively address these privacy concerns while ensuring a guarantee on the expected reward. Through experimentation on multiple benchmark problems, we show that our proposed approach outperforms existing methods in preserving the privacy of reward functions. Code to reproduce the results can be found at: https://github.com/shshnkreddy/DeceptiveRL
Shashank Reddy Chirra, Pradeep Varakantham, Praveen Paruchuri
ECAI3
2024 Interpreting Decision Transformer: Insights from Continuous Control Tasks
Dhillu Thambi, Praveen Paruchuri, Perusha Moodley
ICONIP (1)2
2024 Safety through feedback in Constrained RL
abstract
In safety-critical RL settings, the inclusion of an additional cost function is often favoured over the arduous task of modifying the reward function to ensure the agent's safe behaviour. However, designing or evaluating such a cost function can be prohibitively expensive. For instance, in the domain of self-driving, designing a cost function that encompasses all unsafe behaviours (e.g., aggressive lane changes, risky overtakes) is inherently complex, it must also consider all the actors present in the scene making it expensive to evaluate. In such scenarios, the cost function can be learned from feedback collected offline in between training rounds. This feedback can be system generated or elicited from a human observing the training process. Previous approaches have not been able to scale to complex environments and are constrained to receiving feedback at the state level which can be expensive to collect. To this end, we introduce an approach that scales to more complex domains and extends beyond state-level feedback, thus, reducing the burden on the evaluator. Inferring the cost function in such settings poses challenges, particularly in assigning credit to individual states based on trajectory-level feedback. To address this, we propose a surrogate objective that transforms the problem into a state-level supervised classification task with noisy labels, which can be solved efficiently. Additionally, it is often infeasible to collect feedback for every trajectory generated by the agent, hence, two fundamental questions arise: (1) Which trajectories should be presented to the human? and (2) How many trajectories are necessary for effective learning? To address these questions, we introduce a \textit{novelty-based sampling} mechanism that selectively involves the evaluator only when the the agent encounters a \textit{novel} trajectory, and discontinues querying once the trajectories are no longer \textit{novel}. We showcase the efficiency of our method through experimentation on several benchmark Safety Gymnasium environments and realistic self-driving scenarios. Our method demonstrates near-optimal performance, comparable to when the cost function is known, by relying solely on trajectory-level feedback across multiple domains. This highlights both the effectiveness and scalability of our approach. The code to replicate these results can be found at \href{https://github.com/shshnkreddy/RLSF}{https://github.com/shshnkreddy/RLSF}
Shashank Reddy Chirra, Pradeep Varakantham, Praveen Paruchuri
NeurIPS3
2024 Improving Lane Level Dynamics for EV Traversal: A Reinforcement Learning Approach
Akanksha Tyagi, Meghna Lowalekar, Praveen Paruchuri
VEHITS3
2023 Planning and Learning for Non-markovian Negative Side Effects Using Finite State Controllers
abstract
Autonomous systems are often deployed in the open world where it is hard to obtain complete specifications of objectives and constraints. Operating based on an incomplete model can produce negative side effects (NSEs), which affect the safety and reliability of the system. We focus on mitigating NSEs in environments modeled as Markov decision processes (MDPs). First, we learn a model of NSEs using observed data that contains state-action trajectories and severity of associated NSEs. Unlike previous works that associate NSEs with state-action pairs, our framework associates NSEs with entire trajectories, which is more general and captures non-Markovian dependence on states and actions. Second, we learn finite state controllers (FSCs) that predict NSE severity for a given trajectory and generalize well to unseen data. Finally, we develop a constrained MDP model that uses information from the underlying MDP and the learned FSC for planning while avoiding NSEs. Our empirical evaluation demonstrates the effectiveness of our approach in learning and mitigating Markovian and non-Markovian NSEs.
Aishwarya Srivastava, Sandhya Saisubramanian, Praveen Paruchuri, Akshat Kumar, Shlomo Zilberstein
AAAI3
2023 City-Scale Pollution Aware Traffic Routing by Sampling Max Flows Using MCMC
abstract
A significant cause of air pollution in urban areas worldwide is the high volume of road traffic. Long-term exposure to severe pollution can cause serious health issues. One approach towards tackling this problem is to design a pollution-aware traffic routing policy that balances multiple objectives of i) avoiding extreme pollution in any area ii) enabling short transit times, and iii) making effective use of the road capacities. We propose a novel sampling-based approach for this problem. We give the first construction of a Markov Chain that can sample integer max flow solutions of a planar graph, with theoretical guarantees that the probabilities depend on the aggregate transit length. We designed a traffic policy using diverse samples and simulated traffic on real-world road maps using the SUMO traffic simulator. We observe a considerable decrease in areas with severe pollution when experimented with maps of large cities across the world compared to other approaches.
Shreevignesh Suriyanarayanan, Praveen Paruchuri, Girish Varma
AAAI2
2022 CB+NN Ensemble to Improve Tracking Accuracy in Air Surveillance
abstract
Finding or tracking the location of an object accurately is a crucial problem in defense applications, robotics and computer vision. Radars fall into the spectrum of high-end defense sensors or systems upon which the security and surveillance of the entire world depends. There has been a lot of focus on the topic of Multi Sensor Tracking in recent years, with radars as the sensors. The Indian Air Force uses a Multi Sensor Tracking (MST) system to detect flights pan India, developed and supported by BEL(Bharat Electronics Limited), a defense agency we are working with. In this paper, we describe our Machine Learning approach, which is built on top of the existing system, the Air force uses. For purposes of this work, we trained our models on about 13 million anonymized real Multi Sensor tracking data points provided by radars performing tracking activity across the Indian air space. The approach has shown an increase in the accuracy of tracking by 5 percent from 91 to 96. The model and the corresponding code were transitioned to BEL, which has been tested in their simulation environment with a plan to take forward for ground testing. Our approach comprises of 3 steps: (a) We train a Neural Network model and a CatBoost model and ensemble them using a Logistic Regression model to predict one type of error, namely Splitting error, which can help to improve the accuracy of tracking. (b) We again train a Neural Network model and a CatBoost model and ensemble them using a different Logistic Regression model to predict the second type of error, namely Merging error, which can further improve the accuracy of tracking. (c) We use cosine similarity to find the nearest neighbour and correct the data points, predicted to have Splitting/Merging errors, by predicting the original global track of these data points.
Anoop Karnik Dasika, Praveen Paruchuri
AAAI2
2022 How Private Is Your RL Policy? An Inverse RL Based Analysis Framework
abstract
Reinforcement Learning (RL) enables agents to learn how to perform various tasks from scratch. In domains like autonomous driving, recommendation systems, and more, optimal RL policies learned could cause a privacy breach if the policies memorize any part of the private reward. We study the set of existing differentially-private RL policies derived from various RL algorithms such as Value Iteration, Deep-Q Networks, and Vanilla Proximal Policy Optimization. We propose a new Privacy-Aware Inverse RL analysis framework (PRIL) that involves performing reward reconstruction as an adversarial attack on private policies that the agents may deploy. For this, we introduce the reward reconstruction attack, wherein we seek to reconstruct the original reward from a privacy-preserving policy using the Inverse RL algorithm. An adversary must do poorly at reconstructing the original reward function if the agent uses a tightly private policy. Using this framework, we empirically test the effectiveness of the privacy guarantee offered by the private algorithms on instances of the FrozenLake domain of varying complexities. Based on the analysis performed, we infer a gap between the current standard of privacy offered and the standard of privacy needed to protect reward functions in RL. We do so by quantifying the extent to which each private policy protects the reward function by measuring distances between the original and reconstructed rewards.
Kritika Prakash, Fiza Husain, Praveen Paruchuri, Sujit Gujar
AAAI3
2022 VidyutVanika21: An Autonomous Intelligent Broker for Smart-grids
abstract
An autonomous broker that liaises between retail customers and power-generating companies (GenCos) is essential for the smart grid ecosystem. The efficiency brought in by such brokers to the smart grid setup can be studied through a well-developed simulation environment. In this paper, we describe the design of one such energy broker called VidyutVanika21 (VV21) and analyze its performance using a simulation platform called PowerTAC (PowerTrading Agent Competition). Specifically, we discuss the retail (VV21–RM) and wholesale market (VV21–WM) modules of VV21 that help the broker achieve high net profits in a competitive setup. Supported by game-theoretic analysis, the VV21–RM designs tariff contracts that a) maintain a balanced portfolio of different types of customers; b) sustain an appropriate level of market share, and c) introduce surcharges on customers to reduce energy usage during peak demand times. The VV21–WM aims to reduce the cost of procurement by following the supply curve of the GenCo to identify its lowest ask for a particular auction which is then used to generate suitable bids. We further demonstrate the efficacy of the retail and wholesale strategies of VV21 in PowerTAC 2021 finals and through several controlled experiments.
Sanjay Chandlekar, Bala Suraj Pedasingu, Easwar Subramanian, Sanjay P. Bhat, Praveen Paruchuri, Sujit Gujar
IJCAI5
2021 An Enhanced Advising Model in Teacher-Student Framework using State Categorization
abstract
The teacher-student framework aims to improve the sample efficiency of RL algorithms by deploying an advising mechanism in which a teacher helps a student by guiding its exploration. Prior work in this field has considered an advising mechanism where the teacher advises the student about the optimal action to take in a given state. However, real-world teachers can leverage domain expertise to provide more informative signals. Using this insight, we propose to extend the current advising framework wherein the teacher would provide not only the optimal action but also a qualitative assessment of the state. We introduce a novel architecture, namely Advice Replay Memory (ARM), to effectively reuse the advice provided by the teacher. We demonstrate the robustness of our approach by showcasing our experiments on multiple Atari 2600 games using a fixed set of hyper-parameters. Additionally, we show that a student taking help even from a sub-optimal teacher can achieve significant performance boosts and eventually outperform the teacher. Our approach outperforms the baselines even when provided with comparatively suboptimal teachers and an advising budget, which is smaller by orders of magnitude. The contributions of our paper are 4-fold (a) effectively leveraging a teacher's knowledge by richer advising (b) introduction of ARM to effectively reuse the advice throughout learning (c) ability to achieve significant performance boost even with a coarse state categorization (d) enabling the student to outperform the teacher.
Daksh Anand, Vaibhav Gupta, Praveen Paruchuri, Balaraman Ravindran
AAAI3
2021 A Genetic Algorithm Approach to Compute Mixed Strategy Solutions for General Stackelberg Games
abstract
Stackelberg games have found a role in a number of applications including modeling market competition, identifying traffic equilibrium, developing practical security applications and many others. While a number of solution approaches have been developed for these games in a variety of contexts that use mathematical optimization, analytical analysis or heuristic based solutions, literature has been quite sparse on the usage of Genetic Algorithm (GA) based techniques. In this paper, we develop a GA based solution to compute high quality mixed strategy solution for the leader to commit to in a General Stackelberg Game (GSG) using a normal form game formulation. The leader faces multiple types of followers with discrete utility functions where the mixed strategy of the leader (but not the actual action taken in the round) is known to the follower. Our experiments showcase that the GA developed here performs well in terms of scalability and provides reasonably good solution quality in terms of the average reward obtained. Given that finding the optimal mixed strategy solution for GSGs is NP-hard (and the optimal solution for leader lies in the mixed strategy space), we believe that the solution approach presented here can support further development of practical applications using GSGs.
Srivathsa Gottipati, Praveen Paruchuri
CEC2
2020 Bidding in Smart Grid PDAs: Theory, Analysis and Strategy
abstract
Periodic Double Auctions (PDAs) are commonly used in the real world for trading, e.g. in stock markets to determine stock opening prices, and energy markets to trade energy in order to balance net demand in smart grids, involving trillions of dollars in the process. A bidder, participating in such PDAs, has to plan for bids in the current auction as well as for the future auctions, which highlights the necessity of good bidding strategies. In this paper, we perform an equilibrium analysis of single unit single-shot double auctions with a certain clearing price and payment rule, which we refer to as ACPR, and find it intractable to analyze as number of participating agents increase. We further derive the best response for a bidder with complete information in a single-shot double auction with ACPR. Leveraging the theory developed for single-shot double auction and taking the PowerTAC wholesale market PDA as our testbed, we proceed by modeling the PDA of PowerTAC as an MDP. We propose a novel bidding strategy, namely MDPLCPBS. We empirically show that MDPLCPBS follows the equilibrium strategy for double auctions that we previously analyze. In addition, we benchmark our strategy against the baseline and the state-of-the-art bidding strategies for the PowerTAC wholesale market PDAs, and show that MDPLCPBS outperforms most of them consistently.
Susobhan Ghosh, Sujit Gujar, Praveen Paruchuri, Easwar Subramanian, Sanjay P. Bhat
AAAI3
2020 An Ensemble Learning Approach to Improve Tracking Accuracy of Multi Sensor Fusion
Anoop Karnik Dasika, Praveen Paruchuri
ICONIP (5)2
2020 Towards a Better Management of Emergency Evacuation using Pareto Min Cost Max Flow Approach
Sreeja Kamishetty, Praveen Paruchuri
VEHITS2
2019 Successor Features Based Multi-Agent RL for Event-Based Decentralized MDPs
abstract
Decentralized MDPs (Dec-MDPs) provide a rigorous framework for collaborative multi-agent sequential decisionmaking under uncertainty. However, their computational complexity limits the practical impact. To address this, we focus on a class of Dec-MDPs consisting of independent collaborating agents that are tied together through a global reward function that depends upon their entire histories of states and actions to accomplish joint tasks. To overcome scalability barrier, our main contributions are: (a) We propose a new actor-critic based Reinforcement Learning (RL) approach for event-based Dec-MDPs using successor features (SF) which is a value function representation that decouples the dynamics of the environment from the rewards; (b) We then present Dec-ESR (Decentralized Event based Successor Representation) which generalizes learning for event-based Dec-MDPs using SF within an end-to-end deep RL framework; (c) We also show that Dec-ESR allows useful transfer of information on related but different tasks, hence bootstraps the learning for faster convergence on new tasks; (d) For validation purposes, we test our approach on a large multi-agent coverage problem which models schedule coordination of agents in a real urban subway network and achieves better quality solutions than previous best approaches.
Tarun Gupta 0002, Akshat Kumar, Praveen Paruchuri
AAAI3
2019 VidyutVanika: A Reinforcement Learning Based Broker Agent for a Power Trading Competition
abstract
A smart grid is an efficient and sustainable energy system that integrates diverse generation entities, distributed storage capacity, and smart appliances and buildings. A smart grid brings new kinds of participants in the energy market served by it, whose effect on the grid can only be determined through high fidelity simulations. Power TAC offers one such simulation platform using real-world weather data and complex state-of-the-art customer models. In Power TAC, autonomous energy brokers compete to make profits across tariff, wholesale and balancing markets while maintaining the stability of the grid. In this paper, we design an autonomous broker VidyutVanika, the runner-up in the 2018 Power TAC competition. VidyutVanika relies on reinforcement learning (RL) in the tariff market and dynamic programming in the wholesale market to solve modified versions of known Markov Decision Process (MDP) formulations in the respective markets. The novelty lies in defining the reward functions for MDPs, solving these MDPs, and the application of these solutions to real actions in the market. Unlike previous participating agents, VidyutVanika uses a neural network to predict the energy consumption of various customers using weather data. We use several heuristic ideas to bridge the gap between the restricted action spaces of the MDPs and the much more extensive action space available to VidyutVanika. These heuristics allow VidyutVanika to convert near-optimal fixed tariffs to time-of-use tariffs aimed at mitigating transmission capacity fees, spread out its orders across several auctions in the wholesale market to procure energy at a lower price, more accurately estimate parameters required for implementing the MDP solution in the wholesale market, and account for wholesale procurement costs while optimizing tariffs. We use Power TAC 2018 tournament data and controlled experiments to analyze the performance of VidyutVanika, and illustrate the efficacy of the above strategies.
Susobhan Ghosh, Easwar Subramanian, Sanjay P. Bhat, Sujit Gujar, Praveen Paruchuri
AAAI5
2018 Planning and Learning for Decentralized MDPs With Event Driven Rewards
abstract
Decentralized (PO)MDPs provide a rigorous framework for sequential multiagent decision making under uncertainty. However, their high computational complexity limits the practical impact. To address scalability and real-world impact, we focus on settings where a large number of agents primarily interact through complex joint-rewards that depend on their entire histories of states and actions. Such history-based rewards encapsulate the notion of events or tasks such that the team reward is given only when the joint-task is completed. Algorithmically, we contribute---1) A nonlinear programming (NLP) formulation for such event-based planning model; 2) A probabilistic inference based approach that scales much better than NLP solvers for a large number of agents; 3) A policy gradient based multiagent reinforcement learning approach that scales well even for exponential state-spaces. Our inference and RL-based advances enable us to solve a large real-world multiagent coverage problem modeling schedule coordination of agents in a real urban subway network where other approaches fail to scale.
Tarun Gupta 0002, Akshat Kumar, Praveen Paruchuri
AAAI3
2017 Improving Surveillance Using Cooperative Target Observation
abstract
The Cooperative Target Observation (CTO) problem has been of great interest in the multi-agents and robotics literature due to the problem being at the core of a number of applications including surveillance. In CTO problem, the observer agents attempt to maximize the collective time during which each moving target is being observed by at least one observer in the area of interest. However, most of the prior works for the CTO problem consider the targets movement to be Randomized. Given our focus on surveillance domain, we modify this assumption to make the targets strategic and present two target strategies namely Straight-line strategy and Controlled Randomization strategy. We then modify the observer strategy proposed in the literature based on the K-means algorithm to introduce five variants and provide experimental validation. In surveillance domain, it is often reasonable to assume that the observers may themselves be a subject of observation for a variety of purposes by unknown adversaries whose model may not be known. Randomizing the observers actions can help to make their target observation strategy less predictable. As the fifth variant, we therefore introduce Adjustable Randomization into the best performing observer strategy where the observer can adjust the expected loss in reward due to randomization depending on the situation.
Rashi Aswani, Sai Krishna Munnangi, Praveen Paruchuri
AAAI3
2016 V2V communication for analysis of lane level dynamics for better EV traversal
abstract
Slow moving traffic in heavily populated cities, can many times result in loss of lives due to emergency vehicles not being able to reach their destination hospitals on time. Recent advances in the field of Intelligent Transportation Systems (ITS) makes it increasingly likely that vehicles in the near future will be equipped with advanced systems that allow inter vehicular communication. In this paper, we assume the usage of such a system to optimize the lane level dynamics for an emergency vehicle (EV), traversing a multi lane stretch of road under a variety of traffic settings. In particular, we present the Fixed Lane Strategy (FLS) and the Best Lane Strategy (BLS) for EV traversal and perform an extensive agent based analysis to study their strengths and weaknesses. Through a series of experiments performed using the well-known traffic simulator SUMO, we could show that: (a) BLS performs better than SUMO strategy on all traffic settings we tested. (b) BLS performs better than FLS in settings that capture real-world traffic conditions involving congestion and uncertainties while FLS performs better in well-behaved conditions and (c) BLS was found to be the best strategy for the setting calibrated using real world data (obtained from NYCDOT).
Akash Agarwal, Praveen Paruchuri
Intelligent Vehicles Symposium2
2008 Efficient Algorithms to Solve Bayesian Stackelberg Games for Security Applications
Praveen Paruchuri, Jonathan P. Pearce, Janusz Marecki, Milind Tambe, Fernando Ordóñez, Sarit Kraus
AAAI1
2008 ARMOR Security for Los Angeles International Airport
James Pita, Fernando Ordóñez, Christopher Portway, Milind Tambe, Craig Western, Praveen Paruchuri, Sarit Kraus
AAAI7