EDBT 2026 Demo / reviewers in the wild / expert
Sandhya Saisubramanian
dblp:149/1426
· DBLP profile ↗
15ranked-venue papers
8as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 8 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 3 since 2021Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Reinforcement learning · 36% Planning, search and constraint satisfaction · 35% Trustworthy machine learning · 15% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 100% |
Topics — the 23 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › multi-agent planning
cooperative multi-agent planning |
0.9 | 1 | 2025 | Mitigating Side Effects in Multi-Agent Systems Using Blame Assignment · ICRA 2025 |
Machine learning › Reinforcement learning › markov decision process
constrained markov decision process |
0.7 | 1 | 2023 | Planning and Learning for Non-markovian Negative Side Effects Using Finite State Controllers · AAAI 2023 |
Machine learning › Trustworthy machine learning › interpretability
explanation-based learning |
0.7 | 1 | 2023 | Explanation-Guided Reward Alignment · IJCAI 2023 |
Machine learning › Trustworthy machine learning
interpretability |
0.7 | 1 | 2023 | Explanation-Guided Reward Alignment · IJCAI 2023 |
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
0.7 | 1 | 2023 | Explanation-Guided Reward Alignment · IJCAI 2023 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
planning under uncertainty |
0.7 | 1 | 2023 | Planning and Learning for Reliable Autonomy in the Open World · AAAI 2023 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
reward alignment |
0.7 | 1 | 2023 | Explanation-Guided Reward Alignment · IJCAI 2023 |
Robotics › Motion planning and robot control › motion planning › safe motion planning
safe planning |
0.7 | 1 | 2023 | Planning and Learning for Non-markovian Negative Side Effects Using Finite State Controllers · AAAI 2023 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
metareasoning |
0.6 | 1 | 2022 | Metareasoning for Safe Decision Making in Autonomous Systems · ICRA 2022 |
Machine learning › Reinforcement learning
multi-objective reinforcement learning |
0.4 | 1 | 2020 | A Multi-Objective Approach to Mitigate Negative Side Effects · IJCAI 2020 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
decision making under uncertainty |
0.4 | 1 | 2019 | Adaptive Modeling for Risk-Aware Decision Making · AAAI 2019 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › decision making under uncertainty
risk-aware decision making |
0.4 | 1 | 2019 | Adaptive Modeling for Risk-Aware Decision Making · AAAI 2019 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.3 | 1 | 2025 | Mitigating Side Effects in Multi-Agent Systems Using Blame Assignment · ICRA 2025 |
Mathematical optimization › discrete optimization
mixed integer linear programming |
0.2 | 1 | 2015 | Risk Based Optimization for Improving Emergency Medical Systems · AAAI 2015 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning under uncertainty
finite-state controllers |
0.2 | 1 | 2023 | Planning and Learning for Non-markovian Negative Side Effects Using Finite State Controllers · AAAI 2023 |
Machine learning › Reinforcement learning
markov decision process |
0.2 | 1 | 2014 | STREETS: Game-Theoretic Traffic Patrolling with Exploration and Exploitation · AAAI 2014 |
Knowledge, reasoning and agents › Multi-agent systems
security games |
0.2 | 1 | 2014 | STREETS: Game-Theoretic Traffic Patrolling with Exploration and Exploitation · AAAI 2014 |
Knowledge, reasoning and agents › Multi-agent systems › game theory
stackelberg game |
0.2 | 1 | 2014 | STREETS: Game-Theoretic Traffic Patrolling with Exploration and Exploitation · AAAI 2014 |
Mathematical optimization › multi-objective optimization
bi-criteria optimization |
0.2 | 1 | 2014 | STREETS: Game-Theoretic Traffic Patrolling with Exploration and Exploitation · AAAI 2014 |
Mathematical optimization
multi-objective optimization |
0.2 | 1 | 2014 | STREETS: Game-Theoretic Traffic Patrolling with Exploration and Exploitation · AAAI 2014 |
Robotics › Legged, aerial and field robots › space robotics
planetary rover exploration |
0.2 | 1 | 2022 | Metareasoning for Safe Decision Making in Autonomous Systems · ICRA 2022 |
Robotics › Robot navigation and mapping
safe autonomous systems |
0.2 | 1 | 2022 | Metareasoning for Safe Decision Making in Autonomous Systems · ICRA 2022 |
Smart cities and intelligent transportation › public safety
emergency medical services |
0.1 | 1 | 2015 | Risk Based Optimization for Improving Emergency Medical Systems · AAAI 2015 |
Methods — techniques the papers use, named apart from their topics
decentralized markov decision process · 0.9credit assignment · 0.9blame assignment · 0.9reinforcement learning · 0.7partially observable planning · 0.7inverse reinforcement learning · 0.7finite state controller learning · 0.7explanation-based verification · 0.7constrained MDP · 0.7arbitration algorithm · 0.6mixed integer linear programming · 0.4lagrangian relaxation · 0.4state sampling · 0.2compact game representation · 0.2adversary sampling · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mitigating Side Effects in Multi-Agent Systems Using Blame AssignmentabstractWhen independently trained or designed robots are deployed in a shared environment, their combined actions can lead to unintended negative side effects (NSEs). To ensure safe and efficient operation, robots must optimize task performance while minimizing the penalties associated with NSEs, balancing individual objectives with collective impact. We model the problem of mitigating NSEs in a cooperative multi-agent system as a bi-objective lexicographic decentralized Markov decision process. We assume independence of transitions and rewards with respect to the robots' tasks, but the joint NSE penalty creates a form of dependence in this setting. To improve scalability, the joint NSE penalty is decomposed into individual penalties for each robot using credit assignment, which facilitates decentralized policy computation. We empirically demonstrate, using mobile robots and in simulation, the effectiveness and scalability of our approach in mitigating NSEs. Code: https://tinyurl.com/RECON-NSE-Mitigation Pulkit Rustagi, Sandhya Saisubramanian |
ICRA | 2 |
| 2025 | Multi-Objective Planning with Contextual Lexicographic Reward Preferences
Pulkit Rustagi, Yashwanthi Anand, Sandhya Saisubramanian |
AAMAS | 3 |
| 2023 | Planning and Learning for Reliable Autonomy in the Open WorldabstractSafe and reliable decision-making is critical for long-term deployment of autonomous systems. Despite the recent advances in artificial intelligence, ensuring safe and reliable operation of human-aligned autonomous systems in open-world environments remains a challenge. My research focuses on developing planning and learning algorithms that support reliable autonomy in fully and partially observable environments, in the presence of uncertainty, limited information, and limited resources. This talk covers a summary of my recent research towards reliable autonomy. Sandhya Saisubramanian |
AAAI | 1 |
| 2023 | Planning and Learning for Non-markovian Negative Side Effects Using Finite State ControllersabstractAutonomous systems are often deployed in the open world where it is hard to obtain complete specifications of objectives and constraints. Operating based on an incomplete model can produce negative side effects (NSEs), which affect the safety and reliability of the system. We focus on mitigating NSEs in environments modeled as Markov decision processes (MDPs). First, we learn a model of NSEs using observed data that contains state-action trajectories and severity of associated NSEs. Unlike previous works that associate NSEs with state-action pairs, our framework associates NSEs with entire trajectories, which is more general and captures non-Markovian dependence on states and actions. Second, we learn finite state controllers (FSCs) that predict NSE severity for a given trajectory and generalize well to unseen data. Finally, we develop a constrained MDP model that uses information from the underlying MDP and the learned FSC for planning while avoiding NSEs. Our empirical evaluation demonstrates the effectiveness of our approach in learning and mitigating Markovian and non-Markovian NSEs. Aishwarya Srivastava, Sandhya Saisubramanian, Praveen Paruchuri, Akshat Kumar, Shlomo Zilberstein |
AAAI | 2 |
| 2023 | Explanation-Guided Reward AlignmentabstractAgents often need to infer a reward function from observations to learn desired behaviors. However, agents may infer a reward function that does not align with the original intent because there can be multiple reward functions consistent with its observations. Operating based on such misaligned rewards can be risky. Furthermore, black-box representations make it difficult to verify the learned rewards and prevent harmful behavior. We present a framework for verifying and improving reward alignment using explanations and show how explanations can help detect misalignment and reveal failure cases in novel scenarios. The problem is formulated as inverse reinforcement learning from ranked trajectories. Verification tests created from the trajectory dataset are used to iteratively validate and improve reward alignment. The agent explains its learned reward and a tester signals whether the explanation passes the test. In cases where the explanation fails, the agent offers alternative explanations to gather feedback, which is then used to improve the learned reward. We analyze the efficiency of our approach in improving reward alignment using different types of explanations and demonstrate its effectiveness in five domains. Saaduddin Mahmud, Sandhya Saisubramanian, Shlomo Zilberstein |
IJCAI | 2 |
| 2022 | Metareasoning for Safe Decision Making in Autonomous SystemsabstractAlthough experts carefully specify the high-level decision-making models in autonomous systems, it is infeasible to guarantee safety across every scenario during operation. We therefore propose a safety metareasoning system that optimizes the severity of the system's safety concerns and the interference to the system's task: the system executes in parallel a task process that completes a specified task and safety processes that each address a specified safety concern with a conflict resolver for arbitration. This paper offers a formal definition of a safety metareasoning system, a recommendation algorithm for a safety process, an arbitration algorithm for a conflict resolver, an application of our approach to planetary rover exploration, and a demonstration that our approach is effective in simulation. Justin Svegliato, Connor Basich, Sandhya Saisubramanian, Shlomo Zilberstein |
ICRA | 3 |
| 2022 | Avoiding Negative Side Effects of Autonomous Systems in the Open WorldabstractAutonomous systems that operate in the open world often use incomplete models of their environment. Model incompleteness is inevitable due to the practical limitations in precise model specification and data collection about open-world environments. Due to the limited fidelity of the model, agent actions may produce negative side effects (NSEs) when deployed. Negative side effects are undesirable, unmodeled effects of agent actions on the environment. NSEs are inherently challenging to identify at design time and may affect the reliability, usability and safety of the system. We present two complementary approaches to mitigate the NSE via: (1) learning from feedback, and (2) environment shaping. The solution approaches target settings with different assumptions and agent responsibilities. In learning from feedback, the agent learns a penalty function associated with a NSE. We investigate the efficiency of different feedback mechanisms, including human feedback and autonomous exploration. The problem is formulated as a multi-objective Markov decision process such that optimizing the agent’s assigned task is prioritized over mitigating NSE. A slack parameter denotes the maximum allowed deviation from the optimal expected reward for the agent’s task in order to mitigate NSE. In environment shaping, we examine how a human can assist an agent, beyond providing feedback, and utilize their broader scope of knowledge to mitigate the impacts of NSE. We formulate the problem as a human-agent collaboration with decoupled objectives. The agent optimizes its assigned task and may produce NSE during its operation. The human assists the agent by performing modest reconfigurations of the environment so as to mitigate the impacts of NSE, without affecting the agent’s ability to complete its assigned task. We present an algorithm for shaping and analyze its properties. Empirical evaluations demonstrate the trade-offs in the performance of different approaches in mitigating NSE in different settings. Sandhya Saisubramanian, Ece Kamar, Shlomo Zilberstein |
J. Artif. Intell. Res. | 1 |
| 2021 | Learning to Generate Fair Clusters from DemonstrationsabstractFair clustering is the process of grouping similar entities together, while satisfying a mathematically well-defined fairness metric as a constraint. Due to the practical challenges in precise model specification, the prescribed fairness constraints are often incomplete and act as proxies to the intended fairness requirement. Clustering with proxies may lead to biased outcomes when the system is deployed. We examine how to identify the intended fairness constraint for a problem based on limited demonstrations from an expert. Each demonstration is a clustering over a subset of the data. We present an algorithm to identify the fairness metric from demonstrations and generate clusters using existing off-the-shelf clustering techniques, and analyze its theoretical properties. To extend our approach to novel fairness metrics for which clustering algorithms do not currently exist, we present a greedy method for clustering. Additionally, we investigate how to generate interpretable solutions using our approach. Empirical evaluation on three real-world datasets demonstrates the effectiveness of our approach in quickly identifying the underlying fairness and interpretability constraints, which are then used to generate fair and interpretable clusters. Sainyam Galhotra, Sandhya Saisubramanian, Shlomo Zilberstein |
AIES | 2 |
| 2020 | Balancing the Tradeoff Between Clustering Value and InterpretabilityabstractGraph clustering groups entities -- the vertices of a graph -- based on their similarity, typically using a complex distance function over a large number of features. Successful integration of clustering approaches in automated decision-support systems hinges on the interpretability of the resulting clusters. This paper addresses the problem of generating interpretable clusters, given features of interest that signify interpretability to an end-user, by optimizing interpretability in addition to common clustering objectives. We propose a β-interpretable clustering algorithm that ensures that at least β fraction of nodes in each cluster share the same feature value. The tunable parameter β is user-specified. We also present a more efficient algorithm for scenarios with β\!=\!1$ and analyze the theoretical guarantees of the two algorithms. Finally, we empirically demonstrate the benefits of our approaches in generating interpretable clusters using four real-world datasets. The interpretability of the clusters is complemented by generating simple explanations denoting the feature values of the nodes in the clusters, using frequent pattern mining. Sandhya Saisubramanian, Sainyam Galhotra, Shlomo Zilberstein |
AIES | 1 |
| 2020 | A Multi-Objective Approach to Mitigate Negative Side EffectsabstractAgents operating in unstructured environments often create negative side effects (NSE) that may not be easy to identify at design time. We examine how various forms of human feedback or autonomous exploration can be used to learn a penalty function associated with NSE during system deployment. We formulate the problem of mitigating the impact of NSE as a multi-objective Markov decision process with lexicographic reward preferences and slack. The slack denotes the maximum deviation from an optimal policy with respect to the agent's primary objective allowed in order to mitigate NSE as a secondary objective. Empirical evaluation of our approach shows that the proposed framework can successfully mitigate NSE and that different feedback mechanisms introduce different biases, which influence the identification of NSE. Sandhya Saisubramanian, Ece Kamar, Shlomo Zilberstein |
IJCAI | 1 |
| 2019 | Adaptive Modeling for Risk-Aware Decision MakingabstractThis thesis aims to provide a foundation for risk-aware decision making. Decision making under uncertainty is a core capability of an autonomous agent. A cornerstone for with long-term autonomy and safety is risk-aware decision making. A risk-aware model fully accounts for a known set of risks in the environment, with respect to the problem under consideration, and the process of decision making using such a model is risk-aware decision making. Formulating risk-aware models is critical for robust reasoning under uncertainty, since the impact of using less accurate models may be catastrophic in extreme cases due to overly optimistic view of problems. I propose adaptive modeling, a framework that helps balance the trade-off between model simplicity and risk awareness, for different notions of risks, while remaining computationally tractable. Sandhya Saisubramanian |
AAAI | 1 |
| 2019 | Planning in Stochastic Environments with Goal UncertaintyabstractWe present the Goal Uncertain Stochastic Shortest Path (GUSSP) problem - a general framework to model path planning and decision making in stochastic environments with goal uncertainty. The framework extends the stochastic shortest path (SSP) model to dynamic environments in which it is impossible to determine the exact goal states ahead of plan execution. GUSSPs introduce flexibility in goal specification by allowing a belief over possible goal configurations. The unique observations at potential goals helps the agent identify the true goal during plan execution. The partial observability is restricted to goals, facilitating the reduction to an SSP with a modified state space. We formally define a GUSSP and discuss its theoretical properties. We then propose an admissible heuristic that reduces the planning time using FLARES - a start-of-the-art probabilistic planner. We also propose a determinization approach for solving this class of problems. Finally, we present empirical results on a search and rescue mobile robot and three other problem domains in simulation. Sandhya Saisubramanian, Kyle Hollins Wray, Luis Enrique Pineda, Shlomo Zilberstein |
IROS | 1 |
| 2019 | Adaptive Outcome Selection for Planning with Reduced ModelsabstractReduced models allow autonomous robots to cope with the complexity of planning in stochastic environments by simplifying the model and reducing its accuracy. The solution quality of a reduced model depends on its fidelity. We present 0/1 reduced model that selectively improves model fidelity in certain states by switching between using a simplified deterministic model and the full model, without significantly compromising the run time gains. We measure the reduction impact for a reduced model based on the values of the ignored outcomes and use this as a heuristic for outcome selection. Finally, we present empirical results of our approach on three different domains, including an electric vehicle charging problem using real-world data from a university campus. Sandhya Saisubramanian, Shlomo Zilberstein |
IROS | 1 |
| 2015 | Risk Based Optimization for Improving Emergency Medical SystemsabstractIn emergency medical systems, arriving at the incident locationa few seconds early can save a human life. Thus, this paper is motivated by the need to reduce the response time– time taken to arrive at the incident location after receivingthe emergency call — of Emergency Response Vehicles, ERVs(ex: ambulances, fire rescue vehicles) for as many requests as possible. We expect to achieve this primarily by positioning the ”right” number of ERVs at the ”right” places and at the ”right” times. Given the exponentially large action space(with respect to number of ERVs and their placement) and the stochasticity in location and timing of emergency incidents,this problem is computationally challenging. To that end, ourcontributions building on existing data-driven approaches are three fold:1. Based on real world evaluation metrics, we provide a riskbased optimization criterion to learn from past incident data. Instead of minimizing expected response time, we minimize the largest value of response time such that the risk of finding requests that have a higher value is bounded(ex: Only 10% of requests should have a response time greater than 8 minutes).2. We develop a mixed integer linear optimization formulation to learn and compute an allocation from a set of inputrequests while considering the risk criterion.3. To allow for ”live” reallocation of ambulances, we provide a decomposition method based on Lagrangian Relaxation to significantly reduce the run-time of the optimization formulation.Finally, we provide an exhaustive evaluation on real-world datasets from two asian cities that demonstrates the improvement provided by our approach over current practice and the best known approach from literature. Sandhya Saisubramanian, Pradeep Varakantham, Hoong Chuin Lau |
AAAI | 1 |
| 2014 | STREETS: Game-Theoretic Traffic Patrolling with Exploration and ExploitationabstractTo dissuade reckless driving and mitigate accidents, cities deploy resources to patrol roads. In this paper, we present STREETS, an application developed for the city of Singapore, which models the problem of computing randomized traffic patrol strategies as a defenderattacker Stackelberg game. Previous work on Stackelberg security games has focused extensively on counterterrorism settings. STREETS moves beyond counterterrorism and represents the first use of Stackelberg games for traffic patrolling, in the process providing a novel algorithm for solving such games that addresses three major challenges in modeling and scale-up. First, there exists a high degree of unpredictability in travel times through road networks, which we capture using a Markov Decision Process for planning the patrols of the defender (the police) in the game. Second, modeling all possible police patrols and their interactions with a large number of adversaries (drivers) introduces a significant scalability challenge. To address this challenge we apply a compact game representation in a novel fashion combined with adversary and state sampling. Third, patrol strategies must balance exploitation (minimizing violations) with exploration (maximizing omnipresence), a tradeoff we model by solving a biobjective optimization problem. We present experimental results using real-world traffic data from Singapore. This work is done in collaboration with the Singapore Ministry of Home Affairs and is currently being evaluated by the Singapore Police Force. Matthew Brown 0002, Sandhya Saisubramanian, Pradeep Varakantham, Milind Tambe |
AAAI | 2 |