EDBT 2026 Demo / reviewers in the wild / expert
Prashant Doshi
dblp:d/PrashantDoshi
· DBLP profile ↗
82ranked-venue papers
12as first author
22since 2021 · last 2026
0000-0001-9042-9131ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 9 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 6 first-author · 3 since 2021Software engineering, systems software and programming languages · 12 · 1 first-authorSystems, architecture and hardware · 8 · 4 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decision-theoretic planning and cognitive modeling for active cyber deception
Aditya Shinde, Prashant Doshi |
Artif. Intell. | 2 |
| 2025 | Inferring Hidden Behavioral Signatures of Cyber Adversaries Using Inverse Reinforcement LearningabstractThis paper presents an emerging approach to attacker preference modeling from system-level audit logs using inverse reinforcement learning (IRL). Adversary modeling is an important capability in cybersecurity that lets defenders characterize behaviors of potential attackers, which enables attribution to known cyber adversary groups. Existing approaches rely on documenting an ever-evolving set of attacker tools and techniques to track known threat actors. Although attacks evolve constantly, attacker behavioral preferences are intrinsic and less volatile. Our approach learns the behavioral preferences of cyber adversaries from forensics data on their tools and techniques. We model the attacker as an expert decision-making agent with unknown behavioral preferences situated in a computer host. We leverage attack provenance graphs of audit logs to derive a state-action trajectory of the attack. We test our approach on open datasets of audit logs containing real attack data. Our results demonstrate for the first time that low-level forensics data can automatically reveal an adversary’s subjective preferences, which serves as an additional dimension to modeling and documenting cyber adversaries. Attackers’ preferences tend to be less dynamic despite their different tools and indicate predispositions that are inherent to the attacker. As such, these inferred preferences can potentially serve as unique behavioral signatures of attackers and improve threat attribution. Aditya Shinde, Prashant Doshi |
ECAI | 2 |
| 2025 | A Novel Computational Framework of Robot Trust for Human-Robot TeamsabstractWhen humans collaborate, they form positive or negative experiences with each other. These experiences depend on various factors such as the individual's skills, abilities, and agency. In this paper, we consider human-robot collaborations and present a novel model of an autonomous robot's trust in humans based on the probability of the robot having a positive experience with the human. The model defines a dynamic trust-building process that translates into a computationallyaccessible implementation. We hypothesize predictors of a positive experience with human teammates and derive trust in individual humans. As the interactions continue, team members develop an affinity toward each other. The robot's affinity towards humans can be viewed as kinship, and we also investigate how kinship affects trust and distrust. We present an algorithm for how the robot may use kinship-mediated trust in its decision-making, and demonstrate its use in simulated missions truly requiring human-robot collaboration. Bhavana Nare, John Frericks, Anusha Challa, Prashant Doshi, Kyle Johnsen 0001 |
ICRA | 4 |
| 2025 | Analyzing Human Perceptions of a MEDEVAC Robot in a Simulated Evacuation ScenarioabstractThe use of autonomous systems in medical evacuation (MEDEVAC) scenarios is promising, but existing implementations overlook key insights from human-robot interaction (HRI) research. Studies on human-machine teams demonstrate that human perceptions of a machine teammate are critical in governing the machine’s performance. Consequently, it is essential to identify the factors that contribute to positive human perceptions in human-machine teams. Here, we present a mixed factorial design to assess human perceptions of a MEDEVAC robot in a simulated evacuation scenario. Participants were assigned to the role of casualty (CAS) or bystander (BYS) and subjected to three within-subjects conditions based on the MEDEVAC robot’s operating mode: autonomous-slow (AS), autonomous-fast (AF), and teleoperation (TO). During each trial, a MEDEVAC robot navigated an 11-meter path, acquiring a casualty and transporting them to an ambulance exchange point while avoiding an idle bystander. Following each trial, subjects completed a questionnaire measuring their emotional states, perceived safety, and social compatibility with the robot. Results indicate a consistent main effect of operating mode on reported emotional states and perceived safety. Pairwise analyses suggest that the employment of the AF operating mode negatively impacted perceptions along these dimensions. There were no persistent differences between CAS and BYS responses. Tyson Jordan, Pranav Pandey, Prashant Doshi, Ramviyas Parasuraman, Adam Goodie |
IROS | 3 |
| 2025 | Integrating Perceptions: A Human-Centered Physical Safety Model for Human-Robot InteractionabstractEnsuring safety in human-robot interaction (HRI) is essential to foster user trust and enable the broader adoption of robotic systems. Traditional safety models primarily rely on sensor-based measures, such as relative distance and velocity, to assess physical safety. However, these models often fail to capture subjective safety perceptions, which are shaped by individual traits and contextual factors. In this paper, we introduce and analyze a parameterized general safety model that bridges the gap between physical and perceived safety by incorporating a personalization parameter, ρ, into the safety measurement framework to account for individual differences in safety perception. Through a series of hypothesis-driven human-subject studies in a simulated rescue scenario, we investigate how emotional state, trust, and robot behavior influence perceived safety. Our results show that ρ effectively captures meaningful individual differences, driven by affective responses, trust in task consistency, and clustering into distinct user types. Specifically, our findings confirm that predictable and consistent robot behavior as well as the elicitation of positive emotional states, significantly enhance perceived safety. Moreover, responses cluster into a small number of user types, supporting adaptive personalization based on shared safety models. Notably, participant role significantly shapes safety perception, and repeated exposure reduces perceived safety for participants in the casualty role, emphasizing the impact of physical interaction and experiential change. These findings highlight the importance of adaptive, human-centered safety models that integrate both psychological and behavioral dimensions, offering a pathway toward more trustworthy and effective HRI in safety-critical domains. Pranav Pandey, Ramviyas Parasuraman, Prashant Doshi |
RO-MAN | 3 |
| 2025 | MOHITO: Multi-Agent Reinforcement Learning using Hypergraphs for Task-Open SystemsabstractOpen agent systems are prevalent in the real world, where the sets of agents and tasks change over time. In this paper, we focus on task-open multi-agent systems, exemplified by applications such as ridesharing, where passengers (tasks) appear spontaneously over time and disappear if not attended to promptly. Task-open settings challenge us with an action space which changes dynamically. This renders existing reinforcement learning (RL) methods–intended for fixed state and action spaces–inapplicable. Whereas multi-task learning approaches learn policies generalized to multiple known and related tasks, they struggle to adapt to previously unseen tasks. Conversely, lifelong learning adapts to new tasks over time, but generally assumes that tasks come sequentially from a static and known distribution rather than simultaneously and unpredictably. We introduce a novel category of RL for addressing task openness, modeled using a task-open Markov game. Our approach, MOHITO, is a multi-agent actor-critic schema which represents knowledge about the relationships between agents and changing tasks and actions as dynamically evolving 3-uniform hypergraphs. As popular multi-agent RL testbeds do not exhibit task openness, we evaluate MOHITO on two realistic and naturally task-open domains to establish its efficacy and provide a benchmark for future work in this setting. Gayathri Anil, Prashant Doshi, Daniel Redder, Adam Eck, Leen-Kiat Soh |
UAI | 2 |
| 2025 | Adaptive Human-Robot Collaboration using Type-Based IRLabstractHuman-robot collaboration (HRC) integrates the consistency and precision of robotic systems with the dexterity and cognitive abilities of humans to create synergy. However, human performance may degrade due to various factors (e.g., fatigue, trust) which can manifest unpredictably, and typically results in diminished output and reduced quality. To address this challenge toward successful HRCs, we present a human-aware approach to collaboration using a novel multi-agent decision-making framework. Type-based decentralized Markov decision processes (TB-DecMDP) additionally model latent, causal decision-making factors influencing agent behavior (e.g., fatigue), leading to dynamic agent types. In this framework, agents can switch between types and each maintains a belief about others’ current type based on observed actions while aiming to achieve a shared objective. We introduce a new inverse reinforcement learning (IRL) algorithm, TB-DecAIRL, which uses TB-DecMDP to model complex HRCs. TB-DecAIRL learns a type-contingent reward function and corresponding vector of policies from team demonstrations. Our evaluations in a realistic HRC problem setting establish that modeling human types in TB-DecAIRL improves robot behavior on the default of ignoring human factors, by increasing throughput in a human-robot produce sorting task. Prasanth Sengadu Suresh, Prashant Doshi, Bikramjit Banerjee |
UAI | 2 |
| 2025 | Active legibility in multiagent reinforcement learningabstractA multiagent sequential decision problem has been seen in many critical applications including urban transportation, autonomous driving cars, military operations, etc. Its widely known solution, namely multiagent reinforcement learning, has evolved tremendously in recent years. Among them, the solution paradigm of modeling other agents attracts our interest, which is different from traditional value decomposition or communication mechanisms. It enables agents to understand and anticipate others' behaviors and facilitates their collaboration. Inspired by recent research on the legibility that allows agents to reveal their intentions through their behavior, we propose a multiagent active legibility framework to improve their performance. The legibility-oriented framework drives agents to conduct legible actions so as to help others optimise their behaviors. In addition, we design a series of problem domains that emulate a common legibility-needed scenario and effectively characterize the legibility in multiagent reinforcement learning. The experimental results demonstrate that the new framework is more efficient and requires less training time compared to several multiagent reinforcement learning algorithms. Yanyu Liu, Yinghui Pan, Yifeng Zeng, Biyang Ma, Prashant Doshi |
Artif. Intell. | 5 |
| 2024 | Open Human-Robot Collaboration using Decentralized Inverse Reinforcement LearningabstractThe growing interest in human-robot collaboration (HRC), where humans and robots cooperate towards shared goals, has seen significant advancements over the past decade. While previous research has addressed various challenges, several key issues remain unresolved. Many domains within HRC involve activities that do not necessarily require human presence throughout the entire task. Existing literature typically models HRC as a closed system, where all agents are present for the entire duration of the task. In contrast, an open model offers flexibility by allowing an agent to enter and exit the collaboration as needed, enabling them to concurrently manage other tasks. In this paper, we introduce a novel multiagent framework called oDec-MDP, designed specifically to model open HRC scenarios where agents can join or leave tasks flexibly during execution. We generalize a recent multiagent inverse reinforcement learning method - Dec-AIRL to learn from open systems modeled using the oDec-MDP. Our method is validated through experiments conducted in both a simplified toy firefighting domain and a realistic dyadic human-robot collaborative assembly. Results show that our framework and learning method improves upon its closed system counterpart. Prasanth Sengadu Suresh, Siddarth Jain, Prashant Doshi, Diego Romeres |
IROS | 3 |
| 2024 | An Autoencoder-Like Nonnegative Matrix Co-Factorization for Improved Student Cognitive ModelingabstractStudent cognitive modeling (SCM) is a fundamental task in intelligent education, with applications ranging from personalized learning to educational resource allocation. By exploiting students' response logs, SCM aims to predict their exercise performance as well as estimate knowledge proficiency in a subject. Data mining approaches such as matrix factorization can obtain high accuracy in predicting student performance on exercises, but the knowledge proficiency is unknown or poorly estimated. The situation is further exacerbated if only sparse interactions exist between exercises and students (or knowledge concepts). To solve this dilemma, we root monotonicity (a fundamental psychometric theory on educational assessments) in a co-factorization framework and present an autoencoder-like nonnegative matrix co-factorization (AE-NMCF), which improves the accuracy of estimating the student's knowledge proficiency via an encoder-decoder learning pipeline. The resulting estimation problem is nonconvex with nonnegative constraints. We introduce a projected gradient method based on block coordinate descent with Lipschitz constants and guarantee the method's theoretical convergence. Experiments on several real-world data sets demonstrate the efficacy of our approach in terms of both performance prediction accuracy and knowledge estimation ability, when compared with existing student cognitive models. Shenbao Yu, Yinghui Pan, Yifeng Zeng, Prashant Doshi, Guoquan Liu, Kim-Leng Poh, Mingwei Lin |
NeurIPS | 4 |
| 2024 | IRL for Restless Multi-armed Bandits with Applications in Maternal and Child Health
Gauri Jain, Pradeep Varakantham, Aparna Taneja, Prashant Doshi, Milind Tambe |
PRICAI (5) | 5 |
| 2024 | Robust Individualistic Learning in Many-Agent Systems
Keyang He, Prashant Doshi, Bikramjit Banerjee |
PRIMA | 2 |
| 2024 | Modeling and reinforcement learning in partially observable many-agent systems
Keyang He, Prashant Doshi, Bikramjit Banerjee |
Auton. Agents Multi Agent Syst. | 2 |
| 2022 | Reinforcement learning in many-agent settings under partial observabilityabstractRecent renewed interest in multi-agent reinforcement learning (MARL) has generated an impressive array of techniques that leverage deep RL, primarily actor-critic architectures, and can be applied to a limited range of settings in terms of observability and communication. However, a continuing limitation of much of this work is the curse of dimensionality when it comes to representations based on joint actions, which grow exponentially with the number of agents. In this paper, we squarely focus on this challenge of scalability. We apply the key insight of action anonymity to a recently presented actor-critic based MARL algorithm, interactive A2C. We introduce a Dirichlet-multinomial model for maintaining beliefs over the agent population when agents’ actions are not perfectly observable. We show that the posterior is a mixture of Dirichlet distributions that we approximate as a single component for tractability. We also show that the prediction accuracy of this method increases with more agents. Finally we show empirically that our method can learn optimal behaviors in two recently introduced pragmatic domains with large agent population, and demonstrates robustness in partially observable environments. Keyang He, Prashant Doshi, Bikramjit Banerjee |
UAI | 2 |
| 2022 | Decision-theoretic planning with communication in open multiagent systemsabstractIn open multiagent systems, the set of agents operating in the environment changes over time and in ways that are nontrivial to predict. For example, if collaborative robots were tasked with fighting wildfires, they may run out of suppressants and be temporarily unavailable to assist their peers. Because an agent’s optimal action depends on the actions of others, each agent must not only predict the actions of its peers, but, before that, reason whether they are even present to perform an action. Addressing openness thus requires agents to model each other’s presence, which can be enhanced through agents communicating about their presence in the environment. At the same time, communicative acts can also incur costs (e.g., consuming limited bandwidth), and thus an agent must tradeoff the benefits of enhanced coordination with the costs of communication. We present a new principled, decision-theoretic method in the context provided by the recent communicative interactive POMDP framework for planning in open agent settings that balances this tradeoff. Simulations of multiagent wildfire suppression problems demonstrate how communication can improve planning in open agent environments, as well as how agents tradeoff the benefits and costs of communication under different scenarios. Anirudh Kakarlapudi, Gayathri Anil, Adam Eck, Prashant Doshi, Leen-Kiat Soh |
UAI | 4 |
| 2022 | Marginal MAP estimation for inverse RL under occlusion with observer noiseabstractWe consider the problem of learning the behavioral preferences of an expert engaged in a task from noisy and partially-observable demonstrations. This is motivated by real-world applications such as a line robot learning from observing a human worker, where some observations are occluded by environmental elements. Furthermore, robotic perception tends to be imperfect and noisy. Previous techniques for inverse reinforcement learning (IRL) take the approach of either omitting the missing portions or inferring it as part of expectation-maximization, which tends to be slow and prone to local optima. We present a new method that generalizes the well-known Bayesian maximum-a-posteriori (MAP) IRL method by marginalizing the occluded portions of the trajectory. This is then extended with an observation model to account for perception noise. This novel application of marginal MAP (MMAP) to IRL significantly improves on the previous IRL technique under occlusion in both formative evaluations on a toy problem and in a summative evaluation on a produce sorting line task by a physical robot. Prasanth Sengadu Suresh, Prashant Doshi |
UAI | 2 |
| 2021 | A Novel AI-based Methodology for Identifying Cyber Attacks in Honey PotsabstractWe present a novel AI-based methodology that identifies phases of a host-level cyber attack simply from system call logs. System calls emanating from cyber attacks on hosts such as honey pots are often recorded in audit logs. Our methodology first involves efficiently loading, caching, processing, and querying system events contained in audit logs in support of computer forensics. Output of queries remains at the system call level and is difficult to process. The next step is to infer a sequence of abstracted actions, which we colloquially call a storyline, from the system calls given as observations to a latent-state probabilistic model. These storylines are then accurately identified with class labels using a learned classifier. We qualitatively and quantitatively evaluate methods and models for each step of the methodology using 114 different attack phases collected by logging the attacks of a red team on a server, on some likely benign sequences containing regular user activities, and on traces from a recent DARPA project. The resulting end-to-end system, which we call Cyberian, identifies the attack phases with a high level of accuracy illustrating the benefit that this machine learning-based methodology brings to security forensics. Muhammed AbuOdeh, Christian Adkins, Omid Setayeshfar, Prashant Doshi, Kyu Hyung Lee |
AAAI | 4 |
| 2021 | Min-Max Entropy Inverse RL of Multiple TasksabstractMulti-task IRL recognizes that expert(s) could be switching between multiple ways of solving the same problem, or interleaving demonstrations of multiple tasks. The learner aims to learn the reward functions that individually guide these distinct ways. We present a new method for multi-task IRL that generalizes the well-known maximum entropy approach by combining it with a Dirichlet process based minimum entropy clustering of the observed data. This yields a single nonlinear optimization problem, called MinMaxEnt Multi-task IRL (MME-MTIRL), which can be solved using the Lagrangian relaxation and gradient descent methods. We evaluate MME-MTIRL on the robotic task of sorting onions on a processing line where the expert utilizes multiple ways of detecting and removing blemished onions. The method is able to learn the underlying reward functions to a high level of accuracy and it improves on the previous approaches. Saurabh Arora, Prashant Doshi, Bikramjit Banerjee |
ICRA | 2 |
| 2021 | State-Based Recurrent SPMNs for Decision-Theoretic Planning under Partial ObservabilityabstractThe sum-product network (SPN) has been extended to model sequence data with the recurrent SPN (RSPN), and to decision-making problems with sum-product-max networks (SPMN). In this paper, we build on the concepts introduced by these extensions and present state-based recurrent SPMNs (S-RSPMNs) as a generalization of SPMNs to sequential decision-making problems where the state may not be perfectly observed. As with recurrent SPNs, S-RSPMNs utilize a repeatable template network to model sequences of arbitrary lengths. We present an algorithm for learning compact template structures by identifying unique belief states and the transitions between them through a state matching process that utilizes augmented data. In our knowledge, this is the first data-driven approach that learns graphical models for planning under partial observability, which can be solved efficiently. S-RSPMNs retain the linear solution complexity of SPMNs, and we demonstrate significant improvements in compactness of representation and the run time of structure learning and inference in sequential domains. Layton Hayes, Prashant Doshi, Swaraj Pawar, Hari Teja Tatavarti |
IJCAI | 2 |
| 2021 | I2RL: online inverse reinforcement learning under occlusion
Saurabh Arora, Prashant Doshi, Bikramjit Banerjee |
Auton. Agents Multi Agent Syst. | 2 |
| 2021 | A survey of inverse reinforcement learning: Challenges, methods and progress
Saurabh Arora, Prashant Doshi |
Artif. Intell. | 2 |
| 2021 | PALO bounds for reinforcement learning in partially observable stochastic games
Roi Ceren, Keyang He, Prashant Doshi, Bikramjit Banerjee |
Neurocomputing | 3 |
| 2020 | Scalable Decision-Theoretic Planning in Open and Typed Multiagent SystemsabstractIn open agent systems, the set of agents that are cooperating or competing changes over time and in ways that are nontrivial to predict. For example, if collaborative robots were tasked with fighting wildfires, they may run out of suppressants and be temporarily unavailable to assist their peers. We consider the problem of planning in these contexts with the additional challenges that the agents are unable to communicate with each other and that there are many of them. Because an agent's optimal action depends on the actions of others, each agent must not only predict the actions of its peers, but, before that, reason whether they are even present to perform an action. Addressing openness thus requires agents to model each other's presence, which becomes computationally intractable with high numbers of agents. We present a novel, principled, and scalable method in this context that enables an agent to reason about others' presence in its shared environment and their actions. Our method extrapolates models of a few peers to the overall behavior of the many-agent system, and combines it with a generalization of Monte Carlo tree search to perform individual agent reasoning in many-agent open environments. Theoretical analyses establish the number of agents to model in order to achieve acceptable worst case bounds on extrapolation error, as well as regret bounds on the agent's utility from modeling only some neighbors. Simulations of multiagent wildfire suppression problems demonstrate our approach's efficacy compared with alternative baselines. Adam Eck, Maulik Shah, Prashant Doshi, Leen-Kiat Soh |
AAAI | 3 |
| 2020 | SA-Net: Robust State-Action Recognition for Learning from ObservationsabstractLearning from observation (LfO) offers a new paradigm for transferring task behavior to robots. LfO requires the robot to observe the task being performed and decompose the sensed streaming data into sequences of state-action pairs, which are then input to LfO methods. Thus, recognizing the state-action pairs correctly and quickly in sensed data is a crucial prerequisite. We present SA-Net a deep neural network architecture that recognizes state-action pairs from RGB-D data streams. SA-Net performs well in two replicated robotic applications of LfO - one involving mobile ground robots and another involving a robotic manipulator - which demonstrates that the architecture could generalize well to differing contexts. Comprehensive evaluations including deployment on a physical robot show that SA-Net significantly improves on the accuracy of the previous methods under various conditions. Nihal Soans, Ehsan Asali, Prashant Doshi |
ICRA | 4 |
| 2020 | Recursively modeling other agents for decision making: A research perspective
Prashant Doshi, Piotr J. Gmytrasiewicz, Edmund H. Durfee |
Artif. Intell. | 1 |
| 2019 | Model-Free IRL Using Maximum Likelihood EstimationabstractThe problem of learning an expert’s unknown reward function using a limited number of demonstrations recorded from the expert’s behavior is investigated in the area of inverse reinforcement learning (IRL). To gain traction in this challenging and underconstrained problem, IRL methods predominantly represent the reward function of the expert as a linear combination of known features. Most of the existing IRL algorithms either assume the availability of a transition function or provide a complex and inefficient approach to learn it. In this paper, we present a model-free approach to IRL, which casts IRL in the maximum likelihood framework. We present modifications of the model-free Q-learning that replace its maximization to allow computing the gradient of the Q-function. We use gradient ascent to update the feature weights to maximize the likelihood of expert’s trajectories. We demonstrate on two problem domains that our approach improves the likelihood compared to previous methods. Vinamra Jain, Prashant Doshi, Bikramjit Banerjee |
AAAI | 2 |
| 2019 | Evacuate or Not? A POMDP Model of the Decision Making of Individuals in Hurricane Evacuation Zones
Adithya Raam Sankar, Prashant Doshi, Adam Goodie |
UAI | 2 |
| 2018 | Inverse Learning of Robot Behavior for Collaborative PlanningabstractInverse reinforcement learning (IRL) is an important basis for learning from demonstrations. Observing an agent, human or robotic, perform a task provides information and facilitates learning the task. We show how the agent's preferences learned using IRL can be incorporated in a subject robot's decision making and planning, to enable the robot to spontaneously collaborate with the previously observed agent on the task. We prioritize a real-world application, where a line robot will autonomously collaborate with another robot in sorting ripe and unripe fruit such as oranges. Toward this, our evaluations utilize a colored-ball sorting task as an analog using simulated TurtleBots equipped with Phantom X arms. Our method is comprehensive providing first answers to questions such as how should the robot acquire the complete model for the collaborative planning problem and how should it solve the problem to obtain a plan that permits collaboration without disrupting the line robot's behavior. Maulesh Trivedi, Prashant Doshi |
IROS | 2 |
| 2018 | Online Structure Learning for Feed-Forward and Recurrent Sum-Product NetworksabstractSum-product networks have recently emerged as an attractive representation due to their dual view as a special type of deep neural network with clear semantics and a special type of probabilistic graphical model for which inference is always tractable. Those properties follow from some conditions (i.e., completeness and decomposability) that must be respected by the structure of the network. As a result, it is not easy to specify a valid sum-product network by hand and therefore structure learning techniques are typically used in practice. This paper describes a new online structure learning technique for feed-forward and recurrent SPNs. The algorithm is demonstrated on real-world datasets with continuous features for which it is not clear what network architecture might be best, including sequence datasets of varying length. Agastya Kalra, Abdullah Rashwan, Wei-Shou Hsu, Pascal Poupart, Prashant Doshi, George Trimponias |
NeurIPS | 5 |
| 2018 | Multi-robot inverse reinforcement learning under occlusion with estimation of state transitions
Kenneth D. Bogert, Prashant Doshi |
Artif. Intell. | 2 |
| 2017 | On Markov Games Played by Bayesian and Boundedly-Rational PlayersabstractWe present a new game-theoretic framework in which Bayesian players with bounded rationality engage in a Markov game and each has private but incomplete information regarding other players' types. Instead of utilizing Harsanyi's abstract types and a common prior, we construct intentional player types whose structure is explicit and induces a {\em finite-level} belief hierarchy. We characterize an equilibrium in this game and establish the conditions for existence of the equilibrium. The computation of finding such equilibria is formalized as a constraint satisfaction problem and its effectiveness is demonstrated on two cooperative domains. Muthukumaran Chandrasekaran, Yingke Chen, Prashant Doshi |
AAAI | 3 |
| 2017 | A layered HMM for predicting motion of a leader in multi-robot settingsabstractWe focus on a mobile robot that must learn another robot's motion model from observations to track it in a given map. This problem has several real-world applications such as self-driving cars being electronically towed by other cars and for telepresence robots. Our context is a nested particle filter, a generalization of the traditional particle filter, that allows both self-localization and tracking of another robot simultaneously. While the robot's observations are used to weight nested particles, the problem arises during the propagation step of the nested particles during which a motion model is needed. We introduce a novel layered hidden Markov model for this problem and present an on-line algorithm which learns the HMM parameters from observations gathered during the run. We demonstrate significantly improved tracking accuracy on using this new model to predict the motion of a leading mobile robot, in comparison to pre-defined and random motion models as previously used in literature. Sina Solaimanpour, Prashant Doshi |
ICRA | 2 |
| 2017 | Robust Model Equivalence using Stochastic Bisimulation for N-Agent Interactive DIDs
Muthukumaran Chandrasekaran, Junhuan Zhang, Prashant Doshi, Yifeng Zeng |
UAI | 3 |
| 2017 | Can bounded and self-interested agents be teammates? Application to planning in ad hoc teams
Muthukumaran Chandrasekaran, Prashant Doshi, Yifeng Zeng, Yingke Chen |
Auton. Agents Multi Agent Syst. | 2 |
| 2017 | Decision-Theoretic Planning Under Anonymity in Agent PopulationsabstractWe study the problem of self-interested planning under uncertainty in settings shared with more than a thousand other agents, each of which plans at its own individual level. We refer to such large numbers of agents as an agent population. The decision-theoretic formalism of interactive partially observable Markov decision process (I-POMDP) is used to model the agent's self-interested planning. The first contribution of this article is a method for drastically scaling the finitely-nested I-POMDP to certain agent populations for the first time. Our method exploits two types of structure that is often exhibited by agent populations -- anonymity and context-specific independence. We present a variant called the many-agent I-POMDP that models both these types of structure to plan efficiently under uncertainty in multiagent settings. In particular, the complexity of the belief update and solution in the many-agent I-POMDP is polynomial in the number of agents compared with the exponential growth that challenges the original framework. While exploiting structure helps mitigate the curse of many agents, the well-known curse of history that afflicts I-POMDPs continues to challenge scalability in terms of the planning horizon. The second contribution of this article is an application of the branch-and-bound scheme to reduce the exponential growth of the search tree for look ahead. For this, we introduce new fast-computing upper and lower bounds for the exact value function of the many-agent I-POMDP. This speeds up the look-ahead computations without trading off optimality, and reduces both memory and run time complexity. The third contribution is a comprehensive empirical evaluation of the methods on three new problems domains -- policing large protests, controlling traffic congestion at a busy intersection, and improving the AI for the popular Clash of Clans multiplayer game. We demonstrate the feasibility of exact self-interested planning in these large problems, and that our methods for speeding up the planning are effective. Altogether, these contributions represent a principled and significant advance toward moving self-interested planning under uncertainty to real-world applications. Ekhlas Sonu, Yingke Chen, Prashant Doshi |
J. Artif. Intell. Res. | 3 |
| 2016 | Bayesian Markov Games with Explicit Finite-Level Types
Muthukumaran Chandrasekaran, Yingke Chen, Prashant Doshi |
AAAI | 3 |
| 2016 | Decision Sum-Product-Max NetworksabstractSum-Product Networks (SPNs) were recently proposed as a new class of probabilistic graphical models that guarantee tractable inference, even on models with high-treewidth. In this paper, we propose a new extension to SPNs, called Decision Sum-Product-Max Networks (Decision-SPMNs), that makes SPNs suitable for discrete multi-stage decision problems. We present an algorithm that solves Decision-SPMNs in a time that is linear in the size of the network. We also present algorithms to learn the parameters of the network from data. Mazen Melibari, Pascal Poupart, Prashant Doshi |
AAAI | 3 |
| 2016 | Sum-Product-Max Networks for Tractable Decision Making
Mazen Melibari, Pascal Poupart, Prashant Doshi |
IJCAI | 3 |
| 2016 | Individual Planning in Open and Typed Agent Systems
Muthukumaran Chandrasekaran, Adam Eck, Prashant Doshi, Leen-Kiat Soh |
UAI | 3 |
| 2016 | Approximating behavioral equivalence for scaling solutions of I-DIDs
Yifeng Zeng, Prashant Doshi, Yingke Chen, Yinghui Pan, Hua Mao 0001, Muthukumaran Chandrasekaran |
Knowl. Inf. Syst. | 2 |
| 2015 | Fast Solving of Influence Diagrams for Multiagent Planning on GPU-enabled Architectures
Fadel Adoe, Yingke Chen, Prashant Doshi |
ICAART (2) | 3 |
| 2015 | Toward Estimating Others' Transition Models Under Occlusion for Multi-Robot IRL
Kenneth D. Bogert, Prashant Doshi |
IJCAI | 2 |
| 2015 | Localization and tracking under extreme and persistent sensory occlusionabstractWe focus on a mobile robot who must keep itself localized while closely following another robot or human. This problem has many real-world applications including that of a co-bot engaged in a follow-the-leader behavior or a robot that is participating in a convoy. If the robot is expected to eventually break away and reach its own goal, then the robot must stay self-localized. A key challenge for localization while tailing another is the extreme and persistent occlusion of the robot's sensors by the dynamic obstacle in front of it that is not modeled in its map. Current Monte Carlo localization (MCL) methods use sensor models with random noise, which are inadequate under such occlusion. We utilize a particle filter that simultaneously tracks the subject robot and the leader. We introduce novel particle weighting and adaptive sampling schemes that significantly improve the follower's localization. The result is a robust and adaptive MCL for applications involving persistent occlusion. Kedar Marathe, Prashant Doshi |
IROS | 2 |
| 2015 | Individual Planning in Infinite-Horizon Multiagent Settings: Inference, Structure and ScalabilityabstractThis paper provides the first formalization of self-interested planning in multiagent settings using expectation-maximization (EM). Our formalization in the context of infinite-horizon and finitely-nested interactive POMDPs (I-POMDP) is distinct from EM formulations for POMDPs and cooperative multiagent planning frameworks. We exploit the graphical model structure specific to I-POMDPs, and present a new approach based on block-coordinate descent for further speed up. Forward filtering-backward sampling -- a combination of exact filtering with sampling -- is explored to exploit problem structure. Xia Qu, Prashant Doshi |
NIPS | 2 |
| 2015 | Scalable solutions of interactive POMDPs using generalized and bounded policy iteration
Ekhlas Sonu, Prashant Doshi |
Auton. Agents Multi Agent Syst. | 2 |
| 2014 | A Portal Designed to Learn about Educational Robotics
ChanMin Kim, Prashant Doshi, Chi N. Thai, Jiangmei Yuan |
CogSci | 2 |
| 2014 | Speeding Up Iterative Ontology Alignment using Block-Coordinate DescentabstractIn domains such as biomedicine, ontologies are prominently utilized for annotating data. Consequently, aligning ontologies facilitates integrating data. Several algorithms exist for automatically aligning ontologies with diverse levels of performance. As alignment applications evolve and exhibit online run time constraints, performing the alignment in a reasonable amount of time without compromising the quality of the alignment is a crucial challenge. A large class of alignment algorithms is iterative and often consumes more time than others in delivering solutions of high quality. We present a novel and general approach for speeding up the multivariable optimization process utilized by these algorithms. Specifically, we use the technique of block-coordinate descent (BCD), which exploits the subdimensions of the alignment problem identified using a partitioning scheme. We integrate this approach into multiple well-known alignment algorithms and show that the enhanced algorithms generate similar or improved alignments in significantly less time on a comprehensive testbed of ontology pairs. Because BCD does not overly constrain how we partition or order the parts, we vary the partitioning and ordering schemes in order to empirically determine the best schemes for each of the selected algorithms. As biomedicine represents a key application domain for ontologies, we introduce a comprehensive biomedical ontology testbed for the community in order to evaluate alignment algorithms. Because biomedical ontologies tend to be large, default iterative techniques find it difficult to produce a good quality alignment within a reasonable amount of time. We align a significant number of ontology pairs from this testbed using BCD-enhanced algorithms. Our contributions represent an important step toward making a significant class of alignment techniques computationally feasible. Uthayasanker Thayasivam, Prashant Doshi |
J. Artif. Intell. Res. | 2 |
| 2013 | Bimodal Switching for Online Planning in Multiagent Settings
Ekhlas Sonu, Prashant Doshi |
IJCAI | 2 |
| 2013 | On Modeling Human Learning in Sequential Games with Delayed ReinforcementsabstractWe model human learning in a repeated and sequential game context that provides delayed reinforcements. Our context is signifcantly more complex than previous work in behavioral game theory, which has predominantly focused on repeated single-shot games where the actions of other agent are perfectly observable and provides for an immediate reinforcement. In this complex context, we explore several established reinforcement learning models including temporal difference learning, SARSA and Q-learning. We generalize the default models by introducing behavioral factors that are refective of the cognitive biases observed in human play. We evaluate the model on data gathered from new experiments involving human participants making judgments under uncertainty in a repeated strategic and sequential game. We analyze the descriptive models against their default counterparts and show that modeling human aspects in reinforcement learning signifcantly improves predictive capabilities. This is useful in open and mixed networks of agent and human decision makers. Roi Ceren, Prashant Doshi, Matthew Meisel, Adam Goodie, Dan Hall |
SMC | 2 |
| 2013 | Canonical Forms and Similarity of Complex Concepts for Improved Ontology AlignmentabstractModern ontology languages such as the Web Ontology Language (OWL) allow defining complex concepts that involve restrictions, Boolean combinations, and exhaustive enumeration of individuals.Many of the current ontology alignment algorithms either do not consider the complex concepts in their alignment procedures or model them naively, thereby producing a possibly incomplete alignment. We introduce axiomatic and graphical canonical forms for modeling value and cardinality restrictions and Boolean combinations, and present a way of measuring the similarity between these complex concepts in their canonical forms. We integrate our approach in multiple ontology alignment algorithms. Our results indicate a significant improvement in the F-measure of the alignment. However, this improvement is at the expense of increased run time due to the additional concepts modeled. Tejas Chaudhari, Uthayasanker Thayasivam, Prashant Doshi |
Web Intelligence | 3 |
| 2012 | Improved Convergence of Iterative Ontology Alignment using Block-Coordinate DescentabstractA wealth of ontologies, many of which overlap in their scope, has made aligning ontologies an important problem for the semantic Web. Consequently, several algorithms now exist for automatically aligning ontologies, with mixed success in their performances. Crucial challenges for these algorithms involve scaling to large ontologies, and as applications of ontology alignment evolve, performing the alignment in a reasonable amount of time without compromising on the quality of the alignment. A class of alignment algorithms is iterative and often consumes more time than others while delivering solutions of high quality. We present a novel and general approach for speeding up the multivariable optimization process utilized by these algorithms. Specifically, we use the technique of block-coordinate descent in order to possibly improve the speed of convergence of the iterative alignment techniques. We integrate this approach into three well-known alignment systems and show that the enhanced systems generate similar or improved alignments in significantly less time on a comprehensive testbed of ontology pairs. This represents an important step toward making alignment techniques computationally more feasible. Uthayasanker Thayasivam, Prashant Doshi |
AAAI | 2 |
| 2012 | Exploiting Model Equivalences for Solving Interactive Dynamic Influence DiagramsabstractWe focus on the problem of sequential decision making in partially observable environments shared with other agents of uncertain types having similar or conflicting objectives. This problem has been previously formalized by multiple frameworks one of which is the interactive dynamic influence diagram (I-DID), which generalizes the well-known influence diagram to the multiagent setting. I-DIDs are graphical models and may be used to compute the policy of an agent given its belief over the physical state and others' models, which changes as the agent acts and observes in the multiagent setting. As we may expect, solving I-DIDs is computationally hard. This is predominantly due to the large space of candidate models ascribed to the other agents and its exponential growth over time. We present two methods for reducing the size of the model space and stemming its exponential growth. Both these methods involve aggregating individual models into equivalence classes. Our first method groups together behaviorally equivalent models and selects only those models for updating which will result in predictive behaviors that are distinct from others in the updated model space. The second method further compacts the model space by focusing on portions of the behavioral predictions. Specifically, we cluster actionally equivalent models that prescribe identical actions at a single time step. Exactly identifying the equivalences would require us to solve all models in the initial set. We avoid this by selectively solving some of the models, thereby introducing an approximation. We discuss the error introduced by the approximation, and empirically demonstrate the improved efficiency in solving I-DIDs due to the equivalences. Yifeng Zeng, Prashant Doshi |
J. Artif. Intell. Res. | 2 |
| 2012 | Modeling Human Recursive Reasoning Using Empirically Informed Interactive Partially Observable Markov Decision ProcessesabstractRecursive reasoning of the form what do I think that you think that I think (and so on) arises often while acting in multiagent settings. Previously, multiple experiments studied the level of recursive reasoning generally displayed by humans while playing sequential general-sum and fixed-sum, two-player games. The results show that subjects experiencing a general-sum strategic game display first or second level of recursive thinking with the first level being more prominent. However, if the game is made simpler and more competitive with fixed-sum payoffs, subjects predominantly attributed first-level recursive thinking to opponents thereby acting using second level. In this article, we model the behavioral data obtained from the studies using the interactive partially observable Markov decision process, appropriately simplified and augmented with well-known models simulating human learning and decision. We experiment with data collected at different points in the study to learn the models parameters. Accuracy of the predictions by our models is evaluated by comparing them with the observed study data, and the significance of the fit is demonstrated by comparing the mean squared error of our model predictions with those of a random hypothesis. Accuracy of the predictions by the models suggest that these could be viable ways for computationally modeling strategic behavioral data in a general way. While we do not claim the cognitive plausibility of the models in the absence of more evidence, they represent promising steps toward understanding and computationally simulating strategic human behavior. Prashant Doshi, Xia Qu, Adam Goodie, Diana L. Young |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2011 | Utilizing Partial Policies for Identifying Equivalence of Behavioral ModelsabstractWe present a novel approach for identifying exact and approximate behavioral equivalence between models of agents. This is significant because both decision making and game play in multiagent settings must contend with behavioral models of other agents in order to predict their actions. One approach that reduces the complexity of the model space is to group models that are behaviorally equivalent. Identifying equivalence between models requires solving them and comparing entire policy trees. Because the trees grow exponentially with the horizon, our approach is to focus on partial policy trees for comparison and determining the distance between updated beliefs at the leaves of the trees. We propose a principled way to determine how much of the policy trees to consider, which trades off solution quality for efficiency. We investigate this approach in the context of the interactive dynamic influence diagram and evaluate its performance. Yifeng Zeng, Prashant Doshi, Yinghui Pan, Hua Mao 0001, Muthukumaran Chandrasekaran |
AAAI | 2 |
| 2010 | Risk Sensitive Value of Changed Information for Selective Querying of Web Services
John Harney, Prashant Doshi |
ICSOC | 2 |
| 2010 | Model identification in interactive influence diagrams using mutual informationabstractInteractive influence diagrams (I-IDs) offer a transparent and intuitive representation for the decision-making problem in multiagent settings. They ascribe procedural models such as influence diagrams and I-IDs to model the behavior of other agents. Yifeng Zeng, Prashant Doshi |
Web Intell. Agent Syst. | 2 |
| 2009 | Selective Querying for Adapting Hierarchical Web Service Compositions Using Aggregate VolatilityabstractEnvironments in which Web service compositions (WSC) operate are often dynamic. We address the problem of which service to query for up-to-date information in order to adapt a hierarchical WSC, given that queries are not free. Previously,the value of changed information (VOC) has been proposed to select those services for querying whose revised non-functional information is expected to bring about the most change in the composition. In this paper, we present an approach for utilizing VOC in the context of a WSC composed of services and lower level WSCs, which induces a natural hierarchy over the composition. John Harney, Prashant Doshi |
ICWS | 2 |
| 2009 | Integrating Behavioral Trust in Web Service CompositionsabstractAlgorithms for composing Web services (WS) traditionally utilize the functional and quality-of-service parameters of candidate services to decide which services to include in the composition. Users often have differing experiences with a WS. While trust in a WS is multi-faceted and consists of security and behavioral aspects, our focus in this paper is on the latter. We adopt a formal model for trust in a WS, which meets many of our intuitions about trustworthy WSs. We hypothesize predictors of a positive experience with a WS and conduct a small pilot study to explore correlations between subjects' experiences with WSs in a composition and the predictor values for those WSs. Furthermore, we show how we may derive trust for compositions from trust models of individual services. We conclude by presenting and evaluating a novel framework, called Wisp, that utilizes the trust models and, in combination with any WS composition tool, chooses compositions to deploy that are deemed most trustworthy. Sharon Paradesi, Prashant Doshi, Sonu Swaika |
ICWS | 2 |
| 2009 | Towards Automated RESTful Web Service CompositionabstractEmerging as the popular choice for leading Internet companies to expose internal data and resources, RESTful Web services are attracting increasing attention in the industry. While automating WSDL/SOAP based Web service composition has been extensively studied in the research community, automated RESTful Web service composition in the context of service-oriented architecture (SOA), to the best of our knowledge, is less explored. As an early paper addressing this problem, this paper discusses the challenges of composing RESTful Web services and proposes a formal model for describing individual Web services and automating the composition. It demonstrates our approach by applying it to a real-world RESTful Web service composition problem. This paper represents our initial efforts towards the problem of automated RESTful Web service composition. We are hoping that it will draw interests from the research community on Web services, and engage more researchers in this challenge. Haibo Zhao 0002, Prashant Doshi |
ICWS | 2 |
| 2009 | Speeding Up Exact Solutions of Interactive Dynamic Influence Diagrams Using Action Equivalence
Yifeng Zeng, Prashant Doshi |
IJCAI | 2 |
| 2009 | Graphical models for interactive POMDPs: representations and solutions
Prashant Doshi, Yifeng Zeng, Qiongyu Chen |
Auton. Agents Multi Agent Syst. | 1 |
| 2009 | Monte Carlo Sampling Methods for Approximating Interactive POMDPsabstractPartially observable Markov decision processes (POMDPs) provide a principled framework for sequential planning in uncertain single agent settings. An extension of POMDPs to multiagent settings, called interactive POMDPs (I-POMDPs), replaces POMDP belief spaces with interactive hierarchical belief systems which represent an agents belief about the physical world, about beliefs of other agents, and about their beliefs about others beliefs. This modification makes the difficulties of obtaining solutions due to complexity of the belief and policy spaces even more acute. We describe a general method for obtaining approximate solutions of I-POMDPs based on particle filtering (PF). We introduce the interactive PF, which descends the levels of the interactive belief hierarchies and samples and propagates beliefs at each level. The interactive PF is able to mitigate the belief space complexity, but it does not address the policy space complexity. To mitigate the policy space complexity sometimes also called the curse of history we utilize a complementary method based on sampling likely observations while building the look ahead reachability tree. While this approach does not completely address the curse of history, it beats back the curses impact substantially. We provide experimental results and chart future work. Prashant Doshi, Piotr J. Gmytrasiewicz |
J. Artif. Intell. Res. | 1 |
| 2009 | A hierarchical framework for logical composition of web services
Haibo Zhao 0002, Prashant Doshi |
Serv. Oriented Comput. Appl. | 2 |
| 2009 | Inexact matching of ontology graphs using expectation-maximization
Prashant Doshi, Ravikanth Kolli, Christopher Thomas 0001 |
J. Web Semant. | 1 |
| 2008 | Generalized Point Based Value Iteration for Interactive POMDPs
Prashant Doshi, Dennis Perez |
AAAI | 1 |
| 2008 | Speeding up web service composition with volatile informationabstractThis paper introduces a novel method for composing Web services in the presence of external volatile information. Our approach, which we call the informed-presumptive, is compared to previous state-of-the-art approaches for Web service composition in volatile environments. We show empirically that the informed-presumptive strategy produces compositions in significantly less time than the other strategies with lesser backtracks. John Harney, Prashant Doshi |
WWW | 2 |
| 2008 | Making BPEL flexible: adapting in the context of coordination constraints using WS-BPELabstractAlthough WS-BPEL is emerging as the prominent language for modeling executable business processes, its support for designing flexible processes is limited. Yunzhou Wu, Prashant Doshi |
WWW | 2 |
| 2008 | Selective Querying for Adapting Web Service Compositions Using the Value of Changed InformationabstractWeb service composition (WSC) techniques assume that the parameters used to model the environment remain static and accurate throughout the composition's execution. However, WSCs often operate in environments where the parameters of its component services are volatile. To remain optimal, WSCs must adapt to these changes. Adaptation requires up-to-date knowledge about the revised parameters of each of the services. One way of obtaining this knowledge is to query services for their revised parameters. Querying services for their parameters is time-consuming and expensive. We must therefore carefully manage how queries are conducted. Specifically, an adaptive WSC must know when to query for revised information, and from which service(s) to obtain information. We present a method to selectively query services using the value of changed information (VOC). VOC measures the value of the change that revised information may potentially introduce to the composition. We reduce the complexity of computing the VOC, first by anticipating values of the service parameters that do not change the WSC, and second by using parameter expiration times obtained from predefined service-level agreements. Using two scenarios, we illustrate our approach and demonstrate the computational savings theoretically and experimentally. John Harney, Prashant Doshi |
IEEE Trans. Serv. Comput. | 2 |
| 2007 | Improved State Estimation in Multiagent Settings with Continuous or Large Discrete State Spaces
Prashant Doshi |
AAAI | 1 |
| 2007 | Approximate Solutions of Interactive Dynamic Influence Diagrams Using Model Clustering
Yifeng Zeng, Prashant Doshi, Qiongyu Chen |
AAAI | 2 |
| 2007 | Improved Adaptation of Web Service Compositions Using Value of Changed InformationabstractWorkflows often operate in volatile environments in which the component services' QoS changes frequently. Optimally adapting to these changes becomes an important problem that must be addressed by the Web service composition and execution (WSCE) system being utilized. We adopt the A-WSCE framework that utilizes a three-stage approach for composing and executing Web workflows. The A-WSCE framework offers a way to adapt by defining multiple workflows and switching among them in case of component failure or changes in the QoS parameters. However, the A-WSCE framework suffers from the limitations imposed by a simple strategy of periodically checking the QoS offerings of randomly picked providers in order to decide whether the current workflow is optimal. To address these limitations, we associate the value of changed information (VOC) with each workflow and utilize the VOC to update which workflow to execute. We empirically demonstrate the improved performance of the workflows selected using the new approach in comparison to the original framework. Girish Chafle, Prashant Doshi, John Harney, Sumit Mittal, Biplav Srivastava |
ICWS | 2 |
| 2007 | Haley: A Hierarchical Framework for Logical Composition ofWeb ServicesabstractPrevalent approaches for automatically composing Web services (WSs) into Web processes predominantly utilize planning techniques to achieve the composition. However, many of the planning methods do not scale efficiently to large processes. In addition, they lack the capability to operate directly on the WS descriptions, and specifically on the preconditions and effects which may be represented using methods ground and propositionalize the higher level logic resulting in exponentially many more states. In this paper, we present a new framework for composing Web services into processes, called Haley, that exploits the natural hierarchy often found in Web processes. Haley uses symbolic techniques that operate directly on first order logic based representations of the state space to obtain the compositions. In addition to providing an approach that handles the uncertainty inheret in Web services, Haley guarantees cost-based optimality and offers an approach potentially scalable to large real world processes. Haibo Zhao 0002, Prashant Doshi |
ICWS | 2 |
| 2007 | Speeding up adaptation of web service compositions using expiration timesabstractWeb processes must often operate in volatile environments where the quality of service parameters of the participating service providers change during the life time of the process. In order to remain optimal, the Web process must adapt to these changes. Adaptation requires knowledge about the parameter changes of each of the service providers and using this knowledge to determine whether the Web process should make a different more optimal decision. Previously, we defined a mechanism called the value of changed information which measures the impact of expected changes in the service parameters on the Web process, thereby offering a way to query and incorporate those changes that are useful and cost-efficient. However, computing the value of changed information incurs a substantial computational overhead. In this paper, we use service expiration times obtained from pre-defined service level agreements to reduce the computational overhead of adaptation. We formalize the intuition that services whose parameters have not expired need not be considered for querying for revised information. Using two realistic scenarios, we illustrate our approach and demonstrate the associated computational savings. John Harney, Prashant Doshi |
WWW | 2 |
| 2006 | On the Difficulty of Achieving Equilibrium in Interactive POMDPs
Prashant Doshi, Piotr J. Gmytrasiewicz |
AAAI | 1 |
| 2006 | Inexact Matching of Ontology Graphs Using Expectation-Maximization
Prashant Doshi, Christopher Thomas 0001 |
AAAI | 1 |
| 2006 | Adaptive Web Processes Using Value of Changed Information
John Harney, Prashant Doshi |
ICSOC | 2 |
| 2006 | A Hierarchical Framework for Composing Nested Web Processes
Haibo Zhao 0002, Prashant Doshi |
ICSOC | 2 |
| 2006 | Optimal Adaptation in Web Processes with Coordination ConstraintsabstractWe present methods for optimally adapting Web processes to exogenous events while preserving inter-service constraints that necessitate coordination. For example, in a supply chain process, orders placed by a manufacturer may get delayed in arriving. In response to this event, the manufacturer has the choice of either waiting out the delay or changing the supplier. Additionally, there may be compatibility constraints between the different orders, thereby introducing the problem of coordination between them if the manufacturer chooses to change the suppliers. We focus on formulating the decision making models of the managers, who must adapt to external events while satisfying the coordination constraints, using Markov decision processes. Our methods range from being centralized and globally optimal in their adaptation but not scalable, to decentralized that is suboptimal but scalable to multiple managers. We also develop a hybrid approach that improves on the performance of the decentralized approach with a minimal loss of optimality Kunal Verma, Prashant Doshi, Karthik Gomadam, John A. Miller 0001, Amit P. Sheth |
ICWS | 2 |
| 2005 | A Particle Filtering Based Approach to Approximating Interactive POMDPs
Prashant Doshi, Piotr J. Gmytrasiewicz |
AAAI | 1 |
| 2005 | A Framework for Sequential Planning in Multi-Agent SettingsabstractThis paper extends the framework of partially observable Markov decision processes (POMDPs) to multi-agent settings by incorporating the notion of agent models into the state space. Agents maintain beliefs over physical states of the environment and over models of other agents, and they use Bayesian updates to maintain their beliefs over time. The solutions map belief states to actions. Models of other agents may include their belief states and are related to agent types considered in games of incomplete information. We express the agents autonomy by postulating that their models are not directly manipulable or observable by other agents. We show that important properties of POMDPs, such as convergence of value iteration, the rate of convergence, and piece-wise linearity and convexity of the value functions carry over to our framework. Our approach complements a more traditional approach to interactive settings which uses Nash equilibria as a solution paradigm. We seek to avoid some of the drawbacks of equilibria which may be non-unique and do not capture off-equilibrium behaviors. We do so at the cost of having to represent, process and continuously revise models of other agents. Since the agents beliefs may be arbitrarily nested, the optimal solutions to decision making problems are only asymptotically computable. However, approximate belief updates and approximately optimal plans are computable. We illustrate our framework using a simple application domain, and we show examples of belief updates and value functions. Piotr J. Gmytrasiewicz, Prashant Doshi |
J. Artif. Intell. Res. | 2 |
| 2004 | A Framework for Optimal Sequential Planning in Multiagent Settings
Prashant Doshi |
AAAI | 1 |
| 2004 | Dynamic Workflow Composition using Markov Decision ProcessesabstractThe advent of Web services has made automated workflow composition relevant to Web based applications. One technique, that has received some attention, for automatically composing workflows is AI-based classical planning. However, classical planning suffers from the paradox of first assuming deterministic behavior of Web services, then requiring the additional overhead of execution monitoring to recover from unexpected behavior of services. To address these concerns, we propose using Markov decision processes (MDPs), to model workflow composition. Our method models both, the inherent stochastic nature of Web services, and the dynamic nature of the environment. The resulting workflows are robust to nondeterministic behaviors of Web services and adaptive to a changing environment. Using an example scenario, we demonstrate our method and provide empirical results in its support. Prashant Doshi, Richard Goodwin, Rama Akkiraju, Kunal Verma |
ICWS | 1 |