VLDB 2026 Research / reviewers in the wild / expert
Vaibhav V. Unhelkar
dblp:142/3172 · also Vaibhav Vasant Unhelkar
· DBLP profile ↗
19ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0002-4530-189XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 3 since 2021Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Socratic: Enhancing Human Teamwork via AI-enabled Coaching
Sangwon Seo, Rayan Ebnali Harari, Roger D. Dias, Marco A. Zenati, Eduardo Salas, Vaibhav V. Unhelkar |
AAMAS | 7 |
| 2025 | Hierarchical Imitation Learning of Team Behavior from Heterogeneous Demonstrations
Sangwon Seo, Vaibhav V. Unhelkar |
AAMAS | 2 |
| 2024 | GO-DICE: Goal-Conditioned Option-Aware Offline Imitation Learning via Stationary Distribution Correction EstimationabstractOffline imitation learning (IL) refers to learning expert behavior solely from demonstrations, without any additional interaction with the environment. Despite significant advances in offline IL, existing techniques find it challenging to learn policies for long-horizon tasks and require significant re-training when task specifications change. Towards addressing these limitations, we present GO-DICE an offline IL technique for goal-conditioned long-horizon sequential tasks. GO-DICE discerns a hierarchy of sub-tasks from demonstrations and uses these to learn separate policies for sub-task transitions and action execution, respectively; this hierarchical policy learning facilitates long-horizon reasoning.Inspired by the expansive DICE-family of techniques, policy learning at both the levels transpires within the space of stationary distributions. Further, both policies are learnt with goal conditioning to minimize need for retraining when task goals change. Experimental results substantiate that GO-DICE outperforms recent baselines, as evidenced by a marked improvement in the completion rate of increasingly challenging pick-and-place Mujoco robotic tasks. GO-DICE is also capable of leveraging imperfect demonstration and partial task segmentation when available, both of which boost task performance relative to learning from expert demonstrations alone. Abhinav Jain 0001, Vaibhav V. Unhelkar |
AAAI | 2 |
| 2024 | I-CEE: Tailoring Explanations of Image Classification Models to User ExpertiseabstractEffectively explaining decisions of black-box machine learning models is critical to responsible deployment of AI systems that rely on them. Recognizing their importance, the field of explainable AI (XAI) provides several techniques to generate these explanations. Yet, there is relatively little emphasis on the user (the explainee) in this growing body of work and most XAI techniques generate "one-size-fits-all'' explanations. To bridge this gap and achieve a step closer towards human-centered XAI, we present I-CEE, a framework that provides Image Classification Explanations tailored to User Expertise. Informed by existing work, I-CEE explains the decisions of image classification models by providing the user with an informative subset of training data (i.e., example images), corresponding local explanations, and model decisions. However, unlike prior work, I-CEE models the informativeness of the example images to depend on user expertise, resulting in different examples for different users. We posit that by tailoring the example set to user expertise, I-CEE can better facilitate users' understanding and simulatability of the model. To evaluate our approach, we conduct detailed experiments in both simulation and with human participants (N = 100) on multiple datasets. Experiments with simulated users show that I-CEE improves users' ability to accurately predict the model's decisions (simulatability) compared to baselines, providing promising preliminary results. Experiments with human participants demonstrate that our method significantly improves user simulatability accuracy, highlighting the importance of human-centered XAI. Yao Rong 0001, Peizhu Qian, Vaibhav V. Unhelkar, Enkelejda Kasneci |
AAAI | 3 |
| 2024 | PPS: Personalized Policy Summarization for Explaining Sequential Behavior of Autonomous AgentsabstractAI-enabled agents designed to assist humans are gaining traction in a variety of domains such as healthcare and disaster response. It is evident that, as we move forward, these agents will play increasingly vital roles in our lives. To realize this future successfully and mitigate its unintended consequences, it is imperative that humans have a clear understanding of the agents that they work with. Policy summarization methods help facilitate this understanding by showcasing key examples of agent behaviors to their human users. Yet, existing methods produce “one-size-fits-all” summaries for a generic audience ahead of time. Drawing inspiration from research in pedagogy, we posit that personalized policy summaries can more effectively enhance user understanding. To evaluate this hypothesis, this paper presents and benchmarks a novel technique: Personalized Policy Summarization (PPS). PPS discerns a user’s mental model of the agent through a series of algorithmically generated questions and crafts customized policy summaries to enhance user understanding. Unlike existing methods, PPS actively engages with users to gauge their comprehension of the agent behavior, subsequently generating tailored explanations on the fly. Through a combination of numerical and human subject experiments, we confirm the utility of this personalized approach to explainable AI. Peizhu Qian, Harrison Huang, Vaibhav V. Unhelkar |
AIES (1) | 3 |
| 2024 | RW4T Dataset: Data of Human-Robot Behavior and Cognitive States in Simulated Disaster Response TasksabstractTo forge effective collaborations with humans, robots require the capacity to understand and predict the behaviors of their human counterparts. There is a growing body of computational research on human modeling for human-robot interaction (HRI). However, a key bottleneck in conducting this research is the relative lack of data of cognitive states -- like intent, workload, and trust -- which undeniably affect human behavior. Despite their significance, these states are elusive to measure, making the assembly of datasets a challenge and hindering the progression of human modeling techniques. To help address this, we first introduce Rescue World for Teams (RW4T): a configurable testbed to simulate disaster response scenarios requiring human-robot collaboration. Next, using RW4T, we curate a multimodal dataset of human-robot behavior and cognitive states in dyadic human-robot collaboration. This RW4T dataset includes state, action and reward sequences, and all the necessary data to replay a visual task execution. It further contains psychophysiological metrics like heart rate and pupillometry, complemented by self-reported cognitive state measures. With data from 20 participants, each undertaking five human-robot collaborative tasks, this dataset (comprising of 100 unique trajectories) accompanied with the simulator can serve as a valuable benchmark for human behavior modeling. Liubove Orlov-Savko, Zhiqin Qian, Gregory Gremillion, Catherine Neubauer, Jonroy D. Canady, Vaibhav V. Unhelkar |
HRI | 6 |
| 2024 | Towards Human-Centered Explainable AI: A Survey of User Studies for Model ExplanationsabstractExplainable AI (XAI) is widely viewed as a sine qua non for ever-expanding AI research. A better understanding of the needs of XAI users, as well as human-centered evaluations of explainable models are both a necessity and a challenge. In this paper, we explore how human-computer interaction (HCI) and AI researchers conduct user studies in XAI applications based on a systematic literature review. After identifying and thoroughly analyzing 97 core papers with human-based XAI evaluations over the past five years, we categorize them along the measured characteristics of explanatory methods, namely trust, understanding, usability, and human-AI collaboration performance. Our research shows that XAI is spreading more rapidly in certain application domains, such as recommender systems than in others, but that user evaluations are still rather sparse and incorporate hardly any insights from cognitive or social sciences. Based on a comprehensive discussion of best practices, i.e., common models, design choices, and measures in user studies, we propose practical guidelines on designing and conducting user studies for XAI researchers and practitioners. Lastly, this survey also highlights several open research directions, particularly linking psychological science and human-centered XAI. Yao Rong 0001, Tobias Leemann, Thai-trang Nguyen, Lisa Fiedler, Peizhu Qian, Vaibhav V. Unhelkar, Tina Seidel, Gjergji Kasneci, Enkelejda Kasneci |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Robotic Tutors for Nurse Training: Opportunities for HRI ResearchersabstractAn ongoing nurse labor shortage has the potential to impact patient care well-being in the entire healthcare system. Moreover, more complex and sophisticated nursing care is required today for patients in hospitals forcing hospital-based nurses to carry out frequent training and assessment procedures, both to onboard new nurses and to validate skills of existing staff that guarantees best practices and safety. In this paper we recognize an opportunity for the development and integration of intelligent robot tutoring technology into nursing education to tackle the growing challenges of nurse deficit. To this end, we identify specific research problems in the area of human-robot interaction that will need to be addressed to enable robot tutors for nurse training. Carlos Quintero-Peña, Peizhu Qian, Nicole M. Fontenot, Hsin-Mei Chen, Shannan K. Hamlin, Lydia E. Kavraki, Vaibhav V. Unhelkar |
RO-MAN | 7 |
| 2022 | Human-Guided Motion Planning in Partially Observable EnvironmentsabstractMotion planning is a core problem in robotics, with a range of existing methods aimed to address its diverse set of challenges. However, most existing methods rely on complete knowledge of the robot environment; an assumption that seldom holds true due to inherent limitations of robot perception. To enable tractable motion planning for high-DOF robots under partial observability, we introduce BLIND, an algorithm that leverages human guidance. BLIND utilizes inverse reinforcement learning to derive motion-level guidance from human critiques. The algorithm overcomes the computational challenge of reward learning for high-DOF robots by projecting the robot's continuous configuration space to a motion-planner-guided discrete task model. The learned reward is in turn used as guidance to generate robot motion using a novel motion planner. We demonstrate BLIND using the Fetch robot and perform two simulation experiments with partial observability. Our experiments demonstrate that, despite the challenge of partial observability and high dimensionality, BLIND is capable of generating safe robot motion and outperforms baselines on metrics of teaching efficiency, success rate, and path quality. Carlos Quintero-Peña, Constantinos Chamzas, Zhanyi Sun, Vaibhav V. Unhelkar, Lydia E. Kavraki |
ICRA | 4 |
| 2022 | Semi-Supervised Imitation Learning of Team Policies from Suboptimal DemonstrationsabstractWe present Bayesian Team Imitation Learner (BTIL), an imitation learning algorithm to model the behavior of teams performing sequential tasks in Markovian domains. In contrast to existing multi-agent imitation learning techniques, BTIL explicitly models and infers the time-varying mental states of team members, thereby enabling learning of decentralized team policies from demonstrations of suboptimal teamwork. Further, to allow for sample- and label-efficient policy learning from small datasets, BTIL employs a Bayesian perspective and is capable of learning from semi-supervised demonstrations. We demonstrate and benchmark the performance of BTIL on synthetic multi-agent tasks as well as a novel dataset of human-agent teamwork. Our experiments show that BTIL can successfully learn team policies from demonstrations despite the influence of team members' (time-varying and potentially misaligned) mental states on their behavior. Sangwon Seo, Vaibhav V. Unhelkar |
IJCAI | 2 |
| 2021 | Learning Dense Rewards for Contact-Rich Manipulation TasksabstractRewards play a crucial role in reinforcement learning. To arrive at the desired policy, the design of a suitable reward function often requires significant domain expertise as well as trial-and-error. Here, we aim to minimize the effort involved in designing reward functions for contact-rich manipulation tasks. In particular, we provide an approach capable of extracting dense reward functions algorithmically from robots’ high-dimensional observations, such as images and tactile feedback. In contrast to state-of-the-art high-dimensional reward learning methodologies, our approach does not leverage adversarial training, and is thus less prone to the associated training instabilities. Instead, our approach learns rewards by estimating task progress in a self-supervised manner. We demonstrate the effectiveness and efficiency of our approach on two contact-rich manipulation tasks, namely, peg-in-hole and USB insertion. The experimental results indicate that the policies trained with the learned reward function achieves better performance and faster convergence compared to the baselines. Zheng Wu 0002, Wenzhao Lian, Vaibhav V. Unhelkar, Masayoshi Tomizuka, Stefan Schaal |
ICRA | 3 |
| 2020 | Decision-Making for Bidirectional Communication in Sequential Human-Robot Collaborative TasksabstractCommunication is critical to collaboration; however, too much of it can degrade performance. Motivated by the need for effective use of a robot's communication modalities, in this work, we present a computational framework that decides if, when, and what to communicate during human-robot collaboration. The framework, titled CommPlan, consists of a model specification process and an execution-time POMDP planner. To address the challenge of collecting interaction data, the model specification process is hybrid : where part of the model is learned from data, while the remainder is manually specified. Given the model, the robot's decision-making is performed computationally during interaction and under partial observability of human's mental states. We implement CommPlan for a shared workspace task, in which the robot has multiple communication options and needs to reason within a short time. Through experiments with human participants, we confirm that CommPlan results in the effective use of communication capabilities and improves human-robot collaboration. Vaibhav V. Unhelkar, Shen Li 0003, Julie A. Shah |
HRI | 1 |
| 2019 | Learning Models of Sequential Decision-Making with Partial Specification of Agent BehaviorabstractArtificial agents that interact with other (human or artificial) agents require models in order to reason about those other agents’ behavior. In addition to the predictive utility of these models, maintaining a model that is aligned with an agent’s true generative model of behavior is critical for effective human-agent interaction. In applications wherein observations and partial specification of the agent’s behavior are available, achieving model alignment is challenging for a variety of reasons. For one, the agent’s decision factors are often not completely known; further, prior approaches that rely upon observations of agents’ behavior alone can fail to recover the true model, since multiple models can explain observed behavior equally well. To achieve better model alignment, we provide a novel approach capable of learning aligned models that conform to partial knowledge of the agent’s behavior. Central to our approach are a factored model of behavior (AMM), along with Bayesian nonparametric priors, and an inference approach capable of incorporating partial specifications as constraints for model learning. We evaluate our approach in experiments and demonstrate improvements in metrics of model alignment. Vaibhav V. Unhelkar, Julie A. Shah |
AAAI | 1 |
| 2018 | Learning and Communicating the Latent States of Human-Machine CollaborationabstractArtificial agents (both embodied robots and software agents) that interact with humans are increasing at an exceptional rate. Yet, achieving seamless collaboration between artificial agents and humans in the real world remains an active problem. A key challenge is that the agents need to make decisions without complete information about their shared environment and collaborators. For instance, a human-robot team performing a rescue operation after a disaster may not have an accurate map of their surroundings. Even in structured domains, such as manufacturing, a robot might not know the goals or preferences of its human collaborators. Algorithmically, this challenge manifests itself as a problem of decision-making under uncertainty in which the agent has to reason about the latent states of its environment and human collaborator. However, in practice, quantifying this uncertainty (i.e., the state transition function) and even specifying the features (i.e., the relevant states) of human-machine collaboration is difficult. Thus, the objective of this thesis research is to develop novel algorithms that enable artificial agents to learn and reason about the latent states of human-machine collaboration and achieve fluent interaction. Vaibhav V. Unhelkar, Julie A. Shah |
IJCAI | 1 |
| 2017 | Evaluating Effects of User Experience and System Transparency on Trust in AutomationabstractExisting research assessing human operators' trust in automation and robots has primarily examined trust as a steady-state variable, with little emphasis on the evolution of trust over time. With the goal of addressing this research gap, we present a study exploring the dynamic nature of trust. We defined trust of entirety as a measure that accounts for trust across a human's entire interactive experience with automation, and first identified alternatives to quantify it using real-time measurements of trust. Second, we provided a novel model that attempts to explain how trust of entirety evolves as a user interacts repeatedly with automation. Lastly, we investigated the effects of automation transparency on momentary changes of trust. Our results indicated that trust of entirety is better quantified by the average measure of "area under the trust curve" than the traditional post-experiment trust measure. In addition, we found that trust of entirety evolves and eventually stabilizes as an operator repeatedly interacts with a technology. Finally, we observed that a higher level of automation transparency may mitigate the "cry wolf" effect -- wherein human operators begin to reject an automated system due to repeated false alarms. Xi Jessie Yang, Vaibhav V. Unhelkar, Julie A. Shah |
HRI | 2 |
| 2016 | ConTaCT: Deciding to Communicate during Time-Critical Collaborative Tasks in Unknown, Deterministic DomainsabstractCommunication between agents has the potential to improve team performance of collaborative tasks. However, communication is not free in most domains, requiring agents to reason about the costs and benefits of sharing information. In this work, we develop an online, decentralized communication policy, ConTaCT, that enables agents to decide whether or not to communicate during time-critical collaborative tasks in unknown, deterministic environments. Our approach is motivated by real-world applications, including the coordination of disaster response and search and rescue teams. These settings motivate a model structure that explicitly represents the world model as initially unknown but deterministic in nature, and that de-emphasizes uncertainty about action outcomes. Simulated experiments are conducted in which ConTaCT is compared to other multi-agent communication policies, and results indicate that ConTaCT achieves comparable task performance while substantially reducing communication overhead. Vaibhav V. Unhelkar, Julie A. Shah |
AAAI | 1 |
| 2015 | Human-robot co-navigation using anticipatory indicators of human walking motionabstractMobile, interactive robots that operate in human-centric environments need the capability to safely and efficiently navigate around humans. This requires the ability to sense and predict human motion trajectories and to plan around them. In this paper, we present a study that supports the existence of statistically significant biomechanical turn indicators of human walking motions. Further, we demonstrate the effectiveness of these turn indicators as features in the prediction of human motion trajectories. Human motion capture data is collected with predefined goals to train and test a prediction algorithm. Use of anticipatory features results in improved performance of the prediction algorithm. Lastly, we demonstrate the closed-loop performance of the prediction algorithm using an existing algorithm for motion planning within dynamic environments. The anticipatory indicators of human walking motion can be used with different prediction and/or planning algorithms for robotics; the chosen planning and prediction algorithm demonstrates one such implementation for human-robot co-navigation. Vaibhav V. Unhelkar, Claudia Pérez-D'Arpino, Leia A. Stirling 0001, Julie A. Shah |
ICRA | 1 |
| 2014 | Comparative performance of human and mobile robotic assistants in collaborative fetch-and-deliver tasksabstractThere is an emerging desire across manufacturing industries to deploy robots that support people in their manual work, rather than replace human workers. This paper explores one such opportunity, which is to field a mobile robotic assistant that travels between part carts and the automotive final assembly line, delivering tools and materials to the human workers. We compare the performance of a mobile robotic assistant to that of a human assistant to gain a better understanding of the factors that impact its effectiveness. Statistically significant differences emerge based on type of assistant, human or robot. Interaction times and idle times are statistically significantly higher for the robotic assistant than the human assistant. We report additional differences in participant's subjective response regarding team fluency, situational awareness, comfort and safety. Finally, we discuss how results from the experiment inform the design of a more effective assistant. Vaibhav V. Unhelkar, Ho Chit Siu, Julie A. Shah |
HRI | 1 |
| 2014 | Towards control and sensing for an autonomous mobile robotic assistant navigating assembly linesabstractThere exists an increasing demand to incorporate mobile interactive robots to assist humans in repetitive, non-value added tasks in the manufacturing domain. Our aim is to develop a mobile robotic assistant for fetch-and-deliver tasks in human-oriented assembly line environments. Assembly lines present a niche yet novel challenge for mobile robots; the robot must precisely control its position on a surface which may be either stationary, moving, or split (e.g. in the case that the robot straddles the moving assembly line and remains partially on the stationary surface). In this paper we present a control and sensing solution for a mobile robotic assistant as it traverses a moving-floor assembly line. Solutions readily exist for control of wheeled mobile robots on static surfaces; we build on the open-source Robot Operating System (ROS) software architecture and generalize the algorithms for the moving line environment. Off-the-shelf sensors and localization algorithms are explored to sense the moving surface, and a customized solution is presented using PX4Flow optic flow sensors and a laser scanner-based localization algorithm. Validation of the control and sensing system is carried out both in simulation and in hardware experiments on a customized treadmill. Initial demonstrations of the hardware system yield promising results; the robot successfully maintains its position while on, and while straddling, the moving line. Vaibhav V. Unhelkar, Jorge Perez, Jim Boerkoel, Johannes Bix, Stefan Bartscher, Julie A. Shah |
ICRA | 1 |