VLDB 2026 Research / reviewers in the wild / expert
Julie A. Shah
dblp:30/6778 · also Julie Arnold Shah, Julie Shah
· DBLP profile ↗
89ranked-venue papers
2as first author
35since 2021 · last 2026
0000-0003-1338-8107ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 71 · 1 first-author · 28 since 2021Human-computer interaction and ubiquitous computing · 27 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 5 since 2021Systems, architecture and hardware · 17 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Contextually-Adaptive Rewards via Calibrated FeaturesabstractA key challenge in reward learning from human input is that desired agent behavior often changes based on context. For example, a robot must adapt to avoid a stove once it becomes hot. We observe that while high-level preferences (e.g., prioritizing safety over efficiency) often remain constant, context alters the saliency–or importance–of reward features. For instance, stove heat changes the relevance of the robot’s proximity, not the underlying preference for safety. Moreover, these contextual effects recur across tasks, motivating the need for transferable representations to encode them. Existing multi-task and meta-learning methods simultaneously learn representations and task preferences, at best implicitly capturing contextual effects and requiring substantial data to separate them from task-specific preferences. Instead, we propose explicitly modeling and learning context-dependent feature saliency separately from context-invariant preferences. We introduce calibrated features–modular representations that capture contextual effects on feature saliency–and present specialized paired comparison queries that isolate saliency from preference for efficient learning. Simulated experiments show our method improves sample efficiency, requiring 10x fewer preference queries than baselines to achieve equivalent reward accuracy, with up to 15% better performance in low-data regimes (5–10 queries). An in-person user study (N=12) demonstrates that participants can effectively teach their personal contextual preferences with our method, enabling adaptable and personalized reward learning. Alexandra Forsey-Smerek, Julie A. Shah, Andreea Bobu |
HRI | 2 |
| 2025 | Questioning the Robot: Using Human Non-verbal Cues to Estimate the Need for ExplanationsabstractAs black-box AI systems become increasingly complex, understanding when and how to provide explanations to users is crucial. Multimodal signals, such as facial expressions, offer novel insights into how frequently explanations should be given. This paper explores whether users’ facial features can help estimate the need for explanations in a collaborative robot task. We applied three state-of-the-art eXplainable AI (XAI) methods, addressing how, why, and what-if questions, explaining the robot's failure detection model. Each explanation type conveyed information differently: how-explanations described how the model functions, why-explanations prowided personalised insights into input-feature-related cues, and what-if-explanations explored alternative scenarios. In a mixed-design study ($\mathrm{N}=33$), participants performed a robot-assisted pick-and-place task, receiving different explanation types. Our results show that users responded differently to these explanations, with why-explanations being the most preferred and prompting closer alignment in facial expressions with the robot Contrary to expectations, what-if explanations led to the least alignment and required greater vocal effort. These findings demonstrate how non-verbal cues can guide the frequency and type of explanations (personalised or general) and further highlight the importance of model transparency in human-robot collaboration. Dimosthenis Kontogiorgos, Julie A. Shah |
HRI | 2 |
| 2025 | HDDLGym: A Tool for Studying Multi-Agent Hierarchical Problems Defined in HDDL with OpenAI GymabstractIn recent years, reinforcement learning (RL) methods have been widely tested using tools like OpenAI Gym, though many tasks in these environments could also benefit from hierarchical planning. However, there is a lack of a tool that enables seamless integration of hierarchical planning with RL. Hierarchical Domain Definition Language (HDDL), used in classical planning, introduces a structured approach well-suited for model-based RL to address this gap. To bridge this integration, we introduce HDDLGym, a Python-based tool that automatically generates OpenAI Gym environments from HDDL domains and problems. HDDLGym serves as a link between RL and hierarchical planning, supporting multi-agent scenarios and enabling collaborative planning among agents. This paper provides an overview of HDDLGym’s design and implementation, highlighting the challenges and design choices involved in integrating HDDL with the Gym interface, and applying RL policies to support hierarchical planning. We also provide detailed instructions and demonstrations for using the HDDLGym framework, including how to work with existing HDDL domains and problems from International Planning Competitions, exemplified by the Transport domain. Additionally, we offer guidance on creating new HDDL domains for multi-agent scenarios and demonstrate the practical use of HDDLGym in the Overcooked domain. By leveraging the advantages of HDDL and Gym, HDDLGym aims to be a valuable tool for studying RL in hierarchical planning, particularly in multi-agent contexts. Ngoc La, Ruaridh Mon-Williams, Julie A. Shah |
ICAPS | 3 |
| 2025 | Converting Spatial to Social: Using Persistent Homology to Understand Social GroupsabstractICMI ’25, Canberra, ACT, Australia Valerie K. Chen, Claire Liang, Julie A. Shah, Sean Andrist |
ICMI | 3 |
| 2025 | Inference-Time Policy Steering Through Human InteractionsabstractGenerative policies trained with human demonstrations can autonomously accomplish multimodal, longhorizon tasks. However, during inference, humans are often removed from the policy execution loop, limiting the ability to guide a pre-trained policy towards a specific sub-goal or trajectory shape among multiple predictions. Naive human intervention may inadvertently exacerbate distribution shift, leading to constraint violations or execution failures. To better align policy output with human intent without inducing out-of-distribution errors, we propose an Inference-Time Policy Steering (ITPS) framework that leverages human interactions to bias the generative sampling process, rather than finetuning the policy on interaction data. We evaluate ITPS across three simulated and real-world benchmarks, testing three forms of human interaction and associated alignment distance metrics. Among six sampling strategies, our proposed stochastic sampling with diffusion policy achieves the best trade-off between alignment and distribution shift. Videos are available at https://yanweiw.github.io/itps/. Lirui Wang, Yilun Du, Balakumar Sundaralingam, Xuning Yang, Yu-Wei Chao, Claudia Pérez-D'Arpino, Dieter Fox, Julie A. Shah |
ICRA | 9 |
| 2025 | Inference of Human-derived Specifications of Object Placement via DemonstrationabstractAs robots' manipulation capabilities improve for pick-and-place tasks (e.g., object packing, sorting, and kitting), methods focused on understanding human-acceptable object configurations remain limited expressively with regard to capturing spatial relationships important to humans. To advance robotic understanding of human rules for object arrangement, we introduce positionally-augmented RCC (PARCC), a formal logic framework based on region connection calculus (RCC) for describing the relative position of objects in space. Additionally, we introduce an inference algorithm for learning PARCC specifications via demonstrations. Finally, we present the results from a human study, which demonstrate our framework's ability to capture a human's intended specification and the benefits of learning from demonstration approaches over human-provided specifications. Alex Cuellar, Ho Chit Siu, Julie A. Shah |
IJCAI | 3 |
| 2025 | Versatile Demonstration Interface: Toward More Flexible Robot Demonstration CollectionabstractPrevious methods for Learning from Demonstration leverage several approaches for a human to teach motions to a robot, including teleoperation, kinesthetic teaching, and natural demonstrations. However, little previous work has explored more general interfaces that allow for multiple demonstration types. Given the varied preferences of human demonstrators and task characteristics, a flexible tool that enables multiple demonstration types could be crucial for broader robot skill training. In this work, we propose Versatile Demonstration Interface (VDI), an attachment for collaborative robots that simplifies the collection of three common types of demonstrations. Designed for flexible deployment in industrial settings, our tool requires no additional instrumentation of the environment. Our prototype interface captures human demonstrations through a combination of vision, force sensing, and state tracking (e.g., through the robot proprioception or AprilTag tracking). Through a user study where we deployed our prototype VDI at a local manufacturing innovation center with manufacturing experts, we demonstrated VDI in representative industrial tasks. Interactions from our study highlight the practical value of VDI’s varied demonstration types, expose a range of industrial use cases for VDI, and provide insights for future tool design. Michael Hagenow, Dimosthenis Kontogiorgos, Julie A. Shah |
IROS | 4 |
| 2025 | Analyzing Reluctance to Ask for Help When Cooperating With Robots: Insights to Integrate Artificial Agents in HRCabstractAs robot technology advances, collaboration between humans and robots will become more prevalent in industrial tasks. When humans run into issues in such scenarios, a likely future involves relying on artificial agents or robots for aid. This study identifies key aspects for the design of future user-assisting agents. We analyze quantitative and qualitative data from a user study examining the impact of on-demand assistance received from a remote human in a human-robot collaboration (HRC) assembly task. We study scenarios in which users require help and we assess their experiences in requesting and receiving assistance. Additionally, we investigate participants’ perceptions of future non-human assisting agents and whether assistance should be on-demand or unsolicited. Through a user study, we analyze the impact that such design decisions (human or artificial assistant, on-demand or unsolicited help) can have on elicited emotional responses, productivity, and preferences of humans engaged in HRC tasks. Ane San Martín, Michael Hagenow, Julie A. Shah, Johan Kildal, Elena Lazkano |
RO-MAN | 3 |
| 2025 | Corrections to "On-Manifold Strategies for Reactive Dynamical System Modulation With Nonconvex Obstacles"abstractReferences were removed from the final submission that were part of the accepted paper. There were also two duplicative references. Christopher K. Fourie, Nadia Figueroa, Julie A. Shah |
IEEE Trans. Robotics | 3 |
| 2024 | Understanding Entrainment in Human Groups: Optimising Human-Robot Collaboration from Lessons Learned during Human-Human CollaborationabstractSuccessful entrainment during collaboration positively affects trust, willingness to collaborate, and likeability towards collaborators. In this paper, we present a mixed-method study to investigate characteristics of successful entrainment leading to pair and group-based synchronisation. Drawing inspiration from industrial settings, we designed a fast-paced, short-cycle repetitive task. Using motion tracking, we investigated entrainment in both dyadic and triadic task completion. Furthermore, we utilise audio-video recordings and semi-structured interviews to contextualise participants’ experiences. This paper contributes to the Human-Computer/Robot Interaction (HCI/HRI) literature using a human-centred approach to identify characteristics of entrainment during pair- and group-based collaboration. We present five characteristics related to successful entrainment. These are related to the occurrence of entrainment, leader-follower patterns, interpersonal communication, the importance of the point-of-assembly, and the value of acoustic feedback. Finally, we present three design considerations for future research and design on collaboration with robots. Eike Schneiders, Christopher K. Fourie, Stanley Celestin, Julie A. Shah, Malte F. Jung |
CHI | 4 |
| 2024 | Aligning Human and Robot RepresentationsabstractTo act in the world, robots rely on a representation of salient task aspects: for example, to carry a coffee mug, a robot may consider movement efficiency or mug orientation in its behavior. However, if we want robots to act for and with people, their representations must not be just functional but also reflective of what humans care about, i.e. they must be aligned. We observe that current learning approaches suffer from representation misalignment, where the robot's learned representation does not capture the human's representation. We suggest that because humans are the ultimate evaluator of robot performance, we must explicitly focus our efforts on aligning learned representations with humans, in addition to learning the downstream task. We advocate that current representation learning approaches in robotics should be studied from the perspective of how well they accomplish the objective of representation alignment. We mathematically define the problem, identify its key desiderata, and situate current methods within this formalism. We conclude by suggesting future directions for exploring open challenges. Andreea Bobu, Andi Peng, Pulkit Agrawal 0001, Julie A. Shah, Anca D. Dragan |
HRI | 4 |
| 2024 | Preference-Conditioned Language-Guided AbstractionabstractLearning from demonstrations is a common way for users to teach robots, but it is prone to spurious feature correlations. Recent work constructs state abstractions, i.e. visual representations containing task-relevant features, from language as a way to perform more generalizable learning. However, these abstractions also depend on a user's preference for what matters in a task, which may be hard to describe or infeasible to exhaustively specify using language alone. How do we construct abstractions to capture these latent preferences? We observe that how humans behave reveals how they see the world. Our key insight is that changes in human behavior inform us that there are differences in preferences for how humans see the world, i.e. their state abstractions. In this work, we propose using language models (LMs) to query for those preferences directly given knowledge that a change in behavior has occurred. In our framework, we use the LM in two ways: first, given a text description of the task and knowledge of behavioral change between states, we query the LM for possible hidden preferences; second, given the most likely preference, we query the LM to construct the state abstraction. In this framework, the LM is also able to ask the human directly when uncertain about its own estimate. We demonstrate our framework's ability to construct effective preference-conditioned abstractions in simulated experiments, a user study, as well as on a real Spot robot performing mobile manipulation tasks. Andi Peng, Andreea Bobu, Belinda Z. Li, Theodore R. Sumers, Ilia Sucholutsky, Nishanth Kumar, Thomas L. Griffiths 0001, Julie A. Shah |
HRI | 8 |
| 2024 | Learning with Language-Guided State AbstractionsabstractWe describe a framework for using natural language to design state abstractions for imitation learning.
Generalizable policy learning in high-dimensional observation spaces is facilitated by well-designed state representations, which can surface important features of an environment and hide irrelevant ones.
These state representations are typically manually specified, or derived from other labor-intensive labeling procedures.
Our method, LGA (\textit{language-guided abstraction}), uses a combination of natural language supervision and background knowledge from language models (LMs) to automatically build state representations tailored to unseen tasks.
In LGA, a user first provides a (possibly incomplete) description of a target task in natural language; next, a pre-trained LM translates this task description into a state abstraction function that masks out irrelevant features; finally, an imitation policy is trained using a small number of demonstrations and LGA-generated abstract states.
Experiments on simulated robotic tasks show that LGA yields state abstractions similar to those designed by humans, but in a fraction of the time, and that these abstractions improve generalization and robustness in the presence of spurious correlations and ambiguous specifications.
We illustrate the utility of the learned abstractions on mobile manipulation tasks with a Spot robot. Andi Peng, Ilia Sucholutsky, Belinda Z. Li, Theodore R. Sumers, Thomas L. Griffiths 0001, Jacob Andreas, Julie A. Shah |
ICLR | 7 |
| 2024 | Grounding Language Plans in Demonstrations Through Counterfactual PerturbationsabstractGrounding the common-sense reasoning of Large Language Models in physical domains remains a pivotal yet unsolved problem for embodied AI. Whereas prior works have focused on leveraging LLMs directly for planning in symbolic spaces, this work uses LLMs to guide the search of task structures and constraints implicit in multi-step demonstrations. Specifically, we borrow from manipulation planning literature the concept of mode families, which group robot configurations by specific motion constraints, to serve as an abstraction layer between the high-level language representations of an LLM and the low-level physical trajectories of a robot. By replaying a few human demonstrations with synthetic perturbations, we generate coverage over the demonstrations' state space with additional successful executions as well as counterfactuals that fail the task. Our explanation-based learning framework trains an end-to-end differentiable neural network to predict successful trajectories from failures and as a by-product learns classifiers that ground low-level states and images in mode families without dense labeling. The learned grounding classifiers can further be used to translate language plans into reactive policies in the physical domain in an interpretable manner. We show our approach improves the interpretability and reactivity of imitation learning through 2D navigation and simulated and real robot manipulation tasks. Website: https://yanweiw.github.io/glide/ Tsun-Hsuan Wang, Jiayuan Mao, Michael Hagenow, Julie A. Shah |
ICLR | 5 |
| 2024 | Object Permanence Filter for Robust Tracking with Interactive RobotsabstractObject permanence, which refers to the concept that objects continue to exist even when they are no longer perceivable through the senses, is a crucial aspect of human cognitive development. In this work, we seek to incorporate this understanding into interactive robots by proposing a set of assumptions and rules to represent object permanence in multi-object, multi-agent interactive scenarios. We integrate these rules into the particle filter, resulting in the Object Permanence Filter (OPF). For multi-object scenarios, we propose an ensemble of K interconnected OPFs, where each filter predicts plausible object tracks that are resilient to missing, noisy, and kinematically or dynamically infeasible measurements, thus bringing perceptional robustness. Through several interactive scenarios, we demonstrate that the proposed OPF approach provides robust tracking in human-robot interactive tasks agnostic to measurement type, even in the presence of prolonged and complete occlusion. Webpage: https://opfilter.github.io/. Shaoting Peng, Margaret X. Wang, Julie A. Shah, Nadia Figueroa |
ICRA | 3 |
| 2024 | Enhancing Preference-based Linear Bandits via Human Response TimeabstractInteractive preference learning systems infer human preferences by presenting queries as pairs of options and collecting binary choices. Although binary choices are simple and widely used, they provide limited information about preference strength. To address this, we leverage human response times, which are inversely related to preference strength, as an additional signal. We propose a computationally efficient method that combines choices and response times to estimate human utility functions, grounded in the EZ diffusion model from psychology. Theoretical and empirical analyses show that for queries with strong preferences, response times complement choices by providing extra information about preference strength, leading to significantly improved utility estimation. We incorporate this estimator into preference-based linear bandits for fixed-budget best-arm identification. Simulations on three real-world datasets demonstrate that using response times significantly accelerates preference learning compared to choice-only approaches. Additional materials, such as code, slides, and talk video, are available at https://shenlirobot.github.io/pages/NeurIPS24.html. Shen Li 0003, Zhaolin Ren, Claire Liang, Na Li 0002, Julie A. Shah |
NeurIPS | 6 |
| 2024 | Experimental Assessment of Human-Robot Teaming for Multi-step Remote Manipulation with Expert OperatorsabstractRemote robot manipulation with human control enables applications in which safety and environmental constraints are adverse to humans (e.g., underwater, space robotics and disaster response) or the complexity of the task demands human-level cognition and dexterity (e.g., robotic surgery and manufacturing). These systems typically use direct teleoperation at the motion level and are usually limited to low-DOF arms and two-dimensional (2D) perception. Improving dexterity and situational awareness demands new interaction and planning workflows. We explore the use of human–robot teaming through teleautonomy with assisted planning for remote control of a dual-arm dexterous robot for multi-step manipulation, and conduct a within-subjects experimental assessment (n = 12 expert users) to compare it with direct teleoperation with an imitation controller with 2D and three-dimensional (3D) perception, as well as teleoperation through a teleautonomy interface. The proposed assisted planning approach achieves task times comparable with direct teleoperation while improving other objective and subjective metrics, including re-grasps, collisions, and TLX workload. Assisted planning in the teleautonomy interface achieves faster task execution and removes a significant interaction with the operator’s expertise level, resulting in a performance equalizer across users. Our study protocol, metrics, and models for statistical analysis might also serve as a general benchmarking framework in teleoperation domains. Accompanying video and reference R code: https://people.csail.mit.edu/cdarpino/THRIteleop/ Claudia Pérez-D'Arpino, Rebecca P. Khurshid, Julie A. Shah |
ACM Trans. Hum. Robot Interact. | 3 |
| 2024 | On-Manifold Strategies for Reactive Dynamical System Modulation With Nonconvex ObstaclesabstractIn this work, we present a novel, reactive, modulated control strategy based on dynamical systems (DS) for planning in the context of multiple non-convex obstacles. Our DS modulation strategy leverages an on-manifold planning methodology and provides several methods for real-time on-manifold navigation around non-convex obstacles. We introduce a sample-based obstacle representation for complex, non-convex obstacles, as well as a projection-based method for representing surfaces such as tables, cylinders, and ellipsoids. These representations can be combined to represent multiple obstacles and obstacle types (sample- or projection-based) with a single, continuously differentiable function. We validate our approach in several real-world scenarios, including navigation within (simulated) constrained environments, as well as reactive control of a real 7DoF manipulator with dynamic obstacles (including humans) while utilizing a 1 kHz control loop rate. Using our samplebased representation, we can calculate the obstacle representation function in less than 1 ms with up to 35k points using a CPU implementation, and up to 600k points with a GPU implementation. Christopher K. Fourie, Nadia Figueroa, Julie A. Shah |
IEEE Trans. Robotics | 3 |
| 2023 | The Perils of Trial-and-Error Reward Design: Misdesign through Overfitting and Invalid Task SpecificationsabstractIn reinforcement learning (RL), a reward function that aligns exactly with a task's true performance metric is often necessarily sparse. For example, a true task metric might encode a reward of 1 upon success and 0 otherwise. The sparsity of these true task metrics can make them hard to learn from, so in practice they are often replaced with alternative dense reward functions. These dense reward functions are typically designed by experts through an ad hoc process of trial and error. In this process, experts manually search for a reward function that improves performance with respect to the task metric while also enabling an RL algorithm to learn faster. This process raises the question of whether the same reward function is optimal for all algorithms, i.e., whether the reward function can be overfit to a particular algorithm. In this paper, we study the consequences of this wide yet unexamined practice of trial-and-error reward design. We first conduct computational experiments that confirm that reward functions can be overfit to learning algorithms and their hyperparameters. We then conduct a controlled observation study which emulates expert practitioners' typical experiences of reward design, in which we similarly find evidence of reward function overfitting. We also find that experts' typical approach to reward design---of adopting a myopic strategy and weighing the relative goodness of each state-action pair---leads to misdesign through invalid task specifications, since RL algorithms use cumulative reward rather than rewards for individual state-action pairs as an optimization target. Code, data: github.com/serenabooth/reward-design-perils Serena Booth, W. Bradley Knox, Julie A. Shah, Scott Niekum, Peter Stone 0001, Alessandro Allievi |
AAAI | 3 |
| 2023 | Towards Interpretable Deep Reinforcement Learning with Human-Friendly Prototypes
Eoin M. Kenny, Mycal Tucker, Julie A. Shah |
ICLR | 3 |
| 2023 | Diagnosis, Feedback, Adaptation: A Human-in-the-Loop Framework for Test-Time Policy AdaptationabstractPolicies often fail at test-time due to distribution shifts—changes in the state and reward that occur when an end user deploys the policy in environments different from those seen in training. Data augmentation can help models be more robust to such shifts by varying specific concepts in the state, e.g. object color, that are task-irrelevant and should not impact desired actions. However, designers training the agent don’t often know which concepts are irrelevant a priori. We propose a human-in-the-loop framework to leverage feedback from the end user to quickly identify and augment task-irrelevant visual state concepts. Our framework generates counterfactual demonstrations that allow users to quickly isolate shifted state concepts and identify if they should not impact the desired task, and can therefore be augmented using existing actions. We present experiments validating our full pipeline on discrete and continuous control tasks with real human users. Our method better enables users to (1) understand agent failure, (2) improve sample efficiency of demonstrations required for finetuning, and (3) adapt the agent to their desired reward. Andi Peng, Aviv Netanyahu, Mark K. Ho, Tianmin Shu, Andreea Bobu, Julie A. Shah, Pulkit Agrawal 0001 |
ICML | 6 |
| 2023 | Towards Collaborative Plan Acquisition through Theory of Mind Modeling in Situated DialogueabstractCollaborative tasks often begin with partial task knowledge and incomplete plans from each partner. To complete these tasks, partners need to engage in situated communication with their partners and coordinate their partial plans towards a complete plan to achieve a joint task goal. While such collaboration seems effortless in a human-human team, it is highly challenging for human-AI collaboration. To address this limitation, this paper takes a step towards Collaborative Plan Acquisition, where humans and agents strive to learn and communicate with each other to acquire a complete plan for joint tasks. Specifically, we formulate a novel problem for agents to predict the missing task knowledge for themselves and for their partners based on rich perceptual and dialogue history. We extend a situated dialogue benchmark for symmetric collaborative tasks in a 3D blocks world and investigate computational strategies for plan acquisition. Our empirical results suggest that predicting the partner's missing knowledge is a more viable approach than predicting one's own. We show that explicit modeling of the partner's dialogue moves and mental states produces improved and more stable results than without. These results provide insight for future AI agents that can predict what knowledge their partner is missing and, therefore, can proactively communicate such information to help the partner acquire such missing knowledge toward a common understanding of joint tasks. Cristian-Paul Bara, Ziqiao Ma 0001, Yingzhuo Yu, Julie A. Shah, Joyce Y. Chai |
IJCAI | 4 |
| 2023 | Human-Guided Complexity-Controlled AbstractionsabstractNeural networks often learn task-specific latent representations that fail to generalize to novel settings or tasks. Conversely, humans learn discrete representations (i.e., concepts or words) at a variety of abstraction levels (e.g., "bird" vs. "sparrow'") and use the appropriate abstraction based on tasks. Inspired by this, we train neural models to generate a spectrum of discrete representations, and control the complexity of the representations (roughly, how many bits are allocated for encoding inputs) by tuning the entropy of the distribution over representations. In finetuning experiments, using only a small number of labeled examples for a new task, we show that (1) tuning the representation to a task-appropriate complexity level supports the greatest finetuning performance, and (2) in a human-participant study, users were able to identify the appropriate complexity level for a downstream task via visualizations of discrete representations. Our results indicate a promising direction for rapid model finetuning by leveraging human insight. Andi Peng, Mycal Tucker, Eoin M. Kenny, Noga Zaslavsky, Pulkit Agrawal 0001, Julie A. Shah |
NeurIPS | 6 |
| 2022 | Do Feature Attribution Methods Correctly Attribute Features?abstractFeature attribution methods are popular in interpretable machine learning. These methods compute the attribution of each input feature to represent its importance, but there is no consensus on the definition of "attribution", leading to many competing methods with little systematic evaluation, complicated in particular by the lack of ground truth attribution. To address this, we propose a dataset modification procedure to induce such ground truth. Using this procedure, we evaluate three common methods: saliency maps, rationales, and attentions. We identify several deficiencies and add new perspectives to the growing body of evidence questioning the correctness and reliability of these methods applied on datasets in the wild. We further discuss possible avenues for remedy and recommend new attribution methods to be tested against ground truth before deployment. The code and appendix are available at https://yilunzhou.github.io/feature-attribution-evaluation/. Yilun Zhou, Serena Booth, Marco Túlio Ribeiro, Julie A. Shah |
AAAI | 4 |
| 2022 | Revisiting Human-Robot Teaching and Learning Through the Lens of Human Concept LearningabstractWhen interacting with a robot, humans form con-ceptual models (of varying quality) which capture how the robot behaves. These conceptual models form just from watching or in-teracting with the robot, with or without conscious thought. Some methods select and present robot behaviors to improve human conceptual model formation; nonetheless, these methods and HRI more broadly have not yet consulted cognitive theories of human concept learning. These validated theories offer concrete design guidance to support humans in developing conceptual models more quickly, accurately, and flexibly. Specifically, Analogical Transfer Theory and the Variation Theory of Learning have been successfully deployed in other fields, and offer new insights for the HRI community about the selection and presentation of robot behaviors. Using these theories, we review and contextualize 35 prior works in human-robot teaching and learning, and we assess how these works incorporate or omit the design implications of these theories. From this review, we identify new opportunities for algorithms and interfaces to help humans more easily learn conceptual models of robot behaviors, which in turn can help humans become more effective robot teachers and collaborators. Serena Booth, Sanjana Sharma, Sarah Chung, Julie A. Shah, Elena L. Glassman |
HRI | 4 |
| 2022 | Joint Action, Adaptation, and Entrainment in Human-Robot InteractionabstractResearch in joint action focuses on the psychological, neurological, and physical mechanisms by which humans collabo-rate with other agents, and overlaps with several domains related to human-robot interaction. The development of artificial systems that can support or emulate the requisite aspects of joint action could lead to improved human-robot team performance as well as improvements in subjective metrics (e.g., trust). This workshop highlights theoretical and technical considerations about human-robot joint action and real-time adaptation, with a particular focus on socio-motor entrainment, showing how the emulation of psychological mechanisms (e.g., emotion, intention signaling, mirroring) can lead to improved performance. We will invite speakers with backgrounds in robotics, neuroscience and psychol-ogy, as well as speakers with a focus in adjacent works, such as in human-robot coordinated dance, alignment, or synchronization. We will call for papers that utilize the theory of joint-action in an interactive human-robot context. We will also call for position papers on the application of the theory of joint action to robotics, with a heavy focus on psychological mechanisms that could potentially be emulated or adapted to a human-robot context. Participants will have the opportunity to brainstorm considerations and techniques that would be applicable to joint action inspired works through breakout sessions with the aim to lead to new and improved collaborations across fields. Christopher K. Fourie, Nadia Figueroa, Julie A. Shah, Marta Bienkiewicz, Benoît G. Bardy, Etienne Burdet, Phani-Teja Singamaneni, Rachid Alami 0001, Arianna Curioni, Günther Knoblich, Wafa Johal, Dagmar Sternad, Malte F. Jung |
HRI | 3 |
| 2022 | Prototype Based Classification from Hierarchy to FairnessabstractArtificial neural nets can represent and classify many types of high-dimensional data but are often tailored to particular applications – e.g., for “fair” or “hierarchical” classification. Once an architecture has been selected, it is often difficult for humans to adjust models for a new task; for example, a hierarchical classifier cannot be easily transformed into a fair classifier that shields a protected field. Our contribution in this work is a new neural network architecture, the concept subspace network (CSN), which generalizes existing specialized classifiers to produce a unified model capable of learning a spectrum of multi-concept relationships. We demonstrate that CSNs reproduce state-of-the-art results in fair classification when enforcing concept independence, may be transformed into hierarchical classifiers, or may even reconcile fairness and hierarchy within a single classifier. The CSN is inspired by and matches the performance of existing prototype-based classifiers that promote interpretability. Mycal Tucker, Julie A. Shah |
ICML | 2 |
| 2022 | When Does Syntax Mediate Neural Language Model Performance? Evidence from Dropout ProbesabstractMycal Tucker, Tiwalayo Eisape, Peng Qian, Roger Levy, Julie Shah. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Mycal Tucker, Tiwalayo Eisape, Roger Levy, Julie A. Shah |
NAACL-HLT | 5 |
| 2022 | ExSum: From Local Explanations to Model UnderstandingabstractInterpretability methods are developed to understand the working mechanisms of blackbox models, which is crucial to their responsible deployment.Fulfilling this goal requires both that the explanations generated by these methods are correct and that people can easily and reliably understand them.While the former has been addressed in prior work, the latter is often overlooked, resulting in informal model understanding derived from a handful of local explanations.In this paper, we introduce explanation summary (EXSUM), a mathematical framework for quantifying model understanding, and propose metrics for its quality assessment.On two domains, EXSUM highlights various limitations in the current practice, helps develop accurate model understanding, and reveals easily overlooked properties of the model.We also connect understandability to other properties of explanations such as human alignment, robustness, and counterfactual similarity and plausibility. Yilun Zhou, Marco Túlio Ribeiro, Julie A. Shah |
NAACL-HLT | 3 |
| 2022 | Trading off Utility, Informativeness, and Complexity in Emergent CommunicationabstractEmergent communication (EC) research often focuses on optimizing task-specific utility as a driver for communication. However, there is increasing evidence that human languages are shaped by task-general communicative constraints and evolve under pressure to optimize the Information Bottleneck (IB) tradeoff between the informativeness and complexity of the lexicon. Here, we integrate these two approaches by trading off utility, informativeness, and complexity in EC. To this end, we propose Vector-Quantized Variational Information Bottleneck (VQ-VIB), a method for training neural agents to encode inputs into discrete signals embedded in a continuous space. We evaluate our approach in multi-agent reinforcement learning settings and in color reference games and show that: (1) VQ-VIB agents can continuously adapt to changing communicative needs and, in the color domain, align with human languages; (2) the emergent VQ-VIB embedding spaces are semantically meaningful and perceptually grounded; and (3) encouraging informativeness leads to faster convergence rates and improved utility, both in VQ-VIB and in prior neural architectures for symbolic EC, with VQ-VIB achieving higher utility for any given complexity. This work offers a new framework for EC that is grounded in information-theoretic principles that are believed to characterize human language evolution and that may facilitate human-agent interaction. Mycal Tucker, Roger Levy, Julie A. Shah, Noga Zaslavsky |
NeurIPS | 3 |
| 2022 | The Situation Awareness Framework for Explainable AI (SAFE-AI) and Human Factors Considerations for XAI SystemsabstractRecent advances in artificial intelligence (AI) have drawn attention to the need for AI systems to be understandable to human users. The explainable AI (XAI) literature aims to enhance human understanding and human-AI team performance by providing users with necessary information about AI system behavior. Simultaneously, the human factors literature has long addressed important considerations that contribute to human performance, including how to determine human informational needs, human workload, and human trust in autonomous systems. Drawing from the human factors literature, we propose the Situation Awareness Framework for Explainable AI (SAFE-AI), a three-level framework for the development and evaluation of explanations about AI system behavior. Our proposed levels of XAI are based on the informational needs of human users, which can be determined using the levels of situation awareness (SA) framework from the human factors literature. Based on our levels of XAI framework, we also suggest a method for assessing the effectiveness of XAI systems. We further detail human workload considerations for determining the content and frequency of explanations as well as metrics that can be used to assess human workload. Finally, we discuss the importance of appropriately calibrating user trust in AI systems through explanations along with other trust-related considerations for XAI, and we detail metrics that can be used to evaluate user trust in these systems. Lindsay Sanneman, Julie A. Shah |
Int. J. Hum. Comput. Interact. | 2 |
| 2022 | Latent Space Alignment Using Adversarially Guided Self-PlayabstractWe envision a world in which robots serve as capable partners in heterogeneous teams composed of other robots or humans. A crucial step towards such a world is enabling robots to learn to use the same representations as their partners; with a shared representation scheme, information may be passed among teammates. We define the problem of learning a fixed partner’s representation scheme as that of latent space alignment and propose metrics for evaluating the quality of alignment. While techniques from prior art in other fields may be applied to the latent space alignment problem, they often require interaction with partners during training time or large amounts of training data. We developed a technique, Adversarially Guided Self-Play (ASP), that trains agents to solve the latent space alignment problem with little training data and no access to their pre-trained partners. Simulation results confirmed that, despite using less training data, agents trained by ASP aligned better with other agents than agents trained by other techniques. Subsequent human-participant studies involving hundreds of Amazon Mechanical Turk workers showed how laypeople understood our machines enough to perform well on team tasks and anticipate their machine partner’s successes or failures. Mycal Tucker, Yilun Zhou, Julie A. Shah |
Int. J. Hum. Comput. Interact. | 3 |
| 2021 | Bayes-TrEx: a Bayesian Sampling Approach to Model Transparency by ExampleabstractPost-hoc explanation methods are gaining popularity for interpreting, understanding, and debugging neural networks. Most analyses using such methods explain decisions in response to inputs drawn from the test set. However, the test set may have few examples that trigger some model behaviors, such as high-confidence failures or ambiguous classifications. To address these challenges, we introduce a flexible model inspection framework: Bayes-TrEx. Given a data distribution, Bayes-TrEx finds in-distribution examples which trigger a specified prediction confidence. We demonstrate several use cases of Bayes-TrEx, including revealing highly confident (mis)classifications, visualizing class boundaries via ambiguous examples, understanding novel-class extrapolation behavior, and exposing neural network overconfidence. We use Bayes-TrEx to study classifiers trained on CLEVR, MNIST, and Fashion-MNIST, and we show that this framework enables more flexible holistic model analysis than just inspecting the test set. Code and supplemental material are available at https://github.com/serenabooth/Bayes-TrEx. Serena Booth, Yilun Zhou, Ankit Shah 0003, Julie A. Shah |
AAAI | 4 |
| 2021 | Reactive Task and Motion Planning under Temporal Logic SpecificationsabstractWe present a task-and-motion planning (TAMP) algorithm robust against a human operator's cooperative or adversarial interventions. Interventions often invalidate the current plan and require replanning on the fly. Replanning can be computationally expensive and often interrupts seamless task execution. We introduce a dynamically reconfigurable planning methodology with behavior tree-based control strategies toward reactive TAMP, which takes the advantage of previous plans and incremental graph search during temporal logic-based reactive synthesis. Our algorithm also shows efficient recovery functionalities that minimize the number of replanning steps. Finally, our algorithm produces a robust, efficient, and complete TAMP solution. Our experimental results show the algorithm results in superior manipulation performance in both simulated and real-world tasks. Shen Li 0003, Daehyung Park, Yoonchang Sung, Julie A. Shah, Nicholas Roy |
ICRA | 4 |
| 2021 | Emergent Discrete Communication in Semantic SpacesabstractNeural agents trained in reinforcement learning settings can learn to communicate among themselves via discrete tokens, accomplishing as a team what agents would be unable to do alone. However, the current standard of using one-hot vectors as discrete communication tokens prevents agents from acquiring more desirable aspects of communication such as zero-shot understanding. Inspired by word embedding techniques from natural language processing, we propose neural agent architectures that enables them to communicate via discrete tokens derived from a learned, continuous space. We show in a decision theoretic framework that our technique optimizes communication over a wide range of scenarios, whereas one-hot tokens are only optimal under restrictive assumptions. In self-play experiments, we validate that our trained agents learn to cluster tokens in semantically-meaningful ways, allowing them communicate in noisy environments where other techniques fail. Lastly, we demonstrate both that agents using our method can effectively respond to novel human communication and that humans can understand unlabeled emergent agent communication, outperforming the use of one-hot communication. Mycal Tucker, Huao Li, Siddharth Agrawal, Dana Hughes 0001, Katia P. Sycara, Michael Lewis 0001, Julie A. Shah |
NeurIPS | 7 |
| 2020 | Decision-Making for Bidirectional Communication in Sequential Human-Robot Collaborative TasksabstractCommunication is critical to collaboration; however, too much of it can degrade performance. Motivated by the need for effective use of a robot's communication modalities, in this work, we present a computational framework that decides if, when, and what to communicate during human-robot collaboration. The framework, titled CommPlan, consists of a model specification process and an execution-time POMDP planner. To address the challenge of collecting interaction data, the model specification process is hybrid : where part of the model is learned from data, while the remainder is manually specified. Given the model, the robot's decision-making is performed computationally during interaction and under partial observability of human's mental states. We implement CommPlan for a shared workspace task, in which the robot has multiple communication options and needs to reason within a short time. Through experiments with human participants, we confirm that CommPlan results in the effective use of communication capabilities and improves human-robot collaboration. Vaibhav V. Unhelkar, Shen Li 0003, Julie A. Shah |
HRI | 3 |
| 2020 | Blind Spot Detection for Safe Sim-to-Real TransferabstractAgents trained in simulation may make errors when performing actions in the real world due to mismatches between training and execution environments. These mistakes can be dangerous and difficult for the agent to discover because the agent is unable to predict them a priori. In this work, we propose the use of oracle feedback to learn a predictive model of these blind spots in order to reduce costly errors in real-world applications. We focus on blind spots in reinforcement learning (RL) that occur due to incomplete state representation: when the agent lacks necessary features to represent the true state of the world, and thus cannot distinguish between numerous states. We formalize the problem of discovering blind spots in RL as a noisy supervised learning problem with class imbalance. Our system learns models for predicting blind spots within unseen regions of the state space by combining techniques for label aggregation, calibration, and supervised learning. These models take into consideration noise emerging from different forms of oracle feedback, including demonstrations and corrections. We evaluate our approach across two domains and demonstrate that it achieves higher predictive performance than baseline methods, and also that the learned model can be used to selectively query an oracle at execution time to prevent errors. We also empirically analyze the biases of various feedback types and how these biases influence the discovery of blind spots. Further, we include analyses of our approach that incorporate relaxed initial optimality assumptions. (Interestingly, relaxing the assumptions of an optimal oracle and an optimal simulator policy helped our models to perform better.) We also propose extensions to our method that are intended to improve performance when using corrections and demonstrations data. Ramya Ramakrishnan, Ece Kamar, Debadeepta Dey, Eric Horvitz, Julie A. Shah |
J. Artif. Intell. Res. | 5 |
| 2019 | Overcoming Blind Spots in the Real World: Leveraging Complementary Abilities for Joint ExecutionabstractSimulators are being increasingly used to train agents before deploying them in real-world environments. While training in simulation provides a cost-effective way to learn, poorly modeled aspects of the simulator can lead to costly mistakes, or blind spots. While humans can help guide an agent towards identifying these error regions, humans themselves have blind spots and noise in execution. We study how learning about blind spots of both can be used to manage hand-off decisions when humans and agents jointly act in the real-world in which neither of them are trained or evaluated fully. The formulation assumes that agent blind spots result from representational limitations in the simulation world, which leads the agent to ignore important features that are relevant for acting in the open world. Our approach for blind spot discovery combines experiences collected in simulation with limited human demonstrations. The first step applies imitation learning to demonstration data to identify important features that the human is using but that the agent is missing. The second step uses noisy labels extracted from action mismatches between the agent and the human across simulation and demonstration data to train blind spot models. We show through experiments on two domains that our approach is able to learn a succinct representation that accurately captures blind spot regions and avoids dangerous errors in the real world through transfer of control between the agent and the human. Ramya Ramakrishnan, Ece Kamar, Besmira Nushi, Debadeepta Dey, Julie A. Shah, Eric Horvitz |
AAAI | 5 |
| 2019 | Learning Models of Sequential Decision-Making with Partial Specification of Agent BehaviorabstractArtificial agents that interact with other (human or artificial) agents require models in order to reason about those other agents’ behavior. In addition to the predictive utility of these models, maintaining a model that is aligned with an agent’s true generative model of behavior is critical for effective human-agent interaction. In applications wherein observations and partial specification of the agent’s behavior are available, achieving model alignment is challenging for a variety of reasons. For one, the agent’s decision factors are often not completely known; further, prior approaches that rely upon observations of agents’ behavior alone can fail to recover the true model, since multiple models can explain observed behavior equally well. To achieve better model alignment, we provide a novel approach capable of learning aligned models that conform to partial knowledge of the agent’s behavior. Central to our approach are a factored model of behavior (AMM), along with Bayesian nonparametric priors, and an inference approach capable of incorporating partial specifications as constraints for model learning. We evaluate our approach in experiments and demonstrate improvements in metrics of model alignment. Vaibhav V. Unhelkar, Julie A. Shah |
AAAI | 2 |
| 2019 | Consider the Human Work Experience When Integrating Robotics in the WorkplaceabstractWorldwide, manufacturers are reimagining the future of their workforce and its connection to technology. Rather than replacing humans, Industry 5.0 explores how humans and robots can best complement one another's unique strengths. However, realizing this vision requires an in-depth understanding of how workers view the positive and negative attributes of their jobs, and the place of robots within it. In this paper, we explore the relationship between work attributes and automation goals by engaging in field research at a manufacturing plant. We conducted 50 face-to-face interviews with assembly-line workers (n=50), which we analyzed using discourse analysis and social constructivist methods. We found that the work attributes deemed most positive by participants include social interaction, movement and exercise, (human) autonomy, problem solving, task variety, and building with their hands. The main negative work attributes included health and safety issues, feeling rushed, and repetitive work. We identified several ways robots could help reduce negative work attributes and enhance positive ones, such as reducing work interruptions and cultivating physical and psychological well-being. Based on our findings, we created a set of integration considerations for organizations planning to deploy robotics technology, and discuss how the manufacturing and HRI communities can explore these ideas in the future. Katherine S. Welfare, Matthew R. Hallowell, Julie A. Shah, Laurel D. Riek |
HRI | 3 |
| 2019 | Fast Online Segmentation of Activities from Partial TrajectoriesabstractAugmenting a robot with the capacity to understand the activities of the people it collaborates with in order to then label and segment those activities allows the robot to generate an efficient and safe plan for performing its own actions. In this work, we introduce an online activity segmentation algorithm that can detect activity segments by processing a partial trajectory. We model the transitions through activities as a hidden Markov model, which runs online by implementing an efficient particle-filtering approach to infer the maximum a posteriori estimate of the activity sequence. This process is complemented by an online search process to refine activity segments using task model information about the partial order of activities. We evaluated our algorithm by comparing its performance to two state-of-the-art activity segmentation algorithms on three human activity datasets. The proposed algorithm improved activity segmentation accuracy across all three datasets compared with the other two approaches, with a range from 11.3% to 65.5%, and could accurately recognize an activity through observation alone for 31.6% of the initial trajectory of that activity, on average. We also implemented the algorithm onto an industrial mobile robot during an automotive assembly task in which the robot tracked a human worker's progress and provided the worker with the correct materials at the appropriate time. Tariq Iqbal, Shen Li 0003, Christopher K. Fourie, Bradley Hayes, Julie A. Shah |
ICRA | 5 |
| 2019 | Activity recognition in manufacturing: The roles of motion capture and sEMG+inertial wearables in detecting fine vs. gross motionabstractIn safety-critical environments, robots need to reliably recognize human activity to be effective and trust-worthy partners. Since most human activity recognition (HAR) approaches rely on unimodal sensor data (e.g. motion capture or wearable sensors), it is unclear how the relationship between the sensor modality and motion granularity (e.g. gross or fine) of the activities impacts classification accuracy. To our knowledge, we are the first to investigate the efficacy of using motion capture as compared to wearable sensor data for recognizing human motion in manufacturing settings. We introduce the UCSD-MIT Human Motion dataset, composed of two assembly tasks that entail either gross or fine-grained motion. For both tasks, we compared the accuracy of a Vicon motion capture system to a Myo armband using three widely used HAR algorithms. We found that motion capture yielded higher accuracy than the wearable sensor for gross motion recognition (up to 36.95%), while the wearable sensor yielded higher accuracy for fine-grained motion (up to 28.06%). These results suggest that these sensor modalities are complementary, and that robots may benefit from systems that utilize multiple modalities to simultaneously, but independently, detect gross and fine-grained motion. Our findings will help guide researchers in numerous fields of robotics including learning from demonstration and grasping to effectively choose sensor modalities that are most suitable for their applications. Alyssa Kubota, Tariq Iqbal, Julie A. Shah, Laurel D. Riek |
ICRA | 3 |
| 2019 | Safe and Efficient High Dimensional Motion Planning in Space-Time with Time Parameterized PredictionabstractIn this work, we propose an algorithm that can plan safe and efficient robot trajectories in real time, given time-parameterized motion predictions, in order to avoid fast-moving obstacles in human-robot collaborative environments. Our algorithm is able to reduce the robot configuration space and the time domain significantly by constructing a Lazy Safe Interval Probabilistic Roadmap based on a pre-planned path. The algorithm then plans efficient obstacle-avoidance strategies within the space-time roadmap. We benchmarked our algorithm by evaluating the performance of a simulated 6-joint manipulator attempting to avoid a quickly moving human hand, using a dataset collected from human experiments. We compared our algorithm's performance with those of 8 variations of prior state-of-the-art planners. Results from this empirical evaluation indicate that our method generated safe plans in 97.5% of the evaluated situations, achieved a planning speed 30 times faster than the benchmarked methods that planned in the time domain without space reduction, and accomplished the minimal solution execution time among the benchmarked planners with a similar planning speed. Shen Li 0003, Julie A. Shah |
ICRA | 2 |
| 2019 | Evaluating the Interpretability of the Knowledge Compilation Map: Communicating Logical Statements EffectivelyabstractKnowledge compilation techniques translate propositional theories into equivalent forms to increase their computational tractability. But, how should we best present these propositional theories to a human? We analyze the standard taxonomy of propositional theories for relative interpretability across three model domains: highway driving, emergency triage, and the chopsticks game. We generate decision-making agents which produce logical explanations for their actions and apply knowledge compilation to these explanations. Then, we evaluate how quickly, accurately, and confidently users comprehend the generated explanations. We find that domain, formula size, and negated logical connectives significantly affect comprehension while formula properties typically associated with interpretability are not strong predictors of human ability to comprehend the theory. Serena Booth, Christian J. Muise, Julie A. Shah |
IJCAI | 3 |
| 2019 | Bayesian Inference of Linear Temporal Logic Specifications for Contrastive ExplanationsabstractTemporal logics are useful for providing concise descriptions of system behavior, and have been successfully used as a language for goal definitions in task planning. Prior works on inferring temporal logic specifications have focused on "summarizing" the input dataset - i.e., finding specifications that are satisfied by all plan traces belonging to the given set. In this paper, we examine the problem of inferring specifications that describe temporal differences between two sets of plan traces. We formalize the concept of providing such contrastive explanations, then present BayesLTL - a Bayesian probabilistic model for inferring contrastive explanations as linear temporal logic (LTL) specifications. We demonstrate the robustness and scalability of our model for inferring accurate specifications from noisy data and across various benchmark planning domains. Joseph Kim, Christian J. Muise, Ankit Shah 0003, Shubham Agarwal 0002, Julie A. Shah |
IJCAI | 5 |
| 2019 | A Taxonomy for Characterizing Modes of Interactions in Goal-driven, Human-robot TeamsabstractAs robots and other autonomous agents are increasingly incorporated into complex domains, characterizing interaction within heterogeneous teams that include both humans and machines becomes more necessary. Previous literature has addressed the task of characterizing human-robot interaction from different perspectives and in multiple contexts. However, the numerous factors behind interaction work in conjunction, and the insights gained from one perspective can inadvertently affect another, creating a need for unification of these taxonomies and frameworks within an overarching taxonomy that systematically defines these relationships. In this paper we review existing taxonomies related to human-robot interaction, the behavioral sciences, and social and algorithmic taxonomies, and propose an overarching ontology for the factors from these works. We identify three main components characterizing the structure of an interaction (environment, task, and team), and structure them over two levels: contextual factors and factors driven by local dynamics. Finally, we present an analysis of how these factors affect decisions about levels of robot automation and level of information abstraction in an interaction, and discuss curent gaps in the literature that can motivate future research. Priyam Parashar, Lindsay Sanneman, Julie A. Shah, Henrik I. Christensen |
IROS | 3 |
| 2019 | Predicting ConceptNet Path Quality Using Crowdsourced Assessments of NaturalnessabstractIn many applications, it is important to characterize the way in which two concepts are semantically related. Knowledge graphs such as ConceptNet provide a rich source of information for such characterizations by encoding relations between concepts as edges in a graph. When two concepts are not directly connected by an edge, their relationship can still be described in terms of the paths that connect them. Unfortunately, many of these paths are uninformative and noisy, which means that the success of applications that use such path features crucially relies on their ability to select high-quality paths. In existing applications, this path selection process is based on relatively simple heuristics. In this paper we instead propose to learn to predict path quality from crowdsourced human assessments. Since we are interested in a generic task-independent notion of quality, we simply ask human participants to rank paths according to their subjective assessment of the paths' naturalness, without attempting to define naturalness or steering the participants towards particular indicators of quality. We show that a neural network model trained on these assessments is able to predict human judgments on unseen paths with near optimal performance. Most notably, we find that the resulting path selection method is substantially better than the current heuristic approaches at identifying meaningful paths. Yilun Zhou, Steven Schockaert, Julie A. Shah |
WWW | 3 |
| 2018 | Learning to Infer Final Plans in Human Team PlanningabstractWe envision an intelligent agent that analyzes conversations during human team meetings in order to infer the team’s plan, with the purpose of providing decision support to strengthen that plan. We present a novel learning technique to infer teams' final plans directly from a processed form of their planning conversation. Our method employs reinforcement learning to train a model that maps features of the discussed plan and patterns of dialogue exchange among participants to a final, agreed-upon plan. We employ planning domain models to efficiently search the large space of possible plans, and the costs of candidate plans serve as the reinforcement signal. We demonstrate that our technique successfully infers plans within a variety of challenging domains, with higher accuracy than prior art. With our domain-independent feature set, we empirically demonstrate that our model trained on one planning domain can be applied to successfully infer team plans within a novel planning domain. Joseph Kim, Matthew E. Woicik, Matthew C. Gombolay, Sung-Hyun Son, Julie A. Shah |
IJCAI | 5 |
| 2018 | Learning and Communicating the Latent States of Human-Machine CollaborationabstractArtificial agents (both embodied robots and software agents) that interact with humans are increasing at an exceptional rate. Yet, achieving seamless collaboration between artificial agents and humans in the real world remains an active problem. A key challenge is that the agents need to make decisions without complete information about their shared environment and collaborators. For instance, a human-robot team performing a rescue operation after a disaster may not have an accurate map of their surroundings. Even in structured domains, such as manufacturing, a robot might not know the goals or preferences of its human collaborators. Algorithmically, this challenge manifests itself as a problem of decision-making under uncertainty in which the agent has to reason about the latent states of its environment and human collaborator. However, in practice, quantifying this uncertainty (i.e., the state transition function) and even specifying the features (i.e., the relevant states) of human-machine collaboration is difficult. Thus, the objective of this thesis research is to develop novel algorithms that enable artificial agents to learn and reason about the latent states of human-machine collaboration and achieve fluent interaction. Vaibhav V. Unhelkar, Julie A. Shah |
IJCAI | 2 |
| 2018 | Modality Switching for Mitigation of Sensory Adaptation and Habituation in Personal Navigation SystemsabstractWhen humans need to navigate across terrain accurately and quickly, they often use portable electronic navigation systems for directional guidance. Prior work in this field has focused on selecting either the visual or haptic sensory modality for providing such guidance and has indicated that either option may be preferable depending on the user's specific goals. However, basing the selection of visual or haptic guidance on static criteria of this type discounts important time-varying effects, primarily stimulus-specific adaptation (SSA) and habituation. Here, we propose a navigation system design that mitigates these detrimental effects by periodically switching between visual and haptic navigation guidance. While this is likely to incur an undesirable switching cost, we hypothesize that the long-term benefits of counteracting SSA and habituation will outweigh this cost. In this paper, we describe the design and results of a human-participant study intended to evaluate this hypothesis. Our findings indicate that modality switching results in a transient cost to performance, but also that switching modalities lessens the SSA and habituation effects over time as compared with single-modality systems. The results support the hypothesis that an alternating-modality system would outperform a single-modality system for long-duration navigation tasks. Kyle Kotowick, Julie A. Shah |
IUI | 2 |
| 2018 | Bayesian Inference of Temporal Task Specifications from DemonstrationsabstractWhen observing task demonstrations, human apprentices are able to identify whether a given task is executed correctly long before they gain expertise in actually performing that task. Prior research into learning from demonstrations (LfD) has failed to capture this notion of the acceptability of an execution; meanwhile, temporal logics provide a flexible language for expressing task specifications. Inspired by this, we present Bayesian specification inference, a probabilistic model for inferring task specification as a temporal logic formula. We incorporate methods from probabilistic programming to define our priors, along with a domain-independent likelihood function to enable sampling-based inference. We demonstrate the efficacy of our model for inferring true specifications with over 90% similarity between the inferred specification and the ground truth, both within a synthetic domain and a real-world table setting task. Ankit Shah 0003, Pritish Kamath, Julie A. Shah, Shen Li 0003 |
NeurIPS | 3 |
| 2018 | Effects of an Adaptive Modality Selection Algorithm for Navigation SystemsabstractPortable electronic navigation systems are often used for directional guidance when humans need to navigate terrain quickly and accurately. Prior work in this field has focused on using either the visual or haptic sensory modality for providing such guidance, and results have indicated that either option may be preferable depending upon the user's specific needs. However, conventional methods involve selecting a single modality based on which will work best with the task the user is most likely to perform and using this modality throughout the duration of the navigation. In this paper, we describe the design and results of a study intended to evaluate the effectiveness of an adaptive modality selection algorithm that dynamically selects a navigation system's directional guidance modality while considering both task-specific benefits and the time-varying effects of switching cost, stimulus-specific adaptation, and habituation. Our findings indicate that use of this algorithm can improve user performance in the presence of multiple simultaneous tasks. Kyle Kotowick, Julie A. Shah |
UIST | 2 |
| 2018 | Human-Machine Collaborative Optimization via Apprenticeship SchedulingabstractCoordinating agents to complete a set of tasks with intercoupled temporal and resource constraints is computationally challenging, yet human domain experts can solve these difficult scheduling problems using paradigms learned through years of apprenticeship. A process for manually codifying this domain knowledge within a computational framework is necessary to scale beyond the "single-expert, single-trainee" apprenticeship model. However, human domain experts often have difficulty describing their decision-making processes. We propose a new approach for capturing this decision-making process through counterfactual reasoning in pairwise comparisons. Our approach is model-free and does not require iterating through the state space. We demonstrate that this approach accurately learns multifaceted heuristics on a synthetic and real world data sets. We also demonstrate that policies learned from human scheduling demonstration via apprenticeship learning can substantially improve the efficiency of schedule optimization. We employ this human-machine collaborative optimization technique on a variant of the weapon-to-target assignment problem. We demonstrate that this technique generates optimal solutions up to 9.5 times faster than a state-of-the-art optimization algorithm. Matthew C. Gombolay, Reed Jensen, Jessica Stigile, Toni Golen, Sung-Hyun Son, Julie A. Shah |
J. Artif. Intell. Res. | 7 |
| 2018 | Optimizing Makespan and Ergonomics in Integrating Collaborative Robots Into Manufacturing ProcessesabstractAs collaborative robots begin to appear on factory floors, there is a need to consider how these robots can best help their human partners. In this paper, we propose an optimization framework that generates task assignments and schedules for a human-robot team with the goal of improving both time and ergonomics and demonstrate its use in six real-world manufacturing processes that are currently performed manually. Using the strain index method to quantify human physical stress, we create a set of solutions with assigned priorities on each goal. The resulting schedules provide engineers with insight into selecting the appropriate level of compromise and integrating the robot in a way that best fits the needs of an individual process. Margaret Pearce, Bilge Mutlu, Julie A. Shah, Robert G. Radwin |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2018 | Planning for Manipulation of Interlinked Deformable Linear Objects With Applications to Aircraft AssemblyabstractManipulation of deformable linear objects (DLOs) has potential applications in the fields of aerospace and automotive assembly. In this paper, we introduce a problem formulation for attaching a set of interlinked DLOs to a support structure using a set of clamping points. The formulation describes the manipulation planning problem in terms of known clamp poses; predetermined ideal clamping locations on the cables, called “reference points;” and a set of finite gripping points on the DLOs. We also present a prototype algorithm that generates a solution in terms of primitive manipulation actions. The algorithm guarantees that no interlink constraints are violated at any stage of manipulation. We incorporate gravity in the computation of a DLO shape and propose a property linking geometrically similar cable shapes across the space of cable length and stiffness. This property allows for the computation of solutions for unit length and scaling of these solutions to appropriate length, potentially resulting in faster shape computation. Ankit Shah 0003, Lotta Blumberg, Julie A. Shah |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2018 | Fast Scheduling of Robot Teams Performing Tasks With Temporospatial ConstraintsabstractThe application of robotics to traditionally manual manufacturing processes requires careful coordination between human and robotic agents in order to support safe and efficient coordinated work. Tasks must be allocated to agents and sequenced according to temporal and spatial constraints. Also, systems must be capable of responding on-the-fly to disturbances and people working in close physical proximity to robots. In this paper, we present a centralized algorithm, named “Tercio,” that handles tightly intercoupled temporal and spatial constraints. Our key innovation is a fast, satisficing multi-agent task sequencer inspired by real-time processor scheduling techniques and adapted to leverage a hierarchical problem structure. We use this sequencer in conjunction with a mixed-integer linear program solver and empirically demonstrate the ability to generate near-optimal schedules for real-world problems an order of magnitude larger than those reported in prior art. Finally, we demonstrate the use of our algorithm in a multirobot hardware testbed. Matthew C. Gombolay, Ronald Wilcox, Julie A. Shah |
IEEE Trans. Robotics | 3 |
| 2017 | Collaborative Planning with Encoding of Users' High-Level StrategiesabstractThe generation of near-optimal plans for multi-agent systems with numerical states and temporal actions is computationally challenging. Current off-the-shelf planners can take a very long time before generating a near-optimal solution. In an effort to reduce plan computation time, increase the quality of the resulting plans, and make them more interpretable by humans, we explore collaborative planning techniques that actively involve human users in plan generation. Specifically, we explore a framework in which users provide high-level strategies encoded as soft preferences to guide the low-level search of the planner. Through human subject experimentation, we empirically demonstrate that this approach results in statistically significant improvements to plan quality, without substantially increasing computation time. We also show that the resulting plans achieve greater similarity to those generated by humans with regard to the produced sequences of actions, as compared to plans that do not incorporate user-provided strategies. Joseph Kim, Christopher J. Banks, Julie A. Shah |
AAAI | 3 |
| 2017 | Improving Robot Controller Transparency Through Autonomous Policy ExplanationabstractShared expectations and mutual understanding are critical facets of teamwork. Achieving these in human-robot collaborative contexts can be especially challenging, as humans and robots are unlikely to share a common language to convey intentions, plans, or justifications. Even in cases where human co-workers can inspect a robot's control code, and particularly when statistical methods are used to encode control policies, there is no guarantee that meaningful insights into a robot's behavior can be derived or that a human will be able to efficiently isolate the behaviors relevant to the interaction. We present a series of algorithms and an accompanying system that enables robots to autonomously synthesize policy descriptions and respond to both general and targeted queries by human collaborators. We demonstrate applicability to a variety of robot controller types including those that utilize conditional logic, tabular reinforcement learning, and deep reinforcement learning, synthesizing informative policy descriptions for collaborators and facilitating fault diagnosis by non-experts. Bradley Hayes, Julie A. Shah |
HRI | 2 |
| 2017 | Evaluating Effects of User Experience and System Transparency on Trust in AutomationabstractExisting research assessing human operators' trust in automation and robots has primarily examined trust as a steady-state variable, with little emphasis on the evolution of trust over time. With the goal of addressing this research gap, we present a study exploring the dynamic nature of trust. We defined trust of entirety as a measure that accounts for trust across a human's entire interactive experience with automation, and first identified alternatives to quantify it using real-time measurements of trust. Second, we provided a novel model that attempts to explain how trust of entirety evolves as a user interacts repeatedly with automation. Lastly, we investigated the effects of automation transparency on momentary changes of trust. Our results indicated that trust of entirety is better quantified by the average measure of "area under the trust curve" than the traditional post-experiment trust measure. In addition, we found that trust of entirety evolves and eventually stabilizes as an operator repeatedly interacts with a technology. Finally, we observed that a higher level of automation transparency may mitigate the "cry wolf" effect -- wherein human operators begin to reject an automated system due to repeated false alarms. Xi Jessie Yang, Vaibhav V. Unhelkar, Julie A. Shah |
HRI | 4 |
| 2017 | Interpretable models for fast activity recognition and anomaly explanation during collaborative robotics tasksabstractIn this paper, we present Rapid Activity Prediction Through Object-oriented Regression (RAPTOR), a scalable method for performing rapid, real-time activity recognition and prediction that achieves state-of-the-art classification accuracy on both a generic human activity dataset and two domain-specific collaborative robotics manufacturing datasets. Our approach is designed to be human-interpretable: able to provide explanations for its reasoning such that non-experts can better understand and improve its activity models. We incorporate methods to increase RAPTOR's resilience against confusion due to temporal variations, as well as against learning false correlations between features. We report full and partial trajectory classification results across three datasets and conclude by demonstrating our model's ability to provide interpretable explanations of its reasoning using outlier detection techniques. Bradley Hayes, Julie A. Shah |
ICRA | 2 |
| 2017 | A multiple-predictor approach to human motion predictionabstractThe ability to accurately predict human motion is imperative for any human-robot interaction application in which the human and robot interact in close proximity to one another. Although a variety of human motion prediction approaches have already been developed, they are often designed for specific types of tasks or motions, and thus do not generalize well. Furthermore, it is not always obvious which of these methods is appropriate for a given task, making human motion prediction difficult to implement in practice. We address this problem by introducing a multiple-predictor system (MPS) for human motion prediction. In our approach, the system learns directly from task data in order to determine the most favorable parameters for each implemented prediction method and which combination of these predictors to use. Our implementation consists of three complementary methods: velocity-based position projection, time series classification, and sequence prediction. We describe the process of forming the MPS and our evaluation of its performance against the individual methods in terms of accuracy of predictions of human position over a range of look-ahead time values. We report that our method leads to a reduction in mean error of 18.5%, 28.9%, and 37.3% when compared with the three individual methods, respectively. Przemyslaw A. Lasota, Julie A. Shah |
ICRA | 2 |
| 2017 | C-LEARN: Learning geometric constraints from demonstrations for multi-step manipulation in shared autonomyabstractLearning from demonstrations has been shown to be a successful method for non-experts to teach manipulation tasks to robots. These methods typically build generative models from demonstrations and then use regression to reproduce skills. However, this approach has limitations to capture hard geometric constraints imposed by the task. On the other hand, while sampling and optimization-based motion planners exist that reason about geometric constraints, these are typically carefully hand-crafted by an expert. To address this technical gap, we contribute with C-LEARN, a method that learns multi-step manipulation tasks from demonstrations as a sequence of keyframes and a set of geometric constraints. The system builds a knowledge base for reaching and grasping objects, which is then leveraged to learn multi-step tasks from a single demonstration. C-LEARN supports multi-step tasks with multiple end effectors; reasons about SE(3) volumetric and CAD constraints, such as the need for two axes to be parallel; and offers a principled way to transfer skills between robots with different kinematics. We embed the execution of the learned tasks within a shared autonomy framework, and evaluate our approach by analyzing the success rate when performing physical tasks with a dual-arm Optimas robot, comparing the contribution of different constraints models, and demonstrating the ability of C-LEARN to transfer learned tasks by performing them with a legged dual-arm Atlas robot in simulation. Claudia Pérez-D'Arpino, Julie A. Shah |
ICRA | 2 |
| 2017 | Intelligent Sensory Modality Selection for Electronic Supportive DevicesabstractHumans operating in stressful environments, such as in military or emergency first-responder roles, are subject to high sensory input loads and must often switch their attention between different modalities. Conventional supportive devices that assist users in such situations typically provide information using a single, static sensory modality; however, this carries the risk of overload when the modalities for the primary task and the supportive device overlap. Effective feedback modality selection is essential in order to avoid such a risk. One potential method for accomplishing this is to intelligently select the supportive device's feedback modality based on the user's environment and given task; however, this may result in delayed or lost information due to the performance cost resulting from switching attention from one modality to another. This paper describes the design and results of a human-participant study designed to evaluate the benefits and risks of various intelligent modality-selection strategies. Our findings suggest complex interactions between strategies, sensory input load levels and feedback modalities, with numerous significant effects across many different performance metrics. Kyle Kotowick, Julie A. Shah |
IUI | 2 |
| 2017 | Perturbation Training for Human-Robot TeamsabstractIn this work, we design and evaluate a computational learning model that enables a human-robot team to co-develop joint strategies for performing novel tasks that require coordination. The joint strategies are learned through "perturbation training," a human team-training strategy that requires team members to practice variations of a given task to help their team generalize to new variants of that task. We formally define the problem of human-robot perturbation training and develop and evaluate the first end-to-end framework for such training, which incorporates a multi-agent transfer learning algorithm, human-robot co-learning framework and communication protocol. Our transfer learning algorithm, Adaptive Perturbation Training (AdaPT), is a hybrid of transfer and reinforcement learning techniques that learns quickly and robustly for new task variants. We empirically validate the benefits of AdaPT through comparison to other hybrid reinforcement and transfer learning techniques aimed at transferring knowledge from multiple source tasks to a single target task. We also demonstrate that AdaPT's rapid learning supports live interaction between a person and a robot, during which the human-robot team trains to achieve a high level of performance for new task variants. We augment AdaPT with a co-learning framework and a computational bi-directional communication protocol so that the robot can co-train with a person during live interaction. Results from large-scale human subject experiments (n=48) indicate that AdaPT enables an agent to learn in a manner compatible with a human's own learning process, and that a robot undergoing perturbation training with a human results in a high level of team performance. Finally, we demonstrate that human-robot training using AdaPT in a simulation environment produces effective performance for a team incorporating an embodied robot partner. Ramya Ramakrishnan, Chongjie Zhang, Julie A. Shah |
J. Artif. Intell. Res. | 3 |
| 2016 | ConTaCT: Deciding to Communicate during Time-Critical Collaborative Tasks in Unknown, Deterministic DomainsabstractCommunication between agents has the potential to improve team performance of collaborative tasks. However, communication is not free in most domains, requiring agents to reason about the costs and benefits of sharing information. In this work, we develop an online, decentralized communication policy, ConTaCT, that enables agents to decide whether or not to communicate during time-critical collaborative tasks in unknown, deterministic environments. Our approach is motivated by real-world applications, including the coordination of disaster response and search and rescue teams. These settings motivate a model structure that explicitly represents the world model as initially unknown but deterministic in nature, and that de-emphasizes uncertainty about action outcomes. Simulated experiments are conducted in which ConTaCT is compared to other multi-agent communication policies, and results indicate that ConTaCT achieves comparable task performance while substantially reducing communication overhead. Vaibhav V. Unhelkar, Julie A. Shah |
AAAI | 2 |
| 2016 | Towards manipulation planning for multiple interlinked deformable linear objectsabstractManipulation of deformable linear objects (DLO) has potential applications in aerospace and automotive assembly. The current literature on planning for deformable objects focuses on a single DLO at a time. In this paper, we provide a problem formulation for attaching a set of interlinked DLOs to a support structure through a set of clamping points. We also present a prototype algorithm that generates a solution in terms of primitive manipulation actions. The algorithm guarantees that none of the interlink constraints are violated. Finally, we incorporate gravity in the computation of a DLO shape and propose a property linking geometrically similar cable shapes across the space of cable length and stiffness. This property allows for computation of solutions for unit length and scaling of the solutions to appropriate length, thus potentially making shape computations faster. Ankit Shah 0003, Julie A. Shah |
ICRA | 2 |
| 2016 | Apprenticeship Scheduling: Learning to Schedule from Human Experts
Matthew C. Gombolay, Reed Jensen, Jessica Stigile, Sung-Hyun Son, Julie A. Shah |
IJCAI | 5 |
| 2016 | Fast Motion Prediction for Collaborative Robotics
Claudia Pérez-D'Arpino, Julie A. Shah |
IJCAI | 2 |
| 2016 | Co-Optimizating Multi-Agent Placement with Task Assignment and Scheduling
Chongjie Zhang, Julie A. Shah |
IJCAI | 2 |
| 2016 | Co-optimizing task and motion planningabstractSolutions to robotic manipulation problems can be substantially improved through integrated task and motion planning. Existing approaches typically focus on satisfaction, finding a feasible solution, instead of optimization. We formulate large-scale robotic manipulation problems as multi-level optimization, incorporating task, action, and motion planning. We develop an integrated planning approach for solving this optimization problem and generating a combined motion plan for a robot to optimize a task-level objective. This approach utilizes a combinatorial search algorithm for task planning and incrementally exploits information from lower-level optimization to improve the high-level task plan. Empirical results show that this integrated approach not only significantly outperforms a traditional top-down approach in solution quality, but also avoids infeasible lower-level motion plans. Chongjie Zhang, Julie A. Shah |
IROS | 2 |
| 2016 | Improving Team's Consistency of Understanding in MeetingsabstractUpon concluding a meeting, participants can occasionally leave with different understandings of what had been discussed. Detecting inconsistencies in understanding is a desired capability for an intelligent system designed to monitor meetings and provide feedback to spur stronger shared understanding. In this paper, we present a computational model for the automatic prediction of consistency among team members' understanding of their group's decisions. The model utilizes dialogue features focused on the dynamics of group decision-making. We trained a hidden Markov model using the AMI meeting corpus and achieved a prediction accuracy of 64.2%, as well as robustness across different meeting phases. We, then, implemented our model in an intelligent system that participated in human team planning about a hypothetical emergency response mission. The system suggested topics that the team would derive the most benefit from reviewing with one another. Through an experiment with 30 participants, we evaluated the utility of such a feedback system and observed a statistically significant increase of 17.5% in objective measures of the teams' understanding compared with that obtained using a baseline interactive system. Joseph Kim, Julie A. Shah |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2015 | Scalable and Interpretable Data Representation for High-Dimensional, Complex DataabstractThe majority of machine learning research has been focused on building models and inference techniques with sound mathematical properties and cutting edge performance. Little attention has been devoted to the development of data representation that can be used to improve a user's ability to interpret the data and machine learning models to solve real-world problems. In this paper, we quantitatively and qualitatively evaluate an efficient, accurate and scalable feature-compression method using latent Dirichlet allocation for discrete data. This representation can effectively communicate the characteristics of high-dimensional, complex data points. We show that the improvement of a user's interpretability through the use of a topic modeling-based compression technique is statistically significant, according to a number of metrics, when compared with other representations. Also, we find that this representation is scalable --- it maintains alignment with human classification accuracy as an increasing number of data points are shown. In addition, the learned topic layer can semantically deliver meaningful information to users that could potentially aid human reasoning about data characteristics in connection with compressed topic space. Been Kim, Kayur Patel, Afshin Rostamizadeh, Julie A. Shah |
AAAI | 4 |
| 2015 | On Fairness in Decision-Making under Uncertainty: Definitions, Computation, and ComparisonabstractThe utilitarian solution criterion, which has been extensively studied in multi-agent decision making under uncertainty, aims to maximize the sum of individual utilities. However, as the utilitarian solution often discriminates against some agents, it is not desirable for many practical applications where agents have their own interests and fairness is expected. To address this issue, this paper introduces egalitarian solution criteria for sequential decision-making under uncertainty, which are based on the maximin principle. Motivated by different application domains, we propose four maximin fairness criteria and develop corresponding algorithms for computing their optimal policies. Furthermore, we analyze the connections between these criteria and discuss and compare their characteristics. Chongjie Zhang, Julie A. Shah |
AAAI | 2 |
| 2015 | Efficient Model Learning from Joint-Action Demonstrations for Human-Robot Collaborative TasksabstractWe present a framework for automatically learning human user models from joint-action demonstrations that enables a robot to compute a robust policy for a collaborative task with a human. First, the demonstrated action sequences are clustered into different human types using an unsupervised learning algorithm. A reward function is then learned for each type through the employment of an inverse reinforcement learning algorithm. The learned model is then incorporated into a mixed-observability Markov decision process (MOMDP) formulation, wherein the human type is a partially observable variable. With this framework, we can infer online the human type of a new user that was not included in the training set, and can compute a policy for the robot that will be aligned to the preference of this user. In a human subject experiment (n=30), participants agreed more strongly that the robot anticipated their actions when working with a robot incorporating the proposed framework (p<0.01), compared to manually annotating robot actions. In trials where participants faced difficulty annotating the robot actions to complete the task, the proposed framework significantly improved team efficiency (p<0.01). The robot incorporating the framework was also found to be more responsive to human actions compared to policies computed using a hand-coded reward function by a domain expert (p<0.01). These results indicate that learning human user models from joint-action demonstrations and encoding them in a MOMDP formalism can support effective teaming in human-robot collaborative tasks. Stefanos Nikolaidis, Ramya Ramakrishnan, Keren Gu, Julie A. Shah |
HRI | 4 |
| 2015 | Fast target prediction of human reaching motion for cooperative human-robot manipulation tasks using time series classificationabstractInterest in human-robot coexistence, in which humans and robots share a common work volume, is increasing in manufacturing environments. Efficient work coordination requires both awareness of the human pose and a plan of action for both human and robot agents in order to compute robot motion trajectories that synchronize naturally with human motion. In this paper, we present a data-driven approach that synthesizes anticipatory knowledge of both human motions and subsequent action steps in order to predict in real-time the intended target of a human performing a reaching motion. Motion-level anticipatory models are constructed using multiple demonstrations of human reaching motions. We produce a library of motions from human demonstrations, based on a statistical representation of the degrees of freedom of the human arm, using time series analysis, wherein each time step is encoded as a multivariate Gaussian distribution. We demonstrate the benefits of this approach through offline statistical analysis of human motion data. The results indicate a considerable improvement over prior techniques in early prediction, achieving 70% or higher correct classification on average for the first third of the trajectory (<; 500msec). We also indicate proof-of-concept through the demonstration of a human-robot cooperative manipulation task performed with a PR2 robot. Finally, we analyze the quality of task-level anticipatory knowledge required to improve prediction performance early in the human motion trajectory. Claudia Pérez-D'Arpino, Julie A. Shah |
ICRA | 2 |
| 2015 | Human-robot co-navigation using anticipatory indicators of human walking motionabstractMobile, interactive robots that operate in human-centric environments need the capability to safely and efficiently navigate around humans. This requires the ability to sense and predict human motion trajectories and to plan around them. In this paper, we present a study that supports the existence of statistically significant biomechanical turn indicators of human walking motions. Further, we demonstrate the effectiveness of these turn indicators as features in the prediction of human motion trajectories. Human motion capture data is collected with predefined goals to train and test a prediction algorithm. Use of anticipatory features results in improved performance of the prediction algorithm. Lastly, we demonstrate the closed-loop performance of the prediction algorithm using an existing algorithm for motion planning within dynamic environments. The anticipatory indicators of human walking motion can be used with different prediction and/or planning algorithms for robotics; the chosen planning and prediction algorithm demonstrates one such implementation for human-robot co-navigation. Vaibhav V. Unhelkar, Claudia Pérez-D'Arpino, Leia A. Stirling 0001, Julie A. Shah |
ICRA | 4 |
| 2015 | Mind the Gap: A Generative Approach to Interpretable Feature Selection and ExtractionabstractWe present the Mind the Gap Model (MGM), an approach for interpretable feature extraction and selection. By placing interpretability criteria directly into the model, we allow for the model to both optimize parameters related to interpretability and to directly report a global set of distinguishable dimensions to assist with further data exploration and hypothesis generation. MGM extracts distinguishing features on real-world datasets of animal features, recipes ingredients, and disease co-occurrence. It also maintains or improves performance when compared to related approaches. We perform a user study with domain experts to show the MGM's ability to help with dataset exploration. Been Kim, Julie A. Shah, Finale Doshi-Velez |
NIPS | 2 |
| 2015 | Inferring Team Task Plans from Human Meetings: A Generative Modeling Approach with Logic-Based PriorabstractWe aim to reduce the burden of programming and deploying autonomous systems to work in concert with people in time-critical domains such as military field operations and disaster response. Deployment plans for these operations are frequently negotiated on-the-fly by teams of human planners. A human operator then translates the agreed-upon plan into machine instructions for the robots. We present an algorithm that reduces this translation burden by inferring the final plan from a processed form of the human team's planning conversation. Our hybrid approach combines probabilistic generative modeling with logical plan validation used to compute a highly structured prior over possible plans, enabling us to overcome the challenge of performing inference over a large solution space with only a small amount of noisy data from the team planning session. We validate the algorithm through human subject experimentations and show that it is able to infer a human team's final plan with 86% accuracy on average. We also describe a robot demonstration in which two people plan and execute a first-response collaborative task with a PR2 robot. To the best of our knowledge, this is the first work to integrate a logical planning technique within a generative model to perform plan inference. Been Kim, Caleb M. Chacha, Julie A. Shah |
J. Artif. Intell. Res. | 3 |
| 2014 | Comparative performance of human and mobile robotic assistants in collaborative fetch-and-deliver tasksabstractThere is an emerging desire across manufacturing industries to deploy robots that support people in their manual work, rather than replace human workers. This paper explores one such opportunity, which is to field a mobile robotic assistant that travels between part carts and the automotive final assembly line, delivering tools and materials to the human workers. We compare the performance of a mobile robotic assistant to that of a human assistant to gain a better understanding of the factors that impact its effectiveness. Statistically significant differences emerge based on type of assistant, human or robot. Interaction times and idle times are statistically significantly higher for the robotic assistant than the human assistant. We report additional differences in participant's subjective response regarding team fluency, situational awareness, comfort and safety. Finally, we discuss how results from the experiment inform the design of a more effective assistant. Vaibhav V. Unhelkar, Ho Chit Siu, Julie A. Shah |
HRI | 3 |
| 2014 | A summary of team MIT's approach to the virtual robotics challengeabstractThe paper describes the system developed by researchers from MIT for the Defense Advanced Research Projects Agency's (DARPA) Virtual Robotics Challenge (VRC), held in June 2013. The VRC was the first competition in the DARPA Robotics Challenge (DRC), a program that aims to “develop ground robotic capabilities to execute complex tasks in dangerous, degraded, human-engineered environments”. The VRC required teams to guide a model of Boston Dynamics' humanoid robot, Atlas, through driving, walking, and manipulation tasks in simulation. Team MIT's user interface, the Viewer, provided the operator with a unified representation of all available information. A 3D rendering of the robot depicted its most recently estimated body state with respect to the surrounding environment, represented by point clouds and texture-mapped meshes as sensed by on-board LIDAR and fused over time. Russ Tedrake, Maurice Fallon, Sisir Karumanchi, Scott Kuindersma, Matthew E. Antone, Toby Schneider, Thomas M. Howard, Matthew R. Walter, Hongkai Dai, Robin Deits, Michael Fleder, Dehann Fourie, Riad I. Hammoud, Sachithra Hemachandra, P. Ilardi, Claudia Pérez-D'Arpino, Sudeep Pillai, Andres Valenzuela, Cecilia Cantu, C. Dolan, I. Evans, S. Jorgensen, J. Kristeller, Julie A. Shah, Karl Iagnemma, Seth J. Teller |
ICRA | 24 |
| 2014 | Towards control and sensing for an autonomous mobile robotic assistant navigating assembly linesabstractThere exists an increasing demand to incorporate mobile interactive robots to assist humans in repetitive, non-value added tasks in the manufacturing domain. Our aim is to develop a mobile robotic assistant for fetch-and-deliver tasks in human-oriented assembly line environments. Assembly lines present a niche yet novel challenge for mobile robots; the robot must precisely control its position on a surface which may be either stationary, moving, or split (e.g. in the case that the robot straddles the moving assembly line and remains partially on the stationary surface). In this paper we present a control and sensing solution for a mobile robotic assistant as it traverses a moving-floor assembly line. Solutions readily exist for control of wheeled mobile robots on static surfaces; we build on the open-source Robot Operating System (ROS) software architecture and generalize the algorithms for the moving line environment. Off-the-shelf sensors and localization algorithms are explored to sense the moving surface, and a customized solution is presented using PX4Flow optic flow sensors and a laser scanner-based localization algorithm. Validation of the control and sensing system is carried out both in simulation and in hardware experiments on a customized treadmill. Initial demonstrations of the hardware system yield promising results; the robot successfully maintains its position while on, and while straddling, the moving line. Vaibhav V. Unhelkar, Jorge Perez, Jim Boerkoel, Johannes Bix, Stefan Bartscher, Julie A. Shah |
ICRA | 6 |
| 2014 | The Bayesian Case Model: A Generative Approach for Case-Based Reasoning and Prototype Classification
Been Kim, Cynthia Rudin, Julie A. Shah |
NIPS | 3 |
| 2014 | Fairness in Multi-Agent Sequential Decision-Making
Chongjie Zhang, Julie A. Shah |
NIPS | 2 |
| 2014 | Automatic prediction of consistency among team members' understanding of group decisions in meetingsabstractOccasionally, participants in a meeting can leave with different understandings of what had been discussed. For meetings that require immediate response (such as disaster response planning), the participants must share a common understanding of the decisions reached by the group to ensure successful execution of their mission. In such domains, inconsistency among individuals' understanding of the meeting results would be detrimental, as this can potentially degrade group performance. Thus, detecting the occurrence of inconsistencies in understanding among meeting participants is a desired capability for an intelligent system that would monitor meetings and provide feedback to spur stronger group understanding. In this paper, we seek to predict the consistency among team members' understanding of group decisions. We use self-reported summaries as a representative measure for team members' understanding following meetings, and present a computational model that uses a set of verbal and nonverbal features from natural dialogue. This model focuses on the conversational dynamics between the participants, rather than on what is being discussed. We apply our model to a real-world conversational dataset and show that its features can predict group consistency with greater accuracy than conventional dialogue features. We also show that the combination of verbal and nonverbal features in multimodal fusion improves several performance metrics, and that our results are consistent across different meeting phases. Joseph Kim, Julie A. Shah |
SMC | 2 |
| 2013 | Inferring Robot Task Plans from Human Team Meetings: A Generative Modeling Approach with Logic-Based PriorabstractWe aim to reduce the burden of programming and deploying autonomous systems to work in concert with people in time-critical domains, such as military field operations and disaster response. Deployment plans for these operations are frequently negotiated on-the-fly by teams of human planners. A human operator then translates the agreed upon plan into machine instructions for the robots. We present an algorithm that reduces this translation burden by inferring the final plan from a processed form of the human team's planning conversation. Our approach combines probabilistic generative modeling with logical plan validation used to compute a highly structured prior over possible plans. This hybrid approach enables us to overcome the challenge of performing inference over the large solution space with only a small amount of noisy data from the team planning session. We validate the algorithm through human subject experimentation and show we are able to infer a human team's final plan with 83% accuracy on average. We also describe a robot demonstration in which two people plan and execute a first-response collaborative task with a PR2 robot. To the best of our knowledge, this is the first work that integrates a logical planning technique within a generative model to perform plan inference. Been Kim, Caleb M. Chacha, Julie A. Shah |
AAAI | 3 |
| 2013 | Human-robot cross-training: computational formulation, modeling and evaluation of a human team training strategy
Stefanos Nikolaidis, Julie A. Shah |
HRI | 2 |
| 2011 | Improved human-robot team performance using chaski, a human-inspired plan execution systemabstractWe describe the design and evaluation of Chaski, a robot plan execution system that uses insights from human-human teaming to make human-robot teaming more natural and fluid. Chaski is a task-level executive that enables a robot to collaboratively execute a shared plan with a person. The system chooses and schedules the robot's actions, adapts to the human partner, and acts to minimize the human's idle time. Julie A. Shah, James Wiken, Brian C. Williams, Cynthia Breazeal |
HRI | 1 |
| 2007 | Review and Synthesis of Considerations in Architecting Heterogeneous Teams of Humans and Robots for Optimal Space ExplorationabstractHuman-robot systems will play a critical role in space exploration, should NASA embark on missions to the Moon and Mars. A unified framework to optimally leverage the capabilities of humans and robots in space exploration will be an invaluable tool for mission planning. Although there is a growing body of literature on human-robot interactions, there is not yet a framework that lends itself both to a formal representation of heterogeneous teams of humans and robots, and to an evaluation of such teams across a series of common, task-based metrics. In this paper, we review the literature, and synthesize multiple considerations for architecting heterogeneous teams of humans and robots. We discuss considerations related to formally specifying tasks and representing human--robot systems, enumerating task allocations, and evaluating human--robot systems against common, task-based metrics. Our objective is to lay the foundations of a unified framework for architecting human--robot systems for optimal task performance given a set metrics. Julie A. Shah, Joseph Homer Saleh, Jeffrey A. Hoffman |
IEEE Trans. Syst. Man Cybern. Part C | 1 |
| 2006 | A Preliminary Study of Peer-to-Peer Human-Robot InteractionabstractThe Peer-To-Peer Human-Robot Interaction (P2P-HRI) project is developing techniques to improve task coordination and collaboration between human and robot partners. Our work is motivated by the need to develop effective human-robot teams for space mission operations. A central element of our approach is creating dialogue and interaction tools that enable humans and robots to flexibly support one another. In order to understand how this approach can influence task performance, we recently conducted a series of tests simulating a lunar construction task with a human-robot team. In this paper, we describe the tests performed, discuss our initial results, and analyze the effect of intervention on task performance. Terrence Fong, Jean Scholtz, Julie A. Shah, Lorenzo Flueckiger, Clayton Kunz, David Lees, John Schreiner, Michael D. Siegel, Laura M. Hiatt, Illah R. Nourbakhsh, Reid G. Simmons, Robert O. Ambrose, Robert R. Burridge, Brian Antonishek, Magdalena D. Bugajska, Alan C. Schultz, J. Gregory Trafton |
SMC | 3 |