VLDB 2026 Research / reviewers in the wild / expert
Matthew C. Gombolay
dblp:144/1022 · also Matthew Craig Gombolay, Matthew Gombolay
· DBLP profile ↗
56ranked-venue papers
6as first author
43since 2021 · last 2026
0000-0002-5321-6038ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 5 first-author · 32 since 2021Human-computer interaction and ubiquitous computing · 20 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 5 since 2021Systems, architecture and hardware · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Teaching the Teacher: Live Foundation Model and Augmented Reality Feedback for Human-to-Robot Skill TransferabstractDeploying robots in dynamic, human-populated environments will require techniques for adaptable robot skill acquisition that extend beyond pre-programmed functionality. Learning from demonstration (LfD) methods enable robots to learn skills from human-provided trajectories demonstrated in situ. However, prior work has shown non-expert end-users struggle to provide demonstrations that enable robots to perform complex, multi-step tasks, or to generalize skill knowledge beyond a specific environment and task context. This work enables robots to actively participate in the situated learning interaction by autonomously providing bespoke guidance in response to end-users' demonstrations, thus improving end-users' ability to teach robots useful skills via LfD. We introduce a novel LfD system integrating foundation model (FM)-based textual feedback and augmented reality (AR)-based visual feedback. The FM and AR feedbacks operate synergistically, with FM feedback helping users break tasks down effectively and with AR feedback allowing users to quickly evaluate how well demonstrations perform and generalize. This system provides targeted, actionable guidance throughout the demonstration process: it enhances users' ability to define, decompose, and demonstrate modular, repurposable skills capable of accomplishing complex tasks. We validate our system with a human-subjects experiment in which participants receive bespoke feedback as they teach a robot via kinesthetic demonstrations in a pair of robotic manipulation domains. From this study, we observe positive results demonstrating that the combination of AR and FM feedback improves the quality and generalizability of robot policies, compared to AR feedback alone, FM feedback alone, or a baseline system where learned skills can be played physically on the robot. Nina Moorman, Matthew B. Luebbers, Zhang Xi-Jia, Yee Ching (Marcus) Lau, Yixing Yao, Megan Langwasser, Zulfiqar Zaidi, Letian Chen, Sanne van Waveren, Matthew C. Gombolay |
HRI | 10 |
| 2026 | Asynchronous Training of Mixed-Role Human Actors in a Partially Observable EnvironmentabstractIn cooperative training, humans within a team coordinate on complex tasks, building mental models of their teammates and learning to adapt to teammates’ actions in real-time. To reduce the often prohibitive scheduling constraints associated with cooperative training, this article introduces a paradigm for cooperative asynchronous training of human teams in which trainees practice coordination with autonomous teammates rather than humans. We introduce a novel experimental design for evaluating autonomous teammates for use as training partners in cooperative training. We apply this design to a human-subjects experiment where humans are trained with either another human or an autonomous teammate and are evaluated with a new human subject in a new, partially observable, cooperative game developed for this study. Importantly, we employ an unsupervised sequential clustering methodology to partition teammate trajectories from demonstrations performed in the experiment to form a smaller number of training conditions. This results in a simpler experiment design, enabling us to conduct a complex cooperative training human-subjects study in a reasonable amount of time. Through a demonstration of the proposed experimental design, we provide takeaways and design recommendations for future research in the development of cooperative asynchronous training systems utilizing robot surrogates for human teammates. Kimberlee Chestnut Chang, Reed Jensen, Rohan R. Paleja, Sam L. Polk, Robert Seater, Jackson Steilberg, Curran Schiefelbein, Melissa Scheldrup, Matthew C. Gombolay, Mabel D. Ramirez |
ACM Trans. Hum. Robot Interact. | 9 |
| 2025 | Generating CAD Code with Vision-Language Models for 3D DesignsabstractGenerative AI has transformed the fields of Design and Manufacturing by providing
efficient and automated methods for generating and modifying 3D objects. One
approach involves using Large Language Models (LLMs) to generate Computer-
Aided Design (CAD) scripting code, which can then be executed to render a 3D
object; however, the resulting 3D object may not meet the specified requirements.
Testing the correctness of CAD generated code is challenging due to the complexity
and structure of 3D objects (e.g., shapes, surfaces, and dimensions) that are not
feasible in code. In this paper, we introduce CADCodeVerify, a novel approach to
iteratively verify and improve 3D objects generated from CAD code. Our approach
works by producing ameliorative feedback by prompting a Vision-Language Model
(VLM) to generate and answer a set of validation questions to verify the generated
object and prompt the VLM to correct deviations. To evaluate CADCodeVerify, we
introduce, CADPrompt, the first benchmark for CAD code generation, consisting of
200 natural language prompts paired with expert-annotated scripting code for 3D
objects to benchmark progress. Our findings show that CADCodeVerify improves
VLM performance by providing visual feedback, enhancing the structure of the 3D
objects, and increasing the success rate of the compiled program. When applied to
GPT-4, CADCodeVerify achieved a 7.30% reduction in Point Cloud distance and a
5.0% improvement in success rate compared to prior work. Kamel Alrashedy, Pradyumna Tambwekar, Zulfiqar Zaidi, Megan Langwasser, Wei Xu 0004, Matthew C. Gombolay |
ICLR | 6 |
| 2025 | Generalized Behavior Learning from Diverse DemonstrationsabstractDiverse behavior policies are valuable in domains requiring quick test-time adaptation or personalized human-robot interaction. Human demonstrations provide rich information regarding task objectives and factors that govern individual behavior variations, which can be used to characterize \textit{useful} diversity and learn diverse performant policies.
However, we show that prior work that builds naive representations of demonstration heterogeneity fails in generating successful novel behaviors that generalize over behavior factors.
We propose Guided Strategy Discovery (GSD), which introduces a novel diversity formulation based on a learned task-relevance measure that prioritizes behaviors exploring modeled latent factors.
We empirically validate across three continuous control benchmarks for generalizing to in-distribution (interpolation) and out-of-distribution (extrapolation) factors that GSD outperforms baselines in novel behavior discovery by $\sim$21\%.
Finally, we demonstrate that GSD can generalize striking behaviors for table tennis in a virtual testbed while leveraging human demonstrations collected in the real world.
Code is available at https://github.com/CORE-Robotics-Lab/GSD. Varshith Sreeramdass, Rohan R. Paleja, Letian Chen, Sanne van Waveren, Matthew C. Gombolay |
ICLR | 5 |
| 2025 | ELEMENTAL: Interactive Learning from Demonstrations and Vision-Language Models for Reward Design in RoboticsabstractReinforcement learning (RL) has demonstrated compelling performance in robotic tasks, but its success often hinges on the design of complex, ad hoc reward functions. Researchers have explored how Large Language Models (LLMs) could enable non-expert users to specify reward functions more easily. However, LLMs struggle to balance the importance of different features, generalize poorly to out-of-distribution robotic tasks, and cannot represent the problem properly with only text-based descriptions. To address these challenges, we propose ELEMENTAL (intEractive LEarning froM dEmoNstraTion And Language), a novel framework that combines natural language guidance with visual user demonstrations to align robot behavior with user intentions better. By incorporating visual inputs, ELEMENTAL overcomes the limitations of text-only task specifications, while leveraging inverse reinforcement learning (IRL) to balance feature weights and match the demonstrated behaviors optimally. ELEMENTAL also introduces an iterative feedback-loop through self-reflection to improve feature, reward, and policy learning. Our experiment results demonstrate that ELEMENTAL outperforms prior work by 42.3% on task success, and achieves 41.3% better generalization in out-of-distribution tasks, highlighting its robustness in LfD. Letian Chen, Nina Moorman, Matthew C. Gombolay |
ICML | 3 |
| 2025 | Learning Diverse Robot Striking Motions with Diffusion Models and Kinematically Constrained Gradient GuidanceabstractAdvances in robot learning have enabled robots to generate skills for a variety of tasks. Yet, robot learning is typically sample inefficient, struggles to learn from data sources exhibiting varied behaviors, and does not naturally incorporate constraints. These properties are critical for fast, agile tasks such as playing table tennis. Modern techniques for learning from demonstration improve sample efficiency and scale to diverse data, but are rarely evaluated on agile tasks. In the case of reinforcement learning, achieving good performance requires training on high-fidelity simulators. To overcome these limitations, we develop a novel diffusion modeling approach that is offline, constraint-guided, and expressive of diverse agile behaviors. The key to our approach is a kinematic constraint gradient guidance (KCGG) technique that computes gradients through both the forward kinematics of the robot arm and the diffusion model to direct the sampling process. KCGG minimizes the cost of violating constraints while simultaneously keeping the sampled trajectory in-distribution of the training data. We demonstrate the effectiveness of our approach for time-critical robotic tasks by evaluating KCGG in two challenging domains: simulated air hockey and real table tennis. In simulated air hockey, we achieved a 25.4% increase in block rate, while in table tennis, we achieved a 17.3% increase in success rate compared to imitation learning baselines. Kin Man Lee, Sean Ye, Qingyu Xiao, Zulfiqar Zaidi, David B. D'Ambrosio, Pannag R. Sanketi, Matthew C. Gombolay |
ICRA | 8 |
| 2025 | Learning Wheelchair Tennis Navigation from Broadcast Videos with Domain Knowledge Transfer and Diffusion Motion PlanningabstractIn this paper, we propose a novel and generalizable zero-shot knowledge transfer framework that distills expert sports navigation strategies from web videos into robotic systems with adversarial constraints and out-of-distribution image trajectories. Our pipeline enables diffusion-based imitation learning by reconstructing the full 3D task space from multiple partial views, warping it into 2D image space, closing the planning loop within this 2D space, and transfer constrained motion of interest back to task space. Additionally, we demonstrate that the learned policy can serve as a local planner in conjunction with position control. We apply this framework in the wheelchair tennis navigation problem to guide the wheelchair into the ball-hitting region. Our pipeline achieves a navigation success rate of$\mathbf{9 7. 6 7 \%}$in reaching real-world recorded tennis ball trajectories with a physical robot wheelchair, and achieve a success rate of 68.49% in a real-world, real-time experiment on a full-sized tennis court22Code is at https://github.gatech.edu/MCG-Lab/tennis_gameplay_learning. Zulfiqar Zaidi, Adithya Patil, Qingyu Xiao, Matthew C. Gombolay |
ICRA | 5 |
| 2025 | Learning Dynamics of a Ball with Differentiable Factor Graph and Roto-Translational Invariant RepresentationsabstractRobots in dynamic environments need fast, accurate models of how objects move in their environments to support agile planning. In sports such as ping pong, analytical models often struggle to accurately predict ball trajectories with spins due to complex aerodynamics, elastic behaviors, and the challenges of modeling sliding and rolling friction. On the other hand, despite the promise of data-driven methods, machine learning struggles to make accurate, consistent predictions without precise input. In this paper, we propose an end-to-end learning framework that can jointly train a dynamics model and a factor graph estimator. Our approach leverages a Gram-Schmidt (GS) process to extract roto-translational invariant representations to improve the model performance, which can further reduce the validation error compared to data augmentation method. Additionally, we propose a network architecture that enhances nonlinearity by using self-multiplicative bypasses in the layer connections. By leveraging these novel methods, our proposed approach predicts the ball's position with an RMSE of 37.2 mm at the apex after the first bounce, and 71.5 mm after the second bounce. Qingyu Xiao, Matthew C. Gombolay |
ICRA | 3 |
| 2025 | Diverse Heterogeneous Graph Conditioned Diffusion for Multi-Agent Teaming
Luis Pimentel, Sean Ye, James Ellis Grant Pagan, Matthew C. Gombolay |
AAMAS | 4 |
| 2025 | Heterogeneous Graph Transformers for Simultaneous Mobile Multi-Robot Task Allocation and Scheduling under Temporal ConstraintsabstractCoordinating large teams of heterogeneous mobile agents to perform complex tasks efficiently has scalability bottlenecks in feasible and optimal task scheduling, with critical applications in logistics, manufacturing, and disaster response. Existing task allocation and scheduling methods, including heuristics and optimization-based solvers, often fail to scale and overlook inter-task dependencies and agent heterogeneity. We propose a novel Simultaneous Decision-Making model for Heterogeneous Multi-Agent Task Allocation and Scheduling (HM-MATAS), built on a Residual Heterogeneous Graph Transformer with edge and node-level attention. Our model encodes agent capabilities, travel times, and temporospatial constraints into a rich graph representation and is trainable via reinforcement learning. Trained on small-scale problems (10 agents, 20 tasks), our model generalizes effectively to significantly larger scenarios (up to 40 agents and 200 tasks), enabling fast, one-shot task assignment and scheduling. Our simultaneous model outperforms classical heuristics by assigning 164.10\% more feasible tasks given temporal constraints in 3.83\% of the time, metaheuristics by 201.54\% in 0.01\% of the time and exact solver by 231.73\% in 0.03\% of the time, while achieving $20\times$-to-$250\times$ speedup from prior graph-based methods across scales. Batuhan Altundas, Shivika Singh, Shivangi Deo, Minwoo Cho, Matthew C. Gombolay |
NeurIPS | 6 |
| 2025 | Understanding the effects of humanlike robot motions on unfocused human-robot interaction
Yeseul Kim, Matthew C. Gombolay, Yong K. Cho |
Adv. Eng. Informatics | 3 |
| 2024 | Towards Balancing Preference and Performance through Adaptive Personalized ExplainabilityabstractAs robots and digital assistants are deployed in the real world, these agents must be able to communicate their decision-making criteria to build trust, improve human-robot teaming, and enable collaboration. While the field of explainable artificial intelligence (xAI) has made great strides to enable such communication, these advances often assume that one xAI approach is ideally suited to each problem (e.g., decision trees to explain how to triage patients in an emergency or feature-importance maps to explain radiology reports). This fails to recognize that users have diverse experiences or preferences for interaction modalities. In this work, we present two user-studies set in a simulated autonomous vehicle (AV) domain. We investigate (1) population-level preferences for xAI and (2) personalization strategies for providing robot explanations. We find significant differences between xAI modes (language explanations, feature-importance maps, and decision trees) in both preference (p < 0.01) and performance (p < 0.05). We also observe that a participant's preferences do not always align with their performance, motivating our development of an adaptive personalization strategy to balance the two. We show that this strategy yields significant performance gains (p < 0.05), and we conclude with a discussion of our findings and implications for xAI in human-robot interactions. Andrew Silva, Pradyumna Tambwekar, Mariah Schrum, Matthew C. Gombolay |
HRI | 4 |
| 2024 | Enhancing Safety in Learning from Demonstration Algorithms via Control Barrier Function ShieldingabstractLearning from Demonstration (LfD) is a powerful method for non-roboticists end-users to teach robots new tasks, enabling them to customize the robot behavior. However, modern LfD techniques do not explicitly synthesize safe robot behavior, which limits the deployability of these approaches in the real world. To enforce safety in LfD without relying on experts, we propose a new framework, SElding with Control barrier fUnctions in inverse REinforcement learning (SECURE), which learns a customized Control Barrier Function (CBF) from end-users that prevents robots from taking unsafe actions while imposing little interference with the task completion. We evaluate SECURE in three sets of experiments. First, we empirically validate SECURE learns a high-quality CBF from demonstrations and outperforms conventional LfD methods on simulated robotic and autonomous driving tasks with improvements on safety by up to 100%. Second, we demonstrate that roboticists can leverage SECURE to outperform conventional LfD approaches on a real-world knife-cutting, meal-preparation task by 12.5% in task completion while driving the number of safety violations to zero. Finally, we demonstrate in a user study that non-roboticists can use SECURE to effectively teach the robot safe policies that avoid collisions with the person and prevent coffee from spilling. Yue Yang 0024, Letian Chen, Zulfiqar Zaidi, Sanne van Waveren, Arjun Krishna, Matthew C. Gombolay |
HRI | 6 |
| 2024 | CrossLoco: Human Motion Driven Control of Legged Robots via Guided Unsupervised Reinforcement LearningabstractHuman motion driven control (HMDC) is an effective approach for generating natural and compelling robot motions while preserving high-level semantics. However, establishing the correspondence between humans and robots with different body structures is not straightforward due to the mismatches in kinematics and dynamics properties, which causes intrinsic ambiguity to the problem. Many previous algorithms approach this motion retargeting problem with unsupervised learning, which requires the prerequisite skill sets. However, it will be extremely costly to learn all the skills without understanding the given human motions, particularly for high-dimensional robots. In this work, we introduce CrossLoco, a guided unsupervised reinforcement learning framework that simultaneously learns robot skills and their correspondence to human motions. Our key innovation is to introduce a cycle-consistency-based reward term designed to maximize the mutual information between human motions and robot states. We demonstrate that the proposed framework can generate compelling robot motions by translating diverse human motions, such as running, hopping, and dancing. We quantitatively compare our CrossLoco against the manually engineered and unsupervised baseline algorithms along with the ablated versions of our framework and demonstrate that our method translates human motions with better accuracy, diversity, and user preference. We also showcase its utility in other applications, such as synthesizing robot movements from language input and enabling interactive robot control. Tianyu Li 0005, Hyunyoung Jung 0002, Matthew C. Gombolay, Yong Kwon Cho, Sehoon Ha |
ICLR | 3 |
| 2024 | Multi-Camera Asynchronous Ball Localization and Trajectory Prediction with Factor Graphs and Human PosesabstractThe rapid and precise localization and prediction of a ball are critical for developing agile robots in ball sports, particularly in sports like tennis characterized by high-speed ball movements and powerful spins. The Magnus effect induced by spin adds complexity to trajectory prediction during flight and bounce dynamics upon contact with the ground. In this study, we introduce an innovative approach that combines a multi-camera system with factor graphs for real-time and asynchronous 3D tennis ball localization. Additionally, we estimate hidden states like velocity and spin for trajectory prediction. Furthermore, to enhance spin inference early in the ball’s flight, where limited observations are available, we integrate human pose data using a temporal convolutional network (TCN) to compute spin priors within the factor graph. This refinement provides more accurate spin priors at the beginning of the factor graph, leading to improved early-stage hidden state inference for prediction. Our results show the trained TCN can predict the spin priors with RMSE of 5.27 Hz. Integrating TCN into the factor graph reduces the prediction error of landing positions by over 63.6% compared to a baseline method that utilized an adaptive extended Kalman filter. Qingyu Xiao, Zulfiqar Zaidi, Matthew C. Gombolay |
ICRA | 3 |
| 2024 | Human-Robot Alignment through Interactivity and Interpretability: Don't Assume a "Spherical Human"
Matthew C. Gombolay |
IJCAI | 1 |
| 2024 | Efficient Trajectory Forecasting and Generation with Conditional Flow MatchingabstractTrajectory prediction and generation are crucial for autonomous robots in dynamic environments. While prior research has typically focused on either prediction or generation, our approach unifies these tasks to provide a versatile framework and achieve state-of-the-art performance. While diffusion models excel in trajectory generation, their iterative sampling process is computationally intensive, hindering robotic systems’ dynamic capabilities. We introduce Trajectory Conditional Flow Matching (T-CFM), a novel approach using flow matching techniques to learn a solver time-varying vector field for efficient, fast trajectory generation. T-CFM demonstrates effectiveness in adversarial tracking, real-world aircraft trajectory forecasting, and long-horizon planning, outperforming state-of-the-art baselines with 35% higher predictive accuracy and 142% improved planning performance. Crucially, T-CFM achieves up to 100× speed-up compared to diffusion models without sacrificing accuracy, enabling real-time decision making in robotics. Codebase: https://github.com/CORE-Robotics-Lab/TCFM Sean Ye, Matthew C. Gombolay |
IROS | 2 |
| 2024 | Designs for Enabling Collaboration in Human-Machine Teaming via Interactive and Explainable SystemsabstractCollaborative robots and machine learning-based virtual agents are increasingly entering the human workspace with the aim of increasing productivity and enhancing safety. Despite this, we show in a ubiquitous experimental domain, Overcooked-AI, that state-of-the-art techniques for human-machine teaming (HMT), which rely on imitation or reinforcement learning, are brittle and result in a machine agent that aims to decouple the machine and human’s actions to act independently rather than in a synergistic fashion. To remedy this deficiency, we develop HMT approaches that enable iterative, mixed-initiative team development allowing end-users to interactively reprogram interpretable AI teammates. Our 50-subject study provides several findings that we summarize into guidelines. While all approaches underperform a simple collaborative heuristic (a critical, negative result for learning-based methods), we find that white-box approaches supported by interactive modification can lead to significant team development, outperforming white-box approaches alone, and that black-box approaches are easier to train and result in better HMT performance highlighting a tradeoff between explainability and interactivity versus ease-of-training. Together, these findings present three important future research directions: 1) Improving the ability to generate collaborative agents with white-box models, 2) Better learning methods to facilitate collaboration rather than individualized coordination, and 3) Mixed-initiative interfaces that enable users, who may vary in ability, to improve collaboration. Rohan R. Paleja, Michael Munje, Kimberlee Chestnut Chang, Reed Jensen, Matthew C. Gombolay |
NeurIPS | 5 |
| 2024 | Towards the design of user-centric strategy recommendation systems for collaborative Human-AI tasks
Lakshita Dodeja, Pradyumna Tambwekar, Erin Hedlund-Botti, Matthew C. Gombolay |
Int. J. Hum. Comput. Stud. | 4 |
| 2024 | Trust and Dependence on Robotic Decision SupportabstractThis article investigates people's trust and dependence on robotic decision support systems (DSSs), which provide cognitive assistance through suggestions. Robotic DSSs may not always offer optimal suggestions, requiring people to rely carefully to maximize performance. We analyze user reliance on suboptimal robots for solving instantaneous and sequential decision-making tasks with a math and card game, respectively. In instantaneous tasks, we find that the users' perceived anthropomorphism$(p < . 001$) and the robot's behavior after a decision support failure ($p < . 001$) significantly impact user trust. In a sequential task where the effectiveness of the human–robot team is not revealed until after several decisions, we find that introducing a user-initiated decision proposal before the robot reveals its recommendation can mitigate overreliance ($p < . 05$) and users' task expertise is critical in determining appropriate dependence on the robot's suggestions ($p < . 01$). Combined, these studies are synergistic and the first to jointly examine the influence of various factors on user trust and dependence, offering guidance for designing robotic DSSs to maximize human–robot task performance. Manisha Natarajan, Matthew C. Gombolay |
IEEE Trans. Robotics | 2 |
| 2024 | The Impact of Stress and Workload on Human Performance in Robot Teleoperation TasksabstractAdvances in robot teleoperation have enabled groundbreaking innovations in many fields, such as space exploration, healthcare, and disaster relief. The human operator's performance plays a key role in the success of any teleoperation task, with prior evidence suggesting that operator stress and workload can impact task performance. As robot teleoperation is currently deployed in safety-critical domains, it is essential to analyze how different stress and workload levels impact the operator. We are unaware of any prior work investigating how both stress and workload impact teleoperation performance. We conducted a novel study ($n=24$) to jointly manipulate users' stress and workload and analyze the user's performance through objective and subjective measures. Our results indicate that, as stress increased, over 70% of our participants performed better up to a moderate level of stress; yet, the majority of participants performed worse as the workload increased. Importantly, our experimental design elucidated that stress and workload have related yet distinct impacts on task performance, with workload mediating the effects of distress on performance ($p< .05$). Yi Ting Sam, Erin Hedlund-Botti, Manisha Natarajan, Jamison Heard, Matthew C. Gombolay |
IEEE Trans. Robotics | 5 |
| 2024 | MAVERIC: A Data-Driven Approach to Personalized Autonomous DrivingabstractPersonalization of autonomous vehicles (AVs) may significantly increase acceptance. In particular, we hypothesize that the similarity of an AV's driving style compared to a user's driving style, the level of aggressiveness of the driving style, and other subjective factors (e.g., personality) will have a major impact on user's willingness to use the AV. In this work, we 1) develop a data-driven approach to personalize driving style and calibrate the level of aggressiveness and 2) investigate the subjective factors that impact user preference. Across two human subject studies (n = 54), we demonstrate that our approach can mimic the driving styles and tune the level of aggressiveness. Second, we leverage our framework to investigate the factors that impact homophily. We demonstrate that our approach generates driving styles objectively ($p < .001$) and subjectively ($p = .002$) consistent with end-user styles ($p < .001$) and can effectively isolate and modulate a dimension of style (i.e., aggressiveness) ($p < .001$). Furthermore, we find that personality ($p < .001$), perceived similarity ($p < .001$), and high-velocity driving style ($p = .0031$) significantly modulate the effect of homophily. Mariah Schrum, Emily S. Sumner, Matthew C. Gombolay, Andrew Best |
IEEE Trans. Robotics | 3 |
| 2024 | Heterogeneous Policy Networks for Composite Robot Team Communication and CoordinationabstractHigh-performing human–human teams learn intelligent and efficient communication and coordination strategies to maximize their joint utility. These teams implicitly understand the different roles of heterogeneous team members and adapt their communication protocols accordingly. Multiagent reinforcement learning (MARL) has attempted to develop computational methods for synthesizing such joint coordination–communication strategies, but emulating heterogeneous communication patterns across agents with different state, action, and observation spaces has remained a challenge. Without properly modeling agent heterogeneity, as in prior MARL work that leverages homogeneous graph networks, communication becomes less helpful and can even deteriorate the team's performance. In the past, we proposed heterogeneous policy networks (HetNet) to learn efficient and diverse communication models for coordinating cooperative heterogeneous teams. In this extended work, we extend HetNet to support scaling heterogeneous robot teams. Building on heterogeneous graph-attention networks, we show that HetNet not only facilitates learning heterogeneous collaborative policies, but also enables end-to-end training for learning highly efficient binarized messaging. Our empirical evaluation shows that HetNet sets a new state-of-the-art in learning coordination and communication strategies for heterogeneous multiagent teams by achieving an 5.84% to 707.65% performance improvement over the next-best baseline across multiple domains while simultaneously achieving a 200× reduction in the required communication bandwidth. Esmaeil Seraj, Rohan R. Paleja, Luis Pimentel, Kin Man Lee, Zheyuan Wang, Matthew Sklar, John Z. Zhang, Zahi M. Kakish, Matthew C. Gombolay |
IEEE Trans. Robotics | 10 |
| 2023 | The Effect of Robot Skill Level and Communication in Rapid, Proximate Human-Robot CollaborationabstractAs high-speed, agile robots become more commonplace, these robots will have the potential to better aid and collaborate with humans. However, due to the increased agility and functionality of these robots, close collaboration with humans can create safety concerns that alter team dynamics and degrade task performance. In this work, we aim to enable the deployment of safe and trustworthy agile robots that operate in proximity with humans. We do so by 1) Proposing a novel human-robot doubles table tennis scenario to serve as a testbed for studying agile, proximate human-robot collaboration and 2) Conducting a user-study to understand how attributes of the robot (e.g., robot competency or capacity to communicate) impact team dynamics, perceived safety, and perceived trust, and how these latent factors affect human-robot collaboration (HRC) performance. We find that robot competency significantly increases perceived trust (p < .001), extending skill-to-trust assessments in prior studies to agile, proximate HRC. Furthermore, interestingly, we find that when the robot vocalizes its intention to perform a task, it results in a significant decrease in team performance (p = .037) and perceived safety of the system (p = .009). Kin Man Lee, Arjun Krishna, Zulfiqar Zaidi, Rohan R. Paleja, Letian Chen, Erin Hedlund-Botti, Mariah Schrum, Matthew C. Gombolay |
HRI | 8 |
| 2023 | Impacts of Robot Learning on User Attitude and BehaviorabstractWith an aging population and a growing shortage of caregivers, the need for in-home robots is increasing. However, it is intractable for robots to have all functionalities pre-programmed prior to deployment. Instead, it is more realistic for robots to engage in supplemental, on-site learning about the user's needs and preferences. Such learning may occur in the presence of or involve the user. We investigate the impacts on end-users of in situ robot learning through a series of human-subjects experiments. We examine how different learning methods influence both in-person and remote participants' perceptions of the robot. While we find that the degree of user involvement in the robot's learning method impacts perceived anthropomorphism (p=.001), we find that it is the participants' perceived success of the robot that impacts the participants' trust in (p<.001) and perceived usability of the robot (p<.001) rather than the robot's learning method. Therefore, when presenting robot learning, the performance of the learning method appears more important than the degree of user involvement in the learning. Furthermore, we find that the physical presence of the robot impacts perceived safety (p<.001), trust (p<.001), and usability (p<.014). Thus, for tabletop manipulation tasks, researchers should consider the impact of physical presence on experiment participants. Nina Moorman, Erin Hedlund-Botti, Mariah Schrum, Manisha Natarajan, Matthew C. Gombolay |
HRI | 5 |
| 2023 | Learning Models of Adversarial Agent Behavior Under Partial ObservabilityabstractThe need for opponent modeling and tracking arises in several real-world scenarios, such as professional sports, video game design, and drug-trafficking interdiction. In this work, we present Graph based Adversarial Modeling with Mutual Information (GrAMMI) for modeling the behavior of an adversarial opponent agent. GrAMMI is a novel graph neural network (GNN) based approach that uses mutual information maximization as an auxiliary objective to predict the current and future states of an adversarial opponent with partial observability. To evaluate GrAMMI, we design two large-scale, pursuit-evasion domains inspired by real-world scenarios, where a team of heterogeneous agents is tasked with tracking and interdicting a single adversarial agent, and the adversarial agent must evade detection while achieving its own objectives. With the mutual information formulation, GrAMMI outperforms all baselines in both domains and achieves 31.68% higher log-likelihood on average for future adversarial state predictions across both domains. Sean Ye, Manisha Natarajan, Rohan R. Paleja, Letian Chen, Matthew C. Gombolay |
IROS | 6 |
| 2023 | Mixed-Initiative Multiagent Apprenticeship Learning for Human Training of Robot TeamsabstractExtending recent advances in Learning from Demonstration (LfD) frameworks to multi-robot settings poses critical challenges such as environment non-stationarity due to partial observability which is detrimental to the applicability of existing methods. Although prior work has shown that enabling communication among agents of a robot team can alleviate such issues, creating inter-agent communication under existing Multi-Agent LfD (MA-LfD) frameworks requires the human expert to provide demonstrations for both environment actions and communication actions, which necessitates an efficient communication strategy on a known message spaces. To address this problem, we propose Mixed-Initiative Multi-Agent Apprenticeship Learning (MixTURE). MixTURE enables robot teams to learn from a human expert-generated data a preferred policy to accomplish a collaborative task, while simultaneously learning emergent inter-agent communication to enhance team coordination. The key ingredient to MixTURE's success is automatically learning a communication policy, enhanced by a mutual-information maximizing reverse model that rationalizes the underlying expert demonstrations without the need for human generated data or an auxiliary reward function. MixTURE outperforms a variety of relevant baselines on diverse data generated by human experts in complex heterogeneous domains. MixTURE is the first MA-LfD framework to enable learning multi-robot collaborative policies directly from real human data, resulting in ~44% less human workload, and ~46% higher usability score. Esmaeil Seraj, Jerry Xiong, Mariah Schrum, Matthew C. Gombolay |
NeurIPS | 4 |
| 2023 | Explainable Artificial Intelligence: Evaluating the Objective and Subjective Impacts of xAI on Human-Agent InteractionabstractIntelligent agents must be able to communicate intentions and explain their decision-making processes to build trust, foster confidence, and improve human-agent team dynamics. Recognizing this need, academia and industry are rapidly proposing new ideas, methods, and frameworks to aid in the design of more explainable AI. Yet, there remains no standardized metric or experimental protocol for benchmarking new methods, leaving researchers to rely on their own intuition or ad hoc methods for assessing new concepts. In this work, we present the first comprehensive (n = 286) user study testing a wide range of approaches for explainable machine learning, including feature importance, probability scores, decision trees, counterfactual reasoning, natural language explanations, and case-based reasoning, as well as a baseline condition with no explanations. We provide the first large-scale empirical evidence of the effects of explainability on human-agent teaming. Our results will help to guide the future of explainability research by highlighting the benefits of counterfactual explanations and the shortcomings of confidence scores for explainability. We also propose a novel questionnaire to measure explainability with human participants, inspired by relevant prior work and correlated with human-agent teaming metrics. Andrew Silva, Mariah Schrum, Erin Hedlund-Botti, Nakul Gopalan, Matthew C. Gombolay |
Int. J. Hum. Comput. Interact. | 5 |
| 2023 | Concerning Trends in Likert Scale Usage in Human-robot Interaction: Towards Improving Best PracticesabstractAs robots become more prevalent, the importance of the field of human-robot interaction (HRI) grows accordingly. As such, we should endeavor to employ the best statistical practices in HRI research. Likert scales are commonly used metrics in HRI to measure perceptions and attitudes. Due to misinformation or honest mistakes, many HRI researchers do not adopt best practices when analyzing Likert data. We conduct a review of psychometric literature to determine the current standard for Likert scale design and analysis. Next, we conduct a survey of five years of the International Conference on Human-Robot Interaction (HRIc) (2016 through 2020) and report on incorrect statistical practices and design of Likert scales [ 1 , 2 , 3 , 5 , 7 ]. During these years, only 4 of the 144 papers applied proper statistical testing to correctly designed Likert scales. We additionally conduct a survey of best practices across several venues and provide a comparative analysis to determine how Likert practices differ across the field of Human-robot Interaction. We find that a venue’s impact score negatively correlates with number of Likert-related errors and acceptance rate, and total number of papers accepted per venue positively correlates with the number of errors. We also find statistically significant differences between venues for the frequency of misnomer and design errors. Our analysis suggests there are areas for meaningful improvement in the design and testing of Likert scales. Based on our findings, we provide guidelines and a tutorial for researchers for developing and analyzing Likert scales and associated data. We also detail a list of recommendations to improve the accuracy of conclusions drawn from Likert data. Mariah Schrum, Muyleng Ghuy, Erin Hedlund-Botti, Manisha Natarajan, Michael J. Johnson, Matthew C. Gombolay |
ACM Trans. Hum. Robot Interact. | 6 |
| 2022 | Cross-Loss Influence Functions to Explain Deep Network RepresentationsabstractAs machine learning is increasingly deployed in the real world, it is paramount that we develop the tools necessary to analyze the decision-making of the models we train and deploy to end-users. Recently, researchers have shown that influence functions, a statistical measure of sample impact, can approximate the effects of training samples on classification accuracy for deep neural networks. However, this prior work only applies to supervised learning, where training and testing share an objective function. No approaches currently exist for estimating the influence of unsupervised training examples for deep learning models. To bring explainability to unsupervised and semi-supervised training regimes, we derive the first theoretical and empirical demonstration that influence functions can be extended to handle mismatched training and testing (i.e., "cross-loss") settings. Our formulation enables us to compute the influence in an unsupervised learning setup, explain cluster memberships, and identify and augment biases in language models. Our experiments show that our cross-loss influence estimates even exceed matched-objective influence estimation relative to ground-truth sample impact. Andrew Silva, Rohit Chopra, Matthew C. Gombolay |
AISTATS | 3 |
| 2022 | Machine Learning in Human-Robot Collaboration: Bridging the GapabstractThis workshop aims to bring together researchers to explore and identify ways in which human-robot collaboration can reap the benefits of modern machine learning. The intended outcome is a roadmap that identifies key milestones that will lead us towards fluent effective human-robot teaming. In addition to focus groups and creative brainstorming exercises, this workshop will comprise invited talks, contributed paper talks, a poster session, and a debate. The papers, talks, posters, and roadmap will be made publicly available on our website: https://sites.google.com/view/mlhrc-hri-2022/home. Cynthia Matuszek, Harold Soh, Matthew C. Gombolay, Nakul Gopalan, Reid G. Simmons, Stefanos Nikolaidis |
HRI | 3 |
| 2022 | Personalized Meta-Learning for Domain Agnostic Learning from DemonstrationabstractFor robots to perform novel tasks in the real-world, they must be capable of learning from heterogeneous, non-expert human teachers across various domains. Yet, novice human teachers often provide suboptimal demonstrations, making it difficult for robots to successfully learn. Therefore, to effectively learn from humans, we must develop learning methods that can account for teacher suboptimality and can do so across various robotic platforms. To this end, we introduce Mutual Information Driven Meta-Learning from Demonstration (MIND MELD) [12], [13], a personalized meta-learning framework which meta-learns a mapping from suboptimal human feedback to feedback closer to optimal, conditioned on a learned personalized embedding. In a human subjects study, we demonstrate MIND MELD's ability to improve upon suboptimal demonstrations and learn meaningful, personalized embeddings. We then propose Domain Agnostic MIND MELD, which learns to transfer the personalized embedding learned in one domain to a novel domain, thereby allowing robots to learn from suboptimal humans across disparate platforms (e.g., self-driving car or in-home robot). Mariah Schrum, Erin Hedlund-Botti, Matthew C. Gombolay |
HRI | 3 |
| 2022 | MIND MELD: Personalized Meta-Learning for Robot-Centric Imitation LearningabstractLearning from demonstration (LfD) techniques seek to enable users without computer programming experience to teach robots novel tasks. There are generally two types of LfD: human- and robot-centric. While human-centric learning is intuitive, human centric learning suffers from performance degradation due to covariate shift. Robot-centric approaches, such as Dataset Aggregation (DAgger), address covariate shift but can struggle to learn from suboptimal human teachers. To create a more human-aware version of robot-centric LfD, we present Mutual Information-driven Meta-learning from Demonstration (MIND MELD). MIND MELD meta-learns a mapping from suboptimal and heterogeneous human feedback to optimal labels, thereby improving the learning signal for robot-centric LfD. The key to our approach is learning an informative personalized em-bedding using mutual information maximization via variational inference. The embedding then informs a mapping from human provided labels to optimal labels. We evaluate our framework in a human-subjects experiment, demonstrating that our approach improves corrective labels provided by human demonstrators. Our framework outperforms baselines in terms of ability to reach the goal$(p <. 001)$, average distance from the goal$(p=.006)$, and various subjective ratings$(p=.008)$. Mariah Schrum, Erin Hedlund-Botti, Nina Moorman, Matthew C. Gombolay |
HRI | 4 |
| 2022 | Iterated Reasoning with Mutual Information in Cooperative and Byzantine Decentralized Teaming
Sachin Konan, Esmaeil Seraj, Matthew C. Gombolay |
ICLR | 3 |
| 2022 | Learning Coordination Policies over Heterogeneous Graphs for Human-Robot Teams via Recurrent Neural Schedule PropagationabstractAs human-robot collaboration increases in the workforce, it becomes essential for human-robot teams to coordinate efficiently and intuitively. Traditional approaches for human-robot scheduling either utilize exact methods that are intractable for large-scale problems and struggle to account for stochastic, time varying human task performance, or application-specific heuristics that require expert domain knowledge to develop. We propose a deep learning-based framework, called HybridNet, combining a heterogeneous graph-based encoder with a recurrent schedule propagator for scheduling stochastic human-robot teams under upper- and lower-bound temporal constraints. The HybridNet's encoder leverages Heterogeneous Graph Attention Networks to model the initial environment and team dynamics while accounting for the constraints. By formulating task scheduling as a sequential decision-making process, the HybridNet's recurrent neural schedule propagator leverages Long Short-Term Memory (LSTM) models to propagate forward consequences of actions to carry out fast schedule generation, removing the need to interact with the environment between every taskagent pair selection. The resulting scheduling policy network provides a computationally lightweight yet highly expressive model that is end-to-end trainable via Reinforcement Learning algorithms. We develop a virtual task scheduling environment for mixed human-robot teams in a multi-round setting, capable of modeling the stochastic learning behaviors of human workers. Experimental results showed that HybridNet outperformed other human-robot scheduling solutions across problem sizes for both deterministic and stochastic human performance, with faster runtime compared to pure-GNN-based schedulers. Batuhan Altundas, Zheyuan Wang, Joshua Bishop, Matthew C. Gombolay |
IROS | 4 |
| 2022 | Multi-UAV planning for cooperative wildfire coverage and tracking with quality-of-service guarantees
Esmaeil Seraj, Andrew Silva, Matthew C. Gombolay |
Auton. Agents Multi Agent Syst. | 3 |
| 2022 | Coordinating Human-Robot Teams with Dynamic and Stochastic Task ProficienciesabstractAs robots become ubiquitous in the workforce, it is essential that human-robot collaboration be both intuitive and adaptive. A robot’s ability to coordinate team activities improves based on its ability to infer and reason about the dynamic (i.e., the “learning curve”) and stochastic task performance of its human counterparts. We introduce a novel resource coordination algorithm that enables robots to schedule team activities by (1) actively characterizing the task performance of their human teammates and (2) ensuring the schedule is robust to temporal constraints given this characterization. We first validate our modeling assumptions via user study. From this user study, we create a data-driven prior distribution over human task performance for our virtual and physical evaluations of human-robot teaming. Second, we show that our methods are scalable and produce high-quality schedules. Third, we conduct a between-subjects experiment (n = 90) to assess the effects on a human-robot team of a robot scheduler actively exploring the humans’ task proficiency. Our results indicate that human-robot working alliance ( \( p\lt 0.001 \) ) and human performance ( \( p=0.00359 \) ) are maximized when the robot dedicates more time to exploring the capabilities of human teammates. Ruisen Liu, Manisha Natarajan, Matthew C. Gombolay |
ACM Trans. Hum. Robot Interact. | 3 |
| 2022 | A Hierarchical Coordination Framework for Joint Perception-Action Tasks in Composite Robot TeamsabstractWe propose a collaborative planning and control algorithm to enhance cooperation for composite teams of autonomous robots in dynamic environments. Composite robot teams are groups of agents that perform different tasks according to their respective capabilities in order to accomplish an overarching mission. Examples of such teams include groups of perception agents (can only sense) and action agents (can only manipulate) working together to perform disaster response tasks. Coordinating robots in a composite team is a challenging problem due to the heterogeneity in the robots’ characteristics and their tasks. Here, we propose a coordination framework for composite robot teams. The proposed framework consists of two hierarchical modules: First, A multiagent state-action-reward-time-state-action algorithm in multiagent partially observable semi-Markov decision process as the high-level decision-making module to enable perception agents to learn to surveil in an environment with an unknown number of dynamic targets and second, a low-level coordinated control and planning module that ensures probabilistically guaranteed support for action agents. Simulation and physical robot implementations of our algorithms on a multiagent robot testbed demonstrated the efficacy and feasibility of our coordination framework by reducing the overall operation times in a benchmark wildfire-fighting case study. Esmaeil Seraj, Letian Chen, Matthew C. Gombolay |
IEEE Trans. Robotics | 3 |
| 2021 | Encoding Human Domain Knowledge to Warm Start Reinforcement LearningabstractDeep reinforcement learning has been successful in a variety of tasks, such as game playing and robotic manipulation. However, attempting to learn tabula rasa disregards the logical structure of many domains as well as the wealth of readily available knowledge from domain experts that could help "warm start" the learning process. We present a novel reinforcement learning technique that allows for intelligent initialization of a neural network weights and architecture. Our approach permits the encoding domain knowledge directly into a neural decision tree, and improves upon that knowledge with policy gradient updates. We empirically validate our approach on two OpenAI Gym tasks and two modified StarCraft 2 tasks, showing that our novel architecture outperforms multilayer-perceptron and recurrent architectures. Our knowledge-based framework finds superior policies compared to imitation learning-based and prior knowledge-based approaches. Importantly, we demonstrate that our approach can be used by untrained humans to initially provide >80% increase in expected reward relative to baselines prior to training (p < 0.001), which results in a >60% increase in expected reward after policy optimization (p = 0.011). Andrew Silva, Matthew C. Gombolay |
AAAI | 2 |
| 2021 | The Effects of a Robot's Performance on Human Teachers for Learning from Demonstration TasksabstractLearning from Demonstration (LfD) algorithms seek to enable end-users to teach robots new skills through human demonstration of a task. Previous studies have analyzed how robot failure affects human trust, but not in the context of the human teaching the robot. In this paper, we investigate how human teachers react to robot failure in an LfD setting. We conduct a study in which participants teach a robot how to complete three tasks, using one of three instruction methods, while the robot is pre-programmed to either succeed or fail at the task. We find that when the robot fails, people trust the robot less (p < .001$) and themselves less (p=.004) and they believe that others will trust them less (p < .001$). Human teachers also have a lower impression of the robot and themselves (p < .001) and found the task more difficult when the robot fails (p < .001$). Motion capture was found to be a less difficult instruction method than teleoperation (p=.016), while kinesthetic teaching gave the teachers the lowest impression of themselves compared to teleoperation (p=.017) and motion capture (p < .001). Importantly, a mediation analysis showed that people's trust in themselves is heavily mediated by what they think that others -- including the robot -- think of them (p < .001). These results provide valuable insights to improving the human-robot relationship for LfD. Erin Hedlund-Botti, Michael J. Johnson, Matthew C. Gombolay |
HRI | 3 |
| 2021 | Effects of Social Factors and Team Dynamics on Adoption of Collaborative Robot AutonomyabstractAs automation becomes more prevalent, the fear of job loss due to automation increases [22]. Workers may not be amenable to working with a robotic co-worker due to a negative perception of the technology. The attitudes of workers towards automation are influenced by a variety of complex and multi-faceted factors such as intention to use, perceived usefulness and other external variables [15]. In an analog manufacturing environment, we explore how these various factors influence an individual's willingness to work with a robot over a human co-worker in a collaborative Lego building task. We specifically explore how this willingness is affected by: 1) the level of social rapport established between the individual and his or her human co-worker, 2) the anthropomorphic qualities of the robot, and 3) factors including trust, fluency and personality traits. Our results show that a participant's willingness to work with automation decreased due to lower perceived team fluency (p=0.045), rapport established between a participant and their co-worker (p=0.003), the gender of the participant being male (p=0.041), and a higher inherent trust in people (p=0.018). Mariah Schrum, Glen Neville, Michael J. Johnson, Nina Moorman, Rohan R. Paleja, Karen M. Feigh, Matthew C. Gombolay |
HRI | 7 |
| 2021 | Towards a Comprehensive Understanding and Accurate Evaluation of Societal Biases in Pre-Trained TransformersabstractAndrew Silva, Pradyumna Tambwekar, Matthew Gombolay. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Andrew Silva, Pradyumna Tambwekar, Matthew C. Gombolay |
NAACL-HLT | 3 |
| 2021 | The Utility of Explainable AI in Ad Hoc Human-Machine TeamingabstractRecent advances in machine learning have led to growing interest in Explainable AI (xAI) to enable humans to gain insight into the decision-making of machine learning models. Despite this recent interest, the utility of xAI techniques has not yet been characterized in human-machine teaming. Importantly, xAI offers the promise of enhancing team situational awareness (SA) and shared mental model development, which are the key characteristics of effective human-machine teams. Rapidly developing such mental models is especially critical in ad hoc human-machine teaming, where agents do not have a priori knowledge of others' decision-making strategies. In this paper, we present two novel human-subject experiments quantifying the benefits of deploying xAI techniques within a human-machine teaming scenario. First, we show that xAI techniques can support SA ($p<0.05)$. Second, we examine how different SA levels induced via a collaborative AI policy abstraction affect ad hoc human-machine teaming performance. Importantly, we find that the benefits of xAI are not universal, as there is a strong dependence on the composition of the human-machine team. Novices benefit from xAI providing increased SA ($p<0.05$) but are susceptible to cognitive overhead ($p<0.05$). On the other hand, expert performance degrades with the addition of xAI-based support ($p<0.05$), indicating that the cost of paying attention to the xAI outweighs the benefits obtained from being provided additional information to enhance SA. Our results demonstrate that researchers must deliberately design and deploy the right xAI techniques in the right scenario by carefully considering human-machine team composition and how the xAI method augments SA. Rohan R. Paleja, Muyleng Ghuy, Nadun Ranawaka Arachchige, Reed Jensen, Matthew C. Gombolay |
NeurIPS | 5 |
| 2020 | Optimization Methods for Interpretable Differentiable Decision Trees Applied to Reinforcement LearningabstractDecision trees are ubiquitous in machine learning for their ease of use and interpretability. Yet, these models are not typically employed in reinforcement learning as they cannot be updated online via stochastic gradient descent. We overcome this limitation by allowing for a gradient update over the entire tree that improves sample complexity affords interpretable policy extraction. First, we include theoretical motivation on the need for policy-gradient learning by examining the properties of gradient descent over differentiable decision trees. Second, we demonstrate that our approach equals or outperforms a neural network on all domains and can learn discrete decision trees online with average rewards up to 7x higher than a batch-trained decision tree. Third, we conduct a user study to quantify the interpretability of a decision tree, rule list, and a neural network with statistically significant results (p < 0.001). Andrew Silva, Matthew C. Gombolay, Taylor W. Killian, Ivan Dario Jimenez Jimenez, Sung-Hyun Son |
AISTATS | 2 |
| 2020 | Joint Goal and Strategy Inference across Heterogeneous Demonstrators via Reward Network DistillationabstractReinforcement learning (RL) has achieved tremendous success as a general framework for learning how to make decisions. However, this success relies on the interactive hand-tuning of a reward function by RL experts. On the other hand, inverse reinforcement learning (IRL) seeks to learn a reward function from readily-obtained human demonstrations. Yet, IRL suffers from two major limitations: 1) reward ambiguity - there are an infinite number of possible reward functions that could explain an expert's demonstration and 2) heterogeneity - human experts adopt varying strategies and preferences, which makes learning from multiple demonstrators difficult due to the common assumption that demonstrators seeks to maximize the same reward. In this work, we propose a method to jointly infer a task goal and humans' strategic preferences via network distillation. This approach enables us to distill a robust task reward (addressing reward ambiguity) and to model each strategy's objective (handling heterogeneity). We demonstrate our algorithm can better recover task reward and strategy rewards and imitate the strategies in two simulated tasks and a real-world table tennis task. Letian Chen, Rohan R. Paleja, Muyleng Ghuy, Matthew C. Gombolay |
HRI | 4 |
| 2020 | Effects of Anthropomorphism and Accountability on Trust in Human Robot InteractionabstractThis paper examines how people's trust and dependence on robot teammates providing decision support varies as a function of different attributes of the robot, such as perceived anthropomorphism, type of support provided by the robot, and its physical presence. We conduct a mixed-design user study with multiple robots to investigate trust, inappropriate reliance, and compliance measures in the context of a time-constrained game. We also examine how the effect of human accountability addresses errors due to over-compliance in the context of human robot interaction (HRI). This study is novel as it involves examining multiple attributes at once, thus enabling us to perform multi-way comparisons between different attributes on trust and compliance with the agent. Results from the 4x4x2x2 study show that behavior and anthropomorphism of the agent are the most significant factors in predicting the trust and compliance with the robot. Furthermore, adding a coalition-building preface, where the agent provides context to why it might make errors while giving advice, leads to an increase in trust for specific behaviors of the agent. Manisha Natarajan, Matthew C. Gombolay |
HRI | 2 |
| 2020 | Interpretable and Personalized Apprenticeship Scheduling: Learning Interpretable Scheduling Policies from Heterogeneous User DemonstrationsabstractResource scheduling and coordination is an NP-hard optimization requiring an efficient allocation of agents to a set of tasks with upper- and lower bound temporal and resource constraints. Due to the large-scale and dynamic nature of resource coordination in hospitals and factories, human domain experts manually plan and adjust schedules on the fly. To perform this job, domain experts leverage heterogeneous strategies and rules-of-thumb honed over years of apprenticeship. What is critically needed is the ability to extract this domain knowledge in a heterogeneous and interpretable apprenticeship learning framework to scale beyond the power of a single human expert, a necessity in safety-critical domains. We propose a personalized and interpretable apprenticeship scheduling algorithm that infers an interpretable representation of all human task demonstrators by extracting decision-making criteria via an inferred, personalized embedding non-parametric in the number of demonstrator types. We achieve near-perfect LfD accuracy in synthetic domains and 88.22\% accuracy on a planning domain with real-world data, outperforming baselines. Finally, our user study showed our methodology produces more interpretable and easier-to-use models than neural networks ($p < 0.05$). Rohan R. Paleja, Andrew Silva, Letian Chen, Matthew C. Gombolay |
NeurIPS | 4 |
| 2020 | A Tale of Two Suggestions: Action and Diagnosis Recommendations for Responding to Robot FailureabstractRobots operating without close human supervision might need to rely on a remote call center of operators for assistance in the event of a failure. In this work, we investigate the effects of providing decision support through diagnosis suggestions, as feedback, and action recommendations, as feedforward, to the human operators. We conduct a 10-condition user study involving 200 participants on Amazon Mechanical Turk to evaluate the effects of providing noisy and noise-free diagnosis suggestions and/or action recommendations to operators. We find that although action recommendations (feedforward) have a greater effect on successful error resolution than diagnosis information (feedback), the feedback likely helps ameliorate the deleterious effects of noise. Therefore, we find that error recovery interfaces should display both diagnosis and action recommendations for maximum effectiveness. Siddhartha Banerjee, Matthew C. Gombolay, Sonia Chernova |
RO-MAN | 2 |
| 2019 | Heterogeneous Learning from DemonstrationabstractThe development of human-robot systems able to leverage the strengths of both humans and their robotic counterparts has been greatly sought after because of the foreseen, broad-ranging impact across industry and research. We believe the true potential of these systems cannot be reached unless the robot is able to act with a high level of autonomy, reducing the burden of manual tasking or teleoperation. To achieve this level of autonomy, robots must be able to work fluidly with its human partners, inferring their needs without explicit commands. This inference requires the robot to be able to detect and classify the heterogeneity of its partners. We propose a framework for learning from heterogeneous demonstration based upon Bayesian inference and evaluate a suite of approaches on a real-world dataset of gameplay from StarCraft II. This evaluation provides evidence that our Bayesian approach can outperform conventional methods by up to 12.8%. Rohan R. Paleja, Matthew C. Gombolay |
HRI | 2 |
| 2019 | Human Trust After Robot Mistakes: Study of the Effects of Different Forms of Robot CommunicationabstractCollaborative robots that work alongside humans will experience service breakdowns and make mistakes. These robotic failures can cause a degradation of trust between the robot and the community being served. A loss of trust may impact whether a user continues to rely on the robot for assistance. In order to improve the teaming capabilities between humans and robots, forms of communication that aid in developing and maintaining trust need to be investigated. In our study, we identify four forms of communication which dictate the timing of information given and type of initiation used by a robot. We investigate the effect that these forms of communication have on trust with and without robot mistakes during a cooperative task. Participants played a memory task game with the help of a humanoid robot that was designed to make mistakes after a certain amount of time passed. The results showed that participants' trust in the robot was better preserved when that robot offered advice only upon request as opposed to when the robot took initiative to give advice. Sean Ye, Glen Neville, Mariah Schrum, Matthew C. Gombolay, Sonia Chernova, Ayanna M. Howard |
RO-MAN | 4 |
| 2019 | Machine Learning Techniques for Analyzing Training Behavior in Serious GamingabstractTraining time is a costly, scarce resource across domains such as commercial aviation, healthcare, and military operations. In the context of military applications, serious gaming—the training of warfighters through immersive, real-time environments rather than traditional classroom lectures—offers benefits to improve training not only in its hands-on development and application of knowledge, but also in data analytics via machine learning. In this paper, we explore an array of machine learning techniques that allow teachers to visualize the degree to which training objectives are reflected in actual play. First, we investigate the concept of discovery: learning how warfighters utilize their training tools and develop military strategies within their training environment. Second, we develop machine learning techniques that could assist teachers by automatically predicting player performance, identifying player disengagement, and recommending personalized lesson plans. These methods could potentially provide teachers with insight to assist them in developing better lesson plans and tailored instruction for each individual student. Matthew C. Gombolay, Reed Jensen, Sung-Hyun Son |
IEEE Trans. Games | 1 |
| 2018 | Learning to Infer Final Plans in Human Team PlanningabstractWe envision an intelligent agent that analyzes conversations during human team meetings in order to infer the team’s plan, with the purpose of providing decision support to strengthen that plan. We present a novel learning technique to infer teams' final plans directly from a processed form of their planning conversation. Our method employs reinforcement learning to train a model that maps features of the discussed plan and patterns of dialogue exchange among participants to a final, agreed-upon plan. We employ planning domain models to efficiently search the large space of possible plans, and the costs of candidate plans serve as the reinforcement signal. We demonstrate that our technique successfully infers plans within a variety of challenging domains, with higher accuracy than prior art. With our domain-independent feature set, we empirically demonstrate that our model trained on one planning domain can be applied to successfully infer team plans within a novel planning domain. Joseph Kim, Matthew E. Woicik, Matthew C. Gombolay, Sung-Hyun Son, Julie A. Shah |
IJCAI | 3 |
| 2018 | Human-Machine Collaborative Optimization via Apprenticeship SchedulingabstractCoordinating agents to complete a set of tasks with intercoupled temporal and resource constraints is computationally challenging, yet human domain experts can solve these difficult scheduling problems using paradigms learned through years of apprenticeship. A process for manually codifying this domain knowledge within a computational framework is necessary to scale beyond the "single-expert, single-trainee" apprenticeship model. However, human domain experts often have difficulty describing their decision-making processes. We propose a new approach for capturing this decision-making process through counterfactual reasoning in pairwise comparisons. Our approach is model-free and does not require iterating through the state space. We demonstrate that this approach accurately learns multifaceted heuristics on a synthetic and real world data sets. We also demonstrate that policies learned from human scheduling demonstration via apprenticeship learning can substantially improve the efficiency of schedule optimization. We employ this human-machine collaborative optimization technique on a variant of the weapon-to-target assignment problem. We demonstrate that this technique generates optimal solutions up to 9.5 times faster than a state-of-the-art optimization algorithm. Matthew C. Gombolay, Reed Jensen, Jessica Stigile, Toni Golen, Sung-Hyun Son, Julie A. Shah |
J. Artif. Intell. Res. | 1 |
| 2018 | Fast Scheduling of Robot Teams Performing Tasks With Temporospatial ConstraintsabstractThe application of robotics to traditionally manual manufacturing processes requires careful coordination between human and robotic agents in order to support safe and efficient coordinated work. Tasks must be allocated to agents and sequenced according to temporal and spatial constraints. Also, systems must be capable of responding on-the-fly to disturbances and people working in close physical proximity to robots. In this paper, we present a centralized algorithm, named “Tercio,” that handles tightly intercoupled temporal and spatial constraints. Our key innovation is a fast, satisficing multi-agent task sequencer inspired by real-time processor scheduling techniques and adapted to leverage a hierarchical problem structure. We use this sequencer in conjunction with a mixed-integer linear program solver and empirically demonstrate the ability to generate near-optimal schedules for real-world problems an order of magnitude larger than those reported in prior art. Finally, we demonstrate the use of our algorithm in a multirobot hardware testbed. Matthew C. Gombolay, Ronald Wilcox, Julie A. Shah |
IEEE Trans. Robotics | 1 |
| 2016 | Apprenticeship Scheduling for Human-Robot TeamsabstractResource optimization and scheduling is a costly, challenging problem that affects almost every aspect of our lives. One example that affects each of us is health care: Poor systems design and scheduling of resources can lead to higher rates of patient noncompliance and burnout of health care providers, as highlighted by the Institute of Medicine (Brandenburg et al. 2015). In aerospace manufacturing, every minute re-scheduling in response to dynamic disruptions in the build process of a Boeing 747 can cost up to $100.000. The military is also highly invested in the effective use of resources. In missile defense, for example, operators must =solve a challenging weapon-to-target problem, balancing the cost of expendable, defensive weapons while hedging against uncertainty in adversaries’ tactics. Researchers in artificial intelligence (AI) planning and scheduling strive to develop algorithms to improve resource allocation. However, there are two primary challenges. First, optimal task allocation and sequencing with upper and lower-bound temporal constraints (i.e., deadlines and wait constraints) is NP-Hard (Bertsimas and Weismantel 2005). Approximation techniques for scheduling exist and typically rely on the algorithm designer crafting heuristics based on domain expertise to decompose or structure the scheduling problem and prioritize the manner in which resources are allocated and tasks are sequenced (Tang and Parker 2005; Jones, Dias, and Stentz 2011). The second problem is this aforementioned reliance on crafting clever heuristics based on domain knowledge. Manually capturing domain knowledge within a scheduling algorithm remains a challenging process and leaves much to be desired (Ryan et al. 2013). The aim of my thesis is to develop an autonomous system that 1) learns the heuristics and implicit rules-of-thumb developed by domain experts from years of experience, 2) embeds and leverages this knowledge within a scalable resource optimization framework, and 3) provides decision support in a way that engages users and benefits them in their decision-making process. By intelligently leveraging the ability of humans to learn heuristics and the speed of modern computation, we can improve the ability to coordinate resources in these time and safety-critical domains. Matthew C. Gombolay |
AAAI | 1 |
| 2016 | Apprenticeship Scheduling: Learning to Schedule from Human Experts
Matthew C. Gombolay, Reed Jensen, Jessica Stigile, Sung-Hyun Son, Julie A. Shah |
IJCAI | 1 |