VLDB 2026 Research / reviewers in the wild / expert
Rohan R. Paleja
dblp:237/8623
· DBLP profile ↗
13ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-0773-8054ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Asynchronous Training of Mixed-Role Human Actors in a Partially Observable EnvironmentabstractIn cooperative training, humans within a team coordinate on complex tasks, building mental models of their teammates and learning to adapt to teammates’ actions in real-time. To reduce the often prohibitive scheduling constraints associated with cooperative training, this article introduces a paradigm for cooperative asynchronous training of human teams in which trainees practice coordination with autonomous teammates rather than humans. We introduce a novel experimental design for evaluating autonomous teammates for use as training partners in cooperative training. We apply this design to a human-subjects experiment where humans are trained with either another human or an autonomous teammate and are evaluated with a new human subject in a new, partially observable, cooperative game developed for this study. Importantly, we employ an unsupervised sequential clustering methodology to partition teammate trajectories from demonstrations performed in the experiment to form a smaller number of training conditions. This results in a simpler experiment design, enabling us to conduct a complex cooperative training human-subjects study in a reasonable amount of time. Through a demonstration of the proposed experimental design, we provide takeaways and design recommendations for future research in the development of cooperative asynchronous training systems utilizing robot surrogates for human teammates. Kimberlee Chestnut Chang, Reed Jensen, Rohan R. Paleja, Sam L. Polk, Robert Seater, Jackson Steilberg, Curran Schiefelbein, Melissa Scheldrup, Matthew C. Gombolay, Mabel D. Ramirez |
ACM Trans. Hum. Robot Interact. | 3 |
| 2025 | Generalized Behavior Learning from Diverse DemonstrationsabstractDiverse behavior policies are valuable in domains requiring quick test-time adaptation or personalized human-robot interaction. Human demonstrations provide rich information regarding task objectives and factors that govern individual behavior variations, which can be used to characterize \textit{useful} diversity and learn diverse performant policies.
However, we show that prior work that builds naive representations of demonstration heterogeneity fails in generating successful novel behaviors that generalize over behavior factors.
We propose Guided Strategy Discovery (GSD), which introduces a novel diversity formulation based on a learned task-relevance measure that prioritizes behaviors exploring modeled latent factors.
We empirically validate across three continuous control benchmarks for generalizing to in-distribution (interpolation) and out-of-distribution (extrapolation) factors that GSD outperforms baselines in novel behavior discovery by $\sim$21\%.
Finally, we demonstrate that GSD can generalize striking behaviors for table tennis in a virtual testbed while leveraging human demonstrations collected in the real world.
Code is available at https://github.com/CORE-Robotics-Lab/GSD. Varshith Sreeramdass, Rohan R. Paleja, Letian Chen, Sanne van Waveren, Matthew C. Gombolay |
ICLR | 2 |
| 2024 | Unsupervised Behavior Inference From Human Action SequencesabstractTo effectively coordinate with human teammates in environments where direct communication is limited or unavailable, autonomous agents must be able to infer macro-level strategies from demonstrated micro-level actions in real time. This article introduces UNsupervised Behavior Inference from Action Sequences (UNBIAS): a modular framework for the unsupervised learning of behaviors from sequences of demonstrated human actions. UNBIAS relies on an unsupervised sequential encoder to learn representations of behaviors from sequences of human actions. Behavior representations are then clustered to obtain an unsupervised behavior classification. Real-time strategy prediction is performed using UNBIAS by training a classification layer to learn strategies’ decision boundaries in the behavior representation space and predict cluster labels from encoded incomplete action sequences. In extensive numerical experiments on two real-world datasets of human decision sequences, behavior classifications learned through the UNBIAS framework are shown to better match demonstrator identity groupings than partitions obtained with related unsupervised clustering algorithms, and high-quality online strategy inference is demonstrated. Thus, UNBIAS effectively enables the unsupervised learning of human behaviors from observed actions, enabling autonomous agents to better adapt to human teammates’ strategies in human-machine teams. Sam L. Polk, Eric J. Pabon Cancel, Rohan R. Paleja, Kimberlee Chestnut Chang, Reed Jensen, Mabel D. Ramirez |
CoG | 3 |
| 2024 | Designs for Enabling Collaboration in Human-Machine Teaming via Interactive and Explainable SystemsabstractCollaborative robots and machine learning-based virtual agents are increasingly entering the human workspace with the aim of increasing productivity and enhancing safety. Despite this, we show in a ubiquitous experimental domain, Overcooked-AI, that state-of-the-art techniques for human-machine teaming (HMT), which rely on imitation or reinforcement learning, are brittle and result in a machine agent that aims to decouple the machine and human’s actions to act independently rather than in a synergistic fashion. To remedy this deficiency, we develop HMT approaches that enable iterative, mixed-initiative team development allowing end-users to interactively reprogram interpretable AI teammates. Our 50-subject study provides several findings that we summarize into guidelines. While all approaches underperform a simple collaborative heuristic (a critical, negative result for learning-based methods), we find that white-box approaches supported by interactive modification can lead to significant team development, outperforming white-box approaches alone, and that black-box approaches are easier to train and result in better HMT performance highlighting a tradeoff between explainability and interactivity versus ease-of-training. Together, these findings present three important future research directions: 1) Improving the ability to generate collaborative agents with white-box models, 2) Better learning methods to facilitate collaboration rather than individualized coordination, and 3) Mixed-initiative interfaces that enable users, who may vary in ability, to improve collaboration. Rohan R. Paleja, Michael Munje, Kimberlee Chestnut Chang, Reed Jensen, Matthew C. Gombolay |
NeurIPS | 1 |
| 2024 | Heterogeneous Policy Networks for Composite Robot Team Communication and CoordinationabstractHigh-performing human–human teams learn intelligent and efficient communication and coordination strategies to maximize their joint utility. These teams implicitly understand the different roles of heterogeneous team members and adapt their communication protocols accordingly. Multiagent reinforcement learning (MARL) has attempted to develop computational methods for synthesizing such joint coordination–communication strategies, but emulating heterogeneous communication patterns across agents with different state, action, and observation spaces has remained a challenge. Without properly modeling agent heterogeneity, as in prior MARL work that leverages homogeneous graph networks, communication becomes less helpful and can even deteriorate the team's performance. In the past, we proposed heterogeneous policy networks (HetNet) to learn efficient and diverse communication models for coordinating cooperative heterogeneous teams. In this extended work, we extend HetNet to support scaling heterogeneous robot teams. Building on heterogeneous graph-attention networks, we show that HetNet not only facilitates learning heterogeneous collaborative policies, but also enables end-to-end training for learning highly efficient binarized messaging. Our empirical evaluation shows that HetNet sets a new state-of-the-art in learning coordination and communication strategies for heterogeneous multiagent teams by achieving an 5.84% to 707.65% performance improvement over the next-best baseline across multiple domains while simultaneously achieving a 200× reduction in the required communication bandwidth. Esmaeil Seraj, Rohan R. Paleja, Luis Pimentel, Kin Man Lee, Zheyuan Wang, Matthew Sklar, John Z. Zhang, Zahi M. Kakish, Matthew C. Gombolay |
IEEE Trans. Robotics | 2 |
| 2023 | The Effect of Robot Skill Level and Communication in Rapid, Proximate Human-Robot CollaborationabstractAs high-speed, agile robots become more commonplace, these robots will have the potential to better aid and collaborate with humans. However, due to the increased agility and functionality of these robots, close collaboration with humans can create safety concerns that alter team dynamics and degrade task performance. In this work, we aim to enable the deployment of safe and trustworthy agile robots that operate in proximity with humans. We do so by 1) Proposing a novel human-robot doubles table tennis scenario to serve as a testbed for studying agile, proximate human-robot collaboration and 2) Conducting a user-study to understand how attributes of the robot (e.g., robot competency or capacity to communicate) impact team dynamics, perceived safety, and perceived trust, and how these latent factors affect human-robot collaboration (HRC) performance. We find that robot competency significantly increases perceived trust (p < .001), extending skill-to-trust assessments in prior studies to agile, proximate HRC. Furthermore, interestingly, we find that when the robot vocalizes its intention to perform a task, it results in a significant decrease in team performance (p = .037) and perceived safety of the system (p = .009). Kin Man Lee, Arjun Krishna, Zulfiqar Zaidi, Rohan R. Paleja, Letian Chen, Erin Hedlund-Botti, Mariah Schrum, Matthew C. Gombolay |
HRI | 4 |
| 2023 | Learning Models of Adversarial Agent Behavior Under Partial ObservabilityabstractThe need for opponent modeling and tracking arises in several real-world scenarios, such as professional sports, video game design, and drug-trafficking interdiction. In this work, we present Graph based Adversarial Modeling with Mutual Information (GrAMMI) for modeling the behavior of an adversarial opponent agent. GrAMMI is a novel graph neural network (GNN) based approach that uses mutual information maximization as an auxiliary objective to predict the current and future states of an adversarial opponent with partial observability. To evaluate GrAMMI, we design two large-scale, pursuit-evasion domains inspired by real-world scenarios, where a team of heterogeneous agents is tasked with tracking and interdicting a single adversarial agent, and the adversarial agent must evade detection while achieving its own objectives. With the mutual information formulation, GrAMMI outperforms all baselines in both domains and achieves 31.68% higher log-likelihood on average for future adversarial state predictions across both domains. Sean Ye, Manisha Natarajan, Rohan R. Paleja, Letian Chen, Matthew C. Gombolay |
IROS | 4 |
| 2022 | Mutual Understanding in Human-Machine TeamingabstractCollaborative robots (i.e., "cobots") and machine learning-based virtual agents are increasingly entering the human workspace with the aim of increasing productivity, enhancing safety, and improving the quality of our lives. These agents will dynamically interact with a wide variety of people in dynamic and novel contexts, increasing the prevalence of human-machine teams in healthcare, manufacturing, and search-and-rescue. In this research, we enhance the mutual understanding within a human-machine team by enabling cobots to understand heterogeneous teammates via person-specific embeddings, identifying contexts in which xAI methods can help improve team mental model alignment, and enabling cobots to effectively communicate information that supports high-performance human-machine teaming. Rohan R. Paleja |
AAAI | 1 |
| 2021 | Effects of Social Factors and Team Dynamics on Adoption of Collaborative Robot AutonomyabstractAs automation becomes more prevalent, the fear of job loss due to automation increases [22]. Workers may not be amenable to working with a robotic co-worker due to a negative perception of the technology. The attitudes of workers towards automation are influenced by a variety of complex and multi-faceted factors such as intention to use, perceived usefulness and other external variables [15]. In an analog manufacturing environment, we explore how these various factors influence an individual's willingness to work with a robot over a human co-worker in a collaborative Lego building task. We specifically explore how this willingness is affected by: 1) the level of social rapport established between the individual and his or her human co-worker, 2) the anthropomorphic qualities of the robot, and 3) factors including trust, fluency and personality traits. Our results show that a participant's willingness to work with automation decreased due to lower perceived team fluency (p=0.045), rapport established between a participant and their co-worker (p=0.003), the gender of the participant being male (p=0.041), and a higher inherent trust in people (p=0.018). Mariah Schrum, Glen Neville, Michael J. Johnson, Nina Moorman, Rohan R. Paleja, Karen M. Feigh, Matthew C. Gombolay |
HRI | 5 |
| 2021 | The Utility of Explainable AI in Ad Hoc Human-Machine TeamingabstractRecent advances in machine learning have led to growing interest in Explainable AI (xAI) to enable humans to gain insight into the decision-making of machine learning models. Despite this recent interest, the utility of xAI techniques has not yet been characterized in human-machine teaming. Importantly, xAI offers the promise of enhancing team situational awareness (SA) and shared mental model development, which are the key characteristics of effective human-machine teams. Rapidly developing such mental models is especially critical in ad hoc human-machine teaming, where agents do not have a priori knowledge of others' decision-making strategies. In this paper, we present two novel human-subject experiments quantifying the benefits of deploying xAI techniques within a human-machine teaming scenario. First, we show that xAI techniques can support SA ($p<0.05)$. Second, we examine how different SA levels induced via a collaborative AI policy abstraction affect ad hoc human-machine teaming performance. Importantly, we find that the benefits of xAI are not universal, as there is a strong dependence on the composition of the human-machine team. Novices benefit from xAI providing increased SA ($p<0.05$) but are susceptible to cognitive overhead ($p<0.05$). On the other hand, expert performance degrades with the addition of xAI-based support ($p<0.05$), indicating that the cost of paying attention to the xAI outweighs the benefits obtained from being provided additional information to enhance SA. Our results demonstrate that researchers must deliberately design and deploy the right xAI techniques in the right scenario by carefully considering human-machine team composition and how the xAI method augments SA. Rohan R. Paleja, Muyleng Ghuy, Nadun Ranawaka Arachchige, Reed Jensen, Matthew C. Gombolay |
NeurIPS | 1 |
| 2020 | Joint Goal and Strategy Inference across Heterogeneous Demonstrators via Reward Network DistillationabstractReinforcement learning (RL) has achieved tremendous success as a general framework for learning how to make decisions. However, this success relies on the interactive hand-tuning of a reward function by RL experts. On the other hand, inverse reinforcement learning (IRL) seeks to learn a reward function from readily-obtained human demonstrations. Yet, IRL suffers from two major limitations: 1) reward ambiguity - there are an infinite number of possible reward functions that could explain an expert's demonstration and 2) heterogeneity - human experts adopt varying strategies and preferences, which makes learning from multiple demonstrators difficult due to the common assumption that demonstrators seeks to maximize the same reward. In this work, we propose a method to jointly infer a task goal and humans' strategic preferences via network distillation. This approach enables us to distill a robust task reward (addressing reward ambiguity) and to model each strategy's objective (handling heterogeneity). We demonstrate our algorithm can better recover task reward and strategy rewards and imitate the strategies in two simulated tasks and a real-world table tennis task. Letian Chen, Rohan R. Paleja, Muyleng Ghuy, Matthew C. Gombolay |
HRI | 2 |
| 2020 | Interpretable and Personalized Apprenticeship Scheduling: Learning Interpretable Scheduling Policies from Heterogeneous User DemonstrationsabstractResource scheduling and coordination is an NP-hard optimization requiring an efficient allocation of agents to a set of tasks with upper- and lower bound temporal and resource constraints. Due to the large-scale and dynamic nature of resource coordination in hospitals and factories, human domain experts manually plan and adjust schedules on the fly. To perform this job, domain experts leverage heterogeneous strategies and rules-of-thumb honed over years of apprenticeship. What is critically needed is the ability to extract this domain knowledge in a heterogeneous and interpretable apprenticeship learning framework to scale beyond the power of a single human expert, a necessity in safety-critical domains. We propose a personalized and interpretable apprenticeship scheduling algorithm that infers an interpretable representation of all human task demonstrators by extracting decision-making criteria via an inferred, personalized embedding non-parametric in the number of demonstrator types. We achieve near-perfect LfD accuracy in synthetic domains and 88.22\% accuracy on a planning domain with real-world data, outperforming baselines. Finally, our user study showed our methodology produces more interpretable and easier-to-use models than neural networks ($p < 0.05$). Rohan R. Paleja, Andrew Silva, Letian Chen, Matthew C. Gombolay |
NeurIPS | 1 |
| 2019 | Heterogeneous Learning from DemonstrationabstractThe development of human-robot systems able to leverage the strengths of both humans and their robotic counterparts has been greatly sought after because of the foreseen, broad-ranging impact across industry and research. We believe the true potential of these systems cannot be reached unless the robot is able to act with a high level of autonomy, reducing the burden of manual tasking or teleoperation. To achieve this level of autonomy, robots must be able to work fluidly with its human partners, inferring their needs without explicit commands. This inference requires the robot to be able to detect and classify the heterogeneity of its partners. We propose a framework for learning from heterogeneous demonstration based upon Bayesian inference and evaluate a suite of approaches on a real-world dataset of gameplay from StarCraft II. This evaluation provides evidence that our Bayesian approach can outperform conventional methods by up to 12.8%. Rohan R. Paleja, Matthew C. Gombolay |
HRI | 1 |