Huao Li

dblp:30/783 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
9since 2021 · last 2025
0000-0002-0027-615XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 9 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Adaptively Coordinating with Novel Partners via Learned Latent Strategies
abstract
Adaptation is the cornerstone of effective collaboration among heterogeneous team members. In human-agent teams, artificial agents need to adapt to their human partners in real time, as individuals often have unique preferences and policies that may change dynamically throughout interactions. This becomes particularly challenging in tasks with time pressure and complex strategic spaces, where identifying partner behaviors and selecting suitable responses is difficult. In this work, we introduce a strategy-conditioned cooperator framework that learns to represent, categorize, and adapt to a broad range of potential partner strategies in real-time. Our approach encodes strategies with a variational autoencoder to learn a latent strategy space from agent trajectory data, identifies distinct strategy types through clustering, and trains a cooperator agent conditioned on these clusters by generating partners of each strategy type. For online adaptation to novel partners, we leverage a fixed-share regret minimization algorithm that dynamically infers and adjusts the partner's strategy estimation during interaction. We evaluate our method in a modified version of the Overcooked domain, a complex collaborative cooking environment that requires effective coordination among two players with a diverse potential strategy space. Through these experiments and an online user study, we demonstrate that our proposed agent achieves state of the art performance compared to existing baselines when paired with novel human, and agent teammates.
Benjamin Li, Shuyang Shi, Lucia Romero, Huao Li, Yaqi Xie 0001, Woojun Kim, Stefanos Nikolaidis, Charles Lewis, Katia P. Sycara, Simon Stepputtis
NeurIPS4
2024 Language Grounded Multi-agent Reinforcement Learning with Human-interpretable Communication
abstract
Multi-Agent Reinforcement Learning (MARL) methods have shown promise in enabling agents to learn a shared communication protocol from scratch and accomplish challenging team tasks. However, the learned language is usually not interpretable to humans or other agents not co-trained together, limiting its applicability in ad-hoc teamwork scenarios. In this work, we propose a novel computational pipeline that aligns the communication space between MARL agents with an embedding space of human natural language by grounding agent communications on synthetic data generated by embodied Large Language Models (LLMs) in interactive teamwork scenarios. Our results demonstrate that introducing language grounding not only maintains task performance but also accelerates the emergence of communication. Furthermore, the learned communication protocols exhibit zero-shot generalization capabilities in ad-hoc teamwork scenarios with unseen teammates and novel task states. This work presents a significant step toward enabling effective communication and collaboration between artificial agents and humans in real-world teamwork settings.
Huao Li, Hossein Nourkhiz Mahjoub, Behdad Chalaki, Vaishnav Tadiparthi, Kwonjoon Lee, Ehsan Moradi-Pari, Charles Lewis, Katia P. Sycara
NeurIPS1
2023 Theory of Mind for Multi-Agent Collaboration via Large Language Models
abstract
While Large Language Models (LLMs) have demonstrated impressive accomplishments in both reasoning and planning, their abilities in multi-agent collaborations remains largely unexplored.This study evaluates LLMbased agents in a multi-agent cooperative text game with Theory of Mind (ToM) inference tasks, comparing their performance with Multi-Agent Reinforcement Learning (MARL) and planning-based baselines.We observed evidence of emergent collaborative behaviors and high-order Theory of Mind capabilities among LLM-based agents.Our results reveal limitations in LLM-based agents' planning optimization due to systematic failures in managing long-horizon contexts and hallucination about the task state.We explore the use of explicit belief state representations to mitigate these issues, finding that it enhances task performance and the accuracy of ToM inferences for LLMbased agents.
Huao Li, Yu Quan Chong, Simon Stepputtis, Joseph Campbell, Dana Hughes 0001, Charles Lewis, Katia P. Sycara
EMNLP1
2023 A Framework for Intervention Based Team Support in Time Critical Tasks
abstract
In this paper we describe the intervention framework of ATLAS, an artificial socially intelligent agent that advises teams. The framework treats interventions as atomic components, and manages the lifecycle of each intervention through presentation, as well as followups to interventions. The key benefit of this framework is that it allows for rapid development of scenario-specific Interventions that leverage scenario-agnostic team models. The implementation of this framework is reported for three player teams in a Search and Rescue task simulated in Minecraft. Low competence teams advised by ATLAS improved more between first and second trials than those with a human advisor while the reverse was found for high competence. Four times as many interventions were proposed as were presented. 15 % of advice was withheld to avoid repetitive advice, excessive rate of advice, and needlessly advising high performing teams, while a Theory of Mind model and delay for confirmation mechanism filtered out other unnecessary advice.
Dana Hughes 0001, Huao Li, Max Chis, Ini Oguntola, Simon Stepputtis, Keyang Zheng, Joseph Campbell, Katia P. Sycara, Michael Lewis 0001
SMC2
2023 Personalized Decision Supports based on Theory of Mind Modeling and Explainable Reinforcement Learning
abstract
In this paper, we propose a novel personalized decision support system that combines Theory of Mind (ToM) modeling and explainable Reinforcement Learning (XRL) to provide effective and interpretable interventions. Our method leverages DRL to provide expert action recommendations while incorporating ToM modeling to understand users' mental states and predict their future actions, enabling appropriate timing for intervention. To explain interventions, we use counterfactual explanations based on RL's feature importance and users' ToM model structure. Our proposed system generates accurate and personalized interventions that are easily interpretable by end-users. We demonstrate the effectiveness of our approach through a series of crowd-sourcing experiments in a simulated team decision-making task, where our system outperforms control baselines in terms of task performance. Our proposed approach is agnostic to task environment and RL model structure, therefore has the potential to be generalized to a wide range of applications.
Huao Li, Yao Fan, Keyang Zheng, Michael Lewis 0001, Katia P. Sycara
SMC1
2022 Theory of Mind Modeling in Search and Rescue Teams
abstract
Theory of Mind (ToM) refers to the ability to make inferences about other’s mental states. Such ability is fundamental for human social activities such as empathy, teamwork, and communication. As intelligent agents come to be involved in diverse human-agent teams, they will also be expected to be socially intelligent in order to become effective teammates. In this paper, we describe a computational ToM model which observes team behaviors and infers their mental states in a urban search and rescue (US&R) task. Our modular ToM model approximates human inference by explicitly representing beliefs, belief updates, and action prediction/generation using Deep Neural Networks (DNNs). To validate our model we compare its performance to the gold standard of human observers asked to make the same inferences. The ToM model proved superior to the average judgments of human observers on all four tests of inference and better than 90th percentile observers on three of the four. While the learning bias provided by modularizing belief and prediction proved sufficient for the simple inferences tested, substantial refinement will be needed to replicate the complex nuanced chains of inference observed in human social interaction.
Huao Li, Ini Oguntola, Dana Hughes 0001, Michael Lewis 0001, Katia P. Sycara
RO-MAN1
2021 Hiding Leader's Identity in Leader-Follower Navigation through Multi-Agent Reinforcement Learning
abstract
Leader-follower navigation is a popular class of multi-robot algorithms where a leader robot leads the follower robots in a team. The leader has specialized capabilities or mission critical information (e.g. goal location) that the followers lack, and this makes the leader crucial for the mission’s success. However, this also makes the leader a vulnerability -an external adversary who wishes to sabotage the robot team’s mission can simply harm the leader and the whole robot team’s mission would be compromised. Since robot motion generated by traditional leader-follower navigation algorithms can reveal the identity of the leader, we propose a defense mechanism of hiding the leader’s identity by ensuring the leader moves in a way that behaviorally camouflages it with the followers, making it difficult for an adversary to identify the leader. To achieve this, we combine Multi-Agent Reinforcement Learning, Graph Neural Networks and adversarial training. Our approach enables the multi-robot team to optimize the primary task performance with leader motion similar to follower motion, behaviorally camouflaging it with the followers. Our algorithm outperforms existing work that tries to hide the leader’s identity in a multi-robot team by tuning traditional leader-follower control parameters with Classical Genetic Algorithms. We also evaluated human performance in inferring the leader’s identity and found that humans had lower accuracy when the robot team used our proposed navigation algorithm.
Ankur Deka, Huao Li, Michael Lewis 0001, Katia P. Sycara
IROS3
2021 Emergent Discrete Communication in Semantic Spaces
abstract
Neural agents trained in reinforcement learning settings can learn to communicate among themselves via discrete tokens, accomplishing as a team what agents would be unable to do alone. However, the current standard of using one-hot vectors as discrete communication tokens prevents agents from acquiring more desirable aspects of communication such as zero-shot understanding. Inspired by word embedding techniques from natural language processing, we propose neural agent architectures that enables them to communicate via discrete tokens derived from a learned, continuous space. We show in a decision theoretic framework that our technique optimizes communication over a wide range of scenarios, whereas one-hot tokens are only optimal under restrictive assumptions. In self-play experiments, we validate that our trained agents learn to cluster tokens in semantically-meaningful ways, allowing them communicate in noisy environments where other techniques fail. Lastly, we demonstrate both that agents using our method can effectively respond to novel human communication and that humans can understand unlabeled emergent agent communication, outperforming the use of one-hot communication.
Mycal Tucker, Huao Li, Siddharth Agrawal, Dana Hughes 0001, Katia P. Sycara, Michael Lewis 0001, Julie A. Shah
NeurIPS2
2021 Individualized Mutual Adaptation in Human-Agent Teams
abstract
The ability to collaborate with previously unseen human teammates is crucial for artificial agents to be effective in human-agent teams (HATs). Due to individual differences and complex team dynamics, it is hard to develop a single agent policy to match all potential teammates. In this article, we study both human-human and HAT in a dyadic cooperative task, Team Space Fortress. Results show that the team performance is influenced by both players’ individual skill level and their ability to collaborate with different teammates by adopting complementary policies. Based on human-human team results, we propose an adaptive agent that identifies different human policies and assigns a complementary partner policy to optimize team performance. The adaptation method relies on a novel similarity metric to infer human policy and then selects the most complementary policy from a pretrained library of exemplar policies. We conducted human-agent experiments to evaluate the adaptive agent and examine mutual adaptation in HAT. Results show that both human adaptation and agent adaptation contribute to team performance.
Huao Li, Tianwei Ni, Siddharth Agrawal, Suhas Raja, Yikang Gui, Dana Hughes 0001, Michael Lewis 0001, Katia P. Sycara
IEEE Trans. Hum. Mach. Syst.1
2020 Individual adaptation in teamwork
Huao Li, Dana Hughes 0001, Michael Lewis 0001, Katia P. Sycara
CogSci1
2020 Models of Trust in Human Control of Swarms With Varied Levels of Autonomy
abstract
In this paper, we study human trust and its computational models in supervisory control of swarm robots with varied levels of autonomy (LOA) in a target foraging task. We implement three LOAs: manual, mixed-initiative (MI), and fully autonomous LOA. While the swarm in the MI LOA is controlled by a human operator and an autonomous search algorithm collaboratively, the swarms in the manual and autonomous LOAs are fully directed by the human and the search algorithm, respectively. From user studies, we find that humans tend to make their decisions based on physical characteristics of the swarm rather than its performance since the task performance of swarms is not clearly perceivable by humans. Based on the analysis, we formulate trust as a Markov decision process whose state space includes the factors affecting trust. We develop variations of the trust model for different LOAs. We employ an inverse reinforcement learning algorithm to learn behaviors of the operator from demonstrations where the learned behaviors are used to predict human trust. Compared to an existing model, our models reduce the prediction error by at most 39.6%, 36.5%, and 28.8% in the manual, MI, and auto-LOA, respectively.
Changjoo Nam, Phillip M. Walker, Huao Li, Michael Lewis 0001, Katia P. Sycara
IEEE Trans. Hum. Mach. Syst.3
2019 Perceptions of Domestic Robots' Normative Behavior Across Cultures
abstract
As domestic service robots become more common and widespread, they must be programmed to efficiently accomplish tasks while aligning their actions with relevant norms. The first step to equip domestic robots with normative reasoning competence is understanding the norms that people apply to the behavior of robots in specific social contexts. To that end, we conducted an online survey of Chinese and United States participants in which we asked them to select the preferred normative action a domestic service robot should take in a number of scenarios. The paper makes multiple contributions. Our extensive survey is the first to: (a) collect data on attitudes of people on normative behavior of domestic robots, (b) across cultures and (c) study relative priorities among norms for this domain. We present our findings and discuss their implications for building computational models for robot normative reasoning.
Huao Li, Stephanie Milani, Vigneshram Krishnamoorthy, Michael Lewis 0001, Katia P. Sycara
AIES1
2018 Transparency and Explanation in Deep Reinforcement Learning Neural Networks
abstract
Autonomous AI systems will be entering human society in the near future to provide services and work alongside humans. For those systems to be accepted and trusted, the users should be able to understand the reasoning process of the system, i.e. the system should be transparent. System transparency enables humans to form coherent explanations of the system's decisions and actions. Transparency is important not only for user trust, but also for software debugging and certification. In recent years, Deep Neural Networks have made great advances in multiple application areas. However, deep neural networks are opaque. In this paper, we report on work in transparency in Deep Reinforcement Learning Networks (DRLN). Such networks have been extremely successful in accurately learning action control in image input domains, such as Atari games. In this paper, we propose a novel and general method that (a) incorporates explicit object recognition processing into deep reinforcement learning models, (b) forms the basis for the development of "object saliency maps", to provide visualization of internal states of DRLNs, thus enabling the formation of explanations and (c) can be incorporated in any existing deep reinforcement learning framework. We present computational results and human experiments to evaluate our approach.
Rahul Iyer, Yuezhang Li, Huao Li, Michael Lewis 0001, Ramitha Sundar, Katia P. Sycara
AIES3
2018 Human Interaction Through an Optimal Sequencer to Control Robotic Swarms
abstract
The interaction between swarm robots and human operators is significantly different from the traditional humanrobot interaction due to unique characteristics of the system, such as high cognitive complexity and difficulties in state estimation. In this paper, we concentrated on the method of conveying input from the operator to the swarm. Previous research has shown that control through switching between behaviors offers the greatest flexibility but is particularly difficult for human operators. A recently developed method for finding optimal sequences for composing behaviors offered a potential tool for aiding human operators controlling swarms through behavior switching. This paper compared participants performing a navigation task with and without the availability of the optimal sequencing aid. Results showed that the task of preplanning a sequence of behaviors and durations appeared more difficult for participants than switching between executing behaviors to navigate. Users who used the aid frequently was found to create shorter paths than infrequent users and the control group. In the trails that the aid was used, participants tended to generate more complicated sequences and achieve the first attempt more rapidly, compared to the trails that the aid was not used.
Huao Li, Jaeho Bang, Sasanka Nagavalli, Changjoo Nam, Michael Lewis 0001, Katia P. Sycara
SMC1
2018 Trust of Humans in Supervisory Control of Swarm Robots with Varied Levels of Autonomy
abstract
In this paper, we study trust-related human factors in supervisory control of swarm robots with varied levels of autonomy (LOA) in a target foraging task. We compare three LOAs: manual, mixed-initiative (MI), and fully autonomous LOA. In the manual LOA, the human operator chooses headings for a flocking swarm, issuing new headings as needed. In the fully autonomous LOA, the swarm is redirected automatically by changing headings using a search algorithm. In the mixed-initiative LOA, if performance declines, control is switched from human to swarm or swarm to human. The result of this work extends the current knowledge on human factors in swarm supervisory control. Specifically, the finding that the relationship between trust and performance improved for passively monitoring operators (i.e., improved situation awareness in higher LOAs) is particularly novel in its contradiction of earlier work. We also discover that operators switch the degree of autonomy when their trust in the swarm system is low. Last, our analysis shows that operator's preference for a lower LOA is confirmed for a new domain of swarm control.
Changjoo Nam, Huao Li, Michael Lewis 0001, Katia P. Sycara
SMC2
1997 Modeling of Wavelet Coefficients in Medical Image Compression
abstract
The discrete wavelet transform provides a new framework of multiresolution space-frequency representation. Its preliminary applications in medical image compression are promising. An accurate modeling of the spatial and frequency characteristics of the wavelet coefficients is a key to designing efficient and accurate quantization for wavelet-based source coding. In this study, we investigate the modeling of a finite mixture distribution of the wavelet coefficients, within the context of information theory and statistical model identification. Using a finite generalized Gaussian mixture to model the overall distribution of the coefficients, an unsupervised learning procedure is developed to quantify the histogram through a tripled adaptive algorithm including detection of the number of local kernels, approximation of the shape of local kernels, and estimation of model parameter values. Our preliminary experimental results indicate that the unsupervised and adaptive histogram quantification can efficiently and accurately fit to the overall mixture distribution of the coefficients for any given frequency subband with unknown characteristics.
Yue Joseph Wang, Huao Li, Jianhua Xuan, Shih-Chung Ben Lo, Seong Ki Mun
ICIP (1)2