Yizhe Huang

dblp:178/7261 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0001-7117-6255ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 50% Multi-agent systems · 23% Planning, search and constraint satisfaction · 7%
Network and information security
1 paper
Privacy and data protection · 100%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
multi-agent reinforcement learning
2.642025
Social World Model-Augmented Mechanism Design Policy Learning · NeurIPS 2025
Learning to Balance Altruism and Self-interest Based on Empathy in Mixed-Motive Games · NeurIPS 2024
Efficient Adaptation in Mixed-Motive Environments via Hierarchical Opponent Modeling and Planning · ICML 2024
Machine learning › Reinforcement learning › model-based reinforcement learning
world model
1.722025
World Models Should Prioritize the Unification of Physical and Social Dynamics · NeurIPS 2025
Social World Model-Augmented Mechanism Design Policy Learning · NeurIPS 2025
Robotics › Robot manipulation › robot design
mechanism design
0.912025
Social World Model-Augmented Mechanism Design Policy Learning · NeurIPS 2025
Knowledge, reasoning and agents › Multi-agent systems
social dynamics modeling
0.912025
World Models Should Prioritize the Unification of Physical and Social Dynamics · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot adaptation
0.812024
Efficient Adaptation in Mixed-Motive Environments via Hierarchical Opponent Modeling and Planning · ICML 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
hierarchical planning
0.812024
Efficient Adaptation in Mixed-Motive Environments via Hierarchical Opponent Modeling and Planning · ICML 2024
Machine learning › Reinforcement learning › multi-agent reinforcement learning › markov games
mixed-motive games
0.812024
Learning to Balance Altruism and Self-interest Based on Empathy in Mixed-Motive Games · NeurIPS 2024
Knowledge, reasoning and agents › Multi-agent systems
multi-agent decision making
0.812024
AdaSociety: An Adaptive Environment with Social Structures for Multi-Agent Decision-Making · NeurIPS 2024
Knowledge, reasoning and agents › Multi-agent systems
multi-agent environments
0.812024
AdaSociety: An Adaptive Environment with Social Structures for Multi-Agent Decision-Making · NeurIPS 2024
Machine learning › Reinforcement learning › multi-agent reinforcement learning
opponent modeling
0.812024
Efficient Adaptation in Mixed-Motive Environments via Hierarchical Opponent Modeling and Planning · ICML 2024
Machine learning › Learning theory
generalization bounds
0.512021
Analyzing the Generalization Capability of SGLD Using Properties of Gaussian Channels · NeurIPS 2021
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo › langevin dynamics
stochastic gradient langevin dynamics
0.512021
Analyzing the Generalization Capability of SGLD Using Properties of Gaussian Channels · NeurIPS 2021
Privacy and data protection
differential privacy
0.512021
Analyzing the Generalization Capability of SGLD Using Properties of Gaussian Channels · NeurIPS 2021
Privacy and data protection › differential privacy › differentially private deep learning
DP-SGD
0.512021
Analyzing the Generalization Capability of SGLD Using Properties of Gaussian Channels · NeurIPS 2021
Machine learning › Reinforcement learning
model-based reinforcement learning
0.312025
Social World Model-Augmented Mechanism Design Policy Learning · NeurIPS 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
counterfactual reasoning
0.212024
Learning to Balance Altruism and Self-interest Based on Empathy in Mixed-Motive Games · NeurIPS 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
0.212024
Efficient Adaptation in Mixed-Motive Environments via Hierarchical Opponent Modeling and Planning · ICML 2024

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.6world model · 0.9hierarchical modeling · 0.9perspective-taking · 0.8opponent modeling · 0.8multi-agent reinforcement learning · 0.8monte carlo tree search · 0.8goal-conditioned policy · 0.8counterfactual baseline · 0.8LLM-based agents · 0.8strong data processing inequalities · 0.5gaussian channel · 0.5
YearPublicationVenuePosition
2026 Spike-local field potential-based multi-scale fusion decoding algorithm in intracortical brain-machine interfaces
Xiao Li 0036, Songyang An, Yizhe Huang, Wei Li 0086, Peng Zhang 0106
Eng. Appl. Artif. Intell.6
2026 Unsupervised Feature Selection-Driven Active Learning for Semi-Supervised Automatic ECG Analysis
abstract
Automatic analysis methods of electrocardiograms (ECGs) usually required large-scale annotated training data, but the annotation process is extremely time-consuming. While semi-supervised learning can leverage unlabeled data, its performance depends heavily on the quality of the initial labeled subset. Active learning has been used to identify the most informative samples for annotation, but conventional approaches face three critical limitations: (1) dependency on manual intervention for iterative query design, (2) prohibitive computational costs during sample selection, and (3) limited compatibility with semi-supervised learning frameworks. To address these limitations, we proposed an Unsupervised Active Feature-selective Semi-Supervised Learning (UAFSSL) framework for ECG analysis, including an unsupervised feature selection-based active learning module and a semi-supervised learning module. UAFSSL captures latent data distributions via unsupervised feature extraction, selects diverse and representative samples using pseudo-label clustering, and integrates seamlessly with semi-supervised learning to eliminate human intervention. We validated our algorithm on an ECG waveform segmentation task and an atrial fibrillation detection task. In the waveform segmentation task, our method improved the F1-score for P-wave delineation by 2.4% compared to random sampling, using only 5% of labeled samples. For the atrial fibrillation detection task, we evaluated our method on both the AFDB and a 24-hour dataset collected from 500 atrial fibrillation patients. Using only 200 labeled samples for model training, our method achieved AUC improvements of 2.5% and 2.2% over random sampling in five-fold cross-validation. This is the first study to integrate unsupervised active learning with semi-supervised learning for automatic ECG analysis, offering a robust, automated solution to reduce annotation costs while enhancing clinical applicability.
Xiao Li 0036, Songyang An, Yizhe Huang, Fan Lin, Peng Zhang 0106
IEEE J. Biomed. Health Informatics7
2025 Social World Model-Augmented Mechanism Design Policy Learning
abstract
Designing adaptive mechanisms to align individual and collective interests remains a central challenge in artificial social intelligence. Existing methods often struggle with modeling heterogeneous agents possessing persistent latent traits (e.g., skills, preferences) and dealing with complex multi-agent system dynamics. These challenges are compounded by the critical need for high sample efficiency due to costly real-world interactions. World Models, by learning to predict environmental dynamics, offer a promising pathway to enhance mechanism design in heterogeneous and complex systems. In this paper, we introduce a novel method named SWM-AP (Social World Model-Augmented Mechanism Design Policy Learning), which learns a social world model hierarchically modeling agents' behavior to enhance mechanism design. Specifically, the social world model infers agents' traits from their interaction trajectories and learns a trait-based model to predict agents' responses to the deployed mechanisms. The mechanism design policy collects extensive training trajectories by interacting with the social world model, while concurrently inferring agents' traits online during real-world interactions to further boost policy learning efficiency. Experiments in diverse settings (tax policy design, team coordination, and facility location) demonstrate that SWM-AP outperforms established model-based and model-free RL baselines in cumulative rewards and sample efficiency.
Yizhe Huang, Chengdong Ma, Zhixun Chen, Yali Du 0001, Song-Chun Zhu, Yaodong Yang 0001
NeurIPS2
2025 World Models Should Prioritize the Unification of Physical and Social Dynamics
abstract
World models, which explicitly learn environmental dynamics to lay the foundation for planning, reasoning, and decision-making, are rapidly advancing in predicting both physical dynamics and aspects of social behavior, yet predominantly in separate silos. This division results in a systemic failure to model the crucial interplay between physical environments and social constructs, rendering current models fundamentally incapable of adequately addressing the true complexity of real-world systems where physical and social realities are inextricably intertwined. This position paper argues that the systematic, bidirectional unification of physical and social predictive capabilities is the next crucial frontier for world model development. We contend that comprehensive world models must holistically integrate objective physical laws with the subjective, evolving, and context-dependent nature of social dynamics. Such unification is paramount for AI to robustly navigate complex real-world challenges and achieve more generalizable intelligence. This paper substantiates this imperative by analyzing core impediments to integration, proposing foundational guiding principles (ACE Principles), and outlining a conceptual framework alongside a research roadmap towards truly holistic world models.
Chengdong Ma, Yizhe Huang, Weidong Huang 0008, Siyuan Qi, Song-Chun Zhu, Yaodong Yang 0001
NeurIPS3
2024 Efficient Adaptation in Mixed-Motive Environments via Hierarchical Opponent Modeling and Planning
abstract
Despite the recent successes of multi-agent reinforcement learning (MARL) algorithms, efficiently adapting to co-players in mixed-motive environments remains a significant challenge. One feasible approach is to hierarchically model co-players’ behavior based on inferring their characteristics. However, these methods often encounter difficulties in efficient reasoning and utilization of inferred information. To address these issues, we propose Hierarchical Opponent modeling and Planning (HOP), a novel multi-agent decision-making algorithm that enables few-shot adaptation to unseen policies in mixed-motive environments. HOP is hierarchically composed of two modules: an opponent modeling module that infers others’ goals and learns corresponding goal-conditioned policies, and a planning module that employs Monte Carlo Tree Search (MCTS) to identify the best response. Our approach improves efficiency by updating beliefs about others’ goals both across and within episodes and by using information from the opponent modeling module to guide planning. Experimental results demonstrate that in mixed-motive environments, HOP exhibits superior few-shot adaptation capabilities when interacting with various unseen agents, and excels in self-play scenarios. Furthermore, the emergence of social intelligence during our experiments underscores the potential of our approach in complex multi-agent environments.
Yizhe Huang, Anji Liu, Fanqi Kong, Yaodong Yang 0001, Song-Chun Zhu
ICML1
2024 AdaSociety: An Adaptive Environment with Social Structures for Multi-Agent Decision-Making
abstract
Traditional interactive environments limit agents' intelligence growth with fixed tasks. Recently, single-agent environments address this by generating new tasks based on agent actions, enhancing task diversity. We consider the decision-making problem in multi-agent settings, where tasks are further influenced by social connections, affecting rewards and information access. However, existing multi-agent environments lack a combination of adaptive physical surroundings and social connections, hindering the learning of intelligent behaviors.To address this, we introduce AdaSociety, a customizable multi-agent environment featuring expanding state and action spaces, alongside explicit and alterable social structures. As agents progress, the environment adaptively generates new tasks with social structures for agents to undertake. In AdaSociety, we develop three mini-games showcasing distinct social structures and tasks. Initial results demonstrate that specific social structures can promote both individual and collective benefits, though current reinforcement learning and LLM-based algorithms show limited effectiveness in leveraging social structures to enhance performance. Overall, AdaSociety serves as a valuable research platform for exploring intelligence in diverse physical and social settings. The code is available at https://github.com/bigai-ai/AdaSociety.
Yizhe Huang, Fanqi Kong, Aoyang Qin, Min Tang 0006, Xiaoxi Wang, Song-Chun Zhu, Mingjie Bi, Siyuan Qi
NeurIPS1
2024 Learning to Balance Altruism and Self-interest Based on Empathy in Mixed-Motive Games
abstract
Real-world multi-agent scenarios often involve mixed motives, demanding altruistic agents capable of self-protection against potential exploitation. However, existing approaches often struggle to achieve both objectives. In this paper, based on that empathic responses are modulated by learned social relationships between agents, we propose LASE (**L**earning to balance **A**ltruism and **S**elf-interest based on **E**mpathy), a distributed multi-agent reinforcement learning algorithm that fosters altruistic cooperation through gifting while avoiding exploitation by other agents in mixed-motive games. LASE allocates a portion of its rewards to co-players as gifts, with this allocation adapting dynamically based on the social relationship --- a metric evaluating the friendliness of co-players estimated by counterfactual reasoning. In particular, social relationship measures each co-player by comparing the estimated $Q$-function of current joint action to a counterfactual baseline which marginalizes the co-player's action, with its action distribution inferred by a perspective-taking module. Comprehensive experiments are performed in spatially and temporally extended mixed-motive games, demonstrating LASE's ability to promote group collaboration without compromising fairness and its capacity to adapt policies to various types of interactive co-players.
Fanqi Kong, Yizhe Huang, Song-Chun Zhu, Siyuan Qi
NeurIPS2
2021 Analyzing the Generalization Capability of SGLD Using Properties of Gaussian Channels
abstract
Optimization is a key component for training machine learning models and has a strong impact on their generalization. In this paper, we consider a particular optimization method---the stochastic gradient Langevin dynamics (SGLD) algorithm---and investigate the generalization of models trained by SGLD. We derive a new generalization bound by connecting SGLD with Gaussian channels found in information and communication theory. Our bound can be computed from the training data and incorporates the variance of gradients for quantifying a particular kind of "sharpness" of the loss landscape. We also consider a closely related algorithm with SGLD, namely differentially private SGD (DP-SGD). We prove that the generalization capability of DP-SGD can be amplified by iteration. Specifically, our bound can be sharpened by including a time-decaying factor if the DP-SGD algorithm outputs the last iterate while keeping other iterates hidden. This decay factor enables the contribution of early iterations to our bound to reduce with time and is established by strong data processing inequalities---a fundamental tool in information theory. We demonstrate our bound through numerical experiments, showing that it can predict the behavior of the true generalization gap.
Hao Wang 0063, Yizhe Huang, Rui Gao 0001, Flávio P. Calmon
NeurIPS2