Shuyue Hu

dblp:156/1020 · DBLP profile ↗
← Back
36ranked-venue papers
4as first author
32since 2021 · last 2026
0000-0002-1908-1344ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 3 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 11 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Do We Truly Need So Many Samples? Multi-LLM Repeated Sampling Efficiently Scales Test-Time Compute
abstract
This paper presents a simple, effective, and cost-efficient strategy, named ModelSwitch, to improve LLM performance by scaling test-time compute. ModelSwitch builds upon the repeated-sampling-then-voting framework, with a novel twist: incorporating multiple models, even weaker ones, to leverage their complementary strengths that potentially arise from diverse training data and paradigms. By using sample consistency as a signal, our strategy dynamically switches between models. Theoretical analysis highlights the efficiency and performance advantages of our strategy. Extensive experiments on seven datasets demonstrate that our strategy not only outperforms self-consistency and state-of-the-art multi-agent debate approaches, but also significantly reduces inference costs. Additionally, our strategy requires only a few comparable LLMs to achieve optimal performance and can be extended with verification methods, demonstrating the potential of leveraging multiple LLMs in the generation-verification paradigm.
Jianhao Chen 0001, Zishuo Xun, Bocheng Zhou, Hangfan Zhang, Qiaosheng Zhang 0002, Wei Hu 0007, Yuzhong Qu, Shuyue Hu
AAAI10
2026 Adaptive Theory of Mind for LLM-based Multi-Agent Coordination
abstract
Theory of Mind (ToM) refers to the ability to reason about others’ mental states, and higher-order ToM involves considering that others also possess their own ToM. Equipping large language model (LLM)-driven agents with ToM has long been considered to improve their coordination in multiagent collaborative tasks. However, we find that misaligned ToM orders—mismatches in the depth of ToM reasoning between agents—can lead to insufficient or excessive reasoning about others, thereby impairing their coordination. To address this issue, we design an adaptive ToM (A-ToM) agent, which can align in ToM orders with its partner. Based on prior interactions, the agent estimates the partner’s likely ToM order and leverages this estimation to predict the partner’s action, thereby facilitating behavioral coordination. We conduct empirical evaluations on four multi-agent coordination tasks: a repeated matrix game, two grid navigation tasks and an Overcooked task. The results validate our findings on ToM alignment and demonstrate the effectiveness of our AToM agent. Furthermore, we discuss the generalizability of our A-ToM to non-LLM-based agents, as well as what would diminish the importance of ToM alignment.
Chunjiang Mu, Ya Zeng, Qiaosheng Zhang 0002, Kun Shao, Chen Chu, Danyang Jia, Zhen Wang 0004, Shuyue Hu
AAAI9
2026 ICL-Router: In-Context Learned Model Representations for LLM Routing
abstract
Large language models (LLMs) often exhibit complementary strengths. Model routing harnesses these strengths by dynamically directing each query to the most suitable model, given a candidate model pool. However, routing performance relies on accurate model representations, and adding new models typically requires retraining, limiting scalability. To address these challenges, we propose a novel routing method using in-context vectors to represent model capabilities. The method proceeds in two stages. First, queries are embedded and projected into vectors, with a projector and LLM-based router trained to reconstruct the original queries, aligning vector representations with the router’s semantic space. Second, each candidate model is profiled on a query set, and the router learns---based on in-context vectors of query and model performance---to predict whether each model can correctly answer new queries. Extensive experiments demonstrate that our method achieves state-of-the-art routing performance in both in-distribution and out-of-distribution tasks. Moreover, our method allows for seamless integration of new models without retraining the router.
Hao Li 0069, Linyao Chen, Jianhao Chen 0001, Ping Jian, Qiaosheng Zhang 0002, Shuyue Hu
AAAI8
2026 The Avengers: A Routing Recipe for Collective Intelligence in Language Models
abstract
Proprietary models are increasingly dominating the race for ever-larger language models. Can open-source, smaller models remain competitive across a broad range of tasks? In this paper, we present the Avengers---a lightweight framework that leverages the collective intelligence of these smaller models. The Avengers builds upon four lightweight operations: (i) embedding: encode queries using a text embedding model; (ii) clustering: group queries based on their semantic similarity; (iii) scoring: scores each model's performance within each cluster; and (iv) voting: improve outputs via repeated sampling and voting. At inference time, each query is embedded and assigned to its nearest cluster. The top-performing model(s) within that cluster are selected to generate the response with repeated sampling. Remarkably, with 10 open-source models (~7B parameters each), the Avengers surpasses GPT-4o, 4.1, and 4.5 in average performance across 15 diverse datasets spanning mathematics, coding, logical reasoning, general knowledge, and affective tasks. In particular, it surpasses GPT-4.1 on mathematics tasks by 18.21% and on code tasks by 7.46%. Furthermore, the Avengers delivers superior out-of-distribution generalization, and remains robust across various embedding models, clustering algorithms, ensemble strategies, data efficiency, and values of its sole parameter---the number of clusters.
Hao Li 0069, Linyao Chen, Qiaosheng Zhang 0002, Peng Ye 0006, Shi Feng 0001, Xinrun Wang, Xu Jia 0012, Lei Bai 0001, Shuyue Hu
AAAI11
2026 A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven Enhancement
abstract
Shengji Tang, Jianjian Cao, Weihao Lin, Jiale Hong, Bo Zhang, Shuyue Hu, Lei Bai, Tao Chen, Wanli Ouyang, Peng Ye. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shengji Tang, Jianjian Cao, Weihao Lin 0002, Jiale Hong, Bo Zhang 0069, Shuyue Hu, Lei Bai 0001, Tao Chen 0003, Wanli Ouyang, Peng Ye 0006
ACL (1)6
2026 MTRouter: Cost-Aware Multi-Turn LLM Routing with History-Model Joint Embeddings
abstract
Yiqun Zhang, Hao Li, Zihan Wang, Shi Feng, Xiaocui Yang, Daling Wang, Bo Zhang, Lei Bai, Shuyue Hu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hao Li 0069, Shi Feng 0001, Xiaocui Yang, Daling Wang, Bo Zhang 0069, Lei Bai 0001, Shuyue Hu
ACL (1)9
2026 Nature-Inspired Population-Based Evolution of Large Language Models
abstract
Yiqun Zhang, Peng Ye, Xiaocui Yang, Shi Feng, Shufei Zhang, Lei Bai, Wanli Ouyang, Shuyue Hu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Peng Ye 0006, Xiaocui Yang, Shi Feng 0001, Shufei Zhang, Lei Bai 0001, Wanli Ouyang, Shuyue Hu
ACL (1)8
2026 A successful strategy for iterated Prisoner's dilemma with any number of channels
Zhaoheng Cao, Zhen Wang 0004, Shuyue Hu, Chen Chu
Artif. Intell.4
2026 Dynamics of Q-Learning in Networked Stochastic Games
abstract
Stochastic games form the foundational mathematical framework for describing multiagent interactions and underpin the theoretical foundations of multiagent reinforcement learning (MARL) and optimal decision making. However, previous research has typically focused on either two-agent settings or large-scale well-mixed agent populations, where the considered interaction scenarios were far from realistic. In this article, we consider structured populations where agents can interact with immediate neighbors. By using the pair-approximation method, we develop a new dynamical model to describe the $Q$ -learning dynamics in stochastic games on regular graphs. Through comparisons with agent-based simulation results, we validate the accuracy of our dynamical model across various stochastic games, population structures, and algorithm parameters. Our research thus provides both qualitative and quantitative insights into the effects of state transition rules and graph topologies in population dynamics. In particular, we show that, under certain conditions, state transitions can significantly promote the evolution of cooperation in social dilemmas. We also explored the effects of agent degree on cooperation, and unlike previous findings, we show that this can have either positive or negative implications for cooperation depending on the transition rules.
Guangchen Jiang, Shuyue Hu, Matjaz Perc, Chen Chu, Jinzhuo Liu
IEEE Trans. Neural Networks Learn. Syst.3
2025 Beyond GPT-5: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing
abstract
Balancing performance and efficiency is a central challenge in large language model (LLM) advancement. GPT-5 addresses this with test-time routing, dynamically assigning queries to either an efficient or a high-capacity model during inference. In this work, we present Avengers-Pro, a test-time routing framework that ensembles LLMs of varying capacities and efficiencies, providing a unified solution for all performance-efficiency tradeoffs. The Avengers-Pro embeds and clusters incoming queries, then routes each to the most suitable model based on a performance-efficiency score. Across 6 challenging benchmarks and 8 leading models—including GPT-5-medium, Gemini-2.5-pro, and Claude-opus-4.1—Avengers-Pro achieves state-of-the-art results: by varying a performance-efficiency trade-off parameter, it can surpass the strongest single model (GPT-5-medium) by +7% in average accuracy. Moreover, it can match the average accuracy of the strongest single model at 27% lower cost, and reach ∼ 90% of that performance at 63% lower cost. Last but not least, it achieves a Pareto frontier, consistently yielding the highest accuracy for any given cost, and the lowest cost for any given accuracy, among all single models. Code is available at https://github.com/ZhangYiqun018/AvengersPro.
Hao Li 0069, Jianhao Chen 0001, Hangfan Zhang, Peng Ye 0006, Lei Bai 0001, Shuyue Hu
DAI7
2025 Reinforcement Learning for Large Language Models via Group Preference Reward Shaping
abstract
Huaisheng Zhu, Siyuan Xu, Hangfan Zhang, Teng Xiao, Zhimeng Guo, Shijie Zhou, Shuyue Hu, Vasant G. Honavar. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Huaisheng Zhu, Hangfan Zhang, Teng Xiao, Zhimeng Guo, Shijie Zhou 0008, Shuyue Hu, Vasant G. Honavar
EMNLP7
2025 Graph Attention is Not Always Beneficial: A Theoretical Analysis of Graph Attention Mechanisms via Contextual Stochastic Block Models
abstract
Despite the growing popularity of graph attention mechanisms, their theoretical understanding remains limited. This paper aims to explore the conditions under which these mechanisms are effective in node classification tasks through the lens of Contextual Stochastic Block Models (CSBMs). Our theoretical analysis reveals that incorporating graph attention mechanisms is *not universally beneficial*. Specifically, by appropriately defining *structure noise* and *feature noise* in graphs, we show that graph attention mechanisms can enhance classification performance when structure noise exceeds feature noise. Conversely, when feature noise predominates, simpler graph convolution operations are more effective. Furthermore, we examine the over-smoothing phenomenon and show that, in the high signal-to-noise ratio (SNR) regime, graph convolutional networks suffer from over-smoothing, whereas graph attention mechanisms can effectively resolve this issue. Building on these insights, we propose a novel multi-layer Graph Attention Network (GAT) architecture that significantly outperforms single-layer GATs in achieving *perfect node classification* in CSBMs, relaxing the SNR requirement from $\omega(\sqrt{\log n})$ to $\omega(\sqrt{\log n} / \sqrt[3]{n})$. To our knowledge, this is the first study to delineate the conditions for perfect node classification using multi-layer GATs. Our theoretical contributions are corroborated by extensive experiments on both synthetic and real-world datasets, highlighting the practical implications of our findings.
Zhongtian Ma, Qiaosheng Zhang 0002, Bocheng Zhou, Yexin Zhang, Shuyue Hu, Zhen Wang 0004
ICML5
2025 ReMA: Learning to Meta-Think for LLMs with Multi-agent Reinforcement Learning
abstract
Recent research on Reasoning of Large Language Models (LLMs) has sought to further enhance their performance by integrating meta-thinking—enabling models to monitor, evaluate, and control their reasoning processes for more adaptive and effective problem-solving. However, current single-agent work lacks a specialized design for acquiring meta-thinking, resulting in low efficacy. To address this challenge, we introduce Reinforced Meta-thinking Agents (ReMA), a novel framework that leverages Multi-Agent Reinforcement Learning (MARL) to elicit meta-thinking behaviors, encouraging LLMs to think about thinking. ReMA decouples the reasoning process into two hierarchical agents: a high-level meta-thinking agent responsible for generating strategic oversight and plans, and a low-level reasoning agent for detailed executions. Through iterative reinforcement learning with aligned objectives, these agents explore and learn collaboration, leading to improved generalization and robustness. Empirical results from single-turn experiments demonstrate that ReMA outperforms single-agent RL baselines on complex reasoning tasks, including competitive-level mathematical benchmarks and LLM-as-a-Judge benchmarks. Additionally, we further extend ReMA to multi-turn interaction settings, leveraging turn-level ratio and parameter sharing to improve efficiency. Comprehensive ablation studies further illustrate the evolving dynamics of each distinct agent, providing valuable insights into how the meta-thinking reasoning process enhances the reasoning capabilities of LLMs.
Ziyu Wan, Xiaoyu Wen 0001, Yan Song 0003, Hanjing Wang, Linyi Yang, Mark Schmidt 0001, Jun Wang 0012, Weinan Zhang 0001, Shuyue Hu, Ying Wen 0001
NeurIPS10
2025 Scaling Physical Reasoning with the PHYSICS Dataset
abstract
Large Language Models (LLMs) have achieved remarkable progress on advanced reasoning tasks such as mathematics and coding competitions. Meanwhile, physics, despite being both reasoning-intensive and essential to real-world understanding, received limited academic and industrial attention. This paper introduces PHYSICS, a dataset containing 16,568 high-quality physics problems spanning subjects and difficulty levels, to facilitate this issue. Specifically, PHYSICS is curated with exercises from over 100 textbooks through a carefully designed pipeline for quality control. It covers five major physics domains: Mechanics, Electromagnetism, Thermodynamics, Optics, and Modern Physics. It also spans a wide range of difficulty levels, from high school to graduate-level physics courses. To utilize the data for improving and evaluating the model's physical reasoning capabilities, we split the dataset into training and test sets, and provide reasoning paths generated by powerful reasoning models for the training data to facilitate model training. In addition, for the evaluation part, we find that existing evaluation frameworks exhibit biases in aspects such as units, simplification, and precision in physics domain. To balance efficiency and accuracy, we introduce a Rule+Model evaluation framework tailored to physics problems. Our evaluations on current state-of-the-art open-source and proprietary models highlight the limitations of current models in handling physics-related tasks. We hope that our dataset and evaluation methodology will jointly advance the development of LLMs in the field of physics. The code and data can be found at: https://github.com/Zhengsh123/PHYSICS.
Shenghe Zheng, Qianjia Cheng, Junchi Yao, Mengsong Wu, Ning Ding 0002, Yu Cheng 0001, Shuyue Hu, Lei Bai 0001, Dongzhan Zhou, Ganqu Cui, Peng Ye 0006
NeurIPS8
2025 Provably efficient information-directed sampling algorithms for multi-agent reinforcement learning
Qiaosheng Zhang 0002, Chenjia Bai, Shuyue Hu, Zhen Wang 0004, Xuelong Li 0001
Artif. Intell.3
2025 A formal model for multiagent Q-learning on graphs
Jinzhuo Liu, Guangchen Jiang, Chen Chu, Zhen Wang 0004, Shuyue Hu
Sci. China Inf. Sci.6
2025 Payoff Control in Multichannel Games: Influencing Opponent Learning Evolution
abstract
In this article, we introduce a new theory for payoff control in multichannel learning environments, where agents interact with each other over multiple channels and each channel is a repeated normal form game. We propose two payoff control strategies-partial control and full control-that allow a single agent to set an upper bound to the opponent's expected payoffs summed across all channels, even if the opponent is a reinforcement learning agent. We prove that a partial (or full) control strategy can be obtained by solving a system of inequalities, and characterize the conditions under which such a partial (or full) control strategy exists. We show that by utilizing these control strategies, the agent can influence the opponent's learning evolution and direct it toward a desired viable equilibrium. Our experiments confirm the effectiveness of our theory for payoff control in a wide range of multichannel learning environments.
Chen Chu, Guoxi Fan, Jinzhuo Liu, Zhen Wang 0004, Shuyue Hu
IEEE Trans. Cybern.7
2025 Regret Minimization in Population Network Games: Vanishing Heterogeneity and Convergence to Equilibria
abstract
Understanding and predicting the behavior of large-scale multiagents in games remains a fundamental challenge in multiagent systems. This article examines the role of heterogeneity in equilibrium formation by analyzing how smooth regret matching drives a large number of heterogeneous agents with diverse initial policies toward unified behavior. By modeling the system state as a probability distribution of regrets and analyzing its evolution through the continuity equation, we uncover a key phenomenon in diverse multiagent settings: the variance of the regret distribution diminishes over time, leading to the disappearance of heterogeneity and the emergence of consensus among agents. This universal result enables us to prove convergence to quantal response equilibria in both competitive and cooperative multiagent settings. This work advances the theoretical understanding of multiagent learning and offers a novel perspective on equilibrium selection in diverse game-theoretic scenarios.
Shuyue Hu, Chunjiang Mu, Shiqi Fan, Chen Chu, Jinzhuo Liu, Zhen Wang 0004
IEEE Trans. Neural Networks Learn. Syst.2
2024 Configurable Mirror Descent: Towards a Unification of Decision Making
abstract
Decision-making problems, categorized as single-agent, e.g., Atari, cooperative multi-agent, e.g., Hanabi, competitive multi-agent, e.g., Hold'em poker, and mixed cooperative and competitive, e.g., football, are ubiquitous in the real world. Although various methods have been proposed to address the specific decision-making categories, these methods typically evolve independently and cannot generalize to other categories. Therefore, a fundamental question for decision-making is: *Can we develop **a single algorithm** to tackle **ALL** categories of decision-making problems?* There are several main challenges to address this question: i) different decision-making categories involve different numbers of agents and different relationships between agents, ii) different categories have different solution concepts and evaluation measures, and iii) there lacks a comprehensive benchmark covering all the categories. This work presents a preliminary attempt to address the question with three main contributions. i) We propose the generalized mirror descent (GMD), a generalization of MD variants, which considers multiple historical policies and works with a broader class of Bregman divergences. ii) We propose the configurable mirror descent (CMD) where a meta-controller is introduced to dynamically adjust the hyper-parameters in GMD conditional on the evaluation measures. iii) We construct the GameBench with 15 academic-friendly games across different decision-making categories. Extensive experiments demonstrate that CMD achieves empirically competitive or better outcomes compared to baselines while providing the capability of exploring diverse dimensions of decision making.
Pengdeng Li, Shuxin Li 0001, Xinrun Wang, Shuyue Hu, Xiao Huang 0001, Hau Chan, Bo An 0001
ICML5
2024 A Successful Strategy for Multichannel Iterated Prisoner's Dilemma
Zhen Wang 0004, Zhaoheng Cao, Peican Zhu, Shuyue Hu, Chen Chu
IJCAI5
2024 Emergence of Social Norms in Generative Agent Societies: Principles and Architecture
Siyue Ren, Zhiyao Cui, Zhen Wang 0004, Shuyue Hu
IJCAI5
2024 Multi-agent, human-agent and beyond: A survey on cooperation in social dilemmas
Chunjiang Mu, Chen Shen 0006, Shuyue Hu, Zhen Wang 0004
Neurocomputing6
2023 A Pair-Approximation Method for Modelling the Dynamics of Multi-Agent Stochastic Games
abstract
Developing a dynamical model for learning in games has attracted much recent interest. In stochastic games, agents need to make decisions in multiple states, and transitions between states, in turn, influence the dynamics of strategies. While previous works typically focus either on 2-agent stochastic games or on normal form games under an infinite-agent setting, we aim at formally modelling the learning dynamics in stochastic games under the infinite-agent setting. With a novel use of pair-approximation method, we develop a formal model for myopic Q-learning in stochastic games with symmetric state transition. We verify the descriptive power of our model (a partial differential equation) across various games through comparisons with agent-based simulation results. Based on our proposed model, we can gain qualitative and quantitative insights into the influence of transition probabilities on the dynamics of strategies. In particular, we illustrate that a careful design of transition probabilities can help players overcome the social dilemmas and promote cooperation, even if agents are myopic learners.
Chen Chu, Shuyue Hu, Chunjiang Mu, Zhen Wang 0004
AAAI3
2023 Emergence of Punishment in Social Dilemma with Environmental Feedback
abstract
Altruistic punishment (or punishment) has been extensively shown as an important mechanism for promoting cooperation in human societies. In AI, the emergence of punishment has received much recent interest. In this paper, we contribute with a novel evolutionary game theoretic model to study the impacts of environmental feedback. Whereas a population of agents plays public goods games, there exists a third-party population whose payoffs depend not only on whether to punish or not, but also on the state of the environment (e.g., how cooperative the agents in a social dilemma are). Focusing on one-shot public goods games, we show that environmental feedback, by itself, can lead to the emergence of punishment. We analyze the co-evolution of punishment and cooperation, and derive conditions for their co-presence, co-dominance and co-extinction. Moreover, we show that the system can exhibit bistability as well as cyclic dynamics. Our findings provide a new explanation for the emergence of punishment. On the other hand, our results also alert the need for careful design of implementing punishment in multi-agent systems, as the resulting evolutionary dynamics can be somewhat complex.
Zhen Wang 0004, Zhao Song 0008, Chen Shen 0006, Shuyue Hu
AAAI4
2023 The Best of Both Worlds in Network Population Games: Reaching Consensus and Convergence to Equilibrium
abstract
Reaching consensus and convergence to equilibrium are two major challenges of multi-agent systems. Although each has attracted significant attention, relatively few studies address both challenges at the same time. This paper examines the connection between the notions of consensus and equilibrium in a multi-agent system where multiple interacting sub-populations coexist. We argue that consensus can be seen as an intricate component of intra-population stability, whereas equilibrium can be seen as encoding inter-population stability. We show that smooth fictitious play, a well-known learning model in game theory, can achieve both consensus and convergence to equilibrium in diverse multi-agent settings. Moreover, we show that the consensus formation process plays a crucial role in the seminal thorny problem of equilibrium selection in multi-agent learning.
Shuyue Hu, Harold Soh, Georgios Piliouras
NeurIPS1
2023 Learning by reusing previous advice: a memory-based teacher-student framework
Changxi Zhu, Yi Cai 0001, Shuyue Hu, Ho-fung Leung, Dickson K. W. Chiu
Auton. Agents Multi Agent Syst.3
2022 A Formal Model for Multiagent Q-Learning Dynamics on Regular Graphs
abstract
Modeling the dynamics of multi-agent learning has long been an important research topic. The focus of previous research has been either on 2-agent settings or well-mixed infinitely large agent populations. In this paper, we consider the scenario where n Q-learning agents locate on regular graphs, such that agents can only interact with their neighbors. We examine the local interactions between individuals and their neighbors, and derive a formal model to capture the Q-value dynamics of the entire population. Through comparisons with agent-based simulations on different types of regular graphs, we show that our model describes the agent learning dynamics in an exact manner.
Chen Chu, Jinzhuo Liu, Shuyue Hu, Xuelong Li 0001, Zhen Wang 0004
IJCAI4
2022 Modelling the Dynamics of Multi-Agent Q-learning: The Stochastic Effects of Local Interaction and Incomplete Information
abstract
The theoretical underpinnings of multiagent reinforcement learning has recently attracted much attention. In this work, we focus on the generalized social learning (GSL) protocol --- an agent interaction protocol that is widely adopted in the literature, and aim to develop an accurate theoretical model for the Q-learning dynamics under this protocol. Noting that previous models fail to characterize the effects of local interactions and incomplete information that arise from GSL, we model the Q-values dynamics of each individual agent as a system of stochastic differential equations (SDE). Based on the SDE, we express the time evolution of the probability density function of Q-values in the population with a Fokker-Planck equation. We validate the correctness of our model through extensive comparisons with agent-based simulation results across different types of symmetric games. In addition, we show that as the interactions between agents are more limited and information is less complete, the population can converge to a outcome that is qualitatively different than that with global interactions and complete information.
Chin-Wing Leung, Shuyue Hu, Ho-fung Leung
IJCAI2
2022 Modelling the Dynamics of Regret Minimization in Large Agent Populations: a Master Equation Approach
abstract
Understanding the learning dynamics in multiagent systems is an important and challenging task. Past research on multi-agent learning mostly focuses on two-agent settings. In this paper, we consider the scenario in which a population of infinitely many agents apply regret minimization in repeated symmetric games. We propose a new formal model based on the master equation approach in statistical physics to describe the evolutionary dynamics in the agent population. Our model takes the form of a partial differential equation, which describes how the probability distribution of regret evolves over time. Through experiments, we show that our theoretical results are consistent with the agent-based simulation results.
Zhen Wang 0004, Chunjiang Mu, Shuyue Hu, Chen Chu, Xuelong Li 0001
IJCAI3
2021 Formal Modeling of Reinforcement Learning with Many Agents through Repeated Local Interactions
abstract
Modelling the dynamics of multi-agent reinforcement learning has long been an important research topic. Most of the previous works focus on agents learning under global interactions. In this paper, we investigate learning in a population of agents with local interactions, such that agents learn their policies concurrently by playing with some other agents locally, without the knowledge of the whole population. We derive the stochastic differential equations (SDEs) to describe the Q-values dynamics of each individual agent under the stochastic environment. Applying the Fokker-Planck equation, the time evolution of the probability distribution (PDF) of the population Q-values is worked out. We validate our model through comparisons with agent-based simulations on typical symmetric games with various settings, and the results verify that the model can precisely capture the behaviour of the multi-agent system.
Chin-Wing Leung, Shuyue Hu, Ho-fung Leung
ICTAI2
2021 Gist Trace-based Learning: Efficient Convention Emergence from Multilateral Interactions
abstract
The concept of conventions has attracted much attention in the multi-agent system research. In this article, we study the emergence of conventions from repeated n -player coordination games. Distributed agents learn their policies independently and are capable of observing their neighbours in a network topology. We distinguish two types of information representation about the observations: gist trace and verbatim trace. We conjecture that learning based on the gist trace, which overlooks the details and focuses only on the general choice of action of a neighbourhood, should achieve efficient convention emergence. To this end, a novel learning method that makes use of the gist trace is proposed. The experimental results confirm that the proposed method establishes conventions much faster than the state-of-the-art learning methods across diverse settings of multi-agent systems. In particular, the use of gist trace derived at a low level of abstraction further improves the efficiency of convention emergence.
Shuyue Hu, Chin-Wing Leung, Ho-fung Leung, Jiamou Liu
ACM Trans. Auton. Adapt. Syst.1
2021 A Q-values Sharing Framework for Multi-agent Reinforcement Learning under Budget Constraint
abstract
In teacher-student framework, a more experienced agent (teacher) helps accelerate the learning of another agent (student) by suggesting actions to take in certain states. In cooperative multiagent reinforcement learning (MARL), where agents need to cooperate with one another, a student may fail to cooperate well with others even by following the teachers' suggested actions, as the polices of all agents are ever changing before convergence. When the number of times that agents communicate with one another is limited (i.e., there is budget constraint), the advising strategy that uses actions as advices may not be good enough. We propose a partaker-sharer advising framework (PSAF) for cooperative MARL agents learning with budget constraint. In PSAF, each Q-learner can decide when to ask for Q-values and share its Q-values. We perform experiments in three typical multiagent learning problems. Evaluation results show that our approach PSAF outperforms existing advising methods under both unlimited and limited budget, and we give an analysis of the impact of advising actions and sharing Q-values on agents' learning.
Changxi Zhu, Ho-fung Leung, Shuyue Hu, Yi Cai 0001
ACM Trans. Auton. Adapt. Syst.3
2020 Self-Play or Group Practice: Learning to Play Alternating Markov Game in Multi-Agent System
abstract
The research in reinforcement learning has achieved great success in strategic game playing. These successes are thanks to the incorporation of deep reinforcement learning (DRL) and Monte Carlo Tree Search (MCTS) to the agent trained under the self-play (SP) environment. By self-play, agents are provided with an incrementally more difficult curriculum which in turn facilitates learning. However, recent research suggests that agents trained via self-play may easily lead to getting stuck in local equilibria. In this paper, we consider a population of agents each independently learns to play an alternating Markov game (AMG). We propose a new training framework-group practice- for a population of decentralized RL agents. By group practice (GP), agents are assigned into multiple learning groups during training, for every episode of games, an agent is randomly paired up and practices with another agent in the learning group. The convergence result to the optimal value function and the Nash equilibrium are proved under the GP framework. Experimental study is conducted by applying GP to Q-learning algorithm and the deep Q-learning with Monte-Carlo tree search on the game of Connect Four and the game of Hex. We verify that GP is the more efficient training scheme than SP given the same amount of training. We also show that the learning effectiveness can even be improved when applying local grouping to agents.
Chin-Wing Leung, Shuyue Hu, Ho-fung Leung
ICPR2
2019 Modelling the Dynamics of Multiagent Q-Learning in Repeated Symmetric Games: a Mean Field Theoretic Approach
abstract
Modelling the dynamics of multi-agent learning has long been an important research topic, but all of the previous works focus on 2-agent settings and mostly use evolutionary game theoretic approaches. In this paper, we study an n-agent setting with n tends to infinity, such that agents learn their policies concurrently over repeated symmetric bimatrix games with some other agents. Using mean field theory, we approximate the effects of other agents on a single agent by an averaged effect. A Fokker-Planck equation that describes the evolution of the probability distribution of Q-values in the agent population is derived. To the best of our knowledge, this is the first time to show the Q-learning dynamics under an n-agent setting can be described by a system of only three equations. We validate our model through comparisons with agent-based simulations on typical symmetric bimatrix games and different initial settings of Q-values.
Shuyue Hu, Chin-Wing Leung, Ho-fung Leung
NeurIPS1
2019 Modeling Convention Emergence by Observation with Memorization
Chin-Wing Leung, Shuyue Hu, Ho-fung Leung
PRICAI (1)2
2017 Achieving Coordination in Multi-Agent Systems by Stable Local Conventions under Community Networks
abstract
Recently, the study of social conventions has attracted much attention in the literature. We notice that a type of interesting phenomena, local convention phenomena, may also exist in certain multi-agent systems. When agents are partitioned into compact communities, different local conventions emerge in different communities. In this paper, we provide a definition for local conventions, and propose two metrics measuring their strength and diversity. In our experimental study, we show that agents can achieve coordination via establishing diverse stable local conventions, which indicates a practical way to solve coordination problems other than the traditional global convention emergence. Moreover, we find that with smaller community sizes, denser connections and fewer available actions, diverse local conventions emerge in shorter time.
Shuyue Hu, Ho-fung Leung
IJCAI1