EDBT 2026 Demo / reviewers in the wild / expert
Hangyu Mao
dblp:191/6097
· DBLP profile ↗
33ranked-venue papers
5as first author
28since 2021 · last 2026
0000-0002-4499-7581ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 4 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tool-Star: Empowering Multi-Tool Collaborative Web Agent via Reinforcement Learning
Guanting Dong 0001, Yifei Chen 0001, Xiaoxi Li 0005, Jiajie Jin, Hongjin Qian, Yutao Zhu 0001, Hangyu Mao, Guorui Zhou, Zhicheng Dou, Ji-Rong Wen |
SIGIR | 7 |
| 2026 | Toward Generalized Web Agent Training: A Deep Dive into Entropy-Balanced Reinforcement Learning
Guanting Dong 0001, Licheng Bao, Zhongyuan Wang 0006, Kangzhi Zhao, Xiaoxi Li 0005, Jiajie Jin, Hangyu Mao, Kun Gai, Guorui Zhou, Yutao Zhu 0001, Ji-Rong Wen, Zhicheng Dou |
WWW | 8 |
| 2025 | SkillTree: Explainable Skill-Based Deep Reinforcement Learning for Long-Horizon Control TasksabstractDeep reinforcement learning (DRL) has achieved remarkable success in various domains, yet its reliance on neural networks results in a lack of transparency, which limits its practical applications in safety-critical and human-agent interaction domains. Decision trees, known for their notable explainability, have emerged as a promising alternative to neural networks. However, decision trees often struggle in long-horizon continuous control tasks with high-dimensional observation space due to their limited expressiveness. To address this challenge, we propose SkillTree, a novel hierarchical framework that reduces the complex continuous action space of challenging control tasks into discrete skill space. By integrating the differentiable decision tree within the high-level policy, SkillTree generates discrete skill embeddings that guide low-level policy execution. Furthermore, through distillation, we obtain a simplified decision tree model that improves performance while further reducing complexity. Experiment results validate SkillTree’s effectiveness across various robotic manipulation tasks, providing clear skill-level insights into the decision-making process. The proposed approach not only achieves performance comparable to neural network based methods in complex long-horizon control tasks but also significantly enhances the transparency and explainability of the decision-making process. Yongyan Wen, Siyuan Li 0003, Rongchang Zuo, Hangyu Mao, Peng Liu 0008 |
AAAI | 5 |
| 2025 | DualRAG: A Dual-Process Approach to Integrate Reasoning and Retrieval for Multi-Hop Question AnsweringabstractRong Cheng, Jinyi Liu, Yan Zheng, Fei Ni, Jiazhen Du, Hangyu Mao, Fuzheng Zhang, Bo Wang, Jianye Hao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Rong Cheng, Jinyi Liu 0002, Yan Zheng 0002, Fei Ni 0001, Jiazhen Du, Hangyu Mao, Bo Wang 0027, Jianye Hao |
ACL (1) | 6 |
| 2025 | PET-SQL: A Prompt-Enhanced Two-Round Refinement of Text-to-SQL with Cross-Consistency
Zhishuai Li, Xiang Wang 0012, Sun Yang, Guoqing Du, Xiaoru Hu, Bin Zhang 0052, Yuxiao Ye, Ziyue Li 0002, Hangyu Mao, Rui Zhao 0001 |
DASFAA (2) | 10 |
| 2025 | Reidentify: Context-Aware Identity Generation for Contextual Multi-Agent Reinforcement LearningabstractGeneralizing multi-agent reinforcement learning (MARL) to accommodate variations in problem configurations remains a critical challenge in real-world applications, where even subtle differences in task setups can cause pre-trained policies to fail. To address this, we propose Context-Aware Identity Generation (CAID), a novel framework to enhance MARL performance under the Contextual MARL (CMARL) setting. CAID dynamically generates unique agent identities through the agent identity decoder built on a causal Transformer architecture. These identities provide contextualized representations that align corresponding agents across similar problem variants, facilitating policy reuse and improving sample efficiency. Furthermore, the action regulator in CAID incorporates these agent identities into the action-value space, enabling seamless adaptation to varying contexts. Extensive experiments on CMARL benchmarks demonstrate that CAID significantly outperforms existing approaches by enhancing both sample efficiency and generalization across diverse context variants. Zhiwei Xu 0005, Xin Xin 0003, Weiliang Meng, Yiwei Shi, Hangyu Mao, Bin Zhang 0052, Dapeng Li 0001, Jiangjin Yin |
ICML | 6 |
| 2025 | SheetAgent: Towards a Generalist Agent for Spreadsheet Reasoning and Manipulation via Large Language ModelsabstractSpreadsheets are ubiquitous across the World Wide Web, playing a critical role in enhancing work efficiency across various domains. Large language model (LLM) has been recently attempted for automatic spreadsheet manipulation but has not yet been investigated in complicated and realistic tasks where reasoning challenges exist (e.g., long horizon manipulation with multi-step reasoning and ambiguous requirements). To bridge the gap with the real-world requirements, we introduce SheetRM, a benchmark featuring long-horizon and multi-category tasks with reasoning-dependent manipulation caused by real-life challenges. To mitigate the above challenges, we further propose SheetAgent, a novel autonomous agent that utilizes the power of LLMs. SheetAgent consists of three collaborative modules: Planner, Informer, and Retriever, achieving both advanced reasoning and accurate manipulation over spreadsheets without human interaction through iterative task reasoning and reflection. Extensive experiments demonstrate that SheetAgent delivers 20--40% pass rate improvements on multiple benchmarks over baselines, achieving enhanced precision in spreadsheet manipulation and demonstrating superior table reasoning abilities. More details and visualizations are available at the https://sheetagent.github.io/. The datasets and source code are available at https://anonymous.4open.science/r/SheetAgent. Yibin Chen, Yifu Yuan, Yan Zheng 0002, Jinyi Liu 0002, Fei Ni 0001, Jianye Hao, Hangyu Mao |
WWW | 8 |
| 2025 | Efficient Missing Key Tag Identification in Large-Scale RFID Systems: An Iterative Verification and Selection MethodabstractRadio frequency identification (RFID) system has been extensively employed to track missing items by affixing them with RFID tags. Many practical applications require to efficiently identify missing events for a specific subset of system tags (called key tags) due to their elevated importance. Existing methods primarily aim to identify all tags, which makes it challenging to specifically identify key tags because of interference from other non-key tags (called ordinary tags). In light of this, several key tag identification methods follow a two-step scheme that filters ordinary tags first and then identifies key tags. Nevertheless, this wastes too much time on tag filtering, resulting in low time efficiency. This paper presents a novel missing key tag identification protocol with two creative designs to gain high efficiency. First, we develop a novel verification technique that can rapidly determine the presence or absence of key tags amid the scenarios with both key tags and ordinary ones. By combining the ON-OFF Keying modulation, we could verify multiple key tags in a single slot, thereby reducing the total slots required. Second, we design a new selection technique that efficiently selects the unverified key tags for further verification, while filtering out the verified key tags and irrelevant ordinary tags to avoid redundant data transmission. Additionally, we present an enhancement protocol that leverages a preselection technique to avoid collecting useless tag responses, further boosting efficiency. We carry out rigorous theoretical analysis to optimize the performance of the proposed protocols. Both simulations and practical experiments demonstrate that our method is markedly superior to state-of-the-art solutions. Jiangjin Yin, Xin Xie 0001, Hangyu Mao, Song Guo 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | SQL-to-Schema Enhances Schema Linking in Text-to-SQL
Sun Yang, Qiong Su, Zhishuai Li, Ziyue Li 0002, Hangyu Mao, Chenxi Liu 0003, Rui Zhao 0001 |
DEXA (1) | 5 |
| 2024 | Sequential Asynchronous Action Coordination in Multi-Agent Systems: A Stackelberg Decision Transformer ApproachabstractAsynchronous action coordination presents a pervasive challenge in Multi-Agent Systems (MAS), which can be represented as a Stackelberg game (SG). However, the scalability of existing Multi-Agent Reinforcement Learning (MARL) methods based on SG is severely restricted by network architectures or environmental settings. To address this issue, we propose the Stackelberg Decision Transformer (STEER). It efficiently manages decision-making processes by incorporating the hierarchical decision structure of SG, the modeling capability of autoregressive sequence models, and the exploratory learning methodology of MARL. Our approach exhibits broad applicability across diverse task types and environmental configurations in MAS. Experimental results demonstrate both the convergence of our method towards Stackelberg equilibrium strategies and its superiority over strong baselines in complex scenarios. Bin Zhang 0052, Hangyu Mao, Lijuan Li 0002, Zhiwei Xu 0005, Dapeng Li 0001, Rui Zhao 0001 |
ICML | 2 |
| 2024 | PTDE: Personalized Training with Distilled Execution for Multi-Agent Reinforcement Learning
Yiqun Chen 0004, Hangyu Mao, Jiaxin Mao, Shiguang Wu 0001, Bin Zhang 0052, Wei Yang 0041, Hongxing Chang |
IJCAI | 2 |
| 2024 | X-Light: Cross-City Traffic Signal Control Using Transformer on Transformer as Meta Multi-Agent Reinforcement Learner
Haoyuan Jiang, Ziyue Li 0002, Hua Wei 0001, Xuantang Xiong, Jingqing Ruan, Jiaming Lu, Hangyu Mao, Rui Zhao 0001 |
IJCAI | 7 |
| 2024 | CoSLight: Co-optimizing Collaborator Selection and Decision-making to Enhance Traffic Signal ControlabstractEffective multi-intersection collaboration is pivotal for reinforcement-learning-based traffic signal control to alleviate congestion. Existing work mainly chooses neighboring intersections as collaborators. However, quite a lot of congestion, even some wide-range congestion, is caused by non-neighbors failing to collaborate. To address these issues, we propose to separate the collaborator selection as a second policy to be learned, concurrently being updated with the original signal-controlling policy. Specifically, the selection policy in real-time adaptively selects the best teammates according to phase- and intersection-level features. Empirical results on both synthetic and real-world datasets provide robust validation for the superiority of our approach, offering significant improvements over existing state-of-the-art methods. Code is available at https://github.com/bonaldli/CoSLight. Jingqing Ruan, Ziyue Li 0002, Hua Wei 0001, Haoyuan Jiang, Jiaming Lu, Xuantang Xiong, Hangyu Mao, Rui Zhao 0001 |
KDD | 7 |
| 2024 | Parallel Missing Tag Identification for Anonymous Multiple Users RFID SystemsabstractRadio frequency identification (RFID) system has been widely employed in warehouse management and supply logistics. A fundamental systematic functionality is to determine the presence or absence of tagged items, referred to as missing tag identification. Although this research has attracted exten-sive attention, prior methods predominantly address scenarios involving a single user. In this paper, we extend the research to more common multi-user scenarios, which raise two concerns: time sensitivity and privacy protection. We propose a novel Parallel Missing tag Identification protocol (PaMI) that leverages lightweight and anonymous bit-vector techniques to specify the data and order of tag responses. This effectively prevents both intra- and inter-user tag collisions, enabling the identification of missing tags across multiple users in one shot while safeguarding user privacy. We carry out a comprehensive theoretical analysis to optimize the proposed method's performance. Extensive experiments show that our protocol strikingly outperforms state-of - the-art baselines. Jiangjin Yin, Hangyu Mao, Rongbo Zhu |
SECON | 2 |
| 2024 | WToE: Learning When to Explore in Multiagent Reinforcement LearningabstractExisting multiagent exploration works focus on how to explore in the fully cooperative task, which is insufficient in the environment with nonstationarity induced by agent interactions. To tackle this issue, we propose When to Explore (WToE), a simple yet effective variational exploration method to learn WToE under nonstationary environments. WToE employs an interaction-oriented adaptive exploration mechanism to adapt to environmental changes. We first propose a novel graphical model that uses a latent random variable to model the step-level environmental change resulting from interaction effects. Leveraging this graphical model, we employ the supervised variational auto-encoder (VAE) framework to derive a short-term inferred policy from historical trajectories to deal with the nonstationarity. Finally, agents engage in exploration when the short-term inferred policy diverges from the current actor policy. The proposed approach theoretically guarantees the convergence of the Q -value function. In our experiments, we validate our exploration mechanism in grid examples, multiagent particle environments and the battle game of MAgent environments. The results demonstrate the superiority of WToE over multiple baselines and existing exploration methods, such as MAEXQ, NoisyNets, EITI, and PR2. Shaokang Dong, Hangyu Mao, Shangdong Yang, Shengyu Zhu 0001, Wenbin Li 0006, Jianye Hao, Yang Gao 0001 |
IEEE Trans. Cybern. | 2 |
| 2024 | A General Scenario-Agnostic Reinforcement Learning for Traffic Signal ControlabstractReinforcement learning (RL) can automatically learn a better policy through a trial-and-error paradigm and has been adopted to revolutionize and optimize traditional traffic signal control systems that are usually based on handcrafted methods. However, most existing RL-based models are either based on a single scenario or multiple independent scenarios, where each scenario has a separate simulation environment with predefined road network topology and traffic signal settings. These models implement training and testing in the same scenario, thus being strictly tied up with the specific setting and sacrificing model generalization heavily. While a few recent models could be trained by multiple scenarios, they require a huge amount of manual labor to label the intersection structure, hindering the model’s generalization. In this work, we aim at ageneralframework that could eliminate heavy labeling and model a variety of scenariossimultaneously. To this end, we propose a general Scenario-Agnostic (GESA) reinforcement learning framework for traffic signal control with: (1) A general plug-in module to map all different intersections into a unified structure, freeing us from the heavy manual labor to specify the structure of intersections; (2) A unified state and action space design to keep the model input and output consistently structured; (3) A large-scale co-training with multiple scenarios, leading to a generic traffic signal control algorithm. GESA can automatically handle various structured intersections from various cities without human labeling, and it co-trains a generalist agent to control traffic signals for multiple cities together, which also demonstrates superior transferability in zero-shot settings. In experiments, we demonstrate our algorithm as the first one that can be co-trained with seven different scenarios without manual annotation and gets13.27%higher rewards than baselines. When dealing with a new scenario, our model can still achieve9.39%higher rewards. The code, scenarios, and demos are available https://github.com/bonaldli/GESA. Haoyuan Jiang, Ziyue Li 0002, Zhishuai Li, Lei Bai 0001, Hangyu Mao, Wolfgang Ketter, Rui Zhao 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | Learning to Collaborate by Grouping: A Consensus-Oriented Strategy for Multi-Agent Reinforcement LearningabstractMulti-agent systems require effective coordination between groups and individuals to achieve common goals. However, current multi-agent reinforcement learning (MARL) methods primarily focus on improving individual policies and do not adequately address group-level policies, which leads to weak cooperation. To address this issue, we propose a novel Consensus-oriented Strategy (CoS) that emphasizes group and individual policies simultaneously. Specifically, CoS comprises two main components: (a) the vector quantized group consensus module, which extracts discrete latent embeddings that represent the stable and discriminative group consensus, and (b) the group consensus-oriented strategy, which integrates the group policy using a hypernet and the individual policies using the group consensus, thereby promoting coordination at both the group and individual levels. Through empirical experiments on cooperative navigation tasks with both discrete and continuous spaces, as well as google research football, we demonstrate that CoS outperforms state-of-the-art MARL algorithms and achieves better collaboration, thus providing a promising solution for achieving effective coordination in multi-agent systems. Jingqing Ruan, Xiaotian Hao, Dong Li 0016, Hangyu Mao |
ECAI | 4 |
| 2023 | Boosting Multiagent Reinforcement Learning via Permutation Invariant and Permutation Equivariant Networks
Jianye Hao, Xiaotian Hao, Hangyu Mao, Weixun Wang, Yaodong Yang 0002, Dong Li 0016, Yan Zheng 0002, Zhen Wang 0004 |
ICLR | 3 |
| 2023 | A Dual-Agent Scheduler for Distributed Deep Learning Jobs on Public Cloud via Reinforcement LearningabstractPublic cloud GPU clusters are becoming emerging platforms for training distributed deep learning jobs. Under this training paradigm, the job scheduler is a crucial component to improve user experiences, i.e., reducing training fees and job completion time, which can also save power costs for service providers. However, the scheduling problem is known to be NP-hard. Most existing work divides it into two easier sub-tasks, i.e., ordering task and placement task, which are responsible for deciding the scheduling orders of jobs and placement orders of GPU machines, respectively. Due to the superior adaptation ability, learning-based policies can generally perform better than traditional heuristic-based methods. Nevertheless, there are still two main challenges that have not been well-solved. First, most learning-based methods only focus on ordering or placement policy independently, while ignoring their cooperation. Second, the unbalanced machine performances and resource contention impose huge overhead and uncertainty on job duration, but rarely be considered in existing work. To tackle these issues, this paper presents a dual-agent scheduler framework abstracted from the two sub-tasks to jointly learn the ordering and placement policies and make better-informed scheduling decisions. Specifically, we design an ordering agent with a scalable squeeze-and-communicate strategy for better cooperation; for the placement agent, we propose a novel Random Walk Gaussian Process to learn the performance similarities of GPU machines while being aware of the uncertain performance fluctuation. Finally, the dual-agent is jointly optimized with multi-agent reinforcement learning. Extensive experiments conducted on the real-world production cluster trace demonstrate the superiority of our model. Mingzhe Xing, Hangyu Mao, Shenglin Yin, Lichen Pan, Zhengchao Zhang, Jieyi Long |
KDD | 2 |
| 2022 | What about Inputting Policy in Value Function: Policy Representation and Policy-Extended Value Function ApproximatorabstractWe study Policy-extended Value Function Approximator (PeVFA) in Reinforcement Learning (RL), which extends conventional value function approximator (VFA) to take as input not only the state (and action) but also an explicit policy representation. Such an extension enables PeVFA to preserve values of multiple policies at the same time and brings an appealing characteristic, i.e., value generalization among policies. We formally analyze the value generalization under Generalized Policy Iteration (GPI). From theoretical and empirical lens, we show that generalized value estimates offered by PeVFA may have lower initial approximation error to true values of successive policies, which is expected to improve consecutive value approximation during GPI. Based on above clues, we introduce a new form of GPI with PeVFA which leverages the value generalization along policy improvement path. Moreover, we propose a representation learning framework for RL policy, providing several approaches to learn effective policy embeddings from policy network parameters or state-action pairs. In our experiments, we evaluate the efficacy of value generalization offered by PeVFA and policy representation learning in several OpenAI Gym continuous control tasks. For a representative instance of algorithm implementation, Proximal Policy Optimization (PPO) re-implemented under the paradigm of GPI with PeVFA achieves about 40% performance improvement on its vanilla counterpart in most environments. Hongyao Tang, Zhaopeng Meng, Jianye Hao, Chen Chen 0077, Daniel Graves, Dong Li 0016, Changmin Yu, Hangyu Mao, Wulong Liu, Yaodong Yang 0002, Wenyuan Tao |
AAAI | 8 |
| 2022 | Fast and Fine-grained Autoscaler for Streaming Jobs with Reinforcement LearningabstractOn computing clusters, the autoscaler is responsible for allocating resources for jobs or fine-grained tasks to ensure their Quality of Service. Due to a more precise resource management, fine-grained autoscaling can generally achieve better performance. However, the fine-grained autoscaling for streaming jobs needs intensive computation to model the complicated running states of tasks, and has not been adequately studied previously. In this paper, we propose a novel fine-grained autoscaler for streaming jobs based on reinforcement learning. We first organize the running states of streaming jobs as spatio-temporal graphs. To efficiently make autoscaling decisions, we propose a Neural Variational Subgraph Sampler to sample spatio-temporal subgraphs. Furthermore, we propose a mutual-information-based objective function to explicitly guide the sampler to extract more representative subgraphs. After that, the autoscaler makes decisions based on the learned subgraph representations. Experiments conducted on real-world datasets demonstrate the superiority of our method over six competitive baselines. Mingzhe Xing, Hangyu Mao |
IJCAI | 2 |
| 2022 | Flat-Aware Cross-Stage Distilled Framework for Imbalanced Medical Image Classification
Jinpeng Li 0004, Guangyong Chen, Hangyu Mao, Danruo Deng, Dong Li 0016, Jianye Hao, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (3) | 3 |
| 2022 | Optimizing communication in deep reinforcement learning with XingTianabstractDeep Reinforcement Learning (DRL) achieves great success in various domains. Communication in today's DRL algorithms takes non-negligible time compared to the computation. However, prior DRL frameworks usually focus on computation management while paying little attention to communication optimization, and fail to utilize the opportunity of the communication-computation overlap that hides the communication from the critical path of DRL algorithms. Consequently, communication can take more time than the computation in prior DRL frameworks. In this paper, we present XingTian, a novel DRL framework that co-designs the management of communication and computation in DRL algorithms. XingTian organizes the computation in DRL algorithms in a decentralized way and provides an asynchronous communication channel. XingTian makes the communication execute asynchronously and aggressively and takes advantage of the communication-computation overlapping opportunity from DRL algorithms. Experimental results show that XingTian improves data transmission efficiency and can transmit at least twice as much data per second as the state-of-the-art DRL framework RLLib. DRL algorithms based on XingTian achieve up to 70.71% more throughput than RLLib-based ones with better or similar convergent performance. XingTian maintains high communication efficiency under different scale deployments and the XingTian-based DRL algorithm achieves 91.12% higher throughput than the RLLib-based one when deployed in four machines. XingTian is open-sourced and publicly available at https://github.com/huawei-noah/xingtian. Lichen Pan, Hangyu Mao, Pengze Li |
Middleware | 4 |
| 2022 | Multiagent Q-learning with Sub-Team CoordinationabstractIn many real-world cooperative multiagent reinforcement learning (MARL) tasks, teams of agents can rehearse together before deployment, but then communication constraints may force individual agents to execute independently when deployed. Centralized training and decentralized execution (CTDE) is increasingly popular in recent years, focusing mainly on this setting. In the value-based MARL branch, credit assignment mechanism is typically used to factorize the team reward into each individual’s reward — individual-global-max (IGM) is a condition on the factorization ensuring that agents’ action choices coincide with team’s optimal joint action. However, current architectures fail to consider local coordination within sub-teams that should be exploited for more effective factorization, leading to faster learning. We propose a novel value factorization framework, called multiagent Q-learning with sub-team coordination (QSCAN), to flexibly represent sub-team coordination while honoring the IGM condition. QSCAN encompasses the full spectrum of sub-team coordination according to sub-team size, ranging from the monotonic value function class to the entire IGM function class, with familiar methods such as QMIX and QPLEX located at the respective extremes of the spectrum. Experimental results show that QSCAN’s performance dominates state-of-the-art methods in matrix games, predator-prey tasks, the Switch challenge in MA-Gym. Additionally, QSCAN achieves comparable performances to those methods in a selection of StarCraft II micro-management tasks. Wenhan Huang, Kai Li 0022, Kun Shao, Tianze Zhou, Matthew E. Taylor, Jun Luo 0009, Dongge Wang 0001, Hangyu Mao, Jianye Hao, Jun Wang 0012, Xiaotie Deng |
NeurIPS | 8 |
| 2022 | Common belief multi-agent reinforcement learning based on variational recurrent models
Xianjie Zhang, Yu Liu 0035, Hangyu Mao |
Neurocomputing | 3 |
| 2021 | SEIHAI: A Sample-Efficient Hierarchical AI for the MineRL Competition
Hangyu Mao, Xiaotian Hao, Yihuan Mao, Chengjie Wu, Jianye Hao, Dong Li 0016, Pingzhong Tang |
DAI | 1 |
| 2021 | An Efficient Transfer Learning Framework for Multiagent Reinforcement LearningabstractTransfer Learning has shown great potential to enhance single-agent Reinforcement Learning (RL) efficiency. Similarly, Multiagent RL (MARL) can also be accelerated if agents can share knowledge with each other. However, it remains a problem of how an agent should learn from other agents. In this paper, we propose a novel Multiagent Policy Transfer Framework (MAPTF) to improve MARL efficiency. MAPTF learns which agent's policy is the best to reuse for each agent and when to terminate it by modeling multiagent policy transfer as the option learning problem. Furthermore, in practice, the option module can only collect all agent's local experiences for update due to the partial observability of the environment. While in this setting, each agent's experience may be inconsistent with each other, which may cause the inaccuracy and oscillation of the option-value's estimation. Therefore, we propose a novel option learning algorithm, the successor representation option learning to solve it by decoupling the environment dynamics from rewards and learning the option-value under each agent's preference. MAPTF can be easily combined with existing deep RL and MARL approaches, and experimental results show it significantly boosts the performance of existing methods in both discrete and continuous state spaces. Tianpei Yang, Weixun Wang, Hongyao Tang, Jianye Hao, Zhaopeng Meng, Hangyu Mao, Dong Li 0016, Wulong Liu, Yujing Hu, Changjie Fan, Chengwei Zhang 0001 |
NeurIPS | 6 |
| 2021 | Structural relational inference actor-critic for multi-agent reinforcement learning
Xianjie Zhang, Yu Liu 0035, Xiujuan Xu, Qiong Huang 0003, Hangyu Mao, Chettupally Anil Carie |
Neurocomputing | 5 |
| 2020 | Neighborhood Cognition Consistent Multi-Agent Reinforcement LearningabstractSocial psychology and real experiences show that cognitive consistency plays an important role to keep human society in order: if people have a more consistent cognition about their environments, they are more likely to achieve better cooperation. Meanwhile, only cognitive consistency within a neighborhood matters because humans only interact directly with their neighbors. Inspired by these observations, we take the first step to introduce neighborhood cognitive consistency (NCC) into multi-agent reinforcement learning (MARL). Our NCC design is quite general and can be easily combined with existing MARL methods. As examples, we propose neighborhood cognition consistent deep Q-learning and Actor-Critic to facilitate large-scale multi-agent cooperations. Extensive experiments on several challenging tasks (i.e., packet routing, wifi configuration and Google football player control) justify the superior performance of our methods compared with state-of-the-art MARL approaches. Hangyu Mao, Wulong Liu, Jianye Hao, Jun Luo 0009, Dong Li 0016, Zhengchao Zhang, Jun Wang 0012 |
AAAI | 1 |
| 2020 | Learning Agent Communication under Limited Bandwidth by Message PruningabstractCommunication is a crucial factor for the big multi-agent world to stay organized and productive. Recently, Deep Reinforcement Learning (DRL) has been applied to learn the communication strategy and the control policy for multiple agents. However, the practical limited bandwidth in multi-agent communication has been largely ignored by the existing DRL methods. Specifically, many methods keep sending messages incessantly, which consumes too much bandwidth. As a result, they are inapplicable to multi-agent systems with limited bandwidth. To handle this problem, we propose a gating mechanism to adaptively prune less beneficial messages. We evaluate the gating mechanism on several tasks. Experiments demonstrate that it can prune a lot of messages with little impact on performance. In fact, the performance may be greatly improved by pruning redundant messages. Moreover, the proposed gating mechanism is applicable to several previous methods, equipping them the ability to address bandwidth restricted settings. Hangyu Mao, Zhengchao Zhang, Zhibo Gong, Yan Ni |
AAAI | 1 |
| 2020 | Learning multi-agent communication with double attentional deep reinforcement learning
Hangyu Mao, Zhengchao Zhang, Zhibo Gong, Yan Ni |
Auton. Agents Multi Agent Syst. | 1 |
| 2018 | Topic-Specific Retweet Count Ranking for Weibo
Hangyu Mao, Yuan Wang 0009, Jiakang Wang |
PAKDD (3) | 1 |
| 2016 | Predicting Restaurant Consumption Level through Social Media FootprintsabstractAccurate prediction of user attributes from social media is valuable for both social science analysis and consumer targeting. In this paper, we propose a systematic method to leverage user online social media content for predicting offline restaurant consumption level. We utilize the social login as a bridge and construct a dataset of 8,844 users who have been linked across Dianping (similar to Yelp) and Sina Weibo. More specifically, we construct consumption level ground truth based on user self report spending. We build predictive models using both raw features and, especially, latent features, such as topic distributions and celebrities clusters. The employed methods demonstrate that online social media content has strong predictive power for offline spending. Finally, combined with qualitative feature analysis, we present the differences in words usage, topic interests and following behavior between different consumption level groups. Yuan Wang 0009, Hangyu Mao |
COLING | 3 |