Tenghai Qiu

dblp:258/9926 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
16since 2021 · last 2026
0000-0002-0312-5728ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 1 first-author · 14 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Unreal-MAP: Unreal-Engine-Based General Platform for Multi-agent Reinforcement Learning
abstract
In this paper, we propose Unreal Multi-Agent Playground (Unreal-MAP), an MARL general platform based on the Unreal-Engine (UE). Unreal-MAP allows users to freely create multi-agent tasks using the vast visual and physical resources available in the UE community, and deploy state-of-the-art (SOTA) MARL algorithms within them. Unreal-MAP is user-friendly in terms of deployment, modification, and visualization, and all its components are open-source. We also develop an experimental framework compatible with algorithms ranging from rule-based to learning-based provided by third-party frameworks. Lastly, we deploy several SOTA algorithms in example tasks developed via Unreal-MAP, and conduct corresponding experimental analyses including a sim2real demo. We believe Unreal-MAP can play an important role in the MARL field by closely integrating existing algorithms with user-customized tasks, thus advancing the field of MARL.
Qingxu Fu, Zhiqiang Pu, Tenghai Qiu
AAAI5
2025 Mine-SSD: Dual-threshold set abstraction and radius-adaptive grouping for 3D object detection in open-pit mines
Zhongyu Xie, Yuqian Zhao 0001, Fan Zhang 0106, Biao Luo 0001, Wenliu Hu, Tenghai Qiu
Neurocomputing6
2025 A Policy Resonance Approach to Solve the Problem of Responsibility Diffusion in Multiagent Reinforcement Learning
abstract
State-of-the-art (SOTA) multiagent reinforcement algorithms distinguish themselves in many ways from their single-agent equivalences. However, most of them still totally inherit the single-agent exploration-exploitation strategy. Naively inheriting this strategy from single-agent algorithms causes potential collaboration failures, in which the agents blindly follow mainstream behaviors and reject taking minority responsibility. We name this problem the responsibility diffusion (RD) as it shares similarities with the same-name social psychology effect. In this work, we start by theoretically analyzing the cause of this RD problem, which can be traced back to the exploration-exploitation dilemma of multiagent systems (especially large-scale multiagent systems). We address this RD problem by proposing a policy resonance (PR) approach which modifies the collaborative exploration strategy of agents by refactoring the joint agent policy while keeping individual policies approximately invariant. Next, we show that SOTA algorithms can equip this approach to promote the collaborative performance of agents in complex cooperative tasks. Experiments are performed in multiple test benchmark tasks to illustrate the effectiveness of this approach.
Qingxu Fu, Tenghai Qiu, Jianqiang Yi, Zhiqiang Pu, Xiaolin Ai, Wanmai Yuan
IEEE Trans. Neural Networks Learn. Syst.2
2025 Cognition-Oriented Multiagent Reinforcement Learning
abstract
Inspired by psychological insights into individual behavior, we propose a novel cognition-oriented multiagent reinforcement learning (CORL) framework. CORL equips agents with two distinct types of cognition-situational and self-cognition-derived from local observations. To enhance the informativeness and precision of these cognition types, we introduce two information-theoretical regularizers: one to align situational cognition with the global state and the other to align self-cognition with each agent's identity for improved role differentiation and team coordination. In addition, the centralized training and decentralized execution framework is adopted to train the policy network. Our simulations demonstrate that CORL effectively harnesses local observations for enriched cooperation, leading to pronounced performance improvements, particularly in challenging tasks.
Tenghai Qiu, Shiguang Wu 0001, Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi, Yuqian Zhao 0001, Biao Luo 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Heterogeneous Observation Aggregation Network for Multi-agent Reinforcement Learning
abstract
Learning effective policies is challenging for a multi-agent system in partially observable environments, where agents need to extract relevant features from local observations. Most approaches in multi-agent reinforcement learning (MARL) are limited to feature extraction for homogenous agents. They struggle to deal with local observations in heterogeneous multi-agent scenarios, where agents have different observation spaces and are necessitated to process semantically varied information. To address this issue, we analyze the observational heterogeneity of multi-agent systems, and propose a heterogeneous-graph-based approach for feature extraction in MARL. We model agent observations as heterogeneous graphs, and design a heterogeneous observation aggregation network (HOA-Net) for processing these graph-based observations. HOA-Net is specifically designed to address various forms of observational heterogeneity. It employs class-specific weighting networks and computes across-class attentions for observed entities, effectively reducing the number of learnable parameters. The proposed method is evaluated on SMAC and an Unreal-Engine-based heterogeneous multi-agent testbed. Experimental results demonstrate that our method significantly outperforms other baselines in effectively aggregating an agent’s observation, and finally enhancing the performance of heterogeneous multi-agent systems.
Xiaolin Ai, Zhiqiang Pu, Tenghai Qiu, Jianqiang Yi
IJCNN4
2024 Fuzzy Feedback Multiagent Reinforcement Learning for Adversarial Dynamic Multiteam Competitions
abstract
A large proportion of recent studies on cooperative Multi-Agent Reinforcement Learning (MARL) focus on the policylearning process in scenarios with stationary opponents (or without opponents). This paper, instead, investigates a different challenge of achieving team superiority in dynamic competitions among competitors that evolve dynamically with MARL. We aim to enhance the competitiveness of such MARL learners by enabling them to adjust their own learning settings dynamically, so as to take quick counter-measures against the policy shift of competitor learners, or to learn faster to suppress the opponents. We propose a Competitive Auto-Multiagent Learner with Fuzzy Feedback (CALF) with two essential highlights: (1) CALF establishes feedback controllers to achieve real-time adjustments based on fuzzy logic, using human-readable fuzzy rules to provide significant explainability and flexibility; (2) CALF integrates Bayesian Optimization to search and optimize the feedback fuzzy logic rules automatically. CALF can be used to apply real-time adjustments for MARL hyperparameters and intrinsic rewards. We also give solid empirical results to show that CALF significantly promotes team competitiveness in adversarial competitions, spanning from small-scale tasks involving 2 teams to large-scale tasks involving 3 teams and hundreds of agents. Furthermore, CALF exhibits superior competitiveness when engaging in competition with established competitors like Qmix, Qtran, and Qplex in dynamic competitive environments. Moreover, the experiments also demonstrate that the integration of the fuzzy logic with Bayesian Optimization offers considerable transferability and explainability, enabling a CALF-implemented learner optimized from one scenario to be transferred to other distinct scenarios.
Qingxu Fu, Zhiqiang Pu, Yi Pan 0009, Tenghai Qiu, Jianqiang Yi
IEEE Trans. Fuzzy Syst.4
2023 Learning Superior Cooperative Policy in Adversarial Multi-Team Reinforcement Learning
abstract
Multi-agent Reinforcement Learning (MARL) has become a powerful tool for addressing multi-agent challenges. Existing studies have explored numerous models to use MARL to solve single-team cooperation (competition) problems and adversarial problems with opponents controlled by static knowledge-based policies. However, most studies in the literature often ignore adversarial multi-team problems involving dynamically evolving opponents. We investigate adversarial multi-team problems where all participating teams use MARL learners to learn policies against each other. Two objectives are achieved in this study. Firstly, we design an adversarial team-versus-team learning framework to generate cooperative multi-agent policies to compete against opponents without preprogrammed opponent partners or any supervision. Secondly, we explore the key factors to achieve win-rate superiority during dynamic competitions. Then we put forward a novel FeedBack MARL (FBMARL) algorithm that takes advantage of feedback loops to adjust optimizer hyper-parameters based on real-time game statistics. Finally, the effectiveness of our FBMARL model is tested in a benchmark environment named Multi-Team Decentralized Collective Assault (MT-DCA). The results demonstrate that our feedback MARL model can achieve superior performance over baseline competitor MARL learners in 2-team and 3-team dynamic competitions.
Qingxu Fu, Tenghai Qiu, Zhiqiang Pu, Jianqiang Yi, Xiaolin Ai, Wanmai Yuan
IJCNN2
2023 Deep-Reinforcement-Learning-Based Multitarget Coverage With Connectivity Guaranteed
abstract
Deriving a distributed, time-efficient, and connectivity-guaranteed coverage policy in multitarget environment poses huge challenges for a multirobot team with limited coverage and limited communication. In particular, the robot team needs to cover multiple targets while preserving connectivity. In this article, a novel deep-reinforcement-learning-based approach is proposed to take both multitarget coverage and connectivity preservation into account simultaneously, which consists of four parts: a hierarchical observation attention representation, an interaction attention representation, a two-stage policy learning, and a connectivity-guaranteed policy filtering. The hierarchical observation attention representation is designed for each robot to extract the latent features of the relations from its neighboring robots and the targets. To promote the cooperation behavior among the robots, the interaction attention representation is designed for each robot to aggregate information from its neighboring robots. Moreover, to speed up the training process and improve the performance of the learned policy, the two-stage policy learning is presented using two reward functions based on algebraic connectivity and coverage rate. Furthermore, the learned policy is filtered to strictly guarantee the connectivity based on a model of connectivity maintenance. Finally, the effectiveness of the proposed method is validated by numerous simulations. Besides, our method is further deployed to an experimental platform based on quadrotor unmanned aerial vehicles and omnidirectional vehicles. The experiments illustrate the practicability of the proposed method.
Shiguang Wu 0001, Zhiqiang Pu, Tenghai Qiu, Jianqiang Yi
IEEE Trans. Ind. Informatics3
2023 A Deep Reinforcement Learning Approach Combined With Model-Based Paradigms for Multiagent Formation Control With Collision Avoidance
abstract
Generating collision-free formation control strategy for multiagent systems faces huge challenges in collaborative navigation tasks, especially in a highly dynamic and uncertain environment. Two typical methodologies for solving this problem are the conventional model-based paradigm and the data-driven paradigm, particularly the widely used deep reinforcement learning (DRL) method. However, both the model-based and data-driven paradigms encounter inherent drawbacks. In this paper, we present two novel general schemes that combine these two paradigms together in an online mode. Specifically, the two paradigms are combined in a parallel and a serial structure in these two schemes, respectively. In the parallel scheme, the outputs of the model-based and DRL-based controllers are lumped together. In the serial scheme, the output of the model-based controller is fed as an input of the DRL-based controller. The interpretation of the two combined schemes is suggested from a control-oriented perspective, where the parallel DRL controller is viewed as a complementary uncertainty compensator and the serial DRL controller is taken as an inverse dynamics estimator. Finally, comprehensive simulations are conducted to demonstrate the superiority of the proposed schemes, and the effectiveness is further verified by deploying our schemes to a physical experiment platform based on a set of three-wheeled omnidirectional robots.
Zhiqiang Pu, Xiaolin Ai, Tenghai Qiu, Jianqiang Yi
IEEE Trans. Syst. Man Cybern. Syst.4
2022 Concentration Network for Reinforcement Learning of Large-Scale Multi-Agent Systems
abstract
When dealing with a series of imminent issues, humans can naturally concentrate on a subset of these concerning issues by prioritizing them according to their contributions to motivational indices, e.g., the probability of winning a game. This idea of concentration offers insights into reinforcement learning of sophisticated Large-scale Multi-Agent Systems (LMAS) participated by hundreds of agents. In such an LMAS, each agent receives a long series of entity observations at each step, which can overwhelm existing aggregation networks such as graph attention networks and cause inefficiency. In this paper, we propose a concentration network called ConcNet. First, ConcNet scores the observed entities considering several motivational indices, e.g., expected survival time and state value of the agents, and then ranks, prunes, and aggregates the encodings of observed entities to extract features. Second, distinct from the well-known attention mechanism, ConcNet has a unique motivational subnetwork to explicitly consider the motivational indices when scoring the observed entities. Furthermore, we present a concentration policy gradient architecture that can learn effective policies in LMAS from scratch. Extensive experiments demonstrate that the presented architecture has excellent scalability and flexibility, and significantly outperforms existing methods on LMAS benchmarks.
Qingxu Fu, Tenghai Qiu, Jianqiang Yi, Zhiqiang Pu, Shiguang Wu 0001
AAAI2
2022 A Cooperation Graph Approach for Multiagent Sparse Reward Reinforcement Learning
abstract
Multiagent reinforcement learning (MARL) can solve complex cooperative tasks. However, the efficiency of existing MARL methods relies heavily on well-defined reward functions. Multiagent tasks with sparse reward feedback are especially challenging not only because of the credit distribution problem, but also due to the low probability of obtaining positive reward feedback. In this paper, we design a graph network called Cooperation Graph (CG). The Cooperation Graph is the combination of two simple bipartite graphs, namely, the Agent Clustering subgraph (ACG) and the Cluster Designating subgraph (CDG). Next, based on this novel graph structure, we propose a Cooperation Graph Multiagent Reinforcement Learning (CG-MARL) algorithm, which can efficiently deal with the sparse reward problem in multiagent tasks. In CG-MARL, agents are directly controlled by the Cooperation Graph. And a policy neural network is trained to manipulate this Cooperation Graph, guiding agents to achieve cooperation in an implicit way. This hierarchical feature of CG-MARL provides space for customized cluster-actions, an extensible interface for introducing fundamental cooperation knowledge. In experiments, CG-MARL shows state-of-the-art performance in sparse reward multiagent benchmarks, including the anti-invasion interception task and the multi-cargo delivery task.
Qingxu Fu, Tenghai Qiu, Zhiqiang Pu, Jianqiang Yi, Wanmai Yuan
IJCNN2
2022 Multi-Agent Local Information Reconstruction for Situational Cognition
abstract
Learning an effective strategy is challenging for agents in partially observable environment, where the agents can only observe a part of environment information and make decisions based on local information. Hence, how to effectively utilize local information to achieve efficient cooperation among the agents is particularly important. The agents can establish an understanding of themselves and their surrounding environment based on their historical observation information. However, the understanding lacking global information is local or limited, resulting in low performance in some complex tasks. To solve the problem, a situational cognition learning framework is proposed for each agent based on local information, with which each agent can reconstruct a cognition about itself and its surrounding environment and map it into a high dimensional representation space. In particular, situational cognition is modeled as a random variable under the condition of local trajectory. In addition, an information regularizer is introduced to ensure that the situational cognition is complete and accurate through maximizing the mutual information between the situational cognition and global information, conditioned on the local trajectory of the agent. Various simulations are conducted and show that the proposed framework significantly promotes cooperation among the agents and improves performance.
Shiguang Wu 0001, Zhiqiang Pu, Tenghai Qiu, Jianqiang Yi
IJCNN4
2022 Multi-UAV Cooperative Short-Range Combat via Attention-Based Reinforcement Learning using Individual Reward Shaping
abstract
In this paper, we propose a novel distributed method based on attention-based deep reinforcement learning using individual reward shaping, for multiple unmanned aerial vehicles (UAVs) cooperative short-range combat mission. Specifically, a two-level attention distributed policy, composed of observation-level and communication-level attention networks, is designed to enable each UAV to selectively focus on important environmental features and messages, for enhancing the effectiveness of the cooperative policy. Moreover, due to the high complexity and stochasticity of the UAV combat mission, the learning of UAVs is tricky and low efficient. To embed knowledge to accelerate the policy learning, a potential-based individual reward function is constructed by implicitly translating the individual reward into the specific form of dynamic action potentials. In addition, an actor-critic training algorithm based on the centralized training and decentralized execution framework is adopted to train the policy network of UAV maneuver decision. We build a three-dimensional UAV simulation and training platform based on Unity for multi-UAV short-range combat missions. Simulation results demonstrate the effectiveness of the proposed method and the superiority of the attention policy and individual reward shaping.
Tenghai Qiu, Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi, Jinying Zhu, Ruiguang Hu
IROS2
2021 Multi-target Coverage with Connectivity Maintenance using Knowledge-incorporated Policy Framework
abstract
This paper considers a multi-target coverage problem where a robot team aims to efficiently cover multi-targets while maintaining connectivity in a distributed manner. A novel knowledge-incorporated policy framework is proposed to derive a distributed, efficient, and connectivity guaranteed coverage policy. In particular, a knowledge-guided policy network (KGPnet) is designed, which consists of observation attention representation, interaction attention representation, and knowledge-guided policy learning. Giving credit to the KGPnet, the connectivity guaranteed coverage policy can be applied to different number targets. Moreover, based on the knowledge of the algebraic connectivity and coverage rate, a comprehensive reward is designed to guide the training of the behavior of multi-target coverage with connectivity maintenance. Furthermore, since the policy learned through deep reinforcement learning (DRL) can not guarantee the connectivity of the robot team, a knowledge-nested policy filtering is designed to filter dis-connectivity policies to satisfy the connectivity constraint based on the knowledge model of connectivity maintenance. Various simulations are conducted to verify the effectiveness of the proposed method. Besides, numerous real-world experiments with three-wheel omnidirectional cars and a motion capture system are presented to demonstrate the practicability of the proposed method.
Shiguang Wu 0001, Zhiqiang Pu, Zhen Liu 0020, Tenghai Qiu, Jianqiang Yi
ICRA4
2021 Multi-Agent Cognition Difference Reinforcement Learning for Multi-Agent Cooperation
abstract
Multi-agent cooperation is one of the most attractive research fields in multi-agent systems. There are many attempts made by researchers in this field to promote the cooperation behavior. However, in partially-observable environments, a large number of agents and complex interactions among the agents cause huge difficulty for policy learning. Moreover, redundant communication contents caused by many agents make effective features hard to be extracted, which prevents the policy from converging. To address the limitations above, a novel method called multi-agent cognition difference reinforcement learning (MACD-RL) is proposed in this paper. The key feature of MACD-RL lies in cognition difference network (CDN) and a soft communication network (SCN). CDN is designed to allow each agent to choose its neighbors (communication targets) adaptively with its environment cognition difference. SCN is designed to handle the complex interactions among the agents with soft attention mechanism. The results of simulations including mixed cooperative and competitive tasks demonstrate that the effectiveness and robustness of the proposed model.
Huimu Wang, Tenghai Qiu, Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi, Wanmai Yuan
IJCNN2
2021 Multi-agent Collaborative Learning with Relational Graph Reasoning in Adversarial Environments
abstract
This paper proposes a collaborative policy framework via relational graph reasoning for multi-agent systems to accomplish adversarial tasks. A relational graph reasoning module consisting of an agent graph reasoning module and an opponent graph module, is designed to enable each agent to learn mixture state representation to enhance the effectiveness of the policy. In particular, for each agent, the agent graph reasoning module is designed to infer different underlying influences from different opponents and generate agent-level state representation. The opponent graph reasoning module is creatively designed for the opponents to reason relations from their surrounding objects including the agents and the opponents based on their latent features and then predict the future state of the opponents. It forms an opponent-level state representation. Besides, in order to effectively predict the state of the opponents, an intrinsic reward based on prediction error is designed to motivate the policy learning. Furthermore, interactions among agents are utilized to transmit messages and fuse information to promote the cooperative behaviors among the agents. Finally, various representative simulations on two multi-agent adversarial tasks are conducted to demonstrate the superiority and effectiveness of the proposed framework by comparison with existing methods.
Shiguang Wu 0001, Tenghai Qiu, Zhiqiang Pu, Jianqiang Yi
IROS2