Shiguang Wu 0001

dblp:275/7661-1 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
15since 2021 · last 2025
0000-0001-9091-5236ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Cognition-Oriented Multiagent Reinforcement Learning
abstract
Inspired by psychological insights into individual behavior, we propose a novel cognition-oriented multiagent reinforcement learning (CORL) framework. CORL equips agents with two distinct types of cognition-situational and self-cognition-derived from local observations. To enhance the informativeness and precision of these cognition types, we introduce two information-theoretical regularizers: one to align situational cognition with the global state and the other to align self-cognition with each agent's identity for improved role differentiation and team coordination. In addition, the centralized training and decentralized execution framework is adopted to train the policy network. Our simulations demonstrate that CORL effectively harnesses local observations for enriched cooperation, leading to pronounced performance improvements, particularly in challenging tasks.
Tenghai Qiu, Shiguang Wu 0001, Zhen Liu 0020, Zhiqiang Pu, Jianqiang Yi, Yuqian Zhao 0001, Biao Luo 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Generate Subgoal Images Before Act: Unlocking the Chain-of-Thought Reasoning in Diffusion Model for Robot Manipulation with Multimodal Prompts
abstract
Robotics agents often struggle to understand and follow the multi-modal prompts in complex manipulation scenes which are challenging to be sufficiently and accurately described by text alone. Moreover, for long-horizon manipulation tasks, the deviation from general instruction tends to accumulate if lack of intermediate guidance from high-level subgoals. For this, we consider can we generate subgoal images before act to enhance the instruction following in long-horizon manipulation with multi-modal prompts? Inspired by the great success of diffusion model in image generation tasks, we propose a novel hierarchical framework named as CoTDiffusion that incorporates diffusion model as a high-level planner to convert the general and multimodal prompts into coherent visual subgoal plans, which further guide the low-level policy model before action execution. We design a semantic alignment module that can anchor the progress of generated keyframes along a coherent generation chain, unlocking the chain-of-thought reasoning ability of diffusion model. Additionally, we propose bi-directional generation and frame concat mechanism to further enhance the fidelity of generated subgoal images and the accuracy of instruction following. The experiments cover various robotics manipulation scenarios including visual reasoning, visual rearrange, and visual constraints. CoTDiffusion achieves outstanding performance gain compared to the baselines without explicit subgoal generation, which proves that a subgoal image is worth a thousand words of instruction. The details and visualizations are available at https://cotdiffusion.github.io.
Fei Ni 0001, Jianye Hao, Shiguang Wu 0001, Longxin Kou, Jiashun Liu, Yan Zheng 0002, Bin Wang 0034, Yuzheng Zhuang
CVPR3
2024 PTDE: Personalized Training with Distilled Execution for Multi-Agent Reinforcement Learning
Yiqun Chen 0004, Hangyu Mao, Jiaxin Mao, Shiguang Wu 0001, Bin Zhang 0052, Wei Yang 0041, Hongxing Chang
IJCAI4
2024 PERIA: Perceive, Reason, Imagine, Act via Holistic Language and Vision Planning for Manipulation
abstract
Long-horizon manipulation tasks with general instructions often implicitly encapsulate multiple sub-tasks, posing significant challenges in instruction following. While language planning is a common approach to decompose general instructions into stepwise sub-instructions, text-only guidance may lack expressiveness and lead to potential ambiguity. Considering that humans often imagine and visualize sub-instructions reasoning out before acting, the imagined subgoal images can provide more intuitive guidance and enhance the reliability of decomposition. Inspired by this, we propose **PERIA**(**PE**rceive, **R**eason, **I**magine, **A**ct), a novel framework that integrates holistic language planning and vision planning for long-horizon manipulation tasks with complex instructions, leveraging both logical and intuitive aspects of task decomposition. Specifically, we first perform a lightweight multimodal alignment on the encoding side to empower the MLLM to perceive visual details and language instructions. The MLLM is then jointly instruction-tuned with a pretrained image-editing model to unlock capabilities of simultaneous reasoning of language instructions and generation of imagined subgoals. Furthermore, we introduce a consistency alignment loss to encourage coherent subgoal images and align with their corresponding instructions, mitigating potential hallucinations and semantic conflicts between the two planning manners. Comprehensive evaluations across three task domains demonstrate that PERIA, benefiting from holistic language and vision planning, significantly outperforms competitive baselines in both instruction following accuracy and task success rate on complex manipulation tasks.
Fei Ni 0001, Jianye Hao, Shiguang Wu 0001, Longxin Kou, Yifu Yuan, Zibin Dong, Jinyi Liu 0002, Mingzhi Li, Yuzheng Zhuang, Yan Zheng 0002
NeurIPS3
2024 Multiexperience-Assisted Efficient Multiagent Reinforcement Learning
abstract
Recently, multiagent reinforcement learning (MARL) has shown great potential for learning cooperative policies in multiagent systems (MASs). However, a noticeable drawback of current MARL is the low sample efficiency, which causes a huge amount of interactions with environment. Such amount of interactions greatly hinders the real-world application of MARL. Fortunately, effectively incorporating experience knowledge can assist MARL to quickly find effective solutions, which can significantly alleviate the drawback. In this article, a novel multiexperience-assisted reinforcement learning (MEARL) method is proposed to improve the learning efficiency of MASs. Specifically, monotonicity-constrained reward shaping is innovatively designed using expert experience to provide additional individual rewards to guide multiagent learning efficiently, with the invariance guarantee of the team optimization objective. Furthermore, a reward distribution estimator is specially developed to model an implicated reward distribution of environment by using transition experience from environment, containing collected samples (state-action pair, reward, and next state). This estimator can predict the expectation reward of each agent for the taken action to accurately estimate the state value function and accelerate its convergence. Besides, the performance of MEARL is evaluated on two multiagent environment platforms: our designed unmanned aerial vehicle combat (UAV-C) and StarCraft II Micromanagement (SCII-M). Simulation results demonstrate that the proposed MEARL can greatly improve the learning efficiency and performance of MASs and is superior to the state-of-the-art methods in multiagent tasks.
Zhen Liu 0020, Jianqiang Yi, Shiguang Wu 0001, Zhiqiang Pu
IEEE Trans. Neural Networks Learn. Syst.4
2023 Modal-aware Bias Constrained Contrastive Learning for Multimodal Recommendation
abstract
Multimodal recommendation system has been widely used in short video platform, e-commerce platform and news media. Multimodal data contains information such as product image and product text, which is often used as auxiliary signal to improve the effect of recommendation system significantly. In order to alleviate the problems of data sparsity and noise, some researchers construct data augmentation to use self-supervised learning to help model training. These methods have achieved certain results. However, most of the work is based on data augmentation in random ways, such as random masking and random perturbation. This random method is likely to lose important information and introduce new noise, resulting in biased augmentation data. Therefore, we propose a Modal-aware Bias Constrained Contrastive Learning method (BCCL) to solve the above problems. Specifically, BCCL introduces a bias-constrained data augmentation method to ensure the quality of augmentation samples. Then the multi-modal semantic information is modeled by the designed modal awareness module. Furthermore, we propose a information alignment module to improve the sparse modal feature learning of the model. We conducted a comprehensive experiment on three real-world data sets, and the experimental results showed that the proposed BCCL outperformed all the state-of-art methods. In-depth experiments have verified the effectiveness of our proposed modules.
Wei Yang 0041, Zhengru Fang, Shiguang Wu 0001, Chi Lu 0001
ACM Multimedia4
2023 Deep-Reinforcement-Learning-Based Multitarget Coverage With Connectivity Guaranteed
abstract
Deriving a distributed, time-efficient, and connectivity-guaranteed coverage policy in multitarget environment poses huge challenges for a multirobot team with limited coverage and limited communication. In particular, the robot team needs to cover multiple targets while preserving connectivity. In this article, a novel deep-reinforcement-learning-based approach is proposed to take both multitarget coverage and connectivity preservation into account simultaneously, which consists of four parts: a hierarchical observation attention representation, an interaction attention representation, a two-stage policy learning, and a connectivity-guaranteed policy filtering. The hierarchical observation attention representation is designed for each robot to extract the latent features of the relations from its neighboring robots and the targets. To promote the cooperation behavior among the robots, the interaction attention representation is designed for each robot to aggregate information from its neighboring robots. Moreover, to speed up the training process and improve the performance of the learned policy, the two-stage policy learning is presented using two reward functions based on algebraic connectivity and coverage rate. Furthermore, the learned policy is filtered to strictly guarantee the connectivity based on a model of connectivity maintenance. Finally, the effectiveness of the proposed method is validated by numerous simulations. Besides, our method is further deployed to an experimental platform based on quadrotor unmanned aerial vehicles and omnidirectional vehicles. The experiments illustrate the practicability of the proposed method.
Shiguang Wu 0001, Zhiqiang Pu, Tenghai Qiu, Jianqiang Yi
IEEE Trans. Ind. Informatics1
2023 Attention Enhanced Reinforcement Learning for Multi agent Cooperation
abstract
In this article, a novel method, called attention enhanced reinforcement learning (AERL), is proposed to address issues including complex interaction, limited communication range, and time-varying communication topology for multi agent cooperation. AERL includes a communication enhanced network (CEN), a graph spatiotemporal long short-term memory network (GST-LSTM), and parameters sharing multi-pseudo critic proximal policy optimization (PS-MPC-PPO). Specifically, CEN based on graph attention mechanism is designed to enlarge the agents' communication range and to deal with complex interaction among the agents. GST-LSTM, which replaces the standard fully connected (FC) operator in LSTM with graph attention operator, is designed to capture the temporal dependence while maintaining the spatial structure learned by CEN. PS-MPC-PPO, which extends proximal policy optimization (PPO) in multi agent systems with parameters' sharing to scale to environments with a large number of agents in training, is designed with multi-pseudo critics to mitigate the bias problem in training and accelerate the convergence process. Simulation results for three groups of representative scenarios including formation control, group containment, and predator-prey games demonstrate the effectiveness and robustness of AERL.
Zhiqiang Pu, Huimu Wang, Zhen Liu 0020, Jianqiang Yi, Shiguang Wu 0001
IEEE Trans. Neural Networks Learn. Syst.5
2022 Concentration Network for Reinforcement Learning of Large-Scale Multi-Agent Systems
abstract
When dealing with a series of imminent issues, humans can naturally concentrate on a subset of these concerning issues by prioritizing them according to their contributions to motivational indices, e.g., the probability of winning a game. This idea of concentration offers insights into reinforcement learning of sophisticated Large-scale Multi-Agent Systems (LMAS) participated by hundreds of agents. In such an LMAS, each agent receives a long series of entity observations at each step, which can overwhelm existing aggregation networks such as graph attention networks and cause inefficiency. In this paper, we propose a concentration network called ConcNet. First, ConcNet scores the observed entities considering several motivational indices, e.g., expected survival time and state value of the agents, and then ranks, prunes, and aggregates the encodings of observed entities to extract features. Second, distinct from the well-known attention mechanism, ConcNet has a unique motivational subnetwork to explicitly consider the motivational indices when scoring the observed entities. Furthermore, we present a concentration policy gradient architecture that can learn effective policies in LMAS from scratch. Extensive experiments demonstrate that the presented architecture has excellent scalability and flexibility, and significantly outperforms existing methods on LMAS benchmarks.
Qingxu Fu, Tenghai Qiu, Jianqiang Yi, Zhiqiang Pu, Shiguang Wu 0001
AAAI5
2022 Commander-Soldiers Reinforcement Learning for Cooperative Multi-Agent Systems
abstract
In ball sports, such as basketball, the coach can guide players to better offend and defend from a holistic perspective to win the game. Inspired by such scenarios, we introduce a coach-like concept into the decision-making process of cooperative multi-agent systems. We propose a new framework Commander-Soldiers Reinforcement Learning (CSRL), for Multi-Agent systems. Specifically, we introduce a virtual role, Commander, which can obtain and encode global information every T steps and send the encoded global guidance to Soldiers (real agents). Furthermore, we propose Policy Guidance Network (PGN), which can customize the encoded global guidance from Commander based on observations for each Soldier, providing each Soldier with specified guidance to the decision-making process. The Soldier takes into account not only the local action-observation histories but also the specified guidance from PGN when making decisions. We validate CSRL on the challenging StarCraft II micromanagement benchmark, proving that our approach can take advantage of intermittent global information to improve collaborative performance.
Yiqun Chen 0004, Wei Yang 0041, Shiguang Wu 0001, Hongxing Chang
IJCNN4
2022 Multi-Agent Local Information Reconstruction for Situational Cognition
abstract
Learning an effective strategy is challenging for agents in partially observable environment, where the agents can only observe a part of environment information and make decisions based on local information. Hence, how to effectively utilize local information to achieve efficient cooperation among the agents is particularly important. The agents can establish an understanding of themselves and their surrounding environment based on their historical observation information. However, the understanding lacking global information is local or limited, resulting in low performance in some complex tasks. To solve the problem, a situational cognition learning framework is proposed for each agent based on local information, with which each agent can reconstruct a cognition about itself and its surrounding environment and map it into a high dimensional representation space. In particular, situational cognition is modeled as a random variable under the condition of local trajectory. In addition, an information regularizer is introduced to ensure that the situational cognition is complete and accurate through maximizing the mutual information between the situational cognition and global information, conditioned on the local trajectory of the agent. Various simulations are conducted and show that the proposed framework significantly promotes cooperation among the agents and improves performance.
Shiguang Wu 0001, Zhiqiang Pu, Tenghai Qiu, Jianqiang Yi
IJCNN1
2022 Intrinsic Reward with Peer Incentives for Cooperative Multi-Agent Reinforcement Learning
abstract
In this paper, we propose a novel Intrinsic Reward method with Peer Incentives (IRPI) to promote the inter-agent direct interactions and implicitly address the credit assignment problem in cooperative multi-agent reinforcement learning (MARL). The IRPI method can build mutual incentives between agents by using their causal effect, to realize their advanced cooperation. Specifically, a new intrinsic reward mechanism is conducted, which equips each agent with the ability to reward other agent by using the causal effect between them. Moreover, the mechanism is built through a neural network and learned by using causal effect between the agents. Furthermore, the counterfactual reasoning is used to infer the causal effect between the agents using the joint action-state value function, and then assess the quality of the effect using individual state value function in MARL. Simulational results in Starcraft II Micromanagement demonstrate that the proposed IRPI can enhance cooperation among the RL agents to achieve better performance than some state-of-the-art MARL methods in various cooperative multi-aaent tasks.
Zhen Liu 0020, Shiguang Wu 0001, Zhiqiang Pu, Jianqiang Yi
IJCNN3
2021 Multi-target Coverage with Connectivity Maintenance using Knowledge-incorporated Policy Framework
abstract
This paper considers a multi-target coverage problem where a robot team aims to efficiently cover multi-targets while maintaining connectivity in a distributed manner. A novel knowledge-incorporated policy framework is proposed to derive a distributed, efficient, and connectivity guaranteed coverage policy. In particular, a knowledge-guided policy network (KGPnet) is designed, which consists of observation attention representation, interaction attention representation, and knowledge-guided policy learning. Giving credit to the KGPnet, the connectivity guaranteed coverage policy can be applied to different number targets. Moreover, based on the knowledge of the algebraic connectivity and coverage rate, a comprehensive reward is designed to guide the training of the behavior of multi-target coverage with connectivity maintenance. Furthermore, since the policy learned through deep reinforcement learning (DRL) can not guarantee the connectivity of the robot team, a knowledge-nested policy filtering is designed to filter dis-connectivity policies to satisfy the connectivity constraint based on the knowledge model of connectivity maintenance. Various simulations are conducted to verify the effectiveness of the proposed method. Besides, numerous real-world experiments with three-wheel omnidirectional cars and a motion capture system are presented to demonstrate the practicability of the proposed method.
Shiguang Wu 0001, Zhiqiang Pu, Zhen Liu 0020, Tenghai Qiu, Jianqiang Yi
ICRA1
2021 Multi-agent Collaborative Learning with Relational Graph Reasoning in Adversarial Environments
abstract
This paper proposes a collaborative policy framework via relational graph reasoning for multi-agent systems to accomplish adversarial tasks. A relational graph reasoning module consisting of an agent graph reasoning module and an opponent graph module, is designed to enable each agent to learn mixture state representation to enhance the effectiveness of the policy. In particular, for each agent, the agent graph reasoning module is designed to infer different underlying influences from different opponents and generate agent-level state representation. The opponent graph reasoning module is creatively designed for the opponents to reason relations from their surrounding objects including the agents and the opponents based on their latent features and then predict the future state of the opponents. It forms an opponent-level state representation. Besides, in order to effectively predict the state of the opponents, an intrinsic reward based on prediction error is designed to motivate the policy learning. Furthermore, interactions among agents are utilized to transmit messages and fuse information to promote the cooperative behaviors among the agents. Finally, various representative simulations on two multi-agent adversarial tasks are conducted to demonstrate the superiority and effectiveness of the proposed framework by comparison with existing methods.
Shiguang Wu 0001, Tenghai Qiu, Zhiqiang Pu, Jianqiang Yi
IROS1
2021 Formation Control With Collision Avoidance Through Deep Reinforcement Learning Using Model-Guided Demonstration
abstract
Generating collision-free, time-efficient paths in an uncertain dynamic environment poses huge challenges for the formation control with collision avoidance (FCCA) problem in a leader-follower structure. In particular, the followers have to take both formation maintenance and collision avoidance into account simultaneously. Unfortunately, most of the existing works are simple combinations of methods dealing with the two problems separately. In this article, a new method based on deep reinforcement learning (RL) is proposed to solve the problem of FCCA. Especially, the learning-based policy is extended to the field of formation control, which involves a two-stage training framework: an imitation learning (IL) and later an RL. In the IL stage, a model-guided method consisting of a consensus theory-based formation controller and an optimal reciprocal collision avoidance strategy is designed to speed up training and increase efficiency. In the RL stage, a compound reward function is presented to guide the training. In addition, we design a formation-oriented network structure to perceive the environment. Long short-term memory is adopted to enable the network structure to perceive the information of obstacles of an uncertain number, and a transfer training approach is adopted to improve the generalization of the network in different scenarios. Numerous representative simulations are conducted, and our method is further deployed to an experimental platform based on a multiomnidirectional-wheeled car system. The effectiveness and practicability of our proposed method are validated through both the simulation and experiment results.
Zezhi Sui, Zhiqiang Pu, Jianqiang Yi, Shiguang Wu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2020 Multi-agent Cooperation and Competition with Two-Level Attention Network
Shiguang Wu 0001, Zhiqiang Pu, Jianqiang Yi, Huimu Wang
ICONIP (2)1
2020 Multi-Robot Cooperative Target Encirclement through Learning Distributed Transferable Policy
abstract
Making efficient motion decisions for a multi-robot system is a challenging problem in target encirclement with collision avoidance. Specifically, each robot with local communication has to consider cooperative target encirclement and collision avoidance simultaneously. In this paper, a distributed transferable policy network framework based on deep reinforcement learning is proposed to solve the problem of multi-robot cooperative target encirclement with collision avoidance. The proposed policy network framework is able to process the information of uncertain number of robots and obstacles, which is a desirable property for multi-robot systems. In particular, graph attention communication mechanism is adopted to model multi-robot interactions as a graph and extract cooperative information from the graph. Long short-term memory is used to accept the states of uncertain number of obstacles. In addition, a compound reward is designed to lead the training of the behavior of target encirclement with collision avoidance. Curriculum learning is implemented to speed up the process of this training. Simulation results validate the effectiveness of the proposed algorithm. Moreover, we further show that the learned policy can directly transfer to different scenarios along with good generalization.
Zhen Liu 0020, Shiguang Wu 0001, Zhiqiang Pu, Jianqiang Yi
IJCNN3