Chuxiong Sun

dblp:214/9412 · DBLP profile ↗
← Back
22ranked-venue papers
9as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 7 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Group Causal Policy Optimization for Post-Training Large Language Models
abstract
Recent advances in large language models (LLMs) have broadened their applicability across diverse tasks, yet specialized domains still require targeted post-training. Among existing methods, Group Relative Policy Optimization (GRPO) stands out for its efficiency, leveraging groupwise relative rewards while avoiding costly value function learning. However, GRPO treats candidate responses as independent, overlooking semantic interactions such as complementarity and contradiction. To address this challenge, we first introduce a Structural Causal Model (SCM) that reveals hidden dependencies among candidate responses induced by conditioning on a final integrated output, forming a collider structure. Then, our causal analysis leads to two insights: (1) projecting responses onto a causally-informed subspace improves prediction quality, and (2) this projection yields a better baseline than query-only conditioning. Building on these insights, we propose Group Causal Policy Optimization (GCPO), which integrates causal structure into optimization through two key components: a causally-informed reward adjustment and a novel KL-regularization term that aligns the policy with a causally-projected reference distribution. Comprehensive experimental evaluations on various benchmarks demonstrate that GCPO consistently surpasses existing methods.
Ziyin Gu, Ran Zuo, Chuxiong Sun, Zeen Song, Changwen Zheng, Wenwen Qiang
AAAI4
2026 HTG-GCL: Leveraging Hierarchical Topological Granularity from Cellular Complexes for Graph Contrastive Learning
abstract
Graph contrastive learning (GCL) aims to learn discriminative semantic invariance by contrasting different views of the same graph that share critical topological patterns. However, existing GCL approaches with structural augmentations often struggle to identify task-relevant topological structures, let alone adapt to the varying coarse-to-fine topological granularities required across different downstream tasks. To remedy this issue, we introduce Hierarchical Topological Granularity Graph Contrastive Learning (HTG-GCL), a novel framework that leverages transformations of the same graph to generate multi-scale ring-based cellular complexes, embodying the concept of topological granularity, thereby generating diverse topological views. Recognizing that a certain granularity may contain misleading semantics, we propose a multi-granularity decoupled contrast and apply a granularity-specific weighting mechanism based on uncertainty estimation. Comprehensive experiments on various benchmarks demonstrate the effectiveness of HTG-GCL, highlighting its superior performance in capturing meaningful graph representations through hierarchical topological information.
Qirui Ji, Bin Qin 0001, Yunze Zhao, Chuxiong Sun, Changwen Zheng, Jianwen Cao 0001, Jiangmeng Li
AAAI5
2026 M2I2: Learning Efficient Multi-Agent Communication via Masked State Modeling and Intention Inference
abstract
Communication is essential in coordinating the behaviors of multiple agents. However, existing methods primarily emphasize content, timing, and partners for information sharing, often neglecting the critical aspect of integrating shared information. This gap can significantly impact agents' ability to understand and respond to complex, uncertain interactions, thus affecting overall communication efficiency. To address this issue, we introduce M2I2, a novel framework designed to enhance the agents' capabilities to assimilate and utilize received information effectively. M2I2 equips agents with advanced capabilities for masked state modeling and joint-action prediction, enriching their perception of environmental uncertainties and facilitating the anticipation of teammates' intentions. This approach ensures that agents are furnished with both comprehensive and relevant information, bolstering more informed and synergistic behaviors. Moreover, we propose a Dimensional Rational Network, innovatively trained via a meta-learning paradigm, to identify the importance of dimensional pieces of information, evaluating their contributions to decision-making and auxiliary tasks. Then, we implement an importance-based heuristic for selective information masking and sharing. This strategy optimizes the efficiency of masked state modeling and the rationale behind information sharing. We evaluate M2I2 across diverse multi-agent tasks, the results demonstrate its superior performance, efficiency, and generalization capabilities, over existing state-of-the-art methods in various complex scenarios.
Chuxiong Sun, Qirui Ji, Zehua Zang, Jiangmeng Li, Rui Wang 0079, Wei Wang 0353
AAAI1
2026 TMAE: Learning Targeted Multi-Agent Exploration via Causal Inference
abstract
Exploration in sparse-reward tasks remains a fundamental challenge in multi-agent reinforcement learning (MARL) due to complex inter-agent interactions and the expansive exploration space. To address this issue, we propose Targeted Multi-Agent Exploration (TMAE), a novel framework that uncovers the causal relationships between the state space and the reward function, thereby reducing the exploration space and enabling more targeted exploration. Specifically, we construct a structural causal model (SCM) to model the causality between sub-state variables and sparse rewards, providing a robust analytical foundation for subsequent causal inference. Through counterfactual causal intervention, TMAE identifies the most critical subspaces for discovering rare but pivotal events while filtering out confounders. By incorporating these causal insights into the exploration process, TMAE prioritizes subspaces with stronger causal effects on sparse rewards, significantly enhancing exploration efficiency. We evaluate TMAE on a range of MARL benchmarks featuring sparse rewards, consistently demonstrating superior exploration efficiency compared to state-of-the-art methods. Furthermore, visualized causal insights derived from TMAE reveal its ability to effectively capture intricate dependencies and priorities in targeted exploration, showcasing strong alignment with prior domain knowledge.
Chuxiong Sun, Dunqi Yao, Rui Wang 0079, Wenwen Qiang, Changwen Zheng, Jiangmeng Li
AAAI1
2026 Visual reinforcement learning via sequential consistency preserved policy contrast from optimal transport view
Zehua Zang, Jiangmeng Li, Chuxiong Sun, Rui Wang 0079, Fuchun Sun 0001
Neural Networks3
2025 Hierarchical Cognitive Graph Autoencoder for Multi-Agent Reinforcement Learning
Wei Wang 0353, Chuxiong Sun, Yi Wang 0013
CogSci3
2025 TFS: Revisiting Temporal Language Grounding from Frequency Spiking Perspective
abstract
Temporal Language Grounding (TLG) aims to localize moments in untrimmed videos that are most relevant to natural language queries. While existing weakly-supervised methods have achieved significant success in exploring cross-modal relationships, they still face a critical bottleneck: the interference of task-irrelevant information in query embeddings. To address this issue, we propose TLG Frequency Spiking (TFS), a dimensional mask derived from the frequency domain that models the varying importance specific to different queries. By enhancing the understanding of queries, TFS effectively optimizes the cross-modal alignment of visual and textual modalities. Experimental results show that TFS significantly outperforms state-of-the-art baselines on both the Charades-STA and ActivityNet-Captions datasets.
Hongzhou Wu, Chuxiong Sun
ICASSP4
2025 Revisiting Communication Efficiency in Multi-Agent Reinforcement Learning from the Dimensional Analysis Perspective
Chuxiong Sun, Rui Wang 0079, Changwen Zheng
AAMAS1
2025 Loss of Plasticity: A New Perspective on Solving Multi-Agent Exploration for Sparse Reward Tasks
Zehua Zang, Chuxiong Sun, Fuchun Sun 0001, Changwen Zheng
AAMAS2
2025 Token-Level Accept or Reject: A Micro Alignment Approach for Large Language Models
abstract
With the rapid development of Large Language Models (LLMs), aligning these models with human preferences and values is critical to ensuring ethical and safe applications. However, existing alignment techniques such as RLHF or DPO often require direct fine-tuning on LLMs with billions of parameters, resulting in substantial computational costs and inefficiencies. To address this, we propose Micro token-level Accept-Reject Aligning (MARA) approach designed to operate independently of the language models. MARA simplifies the alignment process by decomposing sentence-level preference learning into token-level binary classification, where a compact three-layer fully-connected network determines whether candidate tokens are “Accepted” or “Rejected” as part of the response. Extensive experiments across seven different LLMs and three open-source datasets show that MARA achieves significant improvements in alignment performance while reducing computational costs. The source code and implementation details are publicly available at https://github.com/IAAR-Shanghai/MARA, and the trained models are released at https://huggingface.co/IAAR-Shanghai/MARA_AGENTS.
Yang Zhang 0072, Yu Yu 0008, Bo Tang 0011, Chuxiong Sun, Wenqiang Wei, Jie Hu 0025, Zipeng Xie, Feiyu Xiong, Edward Chung 0001
IJCAI5
2025 Rethinking Generalizability and Discriminability of Self-Supervised Learning from Evolutionary Game Theory Perspective
Jiangmeng Li, Zehua Zang, Qirui Ji, Chuxiong Sun, Wenwen Qiang, Junge Zhang, Changwen Zheng, Fuchun Sun 0001, Hui Xiong 0001
Int. J. Comput. Vis.4
2025 Supporting vision-language model few-shot inference with confounder-pruned knowledge prompt
Jiangmeng Li, Wenyi Mo, Chuxiong Sun, Wenwen Qiang, Bing Su 0001, Changwen Zheng
Neural Networks4
2025 Scalable and Adaptive Graph Neural Networks with Self-Label-Enhanced Training
Chuxiong Sun, Jie Hu 0025, Hongming Gu, Jinpeng Chen 0001
Pattern Recognit.1
2024 T2MAC: Targeted and Trusted Multi-Agent Communication through Selective Engagement and Evidence-Driven Integration
abstract
Communication stands as a potent mechanism to harmonize the behaviors of multiple agents. However, existing work primarily concentrates on broadcast communication, which not only lacks practicality, but also leads to information redundancy. This surplus, one-fits-all information could adversely impact the communication efficiency. Furthermore, existing works often resort to basic mechanisms to integrate observed and received information, impairing the learning process. To tackle these difficulties, we propose Targeted and Trusted Multi-Agent Communication (T2MAC), a straightforward yet effective method that enables agents to learn selective engagement and evidence-driven integration. With T2MAC, agents have the capability to craft individualized messages, pinpoint ideal communication windows, and engage with reliable partners, thereby refining communication efficiency. Following the reception of messages, the agents integrate information observed and received from different sources at an evidence level. This process enables agents to collectively use evidence garnered from multiple perspectives, fostering trusted and cooperative behaviors. We evaluate our method on a diverse set of cooperative multi-agent tasks, with varying difficulties, involving different scales and ranging from Hallway, MPE to SMAC. The experiments indicate that the proposed model not only surpasses the state-of-the-art methods in terms of cooperative performance and communication efficiency, but also exhibits impressive generalization.
Chuxiong Sun, Zehua Zang, Jiangmeng Li, Rui Wang 0079, Changwen Zheng
AAAI1
2024 Demo:SCDRL: Scalable and Customized Distributed Reinforcement Learning System
abstract
Reinforcement Learning (RL) has marked significant achievements across a variety of complex tasks in real-world scenarios. However, the efficacy of RL predominantly relies on the availability of extensive datasets and considerable training resources. Hence, there is the critical need for a distributed system capable of generating and processing vast amounts of data with efficiency. In this work, we introduce a Scalable and Customized Distributed Reinforcement Learning system (SCDRL). Concretely, we analyze the paradigm of RL and decouple the major RL computations into three main aspects, i.e. environment simulation, policy inference and policy training. Such decouple enables SCDRL to efficiently allocate computing resources (be it CPUs or GPUs of varying computational capabilities) tailored to the specific needs of each component. We demonstrate the effectiveness of our method across several key RL environments, demonstrating that our system not only achieves significant learning outcomes and enhanced throughput but also utilizes computing resources with greater efficiency. Notably, our findings reveal SCDRL's proficiency in optimizing resource use not just in single-machine setups but also in multi-machine configurations, all the while maintaining data efficiency and resource utilization.
Chuxiong Sun, Wenwen Qiang, Jiangmeng Li
ICDCS1
2024 CMIX: Causal Value Decomposition for Cooperative Multi-Agent Reinforcement Learning
abstract
Value decomposition plays a pivotal role in ensuring effective credit assignment within Multi-Agent Reinforcement Learning (MARL), particularly in cooperative multi-agent tasks where agents are limited to accessing team rewards only. However, existing methods treat the mixing network as a black box, implicitly assuming that neural networks can autonomously extract important information and achieve rational credit assignment during policy learning. This approach not only lacks interpretability but may also prove inefficient in complex scenarios. To enhance the interpretability and rationality of value decomposition, we propose an innovative approach called “Causal Value Decomposition”(CMIX). CMIX employs causal inference-based models, introducing a set of metrics beyond environmental rewards to enhance robustness and model interpretability. Specifically, CMIX establishes intricate relational structures among agents in complex environments and leverages causal relationships between agents and their surroundings to address the credit assignment challenge in MARL. By employing do-calculus, CMIX accurately measures the impact of each agent's actions on environmental states, precisely determining their contribution to the collective reward. This approach not only enhances the interpretability of existing black-box models but also improves the accuracy of credit assignment in multi-agent systems. Moreover, CMIX exhibits high scalability and complements existing value decomposition techniques. Its effectiveness and scalability have been rigorously tested across various settings, including MPE, LBF, and SMAC environments.
Dunqi Yao, Chuxiong Sun, Kai Li 0047, Kaijie Zhou, Rui Wang 0079
SMC2
2021 LFMAC: Low-Frequency Multi-Agent Communication
Cong Cong 0004, Chuxiong Sun, Rui Wang 0079
ICONIP (5)2
2021 Reward Space Noise for Exploration in Deep Reinforcement Learning
abstract
A fundamental challenge for reinforcement learning (RL) is how to achieve efficient exploration in initially unknown environments. Most state-of-the-art RL algorithms leverage action space noise to drive exploration. The classical strategies are computationally efficient and straightforward to implement. However, these methods may fail to perform effectively in complex environments. To address this issue, we propose a novel strategy named reward space noise (RSN) for farsighted and consistent exploration in RL. By introducing the stochasticity from reward space, we are able to change agent’s understanding about environment and perturb its behaviors. We find that the simple RSN can achieve consistent exploration and scale to complex domains without intensive computational cost. To demonstrate the effectiveness and scalability of the proposed method, we implement a deep Q-learning agent with reward noise and evaluate its exploratory performance on a set of Atari games which are challenging for the naive [Formula: see text]-greedy strategy. The results show that reward noise outperforms action noise in most games and performs comparably in others. Concretely, we found that in the early training, the best exploratory performance of reward noise is obviously better than action noise, which demonstrates that the reward noise can quickly explore the valuable states and aid in finding the optimal policy. Moreover, the average scores and learning efficiency of reward noise are also higher than action noise through the whole training, which indicates that the reward noise can generate more stable and consistent performance.
Chuxiong Sun, Rui Wang 0079, Qian Li 0003
Int. J. Pattern Recognit. Artif. Intell.1
2020 Learning Effective Value Function Factorization via Attentional Communication
abstract
How to achieve efficient cooperation among agents in partially observed environments remains an overarching problem in multi-agent reinforcement learning (MARL). Value function factorization learning is a promising way as it can efficiently address multi-agent credit assignment problem. However, existing value function factorization methods have been focusing on learning fully decentralized value functions, which are not effective for some complex tasks. To address this limitation, we propose a framework which enhances value function factorization by allowing communication during execution. Communication introduces extra information to help agents understand the complex environment and learn sophisticated factorization. Furthermore, the proposed mechanism of communication differs from existing methods since we additionally design a descriptive key along with the message. By the descriptive key, agents can dynamically measure the importance of different messages and achieve attentional communication. We evaluate our framework on a challenging set of StarCraft II micromanagement tasks, and show that it significantly outperforms existing value function factorization methods.
Xiaoya Yang, Chuxiong Sun, Rui Wang 0079
SMC3
2019 Efficient and Scalable Exploration via Estimation-Error
abstract
Exploring efficiently in complex environments is still a challenging problem in reinforcement learning. Recent exploration algorithms based on "optimism in the face of uncertainty" or intrinsic motivation achieved promising performance in sparse reward settings, but they often rely on additional structures which are hard to build in large scale problems. It renders them impractical and hinders the process of combining with reinforcement learning algorithms. Hence, the most state-of-the-art RL algorithms still use the naive action space noise as exploration strategy. In this paper, we model the uncertainty about environment through agent's ability to estimate the value across state and action space. Then, we parameterize the uncertainty by a neural network and regard it as a reward bonus signal to reward uncertain states. In this way, we generate an end-to-end bonus which can scale to complex environments with less computational cost. In order to prove the effectiveness of our method, we evaluate it on the challenging Atari 2600 games. We observed that our method achieves superior or comparable exploratory performance compared to action space noise in all environments, including environments whose rewards are sparse. The results demonstrate that our exploration method can motivate agent to explore effectively even in complex environments and it generally outperforms the naive action space noise.
Chuxiong Sun, Rui Wang 0079, Ruiying Li
IJCNN1
2019 Cell Segmentation Based on FOPSO Combined With Shape Information Improved Intuitionistic FCM
abstract
Fuzzy c-means (FCM) clustering algorithms have been proved to be effective image segmentation techniques. However, FCM clustering algorithms are sensitive to noises and initialization. They cannot effectively segment cell images with inhomogeneous gray value distributions and complex touching cells. Aiming to overcome these disadvantages, this paper proposes a cell image segmentation algorithm using fractional-order velocity based particle swarm optimization (FOPSO) combined with shape information improved intuitionistic FCM (SI-IFCM) clustering. Iterations are carried out between FOPSO and SI-IFCM to achieve final cell segmentation. Experimental results demonstrate that the proposed algorithm has advantages on cell image segmentation, with the highest recall (90.25%) and lowest false discovery rate (0.28%) compared with the state-of-the-art algorithms.
Xiangzhi Bai, Chuxiong Sun, Changming Sun
IEEE J. Biomed. Health Informatics2
2017 Cell segmentation based on spatial information improved intuitionistic fcm combined with FOPSO
abstract
Fuzzy c-means clustering (FCM) algorithm has been proved to be effective for image segmentation. However, it is sensitive to the noises and initialization. FCM could not effectively segment cell images with inhomogeneity and complicate adhesives. Aimed to overcome these disadvantages, this paper proposes a cell image segmentation algorithm using spatial information improved intuitionistic fuzzy c-means clustering (SI-IFCM) combined with fractional-order velocity based particle swarm optimization (FOPSO). SI-IFCM and FOPSO will iterate alternately with different object functions to obtain the clustering result. Experimental results demonstrate the advantages of our algorithm for cell segmentation comparing with state-of-arts algorithms.
Chuxiong Sun, Xiangzhi Bai
ICIP1