Xingguo Chen

dblp:43/6276 · DBLP profile ↗
← Back
23ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MAPG2: Multiagent Policy Gradient via Potential Game for Multirobot Task Allocation Problems
abstract
Efficient task allocation among multiple UAVs and autonomous robots is critical in modern IoT scenarios. This is typically modeled as a multi-robot task allocation (MRTA) problem, known to be an NP-hard combinatorial optimization problem. Neural sequential modeling combined with reinforcement learning (RL) optimization has emerged as a promising paradigm for solving this problem, owing to its high efficiency during inference. However, most existing methods assume that each robot is capable of performing only a single type of task. The development of sensing technologies has significantly enhanced the functional diversity of robots, thereby challenging the effectiveness and scalability of traditional methods. This paper considers a variant of the MRTA problem, where each robot is capable of handling multiple tasks, and tasks vary in both their types and required resources. To this end, we present a novel game-theoretic multi-agent RL algorithm called multi-agent policy gradient via potential game (MAPG2). The key components of proposed method consist of three parts. Firstly, we utilize graph-based attention model (GAM) to characterize the representations between tasks. Secondly, we formulate the single-step allocation process as a potential game (PG) to guarantee the consistency and soundness of the reward function design. Lastly, our approach sequentially generates allocation strategies through centralized training and decentralized execution (CTDE) framework. Extensive experiments demonstrate that MAPG2achieves a 10% improvement in task completion rate compared to state-of-the-art baselines, validating its effectiveness and robustness.
Shangdong Yang, Hongye Cao, Xingguo Chen, Yansheng Wu, Gongzhi Luo
IEEE Internet Things J.4
2026 Bellman error centering
Xingguo Chen, Jinguo Ye, Shangdong Yang
Neural Networks1
2025 Multi-Agent Reinforcement Learning with Communication-Constrained Priors
abstract
Communication is one of the effective means to improve the learning of cooperative policy in multi-agent systems. However, in most real-world scenarios, lossy communication is a prevalent issue. Existing multi-agent reinforcement learning with communication, due to their limited scalability and robustness, struggles to apply to complex and dynamic real-world environments. To address these challenges, we propose a generalized communication-constrained model to uniformly characterize communication conditions across different scenarios. Based on this, we utilize it as a learning prior to distinguish between lossy and lossless messages for specific scenarios. Additionally, we decouple the impact of lossy and lossless messages on distributed decision-making, drawing on a dual mutual information estimatior, and introduce a communication-constrained multi-agent reinforcement learning framework, quantifying the impact of communication messages into the global reward. Finally, we validate the effectiveness of our approach across several communication-constrained benchmarks.
Guang Yang 0066, Tianpei Yang, Jingwen Qiao, Yanqing Wu, Jing Huo, Xingguo Chen, Yang Gao 0001
NeurIPS6
2025 State Abstraction via Deep Supervised Hash Learning
abstract
State abstraction is a widely used technique in reinforcement learning (RL) that compresses the state space to accelerate learning algorithms. However, designing an effective abstraction function in large-scale or high-dimensional state space problems remains a significant challenge. In this brief, we present a novel state abstraction method based on deep supervised hash learning (DSH) and provide a theoretical analysis of its near-optimal property. Furthermore, by leveraging the DSH-based representation as the optimization objective, we propose a direct and concise optimization method based on the target value. In addition, we construct an auxiliary learning task for state abstraction that can be combined with various RL algorithms. In particular, we apply the DSH-based state abstraction to both deep Q-learning (DQN) and soft actor-critic (SAC). Extensive experiments are conducted on Atari and several classic control benchmarks to evaluate the effectiveness of the DSH-based state abstraction method, showing that our method surpasses existing state abstraction algorithms in performance.
Guang Yang 0066, Jing Huo, Shangdong Yang, Tianyu Ding, Xingguo Chen, Yang Gao 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 Graph Contrastive Learning for Multi-behavior Recommendation
Huihui Wang 0001, Shunmei Meng, Xingguo Chen
ADMA (6)4
2023 Backtracking Exploration for Reinforcement Learning
abstract
Exploration of the behavior policy plays an important role in reinforcement learning as it helps learning algorithms escape local optima. Taking linear value function approximation as an example, exploration directly affects the sampling of states, thereby altering the distribution of states. This distribution is a component of the key matrix, and the magnitude of the smallest eigenvalue of the key matrix is proportional to the convergence speed. However, existing exploration methods are constrained by the MDP chain and require step-by-step backtracking to reach the target policy distribution. This paper breaks the assumption that the action settings of the training environment must be identical to that of the testing environment by introducing state resetting in the training environment and proposes a backtracking exploration algorithm with time window and punishment. This algorithm can be directly combined with existing exploration strategies and value function update rules, and it has the potential to become a new paradigm for the training process in reinforcement learning. Experimental results validate the effectiveness of the proposed algorithm.
Xingguo Chen, Zening Chen, Dingyuanhao Sun, Yang Gao 0001
DAI1
2023 Enhancing OOD Generalization in Offline Reinforcement Learning with Energy-Based Policy Optimization
abstract
Offline Reinforcement Learning (RL) is an important research domain for real-world applications because it can avert expensive and dangerous online exploration. Offline RL is prone to extrapolation errors caused by the distribution shift between offline datasets and states visited by behavior policy. Existing offline RL methods constrain the policy to offline behavior to prevent extrapolation errors. But these methods limit the generalization potential of agents in Out-Of-Distribution (OOD) regions and cannot effectively evaluate OOD generalization behavior. To improve the generalization of the policy in OOD regions while avoiding extrapolation errors, we propose an Energy-Based Policy Optimization (EBPO) method for OOD generalization. An energy function based on the distribution of offline data is proposed for the evaluation of OOD generalization behavior, instead of relying on model discrepancies to constrain the policy. The way of quantifying exploration behavior in terms of energy values can balance the return and risk. To improve the stability of generalization and solve the problem of sparse reward in complex environment, episodic memory is applied to store successful experiences that can improve sample efficiency. Extensive experiments on the D4RL datasets demonstrate that EBPO outperforms the state-of-the-art methods and achieves robust performance on challenging tasks that require OOD generalization.
Hongye Cao, Shangdong Yang, Jing Huo, Xingguo Chen, Yang Gao 0001
ECAI4
2023 Model-Based Offline Policy Optimization with Adversarial Network
abstract
Model-based offline reinforcement learning (RL), which builds a supervised transition model with logging dataset to avoid costly interactions with the online environment, has been a promising approach for offline policy optimization. As the discrepancy between the logging data and online environment may result in a distributional shift problem, many prior works have studied how to build robust transition models conservatively and estimate the model uncertainty accurately. However, the over-conservatism can limit the exploration of the agent, and the uncertainty estimates may be unreliable. In this work, we propose a novel Model-based Offline policy optimization framework with Adversarial Network (MOAN). The key idea is to use adversarial learning to build a transition model with better generalization, where an adversary is introduced to distinguish between in-distribution and out-of-distribution samples. Moreover, the adversary can naturally provide a quantification of the model’s uncertainty with theoretical guarantees. Extensive experiments showed that our approach outperforms existing state-of-the-art baselines on widely studied offline RL benchmarks. It can also generate diverse in-distribution samples, and quantify the uncertainty more accurately.
Junming Yang 0001, Xingguo Chen, Shengyuan Wang 0004, Bolei Zhang
ECAI2
2023 Modified Retrace for Off-Policy Temporal Difference Learning
abstract
Off-policy learning is a key to extend reinforcement learning as it allows to learn a target policy from a different behavior policy that generates the data. However, it is well known as “the deadly triad” when combined with bootstrapping and function approximation. Retrace is an efficient and convergent off-policy algorithm with tabular value functions which employs truncated importance sampling ratios. Unfortunately, Retrace is known to be unstable with linear function approximation. In this paper, we propose modified Retrace to correct the off-policy return, derive a new off-policy temporal difference learning algorithm (TD-MRetrace) with linear function approximation, and obtain a convergence guarantee under standard assumptions. Experimental results on counterexamples and control tasks validate the effectiveness of the proposed algorithm compared with traditional algorithms.
Xingguo Chen, Xingzhou Ma, Guang Yang 0066, Shangdong Yang, Yang Gao 0001
UAI1
2023 Leveraging transition exploratory bonus for efficient exploration in Hard-Transiting reinforcement learning problems
Shangdong Yang, Shaokang Dong, Xingguo Chen
Future Gener. Comput. Syst.4
2023 Online attentive kernel-based temporal difference learning
Xingguo Chen, Guang Yang 0066, Shangdong Yang, Shaokang Dong, Yang Gao 0001
Knowl. Based Syst.1
2021 DHQN: a Stable Approach to Remove Target Network from Deep Q-learning Network
abstract
As the first successful attempt to combine deep neural network and reinforcement learning, Deep Q-learning Network (DQN) draws a lot of attention from reinforcement learning researchers. One of the most important components of DQN is target network, which is used to stabilize learning process. When confront complex network structure, the existence of target network means extra memory resource to preserve the neural network weights and high computing cost to calculate target. Thus, we propose a Deep Hybrid Q-learning Network (DHQN) algorithm, which introduces an alternative approach, Random Hybrid Optimization (RHO), that can simplify DQN and attain a more stable and faster learning without a target network. We illustrate that RHO can decelerate divergence in the classical off-policy counterexample θ → 2θ problem. We also testify the effectiveness of DHQN in several control and Atari domains, which shows DHQN outperforms DQN without a target network and original DQN.
Guang Yang 0006, Di'an Fei, Tian Huang, Qingyun Li, Xingguo Chen
ICTAI6
2020 Multi-Agent Game Abstraction via Graph Attention Neural Network
abstract
In large-scale multi-agent systems, the large number of agents and complex game relationship cause great difficulty for policy learning. Therefore, simplifying the learning process is an important research issue. In many multi-agent systems, the interactions between agents often happen locally, which means that agents neither need to coordinate with all other agents nor need to coordinate with others all the time. Traditional methods attempt to use pre-defined rules to capture the interaction relationship between agents. However, the methods cannot be directly used in a large-scale environment due to the difficulty of transforming the complex interactions between agents into rules. In this paper, we model the relationship between agents by a complete graph and propose a novel game abstraction mechanism based on two-stage attention network (G2ANet), which can indicate whether there is an interaction between two agents and the importance of the interaction. We integrate this detection mechanism into graph neural network-based multi-agent reinforcement learning for conducting game abstraction and propose two novel learning algorithms GA-Comm and GA-AC. We conduct experiments in Traffic Junction and Predator-Prey. The results indicate that the proposed methods can simplify the learning process and meanwhile get better asymptotic performance compared with state-of-the-art algorithms.
Yong Liu 0007, Weixun Wang, Yujing Hu, Jianye Hao, Xingguo Chen, Yang Gao 0001
AAAI5
2019 Accelerating Nash Q-Learning with Graphical Game Representation and Equilibrium Solving
abstract
Traditional Nash Q-learning algorithm generally accepts a fact that agents are tightly coupled, which brings huge computing burden. However, many multi-agent systems in the real world have sparse interactions between agents. In this paper, sparse interactions are divided into two categories: intra-group sparse interactions and inter-group sparse interactions. Previous methods can only deal with one specific type of sparse interactions. Aiming at characterizing the two categories of sparse interactions, we use a novel mathematical model called Markov graphical game. On this basis, graphical game-based Nash Q-learning is proposed to deal with different types of interactions. Experimental results show that our algorithm takes less time per episode and acquires a good policy.
Yunkai Zhuang, Xingguo Chen, Yang Gao 0001, Yujing Hu
ICTAI2
2019 Online learning offloading framework for heterogeneous mobile edge computing system
Jidong Ge, Chifong Wong, Chuanyi Li, Xingguo Chen, Sheng Zhang 0001, Bin Luo 0003, He Zhang 0001, Victor Chang 0001
J. Parallel Distributed Comput.5
2018 Prior Knowledge Guided Gene-Disease Associations Prediction: An Enhanced Inductive Matrix Completion Approach
Lei Chen 0011, Jianyu Pu, Ziwen Yang, Xingguo Chen
PRICAI4
2017 Cost-Sensitive Alternating Direction Method of Multipliers for Large-Scale Classification
Yinghuan Shi, Xingguo Chen, Yang Gao 0001
IDEAL3
2017 The Theory of Modified Rings Game
Yushuang Wu, Xingguo Chen
IDEAL4
2016 Efficient Average Reward Reinforcement Learning Using Constant Shifting Values
abstract
There are two classes of average reward reinforcement learning (RL) algorithms: model-based ones that explicitly maintain MDP models and model-free ones that do not learn such models. Though model-free algorithms are known to be more efficient, they often cannot converge to optimal policies due to the perturbation of parameters. In this paper, a novel model-free algorithm is proposed, which makes use of constant shifting values (CSVs) estimated from prior knowledge. To encourage exploration during the learning process, the algorithm constantly subtracts the CSV from the rewards. A terminating condition is proposed to handle the unboundedness of Q-values caused by such substraction. The convergence of the proposed algorithm is proved under very mild assumptions. Furthermore, linear function approximation is investigated to generalize our method to handle large-scale tasks. Extensive experiments on representative MDPs and the popular game Tetris show that the proposed algorithms significantly outperform the state-of-the-art ones.
Shangdong Yang, Yang Gao 0001, Bo An 0001, Hao Wang 0013, Xingguo Chen
AAAI5
2013 Online Selective Kernel-Based Temporal Difference Learning
abstract
In this paper, an online selective kernel-based temporal difference (OSKTD) learning algorithm is proposed to deal with large scale and/or continuous reinforcement learning problems. OSKTD includes two online procedures: online sparsification and parameter updating for the selective kernel-based value function. A new sparsification method (i.e., a kernel distance-based online sparsification method) is proposed based on selective ensemble learning, which is computationally less complex compared with other sparsification methods. With the proposed sparsification method, the sparsified dictionary of samples is constructed online by checking if a sample needs to be added to the sparsified dictionary. In addition, based on local validity, a selective kernel-based value function is proposed to select the best samples from the sample dictionary for the selective kernel-based value function approximator. The parameters of the selective kernel-based value function are iteratively updated by using the temporal difference (TD) learning algorithm combined with the gradient descent technique. The complexity of the online sparsification procedure in the OSKTD algorithm is O(n). In addition, two typical experiments (Maze and Mountain Car) are used to compare with both traditional and up-to-date O(n) algorithms (GTD, GTD2, and TDC using the kernel-based value function), and the results demonstrate the effectiveness of our proposed algorithm. In the Maze problem, OSKTD converges to an optimal policy and converges faster than both traditional and up-to-date algorithms. In the Mountain Car problem, OSKTD converges, requires less computation time compared with other sparsification methods, gets a better local optima than the traditional algorithms, and converges much faster than the up-to-date algorithms. In addition, OSKTD can reach a competitive ultimate optima compared with the up-to-date algorithms.
Xingguo Chen, Yang Gao 0001, Ruili Wang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2010 RL-DOT: A Reinforcement Learning NPC Team for Playing Domination Games
abstract
In this paper, we describe the design of reinforcement-learning-based domination team (RL-DOT), a nonplayer character (NPC) team for playing Unreal Tournament (UT) Domination games. In RL-DOT, there is a commander NPC and several soldier NPCs. The running process of RL-DOT consists of several decision cycles. In each decision cycle, the commander NPC makes a decision of troop distribution and, according to that decision, sends action orders to other soldier NPCs. Each soldier NPC tries to accomplish its task in a goal-directed way, i.e., decomposing the final ultimate task (attacking or defending a domination point) into basic actions (such as running and shooting) that are directly supported by UT application programming interfaces (APIs). We use a Q-learning-style algorithm to learn the optimal decision-making policy. We carefully choose some opponent policies for our illustrative experiments. In these experiments, RL-DOT shows a distinct learning characteristic, which illustrates its efficiency in playing UT Domination games.
Hao Wang 0013, Yang Gao 0001, Xingguo Chen
IEEE Trans. Comput. Intell. AI Games3
2009 Apply ant colony optimization to Tetris
abstract
Tetris is a falling block game where the player's objective is to arrange a sequence of different shaped tetrominoes smoothly in order to survive. In the intelligence games, agent imitates the real player and chooses the best move based on a linear value function. In this paper, we apply Ant Colony Optimization (ACO) method to learn the weights of the function, trying to search an optimal weight-path in the weight graph. We use dynamic heuristic to prevent premature convergence to local optima. Our experimental result is better than most of traditional reinforcement learning methods.
Xingguo Chen, Hao Wang 0013, Weiwei Wang 0002, Yinghuan Shi, Yang Gao 0001
GECCO1
2007 Improving Transport Layer Performance in Multihop Ad Hoc Networks by Exploiting MAC Layer Information
abstract
The traditional TCP congestion control mechanism encounters a number of new problems and suffers a poor performance when the IEEE 802.11 MAC protocol is used in multihop ad hoc networks. Many of the problems result from medium contention at the MAC layer. In this paper, we first illustrate that severe medium contention and congestion are intimately coupled, and TCP's congestion control algorithm becomes too coarse in its granularity, causing throughput instability and excessively long delay. Further, we illustrate TCP's severe unfairness problem due to the medium contention and the tradeoff between aggregate throughput and fairness. Then, based on the novel use of channel busyness ratio, a more accurate metric to characterize the network utilization and congestion status, we propose a new wireless congestion control protocol (WCCP) to efficiently and fairly support the transport service in multihop ad hoc networks. In this protocol, each forwarding node along a traffic flow exercises the inter-node and intra-node fair resource allocation and determines the MAC layer feedback accordingly. The end-to-end feedback, which is ultimately determined by the bottleneck node along the flow, is carried back to the source to control its sending rate. Extensive simulations show that WCCP significantly outperforms traditional TCP in terms of channel utilization, delay, and fairness, and eliminates the starvation problem
Hong Lin Zhai, Xingguo Chen, Yuguang Fang
IEEE Trans. Wirel. Commun.2