VLDB 2026 Research / reviewers in the wild / expert
Meng Xu 0009
dblp:75/4287-9
· DBLP profile ↗
18ranked-venue papers
15as first author
18since 2021 · last 2026
0000-0003-4857-5439ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 10 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Unified Self-Regulating Training Framework for Federated Deep Reinforcement LearningabstractFederated Deep Reinforcement Learning (FDRL) aims to enable distributed collaborative training of multiple DRL models while preserving privacy. Existing FDRL methods function in static client environments, but real-world scenarios often involve dynamic state transitions, such as noise, which render static model topologies inadequate and result in biased policy loss. This degrades client performance and leads to suboptimal global policies. To address this challenge, we develop a generic solution, referred to as the self-regulating training framework, which can be seamlessly integrated into existing FDRL approaches to address dynamic state transitions. Specifically, we propose a Sparse Training (ST) method that dynamically sparsifies and adjusts the topology of each model during training to maximize model performance and reduce model complexity. Additionally, we introduce an auxiliary model to adaptively regulate the policy loss of client models, mitigating loss bias and facilitating updates that yield improved returns. Experimental results demonstrate that our method enhances six state-of-the-art (SOTA) FDRL approaches across nine tasks in terms of return. Meng Xu 0009, Xinhong Chen 0003, Zhongying Chen, Guanyi Zhao, Jianping Wang 0001 |
AAAI | 1 |
| 2026 | A Unified Experience Replay Framework for Spiking Deep Reinforcement LearningabstractDeep Reinforcement Learning (DRL) methods have shown remarkable success in many applications, yet their high energy consumption limits their practicability. Recent studies incorporated energy-efficient Spiking Neural Networks (SNNs) to build Spiking DRL methods and lower energy consumption by setting a shorter simulation duration for SNNs to compute fewer gradients. However, these existing Spiking DRL methods fail to sample sufficient high-quality samples within a fixed-size replay buffer and perform poorly when the simulation duration is small, introducing the challenging tradeoff between energy consumption and model performance. Motivated by such observations, we develop a generic resilient experience replay method that can be seamlessly integrated into existing spiking DRL methods to effectively address the above tradeoff. Specifically, we allow the replay buffer to dynamically expand as the number of training samples increases, thereby accommodating more potentially valuable candidate samples for policy training. Meanwhile, we introduce an adaptive approach to manage the buffer size by determining when to shrink the replay buffer and removing redundant samples automatically. This strategy prevents the buffer from expanding unnecessarily, thereby mitigating the potential negative impact on model performance. Extensive experimental results demonstrate that our approach significantly enhances the performance of five state-of-the-art (SOTA) spiking DRL methods across various simulation durations in sixteen tasks, in terms of return, without compromising their energy efficiency. Meng Xu 0009, Xinhong Chen 0003, Bingyi Liu, Yi-Rong Lin, Yung-Hui Li, Jianping Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | A Generic Competitive-Cooperative Actor-Critic Framework for Deep Reinforcement LearningabstractIn the field of Deep reinforcement learning (DRL), enhancing exploration capabilities and improving the accuracy of Q-value estimation remain two major challenges. Recently, double-actor DRL methods have emerged as a promising class of DRL approaches, achieving substantial advancements in both exploration and Q-value estimation. However, existing double-actor DRL methods feature actors that operate independently in exploring the environment, lacking mutual learning and collaboration, which leads to suboptimal policies. To address this challenge, this work proposes a generic solution that can be seamlessly integrated into existing double-actor DRL methods by promoting mutual learning among the actors to develop improved policies. Specifically, we calculate the difference in actions output by the actors and minimize this difference as a loss during training to facilitate mutual imitation among the actors. Simultaneously, we also minimize the differences in Q-values output by the various critics as part of the loss, thereby avoiding significant discrepancies in value estimation for the imitated actions. We present two specific implementations of our method and extend these implementations beyond double-actor DRL methods to other DRL approaches to encourage broader adoption. Experimental results demonstrate that our method significantly improves twenty state-of-the-art (SOTA) DRL methods, including SOTA double-actor DRL methods, across eleven tasks, as measured by return and other metrics. Meng Xu 0009, Xinhong Chen 0003, Guanyi Zhao, Jin Huang 0002, Jianping Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Risk-Aware Reinforcement Learning with Group Opinion for Autonomous DrivingabstractTo avoid dangerous situations, such as collisions in dynamic environments, autonomous vehicles must predict the risks of the current scene to take safe actions. Traditional rule-based risk prediction methods and existing reinforcement learning (RL) approaches, which typically rely on manually designed driving decision rules or heuristic reward functions, often fail to capture the complexity of real-world dangerous scenarios, leading to suboptimal and unsafe driving decisions. To address this limitation, we develop a novel RL method, called Group Opinion Risk-Aware Reinforcement Learning (GORA-RL), for safer driving decisions that align with real-world conditions. Specifically, we first introduce surveys of human drivers to assess risk in real-world driving situations. Using these real group opinions as training data, we train a risk prediction model, referred to as the risk prediction model with a Transformer (RPT), that captures the crucial characteristics of these scenarios, resulting in more realistic and reliable risk predictions. This model is then integrated as a reward function to train an RL algorithm for making driving decisions in various scenarios. The experiments validate that our approach outperforms two state-of-the-art (SOTA) methods in challenging congested scenarios, such as merging and intersections, in terms of reward and several other metrics. Project site: https://github.com/naiyisiji/RPT. Guanyi Zhao, Meng Xu 0009, Jianping Wang 0001 |
IROS | 2 |
| 2025 | Policy Correction and State-Conditioned Action Evaluation for Few-Shot Lifelong Deep Reinforcement LearningabstractLifelong deep reinforcement learning (DRL) approaches are commonly employed to adapt continuously to new tasks without forgetting previously acquired knowledge. While current lifelong DRL methods have shown promising advancements in retaining acquired knowledge, they suffer from significant adaptation efforts (i.e., longer training duration) and suboptimal policy when transferring to a new task that significantly deviates from previously learned tasks, a phenomenon known as the few-shot generalization challenge. In this work, we propose a generic approach that equips existing lifelong DRL methods with the capability of few-shot generalization. First, we employ selective experience reuse by leveraging the experience of encountered states, improving adaptation training for new tasks. Then, a relaxed softmax function is applied to the target Q values to improve the accuracy of evaluated Q values, leading to more optimal policies. Finally, we measure and reduce the discrepancy in data distribution between the policy and off-policy samples, resulting in improved adaptation efficiency. Extensive experiments have been conducted on three typical benchmarks to compare our approach with six representative lifelong DRL methods and two state-of-the-art (SOTA) few-shot DRL methods regarding their training speed, episode return, and average return of all episodes. Experimental results substantiate that our method improves the return of six lifelong DRL methods by at least 25%. Meng Xu 0009, Xinhong Chen 0003, Jianping Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | A Novel Topology Adaptation Strategy for Dynamic Sparse Training in Deep Reinforcement LearningabstractDeep reinforcement learning (DRL) has been widely adopted in various applications, yet it faces practical limitations due to high storage and computational demands. Dynamic sparse training (DST) has recently emerged as a prominent approach to reduce these demands during training and inference phases, but existing DST methods achieve high sparsity levels by sacrificing policy performance as they rely on the absolute magnitude of connections for pruning and randomly generating connections. Addressing this, our study presents a generic method that can be seamlessly integrated into existing DST methods in DRL to enhance their policy performance while preserving their sparsity levels. Specifically, we develop a novel method for calculating the importance of connections within the model. Subsequently, we dynamically adjust the sparse network topology by dropping existing connections and introducing new connections based on their respective importance values. Through validation on eight widely used simulation tasks, our method improves two state-of-the-art (SOTA) DST approaches by up to 70% in episode return and average return across all episodes under various sparsity levels. Meng Xu 0009, Xinhong Chen 0003, Jianping Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | A Two-Stage Selective Experience Replay for Double-Actor Deep Reinforcement LearningabstractDeep reinforcement learning (DRL) has been widely applied to various applications, but improving the exploration and the accuracy of Q-value estimation remain key challenges. Recently, the double-actor architecture has emerged as a promising DRL framework that can enhance both exploration and Q-value estimation. Existing double-actor DRL methods sample from the replay buffer to update the two actors; however, the samples used to update each actor are generated by its previous versions and the other actor, resulting in a different data distribution compared with the current actor being updated, which can negatively impact the actor's update and lead to suboptimal policies. To this end, this work proposes a generic solution that can be seamlessly integrated into existing double-actor DRL methods to mitigate the adverse effects of data distribution differences on actor updates, thereby learning better policies. Specifically, we decompose the updates of double-actor DRL methods into two stages, each of which uses the same sampling approach to train a pair of actor-critic. This sampling approach classifies the samples in the replay buffer into distinct categories using a clustering technique, such as K-means, and subsequently employs the Jensen-Shannon (JS) divergence to evaluate the distributional differences between each sample category and the actor currently being updated. Samples are then prioritized from the categories with smaller distribution differences to the current actor to update it. In this way, we can effectively mitigate the distribution difference between the samples and the current actor being updated. Experiments demonstrate that our method enhances the performance of five state-of-the-art (SOTA) double-actor DRL methods and outperforms eight SOTA single-actor DRL methods across eight tasks. Meng Xu 0009, Xinhong Chen 0003, Jianping Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Neighboring State-Aware Policy for Deep Reinforcement LearningabstractDeep reinforcement learning (DRL) methods, which train a policy to obtain the sequence of actions required to complete a task, have achieved remarkable success across diverse applications. It is a long-standing open issue in the DRL community to make the trained policy gradually approach the theoretically globally optimal policy, and existing research has also explored several challenges, such as exploration-exploitation, to improve the quality of the obtained policy. However, most DRL methods rely solely on the current state for decision-making, leading to short-sightedness and suboptimal learning. To overcome this, we propose a neighboring state-aware policy that enhances existing DRL methods by incorporating a neighboring state sequence in the decision-making process. Specifically, our approach saves multiple past and future states and concatenates them as the neighboring state sequence, along with the current state, and inputs them to the actor to generate an action during the training process. This global perspective, provided by neighboring states, is similar to human decision-making and helps the agent better understand state evolution, leading to improved policy learning. We present two specific implementations of our approach and demonstrate through extensive experiments that it effectively enhances ten representative DRL methods across nine tasks, based on three metrics, including return. Meng Xu 0009, Xinhong Chen 0003, Guanyi Zhao, Jianping Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | A Unified Sparse Training Framework for Lightweight Lifelong Deep Reinforcement LearningabstractLifelong deep reinforcement learning (DRL) methods enable continuous adaptation to new tasks and retention of old knowledge. However, these methods often necessitate large model sizes, leading to substantial computational and storage resource requirements during training and inference. Unfortunately, existing research has not yet provided a lightweight solution to address this issue. This work aims to develop a generic method that can be seamlessly integrated into existing lifelong DRL methods to facilitate their achievement of lightweight models while also yielding higher returns. While sparse training (ST) methods have been extensively used in the DRL community to achieve lightweight models, they exacerbate the issue of catastrophic forgetting and compromise generalization when applied in lifelong DRL. To improve generalization, we develop a gradient optimization method that leverages sharpness-aware minimization (SAM) to smooth the gradient surface of the model without introducing excessive computational complexity. In addition, to alleviate catastrophic forgetting and promote model convergence, we introduce a priority-based approach that samples effective past experiences from the replay buffer. Extensive experiments demonstrate that our approach achieves 90% sparsity in five representative lifelong DRL methods while achieving higher episode return and average return (up to 34% improvement) across all episodes compared to the dense models. Meng Xu 0009, Xinhong Chen 0003, Yi-Rong Lin, Yung-Hui Li, Jianping Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2024 | Progressive Hierarchical Deep Reinforcement Learning for defect wafer test
Meng Xu 0009, Xinhong Chen 0003, Yechao She, Jianping Wang 0001 |
Knowl. Based Syst. | 1 |
| 2024 | Strengthening Cooperative Consensus in Multi-Robot ConfrontationabstractMulti-agent reinforcement learning (MARL) has proven effective in training multi-robot confrontation, such as StarCraft and robot soccer games. However, the current joint action policies utilized in MARL have been unsuccessful in recognizing and preventing actions that often lead to failures on our side. This exacerbates the cooperation dilemma, ultimately resulting in our agents acting independently and being defeated individually by their opponents. To tackle this challenge, we propose a novel joint action policy, referred to as the consensus action policy (CAP). Specifically, CAP records the number of times each joint action has caused our side to fail in the past and computes a cooperation tendency, which is integrated with each agent’sQ-value and Nash bargaining solution to determine a joint action. The cooperation tendency promotes team cooperation by selecting joint actions that have a high tendency of cooperation and avoiding actions that may lead to team failure. Moreover, the proposed CAP policy can be extended to partially observable scenarios by combining it with DeepQnetwork or actor-critic–based methods. We conducted extensive experiments to compare the proposed method with seven existing joint action policies, including four commonly used methods and three state-of-the-art methods, in terms of episode rewards, winning rates, and other metrics. Our results demonstrate that this approach holds great promise for multi-robot confrontation scenarios. Meng Xu 0009, Xinhong Chen 0003, Yechao She, Guanyi Zhao, Jianping Wang 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2023 | On-demand Edge Inference Scheduling with Accuracy and Deadline GuaranteeabstractTo meet increasing demands for machine-learning-based applications, pushing inference services to the network edge has been a trend. This work aims to design an on-demand edge inference scheduler with accuracy and deadline guarantee for repetitive tasks. Specifically, we consider an edge server that is preinstalled with multiple early-exit Deep Neural Networks (DNNs), and each DNN-exit pair can provide inference service of different quality. We also consider tasks' diversity in quality of service requirements and related utility. We aim to maximize the system's total utility by optimizing service assignment and time scheduling subject to resource, accuracy, and deadline constraints. We present this problem's integer linear problem formulation and show this problem is NP-hard even for the offline case. This problem is challenging due to the coupled effect of service assignment and time scheduling. To derive low-complexity scheduling solutions, we introduce a task-service graph and convert this problem into a service assignment selection problem with schedulability constraints. Then, we design a polynomial complexity algorithm with$\frac{\rho}{\delta}$-approximation ratio for the offline problem, with$\rho$referring to the task-wise utility ratio,$\delta$referring to the maximum number of concurrent tasks. To handle the online problem, we propose an online heuristic algorithm. Simulation results show that the proposed algorithms outperform the state-of-the-art baseline algorithms. Yechao She, Minming Li, Meng Xu 0009, Jianping Wang 0001, Bin Liu 0001 |
IWQoS | 4 |
| 2023 | A Coupling Approach to Demand Prediction and Repositioning in SAV SystemsabstractIn Shared Autonomous Vehicle (SAV) systems, real-time vehicle repositioning plays a crucial role in meeting time-varying traffic demand, which is normally designed by taking advantage of user demand prediction. Nonetheless, most existing studies only predict traffic demand and schedule SAVs separately, ignoring the tight interaction between the two components, e.g. the potential impact of repositioning results on demand prediction. Such a design lacks a deeply integrated design for both and may lead to inaccurate demand prediction and impaired repositioning performance. To tackle this challenge, we propose DRiVe, a coupling approach to Demand prediction and Repositioning for shared autonomous Vehicle system. Specifically, we consider electric SAVs and adopt model predictive control (MPC) to develop the repositioning strategy with the goal of minimizing the operator’s repositioning costs and passenger dissatisfaction. An online prediction is then introduced which not only implements the traditional demand prediction but also integrates the additional traffic demand generated by repositioning action. The numerical results demonstrate that the proposed DRiVe method achieves better performance in reducing passenger waiting time and idle distance compared to the state-of-the-art repositioning methods. Dongyao Jia, Yechao She, Meng Xu 0009, Shangbo Wang, Jianping Wang 0001 |
VTC Fall | 4 |
| 2023 | Deep Reinforcement Learning for Image-Based Multi-Agent Coverage Path PlanningabstractImage-based Multi-Agent Coverage Path Planning (MACPP) utilizes images as input to control multiple agents touring all nodes in a map, minimizing task duration and node revisiting. State-of-the-art (SOTA) studies have applied Multi-Agent Deep Reinforcement Learning (MADRL) to automate MACPP, primarily focusing on minimizing task duration. However, these approaches overlook the issue of repeated node visits, resulting in longer task durations and limited real-world applicability. To tackle this challenge, we develop a novel MADRL solution, referred to as MADRL with Mask Soft Attention, to minimize task duration and node re-visiting simultaneously. Our method uses mask soft attention to extract key features from raw image observations while masking task-independent features, reducing computational complexity and improving sample efficiency. We also cascade a multi-actor-critic architecture to accommodate even more agents with ease. Each agent is equipped with an actor to learn an action policy, and a shared critic evaluates a state value. To validate our approach, we implement seven SOTA MADRL methods in the MACPP area as baselines. Simulation results show that our method significantly outperforms the baselines regarding task duration and the number of times the node is repeatedly visited. Meng Xu 0009, Yechao She, Jianping Wang 0001 |
VTC Fall | 1 |
| 2023 | Dynamic Weights and Prior Reward in Policy Fusion for Compound Agent LearningabstractIn Deep Reinforcement Learning (DRL) domain, a compound learning task is often decomposed into several sub-tasks in a divide-and-conquer manner, each trained separately and then fused concurrently to achieve the original task, referred to as policy fusion. However, the state-of-the-art (SOTA) policy fusion methods treat the importance of sub-tasks equally throughout the task process, eliminating the possibility of the agent relying on different sub-tasks at various stages. To address this limitation, we propose a generic policy fusion approach, referred to as Policy Fusion Learning with Dynamic Weights and Prior Reward (PFLDWPR), to automate the time-varying selection of sub-tasks. Specifically, PFLDWPR produces a time-varying one-hot vector for sub-tasks to dynamically select a suitable sub-task and mask the rest throughout the entire task process, enabling the fused strategy to optimally guide the agent in executing the compound task. The sub-tasks with the dynamic one-hot vector are then aggregated to obtain the action policy for the original task. Moreover, we collect sub-tasks’s rewards at the pre-training stage as a prior reward, which, along with the current reward, is used to train the policy fusion network. Thus, this approach reduces fusion bias by leveraging prior experience. Experimental results under three popular learning tasks demonstrate that the proposed method significantly improves three SOTA policy fusion methods in terms of task duration, episode reward, and score difference. Meng Xu 0009, Yechao She, Jianping Wang 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2023 | Deep Reinforcement Learning for Parameter Tuning of Robot Visual ServoingabstractRobot visual servoing controls the motion of a robot through real-time visual observations. Kinematics is a key approach to achieving visual servoing. One key challenge of kinematics-based visual servoing is that it requires time-varying parameter configuration throughout the entire process of one task. Parameter tuning is also necessary when applying to different tasks. The existing work on parameter tuning either lacks adaptation or cannot automate the tuning of all parameters. Meanwhile, the transferability of existing methods from one task to another is low. This work develops a Deep Reinforcement Learning (DRL) framework for robot visual servoing, which can automate all parameters tuning for one task and across tasks. In visual servoing, forward kinematics focuses on motion speed, while inverse kinematics focuses on the smoothness of motion. Therefore, we develop two separate modules in the proposed DRL framework. One tunes time-varying Forward Kinematics parameters to accelerate the motion, and the other tunes the Inverse Kinematics parameters to ensure smoothness. Moreover, we customize a knowledge transfer method to generalize the proposed DRL models to various robot tasks without reconstructing the neural network. We verify the proposed method on simulated robot tasks. The experimental results show that the proposed method outperforms the state-of-the-art methods and manual parameter configuration in terms of movement speed and smoothness in one task and across tasks. Meng Xu 0009, Jianping Wang 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2022 | Learning strategy for continuous robot visual control: A multi-objective perspective
Meng Xu 0009, Jianping Wang 0001 |
Knowl. Based Syst. | 1 |
| 2021 | Discounted Sampling Policy Gradient for Robot Multi-objective Visual Control
Meng Xu 0009, Qingfu Zhang 0001, Jianping Wang 0001 |
EMO | 1 |