EDBT 2026 Demo / reviewers in the wild / expert
Chengwei Zhang 0001
dblp:129/0954-1
· DBLP profile ↗
23ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0002-9157-6050ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing traffic signal control through model-based reinforcement learning and policy reuse
Chengwei Zhang 0001, Furui Zhan, Wanting Liu, Kailing Zhou, Longji Zheng |
Expert Syst. Appl. | 2 |
| 2026 | A Two-Stage Framework Based on RL for Truck-Drone Collaborative Delivery ProblemabstractWith the explosive growth of e-commerce, efficient last-mile delivery has emerged as a critical challenge. Truck-drone collaborative delivery has garnered significant attention as a promising solution to this problem. In this work, we formulate the truck-drone collaborative delivery problem as a collaborative optimization problem, aiming to minimize the total completion time for delivering packages to customers by leveraging the complementary strengths of the drone’s speed and the truck’s endurance. We propose a two-stage framework to address this challenge. In the first stage, the Lin-Kernighan Helsgaun (LKH) algorithm is employed to generate a high-quality initial Traveling Salesman Problem (TSP) solution, serving as a robust starting point. In the second stage, a Sequence Allocate Policy (SAPPO), based on Proximal Policy Optimization, refines the TSP solution by optimizing the truck-drone collaborative path using a specially designed action space. Extensive experiments conducted on both random dataset and TSPLIB benchmarks demonstrate that our method significantly outperforms existing algorithms regarding delivery time, while exhibiting improved scalability and less training time. Chengwei Zhang 0001, Wanting Liu, Dou An, Qi Wang 0044 |
IEEE Internet Things J. | 2 |
| 2026 | Structured Sparse Deep Reinforcement Learning for Accelerating Traffic Signal Decision-Making Under Limited Computational Resources
Yujiao Song, Chengwei Zhang 0001, Wanting Liu, Furui Zhan |
IEEE Internet Things J. | 2 |
| 2026 | Using a single actor to output personalized policy for different intersections
Kailing Zhou, Chengwei Zhang 0001, Furui Zhan, Wanting Liu |
Neural Networks | 2 |
| 2026 | Neighbor-Advised Multiagent Reinforcement Learning for Traffic Signal Control in Partially Sensored Networks
Chengwei Zhang 0001, Wanting Liu, Furui Zhan |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | Empowering Non-IID Federated Learning With Data Augmentation and Data-Free Knowledge Distillation
Furui Zhan, Ziyu Deng, Yingxin Liu, Chengwei Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Safe Refined Oil Dispatching via Constrained Multiagent Reinforcement Learning With Hierarchical Action SpacesabstractThe rise in urbanization and the demand for refined oil pose a significant challenge in improving transportation efficiency while ensuring the safety of gas station inventories. Traditional optimization methods, which rely on deterministic models or demand forecasting, enhance efficiency but struggle with uncertainties such as demand fluctuations. Multi-agent reinforcement learning (MARL) presents a promising approach for adaptive cooperative dispatching. This study models the refined oil dispatch problem as a partially observable constrained Markov game (CPOMG), in which agents make cooperative decisions under inventory constraints with limited observation, and proposes Hierarchical Action-Constrained MAPPO (HAC-MAPPO), a novel cooperation MARL algorithm to find the optimal cooperative dispatching policy of the game. Within the CPOMG framework, the HAC-MAPPO agents optimize order generation decisions based on the states of gas station inventories. The subsequent delivery routing for these orders is optimized using an integrated fuel routing tabu search algorithm, which is part of the environment’s state transition logic. HAC-MAPPO incorporates inventory safety thresholds to constrain order generation decisions across fuel types. It employs a dual-critic architecture and a shared hierarchical actor network for efficient decentralized policy learning. Crucially, the Lagrangian relaxation technique is applied to transform the constrained actor optimization objective (for order generation) into an unconstrained form. Evaluations carried out in real-world scenarios (based on actual geographic coordinates) and synthetic scenarios using multiple types of synthetic fuel consumption datasets demonstrate that the proposed model and algorithm significantly reduce inventory violations and improve dispatch efficiency and safety under varying demand patterns, outperforming baseline methods. Chengwei Zhang 0001, Wanting Liu, Qi Wang 0044, Dou An, Furui Zhan |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Vehicle-Level Fairness-Oriented Constrained Multi-Agent Reinforcement Learning for Adaptive Traffic Signal ControlabstractMulti-agent Reinforcement Learning (MARL) has shown considerable promise in enhancing the efficiency of adaptive traffic signal control (ATSC) systems. However, existing MARL approaches primarily focus on optimizing overall traffic flow, often overlooking the issue of fairness in vehicle waiting times. Considering that there is no need to strive for the ultimate fairness, this paper models the ATSC problem as a Constrained Partially Observable Markov Game (CPOMG), where fairness is modeled as a constraint on the maximum waiting time of vehicles on lanes of intersections instead of a reward term that pursues maximization. CPOMG aims to find a cooperative control policy with optimal traffic efficiency within the constrained solution space by multiple agents. On this basis, this paper proposes a new centralized training and decentralized execution cooperative MARL method, i.e., vehicle-level fairness multi-agent proximity policy optimization (VF-MAPPO). VF-MAPPO leverages a centralized trained global Critic Network to estimate the average vehicle traffic efficiency and vehicle maximum waiting time, and an Actor Network shared by all intersections for decentralized execution, which converts the optimization problem with constraints to an unconstrained optimization objective through the Lagrange multiplier method and adopts proximity policy optimization during training. Additionally, VF-MAPPO incorporates spatial-temporal graph attention in the Critic network to efficiently extract state representations in multi-intersection environments. We qualitatively analyzed the monotonic improvement guarantee of VF-MAPPO. Extensive experimental validation across two real-world and one synthetic scenarios substantiates that VF-MAPPO enhances vehicle-level fairness and maintains average traffic efficiency, surpassing state-of-the-art methods. Wanting Liu, Chengwei Zhang 0001, Wanqing Fang, Kailing Zhou, Furui Zhan, Qi Wang 0044, Wanli Xue, Rong Chen 0003 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Sequential Decision MARL for Adaptive Traffic Signal Control With Different Intersections PrioritiesabstractExisting multi-agent reinforcement learning (MARL) in adaptive traffic signal control (ATSC) typically models cooperative control of multiple intersections as a cooperative Markov game, optimizing the average traffic efficiency of intersections with the same emphasis. However, it is insufficient to meet the requirements in real ATSC scenarios when all intersections are treated equally. To this end, this work proposes the Captain-Member Markov Game (CM-MG) that considers the different priorities between intersections. CM-MG categorizes intersections into special and ordinary intersections, controlled by captain agents and member agents to optimize the traffic efficiency of local and overall road networks, respectively. The cooperative requirements of CM-MG are achieved through a sequential decision-making principle. Captains have priority in choosing actions, and members make decisions sequentially, following the breadth-first traversal order in the road network after obtaining their precursors’ intentions. Then, a cooperative MARL algorithm, i.e, Sequential Decision Deep Graph Network (GNSD-Light), is proposed to learn the optimal joint policy that meets the learning goals of both captains and members. To be unrestricted by the scales of intersections, GNSD-Light adopts an autoregressive framework where all agents make decisions sequentially in a predetermined order by sharing the same decision model. In addition, to obtain sufficient state representation, two relative position encoding-based spatiotemporal representation modules are designed for GNSD-Light based on the characteristics of ATSC scenarios. Finally, through adequate experiments and qualitative analysis, we have confirmed that our method effectively balances traffic efficiency among both the overall road network and special intersections. Wanting Liu, Chengwei Zhang 0001, Kailing Zhou, Furui Zhan, Wanli Xue, Rong Chen 0003 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Learning Self-Corrective Network via Adaptive Self-Labeling and Dynamic NMS for High-Performance Long-Term TrackingabstractThis article presents a self-corrective network-based long-term tracker (SCLT) including a self-modulated tracking reliability evaluator (STRE) and a self-adjusting proposal postprocessor (SPPP). The targets in the long-term sequences often suffer from severe appearance variations. Existing long-term trackers often online update their models to adapt the variations, but the inaccurate tracking results introduce cumulative error into the updated model that may cause severe drift issue. To this end, a robust long-term tracker should have the self-corrective capability that can judge whether the tracking result is reliable or not, and then it is able to recapture the target when severe drift happens caused by serious challenges (e.g., full occlusion and out-of-view). To address the first issue, the STRE designs an effective tracking reliability classifier that is built on a modulation subnetwork. The classifier is trained using the samples with pseudo labels generated by an adaptive self-labeling strategy. The adaptive self-labeling can automatically label the hard negative samples that are often neglected in existing trackers according to the statistical characteristics of target state, and the network modulation mechanism can guide the backbone network to learn more discriminative features without extra training data. To address the second issue, after the STRE has been triggered, the SPPP follows it with a dynamic NMS to recapture the target in time and accurately. In addition, the STRE and the SPPP demonstrate good transportability ability, and their performance is improved when combined with multiple baselines. Compared to the commonly used greedy NMS, the proposed dynamic NMS leverages an adaptive strategy to effectively handle the different conditions of in view and out of view, thereby being able to select the most probable object box that is essential to accurately online update the basic tracker. Extensive evaluations on four large-scale and challenging benchmark datasets including VOT2021LT, OxUvALT, TLP, and LaSOT demonstrate superiority of the proposed SCLT to a variety of state-of-the-art long-term trackers in terms of all measures. Source codes and demos can be found at https://github.com/TJUT-CV/SCLT. Wanli Xue, Kaihua Zhang 0001, Bo Liu 0005, Chengwei Zhang 0001, Jingen Liu, Shengyong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Optimistic Exploration Based on Categorical-DQN for Cooperative Markov Games
Chengwei Zhang 0001, Qing Guo 0005, Kangjie Zheng, Wanqing Fang, Xintian Zhao |
DAI | 2 |
| 2022 | Advertising Impression Resource Allocation Strategy with Multi-Level Budget Constraint DQN in Real-Time Bidding
Chengwei Zhang 0001, Kangjie Zheng, Wanli Xue, Tianpei Yang, Dou An, Yongqi Pi, Rong Chen 0003 |
Neurocomputing | 1 |
| 2022 | Data Integrity Attack in Dynamic State Estimation of Smart Grid: Attack Model and CountermeasuresabstractA smart grid integrates advanced sensors, efficient measurement methods, progressive control technologies, and other techniques and devices to achieve safe, efficient and economical operation of the grid system. However, the diversified and open environment of a smart grid makes energy and information of the smart grid vulnerable to malicious attacks. As a representative cyber-physical attack, the data integrity attack has an extremely severe impact on the grid operation for it can bypass the traditional detection mechanisms by adjusting the attack vector. In this paper, we first present the attack strategy against dynamic state estimation of power grid in the perspective of adversary and formulate the data integrity attack detection problem that has the characteristic of sequential decision making as a partially observable Markov decision process. Then, a deep reinforcement learning-based approach is proposed to detect against data integrity attacks, which utilizes the Long Short-Term Memory layer to extract the state features of previous time steps in determining whether the system is currently under attack. Moreover, the noisy networks are employed to ensure effective agent exploration, which prevents the agent from sticking to the non-optimal policy. The principle of a multi-step learning is adopted to increase the estimation accuracy of Q value. To address the sparse rewards problem, the prioritized experience replay is proposed to increase training efficiency. Simulation results demonstrated that the proposed detection approach surpasses the benchmarks in the comparison metrics: delay error rate and false rate.Note to Practitioners—In this paper, we present a deep reinforcement learning-based algorithm to defend against the data integrity attacks of smart grid. Most of the previous works discretized the system states and utilized the current state information to identify whether the system is under attack. For this reason, the detection policy may totally ignored the continuously changing characteristics of the grid states, which will lead to poor detection performance. Moreover, the attacked system states only accounts for a small part of the entire grid operation states, the probability of sampling the experience containing the attack state is extremely small, which limits the learning efficiency of previous RL-based detection approaches. In order to increase the accuracy of detection, we first present the attack strategy against power grid’s dynamic state estimation in the perspective of adversary and formulate the partially observable Markov decision process model of attack detection problem. Moreover, we propose a deep reinforcement learning-based detection approach combining the LSTM network to extract the system state features of the previous time steps to determine whether the system is currently being attacked. To address the sparse rewards problem, the prioritized experience replay is used to increase learning efficiency. The experiments demonstrate the effectiveness of proposed detection scheme compared with benchmarks in terms of detection delay as well as accuracy. In conclusion, the proposed detection scheme is helpful in defending against the data integrity attacks without obtaining the opponent’s strategy in advance and can be conveniently applied to the real-world security management system of smart grid. Dou An, Feiye Zhang, Qingyu Yang 0003, Chengwei Zhang 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2022 | Object-Aware Ghost Identification and Elimination for Dynamic Scene MosaicabstractComposite ghost is a common phenomenon that widely exists in dynamic scene image mosaic and significantly affects the naturalness of mosaic. To remove the ghost effectively and produce visually natural mosaic, we propose a novel image mosaic method by jointly identifying composite ghost and eliminating ghost regions without distorting, splitting, and duplicating objects. Specifically, our main contributions are three-fold:First, we propose themotion-awarecomposite ghost identification to localize the potential composite ghosts in the mosaic region (i.e., overlapping area between two images to be stitched) by detecting the salient-moving objects in two stitched images.Second, we design theobject-awarealternative region selection strategy to produce ghostless regions that can replace the localized composite ghosts while avoiding object distortion, object separation, and object repetition.Third, we realize theimage interpolation-basedcomposite ghost elimination that can generate natural stitched image by eliminating the composite ghost of the initial blending result with the selected image source. We validate the proposed method on challenging datasets and show that our method outperform the state-of-the-art methods. Zhe Zhang 0039, Xuhong Ren, Wanli Xue, Chengwei Zhang 0001, Qing Guo 0005, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Neighborhood Cooperative Multiagent Reinforcement Learning for Adaptive Traffic Signal Control in Epidemic RegionsabstractNowadays, multiagent reinforcement learning (MARL) have shared significant advances in the adaptive traffic signal control (ATSC) problems. For most of the researches, agents are all isomorphic, which disregards the situation in which isomerous intersections cooperative together in a real ATSC scenario, especially in epidemic regions where different intersections have quite different levels of importance. To this end, this paper models the ATSC problem as a networked Markov game (NMG), in which agents take into account information, including traffic conditions of it and its connected neighbors. A cooperative MARL framework named neighborhood cooperative hysteretic DQN (NC-HDQN) is proposed. Specifically, for each NC-HDQN agent in the NMG, first, the framework analyses correlation degrees with their connected neighbors and weighs observations and rewards by these correlations. Second, NC-HDQN agents independently optimize their strategies on the weighted information using hysteretic DQN (HDQN), which is designed to learn optimal joint strategies in cooperative multiagent games. Third, a rule-based NC-HDQN method and a Pearson correlation coefficient based NC-HDQN method, i.e., empirical NC-HDQN (ENC-HDQN) and Pearson NC-HDQN (PNC-HDQN), respectively, are designed. The first method maps the correlation degree between two connected agents according to vehicle numbers on roads between the two agents. In contrast, the second method uses the Pearson correlation coefficient to calculate the correlation degree adaptively. Our methods are empirically evaluated in both a synthetic scenario and two real-world traffic scenarios and give better performances in almost every standard test metric for ATSC. Chengwei Zhang 0001, Wanli Xue, Xiaofei Xie, Tianpei Yang, Rong Chen 0003 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Coordination Between Individual Agents in Multi-Agent Reinforcement LearningabstractThe existing multi-agent reinforcement learning methods (MARL) for determining the coordination between agents focus on either global-level or neighborhood-level coordination between agents. However the problem of coordination between individual agents is remain to be solved. It is crucial for learning an optimal coordinated policy in unknown multi-agent environments to analyze the agent's roles and the correlation between individual agents. To this end, in this paper we propose an agent-level coordination based MARL method. Specifically, it includes two parts in our method. The first is correlation analysis between individual agents based on the Pearson, Spearman, and Kendall correlation coefficients; And the second is an agent-level coordinated training framework where the communication message between weakly correlated agents is dropped out, and a correlation based reward function is built. The proposed method is verified in four mixed cooperative-competitive environments. The experimental results show that the proposed method outperforms the state-of-the-art MARL methods and can measure the correlation between individual agents accurately. Yang Zhang 0097, Qingyu Yang 0003, Dou An, Chengwei Zhang 0001 |
AAAI | 4 |
| 2021 | MARL for Traffic Signal Control in Scenarios with Different Intersection Importance
Liguang Luan, Wanqing Fang, Chengwei Zhang 0001, Wanli Xue, Rong Chen 0003, Chen Sang |
DAI | 4 |
| 2021 | An Efficient Transfer Learning Framework for Multiagent Reinforcement LearningabstractTransfer Learning has shown great potential to enhance single-agent Reinforcement Learning (RL) efficiency. Similarly, Multiagent RL (MARL) can also be accelerated if agents can share knowledge with each other. However, it remains a problem of how an agent should learn from other agents. In this paper, we propose a novel Multiagent Policy Transfer Framework (MAPTF) to improve MARL efficiency. MAPTF learns which agent's policy is the best to reuse for each agent and when to terminate it by modeling multiagent policy transfer as the option learning problem. Furthermore, in practice, the option module can only collect all agent's local experiences for update due to the partial observability of the environment. While in this setting, each agent's experience may be inconsistent with each other, which may cause the inaccuracy and oscillation of the option-value's estimation. Therefore, we propose a novel option learning algorithm, the successor representation option learning to solve it by decoupling the environment dynamics from rewards and learning the option-value under each agent's preference. MAPTF can be easily combined with existing deep RL and MARL approaches, and experimental results show it significantly boosts the performance of existing methods in both discrete and continuous state spaces. Tianpei Yang, Weixun Wang, Hongyao Tang, Jianye Hao, Zhaopeng Meng, Hangyu Mao, Dong Li 0016, Wulong Liu, Yujing Hu, Changjie Fan, Chengwei Zhang 0001 |
NeurIPS | 12 |
| 2019 | SA-IGA: a multiagent reinforcement learning method towards socially optimal outcomes
Chengwei Zhang 0001, Xiaohong Li 0001, Jianye Hao, Siqi Chen 0001, Karl Tuyls, Wanli Xue, Zhiyong Feng 0002 |
Auton. Agents Multi Agent Syst. | 1 |
| 2018 | Achieving Multiagent Coordination Through CALA-rFMQ Learning in Continuous Action Space
Wanshu Liu, Chengwei Zhang 0001, Tianpei Yang, Jianye Hao, Xiaohong Li 0001, Zhijie Bao |
PRICAI | 2 |
| 2017 | Defending Against Man-In-The-Middle Attack in Repeated GamesabstractThe Man-in-the-Middle (MITM) attack has become widespread in networks nowadays. The MITM attack would cause serious information leakage and result in tremendous loss to users. Previous work applies game theory to analyze the MITM attack-defense problem and computes the optimal defense strategy to minimize the total loss. It assumes that all defenders are cooperative and the attacker know defenders' strategies beforehand. However, each individual defender is rational and may not have the incentive to cooperate. Furthermore, the attacker can hardly know defenders' strategies ahead of schedule in practice. To this end, we assume that all defenders are self-interested and model the MITM attack-defense scenario as a simultaneous-move game. Nash equilibrium is adopted as the solution concept which is proved to be always unique. Given the impracticability of computing Nash equilibrium directly, we propose practical adaptive algorithms for the defenders and the attacker to learn towards the unique Nash equilibrium through repeated interactions. Simulation results show that the algorithms are able to converge to Nash equilibrium strategy efficiently. Shuxin Li 0001, Xiaohong Li 0001, Jianye Hao, Bo An 0001, Zhiyong Feng 0002, Kangjie Chen, Chengwei Zhang 0001 |
IJCAI | 7 |
| 2016 | Dynamic analysis of cell interactions in biological environments under multiagent social learning frameworkabstractBiological environment is uncertain and its dynamic is similar to the multiagent environment, thus the research results of the multiagent system area are of great significance and can provide valuable insights to the understanding of biology. Learning in a multiagent environment is highly dynamic since the environment is not stationary anymore and each agent's behavior changes adaptively in response to other coexisting learners, and vice versa. The dynamics becomes more unpredictable when we move from fixed-agent interaction environments to multiagent social learning framework. Analytical understanding of the underlying dynamics is important and challenging. In this work, we consider a social learning framework with homogeneous learners (e.g., Policy Hill Climbing (PHC) learners), to model the behavior of players in the social learning framework as a hybrid dynamical system. By analyzing the dynamical system, we obtain some conditions about convergence or non-convergence. It can be used to predict the convergence of the system. At last, we experimentally verify the predictive power of our model using a number of representative games. Chengwei Zhang 0001, Xiaohong Li 0001, Shuxin Li 0001, Jianye Hao |
BIBM | 1 |
| 2016 | Socially-Aware Multiagent Learning: Towards Socially Optimal OutcomesabstractIn multiagent systems the capability of learning is important for an agent to behave appropriately in face of unknown opponents and a dynamic environment. From the system designer's perspective, it is desirable if the agents can learn to coordinate towards socially optimal outcomes, while also avoiding being exploited by selfish opponents. To this end, we propose a novel gradient ascent based algorithm (SA-IGA) which augments the basic gradient-ascent algorithm by incorporating social awareness into the policy update process. We theoretically analyze the learning dynamics of SA-IGA using dynamical system theory, and SA-IGA is shown to have linear dynamics for a wide range of games including symmetric games. The learning dynamics of two representative games (the prisoner's dilemma game and coordination game) are analyzed in detail. Based on the idea of SA-IGA, we further propose a practical multiagent learning algorithm, called SA-PGA, based on the Q-learning update rule. Simulation results show that an SA-PGA agent can achieve higher social welfare than previous social-optimality oriented Conditional Joint Action Learner (CJAL) and also is robust against individually rational opponents by reaching Nash equilibrium solutions. Xiaohong Li 0001, Chengwei Zhang 0001, Jianye Hao, Karl Tuyls, Siqi Chen 0001, Zhiyong Feng 0002 |
ECAI | 2 |