VLDB 2026 Research / reviewers in the wild / expert
Liangjun Ke
dblp:23/3034
· DBLP profile ↗
29ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-2920-0853ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GRDC: A Unified Graph-Driven Framework for Role Discovery and Communication in Multi-Agent Reinforcement LearningabstractEffective coordination in Multi-Agent Reinforcement Learning (MARL) is particularly challenging under partial observability, where agents must reason about potential collaborators using only local information. Existing methods fall into two categories: communication-based approaches that enable message exchange but often fix or misidentify who the collaborators are, and role-based approaches that encourage specialization based on behavioral similarity. However, both lines of work overlook the task‑induced cooperative dependencies that decide which agents should collaborate, leading to miscommunication or role misassignment under partial observability. We introduce GRDC (Graph‑driven Role Discovery and Communication), a unified framework that approximates these dependencies by dynamically constructing local interaction graphs from trajectory embeddings, then uses these graphs to infer roles via prototype matching and to restrict communication to intra‑role agents with attention-based aggregation. Beyond role inference and communication, GRDC maximizes role entropy, decorrelates prototypes, and dynamically prunes redundant ones to obtain structured yet compact role specialization. Experimental results on Predator Prey, Cooperative Navigation, and SMACv2 demonstrate that GRDC consistently outperforms state-of-the-art communication- and role-based baselines, improving coordination efficiency and training stability across tasks. Zihong Gao, Hongjian Liang, Yuanhui Hao, Liangjun Ke |
AAAI | 5 |
| 2026 | Test-driven Reinforcement Learning in Continuous ControlabstractReinforcement learning (RL) has been recognized as a powerful tool for robot control tasks. RL typically employs reward functions to define task objectives and guide agent learning. However, since the reward function serves the dual purpose of defining the optimal goal and guiding learning, it is challenging to design the reward function manually, which often results in a suboptimal task representation. To tackle the reward design challenge in RL, inspired by the satisficing theory, we propose a Test-driven Reinforcement Learning (TdRL) framework. In the TdRL framework, multiple test functions are used to represent the task objective rather than a single reward function. Test functions can be categorized as pass-fail tests and indicative tests, each dedicated to defining the optimal objective and guiding the learning process, respectively, thereby making defining tasks easier. Building upon such a task definition, we first prove that if a trajectory return function assigns higher returns to trajectories closer to the optimal trajectory set, maximum entropy policy optimization based on this return function will yield a policy that is closer to the optimal policy set. Then, we introduce a lexicographic heuristic approach to compare the relative distance relationship between trajectories and the optimal trajectory set for learning the trajectory return function. Furthermore, we develop an algorithm implementation of TdRL. Experimental results on the DeepMind Control Suite benchmark demonstrate that TdRL matches or outperforms handcrafted reward methods in policy training, with greater design simplicity and inherent support for multi-objective optimization. We argue that TdRL offers a novel perspective for representing task objectives, which could be helpful in addressing the reward design challenges in RL applications. Xiuping Wu, Liangjun Ke |
AAAI | 3 |
| 2026 | Improving generalization in visual reinforcement learning through data distribution
Yv Zhao, Hongjian Liang, Liangjun Ke |
Expert Syst. Appl. | 5 |
| 2025 | Scattered data augmentation for generalization in visual reinforcement learning
Shaonan Zhang, Liangjun Ke |
Neurocomputing | 5 |
| 2025 | Role can be beneficial to mean field multi-agent reinforcement learning
Shaonan Zhang, Liangjun Ke |
Inf. Sci. | 4 |
| 2025 | FedSR: Federated Learning for Image Super-Resolution via detail-assisted contrastive learning
Yue Yang 0022, Liangjun Ke |
Knowl. Based Syst. | 3 |
| 2025 | Communication in Multiagent Reinforcement Learning via Counterfactual Message ValueabstractEffective communication is pivotal for successful team collaboration in cooperative multiagent tasks. However, mainly due to the partially observable nature of the environment, indiscriminate message requests among teammates may lead to confusion for individual agents, impeding effective collaboration and diminishing the overall efficiency of the system. Most previous work has employed gates or attention mechanisms to extract relatively important messages. However, these methods often fail to explicitly assess each message’s value, or the process of calculating is intricate and convoluted. This may result in ineffective communication and cause miscoordination in complex scenarios. To address these challenges, we introduce a novel metric named counterfactual message value (CMV), which quantifies each message’s contribution to an agent, enabling the effective elimination of redundant messages. In addition, we present a practical multiagent reinforcement learning (MARL) algorithm, termed CMV communication (CMVC), which could effectively facilitate the learning of agent–agent communication protocols. It predicts the CMVs of teammates for an agent solely based on its local observation, enabling the agent to initiate communication with those exhibiting positive CMVs. To differentiate message impact levels, we design a message aggregator in CMVC that aggregates messages based on their individual CMVs. We evaluate CMVC in various tasks, including cooperative navigation, predator–prey, and the StarCraft multiagent challenge (SMAC). The results indicate that CMVC is prominent in reducing redundant messages and has the ability to learn more advanced collaboration strategies, in contrast to several existing state-of-the-art methods. Zihong Gao, Changsheng Qu, Yuanhui Hao, Liangjun Ke |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2024 | Exploiting Degradation Prior for Personalized Federated Learning in Real-World Image Super-ResolutionabstractIn this study, we introduce a novel personalized federated learning (pFL) framework for real-world image super-resolution (SR) tasks, named \model, aiming to protect data privacy and construct the customized model conditioned on the client-specific degradation process. Concretely, we propose a local update with degradation prior at the client-side, which exploits the client-specific information (i.e., image degradation) learned from the customized degradation prior generator as prior knowledge, and along with input to facilitate an efficient local model update. Additionally, we propose a personalized aggregation policy on the server-side, which determines the uploaded model components based on their function. Instead of aggregating the entire models (e.g., FedAvg), the model components that have common representation (e.g., image texture) are shared with the server to augment the representation capability of the global model, while the remaining part with client-related information (e.g., semantic concept) is kept locally to ensure model personalization. Extensive experiments conducted on real-world image SR benchmarks demonstrate the superiority of our model in terms of image quality and model performance. Notably, p2FedSR can be seamlessly integrated with various prevalent SR methods, including CNN-based and Transformer-based architectures. Yue Yang 0022, Liangjun Ke |
ICMR | 2 |
| 2024 | Adaptive mean field multi-agent reinforcement learning
Xiaoqiang Wang 0003, Liangjun Ke, Gewei Zhang, Dapeng Zhu |
Inf. Sci. | 2 |
| 2023 | Traffic signal control using a cooperative EWMA-based multi-agent reinforcement learning
Zhimin Qiao, Liangjun Ke, Xiaoqiang Wang 0003 |
Appl. Intell. | 2 |
| 2021 | Ordering-Based Causal Discovery with Reinforcement LearningabstractIt is a long-standing question to discover causal relations among a set of variables in many empirical sciences. Recently, Reinforcement Learning (RL) has achieved promising results in causal discovery from observational data. However, searching the space of directed graphs and enforcing acyclicity by implicit penalties tend to be inefficient and restrict the existing RL-based method to small scale problems. In this work, we propose a novel RL-based approach for causal discovery, by incorporating RL into the ordering-based paradigm. Specifically, we formulate the ordering search problem as a multi-step Markov decision process, implement the ordering generating process with an encoder-decoder architecture, and finally use RL to optimize the proposed model based on the reward mechanisms designed for each ordering. A generated ordering would then be processed using variable selection to obtain the final causal graph. We analyze the consistency and computational complexity of the proposed method, and empirically show that a pretrained model can be exploited to accelerate training. Experimental results on both synthetic and real data sets shows that the proposed method achieves a much improved performance over existing RL-based method. Xiaoqiang Wang 0003, Yali Du 0001, Shengyu Zhu 0001, Liangjun Ke, Zhitang Chen, Jianye Hao, Jun Wang 0012 |
IJCAI | 4 |
| 2021 | Adaptive collaborative optimization of traffic network signal timing based on immune-fireworks algorithm and hierarchical strategy
Zhimin Qiao, Liangjun Ke, Gewei Zhang, Xiaoqiang Wang 0003 |
Appl. Intell. | 2 |
| 2021 | Large-Scale Traffic Signal Control Using a Novel Multiagent Reinforcement LearningabstractFinding the optimal signal timing strategy is a difficult task for the problem of large-scale traffic signal control (TSC). Multiagent reinforcement learning (MARL) is a promising method to solve this problem. However, there is still room for improvement in extending to large-scale problems and modeling the behaviors of other agents for each individual agent. In this article, a new MARL, called cooperative double Q -learning (Co-DQL), is proposed, which has several prominent features. It uses a highly scalable independent double Q -learning method based on double estimators and the upper confidence bound (UCB) policy, which can eliminate the over-estimation problem existing in traditional independent Q -learning while ensuring exploration. It uses mean-field approximation to model the interaction among agents, thereby making agents learn a better cooperative strategy. In order to improve the stability and robustness of the learning process, we introduce a new reward allocation mechanism and a local state sharing method. In addition, we analyze the convergence properties of the proposed algorithm. Co-DQL is applied to TSC and tested on various traffic flow scenarios of TSC simulators. The results show that Co-DQL outperforms the state-of-the-art decentralized MARL algorithms in terms of multiple traffic metrics. Xiaoqiang Wang 0003, Liangjun Ke, Zhimin Qiao, Xinghua Chai |
IEEE Trans. Cybern. | 2 |
| 2020 | Optimal Energy-Delay Scheduling for Energy-Harvesting WSNs With Interference Channel via Negatively Correlated SearchabstractNetwork resource allocation is an important issue for designing energy-harvesting wireless sensor networks (EH-WSNs). This article considers the capacity assignment problem in EH-WSNs with the interference channel for fixed data and energy flow topologies. We focus on the optimal data rates, power allocations, and energy transfers, minimizing the total network delay for the network. We first consider a simplified model where the data flow is fixed on each data link and optimizes transmit power at each sensor node for a single energy harvest in a time slot. However, the optimization problem is nonconvex, making it difficult to find the optimal solution. Unlike the most traditional methods that approximate the original optimization problem as a convex optimization problem by considering the relatively high signal-to-interference-plus-noise ratio (SINR), this article aims to directly solve the original nonconvex formulation by employing a powerful evolutionary algorithm, i.e., negatively correlated search (NCS). Then, we investigate the joint optimization problem of capacity and flow for the entire EH-WSNs, and develop a novel multiobjective NCS algorithm (MOEA/D-NCS) to deal with the complicated nonlinear constraints and optimize the data rates, power allocations, and energy transfer simultaneously, so as to minimize the total network delay. The numerical results demonstrate that solving the nonconvex problem with approximated approach is a good alternative for solving the approximated convex problem with accurate optimization approaches; the joint optimization of capacity and flow is a good solution for EH-WSNs; and the scheme of partial transmission for data flow is an advantage in respect of decreasing the network delay. The solution of this article could also be beneficial to other complex optimization problems in the wireless network design. Dongbin Jiao, Peng Yang 0008, Liqun Fu 0001, Liangjun Ke, Ke Tang 0001 |
IEEE Internet Things J. | 4 |
| 2019 | Signal Control of Urban Traffic Network Based on Multi-Agent Architecture and Fireworks AlgorithmabstractThe application of multi-agent technology in urban traffic network control makes the traffic signal control has ability of adaptive adjustment. With the changing of traffic flow in the road, it can adjust the important parameters such as the offset, green ratio and public cycle of signal lights in real time, which can effectively reduce traffic congestion and improve the vehicle capacity of the urban traffic network. In this paper, a three level multi-agent control framework is used. The objective functions are the optimization model of green ratio delay, public cycle and offset delay. Fireworks algorithm is used to solve the modeling optimization problem. The simulation results show that the adaptive traffic network signal control can significantly reduce the total delay time of the traffic network and improve the road utilization rate with the continuous change of traffic flow. Compared with the traditional traffic signal timing control, the adaptive traffic network signal control has great advantages and overcomes the disadvantages of traditional traffic signal control. At the same time, the fireworks algorithm shows significant performance in solving the optimizing model of the traffic network. Zhimin Qiao, Liangjun Ke, Xiaoqiang Wang 0003 |
CEC | 2 |
| 2019 | Optimal Energy-Delay Scheduling for Energy Harvesting WSNs via Negatively Correlated SearchabstractOptimal energy-delay scheduling for capacity assignment problem in energy harvesting wireless sensor networks (EH-WSNs) with interference channel is addressed for fixed data flows and energy topologies. We formulate the optimization problem for a single time slot and multiple time slots, respectively. We focus on the optimal data rates, power allocations and energy transfers for the optimization problem. The objective is to minimize the total network delay. However, the optimization problem is non-convex, making it difficult to find the optimal solution. Unlike the most traditional methods that approximate the original optimization problem as a convex optimization problem by considering the relatively high Signal-to-Interference-plus-Noise Ratio (SINR), this paper aims to directly solve the original non-convex formulation by employing a powerful evolutionary algorithm, i.e., Negatively Correlated Search (NCS). The simulations under both no-energy-transfer scenario and energy-transfer scenario are carried out, demonstrating that solving the non-convex problem with approximated approach is a good alternative to solving the approximated convex problem with accurate optimization approaches. This idea could also be beneficial to other complex optimization problems in the wireless networks design. Dongbin Jiao, Peng Yang 0008, Liqun Fu 0001, Liangjun Ke, Ke Tang 0001 |
ICC | 4 |
| 2019 | An improved fireworks algorithm for the capacitated vehicle routing problem
Weibo Yang, Liangjun Ke |
Frontiers Comput. Sci. | 2 |
| 2016 | Route planning in a new tourist recommender system: A fireworks algorithm based approachabstractWith the development of online tourist, Tourist Recommender System (TRS) has become a hot research topic. This paper presents a TRS and considers a new tourist trip design problem (TTDP), taking into account the compactness of the trip. To solve this problem, fireworks algorithm (FWA) is adopted. As a meta-heuristic method, FWA has been widely used in continuous problems, while TTDP is a discrete optimization problem with various constraints, the key difference between them lies in the definition of distance. In this paper, operators of the conventional FWA are redesigned to handle with TTDP. Experimental results demonstrate the effectiveness of the proposed FWA. Liangjun Ke, Zihao Geng |
CEC | 2 |
| 2016 | Fireworks algorithm for the satellite link scheduling problem in the navigation constellationabstractGlobal navigation satellite system (GNSS) can provide autonomous geo-spatial positioning and time synchronization services for both civil and military uses. Satellite links in GNSS are used to transmit signal for constellation management and other applications. In this work, we focus on solving the satellite link scheduling problem over dynamic satellite network with the aim of minimizing the number of participant ground-based management stations and the cost of communication between satellites in the background of GNSS networking. Firstly, we assume the navigation constellation has finite states and cope the dynamic topology with Finite State Automation method. Secondly, a Fireworks algorithm (FWA) is designed according to the characteristic of the scheduling problem. Finally, the FWA is compared with ant colony optimization (ACO). The performance analysis of different scenarios is given. The study in this paper provides technical reference for the management of future large-scale satellite network. Liangjun Ke, Jisheng Li, Jingqi Huang |
CEC | 2 |
| 2015 | Fireworks algorithm for the multi-satellite control resource scheduling problemabstractIn this study, fireworks algorithm (FWA) for the multi-satellite control resource scheduling problem (MSCRSP) is presented. FWA is a meta-heuristic method and widely used in continuous problems while MSCRSP is a constrained and large scale combinatorial problem. The key points of FWA are to define a suitable neighborhood structure for launching the local search procedure and to find a metric for quantifying the disparity between solutions. Three kinds of neighborhood structures are presented and the best fitted one is picked. Due to the speciality of this problem, each solution is transformed into a binary vector, and Hamming distance is adopted for defining disparity metric. The experimental results demonstrate the proposed FWA is more competitive than those commonly used methods. Zhenbao Liu, Liangjun Ke |
CEC | 3 |
| 2014 | A cooperative approach between metaheuristic and branch-and-price for the team orienteering problem with time windowsabstractThe team orienteering problem with time windows (TOPTW) is a well studied routing problem. In this paper, a cooperative algorithm is proposed. It collaborates metaheuristic and branch-and-price. A restricted master problem and subproblem are defined. It uses a heuristic to obtain an integral solution for the restricted master problem and a metaheuristic to generate new columns for the subproblem. Experimental study shows that this algorithm can find new better solutions for several instances in short time, which supports the effectiveness of the cooperative mechanism between metaheuristic and branch-and-price. Liangjun Ke, Huimin Guo, Qingfu Zhang 0001 |
IEEE Congress on Evolutionary Computation | 1 |
| 2014 | Hybridization of Decomposition and Local Search for Multiobjective OptimizationabstractCombining ideas from evolutionary algorithms, decomposition approaches, and Pareto local search, this paper suggests a simple yet efficient memetic algorithm for combinatorial multiobjective optimization problems: memetic algorithm based on decomposition (MOMAD). It decomposes a combinatorial multiobjective problem into a number of single objective optimization problems using an aggregation method. MOMAD evolves three populations: 1) population P(L) for recording the current solution to each subproblem; 2) population P(P) for storing starting solutions for Pareto local search; and 3) an external population P(E) for maintaining all the nondominated solutions found so far during the search. A problem-specific single objective heuristic can be applied to these subproblems to initialize the three populations. At each generation, a Pareto local search method is first applied to search a neighborhood of each solution in P(P) to update P(L) and P(E). Then a single objective local search is applied to each perturbed solution in P(L) for improving P(L) and P(E), and reinitializing P(P). The procedure is repeated until a stopping condition is met. MOMAD provides a generic hybrid multiobjective algorithmic framework in which problem specific knowledge, well developed single objective local search and heuristics and Pareto local search methods can be hybridized. It is a population based iterative method and thus an anytime algorithm. Extensive experiments have been conducted in this paper to study MOMAD and compare it with some other state-of-the-art algorithms on the multiobjective traveling salesman problem and the multiobjective knapsack problem. The experimental results show that our proposed algorithm outperforms or performs similarly to the best so far heuristics on these two problems. Liangjun Ke, Qingfu Zhang 0001, Roberto Battiti |
IEEE Trans. Cybern. | 1 |
| 2013 | MOEA/D-ACO: A Multiobjective Evolutionary Algorithm Using Decomposition and AntColonyabstractCombining ant colony optimization (ACO) and the multiobjective evolutionary algorithm (EA) based on decomposition (MOEA/D), this paper proposes a multiobjective EA, i.e., MOEA/D-ACO. Following other MOEA/D-like algorithms, MOEA/D-ACO decomposes a multiobjective optimization problem into a number of single-objective optimization problems. Each ant (i.e., agent) is responsible for solving one subproblem. All the ants are divided into a few groups, and each ant has several neighboring ants. An ant group maintains a pheromone matrix, and an individual ant has a heuristic information matrix. During the search, each ant also records the best solution found so far for its subproblem. To construct a new solution, an ant combines information from its group's pheromone matrix, its own heuristic information matrix, and its current solution. An ant checks the new solutions constructed by itself and its neighbors, and updates its current solution if it has found a better one in terms of its own objective. Extensive experiments have been conducted in this paper to study and compare MOEA/D-ACO with other algorithms on two sets of test problems. On the multiobjective 0-1 knapsack problem,MOEA/D-ACO outperforms the MOEA/D with conventional genetic operators and local search on all the nine test instances. We also demonstrate that the heuristic information matrices in MOEA/D-ACO are crucial to the good performance of MOEA/D-ACO for the knapsack problem. On the biobjective traveling salesman problem, MOEA/D-ACO performs much better than the BicriterionAnt on all the 12 test instances. We also evaluate the effects of grouping, neighborhood, and the location information of current solutions on the performance of MOEA/D-ACO. The work in this paper shows that reactive search optimization scheme, i.e., the "learning while optimizing" principle, is effective in improving multiobjective optimization algorithms. Liangjun Ke, Qingfu Zhang 0001, Roberto Battiti |
IEEE Trans. Cybern. | 1 |
| 2011 | Guidance-solution based ant colony optimization for satellite control resource scheduling problem
Liangjun Ke |
Appl. Intell. | 3 |
| 2011 | PLBP: An effective local binary patterns texture descriptor with pyramid representation
Xueming Qian, Xian-Sheng Hua 0001, Liangjun Ke |
Pattern Recognit. | 4 |
| 2010 | New pheromone trail updating method of ACO for satellite control resource scheduling problemabstractAn ant colony optimization approach for the satellite control resource scheduling problem is presented. Based on the observation that the solution space of the problem is sparse, two pheromone updating methods, i.e., the reinitialize-guidance-updating and current-guidance-updating methods, are proposed to avoid the trapping in local optima. The basic idea of these two methods is to change the distribution of pheromone trails by updating them with a guidance solution once the algorithm stagnates. We compare the proposed algorithm with several other heuristics. The experimental results demonstrate that our approach is competitive in terms of exploration capability of reaching the near-global optimal solution and adaptability to the future situations. Liangjun Ke |
IEEE Congress on Evolutionary Computation | 3 |
| 2009 | Echo State Network for Abrupt Change Detections in Non-stationary SignalsabstractThe issue of abrupt change detection (ACD) in non-stationary time series signal is considered as a signal classification problem in this paper. A novel reservoir-computing based neural network model (RCNN) is proposed. The main component of RCNN is a large size, sparsely and randomly interconnected dynamical reservoir (DR), which is followed by a single layer perceptron (SLP). The signal containing abrupt changes is firstly projected into the high dimensional state space of DR, and then is linearly classified by the SLP. The SLP is trained by the delta leaning rule. The classification brought out by the SLP is the ACD result. Two synthetic non-stationary time series signals, one is non-chaotic, another one is chaotic, are verified on the RCNN respectively. The simulation experiment results show that the ACD performance of the proposed RCNN is comparable with that of the segment function embedded in MATLAB for the non-chaotic signal, and even outperforms for another chaotic signal. It is concluded that RCNN is a more efficient ACD technique. Qingsong Song, Liangjun Ke |
ISDA | 3 |
| 2008 | A fast and efficient ant colony optimization approach for the set covering problemabstractIn this paper, we present an ant colony optimization (ACO) approach to solve the set covering problem. A constraint-oriented solution construction method is proposed. The main difference between it and the existing method is that, while adding a column to the current partial solution, it randomly selects an uncovered row and only considers the columns covering the row, but not all the unselected columns as candidate solution components. This decreases the number of candidate solution components and therefore accelerates the run speed of the algorithm. Moreover, a simple but effective local search procedure, which aims at eliminating redundant columns and replacing some columns with more effective ones, is developed to improve the quality of solutions constructed by ants while keeping their feasibility. The proposed algorithm has been tested on a number of benchmark instances. Computational results indicate that it is capable of producing high quality solutions and performs better than the existing ACO-based algorithms. Liangjun Ke |
IEEE Congress on Evolutionary Computation | 3 |
| 2008 | An efficient ant colony optimization approach to attribute reduction in rough set theory
Liangjun Ke |
Pattern Recognit. Lett. | 1 |