Mingfeng Fan

dblp:267/1635 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0003-4008-5873ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 UHTS-DRL: A deep reinforcement learning framework for integrated agile satellite observation and data transmission scheduling
Mingfeng Fan, Yi Gu 0003, Qizhang Luo, Yalin Wang 0003, Xinwei Wang 0006, Guohua Wu 0001
Inf. Sci.2
2026 Unicorn: A Universal and Collaborative Reinforcement Learning Approach Toward Generalizable Network-Wide Traffic Signal Control
Peizhuo Li, Mingfeng Fan, Guillaume Sartoretti
IEEE Trans. Intell. Transp. Syst.5
2025 Preference-Driven Multi-Objective Combinatorial Optimization with Conditional Computation
abstract
Recent deep reinforcement learning methods have achieved remarkable success in solving multi-objective combinatorial optimization problems (MOCOPs) by decomposing them into multiple subproblems, each associated with a specific weight vector. However, these methods typically treat all subproblems equally and solve them using a single model, hindering the effective exploration of the solution space and thus leading to suboptimal performance. To overcome the limitation, we propose POCCO, a novel plug-and-play framework that enables adaptive selection of model structures for subproblems, which are subsequently optimized based on preference signals rather than explicit reward values. Specifically, we design a conditional computation block that routes subproblems to specialized neural architectures. Moreover, we propose a preference-driven optimization algorithm that learns pairwise preferences between winning and losing solutions. We evaluate the efficacy and versatility of POCCO by applying it to two state-of-the-art neural methods for MOCOPs. Experimental results across four classic MOCOP benchmarks demonstrate its significant superiority and strong generalization.
Mingfeng Fan, Jianan Zhou 0002, Yaoxin Wu, Jinbiao Chen, Guillaume Sartoretti
NeurIPS1
2025 Heterogeneous Attention-Based Graph Convolutional Network for Solving Asymmetric Pickup and Delivery Problem
abstract
In recent years, there has been a notable increase in demand for pickup and delivery services, mainly driven by the expansion of e-commerce. These services can be conceptualized as pickup and delivery problems (PDPs), which are important variants of vehicle routing problems (VRPs). While neural methods based on deep reinforcement learning (DRL) have demonstrated success in solving VRPs, current neural methods for PDP predominantly depend on customer coordinates and Euclidean distances. This heavy reliance can diminish their practical value, given the intricacies of real-world road networks and inherent asymmetries. This paper presents a novel learning-based method, i.e., HA-GCN, that couples heterogeneous attention (HA) and a graph convolutional network (GCN) to tackle the asymmetric pickup and delivery problem (APDP). The HA-GCN model utilizes HA to comprehend the innate pairing and precedence constraints between nodes in APDP. Concurrently, it employs GCN to amalgamate features of nodes and edges, facilitating the extraction of asymmetric attributes. Additionally, to augment the real-world relevance of HA-GCN, we train the model grounded in a dataset drawn from real geographic information. Comprehensive experiments suggest that our proposed method performs favorably against both traditional and learning-based state-of-the-art approaches, while demonstrating robust generalization capabilities. Note to Practitioners—Research on the pickup and delivery problem (PDP) addresses challenges in the rapidly growing field of end-to-end delivery services, such as food and parcel delivery. This study focuses on the asymmetric pickup and delivery problem (APDP), a variant of PDP reflecting real-world road networks with one-way streets and varying distances between locations. However, due to the exponential complexity and asymmetric nature of APDP, existing traditional methods and deep reinforcement learning (DRL) approaches struggle to effectively address this challenge. To address this, we propose HA-GCN, a DRL-based model compling heterogeneous attention and a graph convolutional network. HA-GCN effectively distinguishes pickup and delivery nodes in APDP, capturing differences in asymmetric distances in real-world road networks. Experimental results based on instances of different scales show that HA-GCN outperforms learning-based and traditional baselines, exhibiting satisfactory generalization capabilities. Given these confirmed advantages, our HA-GCN has great potential to not only offer the practitioners effective alternative approaches, but also improve the solution quality for the pickup and delivery in the operation of express hubs or logistics centers in specific regions of the real world.
Guohua Wu 0001, Mingfeng Fan, Zhiguang Cao, Yalin Wang 0003
IEEE Trans Autom. Sci. Eng.3
2025 DL-DRL: A Double-Level Deep Reinforcement Learning Approach for Large-Scale Task Scheduling of Multi-UAV
abstract
Exploiting unmanned aerial vehicles (UAVs) to execute tasks is gaining growing popularity recently. To address the underlying task scheduling problem, conventional exact and heuristic algorithms encounter challenges such as rapidly increasing computation time and heavy reliance on domain knowledge, particularly when dealing with large-scale problems. The deep reinforcement learning (DRL) based methods that learn useful patterns from massive data demonstrate notable advantages. However, their decision space will become prohibitively huge as the problem scales up, thus deteriorating the computation efficiency. To alleviate this issue, we propose a double-level deep reinforcement learning (DL-DRL) approach based on a divide and conquer framework (DCF), where we decompose the task scheduling of multi-UAV into task allocation and route planning. Particularly, we design an encoder-decoder structured policy network in our upper-level DRL model to allocate the tasks to different UAVs, and we exploit another attention-based policy network in our lower-level DRL model to construct the route for each UAV, with the objective to maximize the total value of executed tasks given the maximum flight distance of the UAV. To effectively train the two models, we design an interactive training strategy (ITS), which includes pre-training, intensive training and alternate training. Experimental results show that our DL-DRL performs favorably against the learning-based and conventional baselines including the OR-Tools, in terms of solution quality and computation efficiency. We also verify the generalization performance of our approach by applying it to larger sizes of up to 1500 tasks and to different flight distances of UAVs. Moreover, we also show via an ablation study that our ITS can help achieve a balance between the performance and training efficiency. Our code is publicly available at https://faculty.csu.edu.cn/guohuawu/zh_CN/zdylm/193832/list/ index.htm.Note to Practitioners—Unmanned aerial vehicles (UAVs) are of great practical usage, as they have many real world applications. When a group of UAVs are employed to execute large-scale tasks, a core question is how to scheduling the UAVs, so that they could complete the tasks efficiently. However, it is a computationally hard problem due to the exponentially increasing search space. To solve this problem, we propose a double-level deep reinforcement learning (DL-DRL) approach within a divide-and-conquer framework (DCF), where the upper-level DRL model is responsible for the task allocation, and the lower-level DRL model is responsible for the UAV route planning. To better train the two DRL models who have interplay with each other, we propose a simple yet efficient training strategy, termed interactive training strategy (ITS), which includes pre-training, intensive training and alternate training. The experimental results based on instances of various scales show that our DL-DRL approach outperformed learning-based and conventional baselines, and the designed ITS could strike a good balance between performance and training efficiency. In light of those verified advantages, we believe that our DL-DRL approach has favorable potential to solve the practical task scheduling problem of multi-UAV in real world.
Xiao Mao, Guohua Wu 0001, Mingfeng Fan, Zhiguang Cao, Witold Pedrycz
IEEE Trans Autom. Sci. Eng.3
2025 Conditional Neural Heuristic for Multiobjective Vehicle Routing Problems
abstract
Existing neural heuristics for multiobjective vehicle routing problems (MOVRPs) are primarily conditioned on instance context, which failed to appropriately exploit preference and problem size, thus holding back the performance. To thoroughly unleash the potential, we propose a novel conditional neural heuristic (CNH) that fully leverages the instance context, preference, and size with an encoder-decoder structured policy network. Particularly, in our CNH, we design a dual-attention-based encoder to relate preferences and instance contexts, so as to better capture their joint effect on approximating the exact Pareto front (PF). We also design a size-aware decoder based on the sinusoidal encoding to explicitly incorporate the problem size into the embedding, so that a single trained model could better solve instances of various scales. Besides, we customize the REINFORCE algorithm to train the neural heuristic by leveraging stochastic preferences (SPs), which further enhances the training performance. Extensive experimental results on random and benchmark instances reveal that our CNH could achieve favorable approximation to the whole PF with higher hypervolume (HV) and lower optimality gap (Gap) than those of the existing neural and conventional heuristics. More importantly, a single trained model of our CNH can outperform other neural heuristics that are exclusively trained on each size. In addition, the effectiveness of the key designs is also verified through ablation studies.
Mingfeng Fan, Yaoxin Wu, Zhiguang Cao, Wen Song 0004, Guillaume Sartoretti, Huan Liu 0028, Guohua Wu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 HeteroLight: A General and Efficient Learning Approach for Heterogeneous Traffic Signal Control
abstract
Efficient and scalable adaptive traffic signal control is crucial in reducing congestion, maximizing through-put, and improving mobility experience in ever-expanding cities. Recent advances in multi-agent reinforcement learning (MARL) with parameter sharing have significantly improved the adaptive optimization of large-scale, complex, and dynamic traffic flows. However, the limited model representation capability due to shared parameters impedes the learning of diverse control strategies for intersections with different flows/topologies, posing significant challenges to achieving effective signal control in complex and varied real-world traffic scenarios. To address these challenges, we present a novel MARL-based general traffic signal control framework, called HeteroLight. Specifically, we first introduce a General Feature Extraction (GFE) module, crafted in a decoder-only fashion, where we employ an attention mechanism to facilitate efficient and flexible extraction of traffic dynamics at intersections with varied topologies. Additionally, we incorporate an Intersection Specifics Extraction (ISE) module, designed to identify key latent vectors that represent the unique intersection’s topology and traffic dynamics through variational inference techniques. By integrating the learned intersection-specific information into policy learning, we enhance the parameter-sharing mechanism, improving the model’s representation diversity among different agents and enabling the learning of a more efficient shared control strategy. Through comprehensive evaluations against other state-of-the-art traffic signal control methods on the real-world Monaco traffic network, our empirical findings reveal that HeteroLight consistently outperforms other methods across various evaluation metrics, highlighting its superiority in optimizing traffic flows in heterogeneous traffic networks.
Peizhuo Li, Mingfeng Fan, Guillaume Sartoretti
IROS3
2022 An Autonomous Path Planning Method for Unmanned Aerial Vehicle Based on a Tangent Intersection and Target Guidance Strategy
abstract
Unmanned aerial vehicle (UAV) path planning enables UAVs to avoid obstacles and reach the target efficiently. To generate high-quality paths without obstacle collision for UAVs, this article proposes a novel autonomous path planning algorithm based on a tangent intersection and target guidance strategy (APPATT). Guided by a target, the elliptic tangent graph method is used to generate two sub-paths, one of which is selected based on heuristic rules when confronting an obstacle. The UAV flies along the selected sub-path and repeatedly adjusts its flight path to avoid obstacles through this way until the collision-free path extends to the target. Considering the UAV kinematic constraints, the cubic B-spline curve is employed to smooth the waypoints for obtaining a feasible path. Compared with A*, PRM, RRT and VFH, the experimental results show that APPATT can generate the shortest collision-free path within 0.05 seconds for each instance under static environments. Moreover, compared with VFH and RRTRW, APPATT can generate satisfactory collision-free paths under uncertain environments in a nearly real-time manner. It is worth noting that APPATT has the capability of escaping from simple traps within a reasonable time.
Huan Liu 0028, Xiamiao Li, Mingfeng Fan, Guohua Wu 0001, Witold Pedrycz, Ponnuthurai N. Suganthan
IEEE Trans. Intell. Transp. Syst.3
2021 An Iterative Two-Phase Optimization Method Based on Divide and Conquer Framework for Integrated Scheduling of Multiple UAVs
abstract
Task scheduling of multiple UAVs has become a highly active area of research in recent years. Previous research has generally solved the problem in a whole manner, which makes it hard to efficiently generate high-quality task scheduling schemes due to prohibitive computational complexity. By contrast, the paper constructs a novel divide and conquer framework for multi-UAV task scheduling (DCF), which partitions the original multi-UAV scheduling problem into multiple scheduling sub-problems for all the UAVs. To be specific, DCF includes two phases: one is the task allocation phase which produces multiple scheduling sub-problems and the other is the single UAV scheduling phase which generates the scheduling scheme with sequential tasks for each single UAV considering constraints involving UAV capabilities and task demands. Two phases are iteratively performed until the predefined stopping criteria are met. In the task allocation phase, we propose a tabu-list-based simulated annealing (SATL) algorithm to realize task allocation among multiple UAVs. After obtaining the task allocation scheme, a satisfactory scheduling scheme of each single UAV is generated by variable neighborhood descent (VND) algorithm. Extensive experiments and comparative studies are conducted, demonstrating the efficiency of DCF and the proposed SATL-VND algorithm.
Huan Liu 0028, Xiamiao Li, Guohua Wu 0001, Mingfeng Fan, Rui Wang 0017, Liang Gao 0001, Witold Pedrycz
IEEE Trans. Intell. Transp. Syst.4
2020 Voting-mechanism based ensemble constraint handling technique for real-world single-objective constrained optimization
abstract
Constraint handling techniques are of great significance in efficiently solving constrained optimization problems. This paper proposes a novel ensemble framework for constraint handling techniques based on voting-mechanism, in which four popular constraint handling techniques are included. Each of the constituent constraint handling techniques votes for the solutions at each generation based on its own rules. Solutions getting more votes are regarded as promising individuals and survive to the next generation. This ensemble framework based on voting-mechanism reflects the collective wisdom in decision-making of human beings. In addition, a differential evolution (DE) variant is designed as the search engine, in which four search strategies are combined to generate new individuals and maintain the balance between diversity and convergence of the population. The proposed algorithm has been tested on 57 real world single-objective constraint optimization problems and 7 problems are selected using variable reduction strategy (VRS). The experiment shows that the proposed algorithm achieves competitive performance, indicating that the voting-mechanism ensemble constraint handling technique combine DE algorithm together can effectively deal with constrained optimization problems.
Xupeng Wen, Guohua Wu 0001, Mingfeng Fan, Rui Wang 0017, Ponnuthurai N. Suganthan
CEC3