Wenhan Zhan

dblp:162/6572 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-1851-7185ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 9 · 2 first-author · 6 since 2021Systems, architecture and hardware · 5 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dynamic priority-based area partitioning, trajectory planning, and task scheduling in computing-while-flying UAV networks
Zijia Zhao, Wenhan Zhan, Geyong Min, Xu Jiang 0004, Liang Zhao 0004, Hualong Huang
Future Gener. Comput. Syst.2
2026 Cost-Aware Dependent Task Offloading and Resource Allocation for Satellite Edge Computing: An Asynchronous Deep Reinforcement Learning Approach
abstract
The integration of satellite communications with mobile edge computing (MEC) into space-air-ground integrated networks, known as satellite edge computing (SEC), has become a crucial research field for future communication systems to provide extensive global coverage services. This paper investigates the joint dependent task offloading and resource allocation problem for remote Internet-of-Things (IoT) applications within the SEC architecture. The proposed system leverages unmanned aerial vehicles (UAVs) as mobile access points and edge servers and utilizes low- earth orbit (LEO) satellites and ground stations as cloud computing resources. Multiple applications with dependent tasks from IoT devices (IoTDs) are modeled as directed acyclic graphs (DAGs). To address the challenges of reducing the system cost in UAV-assisted SEC, we first propose a one-to-many matching algorithm to associate IoTDs with UAVs. Then, a multi-application task sequence algorithm is devoted to merging the multiple DAGs and sorting the task order. Finally, a graph-aware asynchronous multi-agent reinforcement learning approach empowers the agents to autonomously discover optimal offloading and resource allocation strategies. Extensive simulations based on real-world datasets demonstrate the effectiveness of the proposed approach in minimizing the system costs while meeting application latency requirements, outperforming other benchmark algorithms.
Hualong Huang, Hancong Duan, Wenhan Zhan, Geyong Min, Kai Peng 0002, Yuchuan Lei
IEEE Trans. Mob. Comput.3
2026 EdgeSD: Efficient Speculative Decoding With Vision-Decoding Disaggregation for MLLM Inference in Edge-Cloud Networks
abstract
The deployment of multimodal large language models (MLLMs) in edge-cloud networks faces critical challenges, including computational resource heterogeneity, memory bottlenecks, and bandwidth constraints. To address these issues, we propose EdgeSD, a novel framework that accelerates MLLM inference by integrating speculative decoding (SD) with edge-cloud collaboration. First, EdgeSD decouples the vision encoding and decoding processes of the draft MLLM across heterogeneous edge servers (ESs). This disaggregation architecture overcomes single-node memory constraints, enabling optimized resource utilization and high-resolution input processing. Second, to resolve the communication bottleneck and computational burden inherent in this distributed architecture, EdgeSD integrates a bandwidth-aware dynamic image token merging (ITM) method. Unlike general pruning techniques, this EdgeSD-specific ITM method focuses on minimizing inter-ES transmission latency for vision-decoding disaggregation while maintaining draft quality. Third, to optimize SD efficiency on consumer-grade ESs, EdgeSD employs an adaptive and scalable token tree structure solved using a parallel delta-stepping algorithm. This structure maximizes the number of accepted tokens under strict edge latency constraints. Extensive experiments on six multimodal datasets and five benchmarks with various MLLM pairs demonstrate that EdgeSD achieves substantial acceleration and throughput gains in edge-cloud collaboration scenarios using a lightweight draft MLLM, achieving 3.04-5.12x speedup compared to baseline methods.
Hualong Huang, Wenhan Zhan, Hancong Duan, Kai Peng 0002, Geyong Min, Zijia Zhao, Zitian Zhao, Yalan Ye
IEEE Trans. Mob. Comput.2
2025 Multiobjective optimization deep reinforcement learning for dependent task scheduling based on spatio-temporal fusion graph neural network
Zhi Wang 0020, Wenhan Zhan, Hancong Duan, Hualong Huang
Eng. Appl. Artif. Intell.2
2025 Dynamic Model Deployment, Batch Scheduling, and Resource Allocation in MLLM-Enabled Edge-Cloud Networks: A Multiagent Two-Timescale DRL Approach
abstract
The deployment of multimodal large language models (MLLMs) on resource-constrained mobile devices poses significant challenges due to their high computational demands. This paper introduces a novel two-timescale optimization framework for efficient MLLM inference in Edge-Cloud networks, addressing the problem of multi-timescale resource management by jointly optimizing slow-timescale MLLMs deployment decisions and fast-timescale batch scheduling, GPU resource allocation, and bandwidth allocation under dynamic network conditions and spatiotemporal request heterogeneity. Our key innovation is a hierarchical twin delayed deep deterministic policy gradient (HALTD3) algorithm that integrates attention mechanisms and long short-term memory networks to optimize slow-timescale MLLMs deployment and fast-timescale resource allocation, minimizing weighted system costs including deployment cost, end-to-end latency, and energy consumption, while meeting stringent quality-of-service requirements. Extensive experiments demonstrate that the HALTD3 algorithm substantially outperforms baseline methods in reducing system costs across diverse MLLM workloads and dynamic network scenarios, validating its effectiveness for practical edge-cloud collaborative inference.
Hualong Huang, Yongkang Du, Wenhan Zhan, Hancong Duan, Kai Peng 0002, Yamin Cheng, Yalan Ye, Zitian Zhao
IEEE Internet Things J.3
2025 Deep-Reinforcement-Learning-Based Continuous Workflows Scheduling in Heterogeneous Environments
abstract
Workflow scheduling plays a critical role in optimizing completion time and throughput in distributed cloud environments, leveraging the parallelism of heterogeneous computing resources. However, existing workflow scheduling algorithms often fall short due to heuristic limitations and the challenges in adaptability within heterogeneous settings, leading to suboptimal scheduling solutions. In this paper, we present a novel deep reinforcement learning (DRL) framework tailored for continuous workflow scheduling in heterogeneous environments. First, we propose an intelligent scheduler that updates the policy network through interactions with a multi-tenant environment, triggered by scheduling events. Next, the framework incorporates a Graph Attention Network (GAT) and a self-attention MultiLayer Perceptron (MLP) to preprocess the workflow topology and embed dynamic features of ready tasks and available processors into the state input at each scheduling step. Additionally, a k-dimensional tree-based k-nearest neighbors (kNN) algorithm is employed to map the output action vector to a pair of executed ready task and processor, facilitating the transition from continuous to discrete action spaces and addressing challenges associated with dynamic action spaces. Experimental results demonstrate that our method converges effectively in continuous workflow scheduling scenarios and significantly outperforms the best-known methods in terms of average makespan and load balancing efficiency.
Zhi Wang 0020, Wenhan Zhan, Hancong Duan, Geyong Min, Hualong Huang
IEEE Internet Things J.2
2024 Battery-Care Resource Allocation and Task Offloading in Multi-Agent Post-Disaster MEC Environment
abstract
Being an up-and-coming application scenario of mobile edge computing (MEC), the post-disaster rescue suffers multitudinous computing-intensive tasks but unstably guaranteed network connectivity. In rescue environments, quality of service (QoS), such as task execution delay, energy consumption and battery state of health (SoH), is of significant meaning. This paper studies a multi-user post-disaster MEC environment with unstable 5G communication, where device-to-device (D2D) link communication and dynamic voltage and frequency scaling (DVFS) are adopted to balance each user's requirement for task delay and energy consumption. A battery degradation evaluation approach to prolong battery lifetime is also presented. The distributed optimization problem is formulated into a mixed cooperative-competitive (MCC) multi-agent Markov decision process (MAMDP) and is tackled with recurrent multi-agent Proximal Policy Optimization (rMAPPO). Extensive simulations and comprehensive comparisons with other representative algorithms clearly demonstrate the effectiveness of the proposed rMAPPO-based offloading scheme.
Yiwei Tang, Hualong Huang, Wenhan Zhan, Geyong Min, Zhekai Duan, Yuchuan Lei
WCNC3
2024 Optimal service caching, pricing and task partitioning in mobile edge computing federation
Hualong Huang, Zhekai Duan, Wenhan Zhan, Geyong Min, Kai Peng 0002
Future Gener. Comput. Syst.3
2024 Mobility-Aware Computation Offloading With Load Balancing in Smart City Networks Using MEC Federation
abstract
Internet-of-Things (IoT) has played a critical role in developing sustainable smart cities and emerging numerous latency-sensitive IoT applications. Mobile edge computing (MEC) federation has the capability to incorporate a transparent resource management approach, which enables the sharing and utilization of MEC services from edge infrastructure providers (EIPs) and provides agile access services to mobile devices (MDs). In this paper, we investigate the joint optimization problem of computation offloading, task migration, and resource allocation in the MEC federation. The objective is to minimize the weighted sum of latency and energy consumption while maintaining load balancing under the constraint of the long-term migration cost budget of EIPs. To address the problem, we decompose it into two sub-problems: 1) the MDs clustering sub-problem and 2) the sub-problem of joint computation offloading, task migration, and resource allocation. Firstly, an MDs clustering matching (MDCM) algorithm is proposed to cluster the MDs in edge servers (ESs) according to the differences in channel gains. Afterward, the second sub-problem is simplified by the Lyapunov optimization technique, and then we propose a Transformer-based mobility prediction model and a decentralized deep deterministic policy gradient (DDPG)-based framework to solve it. Extensive simulation results demonstrate the cost-efficiency of the proposed algorithm.
Hualong Huang, Wenhan Zhan, Geyong Min, Zhekai Duan, Kai Peng 0002
IEEE Trans. Mob. Comput.2
2023 Distributed Dependent Task Offloading in CPU-GPU Heterogenous MEC: A Federated Reinforcement Learning Approach
abstract
Mobile edge computing (MEC) has emerged as a promising paradigm to enable computation-intensive and latency-sensitive mobile applications by offloading tasks to proximal edge servers. This paper proposes a novel federated reinforcement learning framework called Transformer-based Federated Soft Actor-Critic (TFSAC) to address a joint computation offloading and resource scheduling problem in a CPU-GPU heterogeneous MEC network while preserving privacy. Specifically, a graph attention network (GAT) extracts high-dimensional features from the task dependency graph. Rather than simply averaging weights, TFSAC applies transformer encoders to learn contextual relationships between agents and enable selective aggregation of relevant knowledge during federated model training to preserve agents’ privacy. Experiments on real-world trace data demonstrate TFSAC’s superiority over benchmarks in maximizing quality-of-service (QoS) across configurations.
Hualong Huang, Zhekai Duan, Wenhan Zhan, Zhi Wang 0020, Zitian Zhao
TrustCom3
2023 Multi-Scale Human-Object Interaction Detector
abstract
Transformers are transforming the landscape of computer vision, especially for image-level recognition and instance-level detection tasks. Human-object interaction detection transformer (HOI-TR) is the first transformer-based end-to-end learning system for human-object interaction (HOI) detection; vision transformers build a simple multi-stage structure for multi-scale representation with single-scale patch and are the first patch-based transformer architecture for image-level recognition and instance-level detection. In this paper, we build a transformer-based multi-scale human-object interaction detector (MHOI), a novel method to integrate Vision and HOI detection Transformer, instead of directly incorporating two types of transformers, since the vision transformer lacks hierarchical architecture to handle the large variations in the scale of visual entities due to the single-scale patch partitioning. Specifically, MHOI embeds features of the same size (i.e., sequence length) with patches of variable scales simultaneously by utilizing overlapping convolutional patch embedding, then introduces an efficient transformer decoder that designs the query based on anchor points and essential auxiliary techniques to boost the HOI detection performance. Numerically, extensive experiments on several benchmarks demonstrate that our proposed framework outperforms prior existing methods coherently and achieves the impressive performance of 29.67 mAP on HICO-DET and 58.7 mAP on V-COCO, respectively.
Yamin Cheng, Zhi Wang 0020, Wenhan Zhan, Hancong Duan
IEEE Trans. Circuits Syst. Video Technol.3
2022 Dependent Task Offloading for Edge Computing based on Deep Reinforcement Learning
abstract
Edge computing is an emerging promising computing paradigm that brings computation and storage resources to the network edge, hence significantly reducing the service latency and network traffic. In edge computing, many applications are composed of dependent tasks where the outputs of some are the inputs of others. How to offload these tasks to the network edge is a vital and challenging problem which aims to determine the placement of each running task in order to maximize the Quality-of-Service (QoS). Most of the existing studies either design heuristic algorithms that lack strong adaptivity or learning-based methods but without considering the intrinsic task dependency. Different from the existing work, we propose an intelligent task offloading scheme leveraging off-policy reinforcement learning empowered by a Sequence-to-Sequence (S2S) neural network, where the dependent tasks are represented by a Directed Acyclic Graph (DAG). To improve the training efficiency, we combine a specific off-policy policy gradient algorithm with a clipped surrogate objective. We then conduct extensive simulation experiments using heterogeneous applications modelled by synthetic DAGs. The results demonstrate that: 1) our method converges fast and steadily in training; 2) it outperforms the existing methods and approximates the optimal solution in latency and energy consumption under various scenarios.
Jin Wang 0024, Jia Hu 0001, Geyong Min, Wenhan Zhan, Albert Y. Zomaya, Nektarios Georgalas
IEEE Trans. Computers4
2020 Deep-Reinforcement-Learning-Based Offloading Scheduling for Vehicular Edge Computing
abstract
Vehicular edge computing (VEC) is a new computing paradigm that has great potential to enhance the capability of vehicle terminals (VTs) to support resource-hungry in-vehicle applications with low latency and high energy efficiency. In this article, we investigate an important computation offloading scheduling problem in a typical VEC scenario, where a VT traveling along an expressway intends to schedule its tasks waiting in the queue to minimize the long-term cost in terms of a tradeoff between task latency and energy consumption. Due to diverse task characteristics, dynamic wireless environment, and frequent handover events caused by vehicle movements, an optimal solution should take into account both where to schedule (i.e., local computation or offloading) and when to schedule (i.e., the order and time for execution) each task. To solve such a complicated stochastic optimization problem, we model it by a carefully designed Markov decision process (MDP) and resort to deep reinforcement learning (DRL) to deal with the enormous state space. Our DRL implementation is designed based on the state-of-the-art proximal policy optimization (PPO) algorithm. A parameter-shared network architecture combined with a convolutional neural network (CNN) is utilized to approximate both policy and value function, which can effectively extract representative features. A series of adjustments to the state and reward representations are taken to further improve the training efficiency. Extensive simulation experiments and comprehensive comparisons with six known baseline algorithms and their heuristic combinations clearly demonstrate the advantages of the proposed DRL-based offloading scheduling method.
Wenhan Zhan, Chunbo Luo, Jin Wang 0024, Chao Wang 0015, Geyong Min, Hancong Duan, Qingxin Zhu
IEEE Internet Things J.1
2019 Deep Reinforcement Learning-Based Computation Offloading in Vehicular Edge Computing
abstract
Inspired by mobile edge computing (MEC), vehicular edge computing (VEC) enables vehicle terminals to support resource-hungry on-vehicle applications with significantly lower latency and less energy consumption. In this paper, we investigate the computation offloading problem in a typical VEC scenario, where a vehicle offloads its computation tasks to the VEC servers deployed in the road side unit (RSU) to minimize its long-term user cost. The mobility of the vehicle coupled with the high dynamics of the environment makes the problem particularly difficult. To tackle this challenge, a deep reinforcement learning (DRL) based offloading method is proposed, which approximates the offloading policy (OP) by a deep neural network (DNN) and trains the DNN with the proximal policy optimization (PPO) algorithm without a priori knowledge of the environment dynamics. Extensive simulation experiments and comprehensive comparison with six baseline algorithms demonstrate that it can achieve the lowest user cost in most cases.
Wenhan Zhan, Chunbo Luo, Jin Wang 0024, Geyong Min, Hancong Duan
GLOBECOM1
2016 A multi-channel architecture for metadata management in cloud storage systems by binding CPU-cores to disks
abstract
Summary Metadata operations have become dominant file operations in the storage systems. In the scenarios of read‐more and write‐less of massive small files, the current distributed file systems suffer from the unsatisfying performance and scalability of metadata service because of random disk I/O during metadata operations. In this paper, a highly efficient metadata management architecture for cloud storage systems is proposed. The cluster design significantly improves the scalability of the system. A concept of the disk I/O channel is introduced, which is an independent data storage pipe by binding an independent CPU‐core to each physical disk. In addition, a multi‐channel fast key‐value storage engine is proposed to provide the extremely efficient performance for the underlying storage service, which takes full advantages of multi‐core processors and parallel disks I/O. Besides, a new dynamic load‐balancing strategy is proposed to reduce load thrashing and improve the precision of rebalancing among the clusters. Performance measurements under a variety of benchmarks show that the metadata management is capable of handling the massive small files storage and the performance is improved significantly compared to the existing solutions. Copyright © 2016 John Wiley & Sons, Ltd.
Hancong Duan, Xiaoke Xiang, Geyong Min, Wenhan Zhan, Pengcheng Lv
Concurr. Comput. Pract. Exp.5
2015 Distributed in-memory vocabulary tree for real-time retrieval of big data images
Hancong Duan, Yubing Peng, Geyong Min, Xiaoke Xiang, Wenhan Zhan
Ad Hoc Networks5
2015 A high-performance distributed file system for large-scale concurrent HD video streams
abstract
Summary With the rapid development of intelligent transportation technology, high‐definition data storage, and processing of massive amounts of video surveillance have become key issues. When thousands of high‐bit‐rate video streaming concurrently writes, disk I/O throughput becomes a bottleneck. In addition, this leads to serious energy consumption and disk abrasion. To solve these problems, a new distributed file system for high concurrent and high‐bit‐rate writing is designed. It combines an optimized data storage model, efficient metadata management, and exquisite disk schedule mechanism. The optimized data storage model uses a file pre‐allocation strategy and multiple‐stream input modulating technology to convert the randomly concurrent writes on the disk into sequential writes; the metadata management provides an efficient means of retrieving the specified data; and the dual‐partition schedule mechanism can ensure the disk stability with less abrasion. Through this distributed file system, the disk I/O throughput can be saturated in a high concurrent writing environment. The performance evaluation results demonstrate that the I/O throughput of a normal 7200RPM SATA III disk in our scheme can be stabilized at 150MB/s, easily to support 300 concurrent high‐definition video streams (4Mbit/s each). The distributed file system with eight commodity servers can afford the ability of supporting 8000 high‐definition video streams concurrently writing, which is far greater than the existing video surveillance storage solutions. Copyright © 2015 John Wiley & Sons, Ltd.
Hancong Duan, Wenhan Zhan, Geyong Min, Shengmei Luo
Concurr. Comput. Pract. Exp.2