EDBT 2026 Demo / reviewers in the wild / expert
Wentao Fan 0002
dblp:18/7544-2
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0001-7671-7831ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Achieving Efficient and Robust Multi-Job Resource Scheduling in Deep Learning Clusters
Jianfeng Bao, Wentao Fan 0002, Gongming Zhao, Hongli Xu 0001, Peng Yang 0022, Xiaohu Xu |
IEEE Trans. Netw. | 2 |
| 2025 | HiReC: High-Throughput and Reliable Cross-Cluster VPC Communication in CloudsabstractThe increasing demands of tenants are driving the growth of single virtual private cloud (VPC), leading to a trend towards cross-cluster VPC deployments, which fuels an increasing demand for cross-cluster VPC communication. However, with the rapid increase in cross-cluster traffic and its inherent dynamism, existing solutions fail to meet tenants' demands for throughput and reliability, thereby leading to network performance bottlenecks in cross-cluster communication. To address this issue, we present HiReC, a system designed to achieve high-throughput and reliable cross-cluster VPC communication. To improve throughput, HiReC leverages multiple gateways with a rounding-based mapping algorithm for load balancing to forward cross-cluster traffic. Moreover, we further enhance the forwarding capabilities of gateways with the eXpress Data Path (XDP) technology. To enhance reliability, HiReC employs a low-overhead, eBPF-based monitoring module and adaptive load adjustment mechanism to dynamically adjust traffic distribution among gateways, effectively handling gateway node or link failures. We implement our system and evaluate its performance through testbed experiments. The results show that HiReC can effectively improve the throughput of cross-cluster communication and deal with abnormal events. For example, HiReC improves the throughput by$3.88 \times$and reduces the failure recovery latency by$19 \times$compared with state-of-the-art solutions. Baoqing Wang, Gongming Zhao, Hongli Xu 0001, Wentao Fan 0002, Xiaohu Xu |
IWQoS | 6 |
| 2025 | CROP: Efficient and Robust Multi-Job Placement in Deep Learning ClustersabstractDeep learning (DL) has seen a growing dataset, an expanding model scale, and increasing applications in recent years. There is a notable trend of shifting DL training jobs from local computing units to powerful DL clusters built by cloud providers. These clusters allocate physical training nodes to DL jobs through a process referred to as multi-job placement. Existing multi-job placement strategies fail to achieve high efficiency in resource utilization, DL training, and robustness simultaneously, resulting in poor performance when resources are limited or when abnormalities occur in some devices. To tackle these challenges, we present CROP, an approach that performs efficient and robust multi-job placement in DL clusters. We formulate the efficient and robust multi-job placement problem as a non-linear program and prove its NP-hardness. To solve this problem, we present an effective submodular-based algorithm with a tight approximation factor of ($1-1/e$). We evaluate CROP on a small-scale testbed consisting of 8 physical GPUs and a large-scale simulation employing real-world job traces. Experimental results demonstrate that CROP achieves nearoptimal communication overhead while improving the training throughput of the DL cluster by up to$57.5\%$compared to state-of-the-art solutions. Peng Yang 0022, Gongming Zhao, Hongli Xu 0001, Haibo Wang 0004, Wentao Fan 0002, Xiaohu Xu |
IWQoS | 6 |
| 2025 | CARD: Cost-Efficient and Availability-Aware Application Deployment in Geo-Distributed CloudsabstractThe growing reliance on cloud services has made availability critical for global business continuity. To mitigate disruptions caused by cloud outages, many large-scale applications maintain core functionality through multi-instance deployments. To support this, cloud providers enable cross-AZ deployments to deliver high availability. However, pursuing high availability must be balanced against cost efficiency, which presents three key challenges: electricity price disparity, application affinity requirement, and disaster recovery demand. Existing research primarily focuses on single-region optimization, often overlooking the potential benefits of multi-region deployment in cost and availability. While some works explore multi-region deployment strategies, they fail to address application affinity or disaster recovery requirements, resulting in low application availability. To address this issue, we propose CARD, a cost-efficient and availability-aware application deployment scheme in geodistributed clouds. Specifically, we formulates this problem as a mixed-integer nonlinear program and designs an efficient approximation algorithm based on submodular function, achieving an approximation ratio of ($1-1 / e$). Large-scale simulations on realworld datasets demonstrate the algorithm's effectiveness, overall reducing costs by 20% – 50% and improving availability by 90% compared to existing solutions. Gongming Zhao, Hongli Xu 0001, Wentao Fan 0002, Xiaohu Xu |
IWQoS | 5 |
| 2024 | Bi-LSTM/GRU-based anomaly diagnosis for virtual network function instance
Wentao Fan 0002, Shiyuan Cui, Yuehui Tan, Fan Yang 0046, Weihong Wu |
Comput. Networks | 1 |
| 2023 | DRL-Based Service Function Chain Edge-to-Edge and Edge-to-Cloud Joint Offloading in Edge-Cloud NetworkabstractIn this paper, we study service function chain (SFC) offloading in the edge-cloud network. Two offloading options are available for fully-loaded edge nodes: edge to edge (E2E) offloading and edge to cloud (E2C) offloading. Both E2E offloading and E2C offloading have been optimized in existing research, and Deep Reinforcement Learning (DRL) methods were adopted to achieve excellent performances. However, DRL-based SFC E2E and E2C joint offloading is still a research gap. In this paper, we propose a DRL-based SFC E2E and E2C joint offloading optimization algorithm for the edge-cloud network to maximize the utilization efficiency of edge resources. Twin Delayed Deep Deterministic policy gradient (TD3) algorithm is applied to the optimization problem. The simulation results indicate that the proposed algorithm has excellent convergence performance and improves the utilization efficiency of edge resources in edge-cloud network scenarios of diverse scales. Wentao Fan 0002, Fan Yang 0046, Peilong Wang, Mao Miao, Tao Huang 0005 |
IEEE Trans. Netw. Serv. Manag. | 1 |