EDBT 2026 Demo / reviewers in the wild / expert
Gangyi Luo
dblp:150/4653
· DBLP profile ↗
15ranked-venue papers
3as first author
13since 2021 · last 2026
0009-0007-7329-5291ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 8 since 2021Systems, architecture and hardware · 7 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Achieving High-Throughput and Reliable Cross-Cluster VPC Communication in CloudsabstractThe increasing demands of tenants are driving the growth of single virtual private cloud (VPC), leading to a trend towards cross-cluster VPC deployments, which fuels an increasing demand for cross-cluster VPC communication. However, the rapid growth of cross-cluster traffic and its inherently dynamic nature have exposed critical limitations in existing network solutions, which now struggle to maintain required throughput levels and ensure reliable communication. This growing inadequacy has consequently created persistent network performance bottlenecks in cross-cluster communication systems. To address this issue, we present HiReC, a system designed to achieve high-throughput and reliable cross-cluster VPC communication. To optimize throughput performance, HiReC leverages multiple gateways with a rounding-based mapping algorithm that ensures effective load balancing to forward cross-cluster traffic. Furthermore, HiReC augments gateway forwarding efficiency through implementation of the eXpress Data Path (XDP) framework, leveraging kernel-bypass techniques to accelerate packet processing. For reliability enhancement, HiReC employs a low-overhead, eBPF-based monitoring module and adaptive load adjustment mechanism to dynamically adjust traffic distribution among gateways, effectively handling gateway node or link failures. We implement our system and evaluate its performance through testbed experiments and simulation experiments. The results show that HiReC can effectively improve the throughput of cross-cluster communication and deal with abnormal events. For example, HiReC improves the throughput by$3.8\times $and reduces the failure recovery latency by$19\times $compared with state-of-the-art solutions. Gongming Zhao, Hongli Xu 0001, Baoqing Wang, Gangyi Luo |
IEEE Trans. Netw. | 6 |
| 2026 | Resource-Aware Distributed Training Job Placement for GPU Cluster DefragmentationabstractDistributed training (DT) has emerged as a solution to address the growing computational resource demands of training large-scale machine learning models. To meet this need, cloud providers typically build GPU clusters to accommodate DT jobs. For DT job requests, cloud providers need to determine in which GPUs place workers (i.e., job placement). Existing approaches usually place workers on as few idle machines as possible to minimize communication time. However, this scheme will lead to aresource fragmentation problem, which degrades the resource utilization rate of the GPU cluster and increases training costs for cloud providers. In this paper, we propose$\textsf {Titan}$, a novel job placement scheme that mitigates the influence of resource fragmentation by enhancing the utilization of non-idle machines. To further optimize resource allocation, we introduce a dynamic defragmentation algorithm that migrates fragmented jobs to consolidate GPU resources, enabling efficient placement of large-scale training jobs.$\textsf {Titan}$formulates a multi-objective non-linear optimization problem and proves its NP-hardness. To solve this problem,$\textsf {Titan}$presents an effective submodular-based greedy algorithm with a tight approximation ratio ($1-\frac {1}{e}$). We evaluate$\textsf {Titan}$with a large-scale simulation employing real-world job traces and a small-scale testbed consisting of 8 servers with 32 logical GPUs. Experimental results show that$\textsf {Titan}$can achieve near-optimal training throughput while improving the efficiency of the cluster by 74.9% compared to the state-of-the-art solutions. Gongming Zhao, Yichen Dong, Hongli Xu 0001, Baoyi An 0002, Gangyi Luo |
IEEE Trans. Netw. | 7 |
| 2025 | Fossil: A Cost-Effective and Fault-Tolerant Task Placement Scheme for Geo-Distributed Clouds
Gongming Zhao, Baoqing Wang, Jiawei Liu 0007, Hongli Xu 0001, Gangyi Luo |
ICA3PP (2) | 6 |
| 2025 | TAIR: Achieving Tenant Anomaly Isolation with Request Scheduling in Serverless Computing
Junhong Lu, Gongming Zhao, Hongli Xu 0001, Gangyi Luo |
NPC (1) | 5 |
| 2025 | Joint Optimization of Computation and Communication Resources for GPU Allocation in Heterogeneous Clusters
Gongming Zhao, Hongli Xu 0001, Gangyi Luo |
NPC (1) | 5 |
| 2025 | Toward auction-based edge AI: Orchestrating and incentivizing online transfer learning in edge networks
Yang Chen 0001, Lei Jiao 0002, Tuo Cao, Ji Qi 0005, Gangyi Luo, Sheng Zhang 0001, Sanglu Lu, Zhuzhong Qian |
Comput. Networks | 5 |
| 2025 | CPN meets learning: Online scheduling for inference service in Computing Power Network
Mingtao Ji, Ji Qi 0005, Lei Jiao 0002, Gangyi Luo, Hehan Zhao, Xin Li 0017, Zhuzhong Qian |
Comput. Networks | 4 |
| 2025 | ABUV: Adaptive bitrate and upsampling for video streaming on mobile devices
Ji Qi 0005, Sheng Zhang 0001, Gangyi Luo, Andong Zhu 0001, Jie Wu 0001, Zhuzhong Qian |
Comput. Networks | 4 |
| 2025 | Provisioning high precision edge inference with runtime model reconfiguration
Hesheng Sun, Zhuzhong Qian, Andong Zhu 0001, Sheng Zhang 0001, Sanglu Lu, Gangyi Luo |
Comput. Networks | 7 |
| 2025 | Machine-Centric High-Accuracy Multi-Video Analytics With Adaptive Neural CodecsabstractIncreased videos captured by widely deployed cameras are being analyzed by computer vision-based Deep Neural Networks (DNNs) on servers rather than being streamed for humans. Unfortunately, the conventional codecs (e.g., H.26x and MPEG-x) originally designed for video streaming lack content-aware feature extraction and hinder machine-centric video analytics, making it difficult to achieve the required high accuracy with tolerable delay. Neural codecs (e.g., autoencoder) now hold impressive compression performance and have been widely advocated in video streaming. While autoencoder shows transformative potential, the application in video analytics is hampered by low accuracy in detecting small objects of high-resolution videos and the serious challenges posed by multi-video streaming. To this end, we propose AdaStreamer with adaptive neural codecs to enable real machine-centric high-accuracy multi-video analytics. We also investigate how to achieve optimal accuracy under delay constraints via careful scheduling in Compression Ratios (CRs, the ratio of the compressed size to the original data size) and bandwidth allocation, and further propose a Markov-based Adaptive Compression and Bandwidth Allocation algorithm (MACBA). We have practically developed a prototype of AdaStreamer, based on which extensive experiments verify its accuracy improvement (up to 15%) compared to state-of-the-art coding and streaming solutions. Andong Zhu 0001, Ji Qi 0005, Sheng Zhang 0001, Gangyi Luo, Xiaohang Shi 0001, Zhuzhong Qian, Sanglu Lu |
IEEE Trans. Netw. | 4 |
| 2024 | PTMGS: A Cost-Optimal LLM Training Tasks Migration MethodabstractComputing power infrastructure has a significant impact on the training speed and cost during the training process of Large Language Models (LLMs). Due to differences in GPU hardware, network architecture, and training frameworks, the cost and unit price of computing power fluctuate greatly, and even within the same cloud service provider, there are huge differences in prices for different clusters. The feature of regularly saving checkpoints during the training process of large language models makes it possible to dynamically migrate tasks between GPU clusters during the training process. Based on this background, this article proposes a set of greedy strategy heuristic algorithms with the goal of minimizing training costs, which choose the best time and target for task migration, thereby achieving overall optimization of training costs. Through multiple rounds of experimental verification on real task samples and simulation data, the results show that the proposed greedy strategy based preemptive algorithm can reduce the overall training cost by more than 10% . This study provides new solutions and methods for efficient computing power supply for large language model training, which can help reduce training costs for computing infrastructure operators and customers. Gangyi Luo, Siyu He, Genning Zhang, Zhuzhong Qian |
HPCC | 1 |
| 2024 | Low-Carbon Geographically Distributed Cloud-Edge Task Scheduling
Yingjie Zhu, Ji Qi 0005, Shengjie Wei, Tuo Cao, Gangyi Luo, Zhuzhong Qian |
ICA3PP (6) | 7 |
| 2024 | Online Scheduling of Federated Learning with In-Network Aggregation and Flow RoutingabstractContinuously orchestrating in-network model aggregations for federated learning faces fundamental challenges such as the combinatorial nature of traffic reduction, the dynamic trade-offs between system overhead and model convergence, and the unpredictable inputs from uncertain system environments. In this work, we model a nonlinear mixed-integer program to optimize the long-term total cost of federated learning computation overhead, traffic reduction, network delay, and programmable switch reconfigurations over time. To attack the lexicographic minimax, submodular, and online nature of this problem, we propose a polynomial-time algorithmic framework to judiciously designate the timing of reconfigurations, while designing and invoking a linearized transformation for selecting routing paths, a greedy sub-algorithm for selecting aggregation locations, and an online learning sub-algorithm for controlling federated learning convergence. We demonstrate our rigorous mathematical insights behind our algorithms, and prove the competitive ratio as the performance guarantee. Using trace-driven evaluations, we have validated our approach's superiority over existing methods. Mingtao Ji, Lei Jiao 0002, Yitao Fan, Yang Chen 0001, Zhuzhong Qian, Ji Qi 0005, Gangyi Luo |
SECON | 7 |
| 2018 | Improving performance by network-aware virtual machine clustering and consolidation
Gangyi Luo, Zhuzhong Qian, Mianxiong Dong, Kaoru Ota, Sanglu Lu |
J. Supercomput. | 1 |
| 2014 | Network-Aware Re-Scheduling: Towards Improving Network Performance of Virtual Machines in a Data Center
Gangyi Luo, Zhuzhong Qian, Mianxiong Dong, Kaoru Ota, Sanglu Lu |
ICA3PP (1) | 1 |