VLDB 2026 Research / reviewers in the wild / expert
Chengxi Gao
dblp:138/4338
· DBLP profile ↗
34ranked-venue papers
5as first author
30since 2021 · last 2026
0000-0003-1386-7394ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 18 · 1 first-author · 17 since 2021Systems, architecture and hardware · 10 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ConnSched: Selective Connection Offloading Framework for Accelerating Stateful NFs with DPU
Fuliang Li, Chengxi Gao, Man Hou, Jiaxing Shen |
INFOCOM | 4 |
| 2026 | DistSRE: Synthesizing Runtime Executions with Automated Co-Optimization for Distributed LLM Training
Xiuzhu Sha, Chenyang Hei, Fuliang Li, Chengxi Gao, Rongfei Zeng, Xingwei Wang 0001 |
IWQoS | 4 |
| 2026 | UPServe: Backend Agnostic Proxy for Black-box Heterogeneous LLM Scheduling
Haorui Wan, Chenyang Hei, Fuliang Li, Chengxi Gao, Yuhan Jia, Tongrui Liu, Xingwei Wang 0001 |
IWQoS | 4 |
| 2026 | HeteCCL: Synthesizing Near-Optimal Collective Communication Schedules for Heterogeneous GPU Clusters
Chenyang Hei, Jiamin Cao, Chengxi Gao, Xiuzhu Sha, Tongrui Liu, Dengke Zhang, Ennan Zhai, Xingwei Wang 0001 |
NSDI | 4 |
| 2026 | C-Koordinator: Interference-Aware Management for Large-Scale and Co-Located Microservice ClustersabstractABSTRACT Objective Microservices transform traditional monolithic applications into lightweight, loosely coupled application components and have been widely adopted in many enterprises. Cloud platform infrastructure providers enhance the resource utilization efficiency of microservices systems by co‐locating different microservices. However, this approach also introduces resource competition and interference among microservices. Designing interference‐aware strategies for large‐scale, co‐located microservice clusters is crucial for enhancing resource utilization and mitigating competition‐induced interference. These challenges are further exacerbated by unreliable metrics, application diversity, and node heterogeneity. Methods In this paper, we first analyze the characteristics of large‐scale and co‐located microservices clusters at Alibaba and further discuss why cycle per instruction (CPI) is adopted as a metric for interference measurement in large‐scale production clusters, as well as how to achieve accurate prediction of CPI through multi‐dimensional metrics. Based on CPI interference prediction and analysis, we also present the design of the C‐Koordinator platform, an open‐source solution utilized in Alibaba cluster, which incorporates co‐location and interference mitigation strategies. Results The interference prediction models consistently achieve over 90.3% accuracy, enabling precise prediction and rapid mitigation of interference in operational environments. As a result, application latency is reduced and stabilized across all percentiles (P50, P90, P99) response time (RT), achieving improvements ranging from 16.7% to 36.1% under various system loads compared with state‐of‐the‐art system. Conclusion These results demonstrate the system's ability to maintain smooth application performance in co‐located environments. Shengye Song, Minxian Xu, Chengxi Gao, Fansong Zeng, Kejiang Ye, Cheng-Zhong Xu 0001 |
Softw. Pract. Exp. | 4 |
| 2026 | ConfigTransLE: Large Language Models Enhanced Network Configuration Translation
Fuliang Li, Naigong Zheng, Bocheng Liang, Yu Yang 0012, Chengxi Gao, Xingwei Wang 0001, Jiannong Cao 0001 |
IEEE Trans. Netw. | 6 |
| 2025 | Canvas: Scalable Collective Communication Scheduling for Large-Scale GPU ClustersabstractState-of-the-art deep learning models rely on large GPU clusters and various parallelism strategies, which in turn depend on collective communication (CC) operators to synchronize data. While vendor libraries (e.g., NCCL, RCCL) provide standard CC algorithms, they often suffer from bandwidth bottlenecks in imbalanced topologies. Recent synthesis-based methods improve performance but face three key limitations: poor scalability due to the combinatorial explosion of scheduling space, lack of support for multistage execution, and suboptimal communication throughput. We propose Canvas, a scalable and near-optimal CC scheduling framework that addresses these challenges. Canvas introduces: (1) Hierarchical synthesis to decompose the global scheduling problem into tractable subproblems for scalability. (2) Collective decomposition to enable structured, multi-stage algorithm generation. (3) Cross-micro-batch pipeline scheduling to parallelize communication across micro-batches and maximize link utilization. Evaluations show that Canvas achieves up to 1.98× bandwidth speedup over TACCL and 3.56× over TE-CCL, and synthesizes algorithms for 512-GPU topologies within 1.77 hours, whereas TACCL fails to produce results within 24 hours. Chenyang Hei, Fuliang Li, Chengxi Gao, Tongrui Liu, Xiuzhu Sha, Xingwei Wang 0001 |
ICNP | 4 |
| 2025 | TuCCL: Tailored and Unified Configuration Optimizations for High-Performance Collective Communication LibraryabstractModern distributed training systems face escalating communication bottlenecks as GPU clusters scale to accommodate large models. While collective communication libraries and automated synthesizers address algorithmic efficiency, they suffer from three critical limitations including labor-intensive manual intervention requirement, overreliance on predefined input optimization, and suboptimal isolated configuration optimization. To solve these problems, we present TuCCL, a systematic framework that co-optimizes communication algorithms and runtime parameters through three innovations: Topology-Aware Sketch Generation that automatically produces high-performance primitives, Hierarchical Configuration Optimization modeling nonlinear parameter-performance relationships, and Multi-phase Resource-Aware Configuration Optimization enabling joint configuration tuning with adaptive search space pruning. Evaluations demonstrate TuCCL’s superiority over state-of-the-art systems with 1.75x–11.49x bandwidth improvements for AllGather/AllReduce on NVIDIA V100/A100 clusters, 90.2% faster configuration search than grid methods, and 1.22x-2.52x end-to-end training speedups across diverse model scales. Chenyang Hei, Fuliang Li, Tongrui Liu, Chengxi Gao, Xiuzhu Sha, Xingwei Wang 0001 |
ICNP | 5 |
| 2025 | MAZ3: Memory-Assisted ZeRO-3 for Efficient Collective CommunicationabstractLarge Language Models (LLMs) have advanced rapidly, but their growing parameter scales and memory demands pose critical challenges for distributed training. Although GPU memory capacity improves steadily, model sizes expand much faster, causing frequent Out-of-Memory (OOM) errors and rising training costs. Existing memory optimization approaches, such as ZeRO-3 and offloading, alleviate per-GPU memory pressure but introduce excessive collective communication, limited computation–communication overlap, and degraded scalability. We present MAZ3, a distributed training framework that mitigates these limitations through three key techniques: (1) Collaborative CPU–GPU memory management, storing full parameters in CPU memory and broadcasting them within nodes to reduce global synchronization; (2) Fine-grained communication–computation overlap, aligning collective operations with model computation to hide latency; (3) Hierarchical aggregation operators, leveraging intra-node NVLink and inter-node NIC channels concurrently to minimize communication overhead. We implement MAZ3 on a multi-GPU cluster and evaluate it with large-scale models. Results show that MAZ3 reduces inter-node model communication (gradients and parameters) by 33%, improves training efficiency by 40.3%, and increases throughput by 67.9% compared with ZeRO-3. Moreover, MAZ3 retains the memory efficiency of ZeRO-3 while approaching the training efficiency and throughput of ZeRO-2 Offload, achieving a balanced trade-off between memory optimization and performance. Chenyang Hei, Fuliang Li, Chengxi Gao, Xingwei Wang 0001 |
ICNP | 4 |
| 2025 | SmartTC: A Real-Time ML-Based Traffic Classification with SmartnicabstractReal-time network traffic classification plays a crucial role in ensuring Quality of Service and network security, and machine learning (ML) based methods achieve high classification accuracy but induce significant computational overhead. While SmartNIC solutions can offload classification tasks thus reducing CPU burdens, they still suffer from various limitations, manifested in limited computing capabilities, insufficiency in dynamic load handling and high latency from heterogeneous computing architectures. To solve these problems, we propose SmartTC, with three key designs: (1) SmartTC employs hardware-software co-design to optimize SmartNIC processing power, (2) SmartTC adopts a trafficaware dynamic batch submission strategy that adjusts submission policies based on real-time network load, and (3) SmartTC proposes parallel pipeline scheduling that ensures efficient task execution while minimizing communication overhead. Finally, we implement SmartTC on the BlueField-3 DPU and conduct extensive experiments for evaluations, and comparison results demonstrate that SmartTC significantly outperforms existing solutions. For example, it reduces average traffic classification time by up to$\mathbf{1 6. 8 \%}$under low loads and$\mathbf{9 0. 9 \%}$under high loads. Besides, SmartTC does not affect Bluefield-3 network services, and saves host CPU usage by at least two cores. Lingxiang Hu, Chenyang Hei, Fuliang Li, Chengxi Gao, Jiaxing Shen, Xingwei Wang 0001 |
IWQoS | 4 |
| 2025 | ConfAgent: Towards Intelligent Network Configuration Via LLM AgentabstractAs network scale and complexity continue to increase, managing network configurations has become an increasingly challenging task. Existing configuration tools often depend on low-level, abstract intermediate representations, which require users to have substantial technical expertise. This reliance not only increases the learning curve but also heightens the risk of configuration errors. Recent advances in Large Language Models (LLMs) have demonstrated strong potential for automating tasks across various domains. However, their applications to network configuration generation remain limited due to several challenges, including hallucination, restricted context length, and insufficient adaptability to domain-specific requirements. To address these issues, we propose ConfAgent, an advanced network configuration generation system powered by a multi-model intelligent agent. ConfAgent comprises four key components: a conflict detector, an information extractor, a routing algorithm coder, and a formal synthesizer. These components collaborate to accurately interpret complex configuration intents, detect potential conflicts, and generate robust code and network configurations through intuitive natural language interactions. Extensive experiments conducted on the NetConfEval benchmark demonstrate that ConfAgent consistently outperforms existing state-of-the-art methods by margins ranging from 36 % to 100 %, particularly excelling in configuration tasks for large-scale network topologies. Shaowei Li, Zhiwen Gan, Jinyao Liu, Chengxi Gao, Fuliang Li, Si Wu 0003, Pengfei Hu 0001, Feng Li 0002 |
IWQoS | 4 |
| 2025 | ResCCL: Resource-Efficient Scheduling for Collective CommunicationabstractAs distributed deep learning training (DLT) systems scale, collective communication has become a significant performance bottleneck. While current approaches optimize bandwidth utilization and task completion time, existing communication libraries (CCLs) backends fail to efficiently manage GPU resources during algorithm execution, limiting the performance of advanced algorithms. This paper proposes ResCCL, a novel CCL backend designed for Resource-Efficient Scheduling to address key limitations in current systems. ResCCL enhances execution efficiency by optimizing scheduling at the primitive level (e.g., send and recvReduceCopy), enabling flexible thread block (TB) allocation, and generating lightweight communication kernels to minimize runtime overhead. Our approach tackles the global scheduling problem, reduces idle TB resources, and enhances communication bandwidth. Evaluation results demonstrate that ResCCL achieves up to 2.5× improvement in bandwidth performance compared to both NCCL and MSCCL. It reduces SM resource overhead by 77.8% and increases TB utilization by 41.6% while running the same algorithms. In end-to-end DLT, ResCCL boosts Megatron's throughput by up to 39%. Tongrui Liu, Chenyang Hei, Fuliang Li, Chengxi Gao, Jiamin Cao, Ennan Zhai, Xingwei Wang 0001 |
SIGCOMM | 4 |
| 2025 | Distributed learning-based context-aware SFC deployment in the Artificial Intelligence of Things
Wenlin Cheng, Xingwei Wang 0001, Fuliang Li, Bo Yi 0002, Qiang He 0002, Chuangchuang Zhang, Chengxi Gao, Min Huang 0001 |
Comput. Commun. | 7 |
| 2025 | SCC: Synchronization Congestion Control for Multi-Tenant Learning Over Geo-Distributed CloudsabstractDistributed machine learning over geo-distributed clouds enables joint training of data located in different regions, alleviating the burden of transferring large volumes of training datasets, which greatly saves bandwidth. However, the limited capacity of WAN links slows down the inter-cloud communications, which significantly decelerates the synchronization of distributed machine learning over geo-distributed clouds. Besides, the multi-tenancy in clouds results in multiple training tasks running simultaneously, whose synchronizations consistently compete for the limited WAN bandwidth with each other, which further aggravates the training performance of each task. While existing works optimize synchronizations through techniques like gradient compression, multi-resource interleaving and so on, none of them targets at the synchronization congestion especially due to multi-tenant learning, which results in inferior training performance.To solve these problems, we propose a simple but effective scheme, SCC, for fast and efficient multi-tenant learning via synchronization congestion control. SCC monitors the cross-cloud network conditions and evaluates the synchronization congestion level based on the round-trip transmission time for each synchronization. Then SCC alleviates synchronization congestion via controlling the synchronization frequency according to the synchronization congestion level in a probabilistic way. Extensive experiments are conducted within our testbeds consisted of 16 NVIDIA V100 GPUs to evaluate the performance of SCC, and comparison results show that SCC can reduce the average training completion time and makespan by up to 28.6% and 43.2% over SAP-SGD [1]. Targeted experiments are conducted to demonstrate the effectiveness and robustness of SCC. Chengxi Gao, Fuliang Li, Kejiang Ye, Yang Wang 0006, Pengfei Wang 0013, Xingwei Wang 0001, Cheng-Zhong Xu 0001 |
IEEE Trans. Computers | 1 |
| 2024 | Labeled graph partitioning scheme for distributed edge caching
Pengfei Wang 0013, Geng Sun 0001, Changjun Zhou, Chengxi Gao, Sen Qiu, Tiwei Tao, Qiang Zhang 0008 |
Future Gener. Comput. Syst. | 5 |
| 2024 | Distributed Program Deployment for Resource-Aware Programmable SwitchesabstractProgrammable switches allow data plane to program how packets are processed, which enables flexibility for network management tasks, e.g., packet scheduling and flow measurement. Existing studies focus on program deployment at a single switch, while deployment across the whole data plane is still a challenging issue, especially manifested in the difficulty in joint correct implementation of P4 programs, resource load balancing of network devices, and optimization of network performance. In this paper, we present RED, a Resource-Efficient and Distributed program deployment solution for programmable switches. First of all, we analyze data plane programs to estimate the resource utilization and divide them into two categories for further processing. Then, the proposed merging and splitting algorithms are selectively applied to merge or split the pending programs. Finally, we consolidate the scarce resources of the whole data plane for distributed program deployment. Extensive experiments with both testbed and large-scale simulations are conducted and comparison results show that 1) RED achieves network-wide resource balancing in a distributed way and the latency of processing packets within the switch was reduced by 16.7%. 2) RED improves the speedup by two orders of magnitude compared to P4Visor in merging program and merges more 18% tables than SPEED; 3) If the resources required to run a P4 program exceed the resource limit of the switch, it cannot be deployed on the switch. RED makes the overwhelmed programs to be deployed on switches and switch throughput increased by 10.7%. Fuliang Li, Xingxin Jia, Chengxi Gao, Pengfei Wang 0013, Xingwei Wang 0001, Jiannong Cao 0001 |
IEEE Trans. Computers | 4 |
| 2024 | Exploring Intercity Mobility in Urban Agglomeration: Evidence from Private Car Trajectory DataabstractIn this article, we explore intercity mobility in urban agglomerations by surveying people traveling across cities based on private car trajectory data. Specifically, we first adopt the statistical analysis method to mine the intercity mobility in terms of various metrics of travel trips, so as to gain a preliminary understanding of intercity mobility in urban agglomeration. Then, we utilize the tensor decomposition method to conduct in-depth study on the intercity mobility pattern from the perspectives of complexity and multidimensionality. We construct a 4-D tensor based on private car trajectory and point-of-interest (POI) datasets and define the functional similarity and geographic adjacency between regions. Finally, we design an alternating proximal gradient (APG)-based method to resolve the core tensor and factor matrix, leading to the fine-grained discovery of intercity mobility patterns on administrative divisions in the urban agglomeration. Extensive experiments are conducted to evaluate the analysis of intercity mobility, using a real-world dataset containing one-year private car trajectories from five cities in the selected urban agglomeration. The experiments show that the proposed method successfully captures 20 intercity mobility patterns, in which the factor matrices retrieve the patterns from different dimensions with core tensors characterizing correlations between patterns in factor matrices. Besides, the extracted intercity mobility patterns not only cover administrative areas with frequent intercity interactions, but also contain areas with less intercity interactions. It validates that the intercity mobility is consistent with the regional functions in urban agglomeration. Zhu Xiao, Linshan Wu, Hongbo Jiang 0001, Zheng Qin 0001, Chengxi Gao, Hongyang Chen 0001, Jiangchuan Liu |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2024 | Heter-Train: A Distributed Training Framework Based on Semi-Asynchronous Parallel Mechanism for Heterogeneous Intelligent Transportation SystemsabstractTransportation big data (TBD) are increasingly combined with artificial intelligence to mine novel patterns and information due to the powerful representational capabilities of deep neural networks (DNNs), especially for anti-COVID19 applications. The distributed cloud-edge-vehicle training architecture has been applied to accelerate DNNs training while ensuring low latency and high privacy for TBD processing. However, multiple intelligent devices (e.g., intelligent vehicles, edge computing chips at base stations) and different networks in intelligent transportation systems lead to computing power and communication heterogeneity among distributed nodes. Existing parallel training mechanisms perform poorly on heterogeneous cloud-edge-vehicle clusters. The synchronous parallel mechanism may force fast workers to wait for the slowest worker for synchronization, thus wasting their computing power. The asynchronous mechanism has communication bottlenecks and can exacerbate the straggler problem, causing increased training iterations and even incorrect convergence. In this paper, we introduce a distributed training framework, Heter-Train. First, a communication-efficient semi-asynchronous parallel mechanism (SAP-SGD) is proposed, which can take full advantage of acceleration effect of asynchronous strategy on heterogeneous training and constrain the straggler problem by using global interval synchronization. Second, Considering the difference in node bandwidth, we design a solution for heterogeneous communication. Moreover, a novel weighted aggregation strategy is proposed to aggregate the model parameters with different versions. Finally, experimental results show that our proposed strategy can achieve up to$6.74 \times $speedups on training time, with almost no accuracy decrease. Jiawei Geng, Haipeng Jia, Zongwei Zhu, Hai Fang, Chengxi Gao, Cheng Ji 0002, Gangyong Jia, Guangjie Han, Xuehai Zhou |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | Localizing From Classification: Self-Directed Weakly Supervised Object Localization for Remote Sensing ImagesabstractIn recent years, object localization and detection methods in remote sensing images (RSIs) have received increasing attention due to their broad applications. However, most previous fully supervised methods require a large number of time-consuming and labor-intensive instance-level annotations. Compared with those fully supervised methods, weakly supervised object localization (WSOL) aims to recognize object instances using only image-level labels, which greatly saves the labeling costs of RSIs. In this article, we propose a self-directed weakly supervised strategy (SD-WSS) to perform WSOL in RSIs. To specify, we fully exploit and enhance the spatial feature extraction capability of the RSIs' classification model to accurately localize the objects of interest. To alleviate the serious discriminative region problem exhibited by previous WSOL methods, the spatial location information implicit in the classification model is carefully extracted by GradCAM++ to guide the learning procedure. Furthermore, to eliminate the interference from complex backgrounds of RSIs, we design a novel self-directed loss to make the model optimize itself and explicitly tell it where to look. Finally, we review and annotate the existing remote sensing scene classification dataset and create two new WSOL benchmarks in RSIs, named C45V2 and PN2. We conduct extensive experiments to evaluate the proposed method and six mainstream WSOL methods with three backbones on C45V2 and PN2. The results demonstrate that our proposed method achieves better performance when compared with state-of-the-arts. Jing Bai 0003, Junjie Ren, Zhu Xiao, Zheng Chen 0021, Chengxi Gao, Talal Ahmed Ali Ali, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Learning-Based Sketch for Adaptive and High-Performance Network MeasurementabstractWith the development of network measurement technologies, a hybrid measurement architecture can effectively optimize the sketch structure in switches, making it more adaptable to the current complex and volatile network environment. However, current optimization technologies based on hybrid measurement architectures generally suffer from insufficient automation, difficulty of learning effective numerical features, and lack of generality, resulting in poor scalability in real deployment. To solve these problems, we propose theTalentSketchframework, based on which we further developDeepSketchfor effective sketch optimization. First, we useSeq2Seqto automatically identify target flows instead of relying on manual thresholds. Second, we propose a new training strategy that extracts low-precision flows for models with weak learning capabilities. Last, we develop a new sketch optimization framework that can optimize different kinds of sketches only by changing the training data for generality. A large number of experimental results show thatDeepSketchexhibits superior performance. For example: (1) the accuracy of optimized sketches has increased by 20% to 73%, (2) Without replacing the model structure, the accuracy of the optimized sketches can generally reach over 80%. (3) The impact of low sampling rates on accuracy is less than 1% on various sketches. Fuliang Li, Yiming Lv, Yangsheng Yan, Chengxi Gao, Xingwei Wang 0001, Jiannong Cao 0001 |
IEEE/ACM Trans. Netw. | 4 |
| 2024 | Efficient Multi-Task Computation Offloading Game for Mobile Edge ComputingabstractMobile edge computing emerges to serve mobile users with low-latency computation offloading in edge networks, which are resource-constrained with massive users and workloads. However, existing communication and computing resource allocation schemes for offloaded tasks aren't efficient enough, where finished tasks still occupy resources, wasting constrained resources. Besides, the multi-user offloading is usually for scenarios of one task per user, ignoring real-worldmulti-taskoffloading scenarios where each user has multiple tasks, lack generality and flexibility. Meanwhile, local computing resource allocation schemes in multi-task scenarios ignore resource readjustment, causing low resource utilization. To solve these problems, we propose ECO-GAME, an efficient multi-task offloading scheme, which dynamically allocates bandwidth and computing resources to unfinished tasks, resulting in high resource utilization. We initially formulate the multi-task offloading problem as the game minimizing each user's cost, which is NP-hard. Thus we re-formulate the game utilizing potential games to optimize user's objective either locally or globally, and prove the existence of its Nash equilibrium. We then design an efficient multi-task offloading algorithm to obtain an approximate solution in polynomial time, together with computational complexity analysis. We further conduct performance evaluation on ECO-GAME utilizing price of anarchy. Numerical results demonstrate the efficiency of ECO-GAME, and show ECO-GAME reduces 49.2% cost over the state-of-the-art work, and scales well with the increasing number of tasks and users. Shuhui Chu, Chengxi Gao, Minxian Xu, Kejiang Ye, Zhu Xiao, Cheng-Zhong Xu 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2023 | RED: Distributed Program Deployment for Resource-aware Programmable Switches
Xingxin Jia, Fuliang Li, Chengxi Gao, Pengfei Wang 0013, Xingwei Wang 0001 |
INFOCOM | 4 |
| 2023 | NetDiceSyn: Multi-Property Probabilistic Verification of Network ConfigurationsabstractIn recent years, probabilistic network analysis has received significant attention as a crucial aspect of network management. With the increasing use of distributed routing protocols in networks, computing the probabilities of multiple properties becomes a challenging task, as different property violations derive from different failure scenarios. While some previous work has attempted to prune the space of failure scenarios by identifying a set of links whose failure does not change whether a property holds, this approach often requires multiple calls to the system, leading to redundant computation when verifying multiple properties. To solve this problem, we propose NetDiceSyn, a scalable multi-property probabilistic network configuration analyzer. As verifying each property requires exploring a state tree, NetDiceSyn proposes to explore all state trees synchronously and compute a shared data plane for multiple states in different state trees, thus reducing the redundant computation between different state trees and verifying multiple properties simultaneously. Besides, we design a state sharing method to reduce the redundant storage of states during synchronous exploration. Additionally, we propose abstract topology to merge equivalent states within a state tree to speed up probabilistic verification in series link networks. Extensive experiments are conducted with real-world network topologies, and the results show that NetDiceSyn outperforms state-of-the-art methods, providing up to an order of magnitude speedup when verifying hundreds of properties for networks with hundreds of links. Renrui Liu, Fuliang Li, Chengxi Gao, Ce Ji, Xingwei Wang 0001 |
IWQoS | 3 |
| 2023 | Efficient congestion control scheme based on caching strategy in NDN
Dapeng Qu, Chengxi Gao, Haiying Shen, Keqin Li 0001 |
J. Netw. Comput. Appl. | 4 |
| 2023 | Flash: Joint Flow Scheduling and Congestion Control in Data Center NetworksabstractFlow scheduling and congestion control are two important techniques to reduce flow completion time in data center networks. While existing works largely treat them independently, the interactions between flow scheduling and congestion control are in general overlooked which leads to sub-optimal solutions, especially given that the link capacity is increasing faster than the switch port buffer size. In this paper, we presentFlash, a simple yet effective scheme that integrates scheduling and congestion control. Specifically,Flashputs forward a congestion-aware scheduling scheme to determine the priority of flows based on the latest network congestion extent and the flow’s bytes sent. Besides,Flashproposes a priority-based packet dropping scheme in switch port buffers and implements a priority-aware congestion control scheme. Experiment results show thatFlashhas superior performance: (1) it has 35.8% lower tail latency than PIAS and performs similar with pFabric in a 10G network without knowing the flow size, (2) in 100G networks with shallow buffers, the information agnosticFlashhas 6.8% lower average FCT than the information-aware pFabric, (3) it outperforms pFabric by 13.5% in FCT if flow size is also known toFlash. Chengxi Gao, Shuhui Chu, Hong Xu 0001, Minxian Xu, Kejiang Ye, Cheng-Zhong Xu 0001 |
IEEE Trans. Cloud Comput. | 1 |
| 2023 | Bottleneck-Aware Non-Clairvoyant Coflow Scheduling With FaiabstractCoflow scheduling is critical to data-parallel applications in data centers. While schemes like Varys can achieve optimal performance, they require a priori information about coflows which is hard to obtain in practice. Existing non-clairvoyant solutions like Aalo generalize least attained service (LAS) scheduling discipline to address this issue. However, they fail to identify the bottleneck flows in a coflow and tend to allocate excessive bandwidth to the non-bottleneck flows, leading to bandwidth wastage and inferior overall performance. To this end, we present Fai that strives to improve the overall coflow performance by accelerating the bottleneck flows without priori knowledge. Fai employs bottleneck-aware scheduling. It adopts loose coordination to update coflow priority and flow rates based on total bytes sent. In addition, Fai detects bottleneck flows based on a flow’s rate and bytes sent, and de-allocates bandwidth for other flows to match the bottleneck rate without affecting the coflow completion time (CCT). The saved bandwidth is then distributed among coflows according to their priority to improve overall performance. Testbed evaluation on a 40-node cluster shows that Fai improves average (P95) CCT by 1.73× (3.43×), compared to Aalo. Large-scale trace-driven simulations also show that Fai outperforms Aalo substantially. Libin Liu 0001, Chengxi Gao, Peng Wang 0037, Hongming Huang, Jiamin Li 0002, Hong Xu 0001, Wei Zhang 0049 |
IEEE Trans. Cloud Comput. | 2 |
| 2022 | $Q_{C}-DQN$: A Novel Constrained Reinforcement Learning Method for Computation Offloading in Multi-access Edge ComputingabstractIn recent years, multi-access edge computing (MEC) is emerging to provide computation and storage resources to the Internet of things (IoT) devices to assist them in high-performance demanding tasks. Real-time task requests from the IoT devices often have strict delay constraints. However, in practice, the delay requirements of task requests often fail to be satisfied because of the inappropriate computation processing method and inefficient resource allocation in MEC networks. In this article, we present a novel framework for MEC networks with unmanned aerial vehicles (UAVs) and intelligent reflecting surfaces (IRSs) to facilitate computation offloading with delay constraints. In addition, we propose a novel constrained reinforcement learning method with a dynamic balance mechanism named$Q_{c}-DQN$. Finally, we conduct extensive simulations to verify the effectiveness of our proposed method. Compared to the benchmark schemes, our scheme not only improves the overall network performance and reduces the task completion time, but also meets the delay constraints. Shen Zhuang, Chengxi Gao, Ying He 0006, F. Richard Yu, Yuhang Wang 0019, Weike Pan, Zhong Ming 0001 |
IJCNN | 2 |
| 2022 | When Multi-access Edge Computing Meets Multi-area Intelligent Reflecting Surface: A Multi-agent Reinforcement Learning ApproachabstractIn recent years, multi-access edge computing (MEC) is emerging to provide computation and storage capabilities to the Internet of things (IoT) devices to improve the quality of service (QoS) of IoT applications. In addition, intelligent reflecting surface (IRS) techniques have attracted great interests from both academia and industry to improve the communication efficiency. Although existing works leverage the IRS technique in MEC networks, they mainly focus on the single-IRS single-area scenario. However, in practice, multi-IRS will be deployed in multi-area scenarios in future networks. Consequently, considering the single-IRS single-area scenario will have inferior performance. In this paper, to address the aforementioned issue, we propose an efficient resource provisioning scheme for multi-IRS multi-area scenarios in MEC networks. We first model the problem as a cooperative multi-agent reinforcement learning process, where each agent manages one area and all agents share the network bandwidth and computation resources. Then, we propose a multi-agent actor-critic method with an attention mechanism for resource management with latency guarantee. Finally, we conduct extensive simulations to verify the effectiveness of the proposed scheme. Our scheme can reduce the required computation resources by up to 11.84% when compared with the benchmark works. It is also shown that our proposed scheme can improve the efficiency of resource allocation and scale well with the increasing demand from IoT devices. Shen Zhuang, Ying He 0006, F. Richard Yu, Chengxi Gao, Weike Pan, Zhong Ming 0001 |
IWQoS | 4 |
| 2022 | Efficient Multi-Channel Computation Offloading for Mobile Edge Computing: A Game-Theoretic ApproachabstractMobile edge computing is emerging to provide cloud-computing capabilities to mobile users, so that they can offload computation intensive tasks to close proximity for execution. However, most existing works imply that a transmission-finished task still occupies the channel until all users on the same channel finish the transmission, leading to severe channel resource waste. To solve this problem, we propose an efficient computation offloading mechanism which releases the channel resources of transmission-finished tasks for transmission-unfinished tasks, and aims to minimize the response time and energy consumption for each user. Specifically, we formulate the computation offloading problem as a game, analyze its structural properties and show how it possesses a Nash equilibrium and admits the finite improvement property, in the cases of elastic cloud and non-elastic cloud respectively. We then propose aDistributedMulti-channelComputationOffloading (DMCO) algorithm, which can converge to a Nash equilibrium, and find the upper bound of the convergence time. We further evaluate the performance of DMCO using the price of anarchy. Numerical results show that DMCO scales well with the number of users, and outperforms existing works, for example, benefits 13.3 percent more users and reduces cost by 23.7 percent than CO, one of the best existing works. Shuhui Chu, Zhiyi Fang, Shinan Song, Zhanyang Zhang, Chengxi Gao, Cheng-Zhong Xu 0001 |
IEEE Trans. Cloud Comput. | 5 |
| 2021 | D-SRTF: Distributed Shortest Remaining Time First Scheduling for Data Center NetworksabstractMany recent works utilize scheduling to minimize the Flow Completion Time (FCT) in Data Center Networks (DCN), like PIAS using Shortest Job First (SJF) scheduling and pFabric using Shortest Remaining Size First (SRSF) scheduling. However, they only consider the flow size information, without consideration of available bandwidth of the network, leading to inferior performance when the network is congested. Besides, information on flow size is hard to obtain in practice. Moreover, although a centralized scheduler may have optimal scheduling decisions, it suffers from high system overhead. Therefore, a new DCN scheme is expected which is deployment-friendly and implements SRTF scheduling in a distributed manner. In this paper, we propose D-SRTF, a light-weight yet effective DCN scheme to implement SRTF scheduling. D-SRTF determines the remaining time of each flow according to the estimated remaining flow size and the available bandwidth, in order to determine the priority of each flow. Switches perform Strict Priority (SP) scheduling according to the priority of each flow, in order to realize SRTF scheduling. Experiments show that D-SRTF performs better than the currently best implementable scheme, PIAS, and could perform better than pFabric if information on flow size is available. Chengxi Gao, Victor C. S. Lee, Keqin Li 0001 |
IEEE Trans. Cloud Comput. | 1 |
| 2020 | Energy Efficient Algorithms based on VM Consolidation for Cloud Computing: Comparisons and EvaluationsabstractCloud Computing paradigm has revolutionized IT industry and be able to offer computing as the fifth utility. With the pay-as-you-go model, cloud computing enables to offer the resources dynamically for customers anytime. Drawing the attention from both academia and industry, cloud computing is viewed as one of the backbones of the modern economy. However, the high energy consumption of cloud data centers contributes to high operational costs and carbon emission to the environment. Therefore, Green cloud computing is required to ensure energy efficiency and sustainability, which can be achieved via energy efficient techniques. One of the dominant approaches is to apply energy efficient algorithms to optimize resource usage and energy consumption. Currently, various virtual machine consolidation-based energy efficient algorithms have been proposed to reduce the energy of cloud computing environment. However, most of them are not compared comprehensively under the same scenario, and their performance is not evaluated with the same experimental settings. This makes users hard to select the appropriate algorithm for their objectives. To provide insights for existing energy efficient algorithms and help researchers to choose the most suitable algorithm, in this paper, we compare several state-of-the-art energy efficient algorithms in depth from multiple perspectives, including architecture, modelling and metrics. In addition, we also implement and evaluate these algorithms with the same experimental settings in CloudSim toolkit. The experimental results show the performance comparison of these algorithms with comprehensive results. Finally, detailed discussions of these algorithms are provided. Qiheng Zhou, Minxian Xu, Sukhpal Singh, Chengxi Gao, Wenhong Tian, Cheng-Zhong Xu 0001, Rajkumar Buyya |
CCGRID | 4 |
| 2016 | DEME: Decouple packet marking from enqueuing for multiple services in data center networksabstractMost of current Data Center Network (DCN) protocols leverage Explicit Congestion Notification (ECN) for congestion control. However, the majority of them assume single-queue scenario in each switch port, making their performance inferior in multiple-queue scenario. MQECN [1] solves this problem by periodically measuring the round time of queue scheduling, calculating a threshold for individual queue based on its weight and the measured round time, and adopting standard ECN in each queue. However, MQECN incurs non-negligible overhead for frequent round time measurement, and inaccurate round time measurement is unavoidable. To this end, we propose DEME, a light-weight DCN scheme for multiple-queue scenario with no need for round time measurement or per queue threshold setting. The core idea of DEME is to decouple packet marking from enqueuing, which means, when a packet is enqueued and the total queue length exceeds the standard threshold, instead of marking this newly arrived packet, we mark the head packet of the queue whose length exceeds its fair share the most. Experiments show that our light-weight DEME has similar performance with MQECN in terms of average Flow Completion Time and guarantees the fairness. Chengxi Gao, Victor C. S. Lee |
ICNP | 1 |
| 2015 | An Intelligent Economic Approach for Dynamic Resource Allocation in Cloud ServicesabstractWith Inter-Cloud, distributed cloud and open cloud exchange (OCX) emerging, a comprehensive resource allocation approach is fundamental to highly competitive cloud market. Oriented to infrastructure as a service (IaaS), an intelligent economic approach for dynamic resource allocation (IEDA) is proposed with the improved combinatorial double auction protocol devised to enable various kinds of resources traded among multiple consumers and multiple providers at the same time enable task partitioning among multiple providers. To make bidding and asking reasonable in each round of the auction and determine eligible transaction relationship among providers and consumers, a price formation mechanism is proposed, which is consisted of a back propagation neural network (BPNN) based price prediction algorithm and a price matching algorithm. A reputation system is proposed and integrated to exclude dishonest participants from the cloud market. The winner determination problem (WDP) is solved by the improved paddy field algorithm (PFA). Simulation results have shown that IEDA can not only help maximize market surplus and surplus strength but also encourage participants to be honest. Xingwei Wang 0001, Hao Che, Keqin Li 0001, Min Huang 0001, Chengxi Gao |
IEEE Trans. Cloud Comput. | 6 |
| 2013 | A Cloud Resource Allocation Mechanism Based on Mean-Variance Optimization and Double Multi-Attribution Auction
Chengxi Gao, Xingwei Wang 0001, Min Huang 0001 |
NPC | 1 |