Waixi Liu 0001

dblp:19/4199-1 · DBLP profile ↗
← Back
20ranked-venue papers
10as first author
14since 2021 · last 2026
0000-0002-7343-4948ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 15 · 9 first-author · 11 since 2021Systems, architecture and hardware · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 CCGS: A Cross-modal Collaborative Gradient Sparsification for Accelerating Distributed Multimodal Model Training
Shenghui Lu, Waixi Liu 0001, Jinhuang Huang, Qingchun Chen
Euro-Par (2)2
2026 Optimizing When, How, and What to Communicate in Shared ML Clusters
Tianxiang Huang, Waixi Liu 0001, Jun Cai 0002, Shipeng Fan
IPDPS2
2026 Adaptive gradient sparsification with layer and stage-wise for accelerating distributed DNN training
Waixi Liu 0001, Jun Cai 0002, Zhen-Xin Zhang, Kongyang Chen
Comput. Networks1
2026 DRL-BM: Intelligent buffer management in data center network
Waixi Liu 0001, Xin-Jian Zhong, Zhen-xin Zhang, Jun Cai 0002, Zhiquan Liu 0001, Chao-Xuan Zheng, Zhen-zheng Guo
Comput. Networks1
2024 QALL: Distributed Queue-Behavior-Aware Load Balancing Using Programmable Data Planes
abstract
Existing load-balancing methods used in data center networks involve some shortcomings such as excessively large decision delays during reactions to microbursts and large overheads involved in active probing. Programmable data planes have provided new opportunities for local decision-making on switches to address these issues. We observe that queue behavior (i.e., queue occupancy, queuing trend, and dequeue time interval) in switches can reflect the current or future congestion degree on a network. Furthermore, following data-driven experiments, we found an accurate fitting function of congestion degree to queue behavior. Thus, we propose an in-network load-balancing scheme based on a programmable switch, called queue-behavior-aware localized load balancing (QALL). In QALL, each switch independently selects egress ports probabilistically according to fine-grained-measured local queue behavior. The key concept of QALL is to take account the evolutionary process of reaching the current queue state into its decision basis for load balancing. Experimental results under actual DCN workloads (including web search and data mining workloads) demonstrate the effectiveness of QALL. In terms of flow completion time, decision delay, network shock, load sharing accuracy, and packet reordering, QALL outperformed recent perpacket (DRILL), per-flowlet (LetFlow and CONGA), and per-flow (ECMP) load balancers, particularly under heavy load. For example, under asymmetrical topology with 90% load level, the flow completion time of QALL was lower than that of ECMP, LetFlow, CONGA, and DRILL by up to 54.7%, 46.5%, 38.9%, and 18.9%, respectively.
Waixi Liu 0001, Jun Cai 0002, Sen Ling, Jian-Yu Zhang, Qingchun Chen
IEEE Trans. Netw. Serv. Manag.1
2023 CL-SGD: Efficient Communication by Clustering and Local-SGD for Distributed Machine Learning
abstract
Training a deep neural network model requires frequent communications between machines, and heavy communication traffic limits the scalability of the distributed ma-chine learning training. Some works try to reduce the communication traffic by transmitting the clustered gradients. Howe-er, the granularity of gradient clustering in these works is relatively coarse, which may decrease the accuracy and stability of the model. Moreover, our experiments reveal that a type of gradient have a certain degree of correlation, which means that we should cluster the gradients by a fine-grained way. In this article, we propose Cluster and Local Stochastic Gradient De-scent (CL-SGD) scheme, which combines the type-by-type gradient clustering method with local training scheme under the master-slave node architecture. CL-SGD has two key designs: first, fully taking account of the differences from each type of gradient, we propose the type-by-type gradient clustering method, which separately clusters each type of gradient meanwhile combining the local training scheme, to significantly reduce the communication traffic. Second, we use master-slave node architecture to reduce the model accuracy loss caused by clustering gradient. Experiment results show that CL-SGD achieves 1500x compression ratio and reduces training time by up to 51 % than the BSP, Local-SGD, STL-SGD, ClusterGrad.
Haosen Chen, Waixi Liu 0001, Miaoquan Tan, Junming Luo
ICC2
2023 Load balancing inside programmable data planes based on network modeling prediction using a GNN with network behaviors
Waixi Liu 0001, Jun Cai 0002, Yinghao Zhu, Junming Luo, Jin Li 0002
Comput. Networks1
2022 Binary Neural Network with P4 on Programmable Data Plane
abstract
Deploying machine learning (ML) on the programmable data plane (PDP) has some unique advantages, such as quickly responding to network dynamics. However, compared to demands of ML, PDP have limited operations, computing and memory resources. Thus, some works only deploy simple traditional ML approaches (e.g., decision tree, K-means) on PDP, but their performance is not satisfactory. In this article, we propose P4-BNN (Binary Neural Network based on P4), which uses P4 to completely executes binary neural network on PDP. P4-BNN addresses some challenges. First, in order to use shift and simple integer arithmetic operations to replace multiplication, P4-BNN proposes a tailor-made data structure. Second, we use an equivalent replacement programming method to support matrix operation required by ML. Third, we propose a normalization method in PDP which needn't floating-point operations. Fourth, by using register storing the model parameters, the weights of P4-BNN model can be updated without interrupting the P4 program running. Finally, as two use-cases, we deploy P4-BNN on a Netronome SmartNIC (Agilio CX 2x10GbE) to achieve flow classification and anomaly detection. Compared to the N3IC, decision tree and K-means, the accuracy of P4-BNN has 1.7%, 3.4% and 47.7% improvement respectively.
Junming Luo, Waixi Liu 0001, Miaoquan Tan, Haosen Chen
MSN2
2022 Adaptive synchronous strategy for distributed machine learning
abstract
In distributed machine learning training, bulk synchronous parallel (BSP) and asynchronous parallel (ASP) are two main synchronization methods to help achieve gradient aggregation. However, BSP needs longer training time due to “stragglers” problem, while ASP sacrifices the accuracy due to “gradient staleness” problem. In this article, we propose a distributed training paradigm on parameter server framework called adaptive synchronous strategy (A2S) which improves the BSP and ASP paradigms by adaptively adopting different parallel training schemes for workers with different training speeds. Based on the stale value between the fastest and slowest worker, A2S adaptively adds a relaxed synchronous barrier for fast workers to alleviate gradient staleness, where a differentiated weighting gradient aggregation method is used to reduce the impact of slow gradients. Simultaneously, A2S adopts ASP training for slow workers to eliminate stragglers. Hence, A2S not only improves the “gradient staleness” and “stragglers” problems, but also obtains convergence stability and synchronous gain through synchronous and asynchronous parallel, respectively. Specially, we theoretically proved the convergence of A2S by deriving the regret bound. Moreover, experiment results show that A2S improves accuracy by up to 2.64% and accelerates training by up to 41% more than the state-of-the-art synchronization methods BSP, ASP, stale synchronous parallel (SSP), dynamic SSP, and Sync-switch.
Miaoquan Tan, Waixi Liu 0001, Junming Luo, Haosen Chen, Zhenzheng Guo
Int. J. Intell. Syst.2
2022 DRL-PLink: Deep Reinforcement Learning With Private Link Approach for Mix-Flow Scheduling in Software-Defined Data-Center Networks
abstract
In datacenter networks, bandwidth-demanding elephant flows without deadline and delay-sensitive mice flows with strict deadline coexist. They compete with each other for limited network resources, and the effective scheduling of such mix-flows is extremely challenging. We propose a deep reinforcement learning with private link approach (DRL-PLink), which combines the software-defined network and deep reinforcement learning (DRL) to schedule mix-flows. DRL-PLink divides the link bandwidth and establishes some corresponding private-links for different types of flows to isolate them such that the competition among different types of flows can decrease accordingly. DRL is used to adaptively and intelligently allocate bandwidth resources for these private-links. Furthermore, to improve the scheduling policy, DRL-PLink introduces the novel clipped double Q-learning, exploration with noise, and prioritized experience replay technology for DDPG to address function approximation error, to induce lager and more randomness for exploration, as well as more effective and efficient experience replay in DRL respectively. The experiment results under actual datacenter network workloads (including Web search and data mining workload) indicate that DRL-PLink can effectively schedule mix-flows at a small system overhead. Compared with ECMP, pFabric, and Karuna, the average flow completion time of DRL-PLink decreased by 77.79%, 65.61%, and 23.34% respectively, when the deadline meet rate is increased by 16.27%, 0.02%, and 0.836% respectively. Additionally, DRL-PLink can also well achieve load balance between paths.
Waixi Liu 0001, Jinjie Lu, Jun Cai 0002, Yinghao Zhu, Sen Ling, Qingchun Chen
IEEE Trans. Netw. Serv. Manag.1
2021 FullSight: Towards Scalable, High-Coverage, and Fine-grained Network Telemetry
abstract
A variety of network states can better help network operators to manage the whole network. However, the existing network measurement schemes still exhibit some drawbacks, such as excessive bandwidth overhead caused by running packet-level measurement, lack of coverage of variety measurement granularities, and occupying several switch’s memory. This paper presents the FullSight based on the programmable data plane, which provides fine multiple granularities measurement. Based on the programmability of data plane, this paper proposes an intelligent measurement mechanism that can adaptively adjust the measurement frequency according to the network state to greatly reduce the bandwidth overhead of measuring while ensuring a certain measurement accuracy and acceptable processing overhead. Also, the Rotating Memory scheme is proposed to reduce occupying memory of switch when achieving a variety of fine-grained measurements. The simulation results demonstrate the effectiveness of FullSight in terms of the bandwidth overhead reduction, the memory overhead reduction, full coverage of a variety of fine-grained network states. Compared with Netsight, FullSight only suffers from 0. 1% bandwidth overhead which is two orders of magnitude lower than Netsight, and FullSight has taken up no more than 0. 001% memory overhead for different measurement tasks.
Sen Ling, Waixi Liu 0001, Yinghao Zhu, Miaoquan Tan, Jieming Huang, Zhenzheng Guo, Wen-Hong Lin
MSN2
2021 Network Telemetry by Observing and Recording on Programmable Data Plane
abstract
Fine-grained, real-time, and accurate monitoring data can better help detect equipment failure and perform traffic engineering. However, existing in-band network telemetry (INT) implementations still exhibit a few drawbacks such as lack of real-time monitoring, relatively high overheads due to per-packet operation, and limited monitoring range. This paper proposes an INT+PDP-based fine-grained real-time telemetry scheme by observing and recording on the programmable data plane (PDP), referred to as O&R. The key idea lies in designing some registers on data plane to observe the states of packets forwarded by it as well as adding a customized header on a normal data packet to record how it is forwarded on its routing path. Except for measuring some conventional performance parameters such as end-to-end delay, jitter, throughput, and packet loss rate, O&R designs a clock offset elimination algorithm to realize the time synchronization of two adjacent switches, based on which we can complete more fine-grained measurement such as queuing delay, processing delay, transmission delay, and propagation delay on any hop. O&R also can measure the queue state that includes real-time queue depth and how many flows share the queue. Extensive experimental results for the K=4 fat-tree data-center network demonstrate the effectiveness of O&R in terms of higher accuracy, better real-time performance, less overheads, and better fine-graining compared to existing schemes. The measurement accuracy of O&R is 46.3% higher than that of INT-like method. The measurement delay of O&R is ~1 ms, while INT-like method needs ~20 ms. The measurement overhead of O&R is only 2.19% of Pingmesh.
Wen-Hong Lin, Waixi Liu 0001, Gui-Feng Chen, Jin-Jiang Fu, Xing Liang, Sen Ling, Zhitao Chen
Networking2
2021 DRL-R: Deep reinforcement learning approach for intelligent routing in software-defined data-center networks
Waixi Liu 0001, Jun Cai 0002, Qing Chun Chen, Yu Wang 0017
J. Netw. Comput. Appl.1
2021 APPM: Adaptive Parallel Processing Mechanism for Service Function Chains
abstract
By replacing traditional hardware-based middleboxes with software-based Virtual Network Functions (VNFs) running on general-purpose servers, network function virtualization represents a promising technique to reduce the cost of service creation and increase the agility of network operations. Typically, Service Function Chains (SFCs) are adopted to orchestrate dynamical network services and facilitate management of network applications. Recently, SFC parallelism that implements parallel processing of VNFs has been investigated to further improve SFC service quality. However, the unreasonable service graph of parallel processing in existing parallelized SFCs (PSFCs) might cause excessive resource consumption; incoordination between PSFC deployment and scheduling also increases the queuing delay of VNFs and degrades PSFC performance. In this article, an adaptive parallel processing optimization mechanism (APPM) is proposed to self-adaptively adjust the service graph of PSFCs and intelligently solve the joint problem of PSFC deployment and scheduling. Specifically, APPM uses a parallelism optimization algorithm (POA) based on the bin packing problem with soft bin capacity to optimize the structure of the PSFC service graph. Afterward, APPM employs a joint optimization algorithm based on reinforcement learning (JORL) to jointly deploy and schedule the PSFCs optimized by POA via the online perception of environment status. Simulation results showed that POA reduces the SFC parallelism degree and resource consumption by about 35%; JORL lowers SFC delay by reducing the queuing delay and has better overall performance than the state of the art algorithms even with limited resources.
Jun Cai 0002, Zhongwei Huang, Liping Liao, Jian-Zhen Luo, Waixi Liu 0001
IEEE Trans. Netw. Serv. Manag.5
2020 Scheduling mix-flow in SD-DCN based on Deep Reinforcement Learning with Private Link
abstract
In software-defined datacenter networks, there are bandwidth-demanding elephant flows without deadline and delay-sensitive mice flows with strict deadline. They compete with each other for limited network resources, and how to effectively schedule such mix-flow is a huge challenge. We propose DRL-PLink (deep reinforcement learning with private link) that combines software-defined network and deep reinforcement learning (DRL) to schedule mix-flow. It divides the link bandwidth and establishes some corresponding private links for different types of flows respectively to isolate them. DRL is used to adaptively allocate bandwidth resources for these private links. Furthermore, DRL-PLink introduces Clipped Double Q-learning and parameter exploration NoisyNet technology to improve the scheduling policy for overestimated value estimates and action exploration problems in DRL. The simulation results show that DRL-PLink can effectively schedule mix-flow. Compared with ECMP and pFabric, the average flow completion time of DRL-PLink has decreased by 68.87% and 52.18% respectively. At the same time, it maintains a high deadline meet rate (>96.6%) close to pFabric and Karuna very much.
Jinjie Lu, Waixi Liu 0001, Yinghao Zhu, Sen Ling, Zhitao Chen, Jiaqi Zeng
MSN2
2020 AAMcon: an adaptively distributed SDN controller in data center networks
Waixi Liu 0001, Yu Wang 0017, Hongjian Liao, Zhong-Wei Liang, Xiaochu Liu
Frontiers Comput. Sci.1
2020 Fine-grained flow classification using deep learning for software defined data center networks
Waixi Liu 0001, Jun Cai 0002, Yu Wang 0017, Qing Chun Chen, Jia-Qi Zeng
J. Netw. Comput. Appl.1
2019 Intelligent Routing based on Deep Reinforcement Learning in Software-Defined Data-Center Networks
abstract
In software-defined data-center networks, Elephant flow/Mice flow/Coflow coexist and multiple resources (bandwidth, cache and computing) coexist. However, the conventional routing methods cannot overcome the large gap between different performance requirements of flow and efficient resource allocation. Therefore, this paper proposes DRL-R (Deep Reinforcement Learning-based Routing) to bridge this gap. First, we recombine multiple resources (node's cache, link's bandwidth) by quantifying the contribution score of them reducing the delay. This actually converts the performance requirements of flow into resource requirements of this flow, hence, the routing problem can be converted into a job-scheduling problem in resource management. Second, a DRL agent deployed on an SDN controller continually interacts with the network for adaptively performing reasonable routing according to the network state, and optimally allocating network resources for traffic. We employ Deep Q-Network (DQN) and Deep Deterministic Policy Gradient (DDPG) to build the DRL-R. Finally, we demonstrate the effectiveness of DRL-R through extensive simulations. Benefitted from continually learning with a global view, DRL-R can improve throughput highest up to 40% and flow completion time highest up to 47% over OSPF. DRL-R can improve the load balance of link highest up to 18.8% over OSPF. Additionally, DDPG has better performance than DQN.
Waixi Liu 0001
ISCC1
2018 A fog-based privacy-preserving approach for distributed signature-based intrusion detection
Yu Wang 0017, Weizhi Meng 0001, Wenjuan Li 0001, Jin Li 0002, Waixi Liu 0001, Yang Xiang 0001
J. Parallel Distributed Comput.5
2017 Information-centric networking with built-in network coding to achieve multisource transmission at network-layer
Waixi Liu 0001, Shunzheng Yu, Guang Tan, Jun Cai 0002
Comput. Networks1