VLDB 2026 Research / reviewers in the wild / expert
Sen Liu 0002
dblp:91/3699-2
· DBLP profile ↗
61ranked-venue papers
7as first author
49since 2021 · last 2026
0000-0003-2230-7671ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 42 · 2 first-author · 33 since 2021Systems, architecture and hardware · 13 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TPipe: Efficient Spiking Transformer Training with Time Parallelism and Asynchronous Pipeline
Yubing Bao, Zhihui Lu 0002, Qiang Duan 0002, Changze Lv, Xin Du 0002, Zeyi Deng, Jingqi Feng, Sen Liu 0002, Yang Chen 0001, Xin Wang 0002 |
INFOCOM | 8 |
| 2026 | 3D-INA: An Exploration of Integrating In-Network Aggregation into 3D Parallelism for LLM Training
Huifeng Xing, Hao Wang 0231, Yinfan Hu, Xin Ai 0008, Yang Chen 0001, Wanxin Shi, Sen Liu 0002, Yang Xu 0010 |
INFOCOM | 8 |
| 2026 | InfiniFlow: Decoupling Virtual Channel Scalability from Buffer Requirements in Lossless Datacenter Networks
Zerui Tian, Sen Liu 0002, Minkun Xue, Hao Shangguan, Ruyi Yao, Deli Huang, Songchen Xue, Yang Xu 0010 |
SIGCOMM | 2 |
| 2026 | CATS : An enhanced framework in LLM-based tabular data synthesis by correlation augmentationabstractAs AI technology advances, sectors such as finance and healthcare are increasingly adopting AI tools. However, due to privacy concerns that arise with the use of AI and the high cost of real data collection, generating realistic tabular data to replace original data has become an increasingly popular solution to address these limitations. Although the output of existing tabular data generation algorithms seems to match the distribution of original data, they often fail to preserve the correlations present in the original data. This oversight can lead to significant negative impacts on downstream tasks. This paper proposes a tabular data synthesis enhancement framework based on Large Language Model (LLM), i.e., C orrelation A ugmentation T abular data S ynthesis ( CATS ), which emphasizes improving the quality of generated tabular data while preserving the correlations between the features of the original data. We explore the technical details of the CATS framework and demonstrate that CATS achieves state-of-the-art performance compared with baseline methods across seven benchmark datasets. Additionally, it enhances the performance of existing strong LLM-based tabular data generators by an average improvement of over 4.7% across all datasets. Luyu Chen, Mingxuan Jiang, Ziyue Dai, Sen Liu 0002, Hongfeng Chai |
Expert Syst. Appl. | 4 |
| 2026 | Dynamic graph learning for integrating temporal relationships in stock prediction
Ziyue Dai, Qianru Zeng, Nianwang Lin, Hongjie Xia, Keyu Zhao, Sen Liu 0002, Guangnan Ye, Jie Wu 0003, Hongfeng Chai |
Expert Syst. Appl. | 8 |
| 2026 | Graph Self-Supervised Learning via Learnable View Augmentation for Recommender SystemabstractIn the field of recommender systems, graph neural networks (GNNs) have been extensively applied to collaborative filtering to generate personalized recommendations for users. To solve the problem of lack of observed data and contrasting interactions during representation learning, graph contrastive learning as an effective self-supervised learning (SSL) technique is presented to obtain augmented user and item representations. Nevertheless, most self-supervised approaches to generate recommendation either disrupt the graph structure or node embeddings through random augmentations or introduce augmented SSL information from biased data through heuristic methods. To overcome these challenges, we propose a learnable view augmentation model for collaborative filtering (LACF). Specifically, our framework embeds parameterized learnable view generators layer by layer into the automatic augmentation strategy, thus dynamically optimizing the adaptive augmented views of users and items through the backpropagation of weight gradients. In addition, LACF introduces a multiscale learning strategy that guides the view generator with layer-wise aware optimization and graph-level adaptive augmentation, enabling joint learning of representations with topological heterogeneity and semantic similarity from integrated viewpoint, achieving superior view augmentation. Extensive experiments on realworld datasets demonstrate that our LACF outperforms state-of-the-art baselines. In-depth analysis confirms the advantages of LACF in resistance against noise disturbances, alleviating data sparsity, and improving training efficiency. Hengjing Xiang, Yanfeng Xu, Sen Liu 0002, Zhi Liu 0011, Guangnan Ye |
IEEE Trans. Ind. Informatics | 4 |
| 2026 | AIRP: Accelerating Multi-Tenant Distributed Learning With In-Network Resource PoolingabstractThe increasing popularity of large models and datasets has highlighted the significance of distributed training networks. As gradient synchronization generates substantial traffic, in-network aggregation (INA) has emerged as a solution to offload aggregation onto the switch, alleviating network congestion and accelerating distributed training. However, the limited memory capacity of the INA switch becomes a potential bottleneck as computation shifts into the network, especially in multi-tenant scenarios. To address this bottleneck and enhance network throughput, we propose the Aggregation with Innetwork Resource Pooling (AIRP) framework. Unlike existing approaches that optimize individual switches in a localized manner, AIRP takes a holistic view and efficiently pools switch memory resources across the entire network, allocating them to multiple tenants. Evaluation using the ns-3 simulator and P4 testbed demonstrates that AIRP can accelerate the training of various models, including computer vision and language models. The experimental results show that AIRP outperforms existing INA approaches by up to 7 times in terms of network throughput in multi-tenant scenarios, while also achieving great flexibility and efficiency in deployment. Huifeng Xing, Hao Wang 0231, Yang Chen 0001, Yinfan Hu, Xuandong Liu, Zijian Li 0003, Wanxin Shi, Sen Liu 0002, Yang Xu 0010 |
IEEE Trans. Netw. | 9 |
| 2025 | FALI: Fusion Adapter for Multiple LoRA Models Inference
Kexuan Chang, Nianwang Lin, Zhixin Li 0003, Sen Liu 0002, Hongfeng Chai |
ICA3PP (8) | 5 |
| 2025 | Enhancing Equity: A Switch-Assisted Strategy for Improving Fairness in RDMA Networks
Quanwei Sun, Xingbo Gao 0003, Zerui Tian, Sen Liu 0002, Yang Xu 0010, H. Jonathan Chao |
ICA3PP (8) | 4 |
| 2025 | Enhancing In-network Aggregation with Adaptive Gradient Quantization for Multi-tenant LearningabstractWith the increasing popularity of distributed training applications, the growth in network traffic has become an impediment to the communication among worker nodes in the system. In-network aggregation (INA) has emerged as a solution to improve communication efficiency by offloading gradient aggregation to switches. However, in multi-tenant scenarios, INA switch memory capacity has been identified as a main bottleneck, leading to reduced network throughput and slower training processes. To address this, we propose Adaptive Gradient Quantization (AGQ) on the switch. AGQ reduces the quantization bit-width of gradients, allowing for storage of more gradients within the limited switch memory while maintaining training accuracy. Compared to quantization on hosts, AGQ can swiftly adapt to the available memory on switches and offers an improved balance between minimizing precision loss and enhancing training throughput. We implement AGQ on a P4 switch testbed, and experimental results demonstrate that enabling AGQ can achieve an up to 100% increase in training throughput without explicit drop of training accuracy compared with existing INA solutions like ATP and host-based quantization methods like THC. Huifeng Xing, Yinfan Hu, Hao Wang 0231, Yang Chen 0001, Sen Liu 0002, Yang Xu 0010 |
ICDCS | 6 |
| 2025 | Empowering Flowlet Load Balancing in RDMA with Host-Based Flowlet Fine-TuningabstractFlowlet-level load balancing has not demonstrated the expected robust capability in RDMA networks due to insufficient flowlets and the adverse effects of PFC. To delve deeper, we conduct measurements at end hosts and perform a detailed analysis of time gaps between packets. Our investigation reveals that in RDMA networks, the number of time gaps exceeding the flowlet timeout is considerably lower than the number in TCP networks. We also identify a stepwise time gap pattern that predicts the occurrence of PFC. Based on these observations, we propose$\text{HF}^{2} \mathrm{T}$, a host-based time gap adjustment method to improve the effectiveness of flowlet-level load balancing in RDMA networks. The core idea involves delaying a minimal number of specific packets at the host, actively extending the time gaps between them, thereby fostering the generation of sufficient flowlets at the switch and enhancing the utilization of equal-cost links. Incorporating an identification algorithm for the time gap pattern that predicts PFC,$\text{HF}^{2} \mathrm{T}$also leverages the time gap extension to reroute traffic away from potential PFC paths in advance, thus mitigating PFC occurrences. The minor cost of delaying a few packets is vastly offset by the benefits of generating flowlets and reducing PFC. We use DPDK to implement a prototype of$\text{HF}^{2} \mathrm{T}$, and through testbed experiments, we demonstrate that$\text{HF}^{2} \mathrm{T}$, serving as a building block for flowlet load balancing, can enhance the throughput of CONGA by 16.82%. The simulation results also show that$\text{HF}^{2} \mathrm{T}$can reduce the average FCT by 16.38% and the 99-percentile FCT by 21.13% compared to the state-of-the-art RDMA load balancing ConWeave. Chuhao Chen 0001, Deli Huang, Zerui Tian, Ruyi Yao, Sen Liu 0002, Yang Xu 0010 |
IWQoS | 6 |
| 2025 | FSPFL: Mitigating Communication Gap in Personalization Federated Learning Through Flexible Sparsity AllocationabstractFederated learning has emerged as a promising distributed learning paradigm that enables model training across decentralized devices while preserving data privacy. However, two critical challenges hinder its practical deployment: model performance degradation due to client drift in high data heterogeneity scenarios and communication gap due to varying client communication capabilities. In this paper, we provide a comprehensive analysis of these challenges and propose Flexible Sparse Personalized Federated Learning (FSPFL), a novel framework that jointly optimizes model personalization and communication efficiency. FSPFL adaptively allocates local model sparsity while incorporating personalization mechanisms to trade off model performance and communication efficiency. Extensive experiments demonstrate that FSPFL significantly mitigates the communication gap, and outperforms existing methods. Our results show that FSPFL improves communication efficiency by up to$4.1 \times$than the baselines while maintaining similar model accuracy in heterogeneous scenarios where clients have diverse data distributions and communication capabilities. Liuzhi Zhou, Nianwang Lin, Sen Liu 0002, Guangnan Ye, Yimin Yu, Hongfeng Chai |
IWQoS | 5 |
| 2025 | CClinguist: An Expert-Free Framework for Future-Compatible Congestion Control Algorithm IdentificationabstractCongestion control algorithms (CCAs) play a critical role in determining transmission quality. With their rapid evolution during the past few decades, understanding the CCA landscape on the Internet has become increasingly essential for network advancement. Traditional CCA census tools, however, rely heavily on manual configuration and construction, necessitating significant human effort to keep pace with the introduction of new CCAs. Ruyi Yao, Jialin Wei, Ruoshi Sun, Sen Liu 0002, Yang Xu 0010 |
SIGCOMM | 7 |
| 2024 | UniFL: Enabling Loss-tolerant Transmission in Federated LearningabstractAs Distributed Deep Learning (DDL) gains prominence, network constraints have emerged as a critical bottleneck impacting DDL performance. While state-of-the-art loss-tolerant (LT) transmission protocols enhance DDL efficiency, their application in federated learning (FL) environments is hindered by several challenges: (1) LT protocols necessitate client-side modifications, impractical in FL settings; (2) maintaining LT protocol transparency to senders compromises congestion control integrity; (3) LT protocols disrupt stream cipher, which is widely utilized in FL. To address these hurdles, this paper introduces UniFL, an innovative LT protocol tailored for FL applications. UniFL seamlessly integrates with FL architectures by preserving congestion control via a specialized speed limiter and adopting an advanced encryption technique that withstands packet loss, ensuring data integrity. UniFL is implemented within the NS3 for simulation evaluation. UniFL’s efficacy is evaluated across diverse models and datasets, demonstrating substantial performance enhancements in FL operations. In detail, UniFL can bring up to 40x speedup than the original FL with widely used congestion control algorithms and achieves throughput close to the state-of-the-art LT while being transparent to the workers. Yifan Ruan, Sen Liu 0002, Yang Xu 0010 |
APNet | 3 |
| 2024 | HF^2T: Host-Based Flowlet Fine-Tuning for RDMA Load BalancingabstractIn modern data center networks, RDMA is widely applied in scenarios such as high-performance computing, distributed storage and machine learning. In recent studies, it has been observed that flowlet switching load balancers cannot fully unleash their robust capabilities due to an insufficient number of flowlets in RDMA networks. In this paper, we scrutinize the traffic pattern at the end hosts and meticulously analyze time gaps between packets. Our findings reveal that in RDMA, the proportion of time gaps between packets larger than the flowlet threshold is notably scarce, constituting only a fraction of those in TCP, averaging 1/300. Based on this observation, we propose HF2T, a host-based method to improve the effectiveness of flowlet-level load balancing in RDMA. The core idea is to postpone a minimal number of specific packets at the host, actively elongating the time gaps between them, and promoting flowlet generation at the switch. The cost of postponing a minimal number of packets is far outweighed by the benefits of flowlets generation at the switch, improving the network performance. Simulation experiments confirm that HF2T, when deployed in conjunction with the flowlet load balancing, achieves an average reduction of 37.32% in Medium FCT and an average reduction of 28.75% in 99-percentile FCT, compared to deploying the same flowlet load balancing scheme solely at switches. Chuhao Chen 0001, Jiarui Ye, Yongbo Gao, Sen Liu 0002, Yang Xu 0010 |
APNet | 4 |
| 2024 | Halflife: An Adaptive Flowlet-based Load Balancer with Fading Timeout in Data Center NetworksabstractModern data centers (DCs) employ various traffic load balancers to achieve high bisection bandwidth. Among them, flowlet switching has shown remarkable performance in both load balancing and upper-layer protocol (e.g., TCP) friendliness. However, flowlet-based load balancers suffer from the inflexibility of flowlet timeout value (FTV) and result in sub-optimal performance under various application workloads. To this end, we propose Halflife, a novel flowlet-based load balancer that leverages fading FTVs to reroute traffic promptly under different workloads without any prior knowledge. Halflife not only balances traffic better, but also avoids the performance degradation caused by frequent oscillation or shifting of lows between paths. Furthermore, Halflife's fading mechanism is not only compatible with most flowlet-based load balancers, such as CONGA and LetFlow, but also improves their performance when leveraging flowlet switching in RDMA network. Through testbed experiments and simulations, we prove that Halflife improves the performance of CONGA and LetFlow by 10% ~ 150%, and it outperforms other load balancers by 30% ~ 200% across most application workloads. Sen Liu 0002, Yongbo Gao, Jiarui Ye, Furong Liang, Zerui Tian, Quanwei Sun, Zehua Guo 0001, Yang Xu 0010 |
EuroSys | 1 |
| 2024 | Rina: Enhancing Ring-Allreduce with in-Network Aggregation in Distributed Model TrainingabstractParameter Server (PS) and Ring-AllReduce (RAR) are two widely utilized synchronization architectures in multiworker Deep Learning (DL), also referred to as Distributed Deep Learning (DDL). However, PS encounters challenges with the “incast” issue, while RAR struggles with problems caused by the long dependency chain. The emerging In-network Aggregation (INA) has been proposed to integrate with PS to mitigate its incast issue. However, such PS-based INA has poor incremental deployment abilities as it requires replacing all the switches to show significant performance improvement, which is not costeffective. In this study, we present the incorporation of INA capabilities into RAR, called RAR with In-Network Aggregation (Rina), to tackle both the problems above. Rina features its agent-worker mechanism. When an INA-capable ToR switch is deployed, all workers in this rack run as one abstracted worker with the help of the agent, resulting in both excellent incremental deployment capabilities and better throughput. We conducted extensive testbed and simulation evaluations to substantiate the throughput advantages of Rina over existing DDL training synchronization structures. Compared with the state-of-the-art PS-based INA methods ATP, Rina can achieve more than$\mathbf{5 0 \%}$throughput with the same hardware cost. Xuandong Liu, Minglin Li, Yinfan Hu, Huifeng Xing, Hao Wang 0231, Wanxin Shi, Sen Liu 0002, Yang Xu 0010 |
ICNP | 9 |
| 2024 | MUSE: A Runtime Incrementally Reconfigurable Network Adapting to HPC Real-Time TrafficabstractInterconnection network in HPC is becoming a bottleneck due to increasing traffic load. We model adaptive routing mechanisms and prove that even with advanced adaptive routing, static networks like Dragonfly cannot handle non-uniform traffic efficiently, let alone the frequently changing non-uniform traffic. Therefore, it requires architectural changes for network-wide improvements, e.g., reconfigurable networks.Existing reconfigurable networks hardly support agile reaction to traffic changes with little impact on network. Therefore, we propose MUSE1, a Dragonfly-based runtime incrementally reconfigurable network to enable a small number of link adjustments for agility and little impact on transmitting flows during every reconfiguration with optical circuit switch (OCS).Simulations with both synthetic traffic and real-world workloads prove that MUSE can prevent saturation under typical traffic patterns that cause congestion in static Dragonfly. MUSE is 30-55% better than static Dragonfly and Flexfly w.r.t commonly used performance metrics like flow completion time (FCT). We also build a MUSE prototype and demonstrate that MUSE enables 20-30% less application finish time (AFT). Zijian Li 0003, Yiying Tang, Xin Ai 0008, Yuanyi Zhu, Zhigao Zhao, Sen Liu 0002, Bin Liu 0001, Yang Xu 0010 |
IPDPS | 9 |
| 2024 | R-PFC: Enhancing RDMA Network With Restricted And Fine-grained PFCabstractRDMA over Converged Ethernet (RoCE) has been widely used in datacenter networks and it relies on Priority Flow Control (PFC) to ensure a lossless network. However, PFC brings certain side effects, such as Head-of-Line (HoL) blocking, congestion spreading, and deadlock. Existing solutions demonstrate inherent limitations: either fail to completely eliminate the adverse impacts of PFC or introduce extra challenges. In light of these observations, this paper proposes a novel and practical scheme, Restricted Priority Flow Control (R-PFC). R-PFC consists of two parts: one-hop PFC and Virtual Next Output Queue (VNOQ). Instead of passively regarding PFC as a tool to guarantee a lossless network, one-hop PFC proactively employs PFC in a restrictive manner to minimize packet loss while limiting the spread of congestion within one hop. To further enhance the one-hop PFC, the fine-grained VNOQ solves the HoL blocking issue. We theoretically prove that R-PFC does not lead to deadlock and evaluate the performance of R-PFC under typical datacenter network scenarios in ns3 simulations. The results show that R-PFC outperforms both lossless and lossy networks by 43.76% and 39.46% on average. Minglin Li, Xin Ai 0008, Yongbo Gao, Sen Liu 0002, Yang Xu 0010 |
IWQoS | 7 |
| 2024 | TCAMVisor: High-throughput TCAM Virtualization for Multi-tenant Software Defined NetworkingabstractSoftware Defined Networking (SDN) provides users with a unified abstraction of physical networks. To meet the demands of modern data centers, many works have focused on designing network virtualization hypervisors that support multi-tenant SDN. Ternary Content Addressable Memory (TCAM) is widely used in SDN switches for rule storage. While it has extremely high lookup throughput, it also features drawbacks such as small capacity and slow update speed. Faced with multi-tenant scenarios, its limitations are even more pronounced. Existing hypervisors lack consideration for TCAM isolation, leading to slower TCAM updates and the mutual impact of requests from different tenants. Consequently, they fail to provide guaranteed performance to tenants. To solve these problems, we propose TCAMVisor, which further isolates TCAM resources based on traditional SDN hypervisors. Specifically, TCAMVisor provides better allocation mechanisms for TCAM entry and control bandwidth, ensuring inter-tenant isolation while improving resource utilization. Additionally, TCAMVisor improves the update speed of TCAM by delicately placing tenant rules. To the best of our knowledge, TCAMVisor is the first work to effectively achieve tenant isolation in TCAM, with an average throughput improvement of 5.5 times compared to FlowVisor. Ruoshi Sun, Ruyi Yao, Hao Wang 0231, Yiren Zhou, Sen Liu 0002, Yang Xu 0010 |
IWQoS | 7 |
| 2024 | Hierarchical Sketch: An Efficient, Scalable and Latency-aware Content Caching Design for Content Delivery NetworksabstractContent Delivery Networks (CDNs) are designed to reduce user-perceived waiting times and alleviate backbone bandwidth pressure. Since CDN cache servers have limited storage capacity, effective cache replacement policies are needed. However, existing CDN cache replacement policies mainly focus on improving content hit rates. As a result, some content with long origin fetch latency may not be cached, resulting in the long tail latency and degrading user experience. In this paper, we present Hierarchical Sketch, an efficient, scalable, and latency-aware cache replacement algorithm. Our approach leverages hierarchical slicing and voting mechanisms on a modified sketch to optimize content caching, reducing sorting complexity from O(log n) to O(1) with minimal loss of hit rate. Extensive simulations on synthetic and real-life industry CDN traces demonstrate that Hierarchical Sketch outperforms other algorithms in four different scenarios, with up to a 15% improvement. Huifeng Xing, Yuyan Ding, Huiru Huang, Sen Liu 0002, Zehua Guo 0001, Muath Al-Hasan, Mohamed Adel Serhani, Yang Xu 0010 |
IWQoS | 5 |
| 2024 | vPIFO: Virtualized Packet Scheduler for Programmable Hierarchical Scheduling in High-Speed NetworksabstractProgrammable packet scheduling enables the integration of scheduling algorithms into switches without the need for hardware redesign. The Push-In First-Out (PIFO) queue facilitates a programmable packet scheduler, supporting a single scheduling algorithm flexibly. However, hierarchical scheduling required in Multi-Tenant Data Centers (MTDCs) remains non-programmable. Dynamic and diverse hierarchical scheduling algorithms necessitate alterations in both the number of PIFO queues and their connection topology, posing a significant challenge to support them on fixed hardware. Zhiyu Zhang 0012, Shili Chen, Ruyi Yao, Ruoshi Sun, Hao Wang 0231, Gaojian Fang, Yibo Fan, Wanxin Shi, Sen Liu 0002, Yang Xu 0010 |
SIGCOMM | 11 |
| 2024 | TabSAL: Synthesizing Tabular data with Small agent Assisted Language models
Run Qian, Yandan Tan, Zhixin Li 0003, Luyu Chen, Sen Liu 0002, Jie Wu 0003, Hongfeng Chai |
Knowl. Based Syst. | 6 |
| 2023 | SDT: A Low-cost and Topology-reconfigurable Testbed for Network ResearchabstractNetwork experiments are essential to network-related scientific research (e.g., congestion control, QoS, network topology design, and traffic engineering). However, (re)configuring various topologies on a real testbed is expensive, time-consuming, and error-prone. In this paper, we propose Software Defined Topology Testbed (SDT), a method for constructing a user-defined network topology using a few commodity switches. SDT is low-cost, deployment-friendly, and reconfigurable, which can run multiple sets of experiments under different topologies by simply using different topology configuration files at the controller we designed. We implement a prototype of SDT and conduct numerous experiments. Evaluations show that SDT only introduces at most 2% extra overhead than full testbeds on multi-hop latency and is far more efficient than software simulators (reducing the evaluation time by up to 2899x). SDT is more cost-effective and scalable than existing Topology Projection (TP) solutions. Further experiments show that SDT can support various network research experiments at a low cost on topics including but not limited to topology design, congestion control, and traffic engineering. Zhigao Zhao, Zijian Li 0003, Sen Liu 0002, Yang Xu 0010 |
CLUSTER | 5 |
| 2023 | OSP: Boosting Distributed Model Training with 2-stage SynchronizationabstractDistributed deep learning (DDL) is a promising research area, which aims to increase the efficiency of training deep learning tasks with large size of datasets and models. As the computation capability of DDL nodes continues to increase, the network connection between nodes is becoming a major bottleneck. Various methods of gradient compression and improved model synchronization have been proposed to address this bottleneck in Parameter-Server-based DDL. However, these two types of methods can result in accuracy loss due to discarded gradients and have limited enhancement on the throughput of model synchronization, respectively. To address these challenges, we propose a new model synchronization method named Overlapped Synchronization Parallel (OSP), which achieves efficient communication with a 2-stage synchronization approach and uses Local-Gradient-based Parameter correction (LGP) to avoid accuracy loss caused by stale parameters. The prototype of OSP has been implemented using PyTorch and evaluated on commonly used deep learning models and datasets with a 9-node testbed. Evaluation results show that OSP can achieve up to 50% improvement in throughput without accuracy loss compared to popular synchronization models. Lei Shi 0031, Xuandong Liu, Sen Liu 0002, Yang Xu 0010 |
ICPP | 5 |
| 2023 | CoLUE: Collaborative TCAM Update in SDN Switches
Ruyi Yao, Chuhao Chen 0001, Wenjun Li 0004, Ying Wan 0001, Sen Liu 0002, Bin Liu 0001, Yang Xu 0010 |
INFOCOM | 7 |
| 2023 | S-PFC: Enabling Semi-Lossless RDMA Network with Selective Response to PFCabstractRoCEv2 (RDMA over Converged Ethernet version 2) is typically used in a PFC-enabled lossless network for high performance, but PFC can cause side effects, such as head-of-line (HoL) blocking and congestion spreading. Optimizing packet loss recovery mechanisms in lossy networks can enhance RDMA network performance, but packet loss can increase FCT of short flows and waste network resources. This paper proposes a novel concept of semi-lossless networks to exploit the advantages of lossless and lossy networks, and reduce their negative effects. Our proposed solution, Selective PFC (S-PFC), implements semi-lossless networks in two dimensions. First, S-PFC ensures no packet loss for short flows while timely dropping long flows. Second, S-PFC guarantees no packet loss at the network edge to prevent premature packet loss and unnecessary resource waste. Typical data center network scenarios and large-scale simulations show that S-PFC can accommodate different traffic demands effectively. Minglin Li, Sen Liu 0002, Yang Xu 0010 |
ISCC | 4 |
| 2023 | Boosting Distributed Machine Learning Training Through Loss-tolerant Transmission ProtocolabstractDistributed Machine Learning (DML) systems are utilized to enhance the speed of model training in data centers (DCs) and edge nodes. The Parameter Server (PS) communication architecture is commonly employed, but it faces severe long-tail latency caused by many-to-one “incast” traffic patterns, negatively impacting training throughput. To address this challenge, we design the Loss-tolerant Transmission Protocol (LTP), which permits partial loss of gradients during synchronization to avoid unneeded retransmission and contributes to faster synchronization per iteration. LTP implements loss-tolerant transmission through out-of-order transmission and out-of-order Acknowledges (ACKs). LTP employs Early Close to adjust the loss-tolerant threshold based on network conditions and bubble-filling for data correction to maintain training accuracy. LTP is implemented by C++ and integrated into PyTorch. Evaluations on a testbed of 8 worker nodes and one PS node demonstrate that LTP can significantly improve DML training task throughput by up to 30x compared to traditional TCP congestion controls, with no sacrifice to final accuracy. Lei Shi 0031, Xuandong Liu, Xin Ai 0008, Sen Liu 0002, Yang Xu 0010 |
IWQoS | 5 |
| 2023 | RateSheriff: Multipath Flow-aware and Resource Efficient Rate Limiter Placement for Data Center NetworksabstractEmerging cloud services and applications request different Quality of Service (QoS) in Data Center Networks (DCNs). To meet these various requirements, programmable switch-based rate limiters are introduced to provide performance isolation and benefit from easy control and fast deployment. However, existing programmable switch-based rate limiters have two limitations: (1) multipath flows (i.e., MultiPath TCP) cannot be precisely limited, and (2) rate limiter placement solutions in DCNs are missing. These limitations could lead to poor rate limiting performance and low bandwidth utilization. In this paper, we propose RateSheriff to improve rate limiting performance by providing multipath flow-aware and resource efficient rate limiter placement for programmable switch-enabled DCNs. We identify and associate subflows to a multipath flow by extracting and comparing specific packets and header fields. By solving the formulated resource efficient rate limiter placement problem, we can improve rate limiting performance and balance memory utilization among programmable switches in DCNs. Simulation results show that RateSheriff can correctly limit the rate of multipath flows, improve rate limiting performance by up to 46%, and improve memory balancing performance by up to 79% with low computation time, compared with baselines. Songshi Dou, Yongchao He, Sen Liu 0002, Wenfei Wu, Zehua Guo 0001 |
IWQoS | 3 |
| 2023 | Rusen: Rule Semantics Enabler toward Fast TCAM Update for Commodity SDN SwitchesabstractTernary Content Addressable Memory (TCAM) is widely used in Software-Defined Networking (SDN) switches due to its impressive throughput. But its unique circuit design results in long and inconsistent update delays. To overcome this challenge, many TCAM update algorithms based on rule semantics have been proposed. These algorithms eliminate unnecessary order restrictions, thus reducing update delays in theory. However, most commodity switches are semantic-unaware, which maintain rules in strict priority order. These algorithms are therefore not available for practical use. To address this issue, this paper proposes Rusen, a framework that enables the use of many semantic-based algorithms on Semantic-unaware commodity switches. Working as a transparent middle layer, the core idea of Rusen is to express the update scheme derived by semantic-based algorithms as messages that the Semantic-unaware switches can execute. In addition, Rusen optimizes the update scheme based on the specific characteristics of each switch, leading to improved performance of these algorithms. We evaluate the performance of Rusen by enabling several state-of-the-art semantic-based algorithms on commodity SDN switches. Results show that the average update delay can be significantly reduced by 23%∼94% on OpenFlow switches and 39%∼84% on a P4 switch. Ruoshi Sun, Ruyi Yao, Chuhao Chen 0001, Sen Liu 0002, Yang Xu 0010 |
IWQoS | 8 |
| 2023 | BMW Tree: Large-scale, High-throughput and Modular PIFO Implementation using Balanced Multi-Way Sorting TreeabstractPush-In-First-Out (PIFO) queue has been extensively studied as a programmable scheduler. To achieve accurate, large-scale, and high-throughput PIFO implementation, we propose the Balanced Multi-way (BMW) Sorting Tree for real-time packet sorting. The tree is highly modularized, insertion-balanced and pipeline-friendly with autonomous nodes. Ruyi Yao, Zhiyu Zhang 0012, Gaojian Fang, Peixuan Gao, Sen Liu 0002, Yibo Fan, Yang Xu 0010, H. Jonathan Chao |
SIGCOMM | 5 |
| 2023 | Federated Deep Reinforcement Learning-Based Intelligent Dynamic Services in UAV-Assisted MECabstractUnmanned aerial vehicles (UAVs)-assisted multiaccess edge computing (MEC) has emerged as a promising solution in B5G/6G networks. The high flexibility and seamless connectivity of UAVs make them well suited for providing enhanced communications coverage and efficient computing support. Particularly, in situations where ground facilities may be compromised or communication is unreliable. In this article, we study joint dynamic service switching and resource allocation for multiple UAVs in MEC network. We consider the heterogeneity of tasks and UAVs and model the dynamic service process of UAVs as a sequential decision problem based on the Markovian decision process. To enable dynamic and intelligent UAV service, we first propose a centralized dynamic service algorithm DDPG-based centralized (DDBC) based on deep reinforcement learning. However, given the training difficulties of the centralized algorithm, we propose a more promising distributed learning algorithm FLBF, which combines federated learning. We conduct extensive simulations to evaluate the effectiveness and advantages of the proposed algorithms. Our results show that DDBC and FLBF significantly reduce the system cost by 17.99%–35.72% and 12.30%–31.26%, respectively, compared to the comparative algorithms. Furthermore, FLBF can effectively improve the convergence speed with guaranteed learning performance, indicating its suitability for model training in UAV-assisted MEC networks. Peng Hou 0003, Zongshan Wang, Sen Liu 0002, Zhihui Lu 0002 |
IEEE Internet Things J. | 4 |
| 2023 | RaceCC: A rapidly converging explicit congestion control for datacenter networks
Minglin Li, Sen Liu 0002, Bin Liu 0001, Yang Xu 0010 |
J. Netw. Comput. Appl. | 5 |
| 2023 | Congestion-Aware Critical Gradient Scheduling for Distributed Machine Learning in Data Center NetworksabstractDistributed Machine Learning (DML) is proposed not only to accelerate the training of machine learning, but also to solve the inadequate ability for handling a large amount of training data. It adopts multiple computing nodes in data center to collaboratively work in parallel at the cost of high communication overhead. Gradient Compression (GC) is introduced to reduce the communication overhead by reducing the number of synchronized gradients among computing nodes. However, existing GC solutions suffer from varying network congestion. To be specific, when some computing nodes experience high network congestion, their gradient transmission process could be significantly delayed, slowing down the entire training process. To solve the problem, we propose FLASH, a congestion-aware GC solution for DML. FLASH accelerates the training process by jointly considering the iterative approximation of machine learning and dynamic network congestion scenarios. It can maintain good training performance by adaptively adjust and schedule the number of synchronized gradients among computing nodes. We evaluate the effectiveness of FLASH using AlexNet and Resnet18 under different network congestion scenarios. Simulation results show that under the same number of training epochs, FLASH reduces training time 22-71%, maintains good accuracy, and low loss, compared with the existing memory top-K GC solution. Zehua Guo 0001, Sen Liu 0002, Jineng Ren, Yang Xu 0010 |
IEEE Trans. Cloud Comput. | 3 |
| 2023 | Achieving Fast Convergence and High Efficiency using Differential Explicit Feedback in Data CenterabstractSince most flows are short-lived in data center networks, fast convergence becomes very important to help the short flows effectively utilize high bandwidth. Though current explicit feedback-based transport control protocols (TCPs) provide fast convergence via fine-grained congestion information from customized switches, they unavoidably incur large traffic overhead for widely existing small packets in data center applications, resulting in suboptimal network efficiency. To solve this issue, we propose a datacenter TCP based onDifferentialExplicitCongestionNotification, called DECN, to achieve fast convergence without any traffic overhead. Specifically, DECN feeds rate difference between the target and current rate back to the source by using multiple consecutive packets. Besides, we propose an enhanced version DECN* which obtains the optimal number of consecutive packets according to the packet loss rate. The experimental results of NS2 simulation and testbed implementation show that DECN and its enhanced version DECN* achieve comparable fast convergence as XCP without incurring any extra feedback overhead. Compared with the state-of-the-art explicit feedback-based TCPs, they reduce the flow completion time by up to 34% in typical data center applications. Jiawei Huang 0001, Jingling Liu, Sen Liu 0002, Jinbin Hu 0001, Jianxin Wang 0001 |
IEEE Trans. Cloud Comput. | 4 |
| 2023 | SpongeTraining: Achieving High Efficiency and Accuracy for Wireless Edge-Assisted Online Distributed LearningabstractEdge-assisted Distributed Learning (EDL) is a popular machine learning paradigm that uses a set of distributed edge nodes to collaboratively train a machine learning model using training data. Most of existing works implicitly assume that the fixed amount of training data is pre-collected and dispatched from user devices to edge nodes. In real world, however, training data in edge nodes are collected from user devices through wireless networks, and the volume and distribution of training data in edge nodes could exhibit temporal and spatial fluctuations due to varying wireless situations (e.g., network congestion, link capacity variation). In this way, existing solutions suffer from slow convergence and low accuracy. In this paper, we propose SpongeTraining to achieve high efficiency and accuracy for online EDL. To accommodate to fluctuations in training data, SpongeTraining uses a buffer at each worker to store received training data and adaptively adjusts training batch size and learning rate of each worker based on training data extracted from the buffer. Experiment results based on real-world datasets show that SpongeTraining outperforms existing solutions by accelerating the training process up to 50% for reaching the same training accuracy. Zehua Guo 0001, Sen Liu 0002, Jineng Ren, Yang Xu 0010, Yi Wang 0004 |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | An Apprenticeship Learning Approach for Adaptive Video Streaming Based on Chunk Quality and User PreferenceabstractVideo traffic has experienced an exponential increase in current years due to the growing ubiquity of mobile equipment and the constant network improvement. Most commercial players employ adaptive bitrate (ABR) algorithms to dynamically choose bitrate for each chunk based on perceived network capacity and buffer occupancy. Unluckily, even though improving the quality of chunks with dynamic scenes can achieve more QoE gain than static scenes, current ABR algorithms usually strive to maximize the average bitrate instead of perceptual quality, leading to the QoE degradation. To overcome this obstacle, we introduce a dynamic-chunk quality-aware adaptive bitrate algorithm through apprenticeship learning called DAVS (Dynamic-chunk qualityAwareVideoStreaming), where higher quality is selected for the dynamic chunks without reducing the quality of static chunks extravagantly. Furthermore, we take the user’s viewing preference into account to make DAVS adapt to the QoE diversity. The experimental results demonstrate that DAVS ameliorates the quality of dynamic chunks and significantly enhances the QoE compared with several representative ABR algorithms. Weihe Li, Jiawei Huang 0001, Shiqi Wang 0012, Chuliang Wu, Sen Liu 0002, Jianxin Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | ERA: Meeting the Fairness between Sender-driven and Receiver-driven Transmission Protocols in Data Center NetworksabstractThe modern data centers require high throughput and low latency transmission to meet the demands of distributed applications on communication delay. Compared with traditional sender-driven try-and-back-off protocols (e.g., TCP and its variants), receiver-driven protocols (RDPs) achieve the ultra-low transmission latency by reacting to credits or tokens from receivers. However, RDPs face fairness challenges when coexisting with sender-driven protocols (SDPs) in multi-tenant data centers. Their flows barely survive during coexistence with SDP flows since the delicate scheduling of their credits is disrupted and overwhelmed by SDP data packets. To tackle this issue, we propose the Equivalent Rate Adaptor (ERA), a scheme that converts the proactive try-and-back-off mode of SDPs to an RDP-like credit-based reactive mode. ERA leverages the advertised window field in ACK headers at the receiver side to elaborately limit the number of the in-flight packets or bytes in SDPs and thus reduce their impacts on RDPs. Therefore, ERA not only ensures the fairness between two different types of protocols, but also maintains the low latency feature of RDPs. Moreover, ERA is lightweight, flexible, and transparent to tenants by embedding into the prevalent Open vSwitch in the public cloud. The evaluation of both test-bed and NS2 simulation shows that ERA enables SDP flows and RDP flows to maintain good throughput and share the bandwidth fairly, improving the bandwidth stolen by up to 94.29%. Sen Liu 0002, Furong Liang, Zehua Guo 0001, Yang Xu 0010 |
ICDCS | 1 |
| 2022 | ABS: Adaptive Buffer Sizing via Augmented Programmability with Machine LearningabstractProgrammable switches have been proposed in today’s network to enable flexible reconfiguration of devices and reduce time-to-deployment. Buffer sizing, an important factor for network performance, however, has not received enough attention in programmable network. The state-of-the-art buffer sizing solutions usually employ either fixed buffer size or adjust the buffer size heuristically. Without programmability, they suffer from either massive packet drops or large queueing delay in dynamic environment. In this paper, we propose Adaptive Buffer Sizing (ABS), a low-cost and deploy-friendly framework compatible with programmable network. By decoupling the data plane and control plane, ABS-capable switches only need to react to the actions from controller, optimizing network performance in run-time under dynamic traffic. Meanwhile, actions can be programmed by particular Machine Learning (ML) models in the controller to meet different network requirements. In this paper, we address two specific ML models for different scenarios, a reinforcement learning model for relatively stable network with user specific quality requirements, and a supervised learning model for highly dynamic network condition. We implement the ABS framework by integrating the prevalent network simulator NS-2 with ML module. The experiment shows that ABS outperforms state-of-the-art buffer sizing solutions by up to 38.23x under various network environments. Jiaxin Tang, Sen Liu 0002, Yang Xu 0010, Zehua Guo 0001, Junjie Zhang 0001, Peixuan Gao, Yang Chen 0001, Xin Wang 0002, H. Jonathan Chao |
INFOCOM | 2 |
| 2022 | BubbleTCAM: Bubble Reservation in SDN Switches for Fast TCAM UpdateabstractThe unique hardware structure of Ternary Content-Addressable Memory (TCAM) enables its unparalleled lookup throughput but also causes slow update due to the Priority Order Constraint (POC). With the increase of application demands, TCAM update has become a bottleneck in the network. This paper proposes a new TCAM management mechanism named BubbleTCAM to enable fast TCAM update, in which available empty entries are defined as bubbles. The core idea of Bub-bleTCAM is to uniformly distribute bubbles and dependency chains in TCAM, which is beneficial to updates. BubbleTCAM consists of two components: bubble management and rule insertion. Bubble management enables TCAM to have uniformly distributed bubbles at all times through three key procedures: bubble lock reservation, bubble lock release and bubble generation. Rule insertion ensures that dependency chains of rules are uniformly stretched and distributed in TCAM. In addition, BubbleTCAM avoids the reorder problem by pre-sorting. Our evaluation based on the rulesets generated by ClassBench shows that BubbleTCAM effectively reduces the average cost and worst cost (in units of rule movements) during rule updates by at least 48% and 50%, respectively. Especially for the worst cost, the performance can be improved by up to 196x. Chuhao Chen 0001, Ruyi Yao, Ying Wan 0001, Wenjun Li 0004, Sen Liu 0002, Bin Liu 0001, Yang Xu 0010 |
IWQoS | 7 |
| 2022 | ECN-based shared bottleneck detection for multi-path TCP
Jin Ye 0003, Guihao Chen, Sen Liu 0002, Jiawei Huang 0001, Jianxin Wang 0001, Tian He 0001 |
Comput. Commun. | 5 |
| 2022 | Opportunistic Transmission for Video Streaming over Wild InternetabstractThe video streaming system employs adaptive bitrate (ABR) algorithms to optimize a user’s quality of experience. However, it is hard for ABR algorithms to choose the right bitrate consistently under highly dynamic bandwidth fluctuations in wild Internet. In this article, we propose a building block on the client side named Opportunistic Chunk Replacement Mechanism (OCRM) to help existing ABR algorithms make full use of the available bandwidth to improve the network utilization and viewing experience of users. Specifically, the servers take advantages of the spare bandwidth to opportunistically transmit high-quality chunks (called opportunistic chunks ) with low priority to the client, without incurring any extra delay. Then, the client player replaces the low-quality chunks with the opportunistic ones that have high quality. We compare OCRM with state-of-the-art ABR algorithms by using trace-driven experiments spanning a wide variety of quality of experience metrics and network conditions. The test results show that OCRM effectively achieves high network utilization and improves the user’s viewing experience by up to 35%. Jiawei Huang 0001, Qichen Su, Weihe Li, Zhuoran Liu 0003, Tao Zhang 0019, Sen Liu 0002, Ping Zhong 0002, Wanchun Jiang, Jianxin Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2022 | SDNShield: NFV-Based Defense Framework Against DDoS Attacks on SDN Control PlaneabstractSoftware-defined networking (SDN) is increasingly popular in today’s information technology industry, but existing SDN control plane is insufficiently scalable to support on-demand, high-frequency flow requests. Weaknesses along SDN control paths can be exploited by malicious third parties to launch distributed denial-of-service (DDoS) attacks against the SDN control plane. Recently proposed solutions only partially solve the problem, by protecting either the SDN network edges or the centralized controller. We propose SDNShield, a solution based on emerging network function virtualization (NFV) technologies, which enforces more comprehensive defense against potential DDoS attacks on SDN control plane. SDNShield incorporates a three-stage overload control scheme. The first stage statistically identifies legitimate flows with low complexity and performance overhead. The second stage further performs in-depth TCP handshake verification to ensure good flows are eventually served. The third stage intellectually salvages the misclassified legitimate flows that are falsely dropped from the first two stages. Prototype tests and real data-driven simulation results show that SDNShield can achieve high resilience against brute-force attacks, and maintain good flow-level service quality at the same time. Kuan-yin Chen, Sen Liu 0002, Yang Xu 0010, Ishant Kumar Siddhrau, Zehua Guo 0001, H. Jonathan Chao |
IEEE/ACM Trans. Netw. | 2 |
| 2022 | Maintaining Control Resiliency and Flow Programmability in Software-Defined WANs During Controller FailuresabstractProviding resilient network control is a critical concern for deploying Software-Defined Networking (SDN) into Wide-Area Networks (WANs). For performance reasons, a Software-Defined WAN is divided into multiple domains controlled by multiple controllers with a logically centralized view. Under controller failures, we need to remap the control of offline switches from failed controllers to other active controllers. Existing solutions have three limitations: (1) the least flow programmability (e.g., the ability to change paths of flows) cannot be maintained; (2) active controllers could be overloaded, interrupting their normal operations; (3) network performance could be degraded because of the increasing controller-switch communication overhead. In this paper, we propose RetroFlow+ to recover the flow programmability and achieve low communication overhead during controller failures. By intelligently configuring a set of selected offline switches working under the legacy routing mode and several active controllers releasing a few control resources, RetroFlow+ enables active controllers to use the minimum control resource to sustain the flow programmability. RetroFlow+ also smartly transfers the control of offline switches with the SDN routing mode to active controllers to minimize the communication overhead from these offline switches to the active controllers. Simulation results show that RetroFlow+ realizes low communication overhead, recovers all offline flows under one and two controller failures, and improves the flow recovery percentage up to 70% under three controller failures, compared with the state-of-the-art solution. Zehua Guo 0001, Songshi Dou, Sen Liu 0002, Wendi Feng, Wenchao Jiang, Yang Xu 0010, Zhi-Li Zhang |
IEEE/ACM Trans. Netw. | 3 |
| 2021 | Exploring the Impact of Attacks on Ring AllReduceabstractDistributed Machine Learning (DML) is widely used to accelerate the training of the deep learning model. In DML, Parameter-Server (PS) and Ring AllReduce are two typical architectures. Recently, observing that many works address the security problem in PS, whose performance can be greatly degraded by malicious participation during the training process. However, the robustness of Ring AllReduce, which can solve the communication bandwidth problem in PS, to the malicious participant is still unknown. In this paper, we design a series of experiments to explore the security problem in Ring AllReduce, and reveal it can also suffer from the malicious participant. Zehua Guo 0001, Sen Liu 0002 |
APNet | 4 |
| 2021 | Optimizing Flow Completion Time via Adaptive Buffer Management in Data Center NetworksabstractThe traffic of modern data centers exhibits long-tail distribution, in which massive delay-sensitive short flows and a small number of bandwidth-hungry long flows co-exist. These two types of flows could share same bottleneck links in the data center networks but request different or even opposite network requirements. Existing solutions try to realize a trade-off between the requirements of different flows by either prioritizing short flows or limiting the buffer used by long flows at switches or end-hosts. However, they do not consider the dynamic traffic change and suffer from performance degradation, resulted from severe queueing delay and massive packet drops for short flows under current First-In-First- Out (FIFO) queueing mechanism. In this paper, we propose a novel buffer management scheme at switches, called Cut-in Queue (CQ), to achieve both low latency for short flows and high throughput for long flows. Based on network status in real time, CQ prioritizes short flows by dynamically cutting the short flows’ packets into the head of long flows or evicting some enqueued long flows’ packets and enables high throughput for long flows in most of the cases. Evaluation of both DPDK testbed and NS2 simulations show that CQ outperforms state-of-the-art buffer management schemes by reducing flow completion time by up to 73%. Sen Liu 0002, Zehua Guo 0001, Yi Wang 0004, Mohamed Adel Serhani, Yang Xu 0010 |
ICPP | 1 |
| 2021 | Reducing traffic burstiness for MPTCP in data center networks
Sen Liu 0002, Jiawei Huang 0001, Wenchao Jiang, Jianxin Wang 0001 |
J. Netw. Comput. Appl. | 1 |
| 2021 | HybridFlow: Achieving Load Balancing in Software-Defined WANs With Scalable RoutingabstractThe scalability issue hinders the deployment of Software-Defined Networking (SDN) in the Wide Area Networks (WANs). Existing solutions have two issues: (1) network performance relies on complicated controller synchronization, which increases the complexity of network control; (2) fine-grained flow processing enables flexible flow control at the cost of high processing load on the controllers and high flow table occupancy on switches. In this paper, we propose a scalable routing solution named HybridFlow, which achieves a good load balancing performance using a single controller with low control overhead (i.e., flow routing and rerouting overhead). HybridFlow mainly employs two techniques: hybrid routing and crucial flow rerouting. Hybrid routing enabled by commercial SDN switches gives us opportunities to reduce the processing load of the controller by routing flows with the hybrid OpenFlow/OSPF mode. Thus, the majority of flows can be routed by OSPF without involving the controller. Crucial flow rerouting realizes load balancing by dynamically identifying crucial flows based on a new metric called Variation Slope and rerouting these flows with the hybrid OpenFlow/OSPF mode. The simulation based on the real traffic traces and network typologies shows that compared with the optimal solution, HybridFlow can achieve 87% of the optimal load balancing performance by rerouting 36% less flows on average. Zehua Guo 0001, Songshi Dou, Yi Wang 0004, Sen Liu 0002, Wendi Feng, Yang Xu 0010 |
IEEE Trans. Commun. | 4 |
| 2021 | AggreFlow: Achieving Power Efficiency, Load Balancing, and Quality of Service in Data Center NetworksabstractPower-efficient Data Center Networks (DCNs) have been proposed to save power of DCNs using OpenFlow. In these DCNs, the OpenFlow controller adaptively turns on/off links and OpenFlow switches to form a minimum-power subnet that satisfies the traffic demand. As the subnet changes, flows are dynamically routed and rerouted to the routes composed of active switches and links. However, existing flow scheduling schemes could cause undesired results: (1) power inefficiency: due to unbalanced traffic allocation on active routes, extra switches and links may be activated to cater to bursty traffic surges on congested routes, and (2) Quality of Service (QoS) fluctuation: because of the limited flow entry processing ability, switches may not be able to timely install/delete/update flow entries to properly route/reroute flows. In this paper, we propose AggreFlow, a dynamic flow scheduling scheme that achieves power efficiency and QoS improvement using three techniques: Flow-set Routing, Lazy Rerouting, and Adaptive Rerouting. Flow-set Routing achieves load balancing with a small number of flow entry operations by routing flows in a coarse-grained flow-set fashion. Lazy Rerouting spreads rerouting operations over a relatively long period of time, reducing the burstiness of entry operation on switches. Adaptive Rerouting selectively reroutes flow-sets to maintain load balancing. We built an NS3 based fat-tree network simulation platform to evaluate AggreFlow's performance. The simulation results show that AggreFlow reduces power consumption by about 18%, yet achieving load balancing and improved QoS (low packet loss rate and reducing the number of processing entries for flow scheduling by 98%), compared with baseline schemes. Zehua Guo 0001, Yang Xu 0010, Ya-Feng Liu, Sen Liu 0002, H. Jonathan Chao, Zhi-Li Zhang, Yuanqing Xia |
IEEE/ACM Trans. Netw. | 4 |
| 2020 | DAVS: Dynamic-Chunk Quality Aware Adaptive Video Streaming using Apprenticeship LearningabstractTo deliver video in a high quality across various network conditions, adaptive bitrate (ABR) algorithms dynamically select bitrate for each chunk according to perceived network rate and buffer occupancy. Unfortunately, though ameliorating the quality of chunks with dynamic scenes can obtain more QoE gain than the ones with static scenes, current ABR algorithms generally aim to maximize the average bitrate rather than perceptual quality, resulting in the QoE degradation. To address this issue, we propose a dynamic-chunk quality aware adaptive bitrate scheme via apprenticeship learning named DAVS, in which higher quality is chosen for the dynamic chunks without decreasing the quality of static chunks excessively. The experimental results show that DAVS enhances the quality of dynamic chunks and greatly improves the overall QoE compared with the state-of-the-art ABR algorithms. Weihe Li, Jiawei Huang 0001, Shiqi Wang 0012, Sen Liu 0002, Jianxin Wang 0001 |
GLOBECOM | 4 |
| 2020 | QOS-Aware Flow Control for Power-Efficient Data Center Networks with Deep Reinforcement LearningabstractReducing the power consumption and maintaining the Flow Completion Time (FCT) for the Quality of Service (QoS) of applications in Data Center Networks (DCNs) are two major concerns for data center operators. However, existing works either fail in guaranteeing the QoS due to the neglect of the FCT constraints or achieve a less satisfying power efficiency. In this paper, we propose SmartFCT, which employs Software-Defined Networking (SDN) coupled with the Deep Reinforcement Learning (DRL) to improve the power efficiency of DCNs and guarantee the FCT. The DRL agent can generate a dynamic policy to consolidate traffic flows into fewer active switches in the DCN for power efficiency, and the policy also leaves different margins in different active links and switches to avoid FCT violation of unexpected short bursts of flows. Simulation results show that with similar FCT guarantee, SmartFCT can save 8% more of the power consumption compared to the state-of-the-art solutions. Penghao Sun, Zehua Guo 0001, Sen Liu 0002, Julong Lan, Yuxiang Hu 0001 |
ICASSP | 3 |
| 2020 | Achieving Fast Convergence and High Efficiency using Differential Explicit Feedback in Data CenterabstractSince most flows are short-lived in data center networks, fast convergence becomes very important to help the short flows effectively utilize high bandwidth. Though current feedback-based transport control protocols (TCPs) provide fast convergence via fine-grained explicit congestion information from customized switches, they unavoidably incur large traffic overhead for widely existing small packets in data center applications, resulting in suboptimal network efficiency. To solve this issue, we propose a datacenter TCP based on differential feedbacks, called DECN, to achieve fast convergence without any traffic overhead. Specifically, DECN feeds rate difference between the target and current rate back to the source by using multiple consecutive packets. The experimental results of NS2 simulation and testbed implementation show that DECN achieves comparable fast convergence as XCP without incurring any extra feedback overhead. Compared with the state-of-the-art feedback-based TCPs, DECN reduces the flow completion time by up to 34.1% in typical data center applications. Jiawei Huang 0001, Sen Liu 0002, Jinbin Hu 0001, Jianxin Wang 0001 |
ICC | 3 |
| 2020 | Poster: Maintaining Training Efficiency and Accuracy for Edge-assisted Online Federated Learning with ABSabstractThis paper proposes Adaptive Batch Sizing (ABS) for online federated learning. ABS is an iteration process-efficient solution that adaptively adjusts batch size of the training process at edge nodes. Preliminary results show that ABS maintains training efficiency and accuracy, compared with existing iteration round-efficient solutions. Zehua Guo 0001, Sen Liu 0002, Yuanqing Xia |
ICNP | 3 |
| 2020 | SmartFCT: Improving power-efficiency for data center networks with deep reinforcement learning
Penghao Sun, Zehua Guo 0001, Sen Liu 0002, Julong Lan, Yuxiang Hu 0001 |
Comput. Networks | 3 |
| 2019 | DDT: Mitigating the Competitiveness Difference of Data Center TCPsabstractTo achieve better network performance, the cloud service providers are widely deploying the ECN-based transport protocols (i.e., DCTCP) in their data center networks (DCN). In multi-tenant environment, however, the newly introduced ECN-enabled TCP greatly impairs the performance of applications with out-dated and miscon figured TCP stacks. The reason is that the ECN-enabled datacenter switch fails to treat the mixed TCP traffic fairly, causing the distinguished performance gap between the ECN-enabled and ECN-disabled TCPs. This paper proposes DDT (Dual Dynamic Thresholds), an active queue management algorithm (AQM) that aims to achieve the flow-level fairness when the heterogeneous TCP traffic coexists. DDT monitors the switch queue in real time, and dynamically tunes the distance between ECN-marking and packet-dropping thresholds to mitigate the competitiveness difference between the ECN-enabled and ECN-disabled TCP. Our preliminary real implementations and testing results show that DDT elegantly fills the competitiveness gap of heterogeneous TCP traffic without disturbing their own control loops, while only introducing acceptable deployment overhead at the switch. Tao Zhang 0019, Jiawei Huang 0001, Shaojun Zou, Sen Liu 0002, Jinbin Hu 0001, Jingling Liu, Chang Ruan, Jianxin Wang 0001, Geyong Min |
APNet | 4 |
| 2019 | EMPTCP: An ECN Based Approach to Detect Shared Bottleneck in MPTCPabstractThe major challenge of Real Time Protocol is to balance efficiency and fairness over limited bandwidth. MPTCP has proved to be effective for multimedia and real time networks. Ideally, an MPTCP sender should couple the subflows sharing the bottleneck link to provide TCP friendliness. However, existing shared bottleneck detection scheme either utilize end-to-end delay without consideration of multiple bottleneck scenario, or identify subflows on switch at the expense of operation overhead. In this paper, we propose a lightweight yet accurate approach, EMPTCP, to detect shared bottleneck. EMPTCP uses the widely deployed ECN scheme to capture the real congestion state of shared bottleneck, while at the same time can be transparently utilized by various enhanced MPTCP protocols. Through theory analysis, simulation test and real network experiment, we show that EMPTCP achieves higher than 90% accuracy in shared bottleneck detection, thus improving the network efficiency and fairness. Jin Ye 0003, Renzhang Liu, Ziqi Xie, Luting Feng, Sen Liu 0002 |
ICCCN | 5 |
| 2019 | Reducing Flow Completion Time with Replaceable Redundant Packets in Data Center NetworksabstractIn the data center network, a packet-level load balancer such as random packet spraying (RPS) achieves high throughput by spraying data packets to all transmission paths, which easily suffers from the packet out-of-order problem under network asymmetry. While state-of-the-art network coding schemes can mitigate the issue, too many encoded redundant packets introduced by the network coding will cause extra traffic overhead, larger queueing delay and even TCP time out. In this paper, we propose OPportunistic Encoded Redundant (OPER), a middle-layer design upon existing coding schemes to mitigate the curse of redundant packets. Specifically, OPER uses opportunistic redundant packets which are replaceable by the data packets in the switches under heavy congestion. OPER is implemented as a shim layer between TCP and IP layers at end-hosts and a loadable plugin at switches, leaving existing TCP/IP protocols unmodified. The testbed and NS2 experiments show that, OPER reduces the average flow completion time by up to 71% compared with the state-of-the-art multipath coding schemes. Sen Liu 0002, Jiawei Huang 0001, Wenchao Jiang, Jianxin Wang 0001, Tian He 0001 |
ICDCS | 1 |
| 2019 | RetroFlow: maintaining control resiliency and flow programmability for software-defined WANsabstractProviding resilient network control is a critical concern for deploying Software-Defined Networking (SDN) into Wide-Area Networks (WANs). For performance reasons, a Software-Defined WAN is divided into multiple domains controlled by multiple controllers with a logically centralized view. Under controller failures, we need to remap the control of offline switches from failed controllers to other active controllers. Existing solutions could either overload active controllers to interrupt their normal operations or degrade network performance because of increasing the controller-switch communication overhead. In this paper, we propose RetroFlow to achieve low communication overhead without interrupting the normal processing of active controllers during controller failures. By intelligently configuring a set of selected offline switches working under the legacy routing mode, RetroFlow relieves the active controllers from controlling the selected offline switches while maintaining the flow programmability (e.g., the ability to change paths of flows) of SDN. RetroFlow also smartly transfers the control of offline switches with the SDN routing mode to active controllers to minimize the communication overhead from these offline switches to the active controllers. Simulation results show that compared with the baseline algorithm, RetroFlow can reduce the communication overhead up to 52.6% during a moderate controller failure by recovering 100% flows from offline switches and can reduce the communication overhead up to 61.2% during a serious controller failure by setting to recover 90% of flows from offline switches. Zehua Guo 0001, Wendi Feng, Sen Liu 0002, Wenchao Jiang, Yang Xu 0010, Zhi-Li Zhang |
IWQoS | 3 |
| 2019 | RAPID: Avoiding TCP Incast Throughput Collapse in Public Clouds With Intelligent Packet DiscardingabstractMany applications in public clouds require a high fan-in, many-to-one type of data communication (known as TCP incast) in modern Data Center Networks (DCNs). Such communication could cause severe incast congestion in switches and result in TCP throughput collapse, substantially degrading the application performance. The root cause of throughput collapse is the Retransmission Timeouts (RTO) due to packet losses in congested switches. Tenants in public clouds can opt to use a variety of TCP versions. However, the existing solutions rely on modifications of TCP protocols and specific techniques from switches, and thus these existing solutions are not always feasible for public clouds. In this paper, we are inspired by the emerging virtualization and network softwarization technologies to develop a novel scheme called Retransmission timeout Avoidance by Packet Intelligent Discarding (RAPID) using software switches. RAPID considers the number of packets of each incast flow, buffered in the switch to selectively discard some packets, and ensures that the Fast Retransmission/Fast Recovery rather than RTO is invoked at the sender(s) in response to packet loss. Thus, the long idle period of a timeout and the throughput drop are avoided. We prove that, given a predetermined minimum switch buffer space, dedicated to the incast application, RAPID can prevent RTO in all the incast senders. We also present a low-complexity heuristic version of RAPID named RAPID-ED, which combines the principles of RAPID and early detection and is extremely easy to implement on today's software switches. We evaluate the two proposed schemes in a data center network testbed built on NS-3 simulator. The simulation results confirm the theoretical expectation, and show that the RAPID and RAPID-ED perform very well to prevent RTO of TCP incast flows and hence the throughput collapse. Compared with other incast solutions, RAPID and RAPID-ED do not modify TCP protocols and therefore are more suitable in public clouds. Yang Xu 0010, Shikhar Shukla, Zehua Guo 0001, Sen Liu 0002, Adrian Sai-Wah Tam, Kang Xi, H. Jonathan Chao |
IEEE J. Sel. Areas Commun. | 4 |
| 2019 | Task-Aware TCP in Data Center NetworksabstractIn modern data centers, many flow-based and task-based schemes have been proposed to speed up the data transmission in order to provide fast, reliable services for millions of users. However, the existing flow-based schemes treat all flows in isolation, contributing less to or even hurting user experience due to the stalled flows. Other prevalent task-based approaches, such as centralized and decentralized scheduling, are sophisticated or unable to share task information. In this work, we first reveal that the relinquishing bandwidth of leading flows to the stalled ones effectively reduces the task completion time. We further present the design and implementation of a general supporting scheme that shares the flow-tardiness information through a receiver-driven coordination. Our scheme can be flexible and widely integrated with the state-of-the-art TCP protocols designed for data centers in either single stage or multiple stage scenario, while making no modification on switches. Through the testbed experiments and simulations of typical data center applications, we show that in single stage scenario, our scheme reduces the task completion time by 70% and 50% compared with the flow-based protocols (e.g., DCTCP, L2DCT) and task-based scheduling (e.g., Baraat), respectively. Moreover, our scheme also outperforms other approaches by 18%~25% in prevalent topologies of the data center. For multiple stage scenario, our scheme also has up to 50% improvement compared to other schemes. Sen Liu 0002, Jiawei Huang 0001, Yutao Zhou, Jianxin Wang 0001, Tian He 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2017 | Task-aware TCP in Data Center NetworksabstractIn modern data centers, many flow-based and task-based schemes have been proposed to speed up the data transmission in order to provide fast, reliable services for millions of users. However, existing flow-based schemes treat all flows in isolation, contributing less to or even hurting user experience due to the stalled flows. Other prevalent task-based approaches, such as centralized and decentralized scheduling, are sophisticated or unable to share task information. In this work, we first reveal that relinquishing bandwidth of leading flows to the stalled ones effectively reduces the task completion time. We further present the design and implementation of a general supporting scheme that shares the flow-tardiness information through a receiver-driven coordination. Our scheme can be flexibly and widely integrated with the state-of-the-art TCP protocols designed for data centers, while making no modification on switches. Through the testbed experiments and simulations of typical data center applications, we show that our scheme reduces the task completion time by 70% and 50% compared with the flow-based protocols (e.g. DCTCP, L2DCT) and task-based scheduling (e.g. Baraat), respectively. Moreover, our scheme also outperforms other approaches by 18% to 25% in prevalent topologies of data center. Sen Liu 0002, Jiawei Huang 0001, Yutao Zhou, Jianxin Wang 0001, Tian He 0001 |
ICDCS | 1 |