Ran Shu 0001

dblp:06/11343-1 · DBLP profile ↗
← Back
47ranked-venue papers
3as first author
26since 2021 · last 2026
0000-0002-2021-4917ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 31 · 3 first-author · 19 since 2021Systems, architecture and hardware · 16 · 7 since 2021
YearPublicationVenuePosition
2026 Congestion Quarantine in Lossless Ethernet
abstract
Lossless Ethernet uses hop-by-hop backpressure to prevent buffer overflow and has become the mainstream choice for running Remote Direct Memory Access (RDMA) in AI and cloud data centers. Despite preventing congestion-induced drops, lossless networks introduce congestion contagion, which causes head-of-line blocking, congestion spreading, and deadlocks. Congestion control schemes have been introduced to mitigate the drawbacks of lossless Ethernet. However, congestion control mechanisms struggle with bursty traffic, face a dilemma, and can be sidelined by backpressure. In this paper, we propose congestion quarantine (CQ) as a complementary congestion management mechanism for lossless Ethernet. CQ uses a separate queue to quarantine congested flows, preventing congestion contagion, resolving the CC dilemma, and preventing CC from being sidelined. Results show that congestion quarantine can eliminate head-of-line blocking in scenarios with multiple congestion trees. Large-scale simulations demonstrate that CQ reduces the average and 99th percentile FCT slowdown of normal flows by 25–86% and 33.4–90%, respectively, with negligible impact on bursty traffic.
Dongkang Hu, Ran Shu 0001, Wenxue Cheng, Fengyuan Ren
APNet2
2026 OptiFlow: Towards LLM-Driven Optimization of Collective Communication Algorithms
Ziyue Yang 0002, Kaihui Gao, Shuai Wang 0028, Li Chen 0008, Zhixiong Niu, Ran Shu 0001, Wenxue Cheng, Peng Cheng 0005, Yongqiang Xiong, Dan Li 0001
APNet7
2026 PIN: Less Is More for RDMA Load Balancing
abstract
As RDMA becomes increasingly tolerant to out-of-order delivery, fine-grained packet-level multipathing is emerging as a practical design for datacenter fabrics. Packet spraying and related schemes improve load distribution for large flows, but they also force RDMA flows onto multiple paths whose conditions can differ at short timescales due to randomized traffic placement. For short flows, even one packet sent on a temporarily slower path can delay the entire flow. As a result, fine-grained load balancing can hurt, rather than help, small multi-packet flows.
Jichun Wu, Ran Shu 0001, Yongqiang Xiong
APNet2
2026 ECOTE: Priority-Aware Optical Restoration for WAN Traffic Engineering
abstract
Fiber cuts are among the most common and disruptive failures in cloud networks. They can prevent cloud providers from maintaining committed service availability, causing Service Level Agreement (SLA) violations that directly translate into monetary penalties. Existing traffic engineering (TE) approaches enhance failure resilience, and recent systems further incorporate optical restoration to recover lost bandwidth after failures. However, they still treat services largely uniformly and optimize primarily for network throughput rather than the economic impact of heterogeneous SLA penalties. In this paper, we present ECOTE, the first priority-aware TE system with optical restoration that explicitly minimizes revenue loss. Specifically, ECOTE introduces a new optical restoration formulation with a dedicated capacity restoration solver to compute an optimal restoration plan that maximizes restorable capacity under physical constraints. ECOTE also designs a priority-aware TE algorithm that allocates the restored bandwidth capacity according to SLA penalties, thereby reducing monetary cost. We evaluate ECOTE using a production-level WAN testbed and through large-scale simulations. The testbed evaluation demonstrates ECOTE achieves zero loss for high-priority services and more than 10× revenue loss reduction compared to state-of-the-art. Our large-scale simulation results show that ECOTE can support at least 2.5× and 2.0× more demand for different high priority services compared to the state-of-the-art solutions. Meanwhile, ECOTE reduces the revenue loss by at least an order of magnitude less than existing solutions.
Kunling He, Ran Shu 0001, Jilong Wang 0001, Congcong Miao
EuroSys4
2026 SmartNIC-Enabled Live Migration for Storage-Optimized VMs with PYROCUMULUS
Jiechen Zhao 0002, Ran Shu 0001, Ziyue Yang 0002, Rui Ma 0021, Derek Chiou, Natalie D. Enright Jerger, Peng Cheng 0005, Yongqiang Xiong
NSDI2
2026 Nüwa: A Generative Control Plane for AI Network Simulation
abstract
Network simulation plays a critical role in improving the efficiency of large-scale AI clusters for design validation, parameter tuning, and protocol development. However, high-fidelity network simulation becomes prohibitively slow at scale, especially when running large batches of experiments on topologies with tens or hundreds of thousands of accelerators. We observe that a key bottleneck comes from the control plane. Existing network simulators typically compute routes and install forwarding tables at initialization, which can consume hundreds of GB of memory before packet-event execution begins and limit overall simulation throughput. In this paper, we present Nüwa, which views routing as a compilation problem, it leverages the hierarchical and symmetric structure common in AI fabrics and compiles a declarative topology description together with routing policies into compact forwarding artifacts that are fast to generate and efficient to look up. Evaluations show that Nüwa can reduce simulation initialization time from hours to only 25 seconds for a 65,536-GPU cluster. For end-to-end simulation time, Nüwa takes only 20% of that required by existing approaches in a 40K+ GPU cluster, and Nüwa can scale to a 221,184-GPU cluster.
Ran Shu 0001, Peng Zhang 0011, Danfeng Shan, Yongqiang Xiong
SIGCOMM2
2026 STORM: Enabling Traffic Scheduling for RDMA
abstract
Remote Direct Memory Access (RDMA) is increasingly used as a shared communication substrate across datacenter workloads with very different scheduling needs, from request-response services and storage fan-out to AI training collectives. Proper request scheduling can reduce communication time, but in practice, no RDMA flow scheduling is enabled in datacenters, leaving traffic to simple fair sharing. We present STORM, a NIC-level scheduler for all types of RDMA workloads using NIC-only information: the known RDMA request size, and per-queue-pair backlog. STORM converts these signals into a small number of extra priority levels on the wire and prioritizes requests that are either near completion or blocking queued dependent work. STORM requires no application hints and works with both in-order RoCEv2 and newer RDMA stacks that tolerate reordering. We prototype STORM on an FPGA NIC with negligible overhead. Across representative cloud and LLM training workloads, STORM reduces training iteration time by up to 12% and reduces average and P99 flow completion slowdown by up to 90% compared to fair scheduling.
Jichun Wu, Ran Shu 0001, Gianni Antichi, Yongqiang Xiong, Jon Crowcroft
SIGCOMM2
2025 Nüwa: Efficient Generative Control Plane for AI Network Simulation
Ran Shu 0001, Peng Zhang 0011, Yongqiang Xiong
APNet2
2025 Miniature: Fast AI Supercomputer Networks Simulation on FPGAs
Yicheng Qian, Ran Shu 0001, Rui Ma 0021, Yang Wang 0053, Derek Chiou, Nadeen Gebara, Luca Piccolboni, Miriam Leeser, Yongqiang Xiong
APNet2
2025 SRC: A Scalable Reliable Connection for RDMA with Decoupled QPs and Connections
Ran Shu 0001, Yongqiang Xiong
APNet2
2025 Daredevil: Rescue Your Flash Storage from Inflexible Kernel Storage Stack
abstract
Existing kernel storage stacks for NVMe SSDs struggle to address performance interference between I/O requests from tenants with different SLAs, leading to the multi-tenancy issue. Addressing this requires separating their I/O requests within the NVMe I/O queues (NQs). However, our analysis reveals that the static CPU core-NQ bindings of current storage stacks restrict their flexibility to achieve this goal.
Ran Shu 0001, Jiayi Lin 0007, Qingyu Zhang 0005, Ziyue Yang 0002, Jie Zhang 0048, Yongqiang Xiong, Chenxiong Qian
EuroSys2
2025 ScalaTap: Scalable Outbound Rate Limiting in Public Cloud
Zhongjie Chen, Yingchen Fan, Kun Qian 0017, Qingkai Meng 0001, Ran Shu 0001, Bo Wang 0066, Wei Li 0262, Fengyuan Ren
INFOCOM5
2025 Software-based Live Migration for RDMA
abstract
Live migration is critical to ensure services are not interrupted during host maintenance in data centers. On the other hand, RDMA has been widely adopted in data centers, and has attracted both academia and industry for years. However, live migration of RDMA is not supported in today's data centers. Although modifying RDMA NICs (RNICs) to be aware of live migration has been proposed for years, there is no sign of supporting it on commodity RNICs. This paper proposes MigrRDMA, a software-based RDMA live migration that does not rely on any extra hardware support. MigrRDMA provides a software indirection layer to achieve transparent switching to new RDMA communications. Unlike previous RDMA virtualization that provides sharing and isolation, MigrRDMA's indirection layer focuses on keeping the RDMA states on the migration source and destination identical from the perspective of applications. We implemented MigrRDMA prototype over Mellanox RNICs. Our evaluation shows that MigrRDMA adds little downtime when migrating a container with live RDMA connections running at line rate. Besides, the MigrRDMA virtualization layer only adds 3% ~ 9% extra overheads in the data path. When migrating Hadoop tasks, MigrRDMA only incurs an extra 3-second job completion time.
Ran Shu 0001, Yongqiang Xiong, Fengyuan Ren
SIGCOMM2
2025 HyperDrive: Direct Network Telemetry Storage via Programmable Switches
abstract
In cloud datacenter operations, telemetry and logs are indispensable, enabling essential services such as network diagnostics, auditing, and knowledge discovery. The escalating scale of data centers, coupled with increased bandwidth and finer-grained telemetry, results in an overwhelming volume of data. This proliferation poses significant storage challenges for telemetry systems. In this article, we introduce HyperDrive, an innovative system designed to efficiently store large volumes of telemetry and logs in data centers using programmable switches. This in-network approach effectively mitigates bandwidth bottlenecks commonly associated with traditional endpoint-based methods. To our knowledge, we are the first to use a programmable switch to directly control storage, bypassing the CPU to achieve the best performance. With merely 21% of a switch’s resources, our HyperDrive implementation showcases remarkable scalability and efficiency. Through rigorous evaluation, it has demonstrated linear scaling capabilities, efficiently managing 12 SSDs on a single server with minimal host overhead. In an eight-server testbed, HyperDrive achieved an impressive throughput of approximately 730 Gbps, underscoring its potential to transform data center telemetry and logging practices.
Ziyuan Liu 0008, Zhixiong Niu, Ran Shu 0001, Wenxue Cheng, Jacob Nelson 0001, Dan R. K. Ports, Peng Cheng 0005, Yongqiang Xiong
IEEE Trans. Cloud Comput.3
2025 Low-Overhead Intra-Host Container Communication With Hardware Offloading
abstract
Containers are widely embraced for their deployment and performance benefits over virtual machines. Yet, for many data-intensive applications in containerized clouds, bulky data transfers may impose performance issues. In particular, communication across co-located containers on the same host incurs large overheads in memory copy and the kernel’s TCP stack. Existing solutions such as shared-memory networking and RDMA have their own limitations, including insufficient memory isolation and limited scalability. This paper presents PipeDevice, a new system for low overhead intra-host container communication. PipeDevice follows a hardware-software co-design approach — it offloads data forwarding entirely onto hardware, which accesses application data in hugepages on the host, thereby eliminating CPU overhead from memory copy and TCP processing. PipeDevice preserves memory isolation and scales well to connections, making it deployable in public clouds. Isolation is achieved by allocating dedicated memory to each connection from hugepages. To achieve high scalability, PipeDevice stores the connection states entirely in host DRAM and manages them in software. Evaluation with a prototype implementation on commodity FPGA shows that for delivering 80Gbps across containers PipeDevice saves 63.2% CPU compared to kernel TCP stack, and 40.5% over FreeFlow. PipeDevice provides salient benefits to applications. For example, we port baidu-allreduce to PipeDevice and obtain$\sim 2.2\times $gains in allreduce throughput.
Zhixiong Niu, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong, Dongsu Han, Chun Jason Xue, Hong Xu 0001
IEEE Trans. Netw.3
2024 Software-based Live Migration for Containerized RDMA
abstract
Container live migration is critical to ensure services are not interrupted during host maintenance in data centers. On the other hand, RDMA containerization has attracted both academia and industry for years. However, live migration for containerized RDMA is not supported in today’s data centers. Although modifying RDMA NICs (RNICs) to be aware of live migration has been proposed for years, there is no sign of supporting it on commodity RNICs. This paper proposes MigrRDMA, a software-based RDMA live migration for containers, which does not rely on any extra hardware support. MigrRDMA provides a minimum virtualization layer inside the RDMA library loaded in applications, which achieves transparent switching to new RDMA communications. Unlike previous RDMA virtualization that provides sharing and isolation, MigrRDMA’s virtualization layer focuses on keeping the RDMA states on the migration source and destination the same from the perspective of applications. Our evaluation shows that MigrRDMA only adds 0.7 ∼ 12.1 ms downtime to migrate a container with live RDMA connections running at line rate. Besides, the MigrRDMA virtualization layer only adds 3% ∼ 9% overheads in the data path operations.
Ran Shu 0001, Yongqiang Xiong, Fengyuan Ren
APNet2
2024 DockRDMA: Hybrid RDMA Virtualization for Containerized Clouds
abstract
Containers have become the de facto choice for major cloud services. Meanwhile, with demands for extremely high performance, data centers have widely adopted RDMA for their online services. RDMA virtualization is the critical technology that enables RDMA for containers. Hybrid RDMA virtualization leverages the software flexibility in the control path, and keeps the native performance in the data path. Thus, it is the best choice for RDMA virtualization. State-of-the-art hybrid RDMA virtualization cannot address containerspecific problems. This paper proposes DockRDMA, the first hybrid RDMA virtualization solution for containerized clouds. DockRDMA develops several mechanisms, including embedding physical addresses in virtual ones to provide efficient address translation, hybrid network policy enforcement at scale, a general virtual RDMA NIC initialization method to be compatible with all container platforms, and namespace checking to protect the RDMA NIC instances. Evaluation results show that DockRDMA provides bare-metal RDMA performance in the data path, and almost native communication setup time in the control path. Compared with the state-of-the-art hybrid virtualization technology, DockRDMA reduces Hadoop job completion time by 6%. It offers seamless integration with existing container platforms, protects critical information of RDMA NIC instances, and exhibits excellent scalability to meet diverse network policies required by different containers.
Ran Shu 0001, Zhongjie Chen, Xiaohui Luo, Bo Wang 0066, Qingkai Meng 0001, Fengyuan Ren
ICNP2
2024 NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
abstract
The Compute Express Link (CXL) interconnect makes it feasible to integrate diverse types of memory into servers via its byte-addressable SerDes links. Considering the various access latency, harnessing the full potential of CXL-based heterogeneous memory systems requires efficient memory tiering. However, prior work can hardly make a fundamental progress owing to low-resolution and high-overhead memory access profiling techniques. To address this critical challenge, we propose a novel memory tiering solution called NeoMem, which features a hardware/software co-design. NeoMem offloads memory profiling functions to CXL device-side controllers, integrating a dedicated hardware unit called NeoProf. NeoProf readily monitors memory accesses and provides the OS with crucial page hotness statistics and other useful system state information. On the OS kernel side, we design a revamped memory-tiering strategy, enabling accurate and timely hot page promotion based on NeoProf statistics. We implement NeoMem on a real FPGA-based CXL memory platform and Linux kernel v6.3. Comprehensive evaluations demonstrate that NeoMem achieves 32% ~ 67% geomean speedup over several existing memory tiering solutions.
Zhe Zhou 0002, Tao Zhang 0032, Yang Wang 0053, Ran Shu 0001, Shuotao Xu, Peng Cheng 0005, Yongqiang Xiong, Jie Zhang 0048, Guangyu Sun 0003
MICRO5
2023 SlimeMold: Hardware Load Balancer at Scale in Datacenter
abstract
Stateful load balancers (LB) are essential services in cloud data centers, playing a crucial role in enhancing the availability and capacity of applications. Numerous studies have proposed methods to improve the throughput, connections per second, and concurrent flows of single LBs. For instance, with the advancement of programmable switches, hardware-based load balancers (HLB) have become mainstream due to their high efficiency. However, programmable switches still face the issue of limited registers and table entries, preventing them from fully meeting the performance requirements of data centers. In this paper, rather than solely focusing on enhancing individual HLBs, we introduce SlimeMold, which enables HLBs to work collaboratively at scale as an integrated LB system in data centers.
Ziyuan Liu 0008, Zhixiong Niu, Ran Shu 0001, Guohong Lai, Zongying He, Jacob Nelson 0001, Dan R. K. Ports, Peng Cheng 0005, Yongqiang Xiong
APNet3
2023 Polaris: Enhancing CXL-based Memory Expanders with Memory-side Prefetching
Zhe Zhou 0002, Shuotao Xu, Tao Zhang 0032, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong, Guangyu Sun 0003
APPT5
2023 ARK: GPU-driven Code Execution for Distributed Deep Learning
Changho Hwang, KyoungSoo Park, Ran Shu 0001, Xinyuan Qu, Peng Cheng 0005, Yongqiang Xiong
NSDI3
2023 Poster: Meili: Towards SmartNIC as a Service
abstract
The gap between the stagnation of CPU power and the increase in network bandwidth has promoted a shift towards placing more computation on network hardware [16, 17]. Therefore, SmartNICs have become prevalent in data centers to serve various cloud applications, from network functions [15, 17, 22] to high-level applications like distributed applications and storage [14, 16, 18--21, 23].
Shaofeng Wu, Zhixiong Niu, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong, Chun Jason Xue, Zaoxing Liu, Hong Xu 0001
SIGCOMM4
2022 A Disaggregate Data Collecting Approach for Loss-Tolerant Applications
abstract
Datacenter generates operation data at an extremely high rate, and data center operators collect and analyze them for problem diagnosis, resource utilization improvement, and performance optimization. However, existing data collection methods fail to efficiently aggregate and store data at extremely high speed and scale. In this paper, we explore a new approach that leverages programmable switches to aggregate data and directly write data to the destination storage. Our proposed data collection system, ALT, uses programmable switches to control NVMe SSDs on remote hosts without the involvement of a remote CPU. To tolerate loss, ALT uses an elegant data structure to enable efficient data recovery when retrieving the collected data. We implement our system on a Tofino-based programmable switch for a prototype. Our evaluation shows that ALT can saturate SSD’s peak performance without any CPU involvement.
Ziyuan Liu 0008, Zhixiong Niu, Ran Shu 0001, Wenxue Cheng, Peng Cheng 0005, Yongqiang Xiong, Jacob Nelson 0001, Dan R. K. Ports
APNet3
2022 PipeDevice: a hardware-software co-design approach to intra-host container communication
abstract
Containers are prevalently adopted due to the deployment and performance advantages over virtual machines. For many containerized data-intensive applications, however, bulky data transfers may pose performance issues. In particular, communication across co-located containers on the same host incurs large overheads in memory copy and the kernel's TCP stack. Existing solutions such as shared-memory networking and RDMA have their own limitations, including insufficient memory isolation and limited scalability.
Chuanwen Wang, Zhixiong Niu, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong, Dongsu Han, Chun Jason Xue, Hong Xu 0001
CoNEXT4
2022 PilotFish: Harvesting Free Cycles of Cloud Gaming with Deep Learning Training
Wei Zhang 0149, Binghao Chen, Zhenhua Han, Quan Chen 0002, Peng Cheng 0005, Fan Yang 0024, Ran Shu 0001, Yuqing Yang 0001, Minyi Guo
USENIX ATC7
2021 Minimizing Coflow Completion Time in Optical Circuit Switched Networks
abstract
Nowadays, optical circuit switching is becoming an increasingly favored technology in scaling data center networks for its definitive advantages in data rate, power consumption, and device cost. Concurrently, reducing coflow completion time (CCT) is of great significance for improving application-level performance. However, minimizing CCT in circuit switched networks is totally different from that in traditional packet switched networks due to port constraints and circuit reconfiguration delays. To address this issue, this article proposes Grouped Optimization-based Scheduling (GOS), a CCT minimization algorithm for circuit switched networks integrating circuit and coflow scheduling. We first formalize the CCT minimization problem into a 0-1 programming problem, then relax and solve the problem in 2 steps to obtain the coflow order and flow grouping decisions on each circuit. Thus intra-group reconfiguration delays are saved, and small coflows can be prioritized at the group level. Theoretical analysis proves GOS is a 4-approximation algorithm in average CCT. To reduce computing overheads, we further propose a heuristic approximation algorithm. Extensive simulations show that the heuristic algorithm has satisfactory CCT performance (0.12× Varys, 0.36× Sunflow) as well as high throughput (16.74× Varys, 1.32× Sunflow), and well adapts to a wide range of reconfiguration delays and algorithm decision time.
Tong Zhang 0018, Fengyuan Ren, Jiakun Bao, Ran Shu 0001, Wenxue Cheng
IEEE Trans. Parallel Distributed Syst.4
2020 Towards Influence of Chunk Size Variation on Video Streaming in Wireless Networks
abstract
In recent years, the growth in popularity of mobile video streaming services is unbroken. There are tremendous demands for video streaming over wireless networks. Currently, most video streaming is over HTTP. Up to now, HTTP-based adaptive video streaming is standardized as DASH, where a client-side video player can dynamically pick the bitrate level according to the perceived network conditions. Actually, not only the available bandwidth drastically varies due to wireless network properties, but also the chunk sizes in the same bitrate level significantly fluctuate, which also influences the bitrate adaptation. However, existing bitrate adaptation algorithms mostly focus on available bandwidth but do not involve chunk size variation, leading to performance losses. In this paper, we theoretically analyze the influence of chunk size variation on bitrate adaptation performance in wireless networks. Based on DASH system features, we build a general model describing playback buffer evolution. Applying stochastic theories, we respectively analyze the influence of the chunk size variation on rebuffering probability, average bitrate, and bitrate switching interval. Furthermore, based on theoretical insights, we provide several suggestions for algorithm designing and rate encoding, and also design a simple bitrate adaptation algorithm. Extensive simulations verify our insights, suggestions, and designed algorithm effectiveness.
Tong Zhang 0018, Fengyuan Ren, Wenxue Cheng, Xiaohui Luo, Ran Shu 0001
IEEE Trans. Mob. Comput.5
2020 Observing and Mitigating Micro-Burst Traffic in Data Center Networks
abstract
Micro-burst traffic is not uncommon in data centers. It can cause packet dropping, which may result in serious performance degradation (e.g., Incast problem). However, current approaches to mitigate micro-burst is usually ad-hoc and not based on a principled understanding of the underlying behaviors. On the other hand, traditional studies focus on traffic burstiness in a single flow, while micro-burst traffic in the data centers could occur with highly fan-in communication pattern, and its dynamic behavior is still unclear. To this end, in this paper, we re-examine the micro-burst traffic in typical data center scenarios. We find that the evolution of micro-burst is determined by both TCP's self-clocking mechanism and congestion control algorithm. Besides, dynamic behaviors of micro-burst under various scenarios can all be described by the time derivative of queue length evolution.Our observations also implicate that conventional solutions like absorbing and pacing are ineffective to mitigate micro-burst traffic.Instead, senders need to rapidly respond to some explicit signals of the queue buildup caused by the micro-burst traffic rather than independently and ineffectually pacing themselves in isolation. Inspired by the findings and insights from experimental observations, we propose Micro-burst-Aware Transport Control Protocol (MATCP), which leverages characteristic behaviors of micro-burst traffic derived from the time derivative of the queue occupancy. MATCP can suppress the sharp queue length increment by over 2x and reduce the tail query completion time by up to 84.4%.
Danfeng Shan, Fengyuan Ren, Peng Cheng 0005, Ran Shu 0001, Chuanxiong Guo
IEEE/ACM Trans. Netw.4
2019 DLBooster: Boosting End-to-End Deep Learning Workflows with Offloading Data Preprocessing Pipelines
abstract
In recent years, deep learning (DL) has prospered again due to improvements in both computing and learning theory. Emerging studies mostly focus on the acceleration of refining DL models but ignore data preprocessing issues. However, data preprocessing can significantly affect the overall performance of end-to-end DL workflows. Our studies on several image DL workloads show that existing preprocessing backends are quite inefficient: they either perform poorly in throughput (30% degradation) or burn too many (>10) CPU cores. Based on these observations, we propose DLBooster, a high-performance data preprocessing pipeline that selectively offloads key workloads to FPGAs, to fit the stringent demands on data preprocessing for cutting-edge DL applications. Our testbed experiments show that, compared with the existing baselines, DLBooster can achieve 1.35×~2.4× image processing throughput in several DL workloads, but consumes only 1/10 CPU cores. Besides, it also reduces the latency by 1/3 in online image inference.
Dan Li 0001, Binyao Jiang, Xi Fan, Jinkun Geng, Wei Bai 0001, Ran Shu 0001, Peng Cheng 0005, Yongqiang Xiong
ICPP11
2019 Direct Universal Access: Making Data Center Resources Available to FPGA
Ran Shu 0001, Peng Cheng 0005, Guo Chen 0001, Yongqiang Xiong, Derek Chiou, Thomas Moscibroda
NSDI1
2019 Distributed Bottleneck-Aware Coflow Scheduling in Data Centers
abstract
With the booming development of data parallel frameworks, the coflow abstraction has been greatly favored by data center transport designs, for its prominent ability in capturing application-level semantics. To accelerate job completion, coflow completion time (CCT) is a most important metric, and coflow scheduling is the most effective and widely-adopted means of optimizing CCT. However, most existing coflow scheduling mechanisms neglect the ubiquitous in-network bottlenecks and schedule coflows based on non-blocking giant switch hyperthesis. Such a practice is likely to result in undesired link contention inside the fabric, finally impairing CCT performance. To address this problem, we propose the Distributed Bottleneck-Aware coflow scheduling algorithm called DBA, which approximates the minimum remaining time first (MRTF) heuristic on all fabric-wide links. In this way, core link bandwidths are allocated to coflows as expected and the CCT performance will not be violated. As an evolutionary algorithm, DBA enhances the traditional dual decomposition method thus converges to the optimal bandwidth allocation very fast. Extensive simulations verify DBA's outstanding CCT performance as well as high link utilization. Furthermore, DBA introduces very little overhead and is robust to routing strategies, parameter variations and computation delays.
Tong Zhang 0018, Ran Shu 0001, Zhiguang Shan, Fengyuan Ren
IEEE Trans. Parallel Distributed Syst.2
2018 Micro-Burst in Data Centers: Observations, Analysis, and Mitigations
abstract
Micro-burst traffic is not uncommon in data centers. It can cause packet dropping, which results in serious performance degradation (e.g., Incast problem). However, current solutions that attempt to suppress micro-burst traffic are extrinsic and ad hoc, since they lack the comprehensive and essential understanding of micro-burst's root cause and dynamic behavior. On the other hand, traditional studies focus on traffic burstiness in a single flow, while in data centers micro-burst traffic could occur with highly fan-in communication pattern, and its dynamic behavior is still unclear. To this end, in this paper, we re-examine the microburst traffic in typical data center scenarios. We find that evolution of micro-burst is determined by both TCP's self-clocking mechanism and bottleneck link. Besides, dynamic behaviors of micro-burst under various scenarios can all be described by the slope of queue length evolution. Our observations also implicate that conventional solutions like absorbing and pacing are ineffective to mitigate micro-burst traffic. Instead, senders need to slow down as soon as possible. Inspired by the findings and insights from experimental observations, we propose S-ECN policy, which is an ECN marking policy leveraging the slope of queue length evolution. Transport protocols utilizing S-ECN policy can suppress the sharp queue length increment by over 2×, and reduce the average query completion time by ~12-27%.
Danfeng Shan, Fengyuan Ren, Peng Cheng 0005, Ran Shu 0001, Chuanxiong Guo
ICNP4
2018 Scheduling Coflows with Incomplete Information
abstract
In recent years, the coflow abstraction has received significant attentions, for its prominent ability to capture application semantics. On this basis, multiple coflow scheduling mechanisms have been proposed to minimize the coflow completion time (CCT). Currently, existing coflow scheduling mechanisms mainly belong to two categories: information-omniscient and information-agnostic. However, in data center applications, there are still quite a few cases in between where incomplete coflow information is known, and such incomplete information makes great contributions to improving the CCT performance. To address such cases, we propose IICS, a coflow scheduling algorithm based on incomplete coflow information. IICS leverages information of a coflow's arrived parts to deduce the coflow's remaining transmission time, and uses it to approximate the Minimum Remaining Time First (MRTF) heuristic. Besides, IICS allocates bandwidth by monopolization and in a maximal manner, which achieves high bandwidth utilization. Extensive simulations under realistic settings show that IICS achieves the average CCT comparable to that of the information-omniscient algorithm and the 99th percentile CCT much smaller than both information-omniscient and information-agnostic algorithms. Furthermore, IICS holds observably higher throughput and is robust to algorithm parameters.
Tong Zhang 0018, Fengyuan Ren, Ran Shu 0001, Bo Wang 0066
IWQoS3
2018 Analysing and improving convergence of quantized congestion notification in Data Center Ethernet
Ran Shu 0001, Fengyuan Ren, Jiao Zhang 0002, Tong Zhang 0018, Chuang Lin 0002
Comput. Networks1
2018 Towards Stable Flow Scheduling in Data Centers
abstract
At present, soft real-time data center applications are in a booming development and impose stringent delay requirements on internal data transfers. In this context, many recently proposed data center transport protocols share a common goal of minimizing Flow Completion Time (FCT), and the Shortest Remaining Processing Time (SRPT) scheduling algorithm has attracted widespread attentions for its superior performance in average FCT. However, SRPT suffers from the instability problem, incurring more and more flows left uncompleted even if the traffic load is within the fabric capacity, which implies unnecessary bandwidth waste. To solve the problem, this paper proposes a backlog-aware flow scheduling algorithm (BASRPT) for both giant switch and general topologies. Because of taking into account queue backlogs other than flow sizes at scheduling, we prove that BASRPT is stable and still maintains good FCT performance. To overcome the huge computation overhead and enable distributed implementation, a fast and practical approximation algorithm called fast BASRPT is also developed. Extensive flow-level simulations show that fast BASRPT indeed stabilizes the queue length and obtains a higher throughput while being able to push the FCT arbitrarily close to the optimal value in the condition of feasible traffic loads.
Tong Zhang 0018, Fengyuan Ren, Ran Shu 0001
IEEE Trans. Parallel Distributed Syst.3
2018 MPTCP Tunnel: An Architecture for Aggregating Bandwidth of Heterogeneous Access Networks
abstract
Fixed and cellular networks are two typical access networks provided by operators. Fixed access network is widely employed; nevertheless, its bandwidth is sometimes not sufficient enough to meet user bandwidth requirements. Meanwhile, cellular access network owns unique advantages of wider coverage, faster increasing link speed, more flexible deployment, and so forth. Therefore, it is attractive for operators to mitigate the bandwidth shortage by bundling these two. Actually, there have been existing schemes proposed to aggregate the bandwidth of two access networks, whereas they all have their own problems, like packet reordering or extra latency overhead. To address this problem, we design new architecture, MPTCP Tunnel, to aggregate the bandwidth of multiple heterogeneous access networks from the perspective of operators. MPTCP Tunnel uses MPTCP, which solves the reordering problem essentially, to bundle multiple access networks. Besides, MPTCP Tunnel sets up only one MPTCP connection at play which adapts itself to multiple traffic types and TCP flows. Furthermore, MPTCP Tunnel forwards intact IP packets through access networks, maintaining the end‐to‐end TCP semantics. Experimental results manifest that MPTCP Tunnel can efficiently aggregate the bandwidth of multiple access networks and is more adaptable to the increasing heterogeneity of access networks than existing mechanisms.
Danfeng Shan, Ran Shu 0001, Tong Zhang 0018
Wirel. Commun. Mob. Comput.3
2017 Modeling and analyzing the influence of chunk size variation on bitrate adaptation in DASH
abstract
Recently, HTTP-based adaptive video streaming has been widely adopted in the Internet. Up to now, HTTP-based adaptive video streaming is standardized as Dynamic Adaptive Streaming over HTTP (DASH), where a client-side video player can dynamically pick the bitrate level according to the perceived network conditions. Actually, not only the available bandwidth is varying, but also the chunk sizes in the same bitrate level significantly fluctuate, which also influences the bitrate adaptation. However, existing bitrate adaptation algorithms do not accurately involve the chunk size variation, leading to performance losses. In this paper, we theoretically analyze the influence of chunk size variation on bitrate adaptation performance. Based on DASH system features, we build a general model describing the playback buffer evolution. Applying stochastic theories, we respectively analyze the influence of the chunk size variation on rebuffering probability and average bitrate level. Furthermore, based on theoretical insights, we provide several recommendations for algorithm designing and rate encoding, and also propose a simple bitrate adaptation algorithm. Extensive simulations verify our insights as well as the efficiency of the proposed recommendations and algorithm.
Tong Zhang 0018, Fengyuan Ren, Wenxue Cheng, Xiaohui Luo, Ran Shu 0001
INFOCOM5
2016 TFC: token flow control in data center networks
abstract
Services in modern data center networks pose growing performance demands. However, the widely existed special traffic patterns, such as micro-burst, highly concurrent flows, on-off pattern of flow transmission, exacerbate the performance of transport protocols. In this work, an clean-slate explicit transport control mechanism, called Token Flow Control (TFC), is proposed for data center networks to achieve high link utilization, ultra-low latency, fast convergence, and rare packets dropping. TFC uses tokens to represent the link bandwidth resource and define the concept of effective flows to stand for consumers. The total tokens will be explicitly allocated to each consumer every time slot. TFC excludes in-network buffer space from the flow pipeline and thus achieves zero-queueing. Besides, a packet delay function is added at switches to prevent packets dropping with highly concurrent flows. The performance of TFC is evaluated using both experiments on a small real testbed and large-scale simulations. The results show that TFC achieves high throughput, fast convergence, near zero-queuing and rare packets loss in various scenarios.
Jiao Zhang 0002, Fengyuan Ren, Ran Shu 0001, Peng Cheng 0005
EuroSys3
2016 Backlog-Aware SRPT Flow Scheduling in Data Center Networks
abstract
The rapidly developing soft real-time data center applications impose stringent delay requirements on internal data transfers. Therefore many recently emerged network protocols in data center share a common goal of decreasing Flow Completion Time (FCT), in which case the Shortest Remaining Processing Time (SRPT) scheduling discipline has attracted widespread attentions. However, SRPT suffers the instability issue, incurring more and more flows left uncompleted even when traffic load is within network capacity, which implies unnecessary bandwidth waste. To solve the problem, this paper proposes a backlog aware scheduling algorithm (BASRPT) that stabilizes queue length while maintaining relatively low FCT based on Lyapunov optimization. To overcome the huge computational overhead, a fast and practical approximation algorithm called fast BASRPT is also developed. Extensive flow-level simulations show that fast BASRPT indeed stabilizes switch queue and obtains a higher throughput while being able to push FCT arbitrarily close to the optimal value in the condition of feasible traffic load.
Tong Zhang 0018, Fengyuan Ren, Ran Shu 0001
ICDCS3
2016 Guaranteeing Delay of Live Virtual Machine Migration by Determining and Provisioning Appropriate Bandwidth
abstract
The proliferation of cloud services makes virtualization technology more important. One important feature of virtualization is live Virtual Machine (VM) migration. Two main metrics of evaluating a live VM migration mechanism are total migration time and downtime. Most existing literature on live VM migration focus on designing migration mechanisms to shorten the two metrics or making a tradeoff between them. Few of them can be applied to applications with delay requirements, such as a VM backup process that needs to be done in a specific time. This will negatively impact the user experiences and reduce the profit of cloud service providers. Besides, the frequently varied bandwidth required by the widely used pre-copy mechanism is difficult to be provided by current network technologies. In this work, we theoretically analyze how much bandwidth is required to guarantee the total migration time and downtime of a live VM migration, and then propose a novel transport control mechanism to guarantee the computed bandwidth. The experimental results demonstrate that the bandwidth obtained from the proposed reciprocal-based model guarantees the expected total migration time and downtime, and the proposed transport control mechanism ensures that the live VM migration flow obtains the expected bandwidth even if there are background flows.
Jiao Zhang 0002, Fengyuan Ren, Ran Shu 0001, Tao Huang 0005, Yunjie Liu 0001
IEEE Trans. Computers3
2015 Slowing Little Quickens More: Improving DCTCP for Massive Concurrent Flows
abstract
DCTCP is a potential TCP replacement to satisfy the requirements of data center network. It receives wide concerns in both academic and industrial circles. However, DCTCP could only support tens of concurrent flows well and suffers timeouts and throughput collapse facing numerous concurrent flows. This is far from the requirement of data center network. Data centers employing partition/aggregation pattern usually involve hundreds of concurrent flows. In this paper, after tracing DCTCP's dynamic behavior through experiments, we explored two roots for DCTCP's failure under the high fan-in traffic pattern: (1) The regulation mechanism of sending window is ineffective when cwnd is decreased to the minimum size, (2) The bursts induced by synchronized flows with small cwnd cause fatal packet loss leading to severe timeouts. We enhance DCTCP to support massive concurrent flows by regulating the sending time interval and desynchronizing the sending time in particular conditions. The new protocol called DCTCP+ outperforms DCTCP when the number of concurrent flows increases to several hundreds. DCTCP+ can normally work to effectively support the short concurrent query responses in the benchmark from real production clusters, and keep the same good performance with the mixture of background traffic.
Mao Miao, Peng Cheng 0005, Fengyuan Ren, Ran Shu 0001
ICPP4
2015 Sliding Mode Congestion Control for Data Center Ethernet Networks
abstract
Recently, Ethernet is enhanced as the unified switch fabric of data centers, called data center Ethernet. One of the indispensable enhancements is end-to-end congestion management, and currently quantized congestion notification (QCN) has been ratified as the corresponding standard. However, our experiments show that QCN suffers from large oscillations of the queue length at the bottleneck link such that the buffer is emptied frequently and accordingly the link utilization degrades, with certain system parameters and network configurations. This phenomenon is corresponding to our theoretical analysis result that QCN fails to enter into the sliding mode motion (SMM) pattern with certain system parameters and network configurations. Knowing the drawbacks of QCN and realizing the advantage that congestion management system is insensitive to the changes of parameters and network configurations in the SMM pattern, we present sliding mode congestion control (SMCC), which can enter into the SMM pattern under any conditions. SMCC is simple, stable, fair, has short response time, and can be easily used to replace QCN because both of them follow the framework developed by the IEEE 802.1Qau work group. Experiments on the NetFPGA platform show that SMCC is superior to QCN, especially when traffic pattern and network states are variable.
Wanchun Jiang, Fengyuan Ren, Ran Shu 0001, Yongwei Wu 0001, Chuang Lin 0002
IEEE Trans. Computers3
2014 Analysing convergence of Quantized Congestion Notification in Data Center Ethernet
abstract
Enhancing Ethernet as the unified data center fabric to concurrently handle the traffic of Local Area Network (LAN), Storage Area Network (SAN), and High Performance Computing (HPC) has attracted much attention. Congestion management is one critical enhancement to fill the performance gap between traditional Ethernet and the unified data center fabric. Currently, Quantized Congestion Notification (QCN) has been approved as the standard congestion management mechanism. However, lots of work pointed out that QCN suffers from the problem of unfairness among different flows. In this paper, we found that QCN could achieve fairness, merely the convergence time to fairness is quite long. Thus, we build a convergence time model to investigate the reasons of the slow convergence process of QCN. The model indicates that the convergence time of QCN can be decreased if RPs have the same rate increase probability or the rate increase step becomes larger at steady state. We validate the precise of our model by comparing with experimental data on the NetFPGA platform. The results show that it well characterizes the convergence time to fairness of QCN. Based on the proposed model, the impact of QCN parameters, network parameters, and QCN variants on the convergence time is analysed. Finally, enlightened by the analysis, we proposed a mechanism, called QCN-T, which replaces the Byte Counter and Timer at sources with a single modified Timer, to reduce the convergence time of QCN.
Ran Shu 0001, Jiao Zhang 0002, Fengyuan Ren, Chuang Lin 0002
IWQoS1
2014 Catch the Whole Lot in an Action: Rapid Precise Packet Loss Notification in Data Center
Peng Cheng 0005, Fengyuan Ren, Ran Shu 0001, Chuang Lin 0002
NSDI3
2014 Sharing Bandwidth by Allocating Switch Buffer in Data Center Networks
abstract
In today's data centers, the round trip propagation delay is quite small. Therefore, switch buffer sizes are much larger than the Bandwidth Delay Product (BDP). Based on this observation, in this paper we introduce a new transport protocol which provides bandwidth Sharing by Allocating switch Buffer (SAB) for data centers. SAB sets the congestion windows for flows based on the buffer size of the switches along the path. On one hand, as long as the total buffer allocated to all the flows is larger than the BDP, the network bandwidth can be fully utilized. On the other hand, since SAB only allocates the buffer space to flows, the totally injected traffic will not exceed the network capacity. Thus, SAB rarely loses packets. SAB also reduces flow completion time by allowing flows to reach their fair share of bandwidth quickly. The results of a series of experiments and simulations demonstrate that SAB has the features of fast convergence and rare packet loss. It reduces the latency of short flows and solves theTCP Incast and TCP Outcast problems.
Jiao Zhang 0002, Fengyuan Ren, Xin Yue, Ran Shu 0001, Chuang Lin 0002
IEEE J. Sel. Areas Commun.4
2013 Ease the Queue Oscillation: Analysis and Enhancement of DCTCP
abstract
Because of the terrible performance of TCP protocol in data center environment, DCTCP has been proposed as a TCP replacement, which uses a simple marking mechanism at switches and a few amendments at end hosts to adjust congestion window based on the extent of the congestion in networks. Thus, DCTCP can make a proper tradeoff between high throughput and low latency. However, through our observation, we discover that DCTCP causes severe oscillation of queue under some parameters and network configuration. Our perceptual analysis concludes that the rough single-threshold marking mechanism may be the essential reason. Therefore, we propose Double-Threshold DCTCP as an improvement of DCTCP. Then, by applying describing function method in nonlinear control theory, we analyze the stability of both DCTCP and Double-Threshold DCTCP, and theoretically explain why Double-Threshold DCTCP is more stable than DCTCP. At last, we validate theoretical analysis and conclude that the Double- Threshold DCTCP can achieve smaller queue, and the queue length of Double-Threshold DCTCP is less sensitive to the growing number of flows. Further, Double-Threshold DCTCP can postpone the throughput collapse caused by Incast traffic and reduce the tail latency in completion time experiment.
Wen Chen 0026, Peng Cheng 0005, Fengyuan Ren, Ran Shu 0001, Chuang Lin 0002
ICDCS4
2012 Sliding Mode Congestion Control for data center Ethernet networks
abstract
Recently, Ethernet is being enhanced as the unified switch fabric of data centers, called Data Center Ethernet. The end-to-end congestion management is one of the indispensable enhancements, and Quantized Congestion Notification (QCN) has been ratified to be the standard. Our experiments show that QCN suffers from the oscillation of the queue at the bottleneck link. With the changes of system parameters and network configurations, the oscillation may become so serious that the queue is emptied frequently. As a result, the utilization of the bottleneck link degrades. Theoretical analysis shows that QCN approaches to the equilibrium point mainly through the sliding mode motion. But whether QCN enters into the sliding mode motion also depends on both system parameters and network configurations. Hence, we present the Sliding Mode Congestion Control (SMCC) scheme, which can drive the system into the sliding mode motion under any conditions. SMCC benefits from the advantage that the sliding mode motion is insensitive to system parameters and external disturbances. Moreover, SMCC is simple, stable and has short response time. QCN can be replaced by SMCC easily since both of them follow the framework developed by the IEEE 802.1 Qau work group. Experiments on the NetFPGA platform show that SMCC is superior to QCN, especially in the condition that the traffic pattern and the network state are variable.
Wanchun Jiang, Fengyuan Ren, Ran Shu 0001, Chuang Lin 0002
INFOCOM3