Lide Suo

dblp:361/5256 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2025
0009-0001-9370-7878ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021Computer networks · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Fork: A Dual Congestion Control Loop for Small and Large Flows in Datacenters
abstract
Many existing transport designs aim to deliver ultra-low latency and high bandwidth for applications in high-speed datacenter networks. However, almost all of them intertwine the control of small and large flows using the same control entity (e.g., sender or receiver) and congestion feedback signal (e.g., ECN or credit), thus bringing significant performance impairments. By contrast, we seek to decouple the rate control of small flows from that of large ones.
Wenxin Li 0001, Yulong Li 0001, Lide Suo, Xuan Gao 0001, Xin Xie 0001, Sheng Chen 0015, Ziqi Fan, Wenyu Qu, Guyue Liu
EuroSys4
2025 Lark: A Buffer-aware Building Block for Programmable Packet Scheduling in Datacenters
abstract
Programmable packet scheduling enables users to customize scheduling algorithms flexibly without designing new ASICs. Existing schemes prefre to approximate optimal Push-In First-Out (PIFO) using First-In First-Out (FIFO) queues in commodity programmable switches. Despite its availability, these schemes suffer performance degradation due to the unawareness of available switch buffer. To be specific, when the port buffer is drained, existing schemes discard all incoming packets, even though these packets have higher priorities than the enqueued packets. In this paper, we reveal that the problem's key culprit is the lack of coordination between buffer management and packet scheduling in the switch. To fill this gap, we present Lark, a buffer-aware building block for programmable scheduling schemes designed to solve the above problem. Its key idea is to proactively drop the low-priority packets when the allocated buffer is to be drained, thereby admitting the later-arriving high-priority packets. Lark contains two modules, a lightweight gradient-based online prediction module and a simple priority-based decision module. Lark relies the former module to identify whether the allocated buffer is to be drained and uses the later to determine whether to drop the incoming packet. We have integrated Lark into two representative schemes, SP-PIFO and AIFO. Our large-scale evaluations over three realistic workloads show that Lark can significantly optimize their key metrics without sacrificing throughnut.
Song Zhang 0008, Wenxin Li 0001, Yulong Li 0001, Lide Suo, Sheng Chen 0015, Yitao Hu, Laiping Zhao, Keqiu Li
INFOCOM5
2024 WQEFC: A Scalable and Low-Latency RDMA Messages Scheduler for Mixed Messages
abstract
RDMA has been widely deployed to improve the performance of applications with frequently fine-grained remote access. However, restricted on-chip resources result in cache misses under high concurrency that significantly degrade network performance. QPC-aware solutions only focus on the number of concurrent QPs, ignoring the impact of WQE within QPs. SMART limits the number of WQEs in each QP with a credit-based scheme. Nevertheless, we find that equal treatment increases the tail latency of messages ranging from 32 bytes to 1024 bytes by 2× when mixing messages of different sizes. In this paper, we introduce WQEFC, a scalable RDMA message scheduler that provides lower latency and higher throughput for applications with heavily concurrent messages. Our key insight is that there is a significant difference in the sensitivity to cache miss between messages of different sizes. For messages smaller than 32 bytes, which are sensitive to cache misses, we combine the credit limiter and sub-message poller, limiting the number of concurrent wqes to avoid cache miss while ensuring optimal message completion latency. For other messages, which are insensitive to cache miss, we assign them a higher priority and use the sub-message poller to ensure message concurrency while reducing cache miss. We implement WQEFC as a middleware between the driver layer and the application layer for flexible deployment. WQEFC outperforms the state-of-the-art solution Smart by increasing system throughput by 61.3%, and reducing the tail latency of messages smaller than 32 bytes and larger than 32 bytes by 41.6% and 72.6%, respectively.
Yaozhen Li, Lide Suo, Xiancheng Meng, Yiren Pang, Wenxin Li 0001, Keqiu Li, Yitao Hu
HPCC3
2024 PPT: A Pragmatic Transport for Datacenters
abstract
This paper introduces PPT, a pragmatic transport that achieves comparable performance to proactive transports while maintaining good deployability as reactive transports. Our key idea is to run a low-priority control loop to leverage the available bandwidth left by the reactive transports. The main challenge is to send just enough packets to improve performance without harming the primary control loop. We combine two unconventional techniques: an intermittent loop initialization and an exponential window decrease, enabling us to dynamically identify and fill the spare bandwidth. We further complement PPT's design with a buffer-aware flow scheduling scheme to optimize the average FCT of small flows without prior knowledge of flow size information. We have implemented a PPT prototype in the Linux kernel with ~400 lines of code and demonstrated that compared to Homa, it delivers up to 46.3% lower overall average FCT and even 25%/55.5% lower average/tail FCT of small flows in an Memcached workload.
Lide Suo, Yiren Pang, Wenxin Li 0001, Renjie Pei, Keqiu Li, Xiulong Liu 0001, Xin He 0043, Yitao Hu, Guyue Liu
SIGCOMM1
2024 Anole: Scheduling Flows for Fast Datacenter Networks With Packet Re-Prioritization
abstract
Many existing datacenter transports perform one-shot packet priority tagging at end-hosts and leave them fixed during the packet's transmission. In this paper, we experimentally show that: (1) such fixed packet priority is not sufficient for FCT (flow completion time) minimization, and (2) adjusting packet transmission priority in the network requires effective coordination among switches. Building on these insights, we present Anole, a new datacenter transport that advocates packet re-prioritization in near-bottleneck switches to minimize FCT. To this end, Anole integrates three simple-yet-effective techniques. First, it employs an in-network telemetry (INT) based approach to dynamically detect the bottleneck for each flow. Second, it adopts an on-off rate control mechanism for each sender to pause heavily congested flows but send lightly- and non-congested ones. Last, it leverages an altruistic scheduling policy at each switch to let the flows whose next hops are bottleneck switches give way to others. We implement an Anole prototype based on DPDK and show, through both testbed experiments and simulations, that Anole delivers significant performance advantages. For example, compared to EPN, Homa, and Aeolus, it shortens the average FCT of all (small) flows by up to 61.6% (89.1%).
Song Zhang 0008, Lide Suo, Wenxin Li 0001, Yulong Li 0001, Keqiu Li
IEEE Trans. Cloud Comput.2
2024 BRT: Buffer Management for RDMA/TCP Mix-Flows in Datacenter Networks
abstract
The coexistence of RDMA and TCP is prevalent in the datacenter. Despite the sound isolation at the end hosts, they share the same switches in the network. Their different networking behaviors (E.g., in hardware demand and transport protocols) lead to huge differentiated buffer demand for switches. However, existing buffer management schemes ignore these dissimilarities and simply treat such RDMA/TCP mix-flows as the typical multi-class traffic, resulting in inferior isolation and degrading networking performances. This paper presents BRT, a first systematic solution for the buffer management of RDMA/TCP mix-flows in the DCN. BRT’s key insight is to allocate buffer with the awareness of traffic’s networking characteristics while minimally impacting the other’s performance. Guided by this insight, it first employs a traffic characteristics-based window to detect whether queues are in the state of persistent long queues. Then, it adjusts the total allocated buffer for each traffic type based on the number of persistent long queues and the normalized dequeue rates to reduce the buffer occupancy of meaningless queuing. Last, it calculates the buffer threshold for RDMA/TCP queues separately and uses a simple yet effective approach to prioritize the absorption of small flows. Our large-scale packet-level evaluations show that BRT can effectively optimize the networking performances for RDMA/TCP mix-flows. For example, compared to current practice, BRT achieves up to 53.5%, 46.7%, and 48.5% lower average FCT for incast flows, RDMA small flows, and TCP small flows, respectively, without sacrificing the overall throughput.
Song Zhang 0008, Wenxin Li 0001, Lide Suo, Yulong Li 0001, Jien Kato, Keqiu Li
IEEE Trans. Netw. Serv. Manag.3
2023 MiddleCache: Accelerating TCP based In-memory Key-value Stores using eBPF
abstract
In-memory key-value stores are widely used in modern web services to support large-scale user requests by caching popular data. Their performance is critical, and BMC, the state-of-the-art work, builds an in-kernel cache and processes requests before the stack using eBPF to reduce the overhead of the kernel network stack. However, BMC fails to support stateful protocol TCP because pre-stack processing creates TCP state bias between the client and server.TCP is widely used by in-memory key-value stores, is even the only choice for some applications (e.g., Redis), and also suffers from performance issues. In this work, we present MiddleCache, a TCP-enabled in-memory key-value store acceleration design. Our key observation is that the TCP state bias of the client and server can be inferred and eliminated with packet length. The design of MiddleCache has two key parts: (i) A compact TCP state maintenance mechanism that accumulates packet lengths and applies corrections to the packet header, which realize TCP support within the constrains of eBPF. (ii) Lock-free accumulation counters that support high-performance concurrent access by utilizing Receive Side Scaling (RSS). Our experiments show that, compared with Memcached, MiddleCache reduces 56% processing latency on cache hit and achieves a 3.8× throughput improvement on Facebook-like small-size requests workload.
Yiren Pang, Sheng Chen 0015, Wenxin Li 0001, Yulong Li 0001, Xin He 0043, Song Zhang 0008, Zewei Guan, Lide Suo
ICPADS9
2023 Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay Network
abstract
Container overlay network, though being widely adopted to enable communication between containers on different hosts, is a key downside for latency-sensitive applications. The state-of-the-art solution seeks to shorten the data path in packet processing by replacing overlay connection file descriptors with host namespace ones. While promising, it must block each overlay connection until the relevant host connection is set up, thus heavily influencing the request latency. In this paper, we present ShuntFlow, a systematic data delivery framework that seamlessly integrates the host and overlay networks to reduce the application's request-response latency. ShuntFlow first lets all connections flow in the overlay network directly. Then, it adopts a simple-yet-effective syscall-threshold-based mechanism to pick appropriate connections and switches their data delivery to the host network in a blocking-free way using a multi-threading technique. As such, unnecessary connection switches are prevented; yet, the pre-setup phase dilemma is eliminated. We have implemented a ShuntFlow prototype based on Linux and Docker and evaluated it extensively on a 40 Gbps testbed. The results show that ShuntFlow achieves 13%/72% and 19%/69% reductions, in average/tail request-response latency of a web server and an in-memory key-value store, respectively, while incurring less CPU overhead, compared to Slim.
Wenxin Li 0001, Yiren Pang, Renjie Pei, Yitao Hu, Lide Suo, Keqiu Li
IEEE Trans. Parallel Distributed Syst.7