Rui Miao 0001

dblp:62/675-1 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-1436-2374ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 18 · 4 first-author · 6 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Connecting 100K+ GPUs: Building the Communication Stack for Large-Scale LLM Training
abstract
The arrival of 100K+ GPU clusters marks a new frontier in AI infrastructure. Standard communication stack meets new challenges as physical topologies span multiple datacenter buildings, introducing high bandwidth-delay product links where latency increases by up to 30× compared to intra-rack traffic. Furthermore, the transition toward Mixture-of-Experts architectures generating bursty all-to-all patterns that create transient congestion hotspots. These constraints, combined with an operational environment where hardware failures shift from anomalies to frequent occurrences, renders traditionally lightweight operations like initialization and resource management challenging.
Hongyi Zeng, Min Si, Pavan Balaji, Yongzhou Chen, Ching-Hsiang Chu, Adithya Gangidi, Prashanth Kannan, Bingzhe Liu, Saif Hasan, Deep Shah, Ashmitha Jeevaraj Shetty, Gregory R. Steinbrecher, Srikanth Sundaresan, Yulun Wang, Yexin Wu, Mingran Yang, Kenny Yu, Minlan Yu, Cen Zhao, Shengbao Zheng, Wesley Bland, Denis Boyda, Suman Gumudavelli, Subodh Iyengar, Cristian Lumezanu, Rui Miao 0001, Venkat Ramesh, Jingliang Ren, Maxim Samoylov, Jan Seidel, Qiye Tan, Xinfeng Xie, Yimeng Zhao, Shuqiang Zhang, Art Zhu
SIGCOMM29
2024 RDMA over Ethernet for Distributed Training at Meta Scale
abstract
The rapid growth in both computational density and scale in AI models in recent years motivates the construction of an efficient and reliable dedicated network infrastructure. This paper presents the design, implementation, and operation of Meta's Remote Direct Memory Access over Converged Ethernet (RoCE) networks for distributed AI training.
Adithya Gangidi, Rui Miao 0001, Shengbao Zheng, Sai Jayesh Bondu, Guilherme Loch Waltrick Goes, Hany Morsy, Rohit Puri, Mohammad Riftadi, Ashmitha Jeevaraj Shetty, Shuqiang Zhang, Mikel Jimenez Fernandez, Shashidhar Gandham, Hongyi Zeng
SIGCOMM2
2024 Alibaba HPN: A Data Center Network for Large Language Model Training
abstract
This paper presents HPN, Alibaba Cloud's data center network for large language model (LLM) training. Due to the differences between LLMs and general cloud computing (e.g., in terms of traffic patterns and fault tolerance), traditional data center networks are not well-suited for LLM training. LLM training produces a small number of periodic, bursty flows (e.g., 400Gbps) on each host. This characteristic of LLM training predisposes Equal-Cost Multi-Path (ECMP) to hash polarization, causing issues such as uneven traffic distribution. HPN introduces a 2-tier, dual-plane architecture capable of interconnecting 15K GPUs within one Pod, typically accommodated by the traditional 3-tier Clos architecture. Such a new architecture design not only avoids hash polarization but also greatly reduces the search space for path selection. Another challenge in LLM training is that its requirement for GPUs to complete iterations in synchronization makes it more sensitive to singlepoint failure (typically occurring on ToR). HPN proposes a new dual-ToR design to replace the single-ToR in traditional data center networks. HPN has been deployed in our production for more than eight months. We share our experience in designing, and building HPN, as well as the operational lessons of HPN in production.
Kun Qian 0021, Yongqing Xi, Jiamin Cao, Yichi Xu, Yu Guan 0005, Binzhang Fu, Xuemei Shi, Fangbo Zhu, Rui Miao 0001, Peng Wang 0185, Xianlong Zeng, Eddie Ruan, Zhiping Yao, Ennan Zhai, Dennis Cai
SIGCOMM10
2023 Flor: An Open High Performance RDMA Framework Over Heterogeneous RNICs
Qiang Li 0045, Yixiao Gao, Xiaoliang Wang 0001, Haonan Qiu, Yanfang Le, Derui Liu, Qiao Xiang, Bo Li 0061, Jianbo Dong, Lingbo Tang, Hongqiang Harry Liu, Shaozong Liu, Rui Miao 0001, Yaohui Wu, Zhiwu Wu, Zheng Cao 0003, Zhongjie Wu, Chen Tian 0001, Guihai Chen, Dennis Cai, Jiaji Zhu, Jiesheng Wu, Jiwu Shu
OSDI16
2023 Dependable Virtualized Fabric on Programmable Data Plane
abstract
In modern multi-tenant data centers, each tenant desires reassuring dependability from the virtualized network fabric – bandwidth guarantee with work conservation, bounded tail latency and resilient reachability. However, the slow convergence of prior works under network dynamics and uncertainties can hardly provide the dependability for tenants. Further, state-of-the-art load balance schemes are guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. In this paper, we propose vFab, a dependable virtualized fabric framework which can (1) quickly detect network failure in data plane, (2) explicitly select proper paths for all flows, and (3) converge to ideal bandwidth allocation at sub-millisecond. The core idea of vFab is to leverage the programmable data plane to build a fusion of an active edge (e.g., NIC) and an informative core (e.g., switch), where the core sends link status and tenant information to the edge via telemetry to help the latter make a timely and accurate decision on path selection and traffic admission. We fully implement vFab with commodity SmartNICs and programmable switches. Extensive evaluations show that vFab can keep bandwidth guarantee with high bandwidth utilization, low and bounded latency, and resilient reachability under various network scenarios with limited overhead. Application-level experiments show that vFab can improve QPS by$2.4\times $and cut tail latency by$10\times $compared to the alternatives.
Kaihui Gao, Shuai Wang 0028, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Tao Sun 0010
IEEE/ACM Trans. Netw.5
2022 Predictable vFabric on informative data plane
abstract
In multi-tenant data centers, each tenant desires reassuring predictability from the virtual network fabric - bandwidth guarantee, work conservation, and bounded tail latency. Achieving these goals simultaneously relies on rapid and precise traffic admission. However, the slow convergence (tens of milliseconds) of prior works can hardly satisfy the increasingly rigorous performance demand under dynamic traffic patterns. Further, state-of-the-art load balance schemes are all guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works.
Shuai Wang 0028, Kaihui Gao, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Ming Zhang 0005
SIGCOMM5
2022 From luna to solar: the evolutions of the compute-to-storage networks in Alibaba cloud
abstract
This paper presents the two generations of storage network stacks that reduced the average I/O latency of Alibaba Cloud's EBS service by 72% in the last five years: Luna, a user-space TCP stack that corresponds the latency of network to the speed of SSD; and Solar, a storage-oriented UDP stack that enables both storage and network hardware accelerations.
Rui Miao 0001, Lingjun Zhu, Kun Qian 0021, Shujun Zhuang, Bo Li 0061, Shuguang Cheng, Binzhang Fu, Jiaji Zhu, Jiesheng Wu, Dennis Cai, Hongqiang Harry Liu
SIGCOMM1
2020 Flow Event Telemetry on Programmable Data Plane
abstract
Network performance anomalies (NPAs), e.g. long-tailed latency, bandwidth decline, etc., are increasingly crucial to cloud providers as applications are getting more sensitive to performance. The fundamental difficulty to quickly mitigate NPAs lies in the limitations of state-of-the-art network monitoring solutions --- coarse-grained counters, active probing, or packet telemetry either cannot provide enough insights on flows or incur too much overhead. This paper presents NetSeer, a flow event telemetry (FET) monitor which aims to discover and record all performance-critical data plane events, e.g. packet drops, congestion, path change, and packet pause. NetSeer is efficiently realized on the programmable data plane. It has a high coverage on flow events including inter-switch packet drop/corruption which is critical but also challenging to retrieve the original flow information, with novel intra- and inter-switch event detection algorithms running on data plane; NetSeer also achieves high scalability and accuracy with innovative designs of event aggregation, information compression, and message batching that mainly run on data plane, using switch CPU as complement. NetSeer has been implemented on commodity programmable switches and NICs. With real case studies and extensive experiments, we show NetSeer can reduce NPA mitigation time by 61%-99% with only 0.01% overhead of monitoring traffic.
Yu Zhou 0008, Chen Sun 0005, Hongqiang Harry Liu, Rui Miao 0001, Bo Li 0061, Zhilong Zheng, Lingjun Zhu, Yongqing Xi, Dennis Cai, Ming Zhang 0005, Mingwei Xu 0001
SIGCOMM4
2020 Lyra: A Cross-Platform Language and Compiler for Data Plane Programming on Heterogeneous ASICs
abstract
Programmable data plane has been moving towards deployments in data centers as mainstream vendors of switching ASICs enable programmability in their newly launched products, such as Broadcom's Trident-4, Intel/Barefoot's Tofino, and Cisco's Silicon One. However, current data plane programs are written in low-level, chip-specific languages (e.g., P4 and NPL) and thus tightly coupled to the chip-specific architecture. As a result, it is arduous and error-prone to develop, maintain, and composite data plane programs in production networks. This paper presents Lyra, the first cross-platform, high-level language & compiler system that aids the programmers in programming data planes efficiently. Lyra offers a one-big-pipeline abstraction that allows programmers to use simple statements to express their intent, without laboriously taking care of the details in hardware; Lyra also proposes a set of synthesis and optimization techniques to automatically compile this "big-pipeline" program into multiple pieces of runnable chip-specific code that can be launched directly on the individual programmable switches of the target network. We built and evaluated Lyra. Lyra not only generates runnable real-world programs (in both P4 and NPL), but also uses up to 87.5% fewer hardware resources and up to 78% fewer lines of code than human-written programs.
Ennan Zhai, Hongqiang Harry Liu, Rui Miao 0001, Yu Zhou 0008, Bingchuan Tian, Chen Sun 0005, Dennis Cai, Ming Zhang 0005, Minlan Yu
SIGCOMM4
2019 DETER: Deterministic TCP Replay for Performance Diagnosis
Rui Miao 0001, Mohammad Alizadeh, Minlan Yu
NSDI2
2019 HPCC: high precision congestion control
abstract
Congestion control (CC) is the key to achieving ultra-low latency, high bandwidth and network stability in high-speed networks. From years of experience operating large-scale and high-speed RDMA networks, we find the existing high-speed CC schemes have inherent limitations for reaching these goals. In this paper, we present HPCC (High Precision Congestion Control), a new high-speed CC mechanism which achieves the three goals simultaneously. HPCC leverages in-network telemetry (INT) to obtain precise link load information and controls traffic precisely. By addressing challenges such as delayed INT information during congestion and overreac-tion to INT information, HPCC can quickly converge to utilize free bandwidth while avoiding congestion, and can maintain near-zero in-network queues for ultra-low latency. HPCC is also fair and easy to deploy in hardware. We implement HPCC with commodity programmable NICs and switches. In our evaluation, compared to DCQCN and TIMELY, HPCC shortens flow completion times by up to 95%, causing little congestion even under large-scale incasts.
Rui Miao 0001, Hongqiang Harry Liu, Lingbo Tang, Zheng Cao 0003, Ming Zhang 0005, Frank Kelly, Mohammad Alizadeh, Minlan Yu
SIGCOMM2
2017 SilkRoad: Making Stateful Layer-4 Load Balancing Fast and Cheap Using Switching ASICs
abstract
In this paper, we show that up to hundreds of software load balancer (SLB) servers can be replaced by a single modern switching ASIC, potentially reducing the cost of load balancing by over two orders of magnitude. Today, large data centers typically employ hundreds or thousands of servers to load-balance incoming traffic over application servers. These software load balancers (SLBs) map packets destined to a service (with a virtual IP address, or VIP), to a pool of servers tasked with providing the service (with multiple direct IP addresses, or DIPs). An SLB is stateful, it must always map a connection to the same server, even if the pool of servers changes and/or if the load is spread differently across the pool. This property is called per-connection consistency or PCC. The challenge is that the load balancer must keep track of millions of connections simultaneously.
Rui Miao 0001, Hongyi Zeng, Changhoon Kim, Jeongkeun Lee, Minlan Yu
SIGCOMM1
2016 LossRadar: Fast Detection of Lost Packets in Data Center Networks
abstract
Packet losses are common in data center networks, may be caused by a variety of reasons (e.g., congestion, blackhole), and have significant impacts on application performance and network operations. Thus, it is important to provide fast detection of packet losses independent of their root causes. We also need to capture both the locations and packet header information of the lost packets to help diagnose and mitigate these losses. Unfortunately, existing monitoring tools that are generic in capturing all types of network events often fall short in capturing losses fast with enough details and low overhead. Due to the importance of loss in data centers, we propose a specific monitoring system designed for loss detection. We propose LossRadar, a system that can capture individual lost packets and their detailed information in the entire network on a fine time scale. Our extensive evaluation on prototypes and simulations demonstrates that LossRadar is easy to implement in hardware switches, achieves low memory and bandwidth overhead, while providing detailed information about individual lost packets. We also build a loss analysis tool that demonstrates the usefulness of LossRadar with a few example applications.
Rui Miao 0001, Changhoon Kim, Minlan Yu
CoNEXT2
2016 FlowRadar: A Better NetFlow for Data Centers
Rui Miao 0001, Changhoon Kim, Minlan Yu
NSDI2
2015 The Dark Menace: Characterizing Network-based Attacks in the Cloud
abstract
As the cloud computing market continues to grow, the cloud platform is becoming an attractive target for attackers to disrupt services and steal data, and to compromise resources to launch attacks. In this paper, using three months of NetFlow data in 2013 from a large cloud provider, we present the first large-scale characterization of inbound attacks towards the cloud and outbound attacks from the cloud. We investigate nine types of attacks ranging from network-level attacks such as DDoS to application-level attacks such as SQL injection and spam. Our analysis covers the complexity, intensity, duration, and distribution of these attacks, highlighting the key challenges in defending against attacks in the cloud. By characterizing the diversity of cloud attacks, we aim to motivate the research community towards improving future security solutions for cloud systems.
Rui Miao 0001, Rahul Potharaju, Minlan Yu, Navendu Jain
Internet Measurement Conference1
2014 DIBS: just-in-time congestion mitigation for data centers
abstract
Data centers must support a range of workloads with differing demands. Although existing approaches handle routine traffic smoothly, intense hotspots--even if ephemeral--cause excessive packet loss and severely degrade performance. This loss occurs even though congestion is typically highly localized, with spare buffer capacity at nearby switches. In this paper, we argue that switches should share buffer capacity to effectively handle this spot congestion without the monetary hit of deploying large buffers at individual switches. Specifically, we present detour-induced buffer sharing (DIBS), a mechanism that achieves a near lossless network without requiring additional buffers at individual switches. Using DIBS, a congested switch detours packets randomly to neighboring switches to avoid dropping the packets. We implement DIBS in hardware, on software routers in a testbed, and in simulation, and we demonstrate that it reduces the 99th percentile of delay-sensitive query completion time by up to 85%, with very little impact on other traffic.
Kyriakos Zarifis, Rui Miao 0001, Matt Calder, Ethan Katz-Bassett, Minlan Yu, Jitendra Padhye
EuroSys2
2014 NIMBUS: cloud-scale attack detection and mitigation
abstract
No abstract available.
Rui Miao 0001, Minlan Yu, Navendu Jain
SIGCOMM1
2013 Software Defined Traffic Measurement with OpenSketch
Minlan Yu, Lavanya Jose, Rui Miao 0001
NSDI3
2013 Don't drop, detour!
abstract
Today's data centers must support a range of workloads with different demands. While existing approaches handle routine traffic smoothly, ephemeral but intense hotspots cause excessive packet loss and severely degrade performance. This loss occurs even though the congestion is typically highly localized, with spare buffer capacity available at nearby switches.
Matt Calder, Rui Miao 0001, Kyriakos Zarifis, Ethan Katz-Bassett, Minlan Yu, Jitendra Padhye
SIGCOMM2
2013 SIMPLE-fying middlebox policy enforcement using SDN
abstract
Networks today rely on middleboxes to provide critical performance, security, and policy compliance capabilities. Achieving these benefits and ensuring that the traffic is directed through the desired sequence of middleboxes requires significant manual effort and operator expertise. In this respect, Software-Defined Networking (SDN) offers a promising alternative. Middleboxes, however, introduce new aspects (e.g., policy composition, resource management, packet modifications) that fall outside the purvey of traditional L2/L3 functions that SDN supports (e.g., access control or routing).
Zafar Ayyub Qazi, Cheng-Chun Tu, Luis Chiang, Rui Miao 0001, Vyas Sekar, Minlan Yu
SIGCOMM4