Shuhong Zhu

dblp:176/1171 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0000-5404-797XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 4 since 2021
YearPublicationVenuePosition
2026 From Nimitz to NetPila: The Evolution of Production-Scale Container Network
Sheng Cheng 0002, Jiamin Cao, Shuhong Zhu, Ennan Zhai, Dennis Cai
SIGCOMM7
2025 Mitigating Scalability Walls of RDMA-based Container Networks
Wei Liu 0148, Kun Qian 0021, Zhenhua Li 0001, Feng Qian 0001, Tianyin Xu, Yunhao Liu 0001, Yu Guan 0005, Shuhong Zhu, Hongfei Xu, Lanlan Xi, Ennan Zhai
NSDI8
2025 SkeletonHunter: Diagnosing and Localizing Network Failures in Containerized Large Model Training
abstract
The flexibility and portability characteristics have made containers a popular serverless environment for large model training in recent years. Unfortunately, these advantages render the network support for containerized large model training extremely challenging, due to the high dynamics of containers, the complex interplay between underlay and overlay networks, and the stringent requirements on failure detection and localization. Existing data center network debugging tools, which rely on comprehensive or opportunistic monitoring, are either inefficient or inaccurate in this setting.
Wei Liu 0148, Kun Qian 0021, Zhenhua Li 0001, Tianyin Xu, Yunhao Liu 0001, Jiakang Li, Shuhong Zhu, Xue Li 0024, Hongfei Xu, Ennan Zhai
SIGCOMM9
2025 Alibaba Stellar: A New Generation RDMA Network for Cloud AI
abstract
The rapid adoption of Large Language Models (LLMs) in cloud environments has intensified the demand for high-performance AI training and inference, where Remote Direct Memory Access (RDMA) plays a critical role. However, existing RDMA virtualization solutions, such as Single-Root Input/Output Virtualization (SR-IOV), face significant limitations in scalability, performance, and stability. These issues include lengthy container initialization times, hardware resource constraints, and inefficient traffic steering. To address these challenges, we propose Stellar, a new generation RDMA network for cloud AI. Stellar introduces three key innovations: Para-Virtualized Direct Memory Access (PVDMA) for on-demand memory pinning, extended Memory Translation Table (eMTT) for optimized GPU Direct RDMA (GDR) performance, and RDMA Packet Spray for efficient multi-path utilization. Deployed in our large-scale AI clusters, Stellar spins up virtual devices in seconds, reduces container initialization time by 15 times, and improves LLM training speed by up to 14%. Our evaluations demonstrate that Stellar significantly outperforms existing solutions, offering a scalable, stable, and high-performance RDMA network for cloud AI.
Menglei Zheng, Binbin Liao, Suwei Xu, Yongjia Mo, Qinghua Peng, Jilie Luo, Qingxu Li, Zishu Wang, Jianbo Dong, Kunling He, Sheng Cheng 0002, Jiamin Cao, Hairong Jiao, Lingjun Zhu, Yiquan Chen, Wei Wang 0030, Shuhong Zhu, Xingru Li, Qiang Wang 0022, Wei Lin 0016, Ennan Zhai, Jiesheng Wu, Qiang Liu 0036, Binzhang Fu, Dennis Cai
SIGCOMM29
2015 SDN-based TCP congestion control in data center networks
abstract
TCP incast usually happens when a receiver requests the data from multiple senders simultaneously. This many-to-one communication pattern constantly appears in the data center networks due to the data are stored at multiple servers. With Software Defined Networks (SDN), the centralized control methods and the global view of the network can be an effective way to handle this problem. In this paper, we propose a SDN-based TCP (SDTCP) congestion control mechanism at network side. Our approach enables controller to select a long-lived flow to reduce sending rate by adjusting the TCP receive window of ACK packet after OpenFlow-switch triggered a congestion message to controller. The key benefit of SDTCP is that, with global perspective, we can accurately decelerate the rate of long-lived flow to ensure the performance other flows. The experiments indicate that we can achieve almost zero packet loss for TCP incast and guarantee goodput for the high propriety flows.
Yifei Lu 0001, Shuhong Zhu
IPCCC2