Xianneng Zou

dblp:356/7957 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
14since 2021 · last 2026
0009-0005-5370-8162ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 13 · 1 first-author · 13 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EdgePoW: Adaptive Ingress-Aware Defense with Non-Interactive PoW Against Volumetric SYN Floods
abstract
The stability of Internet services is persistently challenged by large volumetric TCP SYN floods, for which conventional defenses such as SYN Cookies preserve server state but still amplify bandwidth pressure. This paper presents EdgePoW, an ingress aware defense architecture that integrates non interactive Proof of Work with an SDN control plane for managed edge networks. The controller monitors per ingress SYN pressure and raises PoW difficulty when flooding is detected. If traffic mainly originates from a stable source region, enforcement is refined to the offending source prefix to reduce overhead on benign co located clients; otherwise, ingress wide enforcement is retained under randomized or spoofed sources. We further design a conservative Difficulty Discovery Protocol that reuses TCP retransmissions and commits difficulty updates only after a successful handshake. Experiments on a custom SDN testbed show restored application QoS under concentrated and spoofed floods, 11.7% higher benign client throughput than ingress only enforcement, and below 0.8% transient false escalations under 2% random loss.
Wenyang Jia, Xianneng Zou, Kai Lei
APNet3
2026 Multipath Collective Communication Beyond Scale-up Networks in GPU Clouds
Yuchen Xu 0003, Jianglong Nie, Baojia Li 0002, Mingzhuo Chen, Guanyu Qu, Zhenchuan Liu, Shuangshuang Yin, Chunzhi He, Yinben Xia, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Congcong Miao, Wenfei Wu
EuroSys16
2026 SwiftEP: Accelerating MoE Inference with Buffer Fusion and TMA Offloading
Xingyi Li 0004, Shangguang Wang, Zhehao Lin, Yinben Xia, Qihang Liu, Xiang Li 0067, Zekun He, Yachen Wang, Xianneng Zou
NSDI16
2026 MirrorNet: High-fidelity and Scalable Network Emulation for Software-defined WAN
Congcong Miao, Yuejie Wang, Xuefeng Ji, Guozhi Shan, Pan Fang, Yanke Zhang, Xianneng Zou, Guyue Liu
NSDI10
2026 Cost-effective and Reliable Global Internet Peering with Programmable Switches
Congcong Miao, Zhiyi Yao, Jianchao Lv, Jinglin Wang, Shihan Lin, Xinyi Zhang 0004, Yunming Xiao, Jiwu Bu, Yachen Wang, Xianneng Zou, Yong Jiang 0001, Marco Canini, Gaogang Xie
NSDI11
2026 A Composable Emulation Framework for Whitebox Switches
Congcong Miao, Xianneng Zou, Chuwen Zhang, Qihang Liu, Zhijie Yan, Yanke Zhang, Yong Jiang 0001, Qiao Xiang, Xin Jin 0008, Zili Meng, Ang Chen 0001
NSDI2
2026 Pegasus: A Data Center Network for Bare-Metal AI Cloud
abstract
Today, AI cloud is key to serving diverse users with AI services, where cloud networking forms the basis. In this paper, we share our experience in designing, deploying, and operating Pegasus, a data center network tailored for the AI cloud, along with operational lessons learned from its deployment. The key designs of Pegasus include: 1) Network virtualization: a DPU-RNIC decoupled collaborative hardware architecture to enable a single DPU to virtualize multiple RNICs while reducing the power consumption. We design two-level flow tables on both DPU and RNICs to support underlay-overlay IP address translation and ensure isolation. For DPU-RNIC communication, we introduce a per-RNIC communication state machine to reduce communication overhead. 2) Network transport: customized and transparent transport offloading in the RNIC for low-latency and high-throughput communication performance for various AI workloads. We carefully offload per-packet load balancing and credit-based congestion control in RNICs, optimizing reorder delay and eliminating the impacts of hardware jitter. Pegasus has been deployed in production for over two years, currently covering 8K GPUs and supporting a wide range of tenants' AI applications.
Xianneng Zou, Zhaoxun Zhou, Xingda Wei, Zhaohe Chen, Yinben Xia, Lizhou Gao, Jiajun Liang, Chunxu Zhao, Jiewei Yang, Yunpeng Guan, Dongbo Gu, Chao Pei, Zekun He, Yachen Wang
SIGCOMM1
2026 Accelerating Hardware/Software Combined Traffic Processing With Fast and Efficient Asynchronous Flow Offloading
abstract
eHardware/software (hw/sw) combined systems are necessary to meet modern clouds’ requirements for processing huge amounts of network traffic by efficiently offloading large flows to hardware. However, existing hw/sw flow offloading systems typically perform traffic statistics collection and large flow selection within a time window—a time-window-based approach. Their offloading decision of large flows issynchronizedin the unit of a time window, which is mismatched to the asynchronous and dynamic nature of each flow’s sending rate. Additionally, the flow measurement and selection for large flows are decoupled in these solutions, leading to memory and CPU inefficiency. In this paper, we introduce TAO, a novel solution to the hw/sw combined flow offloading problem byasynchronouslyselecting and offloading flows based on flow table entries. TAO can reactfasterto the rapid dynamics of flows by taking actions at each table entry and ismore efficientby coupling flow measurement and selection into the entry. We have implemented a full-fledged TAO prototype based on the P4 switch and DPDK. Testbed results demonstrate that TAO can offload ∼16% more traffic to hardware, outperforming existing solutions by achieving 42× lower memory overhead. Meanwhile, it reduces software CPU utilization by 66.7% and cuts tail forwarding latency by 95.59% compared to state-of-the-art methods.
Xijin Yin, Yuanwei Lu, Xin Zhang 0117, Xingtong Lin, Shengli Zheng, Bangwen Deng, Xianneng Zou, Yachen Wang, Guo Chen 0001
IEEE Trans. Netw.7
2025 Holmes: Localizing Irregularities in LLM Training with Mega-scale GPU Clusters
Zhiyi Yao, Pengbo Hu, Congcong Miao, Xuya Jia, Zuning Liang, Yuedong Xu 0001, Chunzhi He, Mingzhuo Chen, Xiang Li 0010, Zekun He, Yachen Wang, Xianneng Zou, Junchen Jiang
NSDI13
2025 Astral: A Datacenter Infrastructure for Large Language Model Training at Scale
abstract
The flourishing of Large Language Models (LLMs) calls for increasingly ultra-scale training. In this paper, we share our experience in designing, deploying, and operating our novel Astral datacenter infrastructure, along with operational lessons and evolutionary insights gained from its production use. Astral has three important innovations: (i) a same-rail interconnection network architecture on tier-2, which enables the scaling of LLM training. To physically deploy this high-density infrastructure, we introduce a distributed high-voltage direct current power system and a new air-liquid integrated cooling system. (ii) a full-stack monitoring system featuring cross-host and hierarchical logging correlation, which diagnoses failures at scale and precisely localizes root causes. (iii) an operator-granular forecasting component Seer that efficiently generates operator execution timelines with acceptable accuracy, aiding in fault diagnosis, model tuning, and network architecture upgrading. Astral infrastructure has been gradually deployed over 18 months, supporting LLM training and inference for multiple customers.
Qingkai Meng 0001, Zhenhui Zhang, ChonLam Lao, Chengyuan Huang, Baojia Li 0002, Weizhen Dang, Zitong Lin, Yuanyuan Gong, Chunzhi He, Xiaoyuan Hu, Yinben Xia, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Kun Yang 0001, Gianni Antichi, Guihai Chen, Chen Tian 0001
SIGCOMM20
2025 PreTE: Traffic Engineering with Predictive Failures
abstract
Fiber links in wide-area networks (WANs) are exposed to complicated environments and hence are vulnerable to failures like fiber cuts. The conventional approach of using static probabilistic failures falls short in fiber-cut scenarios because these fiber cuts are rare but disruptive, making it difficult for network operators to balance network utilization and availability in WAN traffic engineering. Our large-scale measurements of per-second optical-layer data reveal that the fiber's failure probability increases by several orders of magnitude when experiencing a rare and ephemeral degradation state. Therefore, we present a novel traffic engineering (TE) system called PreTE to factor in the dynamic fiber cut probabilities directly into TE systems. At the core of the PreTE system, fiber degradation facilitates failure predictions and traffic tunnels to be proactively updated, followed by traffic allocation optimizations among updated tunnels. We evaluate PreTE using a production-level WAN testbed and large-scale simulations. The testbed evaluation quantifies PreTE's runtime to demonstrate the feasibility to implement in large-scale WANs. Our large-scale simulation results show that PreTE can support up to 2× more demand at the same level of availability as compared to existing TE schemes.
Congcong Miao, Zhizhen Zhong, Arpit Gupta, Ying Zhang 0022, Zekun He, Xianneng Zou, Jilong Wang 0001
SIGCOMM8
2024 Turbo: Efficient Communication Framework for Large-scale Data Processing Cluster
abstract
Big data processing clusters are suffering from a long job completion time due to the inefficient utilization of the RDMA capability. Our production measurement results in a large-scale cluster with hundreds of server nodes to process large-scale jobs have shown that the existing deployment of RDMA technique results in a long-tail job completion time, with some jobs even taking up more than twice the average time to complete. In this paper, we present the design and implementation of Turbo, an efficient communication framework for the large-scale data processing cluster to achieve high performance and scalability. The core of Turbo's approach is to leverage a dynamic block-level flowlet transmission mechanism and a non-blocking communication middleware to improve the network throughput and enhance system's scalability. Furthermore, Turbo ensures high system reliability by utilizing an external shuffle service as well as TCP serving as a backup. We integrate Turbo into Apache Spark and evaluate Turbo in a small-scale testbed and a large-scale cluster consisting of hundreds of server nodes. The small-scale testbed evaluation results show that Turbo improves the network throughput by 15.1% while maintaining high system reliability. The large-scale production results have shown Turbo can reduce the job completion time by 23.9% and increase the job completion rate by 2.03× over the existing RDMA solutions.
Xuya Jia, Zhiyi Yao, Edison Liu, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Chongqing Zhao, Jinhui Chu, Jilong Wang 0001, Congcong Miao
SIGCOMM10
2024 MegaTE: Extending WAN Traffic Engineering to Millions of Endpoints in Virtualized Cloud
abstract
In today's virtualized cloud, containers and virtual machines (VMs) are prevailing methods to deploy applications with different tenant requirements. However, these requirements are at odds with the resource allocation capabilities of conventional networking stacks in wide-area networks (WANs). In particular, existing WAN traffic engineering (TE) systems at the granularity of aggregated traffic flows are not designed to cater to each individual flow. In this paper, we advocate for a radical new approach to extend TE systems to involve millions of virtual instance endpoints. We propose and implement a first-of-its-kind system, called MegaTE, to satisfy the needs of each fine-grained traffic flow at the virtual instance level. At the core of the MegaTE system is the paradigm shift from the top-down centralized control to the bottom-up asynchronous query in the TE control loop, combined with eBPF-based segment routing on the data plane and TE optimization contraction on the control plane. We evaluate MegaTE using flow-level simulations with production traffic traces. Our results show that MegaTE supports 20× more endpoints with the similar algorithm run time compared to prior work. MegaTE has been adopted by large-scale public cloud providers. Notably, Tencent rolled out MegaTE in its cloud WAN since December 2022. Our production analysis shows that MegaTE reduces the packet latency of real-time applications by up to 51%.
Congcong Miao, Zhizhen Zhong, Yunming Xiao, Senkuo Zhang, Yinan Jiang, Zizhuo Bai, Chaodong Lu, Jingyi Geng, Zekun He, Yachen Wang, Xianneng Zou, Chuanchuan Yang
SIGCOMM12
2023 FlexWAN: Software Hardware Co-design for Cost-Effective and Resilient Optical Backbones
abstract
The rising demand for WAN capacity driven by the rapid growth of inter-data center traffic poses new challenges for costly optical networks. Today cloud providers rely on fixed optical backbones, where all hardware devices operate on a rigid spectrum grid, leading to the waste of expensive optical resources and subpar performance in handling failures. In this paper, we introduce FlexWAN, a novel flexible WAN infrastructure designed to provision cost-effective WAN capacity while ensuring resilience to optical failures. FlexWAN achieves this by incorporating spacing-variable hardware at the optical layer, enabling the generated wavelength to optimize the utilization of limited spectrum resources for the WAN capacity. The configuration of spacing-variable hardware in a multi-vendor optical backbone presents challenges related to spectrum management. To address this, FlexWAN leverages a centralized controller to achieve coordinated control of network-wide optical devices in a vendor-agnostic manner. Moreover, the flexibility at the optical layer introduces new algorithmic problems. FlexWAN formulates the problem of provisioning WAN capacity with the goal of minimizing hardware costs. We evaluate the system performance in production and share insights from years of production experience. Compared to existing optical backbones, FlexWAN can save at least 57% of transponders and reduce 36% of spectrum usage while continuing to meet up to 8× the present-day demands using existing hardware and fiber deployments. FlexWAN further incorporates failure resilience that revives 15% more bandwidth capacity in the overloaded optical backbone.
Congcong Miao, Zhizhen Zhong, Ying Zhang 0022, Kunling He, Fangchao Li, Minggang Chen, Xiang Li 0223, Zekun He, Xianneng Zou, Jilong Wang 0001
SIGCOMM10