Zixuan Guan

dblp:305/9753 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0008-3780-1203ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 8 since 2021
YearPublicationVenuePosition
2026 Towards Efficient Serving of Network-intensive LLM Inferences
abstract
Prefix caching has become a key technique for LLM serving, and nowadays the reusable KVCache contents are often hosted on distributed servers. For long-context LLM inferences with high cache hit ratio, cross-server KVCache transmission has become an emerging performance bottleneck; such network-intensive LLM inferences are increasingly prevalent in the coming era of agentic AI. However, existing LLM inference engines are essentially compute-centric; we find that they are highly inefficient when serving such workloads due to compute-stage service blocking and ignorance of KVCache-transfer cost.
Chen Chen 0067, Junxue Zhang 0001, Zhusheng Wang, Zixuan Guan, Qizhen Weng 0001, Minyi Guo
APNet6
2026 Delphinus: Ultra-Fast Link Failure Detection and Recovery for AI Data Center Networks
abstract
High-performance artificial intelligence (AI) applications impose stringent reliability requirements on AI data center networks (DCNs), yet link failures are almost inevitable and can severely disrupt AI workloads such as large language model (LLM) training and inference. Existing deployed link failure detection and recovery mechanisms suffer from slow execution speed and limited failure coverage, failing to meet the demands of production AI DCNs. To address these issues, we propose Delphinus, an ultra-fast link failure detection and recovery solution built on the data-plane of programmable switches. It achieves ultra-fast failure detection via hardware-based port state monitoring, extends recoverable failure coverage through remote failure notification and relay, and enables fast recovery by path switchover. Delphinus can serve as a key generic function of switches, providing host-transparent link failure handling for Ethernet fabrics. We implement Delphinus on commercial hardware switches, and deploy it in large-scale production AI DCNs for over a year. Extensive evaluations demonstrate that Delphinus can complete link failure detection and recovery within sub-milliseconds, with negligible impact on application performance and imperceptible service interruption.
Junye Zhang, Zhigang Ji, Kefei Liu, Rui Zhuang, Ruixue Wang, Weiqiang Cheng, Zixuan Guan
SIGCOMM9
2026 Dragonfly-Ultra: A Scalable, Low-Cost Network Architecture for High-Performance AI Clusters
abstract
Large-scale AI clusters impose higher requirements on network scalability, cost, and communication efficiency. The traditional Clos topology suffers from superlinear cost growth when scaling to over 100k GPUs, while the more cost-effective Dragonfly+ introduces "down-up" detours, deadlock risks, and complex routing design. This paper presents Dragonfly-Ultra, a scalable, low-cost network architecture for high-performance AI clusters. Dragonfly-Ultra can scale to over 260k GPUs with only 82% cost and 81% power consumption of a 3-layer Clos architecture. Dragonfly-Ultra optimizes inter-group connectivity to eliminate intra-group detours entirely. Beyond the topological benefits, Dragonfly-Ultra incorporates three key mechanisms to further improve network performance and optimize collective communication, including lightweight dual-waterline adaptive routing for fast congestion mitigation, virtual-link-based deadlock avoidance with lower hardware overhead, and uniform affinity-aware rank placement for balanced inter-group traffic across all phases. Simulation results on a 4k-node cluster show that, compared to Clos, Dragonfly-Ultra achieves up to 18.8% and 39.2% lower completion time for AllReduce and AlltoAll, respectively. Compared to Dragonfly+, the reductions are up to 27.9% and 62.1%, outperforming current mainstream topologies.
Rui Zhuang, Junye Zhang, Kefei Liu, Weiqiang Cheng, Zixuan Guan, Shengnan Yue, Ruixue Wang, Tong Yang 0003
SIGCOMM10
2026 LLT: Lossless Transmission Using Local Recirculation for WANs
abstract
As distributed applications increasingly span geographically distributed data centers, the demand for high-performance, long-distance transmission has been continuously growing. While intra-data-center networks have employed techniques like remote direct memory access (RDMA) to meet these design goals, extending these techniques toWANs presents unique challenges. WANs notably suffer from inherent packet losses due to buffer overflows in routers and switches, leading to decreased throughput and making distributed applications barely usable. This paper proposes Lossless Transmission (LLT), a novel buffer management scheme for enabling lossless WAN transport. LLT intelligently integrates on-chip switch buffers with an off-chip caching system to absorb traffic bursts that would otherwise cause packet loss. Its data plane logic uses a multi-level threshold system to selectively offload only critical flows during congestion. A closed-loop control protocol, managed by a stateful flow table, ensures these offloaded packets are later re-injected with guaranteed lossless and in-order delivery, effectively protecting latency-sensitive applications from retransmission overhead. We evaluate LLT using both ns-3 simulations and P4-programmable devices. The experimental results show that in typical use cases (RTT > 30ms), LLT improves link bandwidth utilization by 1.9% to 29.5% and reduces the P99 percentile tail latency by 17% to 66% in WANs compared to the state-of-the-art solutions. Overall, LLT provides a scalable, efficient, and reliable framework for long-distance data transmission, addressing critical challenges in WANs. Additionally, LLT eliminates the need for expensive WAN infrastructure modifications.
Junchang Wang, Xin He 0010, Weibei Fan, Zixuan Guan, Xiaolong Zheng 0002, Fu Xiao 0001
IEEE Trans. Netw. Serv. Manag.7
2026 HierCC: Taming Traffic Uncertainty in RDMA Data Centers With Hierarchical Congestion Control
abstract
Existing congestion control schemes for RDMA resolve the dilemma of guaranteeing high throughput and ultra-low latency to some extent from a variety of perspectives. However, they are inefficient in addressing transient large queue build-up and under-utilized bandwidth caused by frequent traffic bursts. In this paper, we argue that traffic uncertainty is the fundamental challenge that limits these schemes from addressing the aforementioned dilemma. Inspired by the investigation that aggregated flows within the same rack are relatively long-lived, we propose HierCC, which aggregates flows destined to the same IP in a rack to ease traffic uncertainty and further provides hierarchically control within the first-hop ToR and between racks. Specifically, the inter-rack rates of aggregate flows are controlled by a credit-based mechanism. Then the bandwidth obtained by the aggregated flow is allocated to the corresponding intra-rack individual flows promptly and accurately. We implement HierCC in a testbed that consists of DPDK-based end-hosts and P4-based Tofino switches. The performance of HierCC is evaluated by comprehensive testbed experiments and SystemC/NS3 simulations. Results indicate that, compared with state-of-the-art, HierCC can mitigate buffer usage by up to$10\times $and reduce the average and 99th percentile FCT by up to 84% and 80%, respectively.
Zirui Wan, Jiao Zhang 0002, Xiaolong Zhong, Zixuan Guan, Haoyu Pan, Tian Pan 0001, Tao Huang 0005
IEEE Trans. Netw.5
2024 PACC: A Proactive CNP Generation Scheme for Datacenter Networks
abstract
The rapid upgrade of link speed and the prosperity of new applications in data center networks (DCNs) lead to a rigorous demand for ultra-low latency and high throughput. To mitigate the overhead of traditional software-based packet processing at end-hosts, RDMA (Remote Direct Memory Access) has been widely adopted in DCNs. Particularly, congestion control (CC) mechanisms designed for RDMA have attracted much attention to avoid performance deterioration when packets lose. However, through comprehensive analysis, we found that existing RDMA CC schemes have limitations of a sluggish response to congestion and unawareness of tiny microbursts due to the long end-to-end control loop. In this paper, we propose PACC, a proactive and accurate switch-driven RDMA CC algorithm with easy deployability. PACC is driven by PI controller-based computation, threshold-based flow discrimination and weight-based allocation at the switch. It leverages real-time queue length to generate accurate congestion feedback proactively and piggybacks it to the corresponding source without modification to end-hosts. We theoretically analyze the stability, convergence and key parameter settings of PACC. Then, we implement PACC in a testbed consisting of DPDK-based end-hosts and Tofino P4 switches. In our evaluation, PACC achieves better fairness, fast reaction, high throughput, and 6$\sim$69% lower FCT (Flow Completion Time) than DCQCN, TIMELY, HPCC and RoCC.
Jiao Zhang 0002, Xiaolong Zhong, Mingxuan Yu, Haoyu Pan, Zixuan Guan, Biyao Che, Zirui Wan, Tian Pan 0001, Tao Huang 0005
IEEE/ACM Trans. Netw.7
2022 PACC: Proactive and Accurate Congestion Feedback for RDMA Congestion Control
abstract
The rapid upgrade of link speed and the prosperity of new applications in data center networks (DCNs) lead to a rigorous demand for ultra-low latency and high throughput. To mitigate the overhead of traditional software-based packet processing at end-hosts, RDMA (Remote Direct Memory Access) has been widely adopted in DCNs. Particularly, congestion control (CC) mechanisms designed for RDMA have attracted much attention to avoid performance deterioration when packets lose. However, through comprehensive analysis, we found that existing RDMA CC schemes have limitations of a sluggish response to congestion and unawareness of tiny microbursts due to the long end-to-end control loop. In this paper, we propose PACC, a switch-driven RDMA CC algorithm with easy deployability. PACC is driven by PI controller-based computation, threshold-based flow discrimination and weight-based allocation at the switch. It leverages real-time queue length to generate accurate congestion feedback proactively and piggybacks it to the corresponding source without modification to end-hosts. We theoretically analyze the stability and key parameter settings of PACC. Then, we conduct both micro-benchmark and large-scale simulations to evaluate the performance of PACC. The results show that PACC achieves fairness, fast reaction, high throughput, and 6~69% lower FCT (Flow Completion Time) than DCQCN, TIMELY and HPCC.
Xiaolong Zhong, Jiao Zhang 0002, Zixuan Guan, Zirui Wan
INFOCOM4
2021 HierCC: Hierarchical RDMA Congestion Control
abstract
RDMA has been increasingly deployed in data centers to decrease latency and CPU utilization. However, existing RDMA congestion control schemes fail to address instantaneous large queue build-up or bandwidth under-utilization associated with frequent traffic bursty. In this paper, we argue that traffic uncertainty is the essential reason that constrains data center congestion control from simultaneously achieving high throughput and deterministic latency. Since aggregated flows within the same rack are relatively long-lived, we propose HierCC, which aggregates flows destined to the same IP in a rack and hierarchically controls the rate of flows. The rate of aggregate flows between racks is controlled by a credit-based congestion control mechanism. Then the bandwidth obtained by an aggregate flow in a rack is allocated to the corresponding individual flows from that rack promptly and accurately. We evaluate HierCC using SystemC and large-scale NS3 simulations. Results indicate that HierCC can significantly mitigate buffer usage and reduce the 99th percentile FCT by up to 20% and 40% compared with HPCC and DCQCN under a realistic workload, respectively.
Jiao Zhang 0002, Zixuan Guan, Zirui Wan, Yinben Xia, Tian Pan 0001, Tao Huang 0005, Dezhi Tang
APNet3