Zhaochen Zhang

dblp:156/2393 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 3 first-author · 8 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 OSCAR: O(1)-Step Convergence and Readily-deployable Congestion Control
Zhaochen Zhang, Feiyang Xue, Rui Ning, Keqiang He, Gianni Antichi, Zhimeng Yin 0001, Rui Li 0020, Zhengqi Cui, Zhehao Lin, Peirui Cao, Guihai Chen, Chen Tian 0001
NSDI1
2026 Anytest: Localizing the Root Cause of Hardware Transport Performance Anomalies
Zhaochen Zhang, Sheng Cheng 0002, Feiyang Xue, Chang Liu 0001, Boliang Liu, Rui Li 0020, Li Wang 0110, Peirui Cao, Qingkai Meng 0001, Guihai Chen, Shuguang Cheng, Yongqing Xi, Binzhang Fu, Dennis Cai, Chen Tian 0001
SIGCOMM1
2026 Revisiting Flow Control in Node-Centric Datacenter Networks
abstract
Node-centric Data Centers (NDCs) are highly flexible, cost-efficient, and failure-resilient, and have gained growing popularity in recent years. However, RDMA technology used in NDC still faces challenges, including high retransmission overhead, Head-of-Line Blocking (HoLB) and deadlock problems. Existing solutions for traditional data centers cannot simultaneously address these issues due to the unique topology and server transmission characteristics of NDC. In this paper, we propose a per-port flow control named PortFC for NDC. PortFC addresses the above problems through the designs of a Pause/Resume control signal, a per-port queue allocation method, an egress-detecting per-port flow control mechanism, and a server-aware queue scheduling method. Our evaluation shows that PortFC is free from retransmission, capable of eliminating HoLB and avoiding deadlocks. PortFC achieves 1.7-8.0 times higher throughput and reduces latency by 11.7%-87.7% compared to the state-of-the-art lossy RDMA based on IRN and the lossless RDMA method based on PFC. In particular, PortFC still demonstrates good performance in a Rail-only NDC with heterogeneous bandwidth domains.
Peirui Cao, Rui Ning, Guangyu Zhao, Zhaochen Zhang, Chang Liu 0001, Yunzhuo Liu, Rui Li 0020, Chengyuan Huang, Tao Sun 0010, Guihai Chen, Baochun Li, Chen Tian 0001
IEEE Trans. Netw.4
2026 UDMP: Unified Delay-Driven Multipath Protocol for AI Clusters
abstract
Distributed AI model training generates bursty, low-entropy elephant flows that challenge existing single-path transport protocols in multi-stage Clos networks, leading to congestion and inefficiency. Multipath transport emerges as a promising solution, leveraging multiple paths to balance traffic and enhance resilience. However, current multipath RDMA solutions suffer from scalability, congestion control, and load-balancing inefficiencies. This paper introduces Unified Delay-driven Multipath Protocol (UDMP), a novel approach that co-designs congestion control and load balancing using network delay as a unified signal. UDMP employs delay-gradient-based congestion control to precisely resolve unavoidable congestion. Moreover, UDMP leverages delay-assisted load balancing to shift traffic across paths with minimal latency adaptively, maintaining throughput when encountering avoidable congestion. A novel Token Pool design integrates these components, eliminating per-path state overhead while achieving fine-grained traffic distribution. Implementations on DPDK and NS3 demonstrate that UDMP achieves up to 2x higher throughput and reduces flow completion times by up to 30% compared to state-of-the-art methods like MPRDMA and QP-Scaling. These results highlight UDMP’s effectiveness in meeting the stringent performance requirements of modern distributed AI training workloads.
Chengyuan Huang, Zhengqi Cui, Jun Xu 0037, Zhaochen Zhang, Li Wang 0110, Peirui Cao, Zhongming Ji, Jilei Chen, Shengju Zhang, Lingkun Meng, Ahmed M. Abdelmoniem, Fu Xiao 0001, Wan-Chun Dou, Guihai Chen, Keqiang He, Chen Tian 0001
IEEE Trans. Netw.4
2026 Rail: ReArranging Inter-GPU Links for GPU-Centric Clusters
abstract
In modern GPU-centric clusters, large-scale AI training relies on two distinct communication domains: a high-bandwidth intra-node domain using proprietary interconnects (e.g., NVLink), and a scale-out inter-node network domain (e.g., RDMA). We observe that the widely-used ring algorithm, often create a significant load imbalance across these domains. This leads to the counter-intuitive scenario where the expensive, high-bandwidth intra-node domain becomes a performance bottleneck, while the inter-node network remains underutilized. This inefficiency is further exacerbated by the disparity in bandwidth provisioning: inter-node network bandwidth is generally more cost-effective and accessible, whereas intra-node bandwidth is often proprietary and more costly to scale. To address this fundamental imbalance, we propose RAIL, aimed at resolving the intra-node bottleneck by strategically rearranging inter-GPU communication paths. This rebalancing ensures that traffic loads are appropriately matched with the distinct transmission capabilities of each domain, thereby maximizing overall communication performance. RAIL incorporates a Load Distributing Strategy (LDS) that can accurately partition physical nodes into logical nodes based on the a transmission capabilities of both domains, shifting excess traffic from the overloaded intra-node domain to the underutilized network domain. Additionally, the Intra-Rail Strategy (IRS) leverages topological characteristics to ensure optimal communication paths through the network domain between logical nodes. Our evaluation demonstrates that RAIL effectively mitigates congestion and achieves a 30.7% average increase in collective communication bus bandwidth compared to the widely-used NCCL solution.
Haixin Nan, Jun Xu 0037, Peirui Cao, Zhaochen Zhang, Yizhi Wang 0004, Zhehao Lin, Yuhang Li 0002, Chengyuan Huang, Xiaohu Xu, Zhongming Ji, Shengju Zhang, Lingkun Meng, Rong Gu 0001, Guihai Chen, Chen Tian 0001
IEEE Trans. Netw.4
2026 Analysis of Pyrrha: Congestion-Root-Based Flow Control Is Most Cost-Effective to Eliminate Head-of-Line Blocking
abstract
In modern datacenters, the effectiveness of end-to-end congestion control (CC) is quickly diminishing with the rapid bandwidth evolution. Per-hop flow control (FC) can react to congestion more promptly. However, a coarse-grained FC can result in Head-Of-Line (HOL) blocking. A fine-grained, per-flow FC can eliminate HOL blocking caused by flow control, however, it does not scale well. This paper presents Pyrrha, a scalable flow control approach that provably eliminates HOL blocking while using a minimum number of queues. In Pyrrha, flow control first takes effect on the root of the congestion, i.e., the port where congestion occurs. And then flows are controlled according to their contributed congestion roots. A prototype of Pyrrha is implemented on Tofino2 switches. Compared with state-of-the-art approaches, the average FCT of uncongested flows is reduced by 42%-98%, and 99th-tail latency can be$1.6\times $-$215\times $lower, without compromising the performance of congested flows.
Zhaochen Zhang, Peirui Cao, Chang Liu 0001, Yizhi Wang 0004, Vamsi Addanki, Stefan Schmid 0001, Qingyue Wang, Xiaoliang Wang 0001, Jiaqi Zheng 0001, Tao Wu 0011, Bingyang Liu, Wan-Chun Dou, Guihai Chen, Chen Tian 0001, Fu Xiao 0001
IEEE Trans. Netw.1
2025 Enabling Virtual Priority in Data Center Congestion Control
abstract
In data center networks, various types of traffic with strict performance requirements operate simultaneously, necessitating effective isolation and scheduling through priority queues. However, most switches support only around ten priority queues. Virtual priority can address this limitation by emulating multi-priority queues on a single physical queue, but existing solutions often require complex switch-level scheduling and hardware changes. Our key insight is that virtual priority can be achieved by carefully managing bandwidth contention in a physical queue, which is traditionally handled by congestion control (CC) algorithms. Hence, the virtual priority mechanism needs to be tightly coupled with CC. In this paper, we propose PrioPlus, a CC enhancement algorithm that can be integrated with existing congestion control schemes to enable virtual priority transmission. PrioPlus assigns specific delay ranges to different priority levels, ensuring that flows transmit only when the delay is within the assigned range, effectively meeting virtual priority requirements. Compared to Swift CC with physical priority queues, PrioPlus provides strict priority for high-priority flows without impacting performance sensibly. Meanwhile, it benefits low-priority flows from 25% to 41% as its priority-aware design enhances CC's ability to fully utilize available bandwidth once higher-priority traffic completes. As a result, in coflow and model training scenarios, PrioPlus improves job completion times by 21% and 33%, respectively, compared to Swift with physical priority queues.
Zhaochen Zhang, Feiyang Xue, Keqiang He, Zhimeng Yin 0001, Gianni Antichi, Yizhi Wang 0004, Rui Ning, Haixin Nan, Xu Zhang 0006, Peirui Cao, Xiaoliang Wang 0001, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
EuroSys1
2025 PortFC: Designing High-performance Deadlock-free BCube Networks
abstract
BCube is a modular data center network.Compared with other topologies, BCube has natural advantages, such as lower deployment costs and stronger failure recovery capabilities.However, RDMA technology used in BCube still faces challenges, including high retransmission overhead, Head-of-Line Blocking (HoLB) and deadlock problems.Existing solutions for traditional data centers cannot simultaneously address these issues due to the unique topology and server transmission characteristics of BCube.In this paper, we propose a per-port flow control named PortFC for BCube.PortFC addresses the above problems through the designs of a Pause/Resume control signal, a per-port queue allocation method, an egress-detecting per-port flow control mechanism, and a serveraware queue scheduling method.Our evaluation shows that PortFC is free from retransmission, capable of eliminating HoLB and avoiding deadlocks.PortFC achieves 1.7-8.0times higher throughput and reduces latency by 11.7%-87.7%compared to the state-of-the-art
Peirui Cao, Rui Ning, Zhaochen Zhang, Chang Liu 0001, Rui Li 0020, Yongqi Yang, Yunzhuo Liu, Chengyuan Huang, Tao Sun 0010, Xiaodong Duan, Guihai Chen, Chen Tian 0001
ICS4
2025 Pyrrha: Congestion-Root-Based Flow Control to Eliminate Head-of-Line Blocking in Datacenter
Zhaochen Zhang, Chang Liu 0001, Yizhi Wang 0004, Vamsi Addanki, Stefan Schmid 0001, Qingyue Wang, Xiaoliang Wang 0001, Jiaqi Zheng 0001, Tao Wu 0011, Bingyang Liu, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
NSDI2
2025 Courier: A Unified Communication Agent to Support Concurrent Flow Scheduling in Cluster Computing
abstract
As one of the pillars in cluster computing frameworks, coflow scheduling algorithms can effectively shorten the network transmission time of cluster computing jobs, thus reducing the job completion times and improving the execution performance. However, most of existing coflow scheduling algorithms failed to consider the influences of concurrent flows, which can degrade their performance under a massive number of concurrent flows. To fill the gap, we propose a unified communication agent named Courier to minimize the number of concurrent flows in cluster computing applications, which is compatible with the mainstream coflow scheduling approaches. To maintain the scheduling order given by the scheduling algorithms, Courier merges multiple flows between each pair of hosts into a unified flow, and determines its order based on that of origin flows. In addition, in order to adapt to various types of topologies, Courier introduces a control mechanism to adjust the number of flows while maintaining the scheduling order. Extensive large-scale trace-driven simulations have shown that Courier is compatible with existing scheduling algorithms, and outperforms the state-of-the-art approaches by about 30% under a variety of workloads and topologies.
Zhaochen Zhang, Xu Zhang 0006, Zhaoxiang Bao, Chaohong Tan, Wan-Chun Dou, Guihai Chen, Chen Tian 0001
IEEE Trans. Parallel Distributed Syst.1
2022 FlyMon: enabling on-the-fly task reconfiguration for network measurement
abstract
Network measurement is important to data center operators. Most existing efforts focus on developing new implementation schemes for measurement tasks. Little attention is paid to on-the-fly task reconfiguration. Due to resource constraints, it is impossible to configure all needed tasks at start-up and dynamically turn on/of them. To support real-time reconfiguration of many different tasks, a key observation is that it is unnecessary to bind a task and its implementation at the compilation phase. We design FlyMon, the first sketch-based measurement system that can make on-the-fly reconfigurations on a large set of measurement tasks. FlyMon introduces the concept of Composable Measurement Units (CMUs), which are general operation units that support reconfigurable implementation for measurement tasks combined from different flow keys and flow attributes. FlyMon maps the design of CMUs to programmable switches' data planes so that the number of compacted CMUs can be maximized. FlyMon also provides dynamic memory management. We prototype FlyMon on Tofino and currently enable four frequently used flow attributes. Each CMU Group (with 3 CMUs) can concurrently perform up to 96 isolated measurement tasks with less than 8.3% hardware resources. The tasks can be deployed with configurable memory size at the millisecond level. By cross-stacking, FlyMon can deploy up to 27 CMUs in one pipeline of Tofino.
Chen Tian 0001, Tong Yang 0003, Chang Liu 0001, Zhaochen Zhang, Wan-Chun Dou, Guihai Chen
SIGCOMM6
2022 SharpNet: A deep learning method for normal vector estimation of point cloud with sharp features
Zhaochen Zhang, Jianhui Nie, Mengjuan Yu
Graph. Model.1
2021 Enhancement of ridge-valley features in point cloud based on position and normal guidance
Jianhui Nie, Zhaochen Zhang, Ye Liu 0005, Hao Gao 0005, Feng Xu 0005, Wenkai Shi
Comput. Graph.2
2021 Bas-relief generation from point clouds based on normal space compression with real-time adjustment on CPU
Jianhui Nie, Wenkai Shi, Ye Liu 0005, Hao Gao 0005, Feng Xu 0005, Zhaochen Zhang
Graph. Model.6