VLDB 2026 Research / reviewers in the wild / expert
Weiqiang Cheng
dblp:284/3901
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0003-1193-1513ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating Link-level Lossless Mechanisms in AI NetworksabstractThe ever-increasing demand for network performance in large language models promotes the advent of many Scale-up networking schemes. To consistently deliver superior low latency and high bandwidth, these schemes have widely adopted Credit-based Flow Control (CBFC) and Link Layer Retry (LLR) to ensure link-level lossless transmission. However, there is currently a lack of evaluations for these mechanisms in Scale-up domains. This paper builds an FPGA-based Scale-up network to evaluate these lossless mechanisms. Evaluation results show that, while these mechanisms consume a small amount of link bandwidth, CBFC can greatly reduce receive buffer utilization, and LLR can substantially mitigate network performance degradation caused by packet corruption. Kefei Liu, Ruixue Wang, Tianrun Jiang, Runlong Hu, Weiqiang Cheng, Danyuan Zhou, Junye Zhang, Xinhong Deng, Rundi Zhai, Cong Qi, Yijian Qi |
SIGCOMM | 6 |
| 2026 | Delphinus: Ultra-Fast Link Failure Detection and Recovery for AI Data Center NetworksabstractHigh-performance artificial intelligence (AI) applications impose stringent reliability requirements on AI data center networks (DCNs), yet link failures are almost inevitable and can severely disrupt AI workloads such as large language model (LLM) training and inference. Existing deployed link failure detection and recovery mechanisms suffer from slow execution speed and limited failure coverage, failing to meet the demands of production AI DCNs. To address these issues, we propose Delphinus, an ultra-fast link failure detection and recovery solution built on the data-plane of programmable switches. It achieves ultra-fast failure detection via hardware-based port state monitoring, extends recoverable failure coverage through remote failure notification and relay, and enables fast recovery by path switchover. Delphinus can serve as a key generic function of switches, providing host-transparent link failure handling for Ethernet fabrics. We implement Delphinus on commercial hardware switches, and deploy it in large-scale production AI DCNs for over a year. Extensive evaluations demonstrate that Delphinus can complete link failure detection and recovery within sub-milliseconds, with negligible impact on application performance and imperceptible service interruption. Junye Zhang, Zhigang Ji, Kefei Liu, Rui Zhuang, Ruixue Wang, Weiqiang Cheng, Zixuan Guan |
SIGCOMM | 8 |
| 2026 | Dragonfly-Ultra: A Scalable, Low-Cost Network Architecture for High-Performance AI ClustersabstractLarge-scale AI clusters impose higher requirements on network scalability, cost, and communication efficiency. The traditional Clos topology suffers from superlinear cost growth when scaling to over 100k GPUs, while the more cost-effective Dragonfly+ introduces "down-up" detours, deadlock risks, and complex routing design. This paper presents Dragonfly-Ultra, a scalable, low-cost network architecture for high-performance AI clusters. Dragonfly-Ultra can scale to over 260k GPUs with only 82% cost and 81% power consumption of a 3-layer Clos architecture. Dragonfly-Ultra optimizes inter-group connectivity to eliminate intra-group detours entirely. Beyond the topological benefits, Dragonfly-Ultra incorporates three key mechanisms to further improve network performance and optimize collective communication, including lightweight dual-waterline adaptive routing for fast congestion mitigation, virtual-link-based deadlock avoidance with lower hardware overhead, and uniform affinity-aware rank placement for balanced inter-group traffic across all phases. Simulation results on a 4k-node cluster show that, compared to Clos, Dragonfly-Ultra achieves up to 18.8% and 39.2% lower completion time for AllReduce and AlltoAll, respectively. Compared to Dragonfly+, the reductions are up to 27.9% and 62.1%, outperforming current mainstream topologies. Rui Zhuang, Junye Zhang, Kefei Liu, Weiqiang Cheng, Zixuan Guan, Shengnan Yue, Ruixue Wang, Tong Yang 0003 |
SIGCOMM | 6 |
| 2026 | RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts TrainingabstractTraining Mixture-of-Experts (MoE) models introduces sparse and highly imbalanced all-to-all communication that dominates iteration time. Conventional load-balancing methods fail to exploit the deterministic topology of Rail architectures, leaving multi-NIC bandwidth underutilized. We present RailS, a distributed load-balancing framework that minimizes all-to-all completion time in MoE training. RailS leverages the Rail topology’s symmetry to prove that uniform sending ensures uniform receiving, transforming global coordination into local scheduling. Each node independently executes a Longest Processing Time First (LPT) spraying scheduler to proactively balance traffic using local information. RailS activates N parallel rails for fine-grained, topology-aware multipath transmission. Across synthetic and real-world MoE workloads, RailS improves bus bandwidth by 20%–78% and reduces completion time by 17%–78%. For Mixtral workloads, it shortens iteration time by 18%–40% and achieves near-optimal load balance, fully exploiting architectural parallelism in distributed training. Chengze Du 0001, Ying Zhou 0017, Weiqiang Cheng, Jialong Li 0006 |
IEEE Trans. Netw. | 7 |
| 2025 | Generalized SRv6 Header Compression Packet Processors based on Multicore Architectures
Weiqiang Cheng, Xiaodong Duan, Han Li 0008, Jialong Li 0006 |
APNet | 1 |