VLDB 2026 Research / reviewers in the wild / expert
Jianxi Ye
dblp:184/0340
· DBLP profile ↗
13ranked-venue papers
0as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 8 since 2021Systems, architecture and hardware · 3 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Handling Network Faults in Distributed AI Training: Failover is Now an OptionabstractDistributed AI training often suffers from network faults. Network faults, especially at the last hop between a switch and a host, result in loss of connectivity, resulting in training job stalls and eventual failure. This is typically managed through a fail-stop mechanism, followed by a restart, incurring significant inefficiencies. We present ReCCL, the first network fault-tolerant collective communication library (CCL) that allows training progress to be preserved by seamlessly failing over to alternate paths when a network fault occurs. During failover, ReCCL keeps communication states synchronized while using dynamic channel load balancing and intra-host GPU routing to improve communication performance. Our evaluations demonstrate that ReCCL can perform failover seamlessly with minimal performance losses. Additionally, our simulations also demonstrate that failover can be effectively used to achieve significant savings in GPU hours for large-scale distributed AI training workloads. Xin Zhe Khooi, Zhuo Jiang, Pan Xie, Zhigang Cui, Meng Wang 0018, Yuze Jin, Pengfei Huo, Lulu Chen, Liaoyuan Feng, Qinlong Wang, Yongcan Wang, Jinshuai Sun, Yingkai Zhao, Haiquan Chen 0002, Yi Li 0098, Jianxi Ye, Mun Choon Chan |
EuroSys | 24 |
| 2026 | BURST: Seeking High-performance, Interoperability and Scalability in Soft-RDMA
Huijun Shen, Zelong Yue, Zhuo Jiang, Lang An, Luochangqi Ding, Xiaolong Zhong, Jianxi Ye, Xijin Yin, Xingyu Guo |
NSDI | 13 |
| 2026 | Odin: Rethinking Congestion Control under All-to-All Traffic
Wanchun Jiang, Jianxi Ye, Huichen Dai |
SIGCOMM | 7 |
| 2025 | UCM: Fast and Maintainable User-space RDMA Connection Setup
Huijun Shen, Zelong Yue, Xingyu Guo, Xijin Yin, Lang An, Jianxi Ye, Guo Chen 0001 |
APNet | 11 |
| 2025 | Minder: Faulty Machine Detection for Large-scale Distributed Model Training
Yangtao Deng, Zhuo Jiang, Xingjian Zhang 0009, Zhang Zhang 0003, Zuquan Song, Gaohong Liu, Fuliang Li, Shuguang Wang, Haibin Lin, Jianxi Ye, Minlan Yu |
NSDI | 14 |
| 2025 | From ATOP to ZCube: Automated Topology Optimization Pipeline and A Highly Cost-Effective Network Topology for Large Model TrainingabstractThe development of large language models (LLMs) poses new challenges in data center network topology design. To assist in exploring topology design, we propose ATOP, an Automated Topology Optimization Pipeline, which models network topology as a set of hyperparameters, enabling the discovery of potential topologies. With various optimization algorithms and customizable optimization objectives, ATOP achieves automated topology optimization on a scale of tens of thousands of GPUs. We apply ATOP on network topologies for 256, 1024, 4096, and 16384 GPUs, optimizing performance under LLMs training traffic patterns, collective communication performance, fault tolerance, and network cost. We also evaluate ATOP in different scenarios: building, optimizing, and expanding a data center. From ATOP's results, we discover a new topology — ZCube, which reaches the highest cost-effectiveness across various GPU scales. Simulation results show that ZCube, compared to the previous state-of-the-art topologies, including Rail-optimized Fat-tree (ROFT), Rail-only, and HPN, improves end-to-end LLM training speed by 3% to 7% and reduces network hardware costs by 26% to 46%. We also construct ZCube on a real-world testbed. Results show that ZCube reduces hardware costs by 25% compared to Rail-Optimized Topology while maintaining the same all-reduce and all-to-all performance. Dan Li 0001, Li Chen 0008, Dian Xiong, Kaihui Gao, Yiwei Zhang 0016, Menglei Zhang, Bochun Zhang, Zhuo Jiang, Jianxi Ye, Haibin Lin |
SIGCOMM | 11 |
| 2025 | MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert ParallelismabstractMixture-of-Experts (MoE) showcases tremendous potential to scale large language models (LLMs) with enhanced performance and reduced computational complexity. However, its sparsely activated architecture shifts feed-forward networks (FFNs) from being compute-intensive to memory-intensive during inference, leading to substantially lower GPU utilization and increased operational costs. Ruidong Zhu, Ziheng Jiang, Chao Jin 0007, Cesar A. Stuardo, Huaping Zhou, Jianzhe Xiao, Lingjun Liu, Haibin Lin, Li-Wen Chang, Jianxi Ye, Xuanzhe Liu, Xin Jin 0008, Xin Liu 0086 |
SIGCOMM | 16 |
| 2025 | Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM TrainingabstractReliability is essential for ensuring efficiency in LLM training. However, many real-world reliability issues remain difficult to resolve, resulting in wasted resources and degraded model performance. Unfortunately, today's collective communication libraries operate as black boxes, hiding critical information needed for effective root cause analysis. Yangtao Deng, Qinlong Wang, Xiaoyun Zhi, Zhuo Jiang, Haohan Xu, Zuquan Song, Gaohong Liu, Shuguang Wang, Wencong Xiao, Jianxi Ye, Minlan Yu, Hong Xu 0001 |
SOSP | 14 |
| 2025 | Barre: Empowering Simplified and Versatile Programmable Congestion Control in High-Speed AI Clusters
Yajuan Peng, Xiaolong Zhong, Haohan Xu, Zhuo Jiang, Jianxi Ye, Xiaoliang Wang 0001, Xiaoming Fu 0001, Huichen Dai |
USENIX ATC | 9 |
| 2024 | MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUs
Ziheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang 0001, Yangrui Chen, Zhi Zhang 0005, Yanghua Peng, Xiang Li 0067, Shibiao Nong, Yulu Jia, Sun He, Hongmin Chen, Zhihao Bai, Qi Hou, Shipeng Yan, Yiyao Sheng, Zhuo Jiang, Haohan Xu, Zhang Zhang 0003, Pengfei Nie, Leqi Zou, Sida Zhao, Zherui Liu, Xiaoying Jia 0001, Jianxi Ye, Xin Jin 0008, Xin Liu 0086 |
NSDI | 30 |
| 2022 | Collie: Finding Performance Anomalies in RDMA Subsystems
Xinhao Kong, Yibo Zhu 0001, Huaping Zhou, Zhuo Jiang, Jianxi Ye, Chuanxiong Guo, Danyang Zhuo |
NSDI | 5 |
| 2020 | EFLOPS: Algorithm and System Co-Design for a High Performance Distributed Training PlatformabstractDeep neural networks (DNNs) have gained tremendous attractions as compelling solutions for applications such as image classification, object detection, speech recognition, and so forth. Its great success comes with excessive trainings to make sure the model accuracy is good enough for those applications. Nowadays, it becomes challenging to train a DNN model because of 1) the model size and data size keep increasing, which usually needs more iterations to train; 2) DNN algorithms evolve rapidly, which requires the training phase to be short for a quick deployment. To address those challenges, distributed training platforms have been proposed to leverage massive server nodes for training, with the hope of significant training time reduction. Therefore, scalability is a critical performance metric to evaluate a distributed training platform. Nevertheless, our analysis reveals that traditional server clusters have poor scalability for training due to the traffic congestions within the server and beyond. The intra-server traffic on the I/O fabric can result in severe congestions and skewed quality of service as high performance devices are competing with each other. Moreover, the traffic congestions on the Ethernet for inter-server communication could also incur significant performance degradation. In this work, we devise a novel distributed training platform, EFLOPS, that adopts an algorithm and system co-design methodology to achieve good scalability. A new server architecture is proposed to alleviate the intra-server congestions. Moreover, a new network topology, BiGraph, is proposed to divide the network into two separate parts, so that there is always a direct connection between any nodes from different parts. Finally, accompany with BiGraph, a topology-aware allreduce algorithm is proposed to eliminate the traffic congestion on the direct connection. The experimental results show that eliminating the congestions on network interface can gain up to 11.3xcommunication speedup. The proposed algorithm and topology can provide further improvement up to 6.08x. The overall performance of ResNet-50 training achieves near-linear scalability, and is competitive to the top-rankings of MLPerf results. Jianbo Dong, Zheng Cao 0003, Jianxi Ye, Shaochuang Wang, Liuyihan Song, Liwei Peng, Yiqun Guo, Xiaowei Jiang, Lingbo Tang, Yin Du, Yingya Zhang |
HPCA | 4 |
| 2016 | RDMA over Commodity Ethernet at ScaleabstractOver the past one and half years, we have been using RDMA over commodity Ethernet (RoCEv2) to support some of Microsoft's highly-reliable, latency-sensitive services. This paper describes the challenges we encountered during the process and the solutions we devised to address them. In order to scale RoCEv2 beyond VLAN, we have designed a DSCP-based priority flow-control (PFC) mechanism to ensure large-scale deployment. We have addressed the safety challenges brought by PFC-induced deadlock (yes, it happened!), RDMA transport livelock, and the NIC PFC pause frame storm problem. We have also built the monitoring and management systems to make sure RDMA works as expected. Our experiences show that the safety and scalability issues of running RoCEv2 at scale can all be addressed, and RDMA can replace TCP for intra data center communications and achieve low latency, low CPU overhead, and high throughput. Chuanxiong Guo, Zhong Deng, Gaurav Soni, Jianxi Ye, Jitendra Padhye, Marina Lipshteyn |
SIGCOMM | 5 |