Hairui Zhao 0002

dblp:313/3426-2 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
17since 2021 · last 2026
0009-0008-7081-5172ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 1 first-author · 12 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training
abstract
Handling communication overhead in large-scale tensor-parallel training remains a critical challenge due to the dense, near-zero distributions of intermediate tensors, which exacerbate errors under frequent communication and introduce significant computational overhead during compression. To this end, we propose TACO (Tensor-parallel Adaptive COmmunication compression), a robust FP8-based framework for compressing TP intermediate tensors. First, we employ a data-driven reshaping strategy combined with an Adaptive Scale–Hadamard Transform to enable high-fidelity FP8 quantization, while its Dual-Scale Quantization mechanism ensures numerical stability throughout training. Second, we design a highly fused compression operator to reduce memory traffic and kernel launch overhead, allowing efficient overlap with communication. Finally, we integrate TACO with existing state-of-the-art methods for Data and Pipeline Parallelism to develop a compression-enabled 3D-parallel training framework. Detailed experiments on GPT models and Qwen model demonstrate up to 1.87 × end-to-end throughput improvement while maintaining near-lossless accuracy, validating the effectiveness and efficiency of TACO in large-scale training.
Xingjian Tian, Bing Lu 0001, Shengkai Lyu, Shengquan Yin, Wenjing Huang 0002, Hairui Zhao 0002, Guangming Tan, Dingwen Tao
HPDC9
2026 A Fully GPU-Accelerated Framework for High-Performance Configuration Interaction Selection with Neural Network Quantum States
abstract
AI-driven methods have demonstrated considerable success in tackling the central challenge of accurately solving the Schrödinger equation for complex many-body systems. Among neural network quantum state (NNQS) approaches, the NNQS-SCI (Selected Configuration Interaction) method stands out as a state-of-the-art technique, recognized for its high accuracy and scalability. However, its application to larger systems is severely constrained by a hybrid CPU-GPU architecture. Specifically, centralized CPU-based global de-duplication creates a severe scalability barrier due to communication bottlenecks, while host-resident coupled-configuration generation induces prohibitive computational overheads. We introduce QiankunNet-cuSCI, a fully GPU-accelerated SCI framework designed to overcome these bottlenecks. It first integrates a distributed, load-balanced global de-duplication algorithm to minimize redundancy and communication overhead at scale. To address compute limitations, it employs specialized, fine-grained CUDA kernels for exact coupled configuration generation. Finally, to break the single-GPU memory barrier exposed by this full acceleration, it incorporates a GPU memory-centric runtime featuring GPU-side pooling, streaming mini-batches, and overlapped offloading. This design enables much larger configuration spaces and shifts the bottleneck from host-side limitations back to on-device inference. Our evaluation demonstrates that our work fundamentally expands the scale of solvable problems. On an NVIDIA A100 cluster with 64 GPUs, our work achieves up to 2.32 × end-to-end speedup over the highly-optimized NNQS-SCI baseline while preserving the same chemical accuracy. Furthermore, it demonstrates excellent distributed performance, maintaining over 90% parallel efficiency in strong scaling tests.
Daran Sun, Bowen Kan, Haoquan Long, Hairui Zhao 0002, Haoxu Li, Ankang Feng, Wenjing Huang 0002, Yida Gu, Honghui Shang, Yunquan Zhang, Dingwen Tao, Ninghui Sun, Guangming Tan
HPDC4
2026 Rehabilitating over Recomputing: A Novel Failure Recovery Method for Large Model Training
Hongliang Li 0003, Jie Wu 0001, Zhewen Xu, Hairui Zhao 0002, Haixiao Xu
INFOCOM5
2026 ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
Jinwu Yang, Jiaan Wu, Xinyang Ma, Hairui Zhao 0002, Yida Gu, Yuanhong Huang, Wenjing Huang 0002, Yili Ma, Zhongzhe Hu, Shaoteng Liu, Jiaxun Lu, Guangming Tan, Dingwen Tao
ISCA5
2026 PRISM: An Efficient GPU-Based Lossy Compression Framework for Progressive Data Retrieval with Multi-Level Interpolation
abstract
With the exponential growth of computing power, large-scale scientific simulations are producing massive volumes of data, leading to critical storage and I/O challenges. Error-bounded lossy compression has become one of the most effective solutions for reducing data size while preserving accuracy. Meanwhile, to achieve high-performance compression on such large datasets, leveraging GPUs has become increasingly essential. GPU-based lossy compressors deliver strong performance, but typically support only single-precision decompression, limiting their ability to meet the diverse accuracy requirements of scientific workflows. Progressive compressors can address this limitation by enabling on-demand precision retrieval. However, existing progressive lossy compressors on GPU still suffer from low throughput. To overcome these challenges, we present PRISM, a GPU-based progressive lossy compressor that achieves both high throughput and multi-precision retrieval, which introduces a high performance progressive framework that integrates the multiple interpolation predictors, efficient bitplane extraction, and an enhanced lossless compression that combines sign-absolute coding with zero-aware parallel algorithms. Evaluations on representative real-world datasets from five scientific domains show that PRISM significantly outperforms state-of-the-art progressive compressors on GPU, reducing retrieval data volume by over 15.6× and achieving up to 20.1× higher throughput on the NVIDIA H100 GPU under the same error bounds.
Bing Lu 0001, Hairui Zhao 0002, Dejun Luo, Wenjing Huang 0002, Yida Gu, Jinyang Liu 0003, Guangming Tan, Dingwen Tao
PPoPP3
2026 CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model Training
abstract
As training scales grow, collective communication libraries (CCL) increasingly face anomalies arising from complex interactions among hardware, software, and environmental factors. These anomalies typically manifest as slow/hang communication, the most frequent and time-consuming category to diagnose. However, traditional diagnostic methods remain inaccurate and inefficient, frequently requiring hours or even days for root cause analysis. To address this, we propose CCL-D, a high-precision diagnostic system designed to detect and locate slow/hang anomalies in large-scale distributed training. CCL-D integrates a rank-level real-time probe with an intelligent decision analyzer. The probe measures cross-layer anomaly metrics using a lightweight distributed tracing framework to monitor communication traffic. The analyzer performs automated anomaly detection and root-cause location, precisely identifying the faulty GPU rank. Deployed on a 4,000-GPU cluster over one year, CCL-D achieved near-complete coverage of known slow/hang anomalies and pinpointed affected ranks within 6 minutes—substantially outperforming existing solutions.
Yida Gu, Fakang Wang, Jianhao Fu, Zhenhang Sun, Qianyu Zhang 0001, Hairui Zhao 0002, Wenjing Huang 0002, Jinwu Yang, Yueyuan Zhou, Qian Zhao 0021, Haoxu Li, Zhan Wang 0003, Guangming Tan, Dingwen Tao
PPoPP6
2026 COCCL: A Collective Communication Library Supporting Easy Integration and Configuration of Customized Compression for Scalable LLM Training
abstract
Collective communication is critical to scaling large language model (LLM) training across various parallelism strategies, including data, tensor, and pipeline parallelism on GPU clusters. However, as model sizes and training scales increase, communication overhead is emerging as a major performance bottleneck. While compression is a promising mitigation strategy, existing solutions often lack user-transparency, hinder deployment and extensibility, and are not co-designed with communication algorithms. To address these limitations, we present COCCL, a high-performance collective communication library built on top of NCCL. COCCL introduces a novel programming model that can easily integrate compression into communication workflows with flexible configurability. It features a suite of compression-aware collective algorithms and runtime overlap mechanisms that mitigate error propagation and reduce computational overhead. We integrate well-established compression techniques into COCCL and tune the compression configurations during 3D-parallel training on GPT and Qwen models with up to 7 billion parameters. Using the optimal configuration (COCCL-3D), we achieve 1.24× throughput improvement while maintaining training accuracy.
Haoran Kong, Hairui Zhao 0002, Shengkai Lyu, Xingjian Tian, Liyang Zhao, Zhuohan Chen, Fakang Wang, Zizhong Chen, Zhan Wang 0003, Guangming Tan, Dingwen Tao
PPoPP3
2026 KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
Xinyang Ma, Dejun Luo, Hairui Zhao 0002, Bing Lu 0001, Wenjing Huang 0002, Yida Gu, Jinyang Liu 0003, Dingwen Tao, Guangming Tan
SIGCOMM4
2026 Sift: Channel-Wise Historical Embedding for High Efficiency Distributed Graph Neural Network Training with Accuracy Guarantee
abstract
Distributed Graph Neural Network (DGNN) is a powerful tool in large-scale graph representation learning. However, high data-transfer overhead among workers in a DGNN training job confines its scalability and thus the overall performance. Vertex-wise historical embedding methods have demonstrated high potential to alleviate the problems, but still suffer from severe accuracy loss and limited performance scalability, which has been attributed to the information loss of critical channels in historical vertices and redundant information in local channels. This article explores the optimization of channel level and construct a quantitative accuracy model for channel-wise historical embedding. We propose Sift, a novel DGNN training framework, supporting channel-wise partial historical embedding with accuracy guarantee. Sift has three components: a historical embedding evaluator with channel-wise quantitative accuracy model, a sawtooth-like matrix rearrangement for accelerating message passing, and a hybrid parallel framework for overlapping communication overhead. Comprehensive experimental results show that Sift achieves near-linear parallel convergence speedup, outperforming the state-of-the-art baselines by up to 72% in total training performance and up to 21% in convergence speed.
Zhewen Xu, Hongliang Li 0003, Junze Han, Hengshan Yue, Hairui Zhao 0002, Dongyuan Tian, Zijian Li 0007, Xiaohui Wei 0002
ACM Trans. Archit. Code Optim.5
2025 ArrayPipe: Introducing Job-Array Pipeline Parallelism for High Throughput Model Exploration
Hairui Zhao 0002, Hongliang Li 0003, Jie Wu 0001, Zhewen Xu, Xiang Li 0197, Haixiao Xu
INFOCOM1
2025 FlexPipe: Maximizing Training Efficiency for Transformer-based Models with Variable-Length Inputs
Hairui Zhao 0002, Hongliang Li 0003, Zizhong Chen
USENIX ATC1
2025 Alleviating straggler impacts for data parallel deep learning with hybrid parameter update
Hongliang Li 0003, Hairui Zhao 0002, Zhewen Xu
Future Gener. Comput. Syst.4
2025 Convergence-aware optimal checkpointing for exploratory deep learning training jobs
Hongliang Li 0003, Hairui Zhao 0002, Xiang Li 0197, Haixiao Xu
Future Gener. Comput. Syst.3
2024 EFNAS: Efficient Federated Neural Architecture Search Across AIoT Devices
abstract
Federated neural architecture search tailors deep learning models to accommodate varied client data in Federated Learning (FL) scenarios. However, the simultaneous optimization of multiple subnetworks leads to substantial GPU memory overhead in differentiable Neural Architecture Search (NAS) methods. Additionally, after each client searches local architecture, the conventional weighted averaging approach may result in a loss of architectural feature information and failure to capture architectural diversity, limiting model performance and expressiveness. To address these challenges, we propose EFNAS, a novel computation-efficient and aggregation-effective Federated NAS framework. Specifically, we propose Single Path Local Search (SPLS) to automatically search for the optimal network architecture with minimal complexity. SPLS uses Gumbel Softmax to continually reparameterize the probability distribution to reduce complexity. Furthermore, we propose Client-Centric Architecture Aggregation (CCAA), considering the aggregated architecture as a graph with subgraphs representing client architectures. CCAA leverages probabilistic distributions extracted from subgraphs that frequently appear across multiple clients to derive a global architecture. Guided by this global architecture, each client compares the accuracy of its locally searched architecture with the global architecture, selecting the architecture best suited for its final model architecture. Comprehensive experiments on various datasets demonstrate that EFNAS achieves excellent performance while guaranteeing efficiency during searching compared to other methods.
Xiaohui Wei 0002, Guanhua Chen 0003, Hairui Zhao 0002, Hengshan Yue
IJCNN4
2024 Visage: Visual-Aware Generation of Adversarial Examples in Black-Box for Text Classification
Hairui Zhao 0002, Hongliang Li 0003
NLPCC (4)1
2024 Interference-aware opportunistic job placement for shared distributed deep learning clusters
Hongliang Li 0003, Hairui Zhao 0002, Xiang Li 0197, Haixiao Xu
J. Parallel Distributed Comput.2
2023 ExplSched: Maximizing Deep Learning Cluster Efficiency for Exploratory Jobs
abstract
Resource management for Deep Learning (DL) clusters is essential for system efficiency and model training quality. Existing schedulers provided by DL frameworks are mostly adaptations from traditional HPC clusters and usually work on jobs’ makespan, assuming that DL training jobs finish completely. Unfortunately, it is reported that a fair amount of training jobs are exploratory jobs and often finish unsuccessfully (over 30%) in production clusters. This is due to the distinct characteristic of Deep Neural Network (DNN) training that it is an exploratory process of frequent user interventions, such as adjusting model structures, tuning hyperparameters, and exploring feature validity. Existing DL cluster schedulers using offline algorithms are not suitable for exploratory jobs when unexpected early terminations can cause noticeable resource waste. Moreover, DL training jobs are iterative and usually yield diminishing returns as they progress. Equally allocating resource among training iterations is not efficient, especially when dealing with exploratory jobs where it can worsen the degradation of system efficiency. The fundamental goal of a DL training job is to gain model quality improvement, usually indicated by the loss reduction (job profit) of a DNN model. This paper introduces a novel scheduling problem for exploratory jobs that seeks to maximize the overall training profit of a DL cluster. We propose ExplSched, an online scheduling solution based on the primal-dual framework, resulting in a competitive ratio of 2α that belongs to O(ln n). It uses a resource price function that emphasizes the importance of job profit to resource consumption ratio to make quick resource allocation decisions. Experimental results show that ExplSched achieved an average system utility improvement of 87.28% compared with other related work.
Hongliang Li 0003, Hairui Zhao 0002, Zhewen Xu, Xiang Li 0197, Haixiao Xu
CLUSTER2