Zhongzhe Hu

dblp:262/1309 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-6708-3942ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EPLoN: Exploiting Efficient Parallelism with Selective Rematerialization for Lightning Attention on Ascend NPU
abstract
The quadratic computational complexity of softmax attention presents a fundamental bottleneck to scaling modern language models to long sequences. While the proposed Lightning Attention mechanism offers a linear-complexity alternative, its state-of-the-art implementations remain predominantly optimized for GPU architectures and fail to fully leverage the capabilities of alternative accelerators such as Ascend NPUs. To bridge this gap, we propose EPLoN (Exploiting Efficient Parallelism with Selective Rematerialization for Lightning Attention on NPU). EPLoN presents a high-performance implementation of Lightning Attention optimized for heterogeneous Ascend NPUs. EPLoN reformulates the algorithm, introducing an efficient parallelism scheme with a rematerialization strategy based on inter- and intra-core that maximizes the utilization of the NPU architecture. In a cross-architectural comparison against the state-of-the-art FlashLinearAttention (FLA) on an Nvidia GPU of comparable computational capacity, our evaluation achieves a speedup of up to 3.39 × and a geometric mean speedup of 1.73 ×, while reducing peak memory consumption by approximately 33%.
Zhenfeng Su, Alexander Setyaev, Stanislav Kamenev, Alexander Gneushev, Junmin Xiao, Anastasiya Bistrigova, Sergey Buzykanov, Evgeny Tetin, Guangming Tan, Boxiao Liu, Xueyi Zou, Zhenhua Dong, Constantine Korikov, Xianzhi Yu, Zhongzhe Hu
ICS18
2026 ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
Jinwu Yang, Jiaan Wu, Xinyang Ma, Hairui Zhao 0002, Yida Gu, Yuanhong Huang, Wenjing Huang 0002, Yili Ma, Zhongzhe Hu, Shaoteng Liu, Jiaxun Lu, Guangming Tan, Dingwen Tao
ISCA15
2026 UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods
abstract
The deployment of Mixture-of-Experts (MoE) models on production high-bandwidth superpods, such as NVIDIA's NVL72/576 and Huawei's CloudMatrix384, introduces critical challenges beyond raw interconnect bandwidth. While these systems provide unified global address spaces and high-bandwidth fabrics, their full potential for sparse MoE communication is hindered by three fundamental bottlenecks: (1) Strict execution serialization imposed by coarse-grained Bulk Synchronous Parallel (BSP) orchestration of interdependent communication phases; (2) Prohibitive synchronization overhead that fails to scale alongside high interconnect bandwidth; and (3) Severe load imbalance resulting from distance-agnostic scheduling of irregular token traffic. To eliminate these bottlenecks, we introduce UBEP (Unified-Bus Expert Parallelism), a production-ready communication library that rethinks MoE's All-to-All primitives for modern superpod architectures. Through large-scale experiments, UBEP reduces All-to-All latency by up to 52.4% and MoE inference Time Per Output Token (TPOT) by up to 11.1%.
Chang Liu 0001, Si Shen, Jiaqi Zheng 0001, Mingfan Li, Yuyang Yang, Guanhua Li, Yuquan Zhang, Zhongzhe Hu, Qihang Duan, Wenkai Ling, Baochuan Yang, Xianzhi Yu, Guihai Chen
SIGCOMM10
2025 Hypertron: Efficiently Scaling Large Models by Exploring High-Dimensional Parallelization Space
abstract
Large models are evolving towards massive scale, diverse model architectures (dense and sparse) and long-context processing, which makes it very challenging to efficiently scale large models on parallel machines. The current widely-used parallelization strategies are often sub-optimal due to their limited parallelization strategy space. To this end, we propose Hypertron, a scalable parallel large-model training framework which incorporates an unprecedented high-dimensional (up to 7D) parallelization space, a holistic scheme for efficient dimension fusion, and a comprehensive performance model to guide the high-dimensional exploration. By exploiting the high-dimensional space to discover the optimal strategy which is not supported by existing frameworks, Hypertron significantly reduces memory and communication cost while improving parallel scalability. Extensive evaluations demonstrate that Hypertron achieves up to 56.7% Model FLOPs Utilization (MFU) on 2,048 new-generation Ascend NPU accelerators (scaling with supernodes) for different large models (such as sparse 141B and dense 310B), with 1.33x speedup over the best configuration of the state-of-the-art frameworks.
Shigang Li 0002, Jingkun Dong, Jihao Chen, Zhongzhe Hu
SC5
2024 ScheMoE: An Extensible Mixture-of-Experts Distributed Training System with Tasks Scheduling
abstract
In recent years, large-scale models can be easily scaled to trillions of parameters with sparsely activated mixture-of-experts (MoE), which significantly improves the model quality while only requiring a sub-linear increase in computational costs. However, MoE layers require the input data to be dynamically routed to a particular GPU for computing during distributed training. The highly dynamic property of data routing and high communication costs in MoE make the training system low scaling efficiency on GPU clusters. In this work, we propose an extensible and efficient MoE training system, ScheMoE, which is equipped with several features. 1) ScheMoE provides a generic scheduling framework that allows the communication and computation tasks in training MoE models to be scheduled in an optimal way. 2) ScheMoE integrates our proposed novel all-to-all collective which better utilizes intra- and inter-connect bandwidths. 3) ScheMoE supports easy extensions of customized all-to-all collectives and data compression approaches while enjoying our scheduling algorithm. Extensive experiments are conducted on a 32-GPU cluster and the results show that ScheMoE outperforms existing state-of-the-art MoE systems, Tutel and Faster-MoE, by 9%-30%.
Shaohuai Shi, Xinglin Pan, Qiang Wang 0022, Chengjian Liu, Xiaozhe Ren, Zhongzhe Hu, Bo Li 0001, Xiaowen Chu 0001
EuroSys6
2022 MegTaiChi: dynamic tensor-based memory management optimization for DNN training
abstract
In real applications, it is common to train deep neural networks (DNNs) on modest clusters. With the continuous increase of model size and batch size, the training of DNNs becomes challenging under restricted memory budget. The tensor partition and tensor rematerialization are two major memory optimization techniques to enable larger model size and batch size within the limited-memory constrain. However, the related algorithms failed to fully extract the memory reduction opportunity, because they ignored the invariable characteristics of dynamic computational graphs and the variation among the same size tensors at different memory locations. In this work, we propose MegTaiChi, a dynamic tensor-based memory management optimization module for the DNN training, which first achieves an efficient coordination of tensor partition and tensor rematerialization. The key feature of MegTaiChi is that it makes memory management decisions based on dynamic tensor access pattern tracked at runtime. This design is motivated by the observation that the access pattern to tensors is regular during training iterations. Based on the identified patterns, MegTaiChi exploits the total memory optimization space and achieves the heuristic, adaptive and fine-grained memory management. The experimental results show, MegTaiChi can reduce the memory footprint by up to 11% for ResNet-50 and 10.5% for GL-base compared with DTR. For the training of 6 representative DNNs, MegTaiChi outperforms MegEngine and Sublinear by 5X and 2.4X of the maximum batch sizes. Compared with FlexFlow, Gshard and ZeRo-3, MegTaiChi achieves 1.2X, 1.8X and 1.5X performance speedups respectively on average. For the million-scale face recognition application, Meg-TaiChi achieves 1.8X speedup compared with the optimal empirical parallelism strategy on 256 GPUs.
Zhongzhe Hu, Junmin Xiao, Zheye Deng, Ninghui Sun, Guangming Tan
ICS1
2022 Fast and accurate variable batch size convolution neural network training on large scale distributed systems
abstract
Abstract Large‐scale distributed convolution neural network (CNN) training brings two performance challenges: model performance and system performance. Large batch size usually leads to model test accuracy loss, which counteracts the benefits of parallel SGD. The existing solutions require massive hyperparameter hand‐tuning. To overcome this difficult, we analyze the training process and find that earlier training stages are more sensitive to batch size. Accordingly, we assert that different stages should use different batch size, and propose a variable batch size strategy. In order to remain high test accuracy under larger batch size cases, we design an auto‐tuning engine for automatic parameter tuning in the proposed variable batch size strategy. Furthermore, we develop a dataflow implementation approach to achieve the high‐throughput CNN training on supercomputer system. Our approach has achieved high generalization performance on SOAT CNN networks. For the ShuffleNet, ResNet‐50, and ResNet‐101 training with ImageNet‐1K dataset, we scale the batch size to 120 K without accuracy loss and to 128 K with only a slight loss. And the dataflow implementation approach achieves 93.5% scaling efficiency on 1024 GPUs compared with the state‐of‐the‐art.
Zhongzhe Hu, Junmin Xiao, Ninghui Sun, Guangming Tan
Concurr. Comput. Pract. Exp.1