Youhui Bai

dblp:205/7130 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
9since 2021 · last 2026
0009-0007-6073-7011ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SMIDT: High-Performance Inference Framework for MoE Models with Dynamic Top-K Routing
abstract
To accelerate Mixture-of-Experts (MoE) inference, the hybrid parallelism paradigm is first applying pipeline parallelism (PP) to vertically divide the model into stages, with each stage further divided horizontally using tensor or expert parallelism. On the algorithm side, dynamic Top-K routing reduces computation by activating fewer experts per token on average. In this paper, we explore the application of dynamic Top-K routing to PP-enabled MoE inference, aiming to fully unleash their combined potential. We identify key performance bottlenecks arising from Top-K value variation across layers, which conflicts with PP's typically uniform stage partitioning, as well as opportunities to optimize memory usage through their integration. To address these challenges, we present SMIDT, an efficient MoE inference framework tailored for dynamic Top-K routing. SMIDT features: (1) an adaptive, module-level uneven partitioning strategy to balance computation across PP stages, (2) a memory-aware expert replication scheme (DPMoE) that reduces communication overhead, and (3) a lightweight search algorithm combining binary search and dynamic programming to generate efficient parallelism plans. We implement SMIDT on SGLang, a state-of-the-art LLM inference framework, evaluate it on 32 A40 GPUs and 16 A100 GPUs, and compare with manually tuned parallelism strategies. Experimental results show that, when co-locating prefill and decoding phases, SMIDT achieves 1.20–3.13x throughput improvements for prefill-only tasks and 1.05–1.89x for prefill-decoding tasks. When disaggregating prefill and decoding tasks, SMIDT improves average and P99 time-to-first-token (TTFT) by 1.10–1.17x and 1.21–1.26x, respectively.
Zewen Jin, Shen Fu, Chengjie Tang, Youhui Bai, Jiaan Zhu, Chizheng Fang, Ping Gong 0009, Cheng Li 0001
AAAI4
2026 FlexUSFL: A Scalable and Memory-Efficient U-Shaped Split Federated Learning Framework for Large Language Model Fine-Tuning
Youhui Bai, Sai Wu
APPT4
2026 nnScaler-M: Constraint-Guided and Placement-Aware Parallelization Plan Generation for Deep Learning Training
abstract
As deep neural networks grow, training increasingly relies on handcrafted search spaces for efficient parallelization plans. However, our study shows existing spaces exclude optimal plans for models like AlphaFold2 and large language models with large embedding tables. We propose nScaler-M, a framework for generating efficient parallelization plans for deep learning training. Instead of searching within predefined spaces, nScaler-M empowers domain experts to compose custom search spaces using three primitives,op-trans, op-assign, andop-order, which capture model transformation, spatial assignment, and temporal scheduling. Besides, nScaler-M captures device placement and communication patterns viap-meshandc-mesh, which enhances the accuracy of communication cost estimation, ultimately supporting the search for optimal plans on heterogeneous networks. To avoid space explosion, nScaler-M allows constraints to be applied to these primitives, effectively pruning the search space. With the proposed primitives and constraints, nScaler-M can compose existing search spaces as well as new ones. Experiments show that nScaler-M can find new parallelization plans that achieve up to 3.5× speedup for popular DNN models. Additionally, equipped withp-meshandc-mesh, nScaler-M can discover optimized parallelization plans achieving up to 2.79× higher throughput when training LLaMA-3 models, and introduce acceptable searching overhead.
Jiaan Zhu, Chizheng Fang, Zewen Jin, Youhui Bai, Cheng Li 0001
IEEE Trans. Parallel Distributed Syst.6
2025 BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and Inference
abstract
The Mixture-of-Experts (MoE) structure scales the Transformer-based large language models (LLMs) and improves their performance with only the sub-linear increase in computation resources. Recently, a fine-grained DeepSeekMoE structure is proposed, which can further improve the computing efficiency of MoE without performance degradation. However, the All-to-All communication introduced by MoE has become a bottleneck, especially for the fine-grained structure, which typically involves and activates more experts, hence contributing to heavier communication overhead. In this paper, we propose a novel MoE structure named BigMac, which is also fine-grained but with high communication efficiency. The innovation of BigMac is mainly due to that we abandon the Communicate-Descend-Ascend-Communicate (CDAC) manner used by fine-grained MoE, which leads to the All-to-All communication always taking place at the highest dimension. Instead, BigMac designs an efficient Descend-Communicate-Communicate-Ascend (DCCA) manner. Specifically, we add a descending and ascending projection at the entrance and exit of the expert, respectively, which enables the communication to perform at a very low dimension. Furthermore, to adapt to DCCA, we re-design the structure of small experts, ensuring that the expert in BigMac has enough complexity to address tokens. Experimental results show that BigMac achieves comparable or even better model quality than fine-grained MoEs with the same number of experts and a similar number of total parameters. Equally importantly, BigMac reduces the end-to-end latency by up to 3.09 x for training and increases the throughput by up to 3.11 x for inference on state-of-the-art AI computing frameworks including Megatron, Tutel, and DeepSpeed-Inference.
Zewen Jin, Jiaan Zhu, Hongrui Zhan, Youhui Bai, Zhenyu Ming
AAAI5
2025 A Generic, High-Performance, Compression-Aware Framework for Data Parallel DNN Training
abstract
Gradient compression is a promising approach to alleviating the communication bottleneck in data parallel deep neural network (DNN) training by significantly reducing the data volume of gradients for synchronization. While gradient compression is being actively adopted by the industry (e.g., Facebook and AWS), our study reveals that there are two critical but often overlooked challenges: 1) inefficient coordination between compression and communication during gradient synchronization incurs substantial overheads, and 2) developing, optimizing, and integrating gradient compression algorithms into DNN systems imposes heavy burdens on DNN practitioners, and ad-hoc compression implementations often yield surprisingly poor system performance. In this paper, we propose a compression-aware gradient synchronization architecture,CaSync, which relies on flexible composition of basic computing and communication primitives. It is general and compatible with any gradient compression algorithms and gradient synchronization strategies and enables high-performance computation-communication pipelining. We further introduce a gradient compression toolkit,CompLL, to enable efficient development and automated integration of on-GPU compression algorithms into DNN systems with little programming burden. Lastly, we build a compression-aware DNN training frameworkHiPresswithCaSyncandCompLL.HiPressis open-sourced and runs on mainstream DNN systems such as MXNet, TensorFlow, and PyTorch. Evaluation via a 16-node cluster with 128 NVIDIA V100 GPUs and a 100 Gbps network shows thatHiPressimproves the training speed over current compression-enabled systems (e.g., BytePS-onebit, Ring-DGC and PyTorch-PowerSGD) by 9.8%-69.5% across six popular DNN models.
Hao Wu 0077, Youhui Bai, Cheng Li 0001, Feng Yan 0001, Ruichuan Chen, Yinlong Xu 0001
IEEE Trans. Parallel Distributed Syst.3
2023 MPress: Democratizing Billion-Scale Model Training on Multi-GPU Servers via Memory-Saving Inter-Operator Parallelism
abstract
It remains challenging to train billion-scale DNN models on a single modern multi-GPU server due to the GPU memory wall. Unfortunately, existing memory-saving techniques such as GPU-CPU swap, recomputation, and ZeRO-Series come at the price of extra computation, communication overhead, or limited memory reduction.We present MPress, a new single-server multi-GPU system that breaks the GPU memory wall of billion-scale model training while minimizing extra cost. MPress first discusses the trade-offs of various memory-saving techniques and offers a holistic solution, which alternatively chooses the inter-operator parallelism with low cross-GPU communication traffics, and combines with recomputation and swap, to balance training performance and sustained model sizes. Additionally, MPress employs a novel, fast D2D swap technique, which simultaneously utilizes multiple high-bandwidth NVLink to swap tensors to light-load GPUs, based on a key observation that inter-operator parallel training may result in imbalanced GPU memory utilization and spare memory space from least used devices plus the high-end interconnects among them have the opportunity to support low-overhead swapping. Finally, we integrate MPress with PipeDream and DAPPLE, two representative inter-operator parallel training systems. Experimental results with two popular DNN models, Bert, and GPT, on two modern GPU servers from the DGX-1 and DGX-2 generation, equipped with 8 V100 or A100 cards, respectively, demonstrate that MPress significantly improves the training throughput over ZeRO-Series with the identical memory reduction, while being able to train larger models than the recomputation baseline.
Cheng Li 0001, Youhui Bai, Feng Yan 0001, Yinlong Xu 0001
HPCA5
2023 A Survey on Auto-Parallelism of Large-Scale Deep Learning Training
abstract
Deep learning (DL) has gained great success in recent years, leading to state-of-the-art performance in research community and industrial fields like computer vision and natural language processing. One of the reasons for this success is the huge amount parameters adopted in DL models. However, it is impractical to train a moderately large model with a large number of parameters on a typical single device. Thus, It is necessary to train DL models in clusters with distributed training algorithms. However, traditional distributed training algorithms are usually sub-optimal and highly customized, which owns the drawbacks to train large-scale DL models in varying computing clusters. To handle the above problem, researchers propose auto-parallelism, which is promising to train large-scale DL models efficiently and practically in various computing clusters. In this survey, we perform a broad and thorough investigation on challenges, basis, and strategy searching methods of auto-parallelism in DL training. First, we abstract basic parallelism schemes with their communication cost and memory consumption in DL training. Further, we analyze and compare a series of current auto-parallelism works and investigate strategies and searching methods which are commonly used in practice. At last, we discuss several trends in auto-parallelism which are promising in further research.
Peng Liang 0017, Xiaoda Zhang, Youhui Bai, Teng Su, Zhiquan Lai, Linbo Qiao, Dongsheng Li 0001
IEEE Trans. Parallel Distributed Syst.4
2021 Gradient Compression Supercharged High-Performance Data Parallel DNN Training
abstract
Gradient compression is a promising approach to alleviating the communication bottleneck in data parallel deep neural network (DNN) training by significantly reducing the data volume of gradients for synchronization. While gradient compression is being actively adopted by the industry (e.g., Facebook and AWS), our study reveals that there are two critical but often overlooked challenges: 1) inefficient coordination between compression and communication during gradient synchronization incurs substantial overheads, and 2) developing, optimizing, and integrating gradient compression algorithms into DNN systems imposes heavy burdens on DNN practitioners, and ad-hoc compression implementations often yield surprisingly poor system performance.
Youhui Bai, Cheng Li 0001, Ping Gong 0009, Feng Yan 0001, Ruichuan Chen, Yinlong Xu 0001
SOSP1
2021 Efficient Data Loader for Fast Sampling-Based GNN Training on Large Graphs
abstract
Emerging graph neural networks (GNNs) have extended the successes of deep learning techniques against datasets like images and texts to more complex graph-structured data. By leveraging GPU accelerators, existing frameworks combine mini-batch and sampling for effective and efficient model training on large graphs. However, this setup faces a scalability issue since loading rich vertex features from CPU to GPU through a limited bandwidth link usually dominates the training cycle. In this article, we propose PaGraph, a novel, efficient data loader that supports general and efficient sampling-based GNN training on single-server with multi-GPU. PaGraph significantly reduces the data loading time by exploiting available GPU resources to keep frequently-accessed graph data with a cache. It also embodies a lightweight yet effective caching policy that takes into account graph structural information and data access patterns of sampling-based GNN training simultaneously. Furthermore, to scale out on multiple GPUs, PaGraph develops a fast GNN-computation-aware partition algorithm to avoid cross-partition access during data-parallel training and achieves better cache efficiency. Finally, it overlaps data loading and GNN computation for further hiding loading costs. Evaluations on two representative GNN models, GCN and GraphSAGE, using two sampling methods, Neighbor and Layer-wise, show that PaGraph could eliminate the data loading time from the GNN training pipeline, and achieve up to 4.8× performance speedup over the state-of-the-art baselines. Together with preprocessing optimization, PaGraph further delivers up to 16.0× end-to-end speedup.
Youhui Bai, Cheng Li 0001, Yufei Wu 0011, Youshan Miao, Yunxin Liu 0001, Yinlong Xu 0001
IEEE Trans. Parallel Distributed Syst.1
2017 PDS: An I/O-Efficient Scaling Scheme for Parity Declustered Data Layout
abstract
Parity declustering is widely deployed in erasure coded storage systems so as to provide fast recovery and high data availability. However, to perform scaling on such RAIDs, it is necessary to preserve the parity declustered data layout so as to guarantee the RAID performance after scaling. Unfortunately, existing scaling algorithms fail to achieve this goal so they can not be applied for scaling RAIDs which have deployed parity declustering. To address this challenge, we develop an efficient scaling algorithm called PDS (Parity Declustering Scaling). In particular, we first employ an auxiliary Balanced Incomplete Block Design (BIBD) to define the data migrations during scaling so as to preserve parity declustered data layout, and then define the addressing algorithm in the scaled system based on the migrations. We provide theoretical proofs to show that PDS preserves the parity declustered data layout, which is the basis for scaling RAIDs with parity declustering, and also theoretically prove that PDS achieves the even distribution of data/parity blocks after scaling and requires only the minimal data migrations. To show the performance of PDS, we implement it in MD in Linux Kernel, and conduct experiments with real-world traces. Results show PDS can reduce 89.70% of data migration time and 24.44% of user response time during scaling on average, compared with the round-robin scheme.
Zhipeng Li 0005, Yinlong Xu 0001, Yongkun Li 0001, Chengjin Tian, Youhui Bai
ICPP5