Zhiquan Lai

dblp:143/4279 · DBLP profile ↗
← Back
42ranked-venue papers
2as first author
38since 2021 · last 2026
0000-0002-3458-4732ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 31 · 2 first-author · 28 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Computer networks · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 DOA: Dataflow Optimization for Attention on Multi-core DSPs with Three-Level Memory Hierarchy
Zhiquan Lai, Shun Ouyang, Zhaoning Zhang 0001, Menghan Jia, Huayou Su, Dongsheng Li 0001
APPT2
2026 DAG-P: Efficient Fine-Grained Partitioning of Open-Source Transformer Models via Dependency-Aligned Planning
Chenke Yi, Zhiquan Lai, Dongsheng Li 0001
Euro-Par (2)2
2026 AdaCheck: An Adaptive Checkpointing System for Efficient LLM Training with Redundancy Utilization
Zhiquan Lai, Ke-shi Ge, Qiaoling Chen, Peng Sun 0006, Dongsheng Li 0001, Kai Lu 0001
FAST3
2026 MeCache: Communication-Efficient Multi-GPU Heterogeneous Graph Neural Network Training
Gongqingjian Jiang, Menghan Jia, Zhiquan Lai, Dongsheng Li 0001
IPDPS4
2026 Di-PS: System-Algorithm Co-Design for Asynchronous and Heterogeneous Cross-cluster LLM Training at Scale
Qiaoling Chen, Zhiquan Lai, Penglong Jiao, Wenwen Qu, Peng Sun 0006, Xingcheng Zhang, Xiaoge Deng, Dongsheng Li 0001, Kai Lu 0001, Tianwei Zhang 0004
NSDI3
2026 Physics-informed residual learning with low-rank adaptation for unsupervised mesh generation
Jiaming Peng, Xinhai Chen 0001, Qingling Wang, Zhiquan Lai, Dongsheng Li 0001, Jie Liu 0002
Comput. Aided Geom. Des.6
2026 Parallelsim: an accurate, generic, and efficient simulator for distributed deep learning
Peng Liang 0017, Linbo Qiao, Zhiquan Lai, Dongsheng Li 0001
CCF Trans. High Perform. Comput.3
2026 Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-Scale MoE Models
abstract
The size of deep learning models has been increasing to enhance model quality. The linear increase in training computation budgets with model size means that training an extremely large-scale model is exceedingly time-consuming. Recently, the Mixture of Experts (MoE) has drawn significant attention as it can scale models to extra-large sizes with a near-stable computation budget. However, inefficient distributed training of large-scale MoE models hinders their broader application. Specifically, a considerable dynamic load imbalance occurs among devices during training, significantly reducing throughput. Several load-balancing works have been proposed to address the challenge. System-level solutions draw more attention for their hardware affinity and non-disruption of model convergence compared to algorithm-level ones. However, they are troubled by high communication costs and poor communication-computation overlap. To address these challenges, we propose a systematic load-balancing method, Pro-Prophet, which consists of a planner and a scheduler for efficient parallel training of large-scale MoE models. To adapt to the dynamic load imbalance, we have profiled training statistics and utilized them to design Pro-Prophet. For lower communication volume, the Pro-Prophet planner determines a series of lightweight load-balancing strategies and efficiently searches for a communication-efficient one for training based on the statistics. For sufficient overlapping of communication and computation, the Pro-Prophet scheduler schedules the data-dependent operations based on the statistics and operation characteristics, further improving the training throughput. We conduct extensive experiments in various clusters and MoE models. The results indicate that Pro-Prophet achieves up to 2.66x speedup on MoE-GPT models compared to two popular MoE frameworks, namely Deepspeed-MoE and FasterMoE. Furthermore, Pro-Prophet has demonstrated a load-balancing improvement of up to 11.01x and speedups on modern MoE models up to 1.22x compared to a representative load-balancing work, FasterMoE.
Zhiquan Lai, Dongsheng Li 0001, Ke-shi Ge, Huayou Su
IEEE Trans. Parallel Distributed Syst.2
2025 Capricorn: Efficient In-Memory Checkpointing for MoE Model Training with Dynamicity Awareness
abstract
Mixture-of-Experts (MoE) has been extensively adopted for its incredible capability to expand model scale with a sub-linear increase in computational requirement. Training MoE models requires substantial computing nodes and extended periods, necessitating reliable distributed training systems. Checkpointing is a common approach to enhance training reliability by periodically saving model states. Current checkpointing optimizations focus on hiding checkpoint overhead in model training computations. However, these approaches overlook the dynamicity inherent in distributed MoE training, leading to an inefficient checkpointing mechanism. In this paper, we propose Capricorn, a dynamicity-aware in-memory checkpointing approach for efficient MoE model training. We observe that the dynamicity impacts computation durations at both the layer and iteration levels. At the layer level, different model layers exhibit various computation durations, while at the iteration level, the computation time of the same layer differs across iterations. To adapt to the layer-level dynamicity, Capricorn employs online profiling at the granularity of individual layers. Based on the profiling results, it strategically partitions checkpoints into chunks and schedules checkpointing communication to overlap with model computations. To deal with the dynamicity across iterations, Capricorn speculatively activates the profiling and partitioning processes utilizing the temporal locality of the experts' load. It can produce an optimal activation for low runtime overhead with high checkpoint partition accuracy. For mainstream MoE models, Capricorn achieves up to$1.56 \times$and$5.98 \times$end-to-end training speedup over Gemini and TorchSnapshot respectively under per-iteration checkpointing.
Wenqian Xie, Zhiquan Lai, Yanqi Hao, Dongsheng Li 0001
CLUSTER2
2025 EAHP: An Efficient Automatic Hybrid Parallelism Approach with Genetic Algorithm
Yichen Gu, Zhiquan Lai, Yinghui Gao
ICA3PP (2)2
2025 HMGraph: Boosting GNN Training on Hierarchical Memory via Coordinated Cache
abstract
The GPU-CPU-SSD hierarchical memory systems are commonly employed for large-scale GNN training. However, existing solutions inefficiently utilize high-bandwidth memory due to coarse-grained memory management and poor data placement that ignores graph access patterns. This paper presents HMGraph, a GNN training system unleashing the full potential of hierarchical memory architectures. The core design of HMGraph is Coordinated Cache, integrating GPU memory and CPU memory as a cache layer for the hierarchical memory system and improving GNN efficiency through fine-grained data and memory management. For this goal, three main designs are proposed. First, we design an automatic cache management mechanism that optimizes cache allocation based on the graph data access pattern to enhance the overall cache hit rate. Second, we propose a dynamic data space management strategy to improve the efficiency of dynamic cache. Third, we develop a hierarchical memory-aware data partitioning strategy that further improves the utilization of high-bandwidth memory. Our evaluation of various large-scale graphs reveals that HMGraph significantly outperforms other state-of-the-art systems by 1.4-36.7 ×.
Menghan Jia, Zhiquan Lai, Qiao Li 0001, Yiming Zhang 0003, Dongsheng Li 0001
ICPP3
2025 FMCC-RT: a scalable and fine-grained all-reduce algorithm for large-scale SMP clusters
Jintao Peng, Jie Liu 0002, Jianbin Fang, Zhiquan Lai, Bo Yang 0023, Chunye Gong, Xinjun Mao, Guo Mao, Jie Ren 0007
Sci. China Inf. Sci.6
2025 Efficient deep neural network training via decreasing precision with layer capacity
abstract
Abstract Low-precision training has emerged as a practical approach, saving the cost of time, memory, and energy during deep neural networks (DNNs) training. Typically, the use of lower precision introduces quantization errors that need to be minimized to maintain model performance, often neglecting to consider the potential benefits of reducing training precision. This paper rethinks low-precision training, highlighting the potential benefits of lowering precision: (1) low precision can serve as a form of regularization in DNN training by constraining excessive variance in the model; (2) layer-wise low precision can be seen as an alternative dimension of sparsity, orthogonal to pruning, contributing to improved generalization in DNNs. Based on these analyses, we propose a simple yet powerful technique–DPC (Decreasing Precision with layer Capacity), which directly assigns different bit-widths to model layers, without the need for an exhaustive analysis of the training process or any delicate low-precision criteria. Thorough extensive experiments on five datasets and fourteen models across various applications consistently demonstrate the effectiveness of the proposed DPC technique in saving computational cost (−16.21%–−44.37%) while achieving comparable or even superior accuracy (up to +0.68%, +0.21% on average). Furthermore, we offer feature embedding visualizations and conduct further analysis with experiments to investigate the underlying mechanisms behind DPC’s effectiveness, enhancing our understanding of low-precision training. Our source code will be released upon paper acceptance.
Zhiquan Lai, Tao Sun 0005, Ke-shi Ge, Dongsheng Li 0001
Frontiers Comput. Sci.2
2025 Accurate and Efficient Fine-Tuning of Quantized Large Language Models Through Optimal Balance in Adaptation
abstract
Abstract Large Language Models (LLMs) have demonstrated impressive performance across various domains. However, the enormous number of model parameters makes fine-tuning challenging, significantly limiting their application and deployment. Existing solutions combine parameter quantization with Low-Rank Adaptation (LoRA), reducing memory usage but causing performance degradation. Additionally, converting fine-tuned models to low-precision representations further degrades performance. In this paper, we identify an imbalance in fine-tuning quantized LLMs with LoRA: overly complex adapter inputs and outputs versus low effective trainability of the adapter, leading to underfitting during fine-tuning. Thus, we propose Quantized LLMs fine-tuning with Balanced Low-Rank Adaptation (Q-BLoRA), which simplifies the adapter’s inputs and outputs while increasing the adapter’s rank to alleviate underfitting during fine-tuning. For low-precision deployment, we propose Quantization-Aware fine-tuning with Balanced Low-Rank Adaptation (QA-BLoRA), which aligns with the block-wise quantization and facilitates quantization-aware fine-tuning of low-rank adaptation based on the parameter merging of Q-BLoRA. Both Q-BLoRA and QA-BLoRA are easily implemented and offer the following optimizations: (i) Q-BLoRA consistently achieves state-of-the-art accuracy compared to baselines and other variants; (ii) QA-BLoRA enables the direct generation of low-precision inference models, which exhibit significant performance improvements over other low-precision models. We validate the effectiveness of Q-BLoRA and QA-BLoRA across various models and scenarios. Code has been made available at https://github.com/xiaocaigou/qbaraqahira.
Zhiquan Lai, Qiang Wang 0006, Xionglve Li, Dongsheng Li 0001
Trans. Assoc. Comput. Linguistics2
2025 AutoPipe-H: A Heterogeneity-Aware Data-Paralleled Pipeline Approach on Commodity GPU Servers
abstract
Recently, the data-parallel pipeline approach has been widely used in training DNN models on commodity GPU servers. However, there are still three challenges for hybrid parallelism on commodity GPU servers: i) a balanced model partition is crucial for efficiency, whereas prior works lack a sound solution to generate a balanced partition automatically; ii) an orchestrated device mapping is essential to reduce communication contention, however, prior works ignore server heterogeneity, exacerbating communication contention; iii) the startup overhead is inevitable and especially significant for deep pipelines, which is an essential source of pipeline bubbles and severely affects pipeline scalability. We proposeAutoPipe-Hto solve these three problems, which contains i) apipeline partitionercomponent for automatically and quickly generating a balanced sub-block partition scheme; ii) adevice mappingcomponent that assigns pipeline stages to devices, considering server heterogeneity, to reduce communication contention; and iii) adistributed training runtimecomponent that reduces pipeline startup overhead by splitting the micro-batch evenly. The experimental results show that AutoPipe-H can accelerate training by up to 1.26x over the hybrid parallelism framework DAPPLE and Piper, with a 2.73x-12.7x improvement in the partition balance and an order-of-magnitude time reduction in partition scheme searching.
Kai Lu 0001, Zhiquan Lai, Ke-shi Ge, Dongsheng Li 0001, Xicheng Lu
IEEE Trans. Computers3
2025 Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
abstract
Deep learning is experiencing a rise in large-scale models. Training large-scale models is costly, prompting researchers to train large-scale models on commodity servers that more researchers can access. The massive number of parameters necessitates the use of model parallelism training methods. Existing studies focus on training with pipeline model parallelism. However, the tensor model parallelism (TMP) is inevitable when the model size keeps increasing, where frequent data-dependent communication and computation operations significantly reduce the training efficiency. In this paper, we present Oases, an automated TMP method with overlapped communication to accelerate large-scale model training on commodity servers. Oases proposes a fine-grained training operation schedule to maximize overlapping communication and computation that have data dependence. Additionally, we design the Oases planner that searches for the best model parameter partition strategy of TMP to achieve further accelerations. Unlike existing methods, Oases planner is tailored to model the cost of overlapped communication-computation operations. We evaluate Oases on various model settings and two commodity clusters, and compare Oases to four state-of-the-art implementations. Experimental results show that Oases achieves speedups of 1.01–1.48 × over the fastest baseline, and speedups of up to 1.95 × over Megatron.
Zhiquan Lai, Dongsheng Li 0001, Yanqi Hao, Ke-shi Ge, Xiaoge Deng, Kai Lu 0001
IEEE Trans. Parallel Distributed Syst.2
2024 Hierarchical Adaptive Pooling by Capturing High-order Dependency for Graph Representation Learning (Extended Abstract)
abstract
Graph pooling technique in GNNs for learning expressive graph-level representation is critical yet still chal-lenging. Existing pooling methods either struggle to capture local substructures or fail to utilize high-order dependency, thus diminishing the expression capability. To solve this problem, we propose HAP, a hierarchical graph-level representation learning framework adaptively sensitive to graph structures. Specifically, HAP utilizes a novel cross-level attention mechanism MOA to naturally focus more on the close neighborhood while effectively capturing higher-order dependency. It also learns a global graph content GCont that extracts the graph pattern properties to stabilize the pre- and post-coarsening graph content, thus providing global guidance in graph coarsening. Experiments show that HAP significantly outperforms the state-of-the-art graph pooling methods.
Ning Liu 0015, Songlei Jian, Dongsheng Li 0001, Yiming Zhang 0003, Zhiquan Lai, Hongzuo Xu
ICDE5
2024 The Self-adaptive and Topology-aware MPI_Bcast leveraging Collective offload on Tianhe Express Interconnect
abstract
Large parallel applications have heavily used MPI (Massage Passing Interface) collectives that support portable and efficient group communication operations. MPI_Bcast is one of the most commonly used MPI collectives that broadcast data to all processes of the communication domain. However, traditional software-based broadcast algorithms fail to fully utilize modern interconnection networks’ advanced features such as offloading collectives to the network hardware for efficient group communications. Besides, the semantic gap between MPI_Bcast and hardware multicast of underlying interconnects presents challenges for offload-based algorithms to accelerate MPI_Bcast for a wide range of message sizes.In this paper, we propose a hardware-software co-design MPI_Bcast by efficiently leveraging the NIC-based collective offload provided by Tianhe-express interconnect, which completely precludes the involvement of CPU to accelerate message broadcast. We detail this broadcast mechanism that can be adaptively tuned to offload MPI_Bcast operations from the CPU to the NIC for various message and system sizes. In addition, we further propose a topology-aware broadcast design in conjunction with this offload method to significantly reduce the broadcast latency by constructing the optimal global inter-node communication tree. We implement and evaluate the proposed Tianhe-Express Offload-based Broadcast (TOB) design on Tianhe-2A and Tianhe-EP supercomputers. Extensive experiments have been conducted to evaluate TOB performance at both microbenchmark and application levels. Our solution offers up to 4.94x significant performance speedup at the microbenchmark level over state-of-the-art MPI libraries. For the application-level evaluation, our technique accelerates scientific applications by a maximum speedup of 1.34x.
Chongshan Liang, Jinbo Xu, Jintao Peng, Weixia Xu 0001, Jie Liu 0002, Zhiquan Lai, Sheng Ma
IPDPS9
2024 HSDP: Accelerating Large-scale Model Training via Efficient Sharded Data Parallelism
abstract
Large deep neural network (DNN) models have demonstrated exceptional performance across diverse downstream tasks. Sharded data parallelism (SDP) has been widely used to reduce the memory footprint of model states. In a DNN training cluster, a device usually has multiple inter-device links that connect to other devices, like NVLink and InfiniBand. However, existing SDP approaches employ a single link at any given time, encountering challenges in efficient training due to significant communication overheads. We observe that the inter-device links can work independently without affecting each other. To reduce the fatal communication overhead of distributed training of large DNNs, this paper introduces HSDP, an efficient SDP training approach that enables the simultaneous utilization of multiple inter-device links. HSDP partitions models in a novel fine-grained manner and orchestrates the communication processes of partitioned parameters while considering inter-device links. This design enables concurrent communication execution and reduces communication overhead. To further optimize the training performance of HSDP, we propose a HSDP planner. The HSDP planner first abstracts the model partition and execution of HSDP into a communication parallel strategy, and builds a cost model to estimate the performance of each strategy. We then formulate the strategy searching as an optimization problem and solve it with an off-the-shelf solver. Evaluations on representative DNN workloads demonstrate that HSDP achieves up to 1.30× speedup compared to the state-of-the-art SDP training approaches.
Yanqi Hao, Zhiquan Lai, Ke-shi Ge, Dongsheng Li 0001
ISPA2
2024 A Memory-Efficient Hybrid Parallel Framework for Deep Neural Network Training
abstract
With the increasing volumes of data samples and deep neural network (DNN) models, efficiently scaling the training of DNN models has become a significant challenge for server clusters with AI accelerators in terms of memory and computing efficiency. Existing parallelism schemes can be broadly classified into three categories: data parallelism (splitting data samples), model parallelism (splitting model parameters), and pipeline model parallelism (splitting model layers). Hybrid approaches split data and models, offering a comprehensive solution for parallel training. However, these methods encounter limitations in efficiently scaling larger models across more computing nodes, as they incur substantial memory constraints that affect training efficiency and overall throughput. In this paper, we proposeHIPPIE, a hybrid parallel training framework designed to enhance memory efficiency and scalability of large DNN training. First, to evaluate the optimization effect more reasonably, we propose an index ofMemory Efficiency(ME) to quantify the tradeoff between throughput and memory overhead. Second, driven by the informed ME optimization objective, we automatically partition the pipeline to balance the throughput and memory. Third, we optimize the model training process via a novel hybrid parallel scheduler that improves the throughput and scalability by informed pipeline scheduling and communication scheduling with gradient-hidden optimization. Experiments on various models show thatHIPPIEachieves above 90% scaling efficiency on a 16-GPU platform. Moreover,HIPPIEincreases throughput by up to 80%, while saving 57% of memory overhead and achieving 4.18× memory-efficiency improvement.
Dongsheng Li 0001, Zhiquan Lai, Yongquan Fu, Xiangyu Ye, Linbo Qiao
IEEE Trans. Parallel Distributed Syst.3
2024 A Multidimensional Communication Scheduling Method for Hybrid Parallel DNN Training
abstract
The transformer-based deep neural network (DNN) models have shown considerable success across diverse tasks, prompting widespread adoption of distributed training methods such as data parallelism and pipeline parallelism. With the increasing parameter number, hybrid parallel training becomes imperative to scale training. The primary bottleneck in scaling remains the communication overhead. The communication scheduling technique, emphasizing the overlap of communication with computation, has demonstrated its benefits in scaling. However, most existing works focus on data parallelism, overlooking the nuances of hybrid parallel training. In this paper, we proposeTriRace, an efficient communication scheduling framework for accelerating communications in hybrid parallel training of asynchronous pipeline parallelism and data parallelism. To achieve effective computation-communication overlap,TriRaceintroduces3D communication scheduling, which adeptly leverages data dependencies between communication and computations, efficiently scheduling AllReduce communication, sparse communication, and peer-to-peer communication in hybrid parallel training. To avoid possible communication contentions,TriRacealso incorporates atopology-aware runtimewhich optimizes the execution of communication operations by considering ongoing communication operations and real-time network status. We have implemented a prototype ofTriRacebased on PyTorch and Pipedream-2BW, and conducted comprehensive evaluations with three representative baselines. Experimental results show thatTriRaceachieves up to 1.07–1.45× speedup compared to the state-of-the-art pipeline parallelism training baseline Pipedream-2BW, and 1.24–1.81× speedup compared to the Megatron.
Kai Lu 0001, Zhiquan Lai, Ke-shi Ge, Dongsheng Li 0001
IEEE Trans. Parallel Distributed Syst.3
2023 Prophet: Fine-grained Load Balancing for Parallel Training of Large-scale MoE Models
abstract
Mixture of Expert (MoE) has received increasing attention for scaling DNN models to extra-large size with negligible increases in computation. The MoE model has achieved the highest accuracy in several domains. However, a significant load imbalance occurs in the device during the training of a MoE model, resulting in significantly reduced throughput. Previous works on load balancing either harm model convergence or suffer from high execution overhead. To address these issues, we present Prophet: a fine-grained load balancing method for parallel training of large-scale MoE models, which consists of a planner and a scheduler. Prophet planner first employs a fine-grained resource allocation method to determine the possible scenarios for the expert placement in a fine-grained manner, and then efficiently searches for a well-balanced expert placement to balance the load without introducing additional overhead. Prophet scheduler exploits the locality of the token distribution to schedule the resource allocation operations using a layer-wise fine-grained schedule strategy to hide their overhead. We conduct extensive experiments in four clusters and five representative models. The results indicate that Prophet gains up to 2.3x speedup compared to the state-of-the-art MoE frameworks including Deepspeed-MoE and FasterMoE. Additionally, Prophet achieves a load balancing enhancement of up to 12.06x when compared to FasterMoE.
Zhiquan Lai, Ke-shi Ge, Dongsheng Li 0001
CLUSTER2
2023 Auto-Divide GNN: Accelerating GNN Training with Subgraph Division
Zhejiang Ran, Ke-shi Ge, Zhiquan Lai, Jingfei Jiang, Dongsheng Li 0001
Euro-Par4
2023 Compressed Collective Sparse-Sketch for Distributed Data-Parallel Training of Deep Learning Models
abstract
Distributed data-parallel training (DDP) is prevalent in large-scale deep learning. To increase the training throughput and scalability, high-performance collective communication methods such as AllReduce have recently proliferated for DDP use. However, these approaches require long communication periods with increasing model sizes. Collective communication transmits many sparse gradient values that can be efficiently compressed to reduce the required training time. State-of-the-art compression approaches do not provide mergeable compression for AllReduce and lack convergence bounds. We present a sparse sketch reducer (S2Reducer), a sparsity-preserving sketch-based collective communication method. S2Reducer preserves gradient sparsity and reduces communication costs via a bitmap informed count sketch structure and adapts to efficient AllReduce operators. We tune the count sketch organization to minimize the hash conflicts in a fixed-size budget. We prove that our method has the same convergence rate as vanilla data-parallel training and a much smaller communication overhead than those of state-of-the-art methods. We implement a GPU-accelerated S2Reducer for the Ring AllReduce-based DDP system. We perform extensive evaluations against four state-of-the-art methods across seven deep learning models. Our results show that S2Reducer converges to the same accuracy as that of state-of-the-art approaches while reducing the sparse communication overhead by up to 86% and achieving a speedup of up to$3.5\times $in distributed training.
Ke-shi Ge, Kai Lu 0001, Yongquan Fu, Xiaoge Deng, Zhiquan Lai, Dongsheng Li 0001
IEEE J. Sel. Areas Commun.5
2023 Accelerating GNN Training by Adapting Large Graphs to Distributed Heterogeneous Architectures
abstract
Graph neural networks (GNNs) have been successfully applied to many important application domains on graph data. As graphs become increasingly large, existing GNN training frameworks typically use mini-batch sampling during feature aggregation to lower resource burdens, which unfortunately suffer from long memory accessing latency and inefficient data transfer of vertex features from CPU to GPU. This paper proposes 2PGraph, a system that addresses these limitations of mini-batch sampling and feature aggregation and supports fast and efficient single-GPU and distributed GNN training. First, 2PGraph presents a locality awareness GNN-training scheduling method that schedules the vertices based on the locality of the graph topology, significantly accelerating the sampling and aggregation, improving the data locality of vertex access, and limiting the range of neighborhood expansion. Second, 2PGraph proposes a GNN-layer-aware feature caching method on available GPU resources with a hit rate up to 100${\bf\%}$, which avoids redundant data transfer between CPU and GPU. Third, 2PGraph presents a self-dependence cluster-based graph partition method, achieving high sampling and cache efficiency for distributed environments. Experimental results on real-world graph datasets show that 2PGraph reduces memory access latency by up to 90${\boldsymbol{\%}}$mini-batch sampling, and data transfer time by up to 99${\boldsymbol{\%}}$. For distributed GNN training over an 8-GPU cluster, 2PGraph achieves up to 8.7$\times$performance speedup over state-of-the-art approaches.
Kai Lu 0001, Zhiquan Lai, Yongquan Fu, Dongsheng Li 0001
IEEE Trans. Computers3
2023 Hierarchical Adaptive Pooling by Capturing High-Order Dependency for Graph Representation Learning
abstract
Graph neural networks (GNN) have been proven to be mature enough for handling graph-structured data on node-level graph representation learning tasks. However, the graph pooling technique for learning expressive graph-level representation is critical yet still challenging. Existing pooling methods either struggle to capture the local substructure or fail to effectively utilize high-order dependency, thus diminishing the expression capability. In this paper we propose HAP, a hierarchical graph-level representation learning framework, which is adaptively sensitive to graph structures, i.e., HAP clusters local substructures incorporating with high-order dependencies. HAP utilizes a novel cross-level attention mechanism MOA to naturally focus more on close neighborhood while effectively capture higher-order dependency that may contain crucial information. It also learns a global graph content GCont that extracts the graph pattern properties to make the pre- and post-coarsening graph content maintain stable, thus providing global guidance in graph coarsening. This novel innovation also facilitates generalization across graphs with the same form of features. Extensive experiments on ten datasets show that HAP significantly outperforms twelve popular graph pooling methods on graph classification task with an maximum accuracy improvement of 20.18%, and exceeds the performance of state-of-the-art graph matching and graph similarity learning algorithms by over 3.42% and 16%.
Ning Liu 0015, Songlei Jian, Dongsheng Li 0001, Yiming Zhang 0003, Zhiquan Lai, Hongzuo Xu
IEEE Trans. Knowl. Data Eng.5
2023 Merak: An Efficient Distributed DNN Training Framework With Automated 3D Parallelism for Giant Foundation Models
abstract
Foundation models are in the process of becoming the dominant deep learning technology. Pretraining a foundation model is always time-consuming due to the large scale of both the model parameter and training dataset. Besides being computing-intensive, the pretraining process is extremely memory- and communication-intensive. These challenges make it necessary to apply 3D parallelism, which integrates data parallelism, pipeline model parallelism, and tensor model parallelism, to achieve high training efficiency. However, current 3D parallelism frameworks still encounter two issues: i) they are not transparent to model developers, requiring manual model modification to parallelize training, and ii) their utilization of computation resources, GPU memory, and network bandwidth is insufficient. We proposeMerak, an automated 3D parallelism deep learning training framework with high resource utilization. Merak automatically deploys 3D parallelism with an automatic model partitioner, which includes a graph-sharding algorithm and proxy node-based model graph. Merak also offers a non-intrusive API to scale out foundation model training with minimal code modification. In addition, we design a high-performance 3D parallel runtime engine that employs several techniques to exploit available training resources, including a shifted critical path pipeline schedule that increases computation utilization, stage-aware recomputation that makes use of idle worker memory, and sub-pipelined tensor model parallelism that overlaps communication and computation. Experiments on 64 GPUs demonstrate Merak's capability to speed up training performance over state-of-the-art 3D parallelism frameworks of models with 1.5, 2.5, 8.3, and 20 billion parameters by up to 1.42, 1.39, 1.43, and 1.61×, respectively.
Zhiquan Lai, Xudong Tang, Ke-shi Ge, Yabo Duan, Linbo Qiao, Dongsheng Li 0001
IEEE Trans. Parallel Distributed Syst.1
2023 A Survey on Auto-Parallelism of Large-Scale Deep Learning Training
abstract
Deep learning (DL) has gained great success in recent years, leading to state-of-the-art performance in research community and industrial fields like computer vision and natural language processing. One of the reasons for this success is the huge amount parameters adopted in DL models. However, it is impractical to train a moderately large model with a large number of parameters on a typical single device. Thus, It is necessary to train DL models in clusters with distributed training algorithms. However, traditional distributed training algorithms are usually sub-optimal and highly customized, which owns the drawbacks to train large-scale DL models in varying computing clusters. To handle the above problem, researchers propose auto-parallelism, which is promising to train large-scale DL models efficiently and practically in various computing clusters. In this survey, we perform a broad and thorough investigation on challenges, basis, and strategy searching methods of auto-parallelism in DL training. First, we abstract basic parallelism schemes with their communication cost and memory consumption in DL training. Further, we analyze and compare a series of current auto-parallelism works and investigate strategies and searching methods which are commonly used in practice. At last, we discuss several trends in auto-parallelism which are promising in further research.
Peng Liang 0017, Xiaoda Zhang, Youhui Bai, Teng Su, Zhiquan Lai, Linbo Qiao, Dongsheng Li 0001
IEEE Trans. Parallel Distributed Syst.6
2022 HPH: Hybrid Parallelism on Heterogeneous Clusters for Accelerating Large-scale DNNs Training
abstract
As the deep learning model grows larger, training model with a single computational resource becomes impractical. To solve this, hybrid parallelism, which combines data and pipeline parallelism emerges to train large models with multiple GPUs. In practice, using heterogeneous GPU clusters to train large models is a common need due to the upgrade of a part of hardware. However, existing hybrid parallelism approaches in the heterogeneous environment do not work well in communication efficacy, workload balance among GPUs and utilizing the memory constrained GPU. To address these problems, we present a parallel DNN training approach, Hybrid Parallelism on Heterogeneous clusters (HPH). In HPH, we propose a topology designer that minimizes the communication time cost. Furthermore, HPH uses a partition algorithm that automatically partitions DNN layers among workers to maximize throughput. Besides, HPH adopts recomputation-aware scheduling to reduce memory consumption and further reschedule the pipeline to eliminate the extra time overhead of recomputation. Our experimental results on a 32-GPU heterogeneous cluster show that HPH achieves up to 1.42x training speed-ups compared with the state-of-the-art approach.
Yabo Duan, Zhiquan Lai, Ke-shi Ge, Peng Liang 0017, Dongsheng Li 0001
CLUSTER2
2022 AutoPipe: A Fast Pipeline Parallelism Approach with Balanced Partitioning and Micro-batch Slicing
abstract
Recently, pipeline parallelism has been widely used in training large DNN models. However, there are still two main challenges for efficient pipeline parallelism: i) a balanced model partition is crucial for pipeline efficiency, whereas prior works lack a sound solution to generate a balanced partition automatically. ii) the startup overhead is inevitable and especially significant for deep pipelines, which is an essential source of pipeline bubbles and severely affects pipeline scalability. We propose AutoPipe to solve these two problems, which contains i) a planner for automatically and quickly generating a balanced pipeline partition scheme with a fine-grained partitioner. This partitioner groups DNN in the sub-layer granularity and finds the balanced scheme with a heuristic search algorithm; and ii) a micro-batch slicer that reduces pipeline startup overhead according to the planner results by splitting the micro-batch evenly. This slicer automatically solves an appropriate number of micro-batches to split. The experimental results show that AutoPipe can accelerate training by up to 1.30x over the state-of-the-art distributed training framework Megatron-LM, with a 50% reduction in startup overhead and an order-of-magnitude reduction in pipeline planning time. Furthermore, AutoPipe Planner improves the partition balance by 2.73x-12.7x compared to DAPPLE Planner and Piper.
Zhiquan Lai, Yabo Duan, Ke-shi Ge, Dongsheng Li 0001
CLUSTER2
2022 S2 Reducer: High-Performance Sparse Communication to Accelerate Distributed Deep Learning
abstract
Distributed stochastic gradient descent (SGD) approach has been widely used in large-scale deep learning, and the gradient collective method is vital to ensure the training scalability of the distributed deep learning system. Collective communication such as AllReduce has been widely adopted for the distributed SGD process to reduce the communication time. However, AllReduce incurs large bandwidth resources while most gradients are sparse in many cases since many gradient values are zeros and should be efficiently compressed for bandwidth saving. To reduce the sparse gradient communication overhead, we propose Sparse-Sketch Reducer (S2 Reducer), a novel sketch-based sparse gradient aggregation method with convergence guarantees. S2 Reducer reduces the communication cost by only compressing the non-zero gradients with count-sketch and bitmap, and enables the efficient AllReduce operators for parallel SGD training. We perform extensive evaluation against four state-of-the-art methods over five training models. Our results show that S2 Reducer converges to the same accuracy, reduces 81% sparse communication overhead, and achieves 1.8× distributed training speedup compared to state-of-the-art approaches.
Ke-shi Ge, Yongquan Fu, Yiming Zhang 0003, Zhiquan Lai, Xiaoge Deng, Dongsheng Li 0001
ICASSP4
2022 EmbRace: Accelerating Sparse Communication for Distributed Training of Deep Neural Networks
abstract
Distributed data-parallel training has been widely adopted for deep neural network (DNN) models. Although current deep learning (DL) frameworks scale well for dense models like image classification models, we find that these DL frameworks have relatively low scalability for sparse models like natural language processing (NLP) models that have highly sparse embedding tables. Most existing works overlook the sparsity of model parameters thus suffering from significant but unnecessary communication overhead. In this paper, we propose EmbRace, an efficient communication framework to accelerate communications of distributed training for sparse models. EmbRace introduces Sparsity-aware Hybrid Communication, which integrates AlltoAll and model parallelism into data-parallel training, so as to reduce the communication overhead of highly sparse parameters. To effectively overlap sparse communication with both backward and forward computation, EmbRace further designs a 2D Communication Scheduling approach which optimizes the model computation procedure, relaxes the dependency of embeddings, and schedules the sparse communications of each embedding row with a priority queue. We have implemented a prototype of EmbRace based on PyTorch and Horovod, and conducted comprehensive evaluations with four representative NLP models. Experimental results show that EmbRace achieves up to 2.41 × speedup compared to the state-of-the-art distributed training baselines.
Zhiquan Lai, Dongsheng Li 0001, Yiming Zhang 0003, Xiangyu Ye, Yabo Duan
ICPP2
2022 BRGraph: An efficient graph neural network training system by reusing batch data on GPU
abstract
Summary With the increasing adoption of graph neural networks (GNNs) in the community, various GPU‐based graph programming systems have been developed to improve the productivity of GNNs. However, sampling‐based GNN training is still inefficient, and we observe that the main bottleneck comes from the data transferring, where vertex features are transferred from host memory to GPU through limited bandwidth. In this article, we propose BRGraph, a sampling‐based GNN training system that supports efficient data transferring. BRGraph leverages the duplicate vertices between mini‐batches by batch reusing (BR) strategy to avoid duplicate data transmission. Furthermore, to reduce the overhead of detecting duplicate vertices, we design an efficient parallel batch reusing algorithm based on GPU. BRGraph also exploits the data reusing potential of the non‐duplicate vertex features by the two‐level batch reusing (two‐level BR) strategy. Comprehensive evaluations on three representative GNN models show that BRGraph reduces data transferring time by up to 60% and delivers up to 1.79 GNN training speedup over the state‐of‐the‐art baselines. Besides, it can save GPU memory by up to 40% while reaching the same training time compared with the static cache strategy. When applying the two‐level BR, BRGraph further reduces 20% of the data transferring time compared with the BR.
Ke-shi Ge, Zhejiang Ran, Zhiquan Lai, Dongsheng Li 0001
Concurr. Comput. Pract. Exp.3
2021 CASQ: Accelerate Distributed Deep Learning with Sketch-Based Gradient Quantization
abstract
Gradient quantization has been widely used in distributed training of deep neural network (DNN) models to reduce communication costs. However, existing quantization methods overlook that gradients have a nonuniform distribution changing over time, which can lead to significant gradient variance that requires a higher number of quantization bits (and consequently higher communication cost) to keep the validation accuracy as high as stochastic gradient descent (SGD). In this paper, we propose Cluster-Aware Sketch Quantization (CASQ), a novel sketch-based gradient quantization method for SGD. CASQ models the nonuniform distribution of gradients via clustering, and adaptively allocates appropriate numbers of hash buckets based on the statistics of different clusters to compress gradients. The extensive evaluation shows that compared to existing quantization methods CASQ-based SGD (i) achieves the same validation accuracy when decreasing quantization level from 3 bits to 2 bits, and (ii) reduces the training time to convergence by up to 43% for the same training loss.
Ke-shi Ge, Yiming Zhang 0003, Yongquan Fu, Zhiquan Lai, Xiaoge Deng, Dongsheng Li 0001
CLUSTER4
2021 2PGraph: Accelerating GNN Training over Large Graphs on GPU Clusters
abstract
Graph neural networks (GNNs) have been emerging as powerful learning tools for unstructured data and successfully applied to many graph-based application domains. Sampling-based graph training is commonly used in existing GNN training frameworks to handle large-scale graphs. However, this type of approach is restricted by the problems of long memory accessing latency, neighborhood explosion during mini-batch sampling, and inefficient loading of vertex features from CPU to GPU. In this paper, we propose 2PGraph, a system that supports high-speed locality-aware mini-batch sampling and GNN layer-aware feature caching. 2PGraph significantly reduces sampling time by vertex-cluster sampling, which improves the locality of vertex access and limits the range of neighborhood expansion. To further reduce the sampling time in a distributed environment, we renumber the vertex numbers in subgraphs after graph partition, which improves the data locality of each partition. 2PGraph also avoids abundant data transfer between CPU and GPU through the feature data caching on available GPU resources with a hit rate of 100%. Furthermore, 2PGraph develops a GNN layer-aware feature caching policy during data parallel training and achieves better cache efficiency and memory utilization. We evaluate 2PGraph against two state-of-the-art industrial GNN frameworks, i.e., PyG and DGL, on a diverse array of benchmarks. Experimental results show that 2PGraph reduces up to 90% mini-batch sampling and 99% data loading time, and achieves up to 8.7 × performance speedup over the state-of-the-art baselines on an 8-GPU cluster.
Zhiquan Lai, Dongsheng Li 0001
CLUSTER2
2021 Hippie: A Data-Paralleled Pipeline Approach to Improve Memory-Efficiency and Scalability for Large DNN Training
abstract
With the increase of both data and parameter volume, it has become a big challenge to efficiently train large-scale DNN models on distributed platforms. Ordinary parallelism modes, i.e., data parallelism, model parallelism and pipeline parallelism, can no longer satisfy the efficient scaling of large DNN model training on multiple nodes. Meanwhile, the problem of too much memory consumption seriously restricts GPU computing efficiency and training throughput. In this paper, we propose Hippie, a hybrid parallel training framework that integrates pipeline parallelism and data parallelism to improve the memory efficiency and scalability of large DNN training. Hippie adopts a hybrid parallel method based on hiding gradient communication, which improves the throughput and scalability of training. Meanwhile, Hippie introduces the last-stage pipeline scheduling and recomputation for specific layers to effectively reduce the memory overhead and ease the difficulties of training large DNN models on memory-constrained devices. To achieve a more reasonable evaluation of the optimization effect, we propose an index of memory efficiency (ME) to represent the tradeoff between throughput and memory overhead. We implement Hippie based on PyTorch and NCCL. Experiments on various models show that Hippie achieves above 90% scaling efficiency on a 16-GPU platform. Moreover, Hippie increases throughput by up to 80% while saving 57% of memory overhead, achieving 4.18 × memory efficiency.
Xiangyu Ye, Zhiquan Lai, Ding Sun, Linbo Qiao, Dongsheng Li 0001
ICPP2
2021 Accelerate Graph Neural Network Training by Reusing Batch Data on GPUs
abstract
With the increasing adoption of graph neural networks (GNNs) in the graph-based deep learning community, various graph programming frameworks and models have been developed to improve the productivity of GNNs. The current GNN frameworks choose GPU as an essential tool to accelerate GNN training. However, it is still challenging to train GNNs on large graphs with limited GPU memory. Unlike traditional neural networks, generating mini-batch data by sampling in GNNs requires some complicated tasks such as traversing the graph to select neighboring nodes and gathering their features. This process takes up most of the training and we find the main bottleneck comes from transferring nodes features from CPU to GPU through limited bandwidth. In this paper, We propose a method Reusing Batch Data for the problem of data transmission. This method utilizes the similarity between adjacent mini-batches to reduce repeated data transmission from CPU to GPU. Furthermore, to reduce the overhead introduced by this method, we design a fast algorithm based on GPU to detect repeated nodes’ data and achieve shorter additional computation time. Evaluations on three representative GNN models show that our method can reduce transmission time by up to 60% and speed the end-to-end GNN training by up to 1.79× over the state-ofthe-art baselines. Besides, Reusing Batch Data can effectively save GPU memory footprint by about 19% to 40% while still reducing the training time compared to the static cache strategy.
Zhejiang Ran, Zhiquan Lai, Dongsheng Li 0001
IPCCC2
2021 Coordinative Scheduling of Computation and Communication in Data-Parallel Systems
abstract
For many data-parallel computing systems like Spark, a job usually consists of multiple computation stages and inter-stage communication (i.e., coflows). Many efforts have been done to schedule coflows and jobs independently. The simple combination of coflow scheduling and job scheduling, however, would prolong the average job completion time (JCT) due to the conflict. For this reason, we propose a new abstraction of scheduling unit, named coBranch, which takes the dependency between computation stages and coflows into consideration, to schedule coflows and jobs jointly. Besides, mainstream coflow schedulers are order-preserving, i.e., all coflows of a high-priority job are prioritized than those of a low-priority job. We observe that the order-preserving constraint incurs low inter-job parallelism. To overcome the problem, we employ an urgency-based mechanism to schedule coBranches, which aims to decrease the average JCT by enhancing the inter-job parallelism. We implement the urgency-based coBranch Scheduling (BS) method on Apache Spark, conduct prototype-based experiments, and evaluate the performance of our method against the shortest-job-first critical-path method and the FIFO method. Results show that our method achieves around 10 and 15 percent reduction in the average JCT, respectively. Large-scale simulations based on the Google trace show that our method performs better and reduces JCT by 23 and 35 percent, respectively.
Dongsheng Li 0001, Zhiyao Hu, Zhiquan Lai, Yiming Zhang 0003, Kai Lu 0001
IEEE Trans. Computers3
2020 ADMMiRNN: Training RNN with Stable Convergence via an Efficient ADMM Approach
Zhigang Kan, Dequan Sun, Linbo Qiao, Zhiquan Lai, Dongsheng Li 0001
ECML/PKDD (2)6
2019 HPDL: Towards a General Framework for High-performance Distributed Deep Learning
abstract
With growing scale of the data volume and neural network size, we have come into the era of distributed deep learning. High-performance training and inference on distributed computing systems has been attracting increasing research attention in both academia and industry. Meanwhile, diversity of existing machine learning frameworks (e.g. TensorFlow, Pytorch and MXNet) and the explosion of deep learning hardwares (e.g. CPUs, GPUs, FPGAs and ASICs) bring more challenges for users to leverage new deep learning technologies and accelerating capability of hardware devices. We firstly search around the state-of-the-art work in the area which open our mind to take a vision upon the future deep learning framework. Then, we propose HPDL, a general framework for high-performance distributed deep learning which is compatible with existing frameworks and adaptive to various hardware architectures. At last, we discuss and foresee the key technologies fulfilling high-performance and large-scale deep learning, including optimization algorithm, hybrid communication mechanism, model parallelization, resource scheduling and single-node execution optimization.
Dongsheng Li 0001, Zhiquan Lai, Ke-shi Ge, Yiming Zhang 0003, Zhaoning Zhang 0001, Huaimin Wang 0001
ICDCS2
2015 Latency-aware DVFS for efficient power state transitions on many-core architectures
Zhiquan Lai, King Tin Lam, Cho-Li Wang, Jinshu Su
J. Supercomput.1
2014 Rhymes: A shared virtual memory system for non-coherent tiled many-core architectures
abstract
The rising core count per processor is pushing chip complexity to a level that hardware-based cache coherency protocols become too hard and costly to scale someday. We need new designs of many-core hardware and software other than traditional technologies to keep up with the ever-increasing scalability demands. A cluster-on-chip architecture, as exemplified by the Intel Single-chip Cloud Computer (SCC), promotes a software-oriented approach instead of hardware support to implementing shared memory coherence. This paper presents a shared virtual memory (SVM) system, dubbed Rhymes, tailored to new processor kinds of non-coherent and hybrid memory architectures. Rhymes features a two-way cache coherence protocol to enforce release consistency for pages allocated in shared physical memory (SPM) and scope consistency for pages in percore private memory. It also supports page remapping on a percore basis to boost data locality. We implement and test Rhymes on the SCC port of the Barrelfish OS. Experimental results show that our SVM outperforms the pure SPM approach used by Intel's software managed coherence (SMC) library by up to 12 times through improved cache utilization for applications with strong data reuse patterns.
King Tin Lam, Jinghao Shi, Dominic Hung, Cho-Li Wang, Zhiquan Lai, Wangbin Zhu, Youliang Yan
ICPADS5