Zeyu Ji

dblp:292/5041 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0003-0362-7506ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FedCED: Consensus enhancement in decentralized federated learning via distillation
Bin Liu 0023, Changle Li, Qi Chu 0013, Banghao Zhai, Zeyu Ji, Keqin Li 0001
Knowl. Based Syst.5
2026 LIBPipe: Efficient Load Imbalance Pipeline Model Parallelism for Large Models Training
abstract
With the increasing size of datasets and the expansion of Deep Neural Networks (DNNs), the training process has become exceedingly time-consuming. Distributed training, specifically the Pipeline Model Parallelism (PMP) method, commonly mitigates this problem but suffers from bubble time delays. This paper proposes LIBPipe, a pipeline training framework that explicitly incorporates a load-imbalance method to reduce bubble time in PMP. Within LIBPipe, a model-unequal-partitioning method is designed from the perspective of load imbalance to reshape the pipeline execution pattern and significantly shorten idle periods during training. On top of this method, a performance-guided unequal-partitioning search algorithm is developed to efficiently identify near-optimal partitioning strategies under memory constraints. The paper theoretically proves the time efficiency of adopting load imbalance in pipeline models. Comprehensive experiments are conducted on an 8-GPU server to evaluate the efficiency of LIBPipe, using the IMDB and mini-ImageNet datasets as well as well-known models such as BERT and ResNet. The BERT-series models achieve a maximum throughput improvement of 60.3%, while the ResNet-series models achieve a maximum improvement of 74.1%.
Bin Liu 0023, Hengzhao Li, Zeyu Ji, Hongming Zhang 0002, Keqin Li 0001
IEEE Trans. Parallel Distributed Syst.4
2025 PRT: An Efficient Pipeline Reuse Technology for Large Models Training
abstract
The rapid evolution of large models and the widespread application of extensive datasets have made the cost of training increasingly prohibitive. While pipeline model parallelism makes it possible to train large models, existing pipeline techniques find it difficult to reduce bubble time due to their strong dependence on the number of GPUs for pipeline depth. This paper introduces a novel pipeline reuse technology, PRT, which breaks the limitation of pipeline depth being dependent on the number of GPUs, allowing for deeper pipelines even when the number of GPUs is limited. This paper also theoretically demonstrates the feasibility of PRT. Furthermore, the high orthogonality of PRT allows it to be implemented in both unidirectional and bidirectional pipelines, further enhancing pipeline efficiency. It is evaluated on a server equipped with 8 GPUs, using the BERT series models and ResNet series models with datasets including the IMDB dataset and the mini-ImageNet dataset. Experimental results show that for the BERT series models, unidirectional and bidirectional pipelines with PRT achieve throughput improvements of up to 54.78% and 30.38%, respectively. For the ResNet series models, the improvements reached up to 76.59% and 26.45%, respectively. Additionally, PRT achieves more balanced memory usage, validating its efficiency.
Zeyu Ji, Banghao Zhai, Qi Chu 0013, Bin Liu 0023
CLUSTER1
2025 GroPipe: A Grouped Pipeline Hybrid Parallel Method for Accelerating DCNNs Training
abstract
Training large Deep Convolutional Neural Networks (DCNNs) with increasingly large datasets to improve model accuracy has become extremely time-consuming. Distributed training methods, such as data parallelism (DP) and pipeline model parallelism (PMP), offer potential solutions but face challenges like load imbalance and significant communication overhead. This paper introduces GroPipe, a novel architecture that synergistically integrates PMP and DP, markedly improving training speeds. GroPipe employs an automatic model partitioning algorithm based on a performance projection technique, ensuring load balance and facilitating quantitative performance evaluation in PMP. Additionally, it adopts a group-based delayed asynchronous communication strategy to efficiently reduce communication overhead in DP. Using the ResNet and VGG models with the ImageNet dataset, extensive experiments are performed on an 8-GPU server and demonstrate GroPipe’s effectiveness. GroPipe achieves substantial improvements in time to accuracy, showing an average improvement of 42.2% and 14.0% on the ResNet series, and 79.2% and 43.9% on the VGG series, without compromising Top-1 accuracy.
Bin Liu 0023, Yongyao Ma, Zeyu Ji, Zhenli He, Keqin Li 0001
IEEE Trans. Computers4
2024 Revisit and Benchmarking of Automated Quantization Toward Fair Comparison
abstract
Automated quantization has emerged as an entirely new design paradigm to automate the optimal configuration of bitwidth for deep neural networks (DNNs), making the DNN more memory-efficient and faster to execute on hardware with limited resources. Reinforcement learning (RL) and differentiable neural architecture search (DNAS) are two main solution paths that have shown their superiority. Yet, there are countless methods with various implementations within each path. It has been hard to comprehend their differences and make a relatively fair comparison due to the lack of a benchmark framework and a clear analysis of which aspects are common, respectively distinct, between different implementations. To this end, we introduce BenQ to pave the way towards fair comparisons in two separate race tracks, i.e., intra-comparison of the RL-based and the DNAS-based methods, respectively. We provide a systematic approach, which helps to reveal relatively vital aspects of different implementations. Finally, we conduct comprehensive experi-ments on VGG, AlexNet, ResNet, GoogleNet, MobileNet-V2, and Vision Transformer (ViT), and the new observations shed light on potential future directions for automated quantization to move forward.
Xingjun Zhang, Zeyu Ji, Jia Wei 0002
IEEE Trans. Computers3
2024 LBB: load-balanced batching for efficient distributed learning on heterogeneous GPU cluster
Feixiang Yao, Zeyu Ji, Bin Liu 0023, Haoyuan Gao
J. Supercomput.3
2023 Leader population learning rate schedule
Jia Wei 0002, Xingjun Zhang, Zhimin Zhuo, Zeyu Ji, Qianyang Li
Inf. Sci.4
2022 BenQ: Benchmarking Automated Quantization on Deep Neural Network Accelerators
abstract
Hardware-aware automated quantization promises to unlock an entirely new algorithm-hardware co-design paradigm for efficiently accelerating deep neural network (DNN) inference by incorporating the hardware cost into the reinforcement learning (RL) -based quantization strategy search process. Existing works usually design an automated quantization algorithm targeting one hardware accelerator with a device-specific performance model or pre-collected data. However, determining the hardware cost is non-trivial for algorithm experts due to their lack of cross-disciplinary knowledge in computer architecture, compiler, and physical chip design. Such a barrier limits reproducibility and fair comparison. Moreover, it is notoriously challenging to interpret the results due to the lack of quantitative metrics. To this end, we first propose BenQ, which includes various RL-based automated quantization algorithms with aligned settings and encapsulates two off-the-shelf performance predictors with standard OpenAI Gym API. Then, we leverage cosine similarity and manhattan distance to interpret the similarity between the searched policies. The experiments show that different automated quantization algorithms can achieve near equivalent optimal trade-offs because of the high similarity between the searched policies, which provides insights for revisiting the innovations in automated quantization algorithms.
Xingjun Zhang, Zeyu Ji, Jia Wei 0002
DATE4
2022 GARLSched: Generative adversarial deep reinforcement learning task scheduling optimization for large-scale high performance computing systems
Xingjun Zhang, Jia Wei 0002, Zeyu Ji
Future Gener. Comput. Syst.4
2022 DPLRS: Distributed Population Learning Rate Schedule
Jia Wei 0002, Xingjun Zhang, Zeyu Ji
Future Gener. Comput. Syst.3
2022 EP4DDL: addressing straggler problem in heterogeneous distributed deep learning
Zeyu Ji, Xingjun Zhang, Jia Wei 0002
J. Supercomput.1
2021 Energy-aware task scheduling optimization with deep reinforcement learning for large-scale heterogeneous systems
Xingjun Zhang, Jia Wei 0002, Zeyu Ji
CCF Trans. High Perform. Comput.5
2021 A tile-fusion method for accelerating Winograd convolutions
Zeyu Ji, Xingjun Zhang, Jia Wei 0002
Neurocomputing1
2021 OKCM: improving parallel task scheduling in high-performance computing systems using online learning
Xingjun Zhang, Zeyu Ji, Xiaoshe Dong, Chenglong Hu
J. Supercomput.4