Yongyao Ma

dblp:409/3177 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
0009-0000-8293-308XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 67% Distributed systems · 33%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
data parallelism
0.912025
GroPipe: A Grouped Pipeline Hybrid Parallel Method for Accelerating DCNNs Training · IEEE Trans. Computers 2025
Distributed systems › distributed machine learning
distributed training
0.912025
GroPipe: A Grouped Pipeline Hybrid Parallel Method for Accelerating DCNNs Training · IEEE Trans. Computers 2025
Parallel and multicore computing
pipeline parallelism
0.912025
GroPipe: A Grouped Pipeline Hybrid Parallel Method for Accelerating DCNNs Training · IEEE Trans. Computers 2025
Machine learning › Deep learning architectures and training › convolutional neural network
convolutional neural network training
0.312025
GroPipe: A Grouped Pipeline Hybrid Parallel Method for Accelerating DCNNs Training · IEEE Trans. Computers 2025

Methods — techniques the papers use, named apart from their topics

performance projection · 1.7model partitioning · 1.7asynchronous communication · 1.7
YearPublicationVenuePosition
2025 GroPipe: A Grouped Pipeline Hybrid Parallel Method for Accelerating DCNNs Training
abstract
Training large Deep Convolutional Neural Networks (DCNNs) with increasingly large datasets to improve model accuracy has become extremely time-consuming. Distributed training methods, such as data parallelism (DP) and pipeline model parallelism (PMP), offer potential solutions but face challenges like load imbalance and significant communication overhead. This paper introduces GroPipe, a novel architecture that synergistically integrates PMP and DP, markedly improving training speeds. GroPipe employs an automatic model partitioning algorithm based on a performance projection technique, ensuring load balance and facilitating quantitative performance evaluation in PMP. Additionally, it adopts a group-based delayed asynchronous communication strategy to efficiently reduce communication overhead in DP. Using the ResNet and VGG models with the ImageNet dataset, extensive experiments are performed on an 8-GPU server and demonstrate GroPipe’s effectiveness. GroPipe achieves substantial improvements in time to accuracy, showing an average improvement of 42.2% and 14.0% on the ResNet series, and 79.2% and 43.9% on the VGG series, without compromising Top-1 accuracy.
Bin Liu 0023, Yongyao Ma, Zeyu Ji, Zhenli He, Keqin Li 0001
IEEE Trans. Computers2