Fan Zhang 0139

dblp:21/3626-139 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
0009-0009-0887-3328ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 61% Memory systems · 30% Performance modeling and evaluation · 9%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory management
DNN training memory management
0.912025
ATP: Achieving Throughput Peak for DNN Training via Smart GPU Memory Management · ACM Trans. Archit. Code Optim. 2025
GPUs and heterogeneous computing
GPU memory
0.912025
ATP: Achieving Throughput Peak for DNN Training via Smart GPU Memory Management · ACM Trans. Archit. Code Optim. 2025
GPUs and heterogeneous computing
GPU memory management
0.912025
ATP: Achieving Throughput Peak for DNN Training via Smart GPU Memory Management · ACM Trans. Archit. Code Optim. 2025
Machine learning › Efficient and distributed learning › large-scale learning
large-scale model training
0.312025
ATP: Achieving Throughput Peak for DNN Training via Smart GPU Memory Management · ACM Trans. Archit. Code Optim. 2025
Performance modeling and evaluation › performance model construction
throughput modeling
0.312025
ATP: Achieving Throughput Peak for DNN Training via Smart GPU Memory Management · ACM Trans. Archit. Code Optim. 2025

Methods — techniques the papers use, named apart from their topics

memory pool optimization · 1.7CUDA stream control · 1.7
YearPublicationVenuePosition
2025 ATP: Achieving Throughput Peak for DNN Training via Smart GPU Memory Management
abstract
Due to the limited GPU memory, the performance of large DNNs training is constrained by the unscalable batch size. Existing studies partially address the issue of GPU memory limit through tensor recomputation and swapping, but overlook the exploration of optimal performance. In response, we propose ATP, a recomputation and swapping based GPU memory management framework that aims to maximize training performance by breaking GPU memory constraints. ATP utilizes a throughput model and we propose to evaluate the theoretical peak performance achievable by DNN training on GPU, and provide the optimum memory size required for recomputation and swapping. We optimize the mechanisms for GPU memory pool and CUDA stream control, employ an optimization method to search for specific tensors requiring recomputation and swapping, thereby bringing the actual DNN training performance on ATP closer to theoretical values. Evaluations with different types of large DNN models indicate that ATP achieve throughput improvements ranging from 1.14∼ 1.49×, while support model training exceeding the GPU memory limit by up to 9.2×.
Weiduo Chen, Xiaoshe Dong, Fan Zhang 0139, Bowen Li 0009, Yufei Wang 0008, Qiang Wang 0062
ACM Trans. Archit. Code Optim.3