EDBT 2026 Demo / reviewers in the wild / expert
Shao-Fu Lin
dblp:314/9336
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
GPUs and heterogeneous computing · 87% Parallel and multicore computing · 13% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU memory management |
0.7 | 1 | 2023 | Tensor Movement Orchestration in Multi-GPU Training Systems · HPCA 2023 |
GPUs and heterogeneous computing › multi-GPU computing
multi-GPU training |
0.7 | 1 | 2023 | Tensor Movement Orchestration in Multi-GPU Training Systems · HPCA 2023 |
Parallel and multicore computing › parallel computing › parallel machine learning
data-parallel training |
0.2 | 1 | 2023 | Tensor Movement Orchestration in Multi-GPU Training Systems · HPCA 2023 |
Methods — techniques the papers use, named apart from their topics
swap command orchestration · 0.7PCIe channel contention mitigation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Tensor Movement Orchestration in Multi-GPU Training SystemsabstractAs deep neural network (DNN) models grow deeper and wider, one of the main challenges for training large-scale neural networks is overcoming limited GPU memory capacity. One common solution is to utilize the host memory as the external memory for swapping tensors in and out of GPU memory. However, the effectiveness of such tensor swapping can be impaired in data-parallel training systems due to contention on the shared PCIe channel to the host. In this paper, we propose the first large-model support framework that coordinates tensor movements among GPUs to alleviate PCIe channel contention. We design two types of coordination mechanisms. In the first mechanism, PCIe channel accesses from different GPUs are interleaved by selecting disjoint swapped-out tensors for each GPU. In the second method, swap commands are orchestrated to avoid contention. The effectiveness of these two methods depends on the model size and how often the GPUs synchronize on gradients. Experimental results show that compared to large-model support that is oblivious to channel contention, the proposed solution achieves average speedups of 38.3% to 31.8% when the memory footprint size is 1.33 to 2 times the GPU memory size. Shao-Fu Lin, Yi-Jung Chen, Hsiang-Yun Cheng, Chia-Lin Yang |
HPCA | 1 |
| 2022 | PUMP: Profiling-free Unified Memory Prefetcher for Large DNN Model SupportabstractModern DNNs are going deeper and wider to achieve higher accuracy. However, existing deep learning frameworks require the whole DNN model to fit into the GPU memory when training with GPUs, which puts an unwanted limitation on training large models. Utilizing NVIDIA Unified Memory (UM) could inherently support training DNN models beyond GPU memory capacity. However, naively adopting UM would suffer a significant performance penalty due to the delay of data transfer. In this paper, we propose PUMP, a Profiling-free Unified Memory Prefetcher. PUMP exploits GPU asynchronous execution for prefetch; that is, there exists a delay between the time that CPU launches a kernel and the time the kernel executes in GPU. PUMP extracts memory blocks accessed by the kernel when launching and swaps these blocks into GPU memory. Experimental results show PUMP achieves about 2x speedup on the average compared to the baseline that naively enables UM. Chung-Hsiang Lin, Shao-Fu Lin, Yi-Jung Chen, En-Yu Jenp, Chia-Lin Yang |
ASP-DAC | 2 |