Hulin Dai

dblp:260/5545 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
GPUs and heterogeneous computing · 74% Performance modeling and evaluation · 20% Parallel and multicore computing · 5%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
GPUs and heterogeneous computing
GPU computing
1.122022
Reveal training performance mystery between TensorFlow and PyTorch in the single GPU environment · Sci. China Inf. Sci. 2022
Feluca: A Two-Stage Graph Coloring Algorithm With Color-Centric Paradigm on GPU · IEEE Trans. Parallel Distributed Syst. 2021
Performance modeling and evaluation
benchmarking
0.612022
Reveal training performance mystery between TensorFlow and PyTorch in the single GPU environment · Sci. China Inf. Sci. 2022
GPUs and heterogeneous computing › deep learning on GPUs
single-GPU training
0.612022
Reveal training performance mystery between TensorFlow and PyTorch in the single GPU environment · Sci. China Inf. Sci. 2022
Graph algorithms and graph theory
graph coloring
0.512021
Feluca: A Two-Stage Graph Coloring Algorithm With Color-Centric Paradigm on GPU · IEEE Trans. Parallel Distributed Syst. 2021
Graph algorithms and graph theory › graph coloring
parallel graph coloring
0.512021
Feluca: A Two-Stage Graph Coloring Algorithm With Color-Centric Paradigm on GPU · IEEE Trans. Parallel Distributed Syst. 2021
GPUs and heterogeneous computing
GPU memory management
0.412020
Capuchin: Tensor-based GPU Memory Management for Deep Learning · ASPLOS 2020
Parallel and multicore computing › parallel algorithms
parallel algorithm design
0.112021
Feluca: A Two-Stage Graph Coloring Algorithm With Color-Centric Paradigm on GPU · IEEE Trans. Parallel Distributed Syst. 2021
Machine learning › Efficient and distributed learning
memory-efficient training
0.112020
Capuchin: Tensor-based GPU Memory Management for Deep Learning · ASPLOS 2020

Methods — techniques the papers use, named apart from their topics

sequential spread · 1.0recursion-based coloring · 1.0color-centric paradigm · 1.0tensor-based memory management · 0.9benchmarking · 0.6
YearPublicationVenuePosition
2022 Reveal training performance mystery between TensorFlow and PyTorch in the single GPU environment
Hulin Dai, Xuanhua Shi, Ligang He, Qian Xiong, Hai Jin 0001
Sci. China Inf. Sci.1
2021 Feluca: A Two-Stage Graph Coloring Algorithm With Color-Centric Paradigm on GPU
abstract
There are great challenges in performing graph coloring on GPU in general. First, the long-tail problem exists in the recursion algorithm because the conflict (i.e., different threads assign the adjacent nodes to the same color) becomes more likely to occur as the number of iterations increases. Second, it is hard to parallelize the sequential spread algorithm because the color allocation depends on the adjoining iteration. Third, the atomic operation is widely used on GPU to maintain the color list, which can greatly reduce the efficiency of GPU threads. In this article, we propose a two-stage high-performance graph coloring algorithm, called Feluca, aiming to address the above challenges. Feluca combines the recursion-based method with the sequential spread-based method. In the first stage, Feluca uses a recursive routine to color a majority of vertices in the graph. Then, it switches to the sequential spread method to color the remaining vertices in order to avoid the conflicts of the recursive algorithm. Moreover, the following techniques are proposed to further improve the graph coloring performance. i) A new method is proposed to eliminate the cycles in the graph; ii) a top-down scheme is developed to avoid the atomic operation originally required for color selection; and iii) a novel color-centric coloring paradigm is designed to improve the degree of parallelism for the sequential spread part. All these newly developed techniques, together with further GPU-specific optimizations such as coalesced memory access, comprise an efficient parallel graph coloring solution in Feluca. We have conducted extensive experiments on NVIDIA GPU. The results show that Feluca can achieve 1.19 - 8.39× speedup over the state-of-the-art algorithms.
Zhigao Zheng 0001, Xuanhua Shi, Ligang He, Hai Jin 0001, Shuo Wei, Hulin Dai
IEEE Trans. Parallel Distributed Syst.6
2020 Capuchin: Tensor-based GPU Memory Management for Deep Learning
abstract
In recent years, deep learning has gained unprecedented success in various domains, the key of the success is the larger and deeper deep neural networks (DNNs) that achieved very high accuracy. On the other side, since GPU global memory is a scarce resource, large models also pose a significant challenge due to memory requirement in the training process. This restriction limits the DNN architecture exploration flexibility.
Xuanhua Shi, Hulin Dai, Hai Jin 0001, Weiliang Ma, Qian Xiong, Fan Yang 0024, Xuehai Qian
ASPLOS3
2020 A parameter-level parallel optimization algorithm for large-scale spatio-temporal data mining
Xuanhua Shi, Ligang He, Dongxiao Yu, Hai Jin 0001, Chen Yu 0003, Hulin Dai, Zezhao Feng
Distributed Parallel Databases7