VLDB 2026 Research / reviewers in the wild / expert
Hanmin Park
dblp:138/9331
· DBLP profile ↗
5ranked-venue papers
1as first author
2since 2021 · last 2022
0000-0001-8252-0840ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Hardware accelerators and domain-specific architectures · 73% Memory systems · 21% Energy-efficient computing · 6% | |
| Artificial intelligence
3 papers |
Deep learning architectures and training · 100% |
Topics — the 9 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
1.1 | 2 | 2022 | ComPreEND: Computation Pruning through Predictive Early Negative Detection for ReLU in a Deep Neural Network Accelerator · IEEE Trans. Computers 2022 GradPIM: A Practical Processing-in-DRAM Architecture for Gradient Descent · HPCA 2021 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN training accelerator |
0.9 | 2 | 2021 | GradPIM: A Practical Processing-in-DRAM Architecture for Gradient Descent · HPCA 2021 Acceleration of DNN Backward Propagation by Selective Computation of Gradients · DAC 2019 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator |
0.6 | 1 | 2022 | ComPreEND: Computation Pruning through Predictive Early Negative Detection for ReLU in a Deep Neural Network Accelerator · IEEE Trans. Computers 2022 |
Memory systems
processing-in-memory |
0.5 | 1 | 2021 | GradPIM: A Practical Processing-in-DRAM Architecture for Gradient Descent · HPCA 2021 |
Memory systems › processing-in-memory
processing-using-DRAM |
0.5 | 1 | 2021 | GradPIM: A Practical Processing-in-DRAM Architecture for Gradient Descent · HPCA 2021 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
0.4 | 1 | 2019 | Acceleration of DNN Backward Propagation by Selective Computation of Gradients · DAC 2019 |
Machine learning › Deep learning architectures and training
activation function |
0.2 | 1 | 2022 | ComPreEND: Computation Pruning through Predictive Early Negative Detection for ReLU in a Deep Neural Network Accelerator · IEEE Trans. Computers 2022 |
Machine learning › Deep learning architectures and training › activation function
ReLU |
0.2 | 1 | 2022 | ComPreEND: Computation Pruning through Predictive Early Negative Detection for ReLU in a Deep Neural Network Accelerator · IEEE Trans. Computers 2022 |
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator |
0.2 | 1 | 2022 | ComPreEND: Computation Pruning through Predictive Early Negative Detection for ReLU in a Deep Neural Network Accelerator · IEEE Trans. Computers 2022 |
Methods — techniques the papers use, named apart from their topics
predictive early detection · 1.1inverted two's complement representation · 1.1bank-group parallelism · 1.0DDR4 SDRAM · 1.0rectified linear unit · 0.8gradient skipping · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | ComPreEND: Computation Pruning through Predictive Early Negative Detection for ReLU in a Deep Neural Network AcceleratorabstractA vast amount of activation values of DNNs are zeros due to ReLU (Rectified Linear Unit), which is one of the most common activation functions used in modern neural networks. Since ReLU outputs zero for all negative inputs, the inputs to ReLU do not need to be determined exactly as long as they are negative. However, many accelerators usually do not consider such aspects of DNNs, losing a huge amount of opportunities for speedups and energy savings. To exploit such opportunities, we propose early negative detection (END), a computation pruning technique that detects the negative results at an early stage. The key to the early negative detection is the adoption of inverted two's complement representation for filter parameters. This ensures that as soon as the intermediate results become negative, the final results are guaranteed to be negative. Upon detection, the remaining computation can be skipped and the following ReLU output can be simply set to zero. We also propose a DNN accelerator architecture (ComPreEND) that takes advantage of such skipping. ComPreEND with END significantly improves both the energy efficiency and the performance according to the evaluation. Compared to the baseline, we obtain 20.5 and 29.3 percent speedup with accurate mode and predictive mode, and energy savings by 28.4 and 41.4 percent, respectively. Namhyung Kim, Hanmin Park, Sungbum Kang, Jinho Lee 0001, Kiyoung Choi |
IEEE Trans. Computers | 2 |
| 2021 | GradPIM: A Practical Processing-in-DRAM Architecture for Gradient DescentabstractIn this paper, we present GradPIM, a processingin-memory architecture which accelerates parameter updates of deep neural networks training. As one of processing-in-memory techniques that could be realized in the near future, we propose an incremental, simple architectural design that does not invade the existing memory protocol. Extending DDR4 SDRAM to utilize bank-group parallelism makes our operation designs in processing-in-memory (PIM) module efficient in terms of hardware cost and performance. Our experimental results show that the proposed architecture can improve the performance of DNN training and greatly reduce memory bandwidth requirement while posing only a minimal amount of overhead to the protocol and DRAM area. Heesu Kim, Hanmin Park, Kwanheum Cho, Eojin Lee, Soojung Ryu, Kiyoung Choi, Jinho Lee 0001 |
HPCA | 2 |
| 2019 | Cell division: weight bit-width reduction technique for convolutional neural network hardware acceleratorsabstractThe datapath bit-width of hardware accelerators for convolutional neural network (CNN) inference is generally chosen to be wide enough, so that they can be used to process upcoming unknown CNNs. Here we introduce the cell division technique, which is a variant of function-preserving transformations. With this technique, it is guaranteed that CNNs that have weights quantized to fixed-point format of arbitrary bit-widths, can be transformed to CNNs with less bit-widths of weights without any accuracy drop (or any accuracy change). As a result, CNN hardware accelerators are released from the weight bit-width constraint, which has been preventing them from having narrower datapaths. In addition, CNNs that have wider weight bit-widths than those assumed by a CNN hardware accelerator can be executed on the accelerator. Experimental results on LeNet-300-100, LeNet-5, AlexNet, and VGG-16 show that weights can be reduced down to 2--5 bits with 2.5X--5.2X decrease in weight storage requirement and of course without any accuracy drop. Hanmin Park, Kiyoung Choi |
ASP-DAC | 1 |
| 2019 | Acceleration of DNN Backward Propagation by Selective Computation of GradientsabstractThe training process of a deep neural network commonly consists of three phases: forward propagation, backward propagation, and weight update. In this paper, we propose a hardware architecture to accelerate the backward propagation. Our approach applies to neural networks that use rectified linear unit. Considering that the backward propagation results in a zero activation gradient when the corresponding activation is zero, we can safely skip the gradient calculation. Based on this observation, we design an efficient hardware accelerator for training deep neural networks by selectively computing gradients. We show the effectiveness of our approach through experiments with various network models. Hanmin Park, Namhyung Kim, Joonsang Yu, Sujeong Jo, Kiyoung Choi |
DAC | 2 |
| 2018 | Training Neural Networks with Low Precision Dynamic Fixed-PointabstractDynamic fixed-point (DFP) is one of the most successful attempts to reduce bit-widths in training neural networks. It has been reported that DFP can reduce the bit-widths of the training operations to 16 bits, except for parameter update operations; parameter updates for general networks need higher precision and are usually done with 32-bit floating-point operations. In this paper, we propose two methods of using 16-bit DFP for all the training operations including parameter updates; weight clipping and gradual batch size increase. Lastly, we combine the two methods to further explore their potentials. We successfully apply 16-bit DFP operations on the parameter updates of LeNet-5 and VGG-16 networks using CIFAR10 and CIFAR100 datasets. Sujeong Jo, Hanmin Park, Kiyoung Choi |
ICCD | 2 |