Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hanmin Park

dblp:138/9331 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
2since 2021 · last 2022
0000-0001-8252-0840ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Hardware accelerators and domain-specific architectures · 73% Memory systems · 21% Energy-efficient computing · 6%
Artificial intelligence
3 papers
Deep learning architectures and training · 100%

Topics — the 9 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.122022
ComPreEND: Computation Pruning through Predictive Early Negative Detection for ReLU in a Deep Neural Network Accelerator · IEEE Trans. Computers 2022
GradPIM: A Practical Processing-in-DRAM Architecture for Gradient Descent · HPCA 2021
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN training accelerator
0.922021
GradPIM: A Practical Processing-in-DRAM Architecture for Gradient Descent · HPCA 2021
Acceleration of DNN Backward Propagation by Selective Computation of Gradients · DAC 2019
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator
0.612022
ComPreEND: Computation Pruning through Predictive Early Negative Detection for ReLU in a Deep Neural Network Accelerator · IEEE Trans. Computers 2022
Memory systems
processing-in-memory
0.512021
GradPIM: A Practical Processing-in-DRAM Architecture for Gradient Descent · HPCA 2021
Memory systems › processing-in-memory
processing-using-DRAM
0.512021
GradPIM: A Practical Processing-in-DRAM Architecture for Gradient Descent · HPCA 2021
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.412019
Acceleration of DNN Backward Propagation by Selective Computation of Gradients · DAC 2019
Machine learning › Deep learning architectures and training
activation function
0.212022
ComPreEND: Computation Pruning through Predictive Early Negative Detection for ReLU in a Deep Neural Network Accelerator · IEEE Trans. Computers 2022
Machine learning › Deep learning architectures and training › activation function
ReLU
0.212022
ComPreEND: Computation Pruning through Predictive Early Negative Detection for ReLU in a Deep Neural Network Accelerator · IEEE Trans. Computers 2022
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator
0.212022
ComPreEND: Computation Pruning through Predictive Early Negative Detection for ReLU in a Deep Neural Network Accelerator · IEEE Trans. Computers 2022

Methods — techniques the papers use, named apart from their topics

predictive early detection · 1.1inverted two's complement representation · 1.1bank-group parallelism · 1.0DDR4 SDRAM · 1.0rectified linear unit · 0.8gradient skipping · 0.8
YearPublicationVenuePosition
2022 ComPreEND: Computation Pruning through Predictive Early Negative Detection for ReLU in a Deep Neural Network Accelerator
abstract
A vast amount of activation values of DNNs are zeros due to ReLU (Rectified Linear Unit), which is one of the most common activation functions used in modern neural networks. Since ReLU outputs zero for all negative inputs, the inputs to ReLU do not need to be determined exactly as long as they are negative. However, many accelerators usually do not consider such aspects of DNNs, losing a huge amount of opportunities for speedups and energy savings. To exploit such opportunities, we propose early negative detection (END), a computation pruning technique that detects the negative results at an early stage. The key to the early negative detection is the adoption of inverted two's complement representation for filter parameters. This ensures that as soon as the intermediate results become negative, the final results are guaranteed to be negative. Upon detection, the remaining computation can be skipped and the following ReLU output can be simply set to zero. We also propose a DNN accelerator architecture (ComPreEND) that takes advantage of such skipping. ComPreEND with END significantly improves both the energy efficiency and the performance according to the evaluation. Compared to the baseline, we obtain 20.5 and 29.3 percent speedup with accurate mode and predictive mode, and energy savings by 28.4 and 41.4 percent, respectively.
Namhyung Kim, Hanmin Park, Sungbum Kang, Jinho Lee 0001, Kiyoung Choi
IEEE Trans. Computers2
2021 GradPIM: A Practical Processing-in-DRAM Architecture for Gradient Descent
abstract
In this paper, we present GradPIM, a processingin-memory architecture which accelerates parameter updates of deep neural networks training. As one of processing-in-memory techniques that could be realized in the near future, we propose an incremental, simple architectural design that does not invade the existing memory protocol. Extending DDR4 SDRAM to utilize bank-group parallelism makes our operation designs in processing-in-memory (PIM) module efficient in terms of hardware cost and performance. Our experimental results show that the proposed architecture can improve the performance of DNN training and greatly reduce memory bandwidth requirement while posing only a minimal amount of overhead to the protocol and DRAM area.
Heesu Kim, Hanmin Park, Kwanheum Cho, Eojin Lee, Soojung Ryu, Kiyoung Choi, Jinho Lee 0001
HPCA2
2019 Cell division: weight bit-width reduction technique for convolutional neural network hardware accelerators
abstract
The datapath bit-width of hardware accelerators for convolutional neural network (CNN) inference is generally chosen to be wide enough, so that they can be used to process upcoming unknown CNNs. Here we introduce the cell division technique, which is a variant of function-preserving transformations. With this technique, it is guaranteed that CNNs that have weights quantized to fixed-point format of arbitrary bit-widths, can be transformed to CNNs with less bit-widths of weights without any accuracy drop (or any accuracy change). As a result, CNN hardware accelerators are released from the weight bit-width constraint, which has been preventing them from having narrower datapaths. In addition, CNNs that have wider weight bit-widths than those assumed by a CNN hardware accelerator can be executed on the accelerator. Experimental results on LeNet-300-100, LeNet-5, AlexNet, and VGG-16 show that weights can be reduced down to 2--5 bits with 2.5X--5.2X decrease in weight storage requirement and of course without any accuracy drop.
Hanmin Park, Kiyoung Choi
ASP-DAC1
2019 Acceleration of DNN Backward Propagation by Selective Computation of Gradients
abstract
The training process of a deep neural network commonly consists of three phases: forward propagation, backward propagation, and weight update. In this paper, we propose a hardware architecture to accelerate the backward propagation. Our approach applies to neural networks that use rectified linear unit. Considering that the backward propagation results in a zero activation gradient when the corresponding activation is zero, we can safely skip the gradient calculation. Based on this observation, we design an efficient hardware accelerator for training deep neural networks by selectively computing gradients. We show the effectiveness of our approach through experiments with various network models.
Hanmin Park, Namhyung Kim, Joonsang Yu, Sujeong Jo, Kiyoung Choi
DAC2
2018 Training Neural Networks with Low Precision Dynamic Fixed-Point
abstract
Dynamic fixed-point (DFP) is one of the most successful attempts to reduce bit-widths in training neural networks. It has been reported that DFP can reduce the bit-widths of the training operations to 16 bits, except for parameter update operations; parameter updates for general networks need higher precision and are usually done with 32-bit floating-point operations. In this paper, we propose two methods of using 16-bit DFP for all the training operations including parameter updates; weight clipping and gradual batch size increase. Lastly, we combine the two methods to further explore their potentials. We successfully apply 16-bit DFP operations on the parameter updates of LeNet-5 and VGG-16 networks using CIFAR10 and CIFAR100 datasets.
Sujeong Jo, Hanmin Park, Kiyoung Choi
ICCD2