VLDB 2026 Research / reviewers in the wild / expert
Duckhwan Kim 0001
dblp:92/10277
· DBLP profile ↗
12ranked-venue papers
4as first author
0since 2021 · last 2019
0000-0002-6494-2182ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Hardware accelerators and domain-specific architectures · 59% Memory systems · 23% Emerging computing paradigms · 5% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% |
Topics — the 15 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
0.8 | 3 | 2019 | Design and Analysis of a Neural Network Inference Engine Based on Adaptive Weight Compression · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 DeepTrain: A Programmable Embedded Platform for Training Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory · ISCA 2016 |
Memory systems
processing-in-memory |
0.6 | 2 | 2018 | DeepTrain: A Programmable Embedded Platform for Training Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory · ISCA 2016 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
inference engine |
0.4 | 1 | 2019 | Design and Analysis of a Neural Network Inference Engine Based on Adaptive Weight Compression · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Hardware accelerators and domain-specific architectures › model compression
weight compression |
0.4 | 1 | 2019 | Design and Analysis of a Neural Network Inference Engine Based on Adaptive Weight Compression · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN training accelerator |
0.3 | 1 | 2018 | DeepTrain: A Programmable Embedded Platform for Training Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
in-memory computing accelerator |
0.3 | 1 | 2018 | DeepTrain: A Programmable Embedded Platform for Training Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Memory systems › processing-in-memory
near-memory processing |
0.3 | 1 | 2017 | A Programmable Hardware Accelerator for Simulating Dynamical Systems · ISCA 2017 |
Hardware accelerators and domain-specific architectures › accelerator architecture
programmable accelerator |
0.3 | 1 | 2017 | A Programmable Hardware Accelerator for Simulating Dynamical Systems · ISCA 2017 |
Memory systems › processing-in-memory
memory-centric computing |
0.2 | 1 | 2016 | Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory · ISCA 2016 |
Emerging computing paradigms
neuromorphic computing |
0.2 | 1 | 2016 | Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory · ISCA 2016 |
Energy-efficient computing › energy-quality tradeoff
energy-accuracy tradeoff |
0.2 | 1 | 2015 | On the Impact of Energy-Accuracy Tradeoff in a Digital Cellular Neural Network for Image Processing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015 |
Integrated circuit design
3d integration |
0.2 | 1 | 2014 | On the Design of Reliable 3D-ICs Considering Charged Device Model ESD Events During Die Stacking · DAC 2014 |
Hardware reliability and fault tolerance › design for reliability
electrostatic discharge protection |
0.2 | 1 | 2014 | On the Design of Reliable 3D-ICs Considering Charged Device Model ESD Events During Die Stacking · DAC 2014 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.1 | 1 | 2016 | Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory · ISCA 2016 |
Integrated circuit design › 3d integration
through-silicon via |
0.1 | 1 | 2014 | On the Design of Reliable 3D-ICs Considering Charged Device Model ESD Events During Die Stacking · DAC 2014 |
Methods — techniques the papers use, named apart from their topics
near-memory processing · 0.6cellular nonlinear network · 0.6precision scaling · 0.4adaptive compression · 0.4JPEG encoding · 0.4spatial computing fabric · 0.3programmable dataflow · 0.3dataflow model · 0.3data-flow model · 0.3state machine · 0.23d memory integration · 0.2voltage over scaling · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Design and Analysis of a Neural Network Inference Engine Based on Adaptive Weight CompressionabstractNeural networks generally require significant memory capacity/bandwidth to store/access a large number of synaptic weights. This paper presents design of an energy-efficient neural network inference engine based on adaptive weight compression using a JPEG image encoding algorithm. To maximize compression ratio with minimum accuracy loss, the quality factor of the JPEG encoder is adaptively controlled depending on the accuracy impact of each block. With 1% accuracy loss, the proposed approach achieves 63.4× compression for multilayer perceptron (MLP) and 31.3× for LeNet-5 with the MNIST dataset, and 15.3× for AlexNet and 10.2× for ResNet-50 with ImageNet. The reduced memory requirement leads to higher throughput and lower energy for neural network inference (3× effective memory bandwidth and 22× lower system energy for MLP). Jong Hwan Ko, Duckhwan Kim 0001, Taesik Na, Saibal Mukhopadhyay |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | DeepTrain: A Programmable Embedded Platform for Training Deep Neural NetworksabstractThis paper presents, DeepTrain, an embedded platform for high-performance and energy-efficient training of deep neural network (DNN). The key architectural concept of DeepTrain is to develop a spatially homogeneous computing (and memory) fabric with temporally heterogeneous programmable data flows to optimize memory mapping and data reuse during different phases of training operation.The DeepTrain is demonstrated as an in-memory accelerator integrated in the logic layer of a 3-D memory module. A programming model and supporting architecture utilizes the flexible data flow to efficiently accelerate training of various types of DNNs. The cycle level simulation and synthesized design in 15 nm FinFET shows power efficiency of 500 GFLOPS/W, and almost similar throughput for a wide range of DNNs, including convolutional, recurrent, and mixed (CNN+RNN) networks. Duckhwan Kim 0001, Taesik Na, Sudhakar Yalamanchili, Saibal Mukhopadhyay |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | Adaptive Precision Cellular Nonlinear Network
Jaeha Kung 0001, Duckhwan Kim 0001, Saibal Mukhopadhyay |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | Adaptive weight compression for memory-efficient neural networksabstractNeural networks generally require significant memory capacity/bandwidth to store/access a large number of synaptic weights. This paper presents an application of JPEG image encoding to compress the weights by exploiting the spatial locality and smoothness of the weight matrix. To minimize the loss of accuracy due to JPEG encoding, we propose to adaptively control the quantization factor of the JPEG algorithm depending on the error-sensitivity (gradient) of each weight. With the adaptive compression technique, the weight blocks with higher sensitivity are compressed less for higher accuracy. The adaptive compression reduces memory requirement, which in turn results in higher performance and lower energy of neural network hardware. The simulation for inference hardware for multilayer perceptron with the MNIST dataset shows up to 42X compression with less than 1% loss of recognition accuracy, resulting in 3X higher effective memory bandwidth and ~19X lower system energy. Jong Hwan Ko, Duckhwan Kim 0001, Taesik Na, Jaeha Kung 0001, Saibal Mukhopadhyay |
DATE | 2 |
| 2017 | A Programmable Hardware Accelerator for Simulating Dynamical SystemsabstractThe fast and energy-efficient simulation of dynamical systems defined by coupled ordinary/partial differential equations has emerged as an important problem. The accelerated simulation of coupled ODE/PDE is critical for analysis of physical systems as well as computing with dynamical systems. This paper presents a fast and programmable accelerator for simulating dynamical systems. The computing model of the proposed platform is based on multilayer cellular nonlinear network (CeNN) augmented with nonlinear function evaluation engines. The platform can be programmed to accelerate wide classes of ODEs/PDEs by modulating the connectivity within the multilayer CeNN engine. An innovative hardware architecture including data reuse, memory hierarchy, and near-memory processing is designed to accelerate the augmented multilayer CeNN. A dataflow model is presented which is supported by optimized memory hierarchy for efficient function evaluation. The proposed solver is designed and synthesized in 15nm technology for the hardware analysis. The performance is evaluated and compared to GPU nodes when solving wide classes of differential equations and the power consumption is analyzed to show orders of magnitude improvement in energy efficiency. Jaeha Kung 0001, Duckhwan Kim 0001, Saibal Mukhopadhyay |
ISCA | 3 |
| 2016 | Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D MemoryabstractThis paper presents a programmable and scalable digital neuromorphic architecture based on 3D high-density memory integrated with logic tier for efficient neural computing. The proposed architecture consists of clusters of processing engines, connected by 2D mesh network as a processing tier, which is integrated in 3D with multiple tiers of DRAM. The PE clusters access multiple memory channels (vaults) in parallel. The operating principle, referred to as the memory centric computing, embeds specialized state-machines within the vault controllers of HMC to drive data into the PE clusters. The paper presents the basic architecture of the Neurocube and an analysis of the logic tier synthesized in 28nm and 15nm process technologies. The performance of the Neurocube is evaluated and illustrated through the mapping of a Convolutional Neural Network and estimating the subsequent power and performance for both training and inference. Duckhwan Kim 0001, Jaeha Kung 0001, Sek M. Chai, Sudhakar Yalamanchili, Saibal Mukhopadhyay |
ISCA | 1 |
| 2016 | Dynamic Approximation with Feedback Control for Energy-Efficient Recurrent Neural Network HardwareabstractThis paper presents methodology of feedback-controlled dynamic approximation to enable energy-accuracy trade-off in digital recurrent neural network (RNN). A low-power digital RNN engine is presented that employs the proposed dynamic approximation. The on-chip feedback controller is realized by utilizing hysteretic or proportional controller. The dynamic adaptation of bit-precisions during the RNN computation is selected as approximation approach. Considering various applications, the digital RNN engine designed in 28nm CMOS shows ~36% average energy saving compared to the baseline case, with only ~4% of accuracy degradation on average. Jaeha Kung 0001, Duckhwan Kim 0001, Saibal Mukhopadhyay |
ISLPED | 2 |
| 2016 | Partitioning Methods for Interface Circuit of Heterogeneous 3-D-ICs Under Process VariationabstractThis paper presents the design of tier-to-tier interface circuits for 3-D-ICs, where different tiers may operate at different voltages and/or frequencies. The design and partitioning methodologies for the tier-to-tier interface circuit are discussed. The footprint, power, and performance of the interface are analyzed considering the effects of tier-to-tier process variations in 3-D-ICs and technology scaling. The simulation results show that dividing the interface circuit evenly between two tiers reduces footprint but increases power dissipation. For heterogeneous systems with different voltages for reading and writing tiers, dividing the interface between tiers provides better performance than the worst case scenario. On the other hand, placing the interface circuit in the reading tier maximizes throughput for a homogeneous system where both tiers operate at the same voltage. In advanced CMOS nodes, placing interface circuit in the reading tier is a better option due to high delay of the level shifters. Duckhwan Kim 0001, Saibal Mukhopadhyay |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | A power-aware digital feedforward neural network platform with backpropagation driven approximate synapsesabstractThis paper proposes a power-aware digital feedforward neural network platform that utilizes the backpropagation algorithm during training to enable energy-quality trade-off. Given a quality constraint, the proposed approach identifies a set of synaptic weights for approximation in a neural network. The approach selects synapses with small impact on output error, estimated by the backpropagation algorithm, for approximation. The approximations are achieved by a coupled software (reduced bit-width) and hardware (approximate multiplication in the processing engine) based design approaches. The full-chip design in 130nm CMOS shows, compared to a baseline accurate design, the proposed approach reduces system power by ~38% with 0.4% lower recognition accuracy in a classification problem. Jaeha Kung 0001, Duckhwan Kim 0001, Saibal Mukhopadhyay |
ISLPED | 2 |
| 2015 | On the Impact of Energy-Accuracy Tradeoff in a Digital Cellular Neural Network for Image ProcessingabstractThis paper studies the opportunities of energy-accuracy tradeoff in cellular neural network (CNN). Algorithmic characteristics of CNN is coupled with hardware-induced error distribution of a digital CNN cell to evaluate energy-accuracy tradeoff for simple image processing tasks as well as a complex application. The analysis shows that errors modulate the cell dynamics and propagate through the network degrading the output quality and increasing the convergence time. The error propagation is determined by the task being performed by the CNN, specifically, the strength of the feedback template. Controlling precision is observed to be a more effective approach for energy-accuracy tradeoff in CNN than voltage over scaling. Jaeha Kung 0001, Duckhwan Kim 0001, Saibal Mukhopadhyay |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | On the Design of Reliable 3D-ICs Considering Charged Device Model ESD Events During Die StackingabstractThis paper studies charged device model electrostatic discharge (CDM-ESD) events in die stacking process of 3D-ICs and investigates CDM-ESD protection circuits for individual TSVs to prevent high voltage stress on transistor connected to TSV. The models for power, area, delay, and signal integrity of TSVs considering ESD protection are presented. The models are used to drive a methodology to design reliable 3D-ICs considering CDM-ESD while minimizing the overheads. We study the impact of ESD protection on die-to-die asynchronous interface circuit. Duckhwan Kim 0001, Saibal Mukhopadhyay |
DAC | 1 |
| 2013 | Pulsed-latch ASIC synthesis in industrial design flowabstractFlip-flop has long been used as a sequencing element of choice in ASIC design; commercial synthesis tools have also been developed in this context. This work has been motivated by a question of whether existing CAD tools can be employed from RTL to layout while pulsed latch replaces flip-flop as a sequencing element. Two important problems have been identified and their solutions are proposed: placement of pulse generators and latches for integrity of pulse shape, and design of special scan latches and their selective use to reduce hold violations. A reference design flow has also been set up using published documents, in order to assess the proposed one. In 40-nm technology, the proposed flow achieves 20% reduction in circuit area and 30% reduction in power consumption, on average of 12 test circuits. Duckhwan Kim 0001, Youngsoo Shin |
ASP-DAC | 2 |