Donghyeon Han

dblp:212/1662 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0002-5212-2072ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 3 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MEGA.mini-S: A 1.9-to-15.0 TOPS/W Generative AI Processor with Bit-serial Weight Scalability and Bank-Conflict-free Sparsity Compression
Taesu Kim, Donghyeon Han
ISCAS2
2026 A 1.12 μJ/pixel Omnidirectional Image Perception Processor Using Planar-figure-based Memory-efficient Icosphere Mapping
Donghyeon Han
ISCAS2
2026 Securing DNN Acceleration From Off-Chip Memory Vulnerabilities With Low-Overhead Authenticated Encryption
abstract
Security vulnerabilities in deep neural network (DNN) accelerators pose risks for high-stakes applications, with off-chip memory attacks representing a critical threat to both data confidentiality and integrity. While general-purpose processors employ comprehensive cryptographic authenticated encryption for memory security, domain-specific DNN accelerators lack adequate protection, particularly against integrity violations. To address this research gap, we present Sorbet, a DNN accelerator equipped with authenticated encryption to defend against both confidentiality and integrity attacks on off-chip memory. Integrity verification introduces complex memory access patterns in DNN accelerators, as the granularity of authentication operations often clashes with the tiling strategies used for efficient off-chip memory access. Our approach tackles this challenge with a secure memory interface (SMI) module that efficiently: 1) translates the accelerator’s tile request to the required data for cryptographic authentication and 2) aligns fetched data with the memory map of the accelerator’s on-chip buffers. Moreover, our design mitigates the area and performance overhead of cryptographic operations by adopting a lightweight cipher while maintaining security requirements against splicing and replay attacks. Our fabricated chip achieves 22% latency overhead across diverse workloads, including convolutions and multihead attentions (MHAs), which can be further reduced with larger on-chip buffer size and double-buffering. It incurs only 7.9% area and 18.3% energy overhead, which is competitive with recent DNN accelerator defenses with weaker off-chip memory protection.
Kyungmi Lee, Gaurab Das, Donghyeon Han, Anantha P. Chandrakasan
IEEE Trans. Very Large Scale Integr. Syst.3
2025 MEGA.mini: A NPU with Novel Heterogeneous AI Processing Architecture Balancing Efficiency, Performance, and Intelligence for the Era of Generative AI
abstract
• NPU with a Novel big.LITTLECore Architecture to Balance 3 Key Aspects of AI Acceleration–Efficiency: > 95% computations w/ Low-precision FXP–Performance: 3 hierarchical solutions @ MEGA+mini–Intelligence: Hybrid IA (FP for < 5% outlier data)
Donghyeon Han, Anantha P. Chandrakasan
HCS1
2025 A 13.8 TOPS/W Polynomial Implicit Neural Representation Accelerator with Tile Similarity Exploitation and LUT-based Matrix Multiplication Reformation
abstract
This paper presents an energy-efficient polynomial implicit neural network (Poly-INR) processor for image generation tasks. Poly-INR can generate high-resolution images with small parameters, but it requires significant computation and has a long inference time, making it unsuitable for mobile device applications. The proposed processor achieves high energy efficiency through the following three key features: 1) Distribution-aware Heterogeneous Tile Processing (DHTP) reduces grid computation by 98.9% with high compression ratio and feature computation by 45.2% with reduced bit precision based on similarity-aware quantization. 2) Reformed Multiply-MAC Core (RMMC) improves core energy efficiency by 3.71× through the reuse of partial products for both grid and feature. 3) Precision-based Tile Reordering Unit (PTRU) reorders the tiles by their precision and similarity for different channels, enhancing throughput by 39.4% with precision-based sorting and an additional 14.0% with similarity consideration. The proposed processor is implemented in 28nm CMOS technology and achieves a peak energy efficiency of 13.8 TOPS/W.
Wonhoon Park, Sanghyuk An, Hoi-Jun Yoo, Donghyeon Han
ISCAS5
2025 A 2.67 mJ/frame Video Mamba Accelerator with Importance-aware Redundancy Elimination and SSM Computing Reformulation
abstract
An energy-efficient video understanding processor, SLYTHERIN, is proposed to accelerate the new AI model, mamba, efficiently on edge devices. Mamba, a state-of-the-art model for in-context learning, is designed to replace the transformer, whose computational complexity increases significantly in video applications. However, the acceleration of video mamba on edge devices presents two main challenges: 1) slow inference due to the iterative operation phase of mamba and 2) the large energy consumption caused by external memory access (EMA). To address these challenges, the SLYTHERIN is proposed with 3 key building blocks. 1) A 6-stage pipelined task allocator minimizes computational complexity by dynamically managing redundant computations with importance-aware prediction. 2) Reformed computing SSM engine increases core efficiency by tackling the overheads in mamba’s iterative stages with reordering and distributed L1 cache. 3) Patch data management unit addresses the large EMA with difference-based bit-sliced data compression. It finally achieves 2.67-mJ/frame system energy efficiency with 274 FPS.
Youngjin Moon, Sangwoo Ha, Junha Ryu, Hoi-Jun Yoo, Donghyeon Han
ISCAS6
2025 A 17.1 TOPS/W FP-INT Transformer Inference Accelerator with Sparsity Boosting and Output Importance-Aware Processing
abstract
This paper presents an energy-efficient FP-INT transformer inference accelerator for diverse applications. The proposed accelerator achieves high energy efficiency with two key features as solutions: 1) Sparsity Boosting Adder Tree (SBAT) to reduce adder tree power by 23.8% by modifying the adder tree structure and booth encoding to maximize sparsity. 2) Output Importance-aware Processing (OIAP) to reduce the Floating-Point Accumulation (FP-ACC) power by 79.6%, dynamically adjusting input block sizes based on the saliency of output channels, thereby reducing the FP-ACC operations by 78.7%. The proposed accelerator is implemented in 28 nm CMOS technology and achieves 17.1 TOPS/W, leveraging the unique FP-INT characteristics on booth encoding and exploiting redundancy based on output saliency in the transformer architecture.
Jeonggyu So, Seongyon Hong, Wooyoung Jo, Hoi-Jun Yoo, Donghyeon Han
ISCAS7
2024 LSPU: A 20.7 ms Low-Latency Point Neural Network-Based 3D Perception and Semantic LiDAR SLAM System-on-Chip for Autonomous Driving System
abstract
Intelligent 3D Interaction with Wide & Dynamic Surroundings
Jueun Jung, Seungbin Kim, Bokyoung Seo, Wuyoung Jang, Jeongmin Shin, Donghyeon Han, Kyuho Jason Lee
HCS7
2024 NeuGPU: A Neural Graphics Processing Unit for Instant Modeling and Real-Time Rendering on Mobile AR/VR Devices
abstract
•Why 3D Modeling using Neural Radiance Field?
Junha Ryu, Hankyul Kwon, Wonhoon Park, Zhiyong Li 0016, Beomseok Kwon, Donghyeon Han, Dongseok Im, Sangyeob Kim, Hyungnam Joo, Hoi-Jun Yoo
HCS6
2024 Space-Mate: A 303.5mW Real-Time NeRF SLAM Processor with Sparse-Mixture-of-Experts-based Acceleration
abstract
NeRF-based SLAM for robotic applications face computation barrier
Seokchan Song, Haoyang Sang, Dongseok Im, Donghyeon Han, Sangyeob Kim, Hongseok Lee, Hoi-Jun Yoo
HCS4
2024 A 422.1 Mpixels/J Tile-based 4K Super Resolution Processor with Variable Bit Compression
abstract
A super resolution (SR) accelerator with variable bit compression method is proposed for 4K restoration with > 60 frames-per-second (fps) in mobile devices. Since 4K SR has huge intermediate feature maps, large external memory access (EMA) bandwidth is required that recent mobile processors cannot support. Previous SR processors quantized feature maps and weights for EMA reduction, but it has a trade-off between data compression rate and performance drop. To facilitate > 60 fps 4K SR on mobile processor, this work proposes two key features: 1) Tile-based distribution-aware statistical encoding that results in 61.2% compression rate without information loss; 2) An energy-efficient SR processor which supports variable bit encoding, achieving 59.1% reduction of EMA. Designed with 28 nm CMOS technology, the proposed system can accelerate ×2 scale 4K image restoration at 68.7 fps. It shows 3.69 TOPS of peak performance and 422.1 Mpixels/J of energy efficiency, achieving 1.35× higher energy efficiency than the previous SR processor.
Wuyoung Jang, Jinhoon Jo, Jueun Jung, Donghyeon Han, Kyuho Jason Lee
ISCAS5
2024 An Energy-Efficient 3D Point Neural Network Accelerator with Fine-grained LiDAR-SoC Pipeline Structure
abstract
3D point neural network (PNN) segmentation using LiDAR data has emerged as a fundamental stage of high-level intelligence algorithms for autonomous applications such as SLAM, path planning, object detection, etc. However, previous processors were not feasible for real-time and low-power 3D PNN systems since they wasted ~100 ms of LiDAR's sensing time and required 107.3 mW of external memory access before PNN processing. Furthermore, their compute-intensive bin partitioning and point sampling methods were not suitable for large-scale outdoor data, causing significant computing power. Therefore, the entire system, from sensing to processing, must be taken into account for 3D PNN processor implementation. This paper proposes L-PNPU, an energy-efficient 3D PNN segmentation processor optimized with the unique mechanical characteristics of LiDAR. It is designed with three key features: 1) Azimuthal bin partitioning to reduce power and latency, 2) Modified PNN algorithm co-optimized with heterogeneous architecture to remove redundant operation and reduce energy, and 3) Fine-grained LiDAR-System-on-Chip (SoC) pipeline structure to enhance the system energy and throughput. At 250 MHz and 1.0V, L-PNPU achieves 1.27M points/s of throughput and 0.51 μJ/point of energy efficiency.
Bokyoung Seo, Jueun Jung, Donghyeon Han, Kyuho Jason Lee
ISLPED3
2022 HNPU-V2: A 46.6 FPS DNN Training Processor for Real-World Environmental Adaptation based Robust Object Detection on Mobile Devices
abstract
■ Smarter DNNs: # of Parameter ▲
Donghyeon Han, Dongseok Im, Gwangtae Park, Seokchan Song, Juhyoung Lee, Hoi-Jun Yoo
HCS1
2022 DSPU: A 281.6mW Real-Time Deep Learning-Based Dense RGB-D Data Acquisition with Sensor Fusion and 3D Perception System-on-Chip
abstract
3D Data in Mobile Platforms
Dongseok Im, Gwangtae Park, Zhiyong Li 0016, Junha Ryu, Donghyeon Han, Jinsu Lee, Wonhoon Park, Hankyul Kwon, Hoi-Jun Yoo
HCS6
2022 An Efficient High-quality FHD Super-resolution Mobile Accelerator SoC with Hybrid-precision and Energy-efficient Cache
abstract
Super-Resolution on Mobile Platform
Zhiyong Li 0016, Dongseok Im, Donghyeon Han, Hoi-Jun Yoo
HCS4
2022 TSUNAMI: Triple Sparsity-Aware Ultra Energy-Efficient Neural Network Training Accelerator With Multi-Modal Iterative Pruning
abstract
This article proposes the TSUNAMI, which supports an energy-efficient deep-neural-network training. The TSUNAMI supports multi-modal iterative pruning to generate zeros in activation and weight. Tile-based dynamic activation pruning unit and weight memory shared pruning unit eliminate additional memory access. Coarse-zero skipping controller skips multiple unnecessary multiply-and-accumulation (MAC) operations at once, and fine-zero skipping controller skips randomly located unnecessary MAC operations. Weight sparsity balancer solves a utilization degradation caused by weight sparsity imbalance, and the workload of each convolution core is allocated by a random channel allocator. The TSUNAMI achieves an energy efficiency of 3.42 TFLOPS/W at 0.78V and 50MHz with floating-point 8-bit activation and weight. Also, it achieves an energy efficiency of 405.96 TFLOPS/W at 90% sparsity condition.
Sangyeob Kim, Juhyoung Lee, Donghyeon Han, Wooyoung Jo, Hoi-Jun Yoo
IEEE Trans. Circuits Syst. I Regul. Pap.4
2021 An Energy-efficient Floating-Point DNN Processor using Heterogeneous Computing Architecture with Exponent-Computing-in-Memory
abstract
Abstract of Proposed FP CIM Processor (1) Heterogeneous FP Computing Arch. : Separate optimization of FP computing: Realize 2 cycles FP MAC w/ CIM (2) Exponent Computing-in-Memory: In-memory AND/NOR + BL charge reusing: Total memory power 46.4% 2) Mantissa Free Exponent Calculation: Removing redundant normalization: Total MAC power 14.4%
Juhyoung Lee, Ji-Hoon Kim 0004, Wooyoung Jo, Sangyeob Kim, Donghyeon Han, Jinsu Lee, Hoi-Jun Yoo
HCS6
2021 OmniDRL: An Energy-Efficient Mobile Deep Reinforcement Learning Accelerators with Dual-mode Weight Compression and Direct Processing of Compressed Data
abstract
Deep Reinforcement Learning (DRL)▪ No Pre-labelled Data ➔ Training with Trial-and-errors!– Sequential decision making problems @ Unknown environments– Applications: gaming agent, autonomous systems, agent adaptation
Juhyoung Lee, Sangyeob Kim, Ji-Hoon Kim 0004, Wooyoung Jo, Donghyeon Han, Hoi-Jun Yoo
HCS6
2019 DT-CNN: Dilated and Transposed Convolution Neural Network Accelerator for Real-Time Image Segmentation on Mobile Devices
abstract
A convolution neural network (CNN) accelerator is proposed for real-time image segmentation on mobile devices. The proposed CNN processor cuts down the redundant zero computations in dilated and transposed convolution for higher throughput. As a result, the overall computations of the image segmentation are reduced by 86.6% and the proposed CNN processor boosts up the throughput 6.7×. Moreover, the proposed processor utilizes RoI (Region of Interest) based image segmentation algorithm to reduce the overall computational requirement significantly. Although RoI based image segmentation degrades the image segmentation accuracy, the proposed dilation rate adjustment compensates for the accuracy degradation and maintains the accuracy of the full-size image segmentation. Finally, the proposed CNN processor is simulated in 65 nm CMOS technology, and it occupies 6.8 mm2. The proposed processor consumes 196 mW and shows 211 frames-per-second (fps) at the image segmentation for human body parts.
Dongseok Im, Donghyeon Han, Sungpill Choi, Hoi-Jun Yoo
ISCAS2
2018 A 141.4 mW Low-Power Online Deep Neural Network Training Processor for Real-time Object Tracking in Mobile Devices
abstract
A low-power online deep neural network (DNN) training processor is proposed for a real-time object tracking in mobile devices. For a real-time object tracking, a homogeneous core architecture is proposed to achieve 1.33× higher throughput than previous DNN training processor. To reduce the external memory access (EMA), a binary feedback alignment (BFA) algorithm and an integral run-length compression (iRLC) decoder are proposed. While the BFA reduces the EMA by 11.4% compared to the conventional back-propagation approach, the iRLC decoder achieves 29.7% EMA reduction without throughput degradation. Finally, a dropout controller is proposed and achieves 43.9% power reduction through clock-gating. Implemented with 65 nm CMOS technology, the 4.4 mm2DNN training processor achieves 141.1 mW power consumption at 30.4 frames-per-second (fps) real-time object tracking in mobile devices.
Donghyeon Han, Jinsu Lee, Jinmook Lee, Sungpill Choi, Hoi-Jun Yoo
ISCAS1