EDBT 2026 Demo / reviewers in the wild / expert
Wonhoon Park
dblp:320/8281
· DBLP profile ↗
10ranked-venue papers
1as first author
10since 2021 · last 2026
0009-0001-9098-8985ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 198.7 μJ/token Block Diffusion LLM Processor with Mask Token Similarity-Based Activation Reuse
Yujin Moon, Seryeong Kim, Wonhoon Park, Wooyoung Jo, Yuseon Choi, Sunjoo Whang, Hoi-Jun Yoo |
ISCAS | 4 |
| 2025 | IRIS: A 8.55 mJ/frame Spatial Computing SoC for Real-time Interactable-Rendering and Surface-aware-Modeling with 3D Gaussian Splatting
Seokchan Song, Seryeong Kim, Wonhoon Park, Jongjun Park, Sanghyuk An, Gwangtae Park, Minseo Kim 0001, Hoi-Jun Yoo |
HCS | 3 |
| 2025 | A 13.8 TOPS/W Polynomial Implicit Neural Representation Accelerator with Tile Similarity Exploitation and LUT-based Matrix Multiplication ReformationabstractThis paper presents an energy-efficient polynomial implicit neural network (Poly-INR) processor for image generation tasks. Poly-INR can generate high-resolution images with small parameters, but it requires significant computation and has a long inference time, making it unsuitable for mobile device applications. The proposed processor achieves high energy efficiency through the following three key features: 1) Distribution-aware Heterogeneous Tile Processing (DHTP) reduces grid computation by 98.9% with high compression ratio and feature computation by 45.2% with reduced bit precision based on similarity-aware quantization. 2) Reformed Multiply-MAC Core (RMMC) improves core energy efficiency by 3.71× through the reuse of partial products for both grid and feature. 3) Precision-based Tile Reordering Unit (PTRU) reorders the tiles by their precision and similarity for different channels, enhancing throughput by 39.4% with precision-based sorting and an additional 14.0% with similarity consideration. The proposed processor is implemented in 28nm CMOS technology and achieves a peak energy efficiency of 13.8 TOPS/W. Wonhoon Park, Sanghyuk An, Hoi-Jun Yoo, Donghyeon Han |
ISCAS | 2 |
| 2025 | A Real-time 4.31 mJ/Frame Neural-3DGS Processor with Voxel Similarity Memory Management and Opacity-based Sparsity GenerationabstractThis work presents an energy-efficient and real-time rendering Neural-3DGS processor for mobile AR/VR devices. While Neural-3DGS shows high-quality and fast rendering, it exhibits low energy efficiency & high latency in mobile implementation. The proposed processor has three key features for the overall processes in Neural-3DGS: 1) Voxel Similarity-aware Memory Management Unit (VSMMU) eliminates redundant operations and achieves 59.5%, 43.3% reduced energy for external memory access and neural-network computation. 2) LUT-based Pre-Sort Unit (LPSU) utilizes pre-computed order to reduce the latency of sorting by 64.3%. 3) Opacity-aware Gaussian Skipping Core (OGSC) exploits sparsity based on opacity and process 63.3% reduced MAC operations. The proposed processor is implemented in 28 nm CMOS technology. It achieves 98.8 FPS for real-time rendering and 4.31 mJ/Frame energy efficiency. Hongseok Lee, Wonhoon Park, Sanghyuk An, Junha Ryu, Hoi-Jun Yoo |
ISCAS | 2 |
| 2024 | NeuGPU: A Neural Graphics Processing Unit for Instant Modeling and Real-Time Rendering on Mobile AR/VR Devicesabstract•Why 3D Modeling using Neural Radiance Field? Junha Ryu, Hankyul Kwon, Wonhoon Park, Zhiyong Li 0016, Beomseok Kwon, Donghyeon Han, Dongseok Im, Sangyeob Kim, Hyungnam Joo, Hoi-Jun Yoo |
HCS | 3 |
| 2024 | A 28.6 mJ/iter Stable Diffusion Processor for Text-to-Image Generation with Patch Similarity-based Sparsity Augmentation and Text-based Mixed-PrecisionabstractThis paper presents an energy-efficient stable diffusion processor for text-to-image generation. While stable diffusion attained attention for high-quality image synthesis results, its inherent characteristics hinder its deployment on mobile platforms. The proposed processor achieves high throughput and energy efficiency with three key features as solutions: 1) Patch similarity-based sparsity augmentation (PSSA) to reduce external memory access (EMA) energy of self-attention score by 60.3 %, leading to 37.8 % total EMA energy reduction. 2) Text-based important pixel spotting (TIPS) to allow 44.8 % of the FFN layer workload to be processed with low-precision activation. 3) Dual-mode bit-slice core (DBSC) architecture to enhance energy efficiency in FFN layers by 43.0 %. The proposed processor is implemented in 28 nm CMOS technology and achieves 3.84 TOPS peak throughput with 225.6 mW average power consumption. In sum, 28.6 mJ/iteration highly energy-efficient text-to-image generation processor can be achieved at MS-COCO dataset. Wooyoung Jo, Seongyon Hong, Beomseok Kwon, Wonhoon Park, Hoi-Jun Yoo |
ISCAS | 5 |
| 2024 | A 3.55 mJ/frame Energy-efficient Mixed-Transformer based Semantic Segmentation Accelerator for Mobile DevicesabstractAn energy-efficient semantic segmentation (SS) processor, achieving 3.55 mJ/frame system energy efficiency, is proposed. To address the challenges posed by Mixed Transformer (MiT)-based SS, including high external memory bandwidth requirement and large on-chip memory footprint, we introduce a novel compression method called Chunk-based Bit Plane Compression (CBPC). CBPC leverages the high inter-token locality of feature maps in MiT-based SS, along with the robustness and compression ratio variations based on bit position to achieve a high compression ratio. To support CBPC, we propose an area and power-efficient CBPC encoder/decoder. In addition, a Similar Token Coarse Skipping (STCS) Core is proposed for high throughput. It enables row-wise clock gating and array-wise coarse skipping to reduce redundant computation. By removing redundant computation, the processor achieves higher throughput and lower computation power. The proposed processor reduces 67.6% of EMA power and accomplishes 19.24 TOPS/W core energy efficiency. The proposed processor achieves 44.3% higher system energy efficiency than the previous processors. Jongjun Park, Seryeong Kim, Wonhoon Park, Seokchan Song, Hoi-Jun Yoo |
ISCAS | 3 |
| 2024 | An Energy-Efficient CNN/Transformer Hybrid Neural Semantic Segmentation Processor With Chunk-Based Bit Plane Data Compression and Similarity-Based Token-Level Skipping ExploitationabstractA novel energy-efficient semantic segmentation (SS) processor is proposed for achieving high system energy efficiency on mobile devices. 1) Excessive external memory access and 2) a large amount of redundant computation hinders energy-efficient SS acceleration. Three key features enable real-time energy-efficient CNN/ViT hybrid SS. A new compression method named Chunk-based Bit Plane Compression (CBPC) reduces the memory footprint and energy consumption due to external memory access. CBPC enhances compression ratio by leveraging the high inter-token similarity of feature maps and applying bit plane compression in sign-magnitude data representation, using chunk-wise low-bit plane shared bias. The proposed CBPC encoder/decoder supports CBPC with minimum area overhead. Additionally, the Similar Token Coarse Skipping (STCS) Core enhances the throughput and reduces the computation power by eliminating redundant computations. STCS core employs Row-wise Line Gating for low-power computation and Array-wise Coarse Skipping to minimize redundant computation. As a result, our proposed processor reduces external memory access energy by 67.6% and achieves a core energy efficiency of 19.24 TOPS/W. Our solution achieves 3.55mJ/frame system-level energy efficiency which is 79.7% higher than the previous SOTA SS processor. Jongjun Park, Seryeong Kim, Wonhoon Park, Seokchan Song, Hoi-Jun Yoo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | A 5.99 TFLOPS/W Heterogeneous CIM-NPU Architecture for an Energy Efficient Floating-Point DNN AccelerationabstractThis work presents an energy-efficient digital-based computing-in-memory (CIM) processor to support floating-point (FP) deep neural network (DNN) acceleration. Previous FP-CIM processors have two limitations. Processors with post-alignment shows low throughput due to serial operation, and the other processor with pre-alignment incurs truncation error. To resolve these problems, we focus on the statistics that outlier exists according to shift amount in pre-alignment-based FP operation. As those outlier decreases energy efficiency due to long operation cycles, it needs to be processed separately. The proposed Hetero-FP-CIM integrates both CIM arrays and shared NPU, so they compute both dense inlier and sparse outlier respectively. It also includes efficient weight caching system to avoid entire weight copy in shared NPU. The proposed Hetero-FP-CIM is simulated in 28 nm CMOS technology and occupies 2.7 mm2. As a result, it achieves 5.99 TOPS/W at ImageNet (ResNet50) with bfloat16 representation. Wonhoon Park, Junha Ryu, Soyeon Um, Wooyoung Jo, Sangyoeb Kim, Hoi-Jun Yoo |
ISCAS | 1 |
| 2022 | DSPU: A 281.6mW Real-Time Deep Learning-Based Dense RGB-D Data Acquisition with Sensor Fusion and 3D Perception System-on-Chipabstract3D Data in Mobile Platforms Dongseok Im, Gwangtae Park, Zhiyong Li 0016, Junha Ryu, Donghyeon Han, Jinsu Lee, Wonhoon Park, Hankyul Kwon, Hoi-Jun Yoo |
HCS | 8 |