VLDB 2026 Research / reviewers in the wild / expert
Junha Ryu
dblp:268/6031
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0001-8147-3085ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 1 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Energy-Efficient High Resolution Vision Transformer Processor Exploiting Token Similarity Beyond Token Merging
Jungjun Oh, Junha Ryu, Byeongcheol Kim, Yuseon Choi, Hoi-Jun Yoo |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | A 51.2 fps Real-Time 3DGS-SLAM Accelerator using Diagonal Feeding with Symmetric Alpha Reuse and Voxel-based 3D Gaussian Cache ManagementabstractThis work presents a high-speed 3D Gaussian Splatting-based SLAM (3DGS-SLAM) accelerator to support dense mapping for mobile devices. 3DGS-SLAM has two main hardware challenges for acceleration: 1) Large α-computation computation. 2) Memory bottleneck caused by irregular memory access and large number of Gaussians. First, diagonal feeding (DF) controller precludes redundant-α computation, and symmetric alpha reuse (SAR) enables reusing computed alpha. This method reduces 35.3% system computation. Second, voxel-based inter-frame caching (VIFC) enables selective inter-frame voxel caching, which reduces 44.0% of external memory access. As a result, the proposed 3DGS-SLAM accelerator achieves 51.2 fps with 0.07µJ/point with support voltage 0.9V, clock frequency 200MHz mapping a high-quality dense-map. Hyungnam Joo, Seryeong Kim, Jongjun Park, Junha Ryu, Hoi-Jun Yoo |
ISCAS | 4 |
| 2025 | A Real-time 4.31 mJ/Frame Neural-3DGS Processor with Voxel Similarity Memory Management and Opacity-based Sparsity GenerationabstractThis work presents an energy-efficient and real-time rendering Neural-3DGS processor for mobile AR/VR devices. While Neural-3DGS shows high-quality and fast rendering, it exhibits low energy efficiency & high latency in mobile implementation. The proposed processor has three key features for the overall processes in Neural-3DGS: 1) Voxel Similarity-aware Memory Management Unit (VSMMU) eliminates redundant operations and achieves 59.5%, 43.3% reduced energy for external memory access and neural-network computation. 2) LUT-based Pre-Sort Unit (LPSU) utilizes pre-computed order to reduce the latency of sorting by 64.3%. 3) Opacity-aware Gaussian Skipping Core (OGSC) exploits sparsity based on opacity and process 63.3% reduced MAC operations. The proposed processor is implemented in 28 nm CMOS technology. It achieves 98.8 FPS for real-time rendering and 4.31 mJ/Frame energy efficiency. Hongseok Lee, Wonhoon Park, Sanghyuk An, Junha Ryu, Hoi-Jun Yoo |
ISCAS | 4 |
| 2025 | A 2.67 mJ/frame Video Mamba Accelerator with Importance-aware Redundancy Elimination and SSM Computing ReformulationabstractAn energy-efficient video understanding processor, SLYTHERIN, is proposed to accelerate the new AI model, mamba, efficiently on edge devices. Mamba, a state-of-the-art model for in-context learning, is designed to replace the transformer, whose computational complexity increases significantly in video applications. However, the acceleration of video mamba on edge devices presents two main challenges: 1) slow inference due to the iterative operation phase of mamba and 2) the large energy consumption caused by external memory access (EMA). To address these challenges, the SLYTHERIN is proposed with 3 key building blocks. 1) A 6-stage pipelined task allocator minimizes computational complexity by dynamically managing redundant computations with importance-aware prediction. 2) Reformed computing SSM engine increases core efficiency by tackling the overheads in mamba’s iterative stages with reordering and distributed L1 cache. 3) Patch data management unit addresses the large EMA with difference-based bit-sliced data compression. It finally achieves 2.67-mJ/frame system energy efficiency with 274 FPS. Youngjin Moon, Sangwoo Ha, Junha Ryu, Hoi-Jun Yoo, Donghyeon Han |
ISCAS | 4 |
| 2024 | NeuGPU: A Neural Graphics Processing Unit for Instant Modeling and Real-Time Rendering on Mobile AR/VR Devicesabstract•Why 3D Modeling using Neural Radiance Field? Junha Ryu, Hankyul Kwon, Wonhoon Park, Zhiyong Li 0016, Beomseok Kwon, Donghyeon Han, Dongseok Im, Sangyeob Kim, Hyungnam Joo, Hoi-Jun Yoo |
HCS | 1 |
| 2023 | Sibia: Signed Bit-slice Architecture for Dense DNN Acceleration with Slice-level Sparsity ExploitationabstractDeep neural networks (DNNs) have achieved high performance in many AI fields such as 1-D language, 2-D image, and 3-D point cloud processing applications. Since recent DNN tasks require dense matrix operations with various bit-precision and non-ReLU activation functions, mobile neural processing units (NPUs) suffer from the acceleration of diverse DNN tasks within their limited hardware resources and power budget. Although bit-slice architectures benefit from slice-level computation and slice-level sparsity exploitation, the conventional bit-slice representation is inefficient in bit-slice architectures resulting in poor dense DNN execution. This paper proposes the efficient signed bit-slice architecture, Sibia, with the signed bit-slice representation (SBR) for efficient dense DNN acceleration. The SBR adds a sign bit to each bit-slice and changes signed 11112bit-slice to 00002by borrowing a value of 1 from its lower order of the bit-slice. This scheme generates large numbers of zero bit-slices in dense DNNs even not relying on accuracy-sensitive pruning methods or retraining processes. Moreover, the SBR balances positive and negative values of 2’s complement data, allowing accurate bit-slice-based output speculation that pre-computes high orders of bit-slices. Sibia integrates the signed multiplier-and-accumulate (MAC) units for efficient signed bit-slice computations, and the flexible zero skipping processing element (PE) supports the zero input bit-slice skipping and output skipping for high throughput and energy-efficiency. Additionally, the dynamic sparsity monitoring unit monitors sparsity ratio between input and weight data and determines the more sparse one for zero bit-slice skipping. The heterogeneous network-on-chip (NoC) benefits from data reusability during bit-slice computation, reducing transmission bandwidth. Finally, Sibia outperforms the previous bit-slice architecture, Bit-fusion, over 3.65× higher area-efficiency, 3.88× higher energy-efficiency, and 5.35× higher throughput. Dongseok Im, Gwangtae Park, Zhiyong Li 0016, Junha Ryu, Hoi-Jun Yoo |
HPCA | 4 |
| 2023 | A 15.9 mW 96.5 fps Memory-Efficient 3D Reconstruction Processor with Dilation-based TSDF Fusion and Block-Projection Cache SystemabstractA real-time dense 3D reconstruction on lightweight AR headsets is challenging since its memory access surpasses the available memory bandwidth. To solve this problem, the proposed processor integrates two key building blocks - Dilation-based TSDF (D-TSDF) fusion and Block-Projection (BP) engine. D-TSDF projects the depth map in the reverse order of voxel-to-pixel coordinate transformation and dilates it, leading to 96.61% External Memory Access (EMA) reduction with minimum map quality degradation. Second, a specialized BP engine compresses high-resolution occupancy grid by decomposing the 3D bitmap into 2D and 1D vectors, achieving$\times \mathbf{166.09}$reduced memory bandwidth. The proposed processor is implemented in 28nm CMOS technology occupying 1.27 mm2area. As a result, 96.45 fps 3D reconstruction is possible while consuming only 15.94 mW power. Hankyul Kwon, Gwangtae Park, Junha Ryu, Wooyoung Jo, Hoi-Jun Yoo |
ISCAS | 3 |
| 2023 | A 5.99 TFLOPS/W Heterogeneous CIM-NPU Architecture for an Energy Efficient Floating-Point DNN AccelerationabstractThis work presents an energy-efficient digital-based computing-in-memory (CIM) processor to support floating-point (FP) deep neural network (DNN) acceleration. Previous FP-CIM processors have two limitations. Processors with post-alignment shows low throughput due to serial operation, and the other processor with pre-alignment incurs truncation error. To resolve these problems, we focus on the statistics that outlier exists according to shift amount in pre-alignment-based FP operation. As those outlier decreases energy efficiency due to long operation cycles, it needs to be processed separately. The proposed Hetero-FP-CIM integrates both CIM arrays and shared NPU, so they compute both dense inlier and sparse outlier respectively. It also includes efficient weight caching system to avoid entire weight copy in shared NPU. The proposed Hetero-FP-CIM is simulated in 28 nm CMOS technology and occupies 2.7 mm2. As a result, it achieves 5.99 TOPS/W at ImageNet (ResNet50) with bfloat16 representation. Wonhoon Park, Junha Ryu, Soyeon Um, Wooyoung Jo, Sangyoeb Kim, Hoi-Jun Yoo |
ISCAS | 2 |
| 2022 | DSPU: A 281.6mW Real-Time Deep Learning-Based Dense RGB-D Data Acquisition with Sensor Fusion and 3D Perception System-on-Chipabstract3D Data in Mobile Platforms Dongseok Im, Gwangtae Park, Zhiyong Li 0016, Junha Ryu, Donghyeon Han, Jinsu Lee, Wonhoon Park, Hankyul Kwon, Hoi-Jun Yoo |
HCS | 4 |