Kyuho Jason Lee

dblp:122/6099 · also Kyuho Lee 0001 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-4047-1013ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 2 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 A Hardware-Software Co-design of Lightweight Polyp Segmentation Network and Ultra-low-power Processor for Capsule Endoscopy
Ghangmin Yun, Jaekyung Lee, Kyungkeon Chung, Jueun Jung, Bokyoung Seo, Hyejin Lee, Junyoung Park 0002, Kyuho Jason Lee
ISCAS9
2026 HotBa: A Heterogeneous Mamba Accelerator with Δ-Guided Early Rejection for Speculative Decoding
abstract
Mamba has emerged as a promising alternative to Transformers for on-device LLM inference, offering linear complexity and eliminating KV-cache. However, autoregressive decoding reloads full model weights every token, accounting for 93.2% of total inference energy, with no existing accelerator resolving this overhead. We present HotBa, a heterogeneous Mamba accelerator that reduces per-token weight transfer by 82% and redundant computation by 59% via Δ-guided early rejection for wide-tree speculative decoding in Mamba, while a heterogeneous INT8/FP16 core and tree management unit achieve 40.4× area efficiency and 5.18× SSM speedup with 0.4% area overhead. Synthesized in 28nm CMOS, HotBa achieves 75.32 tokens/s with 1.65× speedup and 7.55× energy efficiency over the state-of-the-art Mamba accelerator.
Ghangmin Yun, Jueun Jung, Bokyoung Seo, Chaeyoon Kim, Junghyun Yoo, Kyuho Jason Lee
ISLPED6
2026 A Multibit ReRAM Computing-in-Memory Processor With Adaptive Decision Level Nonlinear ADC for Ultra-Low-Energy Keyword Spotting in Mobile Devices
abstract
This paper presents an ultra-low-energy keyword spotting (KWS) processor based on the charge-mode multi-level resistive random-access memory (ReRAM) bitcell and adaptive analog-to-digital converter (ADC) quantization. Previous ReRAM computing-in-memory (CIM) architectures have suffered several challenges, including excessive computing energy due to direct current branches, computation non-linearity, and throughput degradation resulting from analog-to-digital conversion. The proposed processor addresses these challenges by introducing a charge-mode multi-bit ReRAM bitcell (CRB), which achieves a 61.39% reduction in multiply-and-accumulate (MAC) energy. The CRB also enhances MAC linearity by$1.65\times $. Additionally, the 5-bit additive powers-of-two ADC achieves a 65.42% reduction in analog-to-digital conversion energy, an 11.70 percentage points decrease in ADC accuracy loss, and a$1.33\times $increase in conversion throughput. Furthermore, the pipelined layer fusion clusters reduce intermediate data movement energy to 97.2 nJ and enhance KWS system throughput by$1.35\times $. The proposed processor is designed utilizing 45 nm CMOS technology and compatible ReRAM devices. The processor occupies an area of 0.94 mm2with a 68.5 KB ReRAM cell. The processor also achieves$0.74~\mu $J/decision energy consumption, 22.59 TOPS/W energy efficiency, and 92.7% accuracy on the Google Speech Commands Dataset.
Hoichang Jeong, Seungbin Kim, Heein Yoon, Kyuho Jason Lee
IEEE Trans. Circuits Syst. I Regul. Pap.5
2026 CINELL: An Energy-Efficient Compute-In/Near-Memory eDRAM Processor for Sparse Transformer-Based Large Language Models
abstract
Large language models (LLMs) based on transformer architecture have significantly advanced various AI applications. However, the multihead self-attention (MHSA) in a transformer requires extensive computation and memory bandwidth. Techniques such as sparse attention and attention formula reordering have been explored recently to mitigate these challenges. However, existing processors still struggle to process LLMs efficiently. To overcome these limitations, an energy-efficient compute-in/near-memory (CINM) embedded dynamic random access memory (eDRAM) processor CINELL is proposed, incorporating the following three key features: 1) an attention block fusion computation (ABFC), which maximizes data reuse in the attention map, reduces external memory access (EMA) by 85.86%, and improves hardware utilization to 86.1%; 2) a CINM architecture designed to address the imbalance between memory and computation, while the heterogeneous pipeline reduces the system latency by 77.27%; and 3) a compute-in-memory (CIM) array that supports the cross-read operation (CRO) addresses data direction conflicts, achieving a 98.44% reduction in latency, while dual-row computation (DRC) with reduced adder logic (RAL) improves energy efficiency by$1.58\times $. Designed in 28-nm CMOS technology, the CINELL processor achieves 36.28–58.05 TOPS/W and an F1-score of 92.41% on the SQuAD 1.1v using the BigBird-large model.
Sunhong An, Hoichang Jeong, Seungbin Kim, Kyuho Jason Lee
IEEE Trans. Very Large Scale Integr. Syst.4
2025 A 46 TOPS/W In-/Near-Memory Computing Processor for Large Language Model with Extended Sparse Attention
abstract
A highly energy-efficient embedded DRAM (eDRAM)-based heterogeneous-in/near-memory computing (H-INMC) processor for large language model with extended sparse attention is proposed. The proposed H-INMC processor reduces external memory access (EMA) and enhances macro/system energy efficiency through three key features: 1) attention block fusion computation strategy to maximize input and intermediate data reuse, achieving 85.86% EMA reduction; 2) H-INMC architecture for addressing the imbalances in memory and computation intensity across different computation stages, reducing 77.27% of system latency; and 3) cross-read 3T1C bitcell architecture for mitigating read/write datapath conflictions between CIMs and enhancing the energy efficiency up to 46 TOPS/W with dual-row in-memory-computation. Designed with 28 nm CMOS technology, the proposed H-INMC processor achieves a 92.41% F1 score, evaluated under the Bigbird-large model on the SQuAD 1.1v dataset.
Sunhong An, Hoichang Jeong, Seungbin Kim, Keonhee Park, Kyuho Jason Lee
ISCAS6
2025 A Real-time Point Cloud Segmentation System with Optimized Ground Estimation Algorithm and Selective Neural Network
abstract
This paper introduces a fast and accurate point segmentation system for real-time (< 50 ms) 3D-LiDAR semantic segmentation. The real-time application of 3D point-cloud neural networks (PNNs) for semantic segmentation of LiDAR-measured data faces considerable challenges such as memory overhead and processing delay in its implementation on GPUs due to the significant computational requirements. These challenges arise from the large volume of points and their complex spatial relationships, leading to intense computational and memory usage. To facilitate real-time implementation with high accuracy, a selective point segmentation (SPS) system is proposed with 3 key features: 1) Adaptive ground estimation for the surroundings, excluding the ground from PNN inference, thereby reducing latency by 46.0%, and addressing accuracy reductions due to false positives by implementing 2-step bin skipping; 2) Coarse-grained entropy-and-density-based region skipping (RSK) excludes large areas from ground estimation; and 3) Fine-grained bin skipping (BSK) with z-distribution skips non-ground bins within these areas. Together, the system achieves a processing time of 42.24 ms and a 3D semantic segmentation accuracy of 90.69% at the semantic KITTI dataset.
Jihyeon Hwang, Jueun Jung, Kyuho Jason Lee
ISCAS4
2025 A 0.74 μJ/decision and 22.59 TOPS/W Keyword Spotting CIM Processor with Short-Current-Free Multi-level ReRAM and Adaptive-Decision-Level Nonlinear ADC
abstract
An ultra-low energy keyword spotting (KWS) processor based on multi-level ReRAM computing-in-memory (CIM) architecture is proposed in this paper. Previous ReRAM-CIMs suffered from several challenges, such as considerable computation energy due to voltage-based bitcell computation, significant energy consumption and throughput overhead from analog-to-digital (A-to-D) conversion. The proposed processor addresses these challenges by introducing charge-based multi-bit computation ReRAM bitcell, which achieves a 61.39% reduction in MAC computation energy and enhances MAC linearity by 1.65×. Additionally, the proposed 5-bit additive powers-of-two nonlinear ADC leads to a 65.42% reduction in A-to-D conversion energy, a 93.57% decrease in ADC quantization error, and a 1.33× increase in ADC throughput. Furthermore, the pipelined layer fusion clusters reduce the energy consumption for data movement to below 1% and improve the system throughput by 1.35× during keyword spotting inference. The proposed ReRAM-CIM processor is designed in 45 nm technology and occupies a 0.94 mm2area with a 68.5 KB ReRAM cell. The proposed processor also achieves 0.74 μJ/decision energy consumption with 92.7% accuracy on the Google Speech Commands dataset.
Hoichang Jeong, Keonhee Park, Seungbin Kim, Sunhong An, Kyuho Jason Lee
ISCAS6
2025 A 701.7 TOPS/W Compute-in-Memory Processor With Time-Domain Computing for Spiking Neural Network
abstract
Artificial neural networks have led to a higher computational burden, complicating inference tasks on low-power edge devices. Spiking neural network (SNN), which leverages sparse spikes for computation and data transmission, is an effective energy-efficient computing technique. However, the length of spike sequences in SNN varies significantly depending on the input coding method, among which rate coding still results in substantial data movement. A highly energy-efficient SNN accelerator with a time-domain CIM processor is proposed with three key features: 1) time-domain bitcell array for high linearity with lower energy, reducing 58.6% power consumption compared to inverter-chain architecture, 2) time-domain multi-bit accumulate for assisting multi-bit weights without analog-to-digital converter, achieving 47.2% energy reduction of domain-conversion energy, 3) analog precision reconstruction unit for supporting phase coding. The proposed TS-CIM is designed in 65 nm CMOS technology and achieves 701.7 TOPS/W energy efficiency, marking a$1.58\times $enhancement compared to the state-of-the-art SNN CIM.
Keonhee Park, Hoichang Jeong, Seungbin Kim, Jeongmin Shin, Minseo Kim 0001, Kyuho Jason Lee
IEEE Trans. Circuits Syst. I Regul. Pap.6
2025 C2IM-NN: A Low-Power 3D Point Clouds Matching Processor With 1D-CNN Prediction and CAM-Based In-Memory k-NN Searching
abstract
This paper presents a content addressable memory (CAM)-based computing-in-memory (C$^{2}$IM) system designed for energy-efficient k-nearest neighbor (k-NN) searching in 3D point clouds. For autonomous driving applications, an essential process for perceiving the mobile robot’s movements in 3D space is k-NN searching. Especially with the limited hardware resources of mobile processors, the 3D point cloud is too large to upload onto the chip, leading to$O(N^{2})$of external memory accesses and distance calculations. The proposed C$^{2}$IM processor enhances energy efficiency and reduces power consumption through three key features: 1) Dilated 1D-CNN prediction enables voxel-based partitioning, reducing the external memory accesses from$O(N^{2})$to$O(N)$; 2) Vertex clustering reorganizes groups of points into evenly distributed clusters based on the underlying data distribution and reduces the number of points of comparisons by 49.8%; and 3) In-memory k-NN searching with CAM achieves high system energy efficiency while minimizing data transactions between memory and computation logic. Designed with 28 nm CMOS technology, the proposed C$^{2}$IM achieves up to 23.08$\times$energy efficiency, and 48.4% reduction in memory footprint compared to previous ASIC accelerators, and a 99.51% reduction in power consumption compared to state-of-the-art processor implemented in FPGA with high-bandwidth memory.
Jeongmin Shin, Hoichang Jeong, Seungbin Kim, Keonhee Park, Kyuho Jason Lee
IEEE Trans. Circuits Syst. I Regul. Pap.6
2024 LSPU: A 20.7 ms Low-Latency Point Neural Network-Based 3D Perception and Semantic LiDAR SLAM System-on-Chip for Autonomous Driving System
abstract
Intelligent 3D Interaction with Wide & Dynamic Surroundings
Jueun Jung, Seungbin Kim, Bokyoung Seo, Wuyoung Jang, Jeongmin Shin, Donghyeon Han, Kyuho Jason Lee
HCS8
2024 A 422.1 Mpixels/J Tile-based 4K Super Resolution Processor with Variable Bit Compression
abstract
A super resolution (SR) accelerator with variable bit compression method is proposed for 4K restoration with > 60 frames-per-second (fps) in mobile devices. Since 4K SR has huge intermediate feature maps, large external memory access (EMA) bandwidth is required that recent mobile processors cannot support. Previous SR processors quantized feature maps and weights for EMA reduction, but it has a trade-off between data compression rate and performance drop. To facilitate > 60 fps 4K SR on mobile processor, this work proposes two key features: 1) Tile-based distribution-aware statistical encoding that results in 61.2% compression rate without information loss; 2) An energy-efficient SR processor which supports variable bit encoding, achieving 59.1% reduction of EMA. Designed with 28 nm CMOS technology, the proposed system can accelerate ×2 scale 4K image restoration at 68.7 fps. It shows 3.69 TOPS of peak performance and 422.1 Mpixels/J of energy efficiency, achieving 1.35× higher energy efficiency than the previous SR processor.
Wuyoung Jang, Jinhoon Jo, Jueun Jung, Donghyeon Han, Kyuho Jason Lee
ISCAS6
2024 An Energy-Efficient 3D Point Neural Network Accelerator with Fine-grained LiDAR-SoC Pipeline Structure
abstract
3D point neural network (PNN) segmentation using LiDAR data has emerged as a fundamental stage of high-level intelligence algorithms for autonomous applications such as SLAM, path planning, object detection, etc. However, previous processors were not feasible for real-time and low-power 3D PNN systems since they wasted ~100 ms of LiDAR's sensing time and required 107.3 mW of external memory access before PNN processing. Furthermore, their compute-intensive bin partitioning and point sampling methods were not suitable for large-scale outdoor data, causing significant computing power. Therefore, the entire system, from sensing to processing, must be taken into account for 3D PNN processor implementation. This paper proposes L-PNPU, an energy-efficient 3D PNN segmentation processor optimized with the unique mechanical characteristics of LiDAR. It is designed with three key features: 1) Azimuthal bin partitioning to reduce power and latency, 2) Modified PNN algorithm co-optimized with heterogeneous architecture to remove redundant operation and reduce energy, and 3) Fine-grained LiDAR-System-on-Chip (SoC) pipeline structure to enhance the system energy and throughput. At 250 MHz and 1.0V, L-PNPU achieves 1.27M points/s of throughput and 0.51 μJ/point of energy efficiency.
Bokyoung Seo, Jueun Jung, Donghyeon Han, Kyuho Jason Lee
ISLPED4
2024 An Energy-Efficient, Unified CNN Accelerator for Real-Time Multi-Object Semantic Segmentation for Autonomous Vehicle
abstract
An energy-efficient, unified convolutional neural network (CNN) accelerator is proposed with a lightweight RGB-D network to achieve real-time, multi-object semantic segmentation in autonomous electric vehicle system. First, a lightweight Depth-fused Trilateral Network (DTN) is proposed to achieve high accuracy and real-time operation for road and multi-object segmentation at the same time. Optimized with various types of convolution layers and limited hardware resources, the DTN achieves 94.73% accuracy on KITTI Road dataset. Second, the unified CNN processor is designed with dual-mode shift-register-based input reconfiguration units and layer fusion architecture with 2-types of processing elements for depth-wise separable convolution (DSC) to support 5 different types of convolution layers including standard convolution, dilated convolution, transposed convolution, point-wise convolution, and DSC. With flexible architecture, it achieves 17.97$\times$higher throughput with DTN and DSC layer fusion architecture reduces 34.7% of overall external memory access. Implemented with 28nm CMOS technology, the unified CNN processor shows 43.6 mW power consumption and 4.94 TOPS/W energy efficiency. As a result, the proposed system with DTN realizes 40.07 frames-per-second (fps) throughputs in multi-object semantic segmentation application with high resolution driving scenes dataset.
Jueun Jung, Seungbin Kim, Wuyoung Jang, Bokyoung Seo, Kyuho Jason Lee
IEEE Trans. Circuits Syst. I Regul. Pap.5
2024 A Real-Time Sparsity-Aware 3D-CNN Processor for Mobile Hand Gesture Recognition
abstract
A sparsity-aware 3D-convolution neural network (3D-CNN) accelerator is proposed for the real-time mobile hand gesture recognition (HGR) system. The complex computation of 3D-convolution with the video data makes it difficult for real-time operation, especially in a resource-constrained mobile platform. To facilitate real-time implementation of HGR, this paper proposes three key features: 1) Spatio-temporal Variation Encoding and Inter-frame Differential Aware Network for highly sparse and lightweight network, reducing 94.03% parameters with only 2.57% accuracy loss on NvGesture dataset; 2) the ROI-only Computation architecture for utilizing activation sparsity to reduce the number of MAC operations and the external memory bandwidth by 84.3% and 72.3%, respectively; 3) Weight Sparsity-aware PE and Sparsity-distribution-aware Workload Allocation speed up the inference by$19.8\times $. As a result, the low-latency 3D-CNN accelerator utilizes both activation and weight sparsity with data mapping to maximize the reusability of 3D-CNN, achieving$31\times $faster inference than the state-of-the-art. The proposed processor is designed in 65 nm CMOS technology. It consumes 35 mW of power and achieves 46.25 TOPS/W of energy efficiency. As a result, the system realized 1.584 ms inference latency for real-time HGR in a mobile platform.
Seungbin Kim, Jueun Jung, Kyuho Jason Lee
IEEE Trans. Circuits Syst. I Regul. Pap.3
2022 An Energy-Efficient CNN Accelerator for Multi-object Real-Time Semantic Segmentation in Autonomous Vehicle
abstract
An energy-efficient convolutional neural network (CNN) accelerator is proposed for real-time segmentation in autonomous electric vehicle (AEV) system. The computation of semantic segmentation with high-resolution images makes it difficult for real-time operation in time-critical and resource-constrained AEV. To facilitate real-time implementation in AEV, this paper proposes two key features: 1) A compressed multi-object Depth-fused Trilateral Network (DTN) with dilated convolution and depthwise separable convolution that reduces 90% of the overall computation of baseline [1] and achieves 94.73% accuracy on KITTI Road dataset; 2) An energy-efficient CNN accelerator, which supports 5 types of CONV’s, achieving 1.33× higher throughput than the previous processor [2]. Finally, the proposed processor is designed in 28 nm CMOS technology. It consumes 65.7 mW of power and achieves 2.91 TOPS/W of energy efficiency. As a result, the system realizes 72.2 and 37 frames-per-second of semantic segmentation for road and multi-objects with high resolution.
Jueun Jung, Seungbin Kim, Wuyoung Jang, Hoichang Jeong, Kyuho Jason Lee
ISCAS5
2017 A Real-Time and Energy-Efficient Embedded System for Intelligent ADAS with RNN-Based Deep Risk Prediction using Stereo Camera
Kyuho Jason Lee, Gyeongmin Choe, Kyeongryeol Bong, In-So Kweon, Hoi-Jun Yoo
ICVS1
2016 An intelligent ADAS processor with real-time semi-global matching and intention prediction for 720p stereo vision
abstract
Presents a collection of slides covering the following topics: Intelligent ADAS Processor; Stereo Vision; and SoC Architecture.
Kyuho Jason Lee, Kyeongryeol Bong, Hoi-Jun Yoo
Hot Chips Symposium1
2013 A multi-modal and tunable Radial-Basis-Funtion circuit with supply and temperature compensation
abstract
We propose an analog Radial-Basis-Function (RBF) circuit that generates 4 different types of RBFs, which are spline, Gaussian, multi-quadratic, and log-like spline curves. Moreover, the proposed RBF circuit is designed to have high tunability on centers, heights, and widths. The proposed RBF circuit is also robust to both temperature variation (-37~87□C) and supply voltage variation (1~2V). The sum of area and power consumption of each RBF from 3 different previous works is 13, 622μm2and 121μW, respectively. On the other hand, the proposed circuit occupies only 1,050μm2and consumes 10.5μW which are only 13% and 11.5%, respectively. For its verification, an analog/digital mixed-mode RBF Neural Network (RBFNN) classifier is designed which adopted the proposed RBF circuit.
Kyuho Jason Lee, Junyoung Park 0002, Gyeonghoon Kim, Injoon Hong, Hoi-Jun Yoo
ISCAS1