Keonhee Park

dblp:230/8231 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0003-4151-5484ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2025 A 46 TOPS/W In-/Near-Memory Computing Processor for Large Language Model with Extended Sparse Attention
abstract
A highly energy-efficient embedded DRAM (eDRAM)-based heterogeneous-in/near-memory computing (H-INMC) processor for large language model with extended sparse attention is proposed. The proposed H-INMC processor reduces external memory access (EMA) and enhances macro/system energy efficiency through three key features: 1) attention block fusion computation strategy to maximize input and intermediate data reuse, achieving 85.86% EMA reduction; 2) H-INMC architecture for addressing the imbalances in memory and computation intensity across different computation stages, reducing 77.27% of system latency; and 3) cross-read 3T1C bitcell architecture for mitigating read/write datapath conflictions between CIMs and enhancing the energy efficiency up to 46 TOPS/W with dual-row in-memory-computation. Designed with 28 nm CMOS technology, the proposed H-INMC processor achieves a 92.41% F1 score, evaluated under the Bigbird-large model on the SQuAD 1.1v dataset.
Sunhong An, Hoichang Jeong, Seungbin Kim, Keonhee Park, Kyuho Jason Lee
ISCAS4
2025 A 0.74 μJ/decision and 22.59 TOPS/W Keyword Spotting CIM Processor with Short-Current-Free Multi-level ReRAM and Adaptive-Decision-Level Nonlinear ADC
abstract
An ultra-low energy keyword spotting (KWS) processor based on multi-level ReRAM computing-in-memory (CIM) architecture is proposed in this paper. Previous ReRAM-CIMs suffered from several challenges, such as considerable computation energy due to voltage-based bitcell computation, significant energy consumption and throughput overhead from analog-to-digital (A-to-D) conversion. The proposed processor addresses these challenges by introducing charge-based multi-bit computation ReRAM bitcell, which achieves a 61.39% reduction in MAC computation energy and enhances MAC linearity by 1.65×. Additionally, the proposed 5-bit additive powers-of-two nonlinear ADC leads to a 65.42% reduction in A-to-D conversion energy, a 93.57% decrease in ADC quantization error, and a 1.33× increase in ADC throughput. Furthermore, the pipelined layer fusion clusters reduce the energy consumption for data movement to below 1% and improve the system throughput by 1.35× during keyword spotting inference. The proposed ReRAM-CIM processor is designed in 45 nm technology and occupies a 0.94 mm2area with a 68.5 KB ReRAM cell. The proposed processor also achieves 0.74 μJ/decision energy consumption with 92.7% accuracy on the Google Speech Commands dataset.
Hoichang Jeong, Keonhee Park, Seungbin Kim, Sunhong An, Kyuho Jason Lee
ISCAS3
2025 A 701.7 TOPS/W Compute-in-Memory Processor With Time-Domain Computing for Spiking Neural Network
abstract
Artificial neural networks have led to a higher computational burden, complicating inference tasks on low-power edge devices. Spiking neural network (SNN), which leverages sparse spikes for computation and data transmission, is an effective energy-efficient computing technique. However, the length of spike sequences in SNN varies significantly depending on the input coding method, among which rate coding still results in substantial data movement. A highly energy-efficient SNN accelerator with a time-domain CIM processor is proposed with three key features: 1) time-domain bitcell array for high linearity with lower energy, reducing 58.6% power consumption compared to inverter-chain architecture, 2) time-domain multi-bit accumulate for assisting multi-bit weights without analog-to-digital converter, achieving 47.2% energy reduction of domain-conversion energy, 3) analog precision reconstruction unit for supporting phase coding. The proposed TS-CIM is designed in 65 nm CMOS technology and achieves 701.7 TOPS/W energy efficiency, marking a$1.58\times $enhancement compared to the state-of-the-art SNN CIM.
Keonhee Park, Hoichang Jeong, Seungbin Kim, Jeongmin Shin, Minseo Kim 0001, Kyuho Jason Lee
IEEE Trans. Circuits Syst. I Regul. Pap.1
2025 C2IM-NN: A Low-Power 3D Point Clouds Matching Processor With 1D-CNN Prediction and CAM-Based In-Memory k-NN Searching
abstract
This paper presents a content addressable memory (CAM)-based computing-in-memory (C$^{2}$IM) system designed for energy-efficient k-nearest neighbor (k-NN) searching in 3D point clouds. For autonomous driving applications, an essential process for perceiving the mobile robot’s movements in 3D space is k-NN searching. Especially with the limited hardware resources of mobile processors, the 3D point cloud is too large to upload onto the chip, leading to$O(N^{2})$of external memory accesses and distance calculations. The proposed C$^{2}$IM processor enhances energy efficiency and reduces power consumption through three key features: 1) Dilated 1D-CNN prediction enables voxel-based partitioning, reducing the external memory accesses from$O(N^{2})$to$O(N)$; 2) Vertex clustering reorganizes groups of points into evenly distributed clusters based on the underlying data distribution and reduces the number of points of comparisons by 49.8%; and 3) In-memory k-NN searching with CAM achieves high system energy efficiency while minimizing data transactions between memory and computation logic. Designed with 28 nm CMOS technology, the proposed C$^{2}$IM achieves up to 23.08$\times$energy efficiency, and 48.4% reduction in memory footprint compared to previous ASIC accelerators, and a 99.51% reduction in power consumption compared to state-of-the-art processor implemented in FPGA with high-bandwidth memory.
Jeongmin Shin, Hoichang Jeong, Seungbin Kim, Keonhee Park, Kyuho Jason Lee
IEEE Trans. Circuits Syst. I Regul. Pap.4
2018 Let's guide a smart interface for VR HMD and leap motion
abstract
In this paper, we proposed improvements to the Human Computer Interaction in serious contents using Leap Motion and VR HMD. User immersion and interaction could be the core of VR contents technology. Recently, the advance of VR HMD technology is getting attention in VR contents market. Currently used interface methods are mostly uncomfortable for users since the way of interaction stems from the conventional interface layout rather than considering the physical characteristics of VR HMD. To achieve this, the HUD (Head-Up Display) has been used to optimize the interaction of a serious VR content that is suitable for VR HMD and Leap Motion. In addition, to validate the proposed study, we show the comparison of performance tests.
Keonhee Park, Seongah Chin
VRST1