Ji-Hoon Kim 0004

dblp:51/4060-4 · also Jihoon Kim 0004 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-4895-0369ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-author · 8 since 2021
YearPublicationVenuePosition
2026 SCRec: A Scalable Computational Storage System With Statistical Sharding and Tensor-Train Decomposition for Recommendation Models
abstract
Deep Learning Recommendation Models (DLRMs) are essential for personalized content delivery in web applications, but their parameter sizes have grown to terabytes, with memory bandwidth demands exceeding TB/s. Furthermore, the workload intensity within the model varies based on the target mechanism, making it difficult to build an optimized recommendation system. In this paper, we propose SCRec, a scalable computational storage recommendation system that can handle TB-scale industrial DLRMs while guaranteeing high bandwidth requirements. SCRec utilizes a software framework that features a mixed-integer programming (MIP)-based cost model, efficiently fetching data based on data access patterns and adaptively configuring memory-centric and compute-centric cores. Additionally, SCRec integrates hardware acceleration cores to enhance DLRM computations, particularly allowing for the high-performance reconstruction of approximated embedding vectors from extremely compressed tensor-train (TT) format. By combining its software framework and hardware accelerators, while eliminating data communication overhead by being implemented on a single server, SCRec achieves substantial improvements in DLRM inference performance. It delivers up to 55.77× speedup compared to a CPU-DRAM system with no loss in accuracy and up to 13.35× energy efficiency gains over a multi-GPU system.
Ji-Hoon Kim 0004, Joo-Young Kim 0001
IEEE Trans. Computers2
2025 EXION: Exploiting Inter-and Intra-Iteration Output Sparsity for Diffusion Models
Jaehoon Heo, Adiwena Putra, Jieon Yoon, Sungwoong Yune, Hangyeol Lee, Ji-Hoon Kim 0004, Joo-Young Kim 0001
HPCA6
2023 Accelerating Large-Scale Graph-Based Nearest Neighbor Search on a Computational Storage Platform
abstract
$K$-nearest neighbor search is one of the fundamental tasks in various applications and the hierarchical navigable small world (HNSW) has recently drawn attention in large-scale cloud services, as it easily scales up the database while offering fast search. On the other hand, a computational storage device (CSD) that combines programmable logic and storage modules on a single board becomes popular to address the data bandwidth bottleneck of modern computing systems. In this paper, we propose a computational storage platform that can accelerate a large-scale graph-based nearest neighbor search algorithm based on SmartSSD CSD. To this end, we modify the algorithm more amenable on the hardware and implement two types of accelerators using HLS- and RTL-based methodology with various optimization methods. In addition, we scale up the proposed platform to have 4 SmartSSDs and apply graph parallelism to boost the system performance further. As a result, the proposed computational storage platform achieves 75.59 query per second throughput for the SIFT1B dataset at 258.66W power dissipation, which is 12.83x and 17.91x faster and 10.43x and 24.33x more energy efficient than the conventional CPU-based and GPU-based server platform, respectively. With multi-terabyte storage and custom acceleration capability, we believe that the proposed computational storage platform is a promising solution for cost-sensitive cloud datacenters.
Ji-Hoon Kim 0004, Yeo-Reum Park, Jaeyoung Do, Soo-Young Ji, Joo-Young Kim 0001
IEEE Trans. Computers1
2022 A Dual-Mode Similarity Search Accelerator based on Embedding Compression for Online Cross-Modal Image-Text Retrieval
abstract
Image-text retrieval (ITR) that identifies the relevant images for a given text query, or vice versa, is the fundamental task in emerging vision-and-language machine learning applications. Recently, the cross-modal approach that extracts image and text features in separate reasoning pipelines but performs the similarity search on the same embedding representation is proposed for the real-time ITR system. However, the similarity search that finds the most relevant data in huge data embeddings for a given query becomes the bottleneck of the ITR system.In this paper, we propose a dual-mode similarity search accelerator that can solve the computational hurdle for online image-to-text and text-to-image retrieval service. We propose an embedding compression scheme that removes the sparsity in the text embeddings, further eliminating the time-consuming masking operations in the later processing pipeline. Combining with the data quantization from 32-bit floating-point to 8- bit integer, we reduce the target dataset size by 95.1% with less than 0.1% accuracy loss for 1024-dimensional embedding features. In addition, we propose a streamlined similarity search data flow for both query types, which minimizes the required memory bandwidth with maximal data reuse. The query and data embeddings are guaranteed to be fetched only once from the external memory with the optimized data flow. Based on the proposed data representation and flow, we design a scalable similarity search accelerator that includes multiple ITR kernels. Each ITR kernel has modular design, composed of a separate memory access module and a computing module. The computing module supports pipelined operations of the four similarity search tasks: dot product calculation, data reordering, partial score aggregation, and ranking. We double the number of processing operations in the computing module with the DSP packing technique. Finally, we implement the proposed accelerator with six ITR kernels on the Xilinx Alveo U280 FPGA card. It shows 2.98 tera operations per second (TOPS) performance at 186 MHz, achieving 526/144 and 1163/306 queries per second (QPS) performance for image-to-text and text-to-image retrieval on MS-COCO 1K/5K benchmark. It is up to 359.0 × and 13.9 × faster and 503.6 × and 68.7 × more energy-efficient than the baseline and optimized GPU implementation on Nvidia Titan RTX, respectively.
Yeo-Reum Park, Ji-Hoon Kim 0004, Jaeyoung Do, Joo-Young Kim 0001
FCCM2
2022 Trinity: End-to-End In-Database Near-Data Machine Learning Acceleration Platform for Advanced Data Analytics
abstract
Three Important yet Independent Technology Trends
Ji-Hoon Kim 0004, Kwanghyun Park 0001, Soo-Young Ji, Joo-Young Kim 0001
HCS1
2021 Accelerating Large-Scale Nearest Neighbor Search with Computational Storage Device
abstract
K-nearest neighbor algorithm that searches the K closest samples in a high dimensional feature space is one of the most fundamental tasks in machine learning and image retrieval applications. Computational storage device that combines computing unit and storage module on a single board becomes popular to address the data bandwidth bottleneck of the conventional computing system. In this paper, we propose a nearest neighbor search acceleration platform based on computational storage device, which can process a large-scale image dataset efficiently in terms of speed, energy, and cost. We believe that the proposed acceleration platform is promising to be deployed in cloud datacenters for data-intensive applications.
Ji-Hoon Kim 0004, Yeo-Reum Park, Jaeyoung Do, Soo Young Ji, Joo-Young Kim 0001
FCCM1
2021 An Energy-efficient Floating-Point DNN Processor using Heterogeneous Computing Architecture with Exponent-Computing-in-Memory
abstract
Abstract of Proposed FP CIM Processor (1) Heterogeneous FP Computing Arch. : Separate optimization of FP computing: Realize 2 cycles FP MAC w/ CIM (2) Exponent Computing-in-Memory: In-memory AND/NOR + BL charge reusing: Total memory power 46.4% 2) Mantissa Free Exponent Calculation: Removing redundant normalization: Total MAC power 14.4%
Juhyoung Lee, Ji-Hoon Kim 0004, Wooyoung Jo, Sangyeob Kim, Donghyeon Han, Jinsu Lee, Hoi-Jun Yoo
HCS2
2021 OmniDRL: An Energy-Efficient Mobile Deep Reinforcement Learning Accelerators with Dual-mode Weight Compression and Direct Processing of Compressed Data
abstract
Deep Reinforcement Learning (DRL)▪ No Pre-labelled Data ➔ Training with Trial-and-errors!– Sequential decision making problems @ Unknown environments– Applications: gaming agent, autonomous systems, agent adaptation
Juhyoung Lee, Sangyeob Kim, Ji-Hoon Kim 0004, Wooyoung Jo, Donghyeon Han, Hoi-Jun Yoo
HCS3
2019 An Ultra-Low-Power Analog-Digital Hybrid CNN Face Recognition Processor Integrated with a CIS for Always-on Mobile Devices
abstract
An ultra-low-power analog-digital hybrid always-on face recognition (FR) processor integrated with a CMOS image sensor (CIS) is proposed for the wearable mobile devices applications such as user authentication. The proposed processor is the first IC with full process of FR in a single chip. The processor adopts analog-digital hybrid convolution operation for efficient integration of CNN processor with CIS. The analog convolution processor is proposed for the computation of the 1stlayer of CNN and the quantization operation without an ADC that can achieve 15.7% power reduction with 1.3% minimal accuracy loss. In addition, the analog weighted-sum unit with low power (5.18TOPS/W) is proposed with switched-drain regulation (SDR) current mirror which can achieve less than 6% mirroring error. The processor is simulated in 65-nm CMOS technology, 15.84mm2area with 2.5V and 1.2V for analog domain and 0.77-1.1V for digital domain. It consumes 0.6198mW to evaluate one face at 1 fps and achieves 96.18% FR accuracy in LFW dataset.
Ji-Hoon Kim 0004, Kwantae Kim, Hoi-Jun Yoo
ISCAS1