VLDB 2026 Research / reviewers in the wild / expert
Seungkyu Choi
dblp:00/10381
· DBLP profile ↗
18ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-3125-9707ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 6 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | L2Mersit: A Scaling-Free Sub-8-bit Data Format for On-Device Reliable Large Language Model ServingabstractOn-device large language model (LLM) serving drives low-precision computing to address memory and compute limits. This paper presents L2Mersit, a scaling-free, range-adjustable exponent-encoded data format tailored for sub-8-bit LLM quantization. Building upon the Mersit framework, L2Mersit employs dual mode operation, comprising range-expanded and precision-enhanced modes that dynamically adapt to activation distributions with minimal control overhead. The proposed design eliminates on-the-fly scaling and auxiliary computations while effectively preserving range and precision, thereby achieving both superior perplexity and hardware efficiency. Experimental results demonstrate that L2Mersit achieves the highest accuracy among all 6-bit exponent-encoded formats while reducing the hardware complexity of auxiliary units for low-precision computing, resulting in a 62.7% area reduction. Myeongjin Kim, Hyeonseong Kim, Ik Joon Chang, Seungkyu Choi |
ISLPED | 4 |
| 2026 | Q-VESA: Accelerating Quantization-Aware Vector Search for Fast Retrieval in Prompt EngineeringabstractSimilarity search has drawn significant attention due to the growing demand for AI prompt engineering, which leverages retrieval-augmented generation (RAG) systems to maximize the efficiency of large language models (LLMs). Improving the performance of memory-intensive data retrieval processes using approximate nearest neighbor search (ANNS) algorithms has become increasingly crucial for delivering high-quality generative AI services.In this work, we propose Q-VESA, a software-hardware collaborative solution designed to accelerate low-precision graph-based vector search while preserving recall rates comparable to high-precision. We perform a comprehensive analysis of hierarchical navigable small-world (HNSW), one of the most promising graph-based ANNS methods, using recent datasets tailored for RAG systems. Unlike standard vector database datasets, vector data used in LLMs present unique challenges for low-precision search. To address these, we introduce software-oriented precision partitioning techniques that enable mixed-precision computations during graph traversal without compromising hardware performance. The Q-VESA architecture is developed with two key innovations: a database restructuring scheme with vector data segments, and a dedicated accelerator design to maximize throughput in distance computations. Experimental results demonstrate that Q-VESA on a CPU system achieves query-per-second (QPS) speedups of up to 1.81× in the SIMD execution mode. Furthermore, leveraging minimal area overhead, the ASIC implementation delivers an additional speedup of up to 58.3×, 16.5× and 1.5× compared to the CPU, GPU and the state-of-the-art ASIC-based accelerator, respectively. Seongjoon Cho, Moohyeon Nam, Hongchan Roh, Moo-Kyoung Chung, Se-Hyun Yang, Seungkyu Choi |
IEEE Trans. Computers | 8 |
| 2025 | Precon: A Precision-Convertible Architecture for Accelerating Quantized Deep Learning Models across Various Domains Including LLMsabstractThe sensitivity of LLMs to quantization has driven the development of hardware accelerators tailored for specific low-precision configurations such as weight-only quantization and mixed-precision, which can introduce inefficiencies in dedicated hardware architecture. In this work, we propose Precon, a precision-convertible architecture designed to accelerate various quantized deep learning models, particularly LLMs, through a unified processing unit. By enabling on-the-fly switching between half-float (FP16) decoding and integer (INT) decomposition, the design effectively supports INT4-FP16, INT4-INT4, and INT4INT8 arithmetic within shared logic. Precon achieves up to $4.1 \times$ speedup and 81.4% reduction in energy consumption compared to the baseline across various domains, including the support of both accurate and efficient acceleration of quantized LLMs. Hyeonseong Kim, Jiyun Han, Seungkyu Choi |
DAC | 4 |
| 2025 | Accelerating on-device visual task adaptation by exploiting hybrid sparsity in DNN training
Yeonsik Park, Seungkyu Choi |
J. Syst. Archit. | 3 |
| 2024 | MERSIT: A Hardware-Efficient 8-bit Data Format with Enhanced Post-Training Quantization DNN AccuracyabstractPost-training quantization (PTQ) models utilizing conventional 8-bit Integer or floating-point formats still exhibit significant accuracy drops in modern deep neural networks (DNNs), rendering them unreliable. This paper presents MERSIT, a novel 8-bit PTQ data format designed for various DNNs. While leveraging the dynamic configuration of exponent and fraction bits derived from Posit data format, MERSIT demonstrates enhanced hardware efficiency through the proposed merged decoding scheme. Our evaluation indicates that MERSIT yields more reliable 8-bit PTQ models, exhibiting superior accuracy across various DNNs compared to conventional floating-point formats. Furthermore, the proposed processing unit saves 26.6% in area and 22.2% in power consumption compared to the Posit-based unit, while maintaining comparable efficiency to the floating-point-based unit. Nguyen-Dong Ho, Gyujun Jeong, Cheol-Min Kang, Seungkyu Choi, Ik Joon Chang |
DAC | 4 |
| 2024 | ParaBase: A Configurable Parallel Baseband Processor for Ultra-High-Speed Inter-Satellite Optical CommunicationsabstractThis paper presents ParaBase, a configurable baseband processing architecture that efficiently handles parallel sample streams targeting ultra-wide bandwidth inter-satellite optical communications for the target data rates exceeding 100 Gbps. We propose a parallelogram-style systolic accelerator, specifically designed for parallel processing for correlation kernels, preserving the hardware efficiency inherent in the systolic-array architecture. ParaBase supports end-to-end baseband processing through a set of heterogeneous configurable accelerators customized to their respective parallel processing requirements. It outperforms the SIMD-style architecture and surpasses previous baseband processors by ~7.0X in terms of energy efficiency for FIR filtering. The overall architecture reports energy efficiency of up to 2.9 TOPS/W and 121.8 Gbits/J while supporting a wide range of data rates from 1~128 Gbps via fast reconfiguration for various modulation schemes. Seungkyu Choi, Huanshihong Deng, Kuan-Yu Chen 0001, Yufan Yue, David T. Blaauw, Hun-Seok Kim |
ISLPED | 1 |
| 2023 | SONA: An Accelerator for Transform-Domain Neural Networks with Sparse-Orthogonal WeightsabstractRecent advances in model pruning have enabled sparsity-aware deep neural network accelerators that improve the energy-efficiency and performance of inference tasks. We introduce SONA, a novel transform-domain neural network accelerator in which convolution operations are replaced by element-wise multiplications with sparse-orthogonal weights. SONA employs an output stationary dataflow coupled with an energy-efficient memory organization to reduce the overhead of sparse-orthogonal transform-domain kernels that are concurrently processed without any conflicts. Weights in SONA are non-uniformly quantized with bit-sparse canonical-signed-digit representations to reduce multiplications to simple additions. Moreover, for sparse fully-connected layers (FCLs), SONA introduces column-based-block structured pruning, which is integrated into the same architecture that maintains full multiply-and-accumulate (MAC) array utilization. Compared to prior dense and sparse neural networks accelerators, SONA can reduce inference energy by$5.1\times$and$2.4 \times$and increase performance by$5.2\times$and$2.1\times$, respectively, for convolution layers. For sparse FCLs, SONA can reduce inference energy by$2.4\times$and increase performance by$2\times$compared to prior work. Pierre Abillama, Zichen Fan, Yu Chen 0070, Hyochan An, Qirui Zhang 0001, Seungkyu Choi, David T. Blaauw, Dennis Sylvester, Hun-Seok Kim |
ASAP | 6 |
| 2023 | Accelerating On-Device DNN Training Workloads via Runtime Convergence MonitorabstractWith the growing demand for processing deep learning applications on edge devices, on-device DNN training has become a major workload to execute a variety of vision tasks suited for users. Therefore, architectures employed with algorithm co-design to accelerate the training process have been steadily studied. However, previous solutions are mostly supported by extended versions of the inference studies, such as sparsity, data flow, quantization, etc. Moreover, most works examine their schemes on from-the-scratch training that cannot tolerate inaccurate computing. Accordingly, there are still factors that hinder the overall speed of the DNN training process that has not been addressed in practical workloads. In this work, we propose a runtime convergence monitor to achieve massive computational savings in the practical on-device training workloads (i.e., transfer-learning-based task adaptation). By monitoring the network output data, we determine the training intensity of incoming tasks and adaptively detect the convergence in iteration intervals for training diverse datasets. Furthermore, we enable the computation skip of converged images determined by the monitored prediction probability to enhance the training speed within an iteration. As a result, we perform an accurate but fast convergence in model training for the task adaptation with minimal overhead. Unlike the previous approximation methods, our monitoring system enables runtime optimization and can be easily applicable to any type of accelerator attaining significant speedup. Evaluation results on various datasets show geomean of$2.2\times $speedup when applied in any systolic architectures and further enhancement of$3.6\times $when applied in accelerators dedicated for on-device training. Seungkyu Choi, Jaekang Shin, Lee-Sup Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Energy-Efficient CNN Personalized Training by Adaptive Data ReformationabstractTo adopt deep neural networks in resource-constrained edge devices, various energy- and memory-efficient embedded accelerators have been proposed. However, most off-the-shelf networks are well trained with vast amounts of data, but unexplored users’ data or accelerator’s constraints can lead to unexpected accuracy loss. Therefore, a network adaptation suitable for each user and device is essential to make a high confidence prediction in given environment. We propose simple but efficient data reformation methods that can effectively reduce the communication cost with off-chip memory during the adaptation. Our proposal utilizes the data’s zero-centered distribution and spatial correlation to concentrate the sporadically spread bit-level zeros to the units of value. Consequently, we reduced the communication volume by up to 55.6% per task with an area overhead of 0.79% during the personalization training. Youngbeom Jung, Hyeonuk Kim, Seungkyu Choi, Jaekang Shin, Lee-Sup Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Algorithm/architecture co-design for energy-efficient acceleration of multi-task DNNabstractReal-world AI applications, such as augmented reality or autonomous driving, require processing multiple CV tasks simultaneously. However, the enormous data size and the memory footprint have been a crucial hurdle for deep neural networks to be applied in resource-constrained devices. To solve the problem, we propose an algorithm/architecture co-design. The proposed algorithmic scheme, named SqueeD, reduces per-task weight and activation size by 21.9x and 2.1x, respectively, by sharing those data between tasks. Moreover, we design architecture and dataflow to minimize DRAM access by fully utilizing benefits from SqueeD. As a result, the proposed architecture reduces the DRAM access increment and energy consumption increment per task by 2.2x and 1.3x, respectively. Jaekang Shin, Seungkyu Choi, Jongwoo Ra, Lee-Sup Kim |
DAC | 2 |
| 2022 | A Deep Neural Network Training Architecture With Inference-Aware Heterogeneous Data-TypeabstractAs deep learning applications often encounter accuracy degradation due to the distorted inputs from a variety of environmental conditions, training with personal data has become essential for the edge devices. Hence, ‘training on edge’ by supporting a trainable deep learning accelerator has been actively studied. Nevertheless, previous research does not consider the fundamental datapath for training and the importance of retaining the high performance for inference tasks. In this work, we propose NeuroFlix, a deep neural network training accelerator supporting heterogeneous data-type of floating- and fixed-point for input operands. From two perspectives: 1) separate precision decision for each input data, 2) maintenance of high performance on inference, we configure the data with low-bit fixed-point of activation/weight and floating-point based error gradient securing up to half-precision. A novel MAC architecture is designed to compute low- and high-precision modes for the different input combinations. By substituting a high-cost floating-point based addition to brick-level separate accumulations, we realize both area-efficient architecture and high throughput for low-precision computation. Consequently, NeuroFlix outperforms the previous architectures of state-of-the-art configurations proving its high efficiency in both training and inference. By also comparing with the off-the-shelf bfloat16-based accelerator, it achieves 1.2 ×/2.0 × of speedup/energy-efficiency at training and further enhancement of 3.6 ×/4.5 × at inference. Seungkyu Choi, Jaekang Shin, Lee-Sup Kim |
IEEE Trans. Computers | 1 |
| 2022 | Rare Computing: Removing Redundant Multiplications From Sparse and Repetitive Data in Deep Neural NetworksabstractRecent research shows that 4-bit data precision is sufficient for Deep Neural Network (DNN) inference without accuracy degradation. Due to the low bit-width, a large amount of data is repeated. In this article, we propose a hardware architecture, named Rare Computing Architecture (RCA), that skips redundant computations due to repetitive data in the networks. By exploiting redundancy, RCA is not significantly affected by data-sparsity and maintains great improvements in performance and energy efficiency, while the improvements of existing DNN accelerators are vulnerable to variations in sparsity. In the RCA, repeated data in a window for censoring repetition are detected by a Redundancy Censoring Unit (RCU) and processed at a time, achieving high effective throughput. Additionally, we present a dataflow that exploits abundant data-reusability in DNNs, which enables the high-throughput computations to be ceaselessly performed without an increase of bandwidth for data-read. The proposed architecture is evaluated in two ways of exploiting weight- and activation-repetition. In the evaluation, RCA is compared to a value-agnostic computation and UCNN that is the state-of-the-art accelerator exploiting weight-repetition. Additionally, RCA is compared to Bit-pragmatic that exploits bit-level sparsity. Both evaluations demonstrate that the RCA shows steadily high improvements in performance and energy-efficiency. Kangkyu Park, Seungkyu Choi, Yeongjae Choi, Lee-Sup Kim |
IEEE Trans. Computers | 2 |
| 2021 | A Convergence Monitoring Method for DNN Training of On-Device Task AdaptationabstractDNN training has become a major workload in on-device situations to execute various vision tasks with high performance. Accordingly, training architectures accompanying approximate computing have been steadily studied for efficient acceleration. However, most of the works examine their scheme on from-the-scratch training where inaccurate computing is not tolerable. Moreover, previous solutions are mostly provided as an extended version of the inference works, e.g., sparsity/pruning, quantization, dataflow, etc. Therefore, unresolved issues in practical workloads that hinder the total speed of the DNN training process remain still. In this work, with targeting the transfer learning-based task adaptation of the practical on-device training workload, we propose a convergence monitoring method to resolve the redundancy in massive training iterations. By utilizing the network's output value, we detect the training intensity of incoming tasks and monitor the prediction convergence with the given intensity to provide early-exits in the scheduled training iteration. As a result, an accurate approximation over various tasks is performed with minimal overhead. Unlike the sparsity-driven approximation, our method enables runtime optimization and can be easily applicable to off-the-shelf accelerators achieving significant speedup. Evaluation results on various datasets show a geomean of$2.2\times$speedup over baseline and$1.8\times$speedup over the latest convergence-related training method. Seungkyu Choi, Jaekang Shin, Lee-Sup Kim |
ICCAD | 1 |
| 2020 | A Pragmatic Approach to On-device Incremental Learning System with Selective Weight UpdatesabstractIncremental learning is drawing attention to widen capabilities of device-AI. Previous works have researched to reduce numerous computations and memory accesses required for the training process of IL, but they could not show a noticeable improvement in the weight gradient computation (WGC) phase. Therefore, we propose a selective weight update technique that searches for critical weights to be updated by applying the IL algorithm that training per-task binary masks. Also, we introduce a novel dataflow for the implementation of selective WGC on typical NPUs with minimum overheads. On average, our system shows a 2.9× speed up and 2.5× energy efficiency in WGC without degrading training quality. Jaekang Shin, Seungkyu Choi, Yeongjae Choi, Lee-Sup Kim |
DAC | 2 |
| 2019 | An Optimized Design Technique of Low-bit Neural Network Training for Personalization on IoT DevicesabstractPersonalization by incremental learning has become essential for IoT devices to enhance the performance of the deep learning models trained with global datasets. To avoid massive transmission traffic in the network, exploiting on-device learning is necessary. We propose a software/hardware co-design technique that builds an energy-efficient low-bit trainable system: (1) software optimizations by local low-bit quantization and computation freezing to minimize the on-chip storage requirement and computational complexity, (2) hardware design of a bit-flexible multiply-and-accumulate (MAC) array sharing the same resources in inference and training. Our scheme saves 99.2% on on-chip buffer storage and achieves 12.8x higher peak energy efficiency compared to previous trainable accelerators. Seungkyu Choi, Jaekang Shin, Yeongjae Choi, Lee-Sup Kim |
DAC | 1 |
| 2019 | Compressing Sparse Ternary Weight Convolutional Neural Networks for Efficient Hardware AccelerationabstractReducing the bit-width of weights is an attractive solution for decreasing the large size of CNN models embedded in IoT devices. In the most extreme case, the bit per weight can be reduced to 1-bit in binary weight CNNs. However, this network cannot be compressed further due to the lack of sparsity because the weight distribution of the well-trained model is not biased to either +1 or -1. On the other hand, sparse ternary weight CNNs can be compressed to less than 1-bit per weight, maintaining higher accuracy than binary weight CNNs. Therefore, we propose the following for a weight compression methodology in sparse ternary weight CNNs to minimize the model size: (1) an encoding scheme exploiting high sparsity, (2) two elaborate compression techniques based on encoding direction exploration and layer-wise optimization. To verify the efficiency of hardware acceleration, we design an accelerator that fully exploits our compression scheme. Moreover, a layer rearrangement technique is presented to address a load imbalance problem that occurs during hardware acceleration. As a result, we reduce the effective bit per weight to 0.67-0.80 bit and achieve 4.52-7.70x and 1.52-2.21x improvement of performance and energy efficiency respectively, with higher accuracy compared to previous binary weight CNN work. Hyeonwook Wi, Hyeonuk Kim, Seungkyu Choi, Lee-Sup Kim |
ISLPED | 3 |
| 2018 | TrainWare: A Memory Optimized Weight Update Architecture for On-Device Convolutional Neural Network TrainingabstractTraining convolutional neural network on device has become essential where it allows applications to consider user's individual environment. Meanwhile, the weight update operation from the training process is the primary factor of high energy consumption due to its substantial memory accesses. We propose a dedicated weight update architecture with two key features: (1) a specialized local buffer for the DRAM access deduction (2) a novel dataflow and its suitable processing element array structure for weight gradient computation to optimize the energy consumed by internal memories. Our scheme achieves 14.3%-30.2% total energy reduction by drastically eliminating the memory accesses. Seungkyu Choi, Jaehyeong Sim, Myeonggu Kang, Lee-Sup Kim |
ISLPED | 1 |
| 2017 | SENIN: An energy-efficient sparse neuromorphic system with on-chip learningabstractApplying highly accurate neural networks to mobile devices encounters energy problems in battery-limited mobile environments. To resolve these problems, neuromorphic hardware solutions that enable event-driven operation have been proposed. In this work, we present a novel sparse neuromorphic system that implements an E-I Net algorithm to further improve energy efficiency. We introduce a neuron clock-gating technique that significantly reduces energy consumption by predicting future neuron spike activity without any loss of accuracy. We also propose synaptic pruning to save additional energy with minimal impact on classification accuracy. For fast adaptation to a changing environment, a learning algorithm is implemented in the proposed system. Compared to prior studies, our experimental results illustrate that the proposed system achieves 5.3×-11.4× energy efficiency improvement with comparable accuracy. Myung-Hoon Choi, Seungkyu Choi, Jaehyeong Sim, Lee-Sup Kim |
ISLPED | 2 |