EDBT 2026 Demo / reviewers in the wild / expert
Yulhwa Kim
dblp:223/9434
· DBLP profile ↗
19ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0003-3735-821XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 3D Integration of Hybrid IGZO/Si and IGZO eDRAMs for High-Density/High-Performance On-Chip MemoryabstractThe growing need for advanced memory architectures leveraging 3D integration has become increasingly critical in modern computing systems. In particular, memory architectures that match the performance of static random access memory (SRAM) while significantly increasing density are highly impactful. In this paper, we propose a 3D integration-based hybrid InGaZnO(IGZO)/Si embedded dynamic random access memory architecture (Hybrid-3D) and circuit design, which markedly increases on-chip memory density and enhances system performance. The superiority of Hybrid-3D is demonstrated through rigorous validation involving process integration verification, transistor-level modeling, and circuit-level memory design. Detailed evaluations of the vertically stacked memory operation confirm stable operations, enabling a 22× increase in on-chip memory density compared to SRAM. Integrating Hybrid-3D on-chip memory into neural processing unit (NPU) architectures results in substantial improvements in energy efficiency and processing speed. System-level evaluations across vision and natural language processing (NLP) tasks reveal a maximum energy efficiency improvement of 3.2× and a throughput increase of 2.6×. Munhyeon Kim, Sukhyun Choi, Yulhwa Kim, Jae-Joon Kim |
DATE | 3 |
| 2026 | RUnQuant: High-resolution weight quantization via unanchored weight decomposition in column-wise granularity for CIM accelerators
Kang Eun Jeon, Yulhwa Kim, Jong Hwan Ko |
J. Syst. Archit. | 3 |
| 2025 | L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language ModelsabstractDue to the high memory and computational costs associated with large language models (LLMs), model compression techniques such as quantization, which reduces inference costs, and parameter-efficient fine-tuning (PEFT) methods like Low-Rank Adaptation (LoRA), which reduce training costs, have gained significant popularity.This trend has spurred active research into quantization-aware PEFT techniques, aimed at maintaining model accuracy while minimizing memory overhead during both inference and training.Previous quantization-aware PEFT methods typically apply post-training quantization (PTQ) to pre-trained LLMs, followed by PEFT to recover accuracy loss.Meanwhile, this approach has limitations in recovering the accuracy loss.In this paper, we propose L4Q, a method that integrates Quantization-Aware Training (QAT) with LoRA.By employing a memory-optimized layer design, L4Q significantly reduces QAT's memory overhead, making its training cost comparable to LoRA, while preserving the advantage of QAT in producing fully quantized LLMs with high accuracy.Our experiments demonstrate that this combined approach to quantization and fine-tuning achieves superior accuracy compared to decoupled finetuning schemes, particularly in 4-bit and 3-bit quantization, positioning L4Q as an efficient QAT solution.Using the LLaMA and Mistral models with instructional datasets, we showcase L4Q's capabilities in language tasks and few-shot learning. Hyesung Jeon, Yulhwa Kim, Jae-Joon Kim |
ACL (1) | 2 |
| 2025 | Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory AcceleratorsabstractCompute-in-memory (CIM) is an efficient method for implementing deep neural networks (DNNs) but suffers from substantial overhead from analog-to-digital converters (ADCs), especially as ADC precision increases. Low-precision ADCs can reduce this overhead but introduce partial-sum quantization errors degrading accuracy. Additionally, low-bit weight constraints, imposed by cell limitations and the need for multiple cells for higher-bit weights, present further challenges. While fine-grained partial-sum quantization has been studied to lower ADC resolution effectively, weight granularity, which limits overall partial-sum quantized accuracy, remains underexplored. This work addresses these challenges by aligning weight and partial-sum quantization granularities at the column-wise level. Our method improves accuracy while maintaining dequantization overhead, simplifies training by removing two-stage processes, and ensures robustness to memory cell variations via independent column-wise scale factors. We also propose an open-source CIM-oriented convolution framework to handle fine-grained weights and partial-sums efficiently, incorporating a novel tiling method and group convolution. Experimental results on ResNet-20 (CIFAR-10, CIFAR-100) and ResNet-18 (ImageNet) show accuracy improvements of 0.99%, 2.69%, and 1.01%, respectively, compared to the best-performing related works. Additionally, variation analysis reveals the robustness of our method against memory cell variations. These findings highlight the effectiveness of our quantization scheme in enhancing accuracy and robustness while maintaining hardware efficiency in CIM-based DNN implementations. Our code is available at https://github.com/jiyoonkm/ColumnQuant. Kang Eun Jeon, Yulhwa Kim, Jong Hwan Ko |
DATE | 3 |
| 2025 | Compute-in-Memory Array Design Using Stacked Hybrid IGZO/Si eDRAM cellsabstractTo effectively accelerate neural networks in compute-in-memory (CIM) based systems, higher memory cell density is essential to handle the increasing computational workload and number of parameters. While CMOS embedded dynamic random access memory (eDRAM) is being explored as an alternative, improving the short retention time$(t_{ret})$($t_{ret}$(> 100 s), but additional improvements are needed due to its substantial cell variability and slower operating speed compared to CMOS-based cells. This paper proposes a cell and array design for CIM using 3T-based stacked hybrid IGZO/Si eDRAM (Hybrid-3T) and performs a system-level deep neural network (DNN) evaluation. The Hybrid-3T cell, designed based on 7-nm FinFET technology, achieves$t_{ret}$that is 100 s longer compared to IGZO-based 3T eDRAM (IGZO-3T). The proposed Hybrid-3T offers a 3.4 × higher bit cell density compared to 8T SRAM bit cells and a 2 × higher density compared to CMOS-based 3T eDRAM (CMOS-3T), while demonstrating similar throughput and variability levels to CMOS eDRAM and SRAM-based systems. Furthermore, we evaluate DNN inference accuracy for vision and natural language processing (NLP) tasks using the proposed CIM design, examining the impact of improved cell variability and retention time on system-level characteristics. The retention time for ensuring CIM operation accuracy$(t_{ret,CIM})$is 108times longer in Hybrid-3T than CMOS-3T, and the$t_{ret,CIM}$considering variability$(t_{ret,CIM_{v}})$is more than 3 × longer than IGZO-3T eDRAM. As a result, the proposed Hybrid-3T eDRAM CIM leverages the advantages of both CMOS-3T and IGZO-3T CIM designs, enabling the development of high-performance, reliable systems. Munhyeon Kim, Yulhwa Kim, Jae-Joon Kim |
DATE | 2 |
| 2025 | Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM ReasoningabstractRecent reasoning-focused language models achieve high accuracy by generating lengthy intermediate reasoning paths before producing final answers.
While this approach is effective in solving problems that require logical thinking, long reasoning paths significantly increase memory usage and reduce throughput of token generation, limiting the practical deployment of such models.
We propose Reasoning Path Compression (RPC), a training-free method that accelerates inference by leveraging the semantic sparsity of reasoning paths.
RPC periodically compresses the KV cache by retaining cache entries that receive high importance score, which are computed using a selector window composed of recently generated queries.
Experiments show that RPC improves generation throughput of QwQ-32B by up to 1.60$\times$ compared to the inference with full KV cache, with an accuracy drop of 1.2\% on the AIME 2024 benchmark.
Our findings demonstrate that semantic sparsity in reasoning traces can be effectively exploited for compression, offering a practical path toward efficient deployment of reasoning LLMs. Our code is available at https://github.com/jiwonsong-dev/ReasoningPathCompression. Jiwon Song, Dongwon Jo, Yulhwa Kim, Jae-Joon Kim |
NeurIPS | 3 |
| 2024 | FIGNA: Integer Unit-Based Accelerator Design for FP-INT GEMM Preserving Numerical AccuracyabstractThe weight-only quantization has emerged as a promising technique for alleviating the computational burden of large language models (LLMs) by employing low-precision integer (INT) weights, while retaining full-precision floating point (FP) activations to ensure inference quality. Despite the memory footprint reduction achieved through decreased bit-precision of weight parameters, the actual computing performance is often not improved significantly due to FP-INT multiply-accumulation (MAC) operations being performed on the floating point unit (FPU) after de quantizing the INT weight values to FP values, owing to the lack of dedicated FP- INT arithmetic units. In this study, we investigate the impact of introducing a dedicated FP-INT unit on overall performance and find that such specialization does not yield substantial improvements. As an alternative approach, we propose FIGNA, an accelerator based on INT units designed specifically for FP- INT MAC operations. A key feature of FIGNA is its ability to achieve the same numerical accuracy as the FPU while relying solely on the integer-unit, a departure from prior methods that relied on integer-units with numerical approximations for FP arithmetic results, albeit claiming similar inference accuracy through dedicated network training. Through comprehensive experiments on FP- INT quantized networks for LLMs, including OPT and BLOOM, we demonstrate the superior performance of FIGNA compared to conventional FPUs in terms of performance per area ($TOPS/mm^{2}$) and energy efficiency (TOPS/W) across various input and weight precision combinations. For instance, in the FP16-INT4 case, FIGNA shows 6.34x higher$TOPS/ mm^{2}$and 2.19x higher TOPS/W compared to the baseline. Jaeyong Jang, Yulhwa Kim, Juheun Lee, Jae-Joon Kim |
HPCA | 2 |
| 2024 | SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer BlocksabstractLarge language models (LLMs) have proven to be highly effective across various natural language processing tasks. However, their large number of parameters poses significant challenges for practical deployment. Pruning, a technique aimed at reducing the size and complexity of LLMs, offers a potential solution by removing redundant components from the network. Despite the promise of pruning, existing methods often struggle to achieve substantial end-to-end LLM inference speedup. In this paper, we introduce SLEB, a novel approach designed to stream- line LLMs by eliminating redundant transformer blocks. We choose the transformer block as the fundamental unit for pruning, because LLMs exhibit block-level redundancy with high similarity between the outputs of neighboring blocks. This choice allows us to effectively enhance the processing speed of LLMs. Our experimental results demonstrate that SLEB outperforms previous LLM pruning methods in accelerating LLM inference while also maintaining superior perplexity and accuracy, making SLEB as a promising technique for enhancing the efficiency of LLMs. The code is available at: https://github.com/jiwonsong-dev/SLEB. Jiwon Song, Kyungseok Oh, Taesu Kim, Yulhwa Kim, Jae-Joon Kim |
ICML | 5 |
| 2024 | Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language ModelsabstractBinarization, which converts weight parameters to binary values, has emerged as an effective strategy to reduce the size of large language models (LLMs). However, typical binarization techniques significantly diminish linguistic effectiveness of LLMs.
To address this issue, we introduce a novel binarization technique called Mixture of Scales (BinaryMoS). Unlike conventional methods, BinaryMoS employs multiple scaling experts for binary weights, dynamically merging these experts for each token to adaptively generate scaling factors. This token-adaptive approach boosts the representational power of binarized LLMs by enabling contextual adjustments to the values of binary weights. Moreover, because this adaptive process only involves the scaling factors rather than the entire weight matrix, BinaryMoS maintains compression efficiency similar to traditional static binarization methods. Our experimental results reveal that BinaryMoS surpasses conventional binarization techniques in various natural language processing tasks and even outperforms 2-bit quantization methods, all while maintaining similar model size to static binarization techniques. Dongwon Jo, Taesu Kim, Yulhwa Kim, Jae-Joon Kim |
NeurIPS | 3 |
| 2023 | Winning Both the Accuracy of Floating Point Activation and the Simplicity of Integer Arithmetic
Yulhwa Kim, Jaeyong Jang, Jehun Lee, Byeongwook Kim, Baeseong Park, Se Jung Kwon, Dongsoo Lee, Jae-Joon Kim |
ICLR | 1 |
| 2023 | Leveraging Early-Stage Robustness in Diffusion Models for Efficient and High-Quality Image SynthesisabstractWhile diffusion models have demonstrated exceptional image generation capabilities, the iterative noise estimation process required for these models is compute-intensive and their practical implementation is limited by slow sampling speeds. In this paper, we propose a novel approach to speed up the noise estimation network by leveraging the robustness of early-stage diffusion models. Our findings indicate that inaccurate computation during the early-stage of the reverse diffusion process has minimal impact on the quality of generated images, as this stage primarily outlines the image while later stages handle the finer details that require more sensitive information. To improve computational efficiency, we combine our findings with post-training quantization (PTQ) to introduce a method that utilizes low-bit activation for the early reverse diffusion process while maintaining high-bit activation for the later stages. Experimental results show that the proposed method can accelerate the early-stage computation without sacrificing the quality of the generated images. Yulhwa Kim, Dongwon Jo, Hyesung Jeon, Taesu Kim, Daehyun Ahn, Jae-Joon Kim |
NeurIPS | 1 |
| 2022 | Extreme Partial-Sum Quantization for Analog Computing-In-Memory Neural Network AcceleratorsabstractIn Analog Computing-in-Memory (CIM) neural network accelerators, analog-to-digital converters (ADCs) are required to convert the analog partial sums generated from a CIM array to digital values. The overhead from ADCs substantially degrades the energy efficiency of CIM accelerators so that previous works attempted to lower the ADC resolution considering the distribution of the partial sums. Despite the efforts, the required ADC resolution still remains relatively high. In this article, we propose the data-driven partial sum quantization scheme, which exhaustively searches for the optimal quantization range with little computational burden. We also report that analyzing the characteristics of the partial sum distributions at each layer gives an additional information to further reduce the ADC resolution compared to previous works that mostly used the characteristics of the partial sum distributions of the entire network. Based on the finer-level data-driven approach combined with retraining, we present a methodology for extreme partial-sum quantization. Experimental results show that the proposed method can reduce the ADC resolution to 2 to 3 bits for CIFAR-10 dataset, which is the smaller ADC bit resolution than any previous CIM-based NN accelerators. Yulhwa Kim, Jae-Joon Kim |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2021 | Mapping Binary ResNets on Computing-In-Memory Hardware with Low-bit ADCsabstractImplementing binary neural networks (BNNs) on computing-in-memory (CIM) hardware has several attractive features such as small memory requirement and minimal overhead in peripheral circuits such as analog-to-digital converters (ADCs). On the other hand, one of the downsides of using BNNs is that it degrades the classification accuracy. Recently, ResNet-style BNNs are gaining popularity with higher accuracy than conventional BNNs. The accuracy improvement comes from the high-resolution skip connection which binary ResNets use to compensate the information loss caused by binarization. However, the high-resolution skip connection forces the CIM hardware to use high-bit ADCs again so that area and energy overhead becomes larger. In this paper, we demonstrate that binary ResNets can be also mapped on CIM with low-bit ADCs via aggressive partial sum quantization and input-splitting combined with retraining. As a result, the key advantages of BNN CIM such as small area and energy consumption can be preserved with higher accuracy. Yulhwa Kim, Hyunmyung Oh, Jae-Joon Kim |
DATE | 1 |
| 2020 | Algorithm/Hardware Co-Design for In-Memory Neural Network Computing with Minimal Peripheral Circuit OverheadabstractWe propose an in-memory neural network accelerator architecture called MOSAIC which uses minimal form of peripheral circuits; 1-bit word line driver to replace DAC and 1-bit sense amplifier to replace ADC. To map multi-bit neural networks on MOSAIC architecture which has 1-bit precision peripheral circuits, we also propose a bit-splitting method to approximate the original network by separating each bit path of the multi-bit network so that each bit path can propagate independently throughout the network. Thanks to the minimal form of peripheral circuits, MOSAIC can achieve an order of magnitude higher energy and area efficiency than previous in-memory neural network accelerators. Yulhwa Kim, Sungju Ryu, Jae-Joon Kim |
DAC | 2 |
| 2020 | Time-step interleaved weight reuse for LSTM neural network computingabstractIn Long Short-Term Memory (LSTM) neural network models, a weight matrix tends to be repeatedly loaded from DRAM if the size of on-chip storage of the processor is not large enough to store the entire matrix. To alleviate heavy overhead of DRAM access for weight loading in LSTM computations, we propose a weight reuse scheme which utilizes the weight sharing characteristics in two adjacent time-step computations. Experimental results show that the proposed weight reuse scheme reduces the energy consumption by 28.4-57.3% and increases the overall throughput by 110.8% compared to the conventional schemes. Naebeom Park, Yulhwa Kim, Daehyun Ahn, Taesu Kim, Jae-Joon Kim |
ISLPED | 2 |
| 2019 | In-memory batch-normalization for resistive memory based binary neural network hardwareabstractBinary Neural Network (BNN) has a great potential to be implemented on Resistive memory Crossbar Array (RCA)-based hardware accelerators because it requires only 1-bit precision for weights and activations. While general structures to implement convolution or fully-connected layers in RCA-based BNN hardware were actively studied in previous works, Batch-Normalization (BN) layer, which is another key layer of BNN, has not been discussed in depth yet. In this work, we propose in-memory batch-normalization schemes which integrate BN layers on RCA so that area/energy-efficiency of the BNN accelerators can be maximized. In addition, we also show that sense amp error due to device mismatch can be suppressed using the proposed in-memory BN design. Yulhwa Kim, Jae-Joon Kim |
ASP-DAC | 2 |
| 2019 | Effect of Device Variation on Mapping Binary Neural Network to Memristor Crossbar ArrayabstractIn memristor crossbar array (MCA)-based neural network hardware, it is generally assumed that entire word-lines (WLs) are simultaneously enabled for parallel matrix-vector multiplication (MxV) operation. However, the error probability of MxV in a memristor crossbar array (MCA) increases as the resistance ratio (R-ratio) of a memristor decreases and the resistance variation and the number of simultaneously activated WLs increase. In this paper, we analyze the effect of R-ratio and variation of memristor devices on read sense margin and inference accuracy of MCA-based Binary Neural Network (BNN) hardware. We first show that only a limited number of WLs should be enabled to ensure correct MxV output when the R-ratio is small. On the other hand, we also show that, if the resistance variation becomes higher than a certain level, simultaneous activation of large number of WLs produces the higher accuracy even when R-ratio is small. Based on the analysis, we propose the Accuracy Estimation (AE) factor to find the optimal number of word lines that are simultaneously activated. Wooseok Yi, Yulhwa Kim, Jae-Joon Kim |
DATE | 2 |
| 2018 | Input-Splitting of Large Neural Networks for Power-Efficient Accelerator with Resistive Crossbar Memory ArrayabstractResistive Crossbar memory Arrays (RCA) have been gaining interest as a promising platform to implement Convolutional Neural Networks (CNN). One of the major challenges in RCA-based design is that the number of rows in an RCA is often smaller than the number of input neurons in a layer. Previous works used high-resolution Analog-to-Digital Converters (ADCs) to compute the partial weighted sum in each array and merged partial sums from multiple arrays outside the RCAs. However, such approach suffers from significant power consumption due to the need for high-resolution ADCs. In this paper, we propose a methodology to more efficiently construct a large CNN with multiple RCAs. By splitting the input feature map and retraining the CNN with proper initialization, we demonstrate that any CNN model can be represented with multiple arrays without using intermediate partial sums. The experimental results show that the ADC power of the proposed design is 32x smaller and the total chip power of the proposed design is 3x smaller than those of the baseline design. Yulhwa Kim, Daehyun Ahn, Jae-Joon Kim |
ISLPED | 1 |
| 2018 | Compact Convolution Mapping on Neuromorphic Hardware using Axonal DelayabstractMapping Convolutional Neural Network (CNN) to a neuromorphic hardware has been inefficient in synapse memory usage because both kernel/input reuse are not exploited well. We propose a method to enable kernel reuse by utilizing axonal delay, which is a biological parameter for a spiking neuron. Using IBM TrueNorth as a test platform, we demonstrate that the number of cores, neurons, synapses, and synaptic operations per time step can be reduced by up to 20.9x, 27.9x, 88.4x, and 1586x, respectively, compared to the conventional scheme, which raises the possibility of implementing large-scale CNN on neuromorphic hardware. Jinseok Kim 0004, Yulhwa Kim, Jae-Joon Kim |
ISLPED | 2 |