Sunwoo Lee 0005

dblp:56/7811-5 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0001-7760-0168ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 An Area- and Energy-Efficient Point-Based 3-D Object Detection Processor With Feature-Decoupled PointNet for Autonomous Driving
Changyu Seong, Sunwoo Lee 0005, Beomseok Kim, Dongsuk Jeon
IEEE Trans. Circuits Syst. I Regul. Pap.2
2026 An Energy-Efficient Block-Convolution-Based Super-Resolution Processor With Tiling Artifact Reduction
Byeungseok Yoo, Sungjin Park 0003, Sunwoo Lee 0005, Dongsuk Jeon
IEEE Trans. Very Large Scale Integr. Syst.3
2025 PaCA: Partial Connection Adaptation for Efficient Fine-Tuning
abstract
Prior parameter-efficient fine-tuning (PEFT) algorithms reduce memory usage and computational costs of fine-tuning large neural network models by training only a few additional adapter parameters, rather than the entire model. However, the reduction in computational costs due to PEFT does not necessarily translate to a reduction in training time; although the computational costs of the adapter layers are much smaller than the pretrained layers, it is well known that those two types of layers are processed sequentially on GPUs, resulting in significant latency overhead. LoRA and its variants avoid this latency overhead by merging the low-rank adapter matrices with the pretrained weights during inference. However, those layers cannot be merged during training since the pretrained weights must remain frozen while the low-rank adapter matrices are updated continuously over the course of training. Furthermore, LoRA and its variants do not reduce activation memory, as the first low-rank adapter matrix still requires the input activations to the pretrained weights to compute weight gradients. To mitigate this issue, we propose **Pa**rtial **C**onnection **A**daptation (**PaCA**), which fine-tunes randomly selected partial connections within the pretrained weights instead of introducing adapter layers in the model. PaCA not only enhances training speed by eliminating the time overhead due to the sequential processing of the adapter and pretrained layers but also reduces activation memory since only partial activations, rather than full activations, need to be stored for gradient computation. Compared to LoRA, PaCA reduces training time by 22% and total memory usage by 16%, while maintaining comparable accuracy across various fine-tuning scenarios, such as fine-tuning on the MMLU dataset and instruction tuning on the Oasst1 dataset. PaCA can also be combined with quantization, enabling the fine-tuning of large models such as LLaMA3.1-70B. In addition, PaCA enables training with 23% longer sequence and improves throughput by 16\% on both NVIDIA A100 GPU and INTEL Gaudi2 HPU compared to LoRA. The code is available at [https://github.com/WooSunghyeon/paca](https://github.com/WooSunghyeon/paca).
Sunghyeon Woo, Sol Namkung, Sunwoo Lee 0005, Inho Jeong, Beomseok Kim, Dongsuk Jeon
ICLR3
2025 CLAT: A Clustering-Based Attention Transformer Accelerator for Low-Latency Text Generation in LLMs
abstract
Transformer-based large language models (LLMs) excel in text generation but face challenges like memory bandwidth bottlenecks and large key-value (KV) cache sizes as context lengths grow, impacting low-latency performance. Existing accelerators adopt parallelism, model compression, and sparsity exploitation but often fail to fully utilize token-specific sparsity, limiting their effectiveness for long context lengths. CLAT addresses these issues with a low-overhead clustering algorithm that identifies relevant key vector clusters for each token’s query, omitting less relevant vectors with minimal impact. It optimizes memory bandwidth using routing for single-batch inference and introduces scheduling techniques to reduce attention layer latency. Additionally, CLAT compresses model parameters to 4-bit precision and KV caches to 8-bit precision, supported by a multi-precision MAC structure that avoids extra overhead. Validated on Llama2-7B, OPT-6.7B, and Llama3-8B models, CLAT reduces attention layer latency by up to 88.6% and overall text generation latency by up to 34.9%. It improves single-batch text generation throughput by$1.66\times $to$2.42\times $over an A100 GPU, demonstrating significant performance gains.
Sunwoo Lee 0005, Beomseok Kim, Jeongwoo Park 0001, Dongsuk Jeon
IEEE Trans. Circuits Syst. I Regul. Pap.1
2024 ALAM: Averaged Low-Precision Activation for Memory-Efficient Training of Transformer Models
abstract
One of the key challenges in deep neural network training is the substantial amount of GPU memory required to store activations obtained in the forward pass. Various Activation-Compressed Training (ACT) schemes have been proposed to mitigate this issue; however, it is challenging to adopt those approaches in recent transformer-based large language models (LLMs), which experience significant performance drops when the activations are deeply compressed during training. In this paper, we introduce ALAM, a novel ACT framework that utilizes average quantization and a lightweight sensitivity calculation scheme, enabling large memory saving in LLMs while maintaining training performance. We first demonstrate that compressing activations into their group average values minimizes the gradient variance. Employing this property, we propose Average Quantization which provides high-quality deeply compressed activations with an effective precision of less than 1 bit and improved flexibility of precision allocation. In addition, we present a cost-effective yet accurate sensitivity calculation algorithm that solely relies on the L2 norm of parameter gradients, substantially reducing memory overhead due to sensitivity calculation. In experiments, the ALAM framework significantly reduces activation memory without compromising accuracy, achieving up to a 10$\times$ compression rate in LLMs.
Sunghyeon Woo, Sunwoo Lee 0005, Dongsuk Jeon
ICLR2
2022 Toward Efficient Low-Precision Training: Data Format Optimization and Hysteresis Quantization
Sunwoo Lee 0005, Jeongwoo Park 0001, Dongsuk Jeon
ICLR1
2021 Dynamic Block-Wise Local Learning Algorithm for Efficient Neural Network Training
Gwangho Lee, Sunwoo Lee 0005, Dongsuk Jeon
IEEE Trans. Very Large Scale Integr. Syst.2