Janghwan Lee

dblp:27/10012 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
10since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2025 RILQ: Rank-Insensitive LoRA-Based Quantization Error Compensation for Boosting 2-Bit Large Language Model Accuracy
abstract
Low-rank adaptation (LoRA) has become the dominant method for parameter-efficient LLM fine-tuning, with LoRA-based quantization error compensation (LQEC) emerging as a powerful tool for recovering accuracy in compressed LLMs. However, LQEC has underperformed in sub-4-bit scenarios, with no prior investigation into understanding this limitation. We propose RILQ (Rank-Insensitive LoRA-based Quantization Error Compensation) to boost 2-bit LLM accuracy. Based on rank analysis revealing model-wise activation discrepancy loss's rank-insensitive nature, RILQ employs this loss to adjust adapters cooperatively across layers, enabling robust error compensation with low-rank adapters. Evaluations on LLaMA-2 and LLaMA-3 demonstrate RILQ's consistent improvements in 2-bit quantized inference across various state-of-the-art quantizers and enhanced accuracy in task-specific fine-tuning. RILQ maintains computational efficiency comparable to existing LoRA methods, enabling adapter-merged weight-quantized LLM inference with significantly enhanced accuracy, making it a promising approach for boosting 2-bit LLM performance.
Geonho Lee, Janghwan Lee, Sukjin Hong, Euijai Ahn, Du-Seong Chang, Jungwook Choi
AAAI2
2024 Improving Conversational Abilities of Quantized Large Language Models via Direct Preference Alignment
abstract
Janghwan Lee, Seongmin Park, Sukjin Hong, Minsoo Kim, Du-Seong Chang, Jungwook Choi. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Janghwan Lee, Seongmin Park 0003, Sukjin Hong, Du-Seong Chang, Jungwook Choi
ACL (1)1
2024 ISP2DLA: Automated Deep Learning Accelerator Design for On-Sensor Image Signal Processing
abstract
Deep neural network-based image signal processing (ISP-DNN) improves image quality with techniques such as demosaicing, but these models pose substantial computational and memory challenges when implemented on CMOS image sensors, particularly due to the high-resolution inputs that increase memory requirements for activations. Layer fusion reduces memory usage by combining consecutive processing steps, yet it increases computational demands, a critical issue in resource-limited on-sensor environments. To address these challenges, we introduce ISP2DLA, an automated deep learning accelerator design framework that balances computational and memory demands for on-sensor ISP. This framework optimizes hardware designs by adjusting line buffer sizes and the number of MAC units, reducing gate counts by 14-79% across two ISP-DNN models, thus enabling efficient on-sensor ISP model inference within constrained resources.
Dong-eon Won, Yeeun Kim, Janghwan Lee, Jonghyun Bae, Jongjoo Park, Jeongyong Song, Jungwook Choi
ASAP3
2024 SPADE: Sparse Pillar-based 3D Object Detection Accelerator for Autonomous Driving
abstract
3D object detection using point cloud (PC) data is essential for perception pipelines of autonomous driving, where efficient encoding is key to meeting stringent resource and latency requirements. PointPillars, a widely adopted bird's-eye view (BEV) encoding, aggregates 3D point cloud data into 2D pillars for fast and accurate 3D object detection. However, the stateof-the-art methods employing PointPillars overlook the inherent sparsity of pillar encoding where only a valid pillar is encoded with a vector of channel elements, missing opportunities for significant computational reduction. Meanwhile, current sparse convolution accelerators are designed to handle only elementwise activation sparsity and do not effectively address the vector sparsity imposed by pillar encoding. In this paper, we propose SPADE, an algorithm-hardware codesign strategy to maximize vector sparsity in pillar-based 3D object detection and accelerate vector-sparse convolution commensurate with the improved sparsity. SPADE consists of three components: (1) a dynamic vector pruning algorithm balancing accuracy and computation savings from vector sparsity, (2) a sparse coordinate management hardware transforming 2D systolic array into a vector-sparse convolution accelerator, and (3) sparsityaware dataflow optimization tailoring sparse convolution schedules for hardware efficiency. Taped-out with a commercial technology, SPADE saves the amount of computation by 36.3–89.2% for representative 3D object detection networks and benchmarks, leading to 1.3–10.9 × speedup and 1.5–12.6 × energy savings compared to the ideal dense accelerator design. These sparsityproportional performance gains equate to 4.1–28.8 × speedup and 90.2–372.3 × energy savings compared to the counterpart server and edge platforms.
Seongmin Park 0003, Minyong Yoon, Janghwan Lee, Nam Sung Kim, Mingu Kang, Jungwook Choi
HPCA5
2024 Patch-aware Vector Quantized Codebook Learning for Unsupervised Visual Defect Detection
abstract
Unsupervised visual defect detection is essential across various industrial applications. Typically, this involves learning a representation space that captures only the features of normal data, and subsequently identifying defects by measuring deviations from this norm. However, balancing the expressiveness and compactness of this space is challenging. The space must be comprehensive enough to encapsulate all regular patterns of normal data, yet without becoming overly expressive, which leads to wasted computational and storage resources and may cause mode collapse-blurring the distinction between normal and defect data embeddings and impairing detection accuracy. To overcome these issues, we introduce a novel approach using an extended VQ-VAE framework optimized for unsupervised defect detection. Unlike traditional methods that apply a constant and uniform representation capacity across an image, our model employs a patch-aware dynamic code assignment scheme. This approach trains the model to allocate codes of varying resolutions based on the context richness of different image regions, aiming to optimize code usage spatially for each sample in a learnable fashion. We also leverage the learned strategy for code allocation in inference, to enlarge the discrepancy between normal and defective samples, thereby improving detection capabilities. Our extensive testing on the MVTecAD, BTAD, and MTSD datasets demonstrates that our model achieves state-of-the-art performance.
Qisen Cheng, Shuhui Qu, Janghwan Lee
ICTAI3
2023 Range-Invariant Approximation of Non-Linear Operations for Efficient BERT Fine-Tuning
abstract
This paper proposes a range-invariant approximation of non-linear operations for training computations of Transformer-based large language models. The proposed method decomposes the approximation into the scaling and the range-invariant resolution for LUT approximation, covering diverse data ranges of non-linear operations with drastically reduced LUT entries during task-dependent BERT fine-tuning. We demonstrate that the proposed method robustly approximates all the non-linear operations of BERT without score degradation on challenging GLUE benchmarks using only a single-entry LUT, facilitating 52% area savings in hardware implementation.
Janghyeon Kim, Janghwan Lee, Jungwook Choi
DAC2
2023 Enhancing Computation Efficiency in Large Language Models through Weight and Activation Quantization
abstract
Large Language Models (LLMs) are proficient in natural language processing tasks, but their deployment is often restricted by extensive parameter sizes and computational demands.This paper focuses on post-training quantization (PTQ) in LLMs, specifically 4-bit weight and 8-bit activation (W4A8) quantization, to enhance computational efficiency-a topic less explored compared to weight-only quantization.We present two innovative techniques: activation-quantization-aware scaling (AQAS) and sequence-length-aware calibration (SLAC) to enhance PTQ by considering the combined effects on weights and activations and aligning calibration sequence lengths to target tasks.Moreover, we introduce dINT, a hybrid data format combining integer and denormal representations, to address the underflow issue in W4A8 quantization, where small values are rounded to zero.Through rigorous evaluations of LLMs, including OPT and LLaMA, we demonstrate that our techniques significantly boost task accuracies to levels comparable with full-precision models.By developing arithmetic units compatible with dINT, we further confirm that our methods yield a 2× hardware efficiency improvement compared to 8-bit integer MAC unit.
Janghwan Lee, Seungcheol Baek, Seok Joong Hwang, Wonyong Sung, Jungwook Choi
EMNLP1
2023 Finding Optimal Numerical Format for Sub-8-Bit Post-Training Quantization of Vision Transformers
abstract
Vision Transformers (ViTs) have gained significant attention for their exceptional model accuracies on computer vision applications, but their demanding memory requirements and computational complexity have hindered active deployment. Post-training quantization (PTQ) is a practical method to tackle this challenge by directly reducing ViT’s bit-precision. However, diverse data characteristics across different operations of ViT cannot be well captured solely by a single numerical format (fixed or floating-point). This work proposes an analytical framework that optimizes the numerical format of each matrix multiplication of ViTs for mixed-format sub-8bit quantization. The extensive evaluation demonstrates that the proposed method can reduce the PTQ error and achieve state-of-the-art accuracy for popular ViT models.
Janghwan Lee, Youngdeok Hwang, Jungwook Choi
ICASSP1
2023 Token-Scaled Logit Distillation for Ternary Weight Generative Language Models
abstract
Generative Language Models (GLMs) have shown impressive performance in tasks such as text generation, understanding, and reasoning. However, the large model size poses challenges for practical deployment. To solve this problem, Quantization-Aware Training (QAT) has become increasingly popular. However, current QAT methods for generative models have resulted in a noticeable loss of accuracy. To counteract this issue, we propose a novel knowledge distillation method specifically designed for GLMs. Our method, called token-scaled logit distillation, prevents overfitting and provides superior learning from the teacher model and ground truth. This research marks the first evaluation of ternary weight quantization-aware training of large-scale GLMs with less than 1.0 degradation in perplexity and achieves enhanced accuracy in tasks like common-sense QA and arithmetic reasoning as well as natural language understanding. Our code is available at https://github.com/aiha-lab/TSLD.
Sihwa Lee, Janghwan Lee, Sukjin Hong, Du-Seong Chang, Wonyong Sung, Jungwook Choi
NeurIPS3
2021 Efficient Multi-Modal Fusion with Diversity Analysis
abstract
Multi-modal machine learning has been a prominent multi-disciplinary research area since its success in complex real-world problems. Empirically, multi-branch fusion models tend to generate better results when there is a high diversity among each branch of the model. However, such experience alone does not guarantee the fusion model's best performance nor have sufficient theoretical support. We present the theoretical estimation of the fusion models' performance by measuring each branch model's performance and the distance between branches based on the analysis of several most popular fusion methods. The theorem is validated empirically by numerical experiments. We further present a branch model selection framework to identify the candidate branches for fusion models to achieve the optimal multi-modal performance by using the theorem. The framework's effectiveness is demonstrated on various datasets by showing how effectively selecting the combination of branch models to attain superior performance.
Shuhui Qu, Yan Kang 0004, Janghwan Lee
ACM Multimedia3
2008 Achieving throughput fairness in Wireless Mesh Networks based on IEEE 802.11
abstract
We propose a fair bandwidth allocation scheme for multi-radio multi-channel wireless mesh networks (WMNs) using distributed algorithm. Through an extensive simulation, we show that our scheme ensures per node fairness without loss of the total aggregate throughput.
Janghwan Lee, Ikjun Yeom
MASS1