Hyeokjun Kwon

dblp:305/4764 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 IterQuant: Iterative Quantization Framework for Mixed-Precision LLM Compression
abstract
Mixed-precision quantization is a promising approach for compressing large language models (LLMs) while maintaining output quality. However, the design space for selecting resolutions of different layers makes exhaustive search intractable. Existing methods either rely on rigid bit-width allocation schemes or require extensive hyperparameter tuning, often based on inaccurate layer-wise sensitivity metrics. In this work, we propose IterQuant, an iterative quantization framework that efficiently explores the mixed-precision space without requiring exhaustive enumeration. By incorporating momentum-based scoring to reflect historical performance trends and parameter grouping to balance quantization granularity, IterQuant achieves favorable trade-offs between compression and accuracy. Unlike prior approaches that assume bit allocation sensitivity from full-precision models directly transfers to quantized models, IterQuant dynamically updates its quantization decisions as the model evolves, better capturing inter-layer dependencies. Experimental results demonstrate that IterQuant significantly outperforms state-of-the-art mixed-precision quantization approaches by 2.8% near 4 bits in preserving token-level output quality across various LLM benchmarks.
Hyungyo Jeong, Hyeokjun Kwon, Jaeho Lee 0001, Youngjoo Lee 0002
DATE3
2025 FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
abstract
Weight-only quantization has emerged as a promising solution to the deployment challenges of large language models (LLMs). However, it necessitates FP-INT operations, which make implementation on general-purpose hardware like GPUs difficult. In this paper, we propose FIGLUT, an efficient look-up table (LUT)-based GEMM accelerator architecture. Instead of performing traditional arithmetic operations, FIGLUT retrieves precomputed values from an LUT based on weight patterns, significantly reducing the computational complexity. We also introduce a novel LUT design that addresses the limitations of conventional memory architectures. To further improve LUT-based operations, we propose a half-size LUT combined with a dedicated decoding and multiplexing unit. FIGLUT efficiently supports different bit precisions and quantization methods using a single fixed hardware configuration. For the same 3-bit weight precision, FIGLUT demonstrates 59% higher TOPS/W and 20% lower perplexity than state-of-the-art accelerator design. When targeting the same perplexity, FIGLUT achieves $98 \%$ higher TOPS/W by performing 2.4-bit operations.
Gunho Park, Hyeokjun Kwon, Jeongin Bae, Baeseong Park, Dongsoo Lee
HPCA2
2023 Learning Point Cloud Completion without Complete Point Clouds: A Pose-Aware Approach
abstract
Point cloud completion is to restore complete 3D scenes and objects from incomplete observations or limited sensor data. Existing fully-supervised methods rely on paired datasets of incomplete and complete point clouds, which are labor-intensive to obtain. Unpaired methods have been proposed, but still require a set of complete point clouds as a reference. As a remedy, in this paper, we propose a novel point cloud completion framework without using any complete point cloud at all. Our main idea is to generate multiple incomplete point clouds of various poses and integrate them into a complete point cloud. We train our framework based on cycle consistency, to generate an incomplete point cloud such that 1) shares the same object as the input incomplete point cloud and 2) corresponds to an arbitrarily given pose. In addition, we devise a novel projection method conditioned by pose to gather visible features, from a volumetric feature extracted by an encoder. Extensive experiments demonstrate that the proposed method achieves comparable or better results than existing unpaired methods. Further, we show that our method also can be applied to real incomplete point clouds.
Hyeokjun Kwon, Yunseo Yang, Kuk-Jin Yoon
ICCV2
2023 RF SSSL by an Autonomous UAV with Two-Ray Channel Model and Dipole Antenna Patterns
abstract
Advancements in unmanned aerial vehicle (UAV) technology have led to their increased utilization in various commercial and military applications. One such application is signal source search and localization (SSSL) using UAVs, which offers significant benefits over traditional ground-based methods due to improved RF signal reception at higher altitudes and inherent autonomous 3D navigation capabilities. Nevertheless, practical considerations such as propagation models and antenna patterns are frequently neglected in simulation-based studies in the literature. In this work, we address these limitations by using a two-ray channel model and a dipole antenna pattern to develop a simulator that more closely represents real-world radio signal strength (RSS) observations at a UAV. We then examine and compare the performance of previously proposed linear least square (LLS) based localization techniques using UAVs for SSSL. Localization of radio frequency (RF) signal sources is assessed based on two main criteria: 1) achieving the highest possible accuracy and 2) localizing the target as quickly as possible with reasonable accuracy. Various mission types, such as those requiring precise localization like identifying hostile troops, and those demanding rapid localization like search and rescue operations during disasters, have been previously investigated. In this paper, the efficacy of the proposed localization approaches is examined based on these two main localization requirements through computer simulations.
Hyeokjun Kwon, Sung Joon Maeng, Ismail Güvenç
PIMRC1
2023 RF Signal Source Search and Localization Using an Autonomous UAV with Predefined Waypoints
abstract
Localization of a radio frequency (RF) signal source has various use cases, ranging from search and rescue, identification and deactivation of jammers, and tracking hostile activity near borders or on the battlefield. The use of unmanned aerial vehicles (UAVs) for signal source search and localization (SSSL) can have significant advantages when compared to terrestrial-based approaches, due to the ease of capturing RF signals at higher altitudes and the autonomous 3D navigation capabilities of UAVs. However, the limited flight duration of UAVs due to battery constraints, as well as limited computational resources on board of lightweight UAVs introduce challenges for SSSL. In this paper, we study various SSSL techniques using a UAV with predefined waypoints.A linear least square (LLS) based localization scheme is considered with enhanced reference selection due to its relatively lower computational complexity. Five different LLS localization algorithms are proposed and studied for selecting anchor positions to be used for localization as the UAV navigates through an area. The performance of each algorithm is measured in two ways: 1) real-time positioning accuracy during the ongoing UAV flight, and 2) long-term accuracy measured at the end of the UAV flight. We compare and analyze the performance of the proposed approaches using computer simulations in terms of accuracy, UAV flight distance, and reliability.
Hyeokjun Kwon, Ismail Güvenç
VTC2023-Spring1
2022 Algorithm-Hardware Co-Optimization for Cost-Efficient ML-based ISP Accelerator
abstract
In this paper, we present an advanced algorithm-hardware co-optimization method for designing an efficient accelerator architecture for image signal processing (ISP) with deep neural networks (DNNs). Based on the systolic-array structure, for performing the target network model, we newly introduce two evaluation metrics, each of which is dedicated to fairly representing either the processing speed or the energy consumption. Then, the overall evaluation metric is defined to test each systolic array, finding the initial array configuration for the given number of total multipliers. From the initial array, several array-scaling methods are then presented to find the most cost-efficient array structure. In addition, the original ML model is adjusted to further enhance the overall efficiency with subtle quality drops of image outputs. Implementation results in 28nm CMOS technology show that the proposed co-optimization method successfully finds the cost-efficient systolic accelerator architecture for ISP applications, improving the energy efficiency by 51% compared to the straightforward array design.
Dongyoung Rim, Hyeokjun Kwon
ISCAS2
2022 CHAMP: Channel Merging Process for Cost-Efficient Highly-Pruned CNN Acceleration
abstract
This paper presents an advanced offline scheduling scheme to improve the accelerator efficiency, especially for the highly-pruned convolutional neural networks (HP-CNNs). Based on the existing outlier-aware accelerator design, we demonstrate the efficiency drop of HP-CNN processing for the first time, and element-wise channel merging is proposed to make a dense processing sequence even for the highly-pruned model. The dedicated hardware architecture is also presented to process the merged channels with the minimum hardware-level overheads, improving the energy efficiency for handling HP-CNNs by preserving the hardware utilization. We further investigate the optimal size of accumulator and multiplexer, in addition to the number of merged channels, exploiting the attractive energy-performance trade-offs. As a result, unlike the practical HP-CNNs for the on-device solutions, the proposed method enhances the overall efficiency by up to 33% compared to the state-of-the-art schemes.
Hyeokjun Kwon, Younghoon Byun, Seokhyeong Kang, Youngjoo Lee 0002
IEEE Trans. Circuits Syst. I Regul. Pap.1
2021 Neighborhood Reconstructing Autoencoders
abstract
Vanilla autoencoders often produce manifolds that overfit to noisy training data, or have the wrong local connectivity and geometry. Autoencoder regularization techniques, e.g., the denoising autoencoder, have had some success in reducing overfitting, whereas recent graph-based methods that exploit local connectivity information provided by neighborhood graphs have had some success in mitigating local connectivity errors. Neither of these two approaches satisfactorily reduce both overfitting and connectivity errors; moreover, graph-based methods typically involve considerable preprocessing and tuning. To simultaneously address the two issues of overfitting and local connectivity, we propose a new graph-based autoencoder, the Neighborhood Reconstructing Autoencoder (NRAE). Unlike existing graph-based methods that attempt to encode the training data to some prescribed latent space distribution -- one consequence being that only the encoder is the object of the regularization -- NRAE merges local connectivity information contained in the neighborhood graphs with local quadratic approximations of the decoder function to formulate a new neighborhood reconstruction loss. Compared to existing graph-based methods, our new loss function is simple and easy to implement, and the resulting algorithm is scalable and computationally efficient; the only required preprocessing step is the construction of the neighborhood graph. Extensive experiments with standard datasets demonstrate that, compared to existing methods, NRAE improves both overfitting and local connectivity in the learned manifold, in some cases by significant margins. Code for NRAE is available at https://github.com/Gabe-YHLee/NRAE-public.
Yonghyeon Lee, Hyeokjun Kwon, Frank C. Park 0001
NeurIPS2