Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Nam Joon Kim

dblp:285/0903 · also Namjoon Kim · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
7since 2021 · last 2026
0009-0009-8200-039XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Efficient and distributed learning · 94% Deep learning architectures and training · 6%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
3.042025
LRA-QViT: Integrating Low-Rank Approximation and Quantization for Robust and Efficient Vision Transformers · ICML 2025
Trunk Pruning: Highly Compatible Channel Pruning for Convolutional Neural Networks Without Fine-Tuning · IEEE Trans. Multim. 2024
HyQ: Hardware-Friendly Post-Training Quantization for CNN-Transformer Hybrid Networks · IJCAI 2024
Machine learning › Efficient and distributed learning › model compression
quantization
1.622025
LRA-QViT: Integrating Low-Rank Approximation and Quantization for Robust and Efficient Vision Transformers · ICML 2025
HyQ: Hardware-Friendly Post-Training Quantization for CNN-Transformer Hybrid Networks · IJCAI 2024
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
channel pruning
1.422024
Trunk Pruning: Highly Compatible Channel Pruning for Convolutional Neural Networks Without Fine-Tuning · IEEE Trans. Multim. 2024
FP-AGL: Filter Pruning With Adaptive Gradient Learning for Accelerating Deep Convolutional Neural Networks · IEEE Trans. Multim. 2023
Machine learning › Efficient and distributed learning › model compression
low-rank approximation
0.912025
LRA-QViT: Integrating Low-Rank Approximation and Quantization for Robust and Efficient Vision Transformers · ICML 2025
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization
0.812024
HyQ: Hardware-Friendly Post-Training Quantization for CNN-Transformer Hybrid Networks · IJCAI 2024
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.312025
LRA-QViT: Integrating Low-Rank Approximation and Quantization for Robust and Efficient Vision Transformers · ICML 2025
Machine learning › Deep learning architectures and training › transformer
hybrid CNN-transformer architecture
0.212024
HyQ: Hardware-Friendly Post-Training Quantization for CNN-Transformer Hybrid Networks · IJCAI 2024

Methods — techniques the papers use, named apart from their topics

quantization · 1.6low-rank approximation · 0.9knowledge distillation · 0.9fine-tuning · 0.8batch normalization · 0.8taylor-based method · 0.7centripetal stochastic gradient descent · 0.7
YearPublicationVenuePosition
2026 LUT-APP: Dynamic-Precision LUT-based Approximation Unifying Non-Linear Operations in Transformers
abstract
On-device transformer inference faces a growing bottleneck in which non-linear functions (e.g., exponential (EXP), reciprocal, reciprocal square root, GeLU, and SiLU) contribute significantly to inference latency as matrix operations become highly optimized. Existing approximation methods either rely on operator-specific datapaths with poor hardware reusability or exhibit a suboptimal accuracy-resource balance with conventional look-up table (LUT)-based piecewise linear approximation (PWL) under stringent edge constraints. This work presents LUT-APP, a unified dynamic-precision LUT-based PWL approximation framework that reconciles accuracy and hardware efficiency across diverse non-linear operators. First, a dynamic fixed-point format (DFF) adaptively allocates bit-width based on input magnitude and parameter scaling to handle the wide dynamic range of EXP. Second, a genetic adaptive differential evolution (GADE) algorithm synthesizes non-uniform PWL segments to minimize approximation error for a given LUT budget. Third, hardware-efficient DFF processing units enable a unified INT8 multiply-add datapath, allowing a single reusable implementation across functions. Experimental results demonstrate that LUT-APP reduces approximation error by up to 6.87× versus state-of-the-art methods while preserving baseline accuracy in large language models and vision transformers without fine-tuning. Hardware synthesis with a 28nm technology shows 4.19× lower area and 3.26× lower power savings than existing LUT-based PWL approaches, validating LUT-APP as a practical, resource-constrained solution for on-device accelerators. We provide the LUT-APP implementation at https://github.com/IDSL-SeoulTech/LUT-APP
Seokkyu Yoon, Nam Joon Kim, Hyun Kim 0001
DATE2
2025 LRA-QViT: Integrating Low-Rank Approximation and Quantization for Robust and Efficient Vision Transformers
abstract
Recently, transformer-based models have demonstrated state-of-the-art performance across various computer vision tasks, including image classification, detection, and segmentation. However, their substantial parameter count poses significant challenges for deployment in resource-constrained environments such as edge or mobile devices. Low-rank approximation (LRA) has emerged as a promising model compression technique, effectively reducing the number of parameters in transformer models by decomposing high-dimensional weight matrices into low-rank representations. Nevertheless, matrix decomposition inherently introduces information loss, often leading to a decline in model accuracy. Furthermore, existing studies on LRA largely overlook the quantization process, which is a critical step in deploying practical vision transformer (ViT) models. To address these challenges, we propose a robust LRA framework that preserves weight information after matrix decomposition and incorporates quantization tailored to LRA characteristics. First, we introduce a reparameterizable branch-based low-rank approximation (RB-LRA) method coupled with weight reconstruction to minimize information loss during matrix decomposition. Subsequently, we enhance model accuracy by integrating RB-LRA with knowledge distillation techniques. Lastly, we present an LRA-aware quantization method designed to mitigate the large outliers generated by LRA, thereby improving the robustness of the quantized model. To validate the effectiveness of our approach, we conducted extensive experiments on the ImageNet dataset using various ViT-based models. Notably, the Swin-B model with RB-LRA achieved a 31.8\% reduction in parameters and a 30.4\% reduction in GFLOPs, with only a 0.03\% drop in accuracy. Furthermore, incorporating the proposed LRA-aware quantization method reduced accuracy loss by an additional 0.83\% compared to naive quantization.
Beom Jin Kang, Nam Joon Kim, Hyun Kim 0001
ICML2
2025 FAB: FPGA-Accelerated Fully-Pipelined Bottleneck Architecture With Batching for High-Performance MobileNetv2 Inference
abstract
Lightweight neural networks (LWNNs) primarily employ the bottleneck block (BB) introduced in MobileNetv2 or similar architectural structures. However, the channel expansion-reduction process in BB imposes substantial activation memory overhead, a challenge that has not been adequately addressed in prior studies on LWNN accelerators incorporating BB. To overcome this limitation, we propose a fully-pipelined bottleneck architecture (FPB) optimized for the efficient hardware deployment of BB. FPB eliminates the need for intermediate off-chip memory access, effectively addressing deployment challenges associated with BB and enabling an end-to-end accelerator architecture. To enhance hardware efficiency, each FPB core utilizes 2-LUT DSP, Fused-ReLU6, and Q-Residual, optimizing computational performance while minimizing resource consumption. Furthermore, we introduce a batching technique that maximizes the benefits of FPB by ensuring high hardware utilization across FPB cores while enabling the concurrent processing of multiple images. To mitigate the off-chip memory access latency inherently incurred by batching, we propose a stem layer latency hiding technique, which effectively prevents performance degradation. We evaluate the performance of our proposed MobileNetv2 accelerator on the VCU118 board, achieving an energy efficiency of 120.7 GOPS/W at a batch size of 4. This represents an improvement of$1.5\times $to$10.5\times $over prior work. Depending on the batch size configuration, our FAB accelerator achieves a throughput performance ranging from 204.2 GOPS to 772.7 GOPS, demonstrating its high computational efficiency.
Young Chan Kim, Nam Joon Kim, Hyun Kim 0001
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 HyQ: Hardware-Friendly Post-Training Quantization for CNN-Transformer Hybrid Networks
Nam Joon Kim, Hyun Kim 0001
IJCAI1
2024 Trunk Pruning: Highly Compatible Channel Pruning for Convolutional Neural Networks Without Fine-Tuning
abstract
Channel pruning can efficiently reduce the computation and memory footprint within a reasonable accuracy drop by removing unnecessary channels from convolutional neural networks (CNNs). Among the various channel pruning approaches, sparsity training is the most popular because of its convenient implementation and end-to-end training. It automatically identifies the optimal network structures by applying regularization to parameters. Although this sparsity training has achieved a remarkable performance in terms of the trade-off between accuracy and network size reduction, it needs to be accompanied by a time-consuming fine-tuning process. Moreover, although activation functions with high performance are being continuously developed, the existing sparsity training does not display remarkable scalability for these new activation functions. To address these problems, this study proposes a novel pruning method,trunk pruning, which can produce a compact network by minimizing the accuracy drop during inference even without the fine-tuning process. In the proposed method, one kernel of the next convolutional layer absorbs all the information of the kernels to be pruned, considering the effects of the batch normalization (BN) shift parameters remaining after the sparsity training. Therefore, it is possible to eliminate the fine-tuning process because trunk pruning can effectively reproduce the output of the unpruned network after the sparsity training by removing the pruning loss. Furthermore, because trunk pruning is a technique that can effectively control only the shift parameters of the BN in the CONV layer, it has the significant advantage of being compatible with all BN-based sparsity training schemes and can address various activation functions.
Nam Joon Kim, Hyun Kim 0001
IEEE Trans. Multim.1
2023 RepSGD: Channel Pruning Using Reparamerization for Accelerating Convolutional Neural Networks
abstract
Channel pruning is a popular method for compressing convolutional neural networks (CNNs) while maintaining acceptable accuracy. Most existing channel pruning methods use the approach of zeroing unnecessary filters and then removing them. To address the limitations of existing approaches, methods of creating forcibly filter redundancy and then removing redundant filters have been proposed without heuristic knowledge. However, these methods also use a deformed gradient to make filters identical, and performance degradation is inevitable because the parameters cannot be updated using the original gradients. To solve these problems, this study proposes RepSGD, which can compress CNNs simply and efficiently. RepSGD inserts a new point-wise convolution layer after the existing standard convolution layer. Subsequently, only new point-wise convolution layers are trained to produce filter redundancy (i.e., to make the filters identical), whereas the standard convolution layers are trained using the original gradient. After training, RepSGD merges two consecutive convolution layers into one convolution layer. Subsequently, the redundant filters in the merged convolution layer are pruned. Because RepSGD does not change the original architecture of the CNN, additional inference computation is not required, and it is possible to support training from scratch. In addition, using the original gradient in RepSGD optimizes the objective function of the CNNs better. We show that RepSGD outperforms existing pruning methods in various models and datasets through extensive experiments.
Nam Joon Kim, Hyun Kim 0001
ISCAS1
2023 FP-AGL: Filter Pruning With Adaptive Gradient Learning for Accelerating Deep Convolutional Neural Networks
abstract
Filter pruning is a technique that reduces computational complexity, inference time, and memory footprint by removing unnecessary filters in convolutional neural networks (CNNs) with an acceptable drop in accuracy, consequently accelerating the network. Unlike traditional filter pruning methods utilizing zeroing-out filters, we propose two techniques to achieve the effect of pruning more filters with less performance degradation, inspired by the existing research on centripetal stochastic gradient descent (C-SGD), wherein the filters are removed only when the ones that need to be pruned have the same value. First, to minimize the negative effect of centripetal vectors that gradually make filters come closer to each other, we redesign the vectors by considering the effect of each vector on the loss-function using the Taylor-based method. Second, we propose an adaptive gradient learning (AGL) technique that updates weights while adaptively changing the gradients. Through AGL, performance degradation can be mitigated because some gradients maintain their original direction, and AGL also minimizes the accuracy loss by perfectly converging the filters, which require pruning, to a single point. Finally, we demonstrate the superiority of the proposed method on various datasets and networks. In particular, on the ILSVRC-2012 dataset, our method removed 52.09% FLOPs with a negligible 0.15% top-1 accuracy drop on ResNet-50. As a result, we achieve the most outstanding performance compared to those reported in previous studies in terms of the trade-off between accuracy and computational complexity.
Nam Joon Kim, Hyun Kim 0001
IEEE Trans. Multim.1
2007 Two-Bit Transform Based Block Motion Estimation using Second Derivatives
abstract
A novel two-bit transform based block motion estimation (ME) algorithm is presented in this paper. The proposed approach achieves more effective binarization of image frames than the previous 2BT approach by making use of the positive and negative derivative values separately, which are computed from the second derivatives of a local area as the threshold value for the second bit plane. The second derivatives are also used to find the most accurate motion vectors (MVs) and to reduce computational complexity. Experimental results show that the proposed binary motion estimation algorithm improves motion estimation accuracy and furthermore provides faster processing time in flat or background regions with an acceptable bit-rate increase. In applying the proposed 2BT-SD approach in a real video compression standard, a further reduction of ME processing time with reasonably good compression efficiency is achieved by 2BT-SD based integer ME (IME) followed by full resolution fractional ME.
Nam Joon Kim, Sarp Ertürk
ICME1