Bokyeung Lee

dblp:264/5740 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
13since 2021 · last 2025
0000-0002-6826-6732ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Dropout Connects Transformers and CNNs: Transfer General Knowledge for Knowledge Distillation
abstract
Thanks to their long-range dependencies, transformers obtain state-of-the-art performance in diverse research fields such as computer vision and audio processing. In practical scenarios, convolutional neural networks (CNNs) are used more than Transformers due to their low complexity. So, Transformer-to-CNN knowledge distillation (KD) research, where the Transformer is the teacher and the CNN is the student, is in demand and receiving attention. In Transformer-to-CNN KD training, the capacity gap problem arising from structural differences between the teacher and student networks is the main factor of performance degradation of the student network, unlike homogenous architecture KD. However, previous KD studies transfer all of a teacher's knowledge to the student without consid-ering structural differences. They cannot overcome problems caused by structural differences and show poor performance in Transformer-to-CNN KD. In this paper, we iden-tify general and specific knowledge in feature maps of the teacher and student. General and specific knowledge are the generalized and non-generalized feature representation. We propose a novel KD framework DropKD, which extracts general knowledge from the teacher and student while re-moving specific knowledge and then allows general knowledge of the student network to learn general knowledge of the teacher. Our DropKD empowers the student network to achieve generalization by effectively managing general and specific knowledge. Through extensive experiments on challenging image classification datasets, we demonstrate that the proposed method is superior to existing methods.
Bokyeung Lee, Jonghwan Hong, Hyunuk Shin, Bonhwa Ku, Hanseok Ko
WACV1
2024 Prune Channel And Distill: Discriminative Knowledge Distillation For Semantic Segmentation
abstract
The goal of knowledge distillation (KD) for semantic segmentation is to transfer discriminative knowledge, enabling the network to distinguish pixels into each class, from a teacher to a student network. Recent KD studies for semantic segmentation fail to convey discriminative knowledge effectively to the student. Consequently, a student network with previous KD cannot generate segmentation maps that effectively distinguish the boundaries of small objects, unlike a teacher network. In this work, we propose a novel KD learning framework, prune channel and distill (PCD), which consists of channel pruning and distillation processes. To transfer the discriminative knowledge of the teacher to the student network, we propose a discriminative score from the perspective of the difference between class responses and student matching distillation, allowing the student to selectively learn channels of pruned feature maps from the teacher. Our PCD directly provides discriminative knowledge from the teacher to the student. In extensive experiments, PCD outperforms state-of-the-art methods on various semantic segmentation datasets. Representative results demonstrate that the proposed method enhances the granularity of the segmentation maps produced by the student network.
Bokyeung Lee, Kyungdeuk Ko, Jonghwan Hong, Hanseok Ko
ICIP1
2024 Hard Sample-aware Consistency for Low-resolution Facial Expression Recognition
abstract
Facial expression recognition (FER) plays a pivotal role in computer vision applications, encompassing video understanding and human-computer interaction. Despite notable advancements in FER, performance still falters when handling low-resolution facial images encountered in real-world scenarios and datasets. While consistency constraint techniques have garnered attention for generating robust convolutional neural network models that accommodate input variations through augmentation, their efficacy is diminished in the realm of low-resolution FER. This decline in performance can be attributed to augmented samples that networks struggle to extract expressive features. In this paper, we identify hard samples that cause an overfitting problem when considering various degrees of resolution and propose novel hard sample-aware consistency (HSAC) loss functions, which include combined attention consistency and label distribution learning. The combined attention consistency aligns an attention map from multi-scale low-resolution images with an appropriate target attention map by combining activation maps from high-resolution and flipped low-resolution images. We measure the classification difficulty for low-resolution face images and adaptively apply label distribution learning by combining the original target and predictions of high-resolution input. Our HSAC empowers the network to achieve generalization by effectively managing hard samples. Extensive experiments on various FER datasets demonstrate the superiority of our proposed method over existing approaches for multiscale low-resolution images. Furthermore, we achieved a new state-of-the-art performance of 90.97% on the original RAF-DB dataset.
Bokyeung Lee, Kyungdeuk Ko, Jonghwan Hong, Hanseok Ko
WACV1
2024 Noisy label facial expression recognition via face-specific label distribution learning
Hyunuk Shin, Bokyeung Lee, Bonhwa Ku, Hanseok Ko
Image Vis. Comput.2
2024 KFA: Keyword Feature Augmentation for Open Set Keyword Spotting
abstract
In recent years, with the advancement of deep learning technology and the emergence of smart devices, there has been a growing interest in keyword spotting (KWS), which is used to activate AI systems with automatic speech recognition and text-to-speech. However, smart devices with KWS often encounter false alarm errors when inputting unexpected words. To address this issue, existing KWS methods typically train non-target words as anunknownclass. Despite these efforts, there is still a possibility that unseen words not trained as part of theunknownclass could be misclassified as one of the target words. To overcome this limitation, we propose a new method named Keyword Feature Augmentation (KFA) for open-set KWS. KFA performs feature augmentation through adversarial learning to increase the loss. The augmented features are constrained within a limited space using label smoothing. Unlike other generative model-based open set recognition (OSR) methods, KFA does not require any additional training parameters or repeated operation for inference. As a result, KFA has achieved a 0.955 AUROC score and 97.34% target class accuracy for Google Speech Commands V1, and a 0.959 AUROC score and 98.17% target class accuracy for Google Speech Commands V2, which is the highest performance when compared to various OSR methods.
Kyungdeuk Ko, Bokyeung Lee, Jonghwan Hong, Hanseok Ko
IEEE Signal Process. Lett.2
2023 Domain-agnostic single-image super-resolution via a meta-transfer neural architecture search
Bokyeung Lee, Kyungdeuk Ko, Jonghwan Hong, Hanseok Ko
Neurocomputing1
2023 Fast Non-Local Attention network for light super-resolution
Jonghwan Hong, Bokyeung Lee, Kyungdeuk Ko, Hanseok Ko
J. Vis. Commun. Image Represent.2
2023 Channel Shuffle Neural Architecture Search for Key Word Spotting
abstract
The evolution of Network Architecture (NA) allowed Key-Word Spotting (KWS) to exhibit high performance. Generally, NA for KWS is required to have low parameter and computation complexity maintaining high classification performance. Most of the attempts so far have been based on manual approaches, and often the architectures developed from such efforts dwell in the balance of the performance and the network complexity. Then, several KWS models based on Neural Architecture Search (NAS) technique have been proposed. However, these methods do not consider the number of parameters and FLOPs for NA in the search process and manually adjusted the complexity of NA by reducing the number of cells. It may not produce optimized NA with a balance between network complexity and performance. To develop effective network architecture for KWS, network complexity and performance must be considered. In this letter, we propose Channel Shuffle Neural Architecture Search (CSNAS) with channel weights. CSNAS selects whether each channel of the input feature is reflected in the computation or not in the search process and simultaneously controls the number of parameters, FLOPs, and performance. Experiment results show that CSNAS can generate NA that satisfies complexity and performance conditions, and NAs generated by CSNAS outperform state-of-the-art KWS methods.
Bokyeung Lee, Gwantae Kim, Hanseok Ko
IEEE Signal Process. Lett.1
2022 Feature Sparse Coding With CoordConv for Side Scan Sonar Image Enhancement
abstract
In this letter, we propose a learning-based compressive sensing (CS) algorithm for denoising side scan sonar (SSS) images. The proposed method is a deep learning-based CS method with enhanced nonlinearity based on an iterative shrinkage and thresholding algorithm (ISTA). Since noise intensity varies depending on the position within SSS images, the proposed method also incorporates CoordConv, which provides coordinate information to the network to help remove nonhomogeneous noise. Through end-to-end training, both the deep learning module and the CS characteristics can be jointly optimized. Representative experimental results show that the proposed method is better than state-of-art methods in terms of both noise removal and memory requirements.
Bokyeung Lee, Bonhwa Ku, Wan-Jin Kim, Seungil Kim, Hanseok Ko
IEEE Geosci. Remote. Sens. Lett.1
2022 Prototypical Knowledge Distillation for Noise Robust Keyword Spotting
abstract
Keyword Spotting (KWS) is an essential component in contemporary audio-based deep learning systems and should be of minimal design when the system is working in streaming and on-device environments. We presented a robust feature extraction with a single-layer dynamic convolution model in our previous work. In this letter, we expand our earlier study into multi-layers of operation and propose a robust Knowledge Distillation (KD) learning method. Based on the distribution between class-centroids and embedding vectors, we compute three distinct distance metrics for the KD training and feature extraction processes. The results indicate that our KD method shows similar KWS performance over state-of-the-art models in terms of KWS but with low computational costs. Furthermore, our proposed method results in a more robust performance in noisy environments than conventional KD methods.
Gwantae Kim, Bokyeung Lee, Hanseok Ko
IEEE Signal Process. Lett.3
2022 Information Bottleneck Measurement for Compressed Sensing Image Reconstruction
abstract
Image Compressed Sensing (CS) has achieved a lot of performance improvement thanks to advances in deep networks. The CS method is generally composed of a sensing and a decoder. The sensing and decoder networks have a significant impact on the reconstruction performance, and it is obvious that both two networks must be in harmony. However, previous studies have focused on designing the loss function considering only the decoder network. In this paper, we propose a novel training process that can learn sensing and decoder networks simultaneously using Information Bottleneck (IB) theory. By maximizing importance through proposed importance generator, the sensing network is trained to compress important information for image reconstruction of the decoder network. The representative experimental results demonstrate that the proposed method is applied in recently proposed CS algorithms and increases the reconstruction performance with large margin in all CS ratios.
Bokyeung Lee, Kyungdeuk Ko, Jonghwan Hong, Bonhwa Ku, Hanseok Ko
IEEE Signal Process. Lett.1
2021 Injecting Sparsity in Anomaly Detection for Efficient Inference
abstract
Anomaly detection in the video is a challenging problem in computer vision tasks. Deep networks recently have been successfully applied and achieved competitive performance in anomaly detection. Modern deep networks employ many modules which extract important features. The anomaly detection approaches just developed network architecture and inserted additional networks to improve performance, however, these methods generally require a tremendous amount of computational load and training parameters. Because of limitations in the real world such as field equipment, mobile system, etc., reducing the number of trainable parameters and model capacity is an important issue in anomaly detection. Moreover, the method, which improves the performance of the anomaly detection algorithm, should be developed without additional trainable parameters. In this paper, we propose a sparsity injecting module which reinforces the feature representation of the existing model and presents the abnormality score function using sparsity. In experimental results, our sparsity injecting module improves the performance of state-of-the-art methods without additional trainable parameters.
Bokyeung Lee, Hanseok Ko
AVSS1
2021 Deep Degradation Prior for Real-World Super-Resolution
Kyungdeuk Ko, Bokyeung Lee, Jonghwan Hong, David K. Han, Hanseok Ko
BMVC2