VLDB 2026 Research / reviewers in the wild / expert
Kyungdeuk Ko
dblp:229/0089
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2024
0000-0003-3461-8645ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Prune Channel And Distill: Discriminative Knowledge Distillation For Semantic SegmentationabstractThe goal of knowledge distillation (KD) for semantic segmentation is to transfer discriminative knowledge, enabling the network to distinguish pixels into each class, from a teacher to a student network. Recent KD studies for semantic segmentation fail to convey discriminative knowledge effectively to the student. Consequently, a student network with previous KD cannot generate segmentation maps that effectively distinguish the boundaries of small objects, unlike a teacher network. In this work, we propose a novel KD learning framework, prune channel and distill (PCD), which consists of channel pruning and distillation processes. To transfer the discriminative knowledge of the teacher to the student network, we propose a discriminative score from the perspective of the difference between class responses and student matching distillation, allowing the student to selectively learn channels of pruned feature maps from the teacher. Our PCD directly provides discriminative knowledge from the teacher to the student. In extensive experiments, PCD outperforms state-of-the-art methods on various semantic segmentation datasets. Representative results demonstrate that the proposed method enhances the granularity of the segmentation maps produced by the student network. Bokyeung Lee, Kyungdeuk Ko, Jonghwan Hong, Hanseok Ko |
ICIP | 2 |
| 2024 | Hard Sample-aware Consistency for Low-resolution Facial Expression RecognitionabstractFacial expression recognition (FER) plays a pivotal role in computer vision applications, encompassing video understanding and human-computer interaction. Despite notable advancements in FER, performance still falters when handling low-resolution facial images encountered in real-world scenarios and datasets. While consistency constraint techniques have garnered attention for generating robust convolutional neural network models that accommodate input variations through augmentation, their efficacy is diminished in the realm of low-resolution FER. This decline in performance can be attributed to augmented samples that networks struggle to extract expressive features. In this paper, we identify hard samples that cause an overfitting problem when considering various degrees of resolution and propose novel hard sample-aware consistency (HSAC) loss functions, which include combined attention consistency and label distribution learning. The combined attention consistency aligns an attention map from multi-scale low-resolution images with an appropriate target attention map by combining activation maps from high-resolution and flipped low-resolution images. We measure the classification difficulty for low-resolution face images and adaptively apply label distribution learning by combining the original target and predictions of high-resolution input. Our HSAC empowers the network to achieve generalization by effectively managing hard samples. Extensive experiments on various FER datasets demonstrate the superiority of our proposed method over existing approaches for multiscale low-resolution images. Furthermore, we achieved a new state-of-the-art performance of 90.97% on the original RAF-DB dataset. Bokyeung Lee, Kyungdeuk Ko, Jonghwan Hong, Hanseok Ko |
WACV | 2 |
| 2024 | WaveVC: Speech and Fundamental Frequency Consistent Raw Audio Voice ConversionabstractAbstract Voice conversion (VC) is a task for changing the speech of a source speaker to the target voice while preserving linguistic information of the source speech. The existing VC methods typically use mel-spectrogram as both input and output, so a separate vocoder is required to transform mel-spectrogram into waveform. Therefore, the VC performance varies depending on the vocoder performance, and noisy speech can be generated due to problems such as train-test mismatch. In this paper, we propose a speech and fundamental frequency consistent raw audio voice conversion method called WaveVC. Unlike other methods, WaveVC does not require a separate vocoder and can perform VC directly on raw audio waveform using 1D convolution. This eliminates the issue of performance degradation caused by the train-test mismatch of the vocoder. In the training phase, WaveVC employs speech loss and F0 loss to preserve the content of the source speech and generate F0 consistent speech using the pre-trained networks. WaveVC is capable of converting voices while maintaining consistency in speech and fundamental frequency. In the test phase, the F0 feature of the source speech is concatenated with a content embedding vector to ensure the converted speech follows the fundamental frequency flow of the source speech. WaveVC achieves higher performances than baseline methods in both many-to-many VC and any-to-any VC. The converted samples are available online. Kyungdeuk Ko, Kyungseok Oh, Hanseok Ko |
Neural Process. Lett. | 1 |
| 2024 | KFA: Keyword Feature Augmentation for Open Set Keyword SpottingabstractIn recent years, with the advancement of deep learning technology and the emergence of smart devices, there has been a growing interest in keyword spotting (KWS), which is used to activate AI systems with automatic speech recognition and text-to-speech. However, smart devices with KWS often encounter false alarm errors when inputting unexpected words. To address this issue, existing KWS methods typically train non-target words as anunknownclass. Despite these efforts, there is still a possibility that unseen words not trained as part of theunknownclass could be misclassified as one of the target words. To overcome this limitation, we propose a new method named Keyword Feature Augmentation (KFA) for open-set KWS. KFA performs feature augmentation through adversarial learning to increase the loss. The augmented features are constrained within a limited space using label smoothing. Unlike other generative model-based open set recognition (OSR) methods, KFA does not require any additional training parameters or repeated operation for inference. As a result, KFA has achieved a 0.955 AUROC score and 97.34% target class accuracy for Google Speech Commands V1, and a 0.959 AUROC score and 98.17% target class accuracy for Google Speech Commands V2, which is the highest performance when compared to various OSR methods. Kyungdeuk Ko, Bokyeung Lee, Jonghwan Hong, Hanseok Ko |
IEEE Signal Process. Lett. | 1 |
| 2023 | Domain-agnostic single-image super-resolution via a meta-transfer neural architecture search
Bokyeung Lee, Kyungdeuk Ko, Jonghwan Hong, Hanseok Ko |
Neurocomputing | 2 |
| 2023 | Fast Non-Local Attention network for light super-resolution
Jonghwan Hong, Bokyeung Lee, Kyungdeuk Ko, Hanseok Ko |
J. Vis. Commun. Image Represent. | 3 |
| 2022 | Discriminatory and Orthogonal Feature Learning for Noise Robust Keyword SpottingabstractKeyword Spotting (KWS) is an essential component in a smart device for alerting the system when a user prompts it with a command. As these devices are typically constrained by computational and energy resources, the KWS model should be designed with a small footprint. In our previous work, we developed lightweight dynamic filters which extract a robust feature map within a noisy environment. The learning variables of the dynamic filter are jointly optimized with KWS weights by using Cross-Entropy (CE) loss. CE loss alone, however, is not sufficient for high performance when the SNR is low. In order to train the network for more robust performance in noisy environments, we introduce the LOw Variant Orthogonal (LOVO) loss. The LOVO loss is composed of a triplet loss applied on the output of the dynamic filter, a spectral norm-based orthogonal loss, and an inner class distance loss applied in the KWS model. These losses are particularly useful in encouraging the network to extract discriminatory features in unseen noise environments. Kyungdeuk Ko, David K. Han, Hanseok Ko |
IEEE Signal Process. Lett. | 2 |
| 2022 | Information Bottleneck Measurement for Compressed Sensing Image ReconstructionabstractImage Compressed Sensing (CS) has achieved a lot of performance improvement thanks to advances in deep networks. The CS method is generally composed of a sensing and a decoder. The sensing and decoder networks have a significant impact on the reconstruction performance, and it is obvious that both two networks must be in harmony. However, previous studies have focused on designing the loss function considering only the decoder network. In this paper, we propose a novel training process that can learn sensing and decoder networks simultaneously using Information Bottleneck (IB) theory. By maximizing importance through proposed importance generator, the sensing network is trained to compress important information for image reconstruction of the decoder network. The representative experimental results demonstrate that the proposed method is applied in recently proposed CS algorithms and increases the reconstruction performance with large margin in all CS ratios. Bokyeung Lee, Kyungdeuk Ko, Jonghwan Hong, Bonhwa Ku, Hanseok Ko |
IEEE Signal Process. Lett. | 2 |
| 2021 | Deep Degradation Prior for Real-World Super-Resolution
Kyungdeuk Ko, Bokyeung Lee, Jonghwan Hong, David K. Han, Hanseok Ko |
BMVC | 1 |