EDBT 2026 Demo / reviewers in the wild / expert
Jonghwan Hong
dblp:299/6292
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0001-9975-3958ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Less is more: Efficient Scene Graph Generation with reparameterizationabstractScene Graph Generation (SGG) aims to identify objects and their relationships in visual scenes but faces two key challenges: high computational overhead, particularly for real-time applications, and the long-tailed distribution of predicates, which biases models toward frequent relationships. To address these challenges, we propose Reparams-SGG, a lightweight and efficient network architecture composed of multi-path residual blocks. This architecture reduces computational overhead by leveraging a reparameterization strategy that minimizes sequential and parallel processing, making it highly efficient during inference. Moreover, we introduce a dynamic focal loss that dynamically adjusts the temperature scale during training to focus learning on rare predicates, promoting progressively unbiased learning. Additionally, we propose a dynamic distribution loss, compensating for learning limitations solely from one-hot distributions under data imbalance conditions. We evaluate our method on the widely-used Visual Genome and the recent PSG dataset. Reparams-SGG achieves competitive performance with significantly fewer parameters than state-of-the-art models, demonstrating its efficiency and suitability for deployment in resource-constrained environments. Jonghwan Hong, Seonghyeok Noh, Bonhwa Ku, Hanseok Ko |
ICASSP | 1 |
| 2025 | Dropout Connects Transformers and CNNs: Transfer General Knowledge for Knowledge DistillationabstractThanks to their long-range dependencies, transformers obtain state-of-the-art performance in diverse research fields such as computer vision and audio processing. In practical scenarios, convolutional neural networks (CNNs) are used more than Transformers due to their low complexity. So, Transformer-to-CNN knowledge distillation (KD) research, where the Transformer is the teacher and the CNN is the student, is in demand and receiving attention. In Transformer-to-CNN KD training, the capacity gap problem arising from structural differences between the teacher and student networks is the main factor of performance degradation of the student network, unlike homogenous architecture KD. However, previous KD studies transfer all of a teacher's knowledge to the student without consid-ering structural differences. They cannot overcome problems caused by structural differences and show poor performance in Transformer-to-CNN KD. In this paper, we iden-tify general and specific knowledge in feature maps of the teacher and student. General and specific knowledge are the generalized and non-generalized feature representation. We propose a novel KD framework DropKD, which extracts general knowledge from the teacher and student while re-moving specific knowledge and then allows general knowledge of the student network to learn general knowledge of the teacher. Our DropKD empowers the student network to achieve generalization by effectively managing general and specific knowledge. Through extensive experiments on challenging image classification datasets, we demonstrate that the proposed method is superior to existing methods. Bokyeung Lee, Jonghwan Hong, Hyunuk Shin, Bonhwa Ku, Hanseok Ko |
WACV | 2 |
| 2024 | Prune Channel And Distill: Discriminative Knowledge Distillation For Semantic SegmentationabstractThe goal of knowledge distillation (KD) for semantic segmentation is to transfer discriminative knowledge, enabling the network to distinguish pixels into each class, from a teacher to a student network. Recent KD studies for semantic segmentation fail to convey discriminative knowledge effectively to the student. Consequently, a student network with previous KD cannot generate segmentation maps that effectively distinguish the boundaries of small objects, unlike a teacher network. In this work, we propose a novel KD learning framework, prune channel and distill (PCD), which consists of channel pruning and distillation processes. To transfer the discriminative knowledge of the teacher to the student network, we propose a discriminative score from the perspective of the difference between class responses and student matching distillation, allowing the student to selectively learn channels of pruned feature maps from the teacher. Our PCD directly provides discriminative knowledge from the teacher to the student. In extensive experiments, PCD outperforms state-of-the-art methods on various semantic segmentation datasets. Representative results demonstrate that the proposed method enhances the granularity of the segmentation maps produced by the student network. Bokyeung Lee, Kyungdeuk Ko, Jonghwan Hong, Hanseok Ko |
ICIP | 3 |
| 2024 | Hard Sample-aware Consistency for Low-resolution Facial Expression RecognitionabstractFacial expression recognition (FER) plays a pivotal role in computer vision applications, encompassing video understanding and human-computer interaction. Despite notable advancements in FER, performance still falters when handling low-resolution facial images encountered in real-world scenarios and datasets. While consistency constraint techniques have garnered attention for generating robust convolutional neural network models that accommodate input variations through augmentation, their efficacy is diminished in the realm of low-resolution FER. This decline in performance can be attributed to augmented samples that networks struggle to extract expressive features. In this paper, we identify hard samples that cause an overfitting problem when considering various degrees of resolution and propose novel hard sample-aware consistency (HSAC) loss functions, which include combined attention consistency and label distribution learning. The combined attention consistency aligns an attention map from multi-scale low-resolution images with an appropriate target attention map by combining activation maps from high-resolution and flipped low-resolution images. We measure the classification difficulty for low-resolution face images and adaptively apply label distribution learning by combining the original target and predictions of high-resolution input. Our HSAC empowers the network to achieve generalization by effectively managing hard samples. Extensive experiments on various FER datasets demonstrate the superiority of our proposed method over existing approaches for multiscale low-resolution images. Furthermore, we achieved a new state-of-the-art performance of 90.97% on the original RAF-DB dataset. Bokyeung Lee, Kyungdeuk Ko, Jonghwan Hong, Hanseok Ko |
WACV | 3 |
| 2024 | KFA: Keyword Feature Augmentation for Open Set Keyword SpottingabstractIn recent years, with the advancement of deep learning technology and the emergence of smart devices, there has been a growing interest in keyword spotting (KWS), which is used to activate AI systems with automatic speech recognition and text-to-speech. However, smart devices with KWS often encounter false alarm errors when inputting unexpected words. To address this issue, existing KWS methods typically train non-target words as anunknownclass. Despite these efforts, there is still a possibility that unseen words not trained as part of theunknownclass could be misclassified as one of the target words. To overcome this limitation, we propose a new method named Keyword Feature Augmentation (KFA) for open-set KWS. KFA performs feature augmentation through adversarial learning to increase the loss. The augmented features are constrained within a limited space using label smoothing. Unlike other generative model-based open set recognition (OSR) methods, KFA does not require any additional training parameters or repeated operation for inference. As a result, KFA has achieved a 0.955 AUROC score and 97.34% target class accuracy for Google Speech Commands V1, and a 0.959 AUROC score and 98.17% target class accuracy for Google Speech Commands V2, which is the highest performance when compared to various OSR methods. Kyungdeuk Ko, Bokyeung Lee, Jonghwan Hong, Hanseok Ko |
IEEE Signal Process. Lett. | 3 |
| 2023 | Domain-agnostic single-image super-resolution via a meta-transfer neural architecture search
Bokyeung Lee, Kyungdeuk Ko, Jonghwan Hong, Hanseok Ko |
Neurocomputing | 3 |
| 2023 | Fast Non-Local Attention network for light super-resolution
Jonghwan Hong, Bokyeung Lee, Kyungdeuk Ko, Hanseok Ko |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | Information Bottleneck Measurement for Compressed Sensing Image ReconstructionabstractImage Compressed Sensing (CS) has achieved a lot of performance improvement thanks to advances in deep networks. The CS method is generally composed of a sensing and a decoder. The sensing and decoder networks have a significant impact on the reconstruction performance, and it is obvious that both two networks must be in harmony. However, previous studies have focused on designing the loss function considering only the decoder network. In this paper, we propose a novel training process that can learn sensing and decoder networks simultaneously using Information Bottleneck (IB) theory. By maximizing importance through proposed importance generator, the sensing network is trained to compress important information for image reconstruction of the decoder network. The representative experimental results demonstrate that the proposed method is applied in recently proposed CS algorithms and increases the reconstruction performance with large margin in all CS ratios. Bokyeung Lee, Kyungdeuk Ko, Jonghwan Hong, Bonhwa Ku, Hanseok Ko |
IEEE Signal Process. Lett. | 3 |
| 2021 | CPNet: Cross-Parallel Network for Efficient Anomaly DetectionabstractAnomaly detection in video streams is a challenging problem because of the scarcity of abnormal events and the difficulty of accurately annotating them. To alleviate these issues, unsupervised learning-based prediction methods have been previously applied. These approaches train the model with only normal events and predict a future frame from a sequence of preceding frames by use of encoder-decoder architectures so that they result in small prediction errors on normal events but large errors on abnormal events. The architecture, however, comes with the computational burden as some anomaly detection tasks require low computational cost without sacrificing performance. In this paper, Cross-Parallel Network (CPNet) for efficient anomaly detection is proposed here to minimize computations without performance drops. It consists of N smaller parallel U-Net, each of which is designed to handle a single input frame, to make the calculations significantly more efficient. Additionally, an inter-network shift module is incorporated to capture temporal relationships among sequential frames to enable more accurate future predictions. The quantitative results show that our model requires less computational cost than the baseline U-Net while delivering equivalent performance in anomaly detection. Youngsaeng Jin, Jonghwan Hong, David K. Han, Hanseok Ko |
AVSS | 2 |
| 2021 | Deep Degradation Prior for Real-World Super-Resolution
Kyungdeuk Ko, Bokyeung Lee, Jonghwan Hong, David K. Han, Hanseok Ko |
BMVC | 3 |