VLDB 2026 Research / reviewers in the wild / expert
Hyeongjun Kwon
dblp:305/7933
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
0009-0005-2424-7555ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Faster Parameter-Efficient Tuning with Token Redundancy ReductionabstractParameter-efficient tuning (PET) aims to transfer pre-trained foundation models to downstream tasks by learning a small number of parameters. Compared to traditional fine-tuning, which updates the entire model, PET significantly reduces storage and transfer costs for each task regardless of exponentially increasing pre-trained model capacity. However, most PET methods inherit the inference latency of their large backbone models and often introduce additional computational overhead due to additional modules (e.g. adapters), limiting their practicality for compute-intensive applications. In this paper, we propose Faster Parameter-Efficient Tuning (FPET), a novel approach that enhances inference speed and training efficiency while maintaining high storage efficiency. Specifically, we introduce a plug-and-play token redundancy reduction module delicately designed for PET. This module refines tokens from the self-attention layer using an adapter to learn the accurate similarity between tokens and cuts off the tokens through a fully-differentiable token merging strategy, which uses a straight-through estimator for optimal token reduction. Experimental results prove that our FPET achieves faster inference and higher memory efficiency than the pre-trained backbone while keeping competitive performance on par with state-of-the-art PET methods. The code is available at https://github.com/kyk120/fpet. Kwonyoung Kim, Jungin Park, Jin Kim 0005, Hyeongjun Kwon, Kwanghoon Sohn |
CVPR | 4 |
| 2024 | Improving Visual Recognition with Hyperbolical Visual Hierarchy MappingabstractVisual scenes are naturally organized in a hierarchy, where a coarse semantic is recursively comprised of several fine details. Exploring such a visual hierarchy is crucial to recognize the complex relations of visual elements, leading to a comprehensive scene understanding. In this paper, we propose a Visual Hierarchy Mapper (Hi-Mapper), a novel approach for enhancing the structured understanding of the pre-trained Deep Neural Networks (DNNs). Hi-Mapper investigates the hierarchical organization of the visual scene by 1) pre-defining a hierarchy tree through the encapsulation of probability densities; and 2) learning the hierarchical relations in hyperbolic space with a novel hierarchical contrastive loss. The pre-defined hierarchy tree recursively interacts with the visual features of the pre-trained DNNs through hierarchy decomposition and encoding procedures, thereby effectively identifying the visual hierarchy and enhancing the recognition of an entire scene. Extensive experiments demonstrate that Hi-Mapper significantly enhances the representation capability of DNNs, leading to an improved performance on various tasks, including image classification and dense prediction tasks. The code is available at https://github.com/kwonjunn0l/Hi-Mapper. Hyeongjun Kwon, Jinhyun Jang, Jin Kim 0005, Kwonyoung Kim, Kwanghoon Sohn |
CVPR | 1 |
| 2024 | Enhancing Source-Free Domain Adaptive Object Detection with Low-Confidence Pseudo Label Distillation
Ilhoon Yoon, Hyeongjun Kwon, Jin Kim 0005, Hyunsung Jang, Kwanghoon Sohn |
ECCV (84) | 2 |
| 2024 | Layer-wise Auto-Weighting for Non-Stationary Test-Time AdaptationabstractGiven the inevitability of domain shifts during inference in real-world applications, test-time adaptation (TTA) is essential for model adaptation after deployment. However, the real-world scenario of continuously changing target distributions presents challenges including catastrophic forgetting and error accumulation. Existing TTA methods for non-stationary domain shifts, while effective, incur excessive computational load, making them impractical for on-device settings. In this paper, we introduce a layer-wise auto-weighting algorithm for continual and gradual TTA that autonomously identifies layers for preservation or concentrated adaptation. By leveraging the Fisher Information Matrix (FIM), we first design the learning weight to selectively focus on layers associated with log-likelihood changes while preserving unrelated ones. Then, we further propose an exponential min-max scaler to make certain layers nearly frozen while mitigating outliers. This minimizes forgetting and error accumulation, leading to efficient adaptation to non-stationary target distribution. Experiments on CIFAR-10C, CIFAR-100C, and ImageNet-C show our method outperforms conventional continual and gradual TTA approaches while significantly reducing computational load, highlighting the importance of FIM-based learning weight in adapting to continuously or gradually shifting target domains.1 Jin Kim 0005, Hyeongjun Kwon, Ilhoon Yoon, Kwanghoon Sohn |
WACV | 3 |
| 2023 | Probabilistic Prompt Learning for Dense PredictionabstractRecent progress in deterministic prompt learning has become a promising alternative to various downstream vision tasks, enabling models to learn powerful visual representations with the help of pre-trained vision-language models. However, this approach results in limited performance for dense prediction tasks that require handling more complex and diverse objects, since a single and deterministic description cannot sufficiently represent the entire image. In this paper, we present a novel probabilistic prompt learning to fully exploit the vision-language knowledge in dense prediction tasks. First, we introduce learnable class-agnostic attribute prompts to describe universal attributes across the object class. The attributes are combined with class information and visual-context knowledge to define the class-specific textual distribution. Text representations are sampled and used to guide the dense prediction task using the probabilistic pixel-text matching loss, enhancing the stability and generalization capability of the proposed method. Extensive experiments on different dense prediction tasks and ablation studies demonstrate the effectiveness of our proposed method. Hyeongjun Kwon, Taeyong Song, Somi Jeong, Jin Kim 0005, Jinhyun Jang, Kwanghoon Sohn |
CVPR | 1 |
| 2023 | Knowing Where to Focus: Event-aware Transformer for Video GroundingabstractRecent DETR-based video grounding models have made the model directly predict moment timestamps without any hand-crafted components, such as a pre-defined proposal or non-maximum suppression, by learning moment queries. However, their input-agnostic moment queries inevitably overlook an intrinsic temporal structure of a video, providing limited positional information. In this paper, we formulate an event-aware dynamic moment query to enable the model to take the input-specific content and positional information of the video into account. To this end, we present two levels of reasoning: 1) Event reasoning that captures distinctive event units constituting a given video using a slot attention mechanism; and 2) moment reasoning that fuses the moment queries with a given sentence through a gated fusion transformer layer and learns interactions between the moment queries and video-sentence representations to predict moment timestamps. Extensive experiments demonstrate the effectiveness and efficiency of the event-aware dynamic moment queries, outperforming state-of-the-art approaches on several video grounding benchmarks. The code is publicly available at https://github.com/jinhyunj/EaTR. Jinhyun Jang, Jungin Park, Jin Kim 0005, Hyeongjun Kwon, Kwanghoon Sohn |
ICCV | 4 |
| 2022 | Mask-Guided Attention and Episode Adaptive Weights for Few-Shot SegmentationabstractFew-shot segmentation aims to segment objects with novel classes in a query image, given a support set which consists of few annotated support images. A key factor in few-shot segmentation is to effectively exploit information for the target classes from the support set. In addition, we argue that the overall quality of information available in each training episode varies depending on the given support samples. In this paper, we propose Mask-Guided Attention module to extract more beneficial features for few-shot segmentation from the support images. Taking advantage of the support masks, the area correlated to the foreground object is highlighted and enables the support encoder to extract comprehensive support features with contextual information. Furthermore, we propose Episode Adaptive Weight to balance the training between different episodes. It adaptively adjusts loss weight according to the difficulty of each episode determined by self-supervised segmentation loss of support images and encourages the model to pay more attention to more difficult episodes. Extensive experimental results including comparisons with the state-of-the-art methods and ablation studies demonstrate the effectiveness of the proposed method. Hyeongjun Kwon, Taeyong Song, Sunok Kim, Kwanghoon Sohn |
ICIP | 1 |