EDBT 2026 Demo / reviewers in the wild / expert
Sungeun Hong
dblp:135/5718
· DBLP profile ↗
30ranked-venue papers
4as first author
22since 2021 · last 2026
0000-0003-1774-9168ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 1 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Comprehensive survey on advances and challenges in RGB-D semantic segmentation
Soyun Choi, Eunnam Cho, Aecheon Jung, Byung-Cheol Min, Junhong Min, Sungeun Hong |
Pattern Recognit. | 6 |
| 2026 | Connecting the objects: Relational occlusion graphs for amodal segmentation
Huiling Liu 0001, Eunnam Cho, Soyun Choi, Aecheon Jung, Sungeun Hong |
Pattern Recognit. Lett. | 6 |
| 2025 | Question-Aware Gaussian Experts for Audio-Visual Question AnsweringabstractAudio-Visual Question Answering (AVQA) requires not only question-based multimodal reasoning but also precise temporal grounding to capture subtle dynamics for accurate prediction. However, existing methods mainly use question information implicitly, limiting focus on question-specific details. Furthermore, most studies rely on uniform frame sampling, which can miss key question-relevant frames. Although recent Top-K frame selection methods aim to address this, their discrete nature still overlooks fine-grained temporal details. This paper proposes QA-TIGER, a novel framework that explicitly incorporates question information and models continuous temporal dynamics. Our key idea is to use Gaussian-based modeling to adaptively focus on both consecutive and non-consecutive frames based on the question, while explicitly injecting question information and applying progressive refinement. We leverage a Mixture of Experts (MoE) to flexibly implement multiple Gaussian models, activating temporal experts specifically tailored to the question. Extensive experiments on multiple AVQA benchmarks show that QA-TIGER consistently achieves state-of-the-art performance. Code is available at https://aim-skku.github.io/QA-TIGER/ Hongyeob Kim, Inyoung Jung, Dayoon Suh, Youjia Zhang, Sangmin Lee 0001, Sungeun Hong |
CVPR | 6 |
| 2025 | DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow DecodingabstractHuman motion is inherently continuous and dynamic, posing significant challenges for generative models. While discrete generation methods are widely used, they suffer from limited expressiveness and frame-wise noise artifacts. In contrast, continuous approaches produce smoother, more natural motion but often struggle to adhere to conditioning signals due to high-dimensional complexity and limited training data. To resolve this 'discord' between discrete and continuous representations we introduce DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding, a novel method that leverages rectified flow to decode discrete motion tokens in the continuous, raw motion space. Our core idea is to frame token decoding as a conditional generation task, ensuring that DisCoRD captures fine-grained dynamics and achieves smoother, more natural motions. Compatible with any discrete-based framework, our method enhances naturalness without compromising faithfulness to the conditioning signals on diverse settings. Extensive evaluations demonstrate that DisCoRD achieves state-of-the-art performance, with FID of 0.032 on HumanML3D and 0.169 on KIT-ML. These results establish DisCoRD as a robust solution for bridging the divide between discrete efficiency and continuous realism. Project website: https://whwjdqls.github.io/discord-motion/ Jungbin Cho, Junwan Kim, Jisoo Kim 0006, Mingu Kang, Sungeun Hong, Tae-Hyun Oh, Youngjae Yu |
ICCV | 6 |
| 2025 | Task Vector Quantization for Memory-Efficient Model Merging
Youngeun Kim, Aecheon Jung, Bogon Ryu, Sungeun Hong |
ICCV | 5 |
| 2025 | RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual DataabstractVisuo-tactile perception aims to understand an object's tactile properties. However, the field remains underexplored due to the high cost of data collection. We observe that visually distinct objects can exhibit similar surface textures or material properties. For example, a leather sofa and a leather jacket can share similar tactile properties. This implies that tactile understanding can be guided by material cues in visual data, even without direct tactile supervision. In this paper, we introduce RA-Touch, a retrieval-augmented framework that improves visuo-tactile perception by leveraging visual data enriched with tactile semantics. We carefully recaption a large-scale visual dataset with tactile-focused descriptions, enabling the model to access tactile semantics typically absent from conventional visual datasets. A key challenge remains in effectively utilizing these tactile-aware external descriptions. RA-Touch addresses this by retrieving visual-textual representations aligned with tactile inputs and integrating them to focus on relevant textural and material properties. By outperforming prior methods, we demonstrate the potential of retrieval-based visual reuse for tactile understanding. Code is available at https://aim-skku.github.io/RA-Touch. Yoorhim Cho, Hongyeob Kim, Youjia Zhang, Yunseok Choi, Sungeun Hong |
ACM Multimedia | 6 |
| 2025 | PRIMT: Preference-based Reinforcement Learning with Multimodal Feedback and Trajectory Synthesis from Foundation ModelsabstractPreference-based reinforcement learning (PbRL) has emerged as a promising paradigm for teaching robots complex behaviors without reward engineering. However, its effectiveness is often limited by two critical challenges: the reliance on extensive human input and the inherent difficulties in resolving query ambiguity and credit assignment during reward learning. In this paper, we introduce PRIMT, a PbRL framework designed to overcome these challenges by leveraging foundation models (FMs) for multimodal synthetic feedback and trajectory synthesis. Unlike prior approaches that rely on single-modality FM evaluations, PRIMT employs a hierarchical neuro-symbolic fusion strategy, integrating the complementary strengths of vision-language models (VLMs) and large language models (LLMs) in evaluating robot behaviors for more reliable and comprehensive feedback. PRIMT also incorporates foresight trajectory generation to warm-start the trajectory buffer with bootstrapped samples, reducing early-stage query ambiguity, and hindsight trajectory augmentation for counterfactual reasoning with a causal auxiliary loss to improve credit assignment. We evaluate PRIMT on 2 locomotion and 6 manipulation tasks on various benchmarks, demonstrating superior performance over FM-based and scripted baselines. Website at https://primt25.github.io/. Dezhong Zhao, Ziqin Yuan, Tianyu Shao, Dominic Kao, Sungeun Hong, Byung-Cheol Min |
NeurIPS | 7 |
| 2025 | Backpropagation-Free Test-Time Adaptation via Probabilistic Gaussian AlignmentabstractTest-time adaptation (TTA) enhances the zero-shot robustness under distribution shifts by leveraging unlabeled test data during inference. Despite notable advances, several challenges still limit its broader applicability. First, most methods rely on backpropagation or iterative optimization, which limits scalability and hinders real-time deployment. Second, they lack explicit modeling of class-conditional feature distributions. This modeling is crucial for producing reliable decision boundaries and calibrated predictions, but it remains underexplored due to the lack of both source data and supervision at test time. In this paper, we propose ADAPT, an Advanced Distribution-Aware and backPropagation-free Test-time adaptation method. We reframe TTA as a probabilistic inference task by modeling class-conditional likelihoods using gradually updated class means and a shared covariance matrix. This enables closed-form, training-free inference. To correct potential likelihood bias, we introduce lightweight regularization guided by CLIP priors and a historical knowledge bank. ADAPT requires no source data, no gradient updates, and no full access to target data, supporting both online and transductive settings. Extensive experiments across diverse benchmarks demonstrate that our method achieves state-of-the-art performance under a wide range of distribution shifts with superior scalability and robustness. Youjia Zhang, Youngeun Kim, Young-Geun Choi, Hongyeob Kim, Huiling Liu 0001, Sungeun Hong |
NeurIPS | 6 |
| 2025 | CAT-TPT: Class-Agnostic Text-based Test-time Prompt Tuning for Vision-Language Models
Youjia Zhang, Huiling Liu 0001, Youngeun Kim, Sungeun Hong |
Int. J. Comput. Vis. | 4 |
| 2025 | Perceptual metric for face image quality with pixel-level interpretability
Byungho Jo, In Kyu Park, Sungeun Hong |
Neurocomputing | 3 |
| 2025 | Memory-efficient cross-modal attention for RGB-X segmentation and crowd counting
Youjia Zhang, Soyun Choi, Sungeun Hong |
Pattern Recognit. | 3 |
| 2025 | Prototypical class-wise test-time adaptation
Inyoung Jung, Sungeun Hong |
Pattern Recognit. Lett. | 4 |
| 2024 | FVTTS : Face Based Voice Synthesis for Text-to-Speech
Minyoung Lee 0003, Eunil Park, Sungeun Hong |
INTERSPEECH | 3 |
| 2024 | G-TRACE: Grouped temporal recalibration for video object segmentation
Jooho Kim, Sungeun Hong |
Image Vis. Comput. | 3 |
| 2024 | Scale-aware token-matching for transformer-based object detectorabstractOwing to the advancements in deep learning, object detection has made significant progress in estimating the positions and classes of multiple objects within an image. However, detecting objects of various scales within a single image remains a challenging problem. In this study, we suggest a scale-aware token matching to predict the positions and classes of objects for transformer-based object detection. We train a model by matching detection tokens with ground truth considering its size, unlike the previous methods that performed matching without considering the scale during the training process. We divide one detection token set into multiple sets based on scale and match each token set differently with ground truth, thereby, training the model without additional computation costs. The experimental results demonstrate that scale information can be assigned to tokens. Scale-aware tokens can independently learn scale-specific information by using a novel loss function, which improves the detection performance on small objects. Aecheon Jung, Sungeun Hong, Yoonsuk Hyun |
Pattern Recognit. Lett. | 2 |
| 2023 | Intra-inter Modal Attention Blocks for RGB-D Semantic SegmentationabstractIn this paper, we introduce a novel approach to address the challenge of effectively utilizing both RGB and depth information for semantic segmentation. Our approach, Intra-inter Modal Attention (IMA) blocks, considers both intra-modal and inter-modal aspects of the information to produce better results than prior methods which primarily focused on inter-modal relationships. The IMA blocks consist of a cross-modal non-local module and an adaptive channel-wise fusion module. The cross-modal non-local module captures both intra-modal and inter-modal variations at the spatial level through inter-modality parameter sharing, while the adaptive channel-wise fusion module refines the spatially-correlated features. Experimental results on RGB-D benchmark datasets demonstrate consistent performance improvements over various baseline segmentation networks when using the IMA blocks. Our in-depth analysis provides comprehensive results on the impact of intra-, inter-, and intra-inter modal attention on RGB-D segmentation. Soyun Choi, Youjia Zhang, Sungeun Hong |
ICMR | 3 |
| 2023 | IFQA: Interpretable Face Quality AssessmentabstractExisting face restoration models have relied on general assessment metrics that do not consider the characteristics of facial regions. Recent works have therefore assessed their methods using human studies, which is not scalable and involves significant effort. This paper proposes a novel face-centric metric based on an adversarial framework where a generator simulates face restoration and a discriminator assesses image quality. Specifically, our per-pixel discriminator enables interpretable evaluation that cannot be provided by traditional metrics. Moreover, our metric emphasizes facial primary regions considering that even minor changes to the eyes, nose, and mouth significantly affect human cognition. Our face-oriented metric consistently surpasses existing general or facial image quality assessment metrics by impressive margins. We demonstrate the generalizability of the proposed strategy in various architectural designs and challenging scenarios. Interestingly, we find that our IFQA can lead to performance improvement as an objective function. The code and models are available at https://github.com/VCLLab/IFQA. Byungho Jo, Donghyeon Cho, In Kyu Park, Sungeun Hong |
WACV | 4 |
| 2023 | TL-ADA: Transferable Loss-based Active Domain Adaptation
Kyeongtak Han, Youngeun Kim, Dongyoon Han, Sungeun Hong |
Neural Networks | 5 |
| 2022 | Spatio-Channel Attention Blocks for Cross-modal Crowd Counting
Youjia Zhang, Soyun Choi, Sungeun Hong |
ACCV (2) | 3 |
| 2022 | Adaptive Graph Adversarial Networks for Partial Domain AdaptationabstractThis article tackles Partial Domain Adaptation (PDA) where the target label set is a subset of the source label set. A key challenging issue in PDA is to prevent negative transfer by isolating source-private classes. Since there is no label information for a target domain, PDA methods require to estimate a label commonness score between source and target domains. Existing approaches use either class-level or sample-level commonness to alleviate the negative transfer issue. However, class-level methods assign the same label commonness to all samples of the same class without considering each sample’s characteristics. Also, the recently introduced sample-level approaches show better performance but they still suffer from negative transfer due to non-trivial anomaly samples. To address these limitations, we propose Adaptive Graph Adversarial Networks (AGAN) consisting of two specialized modules. The adaptive class-relational graph module is designed to utilize the intra- and inter-domain structures through adaptive feature propagation. Complementarily, the sample-level commonness predictor computes a commonness score of each sample. Extensive experimental results on public PDA benchmark datasets demonstrate that our structure-aware method outperforms state-of-the-art methods. Youngeun Kim, Sungeun Hong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Re-Aging GAN: Toward Personalized Face Age TransformationabstractFace age transformation aims to synthesize past or future face images by reflecting the age factor on given faces. Ideally, this task should synthesize natural-looking faces across various age groups while maintaining identity. However, most of the existing work has focused on only one of these or is difficult to train while unnatural artifacts still appear. In this work, we propose Re-Aging GAN (RAGAN), a novel single framework considering all the critical factors in age transformation. Our framework achieves state-of-the-art personalized face age transformation by compelling the input identity to perform the self-guidance of the generation process. Specifically, RAGAN can learn the personalized age features by using high-order interactions between given identity and target age. Learned personalized age features are identity information that is recalibrated according to the target age. Hence, such features encompass identity and target age information that provides important clues on how an input identity should be at a certain age. Experimental result shows the lowest FID and KID scores and the highest age recognition accuracy compared to previous methods. The proposed method also demonstrates the visual superiority with fewer artifacts, identity preservation, and natural transformation across various age groups. Farkhod Makhmudkhujaev, Sungeun Hong, In Kyu Park |
ICCV | 2 |
| 2021 | Self-Supervised Feature Enhancement Networks for Small Object Detection in Noisy ImagesabstractRecent CNN-based approaches have shown impressive improvements in object detection, but detecting small objects in images is still a challenging task. Small object detection becomes more difficult if the image contains a lot of noise, which is frequent in real environments. The main reason is that the ratio of visual signal to noise on small objects is very low, making it difficult to extract rich features for detection. To address this issue, we propose a feature enhancement network (FEN) that is trained in a self-supervised manner. Specifically, FEN takes features from input images whose values randomly were erased, then predicts the erased values by aggregating neighboring values. This scheme enables FEN to improve features using surrounding values, which have great effects on enriching features from small-object regions during the test phase. To verify the robustness of our method against small object detection from noisy images, we adopt vehicle detection in aerial images as the main target task. The proposed method consistently outperformed the baseline methods in our experiments. We further present a variety of empirical studies, quantitatively and qualitatively, for in-depth analysis. Geonsoo Lee, Sungeun Hong, Donghyeon Cho |
IEEE Signal Process. Lett. | 2 |
| 2020 | Unsupervised Face Domain Transfer for Low-Resolution Face RecognitionabstractLow-resolution face recognition suffers from domain shift due to the different resolution between a high-resolution gallery and a low-resolution probe set. Conventional methods use the pairwise correlation between high-resolution and low-resolution for the same subject, which requires label information for both gallery and probe sets. However, explicitly labeled low-resolution probe images are seldom available, and labeling them is labor-intensive. In this paper, we propose a novel unsupervised face domain transfer for robust low-resolution face recognition. By leveraging the attention mechanism, the proposed generative face augmentation reduces the domain shift at image-level, while spatial resolution adaptation generates domain-invariant and discriminant feature distributions. On public datasets, we demonstrate the complementarity between generative face augmentation at image-level and spatial resolution adaptation at feature-level. The proposed method outperforms the state-of-the-art supervised methods even though we do not use any label information of low-resolution probe set. Sungeun Hong, Jong Bin Ryu |
IEEE Signal Process. Lett. | 1 |
| 2020 | Towards Privacy-Preserving Domain AdaptationabstractThis study suggests a new domain adaptation paradigm that can address potential data-privacy issues. Despite promising results of existing domain adaptation methods, they have a strong constraint where the source and target samples are accessible during a training phase. However, direct usage of source samples possibly causes data-privacy issues especially when each label of source domain acts as an individual's identifier such as biometric information. To address data-privacy problems in conventional domain adaptation, we propose privacy-preserving domain adaptation (PPDA). Our main hypothesis is that if we train our target model initialized from a pre-trained source model in a self-learning manner, we can successfully transfer knowledge from a labeled source domain to an unlabeled target domain. In our preliminary study, we observe that target samples with low self-entropy measured from the pre-trained source model achieves sufficiently high accuracy. From this key observation, we first select the reliable samples based on self-entropy and define them as class prototypes. We then assign pseudo labels to the target samples through the similarity between target samples and class prototypes. To further reduce the uncertainty of the pseudo labeling process, we also introduce a sample-level reweighting scheme. Surprisingly, our PPDA model outperforms conventional domain adaptation methods on public datasets even though we do not directly access any source data. Youngeun Kim, Donghyeon Cho, Sungeun Hong |
IEEE Signal Process. Lett. | 3 |
| 2018 | Scale-Varying Triplet Ranking with Classification Loss for Facial Age Estimation
Woobin Im, Sungeun Hong, Sung-Eui Yoon, Hyun S. Yang |
ACCV (5) | 2 |
| 2018 | CBVMR: Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure ConstraintabstractUp to now, only limited research has been conducted on crossmodal retrieval of suitable music for a specified video or vice versa. Moreover, much of the existing research relies on metadata such as keywords, tags, or description that must be individually produced and attached posterior. This paper introduces a new content-based, cross-modal retrieval method for video and music that is implemented through deep neural networks. We train the network via inter-modal ranking loss such that videos and music with similar semantics end up close together in the embedding space. However, if only the inter-modal ranking constraint is used for embedding, modality-specific characteristics can be lost. To address this problem, we propose a novel soft intra-modal structure loss that leverages the relative distance relationship between intra-modal samples before embedding. We also introduce reasonable quantitative and qualitative experimental protocols to solve the lack of standard protocols for less-mature video-music related tasks. All the datasets and source code can be found in our online repository (https://github.com/csehong/VM-NET). Sungeun Hong, Woobin Im, Hyun Seung Yang |
ICMR | 1 |
| 2018 | D3: Recognizing dynamic scenes with deep dual descriptor based on key frames and key segments
Sungeun Hong, Jong Bin Ryu, Woobin Im, Hyun Seung Yang |
Neurocomputing | 1 |
| 2017 | SSPP-DAN: Deep domain adaptation network for face recognition with single sample per personabstractReal-world face recognition using a single sample per person (SSPP) is a challenging task. The problem is exacerbated if the conditions under which the gallery image and the probe set are captured are completely different. To address these issues from the perspective of domain adaptation, we introduce an SSPP domain adaptation network (SSPP-DAN). In the proposed approach, domain adaptation, feature extraction, and classification are performed jointly using a deep architecture with domain-adversarial training. However, the SSPP characteristic of one training sample per class is insufficient to train the deep architecture. To overcome this shortage, we generate synthetic images with varying poses using a 3D face model. Experimental evaluations using a realistic SSPP dataset show that deep domain adaptation and image synthesis complement each other and dramatically improve accuracy. Experiments on a benchmark dataset using the proposed approach show state-of-the-art performance. Sungeun Hong, Woobin Im, Jong Bin Ryu, Hyun Seung Yang |
ICIP | 1 |
| 2015 | Sorted Consecutive Local Binary Pattern for Texture ClassificationabstractIn this paper, we propose a sorted consecutive local binary pattern (scLBP) for texture classification. Conventional methods encode only patterns whose spatial transitions are not more than two, whereas scLBP encodes patterns regardless of their spatial transition. Conventional methods do not encode patterns on account of rotation-invariant encoding; on the other hand, patterns with more than two spatial transitions have discriminative power. The proposed scLBP encodes all patterns with any number of spatial transitions while maintaining their rotation-invariant nature by sorting the consecutive patterns. In addition, we introduce dictionary learning of scLBP based on kd-tree which separates data with a space partitioning strategy. Since the elements of sorted consecutive patterns lie in different space, it can be generated to a discriminative code with kd-tree. Finally, we present a framework in which scLBPs and the kd-tree can be combined and utilized. The results of experimental evaluation on five texture data sets--Outex, CUReT, UIUC, UMD, and KTH-TIPS2-a--indicate that our proposed framework achieves the best classification rate on the CUReT, UMD, and KTH-TIPS2-a data sets compared with conventional methods. The results additionally indicate that only a marginal difference exists between the best classification rate of conventional methods and that of the proposed framework on the UIUC and Outex data sets. Jong Bin Ryu, Sungeun Hong, Hyun Seung Yang |
IEEE Trans. Image Process. | 2 |
| 2013 | Recursive Bayesian fire recognition using greedy margin-maximizing clustering
Sujung Bae, Sungeun Hong, Yeong-Jae Choi, Hyun Seung Yang |
Mach. Vis. Appl. | 2 |