EDBT 2026 Demo / reviewers in the wild / expert
Suhyeon Lee 0002
dblp:231/5154-2
· DBLP profile ↗
12ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0003-1989-9004ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Extended Scene Description using Relationship Information for Manipulating 3D Reconstructed DataabstractIn this paper, we identify a critical limitation in the MPEG-I Scene Description (SD) format: the absence of a native mechanism for representing semantic relationships between objects. To address this, we propose a relationship extension framework that enables both static and dynamic inter-object semantics within 3D scenes. The proposed solution includes two integration strategies: (1) embedding relationship information within existing extensions such as scene interactivity and scene dynamic, and (2) introducing external relationship extensions that preserve the integrity of the core SD specification. We provide a concrete syntax and demonstrate the utility of our approach through applications in scene reconstruction and ISOBMFF-based integration. The results show that both options enable context-aware manipulation and efficient restoration of inter-object relations, supporting advanced interactive and media synchronization scenarios. We propose further evaluation and adoption of these extensions to enhance interoperability and backward compatibility in future MPEG standards. Chae-yeong Song, Chaewon Moon, Aro Kim, Sanghyo Park 0001, Suhyeon Lee 0002, Sungjei Kim |
DCC | 6 |
| 2026 | GQKD: Greedy Query Knowledge Distillation for Efficient 3D Instance Segmentation
Tae Hyun Jeong, Hyeon-Cheol Moon, Suhyeon Lee 0002 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Correlation Verification for Image Retrieval and Its Memory Footprint OptimizationabstractIn this paper, we propose a novel image retrieval network named Correlation Verification Network (CVNet) to replace the conventional geometric re-ranking with a 4D convolutional neural network that learns diverse geometric matching possibilities. To enable efficient cross-scale matching, we construct feature pyramids and establish cross-scale feature correlations in a single inference, thereby replacing the costly multi-scale inference. Additionally, we employ curriculum learning with the Hide-and-Seek strategy to handle challenging samples. Our proposed CVNet demonstrates state-of-the-art performance on several image retrieval benchmarks by a large margin. From an implementation perspective, however, CVNet has one drawback: it requires high memory usage because it needs to store dense features of all database images. This high memory requirement can be a significant limitation in practical applications. To address this issue, we introduce an extension of CVNet called Dense-to-Sparse CVNet (CVNet), which can significantly reduce memory usage by sparsifying the features of the database images. The sparsification module in CVNet learns to select the relevant parts of image features end-to-end using a Gumbel estimator. Since the sparsification is performed offline, CVNet does not increase online extraction and matching times. CVNet dramatically reduces the memory footprint while preserving performance levels nearly identical to CVNet. Seongwon Lee 0002, Hongje Seong, Suhyeon Lee 0002, Euntai Kim |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | SHUNIT: Style Harmonization for Unpaired Image-to-Image TranslationabstractWe propose a novel solution for unpaired image-to-image (I2I) translation. To translate complex images with a wide range of objects to a different domain, recent approaches often use the object annotations to perform per-class source-to-target style mapping. However, there remains a point for us to exploit in the I2I. An object in each class consists of multiple components, and all the sub-object components have different characteristics. For example, a car in CAR class consists of a car body, tires, windows and head and tail lamps, etc., and they should be handled separately for realistic I2I translation. The simplest solution to the problem will be to use more detailed annotations with sub-object component annotations than the simple object annotations, but it is not possible. The key idea of this paper is to bypass the sub-object component annotations by leveraging the original style of the input image because the original style will include the information about the characteristics of the sub-object components. Specifically, for each pixel, we use not only the per-class style gap between the source and target domains but also the pixel’s original style to determine the target style of a pixel. To this end, we present Style Harmonization for unpaired I2I translation (SHUNIT). Our SHUNIT generates a new style by harmonizing the target domain style retrieved from a class memory and an original source image style. Instead of direct source-to-target style mapping, we aim for source and target styles harmonization. We validate our method with extensive experiments and achieve state-of-the-art performance on the latest benchmark sets. The source code is available online: https://github.com/bluejangbaljang/SHUNIT. Seokbeom Song, Suhyeon Lee 0002, Hongje Seong, Kyoungwon Min, Euntai Kim |
AAAI | 2 |
| 2023 | Revisiting Self-Similarity: Structural Embedding for Image RetrievalabstractDespite advances in global image representation, existing image retrieval approaches rarely consider geometric structure during the global retrieval stage. In this work, we revisit the conventional self-similarity descriptor from a convolutional perspective, to encode both the visual and structural cues of the image to global image representation. Our proposed network, named Structural Embedding Network (SENet), captures the internal structure of the images and gradually compresses them into dense self-similarity descriptors while learning diverse structures from various images. These self-similarity descriptors and original image features are fused and then pooled into global embedding, so that global embedding can represent both geometric and visual cues of the image. Along with this novel structural embedding, our proposed network sets new state-of-the-art performances on several image retrieval benchmarks, convincing its robustness to look-alike distractors. The code and models are available: https://github.com/sungonce/SENet. Seongwon Lee 0002, Suhyeon Lee 0002, Hongje Seong, Euntai Kim |
CVPR | 2 |
| 2023 | Domain Adaptive Video Semantic Segmentation via Cross-Domain Moving Object MixingabstractThe network trained for domain adaptation is prone to bias toward the easy-to-transfer classes. Since the ground truth label on the target domain is unavailable during training, the bias problem leads to skewed predictions, forgetting to predict hard-to-transfer classes. To address this problem, we propose Cross-domain Moving Object Mixing (CMOM) that cuts several objects, including hard-to-transfer classes, in the source domain video clip and pastes them into the target domain video clip. Unlike image-level domain adaptation, the temporal context should be maintained to mix moving objects in two different videos. Therefore, we de-sign CMOM to mix with consecutive video frames, so that unrealistic movements are not occurring. We additionally propose Feature Alignment with Temporal Context (FATC) to enhance target domain feature discriminability. FATC exploits the robust source domain features, which are trained with ground truth labels, to learn discriminative target do-main features in an unsupervised manner by filtering unreliable predictions with temporal consensus. We demonstrate the effectiveness of the proposed approaches through extensive experiments. In particular, our model reaches mIoU of 53.81% on VIPER → Cityscapes-Seq benchmark and mIoU of 56.31% on SYNTHIA-Seq → Cityscapes-Seq benchmark, surpassing the state-of-the-art methods by large margins. Kyusik Cho, Suhyeon Lee 0002, Hongje Seong, Euntai Kim |
WACV | 2 |
| 2023 | Fallen person detection for autonomous driving
Suhyeon Lee 0002, Sangyong Lee, Hongje Seong, Junhyuk Hyun, Euntai Kim |
Expert Syst. Appl. | 1 |
| 2022 | WildNet: Learning Domain Generalized Semantic Segmentation from the WildabstractWe present a new domain generalized semantic segmentation network named WildNet, which learns domain-generalized features by leveraging a variety of contents and styles from the wild. In domain generalization, the low generalization ability for unseen target domains is clearly due to overfitting to the source domain. To address this problem, previous works have focused on generalizing the domain by removing or diversifying the styles of the source domain. These alleviated overfitting to the source-style but overlooked overfitting to the source-content. In this paper, we propose to diversify both the content and style of the source domain with the help of the wild. Our main idea is for networks to naturally learn domain-generalized semantic information from the wild. To this end, we diversify styles by augmenting source features to resemble wild styles and enable networks to adapt to a variety of styles. Further-more, we encourage networks to learn class-discriminant features by providing semantic variations borrowed from the wild to source contents in the feature space. Finally, we regularize networks to capture consistent semantic information even when both the content and style of the source domain are extended to the wild. Extensive experiments on five different datasets validate the effectiveness of our WildNet, and we significantly outperform state-of-the-art methods. The source code and model are available online: https://github.com/suhyeonlee/WildNet. Suhyeon Lee 0002, Hongje Seong, Seongwon Lee 0002, Euntai Kim |
CVPR | 1 |
| 2022 | Correlation Verification for Image RetrievalabstractGeometric verification is considered a de facto solution for the re-ranking task in image retrieval. In this study, we propose a novel image retrieval re-ranking network named Correlation Verification Networks (CVNet). Our proposed network, comprising deeply stacked 4D convolutional layers, gradually compresses dense feature correlation into image similarity while learning diverse geometric matching patterns from various image pairs. To enable cross-scale matching, it builds feature pyramids and constructs cross-scale feature correlations within a single inference, replacing costly multi-scale inferences. In addition, we use curriculum learning with the hard negative mining and Hide-and-Seek strategy to handle hard samples without losing generality. Our proposed re-ranking network shows state-of-the-art performance on several retrieval benchmarks with a significant margin (+12.6% in mAP on ROxford-Hard+1M set) over state-of-the-art methods. The source code and models are available online: ht tps: / /gi thub. com/ sungonce/CVNet. Seongwon Lee 0002, Hongje Seong, Suhyeon Lee 0002, Euntai Kim |
CVPR | 3 |
| 2021 | Unsupervised Domain Adaptation for Semantic Segmentation by Content TransferabstractIn this paper, we tackle the unsupervised domain adaptation (UDA) for semantic segmentation, which aims to segment the unlabeled real data using labeled synthetic data. The main problem of UDA for semantic segmentation relies on reducing the domain gap between the real image and synthetic image. To solve this problem, we focused on separating information in an image into content and style. Here, only the content has cues for semantic segmentation, and the style makes the domain gap. Thus, precise separation of content and style in an image leads to effect as supervision of real data even when learning with synthetic data. To make the best of this effect, we propose a zero-style loss. Even though we perfectly extract content for semantic segmentation in the real domain, another main challenge, the class imbalance problem, still exists in UDA for semantic segmentation. We address this problem by transferring the contents of tail classes from synthetic to real domain. Experimental results show that the proposed method achieves the state-of-the-art performance in semantic segmentation on the major two UDA settings. Suhyeon Lee 0002, Junhyuk Hyun, Hongje Seong, Euntai Kim |
AAAI | 1 |
| 2021 | Hierarchical Memory Matching Network for Video Object SegmentationabstractWe present Hierarchical Memory Matching Network (HMMN) for semi-supervised video object segmentation. Based on a recent memory-based method [33], we propose two advanced memory read modules that enable us to perform memory reading in multiple scales while exploiting temporal smoothness. We first propose a kernel guided memory matching module that replaces the non-local dense memory read, commonly adopted in previous memory-based methods. The module imposes the temporal smoothness constraint in the memory read, leading to accurate memory retrieval. More importantly, we introduce a hierarchical memory matching scheme and propose a top-k guided memory matching module in which memory read on a fine-scale is guided by that on a coarse-scale. With the module, we perform memory read in multiple scales efficiently and leverage both high-level semantic and low-level fine-grained memory features to predict detailed object masks. Our network achieves state-of-the-art performance on the validation sets of DAVIS 2016/2017 (90.8% and 84.7%) and YouTube-VOS 2018/2019 (82.6% and 82.5%), and test-dev set of DAVIS 2017 (78.6%). The source code and model are available online: https://github.com/Hongje/HMMN. Hongje Seong, Seoung Wug Oh, Joon-Young Lee, Seongwon Lee 0002, Suhyeon Lee 0002, Euntai Kim |
ICCV | 5 |
| 2019 | Scene Recognition via Object-to-Scene Class Conversion: End-to-End TrainingabstractWhen a person recognize the scene of an image, contextual understanding from its environmental elements is necessary. These environmental elements are variant and require comprehensive understanding of various situations. Especially, objects are frequently used as environmental elements related with scene. In this paper, we suggest a score level Class Conversion Matrix (CCM) for scene recognition with a great focus on relationship between objects and scene. A lot of existing methods have already build scene recognition systems with consideration of close relationship between object and scenes. However, most of these methods are using the object features directly without any conversions or reconstructions, and it lack confirmation whether these object features are helpful to recognize scenes correctly. To solve this problem, CCM, a matrix converting object feature to scene feature, is suggested. Moreover, CCM can be implemented with neural network layer and end-to-end trainable. Extensive experiments on Places 2 dataset demonstrate the effectiveness of our approach, when it is applied to the existing deep convolutional neural network architectures. The code is available at https://github.com/Hongje/Class_Conversion_Matrix-Places365 Hongje Seong, Junhyuk Hyun, Hyunbae Chang, Suhyeon Lee 0002, Suhan Woo, Euntai Kim |
IJCNN | 4 |