Hongje Seong

dblp:231/5155 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
18since 2021 · last 2025
0000-0001-7221-409XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Correlation Verification for Image Retrieval and Its Memory Footprint Optimization
abstract
In this paper, we propose a novel image retrieval network named Correlation Verification Network (CVNet) to replace the conventional geometric re-ranking with a 4D convolutional neural network that learns diverse geometric matching possibilities. To enable efficient cross-scale matching, we construct feature pyramids and establish cross-scale feature correlations in a single inference, thereby replacing the costly multi-scale inference. Additionally, we employ curriculum learning with the Hide-and-Seek strategy to handle challenging samples. Our proposed CVNet demonstrates state-of-the-art performance on several image retrieval benchmarks by a large margin. From an implementation perspective, however, CVNet has one drawback: it requires high memory usage because it needs to store dense features of all database images. This high memory requirement can be a significant limitation in practical applications. To address this issue, we introduce an extension of CVNet called Dense-to-Sparse CVNet (CVNet), which can significantly reduce memory usage by sparsifying the features of the database images. The sparsification module in CVNet learns to select the relevant parts of image features end-to-end using a Gumbel estimator. Since the sparsification is performed offline, CVNet does not increase online extraction and matching times. CVNet dramatically reduces the memory footprint while preserving performance levels nearly identical to CVNet.
Seongwon Lee 0002, Hongje Seong, Suhyeon Lee 0002, Euntai Kim
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Unveiling the Hidden: Online Vectorized HD Map Construction with Clip-Level Token Interaction and Propagation
abstract
Predicting and constructing road geometric information (e.g., lane lines, road markers) is a crucial task for safe autonomous driving, while such static map elements can be repeatedly occluded by various dynamic objects on the road. Recent studies have shown significantly improved vectorized high-definition (HD) map construction performance, but there has been insufficient investigation of temporal information across adjacent input frames (i.e., clips), which may lead to inconsistent and suboptimal prediction results. To tackle this, we introduce a novel paradigm of clip-level vectorized HD map construction, MapUnveiler, which explicitly unveils the occluded map elements within a clip input by relating dense image representations with efficient clip tokens. Additionally, MapUnveiler associates inter-clip information through clip token propagation, effectively utilizing long- term temporal map information. MapUnveiler runs efficiently with the proposed clip-level pipeline by avoiding redundant computation with temporal stride while building a global map relationship. Our extensive experiments demonstrate that MapUnveiler achieves state-of-the-art performance on both the nuScenes and Argoverse2 benchmark datasets. We also showcase that MapUnveiler significantly outperforms state-of-the-art approaches in a challenging setting, achieving +10.7% mAP improvement in heavily occluded driving road scenes. The project page can be found at https://mapunveiler.github.io.
Nayeon Kim 0006, Hongje Seong, Daehyun Ji, Sujin Jang
NeurIPS2
2023 SHUNIT: Style Harmonization for Unpaired Image-to-Image Translation
abstract
We propose a novel solution for unpaired image-to-image (I2I) translation. To translate complex images with a wide range of objects to a different domain, recent approaches often use the object annotations to perform per-class source-to-target style mapping. However, there remains a point for us to exploit in the I2I. An object in each class consists of multiple components, and all the sub-object components have different characteristics. For example, a car in CAR class consists of a car body, tires, windows and head and tail lamps, etc., and they should be handled separately for realistic I2I translation. The simplest solution to the problem will be to use more detailed annotations with sub-object component annotations than the simple object annotations, but it is not possible. The key idea of this paper is to bypass the sub-object component annotations by leveraging the original style of the input image because the original style will include the information about the characteristics of the sub-object components. Specifically, for each pixel, we use not only the per-class style gap between the source and target domains but also the pixel’s original style to determine the target style of a pixel. To this end, we present Style Harmonization for unpaired I2I translation (SHUNIT). Our SHUNIT generates a new style by harmonizing the target domain style retrieved from a class memory and an original source image style. Instead of direct source-to-target style mapping, we aim for source and target styles harmonization. We validate our method with extensive experiments and achieve state-of-the-art performance on the latest benchmark sets. The source code is available online: https://github.com/bluejangbaljang/SHUNIT.
Seokbeom Song, Suhyeon Lee 0002, Hongje Seong, Kyoungwon Min, Euntai Kim
AAAI3
2023 Revisiting Self-Similarity: Structural Embedding for Image Retrieval
abstract
Despite advances in global image representation, existing image retrieval approaches rarely consider geometric structure during the global retrieval stage. In this work, we revisit the conventional self-similarity descriptor from a convolutional perspective, to encode both the visual and structural cues of the image to global image representation. Our proposed network, named Structural Embedding Network (SENet), captures the internal structure of the images and gradually compresses them into dense self-similarity descriptors while learning diverse structures from various images. These self-similarity descriptors and original image features are fused and then pooled into global embedding, so that global embedding can represent both geometric and visual cues of the image. Along with this novel structural embedding, our proposed network sets new state-of-the-art performances on several image retrieval benchmarks, convincing its robustness to look-alike distractors. The code and models are available: https://github.com/sungonce/SENet.
Seongwon Lee 0002, Suhyeon Lee 0002, Hongje Seong, Euntai Kim
CVPR3
2023 Domain Adaptive Video Semantic Segmentation via Cross-Domain Moving Object Mixing
abstract
The network trained for domain adaptation is prone to bias toward the easy-to-transfer classes. Since the ground truth label on the target domain is unavailable during training, the bias problem leads to skewed predictions, forgetting to predict hard-to-transfer classes. To address this problem, we propose Cross-domain Moving Object Mixing (CMOM) that cuts several objects, including hard-to-transfer classes, in the source domain video clip and pastes them into the target domain video clip. Unlike image-level domain adaptation, the temporal context should be maintained to mix moving objects in two different videos. Therefore, we de-sign CMOM to mix with consecutive video frames, so that unrealistic movements are not occurring. We additionally propose Feature Alignment with Temporal Context (FATC) to enhance target domain feature discriminability. FATC exploits the robust source domain features, which are trained with ground truth labels, to learn discriminative target do-main features in an unsupervised manner by filtering unreliable predictions with temporal consensus. We demonstrate the effectiveness of the proposed approaches through extensive experiments. In particular, our model reaches mIoU of 53.81% on VIPER → Cityscapes-Seq benchmark and mIoU of 56.31% on SYNTHIA-Seq → Cityscapes-Seq benchmark, surpassing the state-of-the-art methods by large margins.
Kyusik Cho, Suhyeon Lee 0002, Hongje Seong, Euntai Kim
WACV3
2023 Fallen person detection for autonomous driving
Suhyeon Lee 0002, Sangyong Lee, Hongje Seong, Junhyuk Hyun, Euntai Kim
Expert Syst. Appl.3
2023 Video Object Segmentation Using Kernelized Memory Network With Multiple Kernels
abstract
Semi-supervised video object segmentation (VOS) is to predict the segment of a target object in a video when a ground truth segmentation mask for the target is given in the first frame. Recently, space-time memory networks (STM) have received significant attention as a promising approach for semi-supervised VOS. However, an important point has been overlooked in applying STM to VOS: The solution (=STM) is non-local, but the problem (=VOS) is predominantly local. To solve this mismatch between STM and VOS, we propose new VOS networks called kernelized memory network (KMN) and KMN with multiple kernels (KMN$^{M}$). Our networks conduct not onlyQuery-to-Memorymatching but alsoMemory-to-Querymatching. InMemory-to-Querymatching, a kernel is employed to reduce the degree of non-localness of the STM. In addition, we present a Hide-and-Seek strategy in pre-training to handle occlusions effectively. The proposed networks surpass the state-of-the-art results on standard benchmarks by a significant margin (+4% in$\mathcal {J_{M}}$on DAVIS 2017 test-dev set). The runtimes of our proposed KMN and KMN$^{M}$on DAVIS 2016 validation set are 0.12 and 0.13 seconds per frame, respectively, and the two networks have similar computation times to STM.
Hongje Seong, Junhyuk Hyun, Euntai Kim
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Iteratively Selecting an Easy Reference Frame Makes Unsupervised Video Object Segmentation Easier
abstract
Unsupervised video object segmentation (UVOS) is a per-pixel binary labeling problem which aims at separating the foreground object from the background in the video without using the ground truth (GT) mask of the foreground object. Most of the previous UVOS models use the first frame or the entire video as a reference frame to specify the mask of the foreground object. Our question is why the first frame should be selected as a reference frame or why the entire video should be used to specify the mask. We believe that we can select a better reference frame to achieve the better UVOS performance than using only the first frame or the entire video as a reference frame. In our paper, we propose Easy Frame Selector (EFS). The EFS enables us to select an "easy" reference frame that makes the subsequent VOS become easy, thereby improving the VOS performance. Furthermore, we propose a new framework named as Iterative Mask Prediction (IMP). In the framework, we repeat applying EFS to the given video and selecting an "easier" reference frame from the video than the previous iteration, increasing the VOS performance incrementally. The IMP consists of EFS, Bi-directional Mask Prediction (BMP), and Temporal Information Updating (TIU). From the proposed framework, we achieve state-of-the-art performance in three UVOS benchmark sets: DAVIS16, FBMS, and SegTrack-V2.
Youngjo Lee 0002, Hongje Seong, Euntai Kim
AAAI2
2022 Graph-Based Point Tracker for 3D Object Tracking in Point Clouds
abstract
In this paper, a new deep learning network named as graph-based point tracker (GPT) is proposed for 3D object tracking in point clouds. GPT is not based on Siamese network applied to template and search area, but it is based on the transfer of target clue from the template to the search area. GPT is end-to-end trainable. GPT has two new modules: graph feature augmentation (GFA) and improved target clue (ITC) module. The key idea of GFA is to exploit one-to-many relationship between template and search area points using a bipartite graph. In GFA, edge features of the bipartite graph are generated by transferring the target clues of template points to search area points through edge convolution. It captures the relationship between template and search area points effectively from the perspective of geometry and shape of two point clouds. The second module is ITC. The key idea of ITC is to embed the information of the center of the target into the edges of the bipartite graph via Hough voting, strengthening the discriminative power of GFA. Both modules significantly contribute to the improvement of GPT by transferring geometric and shape information including target center from target template to search area effectively. Experiments on the KITTI tracking dataset show that GPT achieves state-of-the-art performance and can run in real-time.
Minseong Park, Hongje Seong, Wonje Jang, Euntai Kim
AAAI2
2022 WildNet: Learning Domain Generalized Semantic Segmentation from the Wild
abstract
We present a new domain generalized semantic segmentation network named WildNet, which learns domain-generalized features by leveraging a variety of contents and styles from the wild. In domain generalization, the low generalization ability for unseen target domains is clearly due to overfitting to the source domain. To address this problem, previous works have focused on generalizing the domain by removing or diversifying the styles of the source domain. These alleviated overfitting to the source-style but overlooked overfitting to the source-content. In this paper, we propose to diversify both the content and style of the source domain with the help of the wild. Our main idea is for networks to naturally learn domain-generalized semantic information from the wild. To this end, we diversify styles by augmenting source features to resemble wild styles and enable networks to adapt to a variety of styles. Further-more, we encourage networks to learn class-discriminant features by providing semantic variations borrowed from the wild to source contents in the feature space. Finally, we regularize networks to capture consistent semantic information even when both the content and style of the source domain are extended to the wild. Extensive experiments on five different datasets validate the effectiveness of our WildNet, and we significantly outperform state-of-the-art methods. The source code and model are available online: https://github.com/suhyeonlee/WildNet.
Suhyeon Lee 0002, Hongje Seong, Seongwon Lee 0002, Euntai Kim
CVPR2
2022 Correlation Verification for Image Retrieval
abstract
Geometric verification is considered a de facto solution for the re-ranking task in image retrieval. In this study, we propose a novel image retrieval re-ranking network named Correlation Verification Networks (CVNet). Our proposed network, comprising deeply stacked 4D convolutional layers, gradually compresses dense feature correlation into image similarity while learning diverse geometric matching patterns from various image pairs. To enable cross-scale matching, it builds feature pyramids and constructs cross-scale feature correlations within a single inference, replacing costly multi-scale inferences. In addition, we use curriculum learning with the hard negative mining and Hide-and-Seek strategy to handle hard samples without losing generality. Our proposed re-ranking network shows state-of-the-art performance on several retrieval benchmarks with a significant margin (+12.6% in mAP on ROxford-Hard+1M set) over state-of-the-art methods. The source code and models are available online: ht tps: / /gi thub. com/ sungonce/CVNet.
Seongwon Lee 0002, Hongje Seong, Suhyeon Lee 0002, Euntai Kim
CVPR2
2022 One-Trimap Video Matting
Hongje Seong, Seoung Wug Oh, Brian L. Price, Euntai Kim, Joon-Young Lee
ECCV (29)1
2022 Spatial-Channel Transformer for Scene Recognition
abstract
Despite the great success of attention mechanisms on object recognition, scene recognition remains a challenging problem. The reason is that discriminative regions are not evident in a scene image. For example, a tree in an image can be a cue to recognize a scene, but the tree cannot be the only cue for recognizing the scene. That means several scene categories (e.g. mountain, marsh, and river) can contain a tree. Thus sometimes, overall regions, rather than specific regions, need to be considered for scene recognition. To solve the problem, we propose Spatial-Channel Transformer (SC-Transformer). The SC-Transformer is a simple yet effective module that uses a new attention mechanism by incorporating the importance between the spatial and the channel domain for a given scene image. If the given scene image should be considered only within some specific regions, SC-Transformer turns off the channel attention, and vice versa. Furthermore, the attention mechanism used in our proposed method is advanced from previous approaches. Previous spatial and channel attention mechanisms were designed in a sequential or parallel manner. These mechanisms eventually combine spatial and channel attention together, so spatial and channel attention may often interfere with each other. In contrast to the previous works, we present a new mechanism that simultaneously considers spatial and channel attentions. We validate our approach on a large-scale scene recognition dataset and outperform the previous state-of-the-art spatial-channel attention mechanism. Experimental results demonstrate the efficacy of our attention mechanism for scene recognition.
Seunghyun Baik, Hongje Seong, Youngjo Lee 0002, Euntai Kim
IJCNN2
2022 Indoor Place Category Recognition for a Cleaning Robot by Fusing a Probabilistic Approach and Deep Learning
abstract
Indoor place category recognition for a cleaning robot is a problem in which a cleaning robot predicts the category of the indoor place using images captured by it. This is similar to scene recognition in computer vision as well as semantic mapping in robotics. Compared with scene recognition, the indoor place category recognition considered in this article differs as follows: 1) the indoor places include typical home objects; 2) a sequence of images instead of an isolated image is provided because the images are captured successively by a cleaning robot; and 3) the camera of the cleaning robot has a different view compared with those of cameras typically used by human beings. Compared with semantic mapping, indoor place category recognition can be considered as a component in semantic SLAM. In this article, a new method based on the combination of a probabilistic approach and deep learning is proposed to address indoor place category recognition for a cleaning robot. Concerning the probabilistic approach, a new place-object fusion method is proposed based on Bayesian inference. For deep learning, the proposed place-object fusion method is trained using a convolutional neural network in an end-to-end framework. Furthermore, a new recurrent neural network, called the Bayesian filtering network (BFN), is proposed to conduct time-domain fusion. Finally, the proposed method is applied to a benchmark dataset and a new dataset developed in this article, and its validity is demonstrated experimentally.
Soowook Choe, Hongje Seong, Euntai Kim
IEEE Trans. Cybern.2
2022 Adjacent Feature Propagation Network (AFPNet) for Real-Time Semantic Segmentation
abstract
With the development of deep learning, semantic segmentation has received considerable attention within the robotics community. For semantic segmentation to be applied to mobile robots or autonomous vehicles, real-time processing is essential. In this article, a new real-time semantic segmentation network, called the adjacent feature propagation network (AFPNet), is proposed to achieve high performance and fast inference. AFPNet executes in real time on a commercial embedded GPU. The network includes two new modules. The local memory module (LMM) is the first; it improves the upsampling accuracy by propagating the high-level features to the adjacent grids. The cascaded pyramid pooling module (CPPM) is the second; it reduces computational time by changing the structure of the pyramid pooling module. Using these two modules, the proposed AFPNet achieved 76.4% mean intersection-over-union on the Cityscapes test dataset, outperforming other real-time semantic segmentation networks. Furthermore, AFPNet was successfully deployed on an embedded board Jetson AGX Xavier and applied to the real-world navigation of a mobile robot, proving that AFPNet can be effectively used in a variety of real-time applications.
Junhyuk Hyun, Hongje Seong, Sangki Kim, Euntai Kim
IEEE Trans. Syst. Man Cybern. Syst.2
2021 Unsupervised Domain Adaptation for Semantic Segmentation by Content Transfer
abstract
In this paper, we tackle the unsupervised domain adaptation (UDA) for semantic segmentation, which aims to segment the unlabeled real data using labeled synthetic data. The main problem of UDA for semantic segmentation relies on reducing the domain gap between the real image and synthetic image. To solve this problem, we focused on separating information in an image into content and style. Here, only the content has cues for semantic segmentation, and the style makes the domain gap. Thus, precise separation of content and style in an image leads to effect as supervision of real data even when learning with synthetic data. To make the best of this effect, we propose a zero-style loss. Even though we perfectly extract content for semantic segmentation in the real domain, another main challenge, the class imbalance problem, still exists in UDA for semantic segmentation. We address this problem by transferring the contents of tail classes from synthetic to real domain. Experimental results show that the proposed method achieves the state-of-the-art performance in semantic segmentation on the major two UDA settings.
Suhyeon Lee 0002, Junhyuk Hyun, Hongje Seong, Euntai Kim
AAAI3
2021 Hierarchical Memory Matching Network for Video Object Segmentation
abstract
We present Hierarchical Memory Matching Network (HMMN) for semi-supervised video object segmentation. Based on a recent memory-based method [33], we propose two advanced memory read modules that enable us to perform memory reading in multiple scales while exploiting temporal smoothness. We first propose a kernel guided memory matching module that replaces the non-local dense memory read, commonly adopted in previous memory-based methods. The module imposes the temporal smoothness constraint in the memory read, leading to accurate memory retrieval. More importantly, we introduce a hierarchical memory matching scheme and propose a top-k guided memory matching module in which memory read on a fine-scale is guided by that on a coarse-scale. With the module, we perform memory read in multiple scales efficiently and leverage both high-level semantic and low-level fine-grained memory features to predict detailed object masks. Our network achieves state-of-the-art performance on the validation sets of DAVIS 2016/2017 (90.8% and 84.7%) and YouTube-VOS 2018/2019 (82.6% and 82.5%), and test-dev set of DAVIS 2017 (78.6%). The source code and model are available online: https://github.com/Hongje/HMMN.
Hongje Seong, Seoung Wug Oh, Joon-Young Lee, Seongwon Lee 0002, Suhyeon Lee 0002, Euntai Kim
ICCV1
2021 Universal pooling - A new pooling method for convolutional neural networks
Junhyuk Hyun, Hongje Seong, Euntai Kim
Expert Syst. Appl.2
2020 Kernelized Memory Network for Video Object Segmentation
Hongje Seong, Junhyuk Hyun, Euntai Kim
ECCV (22)1
2019 Scene Recognition via Object-to-Scene Class Conversion: End-to-End Training
abstract
When a person recognize the scene of an image, contextual understanding from its environmental elements is necessary. These environmental elements are variant and require comprehensive understanding of various situations. Especially, objects are frequently used as environmental elements related with scene. In this paper, we suggest a score level Class Conversion Matrix (CCM) for scene recognition with a great focus on relationship between objects and scene. A lot of existing methods have already build scene recognition systems with consideration of close relationship between object and scenes. However, most of these methods are using the object features directly without any conversions or reconstructions, and it lack confirmation whether these object features are helpful to recognize scenes correctly. To solve this problem, CCM, a matrix converting object feature to scene feature, is suggested. Moreover, CCM can be implemented with neural network layer and end-to-end trainable. Extensive experiments on Places 2 dataset demonstrate the effectiveness of our approach, when it is applied to the existing deep convolutional neural network architectures. The code is available at https://github.com/Hongje/Class_Conversion_Matrix-Places365
Hongje Seong, Junhyuk Hyun, Hyunbae Chang, Suhyeon Lee 0002, Suhan Woo, Euntai Kim
IJCNN1