Suhwan Cho

dblp:196/4397 · DBLP profile ↗
← Back
27ranked-venue papers
7as first author
26since 2021 · last 2026
0000-0002-1547-807XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 7 first-author · 22 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Generalizing CLIP prompts for zero-shot anomaly detection
Donghyeong Kim, Suhwan Cho, Hyeonjeong Lim, Sangyoun Lee
Pattern Recognit.3
2026 Bidirectional token-masking autoencoder for Referring Image Segmentation
Minhyeok Lee, Dogyoon Lee, Suhwan Cho, Sangyoun Lee
Pattern Recognit.4
2025 Elevating Flow-Guided Video Inpainting with Reference Generation
abstract
Video inpainting (VI) is a challenging task that requires effective propagation of observable content across frames while simultaneously generating new content not present in the original video. In this study, we propose a robust and practical VI framework that leverages a large generative model for reference generation in combination with an advanced pixel propagation algorithm. Powered by a strong generative model, our method not only significantly enhances frame-level quality for object removal but also synthesizes new content in the missing areas based on user-provided text prompts. For pixel propagation, we introduce a one-shot pixel pulling method that effectively avoids error accumulation from repeated sampling while maintaining sub-pixel precision. To evaluate various VI methods in realistic scenarios, we also propose a high-quality VI benchmark, HQVI, comprising carefully generated videos using alpha matte composition. On public benchmarks and the HQVI dataset, our method demonstrates significantly higher visual quality and metric scores compared to existing solutions. Furthermore, it can process high-resolution videos exceeding 2K resolution with ease, underscoring its superiority for real-world applications.
Suhwan Cho, Seoung Wug Oh, Sangyoun Lee, Joon-Young Lee
AAAI1
2025 Video Diffusion Models Are Strong Video Inpainter
abstract
Propagation-based video inpainting using optical flow at the pixel or feature level has recently garnered significant attention. However, it has limitations such as the inaccuracy of optical flow prediction and the propagation of noise over time. These issues result in non-uniform noise and time consistency problems throughout the video, which are particularly pronounced when the removed area is large and involves substantial movement. To address these issues, we propose a novel First Frame Filling Video Diffusion Inpainting model (FFF-VDI). We design FFF-VDI inspired by the capabilities of pre-trained image-to-video diffusion models that can transform the first frame image into a highly natural video. To apply this to the video inpainting task, we propagate the noise latent information of future frames to fill the masked areas of the first frame's noise latent code. Next, we fine-tune the pre-trained image-to-video diffusion model to generate the inpainted video. The proposed model addresses the limitations of existing methods that rely on optical flow quality, producing much more natural and temporally consistent videos. This proposed approach is the first to effectively integrate image-to-video diffusion models into video inpainting tasks. Through various comparative experiments, we demonstrate that the proposed model can robustly handle diverse inpainting types with high quality.
Minhyeok Lee, Suhwan Cho, Chajin Shin, Sunghun Yang, Sangyoun Lee
AAAI2
2025 CoCoGaussian: Leveraging Circle of Confusion for Gaussian Splatting from Defocused Images
abstract
3D Gaussian Splatting (3DGS) has attracted significant attention for its high-quality novel view rendering, inspiring research to address real-world challenges. While conventional methods depend on sharp images for accurate scene reconstruction, real-world scenarios are often affected by defocus blur due to finite depth of field, making it essential to account for realistic 3D scene representation. In this study, we propose CoCoGaussian, a Circle of Confusion-aware Gaussian Splatting that enables precise 3D scene representation using only defocused images. CoCoGaussian addresses the challenge of defocus blur by modeling the Circle of Confusion (CoC) through a physically grounded approach based on the principles of photographic defocus. Exploiting 3D Gaussians, we compute the CoC diameter from depth and learnable aperture information, generating multiple Gaussians to precisely capture the CoC shape. Furthermore, we introduce a learnable scaling factor to enhance robustness and provide more flexibility in handling unreliable depth in scenes with reflective or refractive surfaces. Experiments on both synthetic and real-world datasets demonstrate that CoCoGaussian achieves state-of-the-art performance across multiple benchmarks.
Suhwan Cho, Taeoh Kim, Ho-Deok Jang, Minhyeok Lee, Geonho Cha, Dongyoon Wee, Dogyoon Lee, Sangyoun Lee
CVPR2
2025 Effective SAM Combination for Open-Vocabulary Semantic Segmentation
abstract
Open-vocabulary semantic segmentation aims to assign pixel-level labels to images across an unlimited range of classes. Traditional methods address this by sequentially connecting a powerful mask proposal generator, such as the Segment Anything Model (SAM), with a pre-trained vision-language model like CLIP. But these two-stage approaches often suffer from high computational costs, memory inefficiencies. In this paper, we propose ESC-Net, a novel one-stage open-vocabulary segmentation model that leverages the SAM decoder blocks for class-agnostic segmentation within an efficient inference framework. By embedding pseudo prompts generated from image-text correlations into SAM’s promptable segmentation framework, ESC-Net achieves refined spatial aggregation for accurate mask predictions. Additionally, a Vision-Language Fusion (VLF) module enhances the final mask prediction through image and text guidance. ESC-Net and PASCAL-Context, outperforming prior methods in both efficiency and accuracy. Comprehensive ablation studies further demonstrate its robustness across challenging conditions.
Minhyeok Lee, Suhwan Cho, Sunghun Yang, Heeseung Choi, Ig-Jae Kim, Sangyoun Lee
CVPR2
2025 CoMoGaussian: Continuous Motion-Aware Gaussian Splatting from Motion-Blurred Images
abstract
3D Gaussian Splatting (3DGS) has gained significant attention due to its high-quality novel view rendering, motivating research to address real-world challenges. A critical issue is the camera motion blur caused by movement during exposure, which hinders accurate 3D scene reconstruction. In this study, we propose CoMoGaussian, a Continuous Motion-Aware Gaussian Splatting that reconstructs precise 3D scenes from motion-blurred images while maintaining real-time rendering speed. Considering the complex motion patterns inherent in real-world camera movements, we predict continuous camera trajectories using neural ordinary differential equations (ODEs). To ensure accurate modeling, we employ rigid body transformations, preserving the shape and size of the object but rely on the discrete integration of sampled frames. To better approximate the continuous nature of motion blur, we introduce a continuous motion refinement (CMR) transformation that refines rigid transformations by incorporating additional learnable parameters. By revisiting fundamental camera theory and leveraging advanced neural ODE techniques, we achieve precise modeling of continuous camera trajectories, leading to improved reconstruction accuracy. Extensive experiments demonstrate state-of-the-art performance both quantitatively and qualitatively on benchmark datasets, which include a wide range of motion blur scenarios, from moderate to extreme blur.
Donghyeong Kim, Dogyoon Lee, Suhwan Cho, Minhyeok Lee, Wonjoon Lee, Taeoh Kim, Dongyoon Wee, Sangyoun Lee
ICCV4
2025 CMTM: Cross-Modal Token Modulation for Unsupervised Video Object Segmentation
abstract
Recent advances in unsupervised video object segmentation have highlighted the potential of two-stream architectures that integrate appearance and motion cues. However, fully leveraging these complementary sources of information requires effectively modeling their interdependencies. In this paper, we introduce cross-modality token modulation, a novel approach designed to strengthen the interaction between appearance and motion cues. Our method establishes dense connections between tokens from each modality, enabling efficient intra-modal and inter-modal information propagation through relation transformer blocks. To improve learning efficiency, we incorporate a token masking strategy that addresses the limitations of relying solely on increased model complexity. Our approach achieves state-of-the-art performance across all public benchmarks, outperforming existing methods. The code is released on https://github.com/InSeokJeon/CMTM
Inseok Jeon, Suhwan Cho, Minhyeok Lee, Donghyeong Kim, Sangyoun Lee
ICIP2
2025 Treating Motion as Option With Output Selection for Unsupervised Video Object Segmentation
abstract
Unsupervised video object segmentation aims to detect the most salient object in a video without any external guidance regarding the object. Salient objects often exhibit distinctive movements compared to the background, and recent methods leverage this by combining motion cues from optical flow maps with appearance cues from RGB images. However, because optical flow maps are often closely correlated with segmentation masks, networks can become overly dependent on motion cues during training, leading to vulnerability when faced with confusing motion cues and resulting in unstable predictions. To address this challenge, we propose a novel motion-as-option network that treats motion cues as an optional component rather than a necessity. During training, we randomly input RGB images into the motion encoder instead of optical flow maps, which implicitly reduces the network’s reliance on motion cues. This design ensures that the motion encoder is capable of processing both RGB images and optical flow maps, leading to two distinct predictions depending on the type of input provided. To make the most of this flexibility, we introduce an adaptive output selection algorithm that determines the optimal prediction during testing. Code and models are available at https://github.com/suhwan-cho/TMO.
Suhwan Cho, Minhyeok Lee, MyeongAh Cho, Seungwook Park, Jaeyeob Kim, Hyunsung Jang, Sangyoun Lee
IEEE Trans. Circuits Syst. Video Technol.1
2024 Dual Prototype Attention for Unsupervised Video Object Segmentation
abstract
Unsupervised video object segmentation (VOS) aims to detect and segment the most salient object in videos. The primary techniques used in unsupervised VOS are 1) the collaboration of appearance and motion information; and 2) temporal fusion between different frames. This paper proposes two novel prototype-based attention mechanisms, inter-modality attention (IMA) and inter-frame attention (IFA), to incorporate these techniques via dense propagation across different modalities and frames. IMA densely in-tegrates context information from different modalities based on a mutual refinement. IFA injects global context of a video to the query frame, enabling a full utilization of useful prop-erties from multiple frames. Experimental results on public benchmark datasets demonstrate that our proposed approach outperforms all existing methods by a substantial margin. The proposed two components are also thoroughly validated via ablative study. Code and models are available at https://github.com/Hydragon516/DPA.
Suhwan Cho, Minhyeok Lee, Seunghoon Lee 0008, Dogyoon Lee, Heeseung Choi, Ig-Jae Kim, Sangyoun Lee
CVPR1
2024 Guided Slot Attention for Unsupervised Video Object Segmentation
abstract
Unsupervised video object segmentation aims to segment the most prominent object in a video sequence. However, the existence of complex backgrounds and multiple foreground objects make this task challenging. To address this issue, we propose a guided slot attention network to reinforce spatial structural information and obtain better foreground-background separation. The foreground and background slots, which are initialized with query guidance, are iteratively refined based on interactions with template information. Furthermore, to improve slot-template interaction and effectively fuse global and local features in the target and reference frames, K-nearest neighbors filtering and a feature aggregation transformer are introduced. The proposed model achieves state-of-the-art performance on two popular datasets. Additionally, we demonstrate the robustness of the proposed model in challenging scenes through various comparative experiments. Code and models are available at https://github.com/Hydragon516/GSANet.
Minhyeok Lee, Suhwan Cho, Dogyoon Lee, Sangyoun Lee
CVPR2
2024 LSHNet: Leveraging Structure-Prior With Hierarchical Features Updates for Salient Object Detection in Optical Remote Sensing Images
abstract
Salient object detection in optical remote sensing images (ORSI-SOD) is a task that detects the most prominent context in optical remote sensing images (ORSIs). ORSI-SOD is challenging due to the diverse sizes and shapes of the targets, their irregular distribution in the scene, and the occlusion caused by surrounding environments. Recently, many deep learning-based models have demonstrated promising performance in ORSI-SOD. However, there still remains considerable potential for addressing these challenges in ORSI-SOD. In this article, we propose a novel architecture, LSHNet. We propose a dual-branch architecture consisting of an edge encoder that leverages structure features using edges as the structure-prior and an image encoder that extracts context features from the image. We propose three modules. Image-structure fusion module (ISFM) integrates the two-stream features extracted from dual branch encoders through intrapatch and internal-patch attention mechanisms to utilize diverse receptive fields. Local-global feature fusion module (LGFM) transfers global features representing the target to local feature maps to discriminate the region of the targets from background clutters. The semantic cues updating module (SCUM) updates the representative features of the target from high-level to low-level. By integrating hierarchical information effectively, global features extracted from multilayers can be rectified. We experiment with the three main evaluated datasets in ORSI-SOD: ORSSD, EORSSD, and ORSI-4199. We demonstrate the promising results on the three datasets and analyze the effectiveness of the proposed modules in the ablation study.
Seunghoon Lee 0008, Suhwan Cho, Seungwook Park, Jaeyeob Kim, Sangyoun Lee
IEEE Trans. Geosci. Remote. Sens.2
2023 FAPM: Fast Adaptive Patch Memory for Real-Time Industrial Anomaly Detection
abstract
Feature embedding-based methods have shown exceptional performance in detecting industrial anomalies by comparing features of target images with normal images. However, some methods do not meet the speed requirements of real-time inference, which is crucial for real-world applications. To address this issue, we propose a new method called Fast Adaptive Patch Memory (FAPM) for real-time industrial anomaly detection. FAPM utilizes patch-wise and layer-wise memory banks that store the embedding features of images at the patch and layer level, respectively, which eliminates unnecessary repetitive computations. We also propose patch-wise adaptive coreset sampling for faster and more accurate detection. FAPM performs well in both accuracy and speed compared to other state-of-the-art methods.
Donghyeong Kim, Suhwan Cho, Sangyoun Lee
ICASSP3
2023 Two-Stream Decoder Feature Normality Estimating Network for Industrial Anomaly Detection
abstract
Image reconstruction-based anomaly detection has recently been in the spotlight because of the difficulty of constructing anomaly datasets. These approaches work by learning to model normal features without seeing abnormal samples during training and then discriminating anomalies at test time based on the reconstructive errors. However, these models have limitations in reconstructing the abnormal samples due to their indiscriminate conveyance of features. Moreover, these approaches are not explicitly optimized for distinguishable anomalies. To address these problems, we propose a two-stream decoder network (TSDN), designed to learn both normal and abnormal features. Additionally, we propose a feature normality estimator (FNE) to eliminate abnormal features and prevent high-quality reconstruction of abnormal regions. Evaluation on a standard benchmark demonstrated performance better than state-of-the-art models.
Minhyeok Lee, Suhwan Cho, Donghyeong Kim, Sangyoun Lee
ICASSP3
2023 Leveraging Spatio-Temporal Dependency for Skeleton-Based Action Recognition
abstract
Skeleton-based action recognition has attracted considerable attention due to its compact representation of the human body’s skeletal sructure. Many recent methods have achieved remarkable performance using graph convolutional networks (GCNs) and convolutional neural networks (CNNs), which extract spatial and temporal features, respectively. Although spatial and temporal dependencies in the human skeleton have been explored separately, spatio-temporal dependency is rarely considered. In this paper, we propose the Spatio-Temporal Curve Network (STC-Net) to effectively leverage the spatio-temporal dependency of the human skeleton. Our proposed network consists of two novel elements: 1) The Spatio-Temporal Curve (STC) module; and 2) Dilated Kernels for Graph Convolution (DK-GC). The STC module dynamically adjusts the receptive field by identifying meaningful node connections between every adjacent frame and generating spatio-temporal curves based on the identified node connections, providing an adaptive spatio-temporal coverage. In addition, we propose DK-GC to consider long-range dependencies, which results in a large receptive field without any additional parameters by applying an extended kernel to the given adjacency matrices of the graph. Our STC-Net combines these two modules and achieves state-of-the-art performance on four skeleton-based action recognition benchmarks. Code is available at https://github.com/Jho-Yonsei/STC-Net.
Minhyeok Lee, Suhwan Cho, Sungmin Woo, Sungjun Jang, Sangyoun Lee
ICCV3
2023 Tsanet: Temporal and Scale Alignment for Unsupervised Video Object Segmentation
abstract
Unsupervised Video Object Segmentation (UVOS) refers to the challenging task of segmenting the prominent object in videos without manual guidance. In recent works, two approaches for UVOS have been discussed that can be divided into: appearance and appearance-motion-based methods, which have limitations respectively. Appearance-based methods do not consider the motion of the target object due to exploiting the correlation information between randomly paired frames. Appearance-motion-based methods have the limitation that the dependency on optical flow is dominant due to fusing the appearance with motion. In this paper, we propose a novel framework for UVOS that can address the aforementioned limitations of the two approaches in terms of both time and scale. Temporal Alignment Fusion aligns the saliency information of adjacent frames with the target frame to leverage the information of adjacent frames. Scale Alignment Decoder predicts the target object mask by aggregating multi-scale feature maps via continuous mapping with implicit neural representation. We present experimental results on public benchmark datasets, DAVIS 2016 and FBMS, which demonstrate the effectiveness of our method. Furthermore, we outperform the state-of-the-art methods on DAVIS 2016.
Seunghoon Lee 0008, Suhwan Cho, Dogyoon Lee, Minhyeok Lee, Sangyoun Lee
ICIP2
2023 Adaptive Graph Convolution Module for Salient Object Detection
abstract
Salient object detection (SOD) is a task that involves identifying and segmenting the most visually prominent object in an image. Existing solutions can accomplish this using a multi-scale feature fusion mechanism to detect the global context of an image. However, as there is no consideration of the structures in the image nor the relations between distant pixels, conventional methods cannot deal with complex scenes effectively. In this paper, we propose an adaptive graph convolution module (AGCM) to overcome these limitations. Prototype features are initially extracted from the input image using a learnable region generation layer that spatially groups features in the image. The prototype features are then refined by propagating information between them based on a graph architecture, where each feature is regarded as a node. Experimental results show that the proposed AGCM dramatically improves the SOD performance both quantitatively and quantitatively.
Minhyeok Lee, Suhwan Cho, Sangyoun Lee
ICIP3
2023 Treating Motion as Option to Reduce Motion Dependency in Unsupervised Video Object Segmentation
abstract
Unsupervised video object segmentation (VOS) aims to detect the most salient object in a video sequence at the pixel level. In unsupervised VOS, most state-of-the-art methods leverage motion cues obtained from optical flow maps in addition to appearance cues to exploit the property that salient objects usually have distinctive movements compared to the background. However, as they are overly dependent on motion cues, which may be unreliable in some cases, they cannot achieve stable prediction. To reduce this motion dependency of existing two-stream VOS methods, we propose a novel motion-as-option network that optionally utilizes motion cues. Additionally, to fully exploit the property of the proposed network that motion is not always required, we introduce a collaborative network learning strategy. On all the public benchmark datasets, our proposed network affords state-of-the-art performance with real-time inference speed. Code and models are available at https://github.com/suhwan-cho/TMO.
Suhwan Cho, Minhyeok Lee, Seunghoon Lee 0008, Donghyeong Kim, Sangyoun Lee
WACV1
2023 Unsupervised Video Object Segmentation via Prototype Memory Network
abstract
Unsupervised video object segmentation aims to segment a target object in the video without a ground truth mask in the initial frame. This challenging task requires extracting features for the most salient common objects within a video sequence. This difficulty can be solved by using motion information such as optical flow, but using only the information between adjacent frames results in poor connectivity between distant frames and poor performance. To solve this problem, we propose a novel prototype memory network architecture. The proposed model effectively extracts the RGB and motion information by extracting superpixel-based component prototypes from the input RGB images and optical flow maps. In addition, the model scores the usefulness of the component prototypes in each frame based on a self-learning algorithm and adaptively stores the most useful prototypes in memory and discards obsolete proto-types. We use the prototypes in the memory bank to predict the next query frame’s mask, which enhances the association between distant frames to help with accurate mask prediction. Our method is evaluated on three datasets, achieving state-of-the-art performance. We prove the effectiveness of the proposed model with various ablation studies.
Minhyeok Lee, Suhwan Cho, Seunghoon Lee 0008, Sangyoun Lee
WACV2
2022 Tackling Background Distraction in Video Object Segmentation
Suhwan Cho, Heansung Lee 0001, Minhyeok Lee, Sungjun Jang, Minjung Kim 0002, Sangyoun Lee
ECCV (22)1
2022 SPSN: Superpixel Prototype Sampling Network for RGB-D Salient Object Detection
Minhyeok Lee, Suhwan Cho, Sangyoun Lee
ECCV (29)3
2022 Occluded Person Re-Identification Via Relational Adaptive Feature Correction Learning
abstract
Occluded person re-identification (Re-ID) in images captured by multiple cameras is challenging because the target person is occluded by pedestrians or objects, especially in crowded scenes. In addition to the processes performed during holistic person Re-ID, occluded person Re-ID involves the removal of obstacles and the detection of partially visible body parts. Most existing methods utilize the off-the-shelf pose or parsing networks as pseudo labels, which are prone to error. To address these issues, we propose a novel Occlusion Correction Network (OCNet) that corrects features through relational-weight learning and obtains diverse and representative features without using external networks. In addition, we present a simple concept of a center feature in order to provide an intuitive solution to pedestrian occlusion scenarios. Furthermore, we suggest the idea of Separation Loss (SL) for focusing on different parts between global features and part features. We conduct extensive experiments on five challenging benchmark datasets for occluded and holistic Re-ID tasks to demonstrate that our method achieves superior performance to state-of-the-art methods especially on occluded scene.
Minjung Kim 0002, MyeongAh Cho, Heansung Lee 0001, Suhwan Cho, Sangyoun Lee
ICASSP4
2022 Detection-Identification Balancing Margin Loss for One-Stage Multi-Object Tracking
abstract
In recent years, one-stage multi-object tracking (MOT) methods, which jointly learn detection and identification in a single network, have attracted extensive attention, due to their efficiency. However, the negative transfer effects caused by the two conflicting objectives of detection and identification have rarely been explored. In this paper, we propose a Detection-Identification Balancing Margin (DIM) loss for minimizing the adverse effects caused by these two different objectives. The proposed DIM loss consists of Detection Margin (DM) loss and Identification Margin (IM) loss. DM loss forces features that are farther from the center of the foreground features than the defined margin due to identification learning to be converged to ensure accurate detection. IM loss enables the various feature representations that are essential for identification by intentionally spreading features that become overly clustered due to detection learning. The proposed DIM loss demonstrates competitive and balanced performance for MOT by providing a positive transfer for features that had a strong negative impact on detection and identification, respectively. (HOTA 61.5, MOTA 75.3, IDF1 75.6 on MOT16, and real-time rates of 25.9 fps were achieved)
Heansung Lee 0001, Suhwan Cho, Sungjun Jang, Sungmin Woo, Sangyoun Lee
ICIP2
2022 Superpixel Group-Correlation Network for Co-Saliency Detection
abstract
Co-saliency detection is a task to segment the occurring salient objects in a group of images. The biggest challenges are distracting objects in the background and ambiguity between the foreground and background. To handle these issues, we propose a novel superpixel group-correlation network (SGCN) architecture that uses a superpixel algorithm to obtain various component features from a group of images and creates a group-correlation matrix to detect the common components of those images. In this way, non-common objects can be effectively excluded from consideration, enabling a clear distinction between foreground and background. Our method outperforms current state-of-the-art methods on three popular benchmark datasets for co-saliency detection, and our extensive experiments thoroughly validate our claimed contributions.
Minhyeok Lee, Suhwan Cho, Sangyoun Lee
ICIP3
2022 Pixel-Level Bijective Matching for Video Object Segmentation
abstract
Semi-supervised video object segmentation (VOS) aims to track the designated objects present in the initial frame of a video at the pixel level. To fully exploit the appearance information of an object, pixel-level feature matching is widely used in VOS. Conventional feature matching runs in a surjective manner, i.e., only the best matches from the query frame to the reference frame are considered. Each location in the query frame refers to the optimal location in the reference frame regardless of how often each reference frame location is referenced. This works well in most cases and is robust against rapid appearance variations, but may cause critical errors when the query frame contains background distractors that look similar to the target object. To mitigate this concern, we introduce a bijective matching mechanism to find the best matches from the query frame to the reference frame and vice versa. Before finding the best matches for the query frame pixels, the optimal matches for the reference frame pixels are first considered to prevent each reference frame pixel from being overly referenced. As this mechanism operates in a strict manner, i.e., pixels are connected if and only if they are the sure matches for each other, it can effectively eliminate background distractors. In addition, we propose a mask embedding module to improve the existing mask propagation method. By embedding multiple historic masks with coordinate information, it can effectively capture the position information of a target object. Code and models are available at https://github.com/suhwan-cho/BMVOS.
Suhwan Cho, Heansung Lee 0001, Minjung Kim 0002, Sungjun Jang, Sangyoun Lee
WACV1
2022 Unsupervised video anomaly detection via normalizing flows with implicit latent features
MyeongAh Cho, Taeoh Kim, Woo Jin Kim, Suhwan Cho, Sangyoun Lee
Pattern Recognit.4
2020 Crvos: Clue Refining Network For Video Object Segmentation
abstract
The encoder-decoder based methods for semi-supervised video object segmentation (Semi-VOS) have received extensive attention due to their superior performances. However, most of them have complex intermediate networks which generate strong specifiers to be robust against challenging scenarios, and this is quite inefficient when dealing with relatively simple scenarios. To solve this problem, we propose a real-time network, Clue Refining Network for Video Object Segmentation (CRVOS), that does not have any intermediate network to efficiently deal with these scenarios. In this work, we propose a simple specifier, referred to as the Clue, which consists of the previous frame’s coarse mask and coordinates information. We also propose a novel refine module which shows the better performance compared with the general ones by using a deconvolution layer instead of a bilinear upsampling layer. Our proposed method shows the fastest speed among the existing methods with a competitive accuracy. On DAVIS 2016 validation set, our method achieves 63.5 fps and $\mathcal{J} \& \mathcal{F}$ score of 81.6%.
Suhwan Cho, MyeongAh Cho, Tae-Young Chung, Heansung Lee 0001, Sangyoun Lee
ICIP1