EDBT 2026 Demo / reviewers in the wild / expert
Chuan Wang 0001
dblp:68/363-1
· DBLP profile ↗
23ranked-venue papers
5as first author
14since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Reconstruction flow recurrent network for compressed video quality enhancement
Zhengning Wang, Xuhang Liu, Chuan Wang 0001, Ting Jiang 0005, Tianjiao Zeng, Zhenni Zeng, Guoqing Wang 0001, Shuaicheng Liu |
Pattern Recognit. | 3 |
| 2024 | Marching Windows: Scalable Mesh Generation for Volumetric Data With Multiple MaterialsabstractVolumetric data abounds in medical imaging and other fields. With the improved imaging quality and the increased resolution, volumetric datasets are getting so large that the existing tools have become inadequate for processing and analyzing the data. Here we consider the problem of computing tetrahedral meshes to represent large volumetric datasets with labeled multiple materials, which are often encountered in medical imaging or microscopy optical slice tomography. Such tetrahedral meshes are a more compact and expressive geometric representation so are in demand for efficient visualization and simulation of the data, which are impossible if the original large volumetric data are used directly due to the large memory requirement. Existing methods for meshing volumetric data are not scalable for handling large datasets due to their sheer demand on excessively large run-time memory or failure to produce a tet-mesh that preserves the multi-material structure of the original volumetric data. In this article we propose a novel approach, called Marching Windows, that uses a moving window and a disk-swap strategy to reduce the run-time memory footprint, devise a new scheme that guarantees to preserve the topological structure of the original dataset, and adopt an error-guided optimization technique to improve both geometric approximation error and mesh quality. Extensive experiments show that our method is capable of processing very large volumetric datasets beyond the capability of the existing methods and producing tetrahedral meshes of high quality. Ya-Ting Yue, Hao Pan 0001, Zhonggui Chen, Chuan Wang 0001, Hanspeter Pfister, Wenping Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | MEFLUT: Unsupervised 1D Lookup Tables for Multi-exposure Image FusionabstractIn this paper, we introduce a new approach for high-quality multi-exposure image fusion (MEF). We show that the fusion weights of an exposure can be encoded into a 1D lookup table (LUT), which takes pixel intensity value as input and produces fusion weight as output. We learn one 1D LUT for each exposure, then all the pixels from different exposures can query 1D LUT of that exposure independently for high-quality and efficient fusion. Specifically, to learn these 1D LUTs, we involve attention mechanism in various dimensions including frame, channel and spatial ones into the MEF task so as to bring us significant quality improvement over the state-of-the-art (SOTA). In addition, we collect a new MEF dataset consisting of 960 samples, 155 of which are manually tuned by professionals as ground-truth for evaluation. Our network is trained by this dataset in an unsupervised manner. Extensive experiments are conducted to demonstrate the effectiveness of all the newly proposed components, and results show that our approach outperforms the SOTA in our and another representative dataset SICE, both qualitatively and quantitatively. Moreover, our 1D LUT approach takes less than 4ms to run a 4K image on a PC GPU. Given its high quality, efficiency and robustness, our method has been shipped into millions of Android mobiles across multiple brands world-wide. Code is available at: https://github.com/Hedlen/MEFLUT. Ting Jiang 0005, Chuan Wang 0001, Xinpeng Li 0002, Ru Li 0002, Haoqiang Fan, Shuaicheng Liu |
ICCV | 2 |
| 2023 | Enhanced Soft Label for Semi-Supervised Semantic SegmentationabstractAs a mainstream framework in the field of semi-supervised learning (SSL), self-training via pseudo labeling and its variants have witnessed impressive progress in semi-supervised semantic segmentation with the recent advance of deep neural networks. However, modern self-training based SSL algorithms use a pre-defined constant threshold to select unlabeled pixel samples that contribute to the training, thus failing to be compatible with different learning difficulties of variant categories and different learning status of the model. To address these issues, we propose Enhanced Soft Label (ESL), a curriculum learning approach to fully leverage the high-value supervisory signals implicit in the untrustworthy pseudo label. ESL believes that pixels with unconfident predictions can be pretty sure about their belonging to a subset of dominant classes though being arduous to determine the exact one. It thus contains a Dynamic Soft Label (DSL) module to dynamically maintain the high probability classes, keeping the label "soft" so as to make full use of the high entropy prediction. However, the DSL itself will inevitably introduce ambiguity between dominant classes, thus blurring the classification boundary. Therefore, we further propose a pixel-to-part contrastive learning method cooperated with an unsupervised object part grouping mechanism to improve its ability to distinguish between different classes. Extensive experimental results on Pascal VOC 2012 and Cityscapes show that our approach achieves remarkable improvements over existing state-of-the-art approaches. Chuan Wang 0001, Yang Liu 0267, Liang Lin 0004, Guanbin Li |
ICCV | 2 |
| 2023 | Hyperspectral Image Mixed Noise Removal via Nonlinear Transform-Based Block-Term Tensor DecompositionabstractRecently, block-term decomposition with rank-(Lr,Lr,1) (termed as LL1 decomposition), which is physically inspired by linear spectral unmixing, has received increasing attention in hyperspectral images (HSIs) denoising. However, due to the intrinsic nonlinear structure of real-world HSIs, the low-rankness of HSIs is usually implicit. Moreover, the essential uniqueness guarantee is usually violated with the low-rank assumption of the abundance maps unsupported in real scenarios, which hampers the successful deployment of LL1 decomposition. Inspired by the nonlinear spectral unmixing, we propose a nonlinear learnable transform-based LL1 decomposition (NT-LL1) for characterizing the implicit low-rank structure of real-world HSIs. More concretely, the nonlinear learnable transform in NT-LL1 decomposition is a composed transform consisting of a linear semi-orthogonal transform and a component-wise nonlinear transform, which collaboratively enhances the low-rankness of the abundance maps. Empowering with the NT-LL1 decomposition, we propose an NT-LL1 decomposition-based model for HSIs denoising. To tackle the resulting model, we develop an efficient proximal alternating minimization-based algorithm with a convergence guarantee. Extensive experimental results including simulated and real data collectively verify the superiority of the proposed method as compared with the competing methods. Chuan Wang 0001, Xi-Le Zhao, Hao Zhang 0103, Ben-Zheng Li, Meng Ding 0002 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2023 | Unsupervised Global and Local Homography Estimation With Motion Basis LearningabstractIn this paper, we introduce a new framework for unsupervised deep homography estimation. Our contributions are 3 folds. First, unlike previous methods that regress 4 offsets for a homography, we propose a homography flow representation, which can be estimated by a weighted sum of 8 pre-defined homography flow bases. Second, considering a homography contains 8 Degree-of-Freedoms (DOFs) that is much less than the rank of the network features, we propose a Low Rank Representation (LRR) block that reduces the feature rank, so that features corresponding to the dominant motions are retained while others are rejected. Last, we propose a Feature Identity Loss (FIL) to enforce the learned image feature warp-equivariant, meaning that the result should be identical if the order of warp operation and feature extraction is swapped. With this constraint, the unsupervised optimization can be more effective and the learned features are more stable. With global-to-local homography flow refinement, we also naturally generalize the proposed method to local mesh-grid homography estimation, which can go beyond the constraint of a single homography. Extensive experiments are conducted to demonstrate the effectiveness of all the newly proposed components, and results show that our approach outperforms the state-of-the-art on the homography benchmark dataset both qualitatively and quantitatively. Code is available at https://github.com/megvii-research/BasesHomo. Shuaicheng Liu, Hai Jiang 0006, Nianjin Ye, Chuan Wang 0001, Bing Zeng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Content-Aware Unsupervised Deep Homography Estimation and its ExtensionsabstractHomography estimation is a basic image alignment method in many applications. It is usually done by extracting and matching sparse feature points, which are error-prone in low-light and low-texture images. On the other hand, previous deep homography approaches use either synthetic images for supervised learning or aerial images for unsupervised learning, both ignoring the importance of handling depth disparities and moving objects in real-world applications. To overcome these problems, in this work, we propose an unsupervised deep homography method with a new architecture design. In the spirit of the RANSAC procedure in traditional methods, we specifically learn an outlier mask to only select reliable regions for homography estimation. We calculate loss with respect to our learned deep features instead of directly comparing image content as did previously. To achieve the unsupervised training, we also formulate a novel triplet loss customized for our network. We verify our method by conducting comprehensive comparisons on a new dataset that covers a wide range of scenes with varying degrees of difficulties for the task. Experimental results reveal that our method outperforms the state-of-the-art, including deep solutions and feature-based solutions. Shuaicheng Liu, Nianjin Ye, Chuan Wang 0001, Jirong Zhang, Lanpeng Jia, Kunming Luo, Jue Wang 0001, Jian Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | UPHDR-GAN: Generative Adversarial Network for High Dynamic Range Imaging With Unpaired DataabstractThe paper proposes a method to effectively fuse multi-exposure inputs and generate high-quality high dynamic range (HDR) images with unpaired datasets. Deep learning-based HDR image generation methods rely heavily on paired datasets. The ground truth images play a leading role in generating reasonable HDR images. Datasets without ground truth are hard to be applied to train deep neural networks. Recently, Generative Adversarial Networks (GAN) have demonstrated their potentials of translating images from source domain$X$to target domain$Y$in the absence of paired examples. In this paper, we propose a GAN-based network for solving such problems while generating enjoyable HDR results, named UPHDR-GAN. The proposed method relaxes the constraint of the paired dataset and learns the mapping from the LDR domain to the HDR domain. Although the pair data are missing, UPHDR-GAN can properly handle the ghosting artifacts caused by moving objects or misalignments with the help of the modified GAN loss, the improved discriminator network and the useful initialization phase. The proposed method preserves the details of important regions and improves the total image perceptual quality. Qualitative and quantitative comparisons against the representative methods demonstrate the superiority of the proposed UPHDR-GAN. Ru Li 0002, Chuan Wang 0001, Jue Wang 0001, Guanghui Liu 0001, Heng-Yu Zhang, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | ASFlow: Unsupervised Optical Flow Learning With Adaptive Pyramid SamplingabstractWe present an unsupervised optical flow estimation method by proposing an adaptive pyramid sampling in the deep pyramid network. Specifically, in the pyramid downsampling, we propose a Content-Aware Pooling (CAP) module, which promotes local feature gathering by avoiding cross region pooling, so that the learned features become more representative. In the pyramid upsampling, we propose an Adaptive Flow Upsampling (AFU) module, where cross edge interpolation can be avoided, producing sharp motion boundaries. Equipped with these two modules, our method achieves the best performance for unsupervised optical flow estimation on multiple leading benchmarks, including MPI-Sintel, KITTI 2012 and KITTI 2015. Particularly, we achieve EPE=1.5 on KITTI 2012 and F1=9.67% KITTI 2015, which outperform the previous state-of-the-art methods by 16.7% and 13.1%, respectively. Shuaicheng Liu, Kunming Luo, Ao Luo, Chuan Wang 0001, Fanman Meng, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Video Vectorization via Bipartite Diffusion Curves Propagation and OptimizationabstractWe propose a new video vectorization approach for converting videos in the raster format to vector representation with the benefits of resolution independence and compact storage. Through classifying extracted curves in each video frame into salient ones and non-salient ones, we introduce a novel bipartite diffusion curves (BDCs) representation in order to preserve both important image features such as sharp boundaries and regions with smooth color variation. This bipartite representation allows us to propagate non-salient curves across frames such that the propagation, in conjunction with geometry optimization and color optimization of salient curves, ensures the preservation of fine details within each frame and across different frames, and meanwhile, achieves good spatial-temporal coherence. Thorough experiments on a variety of videos show that our method is capable of converting videos to the vector representation with low reconstruction errors, low computational cost, and fine details, demonstrating our superior performance over the state of the art. We also show that, when used for video upsampling, our method produces results comparable to video super-resolution. Yuanqi Li, Chuan Wang 0001, Jie Guo 0001, Jue Wang 0001, Yanwen Guo 0001, Wenping Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | UPFlow: Upsampling Pyramid for Unsupervised Optical Flow LearningabstractWe present an unsupervised learning approach for optical flow estimation by improving the upsampling and learning of pyramid network. We design a self-guided upsample module to tackle the interpolation blur problem caused by bilinear upsampling between pyramid levels. Moreover, we propose a pyramid distillation loss to add supervision for intermediate levels via distilling the finest flow as pseudo labels. By integrating these two components together, our method achieves the best performance for unsupervised optical flow learning on multiple leading benchmarks, including MPI-SIntel, KITTI 2012 and KITTI 2015. In particular, we achieve EPE=1.4 on KITTI 2012 and F1=9.38% on KITTI 2015, which outperform the previous state-of-the-art methods by 22.2% and 15.7%, respectively. Kunming Luo, Chuan Wang 0001, Shuaicheng Liu, Haoqiang Fan, Jue Wang 0001, Jian Sun 0001 |
CVPR | 2 |
| 2021 | Motion Basis Learning for Unsupervised Deep Homography Estimation with Subspace ProjectionabstractIn this paper, we introduce a new framework for unsupervised deep homography estimation. Our contributions are 3 folds. First, unlike previous methods that regress 4 offsets for a homography, we propose a homography flow representation, which can be estimated by a weighted sum of 8 pre-defined homography flow bases. Second, considering a homography contains 8 Degree-of-Freedoms (DOFs) that is much less than the rank of the network features, we propose a Low Rank Representation (LRR) block that reduces the feature rank, so that features corresponding to the dominant motions are retained while others are rejected. Last, we propose a Feature Identity Loss (FIL) to enforce the learned image feature warp-equivariant, meaning that the result should be identical if the order of warp operation and feature extraction is swapped. With this constraint, the unsupervised optimization is achieved more effectively and more stable features are learned. Extensive experiments are conducted to demonstrate the effectiveness of all the newly proposed components, and results show that our approach outperforms the state-of-the-art on the homography benchmark datasets both qualitatively and quantitatively. Code is available at https://github.com/megvii-research/BasesHomo Nianjin Ye, Chuan Wang 0001, Haoqiang Fan, Shuaicheng Liu |
ICCV | 2 |
| 2021 | OIFlow: Occlusion-Inpainting Optical Flow Estimation by Unsupervised LearningabstractOcclusion is an inevitable and critical problem in unsupervised optical flow learning. Existing methods either treat occlusions equally as non-occluded regions or simply remove them to avoid incorrectness. However, the occlusion regions can provide effective information for optical flow learning. In this paper, we present OIFlow, an occlusion-inpainting framework to make full use of occlusion regions. Specifically, a new appearance-flow network is proposed to inpaint occluded flows based on the image content. Moreover, a boundary dilated warp is proposed to deal with occlusions caused by displacement beyond the image border. We conduct experiments on multiple leading flow benchmark datasets such as Flying Chairs, KITTI and MPI-Sintel, which demonstrate that the performance is significantly improved by our proposed occlusion handling framework. Shuaicheng Liu, Kunming Luo, Nianjin Ye, Chuan Wang 0001, Jue Wang 0001, Bing Zeng 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Semi-Supervised Pixel-Level Scene Text Segmentation by Mutually Guided NetworkabstractIn this paper we present a new data-driven method for pixel-level scene text segmentation from a single natural image. Although scene text detection, i.e. producing a text region mask, has been well studied in the past decade, pixel-level text segmentation is still an open problem due to the lack of massive pixel-level labeled data for supervised training. To tackle this issue, we incorporate text region mask as an auxiliary data into this task, considering acquiring large-scale of labeled text region mask is commonly less expensive and time-consuming. To be specific, we propose a mutually guided network which produces a polygon-level mask in one branch and a pixel-level text mask in the other. The two branches' outputs serve as guidance for each other and the whole network is trained via a semi-supervised learning strategy. Extensive experiments are conducted to demonstrate the effectiveness of our mutually guided network, and experimental results show our network outperforms the state-of-the-art in pixel-level scene text segmentation. We also demonstrate the mask produced by our network could improve the text recognition performance besides the trivial image editing application. Chuan Wang 0001, Shan Zhao 0010, Li Zhu 0003, Kunming Luo, Yanwen Guo 0001, Jue Wang 0001, Shuaicheng Liu |
IEEE Trans. Image Process. | 1 |
| 2020 | Content-Aware Unsupervised Deep Homography Estimation
Jirong Zhang, Chuan Wang 0001, Shuaicheng Liu, Lanpeng Jia, Nianjin Ye, Jue Wang 0001, Ji Zhou 0001, Jian Sun 0001 |
ECCV (1) | 2 |
| 2020 | Skin Textural Generation via Blue-noise Gabor Filtering based Generative Adversarial NetworkabstractFacial skin texture synthesis is a fundamental problem in high-quality facial image generation and enhancement. The key behind is how to effectively synthesize plausible textured noise for the faces. With the development of CNNs and GANs, most works cast the problem as an image to image translation problem. However, these methods lack an explicit mechanism to simulate the facial noise pattern, so that the generated images are of obvious artifacts. To this end, we propose a new facial noise generation method. Specifically, we utilize the property of blue noise and Gabor filter to implicitly guide the asymmetrical sampling for the face region as a guidance map, where non-uniform point sampling is conducted. Thus we propose a novel Blue-Noise Gabor Module to produce a spatial-variant noisy image. Our proposed two-branch framework combined facial identity enhancing with textures details generation to jointly produce a high-quality facial image. Experimental results demonstrate the superiority of our method compared with the state-of-the-art, which enables the generation of high-quality facial texture based on a 2D image only, without the involvement of any 3D models. Hui Zhang 0027, Chuan Wang 0001, Nenglun Chen, Jue Wang 0001, Wenping Wang 0001 |
ACM Multimedia | 2 |
| 2019 | Video Inpainting by Jointly Learning Temporal Structure and Spatial DetailsabstractWe present a new data-driven video inpainting method for recovering missing regions of video frames. A novel deep learning architecture is proposed which contains two subnetworks: a temporal structure inference network and a spatial detail recovering network. The temporal structure inference network is built upon a 3D fully convolutional architecture: it only learns to complete a low-resolution video volume given the expensive computational cost of 3D convolution. The low resolution result provides temporal guidance to the spatial detail recovering network, which performs imagebased inpainting with a 2D fully convolutional network to produce recovered video frames in their original resolution. Such two-step network design ensures both the spatial quality of each frame and the temporal coherence across frames. Our method jointly trains both sub-networks in an end-to-end manner. We provide qualitative and quantitative evaluation on three datasets, demonstrating that our method outperforms previous learning-based video inpainting methods. Chuan Wang 0001, Xiaoguang Han 0001, Jue Wang 0001 |
AAAI | 1 |
| 2019 | GIF2Video: Color Dequantization and Temporal Interpolation of GIF ImagesabstractGraphics Interchange Format (GIF) is a highly portable graphics format that is ubiquitous on the Internet. Despite their small sizes, GIF images often contain undesirable visual artifacts such as flat color regions, false contours, color shift, and dotted patterns. In this paper, we propose GIF2Video, the first learning-based method for enhancing the visual quality of GIFs in the wild. We focus on the challenging task of GIF restoration by recovering information lost in the three steps of GIF creation: frame sampling, color quantization, and color dithering. We first propose a novel CNN architecture for color dequantization. It is built upon a compositional architecture for multi-step color correction, with a comprehensive loss function designed to handle large quantization errors. We then adapt the SuperSlomo network for temporal interpolation of GIF frames. We introduce two large datasets, namely GIF-Faces and GIF-Moments, for both training and evaluation. Experimental results show that our method can significantly improve the visual quality of GIFs, and outperforms direct baseline and state-of-the-art approaches. Yang Wang 0097, Chuan Wang 0001, Tong He 0002, Jue Wang 0001, Minh Hoai |
CVPR | 3 |
| 2019 | Semi-Supervised Skin Detection by Network With Mutual GuidanceabstractWe present a new data-driven method for robust skin detection from a single human portrait image. Unlike previous methods, we incorporate human body as a weak semantic guidance into this task, considering acquiring large-scale of human labeled skin data is commonly expensive and time-consuming. To be specific, we propose a dual-task neural network for joint detection of skin and body via a semi-supervised learning strategy. The dual-task network contains a shared encoder but two decoders for skin and body separately. For each decoder, its output also serves as a guidance for its counterpart, making both decoders mutually guided. Extensive experiments were conducted to demonstrate the effectiveness of our network with mutual guidance, and experimental results show our network outperforms the state-of-the-art in skin detection. Jiayuan Shi, Chuan Wang 0001, Guanbin Li, Risheng Liu, Jue Wang 0001 |
ICCV | 3 |
| 2019 | Semi-Supervised Video Salient Object Detection Using Pseudo-LabelsabstractDeep learning-based video salient object detection has recently achieved great success with its performance significantly outperforming any other unsupervised methods. However, existing data-driven approaches heavily rely on a large quantity of pixel-wise annotated video frames to deliver such promising results. In this paper, we address the semi-supervised video salient object detection task using pseudo-labels. Specifically, we present an effective video saliency detector that consists of a spatial refinement network and a spatiotemporal module. Based on the same refinement network and motion information in terms of optical flow, we further propose a novel method for generating pixel-level pseudo-labels from sparsely annotated frames. By utilizing the generated pseudo-labels together with a part of manual annotations, our video saliency detector learns spatial and temporal cues for both contrast inference and coherence enhancement, thus producing accurate saliency maps. Experimental results demonstrate that our proposed semi-supervised method even greatly outperforms all the state-of-the-art fully supervised methods across three public benchmarks of VOS, DAVIS, and FBMS. Pengxiang Yan, Guanbin Li, Yuan Xie 0004, Zhen Li 0026, Chuan Wang 0001, Tianshui Chen, Liang Lin 0004 |
ICCV | 5 |
| 2019 | Two-phase Hair Image Synthesis by Self-Enhancing Generative ModelabstractAbstract Generating plausible hair image given limited guidance, such as sparse sketches or low‐resolution image, has been made possible with the rise of Generative Adversarial Networks (GANs). Traditional image‐to‐image translation networks can generate recognizable results, but finer textures are usually lost and blur artifacts commonly exist. In this paper, we propose a two‐phase generative model for high‐quality hair image synthesis. The two‐phase pipeline first generates a coarse image by an existing image translation model, then applies a re‐generating network with self‐enhancing capability to the coarse image. The self‐enhancing capability is achieved by a proposed differentiable layer, which extracts the structural texture and orientation maps from a hair image. Extensive experiments on two tasks, Sketch2Hair and Hair Super‐Resolution, demonstrate that our approach is able to synthesize plausible hair image with finer details, and reaches the state‐of‐the‐art. Haonan Qiu, Chuan Wang 0001, Xiangyu Zhu 0003, Jinjin Gu, Xiaoguang Han 0001 |
Comput. Graph. Forum | 2 |
| 2017 | Video Vectorization via Tetrahedral RemeshingabstractWe present a video vectorization method that generates a video in vector representation from an input video in raster representation. A vector-based video representation offers the benefits of vector graphics, such as compactness and scalability. The vector video we generate is represented by a simplified tetrahedral control mesh over the spatial-temporal video volume, with color attributes defined at the mesh vertices. We present novel techniques for simplification and subdivision of a tetrahedral mesh to achieve high simplification ratio while preserving features and ensuring color fidelity. From an input raster video, our method is capable of generating a compact video in vector representation that allows a faithful reconstruction with low reconstruction errors. Chuan Wang 0001, Yanwen Guo 0001, Wenping Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | Video Object Co-Segmentation via Subspace Clustering and Quadratic Pseudo-Boolean Optimization in an MRF FrameworkabstractMultiple videos may share a common foreground object, for instance a family member in home videos, or a leading role in various clips of a movie or TV series. In this paper, we present a novel method for co-segmenting the common foreground object from a group of video sequences. The issue was seldom touched on in the literature. Starting from over-segmentation of each video into Temporal Superpixels (TSPs), we first propose a new subspace clustering algorithm which segments the videos into consistent spatio-temporal regions with multiple classes, such that the common foreground has consistent labels across different videos. The subspace clustering algorithm exploits the fact that across different videos the common foreground shares similar appearance features, while motions can be used to better differentiate regions within each video, making accurate extraction of object boundaries easier. We further formulate video object co-segmentation as a Markov Random Field (MRF) model which imposes the constraint of foreground model automatically computed or specified with little user effort. The Quadratic Pseudo-Boolean Optimization (QPBO) is used to generate the results. Experiments show that this video co-segmentation framework can achieve good quality foreground extraction results without user interaction for those videos with unrelated background, and with only moderate user interaction for those videos with similar background. Comparisons with previous work also show the superiority of our approach. Chuan Wang 0001, Yanwen Guo 0001, Linbo Wang 0001, Wenping Wang 0001 |
IEEE Trans. Multim. | 1 |