EDBT 2026 Demo / reviewers in the wild / expert
Xiaogang Wang 0005
dblp:91/6236-5
· DBLP profile ↗
26ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-8402-7504ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Para-Roof: High-quality parametric roof reconstruction via primitive proposal extraction from point cloud
Yuncong Liu, Kai Xu 0004, Xiaogang Wang 0005 |
Expert Syst. Appl. | 4 |
| 2025 | Learning 3D Volume Cloud from Single ImageabstractThree-dimensional (3D) cloud modeling plays a pivotal role in advancing atmospheric models and enhancing natural phenomena visualization systems. Nevertheless, the high-quality reconstruction of clouds remains a significant challenge, primarily due to their inherent heterogeneous nature as volumetric media. Image-based modeling approaches offer a promising solution to this challenge. This paper presents a novel two-stage neural network architecture for 3D cloud reconstruction from single image. The first stage introduces an innovative view synthesis network built upon Stable Diffusion, incorporating two specialized modules: a cloud mapper and a viewpoint mapper, which collaboratively generate novel perspective views from a single input image. The second stage implements a physics-based differentiable rendering framework to construct a 3D cloud reconstruction network, leveraging the synthesized multi-view images to optimize a volumetric density grid representation. To enhance the reconstruction fidelity, we integrate real-world cloud density distribution statistics and implement a post-processing refinement using Perlin-Worley noise combined with Fractal Brownian Motion (FBM) for erosion effects. Additionally, to mitigate the inherent limitations of geometric information extraction from single-view images, we developed a comprehensive cloud simulation dataset for pre-training the viewpoint mapper module. This dataset encompasses multi-view cloud images with corresponding camera extrinsic parameters, capturing a diverse range of cloud formations. Extensive quantitative evaluations and qualitative assessments demonstrate the efficacy and potential of our proposed two-stage network in achieving accurate 3D cloud reconstruction from single-view images. Yuhang Cheng, Yu Zhang 0035, Xiaogang Wang 0005 |
ICMR | 3 |
| 2025 | Demystify Transformers & Convolutions in Modern Image Deep NetworksabstractVision transformers have gained popularity recently, leading to the development of new vision backbones with improved features and consistent performance gains. However, these advancements are not solely attributable to novel feature transformation designs; certain benefits also arise from advanced network-level and block-level architectures. This paper aims to identify the real gains of popular convolution and attention operators through a detailed study. We find that the key difference among these feature transformation modules, such as attention or convolution, lies in their spatial feature aggregation approach, known as the "spatial token mixer" (STM). To facilitate an impartial comparison, we introduce a unified architecture to neutralize the impact of divergent network-level and block-level designs. Subsequently, various STMs are integrated into this unified framework for comprehensive comparative analysis. Our experiments on various tasks and an analysis of inductive bias show a significant performance boost due to advanced network-level and block-level designs, but performance differences persist among different STMs. Our detailed analysis also reveals various findings about different STMs, including effective receptive fields, invariance, and adversarial robustness tests. Xiaowei Hu 0001, Min Shi 0004, Weiyun Wang, Sitong Wu, Linjie Xing, Wenhai Wang, Xizhou Zhou, Lewei Lu, Jie Zhou 0001, Xiaogang Wang 0005, Yu Qiao 0001, Jifeng Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2024 | FaceCom: Towards High-fidelity 3D Facial Shape Completion via Optimization and Inpainting GuidanceabstractWe propose FaceCom, a method for 3D facial shape completion, which delivers high-fidelity results for incomplete facial inputs of arbitrary forms. Unlike end-to-end shape completion methods based on point clouds or voxels, our approach relies on a mesh-based generative network that is easy to optimize, enabling it to handle shape completion for irregular facial scans. We first train a shape generator on a mixed 3D facial dataset containing 2405 identities. Based on the incomplete facial input, we fit complete faces using an optimization approach under image inpainting guidance. The completion results are refined through a post-processing step. FaceCom demonstrates the ability to effectively and naturally complete facial scan data with varying missing regions and degrees of missing areas. Our method can be used in medical prosthetic fabrication and the registration of deficient scanning data. Our experimental results demonstrate that FaceCom achieves exceptional performance in fitting and shape completion tasks. The code is available at https://github.com/dragonylee/FaceCom.git. Yinglong Li, Xiaogang Wang 0005, Qingzhao Qin, Yijiao Zhao, Aimin Hao |
CVPR | 3 |
| 2024 | High-Order Differential Regularizing Implicit Surface Representation of Point CloudabstractReconstructing surfaces from diverse raw data in computer graphics poses an enduring challenge. While recent methods deploy neural networks for direct implicit surface reconstruction, they struggle with degraded raw data quality, especially in edge regions. To address this, we advocate for employing high-order total generalized variation (TGV) as a regularization term for implicit surface representation. Acknowledging the non-trivial nature of extending typical image processing methods to implicit surfaces, we present an end-to-end trainable network framework for TGV in implicit surface reconstruction. This approach preserves sharp features, enhances smooth region recovery, and minimizes artificial artifacts. Simultaneously, we address the challenge of increased computational complexity associated with current algorithms by predicting it directly through an implicit neural function. Experimental results demonstrate the efficacy of our technical approach, providing a promising solution for robust implicit surface reconstruction. Yuhang Cheng, Ziyang Fan, Xiaogang Wang 0005 |
ICME | 4 |
| 2024 | Learning to Transfer Heterogeneous Translucent Materials from a 2D Image to 3D ModelsabstractGreat progress has been made in rendering translucent materials in recent years, but automatically estimating parameters for heterogeneous materials such as jade and human skin remains a challenging task, often requiring specialized and expensive physical measurement devices. In this paper, we present a novel approach for estimating and transferring the parameters of heterogeneous translucent materials from a single 2D image to 3D models. Our method consists of four key steps: (1) An efficient viewpoint selection algorithm to minimize redundancy and ensure comprehensive coverage of the model. (2) Initializing a homogeneous translucent material to render initial images for translucent dataset. (3) Edit the rendered translucent images to update the translucent dataset. (4) Optimize the edited translucent results onto material parameters using inverse rendering techniques. Our approach offers a practical and accessible solution that overcomes the limitations of existing methods, which often rely on complex and costly specialized devices. We demonstrate the effectiveness and superiority of our proposed method through extensive experiments, showcasing its ability to transfer and edit high-quality heterogeneous translucent materials on 3D models, surpassing the results achieved by previous techniques in 3D scene editing. Xiaogang Wang 0005, Yuhang Cheng, Ziyang Fan, Kai Xu 0004 |
ACM Multimedia | 1 |
| 2024 | RNNPose: 6-DoF Object Pose Estimation via Recurrent Correspondence Field Estimation and Pose Optimizationabstract6-DoF object pose estimation from a monocular image is a challenging problem, where a post-refinement procedure is generally needed for high-precision estimation. In this paper, we propose a framework, dubbed RNNPose, based on a recurrent neural network (RNN) for object pose refinement, which is robust to erroneous initial poses and occlusions. During the recurrent iterations, object pose refinement is formulated as a non-linear least squares problem based on the estimated correspondence field (between a rendered image and the observed image). The problem is then solved by a differentiable Levenberg-Marquardt (LM) algorithm enabling end-to-end training. The correspondence field estimation and pose refinement are conducted alternately in each iteration to improve the object poses. Furthermore, to improve the robustness against occlusion, we introduce a consistency-check mechanism based on the learned descriptors of the 3D model and observed 2D images, which downweights the unreliable correspondences during pose optimization. We evaluate RNNPose on several public datasets, including LINEMOD, Occlusion-LINEMOD, YCB-Video and TLESS. We demonstrate state-of-the-art performance and strong robustness against severe clutter and occlusion in the scenes. Extensive experiments validate the effectiveness of our proposed method. Besides, the extended system based on RNNPose successfully generalizes to multi-instance scenarios and achieves top-tier performance on the TLESS dataset. Kwan-Yee Lin, Guofeng Zhang 0001, Xiaogang Wang 0005, Hongsheng Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | TCFormer: Visual Recognition via Token Clustering TransformerabstractTransformers are widely used in computer vision areas and have achieved remarkable success. Most state-of-the-art approaches split images into regular grids and represent each grid region with a vision token. However, fixed token distribution disregards the semantic meaning of different image regions, resulting in sub-optimal performance. To address this issue, we propose the Token Clustering Transformer (TCFormer), which generates dynamic vision tokens based on semantic meaning. Our dynamic tokens possess two crucial characteristics: (1) Representing image regions with similar semantic meanings using the same vision token, even if those regions are not adjacent, and (2) concentrating on regions with valuable details and represent them using fine tokens. Through extensive experimentation across various applications, including image classification, human pose estimation, semantic segmentation, and object detection, we demonstrate the effectiveness of our TCFormer. Sheng Jin 0007, Lumin Xu, Wentao Liu 0002, Chen Qian 0006, Wanli Ouyang, Ping Luo 0002, Xiaogang Wang 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | Weak Augmentation Guided Relational Self-Supervised LearningabstractSelf-supervised Learning (SSL) including the mainstream contrastive learning has achieved great success in learning visual representations without data annotations. However, most methods mainly focus on the instance level information (i.e., the different augmented images of the same instance should have the same feature or cluster into the same class), but there is a lack of attention on the relationships between different instances. In this paper, we introduce a novel SSL paradigm, which we term as relational self-supervised learning (ReSSL) framework that learns representations by modeling the relationship between different instances. Specifically, our proposed method employs sharpened distribution of pairwise similarities among different instances as relation metric, which is thus utilized to match the feature embeddings of different augmentations. To boost the performance, we argue that weak augmentations matter to represent a more reliable relation, and leverage momentum strategy for practical efficiency. The designed asymmetric predictor head and an InfoNCE warm-up strategy enhance the robustness to hyper-parameters and benefit the resulting performance. Experimental results show that our proposed ReSSL substantially outperforms the state-of-the-art methods across different network architectures, including various lightweight networks (e.g., EfficientNet and MobileNet). Mingkai Zheng, Shan You, Fei Wang 0032, Chen Qian 0006, Changshui Zhang, Xiaogang Wang 0005, Chang Xu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Parametric Primitive Analysis of CAD Sketches With Vision TransformerabstractThe design and analysis of computer-aided design (CAD) sketches play a crucial role in industrial product design, primarily involving CAD primitives and their interprimitive constraints. To address challenges related to error accumulation in autoregressive models and the complexities associated with self-supervised model design for this task, we propose a two-stage network framework. This framework consists of a primitive network and a constraint network, transforming the sketch analysis task into a set prediction problem to enhance the effective handling of primitives and constraints. By decoupling target types from parameters, the model gains increased flexibility and optimization while reducing complexity. In addition, the constraint network incorporates a pointer module to explicitly indicate the relationship between constraint parameters and primitive indices, enhancing interpretability and performance. Qualitative and quantitative analyzes on two publicly available datasets demonstrate the superiority of this method. Xiaogang Wang 0005, Liang Wang 0001, Guoqiang Xiao 0001, Kai Xu 0004 |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Med-DualNet: Enhancing Medical Image Segmentation with Hard-Patch Mining and Joint Error AdjustmentabstractAutomated segmentation of organs and lesions is crucial for effective treatment design and care planning in the medical field. However, accurately segmenting organs and lesions is challenging due to the complex distribution of lesions within organs, such as cases with early diseases characterized by small volumes and blurred boundaries. In this work, we propose Med-DualNet, a dual-branch segmentation network that incorporates Hard-Patch Mining and Joint Error Adjustment strategies to address these challenges. Branch I focuses on organ segmentation, while Branch II targets challenging lesion segmentation using a two-stage network and dedicated strategies. The Hard-Patch Mining strategy in Stage-one dynamically generates a subset of image patches that highlight challenging aspects of lesion segmentation, including small volumes and blurred boundaries. Furthermore, the Joint Error Adjustment strategy in Stage-two facilitates the learning of challenging lesion features. Experimental results on three public medical datasets demonstrate the improved performance of the proposed model in organ and lesion segmentation, as measured by visual and quantitative metrics. Xiaogang Wang 0005, Xiaoqin Tang |
BIBM | 2 |
| 2023 | Finernet: A Coarse-to-Fine Approach to Learning High-Quality Implicit Surface Reconstruction
Dan Mei, Xiaogang Wang 0005 |
CGI | 2 |
| 2023 | PPI-NET: End-to-End Parametric Primitive Inference
Liang Wang 0001, Xiaogang Wang 0005 |
CGI (4) | 2 |
| 2023 | ZoomNAS: Searching for Whole-Body Human Pose Estimation in the WildabstractThis paper investigates the task of 2D whole-body human pose estimation, which aims to localize dense landmarks on the entire human body including body, feet, face, and hands. We propose a single-network approach, termed ZoomNet, to take into account the hierarchical structure of the full human body and solve the scale variation of different body parts. We further propose a neural architecture search framework, termed ZoomNAS, to promote both the accuracy and efficiency of whole-body pose estimation. ZoomNAS jointly searches the model architecture and the connections between different sub-modules, and automatically allocates computational complexity for searched sub-modules. To train and evaluate ZoomNAS, we introduce the first large-scale 2D human whole-body dataset, namely COCO-WholeBody V1.0, which annotates 133 keypoints for in-the-wild images. Extensive experiments demonstrate the effectiveness of ZoomNAS and the significance of COCO-WholeBody V1.0. Lumin Xu, Sheng Jin 0007, Wentao Liu 0002, Chen Qian 0006, Wanli Ouyang, Ping Luo 0002, Xiaogang Wang 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | Pose for Everything: Towards Category-Agnostic Pose Estimation
Lumin Xu, Sheng Jin 0007, Wentao Liu 0002, Chen Qian 0006, Wanli Ouyang, Ping Luo 0002, Xiaogang Wang 0005 |
ECCV (6) | 8 |
| 2021 | Learning Fine-Grained Segmentation of 3D Shapes Without Part LabelsabstractLearning-based 3D shape segmentation is usually formulated as a semantic labeling problem, assuming that all parts of training shapes are annotated with a given set of tags. This assumption, however, is impractical for learning fine-grained segmentation. Although most off-the-shelf CAD models are, by construction, composed of fine-grained parts, they usually miss semantic tags and labeling those fine-grained parts is extremely tedious. We approach the problem with deep clustering, where the key idea is to learn part priors from a shape dataset with fine-grained segmentation but no part labels. Given point sampled 3D shapes, we model the clustering priors of points with a similarity matrix and achieve part segmentation through minimizing a novel low rank loss. To handle highly densely sampled point sets, we adopt a divide-and-conquer strategy. We partition the large point set into a number of blocks. Each block is segmented using a deep-clustering-based part prior network trained in a category-agnostic manner. We then train a graph convolution network to merge the segments of all blocks to form the final segmentation result. Our method is evaluated with a challenging benchmark of fine-grained segmentation, showing state-of-the-art performance. Xiaogang Wang 0005, Kai Xu 0004 |
CVPR | 1 |
| 2020 | PIE-NET: Parametric Inference of Point Cloud EdgesabstractWe introduce an end-to-end learnable technique to robustly identify feature edges in 3D point cloud data. We represent these edges as a collection of parametric curves (i.e.,~lines, circles, and B-splines). Accordingly, our deep neural network, coined PIE-NET, is trained for parametric inference of edges. The network relies on a "region proposal" architecture, where a first module proposes an over-complete collection of edge and corner points, and a second module ranks each proposal to decide whether it should be considered. We train and evaluate our method on the ABC dataset, a large dataset of CAD models, and compare our results to those produced by traditional (non-learning) processing pipelines, as well as a recent deep learning based edge detector (EC-NET). Our results significantly improve over the state-of-the-art from both a quantitative and qualitative standpoint. Xiaogang Wang 0005, Yuelang Xu, Kai Xu 0004, Andrea Tagliasacchi, Ali Mahdavi-Amiri, Hao (Richard) Zhang |
NeurIPS | 1 |
| 2020 | Unsupervised Video Matting via Sparse and Low-Rank RepresentationabstractA novel method, unsupervised video matting via sparse and low-rank representation, is proposed which can achieve high quality in a variety of challenging examples featuring illumination changes, feature ambiguity, topology changes, transparency variation, dis-occlusion, fast motion and motion blur. Some previous matting methods introduced a nonlocal prior to search samples for estimating the alpha matte, which have achieved impressive results on some data. However, on one hand, searching inadequate or excessive samples may miss good samples or introduce noise; on the other hand, it is difficult to construct consistent nonlocal structures for pixels with similar features, yielding video mattes with spatial and temporal inconsistency. In this paper, we proposed a novel video matting method to achieve spatially and temporally consistent matting result. Toward this end, a sparse and low-rank representation model is introduced to pursue consistent nonlocal structures for pixels with similar features. The sparse representation is used to adaptively select best samples and accurately construct the nonlocal structures for all pixels, while the low-rank representation is used to globally ensure consistent nonlocal structures for pixels with similar features. The two representations are combined to generate spatially and temporally consistent video mattes. We test our method on lots of dataset including the benchmark dataset for image matting and dataset for video matting. Our method has achieved the best performance among all unsupervised matting methods in the public alpha matting evaluation dataset for images. Dongqing Zou, Xiaowu Chen 0001, Guangying Cao, Xiaogang Wang 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Shape2Motion: Joint Analysis of Motion Parts and Attributes From 3D ShapesabstractFor the task of mobility analysis of 3D shapes, we propose joint analysis for simultaneous motion part segmentation and motion attribute estimation, taking a single 3D model as input. The problem is significantly different from those tackled in the existing works which assume the availability of either a pre-existing shape segmentation or multiple 3D models in different motion states. To that end, we develop Shape2Motion which takes a single 3D point cloud as input, and jointly computes a mobility-oriented segmentation and the associated motion attributes. Shape2Motion is comprised of two deep neural networks designed for mobility proposal generation and mobility optimization, respectively. The key contribution of these networks is the novel motion-driven features and losses used in both motion part segmentation and motion attribute estimation. This is based on the observation that the movement of a functional part preserves the shape structure. We evaluate Shape2Motion with a newly proposed benchmark for mobility analysis of 3D shapes. Results demonstrate that our method achieves the state-of-the-art performance both in terms of motion part segmentation and motion attribute estimation. Xiaogang Wang 0005, Yahao Shi, Xiaowu Chen 0001, Qinping Zhao, Kai Xu 0004 |
CVPR | 1 |
| 2019 | Learning semantic abstraction of shape via 3D region of interest
Haiyue Fang, Xiaogang Wang 0005, Zheyuan Cai, Yahao Shi, Shilin Wu |
Graph. Model. | 2 |
| 2018 | Single Image Dehazing Using Ranking Convolutional Neural NetworkabstractSingle image dehazing, which aims to recover the clear image solely from an input hazy or foggy image, is a challenging ill-posed problem. Analyzing existing approaches, the common key step is to estimate the haze density of each pixel. To this end, various approaches oftenheuristically designedhaze-relevant features. Several recent works also automatically learn the features via directly exploiting convolutional neural networks (CNN). However, it may be insufficient to fully capture the intrinsic attributes of hazy images. To obtain effective features for single image dehazing, this paper presents a novel ranking convolutional neural network (Ranking-CNN). In Ranking-CNN, a novel ranking layer is proposed to extend the structure of CNN so that the statistical and structural attributes of hazy images can be simultaneously captured. By training Ranking-CNN in a well-designed manner, powerful haze-relevant features can beautomatically learnedfrom massive hazy image patches. Based on these features, haze can be effectively removed by using a haze density prediction model trained through the random forest regression. Experimental results show that our approach outperforms several previous dehazing approaches on synthetic and real-world benchmark images. Comprehensive analyses are also conducted to interpret the proposed Ranking-CNN from both the theoretical and experimental aspects. Yafei Song 0002, Jia Li 0003, Xiaogang Wang 0005, Xiaowu Chen 0001 |
IEEE Trans. Multim. | 3 |
| 2018 | Learning to group and label fine-grained shape componentsabstractA majority of stock 3D models in modern shape repositories are assembled with many fine-grained components. The main cause of such data form is the component-wise modeling process widely practiced by human modelers. These modeling components thus inherently reflect some function-based shape decomposition the artist had in mind during modeling. On the other hand, modeling components represent an over-segmentation since a functional part is usually modeled as a multi-component assembly. Based on these observations, we advocate that labeled segmentation of stock 3D models should not overlook the modeling components and propose a learning solution to grouping and labeling of the fine-grained components. However, directly characterizing the shape of individual components for the purpose of labeling is unreliable, since they can be arbitrarily tiny and semantically meaningless. We propose to generate part hypotheses from the components based on a hierarchical grouping strategy, and perform labeling on those part groups instead of directly on the components. Part hypotheses are mid-level elements which are more probable to carry semantic information. A multi-scale 3D convolutional neural network is trained to extract context-aware features for the hypotheses. To accomplish a labeled segmentation of the whole shape, we formulate higher-order conditional random fields (CRFs) to infer an optimal label assignment for all components. Extensive experiments demonstrate that our method achieves significantly robust labeling results on raw 3D models from public shape repositories. Our work also contributes the first benchmark for component-wise labeling. Xiaogang Wang 0005, Haiyue Fang, Xiaowu Chen 0001, Qinping Zhao, Kai Xu 0004 |
ACM Trans. Graph. | 1 |
| 2018 | Efficiently consistent affinity propagation for 3D shapes co-segmentation
Xiaogang Wang 0005, Zongji Wang, Dongqing Zou, Xiaowu Chen 0001, Qinping Zhao |
Vis. Comput. | 1 |
| 2016 | 6-DOF Image Localization From Massive Geo-Tagged Reference ImagesabstractThe 6-degrees of freedom (DOF) image localization, which aims to calculate the spatial position and rotation of a camera, is a challenging problem for most location-based services. In existing approaches, this problem is often tackled by finding the matches between 2D image points and 3D structure points so as to derive the location information via direct linear transformation algorithm. However, as these 2D-to-3D-based approaches need to reconstruct the 3D structure points of the scene, they may not be flexible enough to employ massive and increasing geo-tagged data. To this end, this paper presents a novel approach for 6-DOF image localization by fusing candidate poses relative to reference images. In this approach, we propose to localize an input image according to the position and rotation information of multiple geo-tagged images retrieved from a reference dataset. From the reference images, an efficient relative pose estimation algorithm is proposed to derive a set of candidate poses for the input image. Each candidate pose encodes the relative rotation and direction of the input image with respect to a specific reference image. Finally, these candidate poses can be fused together by minimizing a well-defined geometry error so that the 6-DOF location of the input image is effectively derived. Experimental results show that our method can obtain satisfactory localization accuracy. In addition, the proposed relative pose estimation algorithm is much faster than existing work. Yafei Song 0002, Xiaowu Chen 0001, Xiaogang Wang 0005, Yu Zhang 0035, Jia Li 0003 |
IEEE Trans. Multim. | 3 |
| 2015 | Video Matting via Sparse and Low-Rank RepresentationabstractWe introduce a novel method of video matting via sparse and low-rank representation. Previous matting methods [10, 9] introduced a nonlocal prior to estimate the alpha matte and have achieved impressive results on some data. However, on one hand, searching inadequate or excessive samples may miss good samples or introduce noise, on the other hand, it is difficult to construct consistent nonlocal structures for pixels with similar features, yielding spatially and temporally inconsistent video mattes. In this paper, we proposed a novel video matting method to achieve spatially and temporally consistent matting result. Toward this end, a sparse and low-rank representation model is introduced to pursue consistent nonlocal structures for pixels with similar features. The sparse representation is used to adaptively select best samples and accurately construct the nonlocal structures for all pixels, while the low-rank representation is used to globally ensure consistent nonlocal structures for pixels with similar features. The two representations are combined to generate consistent video mattes. Experimental results show that our method has achieved high quality results in a variety of challenging examples featuring illumination changes, feature ambiguity, topology changes, transparency variation, dis-occlusion, fast motion and motion blur. Dongqing Zou, Xiaowu Chen 0001, Guangying Cao, Xiaogang Wang 0005 |
ICCV | 4 |
| 2015 | Cuboids detection in RGB-D images via Maximum Weighted CliqueabstractCuboid detection is an essential step for understanding 3D structure of scenes. As most of indoor scene cuboids are actually objects, we propose in this paper an object-based approach to detect 3D cuboids in indoor RGB-D images. The proposed approach is learning-free and can handle general object classes rather than a limited pre-defined category set. In our approach, we first apply an extended version of the CPMC framework to generate a set of segment hypotheses, and fit a set of cuboid candidates. Given the candidate set, we select several cuboids that can provide plausible interpretations of the images by solving a Maximum Weighted Clique (MWC) problem. With this formulation, a set of ranked mid-level representations of the input image is obtained, and are further re-ranked by Maximal Marginal Relevance (MMR) measure to improve their diversity. Experimental results on NYU-V2 dataset shows that our method significantly outperforms the state-of-the-art, and shows impressive results. Xiaowu Chen 0001, Yu Zhang 0035, Jia Li 0003, Xiaogang Wang 0005 |
ICME | 6 |