VLDB 2026 Research / reviewers in the wild / expert
Xiaowu Chen 0001
dblp:52/2710-1
· DBLP profile ↗
94ranked-venue papers
25as first author
7since 2021 · last 2026
0000-0002-3976-6500ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 71 · 18 first-author · 6 since 2021Artificial intelligence and machine learning · 31 · 6 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-authorHuman-computer interaction and ubiquitous computing · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning When and How to Update Memory for Video Object SegmentationabstractRecent progress in semi-supervised video object segmentation has largely hinged on memory-based methods. However, when faced with increasingly tough challenges emerging in complex scenarios, such as fundamental semantic transformations and severe spatial deformations, the fixed-interval memory update mechanism usually adopted in these memory-based methods is insufficient to align with the pivotal moments of object changes. This inflexible mechanism motivates us to design an adaptive memory update mechanism in response to the semantic-spatial changes of target objects. To this end, we propose a novel Change-Sensitive Network (CSNet) to learn when and how to update memory to effectively address intricate challenges in complex scenarios. Specifically, we first design an Adaptive Perception-Capture module with a hierarchical contrastive learning loss to determine when to update memory moments by measuring the extent of object changes, thus dividing entire videos into different object-change clips. To further extract and highlight object changes to assist in the segmentation of frames after changes occur, we construct Dynamic Memory Update modules to redefine how to update memory by smoothly retaining the object prototypes within clips and dynamically amplifying the object variations across clips. Extensive experiments demonstrate that our proposed CSNet exhibits clear superiority when evaluated on eight datasets covering three kinds: common, complex and long-video datasets. Shengye Qiao, Changqun Xia, Xiaowu Chen 0001, Jia Li 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Rethinking Lightweight Salient Object Detection via Network Depth-Width TradeoffabstractExisting salient object detection methods often adopt deeper and wider networks for better performance, resulting in heavy computational burden and slow inference speed. This inspires us to rethink saliency detection to achieve a favorable balance between efficiency and accuracy. To this end, we design a lightweight framework while maintaining satisfying competitive accuracy. Specifically, we propose a novel trilateral decoder framework by decoupling the U-shape structure into three complementary branches, which are devised to confront the dilution of semantic context, loss of spatial structure and absence of boundary detail, respectively. Along with the fusion of three branches, the coarse segmentation results are gradually refined in structure details and boundary quality. Without adding additional learnable parameters, we further propose Scale-Adaptive Pooling Module to obtain multi-scale receptive field. In particular, on the premise of inheriting this framework, we rethink the relationship among accuracy, parameters and speed via network depth-width tradeoff. With these insightful considerations, we comprehensively design shallower and narrower models to explore the maximum potential of lightweight SOD. Our models are proposed for different application environments: 1) a tiny version CTD-S (1.7M, 125FPS) for resource constrained devices, 2) a fast version CTD-M (12.6M, 158FPS) for speed-demanding scenarios, 3) a standard version CTD-L (26.5M, 84FPS) for high-performance platforms. Extensive experiments validate the superiority of our method, which achieves better efficiency-accuracy balance across five benchmarks. Jia Li 0003, Shengye Qiao, Zhirui Zhao, Xiaowu Chen 0001, Changqun Xia |
IEEE Trans. Image Process. | 5 |
| 2023 | Boosting Broader Receptive Fields for Salient Object DetectionabstractSalient Object Detection has boomed in recent years and achieved impressive performance on regular-scale targets. However, existing methods encounter performance bottlenecks in processing objects with scale variation, especially extremely large- or small-scale objects with asymmetric segmentation requirements, since they are inefficient in obtaining more comprehensive receptive fields. With this issue in mind, this paper proposes a framework named BBRF for Boosting Broader Receptive Fields, which includes a Bilateral Extreme Stripping (BES) encoder, a Dynamic Complementary Attention Module (DCAM) and a Switch-Path Decoder (SPD) with a new boosting loss under the guidance of Loop Compensation Strategy (LCS). Specifically, we rethink the characteristics of the bilateral networks, and construct a BES encoder that separates semantics and details in an extreme way so as to get the broader receptive fields and obtain the ability to perceive extreme large- or small-scale objects. Then, the bilateral features generated by the proposed BES encoder can be dynamically filtered by the newly proposed DCAM. This module interactively provides spacial-wise and channel-wise dynamic attention weights for the semantic and detail branches of our BES encoder. Furthermore, we subsequently propose a Loop Compensation Strategy to boost the scale-specific features of multiple decision paths in SPD. These decision paths form a feature loop chain, which creates mutually compensating features under the supervision of boosting loss. Experiments on five benchmark datasets demonstrate that the proposed BBRF has a great advantage to cope with scale variation and can reduce the Mean Absolute Error over 20% compared with the state-of-the-art methods. Mingcan Ma, Changqun Xia, Xiaowu Chen 0001, Jia Li 0003 |
IEEE Trans. Image Process. | 4 |
| 2022 | Pyramid Grafting Network for One-Stage High Resolution Saliency DetectionabstractRecent salient object detection (SOD) methods based on deep neural network have achieved remarkable performance. However, most of existing SOD models designed for low-resolution input perform poorly on high-resolution images due to the contradiction between the sampling depth and the receptive field size. Aiming at resolving this con-tradiction, we propose a novel one-stage framework called Pyramid Grafting Network (PGNet), using transformer and CNN backbone to extract features from different resolution images independently and then graft the features from transformer branch to CNN branch. An attention-based Cross-Model Grafting Module (CMGM) is proposed to en-able CNN branch to combine broken detailed information more holistically, guided by different source feature during decoding process. Moreover, we design an Attention Guided Loss (AGL) to explicitly supervise the attention matrix generated by CMGM to help the network better interact with the attention from different models. We contribute a new Ultra-High-Resolution Saliency Detection dataset UHRSD, containing 5,920 images at 4K-SK resolutions. To our knowledge, it is the largest dataset in both quantity and resolution for high-resolution SOD task, which can be used for training and testing in future research. Sufficient exper-iments on UHRSD and widely-used SOD datasets demon-strate that our method achieves superior performance compared to the state-of-the-art methods. Changqun Xia, Mingcan Ma, Zhirui Zhao, Xiaowu Chen 0001, Jia Li 0003 |
CVPR | 5 |
| 2022 | Revisiting Stochastic Learning for Generalizable Person Re-identificationabstractGeneralizable person re-identification aims to achieve a well generalization capability on target domains without accessing target data. Existing methods focus on suppressing domain-specific information or simulating unseen environments by meta-learning strategies, which could damage the capture ability on fine-grained visual patterns or lead to overfitting issues by the repetitive training of episodes. In this paper, we revisit the stochastic behaviors from two different perspectives: 1) Stochastic splitting-sliding sampler. It splits domain sources into approximately equal sample-size subsets and selects several subsets from various sources by a sliding window, forcing the model to step out of local minimums under stochastic sources. 2) Variance-varying gradient dropout. Gradients in parts of network are also selected by a sliding window and multiplied by binary masks generated from Bernoulli distribution, making gradients in varying variance and preventing the model from local minimums. By applying these two proposed stochastic behaviors, the model achieves a better generalization performance on unseen target domains without any additional computation costs or auxiliary modules. Extensive experiments demonstrate that our proposed model is effective and outperforms state-of-the-art methods on public domain generalizable person Re-ID benchmarks. Jiajian Zhao, Yifan Zhao 0002, Xiaowu Chen 0001, Jia Li 0003 |
ACM Multimedia | 3 |
| 2021 | Part-Guided Relational Transformers for Fine-Grained Visual RecognitionabstractFine-grained visual recognition is to classify objects with visually similar appearances into subcategories, which has made great progress with the development of deep CNNs. However, handling subtle differences between different subcategories still remains a challenge. In this paper, we propose to solve this issue in one unified framework from two aspects, i.e., constructing feature-level interrelationships, and capturing part-level discriminative features. This framework, namely PArt-guided Relational Transformers (PART), is proposed to learn the discriminative part features with an automatic part discovery module, and to explore the intrinsic correlations with a feature transformation module by adapting the Transformer models from the field of natural language processing. The part discovery module efficiently discovers the discriminative regions which are highly-corresponded to the gradient descent procedure. Then the second feature transformation module builds correlations within the global embedding and multiple part embedding, enhancing spatial interactions among semantic pixels. Moreover, our proposed approach does not rely on additional part branches in the inference time and reaches state-of-the-art performance on 3 widely-used fine-grained object recognition benchmarks. Experimental results and explainable visualizations demonstrate the effectiveness of our proposed approach. Yifan Zhao 0002, Jia Li 0003, Xiaowu Chen 0001, Yonghong Tian 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | RGB-D Salient Object Detection With Ubiquitous Target AwarenessabstractConventional RGB-D salient object detection methods aim to leverage depth as complementary information to find the salient regions in both modalities. However, the salient object detection results heavily rely on the quality of captured depth data which sometimes are unavailable. In this work, we make the first attempt to solve the RGB-D salient object detection problem with a novel depth-awareness framework. This framework only relies on RGB data in the testing phase, utilizing captured depth data as supervision for representation learning. To construct our framework as well as achieving accurate salient detection results, we propose a Ubiquitous Target Awareness (UTA) network to solve three important challenges in RGB-D SOD task: 1) a depth awareness module to excavate depth information and to mine ambiguous regions via adaptive depth-error weights, 2) a spatial-aware cross-modal interaction and a channel-aware cross-level interaction, exploiting the low-level boundary cues and amplifying high-level salient channels, and 3) a gated multi-scale predictor module to perceive the object saliency in different contextual scales. Besides its high performance, our proposed UTA network is depth-free for inference and runs in real-time with 43 FPS. Experimental evidence demonstrates that our proposed network not only surpasses the state-of-the-art methods on five public RGB-D SOD benchmarks by a large margin, but also verifies its extensibility on five public RGB SOD benchmarks. Yifan Zhao 0002, Jia Li 0003, Xiaowu Chen 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Is Depth Really Necessary for Salient Object Detection?abstractSalient object detection (SOD) is a crucial and preliminary task for many computer vision applications, which have made progress with deep CNNs. Most of the existing methods mainly rely on the RGB information to distinguish the salient objects, which faces difficulties in some complex scenarios. To solve this, many recent RGBD-based networks are proposed by adopting the depth map as an independent input and fuse the features with RGB information. Taking the advantages of RGB and RGBD methods, we propose a novel depth-aware salient object detection framework, which has following superior designs: 1) It does not rely on depth data in the testing phase. 2) It comprehensively optimizes SOD features with multi-level depth-aware regularizations. 3) The depth information also serves as error-weighted map to correct the segmentation process. With these insightful designs combined, we make the first attempt in realizing an unified depth-aware framework with only RGB information as input for inference, which not only surpasses the state-of-the-art performance on five public RGB SOD benchmarks, but also surpasses the RGBD-based methods on five benchmarks by a large margin, while adopting less information and implementation light-weighted. Yifan Zhao 0002, Jia Li 0003, Xiaowu Chen 0001 |
ACM Multimedia | 4 |
| 2020 | Human-centric metrics for indoor scene assessment and synthesis
Qiang Fu 0004, Hongbo Fu 0001, Hai Yan, Xiaowu Chen 0001, Xueming Li 0002 |
Graph. Model. | 5 |
| 2020 | Unsupervised Video Matting via Sparse and Low-Rank RepresentationabstractA novel method, unsupervised video matting via sparse and low-rank representation, is proposed which can achieve high quality in a variety of challenging examples featuring illumination changes, feature ambiguity, topology changes, transparency variation, dis-occlusion, fast motion and motion blur. Some previous matting methods introduced a nonlocal prior to search samples for estimating the alpha matte, which have achieved impressive results on some data. However, on one hand, searching inadequate or excessive samples may miss good samples or introduce noise; on the other hand, it is difficult to construct consistent nonlocal structures for pixels with similar features, yielding video mattes with spatial and temporal inconsistency. In this paper, we proposed a novel video matting method to achieve spatially and temporally consistent matting result. Toward this end, a sparse and low-rank representation model is introduced to pursue consistent nonlocal structures for pixels with similar features. The sparse representation is used to adaptively select best samples and accurately construct the nonlocal structures for all pixels, while the low-rank representation is used to globally ensure consistent nonlocal structures for pixels with similar features. The two representations are combined to generate spatially and temporally consistent video mattes. We test our method on lots of dataset including the benchmark dataset for image matting and dataset for video matting. Our method has achieved the best performance among all unsupervised matting methods in the public alpha matting evaluation dataset for images. Dongqing Zou, Xiaowu Chen 0001, Guangying Cao, Xiaogang Wang 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Shape2Motion: Joint Analysis of Motion Parts and Attributes From 3D ShapesabstractFor the task of mobility analysis of 3D shapes, we propose joint analysis for simultaneous motion part segmentation and motion attribute estimation, taking a single 3D model as input. The problem is significantly different from those tackled in the existing works which assume the availability of either a pre-existing shape segmentation or multiple 3D models in different motion states. To that end, we develop Shape2Motion which takes a single 3D point cloud as input, and jointly computes a mobility-oriented segmentation and the associated motion attributes. Shape2Motion is comprised of two deep neural networks designed for mobility proposal generation and mobility optimization, respectively. The key contribution of these networks is the novel motion-driven features and losses used in both motion part segmentation and motion attribute estimation. This is based on the observation that the movement of a functional part preserves the shape structure. We evaluate Shape2Motion with a newly proposed benchmark for mobility analysis of 3D shapes. Results demonstrate that our method achieves the state-of-the-art performance both in terms of motion part segmentation and motion attribute estimation. Xiaogang Wang 0005, Yahao Shi, Xiaowu Chen 0001, Qinping Zhao, Kai Xu 0004 |
CVPR | 4 |
| 2019 | Modeling yarn-level geometry from a single micro-imageabstractDifferent types of cloth show distinctive appearances owing to their unique yarn-level geometrical details. Despite its importance in applications such as cloth rendering and simulation, capturing yarn-level geometry is nontrivial and requires special hardware, e.g., computed tomography scanners, for conventional methods. In this paper, we propose a novel method that can produce the yarn-level geometry of real cloth using a single micro-image, captured by a consumer digital camera with a macro lens. Given a single input image, our method estimates the large-scale yarn geometry by image shading, and the fine-scale fiber details can be recovered via the proposed fiber tracing and generation algorithms. Experimental results indicate that our method can capture the detailed yarn-level geometry of a wide range of cloth and reproduce plausible cloth appearances. Xiaowu Chen 0001, Chen-Xu Zhang, Qinping Zhao |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2019 | Patch-based self-adaptive matting for high-resolution image and video
Guangying Cao, Xiaowu Chen 0001, Zhiqiang He 0002 |
Vis. Comput. | 3 |
| 2018 | Reconstructing non-rigid object with large movement using a single depth camera
Feixiang Lu, Feng Lu 0005, Yu Zhang 0035, Xiaowu Chen 0001, Qinping Zhao |
Comput. Aided Geom. Des. | 5 |
| 2018 | Image-guided 3D model labeling via multiview alignment
Kan Guo, Xiaowu Chen 0001, Qinping Zhao |
Graph. Model. | 2 |
| 2018 | Modeling Garment Seam from a Single Image
Chen-Xu Zhang, Xiaowu Chen 0001 |
J. Comput. Sci. Technol. | 2 |
| 2018 | SymPS: BRDF Symmetry Guided Photometric Stereo for Shape and Light Source EstimationabstractWe propose uncalibrated photometric stereo methods that address the problem due to unknown isotropic reflectance. At the core of our methods is the notion of "constrained half-vector symmetry" for general isotropic BRDFs. We show that such symmetry can be observed in various real-world materials, and it leads to new techniques for shape and light source estimation. Based on the 1D and 2D representations of the symmetry, we propose two methods for surface normal estimation; one focuses on accurate elevation angle recovery for surface normals when the light sources only cover the visible hemisphere, and the other for comprehensive surface normal optimization in the case that the light sources are also non-uniformly distributed. The proposed robust light source estimation method also plays an essential role to let our methods work in an uncalibrated manner with good accuracy. Quantitative evaluations are conducted with both synthetic and real-world scenes, which produce the state-of-the-art accuracy for all of the non-Lambertian materials in MERL database and the real-world datasets. Feng Lu 0005, Xiaowu Chen 0001, Imari Sato, Yoichi Sato 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Semantic Object Segmentation in Tagged Videos via DetectionabstractSemantic object segmentation (SOS) is a challenging task in computer vision that aims to detect and segment all pixels of the objects within predefined semantic categories. In image-based SOS, many supervised models have been proposed and achieved impressive performances due to the rapid advances of well-annotated training images and machine learning theories. However, in video-based SOS it is often difficult to directly train a supervised model since most videos are weakly annotated by tags. To handle such tagged videos, this paper proposes a novel approach that adopts a segmentation-by-detection framework. In this framework, object detection and segment proposals are first generated using the models pre-trained on still images, which provide useful cues to roughly localize the semantic objects. Based on these proposals, we propose an efficient algorithm to initialize object tracks by solving a joint assignment problem. As such tracks provide rough spatiotemporal configurations of the semantic objects, a voting-based refinement algorithm is further proposed to improve their spatiotemporal consistency. Extensive experiments demonstrate that the proposed framework can robustly and effectively segment semantic objects in tagged videos, even when the image-based object detectors provide inaccurate proposals. On various public benchmarks, the proposed approach obtains substantial improvements over the state-of-the-arts. Yu Zhang 0035, Xiaowu Chen 0001, Jia Li 0003, Changqun Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Copy and Paste: Temporally Consistent Stereoscopic Video BlendingabstractWe propose a novel method of stereoscopic video blending, targeted at achieving temporal disparity and color consistency. Video blending is one of the most frequent and important tasks in video editing, which is also true in stereoscopic video editing. However, it is more difficult to achieve temporally consistent blending for stereoscopic videos compared with blending of monocular videos, since there is more channel, namely disparity, to be considered except color channels in stereoscopic videos. Toward this end, two algorithms are proposed in this paper for temporally consistent videos blending. One is a temporally coherent mask propagation mechanism for selecting a source video patch clip from the source stereoscopic video; and the other is a temporal blending algorithm, which seeks to adjust the shape of the source video patch clip so as to keep consistent with disparities of the target stereoscopic video. We show various results on numerous examples to demonstrate the effectiveness and efficiency of our method. Zongji Wang, Xiaowu Chen 0001, Dongqing Zou |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | A Benchmark Dataset and Saliency-Guided Stacked Autoencoders for Video-Based Salient Object DetectionabstractImage-based salient object detection (SOD) has been extensively studied in past decades. However, video-based SOD is much less explored due to the lack of large-scale video datasets within which salient objects are unambiguously defined and annotated. Toward this end, this paper proposes a video-based SOD dataset that consists of 200 videos. In constructing the dataset, we manually annotate all objects and regions over 7650 uniformly sampled keyframes and collect the eye-tracking data of 23 subjects who free-view all videos. From the user data, we find that salient objects in a video can be defined as objects that consistently pop-out throughout the video, and objects with such attributes can be unambiguously annotated by combining manually annotated object/region masks with eye-tracking data of multiple subjects. To the best of our knowledge, it is currently the largest dataset for video-based salient object detection. Based on this dataset, this paper proposes an unsupervised baseline approach for video-based SOD by using saliency-guided stacked autoencoders. In the proposed approach, multiple spatiotemporal saliency cues are first extracted at the pixel, superpixel, and object levels. With these saliency cues, stacked autoencoders are constructed in an unsupervised manner that automatically infers a saliency score for each pixel by progressively encoding the high-dimensional saliency cues gathered from the pixel and its spatiotemporal neighbors. In experiments, the proposed unsupervised approach is compared with 31 state-of-the-art models on the proposed dataset and outperforms 30 of them, including 19 image-based classic (unsupervised or non-deep learning) models, six image-based deep learning models, and five video-based unsupervised models. Moreover, benchmarking results show that the proposed dataset is very challenging and has the potential to boost the development of video-based SOD. Jia Li 0003, Changqun Xia, Xiaowu Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Exploring Weakly Labeled Images for Video Object Segmentation With Submodular Proposal SelectionabstractVideo object segmentation (VOS) is important for various computer vision problems, and handling it with minimal human supervision is highly desired for the large-scale applications. To bring down the supervision, existing approaches largely follow a data mining perspective by assuming the availability of multiple videos sharing the same object categories. It, however, would be problematic for the tasks that consume a single video. To address this problem, this paper proposes a novel approach that explores weakly labeled images to solve video object segmentation. Given a video labeled with a target category, images labeled with the same category are collected, from which noisy object exemplars are automatically discovered. After that the proposed approach extracts a set of region proposals on various frames and efficiently matches them with massive noisy exemplars in terms of appearance and spatial context. We then jointly select the best proposals across the video by solving a novel submodular problem that combines region voting and global region matching. Finally, the localization results are leveraged as strong supervision to guide pixel-level segmentation. Extensive experiments are conducted on two challenging public databases: Youtube-Objects and DAVIS. The results suggest that the proposed approach improves over previous weakly supervised/unsupervised approaches significantly, showing a performance even comparable with the several approaches supervised by the costly manual segmentations. Yu Zhang 0035, Xiaowu Chen 0001, Jia Li 0003, Wei Teng, Haokun Song |
IEEE Trans. Image Process. | 2 |
| 2018 | Single Image Dehazing Using Ranking Convolutional Neural NetworkabstractSingle image dehazing, which aims to recover the clear image solely from an input hazy or foggy image, is a challenging ill-posed problem. Analyzing existing approaches, the common key step is to estimate the haze density of each pixel. To this end, various approaches oftenheuristically designedhaze-relevant features. Several recent works also automatically learn the features via directly exploiting convolutional neural networks (CNN). However, it may be insufficient to fully capture the intrinsic attributes of hazy images. To obtain effective features for single image dehazing, this paper presents a novel ranking convolutional neural network (Ranking-CNN). In Ranking-CNN, a novel ranking layer is proposed to extend the structure of CNN so that the statistical and structural attributes of hazy images can be simultaneously captured. By training Ranking-CNN in a well-designed manner, powerful haze-relevant features can beautomatically learnedfrom massive hazy image patches. Based on these features, haze can be effectively removed by using a haze density prediction model trained through the random forest regression. Experimental results show that our approach outperforms several previous dehazing approaches on synthetic and real-world benchmark images. Comprehensive analyses are also conducted to interpret the proposed Ranking-CNN from both the theoretical and experimental aspects. Yafei Song 0002, Jia Li 0003, Xiaogang Wang 0005, Xiaowu Chen 0001 |
IEEE Trans. Multim. | 4 |
| 2018 | Learning to group and label fine-grained shape componentsabstractA majority of stock 3D models in modern shape repositories are assembled with many fine-grained components. The main cause of such data form is the component-wise modeling process widely practiced by human modelers. These modeling components thus inherently reflect some function-based shape decomposition the artist had in mind during modeling. On the other hand, modeling components represent an over-segmentation since a functional part is usually modeled as a multi-component assembly. Based on these observations, we advocate that labeled segmentation of stock 3D models should not overlook the modeling components and propose a learning solution to grouping and labeling of the fine-grained components. However, directly characterizing the shape of individual components for the purpose of labeling is unreliable, since they can be arbitrarily tiny and semantically meaningless. We propose to generate part hypotheses from the components based on a hierarchical grouping strategy, and perform labeling on those part groups instead of directly on the components. Part hypotheses are mid-level elements which are more probable to carry semantic information. A multi-scale 3D convolutional neural network is trained to extract context-aware features for the hypotheses. To accomplish a labeled segmentation of the whole shape, we formulate higher-order conditional random fields (CRFs) to infer an optimal label assignment for all components. Extensive experiments demonstrate that our method achieves significantly robust labeling results on raw 3D models from public shape repositories. Our work also contributes the first benchmark for component-wise labeling. Xiaogang Wang 0005, Haiyue Fang, Xiaowu Chen 0001, Qinping Zhao, Kai Xu 0004 |
ACM Trans. Graph. | 4 |
| 2018 | Efficiently consistent affinity propagation for 3D shapes co-segmentation
Xiaogang Wang 0005, Zongji Wang, Dongqing Zou, Xiaowu Chen 0001, Qinping Zhao |
Vis. Comput. | 5 |
| 2017 | What is and What is Not a Salient Object? Learning Salient Object Detector by Ensembling Linear Exemplar RegressorsabstractFinding what is and what is not a salient object can be helpful in developing better features and models in salient object detection (SOD). In this paper, we investigate the images that are selected and discarded in constructing a new SOD dataset and find that many similar candidates, complex shape and low objectness are three main attributes of many non-salient objects. Moreover, objects may have diversified attributes that make them salient. As a result, we propose a novel salient object detector by ensembling linear exemplar regressors. We first select reliable foreground and background seeds using the boundary prior and then adopt locally linear embedding (LLE) to conduct manifold-preserving foregroundness propagation. In this manner, a foregroundness map can be generated to roughly pop-out salient objects and suppress non-salient ones with many similar candidates. Moreover, we extract the shape, foregroundness and attention descriptors to characterize the extracted object proposals, and a linear exemplar regressor is trained to encode how to detect salient proposals in a specific image. Finally, various linear exemplar regressors are ensembled to form a single detector that adapts to various scenarios. Extensive experimental results on 5 dataset and the new SOD dataset show that our approach outperforms 9 state-of-art methods. Changqun Xia, Jia Li 0003, Xiaowu Chen 0001, Anlin Zheng, Yu Zhang 0035 |
CVPR | 3 |
| 2017 | Primary Video Object Segmentation via Complementary CNNs and Neighborhood Reversible FlowabstractThis paper proposes a novel approach for segmenting primary video objects by using Complementary Convolutional Neural Networks (CCNN) and neighborhood reversible flow. The proposed approach first pre-trains CCNN on massive images with manually annotated salient objects in an end-to-end manner, and the trained CCNN has two separate branches that simultaneously handle two complementary tasks, i.e., foregroundness and backgroundness estimation. By applying CCNN on each video frame, the spatial foregroundness and backgroundness maps can be initialized, which are then propagated between various frames so as to segment primary video objects and suppress distractors. To enforce efficient temporal propagation, we divide each frame into superpixels and construct neighborhood reversible flow that reflects the most reliable temporal correspondences between superpixels in far-away frames. Within such flow, the initialized foregroundness and backgroundness can be efficiently and accurately propagated along the temporal axis so that primary video objects gradually pop-out and distractors are well suppressed. Extensive experimental results on three video datasets show that the proposed approach achieves impressive performance in comparisons with 18 state-of-the-art models. Jia Li 0003, Anlin Zheng, Xiaowu Chen 0001 |
ICCV | 3 |
| 2017 | Look, Perceive and Segment: Finding the Salient Objects in Images via Two-stream Fixation-Semantic CNNsabstractRecently, CNN-based models have achieved remarkable success in image-based salient object detection (SOD). In these models, a key issue is to find a proper network architecture that best fits for the task of SOD. Toward this end, this paper proposes two-stream fixation-semantic CNNs, whose architecture is inspired by the fact that salient objects in complex images can be unambiguously annotated by selecting the pre-segmented semantic objects that receive the highest fixation density in eye-tracking experiments. In the two-stream CNNs, a fixation stream is pre-trained on eye-tracking data whose architecture well fits for the task of fixation prediction, and a semantic stream is pre-trained on images with semantic tags that has a proper architecture for semantic perception. By fusing these two streams into an inception-segmentation module and jointly fine-tuning them on images with manually annotated salient objects, the proposed networks show impressive performance in segmenting salient objects. Experimental results show that our approach outperforms 10 state-of-the-art models (5 deep, 5 non-deep) on 4 datasets. Xiaowu Chen 0001, Anlin Zheng, Jia Li 0003, Feng Lu 0005 |
ICCV | 1 |
| 2017 | Embedding 3D Geometric Features for Rigid Object Part SegmentationabstractObject part segmentation is a challenging and fundamental problem in computer vision. Its difficulties may be caused by the varying viewpoints, poses, and topological structures, which can be attributed to an essential reason, i.e., a specific object is a 3D model rather than a 2D figure. Therefore, we conjecture that not only 2D appearance features but also 3D geometric features could be helpful. With this in mind, we propose a 2-stream FCN. One stream, named AppNet, is to extract 2D appearance features from the input image. The other stream, named GeoNet, is to extract 3D geometric features. However, the problem is that the input is just an image. To this end, we design a 2D convolution based CNN structure to extract 3D geometric features from 3D volume, which is named VolNet. Then a teacher-student strategy is adopted and VolNet teaches GeoNet how to extract 3D geometric features from an image. To perform this teaching process, we synthesize training data using 3D models. Each training sample consists of an image and its corresponding volume. A perspective voxelization algorithm is further proposed to align them. Experimental results verify our conjecture and the effectiveness of both the proposed 2-stream CNN and VolNet. Yafei Song 0002, Xiaowu Chen 0001, Jia Li 0003, Qinping Zhao |
ICCV | 2 |
| 2017 | Appearance-Based Gaze Estimation via Uncalibrated Gaze Pattern RecoveryabstractAiming at reducing the restrictions due to person/scene dependence, we deliver a novel method that solves appearance-based gaze estimation in a novel fashion. First, we introduce and solve an "uncalibrated gaze pattern" solely from eye images independent of the person and scene. The gaze pattern recovers gaze movements up to only scaling and translation ambiguities, via nonlinear dimension reduction and pixel motion analysis, while no training/calibration is needed. This is new in the literature and enables novel applications. Second, our method allows simple calibrations to align the gaze pattern to any gaze target. This is much simpler than conventional calibrations which rely on sufficient training data to compute person and scene-specific nonlinear gaze mappings. Through various evaluations, we show that: 1) the proposed uncalibrated gaze pattern has novel and broad capabilities; 2) the proposed calibration is simple and efficient, and can be even omitted in some scenarios; and 3) quantitative evaluations produce promising results under various conditions. Feng Lu 0005, Xiaowu Chen 0001, Yoichi Sato 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Learning Discriminative Subspaces on Random Contrasts for Image Saliency AnalysisabstractIn visual saliency estimation, one of the most challenging tasks is to distinguish targets and distractors that share certain visual attributes. With the observation that such targets and distractors can sometimes be easily separated when projected to specific subspaces, we propose to estimate image saliency by learning a set of discriminative subspaces that perform the best in popping out targets and suppressing distractors. Toward this end, we first conduct principal component analysis on massive randomly selected image patches. The principal components, which correspond to the largest eigenvalues, are selected to construct candidate subspaces since they often demonstrate impressive abilities to separate targets and distractors. By projecting images onto various subspaces, we further characterize each image patch by its contrasts against randomly selected neighboring and peripheral regions. In this manner, the probable targets often have the highest responses, while the responses at background regions become very low. Based on such random contrasts, an optimization framework with pairwise binary terms is adopted to learn the saliency model that best separates salient targets and distractors by optimally integrating the cues from various subspaces. Experimental results on two public benchmarks show that the proposed approach outperforms 16 state-of-the-art methods in human fixation prediction. Shu Fang, Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001, Xiaowu Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2017 | Adaptive synthesis of indoor scenes via activity-associated object relation graphsabstractWe present a system for adaptive synthesis of indoor scenes given an empty room and only a few object categories. Automatically suggesting indoor objects and proper layouts to convert an empty room to a 3D scene is challenging, since it requires interior design knowledge to balance the factors like space, path distance, illumination and object relations, in order to insure the functional plausibility of the synthesized scenes. We exploit a database of 2D floor plans to extract object relations and provide layout examples for scene synthesis. With the labeled human positions and directions in each plan, we detect the activity relations and compute the coexistence frequency of object pairs to construct activity-associated object relation graphs. Given the input room and user-specified object categories, our system first leverages the object relation graphs and the database floor plans to suggest more potential object categories beyond the specified ones to make resulting scenes functionally complete, and then uses the similar plan references to create the layout of synthesized scenes. We show various synthesis results to demonstrate the practicability of our system, and validate its usability via a user study. We also compare our system with the state-of-the-art furniture layout and activity-centric scene representation methods, in terms of functional plausibility and user friendliness. Qiang Fu 0004, Xiaowu Chen 0001, Sijia Wen, Hongbo Fu 0001 |
ACM Trans. Graph. | 2 |
| 2017 | Pose-Inspired Shape Synthesis and Functional HybridabstractWe introduce a shape synthesis approach especially for functional hybrid creation that can be potentially used by a human operator under a certain pose. Shape synthesis by reusing parts in existing models has been an active research topic in recent years. However, how to combine models across different categories to design multi-function objects remains challenging, since there is no natural correspondence between models across different categories. We tackle this problem by introducing a human pose to describe object affordance which establishes a bridge between cross-class objects for composite design. Specifically, our approach first identifies groups of candidate shapes which provide affordances desired by an input human pose, and then recombines them as well-connected composite models. Users may control the design process by manipulating the input pose, or optionally specifying one or more desired categories. We also extend our approach to be used by a single operator with multiple poses or by multiple human operators. We show that our approach enables easy creation of nontrivial, interesting synthesized models. Qiang Fu 0004, Xiaowu Chen 0001, Xiaoyu Su, Hongbo Fu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2016 | Local Shape Transfer for Image Co-segmentation
Wei Teng, Yu Zhang 0035, Xiaowu Chen 0001, Jia Li 0003, Zhiqiang He 0002 |
BMVC | 3 |
| 2016 | Modeling interactive furniture from a single image
Xiaowu Chen 0001, Dongqing Zou |
Comput. Graph. | 2 |
| 2016 | Cross-class 3D object synthesis guided by reference examples
Xiaoyu Su, Xiaowu Chen 0001, Qiang Fu 0004, Hongbo Fu 0001 |
Comput. Graph. | 2 |
| 2016 | Structure-adaptive Shape Editing for Man-made ObjectsabstractAbstract One of the challenging problems for shape editing is to adapt shapes with diversified structures for various editing needs. In this paper we introduce a shape editing approach that automatically adapts the structure of a shape being edited with respect to user inputs. Given a category of shapes, our approach first classifies them into groups based on the constituent parts. The group‐sensitive priors, including both inter‐group and intra‐group priors, are then learned through statistical structure analysis and multivariate regression. By using these priors, the inherent characteristics and typical variations of shape structures can be well captured. Based on such group‐sensitive priors, we propose a framework for real‐time shape editing, which adapts the structure of shape to continuous user editing operations. Experimental results show that the proposed approach is capable of both structure‐preserving and structure‐varying shape editing. Qiang Fu 0004, Xiaowu Chen 0001, Xiaoyu Su, Jia Li 0003, Hongbo Fu 0001 |
Comput. Graph. Forum | 2 |
| 2016 | High-level representation sketch for video event retrieval
Yu Zhang 0035, Xiaowu Chen 0001, Liang Lin 0004, Changqun Xia, Dongqing Zou |
Sci. China Inf. Sci. | 2 |
| 2016 | Natural lines inspired 3D shape re-design
Qiang Fu 0004, Xiaowu Chen 0001, Xiaoyu Su, Hongbo Fu 0001 |
Graph. Model. | 2 |
| 2016 | Measuring Visual Surprise Jointly from Intrinsic and Extrinsic Contexts for Image Saliency Estimation
Jia Li 0003, Yonghong Tian 0001, Xiaowu Chen 0001, Tiejun Huang 0001 |
Int. J. Comput. Vis. | 3 |
| 2016 | Person-independent eye gaze prediction from eye images using patch-based features
Feng Lu 0005, Xiaowu Chen 0001 |
Neurocomputing | 2 |
| 2016 | Fuzzy community detection via modularity guided membership-degree propagation
Xiaowu Chen 0001, Jia Li 0003 |
Pattern Recognit. Lett. | 2 |
| 2016 | Learn Sparse Dictionaries for Edit PropagationabstractWith the increasing availability of high-resolution images, videos, and 3D models, the demand for scalable large data processing techniques increases. We introduce a method of sparse dictionary learning for edit propagation of large input data. Previous approaches for edit propagation typically employ a global optimization over the whole set of pixels (or vertexes), incurring a prohibitively high memory and time-consumption for large input data. Rather than propagating an edit pixel by pixel, we follow the principle of sparse representation to obtain a representative and compact dictionary and perform edit propagation on the dictionary instead. The sparse dictionary provides an intrinsic basis for input data, and the coding coefficients capture the linear relationship between all pixels and the dictionary atoms. The learned dictionary is then optimized by a novel scheme, which maximizes the Kullback-Leibler divergence between each atom pair to remove redundant atoms. To enable local edit propagation for images or videos with similar appearance, a dictionary learning strategy is proposed by considering range constraint to better account for the global distribution of pixels in their feature space. We show several applications of the sparsity-based edit propagation, including video recoloring, theme editing, and seamless cloning, operating on both color and texture features. Our approach can also be applied to computer graphics tasks, such as 3D surface deformation. We demonstrate that with an atom-to-pixel ratio in the order of 0.01% signifying a significant reduction on memory consumption, our method still maintains a high degree of visual fidelity. Xiaowu Chen 0001, Dongqing Zou, Qinping Zhao |
IEEE Trans. Image Process. | 1 |
| 2016 | Estimating 3D Gaze Directions Using Unlabeled Eye Images via Synthetic Iris Appearance FittingabstractEstimating three-dimensional (3D) human eye gaze by capturing a single eye image without active illumination is challenging. Although the elliptical iris shape provides a useful cue, existing methods face difficulties in ellipse fitting due to unreliable iris contour detection. These methods may fail frequently especially with low resolution eye images. In this paper, we propose a synthetic iris appearance fitting (SIAF) method that is model-driven to compute 3D gaze direction from iris shape. Instead of fitting an ellipse based on exactly detected iris contour, our method first synthesizes a set of physically possible iris appearances and then optimizes inside this synthetic space to find the best solution to explain the captured eye image. In this way, the solution is highly constrained and guaranteed to be physically feasible. In addition, the proposed advanced image analysis techniques also help the SIAF method be robust to the unreliable iris contour detection. Furthermore, with multiple eye images, we propose a SIAF-joint method that can further reduce the gaze error by half, and it also resolves the binary ambiguity which is inevitable in conventional methods based on simple ellipse fitting. Feng Lu 0005, Yue Gao 0002, Xiaowu Chen 0001 |
IEEE Trans. Multim. | 3 |
| 2016 | 6-DOF Image Localization From Massive Geo-Tagged Reference ImagesabstractThe 6-degrees of freedom (DOF) image localization, which aims to calculate the spatial position and rotation of a camera, is a challenging problem for most location-based services. In existing approaches, this problem is often tackled by finding the matches between 2D image points and 3D structure points so as to derive the location information via direct linear transformation algorithm. However, as these 2D-to-3D-based approaches need to reconstruct the 3D structure points of the scene, they may not be flexible enough to employ massive and increasing geo-tagged data. To this end, this paper presents a novel approach for 6-DOF image localization by fusing candidate poses relative to reference images. In this approach, we propose to localize an input image according to the position and rotation information of multiple geo-tagged images retrieved from a reference dataset. From the reference images, an efficient relative pose estimation algorithm is proposed to derive a set of candidate poses for the input image. Each candidate pose encodes the relative rotation and direction of the input image with respect to a specific reference image. Finally, these candidate poses can be fused together by minimizing a well-defined geometry error so that the 6-DOF location of the input image is effectively derived. Experimental results show that our method can obtain satisfactory localization accuracy. In addition, the proposed relative pose estimation algorithm is much faster than existing work. Yafei Song 0002, Xiaowu Chen 0001, Xiaogang Wang 0005, Yu Zhang 0035, Jia Li 0003 |
IEEE Trans. Multim. | 2 |
| 2015 | Semantic object segmentation via detection in weakly labeled videoabstractSemantic object segmentation in video is an important step for large-scale multimedia analysis. In many cases, however, semantic objects are only tagged at video-level, making them difficult to be located and segmented. To address this problem, this paper proposes an approach to segment semantic objects in weakly labeled video via object detection. In our approach, a novel video segmentation-by-detection framework is proposed, which first incorporates object and region detectors pre-trained on still images to generate a set of detection and segmentation proposals. Based on the noisy proposals, several object tracks are then initialized by solving a joint binary optimization problem with min-cost flow. As such tracks actually provide rough configurations of semantic objects, we thus refine the object segmentation while preserving the spatiotemporal consistency by inferring the shape likelihoods of pixels from the statistical information of tracks. Experimental results on Youtube-Objects dataset and SegTrack v2 dataset demonstrate that our method outperforms state-of-the-arts and shows impressive results. Yu Zhang 0035, Xiaowu Chen 0001, Jia Li 0003, Changqun Xia |
CVPR | 2 |
| 2015 | Conformal and Low-Rank Sparse Representation for Image RestorationabstractObtaining an appropriate dictionary is the key point when sparse representation is applied to computer vision or image processing problems such as image restoration. It is expected that preserving data structure during sparse coding and dictionary learning can enhance the recovery performance. However, many existing dictionary learning methods handle training samples individually, while missing relationships between samples, which result in dictionaries with redundant atoms but poor representation ability. In this paper, we propose a novel sparse representation approach called conformal and low-rank sparse representation (CLRSR) for image restoration problems. To achieve a more compact and representative dictionary, conformal property is introduced by preserving the angles of local geometry formed by neighboring samples in the feature space. Furthermore, imposing low-rank constraint on the coefficient matrix can lead more faithful subspaces and capture the global structure of data. We apply our CLRSR model to several image restoration tasks to demonstrate the effectiveness. Xiaowu Chen 0001, Dongqing Zou, Wei Teng |
ICCV | 2 |
| 2015 | A Data-Driven Metric for Comprehensive Evaluation of Saliency ModelsabstractIn the past decades, hundreds of saliency models have been proposed for fixation prediction, along with dozens of evaluation metrics. However, existing metrics, which are often heuristically designed, may draw conflict conclusions in comparing saliency models. As a consequence, it becomes somehow confusing on the selection of metrics in comparing new models with state-of-the-arts. To address this problem, we propose a data-driven metric for comprehensive evaluation of saliency models. Instead of heuristically designing such a metric, we first conduct extensive subjective tests to find how saliency maps are assessed by the human-being. Based on the user data collected in the tests, nine representative evaluation metrics are directly compared by quantizing their performances in assessing saliency maps. Moreover, we propose to learn a data-driven metric by using Convolutional Neural Network. Compared with existing metrics, experimental results show that the data-driven metric performs the most consistently with the human-being in evaluating saliency maps as well as saliency models. Jia Li 0003, Changqun Xia, Yafei Song 0002, Shu Fang, Xiaowu Chen 0001 |
ICCV | 5 |
| 2015 | Video Matting via Sparse and Low-Rank RepresentationabstractWe introduce a novel method of video matting via sparse and low-rank representation. Previous matting methods [10, 9] introduced a nonlocal prior to estimate the alpha matte and have achieved impressive results on some data. However, on one hand, searching inadequate or excessive samples may miss good samples or introduce noise, on the other hand, it is difficult to construct consistent nonlocal structures for pixels with similar features, yielding spatially and temporally inconsistent video mattes. In this paper, we proposed a novel video matting method to achieve spatially and temporally consistent matting result. Toward this end, a sparse and low-rank representation model is introduced to pursue consistent nonlocal structures for pixels with similar features. The sparse representation is used to adaptively select best samples and accurately construct the nonlocal structures for all pixels, while the low-rank representation is used to globally ensure consistent nonlocal structures for pixels with similar features. The two representations are combined to generate consistent video mattes. Experimental results show that our method has achieved high quality results in a variety of challenging examples featuring illumination changes, feature ambiguity, topology changes, transparency variation, dis-occlusion, fast motion and motion blur. Dongqing Zou, Xiaowu Chen 0001, Guangying Cao, Xiaogang Wang 0005 |
ICCV | 2 |
| 2015 | Cuboids detection in RGB-D images via Maximum Weighted CliqueabstractCuboid detection is an essential step for understanding 3D structure of scenes. As most of indoor scene cuboids are actually objects, we propose in this paper an object-based approach to detect 3D cuboids in indoor RGB-D images. The proposed approach is learning-free and can handle general object classes rather than a limited pre-defined category set. In our approach, we first apply an extended version of the CPMC framework to generate a set of segment hypotheses, and fit a set of cuboid candidates. Given the candidate set, we select several cuboids that can provide plausible interpretations of the images by solving a Maximum Weighted Clique (MWC) problem. With this formulation, a set of ranked mid-level representations of the input image is obtained, and are further re-ranked by Maximal Marginal Relevance (MMR) measure to improve their diversity. Experimental results on NYU-V2 dataset shows that our method significantly outperforms the state-of-the-art, and shows impressive results. Xiaowu Chen 0001, Yu Zhang 0035, Jia Li 0003, Xiaogang Wang 0005 |
ICME | 2 |
| 2015 | Image2Scene: Transforming Style of 3D RoomabstractWe propose a style transformation system to transform a 3D room into one that resembles the style of a photograph. We focus on two major components of interior scene style: layout and color. Using an interior image database, we learn the related style guidelines. Given a reference image and a 3D room of two different interior rooms, we first establish semantic correspondence between the two scenes. The styles of the reference image are then extracted in the form of layout constraints and color schemes. Finally, our framework performs layout rearrangement followed by recoloring of the scene to match the learned style of the reference image. We show style transformation results on numerous examples to demonstrate the effectiveness and efficiency of our system. Xiaowu Chen 0001, Dongqing Zou, Qinping Zhao |
ACM Multimedia | 1 |
| 2015 | Monocular Video Guided Garment Simulation
Xiaowu Chen 0001, Fei-Xiang Lu, Kan Guo, Qiang Fu 0004 |
J. Comput. Sci. Technol. | 2 |
| 2015 | Finding the Secret of Image Saliency in the Frequency DomainabstractThere are two sides to every story of visual saliency modeling in the frequency domain. On the one hand, image saliency can be effectively estimated by applying simple operations to the frequency spectrum. On the other hand, it is still unclear which part of the frequency spectrum contributes the most to popping-out targets and suppressing distractors. Toward this end, this paper tentatively explores the secret of image saliency in the frequency domain. From the results obtained in several qualitative and quantitative experiments, we find that the secret of visual saliency may mainly hide in the phases of intermediate frequencies. To explain this finding, we reinterpret the concept of discrete Fourier transform from the perspective of template-based contrast computation and thus develop several principles for designing the saliency detector in the frequency domain. Following these principles, we propose a novel approach to design the saliency detector under the assistance of prior knowledge obtained through both unsupervised and supervised learning processes. Experimental results on a public image benchmark show that the learned saliency detector outperforms 18 state-of-the-art approaches in predicting human fixations. Jia Li 0003, Ling-Yu Duan, Xiaowu Chen 0001, Tiejun Huang 0001, Yonghong Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Image saliency estimation via random walk guided by informativeness and latent signal correlations
Jia Li 0003, Shu Fang, Yonghong Tian 0001, Tiejun Huang 0001, Xiaowu Chen 0001 |
Signal Process. Image Commun. | 5 |
| 2015 | An Object-Level High-Order Contextual Descriptor Based on Semantic, Spatial, and Scale CuesabstractContext has been playing an increasingly important role in areas such as object detection, scene understanding, and image segmentation. Although many different types of contextual cues have been successfully explored, most of them only consider the pair-wise relationship between objects or parts. Several models utilize the high-order relationship for encoding contextual information. However, they mainly use a single contextual cue. In this paper, we present a novel high-order contextual descriptor (HOOD) to measure the strength of interactions among objects within an image. Heterogeneous contextual cues like semantic, spatial, and scale contexts are jointly integrated into HOOD to define the high-order interactions. The strength of these interactions are inferred by applying Bayes' rule on the pure dependence of the involved objects. As a result, an object-level graph is constructed to represent the contextually consistent interactions. Moreover, we propose a HOOD based object localization framework to verify the effectiveness of HOOD. Experimental results on two benchmark datasets including SUN09 and PASCAL2007 show that our framework outperforms the state-of-the-art context based object localization methods. Finally, we apply HOOD on two multimedia applications: structured image retrieval and out-of-context object detection, which demonstrates the potential usages of HOOD. Xiaochun Cao, Xingxing Wei 0001, Yahong Han, Xiaowu Chen 0001 |
IEEE Trans. Cybern. | 4 |
| 2015 | Learning Templates for Artistic Portrait Lighting AnalysisabstractLighting is a key factor in creating impressive artistic portraits. In this paper, we propose to analyze portrait lighting by learning templates of lighting styles. Inspired by the experience of artists, we first define several novel features that describe the local contrasts in various face regions. The most informative features are then selected with a stepwise feature pursuit algorithm to derive the templates of various lighting styles. After that, the matching scores that measure the similarity between a testing portrait and those templates are calculated for lighting style classification. Furthermore, we train a regression model by the subjective scores and the feature responses of a template to predict the score of a portrait lighting quality. Based on the templates, a novel face illumination descriptor is defined to measure the difference between two portrait lightings. Experimental results show that the learned templates can well describe the lighting styles, whereas the proposed approach can assess the lighting quality of artistic portraits as human being does. Xiaowu Chen 0001, Xin Jin 0015, Qinping Zhao |
IEEE Trans. Image Process. | 1 |
| 2015 | Garment modeling with a depth cameraabstractPrevious garment modeling techniques mainly focus on designing novel garments to dress up virtual characters. We study the modeling of real garments and develop a system that is intuitive to use even for novice users. Our system includes garment component detectors and design attribute classifiers learned from a manually labeled garment image database. In the modeling time, we scan the garment with a Kinect and build a rough shape by KinectFusion from the raw RGBD sequence. The detectors and classifiers will identify garment components (e.g. collar, sleeve, pockets, belt, and buttons) and their design attributes (e.g. falbala collar or lapel collar, hubble-bubble sleeve or straight sleeve ) from the RGB images. Our system also contains a 3D deformable template database for garment components. Once the components and their designs are determined, we choose appropriate templates, stitch them together, and fit them to the initial garment mesh generated by KinectFusion. Experiments on various different garment styles consistently generate high quality results. Xiaowu Chen 0001, Fei-Xiang Lu, Lang Bi |
ACM Trans. Graph. | 1 |
| 2015 | 3D Mesh Labeling via Deep Convolutional Neural NetworksabstractThis article presents a novel approach for 3D mesh labeling by using deep Convolutional Neural Networks (CNNs). Many previous methods on 3D mesh labeling achieve impressive performances by using predefined geometric features. However, the generalization abilities of such low-level features, which are heuristically designed to process specific meshes, are often insufficient to handle all types of meshes. To address this problem, we propose to learn a robust mesh representation that can adapt to various 3D meshes by using CNNs. In our approach, CNNs are first trained in a supervised manner by using a large pool of classical geometric features. In the training process, these low-level features are nonlinearly combined and hierarchically compressed to generate a compact and effective representation for each triangle on the mesh. Based on the trained CNNs and the mesh representations, a label vector is initialized for each triangle to indicate its probabilities of belonging to various object parts. Eventually, a graph-based mesh-labeling algorithm is adopted to optimize the labels of triangles by considering the label consistencies. Experimental results on several public benchmarks show that the proposed approach is robust for various 3D meshes, and outperforms state-of-the-art approaches as well as classic learning algorithms in recognizing mesh labels. Kan Guo, Dongqing Zou, Xiaowu Chen 0001 |
ACM Trans. Graph. | 3 |
| 2014 | Sparse Dictionary Learning for Edit Propagation of High-Resolution ImagesabstractWe introduce a method of sparse dictionary learning for edit propagation of high-resolution images or video. Previous approaches for edit propagation typically employ a global optimization over the whole set of image pixels, incurring a prohibitively high memory and time consumption for high-resolution images. Rather than propagating an edit pixel by pixel, we follow the principle of sparse representation to obtain a compact set of representative samples (or features) and perform edit propagation on the samples instead. The sparse set of samples provides an intrinsic basis for an input image, and the coding coefficients capture the linear relationship between all pixels and the samples. The representative set of samples is then optimized by a novel scheme which maximizes the KL-divergence between each sample pair to remove redundant samples. We show several applications of sparsity-based edit propagation including video recoloring, theme editing, and seamless cloning, operating on both color and texture features. We demonstrate that with a sample-to-pixel ratio in the order of 0.01%, signifying a significant reduction on memory consumption, our method still maintains a high-degree of visual fidelity. Xiaowu Chen 0001, Dongqing Zou, Xiaochun Cao, Qinping Zhao, Hao (Richard) Zhang |
CVPR | 1 |
| 2014 | Image Retrieval and Ranking via Consistently Reconstructing Multi-attribute Queries
Xiaochun Cao, Hua Zhang 0008, Xiaojie Guo 0001, Si Liu 0001, Xiaowu Chen 0001 |
ECCV (1) | 5 |
| 2014 | Canopy-frame interactions for umbrella simulation
Xiaowu Chen 0001, Qinping Zhao |
Comput. Graph. | 2 |
| 2014 | Lighting virtual objects in a single image via coarse scene understanding
Xiaowu Chen 0001, Xin Jin 0015 |
Sci. China Inf. Sci. | 1 |
| 2014 | Structure guided texture inpainting through multi-scale patches and global optimization for image completion
Xiaowu Chen 0001, Qinping Zhao |
Sci. China Inf. Sci. | 1 |
| 2014 | Geodesic Propagation for Semantic LabelingabstractThis paper presents a semantic labeling framework with geodesic propagation (GP). Under the same framework, three algorithms are proposed, including GP, supervised GP (SGP) for image, and hybrid GP (HGP) for video. In these algorithms, we resort to the recognition proposal map and select confident pixels with maximum probability as the initial propagation seeds. From these seeds, the GP algorithm iteratively updates the weights of geodesic distances until the semantic labels are propagated to all pixels. On the contrary, the SGP algorithm further exploits the contextual information to guide the direction of propagation, leading to better performance but higher computational complexity than the GP. For video labeling, we further propose the HGP algorithm, in which the geodesic metric is used in both spatial and temporal spaces. Experiments on four public data sets show that our algorithms outperform several state-of-the-art methods. With the GP framework, convincing results for both image and video semantic labeling can be obtained. Xiaowu Chen 0001, Yafei Song 0002, Yu Zhang 0035, Xin Jin 0015, Qinping Zhao |
IEEE Trans. Image Process. | 2 |
| 2014 | Optimizing neighborhood projection with relaxation factor for inextensible cloth simulation
Xiaowu Chen 0001, Qinping Zhao, Long Quan |
Vis. Comput. | 1 |
| 2013 | Face Image Illumination Transfer through Eye-Relit 3D BasisabstractFace illumination transfer from a single image is still a challenging problem. Most current methods fail when the reference face and the target face differ obviously in pose or facial geometry. In this paper, we propose a face illumination transfer method based on eye-relit 3D basis to conquer this problem. We extract the environment map from one eye of the reference image, preventing the illumination misestimation caused by the reference's pose and facial geometry. By relighting our 3D basis with the map, our transfer result preserves the illumination of the reference image, and the pose and geometry of the target image. Experiments demonstrate our method's good performance in above-mentioned challenging cases. Mengxia Yang, Zhihong Fang, Xiaowu Chen 0001 |
CAD/Graphics | 4 |
| 2013 | Network-Based Clustering and Embedding for High-Dimensional Data VisualizationabstractWe present a novel method to visualize high-dimensional dataset as a landscape. The goal is to provide clear and compact representation to reveal the structure of high-dimensional datasets in a way that the size and distinctiveness of clusters can be easily discerned, and the relationships among single points can be preserved. Our method is network-based, and consists of two main steps: clustering and embedding. First of all, the similarity graph of high-dimensional dataset is constructed based on the Euclidean distances between data points. For clustering, we propose a new network community detection algorithm to calculate the membership-degree of each vertex belonging to each community. For embedding, we bring forward a practical algorithm to obtain an evenly distributed and regularly shaped layout of data points, in a way that the original relationships among single points are preserved. Finally, the landscape-like visualization is produced by assigning altitudes to data points according to their membership-degrees and by inserting control points. In our high-dimensional data visualization, clusters form highlands, and border data points among clusters show up as valleys. The area and altitude of highland indicate the size and distinctiveness of data cluster respectively. Xiaowu Chen 0001 |
CAD/Graphics | 2 |
| 2013 | Image Matting with Local and Nonlocal Smooth PriorsabstractIn this paper we propose a novel alpha matting method with local and nonlocal smooth priors. We observe that the manifold preserving editing propagation [4] essentially introduced a nonlocal smooth prior on the alpha matte. This nonlocal smooth prior and the well known local smooth prior from matting Laplacian complement each other. So we combine them with a simple data term from color sampling in a graph model for nature image matting. Our method has a closed-form solution and can be solved efficiently. Compared with the state-of-the-art methods, our method produces more accurate results according to the evaluation on standard benchmark datasets. Xiaowu Chen 0001, Dongqing Zou, Steven Zhiying Zhou, Qinping Zhao |
CVPR | 1 |
| 2013 | Video Editing with Temporal, Spatial and Appearance ConsistencyabstractGiven an area of interest in a video sequence, one may want to manipulate or edit the area, e.g. remove occlusions from or replace with an advertisement on it. Such a task involves three main challenges including temporal consistency, spatial pose, and visual realism. The proposed method effectively seeks an optimal solution to simultaneously deal with temporal alignment, pose rectification, as well as precise recovery of the occlusion. To make our method applicable to long video sequences, we propose a batch alignment method for automatically aligning and rectifying a small number of initial frames, and then show how to align the remaining frames incrementally to the aligned base images. From the error residual of the robust alignment process, we automatically construct a trimap of the region for each frame, which is used as the input to alpha matting methods to extract the occluding foreground. Experimental results on both simulated and real data demonstrate the accurate and robust performance of our method. Xiaojie Guo 0001, Xiaochun Cao, Xiaowu Chen 0001, Yi Ma 0001 |
CVPR | 3 |
| 2013 | Data-Driven Season Characteristic Enhancement of Natural ImageabstractWe present a system of creating new scenes in different seasons from an input image captured in a particular season by stylizing it according to similar images in our library which includes a vast number of different season scenes and objects. Firstly, we transfer the color appearance of the input scene in accordance with the color style of other seasons scenes by using color transfer approach. Secondly, user scribbles are used to guide the detection of repeated elements in the input image and geometric information of these detected elements are obtained. Then context-sensitive objects of specified class that match most of the required properties are retrieved from our library and edited according to the geometric information of the specified elements in the scene, such as inserting new objects into the scene or replacing objects by new ones. After that, a repeated-elements-based modification duplication scheme is proposed and implemented to semi-automatically propagate the user-specified objects modification. Finally, blending is applied to the inserted objects to achieve consistency with the target scene in illumination. The main contribution of this work is that we present a complete system to create new scenes of different seasons. Dongqing Zou, Xiaowu Chen 0001 |
ICIG | 4 |
| 2013 | Garment Modeling from a Single ImageabstractAbstract Modeling of realistic garments is essential for online shopping and many other applications including virtual characters. Most of existing methods either require a multi‐camera capture setup or a restricted mannequin pose. We address the garment modeling problem according to a single input image. We design an all‐pose garment outline interpretation, and a shading‐based detail modeling algorithm. Our method first estimates the mannequin pose and body shape from the input image. It further interprets the garment outline with an oriented facet decided according to the mannequin pose to generate the initial 3D garment model. Shape details such as folds and wrinkles are modeled by shape‐from‐shading techniques, to improve the realism of the garment model. Our method achieves similar result quality as prior methods from just a single image, significantly improving the flexibility of garment modeling. Xiaowu Chen 0001, Qiang Fu 0004, Kan Guo |
Comput. Graph. Forum | 2 |
| 2013 | Occlusion cues for image scene layering
Xiaowu Chen 0001, Dongyue Zhao, Qinping Zhao |
Comput. Vis. Image Underst. | 1 |
| 2013 | Facial performance illumination transfer from a single video using interpolation in non-skin regionabstractABSTRACT This paper proposes a novel video‐based method to transfer the illumination from a single reference facial performance video to a target one taken under nearly uniform illumination. We first filter the key frames of the reference and the target face videos with an edge‐preserving filter. Then, the illumination component of reference key frame is extracted through dividing the filtered reference key frames by the corresponding filtered target key frames in skin region. The differences in non‐skin region caused by different expressions between the reference and target face may bring about artifacts. Therefore, we interpolate the illumination component of the non‐skin region by that of the surrounded skin region to ensure the spatial smoothness and consistency. After that, the illumination components of key frames are propagated to non‐key frames to ensure the temporal consistency between the two adjacent frames. We obtain convincing results by transferring the illumination effects of a single reference facial performance video to a target one with the spatial and temporal consistencies preserved. Copyright © 2013 John Wiley & Sons, Ltd. Xiaowu Chen 0001, Mengxia Yang, Zhihong Fang |
Comput. Animat. Virtual Worlds | 2 |
| 2013 | Face Illumination Manipulation Using a Single Reference Image by Adaptive Layer DecompositionabstractThis paper proposes a novel image-based framework to manipulate the illumination of human face through adaptive layer decomposition. According to our framework, only a single reference image, without any knowledge of the 3D geometry or material information of the input face, is needed. To transfer the illumination effects of a reference face image to a normal lighting face, we first decompose the lightness layers of the reference and the input images into large-scale and detail layers through weighted least squares (WLS) filter with adaptive smoothing parameters according to the gradient values of the face images. The large-scale layer of the reference image is filtered with the guidance of the input image by guided filter with adaptive smoothing parameters according to the face structures. The relit result is obtained by replacing the largescale layer of the input image with that of the reference image. To normalize the illumination effects of a non-normal lighting face (i.e., face delighting), we introduce similar reflectance prior to the layer decomposition stage by WLS filter, which make the normalized result less affected by the high contrast light and shadow effects of the input face. Through these two procedures, we can change the illumination effects of a non-normal lighting face by first normalizing the illumination and then transferring the illumination of another reference face to it. We acquire convincing relit results of both face relighting and delighting on numerous input and reference face images with various illumination effects and genders. Comparisons with previous papers show that our framework is less affected by geometry differences and can preserve better the identification structure and skin color of the input face. Xiaowu Chen 0001, Xin Jin 0015, Qinping Zhao |
IEEE Trans. Image Process. | 1 |
| 2013 | Deformable model for estimating clothed and naked human shapes from a single image
Xiaowu Chen 0001, Qinping Zhao |
Vis. Comput. | 1 |
| 2012 | Clothed and Naked Human Shapes Estimation from a Single Image
Xiaowu Chen 0001, Qinping Zhao |
CVM | 2 |
| 2012 | Supervised Geodesic Propagation for Semantic Label Transfer
Xiaowu Chen 0001, Yafei Song 0002, Xin Jin 0015, Qinping Zhao |
ECCV (3) | 1 |
| 2012 | Artistic Illumination Transfer for PortraitsabstractAbstract Relighting a portrait in a single image is still a challenging problem, particularly when only a single artistic reference photograph or painting is provided. In this paper, we propose an artistic illumination transfer system for portraits based on a database of portrait images (photographs and paintings) associated with hand‐drawn illumination templates (276) by artists. Users can select a reference portrait image in the database, and the corresponding illumination template is transferred to an input portrait using image warping. Users can also provide reference portrait images those are not in the database. Based on the Face Illumination Descriptor (FID), the system selects from the database the reference image with the closest illumination to that of the user‐provided reference image and adjusts the corresponding illumination template to match the contrast of the user‐provided reference image. Experiments on not only paintings but also photographs, paper‐cuts and sketches demonstrate that convincing illumination transferred results can be rendered by our system. Xiaowu Chen 0001, Xin Jin 0015, Qinping Zhao |
Comput. Graph. Forum | 1 |
| 2012 | Video motion stitching using trajectory and position similarities
Xiaowu Chen 0001, Qinping Zhao |
Sci. China Inf. Sci. | 1 |
| 2012 | Virtual calligraphic carving through smoothness scalar field and brush-pressure distributionabstractABSTRACT Virtual calligraphic carving aims to generate 3D virtual carving works from 2D calligraphic images. This paper proposes a method of virtual calligraphic carving taking calligraphic carving work as a depth map composed of smoothness scalar field and brush‐pressure distribution, which ensures the smoothness of carving surface and interprets the strength of calligraphy, respectively. On one hand, the smoothness scalar field is modeled as a bounded planar scalar field with sources and constructed by applying a series of algorithms. On the other hand, the brush‐pressure distribution is represented by a map created by estimating the touching pressures employed by brush on paper. Then the depth map is composed of the smoothness scalar field and brush‐pressure distribution map and rendered into a virtual calligraphic carving work. The experiments show that this method can generate virtual calligraphic carving works that are vivid and very similar to that of artists. Copyright © 2012 John Wiley & Sons, Ltd. Xiaowu Chen 0001, Qinping Zhao |
Comput. Animat. Virtual Worlds | 1 |
| 2012 | Video event representation and inference on And-Or graphabstractABSTRACT This paper presents an approach for video event inference from dozens of actions performed by multiple players. First, we constructed an And‐Or graph to describe the different configurations of the event category such as shooting in soccer matches. We considered both temporal relations and role relations for the graph and encode them as vector parameters for each pair of graph nodes. Then, we developed an inference algorithm by using bottom‐up and top‐down processes. We found the proposals for each node during the bottom‐up step by considering three terms of energies and refined the proposals during the top‐down step by measuring the action‐labeling similarity and the temporal misplacement penalty. The optimal proposal of the inferring event and its score are obtained as the result. In the experiments, we tested the inference performance of the approach for the shooting events on real soccer match videos. By our approach, we can infer different kinds of shooting events in one scenario and interpret them play‐by‐play in a flexible way. Copyright © 2012 John Wiley & Sons, Ltd. Xiaowu Chen 0001, Yu Zhang 0035, Qinping Zhao |
Comput. Animat. Virtual Worlds | 2 |
| 2012 | Representing and recognizing objects with massive local image patches
Liang Lin 0004, Ping Luo 0002, Xiaowu Chen 0001 |
Pattern Recognit. | 3 |
| 2012 | Integrating Graph Partitioning and Matching for Trajectory Analysis in Video SurveillanceabstractIn order to track moving objects in long range against occlusion, interruption, and background clutter, this paper proposes a unified approach for global trajectory analysis. Instead of the traditional frame-by-frame tracking, our method recovers target trajectories based on a short sequence of video frames, e.g., 15 frames. We initially calculate a foreground map at each frame obtained from a state-of-the-art background model. An attribute graph is then extracted from the foreground map, where the graph vertices are image primitives represented by the composite features. With this graph representation, we pose trajectory analysis as a joint task of spatial graph partitioning and temporal graph matching. The task can be formulated by maximizing a posteriori under the Bayesian framework, in which we integrate the spatio-temporal contexts and the appearance models. The probabilistic inference is achieved by a data-driven Markov chain Monte Carlo algorithm. Given a period of observed frames, the algorithm simulates an ergodic and aperiodic Markov chain, and it visits a sequence of solution states in the joint space of spatial graph partitioning and temporal graph matching. In the experiments, our method is tested on several challenging videos from the public datasets of visual surveillance, and it outperforms the state-of-the-art methods. Liang Lin 0004, Yongyi Lu, Yan Pan 0002, Xiaowu Chen 0001 |
IEEE Trans. Image Process. | 4 |
| 2012 | Manifold preserving edit propagationabstractWe propose a novel edit propagation algorithm for interactive image and video manipulations. Our approach uses the locally linear embedding (LLE) to represent each pixel as a linear combination of its neighbors in a feature space. While previous methods require similar pixels to have similar results, we seek to maintain the manifold structure formed by all pixels in the feature space. Specifically, we require each pixel to be the same linear combination of its neighbors in the result. Compared with previous methods, our proposed algorithm is more robust to color blending in the input data. Furthermore, since every pixel is only related to a few nearest neighbors, our algorithm easily achieves good runtime efficiency. We demonstrate our manifold preserving edit propagation on various applications. Xiaowu Chen 0001, Dongqing Zou, Qinping Zhao |
ACM Trans. Graph. | 1 |
| 2011 | Single Image Based Illumination Estimation for Lighting Virtual Object in Real SceneabstractRendering virtual objects into real scenes with real illumination can greatly increase the realism of virtual objects and the consistency between the virtual and the real. The main challenge lies in illumination estimation from a single image. This article proposes a novel method of single image based illumination estimation for lighting virtual object in real scene. Only a single image, without any knowledge of the 3D geometry or reflectance, is needed, which greatly increases the applicability of the method. We first estimate coarse scene geometry and intrinsic components including shading image and reflectance image. Then the sparse radiance map of the scene is inferred based on the scene geometry and intrinsic components. Finally, the virtual objects are illuminated by the estimated sparse radiance map. Some experimental results show that this method can convincingly light virtual objects into a single real image, without any pre-recorded 3D geometry and reflectance, illumination acquisition equipments or imaging information of the image. Xiaowu Chen 0001, Xin Jin 0015 |
CAD/Graphics | 1 |
| 2011 | Automatic Compositing Soccer Video Highlights with Core-Around Event ModelabstractThis paper presents an automatic video highlighting approach in relation to a soccer match video lasting over ninety minutes, in which only the match video and the target length of the highlight video are required for the composition. Firstly, we propose a core-around event model, which represents the highlight as a complex event and consists of three components: semantic relations, temporal relationship among the activities and local motion appearance of each activity. Secondly, we detect activities involved in a highlight in the soccer match video by the local motion appearance. Thirdly, a candidate highlight video segment is aligned with the model by a specific kernel activity extracted on the basis of the semantic relations, and then assigned a matching score. Finally, we splice together the highlight segments of high matching score. Experimental results demonstrate that the composed highlight videos are attractive episodes of soccer matches, and similar to those edited by professional editors. Xiaowu Chen 0001, Qinping Zhao |
CAD/Graphics | 2 |
| 2011 | Face illumination transfer through edge-preserving filtersabstractThis article proposes a novel image-based method to transfer illumination from a reference face image to a target face image through edge-preserving filters. According to our method, only a single reference image, without any knowledge of the 3D geometry or material information of the target face, is needed. We first decompose the lightness layers of the reference and the target images into large-scale and detail layers through weighted least square (WLS) filter after face alignment. The large-scale layer of the reference image is filtered with the guidance of the target image. Adaptive parameter selection schemes for the edge-preserving filters is proposed in the above two filtering steps. The final relit result is obtained by replacing the large-scale layer of the target image with that of the reference image. We acquire convincing relit result on numerous target and reference face images with different lighting effects and genders. Comparisons with previous work show that our method is less affected by geometry differences and can preserve better the identification structure and skin color of the target face. Xiaowu Chen 0001, Xin Jin 0015, Qinping Zhao |
CVPR | 1 |
| 2011 | Partial similarity based nonparametric scene parsing in certain environmentabstractIn this paper we propose a novel nonparametric image parsing method for the image parsing problem in certain environment. A novel and efficient nearest neighbor matching scheme, the ANN bilateral matching scheme, is proposed. Based on the proposed matching scheme, we first retrieve some partially similar images for each given test image from the training image database. The test image can be well explained by these retrieved images, with similar regions existing in the retrieved images for each region in the test image. Then, we match the test image to the retrieved training images with the ANN bilateral matching scheme, and parse the test image by integrating multiple cues in a markov random field. Experiment on three datasets shows our method achieved promising parsing accuracy and outperformed two state-of-the-art nonparametric image parsing methods. Honghui Zhang, Tian Fang, Xiaowu Chen 0001, Qinping Zhao, Long Quan |
CVPR | 3 |
| 2010 | Learning Artistic Lighting Template from Portrait Photographs
Xin Jin 0015, Mingtian Zhao, Xiaowu Chen 0001, Qinping Zhao, Song-Chun Zhu |
ECCV (4) | 3 |
| 2010 | Automatic Image Inpainting by Heuristic Texture and Structure Completion
Xiaowu Chen 0001 |
MMM | 1 |
| 2010 | Automatic Image Completion with Structure Propagation and Texture SynthesisabstractIn this paper, we present a novel automatic image completion solution in a greedy manner inspired by a primal sketch representation model. Firstly, an image is divided into structure (sketchable) components and texture (non-sketchable) components, and the missing structures, such as curves and corners, are predicted by tensor voting. Secondly, the textures along structural sketches are synthesized with the sampled patches of some known structure components. Then, using the texture completion priorities decided by the confidence term, data term and distance term, the similar image patches of some known texture components are found by selecting a point with the maximum priority on the boundary of hole region. Finally, these image patches inpaint the missing textures of hole region seamlessly through graph cuts. The characteristics of this solution include: (1) introducing the primal sketch representation model to guide completion for visual consistency; (2) achieving fully automatic completion. The experiments on natural images illustrate satisfying image completion results. Xiaowu Chen 0001, Qinping Zhao |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2009 | Accurate semantic image labeling by fast Geodesic PropagationabstractMotivated by recently raised image semantic labeling problem, this paper studies a fast Geodesic Propagation (GP) algorithm that integrates recognition proposal and image compatibility into a graphical representation. Given the recognition proposal map of the image, the initial seeds are selected as confident pixels standing on local proposal peaks by Mean-shift algorithm. The geodesic distance is then defined on a hybrid manifold, combining the color and boundary features with the recognition proposal map. Based on the geodesic distance, the semantic labeling is simultaneously propagated from the initial seeds of all classes to the rest of image pixels. This inference algorithm is capable of multi-labeling an image of 2-mega pixels in one second (with a common PC). In the experiment, we test on 21 generic semantic categories (sky, road, grass ...) on MSRC dataset, and 17 categories on LHI dataset to evaluate the performance. Xiaowu Chen 0001, Dongyue Zhao, Yibiao Zhao, Liang Lin 0004 |
ICIP | 1 |
| 2008 | A Dynamic Awareness Model for Service-Based Collaborative Grid Application in Access Grid
Xiaowu Chen 0001, Qinping Zhao |
GPC | 1 |
| 2006 | Augmented Video Services and Its Applications in an Advanced Access Grid Environment
Xiaowu Chen 0001, Chunmin Xu |
UIC | 2 |
| 2000 | User's Vision Based Multi-resolution Rendering of 3D Models in Distributed Virtual Environment DVENET
Xiaowu Chen 0001, Qinping Zhao |
ICMI | 1 |