Yu Zhang 0035

dblp:50/671-35 · DBLP profile ↗
← Back
31ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-9653-3906ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 17 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Learning 3D Volume Cloud from Single Image
abstract
Three-dimensional (3D) cloud modeling plays a pivotal role in advancing atmospheric models and enhancing natural phenomena visualization systems. Nevertheless, the high-quality reconstruction of clouds remains a significant challenge, primarily due to their inherent heterogeneous nature as volumetric media. Image-based modeling approaches offer a promising solution to this challenge. This paper presents a novel two-stage neural network architecture for 3D cloud reconstruction from single image. The first stage introduces an innovative view synthesis network built upon Stable Diffusion, incorporating two specialized modules: a cloud mapper and a viewpoint mapper, which collaboratively generate novel perspective views from a single input image. The second stage implements a physics-based differentiable rendering framework to construct a 3D cloud reconstruction network, leveraging the synthesized multi-view images to optimize a volumetric density grid representation. To enhance the reconstruction fidelity, we integrate real-world cloud density distribution statistics and implement a post-processing refinement using Perlin-Worley noise combined with Fractal Brownian Motion (FBM) for erosion effects. Additionally, to mitigate the inherent limitations of geometric information extraction from single-view images, we developed a comprehensive cloud simulation dataset for pre-training the viewpoint mapper module. This dataset encompasses multi-view cloud images with corresponding camera extrinsic parameters, capturing a diverse range of cloud formations. Extensive quantitative evaluations and qualitative assessments demonstrate the efficacy and potential of our proposed two-stage network in achieving accurate 3D cloud reconstruction from single-view images.
Yuhang Cheng, Yu Zhang 0035, Xiaogang Wang 0005
ICMR2
2024 Color4E: Event Demosaicing for Full-color Event Guided Image Deblurring
abstract
Neuromorphic event sensors are novel visual cameras that feature high-speed illumination-variation sensing and have found widespread application in guiding frame-based imaging enhancement. This paper focuses on color restoration in the event-guided image deblurring task, we fuse blurry images with mosaic color events instead of mono events to avoid artifacts such as color bleeding. The challenges associated with this approach include demosaicing color events for reconstructing full-resolution sampled signals and fusing bimodal signals to achieve image deblurring. To meet these challenges, we propose a novel network called Color4E to enhance the color restoration quality for the image deblurring task. Color4E leverages an event demosaicing module to upsample the spatial resolution of mosaic color events and a cross-encoding image deblurring module for fusing bimodal signals, a refinement module is designed to fuse full-color events and refine initial deblurred images. Furthermore, to avoid the real-simulated gap of events, we implement a display-filter-camera system that enables mosaic and full-color event data captured synchronously, to collect a real-captured dataset used for network training and validation. The results on the public dataset and our collected dataset show that Color4E enables high-quality event-based image deblurring compared to state-of-the-art methods.
Yi Ma 0001, Peiqi Duan 0002, Yuchen Hong, Chu Zhou, Yu Zhang 0035, Jimmy S. J. Ren, Boxin Shi
ACM Multimedia5
2024 Reliable Event Generation With Invertible Conditional Normalizing Flow
abstract
Event streams provide a novel paradigm to describe visual scenes by capturing intensity variations above specific thresholds along with various types of noise. Existing event generation methods usually rely on one-way mappings using hand-crafted parameters and noise rates, which may not adequately suit diverse scenarios and event cameras. To address this limitation, we propose a novel approach to learn a bidirectional mapping between the feature space of event streams and their inherent parameters, enabling the generation of reliable event streams with enhanced generalization capabilities. We first randomly generate a vast number of parameters and synthesize massive event streams using an event simulator. Subsequently, an event-based normalizing flow network is proposed to learn the invertible mapping between the representation of a synthetic event stream and its parameters. The invertible mapping is implemented by incorporating an intensity-guided conditional affine simulation mechanism, facilitating better alignment between event features and parameter spaces. Additionally, we impose constraints on event sparsity, edge distribution, and noise distribution through novel event losses, further emphasizing event priors in the bidirectional mapping. Our framework surpasses state-of-the-art methods in video reconstruction, optical flow estimation, and parameter estimation tasks on synthetic and real-world datasets, exhibiting excellent generalization across diverse scenes and cameras.
Daxin Gu, Jia Li 0003, Lin Zhu 0012, Yu Zhang 0035, Jimmy S. J. Ren
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Range-nullspace Video Frame Interpolation with Focalized Motion Estimation
abstract
Continuous-time video frame interpolation is a fundamental technique in computer vision for its flexibility in synthesizing motion trajectories and novel video frames at arbitrary intermediate time steps. Yet, how to infer accurate intermediate motion and synthesize high-quality video frames are two critical challenges. In this paper, we present a novel VFI framework with improved treatment for these challenges. To address the former, we propose focalized trajectory fitting, which performs confidence-aware motion trajectory estimation by learning to pay focus to reliable optical flow candidates while suppressing the outliers. The second is range-nullspace synthesis, a novel frame renderer cast as solving an ill-posed problem addressed by learning decoupled components in orthogonal subspaces. The proposed framework sets new records on 7 of 10 public VFI benchmarks.
Zhiyang Yu, Yu Zhang 0035, Dongqing Zou, Xijun Chen, Jimmy S. J. Ren
CVPR2
2023 E2NeRF: Event Enhanced Neural Radiance Fields from Blurry Images
abstract
Neural Radiance Fields (NeRF) achieves impressive rendering performance by learning volumetric 3D representation from several images of different views. However, it is difficult to reconstruct a sharp NeRF from blurry input as often occurred in the wild. To solve this problem, we propose a novel Event-Enhanced NeRF (E2NeRF) by utilizing the combination data of a bio-inspired event camera and a standard RGB camera. To effectively introduce event stream into the learning process of neural volumetric representation, we propose a blur rendering loss and an event rendering loss, which guide the network via modelling real blur process and event generation process, respectively. Moreover, a camera pose estimation framework for real-world data is built with the guidance of event stream to generalize the method to practical applications. In contrast to previous image-based or event-based NeRF, our framework effectively utilizes the internal relationship between events and images. As a result, E2NeRF not only achieves image deblurring but also achieves high-quality novel view image generation. Extensive experiments on both synthetic data and real-world data demonstrate that E2NeRF can effectively learn a sharp NeRF from blurry images, especially in complex and low-light scenes. Our code and datasets are publicly available at https://github.com/iCVTEAM/E2NeRF.
Yunshan Qi, Lin Zhu 0012, Yu Zhang 0035, Jia Li 0003
ICCV3
2023 From Pose to Part: Weakly-Supervised Pose Evolution for Human Part Segmentation
abstract
Human part segmentation is a crucial but challenging task in computer vision. Recent works have achieved progress with the help of pixel-wise annotations. However, annotating pixel-wise masks especially at part-level is a tedious and labor-intensive procedure. To overcome this problem, we propose a part evolution framework to learn reliable predictions from weak pose annotations, which are much easier to collect. Our framework is composed of two essential modules: the first part adaptation module is designed to learn the deep prior knowledge from three related tasks, i.e., pose estimation, part-level and object-level segmentation; the second module is the part evolution module, which refines the part priors from deep predictions with the boundary-aware optimization algorithm. These two modules are conducted iteratively to evolve pose keypoint annotations into reliable part priors. Experimental evidence shows that our weakly-supervised approach generates comparable results with the state-of-the-art strongly-supervised methods on public benchmarks, and also validates the potential of notable improvements when combining weak labels with existing part segmentation masks.
Yifan Zhao 0002, Jia Li 0003, Yu Zhang 0035, Yonghong Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Deep Bayesian Video Frame Interpolation
Zhiyang Yu, Yu Zhang 0035, Xujie Xiang, Dongqing Zou, Xijun Chen, Jimmy S. J. Ren
ECCV (15)2
2021 Informative and Consistent Correspondence Mining for Cross-Domain Weakly Supervised Object Detection
abstract
Cross-domain weakly supervised object detection aims to adapt object-level knowledge from a fully labeled source domain dataset (i.e., with object bounding boxes) to train object detectors for target domains that are weakly labeled (i.e., with image-level tags). Instead of domain-level distribution matching, as popularly adopted in the literature, we propose to learn pixel-wise cross-domain correspondences for more precise knowledge transfer. It is realized through a novel cross-domain co-attention scheme trained as region competition. In this scheme, the cross-domain correspondence module seeks for informative features on the target domain image, which if warped to the source domain image, could best explain its annotations. Meanwhile, a collaborative mask generator competes to mask out the relevant target image region to make the remaining features uninformative. Such competitive learning strives to correlate the full foreground in cross-domain image pairs, revealing the accurate object extent in target domain. To alleviate the ambiguity of inter-domain correspondence learning, a domain-cycle consistency regularizer is further proposed to leverage the more reliable intra-domain correspondence. The proposed approach achieves consistent improvements over existing approaches by a considerable margin, demonstrated by the experiments on various datasets.
Luwei Hou, Yu Zhang 0035, Kui Fu, Jia Li 0003
CVPR2
2021 Training Weakly Supervised Video Frame Interpolation with Events
abstract
Event-based video frame interpolation is promising as event cameras capture dense motion signals that can greatly facilitate motion-aware synthesis. However, training existing frameworks for this task requires high frame-rate videos with synchronized events, posing challenges to collect real training data. In this work we show event-based frame interpolation can be trained without the need of high frame-rate videos. This is achieved via a novel weakly supervised framework that 1) corrects image appearance by extracting complementary information from events and 2) supplants motion dynamics modeling with attention mechanisms. For the latter we propose subpixel attention learning, which supports searching high-resolution correspondence efficiently on low-resolution feature grid. Though trained on low frame-rate videos, our framework outperforms existing models trained with full high frame-rate videos (and events) on both GoPro dataset and a new real event-based dataset. Codes, models and dataset will be made available at: https://github.com/YU-Zhiyang/WEVI.
Zhiyang Yu, Yu Zhang 0035, Deyuan Liu, Dongqing Zou, Xijun Chen, Yebin Liu, Jimmy S. J. Ren
ICCV2
2021 How to Learn a Domain-Adaptive Event Simulator?
abstract
The low-latency streams captured by event cameras have shown impressive potential in addressing vision tasks such as video reconstruction and optical flow estimation. However, these tasks often require massive training event streams, which are expensive to collect and largely bypassed by recently proposed event camera simulators. To align the statistics of synthetic events with that of target event cameras, existing simulators often need to be heuristically tuned with elaborative manual efforts and thus become incompetent to automatically adapt to various domains. To address this issue, this work proposes one of the first learning-based, domain-adaptive event simulator. Given a specific domain, the proposed simulator learns pixel-wise distributions of event contrast thresholds that, after stochastic sampling and paralleled rendering, can generate event representations well aligned with those from the data from realistic event cameras. To achieve such domain-specific alignment, we design a novel divide-and-conquer discrimination scheme that adaptively evaluates the synthetic-to-real consistency of event representations according to the local statistics of images and events. Trained with the data synthesized by the proposed simulator, the performances of state-of-the-art event-based video reconstruction and optical flow estimation approaches are boosted up to 22.9% and 2.8%, respectively. In addition, we show significantly improved domain adaptation capability over existing event simulators and tuning strategies, consistently on three real event datasets.
Daxin Gu, Jia Li 0003, Yu Zhang 0035, Yonghong Tian 0001
ACM Multimedia3
2021 Ordinal Multi-Task Part Segmentation With Recurrent Prior Generation
abstract
Semantic object part segmentation is a fundamental task in object understanding and geometric analysis. The clear understanding of part relationships can be of great use to the segmentation process. In this work, we propose a novel Ordinal Multi-task Part Segmentation (OMPS) approach which explicitly models the part ordinal relationship to guide the segmentation process in a recurrent manner. Quantitative and qualitative experiments are conducted first to explore the mutual impacts among object parts and then an ordinal part inference algorithm is formulated via experimental observations. Specifically, our framework is mainly composed of two modules, the forward module to segment multiple parts as individual subtasks with prior knowledge, and the recurrent module to generate appropriate part priors with the ordinal inference algorithm. These two modules work iteratively to optimize the segmentation performance and the network parameters. Experimental results show that our approach outperforms the state-of-the-art models on human and vehicle part parsing benchmarks. Comprehensive evaluations are conducted to demonstrate the effectiveness of our approach in object part segmentation.
Yifan Zhao 0002, Jia Li 0003, Yu Zhang 0035, Yafei Song 0002, Yonghong Tian 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 Learning Event-Based Motion Deblurring
abstract
Recovering sharp video sequence from a motion-blurred image is highly ill-posed due to the significant loss of motion information in the blurring process. For event-based cameras, however, fast motion can be captured as events at high frame rate, raising new opportunities to exploring effective solutions. In this paper, we start from a sequential formulation of event-based motion deblurring, then show how its optimization can be unfolded with a novel end-toend deep architecture. The proposed architecture is a convolutional recurrent neural network that integrates visual and temporal knowledge of both global and local scales in principled manner. To further improve the reconstruction, we propose a differentiable directional event filtering module to effectively extract rich boundary prior from the evolution of events. We conduct extensive experiments on the synthetic GoPro dataset and a large newly introduced dataset captured by a DAVIS240C camera. The proposed approach achieves state-of-the-art reconstruction quality, and generalizes better to handling real-world motion blur.
Yu Zhang 0035, Dongqing Zou, Jimmy S. J. Ren, Jiancheng Lv 0001, Yebin Liu
CVPR2
2020 Learning to See in the Dark with Events
Yu Zhang 0035, Dongqing Zou, Jimmy S. J. Ren
ECCV (18)2
2020 Model-Guided Multi-Path Knowledge Aggregation for Aerial Saliency Prediction
abstract
As an emerging vision platform, a drone can look from many abnormal viewpoints which brings many new challenges into the classic vision task of video saliency prediction. To investigate these challenges, this paper proposes a large-scale video dataset for aerial saliency prediction, which consists of ground-truth salient object regions of 1,000 aerial videos, annotated by 24 subjects. To the best of our knowledge, it is the first large-scale video dataset that focuses on visual saliency prediction on drones. Based on this dataset, we propose a Model-guided Multi-path Network (MM-Net) that serves as a baseline model for aerial video saliency prediction. Inspired by the annotation process in eye-tracking experiments, MM-Net adopts multiple information paths, each of which is initialized under the guidance of a classic saliency model. After that, the visual saliency knowledge encoded in the most representative paths is selected and aggregated to improve the capability of MM-Net in predicting spatial saliency in aerial scenarios. Finally, these spatial predictions are adaptively combined with the temporal saliency predictions via a spatiotemporal optimization algorithm. Experimental results show that MM-Net outperforms ten state-of-the-art models in predicting aerial video saliency.
Kui Fu, Jia Li 0003, Yu Zhang 0035, Hongze Shen, Yonghong Tian 0001
IEEE Trans. Image Process.3
2020 Efficient Low-Resolution Face Recognition via Bridge Distillation
abstract
Face recognition in the wild is now advancing towards light-weight models, fast inference speed and resolution-adapted capability. In this paper, we propose a bridge distillation approach to turn a complex face model pretrained on private high-resolution faces into a light-weight one for low-resolution face recognition. In our approach, such a cross-dataset resolution-adapted knowledge transfer problem is solved via two-step distillation. In the first step, we conduct cross-dataset distillation to transfer the prior knowledge from private high-resolution faces to public high-resolution faces and generate compact and discriminative features. In the second step, the resolution-adapted distillation is conducted to further transfer the prior knowledge to synthetic low-resolution faces via multi-task learning. By learning low-resolution face representations and mimicking the adapted high-resolution knowledge, a light-weight student model can be constructed with high efficiency and promising accuracy in recognizing low-resolution faces. Experimental results show that the student model performs impressively in recognizing low-resolution faces with only 0.21M parameters and 0.057MB memory. Meanwhile, its speed reaches up to 14,705, 934 and 763 faces per second on GPU, CPU and mobile phone, respectively.
Shiming Ge, Shengwei Zhao, Chenyu Li 0001, Yu Zhang 0035, Jia Li 0003
IEEE Trans. Image Process.4
2019 Structure-Preserving Stereoscopic View Synthesis With Multi-Scale Adversarial Correlation Matching
abstract
This paper addresses stereoscopic view synthesis from a single image. Various recent works solve this task by reorganizing pixels from the input view to reconstruct the target one in a stereo setup. However, purely depending on such photometric-based reconstruction process, the network may produce structurally inconsistent results. Regarding this issue, this work proposes Multi-Scale Adversarial Correlation Matching (MS-ACM), a novel learning framework for structure-aware view synthesis. The proposed framework does not assume any costly supervision signal of scene structures such as depth. Instead, it models structures as self-correlation coefficients extracted from multi-scale feature maps in transformed spaces. In training, the feature space attempts to push the correlation distances between the synthesized and target images far apart, thus amplifying inconsistent structures. At the same time, the view synthesis network minimizes such correlation distances by fixing mistakes it makes. With such adversarial training, structural errors of different scales and levels are iteratively discovered and reduced, preserving both global layouts and fine-grained details. Extensive experiments on the KITTI benchmark show that MS-ACM improves both visual quality and the metrics over existing methods when plugged into recent view synthesis architectures.
Yu Zhang 0035, Dongqing Zou, Jimmy S. J. Ren, Xiaohao Chen
CVPR1
2019 Selectivity or Invariance: Boundary-Aware Salient Object Detection
abstract
Typically, a salient object detection (SOD) model faces opposite requirements in processing object interiors and boundaries. The features of interiors should be invariant to strong appearance change so as to pop-out the salient object as a whole, while the features of boundaries should be selective to slight appearance change to distinguish salient objects and background. To address this selectivity-invariance dilemma, we propose a novel boundary-aware network with successive dilation for image-based SOD. In this network, the feature selectivity at boundaries is enhanced by incorporating a boundary localization stream, while the feature invariance at interiors is guaranteed with a complex interior perception stream. Moreover, a transition compensation stream is adopted to amend the probable failures in transitional regions between interiors and boundaries. In particular, an integrated successive dilation module is proposed to enhance the feature invariance at interiors and transitional regions. Extensive experiments on six datasets show that the proposed approach outperforms 16 state-of-the-art methods.
Jinming Su, Jia Li 0003, Yu Zhang 0035, Changqun Xia, Yonghong Tian 0001
ICCV3
2019 Multi-Class Part Parsing With Joint Boundary-Semantic Awareness
abstract
Object part parsing in the wild, which requires to simultaneously detect multiple object classes in the scene and accurately segments semantic parts within each class, is challenging for the joint presence of class-level and part-level ambiguities. Despite its importance, however, this problem is not sufficiently explored in existing works. In this paper, we propose a joint parsing framework with boundary and semantic awareness to address this challenging problem. To handle part-level ambiguity, a boundary awareness module is proposed to make mid-level features at multiple scales attend to part boundaries for accurate part localization, which are then fused with high-level features for effective part recognition. For class-level ambiguity, we further present a semantic awareness module that selects discriminative part features relevant to a category to prevent irrelevant features being merged together. The proposed modules are lightweight and implementation friendly, improving the performance substantially when plugged into various baseline architectures. Without bells and whistles, the full model sets new state-of-the-art results on the Pascal-Part dataset, in both multi-class and the conventional single-class setting, while running substantially faster than recent high-performance approaches.
Yifan Zhao 0002, Jia Li 0003, Yu Zhang 0035, Yonghong Tian 0001
ICCV3
2019 Cross-Reference Stitching Quality Assessment for 360° Omnidirectional Images
abstract
Along with the development of virtual reality (VR), omnidirectional images play an important role in producing multimedia content with an immersive experience. However, despite various existing approaches for omnidirectional image stitching, how to quantitatively assess the quality of stitched images is still insufficiently explored. To address this problem, we first establish a novel omnidirectional image dataset containing stitched images as well as dual-fisheye images captured from standard quarters of 0$^\circ$, 90$^\circ$, 180$^\circ$, and 270$^\circ$. In this manner, when evaluating the quality of an image stitched from a pair of fisheye images (\eg, 0$^\circ$ and 180$^\circ$), the other pair of fisheye images (\eg, 90$^\circ$ and 270$^\circ$) can be used as the cross-reference to provide ground-truth observations of the stitching regions. Based on this dataset, we propose a set of Omnidirectional Stitching Image Quality Assessment (OS-IQA) metrics. In these metrics, the stitching regions are assessed by exploring the local relationships between the stitched image and its cross-reference with histogram statistics, perceptual hash and sparse reconstruction, while the whole stitched images are assessed by the global indicators of color difference and fitness of blind zones.Qualitative and quantitative experiments show our method outperforms the classic IQA metrics and is highly consistent with human subjective evaluations. To the best of our knowledge, it is the first attempt that assesses the stitching quality of omnidirectional images by using cross-references.
Jia Li 0003, Kaiwen Yu, Yifan Zhao 0002, Yu Zhang 0035, Long Xu 0001
ACM Multimedia4
2018 Reconstructing non-rigid object with large movement using a single depth camera
Feixiang Lu, Feng Lu 0005, Yu Zhang 0035, Xiaowu Chen 0001, Qinping Zhao
Comput. Aided Geom. Des.4
2018 Semantic Object Segmentation in Tagged Videos via Detection
abstract
Semantic object segmentation (SOS) is a challenging task in computer vision that aims to detect and segment all pixels of the objects within predefined semantic categories. In image-based SOS, many supervised models have been proposed and achieved impressive performances due to the rapid advances of well-annotated training images and machine learning theories. However, in video-based SOS it is often difficult to directly train a supervised model since most videos are weakly annotated by tags. To handle such tagged videos, this paper proposes a novel approach that adopts a segmentation-by-detection framework. In this framework, object detection and segment proposals are first generated using the models pre-trained on still images, which provide useful cues to roughly localize the semantic objects. Based on these proposals, we propose an efficient algorithm to initialize object tracks by solving a joint assignment problem. As such tracks provide rough spatiotemporal configurations of the semantic objects, a voting-based refinement algorithm is further proposed to improve their spatiotemporal consistency. Extensive experiments demonstrate that the proposed framework can robustly and effectively segment semantic objects in tagged videos, even when the image-based object detectors provide inaccurate proposals. On various public benchmarks, the proposed approach obtains substantial improvements over the state-of-the-arts.
Yu Zhang 0035, Xiaowu Chen 0001, Jia Li 0003, Changqun Xia
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Exploring Weakly Labeled Images for Video Object Segmentation With Submodular Proposal Selection
abstract
Video object segmentation (VOS) is important for various computer vision problems, and handling it with minimal human supervision is highly desired for the large-scale applications. To bring down the supervision, existing approaches largely follow a data mining perspective by assuming the availability of multiple videos sharing the same object categories. It, however, would be problematic for the tasks that consume a single video. To address this problem, this paper proposes a novel approach that explores weakly labeled images to solve video object segmentation. Given a video labeled with a target category, images labeled with the same category are collected, from which noisy object exemplars are automatically discovered. After that the proposed approach extracts a set of region proposals on various frames and efficiently matches them with massive noisy exemplars in terms of appearance and spatial context. We then jointly select the best proposals across the video by solving a novel submodular problem that combines region voting and global region matching. Finally, the localization results are leveraged as strong supervision to guide pixel-level segmentation. Extensive experiments are conducted on two challenging public databases: Youtube-Objects and DAVIS. The results suggest that the proposed approach improves over previous weakly supervised/unsupervised approaches significantly, showing a performance even comparable with the several approaches supervised by the costly manual segmentations.
Yu Zhang 0035, Xiaowu Chen 0001, Jia Li 0003, Wei Teng, Haokun Song
IEEE Trans. Image Process.1
2018 Real-time 3D scene reconstruction with dynamically moving object using a single depth camera
Feixiang Lu, Yu Zhang 0035, Qinping Zhao
Vis. Comput.3
2017 What is and What is Not a Salient Object? Learning Salient Object Detector by Ensembling Linear Exemplar Regressors
abstract
Finding what is and what is not a salient object can be helpful in developing better features and models in salient object detection (SOD). In this paper, we investigate the images that are selected and discarded in constructing a new SOD dataset and find that many similar candidates, complex shape and low objectness are three main attributes of many non-salient objects. Moreover, objects may have diversified attributes that make them salient. As a result, we propose a novel salient object detector by ensembling linear exemplar regressors. We first select reliable foreground and background seeds using the boundary prior and then adopt locally linear embedding (LLE) to conduct manifold-preserving foregroundness propagation. In this manner, a foregroundness map can be generated to roughly pop-out salient objects and suppress non-salient ones with many similar candidates. Moreover, we extract the shape, foregroundness and attention descriptors to characterize the extracted object proposals, and a linear exemplar regressor is trained to encode how to detect salient proposals in a specific image. Finally, various linear exemplar regressors are ensembled to form a single detector that adapts to various scenarios. Extensive experimental results on 5 dataset and the new SOD dataset show that our approach outperforms 9 state-of-art methods.
Changqun Xia, Jia Li 0003, Xiaowu Chen 0001, Anlin Zheng, Yu Zhang 0035
CVPR5
2016 Local Shape Transfer for Image Co-segmentation
Wei Teng, Yu Zhang 0035, Xiaowu Chen 0001, Jia Li 0003, Zhiqiang He 0002
BMVC2
2016 High-level representation sketch for video event retrieval
Yu Zhang 0035, Xiaowu Chen 0001, Liang Lin 0004, Changqun Xia, Dongqing Zou
Sci. China Inf. Sci.1
2016 6-DOF Image Localization From Massive Geo-Tagged Reference Images
abstract
The 6-degrees of freedom (DOF) image localization, which aims to calculate the spatial position and rotation of a camera, is a challenging problem for most location-based services. In existing approaches, this problem is often tackled by finding the matches between 2D image points and 3D structure points so as to derive the location information via direct linear transformation algorithm. However, as these 2D-to-3D-based approaches need to reconstruct the 3D structure points of the scene, they may not be flexible enough to employ massive and increasing geo-tagged data. To this end, this paper presents a novel approach for 6-DOF image localization by fusing candidate poses relative to reference images. In this approach, we propose to localize an input image according to the position and rotation information of multiple geo-tagged images retrieved from a reference dataset. From the reference images, an efficient relative pose estimation algorithm is proposed to derive a set of candidate poses for the input image. Each candidate pose encodes the relative rotation and direction of the input image with respect to a specific reference image. Finally, these candidate poses can be fused together by minimizing a well-defined geometry error so that the 6-DOF location of the input image is effectively derived. Experimental results show that our method can obtain satisfactory localization accuracy. In addition, the proposed relative pose estimation algorithm is much faster than existing work.
Yafei Song 0002, Xiaowu Chen 0001, Xiaogang Wang 0005, Yu Zhang 0035, Jia Li 0003
IEEE Trans. Multim.4
2015 Semantic object segmentation via detection in weakly labeled video
abstract
Semantic object segmentation in video is an important step for large-scale multimedia analysis. In many cases, however, semantic objects are only tagged at video-level, making them difficult to be located and segmented. To address this problem, this paper proposes an approach to segment semantic objects in weakly labeled video via object detection. In our approach, a novel video segmentation-by-detection framework is proposed, which first incorporates object and region detectors pre-trained on still images to generate a set of detection and segmentation proposals. Based on the noisy proposals, several object tracks are then initialized by solving a joint binary optimization problem with min-cost flow. As such tracks actually provide rough configurations of semantic objects, we thus refine the object segmentation while preserving the spatiotemporal consistency by inferring the shape likelihoods of pixels from the statistical information of tracks. Experimental results on Youtube-Objects dataset and SegTrack v2 dataset demonstrate that our method outperforms state-of-the-arts and shows impressive results.
Yu Zhang 0035, Xiaowu Chen 0001, Jia Li 0003, Changqun Xia
CVPR1
2015 Cuboids detection in RGB-D images via Maximum Weighted Clique
abstract
Cuboid detection is an essential step for understanding 3D structure of scenes. As most of indoor scene cuboids are actually objects, we propose in this paper an object-based approach to detect 3D cuboids in indoor RGB-D images. The proposed approach is learning-free and can handle general object classes rather than a limited pre-defined category set. In our approach, we first apply an extended version of the CPMC framework to generate a set of segment hypotheses, and fit a set of cuboid candidates. Given the candidate set, we select several cuboids that can provide plausible interpretations of the images by solving a Maximum Weighted Clique (MWC) problem. With this formulation, a set of ranked mid-level representations of the input image is obtained, and are further re-ranked by Maximal Marginal Relevance (MMR) measure to improve their diversity. Experimental results on NYU-V2 dataset shows that our method significantly outperforms the state-of-the-art, and shows impressive results.
Xiaowu Chen 0001, Yu Zhang 0035, Jia Li 0003, Xiaogang Wang 0005
ICME3
2014 Geodesic Propagation for Semantic Labeling
abstract
This paper presents a semantic labeling framework with geodesic propagation (GP). Under the same framework, three algorithms are proposed, including GP, supervised GP (SGP) for image, and hybrid GP (HGP) for video. In these algorithms, we resort to the recognition proposal map and select confident pixels with maximum probability as the initial propagation seeds. From these seeds, the GP algorithm iteratively updates the weights of geodesic distances until the semantic labels are propagated to all pixels. On the contrary, the SGP algorithm further exploits the contextual information to guide the direction of propagation, leading to better performance but higher computational complexity than the GP. For video labeling, we further propose the HGP algorithm, in which the geodesic metric is used in both spatial and temporal spaces. Experiments on four public data sets show that our algorithms outperform several state-of-the-art methods. With the GP framework, convincing results for both image and video semantic labeling can be obtained.
Xiaowu Chen 0001, Yafei Song 0002, Yu Zhang 0035, Xin Jin 0015, Qinping Zhao
IEEE Trans. Image Process.4
2012 Video event representation and inference on And-Or graph
abstract
ABSTRACT This paper presents an approach for video event inference from dozens of actions performed by multiple players. First, we constructed an And‐Or graph to describe the different configurations of the event category such as shooting in soccer matches. We considered both temporal relations and role relations for the graph and encode them as vector parameters for each pair of graph nodes. Then, we developed an inference algorithm by using bottom‐up and top‐down processes. We found the proposals for each node during the bottom‐up step by considering three terms of energies and refined the proposals during the top‐down step by measuring the action‐labeling similarity and the temporal misplacement penalty. The optimal proposal of the inferring event and its score are obtained as the result. In the experiments, we tested the inference performance of the approach for the shooting events on real soccer match videos. By our approach, we can infer different kinds of shooting events in one scenario and interpret them play‐by‐play in a flexible way. Copyright © 2012 John Wiley & Sons, Ltd.
Xiaowu Chen 0001, Yu Zhang 0035, Qinping Zhao
Comput. Animat. Virtual Worlds3