Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Peng Wang 0001

dblp:95/4442-1 · DBLP profile ↗
← Back
37ranked-venue papers
11as first author
5since 2021 · last 2026
0000-0002-1265-0233ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 8 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 9 first-author · 3 since 2021Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
29 papers
3D vision · 55% Segmentation and scene understanding · 15% Autonomous driving · 7%
Computer graphics and multimedia
7 papers
Visual content generation and editing · 46% Rendering · 28% Image and video processing · 20%

Topics — the 30 heaviest of 78, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
depth estimation
3.492021
Continual Neural Mapping: Learning An Implicit Scene Representation from Sequential Observations · ICCV 2021
Every Pixel Counts ++: Joint Learning of Geometry and Motion with 3D Holistic Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Learning Depth with Convolutional Spatial Propagation Network · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Computer vision › Segmentation and scene understanding
semantic segmentation
1.552020
The ApolloScape Open Dataset for Autonomous Driving and Its Application · IEEE Trans. Pattern Anal. Mach. Intell. 2020
DeLS-3D: Deep Localization and Segmentation With a 3D Semantic Map · CVPR 2018
Pose-Guided Human Parsing by an AND/OR Graph Using Pose-Context Features · AAAI 2016
Computer vision › 3D vision › motion estimation
optical flow
1.132020
Every Pixel Counts ++: Joint Learning of Geometry and Motion with 3D Holistic Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 2020
UnOS: Unified Unsupervised Optical-Flow and Stereo-Depth Estimation by Watching Videos · CVPR 2019
Occlusion Aware Unsupervised Learning of Optical Flow · CVPR 2018
Computer vision › 3D vision › depth estimation
depth completion
0.922020
Learning Depth with Convolutional Spatial Propagation Network · IEEE Trans. Pattern Anal. Mach. Intell. 2020
CSPN++: Learning Context and Resource Aware Convolutional Spatial Propagation Networks for Depth Completion · AAAI 2020
Machine learning › Generative modeling › diffusion model
image editing
0.912025
HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing · ICLR 2025
Natural language and speech › Language models and text generation
instruction following
0.912025
HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing · ICLR 2025
Visual content generation and editing
image editing
0.912025
HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing · ICLR 2025
Visual content generation and editing › image editing › text-guided image editing
instruction-based image editing
0.912025
HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing · ICLR 2025
Computer vision › 3D vision › depth estimation
self-supervised depth estimation
0.832020
LEGO: Learning Edge With Geometry All at Once by Watching Videos · CVPR 2018
Unsupervised Learning of Geometry From Videos With Edge-Aware Depth-Normal Consistency · AAAI 2018
Omnidirectional Depth Extension Networks · ICRA 2020
Computer vision › Segmentation and scene understanding
instance segmentation
0.822020
3D Part Guided Image Editing for Fine-Grained Object Understanding · CVPR 2020
MaskLab: Instance Segmentation by Refining Object Detection With Semantic and Direction Features · CVPR 2018
Computer vision › 3D vision › motion estimation › optical flow
unsupervised optical flow
0.722019
UnOS: Unified Unsupervised Optical-Flow and Stereo-Depth Estimation by Watching Videos · CVPR 2019
Occlusion Aware Unsupervised Learning of Optical Flow · CVPR 2018
Computer vision › 3D vision
3d scene reconstruction
0.712023
VoGE: A Differentiable Volume Renderer using Gaussian Ellipsoids for Analysis-by-Synthesis · ICLR 2023
Computer vision › 3D vision
analysis-by-synthesis
0.712023
VoGE: A Differentiable Volume Renderer using Gaussian Ellipsoids for Analysis-by-Synthesis · ICLR 2023
Computer vision › 3D vision
surface normal estimation
0.722018
LEGO: Learning Edge With Geometry All at Once by Watching Videos · CVPR 2018
Unsupervised Learning of Geometry From Videos With Edge-Aware Depth-Normal Consistency · AAAI 2018
Rendering › volume rendering
differentiable volume rendering
0.712023
VoGE: A Differentiable Volume Renderer using Gaussian Ellipsoids for Analysis-by-Synthesis · ICLR 2023
Rendering
volume rendering
0.712023
VoGE: A Differentiable Volume Renderer using Gaussian Ellipsoids for Analysis-by-Synthesis · ICLR 2023
Computer vision › Segmentation and scene understanding
human parsing
0.522016
Zoom Better to See Clearer: Human and Object Parsing with Hierarchical Auto-Zoom Net · ECCV (5) 2016
Pose-Guided Human Parsing by an AND/OR Graph Using Pose-Context Features · AAAI 2016
Computer vision › 3D vision
implicit neural representation
0.512021
Continual Neural Mapping: Learning An Implicit Scene Representation from Sequential Observations · ICCV 2021
Computer vision › 3D vision › 3d scene modeling › scene representation
neural scene representation
0.512021
Continual Neural Mapping: Learning An Implicit Scene Representation from Sequential Observations · ICCV 2021
Computer vision › 3D vision › 3d shape representation › implicit surface representation
signed distance function
0.512021
Continual Neural Mapping: Learning An Implicit Scene Representation from Sequential Observations · ICCV 2021
Robotics › Autonomous driving
autonomous driving perception
0.412020
The ApolloScape Open Dataset for Autonomous Driving and Its Application · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Robotics › Autonomous driving › simulation
driving simulation
0.412020
AutoRemover: Automatic Object Removal for Autonomous Driving Videos · AAAI 2020
Robotics › Autonomous driving
perception
0.412020
3D Part Guided Image Editing for Fine-Grained Object Understanding · CVPR 2020
Computer vision › 3D vision
scene flow estimation
0.412020
Every Pixel Counts ++: Joint Learning of Geometry and Motion with 3D Holistic Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Robotics › Robot navigation and mapping
sensor fusion
0.412020
The ApolloScape Open Dataset for Autonomous Driving and Its Application · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Computer vision › 3D vision › stereo vision
stereo matching
0.412020
Learning Depth with Convolutional Spatial Propagation Network · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Visual content generation and editing › image completion
object removal
0.412020
AutoRemover: Automatic Object Removal for Autonomous Driving Videos · AAAI 2020
Image and video processing › video restoration
video inpainting
0.412020
AutoRemover: Automatic Object Removal for Autonomous Driving Videos · AAAI 2020
Computer vision › 3D vision › 3d shape analysis
3d shape understanding
0.412019
ApolloCar3D: A Large 3D Car Instance Understanding Benchmark for Autonomous Driving · CVPR 2019
Computer vision › 3D vision › depth estimation
stereo depth estimation
0.412019
UnOS: Unified Unsupervised Optical-Flow and Stereo-Depth Estimation by Watching Videos · CVPR 2019

Methods — techniques the papers use, named apart from their topics

convolutional neural network · 2.3GPT-4V · 1.7DALL-E 3 · 1.7gaussian ellipsoids · 1.3differentiable rendering · 1.3deep neural network · 1.3multi-task learning · 0.9deep convolutional network · 0.8convolutional spatial propagation network · 0.8conditional random field · 0.5shadow detection network · 0.4geometric frame alignment · 0.4occlusion boundary detection · 0.2deep learning · 0.2superpixel segmentation · 0.2salient object detection · 0.1random forest · 0.1perceptual color naming · 0.1
YearPublicationVenuePosition
2026 RapidMV: Leveraging Spatio-Angular Latent Space for Efficient and Consistent Text-to-Multi-View Synthesis
abstract
Generating consistent multi-view images given a text prompt is an essential bridge to generating synthetic 3D assets. In this work, we introduce RapidMV, a novel text-to-multi-view generative model that can produce 32 multi-view synthetic images in just around 5 seconds. In essence, we introduce a novel spatio-angular latent space, where we encode not only the spatial appearance of a single frame, but also the angular viewpoint deviations across multiple frames into a single latent for improved efficiency and multi-view consistency. We achieve effective training of RapidMV by strategically decomposing our training process into multiple steps. We demonstrate that RapidMV outperforms existing methods in terms of consistency and latency, with competitive quality and text-image alignment.
Seungwook Kim 0005, Yichun Shi, Kejie Li, Minsu Cho, Peng Wang 0001
WACV5
2026 LVM-Lite: Training Large Vision Models with Efficient Sequential Modeling
abstract
Large Vision Models (LVMs) have demonstrated impressive capabilities by leveraging large-scale generative pre-training on visual sequences. However, training these models on massive datasets of single images and image sequences can be computationally expensive and limit their accessibility to researchers without substantial computational resources. This paper introduces LVM-Lite, a new two-stage learning pipeline for more efficient and effective LVM training. In the first stage, the model is pre-trained on a large corpus of single images. Subsequently, in the second stage, the model is fine-tuned on curated long image/video sequences. This decoupled training approach substantially accelerates the training process, achieving up to 2.7× speedup compared to baseline LVM training. Extensive experiments demonstrate that LVM-Lite achieves competitive performance on various generative and discriminative benchmarks while maintaining high training efficiency and strong scalability. https://github.com/UCSC-VLAA/LVM-Lite
Xianhang Li, Hongru Zhu, Sucheng Ren, Peng Wang 0001, Xiaohui Shen, Qing Liu 0017, Cihang Xie
WACV5
2025 HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
abstract
This study introduces HQ-Edit, a high-quality instruction-based image editing dataset with around 200,000 edits. Unlike prior approaches relying on attribute guidance or human feedback on building datasets, we devise a scalable data collection pipeline leveraging advanced foundation models, namely GPT-4V and DALL-E 3. To ensure its high quality, diverse examples are first collected online, expanded, and then used to create high-quality diptychs featuring input and output images with detailed text prompts, followed by precise alignment ensured through post-processing. In addition, we propose two evaluation metrics, Alignment and Coherence, to quantitatively assess the quality of image edit pairs using GPT-4V. HQ-Edits high-resolution images, rich in detail and accompanied by comprehensive editing prompts, substantially enhance the capabilities of existing image editing models. For example, an HQ-Edit finetuned InstructPix2Pix can attain state-of-the-art image editing performance, even surpassing those models fine-tuned with human-annotated data.
Mude Hui, Siwei Yang, Bingchen Zhao, Yichun Shi, Peng Wang 0001, Cihang Xie, Yuyin Zhou
ICLR6
2023 VoGE: A Differentiable Volume Renderer using Gaussian Ellipsoids for Analysis-by-Synthesis
Angtian Wang, Peng Wang 0001, Adam Kortylewski, Alan L. Yuille
ICLR2
2021 Continual Neural Mapping: Learning An Implicit Scene Representation from Sequential Observations
abstract
Recent advances have enabled a single neural network to serve as an implicit scene representation, establishing the mapping function between spatial coordinates and scene properties. In this paper, we make a further step towards continual learning of the implicit scene representation directly from sequential observations, namely Continual Neural Mapping. The proposed problem setting bridges the gap between batch-trained implicit neural representations and commonly used streaming data in robotics and vision communities. We introduce an experience replay approach to tackle an exemplary task of continual neural mapping: approximating a continuous signed distance function (SDF) from sequential depth images as a scene geometry representation. We show for the first time that a single network can represent scene geometry over time continually without catastrophic forgetting, while achieving promising trade-offs between accuracy and efficiency.
Zike Yan, Xuesong Shi, Peng Wang 0001, Hongbin Zha
ICCV5
2020 CSPN++: Learning Context and Resource Aware Convolutional Spatial Propagation Networks for Depth Completion
abstract
Depth Completion deals with the problem of converting a sparse depth map to a dense one, given the corresponding color image. Convolutional spatial propagation network (CSPN) is one of the state-of-the-art (SoTA) methods of depth completion, which recovers structural details of the scene. In this paper, we propose CSPN++, which further improves its effectiveness and efficiency by learning adaptive convolutional kernel sizes and the number of iterations for the propagation, thus the context and computational resource needed at each pixel could be dynamically assigned upon requests. Specifically, we formulate the learning of the two hyper-parameters as an architecture selection problem where various configurations of kernel sizes and numbers of iterations are first defined, and then a set of soft weighting parameters are trained to either properly assemble or select from the pre-defined configurations at each pixel. In our experiments, we find weighted assembling can lead to significant accuracy improvements, which we referred to as "context-aware CSPN", while weighted selection, "resource-aware CSPN" can reduce the computational resource significantly with similar or better accuracy. Besides, the resource needed for CSPN++ can be adjusted w.r.t. the computational budget automatically. Finally, to avoid the side effects of noise or inaccurate sparse depths, we embed a gated network inside CSPN++, which further improves the performance. We demonstrate the effectiveness of CSPN++ on the KITTI depth completion benchmark, where it significantly improves over CSPN and other SoTA methods 1.
Xinjing Cheng, Peng Wang 0001, Chenye Guan, Ruigang Yang
AAAI2
2020 AutoRemover: Automatic Object Removal for Autonomous Driving Videos
abstract
Motivated by the need for photo-realistic simulation in autonomous driving, in this paper we present a video inpainting algorithm AutoRemover, designed specifically for generating street-view videos without any moving objects. In our setup we have two challenges: the first is the shadow, shadows are usually unlabeled but tightly coupled with the moving objects. The second is the large ego-motion in the videos. To deal with shadows, we build up an autonomous driving shadow dataset and design a deep neural network to detect shadows automatically. To deal with large ego-motion, we take advantage of the multi-source data, in particular the 3D data, in autonomous driving. More specifically, the geometric relationship between frames is incorporated into an inpainting deep neural network to produce high-quality structurally consistent video output. Experiments show that our method outperforms other state-of-the-art (SOTA) object removal algorithms, reducing the RMSE by over 19%.
Wei Li 0111, Peng Wang 0001, Chenye Guan, Yuhang Song 0003, Baoquan Chen, Weiwei Xu 0003, Ruigang Yang
AAAI3
2020 Speech2Video Synthesis with 3D Skeleton Regularization and Expressive Body Poses
Miao Liao, Peng Wang 0001, Hao Zhu 0004, Xinxin Zuo, Ruigang Yang
ACCV (5)3
2020 3D Part Guided Image Editing for Fine-Grained Object Understanding
abstract
Holistically understanding an object with its 3D movable parts is essential for visual models of a robot to interact with the world. For example, only by understanding many possible part dynamics of other vehicles (e.g., door or trunk opening, taillight blinking for changing lane), a self-driving vehicle can be success in dealing with emergency cases. However, existing visual models tackle rarely on these situations, but focus on bounding box detection. In this paper, we fill this important missing piece in autonomous driving by solving two critical issues. First, for dealing with data scarcity, we propose an effective training data generation process by fitting a 3D car model with dynamic parts to cars in real images. This allows us to directly edit the real images using the aligned 3D parts, yielding effective training data for learning robust deep neural networks (DNNs). Secondly, to benchmark the quality of 3D part understanding, we collected a large dataset in real driving scenario with cars in uncommon states (CUS), i.e. with door or trunk opened etc., which demonstrates that our trained network with edited images largely outperforms other baselines in terms of 2D detection and instance segmentation accuracy.
Zongdai Liu, Feixiang Lu, Peng Wang 0001, Liangjun Zhang, Ruigang Yang
CVPR3
2020 Omnidirectional Depth Extension Networks
abstract
Omnidirectional 360° camera proliferates rapidly for autonomous robots since it significantly enhances the perception ability by widening the field of view (FoV). However, corresponding 360° depth sensors, which are also critical for the perception system, are still difficult or expensive to have. In this paper, we propose a low-cost 3D sensing system that combines an omnidirectional camera with a calibrated projective depth camera, where the depth from the limited FoV can be automatically extended to the rest of recorded omnidirectional image. To accurately recover the missing depths, we design an omnidirectional depth extension convolutional neural network (ODE-CNN), in which a spherical feature transform layer (SFTL) is embedded at the end of feature encoding layers, and a deformable convolutional spatial propagation network (D-CSPN) is appended at the end of feature decoding layers. The former re-samples the neighborhood of each pixel in the omnidirectional coordination to the projective coordination, which reduce the difficulty of feature learning, and the later automatically finds a proper context to well align the structures in the estimated depths via CNN w.r.t. the reference image, which significantly improves the visual quality. Finally, we demonstrate the effectiveness of proposed ODE-CNN over the popular 360D dataset, and show that ODE-CNN significantly outperforms (relatively 33% reduction in depth error) other state-of-the-art (SoTA) methods.
Xinjing Cheng, Peng Wang 0001, Yanqi Zhou, Chenye Guan, Ruigang Yang
ICRA2
2020 Learning Depth with Convolutional Spatial Propagation Network
abstract
In this paper, we propose the convolutional spatial propagation network (CSPN) and demonstrate its effectiveness for various depth estimation tasks. CSPN is a simple and efficient linear propagation model, where the propagation is performed with a manner of recurrent convolutional operations, in which the affinity among neighboring pixels is learned through a deep convolutional neural network (CNN). Compare to the previous state-of-the-art (SOTA) linear propagation model, i.e., spatial propagation networks (SPN), CSPN is 2 to 5× faster in practice. We concatenate CSPN and its variants to SOTA depth estimation networks, which significantly improve the depth accuracy. Specifically, we apply CSPN to two depth estimation problems: depth completion and stereo matching, in which we design modules which adapts the original 2D CSPN to embed sparse depth samples during the propagation, operate with 3D convolution and be synergistic with spatial pyramid pooling. In our experiments, we show that all these modules contribute to the final performance. For the task of depth completion, our method reduce the depth error over 30 percent in the NYU v2 and KITTI datasets. For the task of stereo matching, our method currently ranks 1st on both the KITTI Stereo 2012 and 2015 benchmarks.
Xinjing Cheng, Peng Wang 0001, Ruigang Yang
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 The ApolloScape Open Dataset for Autonomous Driving and Its Application
abstract
Autonomous driving has attracted tremendous attention especially in the past few years. The key techniques for a self-driving car include solving tasks like 3D map construction, self-localization, parsing the driving road and understanding objects, which enable vehicles to reason and act. However, large scale data set for training and system evaluation is still a bottleneck for developing robust perception models. In this paper, we present the ApolloScape dataset [1] and its applications for autonomous driving. Compared with existing public datasets from real scenes, e.g., KITTI [2] or Cityscapes [3] , ApolloScape contains much large and richer labelling including holistic semantic dense point cloud for each site, stereo, per-pixel semantic labelling, lanemark labelling, instance segmentation, 3D car instance, high accurate location for every frame in various driving videos from multiple sites, cities and daytimes. For each task, it contains at lease 15x larger amount of images than SOTA datasets. To label such a complete dataset, we develop various tools and algorithms specified for each task to accelerate the labelling process, such as joint 3D-2D segment labeling, active labelling in videos etc. Depend on ApolloScape, we are able to develop algorithms jointly consider the learning and inference of multiple tasks. In this paper, we provide a sensor fusion scheme integrating camera videos, consumer-grade motion sensors (GPS/IMU), and a 3D semantic map in order to achieve robust self-localization and semantic segmentation for autonomous driving. We show that practically, sensor fusion and joint learning of multiple tasks are beneficial to achieve a more robust and accurate system. We expect our dataset and proposed relevant algorithms can support and motivate researchers for further development of multi-sensor fusion and multi-task learning in the field of computer vision.
Xinyu Huang 0001, Peng Wang 0001, Xinjing Cheng, Dingfu Zhou, Qichuan Geng, Ruigang Yang
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Every Pixel Counts ++: Joint Learning of Geometry and Motion with 3D Holistic Understanding
abstract
Learning to estimate 3D geometry in a single frame and optical flow from consecutive frames by watching unlabeled videos via deep convolutional network has made significant progress recently. Current state-of-the-art (SoTA) methods treat the two tasks independently. One typical assumption of the existing depth estimation methods is that the scenes contain no independent moving objects. while object moving could be easily modeled using optical flow. In this paper, we propose to address the two tasks as a whole, i.e., to jointly understand per-pixel 3D geometry and motion. This eliminates the need of static scene assumption and enforces the inherent geometrical consistency during the learning process, yielding significantly improved results for both tasks. We call our method as “Every Pixel Counts++” or “EPC++”. Specifically, during training, given two consecutive frames from a video, we adopt three parallel networks to predict the camera motion (MotionNet), dense depth map (DepthNet), and per-pixel optical flow between two frames (OptFlowNet) respectively. The three types of information, are fed into a holistic 3D motion parser (HMP), and per-pixel 3D motion of both rigid background and moving objects are disentangled and recovered. Various loss terms are formulated to jointly supervise the three networks. An effective adaptive training strategy is proposed to achieve better performance and more efficient convergence. Comprehensive experiments were conducted on datasets with different scenes, including driving scenario (KITTI 2012 and KITTI 2015 datasets), mixed outdoor/indoor scenes (Make3D) and synthetic animation (MPI Sintel dataset). Performance on the five tasks of depth estimation, optical flow estimation, odometry, moving object segmentation and scene flow estimation shows that our approach outperforms other SoTA methods, demonstrating the effectiveness of each module of our proposed method. Code will be available at: https://github.com/chenxuluo/EPC.
Chenxu Luo, Zhenheng Yang, Peng Wang 0001, Yang Wang 0046, Wei Xu 0017, Ramakant Nevatia, Alan L. Yuille
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 ApolloCar3D: A Large 3D Car Instance Understanding Benchmark for Autonomous Driving
abstract
Autonomous driving has attracted remarkable attention from both industry and academia. An important task is to estimate 3D properties (e.g. translation, rotation and shape) of a moving or parked vehicle on the road. This task, while critical, is still under-researched in the computer vision community – partially owing to the lack of large scale and fully-annotated 3D car database suitable for autonomous driving research. In this paper, we contribute the first large scale database suitable for 3D car instance understanding – ApolloCar3D. The dataset contains 5,277 driving images and over 60K car instances, where each car is fitted with an industry-grade 3D CAD model with absolute model size and semantically labelled keypoints. This dataset is above 20× larger than PASCAL3D+ and KITTI, the current state-of-the-art. To enable efficient labelling in 3D, we build a pipeline by considering 2D-3D keypoint correspondences for a single instance and 3D relationship among multiple instances. Equipped with such dataset, we build various baseline algorithms with the state-of-the-art deep convolutional neural networks. Specifically, we first segment each car with a pre-trained Mask R-CNN, and then regress towards its 3D pose and shape based on a deformable 3D car model with or without using semantic keypoints. We show that using keypoints significantly improves fitting performance. Finally, we develop a new 3D metric jointly considering 3D pose and 3D shape, allowing for comprehensive evaluation and ablation study.
Xibin Song, Peng Wang 0001, Dingfu Zhou, Chenye Guan, Yuchao Dai, Hongdong Li, Ruigang Yang
CVPR2
2019 UnOS: Unified Unsupervised Optical-Flow and Stereo-Depth Estimation by Watching Videos
abstract
In this paper, we propose UnOS, an unified system for unsupervised optical flow and stereo depth estimation using convolutional neural network (CNN) by taking advantages of their inherent geometrical consistency based on the rigid-scene assumption. UnOS significantly outperforms other state-of-the-art (SOTA) unsupervised approaches that treated the two tasks independently. Specifically, given two consecutive stereo image pairs from a video, UnOS estimates per-pixel stereo depth images, camera ego-motion and optical flow with three parallel CNNs. Based on these quantities, UnOS computes rigid optical flow and compares it against the optical flow estimated from the FlowNet, yielding pixels satisfying the rigid-scene assumption. Then, we encourage geometrical consistency between the two estimated flows within rigid regions, from which we derive a rigid-aware direct visual odometry (RDVO) module. We also propose rigid and occlusion-aware flow-consistency losses for the learning of UnOS. We evaluated our results on the popular KITTI dataset over 4 related tasks, \ie stereo depth, optical flow, visual odometry and motion segmentation.
Yang Wang 0046, Peng Wang 0001, Zhenheng Yang, Chenxu Luo, Yi Yang 0007, Wei Xu 0017
CVPR2
2018 Unsupervised Learning of Geometry From Videos With Edge-Aware Depth-Normal Consistency
abstract
Learning to reconstruct depths from a single image by watching unlabeled videos via deep convolutional network (DCN) is attracting significant attention in recent years, e.g. (Zhou et al. 2017). In this paper, we propose to use surface normal representation for unsupervised depth estimation framework. Our estimated depths are constrained to be compatible with predicted normals, yielding more robust geometry results. Specifically, we formulate an edge-aware depth-normal consistency term, and solve it by constructing a depth-to-normal layer and a normal-to-depth layer inside of the DCN. The depth-to-normal layer takes estimated depths as input, and computes normal directions using cross production based on neighboring pixels. Then given the estimated normals, the normal-to-depth layer outputs a regularized depth map through local planar smoothness. Both layers are computed with awareness of edges inside the image to help address the issue of depth/normal discontinuity and preserve sharp edges. Finally, to train the network, we apply the photometric error and gradient smoothness to supervise both depth and normal predictions. We conducted experiments on both outdoor (KITTI) and indoor (NYUv2) datasets, and showed that our algorithm vastly outperforms state-of-the-art, which demonstrates the benefits of our approach.
Zhenheng Yang, Peng Wang 0001, Wei Xu 0017, Liang Zhao 0006, Ramakant Nevatia
AAAI2
2018 SPG-Net: Segmentation Prediction and Guidance Network for Image Inpainting
Yuhang Song 0003, Chao Yang 0011, Yeji Shen, Peng Wang 0001, Qin Huang 0006, C.-C. Jay Kuo
BMVC4
2018 View Extrapolation of Human Body From a Single Image
abstract
We study how to synthesize novel views of human body from a single image. Though recent deep learning based methods work well for rigid objects, they often fail on objects with large articulation, like human bodies. The core step of existing methods is to fit a map from the observable views to novel views by CNNs; however, the rich articulation modes of human body make it rather challenging for CNNs to memorize and interpolate the data well. To address the problem, we propose a novel deep learning based pipeline that explicitly estimates and leverages the geometry of the underlying human body. Our new pipeline is a composition of a shape estimation network and an image generation network, and at the interface a perspective transformation is applied to generate a forward flow for pixel value transportation. Our design is able to factor out the space of data variation and makes learning at each step much easier. Empirically, we show that the performance for pose-varying objects can be improved dramatically. Our method can also be applied on real data captured by 3D sensors, and the flow generated by our methods can be used for generating high quality results in higher resolution.
Hao Zhu 0004, Peng Wang 0001, Xun Cao, Ruigang Yang
CVPR3
2018 Occlusion Aware Unsupervised Learning of Optical Flow
abstract
It has been recently shown that a convolutional neural network can learn optical flow estimation with unsupervised learning. However, the performance of the unsupervised methods still has a relatively large gap compared to its supervised counterpart. Occlusion and large motion are some of the major factors that limit the current unsupervised learning of optical flow methods. In this work we introduce a new method which models occlusion explicitly and a new warping way that facilitates the learning of large motion. Our method shows promising results on Flying Chairs, MPI-Sintel and KITTI benchmark datasets. Especially on KITTI dataset where abundant unlabeled samples exist, our unsupervised method outperforms its counterpart trained with supervised learning.
Yang Wang 0046, Yi Yang 0007, Zhenheng Yang, Liang Zhao 0006, Peng Wang 0001, Wei Xu 0017
CVPR5
2018 MaskLab: Instance Segmentation by Refining Object Detection With Semantic and Direction Features
abstract
In this work, we tackle the problem of instance segmentation, the task of simultaneously solving object detection and semantic segmentation. Towards this goal, we present a model, called MaskLab, which produces three outputs: box detection, semantic segmentation, and direction prediction. Building on top of the Faster-RCNN object detector, the predicted boxes provide accurate localization of object instances. Within each region of interest, MaskLab performs foreground/background segmentation by combining semantic and direction prediction. Semantic segmentation assists the model in distinguishing between objects of different semantic classes including background, while the direction prediction, estimating each pixel's direction towards its corresponding center, allows separating instances of the same semantic class. Moreover, we explore the effect of incorporating recent successful methods from both segmentation and detection (i.e. atrous convolution and hypercolumn). Our proposed model is evaluated on the COCO instance segmentation benchmark and shows comparable performance with other state-of-art models.
Liang-Chieh Chen, Alexander Hermans, George Papandreou, Florian Schroff, Peng Wang 0001, Hartwig Adam
CVPR5
2018 DeLS-3D: Deep Localization and Segmentation With a 3D Semantic Map
abstract
For applications such as augmented reality, autonomous driving, self-localization/camera pose estimation and scene parsing are crucial technologies. In this paper, we propose a unified framework to tackle these two problems simultaneously. The uniqueness of our design is a sensor fusion scheme which integrates camera videos, motion sensors (GPS/IMU), and a 3D semantic map in order to achieve robustness and efficiency of the system. Specifically, we first have an initial coarse camera pose obtained from consumer-grade GPS/IMU, based on which a label map can be rendered from the 3D semantic map. Then, the rendered label map and the RGB image are jointly fed into a pose CNN, yielding a corrected camera pose. In addition, to incorporate temporal information, a multi-layer recurrent neural network (RNN) is further deployed improve the pose accuracy. Finally, based on the pose from RNN, we render a new label map, which is fed together with the RGB image into a segment CNN which produces perpixel semantic label. In order to validate our approach, we build a dataset with registered 3D point clouds and video camera images. Both the point clouds and the images are semantically-labeled. Each video frame has ground truth pose from highly accurate motion sensors. We show that practically, pose estimation solely relying on images like PoseNet [25] may fail due to street view confusion, and it is important to fuse multiple sensors. Finally, various ablation studies are performed, which demonstrate the effectiveness of the proposed system. In particular, we show that scene parsing and pose estimation are mutually beneficial to achieve a more robust and accurate system.
Peng Wang 0001, Ruigang Yang, Binbin Cao, Wei Xu 0017, Yuanqing Lin
CVPR1
2018 LEGO: Learning Edge With Geometry All at Once by Watching Videos
abstract
Learning to estimate 3D geometry in a single image by watching unlabeled videos via deep convolutional network is attracting significant attention. In this paper, we introduce a "3D as-smooth-as-possible (3D-ASAP)" prior inside the pipeline, which enables joint estimation of edges and 3D scene, yielding results with significant improvement in accuracy for fine detailed structures. Specifically, we define the 3D-ASAP prior by requiring that any two points recovered in 3D from an image should lie on an existing planar surface if no other cues provided. We design an unsupervised framework that Learns Edges and Geometry (depth, normal) all at Once (LEGO). The predicted edges are embedded into depth and surface normal smoothness terms, where pixels without edges in-between are constrained to satisfy the prior. In our framework, the predicted depths, normals and edges are forced to be consistent all the time. We conduct experiments on KITTI to evaluate our estimated geometry and CityScapes to perform edge evaluation. We show that in all of the tasks, i.e. depth, normal and edge, our algorithm vastly outperforms other state-of-the-art (SOTA) algorithms, demonstrating the benefits of our approach.
Zhenheng Yang, Peng Wang 0001, Yang Wang 0046, Wei Xu 0017, Ramakant Nevatia
CVPR2
2018 Depth Estimation via Affinity Learned with Convolutional Spatial Propagation Network
Xinjing Cheng, Peng Wang 0001, Ruigang Yang
ECCV (16)2
2017 Joint Multi-person Pose Estimation and Semantic Part Segmentation
abstract
Human pose estimation and semantic part segmentation are two complementary tasks in computer vision. In this paper, we propose to solve the two tasks jointly for natural multi-person images, in which the estimated pose provides object-level shape prior to regularize part segments while the part-level segments constrain the variation of pose locations. Specifically, we first train two fully convolutional neural networks (FCNs), namely Pose FCN and Part FCN, to provide initial estimation of pose joint potential and semantic part potential. Then, to refine pose joint location, the two types of potentials are fused with a fully-connected conditional random field (FCRF), where a novel segment-joint smoothness term is used to encourage semantic and spatial consistency between parts and joints. To refine part segments, the refined pose and the original part potential are integrated through a Part FCN, where the skeleton feature from pose serves as additional regularization cues for part segments. Finally, to reduce the complexity of the FCRF, we induce human detection boxes and infer the graph inside each box, making the inference forty times faster. Since theres no dataset that contains both part segments and pose labels, we extend the PASCAL VOC part dataset [6] with human pose joints and perform extensive experiments to compare our method against several most recent strategies. We show that our algorithm surpasses competing methods by 10.6% in pose estimation with much faster speed and by 1.5% in semantic part segmentation.
Fangting Xia, Peng Wang 0001, Xianjie Chen, Alan L. Yuille
CVPR2
2016 Pose-Guided Human Parsing by an AND/OR Graph Using Pose-Context Features
abstract
Parsing human into semantic parts is crucial to human-centric analysis. In this paper, we propose a human parsing pipeline that uses pose cues, e.g., estimates of human joint locations, to provide pose-guided segment proposals for semantic parts. These segment proposals are ranked using standard appearance cues, deep-learned semantic feature, and a novel pose feature called pose-context. Then these proposals are selected and assembled using an And-Or graph to output a parse of the person. The And-Or graph is able to deal with large human appearance variability due to pose, choice of clothing, etc. We evaluate our approach on the popular Penn-Fudan pedestrian parsing dataset, showing that it significantly outperforms the state of the art, and perform diagnostics to demonstrate the effectiveness of different stages of our pipeline.
Fangting Xia, Jun Zhu 0001, Peng Wang 0001, Alan L. Yuille
AAAI3
2016 DOC: Deep OCclusion Estimation from a Single Image
Peng Wang 0001, Alan L. Yuille
ECCV (1)1
2016 Zoom Better to See Clearer: Human and Object Parsing with Hierarchical Auto-Zoom Net
Fangting Xia, Peng Wang 0001, Liang-Chieh Chen, Alan L. Yuille
ECCV (5)2
2016 SURGE: Surface Regularized Geometry Estimation from a Single Image
abstract
This paper introduces an approach to regularize 2.5D surface normal and depth predictions at each pixel given a single input image. The approach infers and reasons about the underlying 3D planar surfaces depicted in the image to snap predicted normals and depths to inferred planar surfaces, all while maintaining fine detail within objects. Our approach comprises two components: (i) a fourstream convolutional neural network (CNN) where depths, surface normals, and likelihoods of planar region and planar boundary are predicted at each pixel, followed by (ii) a dense conditional random field (DCRF) that integrates the four predictions such that the normals and depths are compatible with each other and regularized by the planar region and planar boundary information. The DCRF is formulated such that gradients can be passed to the surface normal and depth CNNs via backpropagation. In addition, we propose new planar wise metrics to evaluate geometry consistency within planar surfaces, which are more tightly related to dependent 3D editing applications. We show that our regularization yields a 30% relative improvement in planar consistency on the NYU v2 dataset.
Peng Wang 0001, Xiaohui Shen, Bryan C. Russell, Scott Cohen, Brian L. Price, Alan L. Yuille
NIPS1
2015 Towards unified depth and semantic prediction from a single image
abstract
Depth estimation and semantic segmentation are two fundamental problems in image understanding. While the two tasks are strongly correlated and mutually beneficial, they are usually solved separately or sequentially. Motivated by the complementary properties of the two tasks, we propose a unified framework for joint depth and semantic prediction. Given an image, we first use a trained Convolutional Neural Network (CNN) to jointly predict a global layout composed of pixel-wise depth values and semantic labels. By allowing for interactions between the depth and semantic information, the joint network provides more accurate depth prediction than a state-of-the-art CNN trained solely for depth prediction [6]. To further obtain fine-level details, the image is decomposed into local segments for region-level depth and semantic prediction under the guidance of global layout. Utilizing the pixel-wise global prediction and region-wise local prediction, we formulate the inference problem in a two-layer Hierarchical Conditional Random Field (HCRF) to produce the final depth and semantic map. As demonstrated in the experiments, our approach effectively leverages the advantages of both tasks and provides the state-of-the-art results.
Peng Wang 0001, Xiaohui Shen, Zhe Lin 0001, Scott Cohen, Brian L. Price, Alan L. Yuille
CVPR1
2015 Joint Object and Part Segmentation Using Deep Learned Potentials
abstract
Segmenting semantic objects from images and parsing them into their respective semantic parts are fundamental steps towards detailed object understanding in computer vision. In this paper, we propose a joint solution that tackles semantic object and part segmentation simultaneously, in which higher object-level context is provided to guide part segmentation, and more detailed part-level localization is utilized to refine object segmentation. Specifically, we first introduce the concept of semantic compositional parts (SCP) in which similar semantic parts are grouped and shared among different objects. A two-stream fully convolutional network (FCN) is then trained to provide the SCP and object potentials at each pixel. At the same time, a compact set of segments can also be obtained from the SCP predictions of the network. Given the potentials and the generated segments, in order to explore long-range context, we finally construct an efficient fully connected conditional random field (FCRF) to jointly predict the final object and part labels. Extensive evaluation on three different datasets shows that our approach can mutually enhance the performance of object and part segmentation, and outperforms the current state-of-the-art on both tasks.
Peng Wang 0001, Xiaohui Shen, Zhe Lin 0001, Scott Cohen, Brian L. Price, Alan L. Yuille
ICCV1
2015 Learning an Aesthetic Photo Cropping Cascade
abstract
Cropping is one of the most fundamental and common operations in image processing for improving the aesthetic quality of photographs. Instead of manually designing rules for cropping, in this paper, we propose a generative model that learns an aesthetic photo cropping cascade from a large database of well-composed images and a dataset containing images with crops generated by expert photographers. Specifically, this model includes cropping priori, intuitive likelihood, compositional likelihood and change likelihood. Our learning exploits a spatial pyramid saliency feature and a multi-level foreground segmentation. The inference is done by efficient sub window search (ESS) [10] which is benefited from the bound at conditional distribution in the cascade. Additionally, for extracting attentional subjects and capturing scene composition, we design an iterative saliency method to model the saliency moving paths, which is beyond the typical saliency model predicting a single attentional region. Experiments show that our approach outperforms the state-of-the-art cropping methods by a large margin.
Peng Wang 0001, Zhe Lin 0001, Radomír Mech
WACV1
2015 Error Factor Analysis for Wild Scene Image-Labelling
abstract
PASCAL VOC Segmentation Challenge [10] is currently considered as one of the datasets that reflect the image segmentation difficulties for real world scenarios [29]. However, current evaluation is simply based on a single Inter-section Over Union (IOU) score. In this paper, we try to discover the error factors under the IOU, which makes the results more informative to understand rather than a black box. Specifically, we decompose the error into three error types in terms of object characteristics, i.e. general, appearance and shape. Each error type is composed of respective factors, e.g. size and aspect ratio for general, appearance distinctiveness for appearance, etc. Finally, for each factor and error type, we perform analysis over its impact on and correlation with the final IOU through robust regression. Our experiments show that these error factors have significant relationship with the given IOU accuracy, and the analysis provides practical guidance on further improvement of the given algorithm.
Peng Wang 0001, Alan L. Yuille
WACV1
2013 Supervised Kernel Descriptors for Visual Recognition
abstract
In visual recognition tasks, the design of low level image feature representation is fundamental. The advent of local patch features from pixel attributes such as SIFT and LBP, has precipitated dramatic progresses. Recently, a kernel view of these features, called kernel descriptors (KDES), generalizes the feature design in an unsupervised fashion and yields impressive results. In this paper, we present a supervised framework to embed the image level label information into the design of patch level kernel descriptors, which we call supervised kernel descriptors (SKDES). Specifically, we adopt the broadly applied bag-of-words (BOW) image classification pipeline and a large margin criterion to learn the low-level patch representation, which makes the patch features much more compact and achieve better discriminative ability than KDES. With this method, we achieve competitive results over several public datasets comparing with state-of-the-art methods.
Peng Wang 0001, Jingdong Wang 0001, Weiwei Xu 0003, Hongbin Zha, Shipeng Li 0001
CVPR1
2013 Structure-Sensitive Superpixels via Geodesic Distance
Peng Wang 0001, Rui Gan, Jingdong Wang 0001, Hongbin Zha
Int. J. Comput. Vis.1
2012 Salient object detection for searched web images via global saliency
abstract
In this paper, we deal with the problem of detecting the existence and the location of salient objects for thumbnail images on which most search engines usually perform visual analysis in order to handle web-scale images. Different from previous techniques, such as sliding window-based or segmentation-based schemes for detecting salient objects, we propose to use a learning approach, random forest in our solution. Our algorithm exploits global features from multiple saliency indicators to directly predict the existence and the position of the salient object. To validate our algorithm, we constructed a large image database collected from Bing image search, that contains hundreds of thousands of manually labeled web images. The experimental results using this new database and the resized MSRA database [16] demonstrate that our algorithm outperforms previous state-of-the-art methods.
Peng Wang 0001, Jingdong Wang 0001, Jie Feng 0012, Hongbin Zha, Shipeng Li 0001
CVPR1
2012 Color filter for image search
abstract
Image search relying on surrounding texts can return reliably relevant images to some extent. Most recent efforts are focusing on utilizing visual contents to help users find images with specific visual requirements. In this demonstration, we show a color filter scheme for image search, which enables users to find images containing objects or scenes with their interested color. Color is one of the most crucial cues in describing visual contents and has been frequently used in various applications. The key components in developing our demo include salient object detection and perceptual color naming. The color filter in Microsoft Bing image search is developed using these techniques.
Peng Wang 0001, Dongqing Zhang, Jingdong Wang 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia1
2011 Structure-sensitive superpixels via geodesic distance
abstract
Over-segments (i.e. superpixels) have been commonly used as supporting regions for feature vectors and primitives to reduce computational complexity in various image analysis tasks. In this paper, we describe a structuresensitive over-segmentation technique by exploiting Lloyd's algorithm with a geodesic distance. It generates smaller superpixels to achieve lower under-segmentation in structure-dense regions with high intensity or color variation, and produces larger segments to increase computational efficiency in structure-sparse regions with homogeneous appearance. We adopt geometric flows to compute the geodesic distances amongst pixels, and in the segmentation procedure, the density of over-segments is automatically adjusted according to an energy functional that embeds color homogeneity, structure density and compactness constraints. Comparative experiments with the Berkeley database show that the proposed algorithm outperforms prior arts while offering a comparable computational efficiency with fast methods, such as TurboPixels.
Peng Wang 0001, Jingdong Wang 0001, Rui Gan, Hongbin Zha
ICCV2