Jianping Shi

dblp:00/3188 · DBLP profile ↗
← Back
88ranked-venue papers
9as first author
19since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 74 · 7 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 65 · 7 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Cosine-Initialized MAE for Cross-Domain Few-Shot Recognition in Distributed Fiber-Optic Vibration Sensing Systems
Xiankun Wang, Zhengxian Zhou, Dawei Zhang 0009, Jun Qu, Jianping Shi, Yashuai Han, Xinyan Yang, Songlin Zhuang
IEEE Internet Things J.5
2024 Delving Into the Devils of Bird's-Eye-View Perception: A Review, Evaluation and Recipe
abstract
Learning powerful representations in bird's-eye-view (BEV) for perception tasks is trending and drawing extensive attention both from industry and academia. Conventional approaches for most autonomous driving algorithms perform detection, segmentation, tracking, etc., in a front or perspective view. As sensor configurations get more complex, integrating multi-source information from different sensors and representing features in a unified view come of vital importance. BEV perception inherits several advantages, as representing surrounding scenes in BEV is intuitive and fusion-friendly; and representing objects in BEV is most desirable for subsequent modules as in planning and/or control. The core problems for BEV perception lie in (a) how to reconstruct the lost 3D information via view transformation from perspective view to BEV; (b) how to acquire ground truth annotations in BEV grid; (c) how to formulate the pipeline to incorporate features from different sources and views; and (d) how to adapt and generalize algorithms as sensor configurations vary across different scenarios. In this survey, we review the most recent works on BEV perception and provide an in-depth analysis of different solutions. Moreover, several systematic designs of BEV approach from the industry are depicted as well. Furthermore, we introduce a full suite of practical guidebook to improve the performance of BEV perception tasks, including camera, LiDAR and fusion inputs. At last, we point out the future research directions in this area. We hope this report will shed some light on the community and encourage more research effort on BEV perception.
Hongyang Li 0001, Chonghao Sima, Jifeng Dai, Wenhai Wang, Lewei Lu, Huijie Wang, Jiazhi Yang, Hanming Deng, Hao Tian 0006, Enze Xie, Jiangwei Xie, Li Chen 0008, Tianyu Li 0004, Yang Li 0189, Yulu Gao, Xiaosong Jia, Si Liu 0001, Jianping Shi, Dahua Lin, Yu Qiao 0001
IEEE Trans. Pattern Anal. Mach. Intell.20
2023 PV-RCNN++: Point-Voxel Feature Set Abstraction With Local Vector Representation for 3D Object Detection
abstract
Abstract 3D object detection is receiving increasing attention from both industry and academia thanks to its wide applications in various fields. In this paper, we propose Point-Voxel Region-based Convolution Neural Networks (PV-RCNNs) for 3D object detection on point clouds. First, we propose a novel 3D detector, PV-RCNN, which boosts the 3D detection performance by deeply integrating the feature learning of both point-based set abstraction and voxel-based sparse convolution through two novel steps, i.e. , the voxel-to-keypoint scene encoding and the keypoint-to-grid RoI feature abstraction. Second, we propose an advanced framework, PV-RCNN++, for more efficient and accurate 3D object detection. It consists of two major improvements: sectorized proposal-centric sampling for efficiently producing more representative keypoints, and VectorPool aggregation for better aggregating local point features with much less resource consumption. With these two strategies, our PV-RCNN++ is about $$3\times $$ 3 × faster than PV-RCNN, while also achieving better performance. The experiments demonstrate that our proposed PV-RCNN++ framework achieves state-of-the-art 3D detection performance on the large-scale and highly-competitive Waymo Open Dataset with 10 FPS inference speed on the detection range of $$150m \times 150m$$ 150 m × 150 m .
Shaoshuai Shi, Li Jiang 0009, Jiajun Deng, Zhe Wang 0006, Chaoxu Guo, Jianping Shi, Xiaogang Wang 0001, Hongsheng Li 0001
Int. J. Comput. Vis.6
2023 Context-Aware Mixup for Domain Adaptive Semantic Segmentation
abstract
Unsupervised domain adaptation (UDA) aims to adapt a model of the labeled source domain to an unlabeled target domain. Existing UDA-based semantic segmentation approaches always reduce the domain shifts in pixel level, feature level, and output level. However, almost all of them largely neglect the contextual dependency, which is generally shared across different domains, leading to less-desired performance. In this paper, we propose a novel Context-Aware Mixup (CAMix) framework for domain adaptive semantic segmentation, which exploits this important clue of context-dependency as explicit prior knowledge in a fully end-to-end trainable manner for enhancing the adaptability toward the target domain. Firstly, we present a contextual mask generation strategy by leveraging the accumulated spatial distributions and prior contextual relationships. The generated contextual mask is critical in this work and will guide the context-aware domain mixup on three different levels. Besides, provided the context knowledge, we introduce a significance-reweighted consistency loss to penalize the inconsistency between the mixed student prediction and the mixed teacher prediction, which alleviates the negative transfer of the adaptation, e.g., early performance degradation. Extensive experiments and analysis demonstrate the effectiveness of our method against the state-of-the-art approaches on widely-used UDA benchmarks.
Qianyu Zhou 0001, Zhengyang Feng, Jiangmiao Pang, Xuequan Lu, Jianping Shi, Lizhuang Ma
IEEE Trans. Circuits Syst. Video Technol.7
2022 PersFormer: 3D Lane Detection via Perspective Transformer and the OpenLane Benchmark
Li Chen 0008, Chonghao Sima, Yang Li 0189, Zehan Zheng, Jiajie Xu 0001, Xiangwei Geng, Hongyang Li 0001, Conghui He, Jianping Shi, Yu Qiao 0001, Junchi Yan
ECCV (38)9
2022 Uncertainty-aware consistency regularization for cross-domain semantic segmentation
Qianyu Zhou 0001, Zhengyang Feng, Xuequan Lu, Jianping Shi, Lizhuang Ma
Comput. Vis. Image Underst.6
2022 AdaStereo: An Efficient Domain-Adaptive Stereo Matching Approach
Xiao Song 0002, Guorun Yang, Xinge Zhu, Hui Zhou 0005, Yuexin Ma, Zhe Wang 0006, Jianping Shi
Int. J. Comput. Vis.7
2022 Correction to: AdaStereo: An Efficient Domain-Adaptive Stereo Matching Approach
Xiao Song 0002, Guorun Yang, Xinge Zhu, Hui Zhou 0005, Yuexin Ma, Zhe Wang 0006, Jianping Shi
Int. J. Comput. Vis.7
2022 DMT: Dynamic mutual training for semi-supervised learning
Zhengyang Feng, Qianyu Zhou 0001, Xin Tan 0002, Xuequan Lu, Jianping Shi, Lizhuang Ma
Pattern Recognit.7
2021 MP-Mono: Monocular 3D Detection Using Multiple Priors for Autonomous Driving
abstract
Monocular 3D object detection is an important and challenging task in autonomous driving. Due to the ill-posed nature of 3D detection, recent studies use prior knowledge on object categories to estimate 3D parameters. However, for each object category in real driving scenes, there exist a couple of sub-categories with different shapes (i.e. length, width, and height). For example, vehicle generally contains the sub-categories of car, van, and truck. Obviously, single prior knowledge cannot cover such diverse sub-categories. In this paper, we propose MP-Mono that exploits multiple priors to improve object detection. Specifically, a data-heuristic strategy is presented to generate multiple 3D proposals, in which we leverage the unsupervised algorithm to cluster potential sub-categories from realistic datasets, and a height-guided inference policy is used to determine the initial distances of proposals, reducing the difficulty of network learning. Additionally, we propose a local-ground embedding method that learns local depth information to enhance monocular 3D detection. The experimental results on the KITTI dataset demonstrate that our MP-Mono achieves competitive performances compared to other monocular methods, verifying the effectiveness of multi-prior integration and local-ground embedding.
Guorun Yang, Zhe Wang 0006, Jianping Shi, Zhidong Deng, Yu Qiao 0001
3DV5
2021 Understanding the wiring evolution in differentiable neural architecture search
abstract
Controversy exists on whether differentiable neural architecture search methods discover wiring topology effectively. To understand how wiring topology evolves, we study the underlying mechanism of several existing differentiable NAS frameworks. Our investigation is motivated by three observed searching patterns of differentiable NAS: 1) they search by growing instead of pruning; 2) wider networks are more preferred than deeper ones; 3) no edges are selected in bi-level optimization. To anatomize these phenomena, we propose a unified view on searching algorithms of existing frameworks, transferring the global optimization to local cost minimization. Based on this reformulation, we conduct empirical and theoretical analyses, revealing implicit biases in the cost’s assignment mechanism and evolution dynamics that cause the observed phenomena. These biases indicate strong discrimination towards certain topologies. To this end, we pose questions that future differentiable methods for neural wiring discovery need to confront, hoping to evoke a discussion and rethinking on how much bias has been enforced implicitly in existing NAS methods.
Sirui Xie, Shoukang Hu, Xinjiang Wang, Jianping Shi, Xunying Liu, Dahua Lin
AISTATS5
2021 PointFlow: Flowing Semantics Through Points for Aerial Image Segmentation
abstract
Aerial Image Segmentation is a particular semantic segmentation problem and has several challenging characteristics that general semantic segmentation does not have. There are two critical issues: The one is an extremely foreground-background imbalanced distribution, and the other is multiple small objects along with the complex background. Such problems make the recent dense affinity context modeling perform poorly even compared with baselines due to over-introduced background context. To handle these problems, we propose a point-wise affinity propagation module based on the Feature Pyramid Network (FPN) framework, named PointFlow. Rather than dense affinity learning, a sparse affinity map is generated upon selected points between the adjacent features, which reduces the noise introduced by the background while keeping efficiency. In particular, we design a dual point matcher to select points from the salient area and object boundaries, respectively. Experimental results on three different aerial segmentation datasets suggest that the proposed method is more effective and efficient than state-of-the-art general semantic segmentation methods. Especially, our methods achieve the best speed and accuracy trade-off on three aerial benchmarks. Further experiments on three general semantic segmentation datasets prove the generality of our method. Code and models are made available (https://github.com/lxtGH/PFSegNets).
Xiangtai Li, Hao He 0015, Xia Li 0005, Jianping Shi, Lubin Weng, Yunhai Tong, Zhouchen Lin
CVPR6
2021 AdaStereo: A Simple and Efficient Approach for Adaptive Stereo Matching
abstract
Recently, records on stereo matching benchmarks are constantly broken by end-to-end disparity networks. However, the domain adaptation ability of these deep models is quite poor. Addressing such problem, we present a novel domain-adaptive pipeline called AdaStereo that aims to align multi-level representations for deep stereo matching networks. Compared to previous methods for adaptive stereo matching, our AdaStereo realizes a more standard, complete and effective domain adaptation pipeline. Firstly, we propose a non-adversarial progressive color transfer algorithm for input image-level alignment. Secondly, we design an efficient parameter-free cost normalization layer for internal feature-level alignment. Lastly, a highly related auxiliary task, self-supervised occlusion-aware reconstruction is presented to narrow down the gaps in output space. Our AdaStereo models achieve state-of-the-art cross-domain performance on multiple stereo benchmarks, including KITTI, Middlebury, ETH3D, and DrivingStereo, even outperforming disparity networks finetuned with target-domain ground-truths.
Xiao Song 0002, Guorun Yang, Xinge Zhu, Hui Zhou 0005, Zhe Wang 0006, Jianping Shi
CVPR6
2021 PIT: Position-Invariant Transform for Cross-FoV Domain Adaptation
abstract
Cross-domain object detection and semantic segmentation have witnessed impressive progress recently. Existing approaches mainly consider the domain shift resulting from external environments including the changes of background, illumination or weather, while distinct camera intrinsic parameters appear commonly in different domains and their influence for domain adaptation has been very rarely explored. In this paper, we observe that the Field of View (FoV) gap induces noticeable instance appearance differences between the source and target domains. We further discover that the FoV gap between two domains impairs domain adaptation performance under both the FoV-increasing (source FoV < target FoV) and FoV-decreasing cases. Motivated by the observations, we propose the Position-Invariant Transform (PIT) to better align images in different domains. We also introduce a reverse PIT for mapping the transformed/aligned images back to the original image space, and design a loss reweighting strategy to accelerate the training process. Our method can be easily plugged into existing cross-domain detection/segmentation frameworks, while bringing about negligible computational overhead. Extensive experiments demonstrate that our method can soundly boost the performance on both cross-domain object detection and segmentation for state-of-the-art techniques. Our code is available at https://github.com/sheepooo/PIT-Position-Invariant-Transform.
Qianyu Zhou 0001, Zhengyang Feng, Xuequan Lu, Jianping Shi, Lizhuang Ma
ICCV7
2021 Enhanced Boundary Learning for Glass-like Object Segmentation
abstract
Glass-like objects such as windows, bottles, and mirrors exist widely in the real world. Sensing these objects has many applications, including robot navigation and grasping. However, this task is very challenging due to the arbitrary scenes behind glass-like objects. This paper aims to solve the glass-like object segmentation problem via enhanced boundary learning. In particular, we first propose a novel refined differential module that outputs finer boundary cues. We then introduce an edge-aware point-based graph convolution network module to model the global shape along the boundary. We use these two modules to design a decoder that generates accurate and clean segmentation results, especially on the object contours. Both modules are lightweight and effective: they can be embedded into various segmentation models. In extensive experiments on three recent glass-like object segmentation datasets, including Trans10k, MSD, and GDD, our approach establishes new state-of-the-art results. We also illustrate the strong generalization properties of our method on three generic segmentation datasets, including Cityscapes, BDD, and COCO Stuff. Code and models will be available for further research.
Hao He 0015, Xiangtai Li, Jianping Shi, Yunhai Tong, Gaofeng Meng, Véronique Prinet, Lubin Weng
ICCV4
2021 An FPGA-Based Neural Network Overlay for ADAS Supporting Multi-Model and Multi-Mode
abstract
Advanced Driver-Assistance Systems (ADAS) are complex systems consisting of many computer vision tasks including image classification, object detection and semantic segmentation. FPGA is a feasible solution for deep learning based computer vision accelerator due to its high performance and energy efficiency. However, design a high performance FPGA accelerator requires good understanding of basic hardware concepts and consumes a long compilation time. Overlays can alleviate the above problems by accelerating applications in a software via a hardware architecture and a compiler. In this paper, we propose an FPGA-based neural network overlay processor for ADAS. The overlay architecture contains almost all common computation layers for learning based ADAS. In addition, we design a compiler that can automatically compile the high-level description of neural networks from deep learning framework like Caffe and Tensorflow into FPGA configurable codes, which can be executed by our overlay architecture without reprogramming. Experiments show that our overlay can process learning tasks in ADAS with low latency and low memory usage.
Jiaxi Zhang 0001, Tao Yang 0031, Qingzheng Li, Guojie Luo, Jianping Shi
ISCAS7
2021 Towards Balanced Learning for Instance Recognition
Jiangmiao Pang, Kai Chen 0026, Qi Li 0018, Zhi-hai Xu, Huajun Feng, Jianping Shi, Wanli Ouyang, Dahua Lin
Int. J. Comput. Vis.6
2021 CDTD: A Large-Scale Cross-Domain Benchmark for Instance-Level Image-to-Image Translation and Domain Adaptive Object Detection
Mingyang Huang, Jianping Shi, Zechun Liu, Harsh Maheshwari, Yutong Zheng, Xiangyang Xue 0001, Marios Savvides, Thomas S. Huang
Int. J. Comput. Vis.3
2021 From Points to Parts: 3D Object Detection From Point Cloud With Part-Aware and Part-Aggregation Network
abstract
3D object detection from LiDAR point cloud is a challenging problem in 3D scene understanding and has many practical applications. In this paper, we extend our preliminary work PointRCNN to a novel and strong point-cloud-based 3D object detection framework, the part-aware and aggregation neural network (Part-A2net). The whole framework consists of the part-aware stage and the part-aggregation stage. First, the part-aware stage for the first time fully utilizes free-of-charge part supervisions derived from 3D ground-truth boxes to simultaneously predict high quality 3D proposals and accurate intra-object part locations. The predicted intra-object part locations within the same proposal are grouped by our new-designed RoI-aware point cloud pooling module, which results in an effective representation to encode the geometry-specific features of each 3D proposal. Then the part-aggregation stage learns to re-score the box and refine the box location by exploring the spatial relationship of the pooled intra-object part locations. Extensive experiments are conducted to demonstrate the performance improvements from each component of our proposed framework. Our Part-A2net outperforms all existing 3D detection methods and achieves new state-of-the-art on KITTI 3D object detection dataset by utilizing only the LiDAR point cloud data.
Shaoshuai Shi, Zhe Wang 0006, Jianping Shi, Xiaogang Wang 0001, Hongsheng Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 Every Frame Counts: Joint Learning of Video Segmentation and Optical Flow
abstract
A major challenge for video semantic segmentation is the lack of labeled data. In most benchmark datasets, only one frame of a video clip is annotated, which makes most supervised methods fail to utilize information from the rest of the frames. To exploit the spatio-temporal information in videos, many previous works use pre-computed optical flows, which encode the temporal consistency to improve the video segmentation. However, the video segmentation and optical flow estimation are still considered as two separate tasks. In this paper, we propose a novel framework for joint video semantic segmentation and optical flow estimation. Semantic segmentation brings semantic information to handle occlusion for more robust optical flow estimation, while the non-occluded optical flow provides accurate pixel-level temporal correspondences to guarantee the temporal consistency of the segmentation. Moreover, our framework is able to utilize both labeled and unlabeled frames in the video through joint training, while no additional calculation is required in inference. Extensive experiments show that the proposed model makes the video semantic segmentation and optical flow estimation benefit from each other and outperforms existing methods under the same settings in both tasks.
Mingyu Ding, Zhe Wang 0006, Bolei Zhou, Jianping Shi, Zhiwu Lu 0001, Ping Luo 0002
AAAI4
2020 Learning Depth-Guided Convolutions for Monocular 3D Object Detection
abstract
3D object detection from a single image without LiDAR is a challenging task due to the lack of accurate depth information. Conventional 2D convolutions are unsuitable for this task because they fail to capture local object and its scale information, which are vital for 3D object detection. To better represent 3D structure, prior arts typically transform depth maps estimated from 2D images into a pseudo-LiDAR representation, and then apply existing 3D point-cloud based object detectors. However, their results depend heavily on the accuracy of the estimated depth maps, resulting in suboptimal performance. In this work, instead of using pseudo-LiDAR representation, we improve the fundamental 2D fully convolutions by proposing a new local convolutional network (LCN), termed Depth-guided Dynamic-Depthwise-Dilated LCN (D4LCN), where the filters and their receptive fields can be automatically learned from image-based depth maps, making different pixels of different images have different filters. D4LCN overcomes the limitation of conventional 2D convolutions and narrows the gap between image representation and 3D point cloud representation. Extensive experiments show that D4LCN outperforms existing works by large margins. For example, the relative improvement of D4LCN against the state-of-the-art on KITTI is 9.1\% in the moderate setting. D4LCN ranks 1st on KITTI monocular 3D object detection benchmark at the time of submission (car, December 2019). The code is available at https://github.com/dingmyu/D4LCN
Mingyu Ding, Yuqi Huo, Hongwei Yi, Zhe Wang 0006, Jianping Shi, Zhiwu Lu 0001, Ping Luo 0002
CVPR5
2020 TPNet: Trajectory Proposal Network for Motion Prediction
abstract
Making accurate motion prediction of the surrounding traffic agents such as pedestrians, vehicles, and cyclists is crucial for autonomous driving. Recent data-driven motion prediction methods have attempted to learn to directly regress the exact future position or its distribution from massive amount of trajectory data. However, it remains difficult for these methods to provide multimodal predictions as well as integrate physical constraints such as traffic rules and movable areas. In this work we propose a novel two-stage motion prediction framework, Trajectory Proposal Network (TPNet). TPNet first generates a candidate set of future trajectories as hypothesis proposals, then makes the final predictions by classifying and refining the proposals which meets the physical constraints. By steering the proposal generation process, safe and multimodal predictions are realized. Thus this framework effectively mitigates the complexity of motion prediction problem while ensuring the multimodal output. Experiments on four large-scale trajectory prediction datasets, i.e. the ETH, UCY, Apollo and Argoverse datasets, show that TPNet achieves the state-of-the-art results both quantitatively and qualitatively.
Liangji Fang, Qinhong Jiang, Jianping Shi, Bolei Zhou
CVPR3
2020 DSNAS: Direct Neural Architecture Search Without Parameter Retraining
abstract
If NAS methods are solutions, what is the problem? Most existing NAS methods require two-stage parameter optimization. However, performance of the same architecture in the two stages correlates poorly. In this work, we propose a new problem definition for NAS, task-specific end-to-end, based on this observation. We argue that given a computer vision task for which a NAS method is expected, this definition can reduce the vaguely-defined NAS evaluation to i) accuracy of this task and ii) the total computation consumed to finally obtain a model with satisfying accuracy. Seeing that most existing methods do not solve this problem directly, we propose DSNAS, an efficient differentiable NAS framework that simultaneously optimizes architecture and parameters with a low-biased Monte Carlo estimate. Child networks derived from DSNAS can be deployed directly without parameter retraining. Comparing with two-stage methods, DSNAS successfully discovers networks with comparable accuracy (74.4\%) on ImageNet in 420 GPU hours, reducing the total time by more than 34\%.
Shoukang Hu, Sirui Xie, Hehui Zheng, Jianping Shi, Xunying Liu, Dahua Lin
CVPR5
2020 Graph-Guided Architecture Search for Real-Time Semantic Segmentation
abstract
Designing a lightweight semantic segmentation network often requires researchers to find a trade-off between performance and speed, which is always empirical due to the limited interpretability of neural networks. In order to release researchers from these tedious mechanical trials, we propose a Graph-guided Architecture Search (GAS) pipeline to automatically search real-time semantic segmentation networks. Unlike previous works that use a simplified search space and stack a repeatable cell to form a network, we introduce a novel search mechanism with a new search space where a lightweight model can be effectively explored through the cell-level diversity and latency oriented constraint. Specifically, to produce the cell-level diversity, the cell-sharing constraint is eliminated through the cell-independent manner. Then a graph convolution network (GCN) is seamlessly integrated as a communication mechanism between cells. Finally, a latency-oriented constraint is endowed into the search process to balance the speed and performance. Extensive experiments on Cityscapes and CamVid datasets demonstrate that GAS achieves the new state-of-the-art trade-off between accuracy and speed. In particular, on Cityscapes dataset, GAS achieves the new best performance of 73.5% mIoU with the speed of 108.4 FPS on Titan Xp.
Peiwen Lin, Peng Sun 0011, Sirui Xie, Xi Li 0001, Jianping Shi
CVPR6
2020 PV-RCNN: Point-Voxel Feature Set Abstraction for 3D Object Detection
abstract
We present a novel and high-performance 3D object detection framework, named PointVoxel-RCNN (PV-RCNN), for accurate 3D object detection from point clouds. Our proposed method deeply integrates both 3D voxel Convolutional Neural Network (CNN) and PointNet-based set abstraction to learn more discriminative point cloud features. It takes advantages of efficient learning and high-quality proposals of the 3D voxel CNN and the flexible receptive fields of the PointNet-based networks. Specifically, the proposed framework summarizes the 3D scene with a 3D voxel CNN into a small set of keypoints via a novel voxel set abstraction module to save follow-up computations and also to encode representative scene features. Given the high-quality 3D proposals generated by the voxel CNN, the RoI-grid pooling is proposed to abstract proposal-specific features from the keypoints to the RoI-grid points via keypoint set abstraction. Compared with conventional pooling operations, the RoI-grid feature points encode much richer context information for accurately estimating object confidences and locations. Extensive experiments on both the KITTI dataset and the Waymo Open dataset show that our proposed PV-RCNN surpasses state-of-the-art 3D detection methods with remarkable margins.
Shaoshuai Shi, Chaoxu Guo, Li Jiang 0009, Zhe Wang 0006, Jianping Shi, Xiaogang Wang 0001, Hongsheng Li 0001
CVPR5
2020 Temporal Pyramid Network for Action Recognition
abstract
Visual tempo characterizes the dynamics and the temporal scale of an action. Modeling such visual tempos of different actions facilitates their recognition. Previous works often capture the visual tempo through sampling raw videos at multiple rates and constructing an input-level frame pyramid, which usually requires a costly multi-branch network to handle. In this work we propose a generic Temporal Pyramid Network (TPN) at the feature-level, which can be flexibly integrated into 2D or 3D backbone networks in a plug-and-play manner. Two essential components of TPN, the source of features and the fusion of features, form a feature hierarchy for the backbone so that it can capture action instances at various tempos. TPN also shows consistent improvements over other challenging baselines on several action recognition datasets. Specifically, when equipped with TPN, the 3D ResNet-50 with dense sampling obtains a 2\% gain on the validation set of Kinetics-400. A further analysis also reveals that TPN gains most of its improvements on action classes that have large variances in their visual tempos, validating the effectiveness of TPN.
Ceyuan Yang, Yinghao Xu 0001, Jianping Shi, Bo Dai 0002, Bolei Zhou
CVPR3
2020 TSIT: A Simple and Versatile Framework for Image-to-Image Translation
Liming Jiang 0001, Changxu Zhang, Mingyang Huang, Jianping Shi, Chen Change Loy
ECCV (3)5
2020 Improving Semantic Segmentation via Decoupled Body and Edge Supervision
Xiangtai Li, Xia Li 0005, Li Zhang 0040, Jianping Shi, Zhouchen Lin, Shaohua Tan, Yunhai Tong
ECCV (17)5
2020 Side-Aware Boundary Localization for More Precise Object Detection
Jiaqi Wang 0003, Yuhang Cao, Kai Chen 0026, Jiangmiao Pang, Jianping Shi, Chen Change Loy, Dahua Lin
ECCV (4)7
2020 Search What You Want: Barrier Panelty NAS for Mixed Precision Quantization
Haibao Yu, Jianping Shi
ECCV (9)4
2020 SSN: Shape Signature Networks for Multi-class Object Detection from Point Clouds
Xinge Zhu, Yuexin Ma, Jianping Shi, Dahua Lin
ECCV (25)5
2020 A Winograd-Based CNN Accelerator with a Fine-Grained Regular Sparsity Pattern
abstract
Field-Programmable Gate Array (FPGA) is a high-performance computing platform for Convolution Neural Networks (CNNs) inference. Winograd transformation and weight pruning are widely adopted to reduce the storage and arithmetic overhead in matrix multiplication of CNN on FPGAs. Recent studies strive to prune the weights in the Winograd domain, however, resulting in irregular sparse patterns and leading to low parallelism and reduced utilization of resources. In this paper, we propose a regular sparse pruning pattern in the Winograd-based CNN, namely Sub-Row-Balanced Sparsity (SRBS) pattern, to overcome the above challenge. Then, we develop a 2-step hardware co-optimization approach to improve the model accuracy using the SRBS pattern. Finally, we design an FPGA accelerator that takes advantage of the SRBS pattern to eliminate low-parallelism computation and irregular memory accesses. Experimental results on VGG16 and Resnet-18 with CIFAR-10 and Imagenet show up to 4.4x and 3.06x speedup compared with the state-of-the-art dense Winograd accelerator and 52% (theoretical upper-bound is 72%) performance enhancement compared with the state-of-the-art sparse Winograd accelerator. The resulting sparsity ratio is 80% and 75% and the loss of model accuracy is negligible.
Tao Yang 0031, Yunkun Liao, Jianping Shi, Yun Liang 0001, Naifeng Jing, Li Jiang 0002
FPL3
2020 Outdoor RGBD Instance Segmentation With Residual Regretting Learning
abstract
Indoor semantic segmentation with RGBD input has received decent progress recently, but studies on instance-level objects in outdoor scenarios meet challenges due to the ambiguity in the acquired outdoor depth map. To tackle this problem, we proposed a residual regretting mechanism, incorporated into current flexible, general and solid instance segmentation framework Mask R-CNN in an end-to-end manner. Specifically, regretting cascade is designed to gradually refine and fully unearth useful information in depth maps, acting in a filtering and backup way. Additionally, embedded by a novel residual connection structure, the regretting module combines RGB and depth branches with pixel-level mask robustly. Extensive experiments on the challenging Cityscapes and KITTI dataset manifest the effectiveness of our residual regretting scheme for handling outdoor depth map. Our approach achieves state-of-the-art performance on RGBD instance segmentation, with 13.4% relative improvement over Mask R-CNN on Cityscapes by depth cue.
Zhengtian Xu, Shu Liu 0005, Jianping Shi, Cewu Lu
IEEE Trans. Image Process.3
2019 A2-Net: Molecular Structure Estimation from Cryo-EM Density Volumes
abstract
Constructing of molecular structural models from CryoElectron Microscopy (Cryo-EM) density volumes is the critical last step of structure determination by Cryo-EM technologies. Methods have evolved from manual construction by structural biologists to perform 6D translation-rotation searching, which is extremely compute-intensive. In this paper, we propose a learning-based method and formulate this problem as a vision-inspired 3D detection and pose estimation task. We develop a deep learning framework for amino acid determination in a 3D Cryo-EM density volume. We also design a sequence-guided Monte Carlo Tree Search (MCTS) to thread over the candidate amino acids to form the molecular structure. This framework achieves 91% coverage on our newly proposed dataset and takes only a few minutes for a typical structure with a thousand amino acids. Our method is hundreds of times faster and several times more accurate than existing automated solutions without any human intervention.
Kui Xu 0004, Zhe Wang 0006, Jianping Shi, Hongsheng Li 0001, Qiangfeng Cliff Zhang
AAAI3
2019 Autonomous Driving Towards Mass Production
abstract
Visual recognition technology is very important for autonomous driving especially in direction of mass production. In this talk, we will introduce the algorimic progress for SenseTime in autonoumous driving, as well as our platform foundation for AI technology. Based on this, we illustrate how we make use of these technology into mass production product for autonomous driving.
Jianping Shi
CIKM1
2019 Hybrid Task Cascade for Instance Segmentation
abstract
Cascade is a classic yet powerful architecture that has boosted performance on various tasks. However, how to introduce cascade to instance segmentation remains an open question. A simple combination of Cascade R-CNN and Mask R-CNN only brings limited gain. In exploring a more effective approach, we find that the key to a successful instance segmentation cascade is to fully leverage the reciprocal relationship between detection and segmentation. In this work, we propose a new framework, Hybrid Task Cascade (HTC), which differs in two important aspects: (1) instead of performing cascaded refinement on these two tasks separately, it interweaves them for a joint multi-stage processing; (2) it adopts a fully convolutional branch to provide spatial context, which can help distinguishing hard foreground from cluttered background. Overall, this framework can learn more discriminative features progressively while integrating complementary features together in each stage. Without bells and whistles, a single HTC obtains 38.4% and 1.5% improvement over a strong Cascade Mask R-CNN baseline on MSCOCO dataset. Moreover, our overall system achieves 48.6 mask AP on the test-challenge split, ranking 1st in the COCO 2018 Challenge Object Detection Task. Code is available at https://github.com/open-mmlab/mmdetection.
Kai Chen 0026, Jiangmiao Pang, Jiaqi Wang 0003, Shuyang Sun, Wansen Feng, Ziwei Liu 0002, Jianping Shi, Wanli Ouyang, Chen Change Loy, Dahua Lin
CVPR9
2019 Libra R-CNN: Towards Balanced Learning for Object Detection
abstract
Compared with model architectures, the training process, which is also crucial to the success of detectors, has received relatively less attention in object detection. In this work, we carefully revisit the standard training practice of detectors, and find that the detection performance is often limited by the imbalance during the training process, which generally consists in three levels - sample level, feature level, and objective level. To mitigate the adverse effects caused thereby, we propose Libra R-CNN, a simple but effective framework towards balanced learning for object detection. It integrates three novel components: IoU-balanced sampling, balanced feature pyramid, and balanced L1 loss, respectively for reducing the imbalance at sample, feature, and objective level. Benefitted from the overall balanced design, Libra R-CNN significantly improves the detection performance. Without bells and whistles, it achieves 2.5 points and 2.0 points higher Average Precision (AP) than FPN Faster R-CNN and RetinaNet respectively on MSCOCO.
Jiangmiao Pang, Kai Chen 0026, Jianping Shi, Huajun Feng, Wanli Ouyang, Dahua Lin
CVPR3
2019 Towards Instance-Level Image-To-Image Translation
abstract
Unpaired Image-to-image Translation is a new rising and challenging vision problem that aims to learn a mapping between unaligned image pairs in diverse domains. Recent advances in this field like MUNIT and DRIT mainly focus on disentangling content and style/attribute from a given image first, then directly adopting the global style to guide the model to synthesize new domain images. However, this kind of approaches severely incurs contradiction if the target domain images are content-rich with multiple discrepant objects. In this paper, we present a simple yet effective instance-aware image-to-image translation approach (INIT), which employs the fine-grained local (instance) and global styles to the target image spatially. The proposed INIT exhibits three import advantages: (1) the instance-level objective loss can help learn a more accurate reconstruction and incorporate diverse attributes of objects; (2) the styles used for target domain of local/global areas are from corresponding spatial regions in source domain, which intuitively is a more reasonable mapping; (3) the joint training process can benefit both fine and coarse granularity and incorporates instance information to improve the quality of global translation. We also collect a large-scale benchmark for the new instance-level translation task. We observe that our synthetic images can even benefit real-world vision tasks like generic object detection.
Mingyang Huang, Jianping Shi, Xiangyang Xue 0001, Thomas S. Huang
CVPR3
2019 Not All Areas Are Equal: Transfer Learning for Semantic Segmentation via Hierarchical Region Selection
abstract
The success of deep neural networks for semantic segmentation heavily relies on large-scale and well-labeled datasets, which are hard to collect in practice. Synthetic data offers an alternative to obtain ground-truth labels for free. However, models directly trained on synthetic data often struggle to generalize to real images. In this paper, we consider transfer learning for semantic segmentation that aims to mitigate the gap between abundant synthetic data (source domain) and limited real data (target domain). Unlike previous approaches that either learn mappings to target domain or finetune on target images, our proposed method jointly learn from real images and selectively from realistic pixels in synthetic images to adapt to the target domain. Our key idea is to have weighting networks to score how similar the synthetic pixels are to real ones, and learn such weighting at pixel-, region- and image-levels. We jointly learn these hierarchical weighting networks and segmentation network in an end-to-end manner. Extensive experiments demonstrate that our proposed approach significantly outperforms other existing baselines, and is applicable to scenarios with extremely limited real images.
Ruoqi Sun, Xinge Zhu, Chongruo Wu, Chen Huang 0001, Jianping Shi, Lizhuang Ma
CVPR5
2019 DrivingStereo: A Large-Scale Dataset for Stereo Matching in Autonomous Driving Scenarios
abstract
Great progress has been made on estimating disparity maps from stereo images. However, with the limited stereo data available in the existing datasets and unstable ranging precision of current stereo methods, industry-level stereo matching in autonomous driving remains challenging. In this paper, we construct a novel large-scale stereo dataset named DrivingStereo. It contains over 180k images covering a diverse set of driving scenarios, which is hundreds of times larger than the KITTI Stereo dataset. High-quality labels of disparity are produced by a model-guided filtering strategy from multi-frame LiDAR points. For better evaluations, we present two new metrics for stereo matching in the driving scenes, i.e. a distance-aware metric and a semantic-aware metric. Extensive experiments show that compared with the models trained on FlyingThings3D or Cityscapes, the models trained on our DrivingStereo achieve higher generalization accuracy in real-world driving scenes, while the proposed metrics better evaluate the stereo methods on all-range distances and across different classes. Our dataset and code are available at https://drivingstereo-dataset.github.io.
Guorun Yang, Chaoqin Huang, Zhidong Deng, Jianping Shi, Bolei Zhou
CVPR5
2019 Adapting Object Detectors via Selective Cross-Domain Alignment
abstract
State-of-the-art object detectors are usually trained on public datasets. They often face substantial difficulties when applied to a different domain, where the imaging condition differs significantly and the corresponding annotated data are unavailable (or expensive to acquire). A natural remedy is to adapt the model by aligning the image representations on both domains. This can be achieved, for example, by adversarial learning, and has been shown to be effective in tasks like image classification. However, we found that in object detection, the improvement obtained in this way is quite limited. An important reason is that conventional domain adaptation methods strive to align images as a whole, while object detection, by nature, focuses on local regions that may contain objects of interest. Motivated by this, we propose a novel approach to domain adaption for object detection to handle the issues in ``where to look'' and ``how to align''. Our key idea is to mine the discriminative regions, namely those that are directly pertinent to object detection, and focus on aligning them across both domains. Experiments show that the proposed method performs remarkably better than existing methods with about 4% ~ 6% improvement under various domain-shift scenarios while keeping good scalability.
Xinge Zhu, Jiangmiao Pang, Ceyuan Yang, Jianping Shi, Dahua Lin
CVPR4
2019 CamNet: Coarse-to-Fine Retrieval for Camera Re-Localization
abstract
Camera re-localization is an important but challenging task in applications like robotics and autonomous driving. Recently, retrieval-based methods have been considered as a promising direction as they can be easily generalized to novel scenes. Despite significant progress has been made, we observe that the performance bottleneck of previous methods actually lies in the retrieval module. These methods use the same features for both retrieval and relative pose regression tasks which have potential conflicts in learning. To this end, here we present a coarse-to-fine retrieval-based deep learning framework, which includes three steps, i.e., image-based coarse retrieval, pose-based fine retrieval and precise relative pose regression. With our carefully designed retrieval module, the relative pose regression task can be surprisingly simpler. We design novel retrieval losses with batch hard sampling criterion and two-stage retrieval to locate samples that adapt to the relative pose regression task. Extensive experiments show that our model (CamNet) outperforms the state-of-the-art methods by a large margin on both indoor and outdoor datasets.
Mingyu Ding, Zhe Wang 0006, Jiankai Sun, Jianping Shi, Ping Luo 0002
ICCV4
2019 Prior Guided Dropout for Robust Visual Localization in Dynamic Environments
abstract
Camera localization from monocular images has been a long-standing problem, but its robustness in dynamic environments is still not adequately addressed. Compared with classic geometric approaches, modern CNN-based methods (e.g. PoseNet) have manifested the reliability against illumination or viewpoint variations, but they still have the following limitations. First, foreground moving objects are not explicitly handled, which results in poor performance and instability in dynamic environments. Second, the output for each image is a point estimate without uncertainty quantification. In this paper, we propose a framework which can be generally applied to existing CNN-based pose regressors to improve their robustness in dynamic environments. The key idea is a prior guided dropout module coupled with a self-attention module which can guide CNNs to ignore foreground objects during both training and inference. Additionally, the dropout module enables the pose regressor to output multiple hypotheses from which the uncertainty of pose estimates can be quantified and leveraged in the following uncertainty-aware pose graph optimization to improve the robustness further. We achieve an average accuracy of 9.98m/3.63° on RobotCar dataset, which outperforms the state-of-the-art method by 62.97%/47.08%. The source code of our implementation is available at https://github.com/zju3dv/RVL-Dynamic.
Jianping Shi, Xiaowei Zhou 0001, Hujun Bao, Guofeng Zhang 0001
ICCV3
2019 Switchable Whitening for Deep Representation Learning
abstract
Normalization methods are essential components in convolutional neural networks (CNNs). They either standardize or whiten data using statistics estimated in predefined sets of pixels. Unlike existing works that design normalization techniques for specific tasks, we propose Switchable Whitening (SW), which provides a general form unifying different whitening methods as well as standardization methods. SW learns to switch among these operations in an end-to-end manner. It has several advantages. First, SW adaptively selects appropriate whitening or standardization statistics for different tasks (see Fig.1), making it well suited for a wide range of tasks without manual design. Second, by integrating benefits of different normalizers, SW shows consistent improvements over its counterparts in various challenging benchmarks. Third, SW serves as a useful tool for understanding the characteristics of whitening and standardization techniques. We show that SW outperforms other alternatives on image classification (CIFAR-10/100, ImageNet), semantic segmentation (ADE20K, Cityscapes), domain adaptation (GTA5, Cityscapes), and image style transfer (COCO). For example, without bells and whistles, we achieve state-of-the-art performance with 45.33% mIoU on the ADE20K dataset.
Xingang Pan, Xiaohang Zhan, Jianping Shi, Xiaoou Tang, Ping Luo 0002
ICCV3
2019 Depth Completion From Sparse LiDAR Data With Depth-Normal Constraints
abstract
Depth completion aims to recover dense depth maps from sparse depth measurements. It is of increasing importance for autonomous driving and draws increasing attention from the vision community. Most of the current competitive methods directly train a network to learn a mapping from sparse depth inputs to dense depth maps, which has difficulties in utilizing the 3D geometric constraints and handling the practical sensor noises. In this paper, to regularize the depth completion and improve the robustness against noise, we propose a unified CNN framework that 1) models the geometric constraints between depth and surface normal in a diffusion module and 2) predicts the confidence of sparse LiDAR measurements to mitigate the impact of noise. Specifically, our encoder-decoder backbone predicts the surface normal, coarse depth and confidence of LiDAR inputs simultaneously, which are subsequently inputted into our diffusion refinement module to obtain the final completion results. Extensive experiments on KITTI depth completion dataset and NYU-Depth-V2 dataset demonstrate that our method achieves state-of-the-art performance. Further ablation study and analysis give more insights into the proposed components and demonstrate the generalization capability and stability of our model.
Xinge Zhu, Jianping Shi, Guofeng Zhang 0001, Hujun Bao, Hongsheng Li 0001
ICCV3
2019 Robust Multi-Modality Multi-Object Tracking
abstract
Multi-sensor perception is crucial to ensure the reliability and accuracy in autonomous driving system, while multi-object tracking (MOT) improves that by tracing sequential movement of dynamic objects. Most current approaches for multi-sensor multi-object tracking are either lack of reliability by tightly relying on a single input source (e.g., center camera), or not accurate enough by fusing the results from multiple sensors in post processing without fully exploiting the inherent information. In this study, we design a generic sensor-agnostic multi-modality MOT framework (mmMOT), where each modality (i.e., sensors) is capable of performing its role independently to preserve reliability, and could further improving its accuracy through a novel multi-modality fusion module. Our mmMOT can be trained in an end-to-end manner, enables joint optimization for the base feature extractor of each modality and an adjacency estimator for cross modality. Our mmMOT also makes the first attempt to encode deep representation of point cloud in data association process in MOT. We conduct extensive experiments to evaluate the effectiveness of the proposed framework on the challenging KITTI benchmark and report state-of-the-art performance. Code and models are available at https://github.com/ZwwWayne/mmMOT.
Hui Zhou 0005, Shuyang Sun, Zhe Wang 0006, Jianping Shi, Chen Change Loy
ICCV5
2019 Fast Abnormal Event Detection
Cewu Lu, Jianping Shi, Jiaya Jia
Int. J. Comput. Vis.2
2019 Recognizing road from satellite images by structured neural network
Chongruo Wu, Yu Meng 0002, Jianping Shi, Dongmei Yan
Neurocomputing5
2019 ℛ 2-CNN: Fast Tiny Object Detection in Large-Scale Remote Sensing Images
abstract
Recently, the convolutional neural network has brought impressive improvements for object detection. However, detecting tiny objects in large-scale remote sensing images still remains challenging. First, the extreme large input size makes the existing object detection solutions too slow for practical use. Second, the massive and complex backgrounds cause serious false alarms. Moreover, the ultratiny objects increase the difficulty of accurate detection. To tackle these problems, we propose a unified and self-reinforced network called remote sensing region-based convolutional neural network ($\mathcal {R}^{2}$-CNN), composing of backbone Tiny-Net, intermediate global attention block, and final classifier and detector. Tiny-Net is a lightweight residual structure, which enables fast and powerful features extraction from inputs. Global attention block is built upon Tiny-Net to inhibit false positives. Classifier is then used to predict the existence of target in each patch, and detector is followed to locate them accurately if available. The classifier and detector are mutually reinforced with end-to-end training, which further speed up the process and avoid false alarms. Effectiveness of$\mathcal {R}^{2}$-CNN is validated on hundreds of GF-1 images and GF-2 images that are$18\,000 \times 18\,192$pixels, 2.0-m resolution, and$27\,620 \times 29\,200$pixels, 0.8-m resolution, respectively. Specifically, we can process a GF-1 image in 29.4 s on Titian X just with single thread. According to our knowledge, no previous solution can detect the tiny object on such huge remote sensing images gracefully. We believe that it is a significant step toward practical real-time remote sensing systems.
Jiangmiao Pang, Cong Li 0016, Jianping Shi, Zhi-hai Xu, Huajun Feng
IEEE Trans. Geosci. Remote. Sens.3
2018 Generative Adversarial Frontal View to Bird View Synthesis
abstract
Environment perception is an important task with great practical value and bird view is an essential part for creating panoramas of surrounding environment. Due to the large gap and severe deformation between the frontal view and bird view, generating a bird view image from a single frontal view is challenging. To tackle this problem, we propose the BridgeGAN, i.e., a novel generative model for bird view synthesis. First, an intermediate view, i.e., homography view, is introduced to bridge the large gap. Next, conditioned on the three views (frontal view, homography view and bird view) in our task, a multi-GAN based model is proposed to learn the challenging cross-view translation. Furthermore, to guarantee one-to-one cross-view correspondences and consistent cross-view feature representations, two consistency constraints are designed for our task. Extensive experiments conducted on a synthetic dataset have demonstrated that the images generated by our model are much better than those generated by existing methods, with more consistent global appearance and sharper details. Ablation studies and discussions show its reliability and robustness in some challenging cases.
Xinge Zhu, Zhichao Yin, Jianping Shi, Hongsheng Li 0001, Dahua Lin
3DV3
2018 Spatial as Deep: Spatial CNN for Traffic Scene Understanding
abstract
Convolutional neural networks (CNNs) are usually built by stacking convolutional operations layer-by-layer. Although CNN has shown strong capability to extract semantics from raw pixels, its capacity to capture spatial relationships of pixels across rows and columns of an image is not fully explored. These relationships are important to learn semantic objects with strong shape priors but weak appearance coherences, such as traffic lanes, which are often occluded or not even painted on the road surface as shown in Fig. 1 (a). In this paper, we propose Spatial CNN (SCNN), which generalizes traditional deep layer-by-layer convolutions to slice-by-slice convolutions within feature maps, thus enabling message passings between pixels across rows and columns in a layer. Such SCNN is particular suitable for long continuous shape structure or large objects, with strong spatial relationship but less appearance clues, such as traffic lanes, poles, and wall. We apply SCNN on a newly released very challenging traffic lane detection dataset and Cityscapse dataset. The results show that SCNN could learn the spatial relationship for structure output and significantly improves the performance. We show that SCNN outperforms the recurrent neural network (RNN) based ReNet and MRF+CNN (MRFNet) in the lane detection dataset by 8.7% and 4.6% respectively. Moreover, our SCNN won the 1st place on the TuSimple Benchmark Lane Detection Challenge, with an accuracy of 96.53%.
Xingang Pan, Jianping Shi, Ping Luo 0002, Xiaogang Wang 0001, Xiaoou Tang
AAAI2
2018 Context Encoding for Semantic Segmentation
abstract
Recent work has made significant progress in improving spatial resolution for pixelwise labeling with Fully Convolutional Network (FCN) framework by employing Dilated/Atrous convolution, utilizing multi-scale features and refining boundaries. In this paper, we explore the impact of global contextual information in semantic segmentation by introducing the Context Encoding Module, which captures the semantic context of scenes and selectively highlights class-dependent featuremaps. The proposed Context Encoding Module significantly improves semantic segmentation results with only marginal extra computation cost over FCN. Our approach has achieved new state-of-the-art results 51.7% mIoU on PASCAL-Context, 85.9% mIoU on PASCAL VOC 2012. Our single model achieves a final score of 0.5567 on ADE20K test set, which surpasses the winning entry of COCO-Place Challenge 2017. In addition, we also explore how the Context Encoding Module can improve the feature representation of relatively shallow networks for the image classification on CIFAR-10 dataset. Our 14 layer network has achieved an error rate of 3.45%, which is comparable with state-of-the-art approaches with over 10Ã- more layers. The source code for the complete system are publicly available1.
Hang Zhang 0005, Kristin J. Dana, Jianping Shi, Xiaogang Wang 0001, Ambrish Tyagi, Amit Agrawal 0002
CVPR3
2018 Low-Latency Video Semantic Segmentation
abstract
Recent years have seen remarkable progress in semantic segmentation. Yet, it remains a challenging task to apply segmentation techniques to video-based applications. Specifically, the high throughput of video streams, the sheer cost of running fully convolutional networks, together with the low-latency requirements in many real-world applications, e.g. autonomous driving, present a significant challenge to the design of the video segmentation framework. To tackle this combined challenge, we develop a framework for video semantic segmentation, which incorporates two novel components: (1) a feature propagation module that adaptively fuses features over time via spatially variant convolution, thus reducing the cost of per-frame computation: and (2) an adaptive scheduler that dynamically allocate computation based on accuracy prediction. Both components work together to ensure low latency while maintaining high segmentation quality. On both Cityscapes and CamVid, the proposed framework obtained competitive performance compared to the state of the art, while substantially reducing the latency, from 360 ms to 119 ms.
Yule Li, Jianping Shi, Dahua Lin
CVPR2
2018 Path Aggregation Network for Instance Segmentation
abstract
The way that information propagates in neural networks is of great importance. In this paper, we propose Path Aggregation Network (PANet) aiming at boosting information flow in proposal-based instance segmentation framework. Specifically, we enhance the entire feature hierarchy with accurate localization signals in lower layers by bottom-up path augmentation, which shortens the information path between lower layers and topmost feature. We present adaptive feature pooling, which links feature grid and all feature levels to make useful information in each level propagate directly to following proposal subnetworks. A complementary branch capturing different views for each proposal is created to further improve mask prediction. These improvements are simple to implement, with subtle extra computational overhead. Yet they are useful and make our PANet reach the 1st place in the COCO 2017 Challenge Instance Segmentation task and the 2nd place in Object Detection task without large-batch training. PANet is also state-of-the-art on MVD and Cityscapes.
Shu Liu 0005, Lu Qi 0001, Haifang Qin, Jianping Shi, Jiaya Jia
CVPR4
2018 Eliminating Background-Bias for Robust Person Re-Identification
abstract
Person re-identification is an important topic in intelligent surveillance and computer vision. It aims to accurately measure visual similarities between person images for determining whether two images correspond to the same person. State-of-the-art methods mainly utilize deep learning based approaches for learning visual features for describing person appearances. However, we observe that existing deep learning models are biased to capture too much relevance between background appearances of person images. We design a series of experiments with newly created datasets to validate the influence of background information. To solve the background bias problem, we propose a person-region guided pooling deep neural network based on human parsing maps to learn more discriminative person-part features, and propose to augment training data with person images with random background. Extensive experiments demonstrate the robustness and effectiveness of our proposed method.
Maoqing Tian, Shuai Yi, Hongsheng Li 0001, Xuesen Zhang, Jianping Shi, Xiaogang Wang 0001
CVPR6
2018 GeoNet: Unsupervised Learning of Dense Depth, Optical Flow and Camera Pose
abstract
We propose GeoNet, a jointly unsupervised learning framework for monocular depth, optical flow and egomotion estimation from videos. The three components are coupled by the nature of 3D scene geometry, jointly learned by our framework in an end-to-end manner. Specifically, geometric relationships are extracted over the predictions of individual modules and then combined as an image reconstruction loss, reasoning about static and dynamic scene parts separately. Furthermore, we propose an adaptive geometric consistency loss to increase robustness towards outliers and non-Lambertian regions, which resolves occlusions and texture ambiguities effectively. Experimentation on the KITTI driving dataset reveals that our scheme achieves state-of-the-art results in all of the three tasks, performing better than previously unsupervised methods and comparably with supervised ones.
Zhichao Yin, Jianping Shi
CVPR2
2018 Factorizable Net: An Efficient Subgraph-Based Framework for Scene Graph Generation
Yikang Li 0002, Wanli Ouyang, Bolei Zhou, Jianping Shi, Xiaogang Wang 0001
ECCV (1)4
2018 Two at Once: Enhancing Learning and Generalization Capacities via IBN-Net
Xingang Pan, Ping Luo 0002, Jianping Shi, Xiaoou Tang
ECCV (4)3
2018 Pose Guided Human Video Generation
Ceyuan Yang, Zhe Wang 0006, Xinge Zhu, Chen Huang 0001, Jianping Shi, Dahua Lin
ECCV (10)5
2018 SegStereo: Exploiting Semantic Information for Disparity Estimation
Guorun Yang, Hengshuang Zhao, Jianping Shi, Zhidong Deng, Jiaya Jia
ECCV (7)3
2018 ICNet for Real-Time Semantic Segmentation on High-Resolution Images
Hengshuang Zhao, Xiaojuan Qi 0001, Xiaoyong Shen, Jianping Shi, Jiaya Jia
ECCV (3)4
2018 PSANet: Point-wise Spatial Attention Network for Scene Parsing
Hengshuang Zhao, Yi Zhang 0039, Shu Liu 0005, Jianping Shi, Chen Change Loy, Dahua Lin, Jiaya Jia
ECCV (9)4
2018 Penalizing Top Performers: Conservative Loss for Semantic Segmentation Adaptation
Xinge Zhu, Hui Zhou 0005, Ceyuan Yang, Jianping Shi, Dahua Lin
ECCV (7)4
2018 StripNet: Towards Topology Consistent Strip Structure Segmentation
abstract
In this work, we propose to study a special semantic segmentation problem where the targets are long and continuous strip patterns. Strip patterns widely exist in medical images and natural photos, such as retinal layers in OCT images and lanes on the roads, and segmentation of them has practical significance. Traditional pixel-level segmentation methods largely ignore the structure prior of strip patterns and thus easily suffer from the topological inconformity problem, such as holes and isolated islands in segmentation results. To tackle this problem, we design a novel deep framework, StripNet, that leverages the strong end-to-end learning ability of CNNs to predict the structured outputs as a sequence of boundary locations of the target strips. Specifically, StripNet decomposes the original segmentation problem into more easily solved local boundary-regression problems, and takes account of the topological constraints on the predicted boundaries. Moreover, our framework adopts a coarse-to-fine strategy and uses carefully designed heatmaps for training the boundary localization network. We examine StripNet on two challenging strip pattern segmentation tasks, retinal layer segmentation and lane detection. Extensive experiments demonstrate that StripNet achieves excellent results and outperforms state-of-the-art methods in both tasks.
Guoxiang Qu, Zhe Wang 0006, Xing Dai, Jianping Shi, Junjun He, Xiulan Zhang, Yu Qiao 0001
ACM Multimedia5
2018 Towards Understanding Acceleration Tradeoff between Momentum and Asynchrony in Nonconvex Stochastic Optimization
abstract
Asynchronous momentum stochastic gradient descent algorithms (Async-MSGD) have been widely used in distributed machine learning, e.g., training large collaborative filtering systems and deep neural networks. Due to current technical limit, however, establishing convergence properties of Async-MSGD for these highly complicated nonoconvex problems is generally infeasible. Therefore, we propose to analyze the algorithm through a simpler but nontrivial nonconvex problems --- streaming PCA. This allows us to make progress toward understanding Aync-MSGD and gaining new insights for more general problems. Specifically, by exploiting the diffusion approximation of stochastic optimization, we establish the asymptotic rate of convergence of Async-MSGD for streaming PCA. Our results indicate a fundamental tradeoff between asynchrony and momentum: To ensure convergence and acceleration through asynchrony, we have to reduce the momentum (compared with Sync-MSGD). To the best of our knowledge, this is the first theoretical attempt on understanding Async-MSGD for distributed nonconvex stochastic optimization. Numerical experiments on both streaming PCA and training deep neural networks are provided to support our findings for Async-MSGD.
Jianping Shi, Enlu Zhou, Tuo Zhao
NeurIPS3
2018 Sequential Context Encoding for Duplicate Removal
abstract
Duplicate removal is a critical step to accomplish a reasonable amount of predictions in prevalent proposal-based object detection frameworks. Albeit simple and effective, most previous algorithms utilized a greedy process without making sufficient use of properties of input data. In this work, we design a new two-stage framework to effectively select the appropriate proposal candidate for each object. The first stage suppresses most of easy negative object proposals, while the second stage selects true positives in the reduced proposal set. These two stages share the same network structure, an encoder and a decoder formed as recurrent neural networks (RNN) with global attention and context gate. The encoder scans proposal candidates in a sequential manner to capture the global context information, which is then fed to the decoder to extract optimal proposals. In our extensive experiments, the proposed method outperforms other alternatives by a large margin.
Lu Qi 0001, Shu Liu 0005, Jianping Shi, Jiaya Jia
NeurIPS3
2018 FishNet: A Versatile Backbone for Image, Region, and Pixel Level Prediction
abstract
The basic principles in designing convolutional neural network (CNN) structures for predicting objects on different levels, e.g., image-level, region-level, and pixel-level, are diverging. Generally, network structures designed specifically for image classification are directly used as default backbone structure for other tasks including detection and segmentation, but there is seldom backbone structure designed under the consideration of unifying the advantages of networks designed for pixel-level or region-level predicting tasks, which may require very deep features with high resolution. Towards this goal, we design a fish-like network, called FishNet. In FishNet, the information of all resolutions is preserved and refined for the final task. Besides, we observe that existing works still cannot \emph{directly} propagate the gradient information from deep layers to shallow layers. Our design can better handle this problem. Extensive experiments have been conducted to demonstrate the remarkable performance of the FishNet. In particular, on ImageNet-1k, the accuracy of FishNet is able to surpass the performance of DenseNet and ResNet with fewer parameters. FishNet was applied as one of the modules in the winning entry of the COCO Detection 2018 challenge. The code is available at https://github.com/kevin-ssy/FishNet.
Shuyang Sun, Jiangmiao Pang, Jianping Shi, Shuai Yi, Wanli Ouyang
NeurIPS3
2018 Adaptive Batch Normalization for practical domain adaptation
Yanghao Li, Naiyan Wang, Jianping Shi, Jiaying Liu 0001
Pattern Recognit.3
2017 Face Parsing via Recurrent Propagation
Sifei Liu, Jianping Shi, Ming-Hsuan Yang 0001
BMVC2
2017 Pyramid Scene Parsing Network
abstract
Scene parsing is challenging for unrestricted open vocabulary and diverse scenes. In this paper, we exploit the capability of global context information by different-region-based context aggregation through our pyramid pooling module together with the proposed pyramid scene parsing network (PSPNet). Our global prior representation is effective to produce good quality results on the scene parsing task, while PSPNet provides a superior framework for pixel-level prediction. The proposed approach achieves state-of-the-art performance on various datasets. It came first in ImageNet scene parsing challenge 2016, PASCAL VOC 2012 benchmark and Cityscapes benchmark. A single PSPNet yields the new record of mIoU accuracy 85.4% on PASCAL VOC 2012 and accuracy 80.2% on Cityscapes.
Hengshuang Zhao, Jianping Shi, Xiaojuan Qi 0001, Xiaogang Wang 0001, Jiaya Jia
CVPR2
2017 Zoom-in-Net: Deep Mining Lesions for Diabetic Retinopathy Detection
Zhe Wang 0006, Yanxin Yin, Jianping Shi, Hongsheng Li 0001, Xiaogang Wang 0001
MICCAI (3)3
2017 Distinguishing Cloud and Snow in Satellite Images via Deep Convolutional Network
abstract
Cloud and snow detection has significant remote sensing applications, while they share similar low-level features due to their consistent color distributions and similar local texture patterns. Thus, accurately distinguishing cloud from snow in pixel level from satellite images is always a challenging task with traditional approaches. To solve this shortcoming, in this letter, we proposed a deep learning system to classify cloud and snow with fully convolutional neural networks in pixel level. Specifically, a specially designed fully convolutional network was introduced to learn deep patterns for cloud and snow detection from the multispectrum satellite images. Then, a multiscale prediction strategy was introduced to integrate the low-level spatial information and high-level semantic information simultaneously. Finally, a new and challenging cloud and snow data set was labeled manually to train and further evaluate the proposed method. Extensive experiments demonstrate that the proposed deep model outperforms the state-of-the-art methods greatly both in quantitative and qualitative performances.
Yongjie Zhan, Jianping Shi, Lele Yao
IEEE Geosci. Remote. Sens. Lett.3
2016 Multi-scale Patch Aggregation (MPA) for Simultaneous Detection and Segmentation
abstract
Aiming at simultaneous detection and segmentation (SD-S), we propose a proposal-free framework, which detect and segment object instances via mid-level patches. We design a unified trainable network on patches, which is followed by a fast and effective patch aggregation algorithm to infer object instances. Our method benefits from end-to-end training. Without object proposal generation, computation time can also be reduced. In experiments, our method yields results 62.1% and 61.8% in terms of mAPr on VOC2012 segmentation val and VOC2012 SDS val, which are state-of-the-art at the time of submission. We also report results on Microsoft COCO test-std/test-dev dataset in this paper.
Shu Liu 0005, Xiaojuan Qi 0001, Jianping Shi, Hong Zhang 0009, Jiaya Jia
CVPR3
2016 Augmented Feedback in Semantic Segmentation Under Image Level Supervision
Xiaojuan Qi 0001, Zhengzhe Liu, Jianping Shi, Hengshuang Zhao, Jiaya Jia
ECCV (8)3
2016 Hierarchical Image Saliency Detection on Extended CSSD
abstract
Complex structures commonly exist in natural images. When an image contains small-scale high-contrast patterns either in the background or foreground, saliency detection could be adversely affected, resulting erroneous and non-uniform saliency assignment. The issue forms a fundamental challenge for prior methods. We tackle it from a scale point of view and propose a multi-layer approach to analyze saliency cues. Different from varying patch sizes or downsizing images, we measure region-based scales. The final saliency values are inferred optimally combining all the saliency cues in different scales using hierarchical inference. Through our inference model, single-scale information is selected to obtain a saliency map. Our method improves detection quality on many images that cannot be handled well traditionally. We also construct an extended Complex Scene Saliency Dataset (ECSSD) to include complex but general natural images.
Jianping Shi, Qiong Yan, Li Xu 0001, Jiaya Jia
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 Just noticeable defocus blur detection and estimation
abstract
We tackle a fundamental problem to detect and estimate just noticeable blur (JNB) caused by defocus that spans a small number of pixels in images. This type of blur is common during photo taking. Although it is not strong, the slight edge blurriness contains informative clues related to depth. We found existing blur descriptors based on local information cannot distinguish this type of small blur reliably from unblurred structures. We propose a simple yet effective blur feature via sparse representation and image decomposition. It directly establishes correspondence between sparse edge representation and blur strength estimation. Extensive experiments manifest the generality and robustness of this feature.
Jianping Shi, Li Xu 0001, Jiaya Jia
CVPR1
2015 Semantic Segmentation with Object Clique Potential
abstract
We propose an object clique potential for semantic segmentation. Our object clique potential addresses the misclassified object-part issues arising in solutions based on fully-convolutional networks. Our object clique set, compared to that yielded from segment-proposal-based approaches, is with a significantly smaller size, making our method consume notably less computation. Regarding system design and model formation, our object clique potential can be regarded as a functional complement to local-appearance-based CRF models and works in synergy with these effective approaches for further performance improvement. Extensive experiments verify our method.
Xiaojuan Qi 0001, Jianping Shi, Shu Liu 0005, Renjie Liao 0001, Jiaya Jia
ICCV2
2015 Understanding and Diagnosing Visual Tracking Systems
abstract
Several benchmark datasets for visual tracking research have been created in recent years. Despite their usefulness, whether they are sufficient for understanding and diagnosing the strengths and weaknesses of different trackers remains questionable. To address this issue, we propose a framework by breaking a tracker down into five constituent parts, namely, motion model, feature extractor, observation model, model updater, and ensemble post-processor. We then conduct ablative experiments on each component to study how it affects the overall result. Surprisingly, our findings are discrepant with some common beliefs in the visual tracking research community. We find that the feature extractor plays the most important role in a tracker. On the other hand, although the observation model is the focus of many studies, we find that it often brings no significant improvement. Moreover, the motion model and model updater contain many details that could affect the result. Also, the ensemble post-processor can improve the result substantially when the constituent trackers have high diversity. Based on our findings, we put together some very elementary building blocks to give a basic tracker which is competitive in performance to the state-of-the-art trackers. We believe our framework can provide a solid baseline when conducting controlled experiments for visual tracking research.
Naiyan Wang, Jianping Shi, Dit-Yan Yeung, Jiaya Jia
ICCV2
2015 Break Ames room illusion: depth from general single images
abstract
Photos compress 3D visual data to 2D. However, it is still possible to infer depth information even without sophisticated object learning. We propose a solution based on small-scale defocus blur inherent in optical lens and tackle the estimation problem by proposing a non-parametric matching scheme for natural images. It incorporates a matching prior with our newly constructed edgelet dataset using a non-local scheme, and includes semantic depth order cues for physically based inference. Several applications are enabled on natural images, including geometry based rendering and editing.
Jianping Shi, Xin Tao 0001, Li Xu 0001, Jiaya Jia
ACM Trans. Graph.1
2014 Discriminative Blur Detection Features
abstract
Ubiquitous image blur brings out a practically important question - what are effective features to differentiate between blurred and unblurred image regions. We address it by studying a few blur feature representations in image gradient, Fourier domain, and data-driven local filters. Unlike previous methods, which are often based on restoration mechanisms, our features are constructed to enhance discriminative power and are adaptive to various blur scales in images. To avail evaluation, we build a new blur perception dataset containing thousands of images with labeled ground-truth. Our results are applied to several applications, including blur region segmentation, deblurring, and blur magnification.
Jianping Shi, Li Xu 0001, Jiaya Jia
CVPR1
2014 Scale Adaptive Dictionary Learning
abstract
Dictionary learning has been widely used in many image processing tasks. In most of these methods, the number of basis vectors is either set by experience or coarsely evaluated empirically. In this paper, we propose a new scale adaptive dictionary learning framework, which jointly estimates suitable scales and corresponding atoms in an adaptive fashion according to the training data, without the need of prior information. We design an atom counting function and develop a reliable numerical scheme to solve the challenging optimization problem. Extensive experiments on texture and video data sets demonstrate quantitatively and visually that our method can estimate the scale, without damaging the sparse reconstruction ability.
Cewu Lu, Jianping Shi, Jiaya Jia
IEEE Trans. Image Process.2
2013 Online Robust Dictionary Learning
abstract
Online dictionary learning is particularly useful for processing large-scale and dynamic data in computer vision. It, however, faces the major difficulty to incorporate robust functions, rather than the square data fitting term, to handle outliers in training data. In this paper, we propose a new online framework enabling the use of l1 sparse data fitting term in robust dictionary learning, notably enhancing the usability and practicality of this important technique. Extensive experiments have been carried out to validate our new framework.
Cewu Lu, Jianping Shi, Jiaya Jia
CVPR2
2013 Hierarchical Saliency Detection
abstract
When dealing with objects with complex structures, saliency detection confronts a critical problem - namely that detection accuracy could be adversely affected if salient foreground or background in an image contains small-scale high-contrast patterns. This issue is common in natural images and forms a fundamental challenge for prior methods. We tackle it from a scale point of view and propose a multi-layer approach to analyze saliency cues. The final saliency map is produced in a hierarchical model. Different from varying patch sizes or downsizing images, our scale-based region handling is by finding saliency values optimally in a tree model. Our approach improves saliency detection on many images that cannot be handled well traditionally. A new dataset is also constructed.
Qiong Yan, Li Xu 0001, Jianping Shi, Jiaya Jia
CVPR3
2013 Abnormal Event Detection at 150 FPS in MATLAB
abstract
Speedy abnormal event detection meets the growing demand to process an enormous number of surveillance videos. Based on inherent redundancy of video structures, we propose an efficient sparse combination learning framework. It achieves decent performance in the detection phase without compromising result quality. The short running time is guaranteed because the new method effectively turns the original complicated problem to one in which only a few costless small-scale least square optimization steps are involved. Our method reaches high detection rates on benchmark datasets at a speed of 140-150 frames per second on average when computing on an ordinary desktop PC using MATLAB.
Cewu Lu, Jianping Shi, Jiaya Jia
ICCV2
2013 CoDeL: A Human Co-detection and Labeling Framework
abstract
We propose a co-detection and labeling (CoDeL) framework to identify persons that contain self-consistent appearance in multiple images. Our CoDeL model builds upon the deformable part-based model to detect human hypotheses and exploits cross-image correspondence via a matching classifier. Relying on a Gaussian process, this matching classifier models the similarity of two hypotheses and efficiently captures the relative importance contributed by various visual features, reducing the adverse effect of scattered occlusion. Further, the detector and matching classifier together make our model fit into a semi-supervised co-training framework, which can get enhanced results with a small amount of labeled training data. Our CoDeL model achieves decent performance on existing and new benchmark datasets.
Jianping Shi, Renjie Liao 0001, Jiaya Jia
ICCV1
2013 SCMF: Sparse Covariance Matrix Factorization for Collaborative Filtering
Jianping Shi, Naiyan Wang, Dit-Yan Yeung, Irwin King, Jiaya Jia
IJCAI1
2011 A non-convex relaxation approach to sparse dictionary learning
abstract
Dictionary learning is a challenging theme in computer vision. The basic goal is to learn a sparse representation from an overcomplete basis set. Most existing approaches employ a convex relaxation scheme to tackle this challenge due to the strong ability of convexity in computation and theoretical analysis. In this paper we propose a non-convex online approach for dictionary learning. To achieve the sparseness, our approach treats a so-called minimax concave (MC) penalty as a nonconvex relaxation of the ℓ0penalty. This treatment expects to obtain a more robust and sparse representation than existing convex approaches. In addition, we employ an online algorithm to adaptively learn the dictionary, which makes the non-convex formulation computationally feasible. Experimental results on the sparseness comparison and the applications in image denoising and image inpainting demonstrate that our approach is more effective and flexible.
Jianping Shi, Guang Dai, Jingdong Wang 0001
CVPR1
1999 Smart Avatars in JackMOO
abstract
Creation of compelling 3-dimensional, multi-user virtual worlds for education and training applications requires a high degree of realism in the appearance, interaction, and behavior of avatars within the scene. Our goal is to develop and/or adapt existing 3-dimensional technologies to provide training scenarios across the Internet in a form as close as possible to the appearance and interaction expected of live situations with human participants. We have produced a prototype system, JackMOO, which combines Jack, a virtual human system, and LambdaMOO, a multiuser; network-accessible, programmable, interactive server: Jack provides the visual realization of avatars and other objects. LambdaMOO provides the Web-accessible communication, programmability, and persistent object database. The combined JackMOO allows us to store the richer semantic information necessitated by the scope and range of human actions that an avatar must portray, and to express those actions in the form of imperative sentences. We describe JackMOO, its components, and a prototype application with five virtual human agents.
Jianping Shi, Thomas J. Smith, John P. Granieri, Norman I. Badler
VR1