Zhimin Cao

dblp:46/8653 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
1since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-authorArtificial intelligence and machine learning · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Image recognition and object detection · 51% Face, body and person analysis · 41% Representation and self-supervised learning · 4%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis
face recognition
0.322014
Learning Compact Face Representation: Packing a Face into an int32 · ACM Multimedia 2014
Face recognition with learning-based descriptor · CVPR 2010
Computer vision › Image recognition and object detection › object detection
multi-scale object detection
0.312017
FastMask: Segment Multi-scale Object Candidates in One Shot · CVPR 2017
Computer vision › Image recognition and object detection
pedestrian detection
0.312017
What Can Help Pedestrian Detection? · CVPR 2017
Computer vision › Image recognition and object detection › object detection › object proposal generation
segment proposal
0.312017
FastMask: Segment Multi-scale Object Candidates in One Shot · CVPR 2017
Computer vision › Image recognition and object detection › object detection
bounding box regression
0.212016
UnitBox: An Advanced Object Detection Network · ACM Multimedia 2016
Computer vision › Face, body and person analysis
face detection
0.212016
UnitBox: An Advanced Object Detection Network · ACM Multimedia 2016
Computer vision › Image recognition and object detection
object detection
0.212016
UnitBox: An Advanced Object Detection Network · ACM Multimedia 2016
Computer vision › Face, body and person analysis › face modeling
face hallucination
0.212015
Learning Face Hallucination in the Wild · AAAI 2015
Computer vision › Face, body and person analysis › face recognition
face representation
0.212015
Learning Face Hallucination in the Wild · AAAI 2015
Image and video processing › super-resolution › image super-resolution
face super-resolution
0.212015
Learning Face Hallucination in the Wild · AAAI 2015
Image and video processing
image restoration
0.212015
Learning Face Hallucination in the Wild · AAAI 2015
Image and video processing
super-resolution
0.212015
Learning Face Hallucination in the Wild · AAAI 2015
Machine learning › Representation and self-supervised learning › hierarchical representation › hierarchical representation learning
hierarchical feature learning
0.112017
FastMask: Segment Multi-scale Object Candidates in One Shot · CVPR 2017
Machine learning › Deep learning architectures and training › convolutional neural network › convolutional neural network architecture
fully convolutional network
0.112016
UnitBox: An Advanced Object Detection Network · ACM Multimedia 2016
Machine learning › Efficient and distributed learning
model compression
0.112014
Learning Compact Face Representation: Packing a Face into an int32 · ACM Multimedia 2014

Methods — techniques the papers use, named apart from their topics

deep convolutional network · 0.4bi-channel convolutional neural network · 0.4local homography transformation · 0.3weight-shared residual neck · 0.3scale-tolerant attention · 0.3multi-task training · 0.3feature aggregation · 0.3iou loss · 0.2fully convolutional network · 0.2deep learning frameworks · 0.2
YearPublicationVenuePosition
2026 A Clustering Algorithm Based on Dual-Perspective Adaptive Conical Information Granules
Bin Yu 0012, Zhimin Cao, Yu Fu 0020
IEEE Trans. Knowl. Data Eng.2
2018 GridFace: Face Rectification via Learning Local Homography Transformations
Erjin Zhou, Zhimin Cao, Jian Sun 0001
ECCV (16)2
2017 FastMask: Segment Multi-scale Object Candidates in One Shot
abstract
Objects appear to scale differently in natural images. This fact requires methods dealing with object-centric tasks (e.g. object proposal) to have robust performance over variances in object scales. In the paper, we present a novel segment proposal framework, namely FastMask, which takes advantage of hierarchical features in deep convolutional neural networks to segment multi-scale objects in one shot. Innovatively, we adapt segment proposal network into three different functional components (body, neck and head). We further propose a weight-shared residual neck module as well as a scale-tolerant attentional head module for efficient one-shot inference. On MS COCO benchmark, the proposed FastMask outperforms all state-of-the-art segment proposal methods in average recall being 2~5 times faster. Moreover, with a slight trade-off in accuracy, FastMask can segment objects in near real time (~13 fps) with 800×600 resolution images, demonstrating its potential in practical applications. Our implementation is available on https://github.com/voidrank/FastMask.
Hexiang Hu, Shiyi Lan, Yuning Jiang 0001, Zhimin Cao, Fei Sha
CVPR4
2017 What Can Help Pedestrian Detection?
abstract
Aggregating extra features has been considered as an effective approach to boost traditional pedestrian detection methods. However, there is still a lack of studies on whether and how CNN-based pedestrian detectors can benefit from these extra features. The first contribution of this paper is exploring this issue by aggregating extra features into CNN-based pedestrian detection framework. Through extensive experiments, we evaluate the effects of different kinds of extra features quantitatively. Moreover, we propose a novel network architecture, namely HyperLearner, to jointly learn pedestrian detection as well as the given extra feature. By multi-task training, HyperLearner is able to utilize the information of given features and improve detection performance without extra inputs in inference. The experimental results on multiple pedestrian benchmarks validate the effectiveness of the proposed HyperLearner.
Jiayuan Mao, Tete Xiao, Yuning Jiang 0001, Zhimin Cao
CVPR4
2017 A hierarchical energy minimization method for building roof segmentation from airborne LiDAR data
Yanfeng Gu, Zhimin Cao, Limin Dong
Multim. Tools Appl.2
2016 Building extraction from stereo aerial images based on multi-layer line grouping with height constraint
abstract
This paper introduces a novel multi-layer line grouping method for perceptually building extraction from stereo aerial images. Nowadays, perceptual grouping algorithm for line features obtained from images has been widely investigated, but there are little attentions to be paid to building height information of the line segments applied in existing literature of edge grouping field. In order to enhance accuracy of edge grouping, line segments are grouped into multi-layer with the extracted height information of neighbor matching points as constraint in this paper. In each height layer, line segments are grouped under the guidance of prior knowledges about the boundary style of polygonal buildings, and iteratively verified by the aforementioned group rules or prior knowledges for correctness of the connection and optimal building extraction. Experimental results illustrate a precise building extraction which can be accomplished by the proposed method.
Lechuan Hao, Ye Zhang 0008, Zhimin Cao
IGARSS3
2016 UnitBox: An Advanced Object Detection Network
abstract
In present object detection systems, the deep convolutional neural networks (CNNs) are utilized to predict bounding boxes of object candidates, and have gained performance advantages over the traditional region proposal methods. However, existing deep CNN methods assume the object bounds to be four independent variables, which could be regressed by the l2 loss separately. Such an oversimplified assumption is contrary to the well-received observation, that those variables are correlated, resulting to less accurate localization. To address the issue, we firstly introduce a novel Intersection over Union (IoU) loss function for bounding box prediction, which regresses the four bounds of a predicted box as a whole unit. By taking the advantages of IoU loss and deep fully convolutional networks, the UnitBox is introduced, which performs accurate and efficient localization, shows robust to objects of varied shapes and scales, and converges fast. We apply UnitBox on face detection task and achieve the best performance among all published methods on the FDDB benchmark.
Yuning Jiang 0001, Zhangyang Wang, Zhimin Cao, Thomas S. Huang
ACM Multimedia4
2015 Learning Face Hallucination in the Wild
abstract
Face hallucination method is proposed to generate high-resolution images from low-resolution ones for better visualization. However, conventional hallucination methods are often designed for controlled settings and cannot handle varying conditions of pose, resolution degree, and blur. In this paper, we present a new method of face hallucination, which can consistently improve the resolution of face images even with large appearance variations. Our method is based on a novel network architecture called Bi-channel Convolutional Neural Network (Bi-channel CNN). It extracts robust face representations from raw input by using deep convolutional network, then adaptively integrates two channels of information (the raw input image and face representations) to predict the high-resolution image. Experimental results show our system outperforms the prior state-of-the-art methods.
Erjin Zhou, Haoqiang Fan, Zhimin Cao, Yuning Jiang 0001, Qi Yin
AAAI3
2015 Building LiDAR point cloud denoising processing through sparse representation
abstract
Nowdays, airborne LiDAR comes into a popular way to survey the ground scene, particularly for the application of building reconstruction. However, the LiDAR point cloud acquired is usually polluted by noise for the existence of LiDAR system's inherent error and aircraft's shock. Thus, before LiDAR data is used, a preprocessing such as denoising is needed. This paper focus on the denoising of building LiDAR data. First, the building LiDAR point cloud is rasterized into a two- dimensional image. Then, a dictionary learned from training samples is used to denoise the image according to signal's sparse representation theory. Last, we can get the building's raster image with little noise.
Bingqian Xie, Yanfeng Gu, Zhimin Cao
IGARSS3
2014 Learning Compact Face Representation: Packing a Face into an int32
abstract
This paper addresses the problem of producing very compact representation of a face image for large-scale face search and analysis tasks. In tradition, the compactness of face representation is achieved by a dimension reduction step after representation extraction. However, the dimension reduction usually degrades the discriminative ability of the original representation drastically. In this paper, we present a deep learning framework which optimizes the compactness and discriminative ability jointly. The learnt representation can be as compact as 32 bit (same as the int32) and still produce highly discriminative performance (91.4% on LFW benchmark). Based on the extreme compactness, we show that traditional face analysis tasks (e.g. gender analysis) can be effectively solved by a Look-Up-Table approach given a large-scale face data set.
Haoqiang Fan, Mu Yang, Zhimin Cao, Yuning Jiang 0001, Qi Yin
ACM Multimedia3
2014 Three-Dimensional Reconstruction of Multiplatform Stereo Data With Variance Component Estimation
abstract
In this paper, we address a problem of 3-D reconstruction with generalized stereo data from multiple platforms of remote sensing. Nowadays, rational function model (RFM)-based 3-D reconstruction with stereo images obtained from a single platform of remote sensing like a satellite or an airborne platform has been widely investigated, but there are little attentions to be paid to the problem of 3-D reconstruction with stereo images from multiple platforms in the existing literature. In order to make full use of the generalized stereo images from different platforms with different rigorous sensor models for 3-D reconstruction, we need to form the least squares estimation model of the corresponding RFM-based forward-intersection task after collecting observations from different platforms. However, resolutions of the stereo images from different platforms are greatly different so that the observations in the corresponding least squares problem are mathematically seriously unbalanced. To solve this problem for achieving precise reconstruction, we first model how the spatial resolution of the observation images of different platforms changes pixel by pixel and then embed the variance-component-estimation technique into the RFM-based 3-D reconstruction procedure to adaptively adjust weights for different observations. Experiments are conducted on simulated and real data sets. Experimental results show that the proposed algorithm can efficiently fulfill the 3-D reconstruction task for multiplatform stereo images with noticeable improvement over the classical RFM-based 3-D reconstruction method in terms of precision.
Yanfeng Gu, Zhimin Cao, Ye Zhang 0008
IEEE Trans. Geosci. Remote. Sens.2
2011 MUlti information based Ground Control Points selection method
abstract
Ground Control Points (GCPs) are one of the most important data used in many fields of Remote Sensing. The number and distribution of GCPs are always the key factors for the success of some researches. A GCPs selection method by integrating the three dimensional spatial information (i.e. the earth coordinates (X, Y, Z) ) and the corresponding feature information underlying the data itself was proposed. To testify the performance of this method, a Rational Function Model Resolving experiment is conducted. Experiment results show that the accuracy and time-consuming performance are both improved using GCPs selected by the proposed method.
Yanfeng Gu, Zhimin Cao, Ye Zhang 0008, Xiangrong Zhang
IGARSS2
2010 Face recognition with learning-based descriptor
abstract
We present a novel approach to address the representation issue and the matching issue in face recognition (verification). Firstly, our approach encodes the micro-structures of the face by a new learning-based encoding method. Unlike many previous manually designed encoding methods (e.g., LBP or SIFT), we use unsupervised learning techniques to learn an encoder from the training examples, which can automatically achieve very good tradeoff between discriminative power and invariance. Then we apply PCA to get a compact face descriptor. We find that a simple normalization mechanism after PCA can further improve the discriminative ability of the descriptor. The resulting face representation, learning-based (LE) descriptor, is compact, highly discriminative, and easy-to-extract. To handle the large pose variation in real-life scenarios, we propose a pose-adaptive matching method that uses pose-specific classifiers to deal with different pose combinations (e.g., frontal v.s. frontal, frontal v.s. left) of the matching face pair. Our approach is comparable with the state-of-the-art methods on the Labeled Face in Wild (LFW) benchmark (we achieved 84.45% recognition rate), while maintaining excellent compactness, simplicity, and generalization aability across different datasets.
Zhimin Cao, Qi Yin, Xiaoou Tang, Jian Sun 0001
CVPR1
2010 A Motion Estimation algorithm based on Markov Chain Model
abstract
In this paper, we present a new fast Motion Estimation (ME) algorithm based on Markov Chain Model (MEMCM). Spatial-temporal correlation of video sequence and Markov Chain Model are used in the algorithm, which considers the continuity of video. Experimental simulation shows that the proposed algorithm has better performance in both speed-up and PSNR than PMVFAST.
Zhijie Zhao, Zhimin Cao, Maoliu Lin, Xuesong Jin
ICASSP2