Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xiaoou Tang

dblp:04/5226 · DBLP profile ↗
← Back
355ranked-venue papers
18as first author
3since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 282 · 14 first-authorArtificial intelligence and machine learning · 237 · 5 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-authorSecurity and privacy · 4Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
206 papers
Face, body and person analysis · 26% Video understanding and tracking · 11% Representation and self-supervised learning · 10%
Computer graphics and multimedia
99 papers
Image and video processing · 45% Visual content generation and editing · 14% Geometric modeling and processing · 14%
Databases, data mining, and information retrieval
35 papers
Information retrieval · 58% Data mining · 41% Recommender systems · 1%

Topics — the 30 heaviest of 441, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis
face recognition
5.5512020
Deep Imbalanced Learning for Face Recognition and Attribute Prediction · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Pose-Robust Face Recognition via Deep Residual Equivariant Mapping · CVPR 2018
Hybrid Deep Learning for Face Verification · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Image and video processing
image restoration
2.1112022
Path-Restore: Learning Network Path Selection for Image Restoration · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Deep Network Interpolation for Continuous Imagery Effect Transition · CVPR 2019
Accelerating the Super-Resolution Convolutional Neural Network · ECCV (2) 2016
Machine learning › Generative modeling
generative adversarial network
1.742022
InterFaceGAN: Interpreting the Disentangled Face Representation Learned by GANs · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Interpreting the Latent Space of GANs for Semantic Face Editing · CVPR 2020
Aesthetic-Driven Image Enhancement by Adversarial Learning · ACM Multimedia 2018
Computer vision › Video understanding and tracking
action recognition
1.462019
Temporal Segment Networks for Action Recognition in Videos · IEEE Trans. Pattern Anal. Mach. Intell. 2019
MoFAP: A Multi-level Representation for Action Recognition · Int. J. Comput. Vis. 2016
Temporal Segment Networks: Towards Good Practices for Deep Action Recognition · ECCV (8) 2016
Computer vision › Segmentation and scene understanding
semantic segmentation
1.462019
Deep Learning Markov Random Field for Semantic Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Mix-and-Match Tuning for Self-Supervised Semantic Segmentation · AAAI 2018
Not All Pixels Are Equal: Difficulty-Aware Semantic Segmentation via Deep Layer Cascade · CVPR 2017
Computer vision › Face, body and person analysis
face alignment
1.482016
Learning Deep Representation for Face Alignment with Auxiliary Attributes · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Unconstrained Face Alignment via Cascaded Compositional Learning · CVPR 2016
Face alignment by coarse-to-fine shape searching · CVPR 2015
Computer vision › Face, body and person analysis › face recognition
face verification
1.382016
Hybrid Deep Learning for Face Verification · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Deeply learned face representations are sparse, selective, and robust · CVPR 2015
Surpassing Human-Level Face Verification Performance on LFW with GaussianFace · AAAI 2015
Computer vision › Face, body and person analysis
face detection
1.262018
Faceness-Net: Face Detection through Deep Facial Part Responses · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Recurrent Scale Approximation for Object Detection in CNN · ICCV 2017
WIDER FACE: A Face Detection Benchmark · CVPR 2016
Computer vision › Image recognition and object detection
object detection
1.262017
DeepID-Net: Object Detection with Deformable Part Based Convolutional Neural Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2017
Recurrent Scale Approximation for Object Detection in CNN · ICCV 2017
DeepID-Net: Deformable deep convolutional neural networks for object detection · CVPR 2015
Image and video processing › super-resolution
image super-resolution
1.182016
Image Super-Resolution Using Deep Convolutional Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Accelerating the Super-Resolution Convolutional Neural Network · ECCV (2) 2016
Learning a Deep Convolutional Network for Image Super-Resolution · ECCV (4) 2014
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
1.022022
InterFaceGAN: Interpreting the Disentangled Face Representation Learned by GANs · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Interpreting the Latent Space of GANs for Semantic Face Editing · CVPR 2020
Machine learning › Generative modeling
latent space interpretation
1.022022
InterFaceGAN: Interpreting the Disentangled Face Representation Learned by GANs · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Interpreting the Latent Space of GANs for Semantic Face Editing · CVPR 2020
Machine learning › Deep learning architectures and training
convolutional neural network
1.062018
Spindle Net: Person Re-identification with Human Body Region Guided Feature Decomposition and Fusion · CVPR 2017
Image Super-Resolution Using Deep Convolutional Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Learning a Deep Convolutional Network for Image Super-Resolution · ECCV (4) 2014
Information retrieval
image retrieval
0.992014
Web Image Re-Ranking UsingQuery-Specific Semantic Signatures · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Anchor concept graph distance for web image re-ranking · ACM Multimedia 2013
IntentSearch: Capturing User Intention for One-Click Internet Image Search · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Computer vision › 3D vision › motion estimation
optical flow
0.932021
A Lightweight Optical Flow CNN - Revisiting Data Fidelity and Regularization · IEEE Trans. Pattern Anal. Mach. Intell. 2021
LiteFlowNet: A Lightweight Convolutional Neural Network for Optical Flow Estimation · CVPR 2018
Video Frame Synthesis Using Deep Voxel Flow · ICCV 2017
Computer vision › Video understanding and tracking › action detection
temporal action localization
0.932020
Temporal Action Detection with Structured Segment Networks · Int. J. Comput. Vis. 2020
Temporal Action Detection with Structured Segment Networks · ICCV 2017
Mining Motion Atoms and Phrases for Complex Action Recognition · ICCV 2013
Geometric modeling and processing › 3d reconstruction
3d reconstruction from line drawings
0.892013
Complex 3D General Object Reconstruction from Line Drawings · ICCV 2013
Decomposition of Complex Line Drawings with Hidden Lines for 3D Planar-Faced Manifold Object Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Responses to the Comments on "What the Back of the Object Looks Like: 3D Reconstruction from Line Drawings without Hidden Lines" · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Robotics › Robot navigation and mapping › landmark detection
fashion landmark detection
0.832017
Unconstrained Fashion Landmark Detection via Hierarchical Recurrent Transformer Networks · ACM Multimedia 2017
Fashion Landmark Detection in the Wild · ECCV (2) 2016
DeepFashion: Powering Robust Clothes Recognition and Retrieval with Rich Annotations · CVPR 2016
Data mining
clustering
0.882012
Graph Degree Linkage: Agglomerative Clustering on a Directed Graph · ECCV (1) 2012
Unsupervised Object Segmentation with a Hybrid Graph Model (HGM) · IEEE Trans. Pattern Anal. Mach. Intell. 2010
Isoperimetric cut on a directed graph · CVPR 2010
Computer vision › Image recognition and object detection
image classification
0.852019
Residual Attention Network for Image Classification · CVPR 2017
Pairwise Rotation Invariant Co-Occurrence Local Binary Pattern · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Switchable Whitening for Deep Representation Learning · ICCV 2019
Geometric modeling and processing
3d reconstruction
0.882013
Complex 3D General Object Reconstruction from Line Drawings · ICCV 2013
3-D Modeling From a Single View of a Symmetric Object · IEEE Trans. Image Process. 2012
Decomposition of Complex Line Drawings with Hidden Lines for 3D Planar-Faced Manifold Object Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
subspace learning
0.7122010
Formulating Face Verification With Semidefinite Programming · IEEE Trans. Image Process. 2007
Using Support Vector Machines to Enhance the Performance of Bayesian Face Recognition · IEEE Trans. Inf. Forensics Secur. 2007
A Convengent Solution to Tensor Subspace Learning · IJCAI 2007
Machine learning › Efficient and distributed learning
model compression
0.732022
Sparsifying Neural Network Connections for Face Recognition · CVPR 2016
Face Model Compression by Distilling Knowledge from Neurons · AAAI 2016
Path-Restore: Learning Network Path Selection for Image Restoration · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Machine learning › Learning paradigms
multi-task learning
0.732016
Learning Deep Representation for Face Alignment with Auxiliary Attributes · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Pedestrian detection aided by deep learning semantic tasks · CVPR 2015
Facial Landmark Detection by Deep Multi-task Learning · ECCV (6) 2014
Computer vision › Face, body and person analysis › facial attribute analysis
facial attribute recognition
0.722020
Deep Imbalanced Learning for Face Recognition and Attribute Prediction · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Deep Learning Face Attributes in the Wild · ICCV 2015
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
markov random field
0.632018
Deep Learning Markov Random Field for Semantic Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2018
Deep Markov Random Field for Image Modeling · ECCV (8) 2016
Accurate Face Alignment using Shape Constrained Markov Network · CVPR (1) 2006
Computer vision › Image recognition and object detection
pedestrian detection
0.632015
Deep Learning Strong Parts for Pedestrian Detection · ICCV 2015
Pedestrian detection aided by deep learning semantic tasks · CVPR 2015
Switchable Deep Network for Pedestrian Detection · CVPR 2014
Computer vision › Face, body and person analysis › face recognition › face representation
face representation learning
0.632016
Hybrid Deep Learning for Face Verification · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Deep Learning Face Representation by Joint Identification-Verification · NIPS 2014
Deep Learning Identity-Preserving Face Space · ICCV 2013
Computer vision › 3D vision
3d reconstruction
0.652012
Example-based 3D object reconstruction from line drawings · CVPR 2012
Synthesizing oil painting surface geometry from a single photograph · CVPR 2012
Symmetric piecewise planar object reconstruction from a single image · CVPR 2011
Machine learning › Trustworthy machine learning
interpretability
0.612022
InterFaceGAN: Interpreting the Disentangled Face Representation Learned by GANs · IEEE Trans. Pattern Anal. Mach. Intell. 2022

Methods — techniques the papers use, named apart from their topics

convolutional neural network · 4.5deep learning · 1.6markov random field · 1.4deep convolutional network · 1.2reinforcement learning · 1.1policy mask · 1.1difficulty-regulated reward · 1.1subspace projection · 1.0feature warping · 0.8GAN inversion · 0.6sparse coding · 0.5belief propagation · 0.4linear interpolation · 0.4deep network interpolation · 0.4semidefinite programming · 0.3kernel density estimation · 0.3keyword expansion · 0.3embedding learning · 0.2
YearPublicationVenuePosition
2022 InterFaceGAN: Interpreting the Disentangled Face Representation Learned by GANs
abstract
Although generative adversarial networks (GANs) have made significant progress in face synthesis, there lacks enough understanding of what GANs have learned in the latent representation to map a random code to a photo-realistic image. In this work, we propose a framework called InterFaceGAN to interpret the disentangled face representation learned by the state-of-the-art GAN models and study the properties of the facial semantics encoded in the latent space. We first find that GANs learn various semantics in some linear subspaces of the latent space. After identifying these subspaces, we can realistically manipulate the corresponding facial attributes without retraining the model. We then conduct a detailed study on the correlation between different semantics and manage to better disentangle them via subspace projection, resulting in more precise control of the attribute manipulation. Besides manipulating the gender, age, expression, and presence of eyeglasses, we can even alter the face pose and fix the artifacts accidentally made by GANs. Furthermore, we perform an in-depth face identity analysis and a layer-wise analysis to evaluate the editing results quantitatively. Finally, we apply our approach to real face editing by employing GAN inversion approaches and explicitly training feed-forward models based on the synthetic data established by InterFaceGAN. Extensive experimental results suggest that learning to synthesize faces spontaneously brings a disentangled and controllable face representation.
Yujun Shen, Ceyuan Yang, Xiaoou Tang, Bolei Zhou
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Path-Restore: Learning Network Path Selection for Image Restoration
abstract
Very deep Convolutional Neural Networks (CNNs) have greatly improved the performance on various image restoration tasks. However, this comes at a price of increasing computational burden, hence limiting their practical usages. We observe that some corrupted image regions are inherently easier to restore than others since the distortion and content vary within an image. To leverage this, we propose Path-Restore, a multi-path CNN with a pathfinder that can dynamically select an appropriate route for each image region. We train the pathfinder using reinforcement learning with a difficulty-regulated reward. This reward is related to the performance, complexity and "the difficulty of restoring a region". A policy mask is further investigated to jointly process all the image regions. We conduct experiments on denoising and mixed restoration tasks. The results show that our method achieves comparable or superior performance to existing approaches with less computational cost. In particular, Path-Restore is effective for real-world denoising, where the noise distribution varies across different regions on a single image. Compared to the state-of-the-art RIDNet [1], our method achieves comparable performance and runs 2.7x faster on the realistic Darmstadt Noise Dataset [2]. Models and codes are available on the project page: https://www.mmlab-ntu.com/project/pathrestore/.
Ke Yu 0003, Xintao Wang 0002, Chao Dong 0005, Xiaoou Tang, Chen Change Loy
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 A Lightweight Optical Flow CNN - Revisiting Data Fidelity and Regularization
abstract
Over four decades, the majority addresses the problem of optical flow estimation using variational methods. With the advance of machine learning, some recent works have attempted to address the problem using convolutional neural network (CNN) and have showed promising results. FlowNet2 [1] , the state-of-the-art CNN, requires over 160M parameters to achieve accurate flow estimation. Our LiteFlowNet2 outperforms FlowNet2 on Sintel and KITTI benchmarks, while being 25.3 times smaller in the model size and 3.1 times faster in the running speed. LiteFlowNet2 is built on the foundation laid by conventional methods and resembles the corresponding roles as data fidelity and regularization in variational methods. We compute optical flow in a spatial-pyramid formulation as SPyNet [2] but through a novel lightweight cascaded flow inference. It provides high flow estimation accuracy through early correction with seamless incorporation of descriptor matching. Flow regularization is used to ameliorate the issue of outliers and vague flow boundaries through feature-driven local convolutions. Our network also owns an effective structure for pyramidal feature extraction and embraces feature warping rather than image warping as practiced in FlowNet2 and SPyNet. Comparing to LiteFlowNet [3] , LiteFlowNet2 improves the optical flow accuracy on Sintel Clean by 23.3 percent, Sintel Final by 12.8 percent, KITTI 2012 by 19.6 percent, and KITTI 2015 by 18.8 percent, while being 2.2 times faster. Our network protocol and trained models are made publicly available on https://github.com/twhui/LiteFlowNet2.
Tak-Wai Hui, Xiaoou Tang, Chen Change Loy
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Interpreting the Latent Space of GANs for Semantic Face Editing
abstract
Despite the recent advance of Generative Adversarial Networks (GANs) in high-fidelity image synthesis, there lacks enough understanding of how GANs are able to map a latent code sampled from a random distribution to a photo-realistic image. Previous work assumes the latent space learned by GANs follows a distributed representation but observes the vector arithmetic phenomenon. In this work, we propose a novel framework, called InterFaceGAN, for semantic face editing by interpreting the latent semantics learned by GANs. In this framework, we conduct a detailed study on how different semantics are encoded in the latent space of GANs for face synthesis. We find that the latent code of well-trained generative models actually learns a disentangled representation after linear transformations. We explore the disentanglement between various semantics and manage to decouple some entangled semantics with subspace projection, leading to more precise control of facial attributes. Besides manipulating gender, age, expression, and the presence of eyeglasses, we can even vary the face pose as well as fix the artifacts accidentally generated by GAN models. The proposed method is further applied to achieve real image manipulation when combined with GAN inversion methods or some encoder-involved models. Extensive results suggest that learning to synthesize faces spontaneously brings a disentangled and controllable facial attribute representation.
Yujun Shen, Jinjin Gu, Xiaoou Tang, Bolei Zhou
CVPR3
2020 Temporal Action Detection with Structured Segment Networks
Yue Zhao 0006, Yuanjun Xiong, Limin Wang 0002, Zhirong Wu, Xiaoou Tang, Dahua Lin
Int. J. Comput. Vis.5
2020 Deep Imbalanced Learning for Face Recognition and Attribute Prediction
abstract
Data for face analysis often exhibit highly-skewed class distribution, i.e., most data belong to a few majority classes, while the minority classes only contain a scarce amount of instances. To mitigate this issue, contemporary deep learning methods typically follow classic strategies such as class re-sampling or cost-sensitive training. In this paper, we conduct extensive and systematic experiments to validate the effectiveness of these classic schemes for representation learning on class-imbalanced data. We further demonstrate that more discriminative deep representation can be learned by enforcing a deep network to maintain inter-cluster margins both within and between classes. This tight constraint effectively reduces the class imbalance inherent in the local data neighborhood, thus carving much more balanced class boundaries locally. We show that it is easy to deploy angular margins between the cluster distributions on a hypersphere manifold. Such learned Cluster-based Large Margin Local Embedding (CLMLE), when combined with a simple k-nearest cluster algorithm, shows significant improvements in accuracy over existing methods on both face recognition and face attribute prediction tasks that exhibit imbalanced class distribution.
Chen Huang 0001, Chen Change Loy, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 DeepFashion2: A Versatile Benchmark for Detection, Pose Estimation, Segmentation and Re-Identification of Clothing Images
abstract
Understanding fashion images has been advanced by benchmarks with rich annotations such as DeepFashion, whose labels include clothing categories, landmarks, and consumer-commercial image pairs. However, DeepFashion has nonnegligible issues such as single clothing-item per image, sparse landmarks (4∼8 only), and no per-pixel masks, making it had significant gap from real-world scenarios. We fill in the gap by presenting DeepFashion2 to address these issues. It is a versatile benchmark of four tasks including clothes detection, pose estimation, segmentation, and retrieval. It has 801K clothing items where each item has rich annotations such as style, scale, view- point, occlusion, bounding box, dense landmarks (e.g. 39 for ‘long sleeve outwear’ and 15 for ‘vest’), and masks. There are also 873K Commercial-Consumer clothes pairs. The annotations of DeepFashion2 are much larger than its counterparts such as 8× of FashionAI Global Challenge. A strong baseline is proposed, called Match R- CNN, which builds upon Mask R-CNN to solve the above four tasks in an end-to-end manner. Extensive evaluations are conducted with different criterions in Deep- Fashion2. DeepFashion2 Dataset will be released at : https://github.com/switchablenorms/DeepFashion2
Yuying Ge, Ruimao Zhang, Xiaogang Wang 0001, Xiaoou Tang, Ping Luo 0002
CVPR4
2019 Deep Network Interpolation for Continuous Imagery Effect Transition
abstract
Deep convolutional neural network has demonstrated its capability of learning a deterministic mapping for the desired imagery effect. However, the large variety of user flavors motivates the possibility of continuous transition among different output effects. Unlike existing methods that require a specific design to achieve one particular transition (e.g., style transfer), we propose a simple yet universal approach to attain a smooth control of diverse imagery effects in many low-level vision tasks, including image restoration, image-to-image translation, and style transfer. Specifically, our method, namely Deep Network Interpolation (DNI), applies linear interpolation in the parameter space of two or more correlated networks. A smooth control of imagery effects can be achieved by tweaking the interpolation coefficients. In addition to DNI and its broad applications, we also investigate the mechanism of network interpolation from the perspective of learned filters.
Xintao Wang 0002, Ke Yu 0003, Chao Dong 0005, Xiaoou Tang, Chen Change Loy
CVPR4
2019 Switchable Whitening for Deep Representation Learning
abstract
Normalization methods are essential components in convolutional neural networks (CNNs). They either standardize or whiten data using statistics estimated in predefined sets of pixels. Unlike existing works that design normalization techniques for specific tasks, we propose Switchable Whitening (SW), which provides a general form unifying different whitening methods as well as standardization methods. SW learns to switch among these operations in an end-to-end manner. It has several advantages. First, SW adaptively selects appropriate whitening or standardization statistics for different tasks (see Fig.1), making it well suited for a wide range of tasks without manual design. Second, by integrating benefits of different normalizers, SW shows consistent improvements over its counterparts in various challenging benchmarks. Third, SW serves as a useful tool for understanding the characteristics of whitening and standardization techniques. We show that SW outperforms other alternatives on image classification (CIFAR-10/100, ImageNet), semantic segmentation (ADE20K, Cityscapes), domain adaptation (GTA5, Cityscapes), and image style transfer (COCO). For example, without bells and whistles, we achieve state-of-the-art performance with 45.33% mIoU on the ADE20K dataset.
Xingang Pan, Xiaohang Zhan, Jianping Shi, Xiaoou Tang, Ping Luo 0002
ICCV4
2019 Temporal Segment Networks for Action Recognition in Videos
abstract
We present a general and flexible video-level framework for learning action models in videos. This method, called temporal segment network (TSN), aims to model long-range temporal structure with a new segment-based sampling and aggregation scheme. This unique design enables the TSN framework to efficiently learn action models by using the whole video. The learned models could be easily deployed for action recognition in both trimmed and untrimmed videos with simple average pooling and multi-scale temporal window integration, respectively. We also study a series of good practices for the implementation of the TSN framework given limited training samples. Our approach obtains the state-the-of-art performance on five challenging action recognition benchmarks: HMDB51 (71.0 percent), UCF101 (94.9 percent), THUMOS14 (80.1 percent), ActivityNet v1.2 (89.6 percent), and Kinetics400 (75.7 percent). In addition, using the proposed RGB difference as a simple motion representation, our method can still achieve competitive accuracy on UCF101 (91.0 percent) while running at 340 FPS. Furthermore, based on the proposed TSN framework, we won the video classification track at the ActivityNet challenge 2016 among 24 teams.
Limin Wang 0002, Yuanjun Xiong, Zhe Wang 0013, Yu Qiao 0001, Dahua Lin, Xiaoou Tang, Luc Van Gool
IEEE Trans. Pattern Anal. Mach. Intell.6
2018 Spatial as Deep: Spatial CNN for Traffic Scene Understanding
abstract
Convolutional neural networks (CNNs) are usually built by stacking convolutional operations layer-by-layer. Although CNN has shown strong capability to extract semantics from raw pixels, its capacity to capture spatial relationships of pixels across rows and columns of an image is not fully explored. These relationships are important to learn semantic objects with strong shape priors but weak appearance coherences, such as traffic lanes, which are often occluded or not even painted on the road surface as shown in Fig. 1 (a). In this paper, we propose Spatial CNN (SCNN), which generalizes traditional deep layer-by-layer convolutions to slice-by-slice convolutions within feature maps, thus enabling message passings between pixels across rows and columns in a layer. Such SCNN is particular suitable for long continuous shape structure or large objects, with strong spatial relationship but less appearance clues, such as traffic lanes, poles, and wall. We apply SCNN on a newly released very challenging traffic lane detection dataset and Cityscapse dataset. The results show that SCNN could learn the spatial relationship for structure output and significantly improves the performance. We show that SCNN outperforms the recurrent neural network (RNN) based ReNet and MRF+CNN (MRFNet) in the lane detection dataset by 8.7% and 4.6% respectively. Moreover, our SCNN won the 1st place on the TuSimple Benchmark Lane Detection Challenge, with an accuracy of 96.53%.
Xingang Pan, Jianping Shi, Ping Luo 0002, Xiaogang Wang 0001, Xiaoou Tang
AAAI5
2018 Mix-and-Match Tuning for Self-Supervised Semantic Segmentation
abstract
Deep convolutional networks for semantic image segmentation typically require large-scale labeled data, e.g., ImageNet and MS COCO, for network pre-training. To reduce annotation efforts, self-supervised semantic segmentation is recently proposed to pre-train a network without any human-provided labels. The key of this new form of learning is to design a proxy task (e.g., image colorization), from which a discriminative loss can be formulated on unlabeled data. Many proxy tasks, however, lack the critical supervision signals that could induce discriminative representation for the target image segmentation task. Thus self-supervision’s performance is still far from that of supervised pre-training. In this study, we overcome this limitation by incorporating a "mix-and-match" (M&M) tuning stage in the self-supervision pipeline. The proposed approach is readily pluggable to many self-supervision methods and does not use more annotated samples than the original process. Yet, it is capable of boosting the performance of target image segmentation task to surpass fully-supervised pre-trained counterpart. The improvement is made possible by better harnessing the limited pixel-wise annotations in the target dataset. Specifically, we first introduce the "mix" stage, which sparsely samples and mixes patches from the target set to reflect rich and diverse local patch statistics of target images. A ‘match’ stage then forms a class-wise connected graph, which can be used to derive a strong triplet-based discriminative loss for finetuning the network. Our paradigm follows the standard practice in existing self-supervised studies and no extra data or label is required. With the proposed M&M approach, for the first time, a self-supervision method can achieve comparable or even better performance compared to its ImageNet pretrained counterpart on both PASCAL VOC2012 dataset and CityScapes dataset.
Xiaohang Zhan, Ziwei Liu 0002, Ping Luo 0002, Xiaoou Tang, Chen Change Loy
AAAI4
2018 Pose-Robust Face Recognition via Deep Residual Equivariant Mapping
abstract
Face recognition achieves exceptional success thanks to the emergence of deep learning. However, many contemporary face recognition models still perform relatively poor in processing profile faces compared to frontal faces. A key reason is that the number of frontal and profile training faces are highly imbalanced - there are extensively more frontal training samples compared to profile ones. In addition, it is intrinsically hard to learn a deep representation that is geometrically invariant to large pose variations. In this study, we hypothesize that there is an inherent mapping between frontal and profile faces, and consequently, their discrepancy in the deep representation space can be bridged by an equivariant mapping. To exploit this mapping, we formulate a novel Deep Residual EquivAriant Mapping (DREAM) block, which is capable of adaptively adding residuals to the input deep representation to transform a profile face representation to a canonical pose that simplifies recognition. The DREAM block consistently enhances the performance of profile face recognition for many strong deep networks, including ResNet models, without deliberately augmenting training data of profile faces. The block is easy to use, light-weight, and can be implemented with a negligible computational overhead1.
Kaidi Cao, Yu Rong 0003, Cheng Li 0009, Xiaoou Tang, Chen Change Loy
CVPR4
2018 LiteFlowNet: A Lightweight Convolutional Neural Network for Optical Flow Estimation
abstract
FlowNet2 [14], the state-of-the-art convolutional neural network (CNN) for optical flow estimation, requires over 160M parameters to achieve accurate flow estimation. In this paper we present an alternative network that attains performance on par with FlowNet2 on the challenging Sintel final pass and KITTI benchmarks, while being 30 times smaller in the model size and 1.36 times faster in the running speed. This is made possible by drilling down to architectural details that might have been missed in the current frameworks: (1) We present a more effective flow inference approach at each pyramid level through a lightweight cascaded network. It not only improves flow estimation accuracy through early correction, but also permits seamless incorporation of descriptor matching in our network. (2) We present a novel flow regularization layer to ameliorate the issue of outliers and vague flow boundaries by using a feature-driven local convolution. (3) Our network owns an effective structure for pyramidal feature extraction and embraces feature warping rather than image warping as practiced in FlowNet2. Our code and trained models are available at github.com/twhui/LiteFlowNet.
Tak-Wai Hui, Xiaoou Tang, Chen Change Loy
CVPR2
2018 FaceID-GAN: Learning a Symmetry Three-Player GAN for Identity-Preserving Face Synthesis
abstract
Face synthesis has achieved advanced development by using generative adversarial networks (GANs). Existing methods typically formulate GAN as a two-player game, where a discriminator distinguishes face images from the real and synthesized domains, while a generator reduces its discriminativeness by synthesizing a face of photorealistic quality. Their competition converges when the discriminator is unable to differentiate these two domains. Unlike two-player GANs, this work generates identity-preserving faces by proposing FaceID-GAN, which treats a classifier of face identity as the third player, competing with the generator by distinguishing the identities of the real and synthesized faces (see Fig.1). A stationary point is reached when the generator produces faces that have high quality as well as preserve identity. Instead of simply modeling the identity classifier as an additional discriminator, FaceID-GAN is formulated by satisfying information symmetry, which ensures that the real and synthesized images are projected into the same feature space. In other words, the identity classifier is used to extract identity features from both input (real) and output (synthesized) face images of the generator, substantially alleviating training difficulty of GAN. Extensive experiments show that FaceID-GAN is able to generate faces of arbitrary viewpoint while preserve identity, outperforming recent advanced approaches.
Yujun Shen, Ping Luo 0002, Xiaogang Wang 0001, Xiaoou Tang
CVPR5
2018 Two at Once: Enhancing Learning and Generalization Capacities via IBN-Net
Xingang Pan, Ping Luo 0002, Jianping Shi, Xiaoou Tang
ECCV (4)4
2018 Aesthetic-Driven Image Enhancement by Adversarial Learning
abstract
We introduce EnhanceGAN, an adversarial learning based model that performs automatic image enhancement. Traditional image enhancement frameworks typically involve training models in a fully-supervised manner, which require expensive annotations in the form of aligned image pairs. In contrast to these approaches, our proposed EnhanceGAN only requires weak supervision (binary labels on image aesthetic quality) and is able to learn enhancement operators for the task of aesthetic-based image enhancement. In particular, we show the effectiveness of a piecewise color enhancement module trained with weak supervision, and extend the proposed EnhanceGAN framework to learning a deep filtering-based aesthetic enhancer. The full differentiability of our image enhancement operators enables the training of EnhanceGAN in an end-to-end manner. We further demonstrate the capability of EnhanceGAN in learning aesthetic-based image cropping without any groundtruth cropping pairs. Our weakly-supervised EnhanceGAN reports competitive quantitative results on aesthetic-based color enhancement as well as automatic image cropping, and a user study confirms that our image enhancement results are on par with or even preferred over professional enhancement.
Yubin Deng, Chen Change Loy, Xiaoou Tang
ACM Multimedia3
2018 From Facial Expression Recognition to Interpersonal Relation Prediction
Ping Luo 0002, Chen Change Loy, Xiaoou Tang
Int. J. Comput. Vis.4
2018 Deep Learning Markov Random Field for Semantic Segmentation
abstract
Semantic segmentation tasks can be well modeled by Markov Random Field (MRF). This paper addresses semantic segmentation by incorporating high-order relations and mixture of label contexts into MRF. Unlike previous works that optimized MRFs using iterative algorithm, we solve MRF by proposing a Convolutional Neural Network (CNN), namely Deep Parsing Network (DPN), which enables deterministic end-to-end computation in a single forward pass. Specifically, DPN extends a contemporary CNN to model unary terms and additional layers are devised to approximate the mean field (MF) algorithm for pairwise terms. It has several appealing properties. First, different from the recent works that required many iterations of MF during back-propagation, DPN is able to achieve high performance by approximating one iteration of MF. Second, DPN represents various types of pairwise terms, making many existing models as its special cases. Furthermore, pairwise terms in DPN provide a unified framework to encode rich contextual information in high-dimensional data, such as images and videos. Third, DPN makes MF easier to be parallelized and speeded up, thus enabling efficient inference. DPN is thoroughly evaluated on standard semantic image/video segmentation benchmarks, where a single DPN model yields state-of-the-art segmentation accuracies on PASCAL VOC 2012, Cityscapes dataset and CamVid dataset.
Ziwei Liu 0002, Ping Luo 0002, Chen Change Loy, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.5
2018 Faceness-Net: Face Detection through Deep Facial Part Responses
abstract
We propose a deep convolutional neural network (CNN) for face detection leveraging on facial attributes based supervision. We observe a phenomenon that part detectors emerge within CNN trained to classify attributes from uncropped face images, without any explicit part supervision. The observation motivates a new method for finding faces through scoring facial parts responses by their spatial structure and arrangement. The scoring mechanism is data-driven, and carefully formulated considering challenging cases where faces are only partially visible. This consideration allows our network to detect faces under severe occlusion and unconstrained pose variations. Our method achieves promising performance on popular benchmarks including FDDB, PASCAL Faces, AFW, and WIDER FACE.
Shuo Yang 0003, Ping Luo 0002, Chen Change Loy, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Discriminative Sparse Neighbor Approximation for Imbalanced Learning
abstract
Data imbalance is common in many vision tasks where one or more classes are rare. Without addressing this issue, conventional methods tend to be biased toward the majority class with poor predictive accuracy for the minority class. These methods further deteriorate on small, imbalanced data that have a large degree of class overlap. In this paper, we propose a novel discriminative sparse neighbor approximation (DSNA) method to ameliorate the effect of class-imbalance during prediction. Specifically, given a test sample, we first traverse it through a cost-sensitive decision forest to collect a good subset of training examples in its local neighborhood. Then, we generate from this subset several class-discriminating but overlapping clusters and model each as an affine subspace. From these subspaces, the proposed DSNA iteratively seeks an optimal approximation of the test sample and outputs an unbiased prediction. We show that our method not only effectively mitigates the imbalance issue, but also allows the prediction to extrapolate to unseen data. The latter capability is crucial for achieving accurate prediction on small data set with limited samples. The proposed imbalanced learning method can be applied to both classification and regression tasks at a wide range of imbalance levels. It significantly outperforms the state-of-the-art methods that do not possess an imbalance handling mechanism, and is found to perform comparably or even better than recent deep learning methods by using hand-crafted features only.
Chen Huang 0001, Chen Change Loy, Xiaoou Tang
IEEE Trans. Neural Networks Learn. Syst.3
2017 Not All Pixels Are Equal: Difficulty-Aware Semantic Segmentation via Deep Layer Cascade
abstract
We propose a novel deep layer cascade (LC) method to improve the accuracy and speed of semantic segmentation. Unlike the conventional model cascade (MC) that is composed of multiple independent models, LC treats a single deep model as a cascade of several sub-models. Earlier sub-models are trained to handle easy and confident regions, and they progressively feed-forward harder regions to the next sub-model for processing. Convolutions are only calculated on these regions to reduce computations. The proposed method possesses several advantages. First, LC classifies most of the easy regions in the shallow stage and makes deeper stage focuses on a few hard regions. Such an adaptive and difficulty-aware learning improves segmentation performance. Second, LC accelerates both training and testing of deep network thanks to early decisions in the shallow stage. Third, in comparison to MC, LC is an end-to-end trainable framework, allowing joint learning of all sub-models. We evaluate our method on PASCAL VOC and Cityscapes datasets, achieving state-of-the-art performance and fast speed.
Ziwei Liu 0002, Ping Luo 0002, Chen Change Loy, Xiaoou Tang
CVPR5
2017 ViP-CNN: Visual Phrase Guided Convolutional Neural Network
abstract
As the intermediate level task connecting image captioning and object detection, visual relationship detection started to catch researchers attention because of its descriptive power and clear structure. It detects the objects and captures their pair-wise interactions with a subject-predicate-object triplet, e.g. person-ride-horse. In this paper, each visual relationship is considered as a phrase with three components. We formulate the visual relationship detection as three inter-connected recognition problems and propose a Visual Phrase guided Convolutional Neural Network (ViP-CNN) to address them simultaneously. In ViP-CNN, we present a Phrase-guided Message Passing Structure (PMPS) to establish the connection among relationship components and help the model consider the three problems jointly. Corresponding non-maximum suppression method and model training strategy are also proposed. Experimental results show that our ViP-CNN outperforms the state-of-art method both in speed and accuracy. We further pretrain ViP-CNN on our cleansed Visual Genome Relationship dataset, which is found to perform better than the pretraining on the ImageNet for this task.
Yikang Li 0002, Wanli Ouyang, Xiaogang Wang 0001, Xiaoou Tang
CVPR4
2017 Residual Attention Network for Image Classification
abstract
In this work, we propose Residual Attention Network, a convolutional neural network using attention mechanism which can incorporate with state-of-art feed forward network architecture in an end-to-end training fashion. Our Residual Attention Network is built by stacking Attention Modules which generate attention-aware features. The attention-aware features from different modules change adaptively as layers going deeper. Inside each Attention Module, bottom-up top-down feedforward structure is used to unfold the feedforward and feedback attention process into a single feedforward process. Importantly, we propose attention residual learning to train very deep Residual Attention Networks which can be easily scaled up to hundreds of layers. Extensive analyses are conducted on CIFAR-10 and CIFAR-100 datasets to verify the effectiveness of every module mentioned above. Our Residual Attention Network achieves state-of-the-art object recognition performance on three benchmark datasets including CIFAR-10 (3.90% error), CIFAR-100 (20.45% error) and ImageNet (4.8% single model and single crop, top-5 error). Note that, our method achieves 0.6% top-1 accuracy improvement with 46% trunk depth and 69% forward FLOPs comparing to ResNet-200. The experiment also demonstrates that our network is robust against noisy labels.
Fei Wang 0032, Mengqing Jiang, Chen Qian 0006, Shuo Yang 0003, Cheng Li 0009, Xiaogang Wang 0001, Xiaoou Tang
CVPR8
2017 Spindle Net: Person Re-identification with Human Body Region Guided Feature Decomposition and Fusion
abstract
Person re-identification (ReID) is an important task in video surveillance and has various applications. It is non-trivial due to complex background clutters, varying illumination conditions, and uncontrollable camera settings. Moreover, the person body misalignment caused by detectors or pose variations is sometimes too severe for feature matching across images. In this study, we propose a novel Convolutional Neural Network (CNN), called Spindle Net, based on human body region guided multi-stage feature decomposition and tree-structured competitive feature fusion. It is the first time human body structure information is considered in a CNN framework to facilitate feature learning. The proposed Spindle Net brings unique advantages: 1) it separately captures semantic features from different body regions thus the macro-and micro-body features can be well aligned across images, 2) the learned region features from different semantic regions are merged with a competitive scheme and discriminative features can be well preserved. State of the art performance can be achieved on multiple datasets by large margins. We further demonstrate the robustness and effectiveness of the proposed Spindle Net on our proposed dataset SenseReID without fine-tuning.
Haiyu Zhao, Maoqing Tian, Shuyang Sun, Shuai Yi, Xiaogang Wang 0001, Xiaoou Tang
CVPR8
2017 Learning to Disambiguate by Asking Discriminative Questions
abstract
The ability to ask questions is a powerful tool to gather information in order to learn about the world and resolve ambiguities. In this paper, we explore a novel problem of generating discriminative questions to help disambiguate visual instances. Our work can be seen as a complement and new extension to the rich research studies on image captioning and question answering. We introduce the first large-scale dataset with over 10,000 carefully annotated images-question tuples to facilitate benchmarking. In particular, each tuple consists of a pair of images and 4.6 discriminative questions (as positive samples) and 5.9 non-discriminative questions (as negative samples) on average. In addition, we present an effective method for visual discriminative question generation. The method can be trained in a weakly supervised manner without discriminative images-question tuples but just existing visual question answering datasets. Promising results are shown against representative baselines through quantitative evaluations and user studies.
Chen Huang 0001, Xiaoou Tang, Chen Change Loy
ICCV3
2017 Recurrent Scale Approximation for Object Detection in CNN
abstract
Since convolutional neural network (CNN) lacks an inherent mechanism to handle large scale variations, we always need to compute feature maps multiple times for multiscale object detection, which has the bottleneck of computational cost in practice. To address this, we devise a recurrent scale approximation (RSA) to compute feature map once only, and only through this map can we approximate the rest maps on other levels. At the core of RSA is the recursive rolling out mechanism: given an initial map on a particular scale, it generates the prediction on a smaller scale that is half the size of input. To further increase efficiency and accuracy, we (a): design a scale-forecast network to globally predict potential scales in the image since there is no need to compute maps on all levels of the pyramid. (b): propose a landmark retracing network (LRN) to retrace back locations of the regressed landmarks and generate a confidence score for each landmark; LRN can effectively alleviate false positives due to the accumulated error in RSA. The whole system could be trained end-to-end in a unified CNN framework. Experiments demonstrate that our proposed algorithm is superior against state-of-the-arts on face detection benchmarks and achieves comparable results for generic proposal generation. The source code of our system is available.
Yu Liu 0015, Hongyang Li 0001, Fangyin Wei, Xiaogang Wang 0001, Xiaoou Tang
ICCV6
2017 Video Frame Synthesis Using Deep Voxel Flow
abstract
We address the problem of synthesizing new video frames in an existing video, either in-between existing frames (interpolation), or subsequent to them (extrapolation). This problem is challenging because video appearance and motion can be highly complex. Traditional optical-flow-based solutions often fail where flow estimation is challenging, while newer neural-network-based methods that hallucinate pixel values directly often produce blurry results. We combine the advantages of these two methods by training a deep network that learns to synthesize video frames by flowing pixel values from existing ones, which we call deep voxel flow. Our method requires no human supervision, and any video can be used as training data by dropping, and then learning to predict, existing frames. The technique is efficient, and can be applied at any video resolution. We demonstrate that our method produces results that both quantitatively and qualitatively improve upon the state-of-the-art.
Ziwei Liu 0002, Raymond A. Yeh, Xiaoou Tang, Yiming Liu 0001, Aseem Agarwala
ICCV3
2017 Temporal Action Detection with Structured Segment Networks
abstract
Detecting actions in untrimmed videos is an important yet challenging task. In this paper, we present the structured segment network (SSN), a novel framework which models the temporal structure of each action instance via a structured temporal pyramid. On top of the pyramid, we further introduce a decomposed discriminative model comprising two classifiers, respectively for classifying actions and determining completeness. This allows the framework to effectively distinguish positive proposals from background or incomplete ones, thus leading to both accurate recognition and localization. These components are integrated into a unified network that can be efficiently trained in an end-to-end fashion. Additionally, a simple yet effective temporal action proposal scheme, dubbed temporal actionness grouping (TAG) is devised to generate high quality action proposals. On two challenging benchmarks, THUMOS14 and ActivityNet, our method remarkably outperforms previous state-of-the-art methods, demonstrating superior accuracy and strong adaptivity in handling actions with various temporal structures.
Yue Zhao 0006, Yuanjun Xiong, Limin Wang 0002, Zhirong Wu, Xiaoou Tang, Dahua Lin
ICCV5
2017 Unconstrained Fashion Landmark Detection via Hierarchical Recurrent Transformer Networks
abstract
Fashion landmarks are functional key points defined on clothes, such as corners of neckline, hemline, and cuff. They have been recently introduced [18]as an effective visual representation for fashion image understanding. However, detecting fashion landmarks are challenging due to background clutters, human poses, and scales. To remove the above variations, previous works usually assumed bounding boxes of clothes are provided in training and test as additional annotations, which are expensive to obtain and inapplicable in practice. This work addresses unconstrained fashion landmark detection, where clothing bounding boxes are not provided in both training and test. To this end, we present a novel Deep LAndmark Network (DLAN), where bounding boxes and landmarks are jointly estimated and trained iteratively in an end-to-end manner. DLAN contains two dedicated modules, including a Selective Dilated Convolution for handling scale discrepancies, and a Hierarchical Recurrent Spatial Transformer for handling background clutters. To evaluate DLAN, we present a large-scale fashion landmark dataset, namely Unconstrained Landmark Database (ULD), consisting of 30K images. Statistics show that ULD is more challenging than existing datasets in terms of image scales, background clutters, and human poses. Extensive experiments demonstrate the effectiveness of DLAN over the state-of-the-art methods. DLAN also exhibits excellent generalization across different clothing categories and modalities, making it extremely suitable for real-world fashion analysis.
Sijie Yan, Ziwei Liu 0002, Ping Luo 0002, Xiaogang Wang 0001, Xiaoou Tang
ACM Multimedia6
2017 DeepID-Net: Object Detection with Deformable Part Based Convolutional Neural Networks
abstract
In this paper, we propose deformable deep convolutional neural networks for generic object detection. This new deep learning object detection framework has innovations in multiple aspects. In the proposed new deep architecture, a new deformation constrained pooling (def-pooling) layer models the deformation of object parts with geometric constraint and penalty. A new pre-training strategy is proposed to learn feature representations more suitable for the object detection task and with good generalization capability. By changing the net structures, training strategies, adding and removing some key components in the detection pipeline, a set of models with large diversity are obtained, which significantly improves the effectiveness of model averaging. The proposed approach improves the mean averaged precision obtained by RCNN [16], which was the state-of-the-art, from 31% to 50.3% on the ILSVRC2014 detection test set. It also outperforms the winner of ILSVRC2014, GoogLeNet, by 6.1%. Detailed component-wise analysis is also provided through extensive experimental evaluation, which provides a global view for people to understand the deep learning object detection pipeline.
Wanli Ouyang, Xingyu Zeng, Xiaogang Wang 0001, Ping Luo 0002, Yonglong Tian, Hongsheng Li 0001, Shuo Yang 0003, Zhe Wang 0006, Hongyang Li 0001, Kun Wang 0056, Chen Change Loy, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.14
2016 Reading Scene Text in Deep Convolutional Sequences
abstract
We develop a Deep-Text Recurrent Network (DTRN)that regards scene text reading as a sequence labelling problem. We leverage recent advances of deep convolutional neural networks to generate an ordered highlevel sequence from a whole word image, avoiding the difficult character segmentation problem. Then a deep recurrent model, building on long short-term memory (LSTM), is developed to robustly recognize the generated CNN sequences, departing from most existing approaches recognising each character independently. Our model has a number of appealing properties in comparison to existing scene text recognition methods: (i) It can recognise highly ambiguous words by leveraging meaningful context information, allowing it to work reliably without either pre- or post-processing; (ii) the deep CNN feature is robust to various image distortions; (iii) it retains the explicit order information in word image, which is essential to discriminate word strings; (iv) the model does not depend on pre-defined dictionary, and it can process unknown words and arbitrary strings. It achieves impressive results on several benchmarks, advancing the-state-of-the-art substantially.
Pan He, Yu Qiao 0001, Chen Change Loy, Xiaoou Tang
AAAI5
2016 Face Model Compression by Distilling Knowledge from Neurons
abstract
The recent advanced face recognition systems werebuilt on large Deep Neural Networks (DNNs) or theirensembles, which have millions of parameters. However, the expensive computation of DNNs make theirdeployment difficult on mobile and embedded devices. This work addresses model compression for face recognition,where the learned knowledge of a large teachernetwork or its ensemble is utilized as supervisionto train a compact student network. Unlike previousworks that represent the knowledge by the soften labelprobabilities, which are difficult to fit, we represent theknowledge by using the neurons at the higher hiddenlayer, which preserve as much information as the label probabilities, but are more compact. By leveragingthe essential characteristics (domain knowledge) of thelearned face representation, a neuron selection methodis proposed to choose neurons that are most relevant toface recognition. Using the selected neurons as supervisionto mimic the single networks of DeepID2+ andDeepID3, which are the state-of-the-art face recognition systems, a compact student with simple network structure achieves better verification accuracy on LFW than its teachers, respectively. When using an ensemble of DeepID2+ as teacher, a mimicked student is able to outperform it and achieves 51.6 times compression ratio and 90 times speed-up in inference, making this cumbersome model applicable on portable devices.
Ping Luo 0002, Zhenyao Zhu, Ziwei Liu 0002, Xiaogang Wang 0001, Xiaoou Tang
AAAI5
2016 Learning Deep Representation for Imbalanced Classification
abstract
Data in vision domain often exhibit highly-skewed class distribution, i.e., most data belong to a few majority classes, while the minority classes only contain a scarce amount of instances. To mitigate this issue, contemporary classification methods based on deep convolutional neural network (CNN) typically follow classic strategies such as class re-sampling or cost-sensitive training. In this paper, we conduct extensive and systematic experiments to validate the effectiveness of these classic schemes for representation learning on class-imbalanced data. We further demonstrate that more discriminative deep representation can be learned by enforcing a deep network to maintain both intercluster and inter-class margins. This tighter constraint effectively reduces the class imbalance inherent in the local data neighborhood. We show that the margins can be easily deployed in standard deep learning framework through quintuplet instance sampling and the associated triple-header hinge loss. The representation learned by our approach, when combined with a simple k-nearest neighbor (kNN) algorithm, shows significant improvements over existing methods on both high-and low-level vision classification tasks that exhibit imbalanced class distribution.
Chen Huang 0001, Chen Change Loy, Xiaoou Tang
CVPR4
2016 Unsupervised Learning of Discriminative Attributes and Visual Representations
abstract
Attributes offer useful mid-level features to interpret visual data. While most attribute learning methods are supervised by costly human-generated labels, we introduce a simple yet powerful unsupervised approach to learn and predict visual attributes directly from data. Given a large unlabeled image collection as input, we train deep Convolutional Neural Networks (CNNs) to output a set of discriminative, binary attributes often with semantic meanings. Specifically, we first train a CNN coupled with unsupervised discriminative clustering, and then use the cluster membership as a soft supervision to discover shared attributes from the clusters while maximizing their separability. The learned attributes are shown to be capable of encoding rich imagery properties from both natural images and contour patches. The visual representations learned in this way are also transferrable to other tasks such as object detection. We show other convincing results on the related tasks of image retrieval and classification, and contour detection.
Chen Huang 0001, Chen Change Loy, Xiaoou Tang
CVPR3
2016 DeepFashion: Powering Robust Clothes Recognition and Retrieval with Rich Annotations
abstract
Recent advances in clothes recognition have been driven by the construction of clothes datasets. Existing datasets are limited in the amount of annotations and are difficult to cope with the various challenges in real-world applications. In this work, we introduce DeepFashion1, a large-scale clothes dataset with comprehensive annotations. It contains over 800,000 images, which are richly annotated with massive attributes, clothing landmarks, and correspondence of images taken under different scenarios including store, street snapshot, and consumer. Such rich annotations enable the development of powerful algorithms in clothes recognition and facilitating future researches. To demonstrate the advantages of DeepFashion, we propose a new deep model, namely FashionNet, which learns clothing features by jointly predicting clothing attributes and landmarks. The estimated landmarks are then employed to pool or gate the learned features. It is optimized in an iterative manner. Extensive experiments demonstrate the effectiveness of FashionNet and the usefulness of DeepFashion.
Ziwei Liu 0002, Ping Luo 0002, Xiaogang Wang 0001, Xiaoou Tang
CVPR5
2016 Sparsifying Neural Network Connections for Face Recognition
abstract
This paper proposes to learn high-performance deep ConvNets with sparse neural connections, referred to as sparse ConvNets, for face recognition. The sparse ConvNets are learned in an iterative way, each time one additional layer is sparsified and the entire model is re-trained given the initial weights learned in previous iterations. One important finding is that directly training the sparse ConvNet from scratch failed to find good solutions for face recognition, while using a previously learned denser model to properly initialize a sparser model is critical to continue learning effective features for face recognition. This paper also proposes a new neural correlation-based weight selection criterion and empirically verifies its effectiveness in selecting informative connections from previously learned models in each iteration. When taking a moderately sparse structure (26%-76% of weights in the dense model), the proposed sparse ConvNet model significantly improves the face recognition performance of the previous state-of-the-art DeepID2+ models given the same training data, while it keeps the performance of the baseline model with only 12% of the original parameters.
Xiaogang Wang 0001, Xiaoou Tang
CVPR3
2016 Actionness Estimation Using Hybrid Fully Convolutional Networks
abstract
Actionness was introduced to quantify the likelihood of containing a generic action instance at a specific location. Accurate and efficient estimation of actionness is important in video analysis and may benefit other relevant tasks such as action recognition and action detection. This paper presents a new deep architecture for actionness estimation, called hybrid fully convolutional network (HFCN), which is composed of appearance FCN (A-FCN) and motion FCN (M-FCN). These two FCNs leverage the strong capacity of deep models to estimate actionness maps from the perspectives of static appearance and dynamic motion, respectively. In addition, the fully convolutional nature of H-FCN allows it to efficiently process videos with arbitrary sizes. Experiments are conducted on the challenging datasets of Stanford40, UCF Sports, and JHMDB to verify the effectiveness of H-FCN on actionness estimation, which demonstrate that our method achieves superior performance to previous ones. Moreover, we apply the estimated actionness maps on action proposal generation and action detection. Our actionness maps advance the current state-of-the-art performance of these tasks substantially.
Limin Wang 0002, Yu Qiao 0001, Xiaoou Tang, Luc Van Gool
CVPR3
2016 WIDER FACE: A Face Detection Benchmark
abstract
Face detection is one of the most studied topics in the computer vision community. Much of the progresses have been made by the availability of face detection benchmark datasets. We show that there is a gap between current face detection performance and the real world requirements. To facilitate future face detection research, we introduce the WIDER FACE dataset1, which is 10 times larger than existing datasets. The dataset contains rich annotations, including occlusions, poses, event categories, and face bounding boxes. Faces in the proposed dataset are extremely challenging due to large variations in scale, pose and occlusion, as shown in Fig. 1. Furthermore, we show that WIDER FACE dataset is an effective training source for face detection. We benchmark several representative detection systems, providing an overview of state-of-the-art performance and propose a solution to deal with large scale variation. Finally, we discuss common failure cases that worth to be further investigated.
Shuo Yang 0003, Ping Luo 0002, Chen Change Loy, Xiaoou Tang
CVPR4
2016 Unconstrained Face Alignment via Cascaded Compositional Learning
abstract
We present a practical approach to address the problem of unconstrained face alignment for a single image. In our unconstrained problem, we need to deal with large shape and appearance variations under extreme head poses and rich shape deformation. To equip cascaded regressors with the capability to handle global shape variation and irregular appearance-shape relation in the unconstrained scenario, we partition the optimisation space into multiple domains of homogeneous descent, and predict a shape as a composition of estimations from multiple domain-specific regressors. With a specially formulated learning objective and a novel tree splitting function, our approach is capable of estimating a robust and meaningful composition. In addition to achieving state-of-the-art accuracy over existing approaches, our framework is also an efficient solution (350 FPS), thanks to the on-the-fly domain exclusion mechanism and the capability of leveraging the fast pixel feature.
Shizhan Zhu, Cheng Li 0009, Chen Change Loy, Xiaoou Tang
CVPR4
2016 Accelerating the Super-Resolution Convolutional Neural Network
Chao Dong 0005, Chen Change Loy, Xiaoou Tang
ECCV (2)3
2016 Depth Map Super-Resolution by Deep Multi-Scale Guidance
Tak-Wai Hui, Chen Change Loy, Xiaoou Tang
ECCV (3)3
2016 Human Attribute Recognition by Deep Hierarchical Contexts
Chen Huang 0001, Chen Change Loy, Xiaoou Tang
ECCV (6)4
2016 Fashion Landmark Detection in the Wild
Ziwei Liu 0002, Sijie Yan, Ping Luo 0002, Xiaogang Wang 0001, Xiaoou Tang
ECCV (2)5
2016 Deep Specialized Network for Illuminant Estimation
Wu Shi, Chen Change Loy, Xiaoou Tang
ECCV (4)3
2016 Temporal Segment Networks: Towards Good Practices for Deep Action Recognition
Limin Wang 0002, Yuanjun Xiong, Zhe Wang 0013, Yu Qiao 0001, Dahua Lin, Xiaoou Tang, Luc Van Gool
ECCV (8)6
2016 Deep Markov Random Field for Image Modeling
Zhirong Wu, Dahua Lin, Xiaoou Tang
ECCV (8)3
2016 Joint Face Representation Adaptation and Clustering in Videos
Ping Luo 0002, Chen Change Loy, Xiaoou Tang
ECCV (3)4
2016 Deep Cascaded Bi-Network for Face Hallucination
Shizhan Zhu, Sifei Liu, Chen Change Loy, Xiaoou Tang
ECCV (5)4
2016 Local Similarity-Aware Deep Feature Embedding
abstract
Existing deep embedding methods in vision tasks are capable of learning a compact Euclidean space from images, where Euclidean distances correspond to a similarity metric. To make learning more effective and efficient, hard sample mining is usually employed, with samples identified through computing the Euclidean feature distance. However, the global Euclidean distance cannot faithfully characterize the true feature similarity in a complex visual feature space, where the intraclass distance in a high-density region may be larger than the interclass distance in low-density regions. In this paper, we introduce a Position-Dependent Deep Metric (PDDM) unit, which is capable of learning a similarity metric adaptive to local feature structure. The metric can be used to select genuinely hard samples in a local neighborhood to guide the deep embedding learning in an online and robust manner. The new layer is appealing in that it is pluggable to any convolutional networks and is trained end-to-end. Our local similarity-aware feature embedding not only demonstrates faster convergence and boosted performance on two complex image retrieval datasets, its large margin nature also leads to superior generalization results under the large and open set scenarios of transfer learning and zero-shot learning on ImageNet 2010 and ImageNet-10K datasets.
Chen Huang 0001, Chen Change Loy, Xiaoou Tang
NIPS3
2016 MoFAP: A Multi-level Representation for Action Recognition
Limin Wang 0002, Yu Qiao 0001, Xiaoou Tang
Int. J. Comput. Vis.3
2016 Image Super-Resolution Using Deep Convolutional Networks
abstract
We propose a deep learning method for single image super-resolution (SR). Our method directly learns an end-to-end mapping between the low/high-resolution images. The mapping is represented as a deep convolutional neural network (CNN) that takes the low-resolution image as the input and outputs the high-resolution one. We further show that traditional sparse-coding-based SR methods can also be viewed as a deep convolutional network. But unlike traditional methods that handle each component separately, our method jointly optimizes all layers. Our deep CNN has a lightweight structure, yet demonstrates state-of-the-art restoration quality, and achieves fast speed for practical on-line usage. We explore different network structures and parameter settings to achieve trade-offs between performance and speed. Moreover, we extend our network to cope with three color channels simultaneously, and show better overall reconstruction quality.
Chao Dong 0005, Chen Change Loy, Kaiming He, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Hybrid Deep Learning for Face Verification
abstract
This paper proposes a hybrid convolutional network (ConvNet)-Restricted Boltzmann Machine (RBM) model for face verification. A key contribution of this work is to learn high-level relational visual features with rich identity similarity information. The deep ConvNets in our model start by extracting local relational visual features from two face images in comparison, which are further processed through multiple layers to extract high-level and global relational features. To keep enough discriminative information, we use the last hidden layer neuron activations of the ConvNet as features for face verification instead of those of the output layer. To characterize face similarities from different aspects, we concatenate the features extracted from different face region pairs by different deep ConvNets. The resulting high-dimensional relational features are classified by an RBM for face verification. After pre-training each ConvNet and the RBM separately, the entire hybrid network is jointly optimized to further improve the accuracy. Various aspects of the ConvNet structures, relational features, and face verification classifiers are investigated. Our model achieves the state-of-the-art face verification performance on the challenging LFW dataset under both the unrestricted protocol and the setting when outside data is allowed to be used for training.
Xiaogang Wang 0001, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.3
2016 Learning Deep Representation for Face Alignment with Auxiliary Attributes
abstract
In this study, we show that landmark detection or face alignment task is not a single and independent problem. Instead, its robustness can be greatly improved with auxiliary information. Specifically, we jointly optimize landmark detection together with the recognition of heterogeneous but subtly correlated facial attributes, such as gender, expression, and appearance attributes. This is non-trivial since different attribute inference tasks have different learning difficulties and convergence rates. To address this problem, we formulate a novel tasks-constrained deep model, which not only learns the inter-task correlation but also employs dynamic task coefficients to facilitate the optimization convergence when learning multiple complex tasks. Extensive evaluations show that the proposed task-constrained learning (i) outperforms existing face alignment methods, especially in dealing with faces with severe occlusion and pose variation, and (ii) reduces model complexity drastically compared to the state-of-the-art methods based on cascaded deep model.
Ping Luo 0002, Chen Change Loy, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.4
2016 Bridging Music and Image via Cross-Modal Ranking Analysis
abstract
Human perceptions of music and image are closely related to each other, since both can inspire similar human sensations, such as emotion, motion, and power. This paper aims to explore whether and how music and image can be automatically matched by machines. The main contributions are three aspects. First, we construct a benchmark dataset composed of more than 45 000 music-image pairs. Human labelers are recruited to annotate whether these pairs are well-matched or not. The results show that they generally agree with each other on the matching degree of music-image pairs. Secondly, we investigate suitable semantic representations of music and image for this cross-modal matching task. In particular, we adopt lyrics as a middle-media to connect music and image, and design a set of lyric-based attributes for image representation. Thirdly, we propose cross-modal ranking analysis (CMRA) to learn the semantic similarity between music and image with ranking labeling information. CMRA aims to find the optimal embedding spaces for both music and image in the sense of maximizing the ordinal margin between music-image pairs. The proposed method is able to learn the non-linear relationship between music and image, and to integrate heterogeneous ranking data from different modalities into a unified space. Experimental results demonstrate that the proposed method outperforms state-of-the-art cross-modal methods in the music-image matching task, and achieves a consistency rate of 91.5% with human labelers.
Xixuan Wu, Yu Qiao 0001, Xiaogang Wang 0001, Xiaoou Tang
IEEE Trans. Multim.4
2015 Surpassing Human-Level Face Verification Performance on LFW with GaussianFace
abstract
Face verification remains a challenging problem in very complex conditions with large variations such as pose, illumination, expression, and occlusions. This problemis exacerbated when we rely unrealistically on a singletraining data source, which is often insufficient to coverthe intrinsically complex face variations. This paperproposes a principled multi-task learning approachbased on Discriminative Gaussian Process Latent VariableModel (DGPLVM), named GaussianFace, for faceverification. In contrast to relying unrealistically on asingle training data source, our model exploits additional data from multiple source-domains to improve the generalization performance of face verification inan unknown target-domain. Importantly, our model can adapt automatically to complex data distributions, and therefore can well capture complex face variations inherent in multiple sources. To enhance discriminative power, we introduced a more efficient equivalent form of Kernel Fisher Discriminant Analysis to DGPLVM.To speed up the process of inference and prediction, we exploited the low rank approximation method. Extensive experiments demonstrated the effectiveness of the proposed model in learning from diverse data sources and generalizing to unseen domains. Specifically, the accuracy of our algorithm achieved an impressive accuracyrate of 98.52% on the well-known and challenging Labeled Faces in the Wild (LFW) benchmark. For the first time, the human-level performance in face verification (97.53%) on LFW is surpassed.
Chaochao Lu, Xiaoou Tang
AAAI2
2015 Deep Representation Learning with Target Coding
abstract
We consider the problem of learning deep representation when target labels are available. In this paper, we show that there exists intrinsic relationship between target coding and feature representation learning in deep networks. Specifically, we found that distributed binary acode with error correcting capability is more capable of encouraging discriminative features, in comparison tothe 1-of-K coding that is typically used in supervised deep learning. This new finding reveals additional benefit of using error-correcting code for deep model learning,apart from its well-known error correcting property. Extensive experiments are conducted on popular visual benchmark datasets.
Shuo Yang 0003, Ping Luo 0002, Chen Change Loy, Kenneth W. Shum, Xiaoou Tang
AAAI5
2015 Cascaded hand pose regression
abstract
We extends the previous 2D cascaded object pose regression work [9] in two aspects so that it works better for 3D articulated objects. Our first contribution is 3D pose-indexed features that generalize the previous 2D parameterized features and achieve better invariance to 3D transformations. Our second contribution is a principled hierarchical regression that is adapted to the articulated object structure. It is therefore more accurate and faster. Comprehensive experiments verify the state-of-the-art accuracy and efficiency of the proposed approach on the challenging 3D hand pose estimation problem, on a public dataset and our new dataset.
Xiao Sun 0001, Shuang Liang 0001, Xiaoou Tang, Jian Sun 0001
CVPR4
2015 DeepID-Net: Deformable deep convolutional neural networks for object detection
abstract
In this paper, we propose deformable deep convolutional neural networks for generic object detection. This new deep learning object detection framework has innovations in multiple aspects. In the proposed new deep architecture, a new deformation constrained pooling (def-pooling) layer models the deformation of object parts with geometric constraint and penalty. A new pre-training strategy is proposed to learn feature representations more suitable for the object detection task and with good generalization capability. By changing the net structures, training strategies, adding and removing some key components in the detection pipeline, a set of models with large diversity are obtained, which significantly improves the effectiveness of model averaging. The proposed approach improves the mean averaged precision obtained by RCNN [14], which was the state-of-the-art, from 31% to 50.3% on the ILSVRC2014 detection test set. It also outperforms the winner of ILSVRC2014, GoogLeNet, by 6.1%. Detailed component-wise analysis is also provided through extensive experimental evaluation, which provide a global view for people to understand the deep learning object detection pipeline.
Wanli Ouyang, Xiaogang Wang 0001, Xingyu Zeng, Ping Luo 0002, Yonglong Tian, Hongsheng Li 0001, Shuo Yang 0003, Zhe Wang 0006, Chen Change Loy, Xiaoou Tang
CVPR11
2015 Deeply learned face representations are sparse, selective, and robust
abstract
This paper designs a high-performance deep convolutional network (DeepID2+) for face recognition. It is learned with the identification-verification supervisory signal. By increasing the dimension of hidden representations and adding supervision to early convolutional layers, DeepID2+ achieves new state-of-the-art on LFW and YouTube Faces benchmarks. Through empirical studies, we have discovered three properties of its deep neural activations critical for the high performance: sparsity, selectiveness and robustness. (1) It is observed that neural activations are moderately sparse. Moderate sparsity maximizes the discriminative power of the deep net as well as the distance between images. It is surprising that DeepID2+ still can achieve high recognition accuracy even after the neural responses are binarized. (2) Its neurons in higher layers are highly selective to identities and identity-related attributes. We can identify different subsets of neurons which are either constantly excited or inhibited when different identities or attributes are present. Although DeepID2+ is not taught to distinguish attributes during training, it has implicitly learned such high-level concepts. (3) It is much more robust to occlusions, although occlusion patterns are not included in the training set.
Xiaogang Wang 0001, Xiaoou Tang
CVPR3
2015 Pedestrian detection aided by deep learning semantic tasks
abstract
Deep learning methods have achieved great successes in pedestrian detection, owing to its ability to learn discriminative features from raw pixels. However, they treat pedestrian detection as a single binary classification task, which may confuse positive with hard negative samples (Fig.1 (a)). To address this ambiguity, this work jointly optimize pedestrian detection with semantic tasks, including pedestrian attributes (e.g. `carrying backpack') and scene attributes (e.g. `vehicle', `tree', and `horizontal'). Rather than expensively annotating scene attributes, we transfer attributes information from existing scene segmentation datasets to the pedestrian dataset, by proposing a novel deep model to learn high-level features from multiple tasks and multiple data sources. Since distinct tasks have distinct convergence rates and data from different datasets have different distributions, a multi-task deep model is carefully designed to coordinate tasks and reduce discrepancies among datasets. Extensive evaluations show that the proposed approach outperforms the state-of-the-art on the challenging Caltech [9] and ETH [10] datasets where it reduces the miss rates of previous deep models by 17 and 5.5 percent, respectively.
Yonglong Tian, Ping Luo 0002, Xiaogang Wang 0001, Xiaoou Tang
CVPR4
2015 Action recognition with trajectory-pooled deep-convolutional descriptors
abstract
Visual features are of vital importance for human action understanding in videos. This paper presents a new video representation, called trajectory-pooled deep-convolutional descriptor (TDD), which shares the merits of both hand-crafted features [31] and deep-learned features [24]. Specifically, we utilize deep architectures to learn discriminative convolutional feature maps, and conduct trajectory-constrained pooling to aggregate these convolutional features into effective descriptors. To enhance the robustness of TDDs, we design two normalization methods to transform convolutional feature maps, namely spatiotemporal normalization and channel normalization. The advantages of our features come from (i) TDDs are automatically learned and contain high discriminative capacity compared with those hand-crafted features; (ii) TDDs take account of the intrinsic characteristics of temporal dimension and introduce the strategies of trajectory-constrained sampling and pooling for aggregating deep-learned features. We conduct experiments on two challenging datasets: HMD-B51 and UCF101. Experimental results show that TDDs outperform previous hand-crafted features [31] and deep-learned features [24]. Our method also achieves superior performance to the state of the art on these datasets.
Limin Wang 0002, Yu Qiao 0001, Xiaoou Tang
CVPR3
2015 3D ShapeNets: A deep representation for volumetric shapes
abstract
3D shape is a crucial but heavily underutilized cue in today's computer vision systems, mostly due to the lack of a good generic shape representation. With the recent availability of inexpensive 2.5D depth sensors (e.g. Microsoft Kinect), it is becoming increasingly important to have a powerful 3D shape representation in the loop. Apart from category recognition, recovering full 3D shapes from view-based 2.5D depth maps is also a critical part of visual understanding. To this end, we propose to represent a geometric 3D shape as a probability distribution of binary variables on a 3D voxel grid, using a Convolutional Deep Belief Network. Our model, 3D ShapeNets, learns the distribution of complex 3D shapes across different object categories and arbitrary poses from raw CAD data, and discovers hierarchical compositional part representation automatically. It naturally supports joint object recognition and shape completion from 2.5D depth maps, and it enables active object recognition through view planning. To train our 3D deep learning model, we construct ModelNet - a large-scale 3D CAD model dataset. Extensive experiments show that our 3D deep representation enables significant performance improvement over the-state-of-the-arts in a variety of tasks.
Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu 0001, Linguang Zhang, Xiaoou Tang, Jianxiong Xiao
CVPR6
2015 Recognize complex events from static images by fusing deep channels
abstract
A considerable portion of web images capture events that occur in our personal lives or social activities. In this paper, we aim to develop an effective method for recognizing events from such images. Despite the sheer amount of study on event recognition, most existing methods rely on videos and are not directly applicable to this task. Generally, events are complex phenomena that involve interactions among people and objects, and therefore analysis of event photos requires techniques that can go beyond recognizing individual objects and carry out joint reasoning based on evidences of multiple aspects. Inspired by the recent success of deep learning, we formulate a multi-layer framework to tackle this problem, which takes into account both visual appearance and the interactions among humans and objects, and combines them via semantic fusion. An important issue arising here is that humans and objects discovered by detectors are in the form of bounding boxes, and there is no straightforward way to represent their interactions and incorporate them with a deep network. We address this using a novel strategy that projects the detected instances onto multi-scale spatial maps. On a large dataset with 60, 000 images, the proposed method achieved substantial improvement over the state-of-the-art, raising the accuracy of event recognition by over 10%.
Yuanjun Xiong, Dahua Lin, Xiaoou Tang
CVPR4
2015 A large-scale car dataset for fine-grained categorization and verification
abstract
This paper aims to highlight vision related tasks centered around “car”, which has been largely neglected by vision community in comparison to other objects. We show that there are still many interesting car-related problems and applications, which are not yet well explored and researched. To facilitate future car-related research, in this paper we present our on-going effort in collecting a large-scale dataset, “CompCars”, that covers not only different car views, but also their different internal and external parts, and rich attributes. Importantly, the dataset is constructed with a cross-modality nature, containing a surveillance-nature set and a web-nature set. We further demonstrate a few important applications exploiting the dataset, namely car model classification, car model verification, and attribute prediction. We also discuss specific challenges of the car-related problems and other potential applications that worth further investigations. The latest dataset can be downloaded at http://mmlab.ie.cuhk.edu.hk/ datasets/comp_cars/index.html.
Ping Luo 0002, Chen Change Loy, Xiaoou Tang
CVPR4
2015 Face alignment by coarse-to-fine shape searching
abstract
We present a novel face alignment framework based on coarse-to-fine shape searching. Unlike the conventional cascaded regression approaches that start with an initial shape and refine the shape in a cascaded manner, our approach begins with a coarse search over a shape space that contains diverse shapes, and employs the coarse solution to constrain subsequent finer search of shapes. The unique stage-by-stage progressive and adaptive search i) prevents the final solution from being trapped in local optima due to poor initialisation, a common problem encountered by cascaded regression approaches; and ii) improves the robustness in coping with large pose variations. The framework demonstrates real-time performance and state-of-the-art results on various benchmarks including the challenging 300-W dataset.
Shizhan Zhu, Cheng Li 0009, Chen Change Loy, Xiaoou Tang
CVPR4
2015 Compression Artifacts Reduction by a Deep Convolutional Network
abstract
Lossy compression introduces complex compression artifacts, particularly the blocking artifacts, ringing effects and blurring. Existing algorithms either focus on removing blocking artifacts and produce blurred output, or restores sharpened images that are accompanied with ringing effects. Inspired by the deep convolutional networks (DCN) on super-resolution, we formulate a compact and efficient network for seamless attenuation of different compression artifacts. We also demonstrate that a deeper model can be effectively trained with the features learned in a shallow network. Following a similar "easy to hard" idea, we systematically investigate several practical transfer settings and show the effectiveness of transfer learning in low level vision problems. Our method shows superior performance than the state-of-the-arts both on the benchmark datasets and the real-world use cases (i.e. Twitter).
Chao Dong 0005, Yubin Deng, Chen Change Loy, Xiaoou Tang
ICCV4
2015 Semantic Image Segmentation via Deep Parsing Network
abstract
This paper addresses semantic image segmentation by incorporating rich information into Markov Random Field (MRF), including high-order relations and mixture of label contexts. Unlike previous works that optimized MRFs using iterative algorithm, we solve MRF by proposing a Convolutional Neural Network (CNN), namely Deep Parsing Network (DPN), which enables deterministic end-to-end computation in a single forward pass. Specifically, DPN extends a contemporary CNN architecture to model unary terms and additional layers are carefully devised to approximate the mean field algorithm (MF) for pairwise terms. It has several appealing properties. First, different from the recent works that combined CNN and MRF, where many iterations of MF were required for each training image during back-propagation, DPN is able to achieve high performance by approximating one iteration of MF. Second, DPN represents various types of pairwise terms, making many existing works as its special cases. Third, DPN makes MF easier to be parallelized and speeded up in Graphical Processing Unit (GPU). DPN is thoroughly evaluated on the PASCAL VOC 2012 dataset, where a single DPN model yields a new state-of-the-art segmentation accuracy of 77.5%.
Ziwei Liu 0002, Ping Luo 0002, Chen Change Loy, Xiaoou Tang
ICCV5
2015 Deep Learning Face Attributes in the Wild
abstract
Predicting face attributes in the wild is challenging due to complex face variations. We propose a novel deep learning framework for attribute prediction in the wild. It cascades two CNNs, LNet and ANet, which are fine-tuned jointly with attribute tags, but pre-trained differently. LNet is pre-trained by massive general object categories for face localization, while ANet is pre-trained by massive face identities for attribute prediction. This framework not only outperforms the state-of-the-art with a large margin, but also reveals valuable facts on learning face representation. (1) It shows how the performances of face localization (LNet) and attribute prediction (ANet) can be improved by different pre-training strategies. (2) It reveals that although the filters of LNet are fine-tuned only with image-level attribute tags, their response maps over entire images have strong indication of face locations. This fact enables training LNet for face localization with only image-level annotations, but without face bounding boxes or landmarks, which are required by all attribute recognition works. (3) It also demonstrates that the high-level hidden neurons of ANet automatically discover semantic concepts after pre-training with massive face identities, and such concepts are significantly enriched after fine-tuning with attribute tags. Each attribute can be well explained with a sparse linear combination of these concepts.
Ziwei Liu 0002, Ping Luo 0002, Xiaogang Wang 0001, Xiaoou Tang
ICCV4
2015 Deep Learning Strong Parts for Pedestrian Detection
abstract
Recent advances in pedestrian detection are attained by transferring the learned features of Convolutional Neural Network (ConvNet) to pedestrians. This ConvNet is typically pre-trained with massive general object categories (e.g. ImageNet). Although these features are able to handle variations such as poses, viewpoints, and lightings, they may fail when pedestrian images with complex occlusions are present. Occlusion handling is one of the most important problem in pedestrian detection. Unlike previous deep models that directly learned a single detector for pedestrian detection, we propose DeepParts, which consists of extensive part detectors. DeepParts has several appealing properties. First, DeepParts can be trained on weakly labeled data, i.e. only pedestrian bounding boxes without part annotations are provided. Second, DeepParts is able to handle low IoU positive proposals that shift away from ground truth. Third, each part detector in DeepParts is a strong detector that can detect pedestrian by observing only a part of a proposal. Extensive experiments in Caltech dataset demonstrate the effectiveness of DeepParts, which yields a new state-of-the-art miss rate of 11:89%, outperforming the second best method by 10%.
Yonglong Tian, Ping Luo 0002, Xiaogang Wang 0001, Xiaoou Tang
ICCV4
2015 From Facial Parts Responses to Face Detection: A Deep Learning Approach
abstract
In this paper, we propose a novel deep convolutional network (DCN) that achieves outstanding performance on FDDB, PASCAL Face, and AFW. Specifically, our method achieves a high recall rate of 90.99% on the challenging FDDB benchmark, outperforming the state-of-the-art method [23] by a large margin of 2.91%. Importantly, we consider finding faces from a new perspective through scoring facial parts responses by their spatial structure and arrangement. The scoring mechanism is carefully formulated considering challenging cases where faces are only partially visible. This consideration allows our network to detect faces under severe occlusion and unconstrained pose variation, which are the main difficulty and bottleneck of most existing face detection approaches. We show that despite the use of DCN, our network can achieve practical runtime speed.
Shuo Yang 0003, Ping Luo 0002, Chen Change Loy, Xiaoou Tang
ICCV4
2015 Learning Social Relation Traits from Face Images
abstract
Social relation defines the association, e.g., warm, friendliness, and dominance, between two or more people. Motivated by psychological studies, we investigate if such fine grained and high-level relation traits can be characterised and quantified from face images in the wild. To address this challenging problem we propose a deep model that learns a rich face representation to capture gender, expression, head pose, and age-related attributes, and then performs pairwise-face reasoning for relation prediction. To learn from heterogeneous attribute sources, we formulate a new network architecture with a bridging layer to leverage the inherent correspondences among these datasets. It can also cope with missing target attribute labels. Extensive experiments show that our approach is effective for fine-grained social relation learning in images and videos.
Ping Luo 0002, Chen Change Loy, Xiaoou Tang
ICCV4
2015 MIL: Music Exploration and Visualization via Lyric and Image
abstract
In this paper, we introduce MIL: a music exploration prototype which integrates music (M), image (I), and lyrics (L), for efficiently visualizing and browsing music collections. MIL utilizes a novel structure, music semantic graph (MSG), to organize music collections in a hierarchical way by leveraging lyrics and acoustic cues of music. Each node of MSG corresponds to a music concept and is associated with a cluster of music tracks. MIL offers users a novel way to efficiently explore and scan music collections by using lyrics and image information. In addition, the proposed prototype supplies an easy-to-use interface to visualize MSG hierarchically. The user study shows that our prototype can effectively and efficiently help users to browse, search, and scan music collections.
Xixuan Wu, Yu Qiao 0001, Xiaoou Tang
ACM Multimedia3
2015 Change-Based Image Cropping with Exclusion and Compositional Features
Jianzhou Yan, Stephen Lin 0001, Sing Bing Kang, Xiaoou Tang
Int. J. Comput. Vis.4
2015 Learning Collective Crowd Behaviors with Dynamic Pedestrian-Agents
Bolei Zhou, Xiaoou Tang, Xiaogang Wang 0001
Int. J. Comput. Vis.2
2015 Hierarchical facial landmark localization via cascaded random binary patterns
Wei Zhang 0081, Huijun Ding, Jianzhuang Liu, Xiaoou Tang
Pattern Recognit.5
2015 Depth From Water Reflection
abstract
The scene in a water reflection image often exhibits bilateral symmetry. In this paper, we design a framework to reconstruct the depth from a single water reflection image. This problem can be regarded as a special case of two-view stereo vision. It is challenging to obtain correspondences from the real scene and the mirror scene due to their large appearance difference. We first propose an appearance adaptation method to transform the appearance of the mirror scene so that it is much closer to the real scene. We then present a stereo matching algorithm to obtain the disparity map of the real scene. Compared with other depth-from-symmetry work that deals with man-made objects, our algorithm can recover the depth maps of a variety of scenes, where both natural and man-made objects may exist.
Jianzhuang Liu, Xiaoou Tang
IEEE Trans. Image Process.3
2014 Switchable Deep Network for Pedestrian Detection
abstract
In this paper, we propose a Switchable Deep Network (SDN) for pedestrian detection. The SDN automatically learns hierarchical features, salience maps, and mixture representations of different body parts. Pedestrian detection faces the challenges of background clutter and large variations of pedestrian appearance due to pose and viewpoint changes and other factors. One of our key contributions is to propose a Switchable Restricted Boltzmann Machine (SRBM) to explicitly model the complex mixture of visual variations at multiple levels. At the feature levels, it automatically estimates saliency maps for each test sample in order to separate background clutters from discriminative regions for pedestrian detection. At the part and body levels, it is able to infer the most appropriate template for the mixture models of each part and the whole body. We have devised a new generative algorithm to effectively pretrain the SDN and then fine-tune it with back-propagation. Our approach is evaluated on the Caltech and ETH datasets and achieves the state-of-the-art detection performance.
Ping Luo 0002, Yonglong Tian, Xiaogang Wang 0001, Xiaoou Tang
CVPR4
2014 Realtime and Robust Hand Tracking from Depth
abstract
We present a realtime hand tracking system using a depth sensor. It tracks a fully articulated hand under large viewpoints in realtime (25 FPS on a desktop without using a GPU) and with high accuracy (error below 10 mm). To our knowledge, it is the first system that achieves such robustness, accuracy, and speed simultaneously, as verified on challenging real data. Our system is made of several novel techniques. We model a hand simply using a number of spheres and define a fast cost function. Those are critical for realtime performance. We propose a hybrid method that combines gradient based and stochastic optimization methods to achieve fast convergence and good accuracy. We present new finger detection and hand initialization methods that greatly enhance the robustness of tracking.
Chen Qian 0006, Xiao Sun 0001, Xiaoou Tang, Jian Sun 0001
CVPR4
2014 Deep Learning Face Representation from Predicting 10, 000 Classes
abstract
This paper proposes to learn a set of high-level feature representations through deep learning, referred to as Deep hidden IDentity features (DeepID), for face verification. We argue that DeepID can be effectively learned through challenging multi-class face identification tasks, whilst they can be generalized to other tasks (such as verification) and new identities unseen in the training set. Moreover, the generalization capability of DeepID increases as more face classes are to be predicted at training. DeepID features are taken from the last hidden layer neuron activations of deep convolutional networks (ConvNets). When learned as classifiers to recognize about 10, 000 face identities in the training set and configured to keep reducing the neuron numbers along the feature extraction hierarchy, these deep ConvNets gradually form compact identity-related features in the top layers with only a small number of hidden neurons. The proposed features are extracted from various face regions to form complementary and over-complete representations. Any state-of-the-art classifiers can be learned based on these high-level representations for face verification. 97:45% verification accuracy on LFW is achieved with only weakly aligned faces.
Xiaogang Wang 0001, Xiaoou Tang
CVPR3
2014 A Learning-to-Rank Approach for Image Color Enhancement
abstract
We present a machine-learned ranking approach for automatically enhancing the color of a photograph. Unlike previous techniques that train on pairs of images before and after adjustment by a human user, our method takes into account the intermediate steps taken in the enhancement process, which provide detailed information on the person's color preferences. To make use of this data, we formulate the color enhancement task as a learning-to-rank problem in which ordered pairs of images are used for training, and then various color enhancements of a novel input image can be evaluated from their corresponding rank values. From the parallels between the decision tree structures we use for ranking and the decisions made by a human during the editing process, we posit that breaking a full enhancement sequence into individual steps can facilitate training. Our experiments show that this approach compares well to existing methods for automatic color enhancement.
Jianzhou Yan, Stephen Lin 0001, Sing Bing Kang, Xiaoou Tang
CVPR4
2014 Learning a Deep Convolutional Network for Image Super-Resolution
Chao Dong 0005, Chen Change Loy, Kaiming He, Xiaoou Tang
ECCV (4)4
2014 Robust Scene Text Detection with Convolution Neural Network Induced MSER Trees
Yu Qiao 0001, Xiaoou Tang
ECCV (4)3
2014 Learning the Face Prior for Bayesian Face Recognition
Chaochao Lu, Xiaoou Tang
ECCV (4)2
2014 Video Action Detection with Relational Dynamic-Poselets
Limin Wang 0002, Yu Qiao 0001, Xiaoou Tang
ECCV (5)3
2014 Object Detection and Viewpoint Estimation with Auto-masking Neural Network
Jianzhuang Liu, Xiaoou Tang
ECCV (3)3
2014 Facial Landmark Detection by Deep Multi-task Learning
Ping Luo 0002, Chen Change Loy, Xiaoou Tang
ECCV (6)4
2014 Pedestrian Attribute Recognition At Far Distance
abstract
The capability of recognizing pedestrian attributes, such as gender and clothing style, at far distance, is of practical interest in far-view surveillance scenarios where face and body close-shots are hardly available. We make two contributions in this paper. First, we release a new pedestrian attribute dataset, which is by far the largest and most diverse of its kind. We show that the large-scale dataset facilitates the learning of robust attribute detectors with good generalization performance. Second, we present the benchmark performance by SVM-based method and propose an alternative approach that exploits context of neighboring pedestrian images for improved attribute inference.
Yubin Deng, Ping Luo 0002, Chen Change Loy, Xiaoou Tang
ACM Multimedia4
2014 Fusing Music and Video Modalities Using Multi-timescale Shared Representations
abstract
We propose a deep learning architecture to solve the problem of multimodal fusion of multi-timescale temporal data, using music and video parts extracted from Music Videos (MVs) in particular. We capture the correlations between music and video at multiple levels by learning shared feature representations with Deep Belief Networks (DBN). The shared representations combine information from multiple modalities for decision making tasks, and are used to evaluate matching degrees between modalities and to retrieve matched modalities using single or multiple modalities as input. Moreover, we propose a novel deep architecture to handle temporal data at multiple timescales. When processing long sequences with varying length, we propose to extract hierarchical shared representations by concatenating deep representations at different levels, and to perform decision fusion with a feed forward neural network, which takes input from predictions of local and global classifiers trained with shared representations at each level. The effectiveness of our method is demonstrated through MV classification and retrieval.
Xiaogang Wang 0001, Xiaoou Tang
ACM Multimedia3
2014 Orthogonal Gaussian Process for Automatic Age Estimation
abstract
Age Estimation from facial images has been receiving increasing interest due to its important applications. Among the existing age estimation algorithms, the personalized approaches have been shown to be the most effective ones. However, most of the person-specific approaches (e.g. MTWGP [1], AGES [2], WAS [3]) rely heavily on the availability of training images across different ages for a single subject, which is very difficult to satisfy in practical applications. In order to overcome this problem, we propose a new approach to age estimation, called Orthogonal Gaussian Process (OGP). Compared to standard Gaussian Process, OGP is much more efficient while maintaining the discriminatory power of the standard Gaussian Process. Based on OGP, we further propose an improvement of OGP called anisotropic OGP (A-OGP) to enhance the age estimation performance. Extensive experiments are conducted to demonstrate the state-of-the-art estimation accuracy of our new algorithm on several public-domain face aging datasets: FG-NET face dataset with 82 different subjects, Morph Album 1 dataset with more than 600 subjects, and Morph Album 2 with about 20,000 different subjects.
Dihong Gong, Zhifeng Li 0001, Xiaoou Tang
ACM Multimedia4
2014 Deep Learning Face Representation by Joint Identification-Verification
Xiaogang Wang 0001, Xiaoou Tang
NIPS4
2014 Zeta Hull Pursuits: Learning Nonconvex Data Hulls
Yuanjun Xiong, Wei Liu 0005, Deli Zhao, Xiaoou Tang
NIPS4
2014 Multi-View Perceptron: a Deep Model for Learning Face Identity and View Representations
Zhenyao Zhu, Ping Luo 0002, Xiaogang Wang 0001, Xiaoou Tang
NIPS4
2014 Pairwise Rotation Invariant Co-Occurrence Local Binary Pattern
abstract
Designing effective features is a fundamental problem in computer vision. However, it is usually difficult to achieve a great tradeoff between discriminative power and robustness. Previous works shown that spatial co-occurrence can boost the discriminative power of features. However the current existing co-occurrence features are taking few considerations to the robustness and hence suffering from sensitivity to geometric and photometric variations. In this work, we study the Transform Invariance (TI) of co-occurrence features. Concretely we formally introduce a Pairwise Transform Invariance (PTI) principle, and then propose a novel Pairwise Rotation Invariant Co-occurrence Local Binary Pattern (PRICoLBP) feature, and further extend it to incorporate multi-scale, multi-orientation, and multi-channel information. Different from other LBP variants, PRICoLBP can not only capture the spatial context co-occurrence information effectively, but also possess rotation invariance. We evaluate PRICoLBP comprehensively on nine benchmark data sets from five different perspectives, e.g., encoding strategy, rotation invariance, the number of templates, speed, and discriminative power compared to other LBP variants. Furthermore we apply PRICoLBP to six different but related applications-texture, material, flower, leaf, food, and scene classification, and demonstrate that PRICoLBP is efficient, effective, and of a well-balanced tradeoff between the discriminative power and robustness.
Xianbiao Qi, Rong Xiao 0003, Chun-Guang Li, Yu Qiao 0001, Jun Guo 0002, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.6
2014 Web Image Re-Ranking UsingQuery-Specific Semantic Signatures
abstract
Image re-ranking, as an effective way to improve the results of web-based image search, has been adopted by current commercial search engines such as Bing and Google. Given a query keyword, a pool of images are first retrieved based on textual information. By asking the user to select a query image from the pool, the remaining images are re-ranked based on their visual similarities with the query image. A major challenge is that the similarities of visual features do not well correlate with images' semantic meanings which interpret users' search intention. Recently people proposed to match images in a semantic space which used attributes or reference classes closely related to the semantic meanings of images as basis. However, learning a universal visual semantic space to characterize highly diverse images from the web is difficult and inefficient. In this paper, we propose a novel image re-ranking framework, which automatically offline learns different semantic spaces for different query keywords. The visual features of images are projected into their related semantic spaces to get semantic signatures. At the online stage, images are re-ranked by comparing their semantic signatures obtained from the semantic space specified by the query keyword. The proposed query-specific semantic signatures significantly improve both the accuracy and efficiency of image re-ranking. The original visual features of thousands of dimensions can be projected to the semantic signatures as short as 25 dimensions. Experimental results show that 25-40 percent relative improvement has been achieved on re-ranking precisions compared with the state-of-the-art methods.
Xiaogang Wang 0001, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.4
2014 Measuring Crowd Collectiveness
abstract
Collective motions of crowds are common in nature and have attracted a great deal of attention in a variety of multidisciplinary fields. Collectiveness, which indicates the degree of individuals acting as a union, is a fundamental and universal measurement for various crowd systems. By quantifying the topological structures of collective manifolds of crowd, this paper proposes a descriptor of collectiveness and its efficient computation for the crowd and its constituent individuals. The Collective Merging algorithm is then proposed to detect collective motions from random motions. We validate the effectiveness and robustness of the proposed collectiveness on the system of self-driven particles as well as other real crowd systems such as pedestrian crowds and bacteria colony. We compare the collectiveness descriptor with human perception for collective motion and show their high consistency. As a universal descriptor, the proposed crowd collectiveness can be used to compare different crowd systems. It has a wide range of applications, such as detecting collective motions from crowd clutters, monitoring crowd dynamics, and generating maps of collectiveness for crowded scenes. A new Collective Motion Database, which consists of 413 video clips from 62 crowded scenes, is released to the public.
Bolei Zhou, Xiaoou Tang, Hepeng Zhang, Xiaogang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2014 Multiview Facial Landmark Localization in RGB-D Images via Hierarchical Regression With Binary Patterns
abstract
In this paper, we propose a real-time system of multiview facial landmark localization in RGB-D images. The facial landmark localization problem is formulated into a regression framework, which estimates both the head pose and the landmark positions. In this framework, we propose a coarse-to-fine approach to handle the high-dimensional regression output. At first, 3-D face position and rotation are estimated from the depth observation via a random regression forest. Afterward, the 3-D pose is refined by fusing the estimation from the RGB observation. Finally, the landmarks are located from the RGB observation with gradient boosted decision trees in a pose conditional model. The benefits of the proposed localization framework are twofold: the pose estimation and landmark localization are solved with hierarchical regression, which is different from previous approaches where the pose and landmark locations are iteratively optimized, which relies heavily on the initial pose estimation; due to the different characters of the RGB and depth cues, they are used for landmark localization at different stages and incorporated in a robust manner. In the experiments, we show that the proposed approach outperforms state-of-the-art algorithms on facial landmark localization with RGB-D input.
Wei Zhang 0081, Jianzhuang Liu, Xiaoou Tang
IEEE Trans. Circuits Syst. Video Technol.4
2014 Latent Hierarchical Model of Temporal Structure for Complex Activity Classification
abstract
Modeling the temporal structure of sub-activities is an important yet challenging problem in complex activity classification. This paper proposes a latent hierarchical model (LHM) to describe the decomposition of complex activity into sub-activities in a hierarchical way. The LHM has a tree-structure, where each node corresponds to a video segment (sub-activity) at certain temporal scale. The starting and ending time points of each sub-activity are represented by two latent variables, which are automatically determined during the inference process. We formulate the training problem of the LHM in a latent kernelized SVM framework and develop an efficient cascade inference method to speed up classification. The advantages of our methods come from: 1) LHM models the complex activity with a deep structure, which is decomposed into sub-activities in a coarse-to-fine manner and 2) the starting and ending time points of each segment are adaptively determined to deal with the temporal displacement and duration variation of sub-activity. We conduct experiments on three datasets: 1) the KTH; 2) the Hollywood2; and 3) the Olympic Sports. The experimental results show the effectiveness of the LHM in complex activity classification. With dense features, our LHM achieves the state-of-the-art performance on the Hollywood2 dataset and the Olympic Sports dataset.
Limin Wang 0002, Yu Qiao 0001, Xiaoou Tang
IEEE Trans. Image Process.3
2014 Fast burst images denoising
abstract
This paper presents a fast denoising method that produces a clean image from a burst of noisy images. We accelerate alignment of the images by introducing a lightweight camera motion representation called homography flow . The aligned images are then fused to create a denoised output with rapid per-pixel operations in temporal and spatial domains. To handle scene motion during the capture, a mechanism of selecting consistent pixels for temporal fusion is proposed to "synthesize" a clean, ghost-free image, which can largely reduce the computation of tracking motion between frames. Combined with these efficient solutions, our method runs several orders of magnitude faster than previous work, while the denoising quality is comparable. A smartphone prototype demonstrates that our method is practical and works well on a large variety of real examples.
Ziwei Liu 0002, Lu Yuan 0001, Xiaoou Tang, Matthew Uyttendaele, Jian Sun 0001
ACM Trans. Graph.3
2013 Deep Convolutional Network Cascade for Facial Point Detection
abstract
We propose a new approach for estimation of the positions of facial key points with three-level carefully designed convolutional networks. At each level, the outputs of multiple networks are fused for robust and accurate estimation. Thanks to the deep structures of convolutional networks, global high-level features are extracted over the whole face region at the initialization stage, which help to locate high accuracy key points. There are two folds of advantage for this. First, the texture context information over the entire face is utilized to locate each key point. Second, since the networks are trained to predict all the key points simultaneously, the geometric constraints among key points are implicitly encoded. The method therefore can avoid local minimum caused by ambiguity and data corruption in difficult image samples due to occlusions, large pose variations, and extreme lightings. The networks at the following two levels are trained to locally refine initial predictions and their inputs are limited to small regions around the initial predictions. Several network structures critical for accurate and robust facial point detection are investigated. Extensive experiments show that our approach outperforms state-of-the-art methods in both detection accuracy and reliability.
Xiaogang Wang 0001, Xiaoou Tang
CVPR3
2013 Motionlets: Mid-level 3D Parts for Human Motion Recognition
abstract
This paper proposes motionlet, a mid-level and spatiotemporal part, for human motion recognition. Motion let can be seen as a tight cluster in motion and appearance space, corresponding to the moving process of different body parts. We postulate three key properties of motion let for action recognition: high motion saliency, multiple scale representation, and representative-discriminative ability. Towards this goal, we develop a data-driven approach to learn motion lets from training videos. First, we extract 3D regions with high motion saliency. Then we cluster these regions and preserve the centers as candidate templates for motion let. Finally, we examine the representative and discriminative power of the candidates, and introduce a greedy method to select effective candidates. With motion lets, we present a mid-level representation for video, called motionlet activation vector. We conduct experiments on three datasets, KTH, HMDB51, and UCF50. The results show that the proposed methods significantly outperform state-of-the-art methods.
Limin Wang 0002, Yu Qiao 0001, Xiaoou Tang
CVPR3
2013 Learning the Change for Automatic Image Cropping
abstract
Image cropping is a common operation used to improve the visual quality of photographs. In this paper, we present an automatic cropping technique that accounts for the two primary considerations of people when they crop: removal of distracting content, and enhancement of overall composition. Our approach utilizes a large training set consisting of photos before and after cropping by expert photographers to learn how to evaluate these two factors in a crop. In contrast to the many methods that exist for general assessment of image quality, ours specifically examines differences between the original and cropped photo in solving for the crop parameters. To this end, several novel image features are proposed to model the changes in image content and composition when a crop is applied. Our experiments demonstrate improvements of our method over recent cropping algorithms on a broad range of images.
Jianzhou Yan, Stephen Lin 0001, Sing Bing Kang, Xiaoou Tang
CVPR4
2013 Measuring Crowd Collectiveness
abstract
Collective motions are common in crowd systems and have attracted a great deal of attention in a variety of multidisciplinary fields. Collectiveness, which indicates the degree of individuals acting as a union in collective motion, is a fundamental and universal measurement for various crowd systems. By integrating path similarities among crowds on collective manifold, this paper proposes a descriptor of collectiveness and an efficient computation for the crowd and its constituent individuals. The algorithm of the Collective Merging is then proposed to detect collective motions from random motions. We validate the effectiveness and robustness of the proposed collectiveness descriptor on the system of self-driven particles. We then compare the collectiveness descriptor to human perception for collective motion and show high consistency. Our experiments regarding the detection of collective motions and the measurement of collectiveness in videos of pedestrian crowds and bacteria colony demonstrate a wide range of applications of the collectiveness descriptor.
Bolei Zhou, Xiaoou Tang, Xiaogang Wang 0001
CVPR2
2013 Hidden Factor Analysis for Age Invariant Face Recognition
abstract
Age invariant face recognition has received increasing attention due to its great potential in real world applications. In spite of the great progress in face recognition techniques, reliably recognizing faces across ages remains a difficult task. The facial appearance of a person changes substantially over time, resulting in significant intra-class variations. Hence, the key to tackle this problem is to separate the variation caused by aging from the person-specific features that are stable. Specifically, we propose a new method, called Hidden Factor Analysis (HFA). This method captures the intuition above through a probabilistic model with two latent factors: an identity factor that is age-invariant and an age factor affected by the aging process. Then, the observed appearance can be modeled as a combination of the components generated based on these factors. We also develop a learning algorithm that jointly estimates the latent factors and the model parameters using an EM procedure. Extensive experiments on two well-known public domain face aging datasets: MORPH (the largest public face aging database) and FGNET, clearly show that the proposed method achieves notable improvement over state-of-the-art algorithms.
Dihong Gong, Zhifeng Li 0001, Dahua Lin, Jianzhuang Liu, Xiaoou Tang
ICCV5
2013 Face Recognition Using Face Patch Networks
abstract
When face images are taken in the wild, the large variations in facial pose, illumination, and expression make face recognition challenging. The most fundamental problem for face recognition is to measure the similarity between faces. The traditional measurements such as various mathematical norms, Hausdorff distance, and approximate geodesic distance cannot accurately capture the structural information between faces in such complex circumstances. To address this issue, we develop a novel face patch network, based on which we define a new similarity measure called the random path (RP) measure. The RP measure is derived from the collective similarity of paths by performing random walks in the network. It can globally characterize the contextual and curved structures of the face space. To apply the RP measure, we construct two kinds of networks: the in-face network and the out-face network. The in-face network is drawn from any two face images and captures the local structural information. The out-face network is constructed from all the training face patches, thereby modeling the global structures of face space. The two face networks are structurally complementary and can be combined together to improve the recognition performance. Experiments on the Multi-PIE and LFW benchmarks show that the RP measure outperforms most of the state-of-art algorithms for face recognition.
Chaochao Lu, Deli Zhao, Xiaoou Tang
ICCV3
2013 Pedestrian Parsing via Deep Decompositional Network
abstract
We propose a new Deep Decompositional Network (DDN) for parsing pedestrian images into semantic regions, such as hair, head, body, arms, and legs, where the pedestrians can be heavily occluded. Unlike existing methods based on template matching or Bayesian inference, our approach directly maps low-level visual features to the label maps of body parts with DDN, which is able to accurately estimate complex pose variations with good robustness to occlusions and background clutters. DDN jointly estimates occluded regions and segments body parts by stacking three types of hidden layers: occlusion estimation layers, completion layers, and decomposition layers. The occlusion estimation layers estimate a binary mask, indicating which part of a pedestrian is invisible. The completion layers synthesize low-level features of the invisible part from the original features and the occlusion mask. The decomposition layers directly transform the synthesized visual features to label maps. We devise a new strategy to pre-train these hidden layers, and then fine-tune the entire network using the stochastic gradient descent. Experimental results show that our approach achieves better segmentation accuracy than the state-of-the-art methods on pedestrian images with or without occlusions. Another important contribution of this paper is that it provides a large scale benchmark human parsing dataset that includes 3,673 annotated samples collected from 171 surveillance videos. It is 20 times larger than existing public datasets.
Ping Luo 0002, Xiaogang Wang 0001, Xiaoou Tang
ICCV3
2013 A Deep Sum-Product Architecture for Robust Facial Attributes Analysis
abstract
Recent works have shown that facial attributes are useful in a number of applications such as face recognition and retrieval. However, estimating attributes in images with large variations remains a big challenge. This challenge is addressed in this paper. Unlike existing methods that assume the independence of attributes during their estimation, our approach captures the interdependencies of local regions for each attribute, as well as the high-order correlations between different attributes, which makes it more robust to occlusions and misdetection of face regions. First, we have modeled region interdependencies with a discriminative decision tree, where each node consists of a detector and a classifier trained on a local region. The detector allows us to locate the region, while the classifier determines the presence or absence of an attribute. Second, correlations of attributes and attribute predictors are modeled by organizing all of the decision trees into a large sum-product network (SPN), which is learned by the EM algorithm and yields the most probable explanation (MPE) of the facial attributes in terms of the region's localization and classification. Experimental results on a large data set with 22,400 images show the effectiveness of the proposed approach.
Ping Luo 0002, Xiaogang Wang 0001, Xiaoou Tang
ICCV3
2013 Visual Semantic Complex Network for Web Images
abstract
This paper proposes modeling the complex web image collections with an automatically generated graph structure called visual semantic complex network (VSCN). The nodes on this complex network are clusters of images with both visual and semantic consistency, called semantic concepts. These nodes are connected based on the visual and semantic correlations. Our VSCN with 33,240 concepts is generated from a collection of 10 million web images. A great deal of valuable information on the structures of the web image collections can be revealed by exploring the VSCN, such as the small-world behavior, concept community, in-degree distribution, hubs, and isolated concepts. It not only helps us better understand the web image collections at a macroscopic level, but also has many important practical applications. This paper presents two application examples: content-based image retrieval and image browsing. Experimental results show that the VSCN leads to significant improvement on both the precision of image retrieval (over 200%) and user experience for image browsing.
Xiaogang Wang 0001, Xiaoou Tang
ICCV3
2013 Hybrid Deep Learning for Face Verification
abstract
This paper proposes a hybrid convolutional network (ConvNet)-Restricted Boltzmann Machine (RBM) model for face verification in wild conditions. A key contribution of this work is to directly learn relational visual features, which indicate identity similarities, from raw pixels of face pairs with a hybrid deep network. The deep ConvNets in our model mimic the primary visual cortex to jointly extract local relational visual features from two face images compared with the learned filter pairs. These relational features are further processed through multiple layers to extract high-level and global features. Multiple groups of ConvNets are constructed in order to achieve robustness and characterize face similarities from different aspects. The top-layer RBM performs inference from complementary high-level features extracted from different ConvNet groups with a two-level average pooling hierarchy. The entire hybrid deep network is jointly fine-tuned to optimize for the task of face verification. Our model achieves competitive face verification performance on the LFW dataset.
Xiaogang Wang 0001, Xiaoou Tang
ICCV3
2013 Mining Motion Atoms and Phrases for Complex Action Recognition
abstract
This paper proposes motion atom and phrase as a mid-level temporal ``part'' for representing and classifying complex action. Motion atom is defined as an atomic part of action, and captures the motion information of action video in a short temporal scale. Motion phrase is a temporal composite of multiple motion atoms with an AND/OR structure, which further enhances the discriminative ability of motion atoms by incorporating temporal constraints in a longer scale. Specifically, given a set of weakly labeled action videos, we firstly design a discriminative clustering method to automatically discover a set of representative motion atoms. Then, based on these motion atoms, we mine effective motion phrases with high discriminative and representative power. We introduce a bottom-up phrase construction algorithm and a greedy selection method for this mining task. We examine the classification performance of the motion atom and phrase based representation on two complex action datasets: Olympic Sports and UCF50. Experimental results show that our method achieves superior performance over recent published methods on both datasets.
Limin Wang 0002, Yu Qiao 0001, Xiaoou Tang
ICCV3
2013 Face Recognition via Archetype Hull Ranking
abstract
The archetype hull model is playing an important role in large-scale data analytics and mining, but rarely applied to vision problems. In this paper, we migrate such a geometric model to address face recognition and verification together through proposing a unified archetype hull ranking framework. Upon a scalable graph characterized by a compact set of archetype exemplars whose convex hull encompasses most of the training images, the proposed framework explicitly captures the relevance between any query and the stored archetypes, yielding a rank vector over the archetype hull. The archetype hull ranking is then executed for every block of face images to generate a block wise similarity measure that is achieved by comparing two different rank vectors with respect to the same archetype hull. After integrating block wise similarity measurements with learned importance weights, we accomplish a sensible face similarity measure which can support robust and effective face recognition and verification. We evaluate the face similarity measure in terms of experiments performed on three benchmark face databases Multi-PIE, Pubfig83, and LFW, demonstrating its performance superior to the state-of-the-arts.
Yuanjun Xiong, Wei Liu 0005, Deli Zhao, Xiaoou Tang
ICCV4
2013 Complex 3D General Object Reconstruction from Line Drawings
abstract
An important topic in computer vision is 3D object reconstruction from line drawings. Previous algorithms either deal with simple general objects or are limited to only manifolds (a subset of solids). In this paper, we propose a novel approach to 3D reconstruction of complex general objects, including manifolds, non-manifold solids, and nonsolids. Through developing some 3D object properties, we use the degree of freedom of objects to decompose a complex line drawing into multiple simpler line drawings that represent meaningful building blocks of a complex object. After 3D objects are reconstructed from the decomposed line drawings, they are merged to form a complex object from their touching faces, edges, and vertices. Our experiments show a number of reconstruction examples from both complex line drawings and images with line drawings superimposed. Comparisons are also given to indicate that our algorithm can deal with much more complex line drawings of general objects than previous algorithms.
Jianzhuang Liu, Xiaoou Tang
ICCV3
2013 Deep Learning Identity-Preserving Face Space
abstract
Face recognition with large pose and illumination variations is a challenging problem in computer vision. This paper addresses this challenge by proposing a new learning based face representation: the face identity-preserving (FIP) features. Unlike conventional face descriptors, the FIP features can significantly reduce intra-identity variances, while maintaining discriminative ness between identities. Moreover, the FIP features extracted from an image under any pose and illumination can be used to reconstruct its face image in the canonical view. This property makes it possible to improve the performance of traditional descriptors, such as LBP [2] and Gabor [31], which can be extracted from our reconstructed images in the canonical view to eliminate variations. In order to learn the FIP features, we carefully design a deep network that combines the feature extraction layers and the reconstruction layer. The former encodes a face image into the FIP features, while the latter transforms them to an image in the canonical view. Extensive experiments on the large MultiPIE face database [7] demonstrate that it significantly outperforms the state-of-the-art face recognition methods.
Zhenyao Zhu, Ping Luo 0002, Xiaogang Wang 0001, Xiaoou Tang
ICCV4
2013 AdVisual: a visual-based advertising system
abstract
In this work, we present a visual-based contextual advertising system, AdVisual. It is designed for content service providers to effectively select relevant ads for online videos. First, it will analyze each video and extract high level semantic visual information including specific objects, people, and significant scenes. Then ads highly related to the visual concepts are associated with the corresponding shots. As AdVisual is a user interaction system, it allows users to select favorite ads relevant to the video. By saliency detection, selected ads will be displayed as an overlay window embedded at the non-intrusive part of the shot.
Chao Dong 0005, Shifeng Chen, Xiaoou Tang
ACM Multimedia3
2013 Anchor concept graph distance for web image re-ranking
abstract
Web image re-ranking aims to automatically refine the initial text-based image search results by employing visual information. A strong line of work in image re-ranking relies on building image graphs that requires computing distances between image pairs. In this paper, we present Anchor Concept Graph Distance (ACG Distance), a novel distance measure for image re-ranking. For a given textual query, an Anchor Concept Graph (ACG) is automatically learned from the initial text-based search results. The nodes of the ACG (i.e., anchor concepts) and their correlations well model the semantic structure of the images to be re-ranked. Images are projected to the anchor concepts. The projection vectors undergo a diffusion process over the ACG, and then are used to compute the ACG distance. The ACG distance reduces the semantic gap and better represents distances between images. Experiments on the MSRA-MM and INRIA datasets show that the ACG distance consistently outperforms existing distance measures and significantly improves start-of-the-art methods in image re-ranking.
Xiaogang Wang 0001, Xiaoou Tang
ACM Multimedia3
2013 Facial landmark localization based on hierarchical pose regression with cascaded random ferns
abstract
The main challenge of facial landmark localization in real-world application is that the large changes of head pose and facial expressions cause substantial image appearance variations. To avoid high dimensional regression in the 3D and 2D facial pose spaces simultaneously, we propose a hierarchical pose regression approach, estimating the head rotation, facial components and landmarks hierarchically. The regression process works in a unified cascaded fern framework. We present generalized gradient boosted ferns (GBFs) for the regression framework, which give better performance than traditional ferns. The framework also achieves real time performance. We verify our method on the latest benchmark datasets. The results show that it outperforms state-of-the-art methods in both accuracy and speed.
Wei Zhang 0081, Jianzhuang Liu, Xiaoou Tang
ACM Multimedia4
2013 Guided Image Filtering
abstract
In this paper, we propose a novel explicit image filter called guided filter. Derived from a local linear model, the guided filter computes the filtering output by considering the content of a guidance image, which can be the input image itself or another different image. The guided filter can be used as an edge-preserving smoothing operator like the popular bilateral filter [1], but it has better behaviors near edges. The guided filter is also a more generic concept beyond smoothing: It can transfer the structures of the guidance image to the filtering output, enabling new filtering applications like dehazing and guided feathering. Moreover, the guided filter naturally has a fast and nonapproximate linear time algorithm, regardless of the kernel size and the intensity range. Currently, it is one of the fastest edge-preserving filters. Experiments show that the guided filter is both effective and efficient in a great variety of computer vision and computer graphics applications, including edge-aware smoothing, detail enhancement, HDR compression, image matting/feathering, dehazing, joint upsampling, etc.
Kaiming He, Jian Sun 0001, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Image Transformation Based on Learning Dictionaries across Image Spaces
abstract
In this paper, we propose a framework of transforming images from a source image space to a target image space, based on learning coupled dictionaries from a training set of paired images. The framework can be used for applications such as image super-resolution and estimation of image intrinsic components (shading and albedo). It is based on a local parametric regression approach, using sparse feature representations over learned coupled dictionaries across the source and target image spaces. After coupled dictionary learning, sparse coefficient vectors of training image patch pairs are partitioned into easily retrievable local clusters. For any test image patch, we can fast index into its closest local cluster and perform a local parametric regression between the learned sparse feature spaces. The obtained sparse representation (together with the learned target space dictionary) provides multiple constraints for each pixel of the target image to be estimated. The final target image is reconstructed based on these constraints. The contributions of our proposed framework are three-fold. 1) We propose a concept of coupled dictionary learning based on coupled sparse coding which requires the sparse coefficient vectors of a pair of corresponding source and target image patches to have the same support, i.e., the same indices of nonzero elements. 2) We devise a space partitioning scheme to divide the high-dimensional but sparse feature space into local clusters. The partitioning facilitates extremely fast retrieval of closest local clusters for query patches. 3) Benefiting from sparse feature-based image transformation, our method is more robust to corrupted input data, and can be considered as a simultaneous image restoration and transformation process. Experiments on intrinsic image estimation and super-resolution demonstrate the effectiveness and efficiency of our proposed method.
Kui Jia, Xiaogang Wang 0001, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Learning Semantic Signatures for 3D Object Retrieval
abstract
In this paper, we propose two kinds of semantic signatures for 3D object retrieval (3DOR). Humans are capable of describing an object using attribute terms like “symmetric” and “flyable”, or using its similarities to some known object classes. We convert such qualitative descriptions into attribute signature (AS) and reference set signature (RSS), respectively, and use them for 3DOR. We also show that AS and RSS can be understood as two different quantization methods of the same semantic space of human descriptions of objects. The advantages of the semantic signatures are threefold. First, they are much more compact than low-level shape features yet working with comparable retrieval accuracy. Therefore, the proposed semantic signatures require less storage space and computation cost in retrieval. Second, the high-level signatures are a good complement to low-level shape features. As a result, by incorporating the signatures we can improve the performance of state-of-the-art 3DOR methods by a large margin. To the best of our knowledge, we obtain the best results on two popular benchmarks. Third, the AS enables us to build a user-friendly interface, with which the user can trigger a search by simply clicking attribute bars instead of finding a 3D object as the query. This interface is of great significance in 3DOR considering the fact that while searching, the user usually does not have a 3D query at hand that is similar to his/her targeted objects in the database.
Boqing Gong, Jianzhuang Liu, Xiaogang Wang 0001, Xiaoou Tang
IEEE Trans. Multim.4
2013 Content-Based Photo Quality Assessment
abstract
Automatically assessing photo quality from the perspective of visual aesthetics is of great interest in high-level vision research and has drawn much attention in recent years. In this paper, we propose content-based photo quality assessment using both regional and global features. Under this framework, subject areas, which draw the most attentions of human eyes, are first extracted. Then regional features extracted from both subject areas and background regions are combined with global features to assess photo quality. Since professional photographers adopt different photographic techniques and have different aesthetic criteria in mind when taking different types of photos (e.g., landscape versus portrait), we propose to segment subject areas and extract visual features in different ways according to the variety of photo content. We divide the photos into seven categories based on their visual content and develop a set of new subject area extraction methods and new visual features specially designed for different categories. The effectiveness of this framework is supported by extensive experimental comparisons of existing photo quality assessment approaches as well as our new features on different categories of photos. In addition, we propose an approach of online training an adaptive classifier to combine the proposed features according to the visual content of a test photo without knowing its category. Another contribution of this work is to construct a large and diversified benchmark dataset for the research of photo quality assessment. It includes 17,673 photos with manually labeled ground truth. This new benchmark dataset can be down loaded at http://mmlab.ie.cuhk.edu.hk/CUHKPQ/Dataset.htm.
Xiaoou Tang, Xiaogang Wang 0001
IEEE Trans. Multim.1
2013 Style Transfer Via Image Component Analysis
abstract
Example-based stylization provides an easy way of making artistic effects for images and videos. However, most existing methods do not consider the content and style separately. In this paper, we propose a style transfer algorithm via a novel component analysis approach, based on various image processing techniques. First, inspired by the steps of drawing a picture, an image is decomposed into three components: draft, paint and edge, which describe the content, main style, and strengthened strokes along the boundaries. Then the style is transferred from the template image to the source image in the paint and edge components. Style transfer is formulated as a global optimization problem by using Markov random fields, and a coarse-to-fine belief propagation algorithm is used to solve the optimization problem. To combine the draft component and the obtained style information, the final artistic result can be achieved via a reconstruction step. Compared to other algorithms, our method not only synthesizes the style, but also preserves the image content well. We also extend our algorithm from single image stylization to video personalization, by maintaining the temporal coherence and identifying faces in video sequences. The results indicate that our approach performs excellently in stylization and personalization for images and videos.
Wayne Zhang 0001, Shifeng Chen, Jianzhuang Liu, Xiaoou Tang
IEEE Trans. Multim.5
2012 Synthesizing oil painting surface geometry from a single photograph
abstract
We present an approach to synthesize the subtle 3D relief and texture of oil painting brush strokes from a single photograph. This task is unique from traditional synthesize algorithms due to its mixed modality between the input and output; i.e., our goal is to synthesize surface normals given an intensity image input. To accomplish this task, we propose a framework that first applies intrinsic image decomposition to produce a pair of initial normal maps. These maps are combined into a conditional random field (CRF) optimization framework that incorporates additional information derived from a training set consisting of normals captured using photometric stereo on oil paintings with similar brush styles. Additional constraints are incorporated into the CRF framework to further ensures smoothness and preserve brush stroke edges. Our results show that this approach can produce compelling reliefs that are often indistinguishable from results captured using photometric stereo.
Zheng Lu 0002, Xiaogang Wang 0001, Ying-Qing Xu, Moshe Ben-Ezra, Xiaoou Tang, Michael S. Brown
CVPR6
2012 Hierarchical face parsing via deep learning
abstract
This paper investigates how to parse (segment) facial components from face images which may be partially occluded. We propose a novel face parser, which recasts segmentation of face components as a cross-modality data transformation problem, i.e., transforming an image patch to a label map. Specifically, a face is represented hierarchically by parts, components, and pixel-wise labels. With this representation, our approach first detects faces at both the part- and component-levels, and then computes the pixel-wise label maps (Fig.1). Our part-based and component-based detectors are generatively trained with the deep belief network (DBN), and are discriminatively tuned by logistic regression. The segmentators transform the detected face components to label maps, which are obtained by learning a highly nonlinear mapping with the deep autoencoder. The proposed hierarchical face parsing is not only robust to partial occlusions but also provide richer information for face analysis and face synthesis compared with face keypoint detection and face alignment. The effectiveness of our algorithm is shown through several tasks on 2, 239 images selected from three datasets (e.g., LFW [12], BioID [13] and CUFSF [29]).
Ping Luo 0002, Xiaogang Wang 0001, Xiaoou Tang
CVPR3
2012 Example-based 3D object reconstruction from line drawings
abstract
Recovering 3D geometry from a single 2D line drawing is an important and challenging problem in computer vision. It has wide applications in interactive 3D modeling from images, computer-aided design, and 3D object retrieval. Previous methods of 3D reconstruction from line drawings are mainly based on a set of heuristic rules. They are not robust to sketch errors and often fail for objects that do not satisfy the rules. In this paper, we propose a novel approach, called example-based 3D object reconstruction from line drawings, which is based on the observation that a natural or man-made complex 3D object normally consists of a set of basic 3D objects. Given a line drawing, a graphical model is built where each node denotes a basic object whose candidates are from a 3D model (example) database. The 3D reconstruction is solved using a maximum-a-posteriori (MAP) estimation such that the reconstructed result best fits the line drawing. Our experiments show that this approach achieves much better reconstruction accuracy and are more robust to imperfect line drawings than previous methods.
Tianfan Xue, Jianzhuang Liu, Xiaoou Tang
CVPR3
2012 Understanding collective crowd behaviors: Learning a Mixture model of Dynamic pedestrian-Agents
abstract
In this paper, a new Mixture model of Dynamic pedestrian-Agents (MDA) is proposed to learn the collective behavior patterns of pedestrians in crowded scenes. Collective behaviors characterize the intrinsic dynamics of the crowd. From the agent-based modeling, each pedestrian in the crowd is driven by a dynamic pedestrian-agent, which is a linear dynamic system with its initial and termination states reflecting a pedestrian's belief of the starting point and the destination. Then the whole crowd is modeled as a mixture of dynamic pedestrian-agents. Once the model is unsupervisedly learned from real data, MDA can simulate the crowd behaviors. Furthermore, MDA can well infer the past behaviors and predict the future behaviors of pedestrians given their trajectories only partially observed, and classify different pedestrian behaviors in the scene. The effectiveness of MDA and its applications are demonstrated by qualitative and quantitative experiments on the video surveillance dataset collected from the New York Grand Central Station.
Bolei Zhou, Xiaogang Wang 0001, Xiaoou Tang
CVPR3
2012 Graph Degree Linkage: Agglomerative Clustering on a Directed Graph
Wayne Zhang 0001, Xiaogang Wang 0001, Deli Zhao, Xiaoou Tang
ECCV (1)4
2012 Coherent Filtering: Detecting Coherent Motions from Crowd Clutters
Bolei Zhou, Xiaoou Tang, Xiaogang Wang 0001
ECCV (2)2
2012 Joint semantic segmentation by searching for compatible-competitive references
abstract
This paper presents a framework for semantically segmenting a target image without tags by searching for references in an image database, where all the images are unsegmented but annotated with tags. We jointly segment the target image and its references by optimizing both semantic consistencies within individual images and correspondences between the target image and each of its references. In our framework, we first retrieve two types of references with a semantic-driven scheme: i) the compatible references which share similar global appearance with the target image; and ii) the competitive references which have distinct appearance to the target image but similar tags with one of the compatible references. The two types of references have complementary information for assisting the segmentation of the target image. Then we construct a novel graphical representation, in which the vertices are superpixels extracted from the target image and its references. The segmentation problem is posed as labeling all the vertices with the semantic tags obtained from the references. The method is able to label images without the pixel-level annotation and classifier training, and it outperforms the state-of-the-arts approaches on the MSRC-21 database.
Ping Luo 0002, Xiaogang Wang 0001, Liang Lin 0004, Xiaoou Tang
ACM Multimedia4
2012 Cross matching of music and image
abstract
Human perception of music and image are highly correlated. Both of them can inspire human sensation like emotion and power. This paper investigates how to model the relationship between music and image using 47,888 music-image pairs extracted from music videos. We have two basic observations for this relationship: 1) music space exhibits simpler cluster structure than image space, and 2) the relationship between the two spaces is complex and nonlinear. Based on these observations, we develop Multiple Ranking Canonical Correlation Analysis (MR-CCA) to learn such relationship. MR-CCA clusters the music-image pairs according to their music parts, and then conducts Ranking CCA (R-CCA) for each cluster. Compared with classical CCA, R-CCA takes account of the pairwise ranking information available in our dataset. MR-CCA improves performance and significantly reduce computational cost. Experiment results show that R-CCA outperforms CCA, and MR-CCA has the best performance with a consistency score of 84.52% with human labeling. The proposed method can be generalized to model cross media relationship and has potential applications in video generation, background music recommendation, and joint retrieval of music and image.
Xixuan Wu, Yu Qiao 0001, Xiaogang Wang 0001, Xiaoou Tang
ACM Multimedia4
2012 Automatic music video generation: cross matching of music and image
abstract
Music and image are two most popular media on the Internet. Human perception of music and image are highly correlated. Music video is one of such products, in which music and image are complement to each other. In this paper, we present a system which can automatically generate music video for a given song. The challenge of such system comes from how to select relative images and align them with the song. This paper deals with this challenge by leveraging lyrics (if exists) and the semantic similarity between music and image. We retrieve related image in internet with lyrics keyword as query and use a learning based method to estimate a semantic score between an image and a music segment. Finally we construct a music video after quality filtering and refinement. Our system also allows users to upload their images and re-pick recommended images to personalize the music video.
Xixuan Wu, Yu Qiao 0001, Xiaoou Tang
ACM Multimedia4
2012 IntentSearch: Capturing User Intention for One-Click Internet Image Search
abstract
Web-scale image search engines (e.g., Google image search, Bing image search) mostly rely on surrounding text features. It is difficult for them to interpret users' search intention only by query keywords and this leads to ambiguous and noisy search results which are far from satisfactory. It is important to use visual information in order to solve the ambiguity in text-based image retrieval. In this paper, we propose a novel Internet image search approach. It only requires the user to click on one query image with minimum effort and images from a pool retrieved by text-based search are reranked based on both visual and textual content. Our key contribution is to capture the users' search intention from this one-click query image in four steps. 1) The query image is categorized into one of the predefined adaptive weight categories which reflect users' search intention at a coarse level. Inside each category, a specific weight schema is used to combine visual features adaptive to this kind of image to better rerank the text-based search result. 2) Based on the visual content of the query image selected by the user and through image clustering, query keywords are expanded to capture user intention. 3) Expanded keywords are used to enlarge the image pool to contain more relevant images. 4) Expanded keywords are also used to expand the query image to multiple positive visual examples from which new query specific visual and textual similarity metrics are learned to further improve content-based image reranking. All these steps are automatic, without extra effort from the user. This is critically important for any commercial web-based image search engine, where the user interface has to be extremely simple. Besides this key contribution, a set of visual features which are both effective and efficient in Internet image search are designed. Experimental evaluation shows that our approach significantly improves the precision of top-ranked images and also the user experience.
Xiaoou Tang, Jingyu Cui, Fang Wen 0001, Xiaogang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 3-D Modeling From a Single View of a Symmetric Object
abstract
3-D technologies are considered as the next generation of multimedia applications. Currently, one of the challenges faced by 3-D applications is the shortage of 3-D resources. To solve this problem, many 3-D modeling methods are proposed to directly recover 3-D geometry from 2-D images. However, these methods on single view modeling either require intensive user interaction, or are restricted to a specific kind of object. In this paper, we propose a novel 3-D modeling approach to recover 3-D geometry from a single image of a symmetric object with minimal user interaction. Symmetry is one of the most common properties of natural or manmade objects. Given a single view of a symmetric object, the user marks some symmetric lines and depth discontinuity regions on the image. Our algorithm first finds a set of planes to approximately fit to the object, and then a rough 3-D point cloud is generated by an optimization procedure. The occluded part of the object is further recovered using symmetry information. Experimental results on various indoor and outdoor objects show that the proposed system can obtain 3-D models from single images with only a little user interaction.
Tianfan Xue, Jianzhuang Liu, Xiaoou Tang
IEEE Trans. Image Process.3
2011 A global sampling method for alpha matting
abstract
Alpha matting refers to the problem of softly extracting the foreground from an image. Given a trimap (specifying known foreground/background and unknown pixels), a straightforward way to compute the alpha value is to sample some known foreground and background colors for each unknown pixel. Existing sampling-based matting methods often collect samples near the unknown pixels only. They fail if good samples cannot be found nearby. In this paper, we propose a global sampling method that uses all samples available in the image. Our global sample set avoids missing good samples. A simple but effective cost function is defined to tackle the ambiguity in the sample selection process. To handle the computational complexity introduced by the large number of samples, we pose the sampling task as a correspondence problem. The correspondence search is efficiently achieved by generalizing a randomized algorithm previously designed for patch matching[3]. A variety of experiments show that our global sampling method produces both visually and quantitatively high-quality matting results.
Kaiming He, Christoph Rhemann, Carsten Rother, Xiaoou Tang, Jian Sun 0001
CVPR4
2011 Query-specific visual semantic spaces for web image re-ranking
abstract
Image re-ranking, as an effective way to improve the results of web-based image search, has been adopted by current commercial search engines. Given a query keyword, a pool of images are first retrieved by the search engine based on textual information. By asking the user to select a query image from the pool, the remaining images are re-ranked based on their visual similarities with the query image. A major challenge is that the similarities of visual features do not well correlate with images' semantic meanings which interpret users' search intention. On the other hand, learning a universal visual semantic space to characterize highly diverse images from the web is difficult and inefficient. In this paper, we propose a novel image re-ranking framework, which automatically offline learns different visual semantic spaces for different query keywords through keyword expansions. The visual features of images are projected into their related visual semantic spaces to get semantic signatures. At the online stage, images are re-ranked by comparing their semantic signatures obtained from the visual semantic space specified by the query keyword. The new approach significantly improves both the accuracy and efficiency of image re-ranking. The original visual features of thousands of dimensions can be projected to the semantic signatures as short as 25 dimensions. Experimental results show that 20% - 35% relative improvement has been achieved on re-ranking precisions compared with the state-of-the-art methods.
Xiaogang Wang 0001, Xiaoou Tang
CVPR3
2011 Symmetric piecewise planar object reconstruction from a single image
abstract
Recovering 3D geometry from a single view of an object is an important and challenging problem in computer vision. Previous methods mainly focus on one specific class of objects without large topological changes, such as cars, faces, or human bodies. In this paper, we propose a novel single view reconstruction algorithm for symmetric piece-wise planar objects that are not restricted to some object classes. Symmetry is ubiquitous in manmade and natural objects and provides rich information for 3D reconstruction. Given a single view of a symmetric piecewise planar object, we first find out all the symmetric line pairs. The geometric properties of symmetric objects are used to narrow down the searching space. Then, based on the symmetric lines, a depth map is recovered through a Markov random field. Experimental results show that our algorithm can efficiently recover the 3D shapes of different objects with significant topological variations.
Tianfan Xue, Jianzhuang Liu, Xiaoou Tang
CVPR3
2011 An associate-predict model for face recognition
abstract
Handling intra-personal variation is a major challenge in face recognition. It is difficult how to appropriately measure the similarity between human faces under significantly different settings (e.g., pose, illumination, and expression). In this paper, we propose a new model, called “Associate-Predict” (AP) model, to address this issue. The associate-predict model is built on an extra generic identity data set, in which each identity contains multiple images with large intra-personal variation. When considering two faces under significantly different settings (e.g., non-frontal and frontal), we first “associate” one input face with alike identities from the generic identity date set. Using the associated faces, we generatively “predict” the appearance of one input face under the setting of another input face, or discriminatively “predict” the likelihood whether two input faces are from the same person or not. We call the two proposed prediction methods as “appearance-prediction” and “likelihood-prediction”. By leveraging an extra data set (“memory”) and the “associate-predict” model, the intra-personal variation can be effectively handled. To improve the generalization ability of our model, we further add a switching mechanism - we directly compare the appearances of two faces if they have close intra-personal settings; otherwise, we use the associate-predict model for the recognition. Experiments on two public face benchmarks (Multi-PIE and LFW) demonstrated that our final model can substantially improve the performance of most existing face recognition methods.
Qi Yin, Xiaoou Tang, Jian Sun 0001
CVPR2
2011 Coupled information-theoretic encoding for face photo-sketch recognition
abstract
Automatic face photo-sketch recognition has important applications for law enforcement. Recent research has focused on transforming photos and sketches into the same modality for matching or developing advanced classification algorithms to reduce the modality gap between features extracted from photos and sketches. In this paper, we propose a new inter-modality face recognition approach by reducing the modality gap at the feature extraction stage. A new face descriptor based on coupled information-theoretic encoding is used to capture discriminative local face structures and to effectively match photos and sketches. Guided by maximizing the mutual information between photos and sketches in the quantized feature spaces, the coupled encoding is achieved by the proposed coupled information-theoretic projection tree, which is extended to the randomized forest to further boost the performance. We create the largest face sketch database including sketches of 1, 194 people from the FERET database. Experiments on this large scale dataset show that our approach significantly outperforms the state-of-the-art methods.
Wayne Zhang 0001, Xiaogang Wang 0001, Xiaoou Tang
CVPR3
2011 Random field topic model for semantic region analysis in crowded scenes from tracklets
abstract
In this paper, a Random Field Topic (RFT) model is proposed for semantic region analysis from motions of objects in crowded scenes. Different from existing approaches of learning semantic regions either from optical flows or from complete trajectories, our model assumes that fragments of trajectories (called tracklets) are observed in crowded scenes. It advances the existing Latent Dirichlet Allocation topic model, by integrating the Markov random fields (MRF) as prior to enforce the spatial and temporal coherence between tracklets during the learning process. Two kinds of MRF, pairwise MRF and the forest of randomly spanning trees, are defined. Another contribution of this model is to include sources and sinks as high-level semantic prior, which effectively improves the learning of semantic regions and the clustering of tracklets. Experiments on a large scale data set, which includes 40, 000+ tracklets collected from the crowded New York Grand Central station, show that our model outperforms state-of-the-art methods both on qualitative results of learning semantic regions and on quantitative results of clustering tracklets.
Bolei Zhou, Xiaogang Wang 0001, Xiaoou Tang
CVPR3
2011 Optical flow estimation using learned sparse model
abstract
Optical flow estimation is a fundamental and ill-posed problem in computer vision. To recover a dense flow field, appropriate spatial constraints have to be enforced. Recent advances exploit higher order spatial regularization, and achieve the top performance on the Middlebury benchmark. In this work, we revisit learning-based approach, and propose a learned sparse model to patch-wisely regularize the flow field. In particular, our method is based on multi-scale spatial regularization, which benefits from first-order spatial regularity and our learned, higher order sparse model. To obtain accurate flow estimation, we propose a sequential optimization scheme to solve the corresponding energy minimization problem. Moreover, as the errors in intermediate flow estimates are usually dense with large variations, we further propose flow-driven and image-driven approaches to address the problem of outliers. Experiments on the Middlebury benchmark show that our method is competitive with the state-of-the-art.
Kui Jia, Xiaogang Wang 0001, Xiaoou Tang
ICCV3
2011 Content-based photo quality assessment
abstract
Automatically assessing photo quality from the perspective of visual aesthetics is of great interest in high-level vision research and has drawn much attention in recent years. In this paper, we propose content-based photo quality assessment using regional and global features. Under this framework, subject areas, which draw the most attentions of human eyes, are first extracted. Then regional features extracted from subject areas and the background regions are combined with global features to assess the photo quality. Since professional photographers may adopt different photographic techniques and may have different aesthetic criteria in mind when taking different types of photos (e.g. landscape versus portrait), we propose to segment regions and extract visual features in different ways according to the categorization of photo content. Therefore we divide the photos into seven categories based on their content and develop a set of new subject area extraction methods and new visual features, which are specially designed for different categories. This argument is supported by extensive experimental comparisons of existing photo quality assessment approaches as well as our new regional and global features over different categories of photos. Our new features significantly outperform the state-of-the-art methods. Another contribution of this work is to construct a large and diversified benchmark database for the research of photo quality assessment. It includes 17, 613 photos with manually labeled ground truth.
Xiaogang Wang 0001, Xiaoou Tang
ICCV3
2011 Automatic motion-guided video stylization and personalization
abstract
Video stylization transfers a source video into an artistic version while maintaining temporal coherence between adjacent frames. In this paper, we formulate the unsupervised example-based video stylization with Markov random field model. In our algorithm, we implement an improved optical flow algorithm to maintain temporal coherence while improve the accuracy of estimation along motion boundaries. We also extend our algorithm to the application of video personalization, in which human faces keep clear and distinguishable. A series of techniques are fused in video personalization, including face detection and alignment, motion flow, skin detection, and illumination blending. Given a source video and a style template image, our algorithm produces the stylized and/or personalized video(s) automatically. Experimental results demonstrate that our algorithm performs excellently in both video stylization and personalization.
Shifeng Chen, Wei Zhang 0081, Xiaoou Tang
ACM Multimedia4
2011 3D object retrieval with semantic attributes
abstract
Humans are capable of describing objects using attributes, such as "the object looks circular and is man-made". Motivated by these high-level descriptions, we build a user-friendly 3D object retrieval system, where the user can browse the database and search for targeted objects using semantic attributes. The main advantage of our system is that it does not require the user to find or sketch a 3D object as the query for 3D object retrieval. Besides, to the best of our knowledge, our system has obtained the best retrieval performance on three popular benchmarks.
Boqing Gong, Jianzhuang Liu, Xiaogang Wang 0001, Xiaoou Tang
ACM Multimedia4
2011 Automatic object segmentation from large scale 3D urban point clouds through manifold embedded mode seeking
abstract
This paper presents a system that can automatically segment objects in large scale 3D point clouds obtained from urban ranging images. The system consists of three steps: The first one involves a ground detection process that can detect relatively complex terrain and separate it from other objects. The second step superpixelizes the remaining objects to speed up the segmentation process. In the final step, a manifold embedded mode seeking method is adopted to segment the point clouds. Even though the segmentation of urban objects is a challenging problem in terms of accuracy and problem scale, our system can efficiently generate very good segmentation results. The proposed manifold learning effectively improves the segmentation performance due to the fact that continuous artificial objects often have manifold-like structures.
Zhiding Yu, Chunjing Xu, Jianzhuang Liu, Oscar C. Au, Xiaoou Tang
ACM Multimedia5
2011 Edge-preserving single image super-resolution
abstract
This paper proposes a novel approach to single image super-resolution. First, an image up-sampling scheme is proposed which takes the advantages of both bilateral filtering and mean shift image segmentation. Then we use a shock filter to enhance strong edges in the initial up-sampling result and obtain an intermediate high-resolution image. Finally, we enforce a reconstruction constraint on the high-resolution image so that fine details can be inferred by back projection. Since strong edges in the intermediate result are enhanced, ringing artifacts can be suppressed in the back projection step. We compare our algorithm with several state-of-the-art image super-resolution algorithms. Qualitative and quantitative experimental results demonstrate that our approach performs the best.
Shifeng Chen, Jianzhuang Liu, Xiaoou Tang
ACM Multimedia4
2011 Single Image Haze Removal Using Dark Channel Prior
abstract
In this paper, we propose a simple but effective image prior-dark channel prior to remove haze from a single input image. The dark channel prior is a kind of statistics of outdoor haze-free images. It is based on a key observation-most local patches in outdoor haze-free images contain some pixels whose intensity is very low in at least one color channel. Using this prior with the haze imaging model, we can directly estimate the thickness of the haze and recover a high-quality haze-free image. Results on a variety of hazy images demonstrate the power of the proposed prior. Moreover, a high-quality depth map can also be obtained as a byproduct of haze removal.
Kaiming He, Jian Sun 0001, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.3
2011 Decomposition of Complex Line Drawings with Hidden Lines for 3D Planar-Faced Manifold Object Reconstruction
abstract
Three-dimensional object reconstruction from a single 2D line drawing is an important problem in computer vision. Many methods have been presented to solve this problem, but they usually fail when the geometric structure of a 3D object becomes complex. In this paper, a novel approach based on a divide-and-conquer strategy is proposed to handle the 3D reconstruction of a planar-faced complex manifold object from its 2D line drawing with hidden lines visible. The approach consists of four steps: 1) identifying the internal faces of the line drawing, 2) decomposing the line drawing into multiple simpler ones based on the internal faces, 3) reconstructing the 3D shapes from these simpler line drawings, and 4) merging the 3D shapes into one complete object represented by the original line drawing. A number of examples are provided to show that our approach can handle 3D reconstruction of more complex objects than previous methods.
Jianzhuang Liu, Yu Chen 0009, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.3
2011 Learning to Detect a Salient Object
abstract
In this paper, we study the salient object detection problem for images. We formulate this problem as a binary labeling task where we separate the salient object from the background. We propose a set of novel features, including multiscale contrast, center-surround histogram, and color spatial distribution, to describe a salient object locally, regionally, and globally. A conditional random field is learned to effectively combine these features for salient object detection. Further, we extend the proposed approach to detect a salient object from sequential images by introducing the dynamic salient features. We collected a large image database containing tens of thousands of carefully labeled images by multiple users and a video segment database, and conducted a set of experiments over them to demonstrate the effectiveness of the proposed approach.
Zejian Yuan, Jian Sun 0001, Jingdong Wang 0001, Nanning Zheng 0001, Xiaoou Tang, Harry Shum
IEEE Trans. Pattern Anal. Mach. Intell.6
2011 From Tiger to Panda: Animal Head Detection
abstract
Robust object detection has many important applications in real-world online photo processing. For example, both Google image search and MSN live image search have integrated human face detector to retrieve face or portrait photos. Inspired by the success of such face filtering approach, in this paper, we focus on another popular online photo category--animal, which is one of the top five categories in the MSN live image search query log. As a first attempt, we focus on the problem of animal head detection of a set of relatively large land animals that are popular on the internet, such as cat, tiger, panda, fox, and cheetah. First, we proposed a new set of gradient oriented feature, Haar of Oriented Gradients (HOOG), to effectively capture the shape and texture features on animal head. Then, we proposed two detection algorithms, namely Bruteforce detection and Deformable detection, to effectively exploit the shape feature and texture feature simultaneously. Experimental results on 14,379 well labeled animals images validate the superiority of the proposed approach. Additionally, we apply the animal head detector to improve the image search result through text based online photo search result filtering.
Jian Sun 0001, Xiaoou Tang
IEEE Trans. Image Process.3
2011 Learning Semi-Riemannian Metrics for Semisupervised Feature Extraction
abstract
Discriminant feature extraction plays a central role in pattern recognition and classification. Linear Discriminant Analysis (LDA) is a traditional algorithm for supervised feature extraction. Recently, unlabeled data have been utilized to improve LDA. However, the intrinsic problems of LDA still exist and only the similarity among the unlabeled data is utilized. In this paper, we propose a novel algorithm, called Semisupervised Semi-Riemannian Metric Map (S3RMM), following the geometric framework of semi Riemannian manifolds. S3RMM maximizes the discrepancy of the separability and similarity measures of scatters formulated by using semi-Riemannian metric tensors. The metric tensor of each sample is learned via semisupervised regression. Our method can also be a general framework for proposing new semisupervised algorithms, utilizing the existing discrepancy-criterion-based algorithms. The experiments demonstrated on faces and handwritten digits show that S3RMM is promising for semisupervised feature extraction.
Wayne Zhang 0001, Zhouchen Lin, Xiaoou Tang
IEEE Trans. Knowl. Data Eng.3
2010 Face recognition with learning-based descriptor
abstract
We present a novel approach to address the representation issue and the matching issue in face recognition (verification). Firstly, our approach encodes the micro-structures of the face by a new learning-based encoding method. Unlike many previous manually designed encoding methods (e.g., LBP or SIFT), we use unsupervised learning techniques to learn an encoder from the training examples, which can automatically achieve very good tradeoff between discriminative power and invariance. Then we apply PCA to get a compact face descriptor. We find that a simple normalization mechanism after PCA can further improve the discriminative ability of the descriptor. The resulting face representation, learning-based (LE) descriptor, is compact, highly discriminative, and easy-to-extract. To handle the large pose variation in real-life scenarios, we propose a pose-adaptive matching method that uses pose-specific classifiers to deal with different pose combinations (e.g., frontal v.s. frontal, frontal v.s. left) of the matching face pair. Our approach is comparable with the state-of-the-art methods on the Labeled Face in Wild (LFW) benchmark (we achieved 84.45% recognition rate), while maintaining excellent compactness, simplicity, and generalization aability across different datasets.
Zhimin Cao, Qi Yin, Xiaoou Tang, Jian Sun 0001
CVPR3
2010 Isoperimetric cut on a directed graph
abstract
In this paper, we propose a novel probabilistic view of the spectral clustering algorithm. In our framework, the spectral clustering algorithm can be viewed as assigning class labels to samples to minimize the Bayes classification error rate by using a kernel density estimator (KDE). From this perspective, we propose to construct directed graphs using variable bandwidth KDEs. Such a variable bandwidth KDE based directed graph has the advantage that it encodes the local density information of the data in the graph edge weights. In order to cluster vertices of the directed graph, we develop a directed graph partitioning algorithm which optimizes a random walk isoperimetric ratio. The partitioning result can be obtained efficiently by solving a system of linear equations. We have applied our algorithm to several benchmark data sets and obtained promising results.
Jianzhuang Liu, Xiaoou Tang
CVPR4
2010 Fast matting using large kernel matting Laplacian matrices
abstract
Image matting is of great importance in both computer vision and graphics applications. Most existing state-of-the-art techniques rely on large sparse matrices such as the matting Laplacian. However, solving these linear systems is often time-consuming, which is unfavored for the user interaction. In this paper, we propose a fast method for high quality matting. We first derive an efficient algorithm to solve a large kernel matting Laplacian. A large kernel propagates information more quickly and may improve the matte quality. To further reduce running time, we also use adaptive kernel sizes by a KD-tree trimap segmentation technique. A variety of experiments show that our algorithm provides high quality results and is 5 to 20 times faster than previous methods.
Kaiming He, Jian Sun 0001, Xiaoou Tang
CVPR3
2010 Object cut: Complex 3D object reconstruction through line drawing separation
abstract
This paper proposes an approach called object cut to tackle an important problem in computer vision, 3D object reconstruction from single line drawings. Given a complex line drawing representing a solid object, our algorithm finds the places, called cuts, to separate the line drawing into much simpler ones. The complex 3D object is obtained by first reconstructing the 3D objects from these simpler line drawings and then combining them together. Several propositions and criteria are presented for cut finding. A theorem is given to guarantee the existence and uniqueness of the separation of a line drawing along a cut. Our experiments show that the proposed approach can deal with more complex 3D object reconstruction than state-of-the-art methods.
Tianfan Xue, Jianzhuang Liu, Xiaoou Tang
CVPR3
2010 Guided Image Filtering
Kaiming He, Jian Sun 0001, Xiaoou Tang
ECCV (1)3
2010 Lighting and Pose Robust Face Sketch Synthesis
Wayne Zhang 0001, Xiaogang Wang 0001, Xiaoou Tang
ECCV (6)3
2010 User intention modeling for interactive image retrieval
abstract
We propose three innovative interactive methods to let computer better understand user intention in content-based image retrieval: 1. Smart intention list induces user intention, thereby improves search results by intention-specific search schema; 2. Reference strokes interaction allows user to specify in detail about the intention by pointing out interested regions; 3. Natural user feedback easily collects data of user relevance feedbacks to boost the performance of the system. Systematic user study shows that the proposed interactive mechanism improves search efficiency, reduces user workload, and enhances user experience.
Jingyu Cui, Fang Wen 0001, Xiaoou Tang
ICME3
2010 Fast image rearrangement via multi-scale patch copying
abstract
In this paper, we propose a simple interactive way for a novel type of image synthesis called image rearrangement whose goal is to construct a new image based on some objects cropped from source images. The synthesis results are obtained by copying patches from the source images in a globally consistent way. The patch copying problem is formulated with the Markov random field model, and belief propagation is used as the optimization tool. To speed up our algorithm, a two-step belief propagation and a multi-scale patch copying scheme are taken. Experimental results indicate that our algorithm obtains satisfactory results in both performance and efficiency.
Jiayao Hu, Shifeng Chen, Jianzhuang Liu, Xiaoou Tang
ACM Multimedia4
2010 3D object search through semantic component
abstract
In this paper, we present a novel concept named semantic component for 3D object search which describes a key component that semantically defines a 3D object. In most cases, the semantic component is intra-category stable and therefore can be used to construct an efficient 3D object retrieval scheme. By segmenting an object into segments and learning the similar segments shared by all the objects in the same category, we can summarise what human uses for object recognition, from the analysis of which we develop a method to find the semantic component of an object. In our experiments, the proposed method is justified and the effectiveness of our algorithm is also demonstrated.
Chunjing Xu, Zhengwu Zhang, Jianzhuang Liu, Xiaoou Tang
ACM Multimedia4
2010 Unsupervised Object Segmentation with a Hybrid Graph Model (HGM)
abstract
In this work, we address the problem of performing class-specific unsupervised object segmentation, i.e., automatic segmentation without annotated training images. Object segmentation can be regarded as a special data clustering problem where both class-specific information and local texture/color similarities have to be considered. To this end, we propose a hybrid graph model (HGM) that can make effective use of both symmetric and asymmetric relationship among samples. The vertices of a hybrid graph represent the samples and are connected by directed edges and/or undirected ones, which represent the asymmetric and/or symmetric relationship between them, respectively. When applied to object segmentation, vertices are superpixels, the asymmetric relationship is the conditional dependence of occurrence, and the symmetric relationship is the color/texture similarity. By combining the Markov chain formed by the directed subgraph and the minimal cut of the undirected subgraph, the object boundaries can be determined for each image. Using the HGM, we can conveniently achieve simultaneous segmentation and recognition by integrating both top-down and bottom-up information into a unified process. Experiments on 42 object classes (9,415 images in total) show promising results.
Guangcan Liu, Zhouchen Lin, Yong Yu 0001, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.4
2010 Robust 3D Face Recognition by Local Shape Difference Boosting
abstract
This paper proposes a new 3D face recognition approach, Collective Shape Difference Classifier (CSDC), to meet practical application requirements, i.e., high recognition performance, high computational efficiency, and easy implementation. We first present a fast posture alignment method which is self-dependent and avoids the registration between an input face against every face in the gallery. Then, a Signed Shape Difference Map (SSDM) is computed between two aligned 3D faces as a mediate representation for the shape comparison. Based on the SSDMs, three kinds of features are used to encode both the local similarity and the change characteristics between facial shapes. The most discriminative local features are selected optimally by boosting and trained as weak classifiers for assembling three collective strong classifiers, namely, CSDCs with respect to the three kinds of features. Different schemes are designed for verification and identification to pursue high performance in both recognition and computation. The experiments, carried out on FRGC v2 with the standard protocol, yield three verification rates all better than 97.9 percent with the FAR of 0.1 percent and rank-1 recognition rates above 98 percent. Each recognition against a gallery with 1,000 faces only takes about 3.6 seconds. These experimental results demonstrate that our algorithm is not only effective but also time efficient.
Yueming Wang 0001, Jianzhuang Liu, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.3
2010 Image Segmentation by MAP-ML Estimations
abstract
Image segmentation plays an important role in computer vision and image analysis. In this paper, image segmentation is formulated as a labeling problem under a probability maximization framework. To estimate the label configuration, an iterative optimization scheme is proposed to alternately carry out the maximum a posteriori (MAP) estimation and the maximum likelihood (ML) estimation. The MAP estimation problem is modeled with Markov random fields (MRFs) and a graph cut algorithm is used to find the solution to the MAP estimation. The ML estimation is achieved by computing the means of region features in a Gaussian model. Our algorithm can automatically segment an image into regions with relevant textures or colors without the need to know the number of regions in advance. Its results match image edges very well and are consistent with human perception. Comparing to six state-of-the-art algorithms, extensive experiments have shown that our algorithm performs the best.
Shifeng Chen, Liangliang Cao, Yueming Wang 0001, Jianzhuang Liu, Xiaoou Tang
IEEE Trans. Image Process.5
2010 Misalignment-Robust Face Recognition
abstract
Subspace learning techniques for face recognition have been widely studied in the past three decades. In this paper, we study the problem of general subspace-based face recognition under the scenarios with spatial misalignments and/or image occlusions. For a given subspace derived from training data in a supervised, unsupervised, or semi-supervised manner, the embedding of a new datum and its underlying spatial misalignment parameters are simultaneously inferred by solving a constrained l1 norm optimization problem, which minimizes the l1 error between the misalignment-amended image and the image reconstructed from the given subspace along with its principal complementary subspace. A byproduct of this formulation is the capability to detect the underlying image occlusions. Extensive experiments on spatial misalignment estimation, image occlusion detection, and face recognition with spatial misalignments and/or image occlusions all validate the effectiveness of our proposed general formulation for misalignment-robust face recognition.
Shuicheng Yan, Huan Wang 0001, Jianzhuang Liu, Xiaoou Tang, Thomas S. Huang
IEEE Trans. Image Process.4
2010 A rank-one update algorithm for fast solving kernel Foley-Sammon optimal discriminant vectors
abstract
Discriminant analysis plays an important role in statistical pattern recognition. A popular method is the Foley-Sammon optimal discriminant vectors (FSODVs) method, which aims to find an optimal set of discriminant vectors that maximize the Fisher discriminant criterion under the orthogonal constraint. The FSODVs method outperforms the classic Fisher linear discriminant analysis (FLDA) method in the sense that it can solve more discriminant vectors for recognition. Kernel Foley-Sammon optimal discriminant vectors (KFSODVs) is a nonlinear extension of FSODVs via the kernel trick. However, the current KFSODVs algorithm may suffer from the heavy computation problem since it involves computing the inverse of matrices when solving each discriminant vector, resulting in a cubic complexity for each discriminant vector. This is costly when the number of discriminant vectors to be computed is large. In this paper, we propose a fast algorithm for solving the KFSODVs, which is based on rank-one update (ROU) of the eigensytems. It only requires a square complexity for each discriminant vector. Moreover, we also generalize our method to efficiently solve a family of optimally constrained generalized Rayleigh quotient (OCGRQ) problems which include many existing dimensionality reduction techniques. We conduct extensive experiments on several real data sets to demonstrate the effectiveness of the proposed algorithms.
Wenming Zheng, Zhouchen Lin, Xiaoou Tang
IEEE Trans. Neural Networks3
2009 Single image haze removal using dark channel prior
abstract
In this paper, we propose a simple but effective image prior - dark channel prior to remove haze from a single input image. The dark channel prior is a kind of statistics of the haze-free outdoor images. It is based on a key observation - most local patches in haze-free outdoor images contain some pixels which have very low intensities in at least one color channel. Using this prior with the haze imaging model, we can directly estimate the thickness of the haze and recover a high quality haze-free image. Results on a variety of outdoor haze images demonstrate the power of the proposed prior. Moreover, a high quality depth map can also be obtained as a by-product of haze removal.
Kaiming He, Jian Sun 0001, Xiaoou Tang
CVPR3
2009 Constrained clustering via spectral regularization
abstract
We propose a novel framework for constrained spectral clustering with pairwise constraints which specify whether two objects belong to the same cluster or not. Unlike previous methods that modify the similarity matrix with pairwise constraints, we adapt the spectral embedding towards an ideal embedding as consistent with the pairwise constraints as possible. Our formulation leads to a small semidefinite program whose complexity is independent of the number of objects in the data set and the number of pairwise constraints, making it scalable to large-scale problems. The proposed approach is applicable directly to multi-class problems, handles both must-link and cannot-link constraints, and can effectively propagate pairwise constraints. Extensive experiments on real image data and UCI data have demonstrated the efficacy of our algorithm.
Zhenguo Li, Jianzhuang Liu, Xiaoou Tang
CVPR3
2009 3D reconstruction of curved objects from single 2D line drawings
abstract
An important research area in computer vision is developing algorithms that can reconstruct the 3D surface of an object represented by a single 2D line drawing. Previous work on 3D reconstruction from single 2D line drawings focuses on objects with planar faces. In this paper, we propose a novel approach to the reconstruction of solid objects that have not only planar but also curved faces. Our approach consists of four steps: (1) identifying the curved faces and planar faces in a line drawing, (2) transforming the line drawing into one with straight edges only, (3) reconstructing the 3D wireframe of the curved object from the transformed line drawing and the original line drawing, and (4) generating the curved faces with Bezier patches and triangular meshes. With a number of experimental results, we demonstrate the ability of our approach to perform curved object reconstruction successfully.
Yingze Wang, Yu Chen 0009, Jianzhuang Liu, Xiaoou Tang
CVPR4
2009 Automatic facial expression recognition on a single 3D face by exploring shape deformation
abstract
Facial expression recognition has many applications in multimedia processing and the development of 3D data acquisition techniques makes it possible to identify expressions using 3D shape information. In this paper, we propose an automatic facial expression recognition approach based on a single 3D face. The shape of an expressional 3D face is approximated as the sum of two parts, a basic facial shape component (BFSC) and an expressional shape component (ESC). The BFSC represents the basic face structure and neutral-style shape and the ESC contains shape changes caused by facial expressions. To separate the BFSC and ESC, our method firstly builds a reference face for each input 3D non-neutral face by a learning method, which well represents the basic facial shape. Then, based on the BFSC and the original expressional face, a facial expression descriptor is designed. The surface depth changes are considered in the descriptor. Finally, the descriptor is input into an SVM to recognize the expression. Unlike previous methods which recognize a facial expression with the help of manually labeled key points and/or a neutral face, our method works on a single 3D face without any manual assistance. Extensive experiments are carried out on the BU-3DFE database and comparisons with existing methods are conducted. The experimental results show the effectiveness of our method.
Boqing Gong, Yueming Wang 0001, Jianzhuang Liu, Xiaoou Tang
ACM Multimedia4
2009 Boosting 3D object retrieval by object flexibility
abstract
In this paper, we propose a novel feature, called object flexibility, at a point of a 3D object to describe how the neighborhood of this point is massively connected to the object. We show that this feature is stable to the deformation of objects' articulations, in addition to commonly concerned linear transforms, i.e., translation, scale, and rotation. A shape descriptor is obtained based on this feature using the bag-of-words model. As an application, the descriptor is used to perform 3D object retrieval. Extensive experiments demonstrate its superiority over a variety of existing 3D shape descriptors in the retrieval of articulated objects, as well as its enhancement of other shape descriptors to retrieve generic 3D objects.
Boqing Gong, Chunjing Xu, Jianzhuang Liu, Xiaoou Tang
ACM Multimedia4
2009 Video completion via motion guided spatial-temporal global optimization
abstract
In this paper, a novel global optimization based approach is proposed for video completion whose target is to restore the spatial-temporal missing regions of a video in a visually plausible way. Our algorithm consists of two stages: motion field completion and color completion via global optimization. First, local motions within the missing parts are completed patch-by-patch greedily using pre-computed available motions in the video. Then the missing regions are filled by sampling patches from available parts of the video. We formulate the video completion as a global energy minimization problem by Markov random fields (MRFs). Based on the completed motion field of the video, a well-defined energy function involving both spatial and temporal coherence relationship is constructed. A coarse-to-fine Belief Propagation (BP) is proposed to solve the optimization problem. Experimental results have demonstrated the good performance of our algorithm.
Shifeng Chen, Jianzhuang Liu, Xiaoou Tang
ACM Multimedia4
2009 Performance driven face animation via non-rigid 3d tracking
abstract
In this demo, a performance driven 3D face animation system is proposed. The proposed system consists of two key components: a robust non-rigid 3D tracking module and a MPEG4 compliant facial animation module. Firstly, the facial motion is tracked from source videos which contain both the rigid 3D head motion (6 DOF) and the non-rigid expression variations. Afterward, the tracked facial motion is parameterized via estimating a set of MPEG4 facial animation parameters(FAP). As the final step, these FAP values are transferred to the MPEG4-compliant face model for the animation purpose. The proposed tracking and animation system has a strong generalization ability and can be used in the indoor environment with no additional assumptions.
Wayne Zhang 0001, Qiang Wang 0023, Xiaoou Tang
ACM Multimedia3
2009 Responses to the Comments on "What the Back of the Object Looks Like: 3D Reconstruction from Line Drawings without Hidden Lines"
abstract
Varley (2009) made comments on our paper in (L. Cao et al., 2008) section by section. We answer them in this response paper.
Liangliang Cao, Jianzhuang Liu, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.3
2009 Nonparametric Discriminant Analysis for Face Recognition
abstract
In this paper, we develop a new framework for face recognition based on nonparametric discriminant analysis (NDA) and multi-classifier integration. Traditional LDA-based methods suffer a fundamental limitation originating from the parametric nature of scatter matrices, which are based on the Gaussian distribution assumption. The performance of these methods notably degrades when the actual distribution is Non-Gaussian. To address this problem, we propose a new formulation of scatter matrices to extend the two-class nonparametric discriminant analysis to multi-class cases. Then, we develop two more improved multi-class NDA-based algorithms (NSA and NFA) with each one having two complementary methods based on the principal space and the null space of the intra-class scatter matrix respectively. Comparing to the NSA, the NFA is more effective in the utilization of the classification boundary information. In order to exploit the complementary nature of the two kinds of NFA (PNFA and NNFA), we finally develop a dual NFA-based multi-classifier fusion framework by employing the over complete Gabor representation to boost the recognition performance. We show the improvements of the developed new algorithms over the traditional subspace methods through comparative experiments on two challenging face databases, Purdue AR database and XM2VTS database.
Zhifeng Li 0001, Dahua Lin, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.3
2009 Responses to the Comments on "Plane-Based Optimization for 3D Object Reconstruction from Single Line Drawings"
abstract
We disagree with the comments made by Varley [1] on our previous paper [2]. In this paper, we respond to his comments and show that they are not correct.
Jianzhuang Liu, Liangliang Cao, Zhenguo Li, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.4
2009 Face Photo-Sketch Synthesis and Recognition
abstract
In this paper, we propose a novel face photo-sketch synthesis and recognition method using a multiscale Markov Random Fields (MRF) model. Our system has three components: 1) given a face photo, synthesizing a sketch drawing; 2) given a face sketch drawing, synthesizing a photo; and 3) searching for face photos in the database based on a query sketch drawn by an artist. It has useful applications for both digital entertainment and law enforcement. We assume that faces to be studied are in a frontal pose, with normal lighting and neutral expression, and have no occlusions. To synthesize sketch/photo images, the face region is divided into overlapping patches for learning. The size of the patches decides the scale of local face structures to be learned. From a training set which contains photo-sketch pairs, the joint photo-sketch model is learned at multiple scales using a multiscale MRF model. By transforming a face photo to a sketch (or transforming a sketch to a photo), the difference between photos and sketches is significantly reduced, thus allowing effective matching between the two in face sketch recognition. After the photo-sketch transformation, in principle, most of the proposed face photo recognition approaches can be applied to face sketch recognition in a straightforward way. Extensive experiments are conducted on a face sketch database including 606 faces, which can be downloaded from our Web site (http://mmlab.ie.cuhk.edu.hk/facesketch.html).
Xiaogang Wang 0001, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.2
2009 2D Shape Matching by Contour Flexibility
abstract
In computer vision, shape matching is a challenging problem, especially when articulation and deformation of parts occur. These variations may be insignificant in terms of human recognition, but often cause a matching algorithm to give results that are inconsistent with our perception. In this paper, we propose a novel shape descriptor of planar contours, called contour flexibility, which represents the deformable potential at each point along a contour. With this descriptor, The local and global features can be obtained from the contour. We then present a shape matching scheme based on the features obtained. Experiments with comparisons to recently published algorithms show that our algorithm performs best.
Chunjing Xu, Jianzhuang Liu, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.3
2009 Fast, automatic and fine-grained tampered JPEG image detection via DCT coefficient analysis
Zhouchen Lin, Junfeng He, Xiaoou Tang, Chi-Keung Tang
Pattern Recognit.3
2009 Tensor linear Laplacian discrimination (TLLD) for feature extraction
Wayne Zhang 0001, Zhouchen Lin, Xiaoou Tang
Pattern Recognit.3
2009 Audio-Guided Video-Based Face Recognition
abstract
In this paper, we develop a new video-to-video face recognition algorithm. The major advantage of the video-based method is that more information is available in a video sequence than in a single image. In order to take advantage of the large amount of information in the video sequence and at the same time overcome the processing speed and data size problems, we develop several new techniques including temporal and spatial frame synchronization, multilevel discriminant subspace analysis, and multiclassifier integration for video sequence processing. An aligned video sequence for each person is first obtained by applying temporal and spatial synchronization, which effectively establishes the face correspondence using both audio and video information; then multilevel discriminant subspace analysis or multiclassifier integration is employed for further analysis based on the synchronized sequence. The method preserves most of the temporal-spatial information contained in a video sequence. Extensive experiments on the XM2VTS database clearly show the superiority of our new algorithms with near-perfect classification results (99.3%) obtained.
Xiaoou Tang, Zhifeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2009 Fast algorithm for updating the discriminant vectors of dual-space LDA
abstract
Dual-space linear discriminant analysis (DSLDA) is a popular method for discriminant analysis. The basic idea of the DSLDA method is to divide the whole data space into two complementary subspaces, i.e., the range space of the within-class scatter matrix and its complementary space, and then solve the discriminant vectors in each subspace. Hence, the DSLDA method can take full advantage of the discriminant information of the training samples. However, from the computational point of view, the original DSLDA method may not be suitable for online training problems because of its heavy computational cost. To this end, we modify the original DSLDA method and then propose a data order independent incremental algorithm to accurately update the discriminant vectors of the DSLDA method when new samples are inserted into the training data set. We conduct experiments on the AR face database to confirm the better performance of the proposed algorithms in terms of the recognition accuracy and computational efficiency.
Wenming Zheng, Xiaoou Tang
IEEE Trans. Inf. Forensics Secur.2
2009 A Theory of Phase Singularities for Image Representation and its Applications to Object Tracking and Image Matching
abstract
This paper studies phase singularities (PSs) for image representation. We show that PSs calculated with Laguerre-Gauss filters contain important information and provide a useful tool for image analysis. PSs are invariant to image translation and rotation. We introduce several invariant features to characterize the core structures around PSs and analyze the stability of PSs to noise addition and scale change. We also study the characteristics of PSs in a scale space, which lead to a method to select key scales along phase singularity curves. We demonstrate two applications of PSs: object tracking and image matching. In object tracking, we use the iterative closest point algorithm to determine the correspondences of PSs between two adjacent frames. The use of PSs allows us to precisely determine the motions of tracked objects. In image matching, we combine PSs and scale-invariant feature transform (SIFT) descriptor to deal with the variations between two images and examine the proposed method on a benchmark database. The results indicate that our method can find more correct matching pairs with higher repeatability rates than some well-known methods.
Yu Qiao 0001, Wei Wang 0333, Nobuaki Minematsu, Jianzhuang Liu, Mitsou Takeda, Xiaoou Tang
IEEE Trans. Image Process.6
2009 Correspondence Propagation with Weak Priors
abstract
For the problem of image registration, the top few reliable correspondences are often relatively easy to obtain, while the overall matching accuracy may fall drastically as the desired correspondence number increases. In this paper, we present an efficient feature matching algorithm to employ sparse reliable correspondence priors for piloting the feature matching process. First, the feature geometric relationship within individual image is encoded as a spatial graph, and the pairwise feature similarity is expressed as a bipartite similarity graph between two feature sets; then the geometric neighborhood of the pairwise assignment is represented by a categorical product graph, along which the reliable correspondences are propagated; and finally a closed-form solution for feature matching is deduced by ensuring the feature geometric coherency as well as pairwise feature agreements. Furthermore, our algorithm is naturally applicable for incorporating manual correspondence priors for semi-supervised feature matching. Extensive experiments on both toy examples and real-world applications demonstrate the superiority of our algorithm over the state-of-the-art feature matching techniques.
Huan Wang 0001, Shuicheng Yan, Jianzhuang Liu, Xiaoou Tang, Thomas S. Huang
IEEE Trans. Image Process.4
2009 Synchronized Submanifold Embedding for Person-Independent Pose Estimation and Beyond
abstract
Precise 3-D head pose estimation plays a significant role in developing human-computer interfaces and practical face recognition systems. This task is challenging due to the particular appearance variations caused by pose changes for a certain subject. In this paper, the pose data space is considered as a union of submanifolds which characterize different subjects, instead of a single continuous manifold as conventionally regarded. A novel manifold embedding algorithm dually supervised by both identity and pose information, called synchronized submanifold embedding (SSE), is proposed for person-independent precise 3-D pose estimation, which means that the testing subject may not appear in the model training stage. First, the submanifold of a certain subject is approximated as a set of simplexes constructed using neighboring samples. Then, these simplexized submanifolds from different subjects are embedded by synchronizing the locally propagated poses within the simplexes and at the same time maximizing the intrasubmanifold variances. Finally, the pose of a new datum is estimated as the propagated pose of the nearest point within the simplex constructed by its nearest neighbors in the dimensionality reduced feature space. The experiments on the 3-D pose estimation database, CHIL data for CLEAR07 evaluation, and the extended application for age estimation on FG-NET aging database, demonstrate the superiority of SSE over conventional regression algorithms as well as unsupervised manifold learning algorithms.
Shuicheng Yan, Huan Wang 0001, Yun Fu 0001, Jun Yan 0001, Xiaoou Tang, Thomas S. Huang
IEEE Trans. Image Process.5
2009 Mode-kn Factor Analysis for Image Ensembles
abstract
In this corespondence, we study the extra-factor estimation problem with the assumption that the training image ensemble is expressed as an nth-order tensor with the nth-dimension characterizing all features for an image and other dimensions for different extra factors, such as illuminations, poses, and identities. To overcome the local minimum issue of conventional algorithms designed for this problem, we present a novel statistical learning framework called mode-kn Factor Analysis for obtaining a closed-form solution to estimating the extra factors of any test image. In the learning stage, for the kth [see formula in text] dimension of the data tensor, the mode-kn patterns are constructed by concatenating the feature dimension and the kth extra-factor dimension, and then a mode-kn factor analysis model is learnt based on the mode-kn patterns unfolded from the original data tensor. In the inference stage, for a test image, the mode classification of the kth dimension is performed within a probabilistic framework. The advantages of mode-kn factor analysis over conventional tensor analysis algorithms are twofold: 1) a closed-form solution, instead of iterative sub-optimal solution as conventionally, is derived for estimating the extra-factor mode of any test image; and 2) the classification capability is enhanced by interacting with the process of synthesizing data of all other modes in the k th dimension. Experiments on the Pointing'04 and CMU PIE databases for pose and illumination estimation both validate the superiority of the proposed algorithm over conventional algorithms for extra-factor estimation.
Shuicheng Yan, Huan Wang 0001, Jilin Tu, Xiaoou Tang, Thomas S. Huang
IEEE Trans. Image Process.4
2009 Picture Collage
abstract
In this paper, we address a novel problem of automatically creating a picture collage from a group of images. Picture collage is a kind of visual image summary-to arrange all input images on a given canvas, allowing overlay, to maximize visible visual information. We formulate the picture collage creation problem in a conditional random field model, which integrates image salience, canvas constraint, natural preference, and user interaction. Each image is represented by a group of weighted rectangles, which indicate the salient regions. Then picture collage is resolved by minimizing the energy, guided by the constraints. A two-step optimization method is proposed. First, a quick initialization algorithm based on the proposed 1D collage method is presented. Second, a very efficient Markov chain Monte Carlo method is designed for the refined optimization. We also integrate user interaction in the formulation and optimization to obtain an interactive collage reflecting personalized preference. Visual and quantitative experimental evaluations indicate the efficiency of the proposed collage creation technique.
Jingdong Wang 0001, Jian Sun 0001, Nanning Zheng 0001, Xiaoou Tang, Harry Shum
IEEE Trans. Multim.5
2008 Clustering via Random Walk Hitting Time on Directed Graphs
Jianzhuang Liu, Xiaoou Tang
AAAI3
2008 MQSearch: image search by multi-class query
abstract
Image search is becoming prevalent in web search as the number of digital photos grows exponentially on the internet. For a successful image search system, removing outliers in the top ranked results is a challenging task. Typical content based image search engines take an input image from one class as a query and compute relevance between the query and images in a database. The results often contain a large number of outliers, since these outliers may be similar to the query image in some way. In this paper we present a novel search scheme using query images from multiple classes. Instead of conducting query search for one image class at a time, we conduct multi-class query search jointly. By using several query classes that are similar to each other for multi-class query, we can utilize information across similar classes to fine tune the similarity measure to remove outliers. This strategy can be used for any information search application. In this work, we use content based image search to illustrate the concept.
Yiwen Luo, Wei Liu 0005, Jianzhuang Liu, Xiaoou Tang
CHI4
2008 Sketching in the air: A vision-based system for 3D object design
abstract
3D object design has many applications including flexible 3D sketch input in CAD, computer game, webpage content design, image based object modeling, and 3D object retrieval. Most current 3D object design tools work on a 2D drawing plane such as computer screen or tablet, which is often inflexible with one dimension lost. On the other hand, virtual reality based methods have the drawbacks that there are awkward devices worn by the user and the virtual environment systems are expensive. In this paper, we propose a novel vision-based approach to 3D object design. Our system consists of a PC, a camera, and a mirror. We use the camera and mirror to track a wand so that the user can design 3D objects by sketching in 3D free space directly with out having to wear any cumbersome devices. A number of new techniques are developed for working in this system, including input of object wireframes, gestures for editing and drawing objects, and optimization-based planar and curved surface generation. Our system provides designers a new user interface for designing 3D objects conveniently.
Yu Chen 0009, Jianzhuang Liu, Xiaoou Tang
CVPR3
2008 Transductive object cutout
abstract
In this paper, we address the issue of transducing the object cutout model from an example image to novel image instances. We observe that although object and background are very likely to contain similar colors in natural images, it is much less probable that they share similar color configurations. Motivated by this observation, we propose a local color pattern model to characterize the color configuration in a robust way. Additionally, we propose an edge profile model to modulate the contrast of the image, which enhances edges along object boundaries and attenuates edges inside object or background. The local color pattern model and edge model are integrated in a graph-cut framework. Higher accuracy and improved robustness of the proposed method are demonstrated through experimental comparison with state-of-the-art algorithms.
Jingyu Cui, Qiong Yang, Fang Wen 0001, Qiying Wu, Changshui Zhang, Luc Van Gool, Xiaoou Tang
CVPR7
2008 Misalignment-robust face recognition
abstract
In this paper, we study the problem of subspace-based face recognition under scenarios with spatial misalignments and/or image occlusions. For a given subspace, the embedding of a new datum and the underlying spatial misalignment parameters are simultaneously inferred by solving a constrained ℓ1norm optimization problem, which minimizes the error between the misalignment-amended image and the image reconstructed from the given subspace along with its principal complementary subspace. A byproduct of this formulation is the capability to detect the underlying image occlusions. Extensive experiments on spatial misalignment estimation, image occlusion detection, and face recognition with spatial misalignments and image occlusions all validate the effectiveness of our proposed general formulation.
Huan Wang 0001, Shuicheng Yan, Thomas S. Huang, Jianzhuang Liu, Xiaoou Tang
CVPR5
2008 L1 regularized projection pursuit for additive model learning
abstract
In this paper, we present a L1regularized projection pursuit algorithm for additive model learning. Two new algorithms are developed for regression and classification respectively: sparse projection pursuit regression and sparse Jensen-Shannon Boosting. The introduced L1regularized projection pursuit encourages sparse solutions, thus our new algorithms are robust to overfitting and present better generalization ability especially in settings with many irrelevant input features and noisy data. To make the optimization with L1regularization more efficient, we develop an ldquoinformative feature firstrdquo sequential optimization algorithm. Extensive experiments demonstrate the effectiveness of our proposed approach.
Xiaoou Tang, Harry Shum
CVPR3
2008 Classification via semi-Riemannian spaces
abstract
In this paper, we develop a geometric framework for linear or nonlinear discriminant subspace learning and classification. In our framework, the structures of classes are conceptualized as a semi-Riemannian manifold which is considered as a submanifold embedded in an ambient semi-Riemannian space. The class structures of original samples can be characterized and deformed by local metrics of the semi-Riemannian space. Semi-Riemannian metrics are uniquely determined by the smoothing of discrete functions and the nullity of the semi-Riemannian space. Based on the geometrization of class structures, optimizing class structures in the feature space is equivalent to maximizing the quadratic quantities of metric tensors in the semi-Riemannian space. Thus supervised discriminant subspace learning reduces to unsupervised semi-Riemannian manifold learning. Based on the proposed framework, a novel algorithm, dubbed as Semi-Riemannian Discriminant Analysis (SRDA), is presented for subspace-based classification. The performance of SRDA is tested on face recognition (singular case) and handwritten capital letter classification (nonsingular case) against existing algorithms. The experimental results show that SRDA works well on recognition and classification, implying that semi-Riemannian geometry is a promising new tool for pattern recognition and machine learning.
Deli Zhao, Zhouchen Lin, Xiaoou Tang
CVPR3
2008 Photo and Video Quality Evaluation: Focusing on the Subject
Yiwen Luo, Xiaoou Tang
ECCV (3)2
2008 3D Face Recognition by Local Shape Difference Boosting
Yueming Wang 0001, Xiaoou Tang, Jianzhuang Liu, Gang Pan 0001, Rong Xiao 0003
ECCV (1)2
2008 Cat Head Detection - How to Effectively Exploit Shape and Texture Features
Jian Sun 0001, Xiaoou Tang
ECCV (4)3
2008 Real Time Feature Based 3-D Deformable Face Tracking
Wayne Zhang 0001, Qiang Wang 0023, Xiaoou Tang
ECCV (2)3
2008 Phase singularities for image representation and matching
abstract
Phase features are widely used in image processing and representation due to their stability to deformation and noise. However, phase singularities,where the signals vanish, are generally regarded as harmful and unreliable facts. In this paper, on the contrary, we will show that phase singularities calculated by Laguerre-Gauss filter contain important information of input image and can provide a reliable representation for image matching. We show that the positions of phase singularities are invariant to translation and rotation. Usually, it is possible to recover the input image up to a constant scaling only from the positions of phase singularities. We study phase singularities in scale space, which allows us to determine the "intrinsic scales" of key phase singularities. We introduce three physical measures of the local structures of phase singularities and combine these measures with SIFT descriptor for image matching. We execute experiments on benchmark database to examine the proposed methods. The results indicate that the proposed method can achieve comparable performance with certain well-known methods.
Yu Qiao 0001, Wei Wang 0333, Nobuaki Minematsu, Jianzhuang Liu, Xiaoou Tang
ICASSP5
2008 Pairwise constraint propagation by semidefinite programming for semi-supervised classification
abstract
We consider the general problem of learning from both pairwise constraints and unlabeled data. The pairwise constraints specify whether two objects belong to the same class or not, known as the must-link constraints and the cannot-link constraints. We propose to learn a mapping that is smooth over the data graph and maps the data onto a unit hypersphere, where two must-link objects are mapped to the same point while two cannot-link objects are mapped to be orthogonal. We show that such a mapping can be achieved by formulating a semidefinite programming problem, which is convex and can be solved globally. Our approach can effectively propagate pairwise constraints to the whole data set. It can be directly applied to multi-class classification and can handle data labels, pairwise constraints, or a mixture of them in a unified framework. Promising experimental results are presented for classification tasks on a variety of synthetic and real data sets.
Zhenguo Li, Jianzhuang Liu, Xiaoou Tang
ICML3
2008 Easytoon: an easy and quick tool to personalize a cartoon storyboard using family photo album
abstract
A family photo album based cartoon personalization system, EasyToon, is proposed in this paper. Using state of the art computer vision and graphics technologies and effective UI design, the interactive tool can quickly generate a personalized cartoon storyboard, which naturally blends a real face chosen from the family photo album into a cartoon picture. The personalized cartoon image is easily and quickly obtained in two main steps. First, the best face candidate is selected from the album interactively. Then a personalized cartoon image is automatically synthesized by blending the selected face into the interesting cartoon image. Experiments show that most users express great interest in our system. Without any art background, they can make a personalized cartoon of high quality using the EasyToon within minutes.
Shifeng Chen, Yuandong Tian, Fang Wen 0001, Ying-Qing Xu, Xiaoou Tang
ACM Multimedia5
2008 Real time google and live image search re-ranking
abstract
Nowadays, web-scale image search engines (e.g. Google, Live Image Search) rely almost purely on surrounding text features. This leads to ambiguous and noisy results. We propose to use adaptive visual similarity to re-rank the text-based search results. A query image is first categorized into one of several predefined intention categories, and a specific similarity measure is used inside each category to combine image features for re-ranking based on the query image. Extensive experiments demonstrate that using this algorithm to filter output of Google and Live Image Search is a practical and effective way to dramatically improve the user experience. A real-time image search engine is developed for on-line image search with re-ranking: http://mmlab.ie.cuhk.edu.hk/intentsearch
Jingyu Cui, Fang Wen 0001, Xiaoou Tang
ACM Multimedia3
2008 IntentSearch: interactive on-line image search re-ranking
abstract
In this demo, we present IntentSearch, an interactive system for realtime web based image retrieval. IntentSearch works directly on top of Microsoft Live Image Search, and re-ranks its results according to user specified query image(s) and the automatically inferred user intention. Besides searching in the interface of Microsoft Live Image Search, we also design a more flexible interface to let users browse and play with all the images in the current search session, which makes web image search more efficient and interesting. Please visit http://mmlab.ie.cuhk.edu.hk/intentsearch for the experience.
Jingyu Cui, Fang Wen 0001, Xiaoou Tang
ACM Multimedia3
2008 EasyToon: cartoon personalization using face photos
abstract
In this demo, we present a family photo album based cartoon personalization system, EasyToon. Using the family photo album as the candidate pool, a personalized cartoon image is obtained in two main steps. First, the best face candidate is selected from the album interactively. Then a personalized cartoon image is automatically synthesized by lending the selected face into the target cartoon image. By integrating state of the art computer vision and graphics technologies and effective UI design EasyToon can generate a personalized cartoon storyboard easily and quickly.
Fang Wen 0001, Shifeng Chen, Xiaoou Tang
ACM Multimedia3
2008 Cyclizing Clusters via Zeta Function of a Graph
abstract
Detecting underlying clusters from large-scale data plays a central role in machine learning research. In this paper, we attempt to tackle clustering problems for complex data of multiple distributions and large multi-scales. To this end, we develop an algorithm named Zeta $l$-links, or Zell which consists of two parts: Zeta merging with a similarity graph and an initial set of small clusters derived from local $l$-links of the graph. More specifically, we propose to structurize a cluster using cycles in the associated subgraph. A mathematical tool, Zeta function of a graph, is introduced for the integration of all cycles, leading to a structural descriptor of the cluster in determinantal form. The popularity character of the cluster is conceptualized as the global fusion of variations of the structural descriptor by means of the leave-one-out strategy in the cluster. Zeta merging proceeds, in the agglomerative fashion, according to the maximum incremental popularity among all pairwise clusters. Experiments on toy data, real imagery data, and real sensory data show the promising performance of Zell. The $98.1\%$ accuracy, in the sense of the normalized mutual information, is obtained on the FRGC face data of 16028 samples and 466 facial clusters. The MATLAB codes of Zell will be made publicly available for peer evaluation.
Deli Zhao, Xiaoou Tang
NIPS2
2008 Limits of Learning-Based Superresolution Algorithms
Zhouchen Lin, Junfeng He, Xiaoou Tang, Chi-Keung Tang
Int. J. Comput. Vis.3
2008 A new extension of kernel feature and its application for visual recognition
Qingshan Liu 0001, Hongliang Jin, Xiaoou Tang, Hanqing Lu, Songde Ma
Neurocomputing3
2008 What the Back of the Object Looks Like: 3D Reconstruction from Line Drawings without Hidden Lines
abstract
The human vision system can interpret a single 2D line drawing as a 3D object without much difficulty even if the hidden lines of the object are invisible. Many reconstruction methods have been proposed to emulate this ability, but they cannot recover the complete object if the hidden lines of the object are not shown. This paper proposes a novel approach to reconstructing a complete 3D object, including the shape of the back of the object, from a line drawing without hidden lines. First, we develop theoretical constraints and an algorithm for the inference of the topology of the invisible edges and vertices of an object. Then we present a reconstruction method based on perceptual symmetry and planarity of the object. We show a number of examples to demonstrate the success of our approach.
Liangliang Cao, Jianzhuang Liu, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.3
2008 Plane-Based Optimization for 3D Object Reconstruction from Single Line Drawings
abstract
In previous optimization-based methods of 3D planar-faced object reconstruction from single 2D line drawings, the missing depths of the vertices of a line drawing (and other parameters in some methods) are used as the variables of the objective functions. A 3D object with planar faces is derived by finding values for these variables that minimize the objective functions. These methods work well for simple objects with a small number N of variables. As N grows, however, it is very difficult for them to find expected objects. This is because with the nonlinear objective functions in a space of large dimension N, the search for optimal solutions can easily get trapped into local minima. In this paper, we use the parameters of the planes that pass through the planar faces of an object as the variables of the objective function. This leads to a set of linear constraints on the planes of the object, resulting in a much lower dimensional nullspace where optimization is easier to achieve. We prove that the dimension of this nullspace is exactly equal to the minimum number of vertex depths which define the 3D object. Since a practical line drawing is usually not an exact projection of a 3D object, we expand the nullspace to a larger space based on the singular value decomposition of the projection matrix of the line drawing. In this space, robust 3D reconstruction can be achieved. Compared with two most related methods, our method not only can reconstruct more complex 3D objects from 2D line drawings, but also is computationally more efficient.
Jianzhuang Liu, Liangliang Cao, Zhenguo Li, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.4
2008 Which Components are Important for Interactive Image Searching?
abstract
With many potential industrial applications, content-based image retrieval (CBIR) has recently gained more attention for image management and web searching. As an important tool to capture users' preferences and thus to improve the performance of CBIR systems, a variety of relevance feedback (RF) schemes have been developed in recent years. One key issue in RF is: which features (or feature dimensions) can benefit this human-computer iteration procedure? In this paper, we make theoretical and practical comparisons between principal and complement components of image features in CBIR RF. Most of the previous RF approaches treat the positive and negative feedbacks equivalently although this assumption is not appropriate since the two groups of training feedbacks have very different properties. That is, all positive feedbacks share a homogeneous concept while negative feedbacks do not. We explore solutions to this important problem by proposing an orthogonal complement component analysis. Experimental results are reported on a real-world image collection to demonstrate that the proposed complement components method consistently outperforms the conventional principal components method in both linear and kernel spaces when users want to retrieve images with a homogeneous concept.
Dacheng Tao, Xiaoou Tang, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2008 Regression From Uncertain Labels and Its Applications to Soft Biometrics
abstract
In this paper, we investigate two soft-biometric problems: (1) age estimation and (2) pose estimation, within the scenario where uncertainties exist for the available labels of the training samples. These two tasks are generally formulated as the automatic design of a regressor from training samples with uncertain nonnegative labels. First, the nonnegative label is predicted as the Frobenius norm of a matrix, which is bilinearly transformed from the nonlinear mappings of a set of candidate kernels. Two transformation matrices are then learned for deriving such a matrix by solving two semidefinite programming (SDP) problems, in which the uncertain label of each sample is expressed as two inequality constraints. The objective function of SDP controls the ranks of these two matrices and, consequently, automatically determines the structure of the regressor. The whole framework for the automatic design of a regressor from samples with uncertain nonnegative labels has the following characteristics: (1) the SDP formulation makes full use of the uncertain labels, instead of using conventional fixed labels; (2) regression with the Frobenius norm of matrix naturally guarantees the nonnegativity of the labels, and greater prediction capability is achieved by integrating the squares of the matrix elements, which to some extent act as weak regressors; and (3) the regressor structure is automatically determined by the pursuit of simplicity, which potentially promotes the algorithmic generalization capability. Extensive experiments on two human age databases: (1) FG-NET and (2) Yamaha, and the Pointing'04 head pose database, demonstrate encouraging estimation accuracy improvements over conventional regression algorithms without taking the uncertainties within the labels into account.
Shuicheng Yan, Huan Wang 0001, Xiaoou Tang, Jianzhuang Liu, Thomas S. Huang
IEEE Trans. Inf. Forensics Secur.3
2007 EasyAlbum: an interactive photo annotation system based on face clustering and re-ranking
abstract
Digital photo management is becoming indispensable for the explosively growing family photo albums due to the rapid popularization of digital cameras and mobile phone cameras. In an effective photo management system photo annotation is the most challenging task. In this paper, we develop several innovative interaction techniques for semi-automatic photo annotation. Compared with traditional annotation systems, our approach provides the following new features: "cluster annotation" puts similar faces or photos with similar scene together, and enables user label them in one operation; "contextual re-ranking" boosts the labeling productivity by guessing the user intention; "ad hoc annotation" allows user label photos while they are browsing or searching, and improves system performance progressively through learning propagation. Our results show that these technologies provide a more user friendly interface for the annotation of person name, location, and event, and thus substantially improve the annotation performance especially for a large photo album.
Jingyu Cui, Fang Wen 0001, Rong Xiao 0003, Yuandong Tian, Xiaoou Tang
CHI5
2007 Iterative MAP and ML Estimations for Image Segmentation
abstract
Image segmentation plays an important role in computer vision and image analysis. In this paper, the segmentation problem is formulated as a labeling problem under a probability maximization framework. To estimate the label configuration, an iterative optimization scheme is proposed to alternately carry out the maximum a posteriori (MAP) estimation and the maximum-likelihood (ML) estimation. The MAP estimation problem is modeled with Markov random fields (MRFs). A graph-cut algorithm is used to find the solution to the MAP-MRF estimation. The ML estimation is achieved by finding the means of region features. Our algorithm can automatically segment an image into regions with relevant textures or colors without the need to know the number of regions in advance. In addition, under the same framework, it can be extended to another algorithm that extracts objects of a particular class from a group of images. Extensive experiments have shown the effectiveness of our approach.
Shifeng Chen, Liangliang Cao, Jianzhuang Liu, Xiaoou Tang
CVPR4
2007 Discriminant Mutual Subspace Learning for Indoor and Outdoor Face Recognition
abstract
Outdoor face recognition is among the most challenging problems for face recognition. In this paper, we develop a discriminant mutual subspace learning algorithm for indoor and outdoor face recognition. Unlike traditional algorithms using one subspace to model both indoor and outdoor face images, our algorithm simultaneously learn two related subspaces for indoor and outdoor images respectively thus can better model both. To further improve the recognition performance we develop a DMSL-based multi-classifier fusion framework on Gabor images using a new fusion method called adaptive informative fusion scheme. Experimental results clearly show that this framework can greatly enhance the recognition performance.
Zhifeng Li 0001, Dahua Lin, Helen M. Meng, Xiaoou Tang
CVPR4
2007 A Closed-form Solution to 3D Reconstruction of Piecewise Planar Objects from Single Images
abstract
This paper proposes a new approach to 3D reconstruction of piecewise planar objects based on two image regularities, connectivity and perspective symmetry. First, we formulate the whole shape of the objects in an image as a shape vector consisting of the normals of all the faces of the objects. Then, we impose several linear constraints on the shape vector using connectivity and perspective symmetry of the objects. Finally, we obtain a closed-form solution to the 3D reconstruction problem. We also develop an efficient algorithm to detect a face of perspective symmetry. Experimental results on real images are shown to demonstrate the effectiveness of our approach.
Zhenguo Li, Jianzhuang Liu, Xiaoou Tang
CVPR3
2007 Quality-Driven Face Occlusion Detection and Recovery
abstract
This paper presents a framework to automatically detect and recover the occluded facial region. We first derive a Bayesian formulation unifying the occlusion detection and recovery stages. Then a quality assessment model is developed to drive both the detection and recovery processes, which captures the face priors in both global correlation and local patterns. Based on this formulation, we further propose GraphCut-based detection and confidence-oriented sampling to attain optimal detection and recovery respectively. Compared to traditional works in image repairing, our approach is distinct in three aspects: (1) it frees the user from marking the occlusion area by incorporating an automatic occlusion detector; (2) it learns a face quality model as a criterion to guide the whole procedure; (3) it couples the detection and occlusion stages to simultaneously achieve two goals: accurate occlusion detection and high quality recovery. The comparative experiments show that our method can recover the occluded faces with both the global coherence and local details well preserved.
Dahua Lin, Xiaoou Tang
CVPR2
2007 Learning to Detect A Salient Object
abstract
We study visual attention by detecting a salient object in an input image. We formulate salient object detection as an image segmentation problem, where we separate the salient object from the image background. We propose a set of novel features including multi-scale contrast, center-surround histogram, and color spatial distribution to describe a salient object locally, regionally, and globally. A conditional random field is learned to effectively combine these features for salient object detection. We also constructed a large image database containing tens of thousands of carefully labeled images by multiple users. To our knowledge, it is the first large image database for quantitative evaluation of visual attention algorithms. We validate our approach on this image database, which is public available with this paper.
Jian Sun 0001, Nanning Zheng 0001, Xiaoou Tang, Harry Shum
CVPR4
2007 Offline Signature Verification Using Online Handwriting Registration
abstract
This paper proposes a novel framework for offline signature verification. Different from previous methods, our approach makes use of online handwriting instead of handwritten images for registration. The online registrations enable robust recovery of the writing trajectory from an input offline signature and thus allow effective shape matching between registration and verification signatures. In addition, we propose several new techniques to improve the performance of the new signature verification system: 1. we formulate and solve the recovery of writing trajectory within the framework of conditional random fields; 2. we propose a new shape descriptor, online context, for aligning signatures; 3. we develop a verification criterion which combines the duration and amplitude variances of handwriting. Experiments on a benchmark database show that the proposed method significantly outperforms the well-known offline signature verification methods and achieve comparable performance with online signature verification methods.
Yu Qiao 0001, Jianzhuang Liu, Xiaoou Tang
CVPR3
2007 Flash Cut: Foreground Extraction with Flash and No-flash Image Pairs
abstract
In this paper, we propose a novel approach for foreground layer extraction using flash/no-flash image pairs, which we call flash cut. Flash cut is based on the simple observation that only the foreground is significantly brightened by the flash and the background appearance change is very small, if the background is distant. Changes due to flash, motion, and color information are fused in an MRF framework to produce high quality segmentation results. Flash cut handles some amount of camera shake, and foreground motion, which makes it practical for anyone with a flash-equipped camera to use. We validate our approach on a variety of indoor and outdoor examples.
Jian Sun 0009, Jian Sun 0001, Sing Bing Kang, Zongben Xu, Xiaoou Tang, Harry Shum
CVPR5
2007 A Face Annotation Framework with Partial Clustering and Interactive Labeling
abstract
Face annotation technology is important for a photo management system. In this paper, we propose a novel interactive face annotation framework combining unsupervised and interactive learning. There are two main contributions in our framework. In the unsupervised stage, a partial clustering algorithm is proposed to find the most evident clusters instead of grouping all instances into clusters, which leads to a good initial labeling for later user interaction. In the interactive stage, an efficient labeling procedure based on minimization of both global system uncertainty and estimated number of user operations is proposed to reduce user interaction as much as possible. Experimental results show that the proposed annotation framework can significantly reduce the face annotation workload and is superior to existing solutions in the literature.
Yuandong Tian, Wei Liu 0026, Rong Xiao 0003, Fang Wen 0001, Xiaoou Tang
CVPR5
2007 Trace Ratio vs. Ratio Trace for Dimensionality Reduction
abstract
A large family of algorithms for dimensionality reduction end with solving a Trace Ratio problem in the form of arg maxWTr(WTSPW)/Tr(WT SIW)1, which is generally transformed into the corresponding Ratio Trace form arg maxWTr[ (WTSIW)-1(WTSPW) ] for obtaining a closed-form but inexact solution. In this work, an efficient iterative procedure is presented to directly solve the Trace Ratio problem. In each step, a Trace Difference problem arg maxWTr [WT(SP- lambdaSI) W] is solved with lambda being the trace ratio value computed from the previous step. Convergence of the projection matrix W, as well as the global optimum of the trace ratio value lambda, are proven based on point-to-set map theories. In addition, this procedure is further extended for solving trace ratio problems with more general constraint WTCW=I and providing exact solutions for kernel-based subspace learning problems. Extensive experiments on faces and UCI data demonstrate the high convergence speed of the proposed solution, as well as its superiority in classification capability over corresponding solutions to the ratio trace problem.
Huan Wang 0001, Shuicheng Yan, Dong Xu 0001, Xiaoou Tang, Thomas S. Huang
CVPR4
2007 Linear Laplacian Discrimination for Feature Extraction
abstract
Discriminant feature extraction plays a fundamental role in pattern recognition. In this paper, we propose the linear Laplacian discrimination (LLD) algorithm/or discriminant feature extraction. LLD is an extension of linear discriminant analysis (LDA). Our motivation is to address the issue that LDA cannot work well in cases where sample spaces are non-Euclidean. Specifically, we define the within-class scatter and the between-class scatter using similarities which are based on pairwise distances in sample spaces. Thus the structural information of classes is contained in the within-class and the between-class Laplacian matrices which are free from metrics of sample spaces. The optimal discriminant subspace can be derived by controlling the structural evolution of Laplacian matrices. Experiments are performed on the facial database for FRGC version 2. Experimental results show that LLD is effective in extracting discriminant features.
Deli Zhao, Zhouchen Lin, Rong Xiao 0003, Xiaoou Tang
CVPR4
2007 Exploring Feature Descritors for Face Recognition
abstract
How to encode a face is a widely studied problem in both pattern recognition and psychology literatures. Many feature descriptors, Gabor feature, local binary pattern (LBP), and edge orientation histogram, have been proposed. In this paper, we give a comprehensive study of these descriptors under the framework of principal component analysis (PCA) followed by linear discriminant analysis (LDA), compared on three different popular similarity measures and two different feature correspondence strategies: holistic and local. Moreover, we present a new feature descriptor named multi-radius LBP, and also propose a combination scheme for the LBP and Gabor descriptor. The experiments on the Purdue and CMU PIE databases demonstrate that 1) an obvious recognition boost of LBP is achieved under PCA+LDA framework compared to the direct NN classification; 2) the LBP and Gabor features are comparable as well as mutually complementary, and the combination of these two descriptors brings a significant improvement in classification capability over single ones; and 3) the multi-radius LBP shows to outperform all the state-of-the-art feature descriptors.
Shuicheng Yan, Huan Wang 0001, Xiaoou Tang, Thomas S. Huang
ICASSP (1)3
2007 A Divide-and-Conquer Approach to 3D Object Reconstruction from Line Drawings
abstract
3D object reconstruction from a single 2D line drawing is an important problem in both computer vision and graphics. Many methods have been put forward to solve this problem, but they usually fail when the geometric structure of a 3D object becomes complex. In this paper, a novel approach based on a divide-and-conquer strategy is proposed to handle 3D reconstruction of complex manifold objects from single 2D line drawings. The approach consists of three steps: 1) dividing a complex line drawing into multiple simpler line drawings based on the result efface identification; 2) reconstructing the 3D shapes from these simpler line drawings; and 3) merging the 3D shapes into one complete object represented by the original line drawing. A number of examples are given to show that our approach can handle 3D reconstruction of more complex objects than previous methods.
Yu Chen 0009, Jianzhuang Liu, Xiaoou Tang
ICCV3
2007 Noise Robust Spectral Clustering
abstract
This paper aims to introduce the robustness against noise into the spectral clustering algorithm. First, we propose a warping model to map the data into a new space on the basis of regularization. During the warping, each point spreads smoothly its spatial information to other points. After the warping, empirical studies show that the clusters become relatively compact and well separated, including the noise cluster that is formed by the noise points. In this new space, the number of clusters can be estimated by eigenvalue analysis. We further apply the spectral mapping to the data to obtain a low-dimensional data representation. Finally, the K-means algorithm is used to perform clustering. The proposed method is superior to previous spectral clustering methods in that (i) it is robust against noise because the noise points are grouped into one new cluster; (ii) the number of clusters and the parameters of the algorithm are determined automatically. Experimental results on synthetic and real data have demonstrated this superiority.
Zhenguo Li, Jianzhuang Liu, Shifeng Chen, Xiaoou Tang
ICCV4
2007 Limits of Learning-Based Superresolution Algorithms
abstract
Learning-based superresolution (SR) are popular SR techniques that use application dependent priors to infer the missing details in low resolution images (LRIs). However, their performance still deteriorates quickly when the magnification factor is moderately large. This leads us to an important problem: "Do limits of learning-based SR algorithms exist?" In this paper, we attempt to shed some light on this problem when the SR algorithms are designed for general natural images (GNIs). We first define an expected risk for the SR algorithms that is based on the root mean squared error between the superresolved images and the ground truth images. Then utilizing the statistics of GNIs, we derive a closed form estimate of the lower bound of the expected risk. The lower bound can be computed by sampling real images. By computing the curve of the lower bound w.r.t. the magnification factor, we can estimate the limits of learning-based SR algorithms, at which the lower bound of expected risk exceeds a relatively large threshold. We also investigate the sufficient number of samples to guarantee an accurate estimation of the lower bound.
Zhouchen Lin, Junfeng He, Xiaoou Tang, Chi-Keung Tang
ICCV3
2007 Interactive Offline Tracking for Color Objects
abstract
In this paper, we present an interactive offline tracking system for generic color objects. The system achieves 60- 100 fps on a 320 times 240 video. The user can therefore easily refine the tracking result in an interactive way. To fully exploit user input and reduce user interaction, the tracking problem is addressed in a global optimization framework. The optimization is efficiently performed through three steps. First, from user's input we train a fast object detector that locates candidate objects in the video based on proposed features called boosted color bin. Second, we exploit the temporal coherence to generate multiple object trajectories based on a global best-first strategy. Last, an optimal object path is found by dynamic programming.
Jian Sun 0001, Xiaoou Tang, Harry Shum
ICCV3
2007 Dynamic Cascades for Face Detection
abstract
In this paper, we propose a novel method, called "dynamic cascade", for training an efficient face detector on massive data sets. There are three key contributions. The first is a new cascade algorithm called "dynamic cascade ", which can train cascade classifiers on massive data sets and only requires a small number of training parameters. The second is the introduction of a new kind of weak classifier, called "Bayesian stump", for training boost classifiers. It produces more stable boost classifiers with fewer features. Moreover, we propose a strategy for using our dynamic cascade algorithm with multiple sets of features to further improve the detection performance without significant increase in the detector's computational cost. Experimental results show that all the new techniques effectively improve the detection performance. Finally, we provide the first large standard data set for face detection, so that future researches on the topic can be compared on the same training and testing set.
Rong Xiao 0003, Huaiyi Zhu, He Sun 0003, Xiaoou Tang
ICCV4
2007 Learning Auto-Structured Regressor from Uncertain Nonnegative Labels
abstract
In this paper, we take the human age and pose estimation problems as examples to study automatic designing regressor from training samples with uncertain nonnegative labels. First, the nonnegative label is predicted as the square norm of a matrix, which is bilinearly transformed from the nonlinear mappings of the candidate kernels. Two transformation matrices are then learned for deriving such a matrix by solving a semi definite programming (SDP) problem, in which the uncertain label of each sample is expressed as two inequality constraints. The objective function of SDP controls the ranks of these two matrices, and consequently automatically determines the structure of the regressor. The whole framework for automatic designing regressor from samples with uncertain nonnegative labels has the following characteristics: 1) SDP formulation makes full use of the uncertain labels, instead of using conventional fixed labels; 2) regression with matrix norm naturally guarantees the nonnegativity of the labels, and greater prediction capability is achieved by integrating the squares of the matrix elements, which act as weak regressors; and 3) the regressor structure is automatically determined by the pursuit of simplicity, which potentially promotes the algorithmic generalization capability. Extensive experiments on two human age databases, FG-NET and Yamaha, as well as the Pointing'04 pose database, demonstrate encouraging estimation accuracy improvements over conventional regression algorithms.
Shuicheng Yan, Huan Wang 0001, Xiaoou Tang, Thomas S. Huang
ICCV3
2007 Contextual Distance for Data Perception
abstract
Structural perception of data plays a fundamental role in pattern analysis and machine learning. In this paper, we develop a new structural perception of data based on local contexts. We first identify the contextual set of a point by finding its nearest neighbors. Then the contextual distance between the point and one of its neighbors is defined by the difference between their contribution to the integrity of the geometric structure of the contextual set, which is depicted by a structural descriptor. The centroid and the coding length are introduced as the examples of descriptors of the contextual set. Furthermore, a directed graph (digraph) is built to model the asymmetry of perception. The edges of the digraph are weighted based on the contextual distances. Thus direction is brought to the undirected data. And the structural perception of data can be performed by mining the properties of the digraph. We also present the method for deriving the global digraph Laplacian from the alignment of the local digraph Laplacians. Experimental results on clustering and ranking of toy problems and real data show the superiority of asymmetric perception.
Deli Zhao, Zhouchen Lin, Xiaoou Tang
ICCV3
2007 Laplacian PCA and Its Applications
abstract
Dimensionality reduction plays a fundamental role in data processing, for which principal component analysis (PCA) is widely used. In this paper, we develop the Laplacian PCA (LPCA) algorithm which is the extension of PCA to a more general form by locally optimizing the weighted scatter. In addition to the simplicity of PCA, the benefits brought by LPCA are twofold: the strong robustness against noise and the weak metric-dependence on sample spaces. The LPCA algorithm is based on the global alignment of locally Gaussian or linear subspaces via an alignment technique borrowed from manifold learning. Based on the coding length of local samples, the weights can be determined to capture the local principal structure of data. We also give the exemplary application of LPCA to manifold learning. Manifold unfolding (non-linear dimensionality reduction) can be performed by the alignment of tangential maps which are linear transformations of tangent coordinates approximated by LPCA. The superiority of LPCA to PCA and kernel PCA is verified by the experiments on face recognition (FRGC version 2 face database) and manifold (Scherk surface) unfolding.
Deli Zhao, Zhouchen Lin, Xiaoou Tang
ICCV3
2007 A Two-Stage Fusion Scheme using Multiple Fingerprint Impressions
abstract
In this paper, we propose a two-stage fusion scheme that takes full advantage of the complementary information among multiple fingerprint impressions. While comparing the query fingerprint with a template impression, all the other impressions are also transformed using the 2D warping model to register with the query fingerprint so that the additive matched minutiae pairs can be detected to improve the matching result with a subset combination scheme. Then a matching score level fusion or decision level fusion is performed to integrate the improved matching results corresponding to different impressions. Experiments conducted on FVC2002 show that the proposed method produces a much better performance for fingerprint matching.
Lifeng Sha, Feng Zhao 0004, Xiaoou Tang
ICIP (2)3
2007 Special Effects in Film Making with Object Based Transformations
abstract
The new media initiative project of Ryerson University had been successful in applying image and video processing techniques in creating special effects in film making. Our system utilizes a graph cut image segmentation and snakes active contour approach to obtain object cut outs in 3-D with tracking. It uses exemplar-based inpainting for background filling for missing regions uncovered by object transformation. The two methods combined allows for easy video object editing with reduction in user input. Previously implemented shot detection by twin window amplification method and steerable pyramids texture generation is also part of the system. Lastly, a set of transformation was implemented that serves as a basis for testing the system.
Chun-Hao Wang, Xiaoming Fan, Bruce Elder, Xiaoou Tang, Ling Guan
ICME5
2007 Ranking with Uncertain Labels
abstract
Most techniques for image analysis consider the image labels fixed and without uncertainty. In this paper, we address the problem of ordinal/rank label prediction based on training samples with uncertain labels. First, the core ranking model is designed as the bilinear fusing of multiple candidate kernels. Then, the parameters for feature selection and kernel selection are learned by maximum a posteriori for given samples and uncertain labels. The convergency provable Expectation-Maximization (EM) method is used for inferring these parameters. The effectiveness of the proposed algorithm is finally validated by the extensive experiments on age ranking task. The FG-NET and Yamaha aging database are used for the experiments, and our algorithm significantly outperforms those state-of-the-art algorithms ever reported in literature.
Shuicheng Yan, Huan Wang 0001, Thomas S. Huang, Qiong Yang, Xiaoou Tang
ICME5
2007 Transductive regression piloted by inter-manifold relations
abstract
In this paper, we present a novel semisupervised regression algorithm working on multiclass data that may lie on multiple manifolds. Unlike conventional manifold regression algorithms that do not consider the class distinction of samples, our method introduces the class information to the regression process and tries to exploit the similar configurations shared by the label distribution of multi-class data. To utilize the correlations among data from different classes, we develop a cross-manifold label propagation process and employ labels from different classes to enhance the regression performance. The interclass relations are coded by a set of intermanifold graphs and a regularization item is introduced to impose inter-class smoothness on the possible solutions. In addition, the algorithm is further extended with the kernel trick for predicting labels of the out-of-sample data even without class information. Experiments on both synthesized data and real world problems validate the effectiveness of the proposed framework for semisupervised regression.
Huan Wang 0001, Shuicheng Yan, Thomas S. Huang, Jianzhuang Liu, Xiaoou Tang
ICML5
2007 Directed Graph Embedding
Qiong Yang, Xiaoou Tang
IJCAI3
2007 Bayesian Tensor Inference for Sketch-Based Facial Photo Hallucination
Wei Liu 0005, Xiaoou Tang, Jianzhuang Liu
IJCAI2
2007 A Convengent Solution to Tensor Subspace Learning
Huan Wang 0001, Shuicheng Yan, Thomas S. Huang, Xiaoou Tang
IJCAI4
2007 Image matting using linear optimization
abstract
An image can be assumed to be a composite of the foreground and the background. The foreground and the background of each pixel are linearly combined in terms of this pixel's foreground opacity (called alpha). Image matting is the process of estimating the foreground, the background and the alpha for each pixel. In this paper, we transform the ill-posed image matting problem into two over-determined linear optimization problems by introducing two medium variables and imposing smoothness constraints. Closed form solutions can be obtained from the two problems. Extensive experimental results indicate that our algorithm can generate high-quality matting results.
Shifeng Chen, Zhenguo Li, Jianzhuang Liu, Xiaoou Tang
ACM Multimedia4
2007 Image inpainting by global structure and texture propagation
abstract
Image inpainting is a technique to repair damaged images or modify images in a non-detectable form. In this paper, a novel global algorithm for region filling is proposed for image inpainting. After removing objects from an image, our approach fills the regions using patches taken from the image. The filling process is formulated as an energy minimization problem by Markov random fields (MRFs) and the belief propagation (BP) is utilized to solve the problem. Our energy function includes structure and texture information obtained from the image. One challenge in using BP is that its computational complexity is the square of the number of label candidates. To reduce the large number of label candidates, we present a coarse-to-fine scheme where two BPs run with much smaller numbers of label candidates instead of one BP running with a large number of label candidates. Experimental results demonstrate the excellent performance of our algorithm over other related algorithms.
Huang Ting, Shifeng Chen, Jianzhuang Liu, Xiaoou Tang
ACM Multimedia4
2007 Ranking with uncertain labels and its applications
Shuicheng Yan, Huan Wang 0001, Jianzhuang Liu, Xiaoou Tang, Thomas S. Huang
Frontiers Comput. Sci. China4
2007 Stereo Correspondence with Occlusion Handling in a Symmetric Patch-Based Graph-Cuts Model
abstract
A novel patch-based correspondence model is presented in this paper. Many segment-based correspondence approaches have been proposed in recent years. Untextured pixels and boundaries of discontinuities are imposed with hard constraints by the discontinuity assumption that large disparity variation only happens at the boundaries of segments in the above approaches. Significant improvements on performance of untextured and discontinuity area have been reported. But, the performance near occlusion is not satisfactory because a segmented region in one image may be only partially visible in the other one. To solve this problem, we utilize the observation that the shared edge of a visible area and an occluded area corresponds to the discontinuity in the other image. So, the proposed model conducts color segmentation on both images first and then a segment in one image is further cut into smaller patches corresponding to the boundaries of segments in the other when it is assigned with a disparity. Different visibility of patches in one segment is allowed. The uniqueness constraint in a segment level is used to compute the occlusions. An energy minimization framework using graph-cuts is proposed to find a global optimal configuration including both disparities and occlusions. Besides, some measurements are taken to make our segment-based algorithm suffer less from violation of the discontinuity assumption. Experimental results have shown superior performance of the proposed approach, especially on occlusions, untextured areas, and near discontinuities.
Qiong Yang, Xueyin Lin, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.4
2007 Preprocessing and postprocessing for skeleton-based fingerprint minutiae extraction
Feng Zhao 0004, Xiaoou Tang
Pattern Recognit.2
2007 Using Support Vector Machines to Enhance the Performance of Bayesian Face Recognition
abstract
In this paper, we first develop a direct Bayesian-based support vector machine (SVM) by combining the Bayesian analysis with the SVM. Unlike traditional SVM-based face recognition methods that require one to train a large number of SVMs, the direct Bayesian SVM needs only one SVM trained to classify the face difference between intrapersonal variation and extrapersonal variation. However, the additional simplicity means that the method has to separate two complex subspaces by one hyperplane thus affecting the recognition accuracy. In order to improve the recognition performance, we develop three more Bayesian-based SVMs, including the one-versus-all method, the hierarchical agglomerative clustering-based method, and the adaptive clustering method. Finally, we combine the adaptive clustering method with multilevel subspace analysis to further improve the recognition performance. We show the improvement of the new algorithms over traditional subspace methods through experiments on two face databases - the FERET database and the XM2VTS database
Zhifeng Li 0001, Xiaoou Tang
IEEE Trans. Inf. Forensics Secur.2
2007 A Parameter-Free Framework for General Supervised Subspace Learning
abstract
Supervised subspace learning techniques have been extensively studied in biometrics literature; however, there is little work dedicated to: 1) how to automatically determine the subspace dimension in the context of supervised learning, and 2) how to explicitly guarantee the classification performance on a training set. In this paper, by following our previous work on unified subspace learning framework in our earlier work, we present a general framework, called parameter-free graph embedding (PFGE) to solve the above two problems by posing a general supervised subspace learning task as a semidefinite programming problem. The semipositive feature Gram matrix, namely the product of the transformation matrix and its transpose, is derived by optimizing a trace difference form of an objective function extended from that in our earlier work with the constraints that guarantee the class homogeneity within the neighborhood of each datum. Then, the subspace dimension and the feature weights are simultaneously obtained via the singular value decomposition of the feature Gram matrix. In addition, to alleviate the computational complexity, the Kronecker product approximation of the feature Gram matrix is proposed by taking advantage of the essential matrix form of image pixels. The experiments on simulated data and real-world data demonstrate the capability of the new PFGE framework in estimating the subspace dimension for supervised learning as well as the superiority in classification performance over traditional algorithms for subspace learning
Shuicheng Yan, Jianzhuang Liu, Xiaoou Tang, Thomas S. Huang
IEEE Trans. Inf. Forensics Secur.3
2007 Formulating Face Verification With Semidefinite Programming
abstract
This paper presents a unified solution to three unsolved problems existing in face verification with subspace learning techniques: selection of verification threshold, automatic determination of subspace dimension, and deducing feature fusing weights. In contrast to previous algorithms which search for the projection matrix directly, our new algorithm investigates a similarity metric matrix (SMM). With a certain verification threshold, this matrix is learned by a semidefinite programming approach, along with the constraints of the kindred pairs with similarity larger than the threshold, and inhomogeneous pairs with similarity smaller than the threshold. Then, the subspace dimension and the feature fusing weights are simultaneously inferred from the singular value decomposition of the derived SMM. In addition, the weighted and tensor extensions are proposed to further improve the algorithmic effectiveness and efficiency, respectively. Essentially, the verification is conducted within an affine subspace in this new algorithm and is, hence, called the affine subspace for verification (ASV). Extensive experiments show that the ASV can achieve encouraging face verification accuracy in comparison to other subspace algorithms, even without the need to explore any parameters.
Shuicheng Yan, Jianzhuang Liu, Xiaoou Tang, Thomas S. Huang
IEEE Trans. Image Process.3
2007 Face Verification With Balanced Thresholds
abstract
The process of face verification is guided by a pre-learned global threshold, which, however, is often inconsistent with class-specific optimal thresholds. It is, hence, beneficial to pursue a balance of the class-specific thresholds in the model-learning stage. In this paper, we present a new dimensionality reduction algorithm tailored to the verification task that ensures threshold balance. This is achieved by the following aspects. First, feasibility is guaranteed by employing an affine transformation matrix, instead of the conventional projection matrix, for dimensionality reduction, and, hence, we call the proposed algorithm threshold balanced transformation (TBT). Then, the affine transformation matrix, constrained as the product of an orthogonal matrix and a diagonal matrix, is optimized to improve the threshold balance and classification capability in an iterative manner. Unlike most algorithms for face verification which are directly transplanted from face identification literature, TBT is specifically designed for face verification and clarifies the intrinsic distinction between these two tasks. Experiments on three benchmark face databases demonstrate that TBT significantly outperforms the state-of-the-art subspace techniques for face verification.
Shuicheng Yan, Dong Xu 0001, Xiaoou Tang
IEEE Trans. Image Process.3
2007 Multilinear Discriminant Analysis for Face Recognition
abstract
There is a growing interest in subspace learning techniques for face recognition; however, the excessive dimension of the data space often brings the algorithms into the curse of dimensionality dilemma. In this paper, we present a novel approach to solve the supervised dimensionality reduction problem by encoding an image object as a general tensor of second or even higher order. First, we propose a discriminant tensor criterion, whereby multiple interrelated lower dimensional discriminative subspaces are derived for feature extraction. Then, a novel approach, called k-mode optimization, is presented to iteratively learn these subspaces by unfolding the tensor along different tensor directions. We call this algorithm multilinear discriminant analysis (MDA), which has the following characteristics: 1) multiple interrelated subspaces can collaborate to discriminate different classes, 2) for classification problems involving higher order tensors, the MDA algorithm can avoid the curse of dimensionality dilemma and alleviate the small sample size problem, and 3) the computational cost in the learning stage is reduced to a large extent owing to the reduced data dimensions in k-mode optimization. We provide extensive experiments on ORL, CMU PIE, and FERET databases by encoding face images as second- or third-order tensors to demonstrate that the proposed MDA algorithm based on higher order tensors has the potential to outperform the traditional vector-based subspace learning algorithms, especially in the cases with small sample sizes.
Shuicheng Yan, Dong Xu 0001, Qiang Yang 0001, Lei Zhang 0001, Xiaoou Tang, HongJiang Zhang
IEEE Trans. Image Process.5
2007 Bilinear Analysis for Kernel Selection and Nonlinear Feature Extraction
abstract
This paper presents a unified criterion, Fisher + kernel criterion (FKC), for feature extraction and recognition. This new criterion is intended to extract the most discriminant features in different nonlinear spaces, and then, fuse these features under a unified measurement. Thus, FKC can simultaneously achieve nonlinear discriminant analysis and kernel selection. In addition, we present an efficient algorithm Fisher + kernel analysis (FKA), which utilizes the bilinear analysis, to optimize the new criterion. This FKA algorithm can alleviate the ill-posed problem existed in traditional kernel discriminant analysis (KDA), and usually, has no singularity problem. The effectiveness of our proposed algorithm is validated by a series of face-recognition experiments on several different databases.
Shuicheng Yan, Xiaoou Tang
IEEE Trans. Neural Networks4
2007 Rank-One Projections With Adaptive Margins for Face Recognition
abstract
In supervised dimensionality reduction, tensor representations of images have recently been employed to enhance classification of high dimensional data with small training sets. Previous approaches for handling tensor data have been formulated with tight restrictions on projection directions that, along with convergence issues and the assumption of Gaussian-distributed class data, limit its face-recognition performance. To overcome these problems, we propose a method of rank-one projections with adaptive margins (RPAM) that gives a provably convergent solution for tensor data over a more general class of projections, while accounting for margins between samples of different classes. In contrast to previous margin-based works which determine margin sample pairs within the original high dimensional feature space, RPAM aims instead to maximize the margins defined in the expected lower dimensional feature sub-space by progressive margin refinement after each rank-one projection. In addition to handling tensor data, vector-based variants of RPAM are presented for linear mappings and for nonlinear mappings using kernel tricks. Comprehensive experimental results demonstrate that RPAM brings significant improvement in face recognition over previous subspace learning techniques.
Dong Xu 0001, Stephen Lin 0001, Shuicheng Yan, Xiaoou Tang
IEEE Trans. Syst. Man Cybern. Part B4
2006 Boosting Multi-gabor Subspaces for Face Recognition
Qingshan Liu 0001, Hongliang Jin, Xiaoou Tang, Hanqing Lu, Songde Ma
ACCV (1)3
2006 Space-Time Video Montage
abstract
Conventional video summarization methods focus predominantly on summarizing videos along the time axis, such as building a movie trailer: The resulting video trailer tends to retain much empty space in the background of the video frames while discarding much informative video content due to size limit. In this paper we propose a novel spacetime video summarization method which we call space-time video montage. The method simultaneously analyzes both the spatial and temporal injbrmation distribution in a video sequence, and extracts the visually informative space-time portions of the input videos. The informative video porlions are represented in volumetric layers. The layers are then packrd together in a smull ouzput video volume such that the total amount of visual information in the video volume is maximized. To achieve the packing process, we develop a new algorithm based upon the first-fit and Graph cut optimization techniques. Since our method is uble to cut off spatially und temporally less informative portions, it is uble to generate much more compact yet highly informative output videos. The effecliveness of our method is validated by extensive experiments over a wide variety of videos.
Hong-Wen Kang, Yasuyuki Matsushita, Xiaoou Tang, Xue-Quan Chen
CVPR (2)3
2006 The Design of High-Level Features for Photo Quality Assessment
abstract
We propose a principled method for designing high level features forphoto quality assessment. Our resulting system can classify between high quality professional photos and low quality snapshots. Instead of using the bag of low-level features approach, we first determine the perceptual factors that distinguish between professional photos and snapshots. Then, we design high level semantic features to measure the perceptual differences. We test our features on a large and diverse dataset and our system is able to achieve a classification rate of 72% on this difficult task. Since our system is able to achieve a precision of over 90% in low recall scenarios, we show excellent results in a web image search application.
Yan Ke, Xiaoou Tang
CVPR (1)2
2006 Accurate Face Alignment using Shape Constrained Markov Network
abstract
In this paper, we present a shape constrained Markov network for accurate face alignment. The global face shape is defined as a set of weighted shape samples which are integrated into the Markov network optimization. These weighted samples provide structural constraints to make the Markov network more robust to local image noise. We propose a hierarchical Condensation algorithm to draw the shape samples efficiently. Specifically, a proposal density incorporating the local face shape is designed to generate more samples close to the image features for accurate alignment, based on a local Markov network search. A constrained regularization algorithm is also developed to weigh favorably those points that are already accurately aligned. Extensive experiments demonstrate the accuracy and effectiveness of our proposed approach.
Fang Wen 0001, Ying-Qing Xu, Xiaoou Tang, Harry Shum
CVPR (1)4
2006 Recognize High Resolution Faces: From Macrocosm to Microcosm
abstract
Human faces manifest distinct structures and characteristics when observed in different scales. Traditional face recognition techniques mainly rely on low-resolution face images, leading to the lost of significant information contained in the microscopic traits. In this paper, we introduce a multilayer framework for high resolution face recognition exploiting features in multiple scales. Each face image is factorized into four layers: global appearance, facial organs, skins, and irregular details. We employ Multilevel PCA followed by Regularized LDA to model global appearance and facial organs. However, the description of skin texture and irregular details, for which conventional vector representation are not suitable, brings forth the need of developing novel representations. To address the issue, Discriminative Multiscale Texton Features and SIFT-Activated Pictorial Structure are proposed to describe skin and subtle details respectively. To effectively combine the information conveyed by all layers, we further design an metric fusion algorithm adaptively placing emphasis onto the highly confident layers. Through systematic experiments, we identify different roles played by the layers and convincingly show that by utilizing their complementarities, our framework achieves remarkable performance improvement.
Dahua Lin, Xiaoou Tang
CVPR (2)2
2006 Pursuing Informative Projection on Grassmann Manifold
abstract
Inspired by the underlying relationship between classification capability and the mutual information, in this paper, we first establish a quantitative model to describe the information transmission process from feature extraction to final classification and identify the critical channel in this propagation path, and then propose a Maximum Effective Information Criteria for pursuing the optimal subspace in the sense of preserving maximum information that can be conveyed to final decision. Considering the orthogonality and rotation invariance properties of the solution space, we present a Conjugate Gradient method constrained on a Grassmann manifold to exploit the geometric traits of the solution space for enhancing the efficiency of optimization. Comprehensive experiments demonstrate that the framework integrating the Maximum Effective Information Criteria and Grassmann manifold-based optimization method significantly improves the classification performance.
Dahua Lin, Shuicheng Yan, Xiaoou Tang
CVPR (2)3
2006 Video Completion by Motion Field Transfer
abstract
Existing methods for video completion typically rely on periodic color transitions, layer extraction, or temporally local motion. However, periodicity may be imperceptible or absent, layer extraction is difficult, and temporally local motion cannot handle large holes. This paper presents a new approach for video completion using motion field transfer to avoid such problems. Unlike prior methods, we fill in missing video parts by sampling spatio-temporal patches of local motion instead of directly sampling color. Once the local motion field has been computed within the missing parts of the video, color can then be propagated to produce a seamless hole-free video. We have validated our method on many videos spanning a variety of scenes. We can also use the same approach to perform frame interpolation using motion fields from different videos.
Takaaki Shiratori, Yasuyuki Matsushita, Xiaoou Tang, Sing Bing Kang
CVPR (1)3
2006 Picture Collage
abstract
In this paper, we address a novel problem of automatically creating a picture collage from a group of images. Picture collage is a kind of visual image summary - to arrange all input images on a given canvas, allowing overlay, to maximize visible visual information. We formulate the picture collage creation problem in a Bayesian framework. The salient regions of each image are firstly extracted and represented as a set of weighted rectangles. Then, the image arrangement is formulated as a Maximum a Posterior (MAP) problem such that the output picture collage shows as many visible salient regions (without being overlaid by others) from all images as possible. Moreover, a very efficientMarkov chain Monte Carlo (MCMC) method is designed for the optimization. Applications to desktop image browsing and image search result summarization demonstrate the effectiveness of our approach.
Jingdong Wang 0001, Long Quan, Jian Sun 0001, Xiaoou Tang, Harry Shum
CVPR (1)4
2006 Joint Boosting Feature Selection for Robust Face Recognition
abstract
A fundamental challenge in face recognition lies in determining what facial features are important for the identification of faces. In this paper, a novel face recognition framework is proposed to address this problem. In our framework, 3D face models are used to synthesize a huge database of realistic face images which covers wide appearance variations of faces due to various pose, illumination, and expression changes. A novel feature selection algorithm which we call Joint Boosting is developed to extract discriminative face features using this massive database. The major contributions of this paper are: (1) With the help of 3D face models, a massive database of realistic virtual face images is generated to achieve robust feature selection; (2)Because the huge database covers a wide range of face variations, our feature selection procedure only needs to be trained once, and the selected feature set can be generalized to other face database without re-training; (3) We propose a new learning algorithm, Joint Boosting Algorithm, which is effective and efficient in learning directly from a massive database without having to convert face images to intra-personal and extra-personal difference images. This property is important for applying our algorithm to other general pattern recognition problems. Experimental results show that our method significantly improves recognition performance.
Rong Xiao 0003, Wu-Jun Li, Yuandong Tian, Xiaoou Tang
CVPR (2)4
2006 Rank-one Projections with Adaptive Margins for Face Recognition
abstract
In supervised dimensionality reduction, tensor representations of images have recently been employed to enhance classification of high-dimensional data with small training sets. To handle tensor data, this approach has been formulated with tight restrictions on projection directions that, along with convergence issues and the assumption of Gaussian distributed class data, limits its face recognition performance. To overcome these problems, we propose a method of rank-one projections with adaptive margins (RPAM) that gives a provably convergent solution for tensor data over a more general class of projections, while accounting for margins between samples of different classes. In contrast to previous margin based works which determine margin sample pairs within the original high dimensional space, RPAM instead aims to maximize the margins defined in the expected lower dimensional feature subspace by progressive margin refinement after each rank-one projection. In addition to handling tensor data, vector-based variants of RPAM are presented for linear mappings and for nonlinear mappings using kernel tricks. Comprehensive experimental results demonstrate that RPAM brings significant improvement in face recognition over previous subspace learning techniques.
Dong Xu 0001, Stephen Lin 0001, Shuicheng Yan, Xiaoou Tang
CVPR (1)4
2006 Learning Semantic Patterns with Discriminant Localized Binary Projections
abstract
In this paper, we present a novel approach to learning semantic localized patterns with binary projections in a supervised manner. The pursuit of these binary projections is reformulated into a problem of feature clustering, which optimizes the separability of different classes by taking the members within each cluster as the nonzero entries of a projection vector. An efficient greedy procedure is proposed to incrementally combine the sub-clusters by ensuring the cardinality constraints of the projections and the increase of the objective function. Compared with other algorithms for sparse representations, our proposed algorithm, referred to as Discriminant Localized Binary Projections (dlb), has the following characteristics: 1) dlb is supervised, hence is much more effective than other unsupervised sparse algorithms like Non-negative Matrix Factorization (NMF) in terms of classification power; 2) similar to NMF, dlb can derive spatially localized sparse bases; furthermore, the sparsity of dlb is controllable, and an interesting result is that the bases have explicit semantics in human perception, like eyes and mouth; and 3) classification with dlb is extremely efficient, and only addition operations are required for dimensionality reduction. Extensive experimental results show significant improvements of dlb in sparsity and face recognition accuracy in comparison to the state-of-the-art algorithms for dimensionality reduction and sparse representations.
Shuicheng Yan, Tianqiang Yuan, Xiaoou Tang
CVPR (1)3
2006 An Intensity Similarity Measure in Low-Light Conditions
François Alter, Yasuyuki Matsushita, Xiaoou Tang
ECCV (4)3
2006 Degen Generalized Cylinders and Their Properties
Liangliang Cao, Jianzhuang Liu, Xiaoou Tang
ECCV (1)3
2006 Detecting Doctored JPEG Images Via DCT Coefficient Analysis
Junfeng He, Zhouchen Lin, Lifeng Wang 0001, Xiaoou Tang
ECCV (3)4
2006 An Integrated Model for Accurate Shape Alignment
Fang Wen 0001, Xiaoou Tang, Ying-Qing Xu
ECCV (4)3
2006 Conditional Infomax Learning: An Integrated Framework for Feature Extraction and Fusion
Dahua Lin, Xiaoou Tang
ECCV (1)2
2006 Inter-modality Face Recognition
Dahua Lin, Xiaoou Tang
ECCV (4)2
2006 Spatio-temporal Embedding for Statistical Face Recognition from Video
Wei Liu 0026, Zhifeng Li 0001, Xiaoou Tang
ECCV (2)3
2006 Background Cut
Jian Sun 0001, Xiaoou Tang, Harry Shum
ECCV (2)3
2006 Trace Quotient Problems Revisited
Shuicheng Yan, Xiaoou Tang
ECCV (2)2
2006 Salience Preserving Image Fusion with Dynamic Range Compression
abstract
Gradient conveys important salient features in images. Traditional fusion methods based on gradient generally treat gradients from multichannels as a multi-valued vector, and compute its global statistics under the assumption of identical distribution. However, different source channels may reflect different important salient features, and their gradients are basically non-identically distributed. This prevents existing methods from successful salience preservation. In this paper, we propose to fuse the gradients from multi-channels in the concept of saliency. We first measure the salience map of each channel's gradient, and then use their saliency to weight their contribution in computing the global statistics. Gradients with high saliency are properly highlighted in the target gradient, and thereby salient features in the sources are well preserved. Furthermore, we handle the dynamic range problem by applying range compression on the target gradient, and thereby halo effect is effectively reduced.
Qiong Yang, Xiaoou Tang, Zhongfu Ye
ICIP3
2006 Special Effects in Film/Video Making: A New Media Initiative Project
abstract
We present a system and a set of tools for producing special effects in film/video making by applying image processing and human centered computing techniques. A combination of shot detection, object segmentation, background generation, and image warping techniques are used. The user selects the image frame or the object of interest, and the image warp transformation to be used from the GUI. Wrapping can be performed either on a whole image sequence or on an object of interest in the sequence. In the latter, the object is first segmented and its motion tracked. Object segmentation is achieved either by snakes or graph cuts. Steerable pyramid background generation is then used to fill in the portion cut from the foreground
Chun-Hao Wang, Yongjin Wang, Meifeng Lian, Bruce Elder, Xiaoou Tang, Ling Guan
ICME5
2006 Locally adaptive classification piloted by uncertainty
abstract
Locally adaptive classifiers are usually superior to the use of a single global classifier. However, there are two major problems in designing locally adaptive classifiers. First, how to place the local classifiers, and, second, how to combine them together. In this paper, instead of placing the classifiers based on the data distribution only, we propose a responsibility mixture model that uses the uncertainty associated with the classification at each training sample. Using this model, the local classifiers are placed near the decision boundary where they are most effective. A set of local classifiers are then learned to form a global classifier by maximizing an estimate of the probability that the samples will be correctly classified with a nearest neighbor classifier. Experimental results on both artificial and real-world data sets demonstrate its superiority over traditional algorithms.
Juan Dai, Shuicheng Yan, Xiaoou Tang, James T. Kwok
ICML3
2006 3D object retrieval using 2D line drawing and graph based relevance reedback
abstract
This paper aims to provide a user-friendly interface for 3D object retrieval. In previous 3D retrieval systems, the user mainly uses two methods to input a query: providing an existing 3D objects, or providing partial shape information of desired objects such as text and 2D shapes. The first method fails when the user does not have a similar 3D object in hand, and the second method cannot sufficiently describe 3D shapes of objects. We believe that the best way is to have a good interface that can convert a 2D sketch drawn by the user into a 3D object as the query. A 2D line drawing is easy to be drawn and is the simplest and most direct way of illustrating a 3D object. In this paper, we develop an interface of 3D object reconstruction from line drawings, which allows the user to draw line drawings of objects with both planar and curved surfaces. In addition, in order to refine the retrieved results, we develop a relevance feedback algorithm based on a novel graph discriminant analysis. Compared with recently published relevance feedback algorithms, our algorithm achieves better retrieval performance.
Liangliang Cao, Jianzhuang Liu, Xiaoou Tang
ACM Multimedia3
2006 Shape from regularities for interactive 3D reconstruction of piecewise planar objects from single images
abstract
3D object reconstruction from single 2D images has many applications in multimedia. This paper proposes an approach based on image regularities such as connectivity, parallelism, and orthogonality possessed by the objects with simple user interactions. It is assumed that the objects are piecewise planar. By representing the 3D objects as a shape vector consisting of the normals of the faces of the objects, we impose geometric constraints on this shape vector using the regularities of the objects. We derive a system of equations in terms of the shape vector and the focal length, which we can solve for the shape vector optimally. Experimental results on real images are shown to demonstrate the effectiveness of this method.
Zhenguo Li, Jianzhuang Liu, Xiaoou Tang
ACM Multimedia3
2006 Detecting irregularity in videos using kernel estimation and KD trees
abstract
Automatic event understanding is the ultimate goal for many visual surveillance systems. In this paper, we propose a novel approach for on-line detecting unusual human activities in videos without the need to explicitly define all valid configurations. Within the framework of Bayesian inference, the detection process is formulated as an MAP estimation where we attempt to find whether activities in new video segments have similar activities in a video database. Our approach has three contributions: firstly, we build the statistical representation of normal behaviors in the database using nonparametric kernel density estimation; secondly, local feature descriptors are highly compressed using PCA and stored in a K-D tree structure, making the search for behavior-based similarity fast and effective; thirdly, the K-D trees are used to generate multiple hypotheses which compete for the optimal classification. The approach requires no tracking, no explicit motion estimation, and no predefined class-based templates. Experimental results have validated our approach in many real-world video sequences.
Chunjing Xu, Jianzhuang Liu, Xiaoou Tang
ACM Multimedia4
2006 Progressive cut
abstract
Recently, interactive image cutout technique becomes prevalent for image segmentation problem due to its easy-to-use nature. However, most existing stroke-based interactive object cutout system did not consider the user intention inherent in the user interaction process. Strokes in sequential steps are treated as a collection rather than a process, and only the color information of the additional stroke is used to update the color model in the graph cut framework. Accordingly, unexpected fluctuation effect may occur during the process of interactive object cutout. In fact, each step of user interaction reflects the user's evaluation of previous result and his/her intention. By analyzing the user's intention behind the interaction, we propose a progressive cut algorithm, which explicitly models the user's intention into a graph cut framework for the object cutout task. Three aspects of user intention are utilized: 1) the color of the stroke indicates the kind of change s/he expects, 2) the location of the stroke indicates the region of interest, 3) the relative position between the stroke and the previous result indicates the segmentation error. By incorporating such information into the cutout system, the new algorithm removes the unexpected fluctuation effect of existing stroke-based graph-cut methods, and thus provides the user a more controllable result with fewer strokes and faster visual feedback. Experiments and user study show the strength of progressive cut in accuracy, speed, controllability, and user experience.
Qiong Yang, Xiaoou Tang, Zhongfu Ye
ACM Multimedia4
2006 Maximum unfolded embedding: formulation, solution, and application for image clustering
abstract
In this paper, we present a novel spectral analysis algorithm for image clustering. First, the image manifold is embedded onto a low-dimensional feature space with dual objectives, i.e., maximizing the distances of faraway sample pairs meanwhile preserving the local manifold structure, which essentially results in a Trace Ratio optimization problem. Then an efficient iterative procedure is proposed to directly optimize the trace ratio and finally the clustering process is implemented on the derived low-dimensional embedding. Moreover, the linear approximation is also presented for handling the out-of-sample data. Experimental results show that our algorithm, referred to as Maximum Unfolded Embedding, brings an encouraging improvement in clustering accuracy over the state-of-the-art algorithms, such as K-Means, PCA-Kmeans, normalized cut \cite shi00normalized, and Locality Preserving Clustering [13].
Huan Wang 0001, Shuicheng Yan, Thomas S. Huang, Xiaoou Tang
ACM Multimedia4
2006 Random Sampling for Subspace Face Recognition
Xiaogang Wang 0001, Xiaoou Tang
Int. J. Comput. Vis.2
2006 Full-Frame Video Stabilization with Motion Inpainting
abstract
Video stabilization is an important video enhancement technology which aims at removing annoying shaky motion from videos. We propose a practical and robust approach of video stabilization that produces full-frame stabilized videos with good visual quality. While most previous methods end up with producing smaller size stabilized videos, our completion method can produce full-frame videos by naturally filling in missing image parts by locally aligning image data of neighboring frames. To achieve this, motion inpainting is proposed to enforce spatial and temporal consistency of the completion in both static and dynamic image areas. In addition, image quality in the stabilized video is enhanced with a new practical deblurring algorithm. Instead of estimating point spread functions, our method transfers and interpolates sharper image pixels of neighboring frames to increase the sharpness of the frame. The proposed video completion and deblurring methods enabled us to develop a complete video stabilizer which can naturally keep the original image quality in the stabilized videos. The effectiveness of our method is confirmed by extensive experiments over a wide variety of videos.
Yasuyuki Matsushita, Eyal Ofek, Weina Ge, Xiaoou Tang, Harry Shum
IEEE Trans. Pattern Anal. Mach. Intell.4
2006 Asymmetric Bagging and Random Subspace for Support Vector Machines-Based Relevance Feedback in Image Retrieval
abstract
Relevance feedback schemes based on support vector machines (SVM) have been widely used in content-based image retrieval (CBIR). However, the performance of SVM-based relevance feedback is often poor when the number of labeled positive feedback samples is small. This is mainly due to three reasons: 1) an SVM classifier is unstable on a small-sized training set, 2) SVM's optimal hyperplane may be biased when the positive feedback samples are much less than the negative feedback samples, and 3) overfitting happens because the number of feature dimensions is much higher than the size of the training set. In this paper, we develop a mechanism to overcome these problems. To address the first two problems, we propose an asymmetric bagging-based SVM (AB-SVM). For the third problem, we combine the random subspace method and SVM for relevance feedback, which is named random subspace SVM (RS-SVM). Finally, by integrating AB-SVM and RS-SVM, an asymmetric bagging and random subspace SVM (ABRS-SVM) is built to solve these three problems and further improve the relevance feedback performance.
Dacheng Tao, Xiaoou Tang, Xuelong Li 0001, Xindong Wu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2006 Real-Time Bayesian 3-D Pose Tracking
abstract
In this paper, we propose a novel approach for real-time 3-D tracking of object pose from a single camera. We formulate the 3-D pose tracking task in a Bayesian framework which fuses feature correspondence information from both previous frame and some selected key-frames into the posterior distribution of pose. We also developed an inter-frame motion inference algorithm which can get reliable inter-frame feature correspondences and relative pose. Finally, the maximum a posteriori estimation of pose is obtained via stochastic sampling to achieve stable and drift-free tracking. Experiments show significant improvement of our algorithm over existing algorithms especially in the cases of tracking agile motion, severe occlusion, drastic illumination change, and large object scale change
Qiang Wang 0023, Xiaoou Tang, Harry Shum
IEEE Trans. Circuits Syst. Video Technol.3
2006 Direct Kernel Biased Discriminant Analysis: A New Content-Based Image Retrieval Relevance Feedback Algorithm
abstract
In recent years, a variety of relevance feedback (RF) schemes have been developed to improve the performance of content-based image retrieval (CBIR). Given user feedback information, the key to a RF scheme is how to select a subset of image features to construct a suitable dissimilarity measure. Among various RF schemes, biased discriminant analysis (BDA) based RF is one of the most promising. It is based on the observation that all positive samples are alike, while in general each negative sample is negative in its own way. However, to use BDA, the small sample size (SSS) problem is a big challenge, as users tend to give a small number of feedback samples. To explore solutions to this issue, this paper proposes a direct kernel BDA (DKBDA), which is less sensitive to SSS. An incremental DKBDA (IDKBDA) is also developed to speed up the analysis. Experimental results are reported on a real-world image collection to demonstrate that the proposed methods outperform the traditional kernel BDA (KBDA) and the support vector machine (SVM) based RF algorithms.
Dacheng Tao, Xiaoou Tang, Xuelong Li 0001, Yong Rui
IEEE Trans. Multim.2
2006 Face recognition using kernel scatter-difference-based discriminant analysis
abstract
There are two fundamental problems with the Fisher linear discriminant analysis for face recognition. One is the singularity problem of the within-class scatter matrix due to small training sample size. The other is that it cannot efficiently describe complex nonlinear variations of face images because of its linear property. In this letter, a kernel scatter-difference-based discriminant analysis is proposed to overcome these two problems. We first use the nonlinear kernel trick to map the input data into an implicit feature space F. Then a scatter-difference-based discriminant rule is defined to analyze the data in F. The proposed method can not only produce nonlinear discriminant features but also avoid the singularity problem of the within-class scatter matrix. Extensive experiments show encouraging recognition performance of the new algorithm.
Qingshan Liu 0001, Xiaoou Tang, Hanqing Lu, Songde Ma
IEEE Trans. Neural Networks2
2005 Nonparametric Subspace Analysis for Face Recognition
abstract
Linear discriminant analysis (LDA) is a popular face recognition technique. However, an inherent problem with this technique stems from the parametric nature of the scatter matrix, in which the sample distribution in each class is assumed to be normal distribution. So it tends to suffer in the case of non-normal distribution. In this paper a nonparametric scatter matrix is defined to replace the traditional parametric scatter matrix in order to overcome this problem. Two kinds of nonparametric subspace analysis (NSA): PNSA and NNSA are proposed for face recognition. The former is based on the principal space of intra-personal scatter matrix, while the latter is based on the null space. In addition, based on the complementary nature of PNSA and NNSA, we further develop a dual NSA-based classifier framework using Gabor images to further improve the recognition performance. Experiments achieve near perfect recognition accuracy (99.7%) on the XM2VTS database.
Zhifeng Li 0001, Wei Liu 0026, Dahua Lin, Xiaoou Tang
CVPR (2)4
2005 Detecting Doctored Images Using Camera Response Normality and Consistency
abstract
The advance in image/video editing techniques has facilitated people in synthesizing realistic images/videos that may hard to be distinguished from real ones by visual examination. This poses a problem: how to differentiate real images/videos from doctored ones? This is a serious problem because some legal issues may occur if there is no reliable way for doctored image/video detection when human inspection fails. Digital watermarking cannot solve this problem completely. We propose an approach that computes the response functions of the camera by selecting appropriate patches in different ways. An image may be doctored if the response functions are abnormal or inconsistent to each other. The normality of the response functions is classified by a trained support vector machine (SVM). Experiments show that our method is effective for high-contrast images with many textureless edges.
Zhouchen Lin, Xiaoou Tang, Harry Shum
CVPR (1)3
2005 Hallucinating Faces: TensorPatch Super-Resolution and Coupled Residue Compensation
abstract
In this paper, we propose a new face hallucination framework based on image patches, which integrates two novel statistical super-resolution models. Considering that image patches reflect the combined effect of personal characteristics and patch-location, we first formulate a TensorPatch model based on multilinear analysis to explicitly model the interaction between multiple constituent factors. Motivated by locally linear embedding, we develop an enhanced multilinear patch hallucination algorithm, which efficiently exploits the local distribution structure in the sample space. To better preserve face subtle details, we derive the coupled PCA algorithm to learn the relation between high-resolution residue and low-resolution residue, which is utilized for compensate the error residue in hallucinated images. Experiments demonstrate that our framework on one hand well maintains the global facial structures, on the other hand recovers the detailed facial traits in high quality.
Wei Liu 0026, Dahua Lin, Xiaoou Tang
CVPR (2)3
2005 A Nonlinear Approach for Face Sketch Synthesis and Recognition
abstract
Most face recognition systems focus on photo-based face recognition. In this paper, we present a face recognition system based on face sketches. The proposed system contains two elements: pseudo-sketch synthesis and sketch recognition. The pseudo-sketch generation method is based on local linear preserving of geometry between photo and sketch images, which is inspired by the idea of locally linear embedding. The nonlinear discriminate analysis is used to recognize the probe sketch from the synthesized pseudo-sketches. Experimental results on over 600 photo-sketch pairs show that the performance of the proposed method is encouraging.
Qingshan Liu 0001, Xiaoou Tang, Hongliang Jin, Hanqing Lu, Songde Ma
CVPR (1)2
2005 Full-Frame Video Stabilization
abstract
Video stabilization is an important video enhancement technology which aims at removing annoying shaky motion from videos. We propose a practical and robust approach of video stabilization that produces full-frame stabilized videos with good visual quality. While most previous methods end up with producing low resolution stabilized videos, our completion method can produce full-frame videos by naturally filling in missing image parts by locally aligning image data of neighboring frames. To achieve this, motion inpainting is proposed to enforce spatial and temporal consistency of the completion in both static and dynamic image areas. In addition, image quality in the stabilized video is enhanced with a new practical deblurring algorithm. Instead of estimating point spread functions, our method transfers and interpolates sharper image pixels of neighbouring frames to increase the sharpness of the frame. The proposed video completion and deblurring methods enabled us to develop a complete video stabilizer which can naturally keep the original image quality in the stabilized videos. The effectiveness of our method is confirmed by extensive experiments over a wide variety of videos.
Yasuyuki Matsushita, Eyal Ofek, Xiaoou Tang, Harry Shum
CVPR (1)3
2005 Subspace Analysis Using Random Mixture Models
abstract
In a work by Wang and Tang (2004), three popular subspace face recognition methods, PCA, Bayes, and LDA were analyzed under the same framework and an unified subspace analysis was proposed. However, since they are all based on a single Gaussian model, a global linear subspace often fails to deliver good performance on the data set with complex intrapersonal variation. They also have to face the problem caused by high dimensional face feature vector and the difficulty in finding optimal parameters for subspace analysis. In this paper, we develop a random mixture model to improve Bayes and LDA subspace analysis. By clustering the intrapersonal difference, the complex intrapersonal variation manifold is learned by a set of local linear intrapersonal subspaces. To boost the system performance, we construct multiple low dimensional subspaces by randomly sampling on the high dimensional feature vector and randomly selecting the parameters for subspace analysis. The effectiveness of our method is demonstrated by experiments on the AR face database containing 2340 face images.
Xiaogang Wang 0001, Xiaoou Tang
CVPR (1)2
2005 Discriminant Analysis with Tensor Representation
abstract
In this paper, we present a novel approach to solving the supervised dimensionality reduction problem by encoding an image object as a general tensor of 2nd or higher order. First, we propose a discriminant tensor criterion (DTC), whereby multiple interrelated lower-dimensional discriminative subspaces are derived for feature selection. Then, a novel approach called k-mode cluster-based discriminant analysis is presented to iteratively learn these subspaces by unfolding the tensor along different tensor dimensions. We call this algorithm discriminant analysis with tensor representation (DATER), which has the following characteristics: 1) multiple interrelated subspaces can collaborate to discriminate different classes; 2) for classification problems involving higher-order tensors, the DATER algorithm can avoid the curse of dimensionality dilemma and overcome the small sample size problem; and 3) the computational cost in the learning stage is reduced to a large extent owing to the reduced data dimensions in generalized eigenvalue decomposition. We provide extensive experiments by encoding face images as 2nd or 3rd order tensors to demonstrate that the proposed DATER algorithm based on higher order tensors has the potential to outperform the traditional subspace learning algorithms, especially in the small sample size cases.
Shuicheng Yan, Dong Xu 0001, Qiang Yang 0001, Lei Zhang 0001, Xiaoou Tang, HongJiang Zhang
CVPR (1)5
2005 Fisher+Kernel Criterion for Discriminant Analysis
abstract
We simultaneously approach two tasks of nonlinear discriminant analysis and kernel selection problem by proposing a unified criterion, Fisher+Kernel criterion. In addition, an efficient procedure is derived to optimize this new criterion in an iterative manner. More specifically, original input vector is first transformed into a higher dimensional feature matrix through a battery of nonlinear mappings involved in different kernels. Then, based on the feature matrices, FKC is presented within two coupled projection spaces: one projection space is used to search for the optimal combinations of kernels; while the other encodes the optimal nonlinear discriminating projection directions. Our proposed method is a unified framework for both kernel selection and nonlinear discriminant analysis. Besides, the algorithm potentially alleviates overfitting problem existing in traditional KDA and has no singularity problems in most cases. The effectiveness of our proposed algorithm is validated by extensive face recognition experiments on several datasets.
Shuicheng Yan, Dong Xu 0001, Xiaoou Tang
CVPR (2)4
2005 A Bayesian Mixture Model for Multi-View Face Alignment
abstract
For multi-view face alignment, we have to deal with two major problems: 1) the problem of multi-modality caused by diverse shape variation when the view changes dramatically; 2) the varying number of feature points caused by self-occlusion. Previous works have used nonlinear models or view based methods for multi-view face alignment. However, they either assume all feature points are visible or apply a set of discrete models separately without a uniform criterion. In this paper, we propose a unified framework to solve the problem of multi-view face alignment, in which, both the multi-modality and variable feature points are modeled by a Bayesian mixture model. We first develop a mixture model to describe the shape distribution and the feature point visibility, and then use an efficient EM algorithm to estimate the model parameters and the regularized shape. We use a set of experiments on several datasets to demonstrate the improvement of our method over traditional methods.
Yi Zhou 0020, Wayne Zhang 0001, Xiaoou Tang, Harry Shum
CVPR (2)3
2005 3D Object Reconstruction from a Single 2D Line Drawing without Hidden Lines
abstract
The human vision system can interpret a single 2D line drawing as a 3D object without much difficulty even if the hidden lines of the object are invisible. Several reconstruction approaches have tried to emulate this ability, but they cannot recover the complete object if the hidden lines of the object are not shown. This paper proposes a novel approach for reconstructing complete 3D objects from line drawings without hidden lines. First, we develop some constraints and properties for the inference of the topology of the invisible edges and vertices of an object. Then we present a reconstruction method based on perceptual symmetry and planarity of the object. We give a number of examples to demonstrate the ability of our approach.
Liangliang Cao, Jianzhuang Liu, Xiaoou Tang
ICCV3
2005 A Symmetric Patch-Based Correspondence Model for Occlusion Handling
abstract
Occlusion is one of the challenging problems in stereo. In this paper, we solve the problem in a segment-based style. Both images are segmented, and we propose a novel patch-based stereo algorithm that cuts the segments of one image using the segments of the other, and handles occlusion areas in a proper way. A symmetric graph-cuts optimization framework is used to find correspondence and occlusions simultaneously. The experimental results show superior performance of the proposed algorithm, especially on occlusions, untextured areas and discontinuities
Qiong Yang, Xueyin Lin, Xiaoou Tang
ICCV4
2005 Coupled Space Learning for Image Style Transformation
abstract
In this paper, we present a new learning framework for image style transforms. Considering that the images in different style representations constitute different vector spaces, we propose a novel framework called coupled space learning to learn the relations between different spaces and use them to infer the images from one style to another style. Observing that for each style, only the components correlated to the space of the target style are useful for inference, we first develop the correlative component analysis to pursue the embedded hidden subspaces that best preserve the inter-space correlation information. Then we develop the coupled bidirectional transform algorithm to estimate the transforms between the two embedded spaces, where the coupling between the forward transform and the backward transform is explicitly taken into account. To enhance the capability of modelling complex data, we further develop the coupled Gaussian mixture model to generalize our framework to a mixture-model architecture. The effectiveness of the framework is demonstrated in the applications including face super-resolution and bidirectional portrait style transforms.
Dahua Lin, Xiaoou Tang
ICCV2
2005 Bi-Directional Tracking Using Trajectory Segment Analysis
abstract
In this paper, we present a novel approach to keyframe-based tracking, called bi-directional tracking. Given two object templates in the beginning and ending keyframes, the bi-directional tracker outputs the MAP (maximum a posterior) solution of the whole state sequence of the target object in the Bayesian framework. First, a number of 3D trajectory segments of the object are extracted from the input video, using a novel trajectory segment analysis. Second, these disconnected trajectory segments due to occlusion are linked by a number of inferred occlusion segments. Last, the MAP solution is obtained by trajectory optimization in a coarse-to-fine manner. Experimental results show the robustness of our approach with respect to sudden motion, ambiguity, and short and long periods of occlusion.
Jian Sun 0001, Xiaoou Tang, Harry Shum
ICCV3
2005 Patch Based Blind Image Super Resolution
abstract
In this paper, a novel method for learning based image super resolution (SR) is presented. The basic idea is to bridge the gap between a set of low resolution (LR) images and the corresponding high resolution (HR) image using both the SR reconstruction constraint and a patch based image synthesis constraint in a general probabilistic framework. We show that in this framework, the estimation of the LR image formation parameters is straightforward. The whole framework is implemented via an annealed Gibbs sampling method. Experiments on SR on both single image and image sequence input show that the proposed method provides an automatic and stable way to compute super-resolution and the achieved result is encouraging for both synthetic and real LR images.
Qiang Wang 0023, Xiaoou Tang, Harry Shum
ICCV2
2005 Automatic 3D Face Modeling from Video
abstract
In this paper, we develop an efficient technique for fully automatic recovery of accurate 3D face shape from videos captured by a low cost camera. The method is designed to work with a short video containing a face rotating from frontal view to profile view. The whole approach consists of three components. First, automatic initialization is performed in the first frame with approximately frontal face. Then, to handle the case of low quality image captured by low cost camera, the 2D feature matching, head poses and underlying 3D face shape are estimated and refined iteratively in an efficient way based on image sequence segmentation. Finally, to take advantage of the sparse structure of the proposed algorithm, sparse bundle adjustment technique is further employed to speed up the computation. We demonstrate the accuracy and robustness of the algorithm using a set of experiments
Le Xin, Qiang Wang 0023, Jianhua Tao 0001, Xiaoou Tang, Tieniu Tan, Harry Shum
ICCV4
2005 Extracting micro-structural gabor features for face recognition
abstract
Robustness and discriminability are two key issues in face recognition. In this paper, we propose a new algorithm which extracts micro-structural Gabor feature to achieve good robustness and discriminability simultaneously. We first design a family of directional block partitions to compute the block-level directional projections of the classical Gabor feature. Then we use two statistical kernels, i.e, the mean kernel and the variance kernel, to extract the micro-structural statistics. Analysis of both robustness and discriminability is conducted to show that the new feature is not only more robust to misalignment, but also more discriminative than the classical down-sampling Gabor feature, which is further demonstrated by three groups of experiments on the BANCA dataset.
Dian Gong, Qiong Yang, Xiaoou Tang, Jianhua Lu
ICIP (2)3
2005 Layered local prediction network with dynamic learning for face super-resolution
abstract
In this paper, we propose a novel framework for face super-resolution based on a layered predictor network. In the first layer, multiple predictors are trained online with a dynamic-constructed training set, which is adaptively selected in order to make the trained model tailored to the testing face. When the dynamic training set is obtained, the optimum predictor can be learned based on the resampling-maximum likelihood-model. To further enhance the robustness of prediction and the smoothness of the hallucinated image, additional layers are designed to fuse multiple predictors with the fusion rule learned from the training set. Experiments fully demonstrate the effectiveness of the framework.
Dahua Lin, Wei Liu 0026, Xiaoou Tang
ICIP (1)3
2005 Tensor-based factor decomposition for relighting
abstract
Lighting condition is an important factor in face analysis and synthesis, which has received extensive study in both computer vision and computer graphics. Motivated by the work on multilinear model, we propose a learning-based algorithm for relighting based on tensor framework, which explicitly accounts for the interaction of the identity factor and the lighting factor. The major contribution of our work is that we develop a novel algorithm based on a two-stage decomposition scheme to simultaneously and robustly solve for the identity parameter and the lighting parameter which are both unknown. Equipped with the decomposition algorithm, the capability of the tensor model is significantly extended. Experiment results illustrate the effectiveness of our algorithm.
Dahua Lin, Ying-Qing Xu, Xiaoou Tang, Shuicheng Yan
ICIP (2)3
2005 Feedback-based dynamic generalized LDA for face recognition
abstract
Linear discriminant analysis (LDA) is widely-used in face recognition systems. However, with the traditional formulation, the available information in the training samples is not sufficiently utilized. In this paper, we present a new formulation, called generalized LDA, where the scatter matrices are de defined in a more flexible manner by identifying the fundamental principles of the scatter matrices construction. We further propose a novel framework called feedback-based dynamic generalized LDA. It integrates the generalized LDA and the dynamic feedback strategy for subspace analysis, in which the subspace is iteratively optimized by utilizing the feedback from the previous step. The comparative experiments demonstrate that the new framework achieves encouraging improvement on performances of both the face identification and the face verification.
Dahua Lin, Shuicheng Yan, Xiaoou Tang
ICIP (2)3
2005 Comparative study: face recognition on unspecific persons using linear subspace methods
abstract
Recently many automatic face recognition (AFR) systems were developed for applications with unspecific persons, which is different from conventional pattern recognition problems where all classes are known in the training stage. In this paper, we present a systematic and comprehensive study on linear subspace methods for face recognition on unspecific persons. Over 6700 experiments using different algorithms with different training parameters and testing conditions are conducted on a large scale database (4550 samples) to investigate the compound effect of various influential factors. The observations based on these experiments are expected to provide widely applicable guidelines for designing practical AFR systems.
Dahua Lin, Shuicheng Yan, Xiaoou Tang
ICIP (3)3
2005 Face hallucination through dual associative learning
abstract
In this paper, we propose a novel patch-based face hallucination framework, which employs a dual model to hallucinate different components associated with one facial image. Our model is based on a statistical learning approach: associative learning. It suffices to learn the dependencies between low-resolution image patches and their high-resolution ones with a new concept hidden parameter space as a bridge to connect those patches with different resolutions. To compensate higher frequency information of images, we present a dual associative learning algorithm for orderly inferring main components and high frequency components of faces. The patches can be finally integrated to form a whole high-resolution image. Experiments demonstrate that our approach does render high quality superresolution faces.
Wei Liu 0026, Dahua Lin, Xiaoou Tang
ICIP (1)3
2005 Fingerprint matching using minutiae and interpolation-based square tessellation fingercode
abstract
To improve the overall accuracy, a hybrid fingerprint-matching scheme using both minutiae and square-tessellation-fingercode has been proposed in the literature. However, for identification applications, the matching process is time-consuming since the fingercode of the query fingerprint is repeatedly extracted when it is compared with different template fingerprints in a large database. In addition, the matching accuracy is influenced by nonlinear distortions in fingerprint images. In this paper, we propose a new approach to solve the problem. We extract the fignercode of the query fingerprint only once. When compared with the template fingerprints, the corresponding fingercodes are generated by interpolation and resampling on the extracted fingercode according to the optimally estimated mapping functions with respect to the minutiae matching results. Experimental results on NIST-4 and FVC2002 demonstrate that our algorithm outperforms the original approach in terms of both accuracy and running time.
Lifeng Sha, Feng Zhao 0004, Xiaoou Tang
ICIP (2)3
2005 Largest-eigenvalue-theory for incremental principal component analysis
abstract
In this paper, we present a novel algorithm for incremental principal component analysis. Based on the largest-eigenvalue-theory, i.e. the eigenvector associated with the largest eigenvalue of a symmetry matrix can be iteratively estimated with any initial value, we propose an iterative algorithm, referred as LET-IPCA, to incrementally update the eigenvectors corresponding to the leading eigenvalues. LET-IPCA is covariance matrix free and seamlessly connects the estimations of the leading eigenvectors by cooperatively preserving the most dominating information, as opposed to the state-of-the-art algorithm CCIPCA, in which the estimation of each eigenvector is independent. The experiments on both the MNIST digits database and the CMU PIE face database show that our proposed algorithm is much superior to CCIPCA in both convergency speed and accuracy.
Shuicheng Yan, Xiaoou Tang
ICIP (1)2
2005 Efficient local reflectional symmetries detection
abstract
In this paper, we present a novel framework for efficient multiple reflectional symmetric regions detection in real images. First, we present a fast operator to measure the symmetry. Then based on the extracted edge image, the dilation and erosion operations are applied for potential regions. Finally, the symmetry axes are derived based on weighted principle component analysis (PCA). The experiments based on the proposed algorithm show encouraging results on real images even with small size local symmetric regions and complex backgrounds.
Tianqiang Yuan, Xiaoou Tang
ICIP (3)2
2005 A probabilistic model for robust face alignment in videos
abstract
A new approach for localizing facial structure in videos is proposed in this paper by modeling shape alignment dynamically. The approach makes use of the spatial-temporal continuity of videos and incorporates it into a statistical shape model which is called constrained Bayesian tangent shape model (C-BTSM). Our model includes a prior 2D shape model learnt from labeled examples, an observation model obtained from observation in the current input image, and a constraint model derived from the prediction by the previous frames. By modeling the prior, observation and constraint in a probabilistic framework, the task of aligning shape in each frame of a video is performed as a procedure of MAP parameter estimation, in which the pose and shape parameters are recovered simultaneously. Experiments on low quality videos from web cameras are provided to demonstrate the robustness and accuracy of our algorithm.
Wayne Zhang 0001, Yi Zhou 0020, Xiaoou Tang, Junhui Deng
ICIP (3)3
2005 Binary plankton image classification using random subspace
abstract
In this paper, we implement a random subspace based algorithm to classify the plankton images detected in real time by the shadowed image particle profiling and evaluation recorder. The difficulty of such classification is compounded because the data sets are not only much noisier but the plankton are deformable, projection-variant, and often in partial occlusion. In addition, the images in our experiments are binary thus are lack of texture information. Using random sampling, we construct a set of stable classifiers to take full advantage of nearly all the discriminative information in the feature space of plankton images. The combination of multiple stable classifiers is better than a single classifier. We achieve over 93% classification accuracy on a collection of more than 3000 images, making it comparable with what a trained biologist can achieve by using conventional manual techniques.
Feng Zhao 0004, Xiaoou Tang, Feng Lin 0002, Scott Samson, Andrew Remsen
ICIP (1)2
2005 Learning Local Descriptors for Face Detection
abstract
In this paper, we propose a realtime face detection approach based on local structure and texture of the objects in gray-level images. Our strategy is to map the local spatial structures and image textures of face class into binary patterns, and use these binary patterns as local descriptors. Boosting based face detector is constructed using these local descriptors, and cascade scheme is employed to further improve the efficiency of the face detector. Compared to the existing face detection approaches, our proposed method has two advantages: (1) it is robust to illumination changes to some extend, for the features use the information of local relationship instead of the original gray values; (2) the computational cost is very low, both in training procedure and evaluation step. The experimental results show that the proposed method can meet the demand of realtime applications with a satisfied detection performance.
Hongliang Jin, Qingshan Liu 0001, Xiaoou Tang, Hanqing Lu
ICME3
2005 Neighbor combination and transformation for hallucinating faces
abstract
In this paper, we propose a novel face hallucination framework based on image patches, which exploits local geometry structures of overlapping patches to hallucinate different components associated with one facial image. To achieve local fidelity while preserving smoothness in the target high-resolution image, we develop a neighbor combination super-resolution model for high-resolution patch synthesis. For further enhancing the detailed information, we propose another model, which effectively learns neighbor transformations between low- and high-resolution image patch residuals to compensate modeling errors caused by the first model. Experiments demonstrate that our approach can hallucinate high quality super-resolution faces.
Wei Liu 0026, Dahua Lin, Xiaoou Tang
ICME3
2005 Learning an image-word embedding for image auto-annotation on the nonlinear latent space
abstract
Latent Semantic Analysis (LSA) has shown encouraging performance for the problem of unsupervised image automatic annotation. LSA conducts annotation by keywords propagation on a linear Latent Space, which accounts for the underlying semantic structure of word and image features. In this paper, we formulate a more general nonlinear model, called Nonlinear Latent Space model, to reveal the latent variables of word and visual features more precisely. Instead of the basic propagation strategy, we present a novel inference strategy for image annotation via Image-Word Embedding (IWE). IWE simultaneously embeds images and words and captures the dependencies between them from a probabilistic viewpoint. Experiments show that IWE-based annotation on the nonlinear latent space outperforms previous unsupervised annotation methods.
Wei Liu 0026, Xiaoou Tang
ACM Multimedia2
2005 Evolutionary Search for Faces from Line Drawings
abstract
Single 2D line drawing is a straightforward method to illustrate 3D objects. The faces of an object depicted by a line drawing give very useful information for the reconstruction of its 3D geometry. Two recently proposed methods for face identification from line drawings are based on two steps: finding a set of circuits that may be faces and searching for real faces from the set according to some criteria. The two steps, however, involve two combinatorial problems. The number of the circuits generated in the first step grows exponentially with the number of edges of a line drawing. These circuits are then used as the input to the second combinatorial search step. When dealing with objects having more faces, the combinatorial explosion prevents these methods from finding solutions within feasible time. This paper proposes a new method to tackle the face identification problem by a variable-length genetic algorithm with a novel heuristic and geometric constraints incorporated for local search. The hybrid GA solves the two combinatorial problems simultaneously. Experimental results show that our algorithm can find the faces of a line drawing having more than 30 faces much more efficiently. In addition, simulated annealing for solving the face identification problem is also implemented for comparison.
Jianzhuang Liu, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Video-based handwritten Chinese character recognition
Xiaoou Tang, Feng Lin 0002, Jianzhuang Liu
IEEE Trans. Circuits Syst. Video Technol.1
2005 Insignificant shadow detection for video segmentation
abstract
To prevent moving cast shadows from being misunderstood as part of moving objects in change detection based video segmentation, this paper proposes a novel approach to the cast shadow detection based on the edge and region information in multiple frames. First, an initial change detection mask containing moving objects and cast shadows is obtained. Then a Canny edge map is generated. After that, the shadow region is detected and removed through multiframe integration, edge matching, and region growing. Finally, a post processing procedure is used to eliminate noise and tune the boundaries of the objects. Our approach can be used for video segmentation in indoor environment. The experimental results demonstrate its good performance.
Dong Xu 0001, Jianzhuang Liu, Xuelong Li 0001, Zhengkai Liu, Xiaoou Tang
IEEE Trans. Circuits Syst. Video Technol.5
2005 Hallucinating face by eigentransformation
abstract
In video surveillance, the faces of interest are often of small size. Image resolution is an important factor affecting face recognition by human and computer. In this paper, we propose a new face hallucination method using eigentransformation. Different from most of the proposed methods based on probabilistic models, this method views hallucination as a transformation between different image styles. We use Principal Component Analysis (PCA) to fit the input face image as a linear combination of the low-resolution face images in the training set. The high-resolution image is rendered by replacing the low-resolution training images with high-resolution ones, while retaining the same combination coefficients. Experiments show that the hallucinated face images are not only very helpful for recognition by humans, but also make the automatic recognition procedure easier, since they emphasize the face difference by adding more high-frequency details.
Xiaogang Wang 0001, Xiaoou Tang
IEEE Trans. Syst. Man Cybern. Part C2
2004 Bayesian Face Recognition Using Support Vector Machine and Face Clustering
Zhifeng Li 0001, Xiaoou Tang
CVPR (2)2
2004 Efficient Search of Faces from Complex Line Drawings
Jianzhuang Liu, Xiaoou Tang
CVPR (2)2
2004 Frame Synchronization and Multi-Level Subspace Analysis for Video Based Face Recognition
Xiaoou Tang, Zhifeng Li 0001
CVPR (2)1
2004 Orthogonal Complement Component Analysis for Positive Samples in SVM Based Relevance Feedback Image Retrieval
Dacheng Tao, Xiaoou Tang
CVPR (2)2
2004 Random Sampling Based SVM for Relevance Feedback Image Retrieval
Dacheng Tao, Xiaoou Tang
CVPR (2)2
2004 Random Sampling LDA for Face Recognition
Xiaogang Wang 0001, Xiaoou Tang
CVPR (2)2
2004 Dual-Space Linear Discriminant Analysis for Face Recognition
Xiaogang Wang 0001, Xiaoou Tang
CVPR (2)2
2004 A direct method to solve the biased discriminant analysis in kernel feature space for content based image retrieval
abstract
In recent years, relevance feedback has been widely used to improve the performance of content-based image retrieval. The way in which to select a subset of features from a large-scale feature pool and to construct a suitable dissimilarity measure are key steps in a relevance feedback system. Biased discriminant analysis has been proposed to select features during relevance feedback iterations. However, to solve the BDA, we often encounter the matrix singular problem. In this paper, we propose a kernel-based discriminant analysis, which can overcome the matrix singular problem. The new method is shown to outperform the traditional kernel BDA and constrained support vector machine based relevance feedback algorithms.
Dacheng Tao, Xiaoou Tang
ICASSP (3)2
2004 Content-based SMIL retrieval
abstract
The synchronized multimedia integration language (SML/spl trade/) fulfills the needs of integration, synchronization, and efficient online delivery of different media types such as text, music, speech, image, and video. In this paper, we represent these multimedia elements in a synchronized manner under a unified feature space. An efficient SMIL retrieval scheme based on textual feature and content feature is proposed. Pilot experiments on our SMIL database show that the proposed method can work well on SMIL retrieval.
Kennis Tam, Lam Ching Yu, Dacheng Tao, Hao Liu 0007, Bo Luo, Xiaoou Tang
ICIG6
2004 Combining exclusive and continuous fingerprint classification
abstract
For efficient fingerprint matching, we propose to combine the advantage of exclusive and continuous fingerprint classification. We develop two combined fingerprint classification methods. The first method classifies a fingerprint into arch, left loop, right loop, tented arch and whorl, then performs a continuous fingerprint classification in each class using the global rotation of the ridge flows or the distances between core points and delta points. The second method first classifies a fingerprint into arch, left loop, right loop and whorl using the direction image around the reference point and then performs a continuous fingerprint classification in each class using the FingerCode. Both methods can reduce fingerprint search time by nearly 95%.
Lifeng Sha, Xiaoou Tang
ICIP2
2004 Texture classification of sars infected region in radiographic image
abstract
In this paper, we conduct the first study on SARS radiographic image processing. In order to distinguish SARS infected regions from normal lung regions using texture features, we propose several improvements to the traditional gray-level co-occurrence texture features (R. M. Haralick et al., 1973). We use a multi-level feature selection approach to extract texture features from a multi-resolution region based co-occurrence matrix directly for texture classification. The selected texture features can preserve most of the discriminant information in the texture image. Satisfactory results are obtained on a large set of chest radiographic images of SARS patients.
Xiaoou Tang, Dacheng Tao, George E. Antoniou
ICIP1
2004 Improving indoor and outdoor face recognition using unified subspace analvsis and gabor features
Xiaogang Wang 0001, Xiaoou Tang
ICIP2
2004 SVM-based relevance feedback using random subspace method
abstract
Relevance feedback (RF) schemes based on support vector machines (SVM) have been widely used in content-based image retrieval (CBIR). However, the performance of SVM based RF is often poor when the number of labeled feedback samples is small. In order to solve this problem, we propose an RF algorithm using the random subspace method. The algorithm can overcome the classifier unstable and over-fitting problems that are common to the SVM based RF. Through extensive experiments on 17, 800 images, the proposed algorithm is shown to outperform existing algorithms significantly.
Dacheng Tao, Xiaoou Tang
ICME2
2004 Kernel full-space biased discriminant analysis
abstract
Recently, relevance feedback has been widely used to improve the performance of content-based image retrieval. How to select a subset of features from a large-scale feature pool and to construct a suitable dissimilarity measure are key steps in a relevance feedback system. Biased discriminant analysis (BDA) has been proposed to select features during relevance feedback iterations. However, to solve the BDA, we often encounter the matrix singular problem. Motivated by the direct method and null-space method successfully used in the Fisher linear discriminant analysis for face recognition, we generalize them into the Hilbert space for BDA. Because the direct method and the null-space method may lose some discriminant information, we propose a new full-space method to contain all discriminant information. We also generalize the full-space method into the Hilbert space. All the new methods are demonstrated to outperform the traditional kernel BDA based relevance feedback algorithms based on a statistical experiment in the Corel database with 17, 800 images.
Dacheng Tao, Xiaoou Tang
ICME2
2004 Indoor shadow detection for video segmentation
abstract
To prevent moving cast shadows from being misunderstood as part of moving objects in change detection based video segmentation, this paper proposes a novel approach to cast shadow detection based on the edge and region information in multiple frames. First, an initial change detection mask containing moving objects and cast shadows is obtained. Then, a Canny edge map is generated. After that, the shadow region is detected and removed through multi-frame integration, edge matching and region growing. Finally, a post-processing procedure is used to eliminate noise and tune the boundaries of the objects. Our approach can be used for removing moving cast shadows in indoor environments for better video segmentation. The experimental results demonstrate the good performance of our algorithm.
Dong Xu 0001, Jianzhuang Liu, Zhengkai Liu, Xiaoou Tang
ICME4
2004 Intuitive and effective interfaces for WWW image search engines
abstract
Web image search engine has become an important tool to organize digital images on the Web. However, most commercial search engines still use a list presentation while little effort has been placed on improving their usability. How to present the image search results in a more intuitive and effective way is still an open question to be carefully studied. In this demo, we present iFind, a scalable Web image search engine, in which we integrated two kinds of search result browsing interfaces. User study results have proved that our interfaces are superior to traditional interfaces.
Zhiwei Li 0006, Xing Xie 0001, Hao Liu 0007, Xiaoou Tang, Mingjing Li, Wei-Ying Ma
ACM Multimedia4
2004 K-BOX: a query-by-singing based music retrieval system
abstract
In this paper, we present an efficient query-by-singing based musical retrieval system. We first combine multiple Support Vector Machines by classifier committee learning to segment the sentences from a song automatically. Many new methods in manipulating Mel-Frequency Cepstral Coefficient (MFCC) matrix are studied and compared for optimal feature selection. Experiments show that the 3rd coefficient is the most relevant to music comparison out of 13 coefficients and the proposed simplified MFCC feature is able to achieve a reasonable trade-off between accuracy and efficiency. To improve system efficiency, we re-organize the database by a new two-stage clustering scheme in both time space and feature space. We combine K-means algorithm and dynamic time wrapping similarity measurement for feature space clustering. We also propose a new method for model-selection of K-means algorithm. Experiments show that the proposed approach can achieve more than 30 percent increase in accuracy while speed up more than 16 times in average query time.
Dacheng Tao, Hao Liu 0007, Xiaoou Tang
ACM Multimedia3
2004 A Unified Framework for Subspace Face Recognition
abstract
PCA, LDA, and Bayesian analysis are the three most representative subspace face recognition approaches. In this paper, we show that they can be unified under the same framework. We first model face difference with three components: intrinsic difference, transformation difference, and noise. A unified framework is then constructed by using this face difference model and a detailed subspace analysis on the three components. We explain the inherent relationship among different subspace methods and their unique contributions to the extraction of discriminating information from the face difference. Based on the framework, a unified subspace analysis method is developed using PCA, Bayes, and LDA as three steps. A 3D parameter space is constructed using the three subspace dimensions as axes. Searching through this parameter space, we achieve better recognition performance than standard subspace methods.
Xiaogang Wang 0001, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Guest Editorial Introduction to the Special Issue on Image- and Video-Based Biometrics
Xiaoou Tang, Songde Ma, Lawrence O'Gorman, Massimo Tistarelli
IEEE Trans. Circuits Syst. Video Technol.1
2004 Guest Editorial: Introduction to the Special Issue on Image- and Video-Based Biometrics - Part II
Xiaoou Tang, Songde Ma, Lawrence O'Gorman, Massimo Tistarelli
IEEE Trans. Circuits Syst. Video Technol.1
2004 Face sketch recognition
abstract
Automatic retrieval of face images from police mug-shot databases is critically important for law enforcement agencies. It can effectively help investigators to locate or narrow down potential suspects. However, in many cases, a photo image of a suspect is not available and the best substitute is often a sketch drawing based on the recollection of an eyewitness. We present a novel photo retrieval system using face sketches. By transforming a photo image into a sketch, we reduce the difference between photo and sketch significantly, thus allowing effective matching between the two. Experiments over a data set containing 188 people clearly demonstrate the efficacy of the algorithm.
Xiaoou Tang, Xiaogang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2003 Face recognition using discrete wavelet graph
abstract
We propose a face recognition algorithm using the discrete wavelet face graph. Similar to the Gabor face graph, the method uses the elastic bunch graph matching process to locate fiducial points, but adopts different jet features using a discrete wavelet transform. Following the FERET identification protocol, 2340 face images are used to compare the recognition performance of the two methods. Experimental results show that the DWT face graph has comparable performance as the Gabor face graph. However, with fast computation and compact data structure, the DWT face graph has a unique advantage in real-time applications.
Xiaoou Tang
ICASSP (6)2
2003 An improved Bayesian face recognition algorithm in PCA subspace
abstract
Through modeling the difference between two face images by three components, intrinsic difference (I), transformation difference (T), and random noise (N), we show that the Bayesian algorithm can successfully separate the main disturbing component, T, from the discriminating component, I, however, at a cost of magnified noise, N. To control the noise, we apply PCA on the original image, then carry out the Bayesian analysis in the reduced PCA space. The new method is shown to be more effective than the standard Bayesian algorithm in experiments using 2000+ face images from the Feret database.
Xiaogang Wang 0001, Xiaoou Tang
ICASSP (3)2
2003 Dynamic Stroke Information Analysis for Video-Based Handwritten Chinese Character Recognition
abstract
Video-based handwritten character recognition (VCR) system is a new type of character recognition system with many unique advantages over online character recognition system. Its main problem is to effectively extract stroke dynamic information from video data for character recognition. We propose a new stroke extraction algorithm through dynamic stroke information analysis for a VCR system. The experimental results on over 3000 video character sequences show that our system can extract the Chinese character stroke dynamic information similar to an online system.
Feng Lin 0002, Xiaoou Tang
ICCV2
2003 Face Sketch Synthesis and Recognition
abstract
We propose a novel face photo retrieval system using sketch drawings. By transforming a photo image into a sketch, we reduce the difference between photo and sketch significantly, thus allow effective matching between the two. To improve the synthesis performance, we separate shape and texture information in a face photo, and conduct transformation on them respectively. Finally a Bayesian classifier is used to recognize the probing sketch from the synthesized pseudo-sketches. Experiments on a data set containing 606 people clearly demonstrate the efficacy of the algorithm.
Xiaoou Tang, Xiaogang Wang 0001
ICCV1
2003 Unified Subspace Analysis for Face Recognition
abstract
We propose a face difference model that decomposes face difference into three components, intrinsic difference, transformation difference, and noise. Using the face difference model and a detailed subspace analysis on the three components we develop a unified framework for subspace analysis. Using this framework we discover the inherent relationship among different subspace methods and their unique contributions to the extraction of discriminating information from the face difference. This eventually leads to the construction of a 3D parameter space that uses three subspace dimensions as axis. Within this parameter space, we develop a unified subspace analysis method that achieves better recognition performance than the standard subspace methods on over 2000 face images from the FERET database.
Xiaogang Wang 0001, Xiaoou Tang
ICCV2
2003 Video caption detection and extraction using temporal information
abstract
Video caption detection and extraction is an important step for information retrieval in video databases. In this paper, we extract text information in video by fully utilizing the temporal information contained in the video. First we create a binary abstract sequence from a video segment. By analyzing the statistical pixel changes in the sequence, we can effectively locate the (dis)appealing frames of captions. Finally we extract the captions to create a summary of the video segment.
Bo Luo, Xiaoou Tang, Jianzhuang Liu, HongJiang Zhang
ICIP (1)2
2003 Improved fingercode for filterbank-based fingerprint matching
abstract
FingerCode has been shown to be an effective representation to capture both the local and global information in a fingerprint. However, the performance of fingercode is influenced by the reference point detection process, and the AAD features cannot fully extract the discriminating information in fingerprints. In this paper, we first propose a new rotation-invariant reference point location method, and then combine the direction features with the AAD features to form an oriented fingercode. Experiments conducted on a large fingerprint database (NIST-4) show that the proposed method produces a much improved matching performance.
Lifeng Sha, Feng Zhao 0004, Xiaoou Tang
ICIP (2)3
2002 Video-based handwritten character recognition
abstract
In this paper, we propose a robust stroke-tracing algorithm for a video-based handwritten Chinese character recognition system. By using several error correction techniques, the algorithm works effectively against various shadow and noise problems for video-based stroke-tracing of complex Chinese characters. The stroke temporal information similar to an on-line OCR system is accurately extracted.
Xiaoou Tang, Feng Lin 0002
ICASSP1
2002 Face photo recognition using sketch
abstract
Automatic retrieval of face images from police mug-shot databases is critically important for law enforcement agencies. It can help investigators to locate or narrow down potential suspects efficiently. However, in many cases, the photo image of a suspect is not available and the best substitute is often a sketch drawing based on the recollection of an eyewitness. We present a novel photo retrieval system using face sketches. By transforming a photo image into a sketch, we reduce the difference between photo and sketch significantly, thus allowing effective matching between the two. Experiments over a data set containing 188 people clearly demonstrate the efficacy of the algorithm.
Xiaoou Tang, Xiaogang Wang 0001
ICIP (1)1
2002 Video text extraction using temporal feature vectors
abstract
A new caption text extraction algorithm that takes full advantage of the temporal information in a video sequence is developed. By detecting the (dis)appearance of caption text in a video stream, we first identify the video segment that contains the same caption text. Then using the gray-level vector traced across the segment as the feature vector for a pixel point, we can clearly separate a caption pixel from a background pixel for the entire segment.
Xiaoou Tang, Bo Luo, Xinbo Gao 0001, Edwige E. Pissaloux, HongJiang Zhang
ICME (1)1
2002 Unsupervised video-shot segmentation and model-free anchorperson detection for news video story parsing
abstract
News story parsing is an important and challenging task in a news video library system. We address two important components in a news video story parsing system: shot boundary detection and anchorperson detection. First, an unsupervised fuzzy c-means algorithm is used to detect video-shot boundaries in order to segment a news video into video shots. Then, a graph-theoretical cluster analysis algorithm is implemented to classify the video shots into anchorperson shots and news footage shots. Because of its unsupervised nature, the algorithms require little human intervention. The efficacy of the proposed method is extensively tested on more than five hours of news programs.
Xinbo Gao 0001, Xiaoou Tang
IEEE Trans. Circuits Syst. Video Technol.2
2002 A spatial-temporal approach for video caption detection and recognition
abstract
We present a video caption detection and recognition system based on a fuzzy-clustering neural network (FCNN) classifier. Using a novel caption-transition detection scheme we locate both spatial and temporal positions of video captions with high precision and efficiency. Then employing several new character segmentation and binarization techniques, we improve the Chinese video-caption recognition accuracy from 13% to 86% on a set of news video captions. As the first attempt on Chinese video-caption recognition, our experiment results are very encouraging.
Xiaoou Tang, Xinbo Gao 0001, Jianzhuang Liu, HongJiang Zhang
IEEE Trans. Neural Networks1
2001 Automatic Story Segmentation for Spoken Document Retrieval
abstract
We have been working on speech retrieval based on Cantonese television news programs. Our video archive contains over 20 hours of news programs provided by a local television station. These programs have been hand-segmented into video clips, where each clip is a self-contained news story. The audio tracks in our archive are indexed by Cantonese speech recognition. This is integrated with a vector-space information retrieval model to achieve speech retrieval. This paper proposes an approach for automatic story segmentation from television news programs, intended to replace hand-segmentation as described above. Automatic story segmentation is critical for rapid expansion of our video archive. Our approach relies on the assumption that nearly all the news stories follow the temporal syntax of (begin-story /spl rarr/ anchor shots /spl rarr/ field shots /spl rarr/ end-story). Therefore, our algorithm aims to detect field-to-anchor shot boundaries that should also coincide with the story boundaries. The proposed approach utilizes the video frame information for story boundary detection, and involves such techniques as fuzzy c-means and graph-theoretical clustering. The approach achieved precision and recall values of over 70%, based on a 20-hour video corpus.
Pui-Yu Hui, Xiaoou Tang, Helen M. Meng, Wai Lam, Xinbo Gao 0001
FUZZ-IEEE2
2001 Speech retrieval with video parsing for television news programs
abstract
We have been working on speech retrieval from Chinese (Cantonese) television news programs. The use of automatic speech recognition for audio indexing produces imperfect transcriptions, and recognition errors affect retrieval performance. A news story typically contains a brief report by the anchor person(s) in the studio, as well as news footage from the field. Investigation shows that our recognizer performs better when indexing audio from the studio, compared to that from the field. In order to automatically extract the "reliable" audio segments for speech retrieval, we attempt to detect studio-to-field transitions by means of video parsing. Our study is based on 146 news stories collected from local television Cantonese news programs. We formulated a known-item retrieval task and adopted the average inverse rank (AIR) as our evaluation metric. Retrieval is performed based on syllable bigram units, augmented with skipped syllable bigrams. Retrieval using the entire audio track of each news story gave AIR=0.759. With the incorporation of video parsing, we performed retrieval based only on the studio recordings, which produced AIR=0.768.
Helen M. Meng, Xiaoou Tang, Pui-Yu Hui, Xinbo Gao 0001, Yuk-Chi Li
ICASSP2
2001 Discrete wavelet face graph matching
abstract
We have designed a new face graph model, using the discrete wavelet transform, for fast elastic bunch graph matching. With fiducial points as graph nodes, and local-area power spectrum vectors estimated from a space-frequency tree as node attributes, the DWT graph achieves similar model matching performance to the Gabor graph, but with orders of magnitude faster computation.
Xiaoou Tang
ICIP (2)2
2000 Handwritten Chinese Character Recognition Through a Video Camera
abstract
We propose a handwritten Chinese character recognition system using a video camera. The system combines the advantages of both online and offline approaches. It allows users to write on any regular paper just like using an off-line system. At the same time, using a video camera attached on the computer, the system can capture the stroke temporal information similar to an online system.
Xiaoou Tang, Cheung Yung Chan, Jianzhuang Liu
ICIP1
2000 Automatic News Video Caption Extraction and Recognition
Xinbo Gao 0001, Xiaoou Tang
IDEAL2
2000 Optical and Sonar Image Classification: Wavelet Packet Transform vs Fourier Transform
Xiaoou Tang, W. Kenneth Stewart
Comput. Vis. Image Underst.1
1998 Texture information in run-length matrices
abstract
We use a multilevel dominant eigenvector estimation algorithm to develop a new run-length texture feature extraction algorithm that preserves much of the texture information in run-length matrices and significantly improves image classification accuracy over traditional run-length techniques. The advantage of this approach is demonstrated experimentally by the classification of two texture data sets. Comparisons with other methods demonstrate that the run-length matrices contain great discriminatory information and that a good method of extracting such information is of paramount importance to successful classification.
Xiaoou Tang
IEEE Trans. Image Process.1
1998 Multiple competitive learning network fusion for object classification
abstract
This paper introduces a multiple competitive learning neural network fusion method for pattern recognition. By defining a confidence level measure for the learning vector quantization network classifier, we develop both a serial and a parallel network fusion algorithm to combine the discriminatory ability of different individually trained networks. We use two distinct feature vectors, gray-scale morphological granulometry and Fourier boundary descriptor, to demonstrate the efficacy of the classifier. The algorithms are applied on the classification of more than 8000 underwater plankton images. The classification accuracy for training data and for testing data are over 92% and 85%, respectively.
Xiaoou Tang
IEEE Trans. Syst. Man Cybern. Part B1