VLDB 2026 Research / reviewers in the wild / expert
Xiaobo Wang 0001
dblp:07/6140-1
· DBLP profile ↗
28ranked-venue papers
13as first author
2since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 10 first-author · 1 since 2021Artificial intelligence and machine learning · 20 · 9 first-author · 1 since 2021Security and privacy · 2Human-computer interaction and ubiquitous computing · 2Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Deep learning architectures and training · 27% Face, body and person analysis · 22% Representation and self-supervised learning · 15% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 100% | |
| Network and information security
1 paper |
Biometric security · 100% |
Topics — the 30 heaviest of 40, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Face, body and person analysis
face recognition |
2.3 | 5 | 2022 | RVFace: Reliable Vector Guided Softmax Loss for Face Recognition · IEEE Trans. Image Process. 2022 Loss Function Search for Face Recognition · ICML 2020 Exclusivity-Consistency Regularized Knowledge Distillation for Face Recognition · ECCV (24) 2020 |
Machine learning › Deep learning architectures and training › loss function design
margin-based softmax loss |
1.4 | 3 | 2022 | RVFace: Reliable Vector Guided Softmax Loss for Face Recognition · IEEE Trans. Image Process. 2022 Loss Function Search for Face Recognition · ICML 2020 Mis-Classified Vector Guided Softmax Loss for Face Recognition · AAAI 2020 |
Machine learning › Deep learning architectures and training
loss function design |
1.3 | 3 | 2022 | RVFace: Reliable Vector Guided Softmax Loss for Face Recognition · IEEE Trans. Image Process. 2022 Mis-Classified Vector Guided Softmax Loss for Face Recognition · AAAI 2020 Ensemble Soft-Margin Softmax Loss for Image Classification · IJCAI 2018 |
Machine learning › Representation and self-supervised learning › representation learning › feature extraction
discriminative feature learning |
1.0 | 2 | 2022 | RVFace: Reliable Vector Guided Softmax Loss for Face Recognition · IEEE Trans. Image Process. 2022 Mis-Classified Vector Guided Softmax Loss for Face Recognition · AAAI 2020 |
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels |
0.6 | 2 | 2022 | Co-Mining: Deep Face Recognition With Noisy Labels · ICCV 2019 RVFace: Reliable Vector Guided Softmax Loss for Face Recognition · IEEE Trans. Image Process. 2022 |
Computer vision › Segmentation and scene understanding › image segmentation
boundary-aware segmentation |
0.4 | 1 | 2020 | A New Dataset and Boundary-Attention Semantic Segmentation for Face Parsing · AAAI 2020 |
Computer vision › Segmentation and scene understanding › part parsing
face parsing |
0.4 | 1 | 2020 | A New Dataset and Boundary-Attention Semantic Segmentation for Face Parsing · AAAI 2020 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.4 | 1 | 2020 | Exclusivity-Consistency Regularized Knowledge Distillation for Face Recognition · ECCV (24) 2020 |
Machine learning › Deep learning architectures and training › loss function design
loss function learning |
0.4 | 1 | 2020 | Loss Function Search for Face Recognition · ICML 2020 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.4 | 1 | 2020 | A New Dataset and Boundary-Attention Semantic Segmentation for Face Parsing · AAAI 2020 |
Machine learning › Deep learning architectures and training › normalization
batch normalization |
0.4 | 1 | 2019 | ScratchDet: Training Single-Shot Object Detectors From Scratch · CVPR 2019 |
Computer vision › Image recognition and object detection
object detection |
0.4 | 1 | 2019 | ScratchDet: Training Single-Shot Object Detectors From Scratch · CVPR 2019 |
Computer vision › Image recognition and object detection › object detection
training from scratch |
0.4 | 1 | 2019 | ScratchDet: Training Single-Shot Object Detectors From Scratch · CVPR 2019 |
Biometric security
face anti-spoofing |
0.4 | 1 | 2019 | A Dataset and Benchmark for Large-Scale Multi-Modal Face Anti-Spoofing · CVPR 2019 |
Biometric security › face anti-spoofing
multimodal face anti-spoofing |
0.4 | 1 | 2019 | A Dataset and Benchmark for Large-Scale Multi-Modal Face Anti-Spoofing · CVPR 2019 |
Computer vision › Image recognition and object detection
image classification |
0.3 | 1 | 2018 | Ensemble Soft-Margin Softmax Loss for Image Classification · IJCAI 2018 |
Machine learning › Representation and self-supervised learning
multi-view learning |
0.3 | 1 | 2018 | Latent Semantic Aware Multi-View Multi-Label Classification · AAAI 2018 |
Machine learning › Representation and self-supervised learning › multi-view learning
multi-view representation learning |
0.3 | 1 | 2018 | Latent Semantic Aware Multi-View Multi-Label Classification · AAAI 2018 |
Data mining › predictive modeling › classification
multi-label classification |
0.3 | 1 | 2018 | Latent Semantic Aware Multi-View Multi-Label Classification · AAAI 2018 |
Data mining › predictive modeling › classification › multi-label classification
multi-view multi-label classification |
0.3 | 1 | 2018 | Latent Semantic Aware Multi-View Multi-Label Classification · AAAI 2018 |
Machine learning › Kernel, tree and ensemble methods
ensemble learning |
0.3 | 1 | 2017 | Exclusivity Regularized Machine: A New Ensemble SVM Classifier · IJCAI 2017 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning
ensemble SVM |
0.3 | 1 | 2017 | Exclusivity Regularized Machine: A New Ensemble SVM Classifier · IJCAI 2017 |
Computer vision › Face, body and person analysis
face detection |
0.3 | 1 | 2017 | S^3FD: Single Shot Scale-Invariant Face Detector · ICCV 2017 |
Computer vision › Face, body and person analysis › face detection
scale-invariant face detection |
0.3 | 1 | 2017 | S^3FD: Single Shot Scale-Invariant Face Detector · ICCV 2017 |
Computer vision › Image recognition and object detection › object detection
small object detection |
0.3 | 1 | 2017 | S^3FD: Single Shot Scale-Invariant Face Detector · ICCV 2017 |
Data mining
clustering |
0.3 | 1 | 2017 | Exclusivity-Consistency Regularized Multi-view Subspace Clustering · CVPR 2017 |
Data mining › clustering
multi-view clustering |
0.3 | 1 | 2017 | Exclusivity-Consistency Regularized Multi-view Subspace Clustering · CVPR 2017 |
Data mining › clustering › high-dimensional clustering › subspace clustering
multi-view subspace clustering |
0.3 | 1 | 2017 | Exclusivity-Consistency Regularized Multi-view Subspace Clustering · CVPR 2017 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
dictionary learning |
0.2 | 1 | 2015 | Adaptively Unified Semi-Supervised Dictionary Learning with Active Points · ICCV 2015 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding › dictionary learning
semi-supervised dictionary learning |
0.2 | 1 | 2015 | Adaptively Unified Semi-Supervised Dictionary Learning with Active Points · ICCV 2015 |
Methods — techniques the papers use, named apart from their topics
convolutional neural network · 0.8softmax loss · 0.6semi-hard feature mining · 0.6noisy label detection · 0.6three-branch network · 0.4reinforcement learning · 0.4feature mining · 0.4feature margin · 0.4exclusivity-consistency regularization · 0.4boundary-attention · 0.4multimodal fusion · 0.4feature re-weighting · 0.4matrix factorization · 0.3kernel alignment · 0.3subspace learning · 0.3optimization · 0.3shape cue · 0.2orientation cue · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Decisive vector guided column annotation
Xiaobo Wang 0001, Yanyan Liang 0001, Zhen Lei 0001 |
Pattern Recognit. | 1 |
| 2022 | RVFace: Reliable Vector Guided Softmax Loss for Face RecognitionabstractFace recognition has witnessed significant progress with the advances of deep convolutional neural networks (CNNs), and the central task of which is how to improve the feature discrimination. To this end, several margin-based (e.g., angular, additive and additive angular margins) softmax loss functions have been proposed to increase the feature margin between different classes. However, despite great achievements have been made, they mainly suffer from four issues: 1) They are based on the assumption of well-cleaned training sets, without considering the consequence of noisy labels inherently existing in most of face recognition datasets; 2) They ignore the importance of informative (e.g., semi-hard) features mining for discriminative learning; 3) They encourage the feature margin only from the perspective of ground truth class, without realizing the discriminability from other non-ground truth classes; and 4) They set the feature margin between different classes to be same and fixed, which may not adapt the situation of unbalanced data in different classes very well. To cope with these issues, this paper develops a novel loss function, which explicitly estimates the noisy labels to drop them and adaptively emphasizes the semi-hard feature vectors from the remaining reliable ones to guide the discriminative feature learning. Thus we can address all the above issues and achieve more discriminative features for face recognition. To the best of our knowledge, this is the first attempt to inherit the advantages of feature-based noisy labels detection, feature mining and feature margin into a unified loss function. Extensive experimental results on a variety of face recognition benchmarks have demonstrated the effectiveness of our method over state-of-the-art alternatives. Our source code is available at http://www.cbsr.ia.ac.cn/users/xiaobowang/. Xiaobo Wang 0001, Yanyan Liang 0001, Liang Gu, Zhen Lei 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | A New Dataset and Boundary-Attention Semantic Segmentation for Face ParsingabstractFace parsing has recently attracted increasing interest due to its numerous application potentials, such as facial make up and facial image generation. In this paper, we make contributions on face parsing task from two aspects. First, we develop a high-efficiency framework for pixel-level face parsing annotating and construct a new large-scale Landmark guided face Parsing dataset (LaPa). It consists of more than 22,000 facial images with abundant variations in expression, pose and occlusion, and each image of LaPa is provided with an 11-category pixel-level label map and 106-point landmarks. The dataset is publicly accessible to the community for boosting the advance of face parsing.1 Second, a simple yet effective Boundary-Attention Semantic Segmentation (BASS) method is proposed for face parsing, which contains a three-branch network with elaborately developed loss functions to fully exploit the boundary information. Extensive experiments on our LaPa benchmark and the public Helen dataset show the superiority of our proposed method. Yinglu Liu, Hailin Shi, Yue Si, Xiaobo Wang 0001, Tao Mei 0001 |
AAAI | 5 |
| 2020 | Mis-Classified Vector Guided Softmax Loss for Face RecognitionabstractFace recognition has witnessed significant progress due to the advances of deep convolutional neural networks (CNNs), the central task of which is how to improve the feature discrimination. To this end, several margin-based (e.g., angular, additive and additive angular margins) softmax loss functions have been proposed to increase the feature margin between different classes. However, despite great achievements have been made, they mainly suffer from three issues: 1) Obviously, they ignore the importance of informative features mining for discriminative learning; 2) They encourage the feature margin only from the ground truth class, without realizing the discriminability from other non-ground truth classes; 3) The feature margin between different classes is set to be same and fixed, which may not adapt the situations very well. To cope with these issues, this paper develops a novel loss function, which adaptively emphasizes the mis-classified feature vectors to guide the discriminative feature learning. Thus we can address all the above issues and achieve more discriminative face features. To the best of our knowledge, this is the first attempt to inherit the advantages of feature margin and feature mining into a unified loss function. Experimental results on several benchmarks have demonstrated the effectiveness of our method over state-of-the-art alternatives. Our code is available at http://www.cbsr.ia.ac.cn/users/xiaobowang/. Xiaobo Wang 0001, Tianyu Fu 0001, Hailin Shi, Tao Mei 0001 |
AAAI | 1 |
| 2020 | Exclusivity-Consistency Regularized Knowledge Distillation for Face Recognition
Xiaobo Wang 0001, Tianyu Fu 0001, Shengcai Liao, Zhen Lei 0001, Tao Mei 0001 |
ECCV (24) | 1 |
| 2020 | Loss Function Search for Face RecognitionabstractIn face recognition, designing margin-based (\emph{e.g.}, angular, additive, additive angular margins) softmax loss functions plays an important role to learn discriminative features. However, these hand-crafted heuristic methods may be sub-optimal because they require much effort to explore the large design space. Recently, an AutoML for loss function search method AM-LFS has been derived, which leverages reinforcement learning to search loss functions during the training process. But its search space is complex and unstable that hindering its superiority. In this paper, we first analyze that the key to enhance the feature discrimination is actually \textbf{how to reduce the softmax probability}. We then design a unified formulation for the current margin-based softmax losses. Accordingly, we define a novel search space and develop a reward-guided search method to automatically obtain the best candidate. Experimental results on a variety of face recognition benchmarks have demonstrated the effectiveness of our method over the state-of-the-art alternatives. Xiaobo Wang 0001, Cheng Chi 0003, Tao Mei 0001 |
ICML | 1 |
| 2020 | Listen, Look, and Find the One: Robust Person Search with Multimodality IndexabstractPerson search with one portrait, which attempts to search the targets in arbitrary scenes using one portrait image at a time, is an essential yet unexplored problem in the multimedia field. Existing approaches, which predominantly depend on the visual information of persons, cannot solve problems when there are variations in the person’s appearance caused by complex environments and changes in pose, makeup, and clothing. In contrast to existing methods, in this article, we propose an associative multimodality index for person search with face, body, and voice information. In the offline stage, an associative network is proposed to learn the relationships among face, body, and voice information. It can adaptively estimate the weights of each embedding to construct an appropriate representation. The multimodality index can be built by using these representations, which exploit the face and voice as long-term keys and the body appearance as a short-term connection. In the online stage, through the multimodality association in the index, we can retrieve all targets depending only on the facial features of the query portrait. Furthermore, to evaluate our multimodality search framework and facilitate related research, we construct the Cast Search in Movies with Voice (CSM-V) dataset, a large-scale benchmark that contains 127K annotated voices corresponding to tracklets from 192 movies. According to extensive experiments on the CSM-V dataset, the proposed multimodality person search framework outperforms the state-of-the-art methods. Xiao Wang 0029, Wu Liu 0005, Jun Chen 0001, Xiaobo Wang 0001, Chenggang Yan 0001, Tao Mei 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2019 | A Dataset and Benchmark for Large-Scale Multi-Modal Face Anti-SpoofingabstractFace anti-spoofing is essential to prevent face recognition systems from a security breach. Much of the progresses have been made by the availability of face anti-spoofing benchmark datasets in recent years. However, existing face anti-spoofing benchmarks have limited number of subjects (≤170) and modalities (≤2), which hinder the further development of the academic community. To facilitate face anti-spoofing research, we introduce a large-scale multi-modal dataset, namely CASIA-SURF, which is the largest publicly available dataset for face anti-spoofing in terms of both subjects and visual modalities. Specifically, it consists of 1,000 subjects with 21,000 videos and each sample has 3 modalities (i.e., RGB, Depth and IR). We also provide a measurement set, evaluation protocol and training/validation/testing subsets, developing a new benchmark for face anti-spoofing. Moreover, we present a new multi-modal fusion method as baseline, which performs feature re-weighting to select the more informative channel features while suppressing the less useful ones for each modal. Extensive experiments have been conducted on the proposed dataset to verify its significance and generalization capability. The dataset is available at https://sites.google.com/qq.com/chalearnfacespoofingattackdete/. Xiaobo Wang 0001, Ajian Liu 0001, Jun Wan 0001, Sergio Escalera, Hailin Shi, Stan Z. Li |
CVPR | 2 |
| 2019 | ScratchDet: Training Single-Shot Object Detectors From ScratchabstractCurrent state-of-the-art object objectors are fine-tuned from the off-the-shelf networks pretrained on large-scale classification dataset ImageNet, which incurs some additional problems: 1) The classification and detection have different degrees of sensitivity to translation, resulting in the learning objective bias; 2) The architecture is limited by the classification network, leading to the inconvenience of modification. To cope with these problems, training detectors from scratch is a feasible solution. However, the detectors trained from scratch generally perform worse than the pretrained ones, even suffer from the convergence issue in training. In this paper, we explore to train object detectors from scratch robustly. By analysing the previous work on optimization landscape, we find that one of the overlooked points in current trained-from-scratch detector is the BatchNorm. Resorting to the stable and predictable gradient brought by BatchNorm, detectors can be trained from scratch stably while keeping the favourable performance independent to the network architecture. Taking this advantage, we are able to explore various types of networks for object detection, without suffering from the poor convergence. By extensive experiments and analyses on downsampling factor, we propose the Root-ResNet backbone network, which makes full use of the information from original images. Our ScratchDet achieves the state-of-the-art accuracy on PASCAL VOC 2007, 2012 and MS COCO among all the train-from-scratch detectors and even performs better than several one-stage pretrained methods. Codes will be made publicly available at https://github.com/KimSoybean/ScratchDet. Rui Zhu 0014, Xiaobo Wang 0001, Longyin Wen, Hailin Shi, Liefeng Bo, Tao Mei 0001 |
CVPR | 3 |
| 2019 | Co-Mining: Deep Face Recognition With Noisy LabelsabstractFace recognition has achieved significant progress with the growing scale of collected datasets, which empowers us to train strong convolutional neural networks (CNNs). While a variety of CNN architectures and loss functions have been devised recently, we still have a limited understanding of how to train the CNN models with the label noise inherent in existing face recognition datasets. To address this issue, this paper develops a novel co-mining strategy to effectively train on the datasets with noisy labels. Specifically, we simultaneously use the loss values as the cue to detect noisy labels, exchange the high-confidence clean faces to alleviate the errors accumulated issue caused by the sample-selection bias, and re-weight the predicted clean faces to make them dominate the discriminative model training in a mini-batch fashion. Extensive experiments by training on three popular datasets (\textit{i.e.}, CASIA-WebFace, MS-Celeb-1M and VggFace2) and testing on several benchmarks, including LFW, AgeDB, CFP, CALFW, CPLFW, RFW, and MegaFace, have demonstrated the effectiveness of our new approach over the state-of-the-art alternatives. Xiaobo Wang 0001, Hailin Shi, Jun Wang 0127, Tao Mei 0001 |
ICCV | 1 |
| 2019 | Clustering and Dynamic Sampling Based Unsupervised Domain Adaptation for Person Re-IdentificationabstractPerson Re-Identification (Re-ID) has witnessed great improvements due to the advances of the deep convolutional neural networks (CNN). Despite this, existing methods mainly suffer from the poor generalization ability to unseen scenes because of the different characteristics between different domains. To address this issue, a Clustering and Dynamic Sampling (CDS) method is proposed in this paper, which tries to transfer the useful knowledge of existing labeled source domain to the unlabeled target one. Specifically, to improve the discriminability of CNN model on source domain, we use the commonly shared pedestrian attributes (e.g., gender, hat and clothing color etc.) to enrich the information and resort to the margin-based softmax (e.g., A-Softmax) loss to train the model. For the unlabeled target domain, we iteratively cluster the samples into several centers and dynamically select informative ones from each center to fine-tune the source-domain model. Extensive experiments on DukeMTMC-reID and Market-1501 datasets show that the proposed method greatly improves the state of the arts in unsupervised domain adaptation. Jinlin Wu, Shengcai Liao, Zhen Lei 0001, Xiaobo Wang 0001, Yang Yang 0062, Stan Z. Li |
ICME | 4 |
| 2019 | Faceboxes: A CPU real-time and accurate unconstrained face detector
Xiaobo Wang 0001, Zhen Lei 0001, Stan Z. Li |
Neurocomputing | 2 |
| 2019 | Multi-view subspace clustering with intactness-aware similarity
Xiaobo Wang 0001, Zhen Lei 0001, Xiaojie Guo 0001, Changqing Zhang 0002, Hailin Shi, Stan Z. Li |
Pattern Recognit. | 1 |
| 2018 | Latent Semantic Aware Multi-View Multi-Label ClassificationabstractFor real-world applications, data are often associated with multiple labels and represented with multiple views. Most existing multi-label learning methods do not sufficiently consider the complementary information among multiple views, leading to unsatisfying performance. To address this issue, we propose a novel approach for multi-view multi-label learning based on matrix factorization to exploit complementarity among different views. Specifically, under the assumption that there exists a common representation across different views, the uncovered latent patterns are enforced to be aligned across different views in kernel spaces. In this way, the latent semantic patterns underlying in data could be well uncovered and this enhances the reasonability of the common representation of multiple views. As a result, the consensus multi-view representation is obtained which encodes the complementarity and consistence of different views in latent semantic space. We provide theoretical guarantee for the strict convexity for our method by properly setting parameters. Empirical evidence shows the clear advantages of our method over the state-of-the-art ones. Changqing Zhang 0002, Ziwei Yu, Qinghua Hu, Pengfei Zhu 0001, Xinwang Liu 0002, Xiaobo Wang 0001 |
AAAI | 6 |
| 2018 | Deep Background Subtraction with Guided LearningabstractRecently, convolutional neural networks (CNNs) have been applied in background subtraction (change detection) and gained notable improvements. Two typical methods have been proposed. The first one learns a specific CNN model for each video, but requires manual labeling of training frames on the fly. The other one learns a universal model offline, however, limits its performance in handling various surveillance scenarios. To address these problems, in this paper, a new deep background subtraction method is proposed by introducing a guided learning strategy. The main idea is to learn a specific CNN model for each video to ensure accuracy, but manage to avoid manual labeling. To achieve this, firstly we apply the SubSENSE algorithm [1] to get an initial segmentation, and then an adaptive strategy is designed to select reliable pixels to guide the CNN training. Besides, we also design a simple strategy to automatically select informative frames for guided learning. Experiments on the largest background subtraction benchmark CDnet2014 show that the proposed guided deep learning method outperforms existing state of the arts. Xuezhi Liang, Shengcai Liao, Xiaobo Wang 0001, Wei Liu 0097, Stan Z. Li |
ICME | 3 |
| 2018 | Co-Referenced Subspace ClusteringabstractSubspace clustering refers to the problem of grouping data into their underlying groups. To address this task, spectral clustering based technique is arguably one of the most popular approaches, and its performance largely depends on the constructed similarity. However, most existing works merely employ the primary representation (e.g., sparse or low-rank representation) as the similarity. In this paper, we propose to explore a high-level co-referenced similarity by employing the Hilbert-Schmidt Independence Criterion (HSIC). Moreover, geometry interpretation of the advantage of our co-referenced similarity is provided. Representation-induced kernels such as Mahalanobis metric, can also be easily embedded into the formulation. Extensive experiments on both synthetic and real-world data are conducted to show the superiority of the proposed method over the state-of-the-art alternatives. Xiaobo Wang 0001, Zhen Lei 0001, Hailin Shi, Xiaojie Guo 0001, Xiangyu Zhu 0001, Stan Z. Li |
ICME | 1 |
| 2018 | Ensemble Soft-Margin Softmax Loss for Image ClassificationabstractSoftmax loss is arguably one of the most popular losses to train CNN models for image classification. However, recent works have exposed its limitation on feature discriminability. This paper casts a new viewpoint on the weakness of softmax loss. On the one hand, the CNN features learned using the softmax loss are often inadequately discriminative. We hence introduce a soft-margin softmax function to explicitly encourage the discrmination between different classes. On the other hand, the learned classifier of softmax loss is weak. We propose to assemble multiple these weak classifiers to a strong one, inspired by the recognition that the diversity among weak classifiers is critical to a good ensemble. To achieve the diversity, we adopt the Hilbert-Schmidt Independence Criterion (HSIC). Considering these two aspects in one framework, we design a novel loss, named as Ensemble Soft-Margin Softmax (EM-Softmax). Extensive experiments on benchmark datasets are conducted to show the superiority of our design over the baseline softmax loss and several state-of-the-art alternatives. Xiaobo Wang 0001, Zhen Lei 0001, Si Liu 0001, Xiaojie Guo 0001, Stan Z. Li |
IJCAI | 1 |
| 2018 | Detecting Face with Densely Connected Face Proposal Network
Xiangyu Zhu 0001, Zhen Lei 0001, Xiaobo Wang 0001, Hailin Shi, Stan Z. Li |
Neurocomputing | 4 |
| 2018 | Dependence-Aware Feature Coding for Person Re-IdentificationabstractIn this letter, we focus on how to boost the performance of person re-identification by exploring the discriminative information among person pairs. A novel dependence-aware feature coding framework is proposed for this task. Specifically, we employ the Hilbert–Schmidt independence criterion as the discriminative term, which is to explore the dependence between different kinds of person pairs, i.e., the same person pairs should be dependence maximized, while the different ones should be dependence minimized. Theoretical discussion and analysis on the convexity of the proposed constraint, as well as the convergence of our algorithm, are provided. Experimental results on two benchmark datasets have demonstrated the advantages of our method over the state-of-the-art alternatives. Xiaobo Wang 0001, Zhen Lei 0001, Shengcai Liao, Xiaojie Guo 0001, Yang Yang 0062, Stan Z. Li |
IEEE Signal Process. Lett. | 1 |
| 2017 | Exclusivity-Consistency Regularized Multi-view Subspace ClusteringabstractMulti-view subspace clustering aims to partition a set of multi-source data into their underlying groups. To boost the performance of multi-view clustering, numerous subspace learning algorithms have been developed in recent years, but with rare exploitation of the representation complementarity between different views as well as the indicator consistency among the representations, let alone considering them simultaneously. In this paper, we propose a novel multi-view subspace clustering model that attempts to harness the complementary information between different representations by introducing a novel position-aware exclusivity term. Meanwhile, a consistency term is employed to make these complementary representations to further have a common indicator. We formulate the above concerns into a unified optimization framework. Experimental results on several benchmark datasets are conducted to reveal the effectiveness of our algorithm over other state-of-the-arts. Xiaobo Wang 0001, Xiaojie Guo 0001, Zhen Lei 0001, Changqing Zhang 0002, Stan Z. Li |
CVPR | 1 |
| 2017 | Deep person re-identification with improved embedding and efficient trainingabstractPerson re-identification task has been greatly boosted by deep convolutional neural networks (CNNs) in recent years. The core of which is to enlarge the inter-class distinction as well as reduce the intra-class variance. However, to achieve this, existing deep models prefer to adopt image pairs or triplets to form verification loss, which is inefficient and unstable since the number of training pairs or triplets grows rapidly as the number of training data grows. Moreover, their performance is limited since they ignore the fact that different dimension of embedding may play different importance. In this paper, we propose to employ identification loss with center loss to train a deep model for person re-identification. The training process is efficient since it does not require image pairs or triplets for training while the inter-class distinction and intra-class variance are well handled. To boost the performance, a new feature reweighting (FRW) layer is designed to explicitly emphasize the importance of each embedding dimension, thus leading to an improved embedding. Experiments1on several benchmark datasets have shown the superiority of our method over the state-of-the-art alternatives on both accuracy and speed. Haibo Jin, Xiaobo Wang 0001, Shengcai Liao, Stan Z. Li |
IJCB | 2 |
| 2017 | FaceBoxes: A CPU real-time face detector with high accuracyabstractAlthough tremendous strides have been made in face detection, one of the remaining open challenges is to achieve real-time speed on the CPU as well as maintain high performance, since effective models for face detection tend to be computationally prohibitive.To address this challenge, we propose a novel face detector, named FaceBoxes, with superior performance on both speed and accuracy.Specifically, our method has a lightweight yet powerful network structure that consists of the Rapidly Digested Convolutional Layers (RDCL) and the Multiple Scale Convolutional Layers (MSCL).The RDCL is designed to enable Face-Boxes to achieve real-time speed on the CPU.The MSCL aims at enriching the receptive fields and discretizing anchors over different layers to handle faces of various scales.Besides, we propose a new anchor densification strategy to make different types of anchors have the same density on the image, which significantly improves the recall rate of small faces.As a consequence, the proposed detector runs at 20 FPS on a single CPU core and 125 FPS using a GPU for VGA-resolution images.Moreover, the speed of FaceBoxes is invariant to the number of faces.We comprehensively evaluate this method and present stateof-the-art detection performance on several face detection benchmark datasets, including the AFW, PASCAL face, and FDDB. Xiangyu Zhu 0001, Zhen Lei 0001, Hailin Shi, Xiaobo Wang 0001, Stan Z. Li |
IJCB | 5 |
| 2017 | S^3FD: Single Shot Scale-Invariant Face DetectorabstractThis paper presents a real-time face detector, named Single Shot Scale-invariant Face Detector (S3FD), which performs superiorly on various scales of faces with a single deep neural network, especially for small faces. Specifically, we try to solve the common problem that anchor-based detectors deteriorate dramatically as the objects become smaller. We make contributions in the following three aspects: 1) proposing a scale-equitable face detection framework to handle different scales of faces well. We tile anchors on a wide range of layers to ensure that all scales of faces have enough features for detection. Besides, we design anchor scales based on the effective receptive field and a proposed equal proportion interval principle; 2) improving the recall rate of small faces by a scale compensation anchor matching strategy; 3) reducing the false positive rate of small faces via a max-out background label. As a consequence, our method achieves state-of-the-art detection performance on all the common face detection benchmarks, including the AFW, PASCAL face, FDDB and WIDER FACE datasets, and can run at 36 FPS on a Nvidia Titan X (Pascal) for VGA-resolution images. Xiangyu Zhu 0001, Zhen Lei 0001, Hailin Shi, Xiaobo Wang 0001, Stan Z. Li |
ICCV | 5 |
| 2017 | Soft-Margin Softmax for Deep Classification
Xuezhi Liang, Xiaobo Wang 0001, Zhen Lei 0001, Shengcai Liao, Stan Z. Li |
ICONIP (2) | 2 |
| 2017 | Exclusivity Regularized Machine: A New Ensemble SVM ClassifierabstractThe diversity of base learners is of utmost importance to a good ensemble. This paper defines a novel measurement of diversity, termed as exclusivity. With the designed exclusivity, we further propose an ensemble SVM classifier, namely Exclusivity Regularized Machine (ExRM), to jointly suppress the training error of ensemble and enhance the diversity between bases. Moreover, an Augmented Lagrange Multiplier based algorithm is customized to effectively and efficiently seek the optimal solution of ExRM. Theoretical analysis on convergence, global optimality and linear complexity of the proposed algorithm, as well as experiments are provided to reveal the efficacy of our method and show its superiority over state-of-the-arts in terms of accuracy and efficiency. Xiaojie Guo 0001, Xiaobo Wang 0001, Haibin Ling |
IJCAI | 2 |
| 2017 | Cross-Modality Face Recognition via Heterogeneous Joint BayesianabstractIn many face recognition applications, the modalities of face images between the gallery and probe sets are different, which is known as heterogeneous face recognition. How to reduce the feature gap between images from different modalities is a critical issue to develop a highly accurate face recognition algorithm. Recently, joint Bayesian (JB) has demonstrated superior performance on general face recognition compared to traditional discriminant analysis methods like subspace learning. However, the original JB treats the two input samples equally and does not take into account the modality difference between them and may be suboptimal to address the heterogeneous face recognition problem. In this work, we extend the original JB by modeling the gallery and probe images using two different Gaussian distributions to propose a heterogeneous joint Bayesian (HJB) formulation for cross-modality face recognition. The proposed HJB explicitly models the modality difference of image pairs and, therefore, is able to better discriminate the same/different face pairs accurately. Extensive experiments conducted in the case of visible-near-infrared and ID photo versus spot face recognition problems show the superiority of the HJB over previous methods. Hailin Shi, Xiaobo Wang 0001, Dong Yi, Zhen Lei 0001, Xiangyu Zhu 0001, Stan Z. Li |
IEEE Signal Process. Lett. | 2 |
| 2015 | Adaptively Unified Semi-Supervised Dictionary Learning with Active PointsabstractSemi-supervised dictionary learning aims to construct a dictionary by utilizing both labeled and unlabeled data. To enhance the discriminative capability of the learned dictionary, numerous discriminative terms have been proposed by evaluating either the prediction loss or the class separation criterion on the coding vectors of labeled data, but with rare consideration of the power of the coding vectors corresponding to unlabeled data. In this paper, we present a novel semi-supervised dictionary learning method, which uses the informative coding vectors of both labeled and unlabeled data, and adaptively emphasizes the high confidence coding vectors of unlabeled data to enhance the dictionary discriminative capability simultaneously. By doing so, we integrate the discrimination of dictionary, the induction of classifier to new testing data and the transduction of labels to unlabeled data into a unified framework. To solve the proposed problem, an effective iterative algorithm is designed. Experimental results on a series of benchmark databases show that our method outperforms other state-of-the-art dictionary learning methods in most cases. Xiaobo Wang 0001, Xiaojie Guo 0001, Stan Z. Li |
ICCV | 1 |
| 2014 | Beautifying Fisheye Images using Orientation and Shape CuesabstractFisheye images, due to their wide range of vision, become more and more popular in our daily life. However, the fisheye images usually suffer from misalignment that reduces their visual pleasure. In this paper, we develop a computational method for enhancing the aesthetics of such images by exploiting the orientation and shape cues. More specifically, the orientation cue is based on the observation that cameras are often oriented when taking photos, so that their upvectors are parallel to vertical linear structures in the scene. While the shape one refers to that after repositing the fisheye image, the circular shape should be preserved. By employing these two rules as our basic aesthetic guidelines, our method can correct the rotation angle between the camera coordinate and the world coordinate to make the virtual camera oriented, and complete the missing part. Experimental results on a number of challenging indoor and outdoor fisheye images show the effectiveness of our approach, and demonstrate the superior aesthetics of the proposed method compared to the state-of-the-arts. Xiaobo Wang 0001, Xiaochun Cao, Xiaojie Guo 0001, Zhanjie Song |
ACM Multimedia | 1 |