EDBT 2026 Demo / reviewers in the wild / expert
Hailin Shi
dblp:172/1112
· DBLP profile ↗
35ranked-venue papers
2as first author
15since 2021 · last 2025
0000-0002-3603-2683ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 2 first-author · 14 since 2021Artificial intelligence and machine learning · 21 · 1 first-author · 5 since 2021Computer networks · 3 · 3 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GoHD: Gaze-oriented and Highly Disentangled Portrait Animation with Rhythmic Poses and Realistic ExpressionsabstractAudio-driven talking head generation necessitates seamless integration of audio and visual data amidst the challenges posed by diverse input portraits and intricate correlations between audio and facial motions. In response, we propose a robust framework GoHD designed to produce highly realistic, expressive, and controllable portrait videos from any reference identity with any motion. GoHD innovates with three key modules: Firstly, an animation module utilizing latent navigation is introduced to improve the generalization ability across unseen input styles. This module achieves high disentanglement of motion and identity, and it also incorporates gaze orientation to rectify unnatural eye movements that were previously overlooked. Secondly, a conformer-structured conditional diffusion model is designed to guarantee head poses that are aware of prosody. Thirdly, to estimate lip-synchronized and realistic expressions from the input audio within limited training data, a two-stage training strategy is devised to decouple frequent and frame-wise lip motion distillation from the generation of other more temporally dependent but less audio-related motions, e.g., blinks and frowns. Extensive experiments validate GoHD's advanced generalization capabilities, demonstrating its effectiveness in generating realistic talking face results on arbitrary subjects. Weize Quan, Hailin Shi, Lili Wang 0006, Dong-Ming Yan 0001 |
AAAI | 3 |
| 2025 | C2RL: Content and Context Representation Learning for Gloss-Free Sign Language Translation and RetrievalabstractSign Language Representation Learning (SLRL) is crucial for a range of sign language-related downstream tasks such as Sign Language Translation (SLT) and Sign Language Retrieval (SLRet). Recently, many gloss-based and gloss-free SLRL methods have been proposed, showing promising performance. Among them, the gloss-free approach shows promise for strong scalability without relying on gloss annotations. However, it currently faces suboptimal solutions due to challenges in encoding the intricate, context-sensitive characteristics of sign language videos, mainly struggling to discern essential sign features using a non-monotonic video-text alignment strategy. Therefore, we introduce an innovative pretraining paradigm for gloss-free SLRL, called C2RL, in this paper. Specifically, rather than merely incorporating a non-monotonic semantic alignment of video and text to learn language-oriented sign features, we emphasize two pivotal aspects of SLRL: Implicit Content Learning (ICL) and Explicit Context Learning (ECL). ICL delves into the content of communication, capturing the nuances, emphasis, timing, and rhythm of the signs. In contrast, ECL focuses on understanding the contextual meaning of signs and converting them into equivalent sentences. Despite its simplicity, extensive experiments confirm that the joint optimization of ICL and ECL results in robust sign language representation and significant performance gains in gloss-free SLT and SLRet tasks. Notably, C2RL improves the BLEU-4 score by +5.3 on P14T, +10.6 on CSL-daily, +6.2 on OpenASL, and +1.3 on How2Sign. It also boosts the R@1 score by +8.3 on P14T, +14.4 on CSL-daily, and +5.9 on How2Sign. Additionally, we set a new baseline for the OpenASL dataset in the SLRet task. Benjia Zhou, Jun Wan 0001, Yibo Hu 0001, Hailin Shi, Yanyan Liang 0001, Zhen Lei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | ChatEdit: Towards Multi-turn Interactive Facial Image Editing via DialogueabstractSingle-turn Multi-turn Input Lipstick Pale skin Smiling Black Hair Figure 2: Comparison of previous repeated singleturn editing approaches and our proposed multiturn editing approach.The cascaded errors in the single-turn approach lead to unintended changes in gender and eye makeup. Xing Cui, Zekun Li 0001, Yibo Hu 0001, Hailin Shi, Chunshui Cao, Zhaofeng He 0001 |
EMNLP | 5 |
| 2023 | Introduction to the Special Issue on Trustworthy Multimedia Computing and Applications in Urban ScenesabstractSpecial Issue Part 1 (Issue 3) and Part 2 (Issue 4) of AIEDAM are based on a workshop on Learning and Creativity held at the 2002 conference on Artificial Intelligence in Design, AID '02 (www.cad.strath.ac.uk/AID02_workshop/Workshop_webpage.html; Gero, ... Wu Liu 0005, Hailin Shi, Yunchao Wei, Dan Zeng 0001, Nicu Sebe, Jiebo Luo 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Multi-Agent Semi-Siamese Training for Long-Tail and Shallow Face LearningabstractWith the recent development of deep convolutional neural networks and large-scale datasets, deep face recognition has made remarkable progress and been widely used in various applications. However, unlike the existing public face datasets, in many real-world scenarios of face recognition, the depth of the training dataset is shallow, which means that only two face images are available for each ID. With the non-uniform increase of samples, such issue is converted to a more general case, known as long-tail face learning, which suffers from data imbalance and intra-class diversity dearth simultaneously. These adverse conditions damage the training and result in the decline of model performance. Based on Semi-Siamese Training, we introduce an advanced solution, namedMulti-Agent Semi-Siamese Training(MASST), to address these problems. MASST includes a probe network and multiple gallery agents—the former aims to encode the probe features, and the latter constitutes a stack of networks that encode the prototypes (gallery features). For each training iteration, the gallery network, which is sequentially rotated from the stack, and the probe network form a pair of Semi-Siamese networks. We give the theoretical and empirical analysis that, given the long-tail (or shallow) data and training loss, MASST smooths the loss landscape and satisfies the Lipschitz continuity with the help of multiple agents and the updating gallery queue. The proposed method is out of extra-dependency, and thus can be easily integrated with the existing loss functions and network architectures. It is worth noting that although multiple gallery agents are employed for training, only the probe network is needed for inference, without increasing the inference cost. Extensive experiments and comparisons demonstrate the advantages of MASST for long-tail and shallow face learning. Yichun Tai, Hailin Shi, Dan Zeng 0001, Yibo Hu 0003, Zhijiang Zhang, Tao Mei 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | PetsGAN: Rethinking Priors for Single Image GenerationabstractSingle image generation (SIG), described as generating diverse samples that have the same visual content as the given natural image, is first introduced by SinGAN, which builds a pyramid of GANs to progressively learn the internal patch distribution of the single image. It shows excellent performance in a wide range of image manipulation tasks. However, SinGAN has some limitations. Firstly, due to lack of semantic information, SinGAN cannot handle the object images well as it does on the scene and texture images. Secondly, the independent progressive training scheme is time-consuming and easy to cause artifacts accumulation. To tackle these problems, in this paper, we dig into the single image generation problem and improve SinGAN by fully-utilization of internal and external priors. The main contributions of this paper include: 1) We interpret single image generation from the perspective of the general generative task, that is, to learn a diverse distribution from the Dirac distribution composed of a single image. In order to solve this non-trivial problem, we construct a regularized latent variable model to formulate SIG. To the best of our knowledge, it is the first time to give a clear formulation and optimization goal of SIG, and all the existing methods for SIG can be regarded as special cases of this model. 2) We design a novel Prior-based end-to-end training GAN (PetsGAN), which is infused with internal prior and external prior to overcome the problems of SinGAN. For one thing, we employ the pre-trained GAN model to inject external prior for image generation, which can alleviate the problem of lack of semantic information and generate natural, reasonable and diverse samples, even for the object image. For another, we fully-utilize the internal prior by a differential Patch Matching module and an effective reconstruction network to generate consistent and realistic texture. 3) We construct abundant of qualitative and quantitative experiments on three datasets. The experimental results show our method surpasses other methods on both generated image quality, diversity, and training speed. Moreover, we apply our method to other image manipulation tasks (e.g., style transfer, harmonization) and the results further prove the effectiveness and efficiency of our method. Yinglu Liu, Congying Han, Hailin Shi, Tiande Guo |
AAAI | 4 |
| 2022 | Boosting Semi-Supervised Face Recognition With Noise RobustnessabstractAlthough deep face recognition benefits significantly from large-scale training data, a current bottleneck is the labelling cost. A feasible solution to this problem is semi-supervised learning, exploiting a small portion of labelled data and large amounts of unlabelled data. The major challenge, however, is the accumulated label errors through auto-labelling, compromising the training. In this paper, we present an effective solution to semi-supervised face recognition that is robust to the label noise aroused by the auto-labelling. Specifically, we introduce a multi-agent method, named GroupNet (GN), to endow our solution with the ability to identify the wrongly-labelled samples and preserve the clean samples. We show that GN alone achieves the leading accuracy in traditional supervised face recognition even when the noisy labels take over 50% of the training data. Further, we develop a semi-supervised face recognition solution, named Noise Robust Learning-Labelling (NRoLL), which is based on the robust training ability empowered by GN. It starts with a small amount of labelled data and consequently conducts high-confidence labelling on a large amount of unlabelled data to boost further training. The more data is labelled by NRoLL, the higher confidence is with the label in the dataset. To evaluate the competitiveness of our method, we run NRoLL with a rough condition that only one-fifth of the labelled MSCeleb is available and the rest is used as unlabelled data. On a wide range of benchmarks, our method compares favorably against the state-of-the-art methods. Yuchi Liu, Hailin Shi, Rui Zhu 0014, Jun Wang 0127, Liang Zheng 0001, Tao Mei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Dual Spoof Disentanglement Generation for Face Anti-Spoofing With Depth Uncertainty LearningabstractFace anti-spoofing (FAS) plays a vital role in preventing face recognition systems from presentation attacks. Existing face anti-spoofing datasets lack diversity due to the insufficient identity and insignificant variance, which limits the generalization ability of FAS model. In this paper, we propose Dual Spoof Disentanglement Generation (DSDG) framework to tackle this challenge by “anti-spoofing via generation”. Depending on the interpretable factorized latent disentanglement in Variational Autoencoder (VAE), DSDG learns a joint distribution of the identity representation and the spoofing pattern representation in the latent space. Then, large-scale paired live and spoofing images can be generated from random noise to boost the diversity of the training set. However, some generated face images are partially distorted due to the inherent defect of VAE. Such noisy samples are hard to predict precise depth values, thus may obstruct the widely-used depth supervised optimization. To tackle this issue, we further introduce a lightweight Depth Uncertainty Module (DUM), which alleviates the adverse effects of noisy samples by depth uncertainty learning. DUM is developed without extra-dependency, thus can be flexibly integrated with any depth supervised network for face anti-spoofing. We evaluate the effectiveness of the proposed method on five popular benchmarks and achieve state-of-the-art results under both intra- and inter- test settings. The codes are available athttps://github.com/JDAI-CV/FaceX-Zoo/tree/main/addition_module/DSDG. Hangtong Wu, Dan Zeng 0001, Yibo Hu 0003, Hailin Shi, Tao Mei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | FasterPose: A Faster Simple Baseline for Human Pose EstimationabstractThe performance of human pose estimation depends on the spatial accuracy of keypoint localization. Most existing methods pursue the spatial accuracy through learning the high-resolution (HR) representation from input images. By the experimental analysis, we find that the HR representation leads to a sharp increase of computational cost, while the accuracy improvement remains marginal compared with the low-resolution (LR) representation. In this article, we propose a design paradigm for cost-effective network with LR representation for efficient pose estimation, named FasterPose. Whereas the LR design largely shrinks the model complexity, how to effectively train the network with respect to the spatial accuracy is a concomitant challenge. We study the training behavior of FasterPose and formulate a novel regressive cross-entropy (RCE) loss function for accelerating the convergence and promoting the accuracy. The RCE loss generalizes the ordinary cross-entropy loss from the binary supervision to a continuous range, thus the training of pose estimation network is able to benefit from the sigmoid function. By doing so, the output heatmap can be inferred from the LR features without loss of spatial accuracy, while the computational cost and model size has been significantly reduced. Compared with the previously dominant network of pose estimation, our method reduces 58% of the FLOPs and simultaneously gains 1.3% improvement of accuracy. Extensive experiments show that FasterPose yields promising results on the common benchmarks, i.e., COCO and MPII, consistently validating the effectiveness and efficiency for practical utilization, especially the low-latency and low-energy-budget applications in the non-GPU scenarios. Hanbin Dai, Hailin Shi, Wu Liu 0005, Linfang Wang, Yinglu Liu, Tao Mei 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Dive Into Ambiguity: Latent Distribution Mining and Pairwise Uncertainty Estimation for Facial Expression RecognitionabstractDue to the subjective annotation and the inherent interclass similarity of facial expressions, one of key challenges in Facial Expression Recognition (FER) is the annotation ambiguity. In this paper, we proposes a solution, named DMUE, to address the problem of annotation ambiguity from two perspectives: the latent Distribution Mining and the pairwise Uncertainty Estimation. For the former, an auxiliary multi-branch learning framework is introduced to better mine and describe the latent distribution in the label space. For the latter, the pairwise relationship of semantic feature between instances are fully exploited to estimate the ambiguity extent in the instance space. The proposed method is independent to the backbone architectures, and brings no extra burden for inference. The experiments are conducted on the popular real-world benchmarks and the synthetic noisy datasets. Either way, the proposed DMUE stably achieves leading performance. Jiahui She, Yibo Hu 0003, Hailin Shi, Jun Wang 0127, Qiu Shen, Tao Mei 0001 |
CVPR | 3 |
| 2021 | CM-NAS: Cross-Modality Neural Architecture Search for Visible-Infrared Person Re-IdentificationabstractVisible-Infrared person re-identification (VI-ReID) aims to match cross-modality pedestrian images, breaking through the limitation of single-modality person ReID in dark environment. In order to mitigate the impact of large modality discrepancy, existing works manually design various two-stream architectures to separately learn modality-specific and modality-sharable representations. Such a manual design routine, however, highly depends on massive experiments and empirical practice, which is time consuming and labor intensive. In this paper, we systematically study the manually designed architectures, and identify that appropriately separating Batch Normalization (BN) layers is the key to bring a great boost towards cross-modality matching. Based on this observation, the essential objective is to find the optimal separation scheme for each BN layer. To this end, we propose a novel method, named Cross-Modality Neural Architecture Search (CM-NAS). It consists of a BN-oriented search space in which the standard optimization can be fulfilled subject to the cross-modality task. Equipped with the searched architecture, our method outperforms state-of-the-art counter-parts in both two benchmarks, improving the Rank-1/mAP by 6.70%/6.13% on SYSU-MM01 and by 12.17%/11.23% on RegDB. Code is released at https://github.com/JDAI-CV/CM-NAS. Chaoyou Fu, Yibo Hu 0001, Xiang Wu 0001, Hailin Shi, Tao Mei 0001, Ran He 0001 |
ICCV | 4 |
| 2021 | One-stage Context and Identity Hallucination NetworkabstractFace swapping aims to synthesize a face image, in which the facial identity is well transplanted from the source image and the context (e.g., hairstyle, head posture, facial expression, lighting, and background) keeps consistent with the reference image. The prior work mainly accomplishes the task in two stages, i.e., generating the inner face with the source identity, and then stitching the generation with the complementary part of the reference image by image blending techniques. The blending mask, which is usually obtained by the additional face segmentation model, is a common practice towards photo-realistic face swapping. However, artifacts usually appear at the blending boundary, especially in areas occluded by the hair, eyeglasses, accessories, etc. To address this problem, rather than struggling with the blending mask in the two-stage routine, we develop a novel one-stage context and identity hallucination network, which learns a series of hallucination maps to softly divide the context areas and identity areas. For context areas, the features are fully utilized by a multi-level context encoder. For identity areas, we design a novel two-cascading AdaIN to transfer the identity while retaining the context. Besides, with the help of hallucination maps, we introduce an effectively improved reconstruction loss to utilize unlimited unpaired face images for training. Our network performs well on both context areas and identity areas without any dependency on post-processing. Extensive qualitative and quantitative experiments demonstrate the superiority of our network. Yinglu Liu, Mingcan Xiang, Hailin Shi, Tao Mei 0001 |
ACM Multimedia | 3 |
| 2021 | FaceX-Zoo: A PyTorch Toolbox for Face RecognitionabstractDue to the remarkable progress in recent years, deep face recognition is in great need of public support for practical model production and further exploration. The demands are in three folds, including 1) modular training scheme, 2) standard and automatic evaluation, and 3) groundwork of deployment. To meet these demands, we present a novel open-source project, named FaceX-Zoo, which is constructed with modular and scalable design, and oriented to the academic and industrial community of face-related analysis. FaceX-Zoo provides 1) the training module with various choices of backbone and supervisory head; 2) the evaluation module that enables standard and automatic test on most popular benchmarks; 3) the module of simple yet fully functional face SDK for the validation and primary application of end-to-end face recognition; 4) the additional module that integrates a group of useful tools. Based on these easy-to-use modules, FaceX-Zoo can help the community to easily build stateof-the-art solutions for deep face recognition and, such like the newly-emerged challenge of masked face recognition caused by the worldwide COVID-19 pandemic. Besides, FaceX-Zoo can be easily upgraded and scaled up along with further exploration in face related fields. The source codes and models have been released and received over 900 stars at https://github.com/JDAI-CV/FaceX-Zoo. Jun Wang 0127, Yinglu Liu, Yibo Hu 0003, Hailin Shi, Tao Mei 0001 |
ACM Multimedia | 4 |
| 2021 | Towards NIR-VIS Masked Face RecognitionabstractNear-infrared to visible (NIR-VIS) face recognition is the most common case in heterogeneous face recognition, which aims to match a pair of face images captured from two different modalities. Existing deep learning based methods have made remarkable progress in NIR-VIS face recognition, while it encounters certain newly-emerged difficulties during the pandemic of COVID-19, since people are supposed to wear facial masks to cut off the spread of the virus. We define this task as NIR-VIS masked face recognition, and find it problematic with the masked face in the NIR probe image. First, the lack of masked face data is a challenging issue for the network training. Second, most of the facial parts (cheeks, mouth, nose etc.) are fully occluded by the mask, which leads to a large amount of loss of information. Third, the domain gap still exists in the remaining facial parts. In such scenario, the existing methods suffer from significant performance degradation caused by the above issues. In this paper, we aim to address the challenge of NIR-VIS masked face recognition from the perspectives of training data and training method. Specifically, we propose a novel heterogeneous training method to maximize the mutual information shared by the face representation of two domains with the help of semi-siamese networks. In addition, a 3D face reconstruction based approach is employed to synthesize masked face from the existing NIR image. Resorting to these practices, our solution provides the domain-invariant face representation which is also robust to the mask occlusion. Extensive experiments on three NIR-VIS face datasets demonstrate the effectiveness and cross-dataset-generalization capacity of our method. Hailin Shi, Yinglu Liu, Dan Zeng 0001, Tao Mei 0001 |
IEEE Signal Process. Lett. | 2 |
| 2021 | AGRNet: Adaptive Graph Representation Learning and Reasoning for Face ParsingabstractFace parsing infers a pixel-wise label to each facial component, which has drawn much attention recently. Previous methods have shown their success in face parsing, which however overlook the correlation among facial components. As a matter of fact, the component-wise relationship is a critical clue in discriminating ambiguous pixels in facial area. To address this issue, we propose adaptive graph representation learning and reasoning over facial components, aiming to learn representative vertices that describe each component, exploit the component-wise relationship and thereby produce accurate parsing results against ambiguity. In particular, we devise an adaptive and differentiable graph abstraction method to represent the components on a graph via pixel-to-vertex projection under the initial condition of a predicted parsing map, where pixel features within a certain facial region are aggregated onto a vertex. Further, we explicitly incorporate the image edge as a prior in the model, which helps to discriminate edge and non-edge pixels during the projection, thus leading to refined parsing results along the edges. Then, our model learns and reasons over the relations among components by propagating information across vertices on the graph. Finally, the refined vertex features are projected back to pixel grids for the prediction of the final parsing map. To train our model, we propose a discriminative loss to penalize small distances between vertices in the feature space, which leads to distinct vertices with strong semantics. Experimental results show the superior performance of the proposed model on multiple face parsing datasets, along with the validation on the human parsing task to demonstrate the generalizability of our model. Gusi Te, Wei Hu 0003, Yinglu Liu, Hailin Shi, Tao Mei 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | A New Dataset and Boundary-Attention Semantic Segmentation for Face ParsingabstractFace parsing has recently attracted increasing interest due to its numerous application potentials, such as facial make up and facial image generation. In this paper, we make contributions on face parsing task from two aspects. First, we develop a high-efficiency framework for pixel-level face parsing annotating and construct a new large-scale Landmark guided face Parsing dataset (LaPa). It consists of more than 22,000 facial images with abundant variations in expression, pose and occlusion, and each image of LaPa is provided with an 11-category pixel-level label map and 106-point landmarks. The dataset is publicly accessible to the community for boosting the advance of face parsing.1 Second, a simple yet effective Boundary-Attention Semantic Segmentation (BASS) method is proposed for face parsing, which contains a three-branch network with elaborately developed loss functions to fully exploit the boundary information. Extensive experiments on our LaPa benchmark and the public Helen dataset show the superiority of our proposed method. Yinglu Liu, Hailin Shi, Yue Si, Xiaobo Wang 0001, Tao Mei 0001 |
AAAI | 2 |
| 2020 | Mis-Classified Vector Guided Softmax Loss for Face RecognitionabstractFace recognition has witnessed significant progress due to the advances of deep convolutional neural networks (CNNs), the central task of which is how to improve the feature discrimination. To this end, several margin-based (e.g., angular, additive and additive angular margins) softmax loss functions have been proposed to increase the feature margin between different classes. However, despite great achievements have been made, they mainly suffer from three issues: 1) Obviously, they ignore the importance of informative features mining for discriminative learning; 2) They encourage the feature margin only from the ground truth class, without realizing the discriminability from other non-ground truth classes; 3) The feature margin between different classes is set to be same and fixed, which may not adapt the situations very well. To cope with these issues, this paper develops a novel loss function, which adaptively emphasizes the mis-classified feature vectors to guide the discriminative feature learning. Thus we can address all the above issues and achieve more discriminative face features. To the best of our knowledge, this is the first attempt to inherit the advantages of feature margin and feature mining into a unified loss function. Experimental results on several benchmarks have demonstrated the effectiveness of our method over state-of-the-art alternatives. Our code is available at http://www.cbsr.ia.ac.cn/users/xiaobowang/. Xiaobo Wang 0001, Tianyu Fu 0001, Hailin Shi, Tao Mei 0001 |
AAAI | 5 |
| 2020 | Semi-Siamese Training for Shallow Face Learning
Hailin Shi, Yuchi Liu, Jun Wang 0127, Zhen Lei 0001, Dan Zeng 0001, Tao Mei 0001 |
ECCV (4) | 2 |
| 2020 | Edge-Aware Graph Representation Learning and Reasoning for Face Parsing
Gusi Te, Yinglu Liu, Wei Hu 0003, Hailin Shi, Tao Mei 0001 |
ECCV (12) | 4 |
| 2020 | Dual-Structure Disentangling Variational Generation for Data-Limited Face ParsingabstractDeep learning based face parsing methods have attained state-of-the-art performance in recent years. Their superior performance heavily depends on the large-scale annotated training data. However, it is expensive and time-consuming to construct a large-scale pixel-level manually annotated dataset for face parsing. To alleviate this issue, we propose a novel Dual-Structure Disentangling Variational Generation (D2VG) network. Benefiting from the interpretable factorized latent disentanglement in VAE, D2VG can learn a joint structural distribution of facial image and its corresponding parsing map. Owing to these, it can synthesize large-scale paired face images and parsing maps from a standard Gaussian distribution. Then, we adopt both manually annotated and synthesized data to train a face parsing model in a supervised way. Since there are inaccurate pixel-level labels in synthesized parsing maps, we introduce a coarseness-tolerant learning algorithm, to effectively handle these noisy or uncertain labels. In this way, we can significantly boost the performance of face parsing. Extensive quantitative and qualitative results on HELEN, CelebAMask-HQ and LaPa demonstrate the superiority of our methods. Peipei Li 0002, Yinglu Liu, Hailin Shi, Xiang Wu 0001, Yibo Hu 0001, Ran He 0001, Zhenan Sun |
ACM Multimedia | 3 |
| 2019 | A Dataset and Benchmark for Large-Scale Multi-Modal Face Anti-SpoofingabstractFace anti-spoofing is essential to prevent face recognition systems from a security breach. Much of the progresses have been made by the availability of face anti-spoofing benchmark datasets in recent years. However, existing face anti-spoofing benchmarks have limited number of subjects (≤170) and modalities (≤2), which hinder the further development of the academic community. To facilitate face anti-spoofing research, we introduce a large-scale multi-modal dataset, namely CASIA-SURF, which is the largest publicly available dataset for face anti-spoofing in terms of both subjects and visual modalities. Specifically, it consists of 1,000 subjects with 21,000 videos and each sample has 3 modalities (i.e., RGB, Depth and IR). We also provide a measurement set, evaluation protocol and training/validation/testing subsets, developing a new benchmark for face anti-spoofing. Moreover, we present a new multi-modal fusion method as baseline, which performs feature re-weighting to select the more informative channel features while suppressing the less useful ones for each modal. Extensive experiments have been conducted on the proposed dataset to verify its significance and generalization capability. The dataset is available at https://sites.google.com/qq.com/chalearnfacespoofingattackdete/. Xiaobo Wang 0001, Ajian Liu 0001, Jun Wan 0001, Sergio Escalera, Hailin Shi, Stan Z. Li |
CVPR | 7 |
| 2019 | ScratchDet: Training Single-Shot Object Detectors From ScratchabstractCurrent state-of-the-art object objectors are fine-tuned from the off-the-shelf networks pretrained on large-scale classification dataset ImageNet, which incurs some additional problems: 1) The classification and detection have different degrees of sensitivity to translation, resulting in the learning objective bias; 2) The architecture is limited by the classification network, leading to the inconvenience of modification. To cope with these problems, training detectors from scratch is a feasible solution. However, the detectors trained from scratch generally perform worse than the pretrained ones, even suffer from the convergence issue in training. In this paper, we explore to train object detectors from scratch robustly. By analysing the previous work on optimization landscape, we find that one of the overlooked points in current trained-from-scratch detector is the BatchNorm. Resorting to the stable and predictable gradient brought by BatchNorm, detectors can be trained from scratch stably while keeping the favourable performance independent to the network architecture. Taking this advantage, we are able to explore various types of networks for object detection, without suffering from the poor convergence. By extensive experiments and analyses on downsampling factor, we propose the Root-ResNet backbone network, which makes full use of the information from original images. Our ScratchDet achieves the state-of-the-art accuracy on PASCAL VOC 2007, 2012 and MS COCO among all the train-from-scratch detectors and even performs better than several one-stage pretrained methods. Codes will be made publicly available at https://github.com/KimSoybean/ScratchDet. Rui Zhu 0014, Xiaobo Wang 0001, Longyin Wen, Hailin Shi, Liefeng Bo, Tao Mei 0001 |
CVPR | 5 |
| 2019 | Co-Mining: Deep Face Recognition With Noisy LabelsabstractFace recognition has achieved significant progress with the growing scale of collected datasets, which empowers us to train strong convolutional neural networks (CNNs). While a variety of CNN architectures and loss functions have been devised recently, we still have a limited understanding of how to train the CNN models with the label noise inherent in existing face recognition datasets. To address this issue, this paper develops a novel co-mining strategy to effectively train on the datasets with noisy labels. Specifically, we simultaneously use the loss values as the cue to detect noisy labels, exchange the high-confidence clean faces to alleviate the errors accumulated issue caused by the sample-selection bias, and re-weight the predicted clean faces to make them dominate the discriminative model training in a mini-batch fashion. Extensive experiments by training on three popular datasets (\textit{i.e.}, CASIA-WebFace, MS-Celeb-1M and VggFace2) and testing on several benchmarks, including LFW, AgeDB, CFP, CALFW, CPLFW, RFW, and MegaFace, have demonstrated the effectiveness of our new approach over the state-of-the-art alternatives. Xiaobo Wang 0001, Hailin Shi, Jun Wang 0127, Tao Mei 0001 |
ICCV | 3 |
| 2019 | MetaAdvDet: Towards Robust Detection of Evolving Adversarial AttacksabstractDeep neural networks (DNNs) are vulnerable to the adversarial attack which is maliciously implemented by adding human-imperceptible perturbation to images and thus leads to incorrect prediction. Existing studies have proposed various methods to detect the new adversarial attacks. However, new attack methods keep evolving constantly and yield new adversarial examples to bypass the existing detectors. It needs to collect tens of thousands samples to train detectors, while the new attacks evolve much more frequently than the high-cost data collection. Thus, this situation leads the newly evolved attack samples to remain in small scales. To solve such few-shot problem with the evolving attacks, we propose a meta-learning based robust detection method to detect new adversarial attacks with limited examples. Specifically, the learning consists of a double-network framework: a task-dedicated network and a master network which alternatively learn the detection capability for either seen attack or a new attack. To validate the effectiveness of our approach, we construct the benchmarks with few-shot-fashion protocols based on three conventional datasets, i.e. CIFAR-10, MNIST and Fashion-MNIST. Comprehensive experiments are conducted on them to verify the superiority of our approach with respect to the traditional adversarial attack detection methods. The implementation code is available online. Chen Ma 0003, Hailin Shi, Li Chen 0031, Jun-Hai Yong, Dan Zeng 0001 |
ACM Multimedia | 3 |
| 2019 | Single-Shot Scale-Aware Network for Real-Time Face Detection
Longyin Wen, Hailin Shi, Zhen Lei 0001, Siwei Lyu, Stan Z. Li |
Int. J. Comput. Vis. | 3 |
| 2019 | Large-Scale Bisample Learning on ID Versus Spot Face Recognition
Xiangyu Zhu 0001, Zhen Lei 0001, Hailin Shi, Fan Yang 0062, Dong Yi, Guo-Jun Qi, Stan Z. Li |
Int. J. Comput. Vis. | 4 |
| 2019 | Multi-view subspace clustering with intactness-aware similarity
Xiaobo Wang 0001, Zhen Lei 0001, Xiaojie Guo 0001, Changqing Zhang 0002, Hailin Shi, Stan Z. Li |
Pattern Recognit. | 5 |
| 2018 | Co-Referenced Subspace ClusteringabstractSubspace clustering refers to the problem of grouping data into their underlying groups. To address this task, spectral clustering based technique is arguably one of the most popular approaches, and its performance largely depends on the constructed similarity. However, most existing works merely employ the primary representation (e.g., sparse or low-rank representation) as the similarity. In this paper, we propose to explore a high-level co-referenced similarity by employing the Hilbert-Schmidt Independence Criterion (HSIC). Moreover, geometry interpretation of the advantage of our co-referenced similarity is provided. Representation-induced kernels such as Mahalanobis metric, can also be easily embedded into the formulation. Extensive experiments on both synthetic and real-world data are conducted to show the superiority of the proposed method over the state-of-the-art alternatives. Xiaobo Wang 0001, Zhen Lei 0001, Hailin Shi, Xiaojie Guo 0001, Xiangyu Zhu 0001, Stan Z. Li |
ICME | 3 |
| 2018 | Detecting Face with Densely Connected Face Proposal Network
Xiangyu Zhu 0001, Zhen Lei 0001, Xiaobo Wang 0001, Hailin Shi, Stan Z. Li |
Neurocomputing | 5 |
| 2017 | FaceBoxes: A CPU real-time face detector with high accuracyabstractAlthough tremendous strides have been made in face detection, one of the remaining open challenges is to achieve real-time speed on the CPU as well as maintain high performance, since effective models for face detection tend to be computationally prohibitive.To address this challenge, we propose a novel face detector, named FaceBoxes, with superior performance on both speed and accuracy.Specifically, our method has a lightweight yet powerful network structure that consists of the Rapidly Digested Convolutional Layers (RDCL) and the Multiple Scale Convolutional Layers (MSCL).The RDCL is designed to enable Face-Boxes to achieve real-time speed on the CPU.The MSCL aims at enriching the receptive fields and discretizing anchors over different layers to handle faces of various scales.Besides, we propose a new anchor densification strategy to make different types of anchors have the same density on the image, which significantly improves the recall rate of small faces.As a consequence, the proposed detector runs at 20 FPS on a single CPU core and 125 FPS using a GPU for VGA-resolution images.Moreover, the speed of FaceBoxes is invariant to the number of faces.We comprehensively evaluate this method and present stateof-the-art detection performance on several face detection benchmark datasets, including the AFW, PASCAL face, and FDDB. Xiangyu Zhu 0001, Zhen Lei 0001, Hailin Shi, Xiaobo Wang 0001, Stan Z. Li |
IJCB | 4 |
| 2017 | S^3FD: Single Shot Scale-Invariant Face DetectorabstractThis paper presents a real-time face detector, named Single Shot Scale-invariant Face Detector (S3FD), which performs superiorly on various scales of faces with a single deep neural network, especially for small faces. Specifically, we try to solve the common problem that anchor-based detectors deteriorate dramatically as the objects become smaller. We make contributions in the following three aspects: 1) proposing a scale-equitable face detection framework to handle different scales of faces well. We tile anchors on a wide range of layers to ensure that all scales of faces have enough features for detection. Besides, we design anchor scales based on the effective receptive field and a proposed equal proportion interval principle; 2) improving the recall rate of small faces by a scale compensation anchor matching strategy; 3) reducing the false positive rate of small faces via a max-out background label. As a consequence, our method achieves state-of-the-art detection performance on all the common face detection benchmarks, including the AFW, PASCAL face, FDDB and WIDER FACE datasets, and can run at 36 FPS on a Nvidia Titan X (Pascal) for VGA-resolution images. Xiangyu Zhu 0001, Zhen Lei 0001, Hailin Shi, Xiaobo Wang 0001, Stan Z. Li |
ICCV | 4 |
| 2017 | Cross-Modality Face Recognition via Heterogeneous Joint BayesianabstractIn many face recognition applications, the modalities of face images between the gallery and probe sets are different, which is known as heterogeneous face recognition. How to reduce the feature gap between images from different modalities is a critical issue to develop a highly accurate face recognition algorithm. Recently, joint Bayesian (JB) has demonstrated superior performance on general face recognition compared to traditional discriminant analysis methods like subspace learning. However, the original JB treats the two input samples equally and does not take into account the modality difference between them and may be suboptimal to address the heterogeneous face recognition problem. In this work, we extend the original JB by modeling the gallery and probe images using two different Gaussian distributions to propose a heterogeneous joint Bayesian (HJB) formulation for cross-modality face recognition. The proposed HJB explicitly models the modality difference of image pairs and, therefore, is able to better discriminate the same/different face pairs accurately. Extensive experiments conducted in the case of visible-near-infrared and ID photo versus spot face recognition problems show the superiority of the HJB over previous methods. Hailin Shi, Xiaobo Wang 0001, Dong Yi, Zhen Lei 0001, Xiangyu Zhu 0001, Stan Z. Li |
IEEE Signal Process. Lett. | 1 |
| 2016 | Metric Embedded Discriminative Vocabulary Learning for High-Level Person RepresentationabstractA variety of encoding methods for bag of word (BoW) model have been proposed to encode the local features in image classification. However, most of them are unsupervised and just employ k-means to form the visual vocabulary, thus reducing the discriminative power of the features. In this paper, we propose a metric embedded discriminative vocabulary learning for high-level person representation with application to person re-identification. A new and effective term is introduced which aims at making the same persons closer while different ones farther in the metric space. With the learned vocabulary, we utilize a linear coding method to encode the image-level features (or holistic image features) for extracting high-level person representation. Different from traditional unsupervised approaches, our method can explore the relationship(same or not) among the persons. Since there is an analytic solution to the linear coding, it is easy to obtain the final high-level features. The experimental results on person re-identification demonstrate the effectiveness of our proposed algorithm. Yang Yang 0062, Zhen Lei 0001, Hailin Shi, Stan Z. Li |
AAAI | 4 |
| 2016 | Face Alignment Across Large Poses: A 3D SolutionabstractFace alignment, which fits a face model to an image and extracts the semantic meanings of facial pixels, has been an important topic in CV community. However, most algorithms are designed for faces in small to medium poses (below 45), lacking the ability to align faces in large poses up to 90. The challenges are three-fold: Firstly, the commonly used landmark-based face model assumes that all the landmarks are visible and is therefore not suitable for profile views. Secondly, the face appearance varies more dramatically across large poses, ranging from frontal view to profile view. Thirdly, labelling landmarks in large poses is extremely challenging since the invisible landmarks have to be guessed. In this paper, we propose a solution to the three problems in an new alignment framework, called 3D Dense Face Alignment (3DDFA), in which a dense 3D face model is fitted to the image via convolutional neutral network (CNN). We also propose a method to synthesize large-scale training samples in profile views to solve the third problem of data labelling. Experiments on the challenging AFLW database show that our approach achieves significant improvements over state-of-the-art methods. Xiangyu Zhu 0001, Zhen Lei 0001, Xiaoming Liu 0002, Hailin Shi, Stan Z. Li |
CVPR | 4 |
| 2016 | Embedding Deep Metric for Person Re-identification: A Study Against Large Variations
Hailin Shi, Yang Yang 0062, Xiangyu Zhu 0001, Shengcai Liao, Zhen Lei 0001, Wei-Shi Zheng 0001, Stan Z. Li |
ECCV (1) | 1 |