Shengcai Liao

dblp:16/8313 · DBLP profile ↗
← Back
95ranked-venue papers
11as first author
30since 2021 · last 2026
0000-0001-8941-2295ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 69 · 10 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 65 · 7 first-author · 17 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Security and privacy · 4 · 2 first-authorHuman-computer interaction and ubiquitous computing · 4 · 3 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning
abstract
Parameter-efficient transfer learning (PETL) has emerged as a pivotal paradigm for adapting pre-trained foundation models to downstream tasks, significantly reducing trainable parameters yet suffering from substantial memory overhead caused by gradient backpropagation during fine-tuning. While memory-efficient transfer learning (METL) circumvents this challenge by bypassing backbone gradient computation via lightweight small side networks, its stringent memory constraint severely limits learning capacity of side networks, thereby significantly compromising performance. To address these limitations, we propose a novel Mixed-Precision Interactive Side Mixture-of-Experts framework (MP-ISMoE). Specifically, we first propose an Gaussian Noise Perturbed Iterative Quantization (GNP-IQ) scheme to quantize weights into lower-bits while effectively decreasing quantization errors. By leveraging memory conserved from GNP-IQ, we subsequently employ Interactive Side Mixture-of-Experts (ISMoE) to scale up side networks without sacrificing overall memory efficiency. Different from conventional mixture-of-experts, ISMoE learns to select optimal experts by interacting with salient features from frozen backbones, thus suppressing knowledge forgetting and boosting performance. Extensive experiments across diverse vision-language and language-only tasks demonstrate that MP-ISMoE remarkably promotes accuracy compared to state-of-the-art METL approaches, while maintaining comparable parameter and memory efficiency.
Zimeng Wu, Shengcai Liao, Shujiang Wu, Jiaxin Chen 0002
AAAI3
2026 From Triangles to Squares: A Unified Visual Path to Summation, Induction, and Proof in Teaching Discrete Mathematics
Shengcai Liao
CSEDU (3)1
2026 ConsistentID: Portrait Generation With Multimodal Fine-Grained Identity Preserving
abstract
Diffusion-based technologies have made significant strides, particularly in personalized and customized facial generation. However, existing methods struggle to achieve high-fidelity and detailed identity (ID) consistency. This is mainly due to two challenges: insufficient fine-grained control over specific facial areas and the absence of a comprehensive strategy for ID preservation that accounts for both intricate facial details and the overall facial structure. To address these limitations, we introduce ConsistentID, an innovative method crafted for diverse identity-preserving portrait generation under fine-grained multimodal facial prompts, utilizing only a single reference image. ConsistentID comprises two core components: a multimodal facial prompt generator and an ID-preservation network. The facial prompt generator combines localized facial features, facial feature descriptions, and overall facial descriptions to enhance the precision of facial detail reconstruction. The ID-preservation network, optimized with a facial attention localization strategy, ensures consistent identity preservation across facial regions. Together, these components leverage fine-grained multimodal identity information to improve identity preservation accuracy significantly. To drive ConsistentID's training, we propose a fine-grained portrait dataset, FGID, with over 500,000 facial images, offering greater diversity and comprehensiveness than existing public facial datasets. Experimental results substantiate that our ConsistentID achieves exceptional precision and diversity in personalized facial generation, surpassing existing methods in the MyStyle dataset. In addition, although ConsistentID introduces more multimodal ID information, it still maintains rapid inference speed during the generation process.
Jiehui Huang, Wenhui Song, Zheng Chong, Zhenchao Tang, Yuhao Cheng, Long Chen 0005, Yiqiang Yan, Shengcai Liao, Xiaodan Liang
IEEE Trans. Pattern Anal. Mach. Intell.11
2025 Large Models are Good Annotators for Zero-Shot Learning
abstract
Human-annotated attributes serve as effective semantic label embeddings for zero-shot learning (ZSL); however, their annotation is labor-intensive and difficult to scale. Recent studies have explored weakly supervised semantic label embeddings to reduce human effort, but these methods often fail to capture visual similarity and underperform compared to human-annotated semantics. In this work, we propose a minimally supervised yet effective approach: GPT- and CLIP-powered attributes (GCAtt). Specifically, we introduce a three-step interaction process with ChatGPT-comprising preliminary design, hierarchical refinement, and specific value determination-to generate attributes that are both category-shared and discriminative for classification. Additionally, we develop a method that encodes attributes and their values as potential text pairings, leveraging CLIP's retrieval capabilities for annotation. Experimental results on four widely used benchmarks demonstrate that GCAtt consistently outperforms human-annotated semantics. Code and data are available at https://github.com/RowenaHe/GCAtt.
Qingzhi He, Wentong Li 0001, Shengcai Liao, Rong Quan, Tong Cui, Jie Qin 0004
SIGIR4
2025 Realistic and Efficient Face Swapping: A Unified Approach with Diffusion Models
abstract
Despite promising progress in face swapping task, realistic swapped images remain elusive, often marred by artifacts, particularly in scenarios involving high pose variation, color differences, and occlusion. To address these issues, we propose a novel approach that better harnesses diffusion models for face-swapping by making following core contributions. (a) We propose to reframe the face-swapping task as a self-supervised, train-time inpainting problem, enhancing the identity transfer while blending with the target image. (b) We introduce a multi-step De-noising Diffusion Implicit Model (DDIM) sampling during training, reinforcing identity and perceptual similarities. (c) Third, we introduce CLIP feature disentanglement to extract pose, expression, and lighting information from the target image, improving fidelity. (d) Further, we introduce a mask shuffling technique during inpainting training, which allows us to create a so-called universal model for swapping, with an additional feature of head swapping. Ours can swap hair and even accessories, beyond traditional face swapping. Unlike prior works reliant on multiple off-the-shelf models, ours is a relatively unified approach and so it is resilient to errors in other off-the-shelf models. Extensive experiments on FFHQ and CelebA datasets validate the efficacy and robustness of our approach, show-casing high-fidelity, realistic face-swapping with minimal inference time. Our code is available at REFace.
Sanoojan Baliah, Qinliang Lin, Shengcai Liao, Xiaodan Liang, Muhammad Haris Khan
WACV3
2025 EAMNet: Efficient Adaptive Mamba Network for Infrared Small-Target Detection
abstract
Infrared small target detection (ISTD) is essential for various fields. Recent approaches based on existing network structures including convolutional neural networks (CNNs), Transformers, and diffusion models, still face challenges in balancing accuracy and efficiency. To address this problem, this paper proposes an Efficient Adaptive Mamba Network (EAMNet) based on the advanced Mamba structure, which effectively models long-range dependencies while maintaining linear complexity, enabling EAMNet to achieve superior detection performance while significantly improving efficiency. First, a Mamba-based UNet architecture is introduced, which processes separated features in parallel, making it highly efficient with a low parameter count and computational cost. To better adapt the Mamba-based framework to the unique characteristics of infrared images, such as low contrast and small target sizes, we propose an adaptive filter module (AFM) that applies adaptive filtering by predicting filter parameters through an additional designed sub-network, enhancing the boundaries and visibility of infrared targets. To further enhance model performance and ensure efficient feature fusion, we propose a shared adaptive spatial attention module (SASAM), which enables a more compact and efficient feature representation in generating spatial attention maps, while minimizing additional computational overhead. Extensive experiments on public benchmarks demonstrate the effectiveness of the proposed EAMNet in both improving accuracy and efficiency compared to existing state-of-the-art methods. Besides, ablation experiments verify the effectiveness of each module. The code is available at https://github.com/jiangjin1246/EAMNet.
Jin Jiang 0003, Shengcai Liao, Xiaoyuan Yang 0003, Kangqing Shen
IEEE Trans. Geosci. Remote. Sens.2
2025 RealignDiff: Boosting Text-to-Image Diffusion Model With Coarse-to-Fine Semantic Realignment
abstract
Recent advances in text-to-image diffusion models have achieved remarkable success in generating high-quality, realistic images from textual descriptions. However, these approaches have faced challenges in precisely aligning the generated visual content with the textual concepts described in the prompts. In this article, we propose a two-stage coarse-to-fine semantic realignment method, named RealignDiff, aimed at improving the alignment between text and images in text-to-image diffusion models. In the coarse semantic realignment phase, a novel caption reward, leveraging the BLIP-2 model, is proposed to evaluate the semantic discrepancy between the generated image caption and the given text prompt. Subsequently, the fine semantic realignment stage uses a local dense caption generation module and a reweighting attention modulation module to refine the previously generated images from a local semantic view. Experimental results on the MS-COCO and ViLG-300 datasets demonstrate that the proposed two-stage coarse-to-fine semantic realignment method outperforms other baseline realignment techniques by a substantial margin in both visual quality and semantic similarity with the input prompt.
Zutao Jiang, Guian Fang, Jianhua Han, Guansong Lu, Hang Xu 0004, Shengcai Liao, Xiaojun Chang, Xiaodan Liang
IEEE Trans. Neural Networks Learn. Syst.6
2024 Any-Shift Prompting for Generalization Over Distributions
abstract
Image-language models with prompt learning have shown remarkable advances in numerous downstream vision tasks. Nevertheless, conventional prompt learning methods overfit their training distribution and lose the generalization ability on test distributions. To improve generalization across various distribution shifts, we propose any-shift prompting: a general probabilistic inference framework that considers the relationship between training and test distributions during prompt learning. We explicitly connect training and test distributions in the latent space by constructing training and test prompts in a hierarchical architecture. Within this framework, the test prompt exploits the distribution relationships to guide the generalization of the CLIP image-language model from training to any test distribution. To effectively encode the distribution information and their relationships, we further introduce a transformer inference network with a pseudo-shift training mechanism. The network generates the tailored test prompt with both training and test information in a feed forward pass, avoiding extra training costs at test time. Extensive experiments on twenty-three datasets demonstrate the effectiveness of any-shift prompting on the generalization over various distribution shifts.
Zehao Xiao, Mohammad Mahdi Derakhshani, Shengcai Liao, Cees Snoek
CVPR4
2024 HumanRefiner: Benchmarking Abnormal Human Generation and Refining with Coarse-to-Fine Pose-Reversible Guidance
Guian Fang, Wenbiao Yan, Yuanfan Guo, Jianhua Han, Zutao Jiang, Hang Xu 0004, Shengcai Liao, Xiaodan Liang
ECCV (32)7
2024 Pseudo-Labeling Based Practical Semi-Supervised Meta-Training for Few-Shot Learning
abstract
Most existing few-shot learning (FSL) methods require a large amount of labeled data in meta-training, which is a major limit. To reduce the requirement of labels, a semi-supervised meta-training (SSMT) setting has been proposed for FSL, which includes only a few labeled samples and numbers of unlabeled samples in base classes. However, existing methods under this setting require class-aware sample selection from the unlabeled set, which violates the assumption of unlabeled set. In this paper, we propose a practical semi-supervised meta-training setting with truly unlabeled data to facilitate the applications of FSL in realistic scenarios. To better utilize both the labeled and truly unlabeled data, we propose a simple and effective meta-training framework, called pseudo-labeling based meta-learning (PLML). Firstly, we train a classifier via common semi-supervised learning (SSL) and use it to obtain the pseudo-labels of unlabeled data. Then we build few-shot tasks from labeled and pseudo-labeled data and design a novel finetuning method with feature smoothing and noise suppression to better learn the FSL model from noise labels. Surprisingly, through extensive experiments across two FSL datasets, we find that this simple meta-training framework effectively prevents the performance degradation of various FSL models under limited labeled data, and also significantly outperforms the representative SSMT models. Besides, benefiting from meta-training, our method also improves several representative SSL algorithms as well. We provide the training code and usage examples at https://github.com/ouyangtianran/PLML.
Xingping Dong, Tianran Ouyang, Shengcai Liao, Bo Du 0001, Ling Shao 0001
IEEE Trans. Image Process.3
2023 KD-DLGAN: Data Limited Image Generation via Knowledge Distillation
abstract
Generative Adversarial Networks (GANs) rely heavily on large-scale training data for training high-quality image generation models. With limited training data, the GAN discriminator often suffers from severe overfitting which directly leads to degraded generation especially in generation diversity. Inspired by the recent advances in knowledge distillation (KD), we propose KD-DLGAN, a knowledge-distillation based generation framework that introduces pre-trained vision-language models for training effective data-limited generation models. KD-DLGAN consists of two innovative designs. The first is aggregated generative KD that mitigates the discriminator overfitting by challenging the discriminator with harder learning tasks and distilling more generalizable knowledge from the pre-trained models. The second is correlated generative KD that improves the generation diversity by distilling and preserving the diverse image-text correlation within the pre-trained models. Extensive experiments over multiple benchmarks show that KD-DLGAN achieves superior image generation with limited training data. In addition, KD-DLGAN complements the state-of-the-art with consistent and substantial performance gains. Note that codes will be released.
Kaiwen Cui, Yingchen Yu, Fangneng Zhan, Shengcai Liao, Shijian Lu, Eric P. Xing
CVPR4
2023 Energy-Based Test Sample Adaptation for Domain Generalization
Zehao Xiao, Xiantong Zhen, Shengcai Liao, Cees Snoek
ICLR3
2023 ProtoDiff: Learning to Learn Prototypical Networks by Task-Guided Diffusion
abstract
Prototype-based meta-learning has emerged as a powerful technique for addressing few-shot learning challenges. However, estimating a deterministic prototype using a simple average function from a limited number of examples remains a fragile process. To overcome this limitation, we introduce ProtoDiff, a novel framework that leverages a task-guided diffusion model during the meta-training phase to gradually generate prototypes, thereby providing efficient class representations. Specifically, a set of prototypes is optimized to achieve per-task prototype overfitting, enabling accurately obtaining the overfitted prototypes for individual tasks. Furthermore, we introduce a task-guided diffusion process within the prototype space, enabling the meta-learning of a generative process that transitions from a vanilla prototype to an overfitted prototype. ProtoDiff gradually generates task-specific prototypes from random noise during the meta-test stage, conditioned on the limited samples available for the new task. Furthermore, to expedite training and enhance ProtoDiff's performance, we propose the utilization of residual prototype learning, which leverages the sparsity of the residual prototype. We conduct thorough ablation studies to demonstrate its ability to accurately capture the underlying prototype distribution and enhance generalization. The new state-of-the-art performance on within-domain, cross-domain, and few-task few-shot classification further substantiates the benefit of ProtoDiff.
Yingjun Du, Zehao Xiao, Shengcai Liao, Cees Snoek
NeurIPS3
2023 Efficient Person Search: An Anchor-Free Approach
Yichao Yan, Jinpeng Li 0004, Jie Qin 0004, Peng Zheng 0004, Shengcai Liao, Xiaokang Yang 0001
Int. J. Comput. Vis.5
2023 Center and Scale Prediction: Anchor-free Approach for Pedestrian and Face Detection
Wei Liu 0097, Irtiza Hasan, Shengcai Liao
Pattern Recognit.3
2023 POCE: Pose-Controllable Expression Editing
abstract
Facial expression editing has attracted increasing attention with the advance of deep neural networks in recent years. However, most existing methods suffer from compromised editing fidelity and limited usability as they either ignore pose variations (unrealistic editing) or require paired training data (not easy to collect) for pose controls. This paper presents POCE, an innovative pose-controllable expression editing network that can generate realistic facial expressions and head poses simultaneously with just unpaired training images. POCE achieves the more accessible and realistic pose-controllable expression editing by mapping face images into UV space, where facial expressions and head poses can be disentangled and edited separately. POCE has two novel designs. The first is self-supervised UV completion that allows to complete UV maps sampled under different head poses, which often suffer from self-occlusions and missing facial texture. The second is weakly-supervised UV editing that allows to generate new facial expressions with minimal modification of facial identity, where the synthesized expression could be controlled by either an expression label or directly transplanted from a reference UV map via feature transfer. Extensive experiments show that POCE can learn from unpaired face images effectively, and the learned model can generate realistic and high-fidelity facial expressions under various new poses.
Rongliang Wu, Yingchen Yu, Fangneng Zhan, Shengcai Liao, Shijian Lu
IEEE Trans. Image Process.5
2022 Exploring Visual Context for Weakly Supervised Person Search
abstract
Person search has recently emerged as a challenging task that jointly addresses pedestrian detection and person re-identification. Existing approaches follow a fully supervised setting where both bounding box and identity annotations are available. However, annotating identities is labor-intensive, limiting the practicability and scalability of current frameworks. This paper inventively considers weakly supervised person search with only bounding box annotations. We propose to address this novel task by investigating three levels of context clues (i.e., detection, memory and scene) in unconstrained natural images. The first two are employed to promote local and global discriminative capabilities, while the latter enhances clustering accuracy. Despite its simple design, our CGPS boosts the baseline model by 8.8% in mAP on CUHK-SYSU. Surprisingly, it even achieves comparable performance with several supervised person search models. Our code is available at https://github. com/ljpadam/CGPS.
Yichao Yan, Jinpeng Li 0004, Shengcai Liao, Jie Qin 0004, Bingbing Ni, Ke Lu 0002, Xiaokang Yang 0001
AAAI3
2022 Graph Sampling Based Deep Metric Learning for Generalizable Person Re-Identification
abstract
Recent studies show that, both explicit deep feature matching as well as large-scale and diverse training data can significantly improve the generalization of person reidentification. However, the efficiency of learning deep matchers on large-scale data has not yet been adequately studied. Though learning with classification parameters or class memory is a popular way, it incurs large memory and computational costs. In contrast, pairwise deep metric learning within mini batches would be a better choice. However, the most popular random sampling method, the well-known PKsampler, is not informative and efficient for deep metric learning. Though online hard example mining has improved the learning efficiency to some extent, the mining in mini batches after random sampling is still limited. This inspires us to explore the use of hard example mining earlier, in the data sampling stage. To do so, in this paper, we propose an efficient mini-batch sampling method, called graph sampling (GS), for large-scale deep metric learning. The basic idea is to build a nearest neighbor relationship graph for all classes at the beginning of each epoch. Then, each mini batch is composed of a randomly selected class and its nearest neighboring classes so as to provide informative and challenging examples for learning. Together with an adapted competitive baseline, we improve the state of the art in generalizable person re-identification significantly, by 25.1% in Rank-1 on MSMT17 when trained on RandPerson. Besides, the proposed method also outperforms the competitive baseline, by 6.8% in Rank-1 on CUHK03-NP when trained on MSMT17. Meanwhile, the training time is significantly reduced, from 25.4 hours to 2 hours when trained on RandPerson with 8,000 identities. Code is available at https://github.com/ShengcaiLiao/QAConv.
Shengcai Liao, Ling Shao 0001
CVPR1
2022 Cloning Outfits from Real-World Images to 3D Characters for Generalizable Person Re-Identification
abstract
Recently, large-scale synthetic datasets are shown to be very useful for generalizable person re-identification. However, synthesized persons in existing datasets are mostly cartoon-like and in random dress collocation, which limits their performance. To address this, in this work, an automatic approach is proposed to directly clone the whole outfits from real-world person images to virtual 3D characters, such that any virtual person thus created will appear very similar to its real-world counterpart. Specifically, based on UV texture mapping, two cloning methods are designed, namely registered clothes mapping and homogeneous cloth expansion. Given clothes keypoints detected on person images and labeled on regular UV maps with clear clothes structures, registered mapping applies perspective homography to warp real-world clothes to the counterparts on the UV map. As for invisible clothes parts and irregular UV maps, homogeneous expansion segments a homogeneous area on clothes as a realistic cloth pattern or cell, and expand the cell to fill the UV map. Furthermore, a similarity-diversity expansion strategy is proposed, by clustering person images, sampling images per cluster, and cloning outfits for 3D character generation. This way, virtual persons can be scaled up densely in visual similarity to challenge model learning, and diversely in population to enrich sample distribution. Finally, by rendering the cloned characters in Unity3D scenes, a more realistic virtual dataset called ClonedPerson is created, with 5,621 identities and 887,766 images. Experimental results show that the model trained on ClonedPerson has a better generalization performance, superior to that trained on other popular real-world and synthetic person re-identification datasets. The ClonedPerson project is available at https://github.com/Yanan-Wang-cs/ClonedPerson.
Yanan Wang 0009, Xuezhi Liang, Shengcai Liao
CVPR3
2022 RePFormer: Refinement Pyramid Transformer for Robust Facial Landmark Detection
abstract
This paper presents a Refinement Pyramid Transformer (RePFormer) for robust facial landmark detection. Most facial landmark detectors focus on learning representative image features. However, these CNN-based feature representations are not robust enough to handle complex real-world scenarios due to ignoring the internal structure of landmarks, as well as the relations between landmarks and context. In this work, we formulate the facial landmark detection task as refining landmark queries along pyramid memories. Specifically, a pyramid transformer head (PTH) is introduced to build both homologous relations among landmarks and heterologous relations between landmarks and cross-scale contexts. Besides, a dynamic landmark refinement (DLR) module is designed to decompose the landmark regression into an end-to-end refinement procedure, where the dynamically aggregated queries are transformed to residual coordinates predictions. Extensive experimental results on four facial landmark detection benchmarks and their various subsets demonstrate the superior performance and high robustness of our framework.
Jinpeng Li 0004, Haibo Jin, Shengcai Liao, Ling Shao 0001, Pheng-Ann Heng
IJCAI3
2022 Masked Generative Adversarial Networks are Data-Efficient Generation Learners
abstract
This paper shows that masked generative adversarial network (MaskedGAN) is robust image generation learners with limited training data. The idea of MaskedGAN is simple: it randomly masks out certain image information for effective GAN training with limited data. We develop two masking strategies that work along orthogonal dimensions of training images, including a shifted spatial masking that masks the images in spatial dimensions with random shifts, and a balanced spectral masking that masks certain image spectral bands with self-adaptive probabilities. The two masking strategies complement each other which together encourage more challenging holistic learning from limited training data, ultimately suppressing trivial solutions and failures in GAN training. Albeit simple, extensive experiments show that MaskedGAN achieves superior performance consistently across different network architectures (e.g., CNNs including BigGAN and StyleGAN-v2 and Transformers including TransGAN and GANformer) and datasets (e.g., CIFAR-10, CIFAR-100, ImageNet, 100-shot, AFHQ, FFHQ and Cityscapes).
Jiaxing Huang 0001, Kaiwen Cui, Dayan Guan, Aoran Xiao, Fangneng Zhan, Shijian Lu, Shengcai Liao, Eric P. Xing
NeurIPS7
2022 Urban scene based Semantical Modulation for Pedestrian Detection
Hangzhi Jiang, Shengcai Liao, Jinpeng Li 0004, Véronique Prinet, Shiming Xiang
Neurocomputing2
2022 Attentive WaveBlock: Complementarity-Enhanced Mutual Networks for Unsupervised Domain Adaptation in Person Re-Identification and Beyond
abstract
Unsupervised domain adaptation (UDA) for person re-identification is challenging because of the huge gap between the source and target domain. A typical self-training method is to use pseudo-labels generated by clustering algorithms to iteratively optimize the model on the target domain. However, a drawback to this is that noisy pseudo-labels generally cause trouble in learning. To address this problem, a mutual learning method by dual networks has been developed to produce reliable soft labels. However, as the two neural networks gradually converge, their complementarity is weakened and they likely become biased towards the same kind of noise. This paper proposes a novel light-weight module, the Attentive WaveBlock (AWB), which can be integrated into the dual networks of mutual learning to enhance the complementarity and further depress noise in the pseudo-labels. Specifically, we first introduce a parameter-free module, the WaveBlock, which creates a difference between features learned by two networks by waving blocks of feature maps differently. Then, an attention mechanism is leveraged to enlarge the difference created and discover more complementary features. Furthermore, two kinds of combination strategies, i.e. pre-attention and post-attention, are explored. Experiments demonstrate that the proposed method achieves state-of-the-art performance with significant improvements on multiple UDA person re-identification tasks. We also prove the generality of the proposed method by applying it to vehicle re-identification and image classification tasks. Our codes and models are available at: AWB.
Fang Zhao 0006, Shengcai Liao, Ling Shao 0001
IEEE Trans. Image Process.3
2021 DomainMix: Learning Generalizable Person Re-Identification Without Human Annotations
Shengcai Liao, Fang Zhao 0006, Cuicui Kang, Ling Shao 0001
BMVC2
2021 Generalizable Pedestrian Detection: The Elephant in the Room
abstract
Pedestrian detection is used in many vision based applications ranging from video surveillance to autonomous driving. Despite achieving high performance, it is still largely unknown how well existing detectors generalize to unseen data. This is important because a practical detector should be ready to use in various scenarios in applications. To this end, we conduct a comprehensive study in this paper, using a general principle of direct cross-dataset evaluation. Through this study, we find that existing state-of-the-art pedestrian detectors, though perform quite well when trained and tested on the same dataset, generalize poorly in cross dataset evaluation. We demonstrate that there are two reasons for this trend. Firstly, their designs (e.g. anchor settings) may be biased towards popular benchmarks in the traditional single-dataset training and test pipeline, but as a result largely limit their generalization capability. Secondly, the training source is generally not dense in pedestrians and diverse in scenarios. Under direct cross-dataset evaluation, surprisingly, we find that a general purpose object detector, without pedestrian-tailored adaptation in design, generalizes much better compared to existing state-of-the-art pedestrian detectors. Furthermore, we illustrate that diverse and dense datasets, collected by crawling the web, serve to be an efficient source of pre-training for pedestrian detection. Accordingly, we propose a progressive training pipeline and find that it works well for autonomous-driving oriented pedestrian detection. Consequently, the study conducted in this paper suggests that more emphasis should be put on cross-dataset evaluation for the future design of generalizable pedestrian detectors. Code and models can be accessed at https://github.com/hasanirtiza/Pedestron.
Irtiza Hasan, Shengcai Liao, Jinpeng Li 0004, Saad Ullah Akram, Ling Shao 0001
CVPR2
2021 Anchor-Free Person Search
abstract
Person search aims to simultaneously localize and identify a query person from realistic, uncropped images, which can be regarded as the unified task of pedestrian detection and person re-identification (re-id). Most existing works employ two-stage detectors like Faster-RCNN, yielding encouraging accuracy but with high computational overhead. In this work, we present the Feature-Aligned Person Search Network (AlignPS), the first anchor-free framework to efficiently tackle this challenging task. AlignPS explicitly addresses the major challenges, which we summarize as the misalignment issues in different levels (i.e., scale, region, and task), when accommodating an anchor-free detector for this task. More specifically, we propose an aligned feature aggregation module to generate more discriminative and robust feature embeddings by following a "re-id first" principle. Such a simple design directly improves the baseline anchor-free model on CUHK-SYSU by more than 20% in mAP. Moreover, AlignPS outperforms state-of-the-art two-stage methods, with a higher speed. The code is available at https://github.com/daodaofr/AlignPS.
Yichao Yan, Jinpeng Li 0004, Jie Qin 0004, Song Bai 0001, Shengcai Liao, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001
CVPR5
2021 Learning Anchored Unsigned Distance Functions with Gradient Direction Alignment for Single-view Garment Reconstruction
abstract
While single-view 3D reconstruction has made significant progress benefiting from deep shape representations in recent years, garment reconstruction is still not solved well due to open surfaces, diverse topologies and complex geometric details. In this paper, we propose a novel learn-able Anchored Unsigned Distance Function (AnchorUDF) representation for 3D garment reconstruction from a single image. AnchorUDF represents 3D shapes by predicting unsigned distance fields (UDFs) to enable open garment surface modeling at arbitrary resolution. To capture diverse garment topologies, AnchorUDF not only computes pixel-aligned local image features of query points, but also leverages a set of anchor points located around the surface to enrich 3D position features for query points, which provides stronger 3D space context for the distance function. Furthermore, in order to obtain more accurate point projection direction at inference, we explicitly align the spatial gradient direction of AnchorUDF with the ground-truth direction to the surface during training. Extensive experiments on two public 3D garment datasets, i.e., MGN and Deep Fashion3D, demonstrate that AnchorUDF achieves the state-of-the-art performance on single-view garment reconstruction. Code is available at https://github.com/zhaofang0627/AnchorUDF.
Fang Zhao 0006, Shengcai Liao, Ling Shao 0001
ICCV3
2021 TransMatcher: Deep Image Matching Through Transformers for Generalizable Person Re-identification
abstract
Transformers have recently gained increasing attention in computer vision. However, existing studies mostly use Transformers for feature representation learning, e.g. for image classification and dense predictions, and the generalizability of Transformers is unknown. In this work, we further investigate the possibility of applying Transformers for image matching and metric learning given pairs of images. We find that the Vision Transformer (ViT) and the vanilla Transformer with decoders are not adequate for image matching due to their lack of image-to-image attention. Thus, we further design two naive solutions, i.e. query-gallery concatenation in ViT, and query-gallery cross-attention in the vanilla Transformer. The latter improves the performance, but it is still limited. This implies that the attention mechanism in Transformers is primarily designed for global feature aggregation, which is not naturally suitable for image matching. Accordingly, we propose a new simplified decoder, which drops the full attention implementation with the softmax weighting, keeping only the query-key similarity computation. Additionally, global max pooling and a multilayer perceptron (MLP) head are applied to decode the matching result. This way, the simplified decoder is computationally more efficient, while at the same time more effective for image matching. The proposed method, called TransMatcher, achieves state-of-the-art performance in generalizable person re-identification, with up to 6.1% and 5.7% performance gains in Rank-1 and mAP, respectively, on several popular datasets. Code is available at https://github.com/ShengcaiLiao/QAConv.
Shengcai Liao, Ling Shao 0001
NeurIPS1
2021 Pixel-in-Pixel Net: Towards Efficient Facial Landmark Detection in the Wild
Haibo Jin, Shengcai Liao, Ling Shao 0001
Int. J. Comput. Vis.2
2021 AFAN: Augmented Feature Alignment Network for Cross-Domain Object Detection
abstract
Unsupervised domain adaptation for object detection is a challenging problem with many real-world applications. Unfortunately, it has received much less attention than supervised object detection. Models that try to address this task tend to suffer from a shortage of annotated training samples. Moreover, existing methods of feature alignments are not sufficient to learn domain-invariant representations. To address these limitations, we propose a novel augmented feature alignment network (AFAN) which integrates intermediate domain image generation and domain-adversarial training into a unified framework. An intermediate domain image generator is proposed to enhance feature alignments by domain-adversarial training with automatically generated soft domain labels. The synthetic intermediate domain images progressively bridge the domain divergence and augment the annotated source domain training data. A feature pyramid alignment is designed and the corresponding feature discriminator is used to align multi-scale convolutional features of different semantic levels. Last but not least, we introduce a region feature alignment and an instance discriminator to learn domain-invariant features for object proposals. Our approach significantly outperforms the state-of-the-art methods on standard benchmarks for both similar and dissimilar domain adaptations. Further extensive experiments verify the effectiveness of each component and demonstrate that the proposed network can learn domain-invariant representations.
Hongsong Wang 0001, Shengcai Liao, Ling Shao 0001
IEEE Trans. Image Process.2
2020 Pixel-Aware Deep Function-Mixture Network for Spectral Super-Resolution
abstract
Spectral super-resolution (SSR) aims at generating a hyperspectral image (HSI) from a given RGB image. Recently, a promising direction is to learn a complicated mapping function from the RGB image to the HSI counterpart using a deep convolutional neural network. This essentially involves mapping the RGB context within a size-specific receptive field centered at each pixel to its spectrum in the HSI. The focus thereon is to appropriately determine the receptive field size and establish the mapping function from RGB context to the corresponding spectrum. Due to their differences in category or spatial position, pixels in HSIs often require different-sized receptive fields and distinct mapping functions. However, few efforts have been invested to explicitly exploit this prior.To address this problem, we propose a pixel-aware deep function-mixture network for SSR, which is composed of a new class of modules, termed function-mixture (FM) blocks. Each FM block is equipped with some basis functions, i.e., parallel subnets of different-sized receptive fields. Besides, it incorporates an extra subnet as a mixing function to generate pixel-wise weights, and then linearly mixes the outputs of all basis functions with those generated weights. This enables us to pixel-wisely determine the receptive field size and the mapping function. Moreover, we stack several such FM blocks to further increase the flexibility of the network in learning the pixel-wise mapping. To encourage feature reuse, intermediate features generated by the FM blocks are fused in late stage, which proves to be effective for boosting the SSR performance. Experimental results on three benchmark HSI datasets demonstrate the superiority of the proposed method.
Lei Zhang 0054, Zhiqiang Lang, Peng Wang 0023, Wei Wei 0008, Shengcai Liao, Ling Shao 0001, Yanning Zhang 0001
AAAI5
2020 Unsupervised Adaptation Learning for Hyperspectral Imagery Super-Resolution
abstract
The key for fusion based hyperspectral image (HSI) super-resolution (SR) is to infer the posteriori of a latent HSI using appropriate image prior and likelihood that depends on degeneration. However, in practice the priors of high-dimensional HSIs can be extremely complicated and the degeneration is often unknown. Consequently most existing approaches that assume a shallow hand-crafted image prior and a pre-defined degeneration, fail to well generalize in real applications. To tackle this problem, we present an unsupervised adaptation learning (UAL) framework. Instead of directly modelling the complicated image prior, we propose to first implicitly learn a general image prior using deep networks and then adapt it to a specific HSI. Following this idea, we develop a two-stage SR network that leverages two consecutive modules: a fusion module and an adaptation module, to recover the latent HSI in a coarse-to-fine scheme. The fusion module is pretrained in a supervised manner on synthetic data to capture a spatial-spectral prior that is general across most HSIs. To adapt the learned general prior to the specific HSI under unknown degeneration, we introduce a simple degeneration network to assist learning both the adaptation module and the degeneration in an unsupervised way. In this way, the resultant image-specific prior and the estimated degeneration can benefit the inference of a more accurate posteriori, thereby increasing generalization capacity. To verify the efficacy of UAL, we extensively evaluate it on four benchmark datasets and report strong results that surpass existing approaches.
Lei Zhang 0054, Jiangtao Nie, Wei Wei 0008, Yanning Zhang 0001, Shengcai Liao, Ling Shao 0001
CVPR5
2020 Interpretable and Generalizable Person Re-identification with Query-Adaptive Convolution and Temporal Lifting
Shengcai Liao, Ling Shao 0001
ECCV (11)1
2020 Exclusivity-Consistency Regularized Knowledge Distillation for Face Recognition
Xiaobo Wang 0001, Tianyu Fu 0001, Shengcai Liao, Zhen Lei 0001, Tao Mei 0001
ECCV (24)3
2020 Unsupervised Domain Adaptation with Noise Resistible Mutual-Training for Person Re-identification
Fang Zhao 0006, Shengcai Liao, Guosen Xie, Jian Zhao 0006, Kaihao Zhang, Ling Shao 0001
ECCV (11)2
2020 Box Guided Convolution for Pedestrian Detection
abstract
Occlusions, scale variation and numerous false positives still represent fundamental challenges in pedestrian detection. Intuitively, different sizes of receptive fields and more attention to the visible parts are required for detecting pedestrians with various scales and occlusion levels, respectively. However, these challenges have not been addressed well by existing pedestrian detectors. This paper presents a novel convolutional network, denoted as box guided convolution network (BGCNet), to tackle these challenges simultaneously in a unified framework. In particular, we proposed a box guided convolution (BGC) that can dynamically adjust the sizes of convolution kernels guided by the predicted bounding boxes. In this way, BGCNet provides position-aware receptive fields to address the challenge of large variations of scales. In addition, for the issue of heavy occlusion, the kernel parameters of BGC are spatially localized around the salient and mostly visible key points of a pedestrian, such as the head and foot, to effectively capture high-level semantic features to help detection. Furthermore, a local maximum (LM) loss is introduced to depress false positives and highlight true positives by forcing positives, rather than negatives, as local maximums, without any additional inference burden. We evaluate BGCNet on popular pedestrian detection benchmarks, and achieve the state-of-the-art results, with the significant performance improvement on heavily occluded and small-scale pedestrians.
Jinpeng Li 0004, Shengcai Liao, Hangzhi Jiang, Ling Shao 0001
ACM Multimedia2
2020 Surpassing Real-World Source Training Data: Random 3D Characters for Generalizable Person Re-Identification
abstract
Person re-identification has seen significant advancement in recent years. However, the ability of learned models to generalize to unknown target domains still remains limited. One possible reason for this is the lack of large-scale and diverse source training data, since manually labeling such a dataset is very expensive and privacy sensitive. To address this, we propose to automatically synthesize a large-scale person re-identification dataset following a set-up similar to real surveillance but with virtual environments, and then use the synthesized person images to train a generalizable person re-identification model. Specifically, we design a method to generate a large number of random UV texture maps and use them to create different 3D clothing models. Then, an automatic code is developed to randomly generate various different 3D characters with diverse clothes, races and attributes. Next, we simulate a number of different virtual environments using Unity3D, with customized camera networks similar to real surveillance systems, and import multiple 3D characters at the same time, with various movements and interactions along different paths through the camera networks. As a result, we obtain a virtual dataset, called RandPerson, with 1,801,816 person images of 8,000 identities. By training person re-identification models on these synthesized person images, we demonstrate, for the first time, that models trained on virtual data can generalize well to unseen target images, surpassing the models trained on various real-world datasets, including CUHK03, Market-1501, DukeMTMC-reID, and almost MSMT17. The RandPerson dataset is available at https://github.com/VideoObjectSearch/RandPerson.
Yanan Wang 0009, Shengcai Liao, Ling Shao 0001
ACM Multimedia2
2020 Human Parsing Based Texture Transfer from Single Image to 3D Human via Cross-View Consistency
abstract
This paper proposes a human parsing based texture transfer model via cross-view consistency learning to generate the texture of 3D human body from a single image. We use the semantic parsing of human body as input for providing both the shape and pose information to reduce the appearance variation of human image and preserve the spatial distribution of semantic parts. Meanwhile, in order to improve the prediction for textures of invisible parts, we explicitly enforce the consistency across different views of the same subject by exchanging the textures predicted by two views to render images during training. The perception loss and total variation regularization are optimized to maximize the similarity between rendered and input images, which does not necessitate extra 3D texture supervision. Experimental results on pedestrian images and fashion photos demonstrate that our method can produce higher quality textures with convincing details than other texture generation methods.
Fang Zhao 0006, Shengcai Liao, Kaihao Zhang, Ling Shao 0001
NeurIPS2
2020 Efficient Single-Stage Pedestrian Detector by Asymptotic Localization Fitting and Multi-Scale Context Encoding
abstract
Though Faster R-CNN based two-stage detectors have witnessed significant boost in pedestrian detection accuracy, they are still slow for practical applications. One solution is to simplify this working flow as a single-stage detector. However, current single-stage detectors (e.g. SSD) have not presented competitive accuracy on common pedestrian detection benchmarks. Accordingly, a structurally simple but effective module called Asymptotic Localization Fitting (ALF) is proposed, which stacks a series of predictors to directly evolve the default anchor boxes of SSD step by step to improve detection results. Additionally, combining the advantages from residual learning and multi-scale context encoding, a bottleneck block is proposed to enhance the predictors' discriminative power. On top of the above designs, an efficient single-stage detection architecture is designed, resulting in an attractive pedestrian detector in both accuracy and speed. A comprehensive set of experiments on two of the largest pedestrian detection datasets (i.e. CityPersons and Caltech) demonstrate the superiority of the proposed method, comparing to the state of the arts on both the benchmarks.
Wei Liu 0097, Shengcai Liao, Weidong Hu
IEEE Trans. Image Process.2
2020 Vehicle Re-Identification Using Quadruple Directional Deep Learning Features
abstract
In order to resist the adverse effect of viewpoint variations, we design quadruple directional deep learning networks to extract quadruple directional deep learning features (QD-DLF) of vehicle images for improving vehicle re-identification performance. The quadruple directional deep learning networks are of similar overall architecture, including the same basic deep learning architecture but different directional feature pooling layers. Specifically, the same basic deep learning architecture that is a shortly and densely connected convolutional neural network is utilized to extract the basic feature maps of an input square vehicle image in the first stage. Then, the quadruple directional deep learning networks utilize different directional pooling layers, i.e., horizontal average pooling layer, vertical average pooling layer, diagonal average pooling layer, and anti-diagonal average pooling layer, to compress the basic feature maps into horizontal, vertical, diagonal, and anti-diagonal directional feature maps, respectively. Finally, these directional feature maps are spatially normalized and concatenated together as a quadruple directional deep learning feature for vehicle re-identification. The extensive experiments on both VeRi and VehicleID databases show that the proposed QD-DLF approach outperforms multiple state-of-the-art vehicle re-identification methods.
Jianqing Zhu, Huanqiang Zeng, Jingchang Huang, Shengcai Liao, Zhen Lei 0001, Canhui Cai, Lixin Zheng
IEEE Trans. Intell. Transp. Syst.4
2019 High-Level Semantic Feature Detection: A New Perspective for Pedestrian Detection
abstract
Object detection generally requires sliding-window classifiers in tradition or anchor-based predictions in modern deep learning approaches. However, either of these approaches requires tedious configurations in windows or anchors. In this paper, taking pedestrian detection as an example, we provide a new perspective where detecting objects is motivated as a high-level semantic feature detection task. Like edges, corners, blobs and other feature detectors, the proposed detector scans for feature points all over the image, for which the convolution is naturally suited. However, unlike these traditional low-level features, the proposed detector goes for a higher-level abstraction, that is, we are looking for central points where there are pedestrians, and modern deep models are already capable of such a high-level semantic abstraction. Besides, like blob detection, we also predict the scales of the pedestrian points, which is also a straightforward convolution. Therefore, in this paper, pedestrian detection is simplified as a straightforward center and scale prediction task through convolutions. This way, the proposed method enjoys an anchor-free setting. Though structurally simple, it presents competitive accuracy and good speed on challenging pedestrian detection benchmarks, and hence leading to a new attractive pedestrian detector. Code and models will be available at https://github.com/liuwei16/CSP.
Wei Liu 0097, Shengcai Liao, Weiqiang Ren, Weidong Hu, Yinan Yu
CVPR2
2019 Unsupervised Graph Association for Person Re-Identification
abstract
In this paper, we propose an unsupervised graph association (UGA) framework to learn the underlying view-invariant representations from the video pedestrian tracklets. The core points of UGA are mining the underlying cross-view associations and reducing the damage of noise associations. To this end, UGA is adopts a two-stage training strategy: (1) intra-camera learning stage and (2) intercamera learning stage. The former learns the intra-camera representation for each camera. While the latter builds a cross-view graph (CVG) to associate different cameras. By doing this, we can learn view-invariant representation for all person. Extensive experiments and ablation studies on seven re-id datasets demonstrate the superiority of the proposed UGA over most state-of-the-art unsupervised and domain adaptation re-id methods.
Jinlin Wu, Yang Yang 0062, Zhen Lei 0001, Shengcai Liao, Stan Z. Li
ICCV5
2019 Clustering and Dynamic Sampling Based Unsupervised Domain Adaptation for Person Re-Identification
abstract
Person Re-Identification (Re-ID) has witnessed great improvements due to the advances of the deep convolutional neural networks (CNN). Despite this, existing methods mainly suffer from the poor generalization ability to unseen scenes because of the different characteristics between different domains. To address this issue, a Clustering and Dynamic Sampling (CDS) method is proposed in this paper, which tries to transfer the useful knowledge of existing labeled source domain to the unlabeled target one. Specifically, to improve the discriminability of CNN model on source domain, we use the commonly shared pedestrian attributes (e.g., gender, hat and clothing color etc.) to enrich the information and resort to the margin-based softmax (e.g., A-Softmax) loss to train the model. For the unlabeled target domain, we iteratively cluster the samples into several centers and dynamically select informative ones from each center to fine-tune the source-domain model. Extensive experiments on DukeMTMC-reID and Market-1501 datasets show that the proposed method greatly improves the state of the arts in unsupervised domain adaptation.
Jinlin Wu, Shengcai Liao, Zhen Lei 0001, Xiaobo Wang 0001, Yang Yang 0062, Stan Z. Li
ICME2
2019 Towards accurate tiny vehicle detection in complex scenes
Wei Liu 0097, Shengcai Liao, Weidong Hu
Neurocomputing2
2019 Perceiving Motion From Dynamic Memory for Vehicle Detection in Surveillance Videos
abstract
Most existing video-based object detection methods utilize successful image-based object detector as a base network, and additionally exploit temporal information with either bounding-box post-processing or feature enhancement from multiple frames. However, little work has been done on directly modeling temporal motion in an efficient way for detection in surveillance videos. In this paper, a simple but effective module, denoted as motion-from-memory (MFM), is proposed to encode temporal context for improved detection in surveillance videos. With appearance features extracted from a base CNN, the MFM module maintains a dynamic memory for each input sequence and output motion features on each frame. This module costs minor additional model parameters and computations, but is very helpful for moving object detection, especially in surveillance videos. Thanks to the additional MFM module, the performance of a light-weight MobileNet-based Faster RCNN detector is boosted by 13.93% in mAP, achieving comparable performance to that of strong ResNet-50-based. When MFM is integrated into an even weaker but faster single-stage detector, it ranks the second best one among all published works when submitted to the DEETRAC vehicle detection benchmark, with 69.10% mAP, compared to 69.87% of the best one. However, when running speed is considered, the proposed method is the fastest one, running at 33 FPS with 540×960 surveillance videos on a moderate commercial GPU (NVIDIA GTX 1080Ti), which is about 3 times faster than the second fastest one.
Wei Liu 0097, Shengcai Liao, Weidong Hu
IEEE Trans. Circuits Syst. Video Technol.2
2018 Is Re-ranking Useful for Open-set Person Re-identification?
abstract
Re-ranking algorithms can often boost the performance of close-set person re-identification. However, limited efforts have been devoted to answering whether a similar conclusion could be derived on open-set person re-identification. Considering that open-set scenario is more practical in real applications, in this paper, we try to answer this question and do a benchmark study of re-ranking on open-set person re-identification. Specifically, we evaluate three feature descriptors, namely MB-LBP, LOMO, and IDE, and four distance metrics, namely Euclidean, Cosine, RRDA, and XQDA, with their combinations as baseline algorithms. Then, we evaluate four popular re-ranking algorithms, including k-reciprocal Encoding, ECN-3, ECN-4, and DaF. Through extensive benchmark studies on the OPeRIDv1.0 dataset, the results show that re-ranking algorithms, though useful for closed-set person re-identification, are not generally effective for the open-set person re-identification problem. We argue that this is because re-ranking algorithms change the score distributions per query, and hence disrupt the FAR estimation across all queries. Accordingly, we propose to align the re-ranking scores to the original score via the min-max normalization, which verifies our hypothesis above.
Hongsheng Wang, Shengcai Liao, Zhen Lei 0001, Yang Yang 0062
IEEE BigData2
2018 Semi-automatic Data Annotation Tool for Person Re-identification Across Multi Cameras
abstract
Person re-identification is an important technique towards automatic search of a person's presence in a surveillance video. It is becoming a hot research topic due to its value in both machine learning research and video surveillance applications. Considering the current success of deep learning, having tons of person images with identity labels are important and helpful for learning effective person matchers. However, collecting labeled images for person re-identification is more difficult than other similar tasks such as face recognition due to complex intra-class variations in illumination, pose, viewpoint, blur, low resolution, and occlusion. Although the volume of surveillance videos has become larger and larger today, it is time-consuming and costs lots of human labors in labeling a large dataset for person re-identification. In this paper, we propose a semi-automatic data annotation tool to accelerate annotation of person images across multi cameras. This tool consists of automatic person detection and tracking algorithms for person image collection, and an ad-hoc person matcher for automatic person matching suggestions across multi cameras. Moreover, we further utilize background and video sequence information for identity confirmation during annotation, which is also a good intuition for the future design of person re-identification algorithms.
Shengcai Liao, Zhen Lei 0001
IEEE BigData2
2018 Learning Efficient Single-Stage Pedestrian Detectors by Asymptotic Localization Fitting
Wei Liu 0097, Shengcai Liao, Weidong Hu, Xuezhi Liang, Xiao Chen 0010
ECCV (14)2
2018 Deep Background Subtraction with Guided Learning
abstract
Recently, convolutional neural networks (CNNs) have been applied in background subtraction (change detection) and gained notable improvements. Two typical methods have been proposed. The first one learns a specific CNN model for each video, but requires manual labeling of training frames on the fly. The other one learns a universal model offline, however, limits its performance in handling various surveillance scenarios. To address these problems, in this paper, a new deep background subtraction method is proposed by introducing a guided learning strategy. The main idea is to learn a specific CNN model for each video to ensure accuracy, but manage to avoid manual labeling. To achieve this, firstly we apply the SubSENSE algorithm [1] to get an initial segmentation, and then an adaptive strategy is designed to select reliable pixels to guide the CNN training. Besides, we also design a simple strategy to automatically select informative frames for guided learning. Experiments on the largest background subtraction benchmark CDnet2014 show that the proposed guided deep learning method outperforms existing state of the arts.
Xuezhi Liang, Shengcai Liao, Xiaobo Wang 0001, Wei Liu 0097, Stan Z. Li
ICME2
2018 Improving Tiny Vehicle Detection in Complex Scenes
abstract
Vehicle detection is still a challenge in complex traffic scenes, especially for vehicles of tiny scales. Though RCNN based two-stage detectors have demonstrated considerably good performance, less attention has been paid to the quality of the first stage, where, however, tiny vehicles are very likely to be missed. In this paper, we propose a deep network for accurate vehicle detection, with the main idea of using a relatively large feature map for proposal generation, and keeping ROI feature's spatial layout to represent and detect tiny vehicles. However, large feature maps in lower levels of a deep network generally contain limited discriminant information. To address this, we introduce a backward feature enhancement operation, which absorbs higher level information step by step to enhance the base feature map. By doing so, even with only 100 proposals, the resulting proposal network achieves an encouraging recall over 99%. Furthermore, unlike a common practice which flatten features after ROI pooling, we argue that for a better detection of tiny vehicles, the spatial layout of the ROI features should be preserved and fully integrated. Accordingly, we use a multi-path light-weight processing chain to effectively integrate ROI features, while preserving the spatial layouts. Experiments done on the challenging DETRAC vehicle detection benchmark show that the proposed method largely improves a competitive baseline (ResNet50 based Faster RCNN) by 16.5% mAP, and it outperforms all previously published and unpublished results.
Wei Liu 0097, Shengcai Liao, Weidong Hu, Xuezhi Liang
ICME2
2018 A Shortly and Densely Connected Convolutional Neural Network for Vehicle Re-identification
abstract
In this paper, we propose a shortly and densely connected convolutional neural network (SDC-CNN) for vehicle re-identification. The proposed SDC-CNN mainly consists of short and dense units (SDUs), necessary pooling and normalization layers. The main contribution lies at the design of short and dense connection mechanism, which would effectively improve the feature learning ability. Specifically, in the proposed short and dense connection mechanism, each SDU contains a short list of densely connected convolutional layers and each convolutional layer is of the same appropriate channels. Consequently, the number of connections and the input channel of each convolutional layer are limited in each SDU, and the architecture of SDC-CNN is simple. Extensive experiments on both VeRi and VehicleID datasets show that the proposed SDC-CNN is obviously superior to multiple state-of-the-art vehicle re-identification methods.
Jianqing Zhu, Huanqiang Zeng, Zhen Lei 0001, Shengcai Liao, Lixin Zheng, Canhui Cai
ICPR4
2018 Perceptual hash-based feature description for person re-identification
Hai-Miao Hu, Zihao Hu, Shengcai Liao, Bo Li 0006
Neurocomputing4
2018 Dependence-Aware Feature Coding for Person Re-Identification
abstract
In this letter, we focus on how to boost the performance of person re-identification by exploring the discriminative information among person pairs. A novel dependence-aware feature coding framework is proposed for this task. Specifically, we employ the Hilbert–Schmidt independence criterion as the discriminative term, which is to explore the dependence between different kinds of person pairs, i.e., the same person pairs should be dependence maximized, while the different ones should be dependence minimized. Theoretical discussion and analysis on the convexity of the proposed constraint, as well as the convergence of our algorithm, are provided. Experimental results on two benchmark datasets have demonstrated the advantages of our method over the state-of-the-art alternatives.
Xiaobo Wang 0001, Zhen Lei 0001, Shengcai Liao, Xiaojie Guo 0001, Yang Yang 0062, Stan Z. Li
IEEE Signal Process. Lett.3
2018 Deep Hybrid Similarity Learning for Person Re-Identification
abstract
Person re-identification (Re-ID) aims to match person images captured from two non-overlapping cameras. In this paper, a deep hybrid similarity learning (DHSL) method for person Re-ID based on a convolution neural network (CNN) is proposed. In our approach, a light CNN learning feature pair for the input image pair is simultaneously extracted. Then, both the elementwise absolute difference and multiplication of the CNN learning feature pair are calculated. Finally, a hybrid similarity function is designed to measure the similarity between the feature pair, which is realized by learning a group of weight coefficients to project the elementwise absolute difference and multiplication into a similarity score. Consequently, the proposed DHSL method is able to reasonably assign complexities of feature learning and metric learning in a CNN, so that the performance of person Re-ID is improved. Experiments on three challenging person Re-ID databases, QMUL GRID, VIPeR, and CUHK03, illustrate that the proposed DHSL method is superior to multiple state-of-the-art person Re-ID methods.
Jianqing Zhu, Huanqiang Zeng, Shengcai Liao, Zhen Lei 0001, Canhui Cai, Lixin Zheng
IEEE Trans. Circuits Syst. Video Technol.3
2017 Learning Deep Semantic Embeddings for Cross-Modal Retrieval
abstract
Deep learning methods have been actively researched for cross-modal retrieval, with the softmax cross-entropy loss commonly applied for supervised learning. However, the softmax cross-entropy loss is known to result in large intra-class variances, which is not not very suited for cross-modal matching. In this paper, a deep architecture called Deep Semantic Embedding (DSE) is proposed, which is trained in an end-to-end manner for image-text cross-modal retrieval. With images and texts mapped to a feature embedding space, class labels are used to guide the embedding learning, so that the embedding space has a semantic meaning common for both images and texts. This way, the difference between different modalities is eliminated. Under this framework, the center loss is introduced beyond the commonly used softmax cross-entropy loss to achieve both inter-class separation and intra-class compactness. Besides, a distance based softmax cross-entropy loss is proposed to jointly consider the softmax cross-entropy and center losses in fully gradient based learning. Experiments have been done on three popular image-text cross-modal retrieval databases, showing that the proposed algorithms have achieved the best overall performances.
Cuicui Kang, Shengcai Liao, Zhen Li 0011, Zigang Cao, Gang Xiong 0001
ACML2
2017 Deep person re-identification with improved embedding and efficient training
abstract
Person re-identification task has been greatly boosted by deep convolutional neural networks (CNNs) in recent years. The core of which is to enlarge the inter-class distinction as well as reduce the intra-class variance. However, to achieve this, existing deep models prefer to adopt image pairs or triplets to form verification loss, which is inefficient and unstable since the number of training pairs or triplets grows rapidly as the number of training data grows. Moreover, their performance is limited since they ignore the fact that different dimension of embedding may play different importance. In this paper, we propose to employ identification loss with center loss to train a deep model for person re-identification. The training process is efficient since it does not require image pairs or triplets for training while the inter-class distinction and intra-class variance are well handled. To boost the performance, a new feature reweighting (FRW) layer is designed to explicitly emphasize the importance of each embedding dimension, thus leading to an improved embedding. Experiments1on several benchmark datasets have shown the superiority of our method over the state-of-the-art alternatives on both accuracy and speed.
Haibo Jin, Xiaobo Wang 0001, Shengcai Liao, Stan Z. Li
IJCB3
2017 Soft-Margin Softmax for Deep Classification
Xuezhi Liang, Xiaobo Wang 0001, Zhen Lei 0001, Shengcai Liao, Stan Z. Li
ICONIP (2)4
2017 Multi-label convolutional neural network based pedestrian attribute classification
Jianqing Zhu, Shengcai Liao, Zhen Lei 0001, Stan Z. Li
Image Vis. Comput.2
2016 Large Scale Similarity Learning Using Similar Pairs for Person Verification
abstract
In this paper, we propose a novel similarity measure and then introduce an efficient strategy to learn it by using only similar pairs for person verification. Unlike existing metric learning methods, we consider both the difference and commonness of an image pair to increase its discriminativeness. Under a pairconstrained Gaussian assumption, we show how to obtain the Gaussian priors (i.e., corresponding covariance matrices) of dissimilar pairs from those of similar pairs. The application of a log likelihood ratio makes the learning process simple and fast and thus scalable to large datasets. Additionally, our method is able to handle heterogeneous data well. Results on the challenging datasets of face verification (LFW and Pub-Fig) and person re-identification (VIPeR) show that our algorithm outperforms the state-of-the-art methods.
Yang Yang 0062, Shengcai Liao, Zhen Lei 0001, Stan Z. Li
AAAI2
2016 Embedding Deep Metric for Person Re-identification: A Study Against Large Variations
Hailin Shi, Yang Yang 0062, Xiangyu Zhu 0001, Shengcai Liao, Zhen Lei 0001, Wei-Shi Zheng 0001, Stan Z. Li
ECCV (1)4
2016 A Fast and Accurate Unconstrained Face Detector
abstract
We propose a method to address challenges in unconstrained face detection, such as arbitrary pose variations and occlusions. First, a new image feature called Normalized Pixel Difference (NPD) is proposed. NPD feature is computed as the difference to sum ratio between two pixel values, inspired by the Weber Fraction in experimental psychology. The new feature is scale invariant, bounded, and is able to reconstruct the original image. Second, we propose a deep quadratic tree to learn the optimal subset of NPD features and their combinations, so that complex face manifolds can be partitioned by the learned rules. This way, only a single soft-cascade classifier is needed to handle unconstrained face detection. Furthermore, we show that the NPD features can be efficiently obtained from a look up table, and the detection template can be easily scaled, making the proposed face detector very fast. Experimental results on three public face datasets (FDDB, GENKI, and CMU-MIT) show that the proposed method achieves state-of-the-art performance in detecting unconstrained faces with arbitrary pose variations and occlusions in cluttered scenes.
Shengcai Liao, Anil K. Jain 0001, Stan Z. Li
IEEE Trans. Pattern Anal. Mach. Intell.1
2016 Multicamera Joint Video Synopsis
abstract
Due to an increasing demand for video surveillance, there is an explosive growth of surveillance videos, which causes a big challenge in video storage, browsing, and retrieval. The video synopsis technique is thus developed to extract and rearrange the moving objects so as to handle the massive video browsing challenge. However, the traditional video synopsis (TVS) method only considers the processing videos captured by a single camera, ignoring object interactions in multicamera videos. To address this issue, we propose a novel multicamera joint video synopsis (JVS) algorithm for multicamera surveillance videos. First, a key time stamp (KTS) selection method is designed to find an object's appearing, merging, splitting, and disappearing moments in the frame sequence, called tube, of that object. Second, tubes are rearranged by minimizing a global energy function that involves the overall camera views. Compared with the energy function used in TVS, the proposed global energy function considers the chronological orders of tubes not only in the same camera view but also among different camera views. Moreover, the chronological disorder cost term is formulated based on the KTS labels, and improved by considering the visual similarity between two tubes. Finally, the multicamera synopsis videos are separately generated by stitching together the globally rearranged tubes and background images of the same camera view. Extensive experiments show that the proposed JVS method is better than the traditional single-camera video synopsis method in preserving the chronological orders of moving objects among multicamera synopsis videos.
Jianqing Zhu, Shengcai Liao, Stan Z. Li
IEEE Trans. Circuits Syst. Video Technol.2
2015 Cross-Modal Similarity Learning: A Low Rank Bilinear Formulation
abstract
The cross-media retrieval problem has received much attention in recent years due to the rapid increasing of multimedia data on the Internet. A new approach to the problem has been raised which intends to match features of different modalities directly. In this research, there are two critical issues: how to get rid of the heterogeneity between different modalities and how to match the cross-modal features of different dimensions. Recently metric learning methods show a good capability in learning a distance metric to explore the relationship between data points. However, the traditional metric learning algorithms only focus on single-modal features, which suffer difficulties in addressing the cross-modal features of different dimensions. In this paper, we propose a cross-modal similarity learning algorithm for the cross-modal feature matching. The proposed method takes a bilinear formulation, and with the nuclear-norm penalization, it achieves low-rank representation. Accordingly, the accelerated proximal gradient algorithm is successfully imported to find the optimal solution with a fast convergence rate O(1/t2). Experiments on three well known image-text cross-media retrieval databases show that the proposed method achieves the best performance compared to the state-of-the-art algorithms.
Cuicui Kang, Shengcai Liao, Yonghao He, Jian Wang 0068, Wenjia Niu, Shiming Xiang, Chunhong Pan
CIKM2
2015 Person re-identification by Local Maximal Occurrence representation and metric learning
abstract
Person re-identification is an important technique towards automatic search of a person's presence in a surveillance video. Two fundamental problems are critical for person re-identification, feature representation and metric learning. An effective feature representation should be robust to illumination and viewpoint changes, and a discriminant metric should be learned to match various person images. In this paper, we propose an effective feature representation called Local Maximal Occurrence (LOMO), and a subspace and metric learning method called Cross-view Quadratic Discriminant Analysis (XQDA). The LOMO feature analyzes the horizontal occurrence of local features, and maximizes the occurrence to make a stable representation against viewpoint changes. Besides, to handle illumination variations, we apply the Retinex transform and a scale invariant texture operator. To learn a discriminant metric, we propose to learn a discriminant low dimensional subspace by cross-view quadratic discriminant analysis, and simultaneously, a QDA metric is learned on the derived subspace. We also present a practical computation method for XQDA, as well as its regularization. Experiments on four challenging person re-identification databases, VIPeR, QMUL GRID, CUHK Campus, and CUHK03, show that the proposed method improves the state-of-the-art rank-1 identification rates by 2.2%, 4.88%, 28.91%, and 31.55% on the four databases, respectively.
Shengcai Liao, Xiangyu Zhu 0001, Stan Z. Li
CVPR1
2015 Efficient PSD Constrained Asymmetric Metric Learning for Person Re-Identification
abstract
Person re-identification is becoming a hot research topic due to its value in both machine learning research and video surveillance applications. For this challenging problem, distance metric learning is shown to be effective in matching person images. However, existing approaches either require a heavy computation due to the positive semidefinite (PSD) constraint, or ignore the PSD constraint and learn a free distance function that makes the learned metric potentially noisy. We argue that the PSD constraint provides a useful regularization to smooth the solution of the metric, and hence the learned metric is more robust than without the PSD constraint. Another problem with metric learning algorithms is that the number of positive sample pairs is very limited, and the learning process is largely dominated by the large amount of negative sample pairs. To address the above issues, we derive a logistic metric learning approach with the PSD constraint and an asymmetric sample weighting strategy. Besides, we successfully apply the accelerated proximal gradient approach to find a global minimum solution of the proposed formulation, with a convergence rate of O(1/t^2) where t is the number of iterations. The proposed algorithm termed MLAPG is shown to be computationally efficient and able to perform low rank selection. We applied the proposed method for person re-identification, achieving state-of-the-art performance on four challenging databases (VIPeR, QMUL GRID, CUHK Campus, and CUHK03), compared to existing metric learning methods as well as published results.
Shengcai Liao, Stan Z. Li
ICCV1
2015 Partial Person Re-Identification
abstract
We address a new partial person re-identification (re-id) problem, where only a partial observation of a person is available for matching across different non-overlapping camera views. This differs significantly from the conventional person re-id setting where it is assumed that the full body of a person is detected and aligned. To solve this more challenging and realistic re-id problem without the implicit assumption of manual body-parts alignment, we propose a matching framework consisting of 1) a local patch-level matching model based on a novel sparse representation classification formulation with explicit patch ambiguity modelling, and 2) a global part-based matching model providing complementary spatial layout information. Our framework is evaluated on a new partial person re-id dataset as well as two existing datasets modified to include partial person images. The results show that the proposed method outperforms significantly existing re-id methods as well as other partial visual matching methods.
Wei-Shi Zheng 0001, Xiang Li 0032, Tao Xiang 0002, Shengcai Liao, Jian-Huang Lai, Shaogang Gong
ICCV4
2015 Moving Object Detection Revisited: Speed and Robustness
abstract
The detection of moving objects in videos is very important in many video processing applications, and background modeling is often an indispensable process to achieve this goal. Most of the traditional background modeling methods utilize color or texture information. However, color information is sensitive to illumination variations and texture information cannot be utilized to separate smooth foreground from smooth background in most cases. Achieving good performance in terms of high foreground detection accuracy and low computational cost is also challenging. In this paper, we propose a new integration framework of texture and color information for background modeling, in which the foreground decision equation includes three parts (one part for color information, one part for texture information, and the left part for the integration of color and texture information). This framework is able to combine the advantages of texture and color features while inhibiting their disadvantages as well. Moreover, we propose a block-based method to accelerate the background modeling. In particular, in the texture information modeling process, a single histogram model is established for each block whose bins indicate the occurrence probabilities of different patterns, which is different from the traditional multihistogram model for block-based background modeling, and then dominant background patterns are selected to calculate the background likelihood of new coming blocks. Dynamic background and multimodal problems can be handled through this technique. To evaluate the foreground detection performance reasonably, a new quality measure is proposed. Extensive experiments on various challenging videos validate the effectiveness of the proposed method over state-of-the-art methods.
Hong Han 0001, Jianfei Zhu, Shengcai Liao, Zhen Lei 0001, Stan Z. Li
IEEE Trans. Circuits Syst. Video Technol.3
2015 High-Performance Video Condensation System
abstract
Video synopsis or condensation is a smart solution for fast video browsing and storage. However, most of the existing methods work offline, where two main phases are required. The first phase is to prepare tubes and background images. The second phase is to rearrange tubes and stitch them into backgrounds. However, with a long video sequence, the first phase is memory consuming for data storage, and the second phase is computationally expensive to rearrange all tubes simultaneously. To overcome these problems, we propose a high-performance video condensation system based on an online content-aware framework. The online framework transforms the optimization problem of tube rearrangement into a stepwise optimization problem. Therefore, it can condense video with much less memory and higher speed than the offline framework. With the aid of this transformation, the proposed system can process input videos and produce condensed videos simultaneously. Thus it is suitable for real-time endless surveillance videos. Meanwhile, the online mechanism allows users to directly visit the condensation video that has been generated. Moreover, the content-aware mechanism makes the proposed system able to automatically determine the duration of a condensed video. Finally, the proposed system uses Graphic Processing Unit (GPU) and multicore techniques to improve the speed. Extensive experiments that validate the high efficiency of the system are presented.
Jianqing Zhu, Shikun Feng, Dong Yi, Shengcai Liao, Zhen Lei 0001, Stan Z. Li
IEEE Trans. Circuits Syst. Video Technol.4
2015 Learning Consistent Feature Representation for Cross-Modal Multimedia Retrieval
abstract
The cross-modal feature matching has gained much attention in recent years, which has many practical applications, such as the text-to-image retrieval. The most difficult problem of cross-modal matching is how to eliminate the heterogeneity between modalities. The existing methods (e.g., CCA and PLS) try to learn a common latent subspace, where the heterogeneity between two modalities is minimized so that cross-matching is possible. However, most of these methods require fully paired samples and suffer difficulties when dealing with unpaired data. Besides, utilizing the class label information has been found as a good way to reduce the semantic gap between the low-level image features and high-level document descriptions. Considering this, we propose a novel and effective supervised algorithm, which can also deal with the unpaired data. In the proposed formulation, the basis matrices of different modalities are jointly learned based on the training samples. Moreover, a local group-based priori is proposed in the formulation to make a better use of popular block based features (e.g., HOG and GIST). Extensive experiments are conducted on four public databases: Pascal VOC2007, LabelMe, Wikipedia, and NUS-WIDE. We also evaluated the proposed algorithm with unpaired data. By comparing with existing state-of-the-art algorithms, the results show that the proposed algorithm is more robust and achieves the best performance, which outperforms the second best algorithm by about 5% on both the Pascal VOC2007 and NUS-WIDE databases.
Cuicui Kang, Shiming Xiang, Shengcai Liao, Changsheng Xu, Chunhong Pan
IEEE Trans. Multim.3
2014 Salient Color Names for Person Re-identification
Yang Yang 0062, Jimei Yang, Shengcai Liao, Dong Yi, Stan Z. Li
ECCV (1)4
2014 A benchmark study of large-scale unconstrained face recognition
abstract
Many efforts have been made in recent years to tackle the unconstrained face recognition challenge. For the benchmark of this challenge, the Labeled Faces in theWild (LFW) database has been widely used. However, the standard LFW protocol is very limited, with only 3,000 genuine and 3,000 impostor matches for classification. Today a 97% accuracy can be achieved with this benchmark, remaining a very limited room for algorithm development. However, we argue that this accuracy may be too optimistic because the underlying false accept rate may still be high (e.g. 3%). Furthermore, performance evaluation at low FARs is not statistically sound by the standard protocol due to the limited number of impostor matches. Thereby we develop a new benchmark protocol to fully exploit all the 13,233 LFW face images for large-scale unconstrained face recognition evaluation under both verification and open-set identification scenarios, with a focus at low FARs. Based on the new benchmark, we evaluate 21 face recognition approaches by combining 3 kinds of features and 7 learning algorithms. The benchmark results show that the best algorithm achieves 41.66% verification rates at FAR=0.1%, and 18.07% open-set identification rates at rank 1 and FAR=1%. Accordingly we conclude that the large-scale unconstrained face recognition problem is still largely unresolved, thus further attention and effort is needed in developing effective feature representations and learning algorithms. We thereby release a benchmark tool to advance research in this field.
Shengcai Liao, Zhen Lei 0001, Dong Yi, Stan Z. Li
IJCB1
2014 Multi-camera Trajectory Mining: Database and Evaluation
abstract
In recent years, large-scale video search and mining has been an active research area. Exploring the trajectory of pedestrian of interest in non-overlapping multi-camera network, namely the trajectory mining, is very useful for visual surveillance and criminal investigation. The trajectory mentioned in our work describes the transition of pedestrian among cameras from a macroscopic perspective which is different from the concept in conventional tracking field. In this paper, we collect a database called TMin to promote research and development of trajectory mining. This release of Version 1 contains 1680 images from 30 subjects, all the images are extracted from 6 surveillance videos over two hours, and each subject appears in at least two different cameras. We describe the apparatuses, environments and procedure of the data collection and present baseline performance on the TMin database.
Shengcai Liao, Dong Yi, Zhen Lei 0001, Stan Z. Li
ICPR2
2014 Color Models and Weighted Covariance Estimation for Person Re-identification
abstract
Due to illumination changes, partial occlusions, and object scale differences, person re-identification over disjoint camera views becomes a challenging problem. To address this problem, a variety of image representations have been put forward. In this paper, the illumination invariance and distinctiveness of different color models including the proposed color model are firstly evaluated. Since color distribution is robust to image scales and partial occlusions, color distributions based on different color models are then calculated and fused in the stage of feature extraction. Different color models obtain robustness to different types of illumination and thus fusing them can compensate each other and contribute to better performance. In the stage of feature matching, a weighted KISSME is presented to learn a better distance metric than the original KISSME. Experimental results demonstrate its feasibility and effectiveness. Finally, image pairs are matched based on the learned distance metric. Experiments conducted on two public benchmark datasets (VIPeR and PRID 450S) show that the proposed algorithm outperforms the state-of-the-art methods.
Yang Yang 0062, Shengcai Liao, Zhen Lei 0001, Dong Yi, Stan Z. Li
ICPR2
2014 Deep Metric Learning for Person Re-identification
abstract
Various hand-crafted features and metric learning methods prevail in the field of person re-identification. Compared to these methods, this paper proposes a more general way that can learn a similarity metric from image pixels directly. By using a "siamese" deep neural network, the proposed method can jointly learn the color feature, texture feature and metric in a unified framework. The network has a symmetry structure with two sub-networks which are connected by a cosine layer. Each sub network includes two convolutional layers and a full connected layer. To deal with the big variations of person images, binomial deviance is used to evaluate the cost between similarities and labels, which is proved to be robust to outliers. Experiments on VIPeR illustrate the superior performance of our method and a cross database experiment also shows its good generalization.
Dong Yi, Zhen Lei 0001, Shengcai Liao, Stan Z. Li
ICPR3
2014 Kernel sparse representation with pixel-level and region-level local feature kernels for face recognition
Cuicui Kang, Shengcai Liao, Shiming Xiang, Chunhong Pan
Neurocomputing2
2013 Robust Multi-resolution Pedestrian Detection in Traffic Scenes
abstract
The serious performance decline with decreasing resolution is the major bottleneck for current pedestrian detection techniques. In this paper, we take pedestrian detection in different resolutions as different but related problems, and propose a Multi-Task model to jointly consider their commonness and differences. The model contains resolution aware transformations to map pedestrians in different resolutions to a common space, where a shared detector is constructed to distinguish pedestrians from background. For model learning, we present a coordinate descent procedure to learn the resolution aware transformations and deformable part model (DPM) based detector iteratively. In traffic scenes, there are many false positives located around vehicles, therefore, we further build a context model to suppress them according to the pedestrian-vehicle relationship. The context model can be learned automatically even when the vehicle annotations are not available. Our method reduces the mean miss rate to 60% for pedestrians taller than 30 pixels on the Caltech Pedestrian Benchmark, which noticeably outperforms previous state-of-the-art (71%).
Xucong Zhang, Zhen Lei 0001, Shengcai Liao, Stan Z. Li
CVPR4
2013 Partial Face Recognition: Alignment-Free Approach
abstract
Numerous methods have been developed for holistic face recognition with impressive performance. However, few studies have tackled how to recognize an arbitrary patch of a face image. Partial faces frequently appear in unconstrained scenarios, with images captured by surveillance cameras or handheld devices (e.g., mobile phones) in particular. In this paper, we propose a general partial face recognition approach that does not require face alignment by eye coordinates or any other fiducial points. We develop an alignment-free face representation method based on Multi-Keypoint Descriptors (MKD), where the descriptor size of a face is determined by the actual content of the image. In this way, any probe face image, holistic or partial, can be sparsely represented by a large dictionary of gallery descriptors. A new keypoint descriptor called Gabor Ternary Pattern (GTP) is also developed for robust and discriminative face recognition. Experimental results are reported on four public domain face databases (FRGCv2.0, AR, LFW, and PubFig) under both the open-set identification and verification scenarios. Comparisons with two leading commercial face recognition SDKs (PittPatt and FaceVACS) and two baseline algorithms (PCA+LDA and LBP) show that the proposed method, overall, is superior in recognizing both holistic and partial faces without requiring alignment.
Shengcai Liao, Anil K. Jain 0001, Stan Z. Li
IEEE Trans. Pattern Anal. Mach. Intell.1
2012 Nighttime Face Recognition at Long Distance: Cross-Distance and Cross-Spectral Matching
Hyunju Maeng, Shengcai Liao, Dongoh Kang, Seong-Whan Lee, Anil K. Jain 0001
ACCV (2)2
2012 Kernel Homotopy based sparse representation for object classification
Cuicui Kang, Shengcai Liao, Shiming Xiang, Chunhong Pan
ICPR2
2012 Efficient feature selection for linear discriminant analysis and its application to face recognition
Zhen Lei 0001, Shengcai Liao, Stan Z. Li
ICPR2
2012 Extracting non-negative basis images using pixel dispersion penalty
Wei-Shi Zheng 0001, Jian-Huang Lai, Shengcai Liao, Ran He 0001
Pattern Recognit.3
2012 Coupled Discriminant Analysis for Heterogeneous Face Recognition
abstract
Coupled space learning is an effective framework for heterogeneous face recognition. In this paper, we propose a novel coupled discriminant analysis method to improve the heterogeneous face recognition performance. There are two main advantages of the proposed method. First, all samples from different modalities are used to represent the coupled projections, so that sufficient discriminative information could be extracted. Second, the locality information in kernel space is incorporated into the coupled discriminant analysis as a constraint to improve the generalization ability. In particular, two implementations of locality constraint in kernel space (LCKS)-based coupled discriminant analysis methods, namely LCKS-coupled discriminant analysis (LCKS-CDA) and LCKS-coupled spectral regression (LCKS-CSR), are presented. Extensive experiments on three cases of heterogeneous face matching (high versus low image resolution, digital photo versus video image, and visible light versus near infrared) validate the efficacy of the proposed method.
Zhen Lei 0001, Shengcai Liao, Anil K. Jain 0001, Stan Z. Li
IEEE Trans. Inf. Forensics Secur.2
2011 Partial face recognition: An alignment free approach
abstract
Many approaches have been developed for holistic face recognition with impressive performance. However, few studies have addressed the question of how to recognize an arbitrary image patch of a holistic face. In this paper we ad- dress this problem of partial face recognition. Partial faces frequently appear in unconstrained image capture environments, particularly when faces are captured by surveillance cameras or handheld devices (e.g. mobile phones). The pro- posed approach adopts a variable-size description which represents each face with a set of keypoint descriptors. In this way, we argue that a probe face image, holistic or partial, can be sparsely represented by a large dictionary of gallery descriptors. The proposed method is alignment free and we address large-scale face recognition problems by a fast filtering strategy. Experimental results on three public domain face databases (FRGCv2.0, AR, and LFW) show that the proposed method achieves promising results in recognizing both holistic and partial faces.
Shengcai Liao, Anil K. Jain 0001
IJCB1
2011 Kernel sparse representation with local patterns for face recognition
abstract
In this paper we propose a novel kernel sparse representation classification (SRC) framework and utilize the local binary pattern (LBP) descriptor in this framework for robust face recognition. First we develop a kernel coordinate descent (KCD) algorithm for 11 minimization in the kernel space, which is based on the covariance update technique. Then we extract LBP descriptors from each image and apply two types of kernels (χ2distance based and Hamming distance based) with the proposed KCD algorithm under the SRC framework for face recognition. Experiments on both the Extended Yale B and the PIE face databases show that the proposed method is more robust against noise, occlusion, and illumination variations, even with small number of training samples.
Cuicui Kang, Shengcai Liao, Shiming Xiang, Chunhong Pan
ICIP2
2011 Face Recognition by Exploring Information Jointly in Space, Scale and Orientation
abstract
Information jointly contained in image space, scale and orientation domains can provide rich important clues not seen in either individual of these domains. The position, spatial frequency and orientation selectivity properties are believed to have an important role in visual perception. This paper proposes a novel face representation and recognition approach by exploring information jointly in image space, scale and orientation domains. Specifically, the face image is first decomposed into different scale and orientation responses by convolving multiscale and multiorientation Gabor filters. Second, local binary pattern analysis is used to describe the neighboring relationship not only in image space, but also in different scale and orientation responses. This way, information from different domains is explored to give a good face representation for recognition. Discriminant classification is then performed based upon weighted histogram intersection or conditional mutual information with linear discriminant analysis techniques. Extensive experimental results on FERET, AR, and FRGC ver 2.0 databases show the significant advantages of the proposed method over the existing ones.
Zhen Lei 0001, Shengcai Liao, Matti Pietikäinen, Stan Z. Li
IEEE Trans. Image Process.2
2010 Modeling pixel process with scale invariant local patterns for background subtraction in complex scenes
abstract
Background modeling plays an important role in video surveillance, yet in complex scenes it is still a challenging problem. Among many difficulties, problems caused by illumination variations and dynamic backgrounds are the key aspects. In this work, we develop an efficient background subtraction framework to tackle these problems. First, we propose a scale invariant local ternary pattern operator, and show that it is effective for handling illumination variations, especially for moving soft shadows. Second, we propose a pattern kernel density estimation technique to effectively model the probability distribution of local patterns in the pixel process, which utilizes only one single LBP-like pattern instead of histogram as feature. Third, we develop multimodal background models with the above techniques and a multiscale fusion scheme for handling complex dynamic backgrounds. Exhaustive experimental evaluations on complex scenes show that the proposed method is fast and effective, achieving more than 10% improvement in accuracy compared over existing state-of-the-art algorithms.
Shengcai Liao, Guoying Zhao 0001, Vili Kellokumpu, Matti Pietikäinen, Stan Z. Li
CVPR1
2010 Online Principal Background Selection for Video Synopsis
abstract
Video synopsis provides a means for fast browsing of activities in video. Principal background selection (PBS) is an important step in video synopsis. Existing methods make PBS in an offline way and at a high memory cost. In this paper we propose a novel background selection method, ``online principal background selection'' (OPBS). The OPBS selects n principal backgrounds from N backgrounds in an online fashion with a low memory cost, making it possible to build an efficient online video synopsis system. Another advantage is that, with OPBS, the selected backgrounds are related to not only background changes over time but also video activities. Experimental results demonstrate the advantages of the proposed OPBS.
Shikun Feng, Shengcai Liao, Stan Z. Li
ICPR2
2010 Moving Cast Shadow Removal Based on Local Descriptors
abstract
Moving cast shadow removal is an important yet difficult problem in video analysis and applications. This paper presents a novel algorithm for detection of moving cast shadows, that based on a local texture descriptor called Scale Invariant Local Ternary Pattern (SILTP). An assumption is made that the texture properties of cast shadows bears similar patterns to those of the background beneath them. The likelihood of cast shadows is derived using information in both color and texture. An online learning scheme is employed to update the shadow model adaptively. Finally, the posterior probability of cast shadow region is formulated by further incorporating prior contextual constrains using a Markov Random Field (MRF) model. The optimal solution is found using graph cuts. Experimental results tested on various scenes demonstrate the robustness of the algorithm.
Shengcai Liao, Zhen Lei 0001, Stan Z. Li
ICPR2
2010 Flickr group recommendation based on tensor decomposition
abstract
Over the last few years, Flickr has gained massive popularity and groups in Flickr are one of the main ways for photo diffusion. However, the huge volume of groups brings troubles for users to decide which group to choose. In this paper, we propose a tensor decomposition-based group recommendation model to suggest groups to users which can help tackle this problem. The proposed model measures the latent associations between users and groups by considering both semantic tags and social relations. Experimental results show the usefulness of the proposed model.
Qiudan Li, Shengcai Liao, Leiming Zhang
SIGIR3
2008 Gabor volume based local binary pattern for face representation and recognition
abstract
This paper presents a novel face representation and recognition approach. The face image is first decomposed by multi-scale and multi-orientation Gabor filters and local binary pattern (LBP) analysis is then applied on the derived Gabor magnitude responses. Different from (W.C. Zhang et al., 2005), the present method not only describes the neighboring relationship in spatial domain, but also exploit those between different scales (frequency) and orientations. Specifically, we first reformulate the Gabor magnitude responses as a 3rd-order volume and then apply LBP analysis on three orthogonal planes of the Gabor volume, named GV-LBP-TOP in short, in a hope to encode sufficient information for face representation. Further, a computationally effective version, E-GV-LBP, is proposed to depict the neighboring changes in spatial, frequency and orientation domains simultaneously. Finally, the weighted histogram intersection metric is utilized to measure the dissimilarity of faces. Experimental results on FERET and FRGC ver 2.0 databases show the significant advantages of the proposed method.
Zhen Lei 0001, Shengcai Liao, Ran He 0001, Matti Pietikäinen, Stan Z. Li
FG2
2007 Coarse-to-Fine Statistical Shape Model by Bayesian Inference
Ran He 0001, Stan Z. Li, Zhen Lei 0001, Shengcai Liao
ACCV (1)4
2007 Fusion of Face and Palmprint for Personal Identification Based on Ordinal Features
abstract
In this paper, we present a face and palmprint multimodal biometric identification method and system to improve the identification performance. Effective classifiers based on ordinal features are constructed for faces and palmprints, respectively. Then, the matching scores from the two classifiers are combined using several fusion strategies. Experimental results on a middle-scale data set have demonstrated the effectiveness of the proposed system.
Rufeng Chu, Shengcai Liao, Zhenan Sun, Stan Z. Li, Tieniu Tan
CVPR2
2007 Part-based Face Recognition Using Near Infrared Images
abstract
Recently, the authors developed NIR based face recognition for highly accurate face recognition under illumination variations. In this paper, we present a part-based method for improving its robustness with respect to pose variations. An NIR face is decomposed into parts. A part classifier is built for each part, using the most discriminative LBP histogram features selected by AdaBoost learning. The outputs of part classifiers are fused to give the final score. Experiments show that the present method outperforms the whole face-based method by 4.53%.
Shengcai Liao, Stan Z. Li, Peiren Zhang
CVPR2
2007 On Constrained Sparse Matrix Factorization
abstract
Various linear subspace methods can be formulated in the notion of matrix factorization in which a cost function is minimized subject to some constraints. Among them, constraints on sparseness have received much attention recently. Some popular constraints such as non-negativity, lasso penalty, and (plain) orthogonality etc have been so far applied to extract sparse features. However, little work has been done to give theoretical and experimental analyses on the differences of the impacts of different constraints within a framework. In this paper, we analyze the problem in a more general framework called Constrained Sparse Matrix Factorization (CSMF). In CSMF, a particular case called CSMF with non-negative components (CSMFnc) is further discussed. Unlike NMF, CSMFnc allows not only additive but also subtractive combinations of non-negative sparse components. It is useful to produce much sparser features than those produced by NMF and meanwhile have better reconstruction ability, achieving a trade-off between sparseness and low MSE value. Moreover, for optimization, an alternating algorithm is developed and a gentle update strategy is further proposed for handling the alternating process. Experimental analyses are performed on the Swimmer data set and CBCLface database. In particular, CSMF can successfully extract all the proper components without any ghost on Swimmer, gaining a significant improvement over the compared well-known algorithms.
Wei-Shi Zheng 0001, Stan Z. Li, Jian-Huang Lai, Shengcai Liao
ICCV4
2007 Illumination Invariant Face Recognition Using Near-Infrared Images
abstract
Most current face recognition systems are designed for indoor, cooperative-user applications. However, even in thus-constrained applications, most existing systems, academic and commercial, are compromised in accuracy by changes in environmental illumination. In this paper, we present a novel solution for illumination invariant face recognition for indoor, cooperative-user applications. First, we present an active near infrared (NIR) imaging system that is able to produce face images of good condition regardless of visible lights in the environment. Second, we show that the resulting face images encode intrinsic information of the face, subject only to a monotonic transform in the gray tone; based on this, we use local binary pattern (LBP) features to compensate for the monotonic transform, thus deriving an illumination invariant face representation. Then, we present methods for face recognition using NIR images; statistical learning algorithms are used to extract most discriminative features from a large pool of invariant LBP features and construct a highly accurate face matching engine. Finally, we present a system that is able to achieve accurate and fast face recognition in practice, in which a method is provided to deal with specular reflections of active NIR lights on eyeglasses, a critical issue in active NIR image-based face recognition. Extensive, comparative results are provided to evaluate the imaging hardware, the face and eye detection algorithms, and the face recognition algorithms and systems, with respect to various factors, including illumination, eyeglasses, time lapse, and ethnic groups.
Stan Z. Li, Rufeng Chu, Shengcai Liao, Lun Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3