Qiangchang Wang

dblp:210/2221 · DBLP profile ↗
← Back
29ranked-venue papers
8as first author
24since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 10 since 2021Security and privacy · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Recent advances of local mechanisms in vision foundation models: A survey and outlook
Qiangchang Wang, Jing Li 0175, Yilong Yin, Huimin Lu 0001
Comput. Vis. Image Underst.1
2026 Progressive intra- and inter-modality relation learning for multimodal sentiment analysis
Jing Li 0175, Zhibin Wu, Qiangchang Wang
Knowl. Based Syst.3
2026 Multi-feature embedding, fusion and enhancement for partial finger vein recognition
Enyan Li, Lu Yang 0005, Qiangchang Wang, Yilong Yin
Pattern Recognit.4
2025 Disparity-Guided Cross-View Transformer For Stereo Image Super-Resolution
abstract
Although transformer-based methods excel in stereo image super-resolution, the full potential of the distinctive, complementary information inherent in stereo images has not been fully utilized. We propose a Disparity-Guided Cross-View Transformer (DCT) to extract features across dimensions and views, achieving a more comprehensive feature representation. The proposed method introduces mutual attention within the transformer architecture, establishing the difference between left and right views through cross-view interaction. The proposed algorithm effectively harnesses the complementary information present in stereo image pairs, enhancing the restoration performance. Furthermore, we propose a disparity-guided cross-modal residual fusion module that leverages disparity information as prior knowledge to substantially improve image reconstruction. This module significantly complements the missing information in stereo images, enabling the network to comprehend more effectively and reconstruct the image content with greater accuracy. Extensive experimental results and ablation studies demonstrate the effectiveness of our method.
Bingting Li, Wenjing Shang, Yongshun Gong, Qiangchang Wang, Xinxin Zhang 0004, Yilong Yin
ICASSP4
2025 LOFI: Harnessing Attention Dynamics for Facial Expression Recognition with Noisy Labels
abstract
Facial expression recognition (FER) faces unique challenges from expression ambiguity and noisy labels, degrading performance in real-world applications. While leveraging attention, existing methods frequently neglect attention dynamic mechanism of dispersion followed by focus and the spatially structural knowledge essential for effectively guiding this dynamic dispersion of attention. To address this, we propose the Last fOcus First dIsperse (LOFI), which harnesses attention dynamics and spatial structure information dynamically refining focus during classification to mitigate the impact of noise labels. LOFI comprises two modules: Spatial Keypoint-enhanced Fused Attention (SKFA), which disperses focus on subtle, critical features near facial landmarks, and Hybrid Consistency-Calibrated Loss (HCCL), which employs consistency and re-weighting strategies focusing attention to boost performance. The synergy between these modules enables LOFI to adapt to various noise levels and challenging classes. Extensive experiments demonstrate that LOFI outperforms existing state-of-the-art (SOTA) methods in noisy FER, offering a robust solution for real-world applications.
Qiangchang Wang, Xinxin Zhang 0004, Yilong Yin
ICASSP2
2025 Semantic-Aware Adaptation with Hierarchical Multimodal Prompts for Few-Shot Learning
abstract
Few-shot learning aims to recognize novel classes with limited labeled samples. Existing methods often utilize semantic information from natural language but integrate it after visual feature extraction, overlooking fine-grained cross-modal interactions. Moreover, they struggle with spatial variations, as target objects often appear in varying regions. To address these limitations, we propose Semantic-Aware Adaptation (SAA) which consists of Hierarchical Multimodal Prompts (HMP) and Global-Local Adaptation (GLA). SAA leverages textual prompts encoded by CLIP and adaptively modulated by learnable visual prompts to better align text and vision feature distributions. During visual extraction, these fused prompts are integrated with visual patches in channel and spatial dimensions, dynamically enhancing visual features, while a consistency loss is used to regularize them to prevent bias and overfitting. By the deep cross-modal interactions at different scales, HMP is thus constructed for improving robustness to spatial variations. In GLA, patch-level soft labels based on rich semantics further emphasize class-specific visual patches, improving token dependency learning. Experiments on five benchmarks demonstrate the effectiveness of our SAA, with an average accuracy improvement of 1.6% on challenging 1-shot tasks.
Wenhao Li 0011, Qiangchang Wang, Jing Li 0175, Mindi Ruan, Yilong Yin
ICME2
2025 Video Instance Segmentation by Weighted Structure Inference
abstract
Video instance segmentation presents significant challenges in complex and dynamic environments, where instances experience progressive occlusion, either from objects obstructing each other or due to changes in the camera's viewpoint. Current state-of-the-art methods rely on memory bank mechanisms, but we still look forward to new paradigms that have the ability to capture and utilize structural information, the ability to model complex relationships, and the flexibility to adapt to dynamic scenarios. To this end, we propose the Weighted Structure Inference method for Video Instance Segmentation. We build on high-order structural relationships by constructing hypergraphs for each video frame, enabling the capture of complex interactions that go beyond traditional pairwise methods. To model intricate dynamics, we introduce Weighted Sheaf Hypergraph Convolution, which enhances the hierarchical and structural information embedded in the hypergraph. Furthermore, we ensure spatio-temporal consistency by employing a dynamic inference mechanism based on Weighted Sliced Wasserstein distance to compare structural features across adjacent frames. Our method preserves the topological characteristics of occlusion instances and improves the reliability of instance tracking across frames. Experimental results demonstrate that our method outperforms existing video instance segmentation frameworks in both Video Instance and Panoptic Segmentation tasks.
Zheyun Qin, Deng Yu, Qiangchang Wang, Zhumin Chen
ACM Multimedia4
2025 VT-FSL: Bridging Vision and Text with LLMs for Few-Shot Learning
abstract
Few-shot learning (FSL) aims to recognize novel concepts from only a few labeled support samples. Recent studies enhance support features by incorporating additional semantic information (e.g., class descriptions) or designing complex semantic fusion modules. However, these methods still suffer from hallucinating semantics that contradict the visual evidence due to the lack of grounding in actual instances, resulting in noisy guidance and costly corrections. To address these issues, we propose a novel framework, bridging Vision and Text with LLMs for Few-Shot Learning (VT-FSL), which constructs precise cross-modal prompts conditioned on Large Language Models (LLMs) and support images, seamlessly integrating them through a geometry-aware alignment mechanism. It mainly consists of Cross-modal Iterative Prompting (CIP) and Cross-modal Geometric Alignment (CGA). Specifically, the CIP conditions an LLM on both class names and support images to generate precise class descriptions iteratively in a single structured reasoning pass. These descriptions not only enrich the semantic understanding of novel classes but also enable the zero-shot synthesis of semantically consistent images. The descriptions and synthetic images act respectively as complementary textual and visual prompts, providing high-level class semantics and low-level intra-class diversity to compensate for limited support data. Furthermore, the CGA jointly aligns the fused textual, support, and synthetic visual representations by minimizing the kernelized volume of the 3-dimensional parallelotope they span. It captures global and nonlinear relationships among all representations, enabling structured and consistent multimodal integration. The proposed VT-FSL method establishes new state-of-the-art performance across ten diverse benchmarks, including standard, cross-domain, and fine-grained few-shot learning scenarios. Code is available at https://github.com/peacelwh/VT-FSL.
Wenhao Li 0011, Qiangchang Wang, Xianjing Meng, Zhibin Wu, Yilong Yin
NeurIPS2
2025 SPADesc: Semantic and parallel attention with feature description
Haijun Meng, Huimin Lu 0001, Bozhi Ding, Qiangchang Wang
Neurocomputing4
2025 A noise-robust and generalizable framework for facial expression recognition
Qiangchang Wang, Jing Li 0175, Yilong Yin
Inf. Sci.2
2025 Diverse Information Aggregation with Adaptive Graph Construction and prompts for deepfake detection
Zhenhua Bai, Qiangchang Wang, Lu Yang 0005, Xinxin Zhang 0004, Yanbo Gao, Yilong Yin
Image Vis. Comput.2
2024 Two-phase Parametric Registration for Retinal Images
abstract
We propose a two-phase parametric registration algorithm for retinal images. Our algorithm focuses on dealing with the geometric transformation and the intensity transformation in the retinal image registration problem. In the first phase, we efficiently detect only one pair of feature points in the source and the target retinal images to estimate a translation transformation and get a warped source image. In the second phase, we estimate both the intensity and the geometric transformations between the target image and the warped source image by fitting parametric expressions. The displacement field is generated by a super fast and accurate coarse-to-fine elastic registration algorithm—local all-pass filters algorithm (LAP). At each iteration of the LAP, the elastic displacement field and the intensity difference take turns being fitted by two different low-order polynomial functions. The fitting steps are performed by solving linear systems of equations efficiently. Experiments on real retinal image datasets demonstrated the high accuracy and computational efficiency of the proposed retinal image registration method.
Xinxin Zhang 0004, Xiankai Lu, Jizhou Li, Yongshun Gong, Qiangchang Wang, Yilong Yin
ICME5
2024 Hypergraph-guided Intra- and Inter-category Relation Modeling for Fine-grained Visual Recognition
abstract
Fine-grained Visual Recognition (FGVR) aims to distinguish objects within similar subcategories. Humans adeptly perform this challenging task by leveraging both intra-category distinctiveness and inter-category similarity. However, previous methods fail to combine these two complementary dimensions and mine the intrinsic relations among various semantic features. To address these limitations, we propose HI2R, a Hypergraph-guided Intra- and Inter-category Relation Modeling approach, which simultaneously extracts the intra-category structural information and inter-category relation for more precise reasoning. Specifically, we exploit a Hypergraph-guided Structure Learning (HSL) module, which employs hypergraphs to capture high-order structural relations, transcending traditional graph-based methods that are limited to pairwise linkages. This advancement allows the model to adapt to significant intra-category variations. Additionally, we propose an Inter-category Relation Perception (IRP) module to improve feature discrimination across categories by extracting and analyzing semantic relations among them. Our objective is to alleviate the robustness issue associated with exclusive reliance on intra-category discriminative features. Furthermore, a random semantic consistency (RSC) loss is introduced to direct the model's attention to commonly overlooked yet distinctive regions, indirectly enhancing the representation ability of both HSL and IRP modules. Both qualitative and quantitative results demonstrate the effectiveness and usefulness of HI2R.
Qiangchang Wang, Yilong Yin
ACM Multimedia2
2024 KNN Transformer with Pyramid Prompts for Few-Shot Learning
abstract
Few-Shot Learning (FSL) aims to recognize new classes with limited labeled data. Recent studies have attempted to address the challenge of rare samples with textual prompts to modulate visual features. However, they usually struggle to capture complex semantic relationships between textual and visual features. Moreover, vanilla self-attention is heavily affected by useless information in images, severely constraining the potential of semantic priors in FSL due to the confusion of numerous irrelevant tokens during interaction. To address these aforementioned issues, a K-NN Transformer with Pyramid Prompts (KTPP) is proposed to select discriminative information with K-NN Context Attention (KCA) and adaptively modulate visual features with Pyramid Cross-modal Prompts (PCP). First, for each token, the KCA only selects the K most relevant tokens to compute the self-attention matrix and incorporates the mean of all tokens as the context prompt to provide the global context in three cascaded stages. As a result, irrelevant tokens can be progressively suppressed. Secondly, pyramid prompts are introduced in the PCP to emphasize visual features via interactions between text-based class-aware prompts and multi-scale visual features. This allows the ViT to dynamically adjust the importance weights of visual features based on rich semantic information at different scales, making models robust to spatial variations. Finally, augmented visual features and class-aware prompts are interacted via the KCA to extract class-specific features. Consequently, our model further enhances noise-free visual representations via deep cross-modal interactions, extracting generalized visual representation in scenarios with few labeled samples. Extensive experiments on four benchmark datasets demonstrate significant gains over the state-of-the-art methods, especially for the 1-shot task with 2.28% improvement on average due to semantically enhanced visual representations.
Wenhao Li 0011, Qiangchang Wang, Peng Zhao 0016, Yilong Yin
ACM Multimedia2
2024 Cascaded Cross-modal Alignment for Visible-Infrared Person Re-Identification
Qiangchang Wang, Xinxin Zhang 0004, Yilong Yin
Knowl. Based Syst.2
2024 Spatio-Temporal Multi-Image Reflection Removal
abstract
In this letter, we propose a precise algorithm to eliminate reflections from two images by utilizing temporal and spatial priors. For the temporal prior, we compute the motion information between reflection layers in the two input reflection-contaminated images. Different from numerous popular multi-image reflection removal methods, our proposed algorithm does not assume that two input images are captured under similar lighting conditions and the same camera settings. Furthermore, the proposed algorithm is robust to the difference between the two reflection layers, such as moving objects and different reflections. For the spatial term, a sparsity gradient regularization is adopted to enforce the spatial smoothness of transmission layers and reflection layers. Importantly, the proposed algorithm does not rely on additional training data or high-performance computing devices. Experimental results on both synthetic images and real-world photographs demonstrate that the proposed algorithm achieves State-of-the-Art performance.
Xinxin Zhang 0004, Wenjing Shang, Qiangchang Wang, Yongshun Gong
IEEE Signal Process. Lett.3
2024 Characterizing Hierarchical Semantic-Aware Parts With Transformers for Generalized Zero-Shot Learning
abstract
This paper presents a novel Transformer architecture for zero-shot learning (ZSL), termed TransZSL, which can characterize hierarchical semantic-aware parts. It consists of an adaptive token refinement (ATR), a hierarchical token aggregation (HTA), and semantic-aware prototypes (SAP). Firstly, the ViT is used as the backbone that provides comprehensive local information without missing details. To address the different degrees of noise caused by large appearance variations, the ATR is proposed to highlight important tokens and suppress useless ones adaptively. However, due to the complex image structure, some important tokens may be incorrectly discarded. Therefore, a random perturbation is proposed to reactivate discarded tokens randomly, reducing the risk of missing discriminative information. Secondly, dataset descriptions contain both low- and high-level attributes. To this end, the HTA aggregates complementary hierarchical tokens from multiple ViT layers. Thirdly, semantically similar content may be distributed in different tokens. To overcome this issue, the SAP is proposed to group semantically identical tokens into one prototype, focusing on semantic-aware parts. Besides, diversity loss is used to encourage networks to learn diverse prototypes that discover diverse parts. Both qualitative and quantitative results on several challenging tasks demonstrate the usefulness and effectiveness of our proposed methods.
Peng Zhao 0016, Xiaoming Xi, Qiangchang Wang, Yilong Yin
IEEE Trans. Circuits Syst. Video Technol.3
2023 M3R: Masked Token Mixup and Cross-Modal Reconstruction for Zero-Shot Learning
abstract
In the zero-shot learning (ZSL), learned representation spaces are often biased toward seen classes, thus limiting the ability to predict previously unseen classes. In this paper, we propose Masked token Mixup and cross-Modal Reconstruction for zero-shot learning, termed as M3R, which can significantly alleviate the bias toward seen classes. The M3R mainly consists of Random Token Mixup (RTM), Unseen Class Detection (UCD), and Hard Cross-modal Reconstruction (HCR). Firstly, mappings without proper adaptations to unseen classes would cause the bias toward seen classes. To address this issue, the RTM is introduced to generate diverse unseen class agents, thereby broadening the representation space to cover unknown classes. It is applied at a randomly selected layer in the Vision Transformer, producing smooth low- and high-level representation space boundaries to cover rich attributes. Secondly, it should be noted that unseen class agents generated by the RTM may be mixed with seen class samples. To overcome this challenge, the UCD is designed to generate greater entropy values for unseen classes, thereby distinguishing seen classes from unseen classes. Thirdly, to further mitigate the bias toward seen classes and explore associations between semantics and visual images, the HCR is proposed, which can reconstruct masked pixels based on few discriminative tokens and attribute embeddings. This approach can enable models to have a deep understanding of image contents and build powerful connections between semantic attributes and visual information. Both qualitative and quantitative results demonstrate the effectiveness and usefulness of our proposed M3R model.
Peng Zhao 0016, Qiangchang Wang, Yilong Yin
ACM Multimedia2
2023 Vision Transformer With Attentive Pooling for Robust Facial Expression Recognition
abstract
Facial Expression Recognition (FER) in the wild is an extremely challenging task. Recently, some Vision Transformers (ViT) have been explored for FER, but most of them perform inferiorly compared to Convolutional Neural Networks (CNN). This is mainly because the new proposed modules are difficult to converge well from scratch due to lacking inductive bias and easy to focus on the occlusion and noisy areas. TransFER, a representative transformer-based method for FER, alleviates this with multi-branch attention dropping but brings excessive computations. On the contrary, we present two attentive pooling (AP) modules to pool noisy features directly. The AP modules include Attentive Patch Pooling (APP) and Attentive Token Pooling (ATP). They aim to guide the model to emphasize the most discriminative features while reducing the impacts of less relevant features. The proposed APP is employed to select the most informative patches on CNN features, and ATP discards unimportant tokens in ViT. Being simple to implement and without learnable parameters, the APP and ATP intuitively reduce the computational cost while boosting the performance by ONLY pursuing the most discriminative features. Qualitative results demonstrate the motivations and effectiveness of our attentive poolings. Besides, quantitative results on six in-the-wild datasets outperform other state-of-the-art methods.
Fanglei Xue, Qiangchang Wang, Zichang Tan, Zhongsong Ma, Guodong Guo
IEEE Trans. Affect. Comput.2
2022 CQA-Face: Contrastive Quality-Aware Attentions for Face Recognition
abstract
Few existing face recognition (FR) models take local representations into account. Although some works achieved this by extracting features on cropped parts around face landmarks, landmark detection may be inaccurate or even fail in some extreme cases. Recently, without relying on landmarks, attention-based networks can focus on useful parts automatically. However, there are two issues: 1) It is noticed that these approaches focus on few facial parts, while missing other potentially discriminative regions. This can cause performance drops when emphasized facial parts are invisible under heavy occlusions (e.g. face masks) or large pose variations; 2) Different facial parts may appear at various quality caused by occlusion, blur, or illumination changes. In this paper, we propose contrastive quality-aware attentions, called CQA-Face, to address these two issues. First, a Contrastive Attention Learning (CAL) module is proposed, pushing models to explore comprehensive facial parts. Consequently, more useful parts can help identification if some facial parts are invisible. Second, a Quality-Aware Network (QAN) is developed to emphasize important regions and suppress noisy parts in a global scope. Thus, our CQA-Face model is developed by integrating the CAL with QAN, which extracts diverse quality-aware local representations. It outperforms the state-of-the-art methods on several benchmarks, demonstrating its effectiveness and usefulness.
Qiangchang Wang, Guodong Guo
AAAI1
2022 Learning Multi-Granularity Temporal Characteristics for Face Anti-Spoofing
abstract
Face anti-spoofing (FAS) is essential for securing face recognition systems. Despite the decent performance, few existing works fully leverage temporal information. This would inevitably lead to inferior performance because real and fake faces tend to share highly similar spatial appearances, while important temporal features between consecutive frames are neglected. In this work, we propose a temporal transformer network (TTN) to learn multi-granularity temporal characteristics for FAS. It mainly consists of temporal difference attentions (TDA), a pyramid temporal aggregation (PTA), and a temporal depth difference loss (TDL). Firstly, the vision transformer (ViT) is used as the backbone where comprehensive local patches are utilized to provide subtle differences between live and spoof faces. Then, instead of learning temporal features on global faces which may miss some important local cues, the TDA is developed to extract motion-sensitive cues on each of the comprehensive local patches. Moreover, the TDA is inserted into different layers of the ViT, learning multi-scale motion-sensitive local cues to improve the FAS performance. Secondly, it is observed that different subjects may have different visual tempos in some actions, making it necessary to model different temporal speeds. Our PTA aggregates temporal features at various tempos, which could build short-range and long-range relations among multiple frames. Thirdly, depth maps for real parts may change continuously, while they remain zeros for spoof regions. In order to locate motion features on facial parts, the TDL is proposed to guide the network to locate spoof facial parts where motion patterns between neighboring frames are set as the ground truth. To the best of our knowledge, this work is the first attempt to learn temporal characteristics via transformers. Both qualitative and quantitative results on several challenging tasks demonstrate the usefulness and effectiveness of our proposed methods.
Qiangchang Wang, Weihong Deng, Guodong Guo
IEEE Trans. Inf. Forensics Secur.2
2021 TransFER: Learning Relation-aware Facial Expression Representations with Transformers
abstract
Facial expression recognition (FER) has received increasing interest in computer vision. We propose the Trans-FER model which can learn rich relation-aware local representations. It mainly consists of three components: Multi-Attention Dropping (MAD), ViT-FER, and Multi-head Self-Attention Dropping (MSAD). First, local patches play an important role in distinguishing various expressions, however, few existing works can locate discriminative and diverse local patches. This can cause serious problems when some patches are invisible due to pose variations or viewpoint changes. To address this issue, the MAD is proposed to randomly drop an attention map. Consequently, models are pushed to explore diverse local patches adaptively. Second, to build rich relations between different local patches, the Vision Transformers (ViT) are used in FER, called ViT-FER. Since the global scope is used to reinforce each local patch, a better representation is obtained to boost the FER performance. Thirdly, the multi-head self-attention allows ViT to jointly attend to features from different information subspaces at different positions. Given no explicit guidance, however, multiple self-attentions may extract similar relations. To address this, the MSAD is proposed to randomly drop one self-attention module. As a result, models are forced to learn rich relations among diverse local patches. Our proposed TransFER model outperforms the state-of-the-art methods on several FER benchmarks, showing its effectiveness and usefulness.
Fanglei Xue, Qiangchang Wang, Guodong Guo
ICCV2
2021 DSA-Face: Diverse and Sparse Attentions for Face Recognition Robust to Pose Variation and Occlusion
abstract
Learning local representations is important for face recognition (FR). Recent attention-based networks emphasize few facial parts, while ignoring other potentially discriminative ones. This is more serious when there are large pose variations, occlusions (e.g. face masks), or other image quality changes. To address this, we propose Diverse and Sparse Attentions, called DSA-Face. First, a divergence loss is designed to explicitly encourage the diversity among multiple attention maps by maximizing the Euclidean distance between every pair attention maps. As a result, a Pairwise Self-Contrastive Attention (PSCA) is developed to locate diverse facial parts which provide comprehensive descriptions. Second, an Attention Sparsity Loss (ASL) is proposed to encourage sparse responses in attention maps where only discriminative parts are emphasized while distracted regions (e.g. background or face masks) are discouraged. Built upon the PSCA and ASL, the DSA-Face model is developed to learn diverse and sparse attentions, which can extract diverse discriminative local representations and suppress the focus on noisy regions. Due to the pandemic of the COVID-19, the task of masked face matching is now very important, and our model can handle this much better than previous methods, demonstrating its effectiveness and usefulness. Moreover, our model outperforms the state-of-the-art methods on several other FR benchmarks, showing that it is also general to address various challenges in FR.
Qiangchang Wang, Guodong Guo
IEEE Trans. Inf. Forensics Secur.1
2021 AAN-Face: Attention Augmented Networks for Face Recognition
abstract
Convolutional neural networks are capable of extracting powerful representations for face recognition. However, they tend to suffer from poor generalization due to imbalanced data distributions where a small number of classes are over-represented (e.g. frontal or non-occluded faces) and some of the remaining rarely appear (e.g. profile or heavily occluded faces). This is the reason why the performance is dramatically degraded in minority classes. For example, this issue is serious for recognizing masked faces in the scenario of ongoing pandemic of the COVID-19. In this work, we propose an Attention Augmented Network, called AAN-Face, to handle this issue. First, an attention erasing (AE) scheme is proposed to randomly erase units in attention maps. This well prepares models towards occlusions or pose variations. Second, an attention center loss (ACL) is proposed to learn a center for each attention map, so that the same attention map focuses on the same facial part. Consequently, discriminative facial regions are emphasized, while useless or noisy ones are suppressed. Third, the AE and the ACL are incorporated to form the AAN-Face. Since the discriminative parts are randomly removed by the AE, the ACL is encouraged to learn different attention centers, leading to the localization of diverse and complementary facial parts. Comprehensive experiments on various test datasets, especially on masked faces, demonstrate that our AAN-Face models outperform the state-of-the-art methods, showing the importance and effectiveness.
Qiangchang Wang, Guodong Guo
IEEE Trans. Image Process.1
2020 Hierarchical Pyramid Diverse Attention Networks for Face Recognition
abstract
Deep learning has achieved a great success in face recognition (FR), however, few existing models take hierarchical multi-scale local features into consideration. In this work, we propose a hierarchical pyramid diverse attention (HPDA) network. First, it is observed that local patches would play important roles in FR when the global face appearance changes dramatically. Some recent works apply attention modules to locate local patches automatically without relying on face landmarks. Unfortunately, without considering diversity, some learned attentions tend to have redundant responses around some similar local patches, while neglecting other potential discriminative facial parts. Meanwhile, local patches may appear at different scales due to pose variations or large expression changes. To alleviate these challenges, we propose a pyramid diverse attention (PDA) to learn multi-scale diverse local representations automatically and adaptively. More specifically, a pyramid attention is developed to capture multi-scale features. Meanwhile, a diverse learning is developed to encourage models to focus on different local patches and generate diverse local features. Second, almost all existing models focus on extracting features from the last convolutional layer, lacking of local details or small-scale face parts in lower layers. Instead of simple concatenation or addition, we propose to use a hierarchical bilinear pooling (HBP) to fuse information from multiple layers effectively. Thus, the HPDA is developed by integrating the PDA into the HBP. Experimental results on several datasets show the effectiveness of the HPDA, compared to the state-of-the-art methods.
Qiangchang Wang, He Zheng, Guodong Guo
CVPR1
2020 Face presentation attack detection in mobile scenarios: A comprehensive evaluation
Shan Jia, Guodong Guo, Zhengquan Xu, Qiangchang Wang
Image Vis. Comput.4
2020 LS-CNN: Characterizing Local Patches at Multiple Scales for Face Recognition
abstract
Faces in the wild may contain pose variations, age changes, and with different qualities which significantly enlarge the intra-class variations. Although great progresses have been made in face recognition, few existing works could learn local and multi-scale representations together. In this work, we propose a new model, called Local and multi-Scale Convolutional Neural Networks (LS-CNN). First, since similar discriminative face regions may occur at different scales, it is necessary to learn multi-scale features. To this aim, we introduce a new backbone network, namely Harmonious multi-Scale Network (HSNet), which extracts rich multi-scale features from two harmonious perspectives: utilization of different kernel sizes in a single layer, and concatenation of multi-scale feature maps from different layers. Second, identifying similar local patches is important when global face appearances have dramatic changes. Meanwhile, different face regions have different discriminative abilities. To capture critical local similarities and weigh adaptively on different local patches, a spatial attention is proposed. Third, channels have different convolutional kernels which can detect different features with various importance. Besides, hierarchical channels concatenated from different layers contain diverse information: channels from low layers describe local details or small-scale parts, and channels in high layers represent high-level abstraction or large-scale parts. To emphasize important channels and suppress less informative ones automatically, channel attention is used. Due to the complementary characteristics of channel attention and spatial attention, they are fused to form the Dual Face Attentions (DFA). To the best of our knowledge, this is the first effort to employ attentions for the general face recognition task. The LS-CNN is developed by incorporating DFA into HSNet model. Experimental results on various face matching tasks show its capability of learning complex data distributions.
Qiangchang Wang, Guodong Guo
IEEE Trans. Inf. Forensics Secur.1
2019 Benchmarking deep learning techniques for face recognition
Qiangchang Wang, Guodong Guo
J. Vis. Commun. Image Represent.1
2018 Multiscale Rotation-Invariant Convolutional Neural Networks for Lung Texture Classification
abstract
We propose a new multiscale rotation-invariant convolutional neural network (MRCNN) model for classifying various lung tissue types on high-resolution computed tomography. MRCNN employs Gabor-local binary pattern that introduces a good property in image analysis-invariance to image scales and rotations. In addition, we offer an approach to deal with the problems caused by imbalanced number of samples between different classes in most of the existing works, accomplished by changing the overlapping size between the adjacent patches. Experimental results on a public interstitial lung disease database show a superior performance of the proposed method to state of the art.
Qiangchang Wang, Yuanjie Zheng, Gongping Yang 0001, Weidong Jin, Xinjian Chen 0001, Yilong Yin
IEEE J. Biomed. Health Informatics1