VLDB 2026 Research / reviewers in the wild / expert
Jie Zhang 0071
dblp:84/6889-71
· DBLP profile ↗
65ranked-venue papers
11as first author
48since 2021 · last 2026
0000-0002-8899-3996ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 6 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 5 first-author · 21 since 2021Security and privacy · 9 · 3 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMsabstractDespite rapid progress, Video Large Language Models (Video-LLMs) remain unreliable due to hallucinations, which are outputs that contradict either video evidence (faithfulness) or verifiable world knowledge (factuality).Existing benchmarks provide limited coverage of factuality hallucinations and predominantly evaluate models only in clean settings.We introduce INFACT, a diagnostic benchmark comprising 9,800 QA instances with fine-grained taxonomies for faithfulness and factuality, spanning real and synthetic videos.INFACT evaluates models in four modes: Base (clean), Visual Degradation, Evidence Corruption, and Temporal Intervention for order-sensitive items.Reliability under induced modes is quantified using Resist Rate (RR) and Temporal Sensitivity Score (TSS).Experiments on 14 representative Video-LLMs reveal that higher Base-mode accuracy does not reliably translate to higher reliability in the induced modes, with evidence corruption reducing stability and temporal intervention yielding the largest degradation.Notably, many open-source baselines exhibit nearzero TSS on factuality, indicating pronounced temporal inertia on order-sensitive questions. Junqi Yang, Yuecong Min, Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001 |
ACL (1) | 3 |
| 2026 | Leveraging auxiliary-tasks for height and weight estimation with pose-disentanglement
Jie Zhang 0071, Shiguang Shan |
Frontiers Comput. Sci. | 2 |
| 2026 | A Survey of Multimodal Hallucination Evaluation and Detection
Yuecong Min, Jie Zhang 0071, Bei Yan, Shiguang Shan |
Int. J. Comput. Vis. | 3 |
| 2026 | Anonymization Prompt Learning for Facial Privacy-Preserving Text-to-Image Generation
Liang Shi 0002, Jie Zhang 0071, Shiguang Shan |
Int. J. Comput. Vis. | 2 |
| 2026 | Revisiting Face Forgery Detection: From Facial Representation to Forgery DetectionabstractFace Forgery Detection (FFD), or Deepfake detection, aims to determine whether a digital face is real or fake. Due to different face synthesis algorithms with diverse forgery patterns, FFD models often overfit specific patterns in training datasets, resulting in poor generalization to other unseen forgeries. Existing FFD methods primarily leverage pre-trained backbones with general image representation capabilities and fine-tune them to identify facial forgery cues. However, these backbones lack domain-specific facial knowledge and insufficiently capture complex facial features, thus hindering effective implicit forgery cue identification and limiting generalization. Therefore, it is essential to revisit FFD workflow across the pre-training and fine-tuning stages, achieving an elaborate integration from facial representation to forgery detection to improve generalization. Specifically, we develop an FFD-specific pre-trained backbone with superior facial representation capabilities through self-supervised pre-training on real faces. We then propose a competitive fine-tuning framework that stimulates the backbone to identify implicit forgery cues through a competitive learning mechanism. Moreover, we devise a threshold optimization mechanism that utilizes prediction confidence to improve the inference reliability. Comprehensive experiments demonstrate that our method achieves excellent performance in FFD and extra face-related tasks, i.e., presentation attack detection. Zonghui Guo, Jie Zhang 0071, Haiyong Zheng, Shiguang Shan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | VLBiasBench: A Comprehensive Benchmark for Evaluating Bias in Large Vision-Language ModelabstractThe emergence of Large Vision-Language Models (LVLMs) marks significant strides towards achieving general artificial intelligence. However, these advancements are accompanied by concerns about biased outputs, a challenge that has yet to be thoroughly explored. Existing benchmarks are not sufficiently comprehensive in evaluating biases due to their limited data scale, single questioning format and narrow sources of bias. To address this problem, we introduce VLBiasBench, a comprehensive benchmark designed to evaluate biases in LVLMs. VLBiasBench features a dataset that covers nine distinct categories of social biases, including age, disability status, gender, nationality, physical appearance, race, religion, profession, social economic status, as well as two intersectional bias categories: race × gender and race × social economic status. To build a large-scale dataset, we use Stable Diffusion XL model to generate 46,848 high-quality images, which are combined with various questions to create 128,342 samples. These questions are divided into open-ended and close-ended types, ensuring thorough consideration of bias sources and a comprehensive evaluation of LVLM biases from multiple perspectives. We conduct extensive evaluations on 15 open-source models as well as two advanced closed-source models, yielding new insights into the biases present in these models. Sibo Wang 0012, Xiangkui Cao, Jie Zhang 0071, Zheng Yuan 0005, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Dynamic Attention Analysis for Backdoor Detection in Text-to-Image Diffusion ModelsabstractRecent studies have revealed that text-to-image diffusion models are vulnerable to backdoor attacks, where attackers implant stealthy textual triggers to manipulate model outputs. Previous backdoor detection methods primarily focus on the static features of backdoor samples. However, a vital property of diffusion models is their inherent dynamism. This study introduces a novel backdoor detection perspective named Dynamic Attention Analysis (DAA), showing that these dynamic characteristics serve as better indicators for backdoor detection. Specifically, by examining the dynamic evolution of cross-attention maps, we observe that backdoor samples exhibit distinct feature evolution patterns at the $< $ $> token compared to benign samples. To quantify these dynamic anomalies, we first introduce DAA-I, which treats the tokens' attention maps as spatially independent and measures dynamic feature using the Frobenius norm. Furthermore, to better capture the interactions between attention maps and refine the feature, we propose a dynamical system-based approach, referred to as DAA-S. This model formulates the spatial correlations among attention maps using a graph-based state equation and we theoretically analyze the global asymptotic stability of this method. Extensive experiments across six representative backdoor attack scenarios demonstrate that our approach significantly surpasses existing detection methods, achieving an average F1 Score of 79.27% and an AUC of 86.27%. Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Toward Transferable Defense Against Malicious Image EditsabstractRecent approaches employing imperceptible perturbations in input images have demonstrated promising potential to counter malicious manipulations in diffusion-based image editing systems. However, existing methods suffer from limited transferability in cross-model evaluations. To address this, we propose Transferable Defense Against Malicious Image Edits (TDAE), a novel bimodal framework that enhances image immunity against malicious edits through coordinated image-text optimization. Specifically, at the visual defense level, we introduce FlatGrad Defense Mechanism (FDM), which incorporates gradient regularization into the adversarial objective. By explicitly steering the perturbations toward flat minima, FDM amplifies immune robustness against unseen editing models. For textual enhancement protection, we propose an adversarial optimization paradigm named Dynamic Prompt Defense (DPD), which periodically refines text embeddings to align the editing outcomes of immunized images with those of the original images, then updates the images under optimized embeddings. Through iterative adversarial updates to diverse embeddings, DPD enforces the generation of immunized images that seek a broader set of immunity-enhancing features, thereby achieving cross-model transferability. Extensive experimental results demonstrate that our TDAE achieves state-of-the-art performance in mitigating malicious edits under both intra- and cross-model evaluations. Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | MM-MoralBench: A multimodal moral evaluation benchmark for large vision-language models
Bei Yan, Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001 |
Pattern Recognit. | 2 |
| 2026 | BIMM: Brain-Inspired Masked Modeling for Video Representation LearningabstractThe visual pathway of human brain includes two sub-pathways,i.e., the ventral pathway and the dorsal pathway, which focus on object identification and dynamic information modeling, respectively. Both pathways comprise multi-layer structures, with each layer responsible for processing different aspects of visual information. Inspired by the human visual information processing mechanism, we propose the Brain Inspired Masked Modeling (BIMM) framework, aiming to learn comprehensive representations from videos. Specifically, our approach consists of ventral and dorsal branches, which learn image and video representations, respectively. Both branches employ the Vision Transformer (ViT) as their backbone and are trained through a masked modeling method. To emulate the distinct functions of the visual cortices, we segment the encoder of each branch into three intermediate blocks and reconstruct progressive prediction targets with light weight decoders. Furthermore, drawing inspiration from the information-sharing mechanism in the brain’s visual pathways, we introduce a partial parameter sharing strategy between the branches during training. Extensive experiments demonstrate that BIMM achieves superior performance compared to the state-of-the-art methods. Jie Zhang 0071, Zhifan Wan, Sen Nie, Changzhen Li, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Dual Attention Guided Defense Against Malicious EditsabstractRecent progress in text-to-image diffusion models has transformed image editing via text prompts, yet this also introduces significant ethical challenges from potential misuse in creating deceptive or harmful content. While current defenses seek to mitigate this risk by embedding imperceptible perturbations, their effectiveness is limited against malicious tampering. To address this issue, we propose a Dual Attention-Guided Noise Perturbation (DANP) immunization method that adds imperceptible perturbations to disrupt the model’s semantic understanding and generation process. DANP functions over multiple timesteps to manipulate both cross-attention maps and the noise prediction process, using a dynamic threshold to generate masks that identify text-relevant and irrelevant regions. It then reduces attention in relevant areas while increasing it in irrelevant ones, thereby misguides the edit towards incorrect regions and preserves the intended targets. Additionally, our method maximizes the discrepancy between the injected noise and the model’s predicted noise to further interfere with the generation. By targeting both attention and noise prediction mechanisms, DANP exhibits impressive immunity against malicious edits, and extensive experiments confirm that our method achieves state-of-the-art performance. Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2026 | Trigger Without Trace: Toward Stealthy Backdoor Attack on Text-to-Image Diffusion ModelsabstractBackdoor attacks targeting text-to-image diffusion models have advanced rapidly. However, current backdoor samples often exhibit two key abnormalities compared to benign samples: 1) Semantic Consistency, where backdoor prompts tend to generate images with similar semantic content even with significant textual variations to the prompts; 2) Attention Consistency, where the trigger induces consistent structural responses in the crossattention maps. These consistencies leave detectable traces for defenders, making backdoors easier to identify. In this paper, toward stealthy backdoor samples, we propose Trigger without Trace (TwT) by explicitly mitigating these consistencies. Specifically, our approach leverages syntactic structures as backdoor triggers to amplify the sensitivity to textual variations, effectively breaking down the semantic consistency. Besides, a regularization method based on Kernel Maximum Mean Discrepancy (KMMD) is proposed to align the distribution of cross-attention responses between backdoor and benign samples, thereby disrupting attention consistency. Extensive experiments demonstrate that our method achieves a 97.5% attack success rate while exhibiting stronger resistance to defenses. It achieves an average of over 98% backdoor samples bypassing three state-of-the-art detection mechanisms, revealing the vulnerabilities of current backdoor defense methods. The code is available at https://github.com/Robin-WZQ/TwT. Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Face Forgery Video Detection via Temporal Forgery Cue UnravelingabstractFace Forgery Video Detection (FFVD) is a critical yet challenging task in determining whether a digital facial video is authentic or forged. Existing FFVD methods typically focus on isolated spatial or coarsely fused spatiotemporal information, failing to leverage temporal forgery cues thus resulting in unsatisfactory performance. We strive to unravel these cues across three progressive levels: momentary anomaly, gradual inconsistency, and cumulative distortion. Accordingly, we design a consecutive correlate module to capture momentary anomaly cues by correlating interactions among consecutive frames. Then, we devise a future guide module to unravel inconsistency cues by iteratively aggregating historical anomaly cues and gradually propagating them into future frames. Finally, we introduce a historical review module that unravels distortion cues via momentum accumulation from future to historical frames. These three modules form our Temporal Forgery Cue Unraveling (TFCU) framework, sequentially highlighting spatial discriminative features by unraveling temporal forgery cues bidirectionally between historical and future frames. Extensive experiments and ablation studies demonstrate the effectiveness of our TFCU method, achieving state-of-the-art performance across diverse unseen datasets and manipulation methods. Code is available at https://github.com/zhenglab/TFCU. Zonghui Guo, Jie Zhang 0071, Haiyong Zheng, Shiguang Shan |
CVPR | 3 |
| 2025 | Evaluating Cognitive-Behavioral Fixation via Multimodal User Viewing Patterns on Social MediaabstractDigital social media platforms frequently contribute to cognitive-behavioral fixation, a phenomenon in which users exhibit sustained and repetitive engagement with narrow content domains. While cognitive-behavioral fixation has been extensively studied in psychology, methods for computationally detecting and evaluating such fixation remain underexplored. To address this gap, we propose a novel framework for assessing cognitive-behavioral fixation by analyzing users’ multimodal social media engagement patterns. Specifically, we introduce a multimodal topic extraction module and a cognitive-behavioral fixation quantification module that collaboratively enable adaptive, hierarchical, and interpretable assessment of user behavior. Experiments on existing benchmarks and a newly curated multimodal dataset demonstrate the effectiveness of our approach, laying the groundwork for scalable computational analysis of cognitive fixation. All code in this project is publicly available for research purposes at https://github.com/Liskie/cognitive-fixation-evaluation. Yunwei Zhao, Shiguang Shan, Jie Zhang 0071 |
EMNLP | 6 |
| 2025 | Dual-Branch Partial Annotation Learning for Facial Attributes RecognitionabstractFacial attribute recognition (FAR) aims to identify the attributes of a given face image. As a multi-label classification problem, conventional methods typically rely on large-scale fully annotated datasets. However, annotating all the facial attributes extensively is challenging and expensive, as totally tens of facial attributes can be defined and many of them are subtle or even vague thus requiring expertise for annotation. In contrast, it is much easier to annotate few (even one) most prominent attributes per face image, thus resulting in partially annotated dataset. To fully leverage this kind of datasets, we propose a novel pseudo-label based method named Dual-Branch Partial Annotation Learning (DB-PAL), in which two predicting branches respectively generate positive and negative annotations via loss-based ranking and validate each other to obtain better pseudo labels for training set augmentation. Extensive experiments on the CelebA and LFWA datasets demonstrate the superiority of our method. Jie Zhang 0071, Shiguang Shan |
FG | 3 |
| 2025 | Dysca: A Dynamic and Scalable Benchmark for Evaluating Perception Ability of LVLMsabstractCurrently many benchmarks have been proposed to evaluate the perception ability of the Large Vision-Language Models (LVLMs).
However, most benchmarks conduct questions by selecting images from existing datasets, resulting in the potential data leakage. Besides, these benchmarks merely focus on evaluating LVLMs on the realistic style images and clean scenarios, leaving the multi-stylized images and noisy scenarios unexplored. In response to these challenges, we propose a dynamic and scalable benchmark named Dysca for evaluating LVLMs by leveraging synthesis images. Specifically, we leverage Stable Diffusion and design a rule-based method to dynamically generate novel images, questions and the corresponding answers. We consider 51 kinds of image styles and evaluate the perception capability in 20 subtasks. Moreover, we conduct evaluations under 4 scenarios (i.e., Clean, Corruption, Print Attacking and Adversarial Attacking) and 3 question types (i.e., Multi-choices, True-or-false and Free-form). Thanks to the generative paradigm, Dysca serves as a scalable benchmark for easily adding new subtasks and scenarios. A total of 24 advanced open-source LVLMs and 2 close-source LVLMs are evaluated on Dysca, revealing the drawbacks of current LVLMs. The benchmark is released in anonymous github page \url{https://github.com/Benchmark-Dysca/Dysca}. Jie Zhang 0071, Mengqi Lei, Zheng Yuan 0005, Bei Yan, Shiguang Shan, Xilin Chen 0001 |
ICLR | 1 |
| 2025 | SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMsabstractDespite rapid advances, Large Vision-Language Models (LVLMs) still suffer from hallucinations, i.e., generating content inconsistent with input or established world knowledge, which correspond to faithfulness and factuality hallucinations, respectively. Prior studies primarily evaluate faithfulness hallucination at a rather coarse level (e.g., object-level) and lack fine-grained analysis. Additionally, existing benchmarks often rely on costly manual curation or reused public datasets, raising concerns about scalability and data leakage. To address these limitations, we propose an automated data construction pipeline that produces scalable, controllable, and diverse evaluation data. We also design a hierarchical hallucination induction framework with input perturbations to simulate realistic noisy scenarios. Integrating these designs, we construct SHALE, a Scalable HALlucination Evaluation benchmark designed to assess both faithfulness and factuality hallucinations via a fine-grained hallucination categorization scheme. SHALE comprises over 30K image-instruction pairs spanning 12 representative visual perception aspects for faithfulness and 6 knowledge domains for factuality, considering both clean and noisy scenarios. Extensive experiments on over 20 mainstream LVLMs reveal significant factuality hallucinations and high sensitivity to semantic perturbations. Bei Yan, Yuecong Min, Jie Zhang 0071, Shiguang Shan |
ACM Multimedia | 4 |
| 2025 | SafetyQuizzer: Timely and Dynamic Evaluation on the Safety of LLMsabstractZhichao Shi, Shaoling Jing, Yi Cheng, Hao Zhang, Yuanzhuo Wang, Jie Zhang, Huawei Shen, Xueqi Cheng. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zhichao Shi 0001, Shaoling Jing, Hao Zhang 0048, Yuanzhuo Wang, Jie Zhang 0071, Huawei Shen, Xueqi Cheng 0001 |
NAACL (Long Papers) | 6 |
| 2025 | Generalized Face Liveness Detection via De-Fake Face GeneratorabstractPrevious Face Anti-spoofing (FAS) methods face the challenge of generalizing to unseen domains, mainly because most existing FAS datasets are relatively small and lack data diversity. Thanks to the development of face recognition in the past decade, numerous real face images are available publicly, which are however neglected previously by the existing literature. In this paper, we propose an Anomalous cue Guided FAS (AG-FAS) method, which can effectively leverage large-scale additional real faces for improving model generalization via a De-fake Face Generator (DFG). Specifically, by training on a large-scale real face only dataset, the generator obtains the knowledge of what a real face should be like, and thus has the capability of generating a "real" version of any input face image. Consequently, the difference between the input face and the generated "real" face can be treated as cues of attention for the fake feature learning. With the above ideas, an Off-real Attention Network (OA-Net) is proposed which allocates its attention to the spoof region of the input according to the anomalous cue. Extensive experiments on a total of nine public datasets show our method achieves state-of-the-art results under cross-domain evaluations with unseen scenarios and unknown presentation attacks. Besides, we provide theoretical analysis demonstrating the effectiveness of the proposed anomalous cues. Xingming Long, Jie Zhang 0071, Shiguang Shan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Real face foundation representation learning for generalized deepfake detection
Liang Shi 0002, Jie Zhang 0071, Zhilong Ji, Jinfeng Bai, Shiguang Shan |
Pattern Recognit. | 2 |
| 2025 | Confidence Aware Learning for Reliable Face Anti-SpoofingabstractCurrent Face Anti-spoofing (FAS) models tend to make overly confident predictions even when encountering unfamiliar scenarios or unknown presentation attacks, which leads to serious potential risks. To solve this problem, we propose a Confidence Aware Face Anti-spoofing (CA-FAS) model, which is aware of its capability boundary, thus achieving reliable liveness detection within this boundary. To enable the CA-FAS to “know what it doesn’t know”, we propose to estimate its confidence during the prediction of each sample. Specifically, we build Gaussian distributions for both the live faces and the known attacks. The prediction confidence for each sample is subsequently assessed using the Mahalanobis distance between the sample and the Gaussians for the “known data”. We further introduce the Mahalanobis distance-based triplet mining to optimize the parameters of both the model and the constructed Gaussians as a whole. Extensive experiments show that the proposed CA-FAS can effectively recognize samples with low prediction confidence and thus achieve much more reliable performance than other FAS models by filtering out samples that are beyond its reliable range. Xingming Long, Jie Zhang 0071, Shiguang Shan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Collaboratively Self-Supervised Video Representation Learning for Action RecognitionabstractConsidering the close connection between action recognition and human pose estimation, we design a Collaboratively Self-supervised Video Representation (CSVR) learning framework specific to action recognition by jointly factoring in generative pose prediction and discriminative context matching as pretext tasks. Specifically, our CSVR consists of three branches: a generative pose prediction branch, a discriminative context matching branch, and a video generating branch. Among them, the first one encodes dynamic motion feature by utilizing Conditional-GAN to predict the human poses of future frames, and the second branch extracts static context features by contrasting positive and negative video feature and I-frame feature pairs. The third branch is designed to generate both current and future video frames, for the purpose of collaboratively improving dynamic motion features and static context features. Extensive experiments demonstrate that our method achieves state-of-the-art performance on multiple popular video datasets. Jie Zhang 0071, Zhifan Wan, Lanqing Hu, Shuzhe Wu, Shiguang Shan |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | FullLoRA: Efficiently Boosting the Robustness of Pretrained Vision TransformersabstractIn recent years, the Vision Transformer (ViT) model has gradually become mainstream in various computer vision tasks, and the robustness of the model has received increasing attention. However, existing large models tend to prioritize performance during training, potentially neglecting the robustness, which may lead to serious security concerns. In this paper, we establish a new challenge: exploring how to use a small number of additional parameters for adversarial finetuning to quickly and effectively enhance the adversarial robustness of a standardly trained model. To address this challenge, we develop novel LNLoRA module, incorporating a learnable layer normalization before the conventional LoRA module, which helps mitigate magnitude differences in parameters between the adversarial and standard training paradigms. Furthermore, we propose the FullLoRA framework by integrating the learnable LNLoRA modules into all key components of ViT-based models while keeping the pretrained model frozen, which can significantly improve the model robustness via adversarial finetuning in a parameter-efficient manner. Extensive experiments on several datasets demonstrate the superiority of our proposed FullLoRA framework. It achieves comparable robustness with full finetuning while only requiring about 5% of the learnable parameters. This also effectively addresses concerns regarding extra model storage space and enormous training time caused by adversarial finetuning. Zheng Yuan 0005, Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | Pre-Trained Model Guided Fine-Tuning for Zero-Shot Adversarial RobustnessabstractLarge-scale pre-trained vision-language models like CLIP have demonstrated impressive performance across various tasks, and exhibit remarkable zero-shot generalization capability, while they are also vulnerable to impercep-tible adversarial examples. Existing works typically em-ploy adversarial training (fine-tuning) as a defense method against adversarial examples. However, direct application to the CLIP model may result in overfitting, compromising the model's capacity for generalization. In this paper, we propose Pre-trained Model Guided Adversarial Fine-Tuning (PMG-AFT) method, which leverages supervision from the original pre-trained model by carefully designing an auxiliary branch, to enhance the model's zero-shot ad-versarial robustness. Specifically, PMG-AFT minimizes the distance between the features of adversarial examples in the target model and those in the pre-trained model, aiming to preserve the generalization features already captured by the pre-trained model. Extensive Experiments on 15 zero-shot datasets demonstrate that PMG-AFT significantly outper-forms the state-of-the-art method, improving the top-1 ro-bust accuracy by an average of 4.99%. Furthermore, our approach consistently improves clean accuracy by an aver-age of 8.72%. Our code is available at here.1 Sibo Wang 0012, Jie Zhang 0071, Zheng Yuan 0005, Shiguang Shan |
CVPR | 2 |
| 2024 | Video Harmonization with Triplet Spatio-Temporal Variation PatternsabstractVideo harmonization is an important and challenging task that aims to obtain visually realistic composite videos by automatically adjusting the foreground's appearance to harmonize with the background. Inspired by the short-term and long-term gradual adjustment process of manual har-monization, we present a Video Triplet Transformer frame-work to model three spatio-temporal variation patterns within videos, i.e., short-term spatial as well as long-term global and dynamic, for video-to-video tasks like video har-monization. Specifically, for short-term harmonization, we adjust foreground appearance to consist with background in spatial dimension based on the neighbor frames; for long-term harmonization, we not only explore global ap-pearance variations to enhance temporal consistency but also alleviate motion offset constraints to align similar con-textual appearances dynamically. Extensive experiments and ablation studies demonstrate the effectiveness of our method, achieving state-of-the-art performance in video harmonization, video enhancement, and video demoireing tasks. We also propose a temporal consistency metric to better evaluate the harmonized videos. Code is available at https://github.com/zhenglablVideoTripletTransformer. Zonghui Guo, Jie Zhang 0071, Shiguang Shan, Haiyong Zheng |
CVPR | 3 |
| 2024 | T2IShield: Defending Against Backdoors on Text-to-Image Diffusion Models
Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001 |
ECCV (85) | 2 |
| 2024 | Rethinking the Evaluation of Out-of-Distribution Detection: A Sorites ParadoxabstractMost existing out-of-distribution (OOD) detection benchmarks classify samples with novel labels as the OOD data. However, some marginal OOD samples actually have close semantic contents to the in-distribution (ID) sample, which makes determining the OOD sample a Sorites Paradox. In this paper, we construct a benchmark named Incremental Shift OOD (IS-OOD) to address the issue, in which we divide the test samples into subsets with different semantic and covariate shift degrees relative to the ID dataset. The data division is achieved through a shift measuring method based on our proposed Language Aligned Image feature Decomposition (LAID). Moreover, we construct a Synthetic Incremental Shift (Syn-IS) dataset that contains high-quality generated images with more diverse covariate contents to complement the IS-OOD benchmark. We evaluate current OOD detection methods on our benchmark and find several important insights: (1) The performance of most OOD detection methods significantly improves as the semantic shift increases; (2) Some methods like GradNorm may have different OOD detection mechanisms as they rely less on semantic shifts to make decisions; (3) Excessive covariate shifts in the image are also likely to be considered as OOD for some methods. Our code and data are released in https://github.com/qqwsad5/IS-OOD. Xingming Long, Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001 |
NeurIPS | 2 |
| 2024 | Hierarchical compositional representations for few-shot action recognition
Changzhen Li, Jie Zhang 0071, Shuzhe Wu, Xin Jin 0004, Shiguang Shan |
Comput. Vis. Image Underst. | 2 |
| 2024 | Towards Robust Semantic Segmentation against Patch-Based Attack via Attention Refinement
Zheng Yuan 0005, Jie Zhang 0071, Yude Wang, Shiguang Shan, Xilin Chen 0001 |
Int. J. Comput. Vis. | 2 |
| 2024 | Adaptive Perturbation for Adversarial AttackabstractIn recent years, the security of deep learning models achieves more and more attentions with the rapid development of neural networks, which are vulnerable to adversarial examples. Almost all existing gradient-based attack methods use the sign function in the generation to meet the requirement of perturbation budget on$L_\infty$norm. However, we find that the sign function may be improper for generating adversarial examples since it modifies the exact gradient direction. Instead of using the sign function, we propose to directly utilize the exact gradient direction with a scaling factor for generating adversarial perturbations, which improves the attack success rates of adversarial examples even with fewer perturbations. At the same time, we also theoretically prove that this method can achieve better black-box transferability. Moreover, considering that the best scaling factor varies across different images, we propose an adaptive scaling factor generator to seek an appropriate scaling factor for each image, which avoids the computational cost for manually searching the scaling factor. Our method can be integrated with almost all existing gradient-based attack methods to further improve their attack success rates. Extensive experiments on the CIFAR10 and ImageNet datasets show that our method exhibits higher transferability and outperforms the state-of-the-art methods. Zheng Yuan 0005, Jie Zhang 0071, Zhaoyan Jiang, Shiguang Shan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Dual Sampling Based Causal Intervention for Face Anti-Spoofing With Identity DebiasingabstractImproving generalization to unseen scenarios is one of the greatest challenges in Face Anti-spoofing (FAS). Most previous FAS works focus on domain debiasing to eliminate the distribution discrepancy between training and test data. However, a crucial but usually neglected bias factor is the face identity. Generally, the identity distribution varies across the FAS datasets as the participants in these datasets are from different regions, which will lead to serious identity bias in the cross-dataset FAS tasks. In this work, we resort to causal learning and propose Dual Sampling based Causal Intervention (DSCI) for face anti-spoofing, which improves the generalization of the FAS model by eliminating the identity bias. DSCI treats the bias as a confounder and applies the backdoor adjustment through the proposed dual sampling on the face identity and the FAS feature. Specifically, we first sample the data uniformly on the identity distribution that is obtained by a pretrained face recognition model. By feeding the sampled data into a network, we can get an estimated FAS feature distribution and sample the FAS feature on it. Sampling the FAS feature from a complete estimated distribution can include potential counterfactual features in the training, which effectively expands the training data. The dual sampling process helps the model learn the real causality between the FAS feature and the input liveness, allowing the model to perform more stably across various identity distributions. Extensive experiments demonstrate our proposed method outperforms the state-of-the-art methods on both intra- and cross-dataset evaluations. Xingming Long, Jie Zhang 0071, Shuzhe Wu, Xin Jin 0004, Shiguang Shan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Enhancing Face Recognition With Detachable Self-Supervised Bypass NetworksabstractAttributed to the development of deep networks and abundant data, automatic face recognition (FR) has quickly reached human-level capacity in the past few years. However, the FR problem is not perfectly solved in case of large poses and uncontrolled occlusions. In this paper, we propose a novel bypass enhanced representation learning (BERL) method to improve face recognition under unconstrained scenarios. The proposed method integrates self-supervised learning and supervised learning together by attaching two auxiliary bypasses, a 3D reconstruction bypass and a blind inpainting bypass, to assist robust feature learning for face recognition. Among them, the 3D reconstruction bypass enforces the face recognition network to encode pose independent 3D facial information, which enhances the robustness to various poses. The blind inpainting bypass enforces the face recognition network to capture more facial context information for face inpainting, which enhances the robustness to occlusions. The whole framework is trained in end-to-end manner with two self-supervised tasks above and the classic supervised face identification task. During inference, the two auxiliary bypasses can be detached from the face recognition network, avoiding any additional computational overhead. Extensive experimental results on various face recognition benchmarks show that, without any cost of extra annotations and computations, our method outperforms state-of-the-art methods. Moreover, the learnt representations can also well generalize to other face-related downstream tasks such as the facial attribute recognition with limited labeled data. Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Self-supervised Learning for Fine-grained Ethnicity Classification under Limited Labeled DataabstractHuman faces are always determined by genes and other external causes, such as geographical environment, which makes it possible for us to predict ethnicity according to the faces. However, it remains a challenging task due to the tiny differences in faces for various ethnicities, which is hard for human beings to tell, especially for ethnicities on the same continent, e.g., East Asia. Although some strongly-supervised methods have demonstrated their feasibility in this task, they cease to be effective when suffering from data-hungry issues in practice. This paper proposes a novel self-supervised model with a polynomial stacked attention mechanism to well excavate distinctions across different nations under limited labeled data. And we also construct a new ethnicity dataset named Cupid which observably extends the scale and categories of ethnic data compared to the existing datasets. Extensive experiments confirm that our method achieves the state-of-the-art results on both the Asian Face dataset and our proposed Cupid dataset. Kunyan Li, Jie Zhang 0071, Shiguang Shan |
FG | 2 |
| 2023 | Adaptive Adversarial Patch Attack on Face Recognition ModelsabstractFace recognition models have become widely used for identity authentication in scenarios such as cell phone unlocking and financial payment, but they are vulnerable to adversarial examples. Due to the realizability in the physical world, adversarial patch attack has emerged as a significant security threat. However, most existing adversarial patch attack methods focus on only one aspect of patch generation, such as patch location or shape. To overcome this limitation, we propose a novel unified Adaptive Adversarial Patch (AAP) attack framework for targeted attack on face recognition models. Our method comprehensively considers various factors during patch generation, including location, shape, and number. Our approach adaptively selects patch location and number based on saliency map and clustering, while simultaneously deforming patch shape and optimizing perturbations. Extensive experiments under both white-box and black-box settings demonstrate that our proposed method achieves higher attack success rates compared to SOTA methods. Bei Yan, Jie Zhang 0071, Zheng Yuan 0005, Shiguang Shan |
IJCB | 2 |
| 2023 | CCLAP: Controllable Chinese Landscape Painting Generation Via Latent Diffusion ModelabstractWith the development of deep generative models, recent years have seen great success of Chinese landscape painting generation. However, few works focus on controllable Chinese landscape painting generation due to the lack of data and limited modeling capabilities. In this work, we propose a controllable Chinese landscape painting generation method named CCLAP, which can generate painting with specific content and style based on Latent Diffusion Model. Specifically, it consists of two cascaded modules, i.e., content generator and style aggregator. The content generator module guarantees the content of generated paintings specific to the input text. While the style aggregator module is to generate paintings of a style corresponding to a reference image. Moreover, a new dataset of Chinese landscape paintings named CLAP is collected for comprehensive evaluation. Both the qualitative and quantitative results demonstrate that our method achieves state-of-the-art performance, especially in artfully-composed and artistic conception. Codes are available at https://github.com/Robin-WZQ/CCLAP. Jie Zhang 0071, Zhilong Ji, Jinfeng Bai, Shiguang Shan |
ICME | 2 |
| 2023 | Data-Efficient Masked Video Modeling for Self-supervised Action RecognitionabstractRecently, self-supervised video representation learning based on Masked Video Modeling (MVM) has demonstrated promising results for action recognition. However, existing methods face two significant challenges: (1) video actions involve a crucial temporal dimension, yet current masking strategies adopt inefficient random approaches that undermine low-density dynamic motion clues in videos; (2) pre-training requires large-scale datasets and significant computing resources (including large batch sizes and enormous iterations). To address these issues, we propose a novel method named Data-Efficient Masked Video Modeling (DEMVM) for self-supervised action recognition. Specifically, a novel masking strategy named Flow-Guided Dense Masking (FGDM) is proposed to facilitate efficient learning by focusing more on the action-related temporal clues, which applies dense masking to dynamic regions based on optical flow priors, while sparse masking to background regions. Furthermore, DEMVM introduces a 3D video tokenizer to enhance the modeling of temporal clues. Finally, Progressive Masking Ratio (PMR) and 2D initialization strategies are presented to enable the model to adapt to the characteristics of the MVM paradigm during different training stages. Extensive experiments on multiple benchmarks, UCF101, HMDB51, and Mimetics, demonstrate that our method achieves state-of-the-art performance in the downstream action recognition task with both efficient data and low computational cost. More interestingly, the few-shot experiment on the Mimetics dataset shows that DEMVM can accurately recognize actions even in the presence of context bias. Qiankun Li 0004, Xiaolong Huang 0001, Zhifan Wan, Lanqing Hu, Shuzhe Wu, Jie Zhang 0071, Shiguang Shan, Zengfu Wang |
ACM Multimedia | 6 |
| 2023 | BLPSeg: Balance the Label Preference in Scribble-Supervised Semantic SegmentationabstractScribble-supervised semantic segmentation is an appealing weakly supervised technique with low labeling cost. Existing approaches mainly consider diffusing the labeled region of scribble by low-level feature similarity to narrow the supervision gap between scribble labels and mask labels. In this study, we observe an annotation bias between scribble and object mask, i.e., label workers tend to scribble on the spacious region instead of corners. This label preference makes the model learn well on those frequently labeled regions but poor on rarely labeled pixels. Therefore, we propose BLPSeg to balance the label preference for complete segmentation. Specifically, the BLPSeg first predicts an annotation probability map to evaluate the rarity of labels on each image, then utilizes a novel BLP loss to balance the model training by up-weighting those rare annotations. Additionally, to further alleviate the impact of label preference, we design a local aggregation module (LAM) to propagate supervision from labeled to unlabeled regions in gradient backpropagation. We conduct extensive experiments to illustrate the effectiveness of our BLPSeg. Our single-stage method even outperforms other advanced multi-stage methods and achieves state-of-the-art performance. Yude Wang, Jie Zhang 0071, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Enhancing Face Recognition with Self-Supervised 3D ReconstructionabstractAttributed to both the development of deep networks and abundant data, automatic face recognition (FR) has quickly reached human-level capacity in the past few years. However, the FR problem is not perfectly solved in case of uncontrolled illumination and pose. In this paper, we propose to enhance face recognition with a bypass of self-supervised 3D reconstruction, which enforces the neural backbone to focus on the identity-related depth and albedo information while neglects the identity-irrelevant pose and illumination information. Specifically, inspired by the physical model of image formation, we improve the backbone FR network by introducing a 3D face reconstruction loss with two auxiliary networks. The first one estimates the pose and illumination from the input face image while the second one decodes the canonical depth and albedo from the intermediate feature of the FR backbone network. The whole network is trained in end-to-end manner with both classic face identification loss and the loss of 3D face reconstruction with the physical parameters. In this way, the self-supervised reconstruction acts as a regularization that enables the recognition network to understand faces in 3D view, and the learnt features are forced to encode more information of canonical facial depth and albedo, which is more intrinsic and beneficial to face recognition. Extensive experimental results on various face recognition benchmarks show that, without any cost of extra annotations and computations, our method outperforms state-of-the-art ones. Moreover, the learnt representations can also well generalize to other face-related downstream tasks such as the facial attribute recognition with limited labeled data. Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2022 | Adaptive Image Transformations for Transfer-Based Adversarial Attack
Zheng Yuan 0005, Jie Zhang 0071, Shiguang Shan |
ECCV (5) | 2 |
| 2022 | Polynomial stacked-attention network for nationality classification
Kunyan Li, Jie Zhang 0071, Shiguang Shan |
Frontiers Comput. Sci. | 2 |
| 2022 | Learning pseudo labels for semi-and-weakly supervised semantic segmentation
Yude Wang, Jie Zhang 0071, Meina Kan, Shiguang Shan |
Pattern Recognit. | 2 |
| 2022 | Dual-Branch Meta-Learning Network With Distribution Alignment for Face Anti-SpoofingabstractExisting face anti-spoofing (FAS) methods fail to generalize well to unseen domains with different data distribution from the training domains, due to the distribution discrepancies between various domains. To extract domain-invariant features for unseen domains, this work proposes a Dual-Branch Meta-learning Network (DBMNet) with distribution alignment for face anti-spoofing. Specifically, DBMNet consists of a feature embedding (FE) branch and a depth estimating (DE) branch for real and fake face discrimination. Each branch acts as a meta-learner and is optimized by step-adjusted meta-learning that can adaptively select the best number of meta-train steps. In order to mitigate distribution discrepancies between domains, we introduce two distribution alignment losses to directly regularize the two meta-learners,i.e., the triplet loss for FE branch and the depth loss for DE branch, respectively. Both of them are designed as part of the meta-train and meta-test objectives, which contribute to higher-order derivatives on the parameters during the meta-optimization for further seeking domain-invariant features. Extensive ablation studies and comparisons with the state-of-the-art methods show the effectiveness of our method for better generalization. Yunpei Jia, Jie Zhang 0071, Shiguang Shan |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | Locality-Aware Channel-Wise Dropout for Occluded Face RecognitionabstractFace recognition remains a challenging task in unconstrained scenarios, especially when faces are partially occluded. To improve the robustness against occlusion, augmenting the training images with artificial occlusions has been proved as a useful approach. However, these artificial occlusions are commonly generated by adding a black rectangle or several object templates including sunglasses, scarfs and phones, which cannot well simulate the realistic occlusions. In this paper, based on the argument that the occlusion essentially damages a group of neurons, we propose a novel and elegant occlusion-simulation method via dropping the activations of a group of neurons in some elaborately selected channel. Specifically, we first employ a spatial regularization to encourage each feature channel to respond to local and different face regions. Then, the locality-aware channel-wise dropout (LCD) is designed to simulate occlusions by dropping out a few feature channels. The proposed LCD can encourage its succeeding layers to minimize the intra-class feature variance caused by occlusions, thus leading to improved robustness against occlusion. In addition, we design an auxiliary spatial attention module by learning a channel-wise attention vector to reweight the feature channels, which improves the contributions of non-occluded regions. Extensive experiments on various benchmarks show that the proposed method outperforms state-of-the-art methods with a remarkable improvement. Jie Zhang 0071, Shiguang Shan, Xiao Liu 0040, Zhongqin Wu, Xilin Chen 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Learning Shape-Appearance Based Attributes Representation for Facial Attribute Recognition with Limited Labeled DataabstractThe Facial Attribute Recognition (FAR) is a challenging task especially when there exists limited labeled data, which may lead the mainstream fully-supervised FAR methods to be no longer in force. To tackle this problem, we propose a novel unsupervised learning framework named Shape-Appearance Based Attributes Representation Learning (SABAL) by leveraging large-scale unlabeled face data. Considering face attributes are mainly determined by 3D shape and facial appearance, we decouple a face image into 3D shape and appearance features by two branch networks, i.e., 3D Shape Branch and Facial Appearance Branch. 3D Shape Branch and Facial Appearance Branch are jointly trained with orthogonal loss and 2D face reconstruction loss to obtain robust facial representations containing 3D-geometry and texture information, which are beneficial for attributes recognition. Finally, the unsupervised learnt features are transferred to the FAR task by fine-tuning on limited labeled data from CelebA. Extensive experiments show that we achieve comparable results to state-of-the-art methods. Kunyan Li, Jie Zhang 0071, Shiguang Shan |
FG | 2 |
| 2021 | Unknown Aware Feature Learning for Face Forgery DetectionabstractThe face forgery detection problem has attracted wide attention in recent years. Although the vanilla convolutional neural network achieves promising results under the intra-domain testing scenario, it always fails to generalize to unseen scenarios. To address this problem, we propose Generalized Feature Space Learning (GFSL) with unknown forgery awareness, which leverages domain generalization to utilize face images forged with various methods. Considering that the true distribution of fake samples is harder to predict than the real samples, we regularize the model with an asymmetric triplet loss, aggregating only the real samples to learn an accurate real-image distribution, which forms an classification boundary that surrounds the real samples and generalizes well to unknown fake samples. Moreover, we apply Representation Self-Challenging (RSC) to perform selective dropout on features, which forces the model to learn more completed features rather than one or a few of the most prominent features, leading to better generalization ability. Extensive experiments show that our method consistently outperforms baseline models under various cross-manipulation-method tests and achieves comparable performance to the state-of-the-art methods on both intra- and cross-dataset evaluations. Liang Shi 0002, Jie Zhang 0071, Chenyue Liang, Shiguang Shan |
FG | 2 |
| 2021 | MFR 2021: Masked Face Recognition CompetitionabstractThis paper presents a summary of the Masked Face Recognition Competitions (MFR) held within the 2021 International Joint Conference on Biometrics (IJCB 2021). The competition attracted a total of 10 participating teams with valid submissions. The affiliations of these teams are diverse and associated with academia and industry in nine different countries. These teams successfully submitted 18 valid solutions. The competition is designed to motivate solutions aiming at enhancing the face recognition accuracy of masked faces. Moreover, the competition considered the deployability of the proposed solutions by taking the compactness of the face recognition models into account. A private dataset representing a collaborative, multisession, real masked, capture scenario is used to evaluate the submitted solutions. In comparison to one of the topperforming academic face recognition solutions, 10 out of the 18 submitted solutions did score higher masked face verification accuracy. Fadi Boutros, Naser Damer, Jan Niklas Kolf, Kiran B. Raja, Florian Kirchbuchner, Ramachandra Raghavendra, Arjan Kuijper, Pengcheng Fang, Fei Wang 0032, David Montero 0002, Naiara Aginako, Basilio Sierra, Marcos Nieto Doncel, Mustafa Ekrem Erakin, Ugur Demir, Hazim Kemal Ekenel, Asaki Kataoka, Kohei Ichikawa, Shizuma Kubo, Jie Zhang 0071, Shiguang Shan, Klemen Grm, Vitomir Struc, Sachith Seneviratne, Nuran Kasthuriarachchi, Sanka Rasnayaka, Pedro C. Neto, Ana Filipa Sequeira, João Ribeiro Pinto, Mohsen Saffari, Jaime S. Cardoso 0001 |
IJCB | 21 |
| 2021 | Meta Gradient Adversarial AttackabstractIn recent years, research on adversarial attacks has be-come a hot spot. Although current literature on the transfer-based adversarial attack has achieved promising results for improving the transferability to unseen black-box models, it still leaves a long way to go. Inspired by the idea of meta-learning, this paper proposes a novel architecture called Meta Gradient Adversarial Attack (MGAA), which is plug-and-play and can be integrated with any existing gradient-based attack method for improving the cross-model transferability. Specifically, we randomly sample multiple models from a model zoo to compose different tasks and iteratively simulate a white-box attack and a black-box attack in each task. By narrowing the gap between the gradient directions in white-box and black-box attacks, the transfer-ability of adversarial examples on the black-box setting can be improved. Extensive experiments on the CIFAR10 and ImageNet datasets show that our architecture outperforms the state-of-the-art methods for both black-box and white-box attack settings. Zheng Yuan 0005, Jie Zhang 0071, Yunpei Jia, Chuanqi Tan, Shiguang Shan |
ICCV | 2 |
| 2021 | Unified unsupervised and semi-supervised domain adaptation network for cross-scenario face anti-spoofing
Yunpei Jia, Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001 |
Pattern Recognit. | 2 |
| 2020 | Single-Side Domain Generalization for Face Anti-SpoofingabstractExisting domain generalization methods for face anti-spoofing endeavor to extract common differentiation features to improve the generalization. However, due to large distribution discrepancies among fake faces of different domains, it is difficult to seek a compact and generalized feature space for the fake faces. In this work, we propose an end-to-end single-side domain generalization framework (SSDG) to improve the generalization ability of face anti-spoofing. The main idea is to learn a generalized feature space, where the feature distribution of the real faces is compact while that of the fake ones is dispersed among domains but compact within each domain. Specifically, a feature generator is trained to make only the real faces from different domains undistinguishable, but not for the fake ones, thus forming a single-side adversarial learning. Moreover, an asymmetric triplet loss is designed to constrain the fake faces of different domains separated while the real ones aggregated. The above two points are integrated into a unified framework in an end-to-end training manner, resulting in a more generalized class boundary, especially good for samples from novel domains. Feature and weight normalization is incorporated to further improve the generalization ability. Extensive experiments show that our proposed approach is effective and outperforms the state-of-the-art methods on four public databases. The code is released online. Yunpei Jia, Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2020 | Self-Supervised Equivariant Attention Mechanism for Weakly Supervised Semantic SegmentationabstractImage-level weakly supervised semantic segmentation is a challenging problem that has been deeply studied in recent years. Most of advanced solutions exploit class activation map (CAM). However, CAMs can hardly serve as the object mask due to the gap between full and weak supervisions. In this paper, we propose a self-supervised equivariant attention mechanism (SEAM) to discover additional supervision and narrow the gap. Our method is based on the observation that equivariance is an implicit constraint in fully supervised semantic segmentation, whose pixel-level labels take the same spatial transformation as the input images during data augmentation. However, this constraint is lost on the CAMs trained by image-level supervision. Therefore, we propose consistency regularization on predicted CAMs from various transformed images to provide self-supervision for network learning. Moreover, we propose a pixel correlation module (PCM), which exploits context appearance information and refines the prediction of current pixel by its similar neighbors, leading to further improvement on CAMs consistency. Extensive experiments on PASCAL VOC 2012 dataset demonstrate our method outperforms state-of-the-art methods using the same level of supervision. The code is released online. Yude Wang, Jie Zhang 0071, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2020 | PAS-Net: Pose-based and Appearance-based Spatiotemporal Networks Fusion for Action RecognitionabstractHuman poses play important roles in action analysis. However, most state-of-the-art approaches in action recognition ignore the importance of human poses and rarely leverage the pose information for further improving the recognition performance. In this paper, we propose a novel network architecture, which simultaneously considers the appearance information and pose knowledge for robust action recognition. We explore various architectures for fusing the appearance and pose information rather than simply averaging scores at the final layer. Moreover, a novel training strategy is proposed to reduce the influence of overfitting for limited training data. Extensive experiments show that our method achieves competitive performance on the popular benchmarks, i.e., UCF-101 and HMDB-51. Changzhen Li, Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001 |
FG | 2 |
| 2020 | Noise Robust Hard Example Mining for Human Detection with Efficient Depth-Thermal FusionabstractIdentity-preserving human detection is important for the privacy-protecting applications. IPHD [1] is a newly collected identity-preserving dataset that only contains depth and thermal images, which have much less information than RGB images. While less information and weakly labeled ground-truth boxes make it difficult to locate the objects correctly. In this paper, we adopt an efficient depth-thermal fusion approach to combine these two different inputs and enhance the representation. Moreover, a noise robust hard example mining algorithm is proposed to deal with weakly labeled data. The experiments show that our single model with single scale testing can get the AP=88.1 at IoU=0.5, which is a significant improvement compared with other competition results. Jie Zhang 0071, Shiguang Shan |
FG | 2 |
| 2020 | Leveraging Auxiliary Tasks for Height and Weight Estimation by Multi Task LearningabstractHeight and weight, two of the most important biological characteristics of human body, play crucial roles in physical condition estimation. Height and weight estimation with single face image via deep convolutional neural network suffers from poor performance due to lack of labeled data. To address this issue, inspired by the relevance of gender, age, height and weight, we propose an auxiliary-task learning framework, employing multiple relevant tasks to improve the performance of primary tasks. Specifically, gender prediction and age estimation are utilized as auxiliary tasks to assist primary tasks (i.e., height and weight estimation) learning via deep residual auxiliary block. Experiments are conducted on the public VIP-attributes datasets and our private VIPL-MumoFace- WH datasets. Our method outperforms the baseline methods of hard parameter sharing in multi-task learning, demonstrating the effectiveness of auxiliary-task learning framework for height and weight estimation. Jie Zhang 0071, Shiguang Shan |
IJCB | 2 |
| 2020 | Attributes Aware Face Generation with Generative Adversarial NetworksabstractRecent studies have shown remarkable success in face image generations. However, most of the existing methods only generate face images from random noise, and cannot generate face images according to the specific attributes. In this paper, we focus on the problem of face synthesis from attributes, which aims at generating faces with specific characteristics corresponding to the given attributes. To this end, we propose a novel attributes aware face image generator method with generative adversarial networks called AFGAN. Specifically, we firstly propose a two-path embedding layer and self-attention mechanism to convert binary attribute vector to rich attribute features. Then three stacked generators generate 64 × 64, 128 × 128 and 256 × 256 resolution face images respectively by taking the attribute features as input. In addition, an image-attribute matching loss is proposed to enhance the correlation between the generated images and input attributes. Extensive experiments on CelebA demonstrate the superiority of our AFGAN in terms of both qualitative and quantitative evaluations. Zheng Yuan 0005, Jie Zhang 0071, Shiguang Shan, Xilin Chen 0001 |
ICPR | 2 |
| 2020 | Deformable face net for pose invariant face recognition
Jie Zhang 0071, Shiguang Shan, Meina Kan, Xilin Chen 0001 |
Pattern Recognit. | 2 |
| 2019 | Deformable Face Net: Learning Pose Invariant Feature with Pose Aware Feature Alignment for Face RecognitionabstractFace recognition plays an important role in computer vision. It still remains a challenging task due to pose, expression, illumination, partial occlusion, etc. In this work, we propose a novel Deformable Face Net (DFN) to handle the pose variations in face recognition. The Deformable Face Net introduces deformable convolution modules to simultaneously learn face recognition oriented alignment and feature extraction. Specifically, two loss functions, namely displacement consistency loss (DCL) and identity consistency loss (ICL) are designed to minimize the intra-class feature variation caused by different poses. These two loss functions jointly learn pose-aware displacement fields for deformable convolutions in the DFN. Different from the existing methods, the DFN focuses on aligning features across different poses rather than frontalizing the input faces. Extensive experiments show that the proposed DFN outperforms the state-of-the-art methods, especially on the datasets with large poses. Jie Zhang 0071, Shiguang Shan, Meina Kan, Xilin Chen 0001 |
FG | 2 |
| 2019 | DFT-Net: Disentanglement of Face Deformation and Texture Synthesis for Expression EditingabstractThis paper presents a novel deep architecture DFT-Net that combines the advantages of Generative Adversarial Networks (GANs) and warp mechanisms for expression editing. Recent generative models leverage Action Units as annotations and show more flexible expression manipulation than previous approaches using other guiding information. However, those methods bring inevitable artifacts where facial components deform (e.g. eyes from open to close), for the structural defect in modeling shape variations without geometric guidance such as facial landmarks. Our approach explicitly disentangles face deformations and appearance details by constructing two parallel networks, one that learns an appearance flow for 2D warps and the other generates corresponding texture and hallucinates hidden regions such as mouth interiors. Experimental results show our method outperforms the state-of-the-art on various expression editing tasks. Jie Zhang 0071, Zijia Lu, Shiguang Shan |
ICIP | 2 |
| 2019 | Locality-constrained framework for face alignment
Jie Zhang 0071, Meina Kan, Shiguang Shan, Xiujuan Chai, Xilin Chen 0001 |
Frontiers Comput. Sci. | 1 |
| 2018 | A Three-Category Face Detector with Contextual Information on Finding Tiny FacesabstractGreat progresses have been achieved on object detection in the wild. However, it still remains a challenging problem due to tiny objects. In this paper, we present a Three-category Classification Neural Network to find tiny faces under complex environments by leveraging contextual information around faces. Tiny faces (within 20×20 pixels) are so fuzzy that the facial patterns are not clear or even ambiguous for detection. To solve this problem, instead of formulating the face detection as a two-category classification task, a novel face detection network is proposed for three-category classification, i.e., normal face, tiny face and background. Moreover, we take full advantage of contextual information around faces and pick good prior anchors to predict good detection on tiny faces. Extensive experiments on two challenging face detection benchmarks, FDDB and WIDER FACE, demonstrate the effectiveness of our method. Jie Zhang 0071, Yuanqing Xia, Shiguang Shan |
ICIP | 2 |
| 2018 | Efficient Weighted Kernel Sharing Convolutional Neural NetworksabstractTo lessen the redundancy of convolutional kernels, this paper proposes a new convolutional structure, i.e., weighted kernel sharing convolution (WKSC), which gathers the inputs with the same kernel, so the inputs in each group can share the same convolutional kernel. Also, an extra weighting is imposed for each input channel before the sharing process to manifest its diversity. As a consequence, the number of kernels can be greatly reduced, leading to a reduction of model parameters and the speedup of inference. Moreover, WKSC can be combined with other existing compression models such as depthwise separable convolutions, resulting in a more compressed architecture. Extensive experiments on CIFAR-100 and ImageNet classification demonstrate the effectiveness of the new approach in both computation cost and the parameters required compared with the state-of-the-art works. Helong Zhou, Yie-Tarng Chen, Jie Zhang 0071, Wen-Hsien Fang |
VCIP | 3 |
| 2017 | A Fully End-to-End Cascaded CNN for Facial Landmark DetectionabstractFacial landmark detection plays an important role in computer vision. It is a challenging problem due to various poses, exaggerated expressions and partial occlusions. In this work, we propose a Fully End-to-End Cascaded Convolutional Neural Network (FEC-CNN) for more promising facial landmark detection. Specifically, FEC-CNN includes several sub- CNNs, which progressively refine the shape prediction via finer and finer modeling, and the overall network is optimized fully end-to-end. Experiments on three challenging datasets, IBUG, 300W competition and AFLW, demonstrate that the proposed method is robust to large poses, exaggerated expressions and partial occlusions. The proposed FEC-CNN significantly improves the accuracy of landmark prediction. Zhenliang He, Meina Kan, Jie Zhang 0071, Xilin Chen 0001, Shiguang Shan |
FG | 3 |
| 2016 | Occlusion-Free Face Alignment: Deep Regression Networks Coupled with De-Corrupt AutoEncodersabstractFace alignment or facial landmark detection plays an important role in many computer vision applications, e.g., face recognition, facial expression recognition, face animation, etc. However, the performance of face alignment system degenerates severely when occlusions occur. In this work, we propose a novel face alignment method, which cascades several Deep Regression networks coupled with De-corrupt Autoencoders (denoted as DRDA) to explicitly handle partial occlusion problem. Different from the previous works that can only detect occlusions and discard the occluded parts, our proposed de-corrupt autoencoder network can automatically recover the genuine appearance for the occluded parts and the recovered parts can be leveraged together with those non-occluded parts for more accurate alignment. By coupling de-corrupt autoencoders with deep regression networks, a deep alignment model robust to partial occlusions is achieved. Besides, our method can localize occluded regions rather than merely predict whether the landmarks are occluded. Experiments on two challenging occluded face datasets demonstrate that our method significantly outperforms the state-of-the-art methods. Jie Zhang 0071, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
CVPR | 1 |
| 2015 | Leveraging Datasets with Varying Annotations for Face Alignment via Deep Regression NetworkabstractFacial landmark detection, as a vital topic in computer vision, has been studied for many decades and lots of datasets have been collected for evaluation. These datasets usually have different annotations, e.g., 68-landmark markup for LFPW dataset, while 74-landmark markup for GTAV dataset. Intuitively, it is meaningful to fuse all the datasets to predict a union of all types of landmarks from multiple datasets (i.e., transfer the annotations of each dataset to all other datasets), but this problem is nontrivial due to the distribution discrepancy between datasets and incomplete annotations of all types for each dataset. In this work, we propose a deep regression network coupled with sparse shape regression (DRN-SSR) to predict the union of all types of landmarks by leveraging datasets with varying annotations, each dataset with one type of annotation. Specifically, the deep regression network intends to predict the union of all landmarks, and the sparse shape regression attempts to approximate those undefined landmarks on each dataset so as to guide the learning of the deep regression network for face alignment. Extensive experiments on two challenging datasets, IBUG and GLF, demonstrate that our method can effectively leverage the multiple datasets with different annotations to predict the union of all types of landmarks. Jie Zhang 0071, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
ICCV | 1 |
| 2014 | Topic-Aware Deep Auto-Encoders (TDA) for Face Alignment
Jie Zhang 0071, Meina Kan, Shiguang Shan, Xilin Chen 0001 |
ACCV (3) | 1 |
| 2014 | Coarse-to-Fine Auto-Encoder Networks (CFAN) for Real-Time Face Alignment
Jie Zhang 0071, Shiguang Shan, Meina Kan, Xilin Chen 0001 |
ECCV (2) | 1 |