EDBT 2026 Demo / reviewers in the wild / expert
Xiangyu Zhu 0001
dblp:19/10065-1
· DBLP profile ↗
83ranked-venue papers
9as first author
57since 2021 · last 2026
0000-0002-4636-9677ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 69 · 5 first-author · 46 since 2021Artificial intelligence and machine learning · 54 · 9 first-author · 33 since 2021Security and privacy · 14 · 13 since 2021Human-computer interaction and ubiquitous computing · 11 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unifying Locality of KANs and Feature Drift Compensation Projection for Data-Free Replay Based Continual Face Forgery DetectionabstractThe rapid advancements in face forgery techniques necessitate that detectors continuously adapt to new forgery methods, thus situating face forgery detection within a continual learning paradigm. However, when detectors learn new forgery types, their performance on previous types often degrades rapidly, a phenomenon known as catastrophic forgetting. Kolmogorov-Arnold Networks (KANs) utilize locally plastic splines as their activation functions, enabling them to learn new tasks by modifying only local regions of the functions while leaving other areas unaffected. Therefore, they are naturally suitable for addressing catastrophic forgetting. However, KANs have two significant limitations: 1) the splines are ineffective for modeling high-dimensional images, while alternative activation functions that are suitable for images lack the essential property of locality; 2) in continual learning, when features from different domains overlap, the mapping of different domains to distinct curve regions always collapses due to repeated modifications of the same regions. In this paper, we propose a KAN-based Continual Face Forgery Detection (KAN-CFD) framework, which includes a Domain-Group KAN Detector (DG-KD) and a data-free replay Feature Separation strategy via KAN Drift Compensation Projection (FS-KDCP). DG-KD enables KANs to fit high-dimensional image inputs while preserving locality and local plasticity. FS-KDCP avoids the overlap of the KAN input spaces without using data from prior tasks. Experimental results demonstrate that the proposed method achieves superior performance while notably reducing forgetting. Tianshuo Zhang, Siran Peng, Xiangyu Zhu 0001, Zhen Lei 0001 |
AAAI | 5 |
| 2026 | AdaField: Generalizable Surface Pressure Modeling with Physics-Informed Pre-training and Flow-Conditioned AdaptationabstractThe surface pressure field of transportation systems, including cars, trains, and aircraft, is critical for aerodynamic analysis and design. In recent years, deep neural networks have emerged as promising and efficient methods for modeling surface pressure field, being alternatives to computationally expensive CFD simulations. Currently, large-scale public datasets are available for domains such as automotive aerodynamics. However, in many specialized areas, such as high-speed trains, data scarcity remains a fundamental challenge in aerodynamic modeling, severely limiting the effectiveness of standard neural network approaches. To address this limitation, we propose the Adaptive Field Learning Framework (AdaField), which pre-trains the model on public large-scale datasets to improve generalization in sub-domains with limited data. AdaField comprises two key components. First, we design the Semantic Aggregation Point Transformer (SAPT) as a high-performance backbone that efficiently handles large-scale point clouds for surface pressure prediction. Second, regarding the substantial differences in flow conditions and geometric scales across different aerodynamic subdomains, we propose Flow-Conditioned Adapter (FCA) and Physics-Informed Data Augmentation (PIDA). FCA enables the model to flexibly adapt to different flow conditions with a small set of trainable parameters, while PIDA expands the training data distribution to better cover variations in object scale and velocity. Our experiments show that AdaField achieves SOTA performance on the DrivAerNet++ dataset and can be effectively transferred to train and aircraft scenarios with minimal fine-tuning. These results highlight AdaField’s potential as a generalizable and transferable solution for surface pressure field modeling, supporting efficient aerodynamic design across a wide range of transportation systems. Junhong Zou, Zhenxu Sun, Zhaoxiang Zhang 0001, Xiangyu Zhu 0001 |
AAAI | 6 |
| 2026 | CPG-PAD: Concept-Informed Prompts Guided Presentation Attack Detection
Xiangyu Zhu 0001, Ajian Liu 0001, Siran Peng, Zhen Lei 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2026 | Adaptive 3D Convolution for Remote Sensing Image FusionabstractRemote sensing image fusion aims to create a high-resolution multi/hyper-spectral image from a high-resolution image with limited spectral information and a low-resolution image with abundant spectral data. Recently, deep learning (DL) techniques have shown significant effectiveness in this area. Most DL-based methods approach image fusion as a 2D problem by encoding spectral information into feature map channels. However, our research suggests that this strategy introduces notable spectral distortions. In contrast, some methods consider spectral data as an additional dimension, utilizing standard 3D convolutions to preserve spectral information. Nevertheless, in a standard 3D convolutional layer, the same set of kernels is applied across all input regions, which we have found to be sub-optimal for image fusion. Furthermore, standard 3D convolutions necessitate substantial computational resources. To address these challenges, we propose a novel convolutional paradigm called Adaptive 3D Convolution (Ada3D) for remote sensing image fusion. Ada3D applies a unique set of 3D kernels to each input voxel, enabling the capture of fine-grained details. These adaptive kernels are generated through a two-step process: 1) spatial and spectral kernels are derived from their respective image sources and 2) these two types of kernels are then combined to form content-aware 3D kernels that effectively integrate spatial and spectral information. Additionally, adaptive biases are introduced to enhance the convolutional outcome at the voxel level. Furthermore, we incorporate the group convolution technique to reduce computational complexity. As a result, Ada3D offers full adaptivity in an efficient manner. Evaluation results across five datasets demonstrate that our method achieves state-of-the-art (SOTA) performance, underscoring the superiority of Ada3D. The code is available at https://github.com/PSRben/Ada3D. Siran Peng, Xiangyu Zhu 0001, Shangqi Deng, Liang-Jian Deng, Zhen Lei 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Progressive Rendering Distillation: Adapting Stable Diffusion for Instant Text-to-Mesh Generation without 3D DataabstractIt is highly desirable to obtain a model that can generate high-quality 3D meshes from text prompts in just seconds. While recent attempts have adapted pre-trained text-to-image diffusion models, such as Stable Diffusion (SD), into generators of 3D representations (e.g., Triplane), they often suffer from poor quality due to the lack of sufficient high-quality 3D training data. Aiming at overcoming the data shortage, we propose a novel training scheme, termed as Progressive Rendering Distillation (PRD), eliminating the need for 3D ground-truths by distilling multi-view diffusion models and adapting SD into a native 3D generator. In each iteration of training, PRD uses the U-Net to progressively denoise the latent from random noise for a few steps, and in each step it decodes the denoised latent into 3D output. Multi-view diffusion models, including MVDream and RichDreamer, are used in joint with SD to distill text-consistent textures and geometries into the 3D outputs through score distillation. Since PRD supports training without 3D ground-truths, we can easily scale up the training data and improve generation quality for challenging text prompts with creative concepts. Meanwhile, PRD can accelerate the inference speed of the generation model in just a few steps. With PRD, we train a Triplane generator, namely TriplaneTurbo, which adds only 2.5% trainable parameters to adapt SD for Triplane generation. TriplaneTurbo outperforms previous text-to-3D generators in both efficiency and quality. Specifically, it can produce high-quality 3D meshes in 1.2 seconds and generalize well for challenging text input. The code is available at github.com/theEricMa/TriplaneTurbo. Zhiyuan Ma 0002, Rongyuan Wu, Xiangyu Zhu 0001, Zhen Lei 0001, Lei Zhang 0006 |
CVPR | 4 |
| 2025 | MVBoost: Boost 3D Reconstruction with Multi-View RefinementabstractRecent advancements in 3D object reconstruction have been remarkable, yet most current 3D models rely heavily on existing 3D datasets. The scarcity of diverse 3D datasets results in limited generalization capabilities of 3D reconstruction models. In this paper, we propose a novel framework for boosting 3D reconstruction with multi-view refinement (MVBoost) by generating pseudo-GT data. The key of MVBoost is combining the advantages of the high accuracy of the multi-view generation model and the consistency of the 3D reconstruction model to create a reliable data source. Specifically, given a single-view input image, we employ a multi-view diffusion model to generate multiple views, followed by a large 3D reconstruction model to produce consistent 3D data. MVBoost then adaptively refines these multi-view images, rendered from the consistent 3D data, to build a large-scale multi-view dataset for training a feed-forward 3D reconstruction model. Additionally, the input view optimization is designed to optimize the corresponding viewpoints based on the user’s input image, ensuring that the most important viewpoint is accurately tailored to the user’s needs. Extensive evaluations demonstrate that our method achieves superior reconstruction results and robust generalization compared to prior works. Zhiyuan Ma 0002, Xiangyu Zhu 0001, Zhen Lei 0001 |
CVPR | 4 |
| 2025 | Diffusion Models are Zero-Shot Generative Text-Vision RetrieversabstractLarge-scale text-to-image diffusion models have demonstrated impressive capabilities for downstream tasks by leveraging strong vision-language alignment from generative pre-training. Recently, a number of works have explored how to use the power of text-to-image diffusion models for text-image matching tasks. While previous generative text-image matching methods have shown potential in retrieving the most challenging candidates, they still suffer from extremely slow retrieval speeds and lack the ability to handle temporal dimensions, making them impractical for text-video retrieval. In this paper, we propose ZSGenRet, a simple yet effective zero-shot generative text-vision retrieval framework for both text-image and text-video tasks, based on pre-trained text-to-image diffusion models. We further incorporate inversion saliency detection to identify key frames in videos and enhance the semantic representation of the vision encoder. Experimental results demonstrate that ZSGenRet significantly improves text-video retrieval performance and achieves competitive results on text-image retrieval while remarkably improving efficiency. To the best of our knowledge, the proposed ZSGenRet is the first to explore zero-shot text-video retrieval based on diffusion models. Zeke Xie, Xiangyu Zhu 0001, Zhen Lei 0001 |
ICASSP | 4 |
| 2025 | DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion TransformerabstractSpeech-driven 3D facial animation is important for many multimedia applications. Recent work has shown promise in using either Diffusion models or Transformer architectures for this task. However, their mere aggregation does not lead to improved performance. We suspect this is due to a shortage of paired audio-4D data, which is crucial for the Transformer to effectively perform as a denoiser within the Diffusion framework. To tackle this issue, we present DiffSpeaker, a Transformer-based network equipped with novel biased conditional attention modules. These modules serve as substitutes for the traditional self/cross-attention in standard Transformers, incorporating thoughtfully designed biases that steer the attention mechanisms to concentrate on both the relevant task-specific and diffusion-related conditions. We also explore the trade-off between accurate lip synchronization and non-verbal facial expressions within the Diffusion paradigm. Experiments show our model achieves state-of-the-art performance on existing benchmarks, and fast inference speed owing to its ability to generate facial motions in parallel. Our code is avalable at https://github.com/theEricMa/DiffSpeaker. Zhiyuan Ma 0002, Xiangyu Zhu 0001, Chen Qian 0006, Shukai Chen, Guo-Jun Qi, Zhaoxiang Zhang 0001, Zhen Lei 0001 |
IJCB | 2 |
| 2025 | Explaining Convolutional Neural Networks via a Concise and Hierarchical ApproachabstractExplainable artificial intelligence (XAI) aims to bring transparency to black-box neural networks. Many innovative explainable methods provide rich and multifaceted explanations. However, these methods often have complex mechanisms or a high complexity of explanations, making them difficult for people to understand. To address this challenge, we introduce a concept-based explainable method which can improve the readability of deep neural network explanations by obtaining a concise set of high-quality explanations. We reduce the explanation redundancy by weighted hierarchical clustering, thereby obtaining a set of explanations that completely describe the input image and are crucial to the network’s decision-making. Compared with existing XAI methods, our approach reduces the complexity while ensuring the integrity and comprehensibility of the explanation. In addition, we show that our approach can be traced back to the neuron-level explanation, which can also provide inspiration for model researchers to interpret the model. We validated the effectiveness of our approach through experiments in bird classification and facial recognition tasks. Specifically, we employed XAI to investigate the mechanisms behind face recognition, identifying critical neurons that correspond to key concepts in the recognition process. Xiangyu Zhu 0001, Stan Z. Li, Zhen Lei 0001 |
IJCB | 2 |
| 2025 | StreamWMR: A Streaming Framework for Real-time 3D Whole-body Mesh Recoveryabstract3D whole-body mesh recovery aims to extract parameters for the human body, hands, and head from a single human image. Most applications related to human mesh recovery, such as physical fitness motion capture and operating room motion capture, necessitate real-time video stream processing. However, existing methods ignore the video processing and often require significant computational resources, making real-time performance unattainable and greatly limiting their practicality. Moreover, noticeable misalignments are often observed when concatenating them back to the body and reprojecting them onto the image. In this paper, we propose a streaming framework for whole-body mesh recovery in the video. First, we simplify pose regression by leveraging the root nodes of the hands and head to locate each component. Second, for temporal optimization, we incorporate attention mechanisms related to keypoint velocity to incorporate information from previous frames and achieve more stable and smooth motions. Finally, we propose a multi-view projection loss to eliminate the ambiguity caused by inaccurate 3D regression and pose estimation in computing reprojection errors. The combination of these enables our method to achieve real-time inference speed while maintaining accuracy and stability. Xiangyu Zhu 0001, Jinlin Wu, Zidu Wang, Shukai Chen, Dong Yi, Zhen Lei 0001 |
IJCB | 2 |
| 2025 | GenFIQA: Generative Face Image Quality Assessment via Identity-conditioned Diffusion ModelabstractFace recognition (FR) systems are widely deployed but often struggle due to unconstrained image-capturing conditions. Face image quality assessment (FIQA), applied before recognition, mitigates these challenges by filtering out unreliable samples. Current leading FIQA methods evaluate image quality based on the characteristics observed within the FR model pipeline. However, they leave out the inherent differences in identity embeddings between high-and low-quality face images. To this end, we propose Gen-FIQA, which utilizes a generative model to probe and amplify this difference. Specifically, we extract the identity embedding from an input image using a pre-trained FR model, and then use it as a conditioning signal to generate several face images of the same identity. This generation process leverages the inherent prior in the generative model to translate the difference in identity embedding space back to pixel space. To quantify these differences, the quality score is computed as the average cosine similarity between embeddings from the original and generated images. To improve computational efficiency, we further distill GenFIQA into a lightweight regression-based variant, GenFIQA(R). Extensive experiments across five benchmark datasets and four FR models demonstrate the superiority of our methods over thirteen state-of-the-art FIQA methods. Zheyu Yan, Weisong Zhao, Kai Pang, Xiangyu Zhu 0001, Xiaoyu Zhang 0002, Zhen Lei 0001 |
IJCB | 5 |
| 2025 | Exploiting Facial Discomfort Clues with Vision-Language Model for Generalizable Face Forgery DetectionabstractFace forgery detection is a challenging problem due to the diversity and rapid iteration of face manipulation methods, especially in detecting unknown forgery types. To address this challenge, we explore the common features shared among various forgery types. We find that even though multiple manipulation methods leave different invisible forgery traces, the fake faces often exhibit a similar overall pattern of discomfort. Such discomfort can serve as a universal clue across multiple forgery types, thereby possessing the potential to achieve strong generalization in face forgery detection. To this end, we utilize Vision-Language Models (VLMs) to simulate the cognitive process from perceiving the image to generating the sense of discomfort and propose a multi-task Forgery-Discomfort Joint Learning (FDJL) framework to leverage VLMs to perceive and identify fake faces by integrating discomfort cues. Specifically, we collect a Facial Discomfort dataset guided by the uncanny valley theory, enabling the model to extract and learn discomfort features. Extensive experiments demonstrate that our method achieves state-of-the-art performance and exhibits the best generalization for unknown forgery types. Tianshuo Zhang, Xiangyu Zhu 0001, Kai Pang, Shukai Chen, Zhen Lei 0001 |
IJCB | 2 |
| 2025 | Spoof Trace Discovery for Deep Learning Based Explainable Face Anti-SpoofingabstractWith the rapid growth usage of face recognition in people’s daily life, face anti-spoofing becomes increasingly important to avoid malicious attacks. Recent face anti-spoofing models can reach a high classification accuracy on multiple datasets but these models can only tell people "this face is fake" while lacking the explanation to answer "why it is fake". Such a system undermines trustworthiness and causes user confusion, as it denies their requests without providing any explanations. In this paper, we incorporate XAI into face anti-spoofing and propose a new problem termed X-FAS (eXplainable Face Anti-Spoofing) empowering face anti-spoofing models to provide an explanation. We propose SPTD (SPoof Trace Discovery), an X-FAS method which can discover spoof concepts and provide reliable explanations on the basis of discovered concepts. To evaluate the quality of X-FAS methods, we propose an X-FAS benchmark with annotated spoof traces by experts. We analyze SPTD explanations on face anti-spoofing dataset and compare SPTD quantitatively and qualitatively with previous XAI methods on proposed X-FAS benchmark. Experimental results demonstrate SPTD’s ability to generate reliable explanations. Xiangyu Zhu 0001, Kai Pang, Guoying Zhao 0001, Zhen Lei 0001 |
IJCB | 2 |
| 2025 | ET-Talk: Effective Training Strategy to Enhance Synchrony and Fidelity for Talking Face GenerationabstractRecently, significant advancements have been made in audio-driven talking face generation. While GAN-based methods are widely used in this task, they struggle to achieve simultaneous lip accuracy and high-fidelity. Generated lip shapes tend to be overly influenced by the lip of reference images that provide identity information, leading to unstable and unsynchronized results. Moreover, the synthesized face frequently suffers from blurred teeth, skin textures, and compromised facial identity. To address these challenges, we propose an effective and innovative training strategy that simultaneously ensures lip synchrony and facial fidelity. First, we adaptively select the reference image using a hard-mining based strategy to prevent the network from simply copying the reference lip, enhancing the stability and synchronicity of lip movements. Second, we incorporate high-resolution facial images in training a quality discriminator within the GAN loss, improving the generated faces’ fidelity. Third, a global-to-detail training strategy is employed, starting with strengthening synchrony and then image quality to preserve identity and visual details. Experiments on the HDTF dataset demonstrate that our method achieves state-of-the-art performance in both lip accuracy and image quality. Baiqin Wang, Xiangyu Zhu 0001, Shukai Chen, Zhen Lei 0001 |
ICME | 2 |
| 2025 | Top-Down Guidance for Learning Object-Centric RepresentationsabstractHumans' innate ability to decompose scenes into objects allows for efficient understanding, predicting, and planning. In light of this, Object-Centric Learning (OCL) attempts to endow networks with similar capabilities, learning to represent scenes with the composition of objects. However, existing OCL models only learn through reconstructing the input images, which does not assist the model in distinguishing objects, resulting in suboptimal object-centric representations. This flaw limits current object-centric models to relatively simple downstream tasks. To address this issue, we draw on humans’ top-down vision pathway and propose Top-Down Guided Network (TDGNet), which includes a top-down pathway to improve object-centric representations. During training, the top-down pathway constructs guidance with high-level object-centric representations to optimize low-level grid features output by the backbone. While during inference, it refines object-centric representations by detecting and solving conflicts between low- and high-level features. We show that TDGNet outperforms current object-centric models on multiple datasets of varying complexity. In addition, we expand the downstream task scope of object-centric representations by applying TDGNet to the field of robotics, validating its effectiveness in downstream tasks including video prediction and visual planning. Code will be available at https://github.com/zoujunhong/RHGNet. Junhong Zou, Xiangyu Zhu 0001, Zhaoxiang Zhang 0001, Zhen Lei 0001 |
IJCAI | 2 |
| 2025 | Reconstructing 3D Hand-Instrument Interaction from a Single 2D Image in Medical Scenes
Xiangyu Zhu 0001, Jinlin Wu, Ming Feng, Zelin Zang, Hongbin Liu 0001, Zhen Lei 0001 |
MICCAI (10) | 2 |
| 2025 | WMamba: Wavelet-based Mamba for Face Forgery DetectionabstractThe rapid evolution of deepfake generation technologies necessitates the development of robust face forgery detection algorithms. Recent studies have demonstrated that wavelet analysis can enhance the generalization abilities of forgery detectors. Wavelets effectively capture key facial contours, often slender, fine-grained, and globally distributed, that may conceal subtle forgery artifacts imperceptible in the spatial domain. However, current wavelet-based approaches fail to fully exploit the distinctive properties of wavelet data, resulting in sub-optimal feature extraction and limited performance gains. To address this challenge, we introduce WMamba, a novel wavelet-based feature extractor built upon the Mamba architecture. WMamba maximizes the utility of wavelet information through two key innovations. First, we propose Dynamic Contour Convolution (DCConv), which employs specially crafted deformable kernels to adaptively model slender facial contours. Second, by leveraging the Mamba architecture, our method captures long-range spatial relationships with linear complexity. This efficiency allows for the extraction of fine-grained, globally distributed forgery artifacts from small image patches. Extensive experiments show that WMamba achieves state-of-the-art (SOTA) performance, highlighting its effectiveness in face forgery detection. Siran Peng, Tianshuo Zhang, Xiangyu Zhu 0001, Kai Pang, Zhen Lei 0001 |
ACM Multimedia | 4 |
| 2025 | DevFD : Developmental Face Forgery Detection by Learning Shared and Orthogonal LoRA SubspacesabstractThe rise of realistic digital face generation and manipulation poses significant social risks. The primary challenge lies in the rapid and diverse evolution of generation techniques, which often outstrip the detection capabilities of existing models. To defend against the ever-evolving new types of forgery, we need to enable our model to quickly adapt to new domains with limited computation and data while avoiding forgetting previously learned forgery types. In this work, we posit that genuine facial samples are abundant and relatively stable in acquisition methods, while forgery faces continuously evolve with the iteration of manipulation techniques. Given the practical infeasibility of exhaustively collecting all forgery variants, we frame face forgery detection as a continual learning problem and allow the model to develop as new forgery types emerge. Specifically, we employ a Developmental Mixture of Experts (MoE) architecture that uses LoRA models as its individual experts. These experts are organized into two groups: a Real-LoRA to learn and refine knowledge of real faces, and multiple Fake-LoRAs to capture incremental information from different forgery types. To prevent catastrophic forgetting, we ensure that the learning direction of Fake-LoRAs is orthogonal to the established subspace. Moreover, we integrate orthogonal gradients into the orthogonal loss of Fake-LoRAs, preventing gradient interference throughout the training process of each task. Experimental results under both the datasets and manipulation types incremental protocols demonstrate the effectiveness of our method. Tianshuo Zhang, Siran Peng, Xiangyu Zhu 0001, Zhen Lei 0001 |
NeurIPS | 4 |
| 2025 | ROLA: real-world object-centric learning with attention optimization
Qu Tang, Xiangyu Zhu 0001, Zhen Lei 0001, Zhaoxiang Zhang 0001 |
Sci. China Inf. Sci. | 3 |
| 2025 | Global Cross-Entropy Loss for Deep Face RecognitionabstractContemporary deep face recognition techniques predominantly utilize the Softmax loss function, designed based on the similarities between sample features and class prototypes. These similarities can be categorized into four types: in-sample target similarity, in-sample non-target similarity, out-sample target similarity, and out-sample non-target similarity. When a sample feature from a specific class is designated as the anchor, the similarity between this sample and any class prototype is referred to as in-sample similarity. In contrast, the similarity between samples from other classes and any class prototype is known as out-sample similarity. The terms target and non-target indicate whether the sample and the class prototype used for similarity calculation belong to the same identity or not. The conventional Softmax loss function promotes higher in-sample target similarity than in-sample non-target similarity. However, it overlooks the relation between in-sample and out-sample similarity. In this paper, we propose Global Cross-Entropy loss (GCE), which promotes 1) greater in-sample target similarity over both the in-sample and out-sample non-target similarity, and 2) smaller in-sample non-target similarity to both in-sample and out-sample target similarity. In addition, we propose to establish a bilateral margin penalty for both in-sample target and non-target similarity, so that the discrimination and generalization of the deep face model are improved. To bridge the gap between training and testing of face recognition, we adapt the GCE loss into a pairwise framework by randomly replacing some class prototypes with sample features. We designate the model trained with the proposed Global Cross-Entropy loss as GFace. Extensive experiments on several public face benchmarks, including LFW, CALFW, CPLFW, CFP-FP, AgeDB, IJB-C, IJB-B, MFR-Ongoing, and MegaFace, demonstrate the superiority of GFace over other methods. Additionally, GFace exhibits robust performance in general visual recognition task. Weisong Zhao, Xiangyu Zhu 0001, Haichao Shi, Xiaoyu Zhang 0002, Guoying Zhao 0001, Zhen Lei 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Weakly Aligned Feature Fusion for Multimodal Object DetectionabstractTo achieve accurate and robust object detection in the real-world scenario, various forms of images are incorporated, such as color, thermal, and depth. However, multimodal data often suffer from the position shift problem, i.e., the image pair is not strictly aligned, making one object has different positions in different modalities. For the deep learning method, this problem makes it difficult to fuse multimodal features and puzzles the convolutional neural network (CNN) training. In this article, we propose a general multimodal detector named aligned region CNN (AR-CNN) to tackle the position shift problem. First, a region feature (RF) alignment module with adjacent similarity constraint is designed to consistently predict the position shift between two modalities and adaptively align the cross-modal RFs. Second, we propose a novel region of interest (RoI) jitter strategy to improve the robustness to unexpected shift patterns. Third, we present a new multimodal feature fusion method that selects the more reliable feature and suppresses the less useful one via feature reweighting. In addition, by locating bounding boxes in both modalities and building their relationships, we provide novel multimodal labeling named KAIST-Paired. Extensive experiments on 2-D and 3-D object detection, RGB-T, and RGB-D datasets demonstrate the effectiveness and robustness of our method. Lu Zhang 0054, Zhiyong Liu 0001, Xiangyu Zhu 0001, Zhan Song, Xu Yang 0004, Zhen Lei 0001, Hong Qiao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | SyncTalk: The Devil is in the Synchronization for Talking Head SynthesisabstractAchieving high synchronization in the synthesis of realistic, speech-driven talking head videos presents a significant challenge. Traditional Generative Adversarial Networks (GAN) struggle to maintain consistent facial identity, while Neural Radiance Fields (NeRF) methods, although they can address this issue, often produce mismatched lip movements, inadequate facial expressions, and unstable head poses. A lifelike talking head requires synchronized coordination of subject identity, lip movements, facial expression, and head poses. The absence of these synchronizations is a fundamental flaw, leading to unrealistic and artificial outcomes. To address the critical issue of synchronization, identified as the “devil” in creating realistic talking heads, we introduce SyncTalk. This NeRF-based method effectively maintains subject identity, enhancing synchronization and realism in talking head synthesis. SyncTalk employs a Face-Sync Controller to align lip movements with speech and innovatively uses a 3D facial blendshape model to capture accurate facial expressions. Our Head-Sync Stabilizer optimizes head poses, achieving more natural head movements. The Portrait-Sync Generator restores hair details and blends the generated head with the torso for a seamless visual experience. Extensive experiments and user studies demonstrate that SyncTalk outperforms state-of-the-art methods in synchronization and realism. We recommend watching the supplementary video: https://ziqiaopeng.github.io/synctalk Ziqiao Peng, Xiangyu Zhu 0001, Hao Zhao 0002, Jun He 0008, Hongyan Liu 0002, Zhaoxin Fan |
CVPR | 4 |
| 2024 | 3D Face Reconstruction with the Geometric Guidance of Facial Part Segmentationabstract3D Morphable Models (3DMMs) provide promising 3D face reconstructions in various applications. However, existing methods struggle to reconstruct faces with extreme expressions due to deficiencies in supervisory signals, such as sparse or inaccurate landmarks. Segmentation information contains effective geometric contexts for face reconstruction. Certain attempts intuitively depend on differentiable renderers to compare the rendered silhouettes of reconstruction with segmentation, which is prone to issues like local optima and gradient instability. In this paper, we fully utilize the facial part segmentation geometry by introducing Part Re-projection Distance Loss (PRDL). Specifically, PRDL transforms facial part segmentation into 2D points and re-projects the reconstruction onto the image plane. Subsequently, by introducing grid anchors and computing different statistical distances from these anchors to the point sets, PRDL establishes geometry descriptors to optimize the distribution of the point sets for face reconstruction. PRDL exhibits a clear gradient compared to the renderer-based methods and presents state-of-the-art reconstruction performance in extensive quantitative and qualitative experiments. Our project is available at https://github.com/wang-zidu/3DDFA-V3. Zidu Wang, Xiangyu Zhu 0001, Tianshuo Zhang, Baiqin Wang, Zhen Lei 0001 |
CVPR | 2 |
| 2024 | ScaleDreamer: Scalable Text-to-3D Synthesis with Asynchronous Score Distillation
Zhiyuan Ma 0002, Yuxiang Wei 0001, Yabin Zhang 0001, Xiangyu Zhu 0001, Zhen Lei 0001, Lei Zhang 0006 |
ECCV (7) | 4 |
| 2024 | BLIP-Adapter: Bridging Vision-Language Models with Adapters for Generalizable Face Anti-spoofingabstractFace anti-spoofing is essential for ensuring the security of facial recognition systems against spoofing attacks. Recent methods have transferred Vision-Language models to face anti-spoofing (e.g., FLIP and CLIPC8), demonstrating that learning perception from supervision in natural language can enhance the model’s detection performance. However, such methods exhibit limited depth in the interaction between images and texts, resulting in poor performance on fine-grained understanding tasks such as face anti-spoofing. Besides, the lack of diversity in image-text pairs for face anti-spoofing further hinders such methods from playing their best. To address these issues, we propose a novel fine-tuning strategy for Vision-Language models in face anti-spoofing. This strategy introduces the Bootstrapping Language-Image Pre-training model (BLIP), known for its novel interaction mechanisms and superior image-text comprehension, to construct a more generalized feature representation for face anti-spoofing. Furthermore, we propose an Adapter module for the text branch to reduce the negative impact of insufficient data diversity and catastrophic forgetting. Extensive experiments conducted on various cross-domain testing benchmarks demonstrate the significant superiority of our method over the state-of-the-art, highlighting its effectiveness and robustness. Xiangyu Zhu 0001, Xiaoyu Zhang 0002, Zhen Lei 0001 |
IJCB | 2 |
| 2024 | CPL-CLIP: Compound Prompt Learning for Flexible-Modal Face Anti-SpoofingabstractFace anti-spoofing (FAS) is pivotal in safeguarding the integrity of face recognition systems. Flexible-modal FAS utilizes multi-modal data and trains a unified model adaptable to any single-modal testing scenario. This innovation addresses the shortcomings of conventional multi-modal FAS approaches, which typically demand separate model training and deployment for each modality. However, existing flexible-modal FAS approaches activate specific network branches based on the modality of the tested sample. This not only increases the model’s parameters but also necessitates the provision of the image’s modality for testing, thereby constraining deployment flexibility. To address the issue, we present Compound Prompt Learning CLIP (CPL-CLIP), a novel method for flexible-modal FAS. This approach capitalizes on a learned textual prompt that is nearly independent of modality, thus bolstering class-based classification across arbitrary modalities. Specifically, our CPL-CLIP introduces a Dual-Branch Prompt (DBP), consisting of class and modal prompts that describe and guide classification, where each prompt is composed of learnable vectors and fixed templates. To further render the class prompt as modality-agnostic as possible, a Cosine Similarity Loss (CSL) is proposed to facilitate the maximal separation of the class prompt from the modality prompt. With only the class prompt utilized during testing, CPL-CLIP enables deployment in diverse modal testing scenarios without the necessity of the test image’s modality to be known. Extensive experiments demonstrate CPL-CLIP’s superiority over existing methods on several flexible-modal FAS benchmarks. Xiangyu Zhu 0001, Ajian Liu 0001, Xun Lin, Jun Wan 0001, Zhen Lei 0001 |
IJCB | 2 |
| 2024 | S2TD-Face: Reconstruct a Detailed 3D Face with Controllable Texture from a Single Sketchabstract3D textured face reconstruction from sketches applicable in many scenarios such as animation, 3D avatars, artistic design, missing people search, etc., is a highly promising but underdeveloped research topic.On the one hand, the stylistic diversity of sketches leads to existing sketch-to-3D-face methods only being able to handle pose-limited and realistically shaded sketches.On the other hand, texture plays a vital role in representing facial appearance, yet sketches lack this information, necessitating additional texture control in the reconstruction process.This paper proposes a novel method for reconstructing controllable textured and detailed 3D faces from sketches, named S2TD-Face.S2TD-Face introduces a two-stage geometry reconstruction framework that directly reconstructs detailed geometry from the input sketch.To keep geometry consistent with the delicate strokes of the sketch, we propose a novel sketch-to-geometry loss that ensures the reconstruction accurately fits the input features like dimples and wrinkles.Our training strategies do not rely on hard-to-obtain 3D face scanning data or labor-intensive hand-drawn sketches.Furthermore, S2TD-Face introduces a texture control module utilizing text prompts to select the most suitable textures from a library and seamlessly integrate them into the geometry, resulting in a 3D detailed face with controllable texture.S2TD-Face surpasses existing state-of-the-art methods in extensive quantitative and qualitative experiments.Our project is available at https://github.com/wang-zidu/S2TD-Face. Zidu Wang, Xiangyu Zhu 0001, Tianshuo Zhang, Zhen Lei 0001 |
ACM Multimedia | 2 |
| 2024 | Spatiotemporal Fine-grained Video Description for Short VideosabstractIn the mobile internet era, short videos are inundating people's lives. However, research on visual language models specifically designed for short videos has not yet received sufficient attention. Short videos are not just videos of limited duration. The prominent visual details and high information density of short videos differentiate them to long videos. In this paper, we propose the SpatioTemporal Fine-grained Description (STFVD) emphasizing on the uniqueness of short videos, which entails capturing the intricate details of the main subject and fine-grained movements. To this end, we create a comprehensive Short Video Advertisements Description (SVAD) dataset, comprising 34,930 clips from 5,046 videos. The dataset covers a range of topics, including 191 sub-industries, 649 popular products, and 470 trending games. Various efforts have been made in the data annotation process to ensure the inclusion of fine-grained spatiotemporal information, resulting in 34,930 high-quality annotations. Compared to existing datasets, samples in SVAD exhibit a superior text information density, suggesting that SVAD is more appropriate for the analysis of short videos. Based on the SVAD dataset, we develop a visual language model (SVAD-VLM) to generate spatiotemporal fine-grained description for short videos. We use a prompt-guided keyword generation task to efficiently learn key visual information. Moreover, we also utilize dual visual alignment to exploit the advantage of mixed-datasets training. Experiments on SVAD dataset demonstrate the challenge of STFVD and the competitive performance of proposed method compared to previous ones. Te Yang, Jian Jia, Bo Wang 0071, Yanhua Cheng, Yan Li 0043, Dongze Hao, Xipeng Cao, Quan Chen 0006, Han Li 0005, Peng Jiang 0002, Xiangyu Zhu 0001, Zhen Lei 0001 |
ACM Multimedia | 11 |
| 2024 | STODINE: Decompose video to Object-centric Spatial-Temporal Slots for physical reasoning
Xiangyu Zhu 0001, Qu Tang, Zhaoxiang Zhang 0001, Zhen Lei 0001 |
MMAsia | 2 |
| 2024 | FusionMamba: Efficient Remote Sensing Image Fusion With State Space ModelabstractRemote sensing image fusion aims to generate a high-resolution multi/hyperspectral image by combining a high-resolution image with limited spectral data and a low-resolution image rich in spectral information. Current deep learning (DL) methods typically employ convolutional neural networks (CNNs) or Transformers for feature extraction and information integration. While CNNs are efficient, their limited receptive fields restrict their ability to capture global context. Transformers excel at learning global information but are computationally expensive. Recent advancements in the state space model (SSM), particularly Mamba, present a promising alternative by enabling global perception with low complexity. However, the potential of SSM for information integration remains largely unexplored. Therefore, we propose FusionMamba, an innovative method for efficient remote sensing image fusion. Our contributions are twofold. First, to effectively merge spatial and spectral features, we expand the single-input Mamba block to accommodate dual inputs, creating the FusionMamba block, which serves as a plug-and-play solution for information integration. Second, we incorporate Mamba and FusionMamba blocks into an interpretable network architecture tailored for remote sensing image fusion. Our designs utilize two U-shaped network branches, each primarily composed of four-directional (FD) Mamba blocks, to extract spatial and spectral features separately and hierarchically. The resulting feature maps are sufficiently merged in an auxiliary network branch constructed with FusionMamba blocks. Furthermore, we improve the representation of spectral information through an enhanced channel attention module. Quantitative and qualitative valuation results across six datasets demonstrate that our method achieves the state-of-the-art (SOTA) performance, underscoring the effectiveness of FusionMamba. The code is available athttps://github.com/PSRben/FusionMamba. Siran Peng, Xiangyu Zhu 0001, Liang-Jian Deng, Zhen Lei 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Masked Face TransformerabstractThe COVID-19 pandemic makes wearing masks mandatory. Existing CNN-based face recognition (FR) systems suffer from severe performance degradation as masks occlude the vital facial regions. Recently, Vision Transformers have shown promising performance in various vision tasks with quadratic computation costs. Swin Transformer first proposes a successive window attention mechanism allowing the cross-window connection and more computational efficiency. Despite its potential, the deployment of Swin Transformer in masked face recognition encounters two challenges: 1) the attention range is insufficient to capture locally compatible face regions. 2) Masked face recognition can be defined as an occlusion-robust classification task with a known occlusion position, i.e., the position of the mask is minor-varying, which is overlooked but efficient in improving the model’s recognition accuracy. To alleviate the above problem, we propose a Masked Face Transformer (MFT) with Masked Face-compatible Attention (MFA). The proposed MFA 1) introduces two additional window partition configurations, e.g., row shift and column shift, to enlarge the attention range in Swin with invariant computation costs, and 2) suppresses the interaction between the masked and non-masked regions to retain their discrepancies. Additionally, as mask occlusion leads to a separation between the masked and non-masked samples of the same identity, we propose to explore the relationship between them by a ClassFormer module to enhance intra-class aggregation. Extensive experiments show that MFT outperforms state-of-the-art masked face recognition methods in both simulated and real masked face testing datasets. Weisong Zhao, Xiangyu Zhu 0001, Haichao Shi, Xiaoyu Zhang 0002, Zhen Lei 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Multi-Level Pixel-Wise Correspondence Learning for 6DoF Face Pose EstimationabstractIn this paper, we focus on estimating six degrees of freedom (6DoF) pose of a face from a single RGB image, which is an important but under-investigated problem in 3D face applications such as face reconstruction, forgery detection and virtual try-on. This problem is different from traditional face pose estimation and 3D face reconstruction since the distance from camera to face should be estimated, which can not be directly regressed due to the non-linearity of the pose space. To solve the problem, we follow Perspective-n-Point (PnP) and predict the correspondences between 3D points in canonical space and 2D facial pixels on the input image to solve the 6DoF pose parameters. In this framework, the central problem of 6DoF estimation is building the correspondence matrix between a set of sampled 2D pixels and 3D points, and we propose a Correspondence Learning Transformer (CLT) to achieve this goal. Specifically, we build the 2D and 3D features with local, global, and semantic information, and employ self-attention to make the 2D and 3D features interact with each other and build the 2D–3D correspondence. Besides, we argue that 6DoF estimation is not only related with face appearance itself but also the facial external context, which contains rich information about the distance to camera. Therefore, we extract global-and-local features from the integration of face and context, where the cropped face image with smaller receptive fields concentrates on the small distortion by perspective projection, and the whole image with large receptive field provides shoulder and environment information. Experiments show that our method achieves a 2.0% improvement of$MAE_{r}$and$ADD$on ARKitFace and a 4.0%/0.7% improvement of$MAE_{t}$on ARKitFace/BIWI. Xiangyu Zhu 0001, Yueying Kao, Zhiwen Chen 0002, Jiangjing Lyu, Zhen Lei 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Grouped Knowledge Distillation for Deep Face RecognitionabstractCompared with the feature-based distillation methods, logits distillation can liberalize the requirements of consistent feature dimension between teacher and student networks, while the performance is deemed inferior in face recognition. One major challenge is that the light-weight student network has difficulty fitting the target logits due to its low model capacity, which is attributed to the significant number of identities in face recognition. Therefore, we seek to probe the target logits to extract the primary knowledge related to face identity, and discard the others, to make the distillation more achievable for the student network. Specifically, there is a tail group with near-zero values in the prediction, containing minor knowledge for distillation. To provide a clear perspective of its impact, we first partition the logits into two groups, i.e., Primary Group and Secondary Group, according to the cumulative probability of the softened prediction. Then, we reorganize the Knowledge Distillation (KD) loss of grouped logits into three parts, i.e., Primary-KD, Secondary-KD, and Binary-KD. Primary-KD refers to distilling the primary knowledge from the teacher, Secondary-KD aims to refine minor knowledge but increases the difficulty of distillation, and Binary-KD ensures the consistency of knowledge distribution between teacher and student. We experimentally found that (1) Primary-KD and Binary-KD are indispensable for KD, and (2) Secondary-KD is the culprit restricting KD at the bottleneck. Therefore, we propose a Grouped Knowledge Distillation (GKD) that retains the Primary-KD and Binary-KD but omits Secondary-KD in the ultimate KD loss calculation. Extensive experimental results on popular face recognition benchmarks demonstrate the superiority of proposed GKD over state-of-the-art methods. Weisong Zhao, Xiangyu Zhu 0001, Xiaoyu Zhang 0002, Zhen Lei 0001 |
AAAI | 2 |
| 2023 | High-Fidelity Clothed Avatar Reconstruction from a Single ImageabstractThis paper presents a framework for efficient 3D clothed avatar reconstruction. By combining the advantages of the high accuracy of optimization-based methods and the efficiency of learning-based methods, we propose a coarse-to-fine way to realize a high-fidelity clothed avatar reconstruction (CAR) from a single image. At the first stage, we use an implicit model to learn the general shape in the canonical space of a person in a learning-based way, and at the second stage, we refine the surface detail by estimating the non-rigid deformation in the posed space in an optimization way. A hyper-network is utilized to generate a good initialization so that the convergence of the optimization process is greatly accelerated. Extensive experiments on various datasets show that the proposed CAR successfully produces high-fidelity avatars for arbitrarily clothed humans in real scenes. The codes will be released in https://github.com/TingtingLiao/CAR. Tingting Liao, Yuliang Xiu, Hongwei Yi, Xudong Liu 0006, Guo-Jun Qi, Yong Zhang 0034, Xuan Wang 0009, Xiangyu Zhu 0001, Zhen Lei 0001 |
CVPR | 9 |
| 2023 | OTAvatar: One-Shot Talking Face Avatar with Controllable Tri-Plane RenderingabstractControllability, generalizability and efficiency are the major objectives of constructing face avatars represented by neural implicit field. However, existing methods have not managed to accommodate the three requirements simultaneously. They either focus on static portraits, restricting the representation ability to a specific subject, or suffer from substantial computational cost, limiting their flexibility. In this paper, we propose One-shot Talking face Avatar (OTAvatar), which constructs face avatars by a generalized controllable tri-plane rendering solution so that each personalized avatar can be constructed from only one portrait as the reference. Specifically, OTAvatar first inverts a portrait image to a motion-free identity code. Second, the identity code and a motion code are utilized to modulate an efficient CNN to generate a tri-plane formulated volume, which encodes the subject in the desired motion. Finally, volume rendering is employed to generate an image in any view. The core of our solution is a novel decoupling-by-inverting strategy that disentangles identity and motion in the latent code via optimization-based inversion. Benefiting from the efficient tri-plane representation, we achieve controllable rendering of generalized face avatar at 35 FPS on AIOO. Experiments show promising performance of crossidentity reenactment on subjects out of the training set and better 3D consistency. The code is available at https://github.com/theEricMaIOTAvatar. Zhiyuan Ma 0002, Xiangyu Zhu 0001, Guo-Jun Qi, Zhen Lei 0001, Lei Zhang 0006 |
CVPR | 2 |
| 2023 | Intrinsic Physical Concepts Discovery with Object-Centric Predictive ModelsabstractThe ability to discover abstract physical concepts and understand how they work in the world through observing lies at the core of human intelligence. The acquisition of this ability is based on compositionally perceiving the environment in terms of objects and relations in an unsupervised manner. Recent approaches learn object-centric represen-tations and capture visually observable concepts of objects, e.g., shape, size, and location. In this paper, we take a step forward and try to discover and represent intrinsic physical concepts such as mass and charge. We introduce the PHYsi-cal Concepts Inference NEtwork (PHYCINE), a system that infers physical concepts in different abstract levels with-out supervision. The key insights underlining PHYCINE are two-fold, commonsense knowledge emerges with pre-diction, and physical concepts of different abstract levels should be reasoned in a bottom-up fashion. Empirical eval-uation demonstrates that variables inferred by our system work in accordance with the properties of the corresponding physical concepts. We also show that object representations containing the discovered physical concepts variables could help achieve better performance in causal reasoning tasks, i.e., ComPhy. Qu Tang, Xiangyu Zhu 0001, Zhen Lei 0001, Zhaoxiang Zhang 0001 |
CVPR | 2 |
| 2023 | Graphics Capsule: Learning Hierarchical 3D Face Representations from 2D ImagesabstractThe function of constructing the hierarchy of objects is important to the visual process of the human brain. Previous studies have successfully adopted capsule networks to decompose the digits and faces into parts in an unsupervised manner to investigate the similar perception mechanism of neural networks. However, their descriptions are restricted to the 2D space, limiting their capacities to imitate the intrinsic 3D perception ability of humans. In this paper, we propose an Inverse Graphics Capsule Network (IGC-Net) to learn the hierarchical 3D face representations from large-scale unlabeled images. The core of IGC-Net is a new type of capsule, named graphics capsule, which represents 3D primitives with interpretable parameters in computer graphics (CG), including depth, albedo, and 3D pose. Specifically, IGC-Net first decomposes the objects into a set of semantic-consistent part-level descriptions and then assembles them into object-level descriptions to build the hierarchy. The learned graphics capsules reveal how the neural networks, oriented at visual perception, understand faces as a hierarchy of 3D models. Besides, the discovered parts can be deployed to the unsupervised face segmentation task to evaluate the semantic consistency of our method. Moreover, the part-level descriptions with explicit physical meanings provide insight into the face analysis that originally runs in a black box, such as the importance of shape and texture for face recognition. Experiments on CelebA, BP4D, and Multi-PIE demonstrate the characteristics of our IGC-Net. Chang Yu 0001, Xiangyu Zhu 0001, Zhaoxiang Zhang 0001, Zhen Lei 0001 |
CVPR | 2 |
| 2023 | Modeling Spoof Noise by De-spoofing Diffusion and its Application in Face Anti-spoofingabstractFace anti-spoofing is crucial for ensuring the security and reliability of face recognition systems. Several existing face anti-spoofing methods utilize GAN-like networks to detect presentation attacks by estimating the noise pattern of a spoof image and recovering the corresponding genuine image. But GAN’s limited face appearance space results in the denoised faces cannot cover the full data distribution of genuine faces, thereby undermining the generalization performance of such methods. In this work, we present a pioneering attempt to employ diffusion models to denoise a spoof image and restore the genuine image. The difference between these two images is considered as the spoof noise, which can serve as a discriminative cue for face anti-spoofing. We evaluate our proposed method on several intra-testing and inter-testing protocols, where the experimental results showcase the effectiveness of our method in achieving competitive performance in terms of both accuracy and generalization. Xiangyu Zhu 0001, Xiaoyu Zhang 0002, Zhen Lei 0001 |
IJCB | 2 |
| 2023 | EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face AnimationabstractSpeech-driven 3D face animation aims to generate realistic facial expressions that match the speech content and emotion. However, existing methods often neglect emotional facial expressions or fail to disentangle them from speech content. To address this issue, this paper proposes an end-to-end neural network to disentangle different emotions in speech so as to generate rich 3D facial expressions. Specifically, we introduce the emotion disentangling encoder (EDE) to disentangle the emotion and content in the speech by cross-reconstructed speech signals with different emotion labels. Then an emotion-guided feature fusion decoder is employed to generate a 3D talking face with enhanced emotion. The decoder is driven by the disentangled identity, emotional, and content embeddings so as to generate controllable personal and emotional styles. Finally, considering the scarcity of the 3D emotional talking face data, we resort to the supervision of facial blendshapes, which enables the reconstruction of plausible 3D faces from 2D emotional data, and contribute a large-scale 3D emotional talking face dataset (3D-ETF) to train the network. Our experiments and user studies demonstrate that our approach outperforms state-of-the-art methods and exhibits more diverse facial movements. We recommend watching the supplementary video: https://ziqiaopeng.github.io/emotalk Ziqiao Peng, Zhenbo Song, Xiangyu Zhu 0001, Jun He 0008, Hongyan Liu 0002, Zhaoxin Fan |
ICCV | 5 |
| 2023 | SelfTalk: A Self-Supervised Commutative Training Diagram to Comprehend 3D Talking FacesabstractSpeech-driven 3D face animation technique, extending its applications to various multimedia fields.Previous research has generated promising realistic lip movements and facial expressions from audio signals. However, traditional regression models solely driven by data face several essential problems, such as difficulties in accessing precise labels and domain gaps between different modalities, leading to unsatisfactory results lacking precision and coherence.To enhance the visual accuracy of generated lip movement while reducing the dependence on labeled data, we propose a novel framework SelfTalk, by involving self-supervision in a cross-modals network system to learn 3D talking faces. The framework constructs a network system consisting of three modules: facial animator, speech recognizer, and lip-reading interpreter. The core of SelfTalk is a commutative training diagram that facilitates compatible features exchange among audio, text, and lip shape, enabling our models to learn the intricate connection between these factors. The proposed framework leverages the knowledge learned from the lip-reading interpreter to generate more plausible lip shapes. Extensive experiments and user studies demonstrate that our proposed approach achieves state-of-the-art performance both qualitatively and quantitatively. We recommend watching the supplementary video. Ziqiao Peng, Yihao Luo, Xiangyu Zhu 0001, Hongyan Liu 0002, Jun He 0008, Zhaoxin Fan |
ACM Multimedia | 5 |
| 2023 | Cross-Architecture Distillation for Face RecognitionabstractTransformers have emerged as the superior choice for face recognition tasks, but their insufficient platform acceleration hinders their application on mobile devices. In contrast, Convolutional Neural Networks (CNNs) capitalize on hardware-compatible acceleration libraries. Consequently, it has become indispensable to preserve the distillation efficacy when transferring knowledge from a Transformer-based teacher model to a CNN-based student model, known as Cross-Architecture Knowledge Distillation (CAKD). Despite its potential, the deployment of CAKD in face recognition encounters two challenges: 1) the teacher and student share disparate spatial information for each pixel, obstructing the alignment of feature space, and 2) the teacher network is not trained in the role of a teacher, lacking proficiency in handling distillation-specific knowledge. To surmount these two constraints, 1) we first introduce a Unified Receptive Fields Mapping module (URFM) that maps pixel features of the teacher and student into local features with unified receptive fields, thereby synchronizing the pixel-wise spatial information of teacher and student. Subsequently, 2) we develop an Adaptable Prompting Teacher network (APT) that integrates prompts into the teacher, enabling it to manage distillation-specific knowledge while preserving the model's discriminative capacity. Extensive experiments on popular face benchmarks and two large-scale verification sets demonstrate the superiority of our method. Weisong Zhao, Xiangyu Zhu 0001, Zhixiang He, Xiaoyu Zhang 0002, Zhen Lei 0001 |
ACM Multimedia | 2 |
| 2023 | Spoof-Guided Image Decomposition for Face Anti-spoofing
Xiangyu Zhu 0001, Xiaoyu Zhang 0002, Shukai Chen, Peng Li 0035, Zhen Lei 0001 |
PRCV (5) | 2 |
| 2023 | Face Forgery Detection by 3D Decomposition and Composition SearchabstractDetecting digital face manipulation has attracted extensive attention due to fake media's potential risks to the public. However, recent advances have been able to reduce the forgery signals to a low magnitude. Decomposition, which reversibly decomposes an image into several constituent elements, is a promising way to highlight the hidden forgery details. In this paper, we investigate a novel 3D decomposition based method that considers a face image as the production of the interaction between 3D geometry and lighting environment. Specifically, we disentangle a face image into four graphics components including 3D shape, lighting, common texture, and identity texture, which are respectively constrained by 3D morphable model, harmonic reflectance illumination, and PCA texture model. Meanwhile, we build a fine-grained morphing network to predict 3D shapes with pixel-level accuracy to reduce the noise in the decomposed elements. Moreover, we propose a composition search strategy that enables an automatic construction of an architecture to mine forgery clues from forgery-relevant components. Extensive experiments validate that the decomposed components highlight forgery artifacts, and the searched architecture extracts discriminative forgery features. Thus, our method achieves the state-of-the-art performance. Xiangyu Zhu 0001, Hongyan Fei, Tianshuo Zhang, Xiaoyu Zhang 0002, Stan Z. Li, Zhen Lei 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Beyond 3DMM: Learning to Capture High-Fidelity 3D Face Shapeabstract3D Morphable Model (3DMM) fitting has widely benefited face analysis due to its strong 3D priori. However, previous reconstructed 3D faces suffer from degraded visual verisimilitude due to the loss of fine-grained geometry, which is attributed to insufficient ground-truth 3D shapes, unreliable training strategies and limited representation power of 3DMM. To alleviate this issue, this paper proposes a complete solution to capture the personalized shape so that the reconstructed shape looks identical to the corresponding person. Specifically, given a 2D image as the input, we virtually render the image in several calibrated views to normalize pose variations while preserving the original image geometry. A many-to-one hourglass network serves as the encode-decoder to fuse multiview features and generate vertex displacements as the fine-grained geometry. Besides, the neural network is trained by directly optimizing the visual effect, where two 3D shapes are compared by measuring the similarity between the multiview images rendered from the shapes. Finally, we propose to generate the ground-truth 3D shapes by registering RGB-D images followed by pose and shape augmentation, providing sufficient data for network training. Experiments on several challenging protocols demonstrate the superior reconstruction accuracy of our proposal on the face shape. Xiangyu Zhu 0001, Chang Yu 0001, Di Huang 0001, Zhen Lei 0001, Hao Wang 0074, Stan Z. Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Toward 3D Face Reconstruction in Perspective Projection: Estimating 6DoF Face Pose From Monocular ImageabstractIn 3D face reconstruction, orthogonal projection has been widely employed to substitute perspective projection to simplify the fitting process. This approximation performs well when the distance between camera and face is far enough. However, in some scenarios that the face is very close to camera or moving along the camera axis, the methods suffer from the inaccurate reconstruction and unstable temporal fitting due to the distortion under the perspective projection. In this paper, we aim to address the problem of single-image 3D face reconstruction under perspective projection. Specifically, a deep neural network, Perspective Network (PerspNet), is proposed to simultaneously reconstruct 3D face shape in canonical space and learn the correspondence between 2D pixels and 3D points, by which the 6DoF (6 Degrees of Freedom) face pose can be estimated to represent perspective projection. Besides, we contribute a large ARKitFace dataset to enable the training and evaluation of 3D face reconstruction solutions under the scenarios of perspective projection, which has 902,724 2D facial images with ground-truth 3D face mesh and annotated 6DoF pose parameters. Experimental results show that our approach outperforms current state-of-the-art methods by a significant margin. The code and data are available at https://github.com/cbsropenproject/6dof_face. Yueying Kao, Bowen Pan, Jiangjing Lyu, Xiangyu Zhu 0001, Yuanzhang Chang, Zhen Lei 0001 |
IEEE Trans. Image Process. | 5 |
| 2023 | Human Parsing With Part-Aware Relation ModelingabstractIn this paper, a Part-aware Relation Modeling (PRM) is developed to handle the task of human parsing. For pixel-level recognition, it is essential to generate features with adaptive context for various sizes and shapes of human parts. To address the issue, we adaptively capture contexts based on the part-aware relation mechanism. PRM mainly consists of three modules, including a part class module, a part-relation aggregation module, and a part-relation dispersion module. The part class module selectively enhances spatial details of the high-level features to obtain enhanced original features, and then extracts the high-level representations of every human part from a categorical perspective. The part-relation aggregation module is developed to extract the representative global context by exploring associated semantics of human parts, adaptively augmenting the context for human parts. The part-relation dispersion module is designed to generate the discriminative and effective local context and neglect the distracting one by making the affinity of human parts disperse. It ensures that features of the same class will be close to each other and away from those of different classes. By fusing the outputs of the two part-relation modules and the first outputs of the part class module, our PRM produces adaptive contextual features for diverse sizes of human parts, boosting the parsing accuracy. Extensive experiments are conducted to validate the effectiveness of our network, and a new state-of-the-art segmentation performance is achieved on three challenging human parsing datasets,i.e., PASCAL-Person-Part, LIP, and CIHP. PRM is also extended to other tasks like animal parsing, and exhibits its generality. Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang, Xiangyu Zhu 0001, Zhen Lei 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | Deconfounding Physical Dynamics with Global Causal Relation and Confounder Transmission for Counterfactual PredictionabstractDiscovering the underneath causal relations is the fundamental ability for reasoning about the surrounding environment and predicting the future states in the physical world. Counterfactual prediction from visual input, which requires simulating future states based on unrealized situations in the past, is a vital component in causal relation tasks. In this paper, we work on the confounders that have effect on the physical dynamics, including masses, friction coefficients, etc., to bridge relations between the intervened variable and the affected variable whose future state may be altered. We propose a neural network framework combining Global Causal Relation Attention (GCRA) and Confounder Transmission Structure (CTS). The GCRA looks for the latent causal relations between different variables and estimates the confounders by capturing both spatial and temporal information. The CTS integrates and transmits the learnt confounders in a residual way, so that the estimated confounders can be encoded into the network as a constraint for object positions when performing counterfactual prediction. Without any access to ground truth information about confounders, our model outperforms the state-of-the-art method on various benchmarks by fully utilizing the constraints of confounders. Extensive experiments demonstrate that our model can generalize to unseen environments and maintain good performance. Zongzhao Li, Xiangyu Zhu 0001, Zhen Lei 0001, Zhaoxiang Zhang 0001 |
AAAI | 2 |
| 2022 | HP-Capsule: Unsupervised Face Part Discovery by Hierarchical Parsing Capsule NetworkabstractCapsule networks are designed to present the objects by a set of parts and their relationships, which provide an insight into the procedure of visual perception. Although recent works have shown the success of capsule networks on simple objects like digits, the human faces with homologous structures, which are suitable for capsules to describe, have not been explored. In this paper, we propose a Hierarchical Parsing Capsule Network (HP-Capsule) for unsupervised face subpart-part discovery. When browsing large-scale face images without labels, the network first encodes the frequently observed patterns with a set of explainable subpart capsules. Then, the subpart capsules are assembled into part-level capsules through a Transformer-based Parsing Module (TPM) to learn the compositional relations between them. During training as the face hierarchy is progressively built and refined, the part capsules adaptively encode the face parts with semantic consistency. HP-Capsule extends the application of capsule networks from digits to human faces and takes a step forward to show how the neural networks understand homologous objects without human intervention. Besides, HP-Capsule gives unsupervised face segmentation results by the covered regions of part capsules, enabling qualitative and quantitative evaluation. Experiments on BP4D and Multi-PIE datasets show the effectiveness of our method. Chang Yu 0001, Xiangyu Zhu 0001, Zidu Wang, Zhaoxiang Zhang 0001, Zhen Lei 0001 |
CVPR | 2 |
| 2022 | Object Dynamics Distillation for Scene Decomposition and Representation
Qu Tang, Xiangyu Zhu 0001, Zhen Lei 0001, Zhaoxiang Zhang 0001 |
ICLR | 2 |
| 2022 | Caged Monkey Dataset: A New Benchmark for Caged Monkey Pose Estimation
Xiangyu Zhu 0001, Zhen Lei 0001, Xibo Ma |
PRCV (4) | 2 |
| 2022 | Consistent Sub-Decision Network for Low-Quality Masked Face RecognitionabstractThe COVID-19 pandemic makes wearing masks mandatory in supermarkets, pharmacies, public transport, etc. Existing facial recognition systems encounter severe performance degradation as the masks occlude key facial regions. Recently, simulation-based methods are proposed to generate masked faces from unmasked faces. However, among simulated faces, there are low-quality samples with negative occlusion, which leads to ambiguous or absent facial features. In this paper, we propose a consistent sub-decision network to obtain sub-decisions that correspond to different facial regions and constrain sub-decisions by weighted bidirectional KL divergence to make the network concentrate on the upper faces without occlusion. In addition, we perform knowledge distillation to drive the masked face embeddings towards an approximation of the original data distribution to mitigate the information loss. Experiments show that the proposed method performs better than the baseline on public masked face recognition datasets, i.e., RMFD, MFR2, and MLFW. Weisong Zhao, Xiangyu Zhu 0001, Haichao Shi, Xiaoyu Zhang 0002, Zhen Lei 0001 |
IEEE Signal Process. Lett. | 2 |
| 2021 | Face Forgery Detection by 3D DecompositionabstractDetecting digital face manipulation has attracted extensive attention due to fake media’s potential harms to the public. However, recent advances have been able to reduce the forgery signals to a low magnitude. Decomposition, which reversibly decomposes an image into several constituent elements, is a promising way to highlight the hidden forgery details. In this paper, we consider a face image as the production of the intervention of the underlying 3D geometry and the lighting environment, and decompose it in a computer graphics view. Specifically, by disentangling the face image into 3D shape, common texture, identity texture, ambient light, and direct light, we find the devil lies in the direct light and the identity texture. Based on this observation, we propose to utilize facial detail, which is the combination of direct light and identity texture, as the clue to detect the subtle forgery patterns. Besides, we highlight the manipulated region with a supervised attention mechanism and introduce a two-stream structure to exploit both face image and facial detail together as a multi-modality task. Extensive experiments indicate the effectiveness of the extra features extracted from the facial detail, and our method achieves the state-of-the-art performance. Xiangyu Zhu 0001, Hao Wang 0074, Hongyan Fei, Zhen Lei 0001, Stan Z. Li |
CVPR | 1 |
| 2021 | SADet: Learning An Efficient and Accurate Pedestrian DetectorabstractAlthough the anchor-based detectors have taken a big step forward in pedestrian detection, the overall performance of algorithm still needs further improvement for practical applications, e.g., a good trade-off between the accuracy and efficiency. To this end, this paper proposes a series of systematic optimization strategies for the detection pipeline of one-stage detector, forming a single shot anchor-based detector (SADet) for efficient and accurate pedestrian detection, which includes three main improvements. Firstly, we optimize the sample generation process by assigning soft labels to the outlier samples to generate semi-positive samples with continuous tag value between 0 and 1. Secondly, a novel Center-IoU loss is applied as a new regression loss for bounding box regression, which not only retains the good characteristics of IoU loss, but also solves some defects of it. Thirdly, we also design Cosine-NMS for the post-processing of predicted bounding boxes, and further propose adaptive anchor matching to enable the model to adaptively match the anchor boxes to full or visible bounding boxes according to the degree of occlusion. Though structurally simple, it presents state-of-the-art result and real-time speed of 20 FPS for VGA-resolution images (640×480) tested on one GeForce GTX 1080Ti GPU on challenging pedestrian detection benchmarks, i.e., CityPersons, Caltech, and human detection benchmark CrowdHuman, leading to a new attractive pedestrian detector. Chubin Zhuang, Zongzhao Li, Xiangyu Zhu 0001, Zhen Lei 0001, Stan Z. Li |
IJCB | 3 |
| 2021 | Multi-initialization Optimization Network for Accurate 3D Human Pose and Shape Estimationabstract3D human pose and shape recovery from a monocular RGB image is a challenging task. Existing learning based methods highly depend on weak supervision signals, e.g. 2D and 3D joint location, due to the lack of in-the-wild paired 3D supervision. However, considering the 2D-to-3D ambiguities existed in these weak supervision labels, the network is easy to get stuck in local optima when trained with such labels. In this paper, we reduce the ambituity by optimizing multiple initializations. Specifically, we propose a three-stage framework named Multi-Initialization Optimization Network (MION). In the first stage, we strategically select different coarse 3D reconstruction candidates which are compatible with the 2D keypoints of input sample. Each coarse reconstruction can be regarded as an initialization leads to one optimization branch. In the second stage, we design a mesh refinement transformer (MRT) to respectively refine each coarse reconstruction result via a self-attention mechanism. Finally, a Consistency Estimation Network (CEN) is proposed to find the best result from mutiple candidates by evaluating if the visual evidence in RGB image matches a given 3D reconstruction. Experiments demonstrate that our Multi-Initialization Optimization Network outperforms existing 3D mesh based methods on multiple public benchmarks. Zhiwei Liu 0004, Xiangyu Zhu 0001, Lu Yang 0006, Ming Tang 0001, Zhen Lei 0001, Guibo Zhu, Xuetao Feng, Yan Wang 0068, Jinqiao Wang |
ACM Multimedia | 2 |
| 2021 | Prior-knowledge and attention based meta-learning for few-shot learning
Yunxiao Qin, Xiangyu Zhu 0001, Jingping Shi, Guo-Jun Qi, Zhen Lei 0001 |
Knowl. Based Syst. | 5 |
| 2021 | Fast Adapting Without Forgetting for Face RecognitionabstractAlthough face recognition has made dramatic improvements in recent years, there are still many challenges in real-world applications such as face recognition for the elderly and children, for the surveillance scenes and for Near infrared vs. Visible light (NIR-VIS) heterogeneous scene, etc. Due to the existence of these challenges, there are usually domain gaps between training (source domain) and test (target domain). A common way to improve the performance on the target domain is fine-tuning the base model trained on source domain using target data. However, it will severely degrade performance on the source domain. Another way which jointly trains models using both source and target data, suffers from the heavy computations and large data storage, especially when we continue to encounter new domains. In response to these problems, we introduce a new challenging task: Single Exemplar Domain Incremental Learning (SE-DIL), which utilizes the target domain data and just one exemplar per identity from source domain data to quickly improve the performance on the target domain while keeping the performance on the source domain. To deal with SE-DIL, we propose our Fast Adapting without Forgetting (FAwF) method with three components: margin-based exemplar selection, prototype-based class extension and hard&soft knowledge distillation. Through FAwF, we can well maintain the source domain performance with only one sample per source domain class, greatly reducing the fine-tuning time-cost and data storage. Besides, we collected a large-scale children face dataset KidsFace with 12 K identities for studying the SE-DIL in face recognition. Extensive analysis and experiments on our KidsFace-Test protocol and other challenging face test sets show that our method performs better than the state-of-the-art methods on both target and source domain. Xiangyu Zhu 0001, Zhen Lei 0001, Stan Z. Li |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Decomposed Meta Batch Normalization for Fast Domain Adaptation in Face RecognitionabstractFace recognition systems are sometimes deployed to a target domain with limited unlabeled samples available. For instance, a model trained on the large-scale webfaces may be required to adapt to a NIR-VIS scenario via very limited unlabeled faces. This situation poses a great challenge to Unsupervised Domain Adaptation with Limited samples for Face Recognition (UDAL-FR), which is less studied in previous works. In this paper, with deep learning methods, we propose a novel training remedy by decomposing the model into the weight parameters and the BN statistics in the training phase. Based on decomposing, we design a novel framework via meta-learning, calledDecomposed Meta Batch Normalization(DMBN) for fast domain adaptation in face recognition. DMBN trains the network such that domain-invariant information is prone to store in the weight parameters and domain-specific knowledge tends to be represented by the BN statistics. Specifically, DMBN constructs distribution-shifted tasks via domain-aware sampling, on which several meta-gradients are obtained by optimizing discriminative representations across different BNs. Finally, the weight parameters are updated with these meta-gradients for better consistency across different BNs. With the learned weight parameters, the adaptation is very fast since only the BN updating on limited data is needed. We propose two UDAL-FR benchmarks to evaluate the domain-adaptive ability of a model with limited unlabeled samples. Extensive experiments validate the efficacy of our proposed DMBN. Jianzhu Guo, Xiangyu Zhu 0001, Zhen Lei 0001, Stan Z. Li |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | Learning Meta Model for Zero- and Few-Shot Face Anti-SpoofingabstractFace anti-spoofing is crucial to the security of face recognition systems. Most previous methods formulate face anti-spoofing as a supervised learning problem to detect various predefined presentation attacks, which need large scale training data to cover as many attacks as possible. However, the trained model is easy to overfit several common attacks and is still vulnerable to unseen attacks. To overcome this challenge, the detector should: 1) learn discriminative features that can generalize to unseen spoofing types from predefined presentation attacks; 2) quickly adapt to new spoofing types by learning from both the predefined attacks and a few examples of the new spoofing types. Therefore, we define face anti-spoofing as a zero- and few-shot learning problem. In this paper, we propose a novel Adaptive Inner-update Meta Face Anti-Spoofing (AIM-FAS) method to tackle this problem through meta-learning. Specifically, AIM-FAS trains a meta-learner focusing on the task of detecting unseen spoofing types by learning from predefined living and spoofing faces and a few examples of new attacks. To assess the proposed approach, we propose several benchmarks for zero- and few-shot FAS. Experiments show its superior performances on the presented benchmarks to existing methods in existing zero-shot FAS protocols. Yunxiao Qin, Xiangyu Zhu 0001, Zitong Yu, Tianyu Fu 0001, Jingping Shi, Zhen Lei 0001 |
AAAI | 3 |
| 2020 | Domain Balancing: Face Recognition on Long-Tailed DomainsabstractLong-tailed problem has been an important topic in face recognition task. However, existing methods only concentrate on the long-tailed distribution of classes. Differently, we devote to the long-tailed domain distribution problem, which refers to the fact that a small number of domains frequently appear while other domains far less existing. The key challenge of the problem is that domain labels are too complicated (related to race, age, pose, illumination, etc.) and inaccessible in real applications. In this paper, we propose a novel Domain Balancing (DB) mechanism to handle this problem. Specifically, we first propose a Domain Frequency Indicator (DFI) to judge whether a sample is from head domains or tail domains. Secondly, we formulate a light-weighted Residual Balancing Mapping (RBM) block to balance the domain distribution by adjusting the network according to DFI. Finally, we propose a Domain Balancing Margin (DBM) in the loss function to further optimize the feature space of the tail domains to improve generalization. Extensive analysis and experiments on several face recognition benchmarks demonstrate that the proposed method effectively enhances the generalization capacities and achieves superior performance. Xiangyu Zhu 0001, Jianzhu Guo, Zhen Lei 0001 |
CVPR | 2 |
| 2020 | Learning Meta Face Recognition in Unseen DomainsabstractFace recognition systems are usually faced with unseen domains in real-world applications and show unsatisfactory performance due to their poor generalization. For example, a well-trained model on webface data cannot deal with the ID vs. Spot task in surveillance scenario. In this paper, we aim to learn a generalized model that can directly handle new unseen domains without any model updating. To this end, we propose a novel face recognition method via meta-learning named Meta Face Recognition (MFR). MFR synthesizes the source/target domain shift with a meta-optimization objective, which requires the model to learn effective representations not only on synthesized source domains but also on synthesized target domains. Specifically, we build domain-shift batches through a domain-level sampling strategy and get back-propagated gradients/meta-gradients on synthesized source/target domains by optimizing multi-domain distributions. The gradients and meta-gradients are further combined to update the model to improve generalization. Besides, we propose two benchmarks for generalized face recognition evaluation. Experiments on our benchmarks validate the generalization of our method compared to several baselines and other state-of-the-arts. The proposed benchmarks and code will be available at https://github.com/cleardusk/MFR. Jianzhu Guo, Xiangyu Zhu 0001, Zhen Lei 0001, Stan Z. Li |
CVPR | 2 |
| 2020 | Deep Spatial Gradient and Temporal Depth Learning for Face Anti-SpoofingabstractFace anti-spoofing is critical to the security of face recognition systems. Depth supervised learning has been proven as one of the most effective methods for face anti-spoofing. Despite the great success, most previous works still formulate the problem as a single-frame multi-task one by simply augmenting the loss with depth, while neglecting the detailed fine-grained information and the interplay between facial depths and moving patterns. In contrast, we design a new approach to detect presentation attacks from multiple frames based on two insights: 1) detailed discriminative clues (e.g., spatial gradient magnitude) between living and spoofing face may be discarded through stacked vanilla convolutions, and 2) the dynamics of 3D moving faces provide important clues in detecting the spoofing faces. The proposed method is able to capture discriminative details via Residual Spatial Gradient Block (RSGB) and encode spatio-temporal information from Spatio-Temporal Propagation Module (STPM) efficiently. Moreover, a novel Contrastive Depth Loss is presented for more accurate depth supervision. To assess the efficacy of our method, we also collect a Double-modal Anti-spoofing Dataset (DMAD) which provides actual depth for each sample. The experiments demonstrate that the proposed approach achieves state-of-the-art results on five benchmark datasets including OULU-NPU, SiW, CASIA-MFSD, Replay-Attack, and the new DMAD. Codes will be available at https://github.com/clks-wzz/FAS-SGTD. Zitong Yu, Xiangyu Zhu 0001, Yunxiao Qin, Qiusheng Zhou, Zhen Lei 0001 |
CVPR | 4 |
| 2020 | Towards Fast, Accurate and Stable 3D Dense Face Alignment
Jianzhu Guo, Xiangyu Zhu 0001, Yang Yang 0062, Fan Yang 0062, Zhen Lei 0001, Stan Z. Li |
ECCV (19) | 2 |
| 2020 | Beyond 3DMM Space: Towards Fine-Grained 3D Face Reconstruction
Xiangyu Zhu 0001, Fan Yang 0062, Di Huang 0001, Chang Yu 0001, Hao Wang 0074, Jianzhu Guo, Zhen Lei 0001, Stan Z. Li |
ECCV (8) | 1 |
| 2020 | Out-of-Distribution Detection for Reliable Face RecognitionabstractIn real applications, face recognition systems are always faced with non-face inputs and low-quality faces due to the complicated conditions like mis-detections by face detectors. However, in deep learning based methods, these outliers are always ignored during training phase and the models tend to make unreasonable decisions on these images. For example, matching a texture-rich patch to an old-man face overconfidently. We formulate this challenge on the task of out-of-distribution detection (OOD), where a network must determine whether or not an input is outside of the set on which the network can safely perform. In this paper, we propose to detect out-of-distribution samples based on uncertainty prediction and the L2-norm of features, so as to effectively filter out non-face and low-quality faces. We demonstrate that the proposed method can reliably detect out-of-distribution samples and improve the performance of face recognition, without the need of labelled OOD data. Chang Yu 0001, Xiangyu Zhu 0001, Zhen Lei 0001, Stan Z. Li |
IEEE Signal Process. Lett. | 2 |
| 2019 | 3DMA: A Multi-modality 3D Mask Face Anti-spoofing DatabaseabstractBenefiting from publicly available databases, face anti-spoofing has recently gained extensive attention in the academic community. However, most of the existing databases focus on the 2D object attacks, including photo and video attacks. The only two public 3D mask face anti-spoofing database are very small. In this paper, we release a multi-modality 3D mask face anti-spoofing database named 3DMA, which contains 920 videos of 67 genuine subjects wearing 48 kinds of 3D masks, captured in visual (VIS) and near-infrared (NIR) modalities. To simulate the real world scenarios, two illumination and four capturing distance settings are deployed during the collection process. To the best of our knowledge, the proposed database is currently the most extensive public database for 3D mask face anti-spoofing. Furthermore, we build three protocols for performance evaluation under different illumination conditions and distances. Experimental results with Convolutional Neural Network (CNN) and LBP-based methods reveal that our proposed 3DMA is indeed a challenge for face anti-spoofing. This database is available at http://www.cbsr.ia.ac.cn/english/3DMA.html. We hope our public 3DMA database can help to pave the way for further research on 3D mask face anti-spoofing. Jinchuan Xiao, Yinhang Tang, Jianzhu Guo, Yang Yang 0062, Xiangyu Zhu 0001, Zhen Lei 0001, Stan Z. Li |
AVSS | 5 |
| 2019 | Semantic Alignment: Finding Semantically Consistent Ground-Truth for Facial Landmark DetectionabstractRecently, deep learning based facial landmark detection has achieved great success. Despite this, we notice that the semantic ambiguity greatly degrades the detection performance. Specifically, the semantic ambiguity means that some landmarks (e.g. those evenly distributed along the face contour) do not have clear and accurate definition, causing inconsistent annotations by annotators. Accordingly, these inconsistent annotations, which are usually provided by public databases, commonly work as the ground-truth to supervise network training, leading to the degraded accuracy. To our knowledge, little research has investigated this problem. In this paper, we propose a novel probabilistic model which introduces a latent variable, i.e. the `real' ground-truth which is semantically consistent, to optimize. This framework couples two parts (1) training landmark detection CNN and (2) searching the `real' ground-truth. These two parts are alternatively optimized: the searched `real' ground-truth supervises the CNN training; and the trained CNN assists the searching of `real' ground-truth. In addition, to recover the unconfidently predicted landmarks due to occlusion and low quality, we propose a global heatmap correction unit (GHCU) to correct outliers by considering the global face shape as a constraint. Extensive experiments on both image-based (300W and AFLW) and video-based (300-VW) databases demonstrate that our method effectively improves the landmark detection accuracy and achieves the state of the art performance. Zhiwei Liu 0004, Xiangyu Zhu 0001, Guosheng Hu, Haiyun Guo, Ming Tang 0001, Zhen Lei 0001, Neil Robertson 0002, Jinqiao Wang |
CVPR | 2 |
| 2019 | AdaptiveFace: Adaptive Margin and Sampling for Face RecognitionabstractTraining large-scale unbalanced data is the central topic in face recognition. In the past two years, face recognition has achieved remarkable improvements due to the introduction of margin based Softmax loss. However, these methods have an implicit assumption that all the classes possess sufficient samples to describe its distribution, so that a manually set margin is enough to equally squeeze each intra-class variations. However, real face datasets are highly unbalanced, which means the classes have tremendously different numbers of samples. In this paper, we argue that the margin should be adapted to different classes. We propose the Adaptive Margin Softmax to adjust the margins for different classes adaptively. In addition to the unbalance challenge, face data always consists of large-scale classes and samples. Smartly selecting valuable classes and samples to participate in the training makes the training more effective and efficient. To this end, we also make the sampling process adaptive in two folds: Firstly, we propose the Hard Prototype Mining to adaptively select a small number of hard classes to participate in classification. Secondly, for data sampling, we introduce the Adaptive Data Sampling to find valuable samples for training adaptively. We combine these three parts together as AdaptiveFace. Extensive analysis and experiments on LFW, LFW BLUFR and MegaFace show that our method performs better than state-of-the-art methods using the same network architecture and training dataset. Code is available at https://github.com/haoliu1994/AdaptiveFace. Xiangyu Zhu 0001, Zhen Lei 0001, Stan Z. Li |
CVPR | 2 |
| 2019 | Weakly Aligned Cross-Modal Learning for Multispectral Pedestrian DetectionabstractMultispectral pedestrian detection has shown great advantages under poor illumination conditions, since the thermal modality provides complementary information for the color image. However, real multispectral data suffers from the position shift problem, i.e. the color-thermal image pairs are not strictly aligned, making one object has different positions in different modalities. In deep learning based methods, this problem makes it difficult to fuse the feature maps from both modalities and puzzles the CNN training. In this paper, we propose a novel Aligned Region CNN (AR-CNN) to handle the weakly aligned multispectral data in an end-to-end way. Firstly, we design a Region Feature Alignment (RFA) module to capture the position shift and adaptively align the region features of the two modalities. Secondly, we present a new multimodal fusion method, which performs feature re-weighting to select more reliable features and suppress the useless ones. Besides, we propose a novel RoI jitter strategy to improve the robustness to unexpected shift patterns of different devices and system settings. Finally, since our method depends on a new kind of labelling: bounding boxes that match each modality, we manually relabel the KAIST dataset by locating bounding boxes in both modalities and building their relationships, providing a new KAIST-Paired Annotation. Extensive experimental validations on existing datasets are performed, demonstrating the effectiveness and robustness of the proposed method. Code and data are available at https://github.com/luzhang16/AR-CNN. Lu Zhang 0054, Xiangyu Zhu 0001, Xu Yang 0004, Zhen Lei 0001, Zhiyong Liu 0001 |
ICCV | 2 |
| 2019 | Pose-Weighted Gan for Photorealistic Face FrontalizationabstractFace recognition methods have achieved high accuracy when faces are captured in frontal pose and constrained scenes. However, severe drop in accuracy is observed when large pose variations exist. The main reason is that the large yaw angle leads to ID information loss. In this paper, we intend to solve the large pose variations in a generation manner. Specifically, we propose a Pose-Weighted Generative Adversarial Network (PW-GAN) for photorealistic frontal view synthesis. We find frontalizing the faces in large poses (yaw angle larger than 60°) is so difficult that the results are not photorealistic and the ID information is lost. To simplify the problem, we first frontalize the face image through 3D face model, which is then used to guide the network predicting. Second, we refine the pose code in the loss function to make the network pay more attention to large poses. Quantitative and qualitative experimental results on the Multi-PIE and LFW demonstrate our method achieves state of the art. Su-Fang Zhang, Qinghai Miao, Min Huang 0009, Xiangyu Zhu 0001, Yingying Chen 0003, Zhen Lei 0001, Jinqiao Wang |
ICIP | 4 |
| 2019 | Large-Scale Bisample Learning on ID Versus Spot Face Recognition
Xiangyu Zhu 0001, Zhen Lei 0001, Hailin Shi, Fan Yang 0062, Dong Yi, Guo-Jun Qi, Stan Z. Li |
Int. J. Comput. Vis. | 1 |
| 2019 | Face Alignment in Full Pose Range: A 3D Total SolutionabstractFace alignment, which fits a face model to an image and extracts the semantic meanings of facial pixels, has been an important topic in the computer vision community. However, most algorithms are designed for faces in small to medium poses (yaw angle is smaller than 45 degree), which lack the ability to align faces in large poses up to 90 degree. The challenges are three-fold. First, the commonly used landmark face model assumes that all the landmarks are visible and is therefore not suitable for large poses. Second, the face appearance varies more drastically across large poses, from the frontal view to the profile view. Third, labelling landmarks in large poses is extremely challenging since the invisible landmarks have to be guessed. In this paper, we propose to tackle these three challenges in an new alignment framework termed 3D Dense Face Alignment (3DDFA), in which a dense 3D Morphable Model (3DMM) is fitted to the image via Cascaded Convolutional Neural Networks. We also utilize 3D information to synthesize face images in profile views to provide abundant samples for training. Experiments on the challenging AFLW database show that the proposed approach achieves significant improvements over the state-of-the-art methods. Xiangyu Zhu 0001, Xiaoming Liu 0002, Zhen Lei 0001, Stan Z. Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Co-Referenced Subspace ClusteringabstractSubspace clustering refers to the problem of grouping data into their underlying groups. To address this task, spectral clustering based technique is arguably one of the most popular approaches, and its performance largely depends on the constructed similarity. However, most existing works merely employ the primary representation (e.g., sparse or low-rank representation) as the similarity. In this paper, we propose to explore a high-level co-referenced similarity by employing the Hilbert-Schmidt Independence Criterion (HSIC). Moreover, geometry interpretation of the advantage of our co-referenced similarity is provided. Representation-induced kernels such as Mahalanobis metric, can also be easily embedded into the formulation. Extensive experiments on both synthetic and real-world data are conducted to show the superiority of the proposed method over the state-of-the-art alternatives. Xiaobo Wang 0001, Zhen Lei 0001, Hailin Shi, Xiaojie Guo 0001, Xiangyu Zhu 0001, Stan Z. Li |
ICME | 5 |
| 2018 | Detecting Face with Densely Connected Face Proposal Network
Xiangyu Zhu 0001, Zhen Lei 0001, Xiaobo Wang 0001, Hailin Shi, Stan Z. Li |
Neurocomputing | 2 |
| 2017 | Multi-modality Network with Visual and Geometrical Information for Micro Emotion RecognitionabstractMicro emotion recognition is a very challenging problem because of the subtle appearance variants among different facial expression classes. To deal with the mentioned problem, we proposed a multi-modality convolutional neural networks (CNNs) based on visual and geometrical information in this paper. The visual face image and structured geometry are embedded into a unified network and the recognition accuracy can be benefic from the fused information. The proposed network includes two branches. The first branch is used to extract visual feature from color face images, and another branch is used to extract the geometry feature from 68 facial landmarks. Then, both visual and geometry features are concatenated into a long vector. Finally, the concatenated vector is fed to the hinge loss layer. Compared with the CNN architecture only used face images, our method is more effective and has got better performance. In the final testing phase of Micro Emotion Challenge1, our method has got the first place with the misclassification of 80.212137. Jianzhu Guo, Jinlin Wu, Jun Wan 0001, Xiangyu Zhu 0001, Zhen Lei 0001, Stan Z. Li |
FG | 5 |
| 2017 | FaceBoxes: A CPU real-time face detector with high accuracyabstractAlthough tremendous strides have been made in face detection, one of the remaining open challenges is to achieve real-time speed on the CPU as well as maintain high performance, since effective models for face detection tend to be computationally prohibitive.To address this challenge, we propose a novel face detector, named FaceBoxes, with superior performance on both speed and accuracy.Specifically, our method has a lightweight yet powerful network structure that consists of the Rapidly Digested Convolutional Layers (RDCL) and the Multiple Scale Convolutional Layers (MSCL).The RDCL is designed to enable Face-Boxes to achieve real-time speed on the CPU.The MSCL aims at enriching the receptive fields and discretizing anchors over different layers to handle faces of various scales.Besides, we propose a new anchor densification strategy to make different types of anchors have the same density on the image, which significantly improves the recall rate of small faces.As a consequence, the proposed detector runs at 20 FPS on a single CPU core and 125 FPS using a GPU for VGA-resolution images.Moreover, the speed of FaceBoxes is invariant to the number of faces.We comprehensively evaluate this method and present stateof-the-art detection performance on several face detection benchmark datasets, including the AFW, PASCAL face, and FDDB. Xiangyu Zhu 0001, Zhen Lei 0001, Hailin Shi, Xiaobo Wang 0001, Stan Z. Li |
IJCB | 2 |
| 2017 | S^3FD: Single Shot Scale-Invariant Face DetectorabstractThis paper presents a real-time face detector, named Single Shot Scale-invariant Face Detector (S3FD), which performs superiorly on various scales of faces with a single deep neural network, especially for small faces. Specifically, we try to solve the common problem that anchor-based detectors deteriorate dramatically as the objects become smaller. We make contributions in the following three aspects: 1) proposing a scale-equitable face detection framework to handle different scales of faces well. We tile anchors on a wide range of layers to ensure that all scales of faces have enough features for detection. Besides, we design anchor scales based on the effective receptive field and a proposed equal proportion interval principle; 2) improving the recall rate of small faces by a scale compensation anchor matching strategy; 3) reducing the false positive rate of small faces via a max-out background label. As a consequence, our method achieves state-of-the-art detection performance on all the common face detection benchmarks, including the AFW, PASCAL face, FDDB and WIDER FACE datasets, and can run at 36 FPS on a Nvidia Titan X (Pascal) for VGA-resolution images. Xiangyu Zhu 0001, Zhen Lei 0001, Hailin Shi, Xiaobo Wang 0001, Stan Z. Li |
ICCV | 2 |
| 2017 | Cross-Modality Face Recognition via Heterogeneous Joint BayesianabstractIn many face recognition applications, the modalities of face images between the gallery and probe sets are different, which is known as heterogeneous face recognition. How to reduce the feature gap between images from different modalities is a critical issue to develop a highly accurate face recognition algorithm. Recently, joint Bayesian (JB) has demonstrated superior performance on general face recognition compared to traditional discriminant analysis methods like subspace learning. However, the original JB treats the two input samples equally and does not take into account the modality difference between them and may be suboptimal to address the heterogeneous face recognition problem. In this work, we extend the original JB by modeling the gallery and probe images using two different Gaussian distributions to propose a heterogeneous joint Bayesian (HJB) formulation for cross-modality face recognition. The proposed HJB explicitly models the modality difference of image pairs and, therefore, is able to better discriminate the same/different face pairs accurately. Extensive experiments conducted in the case of visible-near-infrared and ID photo versus spot face recognition problems show the superiority of the HJB over previous methods. Hailin Shi, Xiaobo Wang 0001, Dong Yi, Zhen Lei 0001, Xiangyu Zhu 0001, Stan Z. Li |
IEEE Signal Process. Lett. | 5 |
| 2016 | Face Alignment Across Large Poses: A 3D SolutionabstractFace alignment, which fits a face model to an image and extracts the semantic meanings of facial pixels, has been an important topic in CV community. However, most algorithms are designed for faces in small to medium poses (below 45), lacking the ability to align faces in large poses up to 90. The challenges are three-fold: Firstly, the commonly used landmark-based face model assumes that all the landmarks are visible and is therefore not suitable for profile views. Secondly, the face appearance varies more dramatically across large poses, ranging from frontal view to profile view. Thirdly, labelling landmarks in large poses is extremely challenging since the invisible landmarks have to be guessed. In this paper, we propose a solution to the three problems in an new alignment framework, called 3D Dense Face Alignment (3DDFA), in which a dense 3D face model is fitted to the image via convolutional neutral network (CNN). We also propose a method to synthesize large-scale training samples in profile views to solve the third problem of data labelling. Experiments on the challenging AFLW database show that our approach achieves significant improvements over state-of-the-art methods. Xiangyu Zhu 0001, Zhen Lei 0001, Xiaoming Liu 0002, Hailin Shi, Stan Z. Li |
CVPR | 1 |
| 2016 | Embedding Deep Metric for Person Re-identification: A Study Against Large Variations
Hailin Shi, Yang Yang 0062, Xiangyu Zhu 0001, Shengcai Liao, Zhen Lei 0001, Wei-Shi Zheng 0001, Stan Z. Li |
ECCV (1) | 3 |
| 2015 | Person re-identification by Local Maximal Occurrence representation and metric learningabstractPerson re-identification is an important technique towards automatic search of a person's presence in a surveillance video. Two fundamental problems are critical for person re-identification, feature representation and metric learning. An effective feature representation should be robust to illumination and viewpoint changes, and a discriminant metric should be learned to match various person images. In this paper, we propose an effective feature representation called Local Maximal Occurrence (LOMO), and a subspace and metric learning method called Cross-view Quadratic Discriminant Analysis (XQDA). The LOMO feature analyzes the horizontal occurrence of local features, and maximizes the occurrence to make a stable representation against viewpoint changes. Besides, to handle illumination variations, we apply the Retinex transform and a scale invariant texture operator. To learn a discriminant metric, we propose to learn a discriminant low dimensional subspace by cross-view quadratic discriminant analysis, and simultaneously, a QDA metric is learned on the derived subspace. We also present a practical computation method for XQDA, as well as its regularization. Experiments on four challenging person re-identification databases, VIPeR, QMUL GRID, CUHK Campus, and CUHK03, show that the proposed method improves the state-of-the-art rank-1 identification rates by 2.2%, 4.88%, 28.91%, and 31.55% on the four databases, respectively. Shengcai Liao, Xiangyu Zhu 0001, Stan Z. Li |
CVPR | 3 |
| 2015 | Object detection by labeling superpixelsabstractObject detection is often conducted by object proposal generation and classification sequentially. This paper handles object detection in a superpixel oriented manner instead of the proposal oriented. Specially, this paper takes object detection as a multi-label superpixel labeling problem by minimizing an energy function. It uses the data cost term to capture the appearance, smooth cost term to encode the spatial context and label cost term to favor compact detection. The data cost is learned through a convolutional neural network and the parameters in the labeling model are learned through a structural SVM. Compared with proposal generation and classification based methods, the proposed superpixel labeling method can naturally detect objects missed by proposal generation step and capture the global image context to infer the overlapping objects. The proposed method shows its advantage in Pascal VOC and ImageNet. Notably, it performs better than the ImageNet ILSVRC2014 winner GoogLeNet (45.0% V.S. 43.9% in mAP) with much shallower and fewer CNNs. Yinan Yu, Xiangyu Zhu 0001, Zhen Lei 0001, Stan Z. Li |
CVPR | 3 |
| 2015 | High-fidelity Pose and Expression Normalization for face recognition in the wildabstractPose and expression normalization is a crucial step to recover the canonical view of faces under arbitrary conditions, so as to improve the face recognition performance. An ideal normalization method is desired to be automatic, database independent and high-fidelity, where the face appearance should be preserved with little artifact and information loss. However, most normalization methods fail to satisfy one or more of the goals. In this paper, we propose a High-fidelity Pose and Expression Normalization (HPEN) method with 3D Morphable Model (3DMM) which can automatically generate a natural face image in frontal pose and neutral expression. Specifically, we firstly make a landmark marching assumption to describe the non-correspondence between 2D and 3D landmarks caused by pose variations and propose a pose adaptive 3DMM fitting algorithm. Secondly, we mesh the whole image into a 3D object and eliminate the pose and expression variations using an identity preserving 3D transformation. Finally, we propose an inpainting method based on Possion Editing to fill the invisible region caused by self occlusion. Extensive experiments on Multi-PIE and LFW demonstrate that the proposed method significantly improves face recognition performance and outperforms state-of-the-art methods in both constrained and unconstrained environments. Xiangyu Zhu 0001, Zhen Lei 0001, Dong Yi, Stan Z. Li |
CVPR | 1 |
| 2014 | Robust 3D Morphable Model Fitting by Sparse SIFT Flowabstract3D Morph able Model (3DMM) has been widely used in face analysis for many years. The most challenging part of 3DMM is to find the correspondences between 3D points and 2D pixels. Existing methods only use key points, edges, specular highlights and image pixels to complete the task, which are not accurate or robust. This paper proposes a new algorithm called Sparse SIFT Flow (SSF) to improve the reconstruction accuracy. We mark a set of salient points to control the shape of facial components and use SSF to find their corresponding pixels on the input image. We also incorporate SSF into Multi-Features Framework to construct a robust 3DMM fitting algorithm. Compared with the state-of-the art, our approach significantly improves the fitting results in facial component area. Xiangyu Zhu 0001, Dong Yi, Zhen Lei 0001, Stan Z. Li |
ICPR | 1 |