VLDB 2026 Research / reviewers in the wild / expert
Jun Wan 0001
dblp:69/6563-1
· DBLP profile ↗
94ranked-venue papers
7as first author
64since 2021 · last 2026
0000-0002-4735-2885ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 6 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 52 · 1 first-author · 35 since 2021Security and privacy · 13 · 12 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniAttack: Unified Physical-Digital Face Attack Detection
Shunxin Chen, Ajian Liu 0001, Haocheng Yuan, Junze Zheng, Dingheng Zeng, Jiankang Deng, Sergio Escalera, Xiaoming Liu 0002, Jun Wan 0001, Zhen Lei 0001 |
Int. J. Comput. Vis. | 11 |
| 2026 | ICPE-FAS: Instance and Category Prompts Engineering for Generalizable Face Anti-Spoofing
Ajian Liu 0001, Xun Lin, Hui Ma 0018, Xinxing Yu, Jiabao Guo, Zitong Yu, Jun Wan 0001, Zhanchuan Cai, Zhen Lei 0001, Yanyan Liang 0001 |
Int. J. Comput. Vis. | 7 |
| 2026 | DGPDL: Domain-Guided Prompt Distribution Learning for Generalizable Face Anti-SpoofingabstractThe overfitting of domain signals results in poor domain generalization of face anti-spoofing. The current methods usually improve the diversity of source domains to alleviate this overfitting. However, this benefit is minimal, as even the most diverse domain signals will also be absent in the target domain. In this work, we propose a Domain-Guided Prompt Distribution Learning (DGPDL) built on Vision-Language Models like CLIP, which explores a unified representation of domain signals as a prompt across the source and target domain to alleviate the understanding bias caused by domain gaps. Specifically, we first define a learnable Domain-Specific Distribution (DSD) that covers as many domain elements as possible, such as image quality, color tone, camera settings, etc., which establish connections between different domains and linearly combinable prompt in any domain; Then, based on the style statistics of the given sample, we construct its optimal Domain-Specific Prompts (DSPs) from the defined DSD through the designed Prompt Assemble Attention (PAA) with the similarity matching; Finally, the assembled DSPs will act as carrier or agent to perform on both the vision and language branches, synergistically improving the model's recognition of domain signals. By using the prompt to represent domain signals uniformly, if the model can be robust to DSPs in the source domain, it should be applicable to target domain, as they share the same DSD. By representing domain signals as prompts rather than instantiation features, DGPDL effectively reduces the reliance on specific domain appearances. This design enables the model to dynamically adapt to unseen target domains without the need for retraining. Extensive experiments show that the DGPDL is effective and outperforms the state-of-the-art methods on several cross-domain benchmarks. Ajian Liu 0001, Xun Lin, Ruicong Zhi, Yanyan Liang 0001, Xinshan Zhu, Zhanchuan Cai, Jun Wan 0001, Sergio Escalera, Zhen Lei 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | AU-EMO Correlation based zero-shot facial expression recognition with graph convolutional network
Ruicong Zhi, Jun Wan 0001 |
Pattern Recognit. | 3 |
| 2026 | Accurate Absolute Scale With Pseudo Depth Constraint for 3-D Sparse Construction of Monocular RGB Imageabstract3D Sparse reconstruction and camera pose estimation from monocular sequences are inherently plagued by scale ambiguity, which prevents the recovery of the scene’s absolute scale. Existing solutions are primarily limited in two aspects: (1) traditional geometric methods often depend on additional sensors or restrictive pre-calibration, lacking general applicability; (2) learning-based approaches frequently suffer from scale biases when faced with domain shifts. To overcome these challenges, this paper introduces a unified framework that deeply integrates monocular depth estimation into a traditional reconstruction pipeline. Our method leverages the robustness of modern depth prediction networks to eliminate cumbersome prior assumptions while employing iterative geometric refinement to ensure precise scale accuracy. The proposed framework begins by estimating depth maps via RGB images. First, an absolute scale factor is initialized from the depth priors to calibrate the initial two-view reconstruction, effectively aligning the model into a metric space. Second, during incremental reconstruction, new points and poses are registered via depth-constrained triangulation, with an iterative weighting strategy to enhance stability. Third, we propose a joint bundle adjustment that harmonizes reprojection and geometric errors using an adaptive Mahalanobis distance, which implicitly optimizes the global scale and refines the scene geometry concurrently. Comprehensive evaluations on the ETH3D, TUM-RGBD, and ICL-NUIM benchmarks demonstrate that our framework achieves state-of-the-art performance in absolute scale recovery. The results confirm that our method successfully maintains high geometric fidelity across diverse scenarios, proving its effectiveness and robustness. Erjie Jiao, Hengyou Wang, Jun Wan 0001, Yihong Wu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Progressive Curriculum Learning With Teacher-Student Collaboration for Source-Free Unsupervised Domain AdaptationabstractIn the present environment where privacy protection is increasingly emphasized, source-free unsupervised domain adaptation (SFUDA) has garnered more attention compared to standard unsupervised domain adaptation (UDA). It concentrates on transferring knowledge directly from well-trained source models to unlabeled target domains without requiring the involvement of source domain like UDA, greatly enhancing data protection capabilities. Many existing methods employ pseudo-labeling to guide this process, but due to domain shift, pseudo-labels often introduce significant noise. Although there are methods to filter out this noise and mitigate its impact, they may also result in the loss of crucial sample knowledge, leading to performance deterioration. In contrast, we propose a novel approach called Progressive Curriculum Learning with Teacher-Student Collaboration (PCTSC) method to mitigate the adverse influence of noisy labels in SFUDA. Inspired by curriculum learning, PCTSC assesses samples’ learning difficulty and trains models in an incremental manner from easy to hard, thereby enhancing the capability of model to against noise. Furthermore, PCTSC employs a two-stage learning approach: initially, a teacher model directs the student model, and later, the student model transitions to independent learning. We assess the effectiveness of PCTSC by conducting extensive experiments across three benchmark datasets, demonstrating its robustness against pseudo-label noise in SFUDA setting. Qing Tian 0001, Junyu Shen, Lulu Kang, Weihua Ou, Jun Wan 0001, Zhen Lei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Flexible Modal Mixture-of-Experts With Inter-Modal Knowledge Distillation for Face Anti-Spoofing
Hui Ma 0018, Ajian Liu 0001, Ning Li 0035, Boyun Wang, Hang Zou 0002, Yuan Zhang 0023, Jing Huang 0017, Zhiqiang Pu, Jun Wan 0001, Zhanchuan Cai, Zhen Lei 0001, Yanyan Liang 0001 |
IEEE Trans. Inf. Forensics Secur. | 10 |
| 2026 | HySpeFAS: A Hyperspectral Face Anti-Spoofing Dataset Based on Snapshot Compressive ImagingabstractFace anti-spoofing, which aims to prevent the attacks of widely-used face recognition systems, is highly related to personal privacy and property security. However, existing benchmarks on face anti-spoofing mainly focus on RGB images, further challenged by consistently developed 3D high-fidelity (HiFi) masks. To facilitate the research on multimodal face anti-spoofing, we construct the HyperSpectral Face Anti-Spoofing (HySpeFAS) dataset. We introduce the newly-developed snapshot spectral imaging (SSI) technology to capture real and spoof faces, as well as identify unknown HiFi masks. Specifically, hyperspectral images (HSIs) acquired by SSI sensor contain rich information about the chemical composition of the targets, which can be used to effectively distinguish live human skin and various spoof materials. The HySpeFAS dataset contains 22,368 multimodal images (i.e., RGB, SSI, HSI) of 17 live subjects and 60 spoof subjects. Moreover, extensive experiments with baseline deep learning models validate the special features of the SSI images and the potential of SSI in FAS. By publishing the dataset as well as the baseline models, we encourage the community to foster the algorithm study associated with hyperspectral images and the development of SSI-equipped intelligent systems. Shijie Rao, Yidong Huang, Xueqian Zhang, Ajian Liu 0001, Jun Wan 0001, Kaiyu Cui, Yali Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Mixture-of-Attack-Experts with Class Regularization for Unified Physical-Digital Face Attack DetectionabstractUnified detection of digital and physical attacks in facial recognition systems has become a focal point of research in recent years. However, current multi-modal methods typically ignore the intra-class and inter-class variability across different types of attacks, leading to degraded performance. To address this limitation, we propose MoAE-CR, a framework that effectively leverages class-aware information for improved attack detection. Our improvements manifest at two levels, i.e., the feature and loss level. At the feature level, we propose Mixture-of-Attack-Experts (MoAEs) to capture more subtle differences among various types of fake faces. At the loss level, we introduce Class Regularization (CR) through the Disentanglement Module (DM) and the Cluster Distillation Module (CDM). The DM enhances class separability by increasing the distance between the centers of live and fake face classes. However, center-to-center constraints alone are insufficient to ensure distinctive representations for individual features. Thus, we propose the CDM to further cluster features around their class centers while maintaining separation from other classes. Moreover, specific attacks that significantly deviate from common attack patterns are often overlooked. To address this issue, our distance calculation prioritizes more distant features. Extensive experiments on two unified physical-digital attack datasets demonstrate the state-of-the-art performance of the proposed method. Shunxin Chen, Ajian Liu 0001, Junze Zheng, Jun Wan 0001, Kailai Peng, Sergio Escalera, Zhen Lei 0001 |
AAAI | 4 |
| 2025 | Recover and Match: Open-Vocabulary Multi-Label Recognition through Knowledge-Constrained Optimal TransportabstractIdentifying multiple novel classes in an image, known as open-vocabulary multi-label recognition, is a challenging task in computer vision. Recent studies explore the transfer of powerful vision-language models such as CLIP. However, these approaches face two critical challenges: (1) The local semantics of CLIP are disrupted due to its global pre-training objectives, resulting in unreliable regional predictions. (2) The matching property between image regions and candidate labels has been neglected, relying instead on naive feature aggregation such as average pooling, which leads to spurious predictions from irrelevant regions. In this paper, we present RAM (Recover And Match), a novel framework that effectively addresses the above issues. To tackle the first problem, we propose Ladder Local Adapter (LLA) to enforce refocusing on local regions, recovering local semantics in a memory-friendly way. For the second issue, we propose Knowledge-Constrained Optimal Transport (KCOT) to suppress meaningless matching to non-GT labels by formulating the task as an optimal transport problem. As a result, RAM achieves state-of-the-art performance on various datasets from three distinct domains, and shows great potential to boost the existing methods. Code: https://github.com/EricTan7/RAM. Zichang Tan, Jun Li 0033, Ajian Liu 0001, Jun Wan 0001, Zhen Lei 0001 |
CVPR | 5 |
| 2025 | PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional Styles
Tianshun Han, Benjia Zhou, Ajian Liu 0001, Yanyan Liang 0001, Zhen Lei 0001, Jun Wan 0001 |
ACM Multimedia | 7 |
| 2025 | Guest Editorial: Special Issue on Biometrics Security and Privacy
Jun Wan 0001, Arun Ross, Sergio Escalera |
Int. J. Comput. Vis. | 1 |
| 2025 | Enhance the old representations' adaptability dynamically for exemplar-free continual learning
Kunchi Li, Chaoyue Ding, Jun Wan 0001 |
Neurocomputing | 3 |
| 2025 | Reliable and Balanced Transfer Learning for Generalized Multimodal Face Anti-SpoofingabstractFace Anti-Spoofing (FAS) is essential for securing face recognition systems against presentation attacks. Recent advances in sensor technology and multimodal learning have enabled the development of multimodal FAS systems. However, existing methods often struggle to generalize to unseen attacks and diverse environments due to two key challenges: (1) Modality unreliability, where sensors such as depth and infrared suffer from severe domain shifts, impairing the reliability of cross-modal fusion; and (2) Modality imbalance, where over-reliance on a dominant modality weakens the model's robustness against attacks that affect other modalities. To overcome these issues, we propose MMDG++, a multimodal domain-generalized FAS framework built upon the vision-language model CLIP. In MMDG++, we design the Uncertainty-Guided Cross-Adapter++ (U-Adapter++) to filter out unreliable regions within each modality, enabling more reliable multimodal interactions. Additionally, we introduce Rebalanced Modality Gradient Modulation (ReGrad) for adaptive gradient modulation to balance modality convergence. To further enhance generalization, propose Asymmetric Domain Prompts (ADPs) that leverage CLIP's language priors to learn generalized decision boundaries across modalities. We also develop a novel multimodal FAS benchmark to evaluate generalizability under various deployment conditions. Extensive experiments across this benchmark show our method outperforms state-of-the-art FAS methods, demonstrating superior generalization capability. Xun Lin, Ajian Liu 0001, Zitong Yu, Rizhao Cai, Shuai Wang 0049, Yi Yu 0011, Jun Wan 0001, Zhen Lei 0001, Xiaochun Cao, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Collaborative Adapter Experts for Class-Incremental LearningabstractPre-trained models (PTMs) with parameter-efficient fine-tuning (PEFT) techniques have been extensively utilized in class-incremental learning (CIL) scenarios. However, they still remain susceptible to performance degradation as the individual PEFT module operates as an independent learning entity during the incremental process. To this end, this work proposes a novel class-incremental collaborative adapter experts (CICAE) model, which incorporates multiple adapters operating collaboratively to facilitate CIL. Specifically, our model primarily consists of two phases. Initially, multiple adapters are employed to establish a multi-expert system aimed at acquiring diverse incremental knowledge. Through the collaborative knowledge sharing (CKS) mechanism, the expertise of each adapter expert is transferable, promoting collaborative development and mutual advancement. Subsequently, with the category prototype distributions, collaborative classifier alignment (CCA) is proposed to further align the classifiers with the representation space in a cooperative manner. Extensive experiments on CIL benchmarks validate the superior performance of our model. Sunyuan Qiang, Xinxing Yu, Yanyan Liang 0001, Jun Wan 0001 |
IEEE Signal Process. Lett. | 4 |
| 2025 | C2RL: Content and Context Representation Learning for Gloss-Free Sign Language Translation and RetrievalabstractSign Language Representation Learning (SLRL) is crucial for a range of sign language-related downstream tasks such as Sign Language Translation (SLT) and Sign Language Retrieval (SLRet). Recently, many gloss-based and gloss-free SLRL methods have been proposed, showing promising performance. Among them, the gloss-free approach shows promise for strong scalability without relying on gloss annotations. However, it currently faces suboptimal solutions due to challenges in encoding the intricate, context-sensitive characteristics of sign language videos, mainly struggling to discern essential sign features using a non-monotonic video-text alignment strategy. Therefore, we introduce an innovative pretraining paradigm for gloss-free SLRL, called C2RL, in this paper. Specifically, rather than merely incorporating a non-monotonic semantic alignment of video and text to learn language-oriented sign features, we emphasize two pivotal aspects of SLRL: Implicit Content Learning (ICL) and Explicit Context Learning (ECL). ICL delves into the content of communication, capturing the nuances, emphasis, timing, and rhythm of the signs. In contrast, ECL focuses on understanding the contextual meaning of signs and converting them into equivalent sentences. Despite its simplicity, extensive experiments confirm that the joint optimization of ICL and ECL results in robust sign language representation and significant performance gains in gloss-free SLT and SLRet tasks. Notably, C2RL improves the BLEU-4 score by +5.3 on P14T, +10.6 on CSL-daily, +6.2 on OpenASL, and +1.3 on How2Sign. It also boosts the R@1 score by +8.3 on P14T, +14.4 on CSL-daily, and +5.9 on How2Sign. Additionally, we set a new baseline for the OpenASL dataset in the SLRet task. Benjia Zhou, Jun Wan 0001, Yibo Hu 0001, Hailin Shi, Yanyan Liang 0001, Zhen Lei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Cross-Attention With Conditional Matching for Multi-Target Domain AdaptationabstractAs an emerging direction of machine learning, multi-target domain adaptation (MTDA) aims to address the challenges of adapting models to multiple target domains. However, existing studies often focus on single-target domain adaptation or fail to delve into the complexities associated with multiple target domains. So there is a notable lack of comprehensive research and exploration in MTDA. Consequently, we propose a cross-attention with conditional matching for MTDA that intends to overcome the challenges posed by domain discrepancy, multi-target domain heterogeneity, and scalability. Foremost, we design a novel multi-target conditional matching that aims to align the sample distribution by leveraging nearest neighbor principle. This strategy takes into account the unique characteristics of each target domain, facilitating adaptive adaptation across multiple domains. Furthermore, we use the transformer module and well-design a cross-attention mechanism to facilitate the alignment of distributions across the source and target domains, as well as among the target domains, thus mitigating discrepancies among multiple domains. Through integrating the cross-attention mechanism into the training phase, attaining effective alignment of cross-domain distributions, we improve the adaptability and performance of the method. By the end, our approach demonstrates effective and superior experimental results indicating the significance of our work. Qing Tian 0001, Yuhui Zheng, Jun Wan 0001, Zhen Lei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | FA3-CLIP: Frequency-Aware Cues Fusion and Attack-Agnostic Prompt Learning for Unified Face Attack DetectionabstractFacial recognition systems are vulnerable to physical (e.g., printed photos) and digital (e.g., DeepFake) face attacks. Existing methods struggle to simultaneously detect physical and digital attacks due to: 1) significant intra-class variations between these attack types, and 2) the inadequacy of spatial information alone to comprehensively capture live and fake cues. To address these issues, we propose a unified attack detection model termed Frequency-Aware and Attack-Agnostic CLIP (FA3-CLIP), which introduces attack-agnostic prompt learning to express generic live and fake cues derived from the fusion of spatial and frequency features, enabling unified detection of live faces and all categories of attacks. Specifically, the attack-agnostic prompt module generates generic live and fake prompts within the language branch to extract corresponding generic representations from both live and fake faces, guiding the model to learn a unified feature space for unified attack detection. Meanwhile, the module adaptively generates the live/fake conditional bias from the original spatial and frequency information to optimize the generic prompts accordingly, reducing the impact of intra-class variations. We further propose a dual-stream cues fusion framework in the vision branch, which leverages frequency information to complement subtle cues that are difficult to capture in the spatial domain. In addition, a frequency compression block is utilized in the frequency stream, which reduces redundancy in frequency features while preserving the diversity of crucial cues. We also establish new challenging protocols to facilitate unified face attack detection effectiveness. Experimental results on multiple benchmarks demonstrate that FA3-CLIP significantly improves performance, reducing ACER by over 1.2% on UniAttackData, and increasing AUC by more than 3% as well as reducing EER by over 4% on the JFSFDB dataset. Yongze Li, Ning Li 0035, Ajian Liu 0001, Hui Ma 0018, Xihong Chen, Zhiyao Liang, Yanyan Liang 0001, Jun Wan 0001, Zhen Lei 0001 |
IEEE Trans. Inf. Forensics Secur. | 9 |
| 2025 | Unsupervised Domain Adaptation Person Re-Identification: Bridged by Feature Fusion Transitional DomainabstractThe goal of unsupervised domain adaptation person re-identification (UDA Reid) is to achieve feature space alignment between the source domain and the target domain, so that the Reid model can effectively match pedestrians in the target domain. Creating the transitional domain is an effective approach, but existing models often have difficulty synthesizing transitional domains with sufficiently public features. To tackle this challenge, we propose an innovative approach named feature fusion transitional domain (F2TD-Reid), which comprises two essential components: the dictionary fusion module (DFM) and the transitional domain attention module (TDAM). Among them, the DFM utilizes a feature fusion to extract and reconstruct pedestrian images from instances, focusing on capturing the essential visual elements within the images. For the TDAM, it further refines the feature extraction of instance points through an innovative weighted attention mechanism. These two modules optimize the generation process of scaling factors, thereby facilitating the transfer of knowledge between the source domain and the target domain. Through a series of comparative experiments, we verify the superiority of the F2TD-Reid method in solving UDA Reid. The code is available at https://github.com/1x-x/F2TD-Reid. Qing Tian 0001, Jixin Sun, Jun Wan 0001, Zhen Lei 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Toward Generalized Iris Presentation Attack Detection: A Mask-and-Distill Mixture of Experts ApproachabstractIris Presentation Attack Detection (PAD) is critical for securing recognition systems, yet its practical deployment is severely hindered by the poor generalization of models across different acquisition devices and diverse datasets. To address this persistent cross-domain challenge, we first introduce a comprehensive evaluation framework, the Iris Presentation Attack Detection Cross-Domain-Testing (IPAD-CDT) Protocol, designed to evaluate the model robustness in these scenarios. Our core contribution is a novel Masked Mixture-of-Experts (MMoE) method, which enhances the generalization of Transformer-based architectures. MMoE introduces a structured information asymmetry, where "student" Experts learn robust features from masked inputs by distilling knowledge from an unmasked "teacher" Expert via a cosine distance loss. This mask-and-distill mechanism effectively mitigates overfitting and guides the model to learn domain-invariant cues. By integrating MMoE into a CLIP-based model, we conduct extensive experiments on our IPAD-CDT protocol. The results demonstrate that our method sets a new state-of-the-art, significantly outperforming existing models, especially in the challenging cross-dataset and cross-device settings. Hang Zou 0002, Chenxi Du, Ajian Liu 0001, Yuan Zhang 0023, Jing Liu 0062, Jun Wan 0001, Hui Zhang 0061, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | PMMTalk$:$ Speech-Driven 3D Facial Animation From Complementary Pseudo Multi-Modal FeaturesabstractSpeech-driven 3D facial animation has improved a lot recently while most related works only utilize acoustic modality and neglect the influence of visual and textual cues, leading to unsatisfactory results in terms of precision and coherence. We argue that visual and textual cues are not trivial information. Therefore, we present a novel framework, namely PMMTalk, using complementaryPseudoMulti-Modal features for improving the accuracy of facial animation. The framework entails three modules: PMMTalk encoder, cross-modal alignment module, and PMMTalk decoder. Specifically, the PMMTalk encoder employs the off-the-shelf talking head generation architecture and speech recognition technology to extract visual and textual information from speech, respectively. Following this, the cross-modal alignment module aligns the audio-image-text features at temporal and semantic levels. Subsequently, the PMMTalk decoder is employed to predict lip-syncing facial blendshape coefficients. Contrary to prior methods, PMMTalk only requires an additional random reference face image but yields more accurate results. Additionally, it is artist-friendly as it seamlessly integrates into standard animation production workflows by introducing facial blendshape coefficients. Finally, given the scarcity of 3D talking face datasets, we introduce a large-scale3DChineseAudio-VisualFacialAnimation (3D-CAVFA) dataset. Extensive experiments and user studies show that our approach outperforms the state of the art. Codes and datasets are available at PMMTalk. Tianshun Han, Shengnan Gui, Baihui Li, Lijian Liu, Benjia Zhou, Ruicong Zhi, Yanyan Liang 0001, Jun Wan 0001 |
IEEE Trans. Multim. | 12 |
| 2025 | Vision Transformer With Relation Exploration for Pedestrian Attribute RecognitionabstractPedestrian attribute recognition has achieved high accuracy by exploring the relations between image regions and attributes. However, existing methods typically adopt features directly extracted from the backbone or utilize a single structure (e.g., transformer) to explore the relations, leading to inefficient and incomplete relation mining. To overcome these limitations, this paper proposes a comprehensive relationship framework called Vision Transformer with Relation Exploration (ViT-RE) for pedestrian attribute recognition, which includes two novel modules, namely Attribute and Contextual Feature Projection (ACFP) and Relation Exploration Module (REM). In ACFP, attribute-specific features and contextual-aware features are learned individually to capture discriminative information tailored for attributes and image regions, respectively. Then, REM employs Graph Convolutional Network (GCN) Blocks and Transformer Blocks to concurrently explore attribute, contextual, and attribute-contextual relations. To enable fine-grained relation mining, a Dynamic Adjacency Module (DAM) is further proposed to construct instance-wise adjacency matrix for the GCN Block. Equipped with comprehensive relation information, ViT-RE achieves promising performance on three popular benchmarks, including PETA, RAP, and PA-100 K datasets. Moreover, ViT-RE achieves the first place in theWACV 2023 UPAR Challenge. Zichang Tan, Dunfang Weng, Ajian Liu 0001, Jun Wan 0001, Zhen Lei 0001, Stan Z. Li |
IEEE Trans. Multim. | 5 |
| 2025 | CKDF-V2: Effectively Alleviating Representation Shift for Continual Learning With Small MemoryabstractIn continual learning (CL), the newly arrived data are often out-of-distribution from the previous ones, causing drastic representation shift (RS) when updating the old model on the new data, leading to catastrophic forgetting. In this work, we propose feature boosting calibration (FBC) to tackle this problem. Specifically, an expanded module is trained to learn all the classes, including the old and new classes, discovering critical features missed by the original/old model. Then, an FBC network (FBCN) is trained to exploit these missed features to calibrate the old representations. As the missed features increase the information needed for distinguishing between the old and new classes, FBCN generates the calibrated ones with more transferable features, thus alleviating the RS. Next, given the limited memory to store samples of the old/learned classes, the data are severely imbalanced between the old and new classes. To cope with this problem, we propose blockwise knowledge distillation (BWKD), which splits the softmax layer into blocks according to class frequency and then distills each block separately, resolving data imbalance effectively. Building upon the two improvements, we propose a two-stage training framework for CL, named CKDF-V2, providing an enhanced version of the cascaded knowledge distillation framework (CKDF). Furthermore, we integrate it with a task-token expansion method to develop a novel approach for CL based on the vision transformer (ViT). Extensive experiments show that both a convolutional neural network (CNN) and ViT-based CKDF-V2 obtain favorable results across multiple CL benchmarks. Kunchi Li, Hongyang Chen 0001, Jun Wan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | UBG: An Unreal BattleGround Benchmark With Object-Aware Hierarchical Proximal Policy OptimizationabstractThe deep reinforcement learning (DRL) has made significant progress in various simulation environments. However, applying DRL methods to real-world scenarios poses certain challenges due to limitations in visual fidelity, scene complexity, and task diversity within existing environments. To address limitations and explore the potential ability of DRL, we developed a 3-D open-world first-person shooter (FPS) game called Unreal BattleGround (UBG) using the unreal engine (UE). UBG provides a realistic 3-D environment with variable complexity, random scenes, diverse tasks, and multiple scene interaction methods. This benchmark involves far more complex state-action spaces than classic pseudo-3-D FPS games (e.g., ViZDoom), making it challenging for DRL to learn human-level decision sequences. Then, we propose the object-aware hierarchically proximal policy optimization (OaH-PPO) method in the UBG. It involves a two-level hierarchy, where the high-level controller is tasked with learning option control, and the low-level workers focus on mastering subtasks. To boost the learning of subtasks, we propose three modules: an object-aware module for extracting depth detection information from the environment, potential-based intrinsic reward shaping for efficient exploration, and annealing imitation learning (IL) to guide the initialization. Experimental results have demonstrated the broad applicability of the UBG and the effectiveness of the OaH-PPO. We will release the code of the UBG and OaH-PPO after publication. Longyu Niu, Baihui Li, Xingjian Fan, Jun Li 0033, Junliang Xing, Jun Wan 0001, Zhen Lei 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | Dual Balanced Class-Incremental Learning With im-Softmax and Angular RectificationabstractOwing to the superior performances, exemplar-based methods with knowledge distillation (KD) are widely applied in class incremental learning (CIL). However, it suffers from two drawbacks: 1) data imbalance between the old/learned and new classes causes the bias of the new classifier toward the head/new classes and 2) deep neural networks (DNNs) suffer from distribution drift when learning sequence tasks, which results in narrowed feature space and deficient representation of old tasks. For the first problem, we analyze the insufficiency of softmax loss when dealing with the problem of data imbalance in theory and then propose the imbalance softmax (im-softmax) loss to relieve the imbalanced data learning, where we re-scale the output logits to underfit the head/new classes. For another problem, we calibrate the feature space by incremental-adaptive angular margin (IAAM) loss. The new classes form a complete distribution in feature space yet the old are squeezed. To recover the old feature space, we first compute the included angle of normalized features and normalized anchor prototypes, and use the angle distribution to represent the class distribution, then we replenish the old distribution with the deviation from the new. Each anchor prototype is predefined as a learnable vector for a designated class. The proposed im-softmax reduces the bias in the linear classification layer. IAAM rectifies the representation learning, reduces the intra-class distance, and enlarges the inter-class margin. Finally, we seamlessly combine the im-softmax and IAAM in an end-to-end training framework, called the dual balanced class incremental learning (DBL), for further improvements. Experiments demonstrate the proposed method achieves state-of-the-art (SOTA) performances on several benchmarks, such as CIFAR10, CIFAR100, Tiny-ImageNet, and ImageNet-100. Ruicong Zhi, Yicheng Meng, Junyi Hou, Jun Wan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Compound Text-Guided Prompt Tuning via Image-Adaptive CuesabstractVision-Language Models (VLMs) such as CLIP have demonstrated remarkable generalization capabilities to downstream tasks. However, existing prompt tuning based frameworks need to parallelize learnable textual inputs for all categories, suffering from massive GPU memory consumption when there is a large number of categories in the target dataset. Moreover, previous works require to include category names within prompts, exhibiting subpar performance when dealing with ambiguous category names. To address these shortcomings, we propose Compound Text-Guided Prompt Tuning (TGP-T) that significantly reduces resource demand while achieving superior performance. We introduce text supervision to the optimization of prompts, which enables two benefits: 1) releasing the model reliance on the pre-defined category names during inference, thereby enabling more flexible prompt generation; 2) reducing the number of inputs to the text encoder, which decreases GPU memory consumption significantly. Specifically, we found that compound text supervisions, i.e., category-wise and content-wise, is highly effective, since they provide inter-class separability and capture intra-class variations, respectively. Moreover, we condition the prompt generation on visual features through a module called Bonder, which facilitates the alignment between prompts and visual features. Extensive experiments on few-shot recognition and domain generalization demonstrate that TGP-T achieves superior performance with consistently lower training costs. It reduces GPU memory usage by 93% and attains a 2.5% performance gain on 16-shot ImageNet. The code is available at https://github.com/EricTan7/TGP-T. Jun Li 0033, Yizhuang Zhou, Jun Wan 0001, Zhen Lei 0001, Xiangyu Zhang 0005 |
AAAI | 4 |
| 2024 | Factorized Learning Assisted with Large Language Model for Gloss-free Sign Language TranslationabstractPrevious Sign Language Translation (SLT) methods achieve superior performance by relying on gloss annotations. However, labeling high-quality glosses is a labor-intensive task, which limits the further development of SLT. Although some approaches work towards gloss-free SLT through jointly training the visual encoder and translation network, these efforts still suffer from poor performance and inefficient use of the powerful Large Language Model (LLM). Most seriously, we find that directly introducing LLM into SLT will lead to insufficient learning of visual representations as LLM dominates the learning curve. To address these problems, we propose Factorized Learning assisted with Large Language Model (FLa-LLM) for gloss-free SLT. Concretely, we factorize the training process into two stages. In the visual initialing stage, we employ a lightweight translation model after the visual encoder to pre-train the visual encoder. In the LLM fine-tuning stage, we freeze the acquired knowledge in the visual encoder and integrate it with a pre-trained LLM to inspire the LLM’s translation potential. This factorized training strategy proves to be highly effective as evidenced by significant improvements achieved across three SLT datasets which are all conducted under the gloss-free setting. Benjia Zhou, Jun Li 0033, Jun Wan 0001, Zhen Lei 0001 |
LREC/COLING | 4 |
| 2024 | CFPL-FAS: Class Free Prompt Learning for Generalizable Face Anti-SpoofingabstractDomain generalization (DG) based Face Anti-Spoofing (FAS) aims to improve the model's performance on unseen domains. Existing methods either rely on domain labels to align domain-invariant feature spaces, or disentangle generalizable features from the whole sample, which inevitably lead to the distortion of semantic feature structures and achieve limited generalization. In this work, we make use of large-scale VLMs like CLIP and leverage the textual feature to dynamically adjust the classifier's weights for exploring generalizable visual features. Specifically, we propose a novel Class Free Prompt Learning (CFPL) paradigm for DG FAS, which utilizes two lightweight transformers, namely Content Q-Former (CQF) and Style Q-Former (SQF), to learn the different semantic prompts conditioned on content and style features by using a set of learnable query vectors, respectively. Thus, the generalizable prompt can be learned by two improvements: (1) A Prompt-Text Matched (PTM) supervision is introduced to ensure CQF learns visual representation that is most informative of the content description. (2) A Diversified Style Prompt (DSP) technology is proposed to diversify the learning of style prompts by mixing feature statistics between instance-specific styles. Finally, the learned text features modulate visual features to generalization through the designed Prompt Modulation (PM). Extensive experiments show that the CFPL is effective and outperforms the state-of-the-art methods on several cross-domain datasets. Ajian Liu 0001, Jianwen Gan, Jun Wan 0001, Yanyan Liang 0001, Jiankang Deng, Sergio Escalera, Zhen Lei 0001 |
CVPR | 4 |
| 2024 | VL-FAS: Domain Generalization via Vision-Language Model For Face Anti-SpoofingabstractRecent approaches have demonstrated the effectiveness of Vision Transformer (ViT) with attention mechanisms for domain generalization of Face Anti-Spoofing (FAS). However, current attention algorithms highlight all the salient objects (e.g., background objects, hair, glasses), which results in the feature learned by the model containing face-irrelevant noisy information. Inspired by existing Vision-language works, we propose the VL-FAS to extract more generalized and cleaner discriminative features. Specifically, we leverage fine-grained natural language descriptions of the face region to act as a task-oriented teacher, directing the model’s attention towards the face region through top-down attention regulation. Furthermore, to enhance the domain generalization ability of the model, we propose a Sample-Level Vision-Text optimization module (SLVT). SLVT uses sample-level image-text pairs for contrastive learning, allowing the visual coder to comprehend the intrinsic semantics of each image sample, thereby reducing the dependence on domain information. Extensive experiments show that our approach significantly outperforms the state-of-the-art and improves the performance of the ViT by about twice. Ajian Liu 0001, Jun Wan 0001 |
ICASSP | 6 |
| 2024 | CPL-CLIP: Compound Prompt Learning for Flexible-Modal Face Anti-SpoofingabstractFace anti-spoofing (FAS) is pivotal in safeguarding the integrity of face recognition systems. Flexible-modal FAS utilizes multi-modal data and trains a unified model adaptable to any single-modal testing scenario. This innovation addresses the shortcomings of conventional multi-modal FAS approaches, which typically demand separate model training and deployment for each modality. However, existing flexible-modal FAS approaches activate specific network branches based on the modality of the tested sample. This not only increases the model’s parameters but also necessitates the provision of the image’s modality for testing, thereby constraining deployment flexibility. To address the issue, we present Compound Prompt Learning CLIP (CPL-CLIP), a novel method for flexible-modal FAS. This approach capitalizes on a learned textual prompt that is nearly independent of modality, thus bolstering class-based classification across arbitrary modalities. Specifically, our CPL-CLIP introduces a Dual-Branch Prompt (DBP), consisting of class and modal prompts that describe and guide classification, where each prompt is composed of learnable vectors and fixed templates. To further render the class prompt as modality-agnostic as possible, a Cosine Similarity Loss (CSL) is proposed to facilitate the maximal separation of the class prompt from the modality prompt. With only the class prompt utilized during testing, CPL-CLIP enables deployment in diverse modal testing scenarios without the necessity of the test image’s modality to be known. Extensive experiments demonstrate CPL-CLIP’s superiority over existing methods on several flexible-modal FAS benchmarks. Xiangyu Zhu 0001, Ajian Liu 0001, Xun Lin, Jun Wan 0001, Zhen Lei 0001 |
IJCB | 5 |
| 2024 | La-SoftMoE CLIP for Unified Physical-Digital Face Attack DetectionabstractFacial recognition systems are susceptible to both physical and digital attacks, posing significant security risks. Traditional approaches often treat these two attack types separately due to their distinct characteristics. Thus, when being combined attacked, almost all methods could not deal. Some studies attempt to combine the sparse data from both types of attacks into a single dataset and try to find a common feature space, which is often impractical due to the space is difficult to be found or even non-existent. To overcome these challenges, we propose a novel approach that uses the sparse model to handle sparse data, utilizing different parameter groups to process distinct regions of the sparse feature space. Specifically, we employ the Mixture of Experts (MoE) framework in our model, expert parameters are matched to tokens with varying weights during training and adaptively activated during testing. However, the traditional MoE struggles with the complex and irregular classification boundaries of this problem. Thus, we introduce a flexible self-adapting weighting mechanism, enabling the model to better fit and adapt. In this paper, we proposed La-SoftMoE CLIP, which allows for more flexible adaptation to the Unified Attack Detection (UAD) task, significantly enhancing the model’s capability to handle diversity attacks. Experiment results demonstrate that our proposed method has SOTA performance. Hang Zou 0002, Chenxi Du, Hui Zhang 0061, Yuan Zhang 0023, Ajian Liu 0001, Jun Wan 0001, Zhen Lei 0001 |
IJCB | 6 |
| 2024 | Unified Physical-Digital Face Attack Detection
Ajian Liu 0001, Haocheng Yuan, Junze Zheng, Dingheng Zeng, Jiankang Deng, Sergio Escalera, Xiaoming Liu 0002, Jun Wan 0001, Zhen Lei 0001 |
IJCAI | 10 |
| 2024 | FM-CLIP: Flexible Modal CLIP for Face Anti-SpoofingabstractIn this work, borrowing a solution from the large-scale vision-language models (VLMs) instead of directly removing modality-specific signals from visual features, we propose a novel Flexible Modal CLIP (FM-CLIP) for flexible modal FAS, that can utilize text features to dynamically adjust visual features to be modality independent. In the visual branch, considering the huge visual differences of the same attack in different modalities, which makes it difficult for classifiers to flexibly identify subtle spoofing clues in different test modalities, we propose Cross-Modal Spoofing Enhancer (CMS-Enhancer). It includes a Frequency Extractor (FE) and Cross-Modal Interactor (CMI), aiming to map different modal attacks in a shared frequency space to reduce interference from modality-specific signals and enhance spoofing clues by leveraging cross-modal learning from the shared frequency space. In the text branch, we introduce a Language-Guided Patch Alignment (LGPA) based on prompt learning, which further guides the image encoder to focus on patch-level spoofing representations through dynamic weighting by text features. Thus, our FM-CLIP can flexibly test different modal samples by identifying and enhancing modality-agnostic spoofing cues. Finally, extensive experiments show that FM-CLIP is effective and outperforms state-of-the-art methods on multiple multi-modal datasets. Ajian Liu 0001, Hui Ma 0018, Junze Zheng, Haocheng Yuan, Xiaoyuan Yu, Yanyan Liang 0001, Sergio Escalera, Jun Wan 0001, Zhen Lei 0001 |
ACM Multimedia | 8 |
| 2024 | Adapt and Refine: A Few-Shot Class-Incremental Learner via Pre-Trained Models
Sunyuan Qiang, Zhu Xiong, Yanyan Liang 0001, Jun Wan 0001 |
PRCV (1) | 4 |
| 2024 | NCL++: Nested Collaborative Learning for long-tailed visual recognition
Zichang Tan, Jun Li 0033, Jinhao Du, Jun Wan 0001, Zhen Lei 0001, Guodong Guo |
Pattern Recognit. | 4 |
| 2024 | ESDB: Expand the Shrinking Decision Boundary via One-to-Many Information Matching for Continual Learning With Small MemoryabstractRehearsal methods based on knowledge distillation (KD) have been widely used in continual learning (CL). However, given memory constraints, few exemplars contain limited variations of previously learned tasks, impeding the effectiveness of KD in retaining long-term knowledge. The decision boundaries learned by the typical KD strategy overfit the limited exemplars, leading to “shrunk boundaries" of the old classes. To tackle this problem, we propose a novel KD strategy, called One-to-Many Information Matching method (O2MIM), which generates interpolated data by mixing samples between old and new classes, disentangles the supervision information from them and assigns supervision information to them in favor of the old classes. By doing so, the supervision information from a single exemplar can be matched with multiple information from different interpolated images. Moreover, O2MIM utilizes one trainable parameter to create an adaptive KD loss, thereby facilitating a flexible matching process with the designated supervision information. Consequently, O2MIM exploits the exemplar corset more effectively, expanding the shrunk decision boundaries towards the new classes. Next, to incorporate new classes into our classification model, we apply an effective classification training strategy to train a debiased classifier. Combining it with O2MIM, we propose the method of Expanding the Shrinking Decision Boundaries (ESDB), which simultaneously transfers knowledge from the old model via O2MIM and learns new classes by the classification training strategy. Extensive experiments demonstrate that ESDB achieves state-of-the-art performance on diverse CL benchmarks. We also confirm that O2MIM can be used with various label-mixing methods to improve overall performance in CL. The code is available at: https://github.com/CSTiger77/ESDB. Kunchi Li, Hongyang Chen 0001, Jun Wan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | DCL: Dipolar Confidence Learning for Source-Free Unsupervised Domain AdaptationabstractSource-free unsupervised domain adaptation (SFUDA) aims to conduct prediction on the target domain by leveraging knowledge from the well-trained source model. Due to the absence of source data in the SFUDA setting, the existing methods mainly build the target classifier by fine-tuning the source model incorporated with empirical adaptation losses. Although these methods have achieved somewhat promising results, nearly all of them typically suffer from the closed-fitting dilemma that their models are dominantly affected by these easy-to-distinguish instances than those hard-to-distinguish ones, resulting from the absence of the labeled source data. To address aforementioned issues, we propose the Dipolar Confidence Learning (DCL) for SFUDA. Specifically, we conduct positive confidence learning on the samples with standard outputs to avoid overfitting of the model to these samples. In contrast, we perform negative confidence learning for the samples with abnormal outputs to optimize the complementary label, which forces the network to pay more attention to these confusing samples. Furthermore, to achieve more generalized domain alignment, both the confidence-based fuzzy mixup and rotation-based self-supervised learning are respectively constructed to boost the representation ability of the target model. Finally, extensive experiments are conducted to demonstrate the effectiveness and performance superiority of the proposed method. Qing Tian 0001, Heyang Sun, Shun Peng, Yuhui Zheng, Jun Wan 0001, Zhen Lei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Surveillance Face Anti-SpoofingabstractFace Anti-spoofing (FAS) is essential to secure face recognition systems from various physical attacks. However, recent research generally focuses on short-distance applications (i.e., phone unlocking) while lacking consideration of long-distance scenes (i.e., surveillance security checks). In order to promote relevant research and fill this gap in the community, we collect a large-scale Su rveillance Hi gh-Fi delity Mask (SuHiFiMask) dataset captured under 40 surveillance scenes, which has 101 subjects from different age groups with$232~3\text{D}$attacks (high-fidelity masks),$200~2\text{D}$attacks (posters, portraits, and screens), and 2 adversarial attacks. In this scene, low image resolution and noise interference are new challenges faced in surveillance FAS. Together with the SuHiFiMask dataset, we propose a Contrastive Quality-Invariance Learning (CQIL) network to alleviate the performance degradation caused by image quality from three aspects: 1) An Image Quality Variable module (IQV) is introduced to recover image information associated with discrimination by combining the super-resolution network. 2) Using generated sample pairs to simulate quality variance distributions to help contrastive learning strategies obtain robust feature representation under quality variation. 3) A Separate Quality Network (SQN) is designed to learn discriminative features independent of image quality. Finally, a large number of experiments verify the quality of the SuHiFiMask dataset and the superiority of the proposed CQIL. Ajian Liu 0001, Jun Wan 0001, Sergio Escalera, Stan Z. Li, Zhen Lei 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Unsupervised Multitarget Domain Adaptation With Dictionary-Bridged Knowledge ExploitationabstractUnsupervised domain adaptation (UDA) is an emerging learning paradigm that models on unlabeled datasets by leveraging model knowledge built on other labeled datasets, in which the statistical distributions of these datasets are usually not identical. Formally, UDA is to leverage knowledge from a labeled source domain to promote an unlabeled target domain. Although there have been a variety of methods proposed to address the UDA problem, most of them are dedicated to single-source-to-single-target domain, while the works on single-source-to-multitarget domain are relatively rare. Compared to the single-source domain with single-target domain scenario, the UDA from single-source domain to multitarget domain is more challenging since it needs to consider not only the relationships between the source and the target domains but also those among the target domains. To this end, this article proposes a kind of dictionary learning-based unsupervised multitarget domain adaptation method (DL-UMTDA). In DL-UMTDA, a common dictionary is constructed to correlate the single-source and multitarget domains, while individual dictionaries are designed to exploit the private knowledge for the target domains. Through learning the corresponding dictionary representation coefficients in the UDA process, the correlations from the source to the target domains as well as these potential relationships between the target domains can be effectively exploited. In addition, we design an alternating algorithm to solve the DL-UMTDA model with theoretical convergence guarantee. Finally, extensive experiments on benchmark (Office + Caltech) and real datasets (AgeDB, Morph, and CACD) validate the superiority of the proposed method. Qing Tian 0001, Meng Cao 0005, Jun Wan 0001, Zhen Lei 0001, Songcan Chen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Mixture Uniform Distribution Modeling and Asymmetric Mix Distillation for Class Incremental LearningabstractExemplar rehearsal-based methods with knowledge distillation (KD) have been widely used in class incremental learning (CIL) scenarios. However, they still suffer from performance degradation because of severely distribution discrepancy between training and test set caused by the limited storage memory on previous classes. In this paper, we mathematically model the data distribution and the discrepancy at the incremental stages with mixture uniform distribution (MUD). Then, we propose the asymmetric mix distillation method to uniformly minimize the error of each class from distribution discrepancy perspective. Specifically, we firstly promote mixup in CIL scenarios with the incremental mix samplers and incremental mix factor to calibrate the raw training data distribution. Next, mix distillation label augmentation is incorporated into the data distribution to inherit the knowledge information from the previous models. Based on the above augmented data distribution, our trained model effectively alleviates the performance degradation and extensive experimental results validate that our method exhibits superior performance on CIL benchmarks. Sunyuan Qiang, Jiayi Hou, Jun Wan 0001, Yanyan Liang 0001, Zhen Lei 0001 |
AAAI | 3 |
| 2023 | Gloss-free Sign Language Translation: Improving from Visual-Language PretrainingabstractSign Language Translation (SLT) is a challenging task due to its cross-domain nature, involving the translation of visual-gestural language to text. Many previous methods employ an intermediate representation, i.e., gloss sequences, to facilitate SLT, thus transforming it into a two-stage task of sign language recognition (SLR) followed by sign language translation (SLT). However, the scarcity of gloss-annotated sign language data, combined with the information bottleneck in the mid-level gloss representation, has hindered the further development of the SLT task. To address this challenge, we propose a novel Gloss-Free SLT based on Visual-Language Pretraining (GFSLT-VLP), which improves SLT by inheriting language-oriented prior knowledge from pre-trained models, without any gloss annotation assistance. Our approach involves two stages: (i) integrating Contrastive Language-Image Pre-training (CLIP) with masked self-supervised learning to create pre-tasks that bridge the semantic gap between visual and textual representations and restore masked sentences, and (ii) constructing an end-to-end architecture with an encoder-decoder-like structure that inherits the parameters of the pre-trained Visual Encoder and Text Decoder from the first stage. The seamless combination of these novel designs forms a robust sign language representation and significantly improves gloss-free sign language translation. In particular, we have achieved unprecedented improvements in terms of BLEU-4 score on the PHOENIX14T dataset (≥+5) and the CSL-Daily dataset (≥+3) compared to state-of-the-art gloss-free SLT methods. Furthermore, our approach also achieves competitive results on the PHOENIX14T dataset when compared with most of the gloss-based methods1. Benjia Zhou, Albert Clapés, Jun Wan 0001, Yanyan Liang 0001, Sergio Escalera, Zhen Lei 0001 |
ICCV | 4 |
| 2023 | Combining Self-Supervised and Supervised Learning with Noisy LabelsabstractSince convolutional neural networks (CNNs) can easily overfit noisy labels, which are ubiquitous in visual classification tasks, it has been a great challenge to train CNNs against them robustly. Various methods have been proposed for this challenge. However, none of them pay attention to the difference between representation and classifier learning of CNNs. Thus, inspired by the observation that classifier is more robust to noisy labels while representation is much more fragile, and by the recent advances of self-supervised representation learning (SSRL) technologies, we design a new method, i.e., CS3NL, to obtain representation by SSRL without labels and train the classifier directly with noisy labels. Extensive experiments are performed on both synthetic and real benchmark datasets. Results demonstrate that the proposed method can beat the state-of-the-art ones by a large margin, especially under a high noisy level. Hui Zhang 0085, Quanming Yao, Jun Wan 0001 |
ICIP | 4 |
| 2023 | Deep domain-invariant learning for facial age estimation
Zenghao Bao, Yutian Luo, Zichang Tan, Jun Wan 0001, Xibo Ma, Zhen Lei 0001 |
Neurocomputing | 4 |
| 2023 | A Unified Multimodal De- and Re-Coupling Framework for RGB-D Motion RecognitionabstractMotion recognition is a promising direction in computer vision, but the training of video classification models is much harder than images due to insufficient data and considerable parameters. To get around this, some works strive to explore multimodal cues from RGB-D data. Although improving motion recognition to some extent, these methods still face sub-optimal situations in the following aspects: (i) Data augmentation, i.e., the scale of the RGB-D datasets is still limited, and few efforts have been made to explore novel data augmentation strategies for videos; (ii) Optimization mechanism, i.e., the tightly space-time-entangled network structure brings more challenges to spatiotemporal information modeling; And (iii) cross-modal knowledge fusion, i.e., the high similarity between multimodal representations leads to insufficient late fusion. To alleviate these drawbacks, we propose to improve RGB-D-based motion recognition both from data and algorithm perspectives in this article. In more detail, firstly, we introduce a novel video data augmentation method dubbed ShuffleMix, which acts as a supplement to MixUp, to provide additional temporal regularization for motion recognition. Secondly, a Unified Multimodal De-coupling and multi-stage Re-coupling framework, termed UMDR, is proposed for video representation learning. Finally, a novel cross-modal Complement Feature Catcher (CFCer) is explored to mine potential commonalities features in multimodal information as the auxiliary fusion stream, to improve the late fusion results. The seamless combination of these novel designs forms a robust spatiotemporal representation and achieves better performance than state-of-the-art methods on four public motion datasets. Specifically, UMDR achieves unprecedented improvements of ↑ 4.5% on the Chalearn IsoGD dataset. Benjia Zhou, Pichao Wang, Jun Wan 0001, Yanyan Liang 0001, Fan Wang 0019 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Divergence-Driven Consistency Training for Semi-Supervised Facial Age EstimationabstractFacial age estimation has attracted considerable attention owing to its great potential in applications. However, it still falls short of reliable age estimation due to the lack of sufficient training data with accurate age labels. Using conventional semi-supervised methods to exploit unlabeled data appears to be a good solution, but it does not yield sufficient performance gains while significantly increasing training time. Therefore, to tackle these problems, we present a Divergence-driven Consistency Training (DCT) method for enhancing both efficiency and performance in this paper. Following the idea of pseudo-labeling and consistency regularization, we assign pseudo labels predicted by the teacher model to unlabeled samples and then train the student model on labeled and unlabeled samples based on consistency regularization. Based on this, we propose two main promotions. The first is the Efficient Sample Selection (ESS) strategy, which is based on the Divergence Score to select effective samples from massive unlabeled images to reduce the training time and improve efficiency. The second is Identity Consistency (IC) regularization as the additional loss function, which introduces a high dependency of aging traits on a person. Moreover, we propose Local Prediction (LP), which is a plug-and-play component, to capture local semantics. Extensive experiments on multiple age benchmark datasets, including CACD, Morph II, MIVIA, and Chalearn LAP 2015, indicate DCT outperforms the state-of-the-art approaches significantly. Zenghao Bao, Zichang Tan, Jun Wan 0001, Xibo Ma, Guodong Guo, Zhen Lei 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | FM-ViT: Flexible Modal Vision Transformers for Face Anti-SpoofingabstractThe availability of handy multi-modal (i.e., RGB-D) sensors has brought about a surge of face anti-spoofing research. However, the current multi-modal face presentation attack detection (PAD) has two defects: (1) The framework based on multi-modal fusion requires providing modalities consistent with the training input, which seriously limits the deployment scenario. (2) The performance of ConvNet-based model on high fidelity datasets is increasingly limited. In this work, we present a pure transformer-based framework, dubbed the Flexible Modal Vision Transformer (FM-ViT), for face anti-spoofing to flexibly target any single-modal (i.e., RGB) attack scenarios with the help of available multi-modal data. Specifically, FM-ViT retains a specific branch for each modality to capture different modal information and introduces the Cross-Modal Transformer Block (CMTB), which consists of two cascaded attentions named Multi-headed Mutual-Attention (MMA) and Fusion-Attention (MFA) to guide each modal branch to mine potential features from informative patch tokens, and to learn modality-agnostic liveness features by enriching the modal information of own CLS token, respectively. Experiments demonstrate that the single model trained based on FM-ViT can not only flexibly evaluate different modal samples, but also outperforms existing single-modal frameworks by a large margin, and approaches the multi-modal frameworks introduced with smaller FLOPs and model parameters. Ajian Liu 0001, Zichang Tan, Zitong Yu, Jun Wan 0001, Yanyan Liang 0001, Zhen Lei 0001, Stan Z. Li, Guodong Guo |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | General vs. Long-Tailed Age Estimation: An Approach to Kill Two Birds With One StoneabstractFacial age estimation has received a lot of attention for its diverse application scenarios. Most existing studies treat each sample equally and aim to reduce the average estimation error for the entire dataset, which can be summarized as General Age Estimation. However, due to the long-tailed distribution prevalent in the dataset, treating all samples equally will inevitably bias the model toward the head classes (usually the adult with a majority of samples). Driven by this, some works suggest that each class should be treated equally to improve performance in tail classes (with a minority of samples), which can be summarized as Long-tailed Age Estimation. However, Long-tailed Age Estimation usually faces a performance trade-off, i.e., achieving improvement in tail classes by sacrificing the head classes. In this paper, our goal is to design a unified framework to perform well on both tasks, killing two birds with one stone. To this end, we propose a simple, effective, and flexible training paradigm named GLAE, which is two-fold. First, we propose Feature Rearrangement (FR) and Pixel-level Auxiliary learning (PA) for better feature utilization to improve the overall age estimation performance. Second, we propose Adaptive Routing (AR) for selecting the appropriate classifier to improve performance in the tail classes while maintaining the head classes. Moreover, we introduce a new metric, named Class-wise Mean Absolute Error (CMAE), to equally evaluate the performance of all classes. Our GLAE provides a surprising improvement on Morph II, reaching the lowest MAE and CMAE of 1.14 and 1.27 years, respectively. Compared to the previous best method, MAE dropped by up to 34%, which is an unprecedented improvement, and for the first time, MAE is close to 1 year old. Extensive experiments on other age benchmark datasets, including CACD, MIVIA, and Chalearn LAP 2015, also indicate that GLAE outperforms the state-of-the-art approaches significantly. Zenghao Bao, Zichang Tan, Jun Li 0033, Jun Wan 0001, Xibo Ma, Zhen Lei 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | Nested Collaborative Learning for Long-Tailed Visual RecognitionabstractThe networks trained on the long-tailed dataset vary remarkably, despite the same training settings, which shows the great uncertainty in long-tailed learning. To alleviate the uncertainty, we propose a Nested Collaborative Learning (NCL), which tackles the problem by collaboratively learning multiple experts together. NCL consists of two core components, namely Nested Individual Learning (NIL) and Nested Balanced Online Distillation (NBOD), which focus on the individual supervised learning for each single expert and the knowledge transferring among multiple experts, respectively. To learn representations more thoroughly, both NIL and NBOD are formulated in a nested way, in which the learning is conducted on not just all categories from a full perspective but some hard categories from a partial perspective. Regarding the learning in the partial perspective, we specifically select the negative categories with high predicted scores as the hard categories by using a proposed Hard Category Mining (HCM). In the NCL, the learning from two perspectives is nested, highly related and complementary, and helps the network to capture not only global and robust features but also meticulous distinguishing ability. Moreover, self-supervision is further utilized for feature enhancement. Extensive experiments manifest the superiority of our method with outperforming the state-of-the-art whether by using a single model or an ensemble. Code is available at https://github.com/Bazinga699/NCL Jun Li 0033, Zichang Tan, Jun Wan 0001, Zhen Lei 0001, Guodong Guo |
CVPR | 3 |
| 2022 | Decoupling and Recoupling Spatiotemporal Representation for RGB-D-based Motion RecognitionabstractDecoupling spatiotemporal representation refers to decomposing the spatial and temporal features into dimension-independent factors. Although previous RGB-D-based motion recognition methods have achieved promising performance through the tightly coupled multi-modal spatiotemporal representation, they still suffer from (i) optimization difficulty under small data setting due to the tightly spatiotemporal-entangled modeling; (ii) information redundancy as it usually contains lots of marginal information that is weakly relevant to classification; and (iii) low interaction between multi-modal spatiotemporal information caused by insufficient late fusion. To alleviate these drawbacks, we propose to decouple and recouple spatiotemporal representation for RGB-D-based motion recognition. Specifically, we disentangle the task of learning spatiotemporal representation into 3 sub-tasks: (1) Learning high-quality and dimension independent features through a decoupled spatial and temporal modeling network. (2) Recoupling the decoupled representation to establish stronger space-time dependency. (3) Introducing a Cross-modal Adaptive Posterior Fusion (CAPF) mechanism to capture cross-modal spatiotemporal information from RGB-D data. Seamless combination of these novel designs forms a robust spatiotemporal representation and achieves better performance than state-of-the-art methods on four public motion datasets. Our code is available at https://github.com/damo-cv/MotionRGBD. Benjia Zhou, Pichao Wang, Jun Wan 0001, Yanyan Liang 0001, Fan Wang 0019, Zhen Lei 0001, Hao Li 0030, Rong Jin 0001 |
CVPR | 3 |
| 2022 | Disentangling Facial Pose and Appearance Information for Face Anti-spoofingabstractFace Anti-spoofing aims to determine whether the captured face from a face recognition system is real or fake. However, the facial pose and local significant spoofing traces (i.e., the boundary and reflection spot in presentation attack instruments) seriously affects the performance and stability of the current algorithms. Due to they regard the face image as an indivisible unit, and process it holistically, rarely consider excluding these liveness-irrelated factors. Unlike it, we design a Pose-Independent Face Anti-Spoofing (PIFAS) framework to disentangle face into an appearance information and a pose code to capture liveness and liveness-irrelated features, respectively. Specifically, the PIFAS consists of an Unsupervised Pose Switching (UPS) module and a Mutual Information Averaged Defense (MIAD) module, which are used to control the facial pose and suppress the local significant attack traces by averaging the local and global knowledge. Extensive experimental evaluations on multiple face anti-spoofing datasets verify that the proposed method can improve the generalization and stabilize the performance of each testing video through alleviating the interference from liveness-irrelated factors. Ajian Liu 0001, Jun Wan 0001, Yanyan Liang 0001 |
ICPR | 2 |
| 2022 | ChaLearn Looking at People: IsoGD and ConGD Large-Scale RGB-D Gesture RecognitionabstractThe ChaLearn large-scale gesture recognition challenge has run twice in two workshops in conjunction with the International Conference on Pattern Recognition (ICPR) 2016 and International Conference on Computer Vision (ICCV) 2017, attracting more than 200 teams around the world. This challenge has two tracks, focusing on isolated and continuous gesture recognition, respectively. It describes the creation of both benchmark datasets and analyzes the advances in large-scale gesture recognition based on these two datasets. In this article, we discuss the challenges of collecting large-scale ground-truth annotations of gesture recognition and provide a detailed analysis of the current methods for large-scale isolated and continuous gesture recognition. In addition to the recognition rate and mean Jaccard index (MJI) as evaluation metrics used in previous challenges, we introduce the corrected segmentation rate (CSR) metric to evaluate the performance of temporal segmentation for continuous gesture recognition. Furthermore, we propose a bidirectional long short-term memory (Bi-LSTM) method, determining video division points based on skeleton points. Experiments show that the proposed Bi-LSTM outperforms state-of-the-art methods with an absolute improvement of 8.1% (from 0.8917 to 0.9639) of CSR. Jun Wan 0001, Chi Lin 0002, Longyin Wen, Yunan Li 0001, Qiguang Miao, Sergio Escalera, Gholamreza Anbarjafari, Isabelle Guyon, Guodong Guo, Stan Z. Li |
IEEE Trans. Cybern. | 1 |
| 2022 | Contrastive Context-Aware Learning for 3D High-Fidelity Mask Face Presentation Attack DetectionabstractFace presentation attack detection (PAD) is essential to secure face recognition systems primarily from high-fidelity mask attacks. Most existing 3D mask PAD benchmarks suffer from several drawbacks: 1) a limited number of mask identities, types of sensors, and a total number of videos; 2) low-fidelity quality of facial masks. Basic deep models and remote photoplethysmography (rPPG) methods achieved acceptable performance on these benchmarks but still far from the needs of practical scenarios. To bridge the gap to real-world applications, we introduce a large-scale High-Fidelity Mask dataset, namely HiFiMask. Specifically, a total amount of 54,600 videos are recorded from 75 subjects with 225 realistic masks by 7 new kinds of sensors. Along with the dataset, we propose a novel Contrastive Context-aware Learning (CCL) framework. CCL is a new training methodology for supervised PAD tasks, which is able to learn by leveraging rich contexts accurately (e.g., subjects, mask material and lighting) among pairs of live faces and high-fidelity mask attacks. Extensive experimental evaluations on HiFiMask and three additional 3D mask datasets demonstrate the effectiveness of our method. The codes and dataset will be released soon. Ajian Liu 0001, Zitong Yu, Jun Wan 0001, Anyang Su, Zichang Tan, Sergio Escalera, Junliang Xing, Yanyan Liang 0001, Guodong Guo, Zhen Lei 0001, Stan Z. Li |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | CKDF: Cascaded Knowledge Distillation Framework for Robust Incremental LearningabstractRecently, owing to the superior performances, knowledge distillation-based (kd-based) methods with the exemplar rehearsal have been widely applied in class incremental learning (CIL). However, we discover that they suffer from the feature uncalibration problem, which is caused by directly transferring knowledge from the old model immediately to the new model when learning a new task. As the old model confuses the feature representations between the learned and new classes, the kd loss and the classification loss used in kd-based methods are heterogeneous. This is detrimental if we learn the existing knowledge from the old model directly in the way as in typical kd-based methods. To tackle this problem, the feature calibration network (FCN) is proposed, which is used to calibrate the existing knowledge to alleviate the feature representation confusion of the old model. In addition, to relieve the task-recency bias of FCN caused by the limited storage memory in CIL, we propose a novel image-feature hybrid sample rehearsal strategy to train FCN by splitting the memory budget to store the image-and-feature exemplars of the previous tasks. As feature embeddings of images have much lower-dimensions, this allows us to store more samples to train FCN. Based on these two improvements, we propose the Cascaded Knowledge Distillation Framework (CKDF) including three main stages. The first stage is used to train FCN to calibrate the existing knowledge of the old model. Then, the new model is trained simultaneously by transferring knowledge from the calibrated teacher model through the knowledge distillation strategy and learning new classes. Finally, after completing the new task learning, the feature exemplars of previous tasks are updated. Importantly, we demonstrate that the proposed CKDF is a general framework that can be applied to various kd-based methods. Experimental results show that our method achieves state-of-the-art performances on several CIL benchmarks. Kunchi Li, Jun Wan 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Cross-Batch Hard Example Mining With Pseudo Large Batch for ID vs. Spot Face RecognitionabstractIn our daily life, a large number of activities require identity verification, e.g., ePassport gates. Most of those verification systems recognize who you are by matching the ID document photo (ID face) to your live face image (spot face). The ID vs. Spot (IvS) face recognition is different from general face recognition where each dataset usually contains a small number of subjects and sufficient images for each subject. In IvS face recognition, the datasets usually contain massive class numbers (million or more) while each class only has two image samples (one ID face and one spot face), which makes it very challenging to train an effective model (e.g., excessive demand on GPU memory if conducting the classification on such massive classes, hardly capture the effective features for bisample data of each identity, etc.). To avoid the excessive demand on GPU memory, a two-stage training method is developed, where we first train the model on the dataset in general face recognition (e.g., MS-Celeb-1M) and then employ the metric learning losses (e.g., triplet and quadruplet losses) to learn the features on IvS data with million classes. To extract more effective features for IvS face recognition, we propose two novel algorithms to enhance the network by selecting harder samples for training. Firstly, a Cross-Batch Hard Example Mining (CB-HEM) is proposed to select the hard triplets from not only the current mini-batch but also past dozens of mini-batches (for convenience, we use batch to denote a mini-batch in the following), which can significantly expand the space of sample selection. Secondly, a Pseudo Large Batch (PLB) is proposed to virtually increase the batch size with a fixed GPU memory. The proposed PLB and CB-HEM can be employed simultaneously to train the network, which dramatically expands the selecting space by hundreds of times, where the very hard sample pairs especially the hard negative pairs can be selected for training to enhance the discriminative capability. Extensive comparative evaluations conducted on multiple IvS benchmarks demonstrate the effectiveness of the proposed method. Zichang Tan, Ajian Liu 0001, Jun Wan 0001, Hao Li 0030, Zhen Lei 0001, Guodong Guo, Stan Z. Li |
IEEE Trans. Image Process. | 3 |
| 2022 | Frequency Feature Pyramid Network With Global-Local Consistency Loss for Crowd-and-Vehicle Counting in Congested ScenesabstractContext prediction plays a crucial role in implementing autonomous driving applications. As one of important context-prediction tasks, crowd-and-vehicle counting is critical for achieving real-time traffic and crowd analysis, consequently facilitating decision-making processes for autonomous vehicles. However, the completion of crowd-and-vehicle counting also faces challenges, such as large-scale variations, imbalanced data distribution, and insufficient local patterns. To tackle these challenges, we put forth a novel frequency feature pyramid network (FFPNet) in this paper. Our proposed FFPNet extracts the multi-scale information by frequency feature pyramid module, which can tackle the issue of large-scale variations. Meanwhile, the frequency feature pyramid module uses different frequency branches to obtain different scale information. We also adopt the attention mechanism to strength the extraction of different scale information. Moreover, we devise a novel loss function, namely global-local consistency loss, to address the existing problems of imbalanced data distribution and insufficient local patterns. Furthermore, we conduct extensive experiments on six datasets to evaluate our proposed FFPNet. It is worth mentioning that we also construct a novel crowd-and-vehicle dataset (CROVEH), which is the only dataset that contains both crowd-and-vehicle annotations. The experimental results show that FFPNet achieves the best performance on different backbones, e.g., 52.69 mean absolute error (MAE) on P2PNet with FFP module. The codes are available at:https://github.com/MUST-AI-Lab/FFPNet. Xiaoyuan Yu, Yanyan Liang 0001, Xuxin Lin, Jun Wan 0001, Tian Wang 0001, Hongning Dai |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Regional Attention with Architecture-Rebuilt 3D Network for RGB-D Gesture RecognitionabstractHuman gesture recognition has drawn much attention in the area of computer vision. However, the performance of gesture recognition is always influenced by some gesture-irrelevant factors like the background and the clothes of performers. Therefore, focusing on the regions of hand/arm is important to the gesture recognition. Meanwhile, a more adaptive architecture-searched network structure can also perform better than the block-fixed ones like ResNet since it increases the diversity of features in different stages of the network better. In this paper, we propose a regional attention with architecture-rebuilt 3D network (RAAR3DNet) for gesture recognition. We replace the fixed Inception modules with the automatically rebuilt structure through the network via Neural Architecture Search (NAS), owing to the different shape and representation ability of features in the early, middle, and late stage of the network. It enables the network to capture different levels of feature representations at different layers more adaptively. Meanwhile, we also design a stackable regional attention module called Dynamic-Static Attention (DSA), which derives a Gaussian guidance heatmap and dynamic motion map to highlight the hand/arm regions and the motion information in the spatial and temporal domains, respectively. Extensive experiments on two recent large-scale RGB-D gesture datasets validate the effectiveness of the proposed method and show it outperforms state-of-the-art methods. The codes of our method are available at: https://github.com/zhoubenjia/RAAR3DNet. Benjia Zhou, Yunan Li 0001, Jun Wan 0001 |
AAAI | 3 |
| 2021 | LAE : Long-Tailed Age Estimation
Zenghao Bao, Zichang Tan, Yu Zhu 0006, Jun Wan 0001, Xibo Ma, Zhen Lei 0001, Guodong Guo |
CAIP (2) | 4 |
| 2021 | CASIA-SURF CeFA: A Benchmark for Multi-modal Cross-ethnicity Face Anti-spoofingabstractThe issue of ethnic bias has proven to affect the performance of face recognition in previous works, while it still remains to be vacant in face anti-spoofing. Therefore, in order to study the ethnic bias for face anti-spoofing, we introduce the largest CASIA-SURF Cross-ethnicity Face Anti-spoofing (CeFA) dataset, covering 3 ethnicities, 3 modalities, 1,607 subjects, and 2D plus 3D attack types. Five protocols are introduced to measure the affect under varied evaluation conditions, such as cross-ethnicity, unknown spoofs or both of them. As our knowledge, CASIA-SURF CeFA is the first dataset including explicit ethnic labels in current released datasets. Then, we propose a novel multi-modal fusion method as a strong baseline to alleviate the ethnic bias, which employs a partially shared fusion strategy to learn complementary information from multiple modalities. Extensive experiments have been conducted on the proposed dataset to verify its significance and generalization capability for other existing datasets, i.e., CASIA-SURF, OULU-NPU and SiW datasets. The dataset is available at https://sites.google.com/qq.com/face-anti-spoofing/welcome/challengecvpr2020?authuser=0. Ajian Liu 0001, Zichang Tan, Jun Wan 0001, Sergio Escalera, Guodong Guo, Stan Z. Li |
WACV | 3 |
| 2021 | Cascaded Split-and-Aggregate Learning with Feature Recombination for Pedestrian Attribute Recognition
Yang Yang 0062, Zichang Tan, Prayag Tiwari, Hari Mohan Pandey, Jun Wan 0001, Zhen Lei 0001, Guodong Guo, Stan Z. Li |
Int. J. Comput. Vis. | 5 |
| 2021 | NAS-FAS: Static-Dynamic Central Difference Network Search for Face Anti-SpoofingabstractFace anti-spoofing (FAS) plays a vital role in securing face recognition systems. Existing methods heavily rely on the expert-designed networks, which may lead to a sub-optimal solution for FAS task. Here we propose the first FAS method based on neural architecture search (NAS), called NAS-FAS, to discover the well-suited task-aware networks. Unlike previous NAS works mainly focus on developing efficient search strategies in generic object classification, we pay more attention to study the search spaces for FAS task. The challenges of utilizing NAS for FAS are in two folds: the networks searched on 1) a specific acquisition condition might perform poorly in unseen conditions, and 2) particular spoofing attacks might generalize badly for unseen attacks. To overcome these two issues, we develop a novel search space consisting of central difference convolution and pooling operators. Moreover, an efficient static-dynamic representation is exploited for fully mining the FAS-aware spatio-temporal discrepancy. Besides, we propose Domain/Type-aware Meta-NAS, which leverages cross-domain/type knowledge for robust searching. Finally, in order to evaluate the NAS transferability for cross datasets and unknown attack types, we release a large-scale 3D mask dataset, namely CASIA-SURF 3DMask, for supporting the new 'cross-dataset cross-type' testing protocol. Experiments demonstrate that the proposed NAS-FAS achieves state-of-the-art performance on nine FAS benchmark datasets with four testing protocols. Zitong Yu, Jun Wan 0001, Yunxiao Qin, Stan Z. Li, Guoying Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Face Anti-Spoofing via Adversarial Cross-Modality TranslationabstractFace Presentation Attack Detection (PAD) approaches based on multi-modal data have been attracted increasingly by the research community. However, they require multi-modal face data consistently involved in both the training and testing phases. It would severely limit the applicability due to the most Face Anti-spoofing (FAS) systems are only equipped with Visible (VIS) imaging devices, i.e., RGB cameras. Therefore, how to use other modality (i.e., Near-Infrared (NIR)) to assist the performance improvement of VIS-based PAD is significant for FAS. In this work, we first discuss the big gap of performances among different modalities even though the same backbone network is applied. Then, we propose a novel Cross-modal Auxiliary (CMA) framework for the VIS-based FAS task. The main trait of CMA is that the performance can be greatly improved with the help of other modality while no other modality is required in the testing stage. The proposed CMA consists of a Modality Translation Network (MT-Net) and a Modality Assistance Network (MA-Net). The former aims to close the visible gap between different modalities via a generative model that maps inputs from one modality (i.e., RGB) to another (i.e., NIR). The latter focuses on how to use the translated modality (i.e., target modality) and RGB modality (i.e., source modality) together to train a discriminative PAD model. Extensive experiments are conducted to demonstrate that the proposed framework can push the state-of-the-art (SOTA) performances on both multi-modal datasets (i.e., CASIA-SURF, CeFA, and WMCA) and RGB-based datasets (i.e., OULU-NPU, and SiW). Ajian Liu 0001, Zichang Tan, Jun Wan 0001, Yanyan Liang 0001, Zhen Lei 0001, Guodong Guo, Stan Z. Li |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Searching Multi-Rate and Multi-Modal Temporal Enhanced Networks for Gesture RecognitionabstractGesture recognition has attracted considerable attention owing to its great potential in applications. Although the great progress has been made recently in multi-modal learning methods, existing methods still lack effective integration to fully explore synergies among spatio-temporal modalities effectively for gesture recognition. The problems are partially due to the fact that the existing manually designed network architectures have low efficiency in the joint learning of multi-modalities. In this paper, we propose the first neural architecture search (NAS)-based method for RGB-D gesture recognition. The proposed method includes two key components: 1) enhanced temporal representation via the proposed 3D Central Difference Convolution (3D-CDC) family, which is able to capture rich temporal context via aggregating temporal difference information; and 2) optimized backbones for multi-sampling-rate branches and lateral connections among varied modalities. The resultant multi-modal multi-rate network provides a new perspective to understand the relationship between RGB and depth modalities and their temporal dynamics. Comprehensive experiments are performed on three benchmark datasets (IsoGD, NvGesture, and EgoGesture), demonstrating the state-of-the-art performance in both single- and multi-modality settings. The code is available at https://github.com/ZitongYu/3DCDC-NAS. Zitong Yu, Benjia Zhou, Jun Wan 0001, Pichao Wang, Haoyu Chen 0001, Xin Liu 0012, Stan Z. Li, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Shared Low-Rank Correlation Embedding for Multiple Feature FusionabstractThe diversity of multimedia data in the real world usually forms heterogeneous types of feature sets. How to explore the structure information and the relationships among multiple features is still an open problem. In this paper, we propose an unsupervised subspace learning method, named the shared low-rank correlation embedding (SLRCE) for multiple feature fusion. First, in the learned subspace, we implement the low-rank representation on each feature set and enforce a shared low-rank constraint to uncover the common structure information of multiple features. Second, we develop an enhanced correlation analysis in the learned subspace for simultaneously removing the redundancy of each feature set and exploring the correlation of multiple features. Finally, we incorporate the shared low-rank representation and the correlation analysis into a unified framework. The shared low-rank constraint not only depicts the data distribution consistency among multiple features, but also assists robust subspace learning. Our method is robust to noise in practice and can be extended to the kernel case to handle the nonlinear feature fusion. Experimental results on several typical datasets demonstrate the superior performance of the proposed methods. Zhan Wang 0007, Lizhi Wang 0001, Jun Wan 0001, Hua Huang 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | Adaptive cross-fusion learning for multi-modal gesture recognitionabstractGesture recognition has attracted significant attention because of its wide range of potential applications. Although multi-modal gesture recognition has made significant progress in recent years, a popular method still is simply fusing prediction scores at the end of each branch, which often ignores complementary features among different modalities in the early stage and does not fuse the complementary features into a more discriminative feature. This paper proposes an Adaptive Cross-modal Weighting (ACmW) scheme to exploit complementarity features from RGB-D data in this study. The scheme learns relations among different modalities by combining the features of different data streams. The proposed ACmW module contains two key functions: (1) fusing complementary features from multiple streams through an adaptive one-dimensional convolution; and (2) modeling the correlation of multi-stream complementary features in the time dimension. Through the effective combination of these two functional modules, the proposed ACmW can automatically analyze the relationship between the complementary features from different streams, and can fuse them in the spatial and temporal dimensions. Extensive experiments validate the effectiveness of the proposed method, and show that our method outperforms state-of-the-art methods on IsoGD and NVGesture. Benjia Zhou, Jun Wan 0001, Yanyan Liang 0001, Guodong Guo |
Virtual Real. Intell. Hardw. | 2 |
| 2020 | Relation-Aware Pedestrian Attribute Recognition with Graph Convolutional NetworksabstractIn this paper, we propose a new end-to-end network, named Joint Learning of Attribute and Contextual relations (JLAC), to solve the task of pedestrian attribute recognition. It includes two novel modules: Attribute Relation Module (ARM) and Contextual Relation Module (CRM). For ARM, we construct an attribute graph with attribute-specific features which are learned by the constrained losses, and further use Graph Convolutional Network (GCN) to explore the correlations among multiple attributes. For CRM, we first propose a graph projection scheme to project the 2-D feature map into a set of nodes from different image regions, and then employ GCN to explore the contextual relations among those regions. Since the relation information in the above two modules is correlated and complementary, we incorporate them into a unified framework to learn both together. Experiments on three benchmarks, including PA-100K, RAP, PETA attribute datasets, demonstrate the effectiveness of the proposed JLAC. Zichang Tan, Yang Yang 0062, Jun Wan 0001, Guodong Guo, Stan Z. Li |
AAAI | 3 |
| 2020 | 3DPC-Net: 3D Point Cloud Network for Face Anti-spoofingabstractFace anti-spoofing plays a vital role in face recognition systems. Most deep learning-based methods directly use 2D images assisted with temporal information (i.e., motion, rPPG) or pseudo-3D information (i.e., Depth). The main drawback of the mentioned methods is that another extra network is needed to generate the depth/rPPG information to assist the backbone network for face anti-spoofing. Different from these methods, we propose a novel method named 3D Point Cloud Network (3DPC-Net). It is an encoder-decoder network that can predict the 3DPC maps to discriminate live faces from spoofing ones. The main traits of the proposed method are that: 1) It is the first time that 3DPC is used for face anti-spoofing; 2) 3DPC-Net is simple and effective and it only relies on 3DPC supervision. Extensive experiments on four databases (i.e., Oulu-NPU, SiW, CASIA-FASD, Replay Attack) have demonstrated that the 3DPC-Net is comparative to the state-of-the-art methods. Jun Wan 0001, Yi Jin 0001, Ajian Liu 0001, Guodong Guo, Stan Z. Li |
IJCB | 2 |
| 2020 | CR-Net: A Deep Classification-Regression Network for Multimodal Apparent Personality Analysis
Yunan Li 0001, Jun Wan 0001, Qiguang Miao, Sergio Escalera, Huijuan Fang, Huizhou Chen, Xiangda Qi, Guodong Guo |
Int. J. Comput. Vis. | 2 |
| 2020 | Guest Editorial: Image and Video Inpainting and DenoisingabstractThe papers in this special issue comprise all aspects of computer vision and pattern recognition devoted to image and video inpainting, including related tasks like denoising, debluring, sampling, super-resolutkon enhancement, restoration, hallucination, etc. The special issue was associated to the 2018 Chalearn Looking at People Satellite ECCV Workshop1 and the 2018 ChaLearn Challenges on Image and Video Inpainting. Sergio Escalera, Hugo Jair Escalante, Xavier Baró, Isabelle Guyon, Meysam Madadi, Jun Wan 0001, Stéphane Ayache, Yagmur Güçlütürk, Umut Güçlü |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2020 | Task-Oriented Feature-Fused Network With Multivariate Dataset for Joint Face AnalysisabstractDeep multitask learning for face analysis has received increasing attentions. From literature, most existing methods focus on optimizing a main task by jointly learning several auxiliary tasks. It is challenging to consider the performance of each task in a multitask framework due to the following reasons: 1) different face tasks usually rely on different levels of semantic features; 2) each task has different learning convergence rate, which could affect the whole performance when joint training; and 3) multitask model needs rich label information for efficient training, but existing facial datasets provide limited annotations. To address these issues, we propose a task-oriented feature-fused network (TFN) for simultaneously solving face detection, landmark localization, and attribute analysis. In this network, a task-oriented feature-fused block is designed to learn task-specific feature combinations; then, an alternative multitask training scheme is presented to optimize each task with considering of their different learning capacities. We also present a large-scale face dataset called JFA in support of proposed method, which provides multivariate labels, including face bounding box, 68 facial landmarks, and 3 attribute labels (i.e., apparent age, gender, and ethnicity). The experimental results suggest that the TFN outperforms several multitask models on the JFA dataset. Furthermore, our approach achieves competitive performances on WIDER FACE and 300W dataset, and obtains state-of-the-art results for gender recognition on the MORPH II dataset. Xuxin Lin, Jun Wan 0001, Yiliang Xie, Chi Lin 0002, Yanyan Liang 0001, Guodong Guo, Stan Z. Li |
IEEE Trans. Cybern. | 2 |
| 2020 | WiderPerson: A Diverse Dataset for Dense Pedestrian Detection in the WildabstractPedestrian detection has achieved significant progress with the availability of existing benchmark datasets. However, there is a gap in the diversity and density between real world requirements and current pedestrian detection benchmarks: first, most existing datasets are taken from a vehicle driving through the regular traffic scenario, usually leading to insufficient diversity; second, crowd scenarios with highly occluded pedestrians are still underrepresented, resulting in low density. To narrow this gap and facilitate future pedestrian detection research, we introduce a large and diverse dataset named WiderPerson for dense pedestrian detection in the wild. This dataset involves five types of annotations in a wide range of scenarios, no longer limited to the traffic scenario. There are a total of 13 382 images with 399 786 annotations, that is, 29.87 annotations per image, which means this dataset contains dense pedestrians with various kinds of occlusions. Hence, pedestrians in the proposed dataset are extremely challenging due to large variations in the scenario and occlusion, which is suitable to evaluate pedestrian detectors in the wild. We introduce an improved Faster R-CNN and the vanilla RetinaNet to serve as baselines for the new pedestrian detection benchmark. Several experiments are conducted on previous datasets including Caltech-USA and CityPersons to analyze the generalization capabilities of the proposed dataset, and we achieve state-of-the-art performances on these previous datasets without bells and whistles. Finally, we analyze common failure cases and find the classification ability of pedestrian detector needs to be improved to reduce false alarm and misdetection rates. The proposed dataset is available at http://www.cbsr.ia.ac.cn/users/sfzhang/WiderPerson. Yiliang Xie, Jun Wan 0001, Hansheng Xia, Stan Z. Li, Guodong Guo |
IEEE Trans. Multim. | 3 |
| 2019 | A Dataset and Benchmark for Large-Scale Multi-Modal Face Anti-SpoofingabstractFace anti-spoofing is essential to prevent face recognition systems from a security breach. Much of the progresses have been made by the availability of face anti-spoofing benchmark datasets in recent years. However, existing face anti-spoofing benchmarks have limited number of subjects (≤170) and modalities (≤2), which hinder the further development of the academic community. To facilitate face anti-spoofing research, we introduce a large-scale multi-modal dataset, namely CASIA-SURF, which is the largest publicly available dataset for face anti-spoofing in terms of both subjects and visual modalities. Specifically, it consists of 1,000 subjects with 21,000 videos and each sample has 3 modalities (i.e., RGB, Depth and IR). We also provide a measurement set, evaluation protocol and training/validation/testing subsets, developing a new benchmark for face anti-spoofing. Moreover, we present a new multi-modal fusion method as baseline, which performs feature re-weighting to select the more informative channel features while suppressing the less useful ones for each modal. Extensive experiments have been conducted on the proposed dataset to verify its significance and generalization capability. The dataset is available at https://sites.google.com/qq.com/chalearnfacespoofingattackdete/. Xiaobo Wang 0001, Ajian Liu 0001, Jun Wan 0001, Sergio Escalera, Hailin Shi, Stan Z. Li |
CVPR | 5 |
| 2019 | Deeply-learned Hybrid Representations for Facial Age EstimationabstractIn this paper, we propose a novel unified network named Deep Hybrid-Aligned Architecture for facial age estimation. It contains global, local and global-local branches. They are jointly optimized and thus can capture multiple types of features with complementary information. In each branch, we employ a separate loss for each sub-network to extract the independent features and use a recurrent fusion to explore correlations among those region features. Considering that the pose variations may lead to misalignment in different regions, we design an Aligned Region Pooling operation to generate aligned region features. Moreover, a new large age dataset named Web-FaceAge owning more than 120K samples is collected under diverse scenes and spanning a large age range. Experiments on five age benchmark datasets, including Web-FaceAge, Morph, FG-NET, CACD and Chalearn LAP 2015, show that the proposed method outperforms the state-of-the-art approaches significantly. Zichang Tan, Yang Yang 0062, Jun Wan 0001, Guodong Guo, Stan Z. Li |
IJCAI | 3 |
| 2019 | Guest editorial: special issue on human abnormal behavioural analysis
Gholamreza Anbarjafari, Sergio Escalera, Kamal Nasrollahi, Hugo Jair Escalante, Xavier Baró, Jun Wan 0001, Thomas B. Moeslund |
Mach. Vis. Appl. | 6 |
| 2019 | FERLrTc: 2D+3D facial expression recognition via low-rank tensor completion
Yunfang Fu, Qiuqi Ruan, Ziyan Luo, Yi Jin 0001, Gaoyun An, Jun Wan 0001 |
Signal Process. | 6 |
| 2019 | Attention-Based Pedestrian Attribute AnalysisabstractRecognizing the pedestrian attributes in surveillance scenes is an inherently challenging task, especially for the pedestrian images with large pose variations, complex backgrounds, and various camera viewing angles. To select important and discriminative regions or pixels against the variations, three attention mechanisms are proposed, including parsing attention, label attention, and spatial attention. Those attentions aim at accessing effective information by considering problems from different perspectives. To be specific, the parsing attention extracts discriminative features by learning not only where to turn attention to but also how to aggregate features from different semantic regions of human bodies, e.g., head and upper body. The label attention aims at targetedly collecting the discriminative features for each attribute. Different from the parsing and label attention mechanisms, the spatial attention considers the problem from a global perspective, aiming at selecting several important and discriminative image regions or pixels for all attributes. Then, we propose a joint learning framework formulated in a multi-task-like way with these three attention mechanisms learned concurrently to extract complementary and correlated features. This joint learning framework is named Joint Learning of Parsing attention, Label attention, and Spatial attention for Pedestrian Attributes Analysis (JLPLS-PAA, for short). Extensive comparative evaluations conducted on multiple large-scale benchmarks, including PA-100K, RAP, PETA, Market-1501, and Duke attribute datasets, further demonstrate the effectiveness of the proposed JLPLS-PAA framework for pedestrian attribute analysis. Zichang Tan, Yang Yang 0062, Jun Wan 0001, Hanyuan Hang, Guodong Guo, Stan Z. Li |
IEEE Trans. Image Process. | 3 |
| 2019 | Region-Based Context Enhanced Network for Robust Multiple Face AlignmentabstractThe recent studies for face alignment have involved developing an isolated algorithm on well-cropped face images. It is difficult to obtain the expected input by using an off-the-shelf face detector in practical applications. In this paper, we attempt to bridge between face detection and face alignment by establishing a novel joint multi-task model, which allows us to simultaneously detect multiple faces and their landmarks on a given scene image. In contrast to the pipeline-based framework by cascading separate models, we aim to propose an end-to-end convolutional network by sharing and transform feature representations between the task-specific modules. To learn a robust landmark estimator for unconstrained face alignment, three types of context enhanced blocks are designed to encode feature maps with multi-level context, multi-scale context, and global context. In the post-processing step, we develop a shape reconstruction algorithm based on point distribution model to refine the landmark outliers. Extensive experiments demonstrate that our results are robust for the landmark location task and insensitive to the location of estimated face regions. Furthermore, our method significantly outperforms recent state-of-the-art methods on several challenging datasets including 300 W, AFLW, and COFW. Xuxin Lin, Yanyan Liang 0001, Jun Wan 0001, Chi Lin 0002, Stan Z. Li |
IEEE Trans. Multim. | 3 |
| 2018 | Cooperative Training of Deep Aggregation Networks for RGB-D Action RecognitionabstractA novel deep neural network training paradigm that exploits the conjoint information in multiple heterogeneous sources is proposed. Specifically, in a RGB-D based action recognition task, it cooperatively trains a single convolutional neural network (named c-ConvNet) on both RGB visual features and depth features, and deeply aggregates the two kinds of features for action recognition. Differently from the conventional ConvNet that learns the deep separable features for homogeneous modality-based classification with only one softmax loss function, the c-ConvNet enhances the discriminative power of the deeply learned features and weakens the undesired modality discrepancy by jointly optimizing a ranking loss and a softmax loss for both homogeneous and heterogeneous modalities. The ranking loss consists of intra-modality and cross-modality triplet losses, and it reduces both the intra-modality and cross-modality feature variations. Furthermore, the correlations between RGB and depth data are embedded in the c-ConvNet, and can be retrieved by either of the modalities and contribute to the recognition in the case even only one of the modalities is available. The proposed method was extensively evaluated on two large RGB-D action recognition datasets, ChaLearn LAP IsoGD and NTU RGB+D datasets, and one small dataset, SYSU 3D HOI, and achieved state-of-the-art results. Pichao Wang, Wanqing Li 0001, Jun Wan 0001, Philip Ogunbona, Xinwang Liu 0002 |
AAAI | 3 |
| 2018 | Large-Scale Isolated Gesture Recognition Using a Refined Fused Model Based on Masked Res-C3D Network and Skeleton LSTMabstractIn this paper, we focus on large-scale isolated gesture recognition for RGB-D videos. We develop a novel ensemble method to explore deep spatio-temporal features using 3D Convolutional Neural Networks (CNNs) with residual architecture (Res-C3D) and build a time-series model with skeleton information based on Long Short Term Memory network (LSTM). First, relative positions and angles of different keypoints are extracted and used to build time-series model in LSTM. Obtaining the skeleton information (keypoints) of body and reserving arm regions with discarding other parts, masked Res-C3D is obtained, which decreases the effect of the background and other variations, as gestures are mainly derived from the arm or hand movements. Moreover, the weights of each voting sub-classifier being of advantage to a certain class in our ensemble model are adaptively obtained by training in place of fixed weights. Our experimental results show that the proposed method has obtained a state-of-the-art performance with accuracy 0.6842 in the IsoGD dataset. Chi Lin 0002, Jun Wan 0001, Yanyan Liang 0001, Stan Z. Li |
FG | 2 |
| 2018 | RGB-D-based human motion recognition with deep learning: A survey
Pichao Wang, Wanqing Li 0001, Philip Ogunbona, Jun Wan 0001, Sergio Escalera |
Comput. Vis. Image Underst. | 4 |
| 2018 | Efficient Group-n Encoding and Decoding for Facial Age EstimationabstractDifferent ages are closely related especially among the adjacent ages because aging is a slow and extremely non-stationary process with much randomness. To explore the relationship between the real age and its adjacent ages, an age group-n encoding (AGEn) method is proposed in this paper. In our model, adjacent ages are grouped into the same group and each age corresponds to n groups. The ages grouped into the same group would be regarded as an independent class in the training stage. On this basis, the original age estimation problem can be transformed into a series of binary classification sub-problems. And a deep Convolutional Neural Networks (CNN) with multiple classifiers is designed to cope with such sub-problems. Later, a Local Age Decoding (LAD) strategy is further presented to accelerate the prediction process, which locally decodes the estimated age value from ordinal classifiers. Besides, to alleviate the imbalance data learning problem of each classifier, a penalty factor is inserted into the unified objective function to favor the minority class. To compare with state-of-the-art methods, we evaluate the proposed method on FG-NET, MORPH II, CACD and Chalearn LAP 2015 databases and it achieves the best performance. Zichang Tan, Jun Wan 0001, Zhen Lei 0001, Ruicong Zhi, Guodong Guo, Stan Z. Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Articulated motion and deformable objects
Jun Wan 0001, Sergio Escalera, Francisco José Perales López, Josef Kittler |
Pattern Recognit. | 1 |
| 2018 | Auxiliary Demographic Information Assisted Age Estimation With Cascaded StructureabstractOwing to the variations including both intrinsic and extrinsic factors, age estimation remains a challenging problem. In this paper, five cascaded structure frameworks are proposed for age estimation based on convolutional neural networks. All frameworks are learned and guided by auxiliary demographic information, since other demographic information (i.e., gender and race) is beneficial for age prediction. Each cascaded structure framework is embodied in a parent network and several subnetworks. For example, one of the applied framework is a gender classifier trained by gender information, and then two subnetworks are trained by the male and female samples, respectively. Furthermore, we use the features extracted from the cascaded structure frameworks with Gaussian process regression that can boost the performance further for age estimation. Experimental results on the MORPH II and CACD datasets have gained superior performances compared to the state-of-the-art methods. The mean absolute error is significantly reduced from 3.63 to 2.93 years under the same test protocol on the MORPH II dataset. Jun Wan 0001, Zichang Tan, Zhen Lei 0001, Guodong Guo, Stan Z. Li |
IEEE Trans. Cybern. | 1 |
| 2018 | Label Distribution-Based Facial Attractiveness Computation by Deep Residual LearningabstractTwo key challenges lie in the facial attractiveness computation research: the lack of discriminative face representations, and the scarcity of sufficient and complete training data. Motivated by recent promising work in face recognition using deep neural networks to learn effective features, the first challenge is expected to be addressed from a deep learning point of view. A very deep residual network is utilized to enable automatic learning of hierarchical aesthetics representation. The inspiration to deal with the second challenge comes from the natural representation of the training data, where each training face can be associated with a label (score) distribution given by human raters rather than a single label (average score). This paper, therefore, recasts facial attractiveness computation as a label distribution learning problem. Integrating these two ideas, an end-to-end attractiveness learning framework is established. We also perform feature-level fusion by incorporating the low-level geometric features to further improve the computational performance. Extensive experiments are conducted on a standard benchmark, the SCUT-FBP dataset, where our approach shows significant advantages over the other state-of-the-art work. Yangyu Fan, Shu Liu 0002, Bo Li 0090, Ashok Samal, Jun Wan 0001, Stan Z. Li |
IEEE Trans. Multim. | 6 |
| 2018 | A Unified Framework for Multi-Modal Isolated Gesture RecognitionabstractIn this article, we focus on isolated gesture recognition and explore different modalities by involving RGB stream, depth stream, and saliency stream for inspection. Our goal is to push the boundary of this realm even further by proposing a unified framework that exploits the advantages of multi-modality fusion. Specifically, a spatial-temporal network architecture based on consensus-voting has been proposed to explicitly model the long-term structure of the video sequence and to reduce estimation variance when confronted with comprehensive inter-class variations. In addition, a three-dimensional depth-saliency convolutional network is aggregated in parallel to capture subtle motion characteristics. Extensive experiments are done to analyze the performance of each component and our proposed approach achieves the best results on two public benchmarks, ChaLearn IsoGD and RGBD-HuDaAct, outperforming the closest competitor by a margin of over 10% and 15%, respectively. Our project and codes will be released at https://davidsonic.github.io/index/acm_tomm_2017.html. Jiali Duan, Jun Wan 0001, Xiaoyuan Guo, Stan Z. Li |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2017 | Multi-Region Ensemble Convolutional Neural Network for High Accuracy Age Estimation
Yiliang Chen, Zichang Tan, Alex Po Leung, Jun Wan 0001 |
BMVC | 4 |
| 2017 | Multi-modality Network with Visual and Geometrical Information for Micro Emotion RecognitionabstractMicro emotion recognition is a very challenging problem because of the subtle appearance variants among different facial expression classes. To deal with the mentioned problem, we proposed a multi-modality convolutional neural networks (CNNs) based on visual and geometrical information in this paper. The visual face image and structured geometry are embedded into a unified network and the recognition accuracy can be benefic from the fused information. The proposed network includes two branches. The first branch is used to extract visual feature from color face images, and another branch is used to extract the geometry feature from 68 facial landmarks. Then, both visual and geometry features are concatenated into a long vector. Finally, the concatenated vector is fed to the hinge loss layer. Compared with the CNN architecture only used face images, our method is more effective and has got better performance. In the final testing phase of Micro Emotion Challenge1, our method has got the first place with the misclassification of 80.212137. Jianzhu Guo, Jinlin Wu, Jun Wan 0001, Xiangyu Zhu 0001, Zhen Lei 0001, Stan Z. Li |
FG | 4 |
| 2017 | Principal motion components for one-shot gesture recognition
Hugo Jair Escalante, Isabelle Guyon, Vassilis Athitsos, Pat Jangyodsuk, Jun Wan 0001 |
Pattern Anal. Appl. | 5 |
| 2016 | Age Estimation Based on a Single Network with Soft Softmax of Aging Modeling
Zichang Tan, Jun Wan 0001, Zhen Lei 0001, Stan Z. Li |
ACCV (3) | 3 |
| 2016 | ChaLearn Joint Contest on Multimedia Challenges Beyond Visual Analysis: An overviewabstractThis paper provides an overview of the Joint Contest on Multimedia Challenges Beyond Visual Analysis. We organized an academic competition that focused on four problems that require effective processing of multimodal information in order to be solved. Two tracks were devoted to gesture spotting and recognition from RGB-D video, two fundamental problems for human computer interaction. Another track was devoted to a second round of the first impressions challenge of which the goal was to develop methods to recognize personality traits from short video clips. For this second round we adopted a novel collaborative-competitive (i.e., coopetition) setting. The fourth track was dedicated to the problem of video recommendation for improving user experience. The challenge was open for about 45 days, and received outstanding participation: almost 200 participants registered to the contest, and 20 teams sent predictions in the final stage. The main goals of the challenge were fulfilled: the state of the art was advanced considerably in the four tracks, with novel solutions to the proposed problems (mostly relying on deep learning). However, further research is still required. The data of the four tracks will be available to allow researchers to keep making progress in the four tracks. Hugo Jair Escalante, Víctor Ponce-López, Jun Wan 0001, Michael Riegler 0001, Albert Clapés, Sergio Escalera, Isabelle Guyon, Xavier Baró, Pål Halvorsen, Henning Müller, Martha A. Larson |
ICPR | 3 |
| 2016 | Explore Efficient Local Features from RGB-D Data for One-Shot Learning Gesture RecognitionabstractAvailability of handy RGB-D sensors has brought about a surge of gesture recognition research and applications. Among various approaches, one shot learning approach is advantageous because it requires minimum amount of data. Here, we provide a thorough review about one-shot learning gesture recognition from RGB-D data and propose a novel spatiotemporal feature extracted from RGB-D data, namely mixed features around sparse keypoints (MFSK). In the review, we analyze the challenges that we are facing, and point out some future research directions which may enlighten researchers in this field. The proposed MFSK feature is robust and invariant to scale, rotation and partial occlusions. To alleviate the insufficiency of one shot training samples, we augment the training samples by artificially synthesizing versions of various temporal scales, which is beneficial for coping with gestures performed at varying speed. We evaluate the proposed method on the Chalearn gesture dataset (CGD). The results show that our approach outperforms all currently published approaches on the challenging data of CGD, such as translated, scaled and occluded subsets. When applied to the RGB-D datasets that are not one-shot (e.g., the Cornell Activity Dataset-60 and MSR Daily Activity 3D dataset), the proposed feature also produces very promising results under leave-one-out cross validation or one-shot learning. Jun Wan 0001, Guodong Guo, Stan Z. Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Dimensionality reduction using graph-embedded probability-based semi-supervised discriminant analysis
Wei Li 0162, Qiuqi Ruan, Jun Wan 0001 |
Neurocomputing | 3 |
| 2014 | CSMMI: Class-Specific Maximization of Mutual Information for Action and Gesture RecognitionabstractIn this paper, we propose a novel approach called class-specific maximization of mutual information (CSMMI) using a submodular method, which aims at learning a compact and discriminative dictionary for each class. Unlike traditional dictionary-based algorithms, which typically learn a shared dictionary for all of the classes, we unify the intraclass and interclass mutual information (MI) into an single objective function to optimize class-specific dictionary. The objective function has two aims: 1) maximizing the MI between dictionary items within a specific class (intrinsic structure) and 2) minimizing the MI between the dictionary items in a given class and those of the other classes (extrinsic structure). We significantly reduce the computational complexity of CSMMI by introducing an novel submodular method, which is one of the important contributions of this paper. This paper also contributes a state-of-the-art end-to-end system for action and gesture recognition incorporating CSMMI, with feature extraction, learning initial dictionary per each class by sparse coding, CSMMI via submodularity, and classification based on reconstruction errors. We performed extensive experiments on synthetic data and eight benchmark data sets. Our experimental results show that CSMMI outperforms shared dictionary methods and that our end-to-end system is competitive with other state-of-the-art approaches. Jun Wan 0001, Vassilis Athitsos, Pat Jangyodsuk, Hugo Jair Escalante, Qiuqi Ruan, Isabelle Guyon |
IEEE Trans. Image Process. | 1 |
| 2013 | Graph-preserving shortest feature line segment for dimensionality reduction
Wei Li 0162, Qiuqi Ruan, Jun Wan 0001 |
Neurocomputing | 3 |
| 2013 | One-shot learning gesture recognition from RGB-D data using bag of features
Jun Wan 0001, Qiuqi Ruan, Wei Li 0162, Shuang Deng |
J. Mach. Learn. Res. | 1 |