Zhenhua Feng 0001

dblp:348/7584-1 · DBLP profile ↗
← Back
95ranked-venue papers
8as first author
68since 2021 · last 2026
0000-0002-4485-4249ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 59 · 5 first-author · 42 since 2021Graphics, computer vision, multimedia, augmented reality and games · 46 · 7 first-author · 29 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Security and privacy · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Zero-Shot Evolutionary Architecture Search for Low-Rank Adaptation
abstract
Fine-tuning large-scale Transformer-based models is computationally expensive due to the enormous parameter space. Low-Rank Adaptation (LoRA) substantially reduces the number of trainable parameters while maintaining performance; however, identifying the optimal LoRA configuration — such as rank r, scaling factor [Formula: see text], and insertion positions — remains challenging. To address this issue, we propose a zero-shot proxy metric, termed Gradient Projection Score (GPS), which enables rapid evaluation of candidate configurations using only a few forward and backward passes. Building upon this metric, we further introduce EvoLoRA, a zero-shot evolutionary architecture search method that jointly optimizes three objectives: performance proxy, evaluation stability, and trainable parameter size. EvoLoRA automatically discovers effective LoRA configurations across different models and datasets. Experimental results demonstrate that GPS is strongly correlated with final model performance; moreover, on tasks such as image classification and object detection, EvoLoRA markedly reduces search and training costs while generally outperforming other fine-tuning methods and manually designed LoRA configurations.
Pengjin Wu, Ferrante Neri, Zhenhua Feng 0001
Int. J. Neural Syst.3
2026 Interactive image-to-video transfer learning
Cong Wu 0006, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
Neural Networks3
2026 Improving Generalized Visual Grounding With Instance-Aware Joint Learning
abstract
Generalized visual grounding tasks, including Generalized Referring Expression Comprehension (GREC) and Segmentation (GRES), extend the classical visual grounding paradigm by accommodating multi-target and non-target scenarios. Specifically, GREC focuses on accurately identifying all referential objects at the coarse bounding box level, while GRES aims for achieve fine-grained pixel-level perception. However, existing approaches typically treat these tasks independently, overlooking the benefits of jointly training GREC and GRES to ensure consistent multi-granularity predictions and streamline the overall process. Moreover, current methods often treat GRES as a semantic segmentation task, neglecting the crucial role of instance-aware capabilities and the necessity of ensuring consistent predictions between instance-level boxes and masks. To address these limitations, we propose InstanceVG, a multi-task generalized visual grounding framework equipped with instance-aware capabilities, which leverages instance queries to unify the joint and consistency predictions of instance-level boxes and masks. To the best of our knowledge, InstanceVG is the first framework to simultaneously tackle both GREC and GRES while incorporating instance-aware capabilities into generalized visual grounding. To instantiate the framework, we assign each instance query a prior reference point, which also serves as an additional basis for target matching. This design facilitates consistent predictions of points, boxes, and masks for the same instance. Extensive experiments obtained on ten datasets across four tasks demonstrate that InstanceVG achieves state-of-the-art performance, significantly surpassing the existing methods in various evaluation metrics. The code and model will be publicly available at https://github.com/Dmmm1997/InstanceVG.
Wenxuan Cheng, Jiang-Jiang Liu 0001, Lingfeng Yang, Zhenhua Feng 0001, Wankou Yang, Jingdong Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Probabilistically Aligned View-Unaligned Clustering With Adaptive Template Selection
abstract
In most existing multi-view modeling scenarios, cross-view correspondence (CVC) between instances of the same target from different views, like paired image-text data, is a crucial prerequisite for effortlessly deriving a consistent representation. Nevertheless, this premise is frequently compromised in certain applications, where each view is organized and transmitted independently, resulting in the view-unaligned problem (VuP). Restoring CVC of unaligned multi-view data is a challenging and highly demanding task that has received limited attention from the research community. To tackle this practical challenge, we propose to integrate the permutation derivation procedure into the bipartite graph paradigm for view-unaligned clustering, termed Probabilistically Aligned View-unaligned Clustering with Adaptive Template Selection (PAVuC-ATS). Specifically, we learn consistent anchors and view-specific graphs by the bipartite graph, and derive permutations applied to the unaligned graphs by reformulating the alignment between two latent representations as a 2-step transition of a Markov chain with adaptive template selection, thereby achieving the probabilistic alignment. The convergence of the resultant optimization problem is validated both experimentally and theoretically. Extensive experiments on six benchmark datasets demonstrate the superiority of the proposed PAVuC-ATS over the baseline methods.
Wenhua Dong, Xiaojun Wu 0001, Zhenhua Feng 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 DRL: An efficient heterogeneous spatial feature interaction framework for UAV self-localization
Enhui Zheng, Wenxuan Cheng, Zhenhua Feng 0001, Wankou Yang
Pattern Recognit.5
2026 GC3VG: Generalized Multi-Task Visual Grounding With Coarse-to-Fine Consistency Constraints
abstract
In this work, we propose an efficient and streamlined paradigm to address the challenge of consistency prediction in generalized multi-task visual grounding. While most existing approaches primarily focus on integrating multi-modal information and employing multi-task learning to enhance both visual and linguistic understanding, they often rely on joint supervision at the region and pixel levels to exploit task complementarities. In contrast, C3VG explores the relatively under-addressed problem ofconsistency across multi-task predictions. To this end, a multi-task visual grounding framework based on a coarse-to-fine architecture is introduced. Empirical studies demonstrate that the incorporation of both implicit and explicit consistency constraints substantially enhances the coherence between detection and segmentation outputs. However, C3VG is restricted to single-referent visual grounding scenarios and exhibits limited generalizability to real-world applications, which often involve multi-referents or even absent referent. To overcome these limitations, we proposeGC3VG, which incorporates three key advancements: (1) extension to generalized scenarios, including both multi-referent and non-referent cases; (2) aUnified Coherent Refinement Modulethat implicitly encodes region- and instance-level features while explicitly modeling their relational alignment through an IoUbased constraint; and (3) aGranularity-aware Hard-mining Alignmentstrategy that enforces prediction consistency in the feature space and simultaneously enhances the discriminative power of visual and linguistic representations. Extensive experiments on RefCOCO/+/g and gRefCOCO demonstrate the effectiveness and generalizability of the proposed framework.
Kai Chen 0037, Wenxuan Cheng, Jiedong Zhuang, Zhenhua Feng 0001, Pengfei Zhu 0001, Wankou Yang
IEEE Trans. Circuits Syst. Video Technol.5
2025 R-DTI: Drug Target Interaction Prediction Based on Second-Order Relevance Exploration
abstract
Drug Target Interaction (DTI) prediction has witnessed promising performance boosts accompanied by advanced multimodal feature extraction. However, existing approaches suffer from two main difficulties. First, the complex protein structures cannot be well represented by current protein-sequence-based feature extractors. Second, the gap between protein and drug features increases the vulnerability of the obtained classifier thus degrading the prediction robustness. To address these issues, we propose a novel R-DTI method by exploring the second-order relevance in both protein structural feature extraction and DTI prediction phases. Specifically, we construct a pre-trained structural feature extractor that mines the atomic relevance of each amino acid. Then, an inter-feature structure-preserved Riemannian network is designed to expand the existing protein extraction patterns. To improve the prediction robustness, we also develop a Riemannian classifier that uses the second-order protein-drug relevance with a unified feature space. Extensive experimental results demonstrate the merits and superiority of our R-DTI against the state-of-the-art, achieving 1.4% and 1.9% higher AUC-ROC on the BindingDB and DrugBank datasets, respectively.
Yang Hua 0002, Tianyang Xu 0001, Xiaoning Song, Zhenhua Feng 0001, Rui Wang 0050, Wenjie Zhang 0009, Xiaojun Wu 0001
AAAI4
2025 One Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image Fusion
abstract
Advanced image fusion methods mostly prioritise high-level missions, where task interaction struggles with semantic gaps, requiring complex bridging mechanisms. In contrast, we propose to leverage low-level vision tasks from digital photography fusion, allowing for effective feature interaction through pixel-level supervision. This new paradigm provides strong guidance for unsupervised multimodal fusion without relying on abstract semantics, enhancing task-shared feature learning for broader applicability. Owning to the hybrid image features and enhanced universal representations, the proposed GIFNet supports diverse fusion tasks, achieving high performance across both seen and unseen scenarios with a single model. Uniquely, experimental results reveal that our framework also supports single-modality enhancement, offering superior flexibility for practical applications. Our code will be available at https://github.com/AWCXV/GIFNet.
Chunyang Cheng, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Zhangyong Tang, Hui Li 0037, Zeyang Zhang 0002, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler
CVPR3
2025 Text Augmented Correlation Transformer For Few-shot Classification & Segmentation
abstract
Foundation models like CLIP and ALIGN have transformed few-shot and zero-shot vision applications by fusing visual and textual data, yet the integrative few-shot classification and segmentation (FS-CS) task primarily leverages visual cues, overlooking the potential of textual support. In FS-CS scenarios, ambiguous object boundaries and overlapping classes often hinder model performance, as limited visual data struggles to fully capture high-level semantics. To bridge this gap, we present a novel multi-modal FS-CS framework that integrates textual cues into support data, facilitating enhanced semantic disambiguation and fine-grained segmentation. Our approach first investigates the unique contributions of exclusive text-based support, using only class labels to achieve FS-CS. This strategy alone achieves performance competitive with vision-only methods on FS-CS tasks, underscoring the power of textual cues in few-shot learning. Building on this, we introduce a dualmodal prediction mechanism that synthesizes insights from both textual and visual support sets, yielding robust multimodal predictions. This integration significantly elevates FS-CS performance, with classification and segmentation improvements of +3.7/6.6% (1-way 1-shot) and +8.0/6.5% (2-way 1-shot) on COCO-20i, and +2.2/3.8% (1-way 1shot) and +4.3/4.0% (2-way 1-shot) on Pascal-5i. Additionally, in weakly supervised FS-CS settings, our method surpasses visual-only benchmarks using textual support exclusively, further enhanced by our dual-modal predictions. By rethinking the role of text in FS-CS, our work establishes new benchmarks for multi-modal few-shot learning and demonstrates the efficacy of textual cues for improving model generalization and segmentation accuracy.
Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001
CVPR3
2025 Enhanced Weakly Supervised Few-shot Classification & Segmentation
abstract
The emergence of vision-language foundation models has enabled the integration of textual information into vision-based applications. However, in few-shot classification and segmentation (FS-CS), this potential remains underutilised. Commonly, self-supervised vision models have been employed, particularly in weakly-supervised scenarios, to generate pseudo-segmentation masks, as ground truth masks are typically unavailable and only target classification is provided. Despite their success, such models find it difficult to capture accurate semantics when compared to vision-language models. To address this limitation, we propose a novel FS-CS approach that leverages the rich semantic alignment of vision-language models to generate more precise pseudo ground-truth masks. While current vision-language models excel in global visual-text alignment, they struggle with finer, patch-level alignment, which is crucial for detailed segmentation tasks. To overcome this, we introduce a method that enhances patch-level alignment without requiring additional training. In addition, existing FS-CS frameworks typically lacks multi-scale information, limiting their ability to capture fine and coarse features simultaneously. To overcome this, we incorporate a module based on atrous convolutions to inject multi-scale information into the feature maps. Together, these contributions - text enhanced pseudo-mask generation and improved multi-scale feature representation - significantly boost the performance of our model in weakly-supervised settings, surpassing state-of-the-art methods and demonstrating the importance of integrating multi-modal information for robust FS-CS solutions.
Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001
ICASSP3
2025 PropVG: End-To-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
abstract
Recent advances in visual grounding have largely shifted away from traditional proposal-based two-stage frameworks due to their inefficiency and high computational complexity, favoring end-to-end direct reference paradigms. However, these methods rely exclusively on the referred target for supervision, overlooking the potential benefits of prominent prospective targets. Moreover, existing approaches often fail to incorporate multi-granularity discrimination, which is crucial for robust object identification in complex scenarios. To address these limitations, we propose PropVG, an end-to-end proposal-based framework that, to the best of our knowledge, is the first to seamlessly integrate foreground object proposal generation with referential object comprehension without requiring additional detectors. Furthermore, we introduce a Contrastive-based Refer Scoring (CRS) module, which employs contrastive learning at both sentence and word levels to enhance the capability in understanding and distinguishing referred objects. Additionally, we design a Multi-granularity Target Discrimination (MTD) module that fuses object- and semantic-level information to improve the recognition of absent targets. Extensive experiments on gRefCOCO (GREC/GRES), Ref-ZOM, R-RefCOCO, and RefCOCO (REC/RES) benchmarks demonstrate the effectiveness of PropVG. The codes and models are available at https://github.com/Dmmm1997/PropVG.
Wenxuan Cheng, Jiedong Zhuang, Jiang-jiang Liu, Hongshen Zhao, Zhenhua Feng 0001, Wankou Yang
ICCV6
2025 Catching Inter-Modal Artifacts: A Cross-Modal Framework for Temporal Forgery Localization
Yuhan Cai, Yang Hua 0002, Wenjie Zhang 0009, Xiaoning Song, Zhenhua Feng 0001
ICIC (6)5
2025 DASViT: Differentiable Architecture Search for Vision Transformer
abstract
Designing effective neural networks is a cornerstone of deep learning, and Neural Architecture Search (NAS) has emerged as a powerful tool for automating this process. Among the existing NAS approaches, Differentiable Architecture Search (DARTS) has gained prominence for its efficiency and ease of use, inspiring numerous advancements. Since the rise of Vision Transformers (ViT), researchers have applied NAS to explore ViT architectures, often focusing on macro-level search spaces and relying on discrete methods like evolutionary algorithms. While these methods ensure reliability, they face challenges in discovering innovative architectural designs, demand extensive computational resources, and are time-intensive. To address these limitations, we introduce Differentiable Architecture Search for Vision Transformer (DASViT), which bridges the gap in differentiable search for ViTs and uncovers novel designs. Experiments show that DASViT delivers architectures that break traditional Transformer encoder designs, outperform ViT-B/16 on multiple datasets, and achieve superior efficiency with fewer parameters and FLOPs.
Pengjin Wu, Ferrante Neri, Zhenhua Feng 0001
IJCNN3
2025 Revisiting Generative Infrared and Visible Image Fusion Based on Human Cognitive Laws
abstract
Existing infrared and visible image fusion methods often face the dilemma of balancing modal information. Generative fusion methods reconstruct fused images by learning from data distributions, but their generative capabilities remain limited. Moreover, the lack of interpretability in modal information selection further affects the reliability and consistency of fusion results in complex scenarios. This manuscript revisits the essence of generative image fusion under the inspiration of human cognitive laws and proposes a novel infrared and visible image fusion method, termed HCLFuse. First, HCLFuse investigates the quantification theory of information mapping in unsupervised fusion networks, which leads to the design of a multi-scale mask-regulated variational bottleneck encoder. This encoder applies posterior probability modeling and information decomposition to extract accurate and concise low-level modal information, thereby supporting the generation of high-fidelity structural details. Furthermore, the probabilistic generative capability of the diffusion model is integrated with physical laws, forming a time-varying physical guidance mechanism that adaptively regulates the generation process at different stages, thereby enhancing the ability of the model to perceive the intrinsic structure of data and reducing dependence on data quality. Experimental results show that the proposed method achieves state-of-the-art fusion performance in qualitative and quantitative evaluations across multiple datasets and significantly improves semantic segmentation metrics. This fully demonstrates the advantages of this generative image fusion method, drawing inspiration from human cognition, in enhancing structural consistency and detail quality.
Xiaoqing Luo, Zhancheng Zhang, Hui Li 0037, Rui Wang 0050, Zhenhua Feng 0001, Xiaoning Song
NeurIPS7
2025 Fastere: a fast framework for entity relation extractions
Wenjie Zhang 0009, Tianyang Xu 0001, Yang Hua 0002, Zhenhua Feng 0001, Xiaoning Song
Data Min. Knowl. Discov.4
2025 Investigating Self-Supervised Methods for Label-Efficient Learning
abstract
Abstract Vision transformers combined with self-supervised learning have enabled the development of models which scale across large datasets for several downstream tasks, including classification, segmentation, and detection. However, the potential of these models for low-shot learning across several downstream tasks remains largely under explored. In this work, we conduct a systematic examination of different self-supervised pretext tasks, namely contrastive learning, clustering, and masked image modelling, to assess their low-shot capabilities by comparing different pretrained models. In addition, we explore the impact of various collapse avoidance techniques, such as centring, ME-MAX, and sinkhorn, on these downstream tasks. Based on our detailed analysis, we introduce a framework that combines mask image modelling and clustering as pretext tasks. This framework demonstrates superior performance across all examined low-shot downstream tasks, including multi-class classification, multi-label classification and semantic segmentation. Furthermore, when testing the model on large-scale datasets, we show performance gains in various tasks.
Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001
Int. J. Comput. Vis.3
2025 Correction: Investigating Self-Supervised Methods for Label-Efficient Learning
Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001
Int. J. Comput. Vis.3
2025 Learning Structure-Supporting Dependencies via Keypoint Interactive Transformer for General Mammal Pose Estimation
Tianyang Xu 0001, Jiyong Rao, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001
Int. J. Comput. Vis.4
2025 MMDG-DTI: Drug-target interaction prediction via multimodal feature fusion and domain generalization
Yang Hua 0002, Zhenhua Feng 0001, Xiaoning Song, Xiaojun Wu 0001, Josef Kittler
Pattern Recognit.2
2025 Adaptive Colour-Depth Aware Attention for RGB-D Object Tracking
abstract
Recent advances in RGB-D tracking have been driven by the synergistic combination of high-performing RGB-only trackers and auxiliary depth information. However, most existing methods rely on visual feature descriptors to extract depth features, which are then fused with vision features. This pipeline may lead to performance degradation due to the incongruence between the RGB and depth modalities. In this letter, we propose an efficient and effective transformer-based framework, that explicitly models colour and depth information for RGB-D tracking. Specifically, we first statistically code the colour and depth information of the foreground and background for the template. Then, the spatial attention maps of the search region are obtained using these colour-depth statistical models, enhancing the visual features of the search region for improved object localisation accuracy. The comprehensive experimental results obtained on multiple benchmarks demonstrate the effectiveness and merits of the proposed approach in explicit colour-depth coding for RGB-D tracking. The code and the models are publicly accessible athttps://github.com/xuefeng-zhu5/CDAAT.
Xuefeng Zhu 0003, Tianyang Xu 0001, Xiaojun Wu 0001, Zhenhua Feng 0001, Josef Kittler
IEEE Signal Process. Lett.4
2025 Revisiting RGBT Tracking Benchmarks From the Perspective of Modality Validity: A New Benchmark, Problem, and Solution
abstract
RGBT tracking draws increasing attention because of its robustness in multi-modal warranting (MMW) scenarios, such as nighttime and adverse weather conditions, where relying on a single sensing modality fails to ensure stable tracking results. However, existing benchmarks predominantly contain videos collected in common scenarios where both RGB and thermal infrared (TIR) information are of sufficient quality. This weakens the representativeness of existing benchmarks in severe imaging conditions, leading to tracking failures in MMW scenarios. To bridge this gap, we present a new benchmark considering the modality validity, MV-RGBT, captured specifically from MMW scenarios where either RGB (extreme illumination) or TIR (thermal truncation) modality is invalid. Hence, it is further divided into two subsets according to the valid modality, offering a new compositional perspective for evaluation and providing valuable insights for future designs. Moreover, MV-RGBT is the most diverse benchmark of its kind, featuring 36 different object categories captured across 19 distinct scenes. Furthermore, considering severe imaging conditions in MMW scenarios, a new problem is posed in RGBT tracking, named 'when to fuse', to stimulate the development of fusion strategies for such scenarios. To facilitate its discussion, we propose a new solution with a mixture of experts, named MoETrack, where each expert generates independent tracking results along with a confidence score. Extensive results demonstrate the significant potential of MV-RGBT in advancing RGBT tracking and elicit the conclusion that fusion is not always beneficial, especially in MMW scenarios. Besides, MoETrack achieves state-of-the-art results on several benchmarks, including MV-RGBT, GTOT, and LasHeR. Source codes and benchmarks are available at https://github.com/Zhangyong-Tang/MVRGBT.
Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Xuefeng Zhu 0003, Chunyang Cheng, Zhenhua Feng 0001, Josef Kittler
IEEE Trans. Image Process.6
2025 ATMNet: Adaptive Two-Stage Modular Network for Accurate Video Captioning
abstract
In recent years, pretrained language-image models (PLIMs) have delivered advances in video captioning. However, existing PLIMs primarily focus on extracting global feature representations from still images and text sequences, while neglecting fine-grained semantic alignment and temporal variations between vision and text pairs. To this end, we propose a global-local alignment module and a temporal parsing module to reflect the detailed correspondence and temporal perception between the two modalities, respectively. In particular, the global-local alignment module enables cross-modal registration at two levels, i.e., the sentence-video level and the word-frame level, to obtain mixed-granularity semantic video features. The temporal parsing module is a dedicated self-attention structure that highlights temporal order cues along video frames, compensating for the limited temporal capacity of PLIMs. In addition, an adaptive two-stage gating structure is designed to leverage the linguistic predictions further. The linguistic information derived from the first stage prediction is dynamically routed through an adaptive decision gate, allowing for quality assessment of whether the information should proceed to the second stage. This structure can effectively reduce the computational burden for easy samples and further improve the accuracy of the prediction results. The experimental results obtained on several benchmark datasets demonstrate the effectiveness of the proposed solution, with improved performance compared to state-of-the-art methods.
Tianyang Xu 0001, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001
IEEE Trans. Syst. Man Cybern. Syst.4
2024 SCD-Net: Spatiotemporal Clues Disentanglement Network for Self-Supervised Skeleton-Based Action Recognition
abstract
Contrastive learning has achieved great success in skeleton-based action recognition. However, most existing approaches encode the skeleton sequences as entangled spatiotemporal representations and confine the contrasts to the same level of representation. Instead, this paper introduces a novel contrastive learning framework, namely Spatiotemporal Clues Disentanglement Network (SCD-Net). Specifically, we integrate the decoupling module with a feature extractor to derive explicit clues from spatial and temporal domains respectively. As for the training of SCD-Net, with a constructed global anchor, we encourage the interaction between the anchor and extracted clues. Further, we propose a new masking strategy with structural constraints to strengthen the contextual associations, leveraging the latest development from masked image modelling into the proposed SCD-Net. We conduct extensive evaluations on the NTU-RGB+D (60&120) and PKU-MMD (I&II) datasets, covering various downstream tasks such as action recognition, action retrieval, transfer learning, and semi-supervised learning. The experimental results demonstrate the effectiveness of our method, which outperforms the existing state-of-the-art (SOTA) approaches significantly. Our code and supplementary material can be found at https://github.com/cong-wu/SCD-Net.
Cong Wu 0006, Xiaojun Wu 0001, Josef Kittler, Tianyang Xu 0001, Muhammad Awais 0001, Zhenhua Feng 0001
AAAI7
2024 LabelPrompt: Effective prompt-based learning for relation classification
Wenjie Zhang 0009, Xiaoning Song, Zhenhua Feng 0001, Tianyang Xu 0001, Xiaojun Wu 0001
ACML3
2024 Pseudo Labelling for Enhanced Masked Auto Encoders
Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001
BMVC3
2024 C2C: Component-to-Composition Learning for Zero-Shot Compositional Action Recognition
Rongchang Li 0001, Zhenhua Feng 0001, Tianyang Xu 0001, Linze Li 0002, Xiaojun Wu 0001, Muhammad Awais 0001, Sara Atito Ali Ahmed, Josef Kittler
ECCV (38)2
2024 Efficient Few-Shot Action Recognition via Multi-level Post-reasoning
Cong Wu 0006, Xiaojun Wu 0001, Linze Li 0002, Tianyang Xu 0001, Zhenhua Feng 0001, Josef Kittler
ECCV (3)5
2024 Investigating Self-Supervised Methods for Label-Efficient Learning
abstract
Vision transformers combined with self-supervised learning have enabled the development of models which scale across large datasets for several downstream tasks like classification, segmentation and detection. The low-shot learning capability of these models, across several low-shot downstream tasks, has been largely under explored. We perform a system level study of different self supervised pretext tasks, namely contrastive learning, clustering, and masked image modelling for their low-shot capabilities by comparing the pretrained models. In addition we also study the effects of collapse avoidance methods, namely centring, ME-MAX, sinkhorn, on these downstream tasks. Based on our detailed analysis, we introduce a framework involving both mask image modelling and clustering as pretext tasks, which performs better across all low-shot downstream tasks, including multi-class classification, multi-label classification and semantic segmentation. Furthermore, when testing the model on full scale datasets, we show performance gains in multi-class classification, multi-label classification and semantic segmentation.
Srinivasa Rao Nandam, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001
ICIP3
2024 Masked Momentum Contrastive Learning for Semantic Understanding by Observation
abstract
Large language models (LLMs) have shown excellent performance in zero-shot learning using natural language prompts. However, in the domain of computer vision (CV), the paradigm of pretraining followed by finetuning remains dominant. The aim of this study is to reduce this gap by utilizing the capability of Self-Supervised Learning (SSL) in semantic understanding for zero-shot segmentation, without relying on human-provided labels or vision-language supervision. We introduce a novel evaluation framework that employs visual prompts, including a threshold and a query patch. This framework evaluates the ability of SSL models to derive concepts from observational data. Through this evaluation, we identify the strengths and limitations of SSL models in understanding semantics. Building on the insights from various SSL methods, we further propose the MMC approach to enhance the representations for objects, which integrates Masked image modeling, Momentum-based self-distillation, and global Contrastive learning. MMC achieves a better balance between the inter-object discriminability and the intra-object compactness of learned features. Our experiments on COCO, DAVIS-2017, PASCAL VOC, and ADE20K demonstrate outstanding performance of MMC’s representations.
Jiantao Wu, Shentong Mo, Sara Atito Ali Ahmed, Zhenhua Feng 0001, Josef Kittler, Syed Sameed Husain, Muhammad Awais 0001
ICIP4
2024 StableTalk: Advancing Audio-to-Talking Face Generation with Stable Diffusion and Vision Transformer
Fatemeh Nazarieh, Josef Kittler, Muhammad Awais 0001, Diptesh Kanojia, Zhenhua Feng 0001
ICPR (6)5
2024 A Novel Loss for Contrastive Deep Supervision
Zhengming Ye, Yang Hua 0002, Wenjie Zhang 0009, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001
ICPR (25)5
2024 Human-Aligned Longitudinal Control for Occluded Pedestrian Crossing With Visual Attention
abstract
Reinforcement Learning (RL) has been widely used to create generalizable autonomous vehicles. However, they rely on fixed reward functions that struggle to balance values like safety and efficiency. How can autonomous vehicles balance different driving objectives and human values in a constantly changing environment? To bridge this gap, we propose an adaptive reward function that utilizes visual attention maps to detect pedestrians in the driving scene and dynamically switch between prioritizing safety or efficiency depending on the current observation. The visual attention map is used to provide spatial attention to the RL agent to boost the training efficiency of the pipeline. We evaluate the pipeline against variants of an occluded pedestrian crossing scenario in the CARLA Urban Driving simulator. Specifically, the proposed pipeline is compared against a modular setup that combines the well-established object detection model, YOLO, with a Proximal Policy Optimization (PPO) agent. The results indicate that the proposed approach can compete with the modular setup while yielding greater training efficiency. The trajectories collected with the approach confirm the effectiveness of the proposed adaptive reward function.
Vinal Asodia, Zhenhua Feng 0001, Saber Fallah
ICRA2
2024 SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion
abstract
Visual grounding is a common vision task that involves grounding descriptive sentences to the corresponding regions of an image. Most existing methods use independent image-text encoding and apply complex hand-crafted modules or encoder-decoder architectures for modal interaction and query reasoning. However, their performance significantly drops when dealing with complex textual expressions. This is because the former paradigm only utilizes limited downstream data to fit the multi-modal feature fusion. Therefore, it is only effective when the textual expressions are relatively simple. In contrast, given the wide diversity of textual expressions and the uniqueness of downstream training data, the existing fusion module, which extracts multimodal content from a visual-linguistic context, has not been fully investigated. In this paper, we present a simple yet robust transformer-based framework, SimVG, for visual grounding. Specifically, we decouple visual-linguistic feature fusion from downstream tasks by leveraging existing multimodal pre-trained models and incorporating additional object tokens to facilitate deep integration of downstream and pre-training tasks. Furthermore, we design a dynamic weight-balance distillation method in the multi-branch synchronous learning process to enhance the representation capability of the simpler branch. This branch only consists of a lightweight MLP, which simplifies the structure and improves reasoning speed. Experiments on six widely used VG datasets, i.e., RefCOCO/+/g, ReferIt, Flickr30K, and GRefCOCO, demonstrate the superiority of SimVG. Finally, the proposed method not only achieves improvements in efficiency and convergence speed but also attains new state-of-the-art performance on these benchmarks. Codes and models are available at https://github.com/Dmmm1997/SimVG.
Lingfeng Yang, Zhenhua Feng 0001, Wankou Yang
NeurIPS4
2024 Learning Feature Restoration Transformer for Robust Dehazing Visual Object Tracking
Tianyang Xu 0001, Yifan Pan, Zhenhua Feng 0001, Xuefeng Zhu 0003, Chunyang Cheng, Xiaojun Wu 0001, Josef Kittler
Int. J. Comput. Vis.3
2024 Guest Editorial: Special Issue on the British Machine Vision Conference 2022
Guang Yang 0006, Angelica I. Avilés-Rivero, Yingying Fang, Zhenhua Feng 0001, Gianluigi Ciocca, Yulia Hicks, Constantino Carlos Reyes-Aldasoro
Int. J. Comput. Vis.4
2024 View-shuffled clustering via the modified Hungarian algorithm
Wenhua Dong, Xiaojun Wu 0001, Tianyang Xu 0001, Zhenhua Feng 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler
Neural Networks4
2024 Attention-based investigation and solution to the trade-off issue of adversarial training
Chang-Bin Shao, Wenbin Li 0006, Jing Huo, Zhenhua Feng 0001, Yang Gao 0001
Neural Networks4
2024 Self-supervised learning for RGB-D object tracking
Xuefeng Zhu 0003, Tianyang Xu 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Xiaojun Wu 0001, Zhenhua Feng 0001, Josef Kittler
Pattern Recognit.6
2024 Towards accurate unsupervised video captioning with implicit visual feature injection and explicit
Tianyang Xu 0001, Xiaoning Song, Xuefeng Zhu 0003, Zhenhua Feng 0001, Xiaojun Wu 0001
Pattern Recognit. Lett.5
2024 APMG: 3D Molecule Generation Driven by Atomic Chemical Properties
abstract
Recently, mask-fill-based 3D Molecular Generation (MG) methods have become very popular in virtual drug design. However, the existing MG methods ignore the chemical properties of atoms and contain inappropriate atomic position training data, which limits their generation capability. To mitigate the above issues, this paper presents a novel mask-fill-based 3D molecule generation model driven by atomic chemical properties (APMG). Specifically, we construct a new attention-MPNN-based encoder and introduce the electronic information into atom representations to enrich chemical properties. Also, a multi-functional classifier is designed to predict the electronic information of each generated atom, guiding the type prediction of elements and bonds. By design, the proposed method uses the chemical properties of atoms and their correlations for high-quality molecule generation. Second, to optimize the atomic position training data, we propose a novel atomic training position generation approach using the Chi-Square distribution. We evaluate our APMG method on the CrossDocked dataset and visualize the docking states of the pockets and generated molecules. The obtained results demonstrate the superiority and merits of APMG over the state-of-the-art approaches.
Yang Hua 0002, Zhenhua Feng 0001, Xiaoning Song, Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Dongjun Yu
IEEE ACM Trans. Comput. Biol. Bioinform.2
2024 A Survey of Cross-Modal Visual Content Generation
abstract
Cross-modal content generation has become very popular in recent years. To generate high-quality and realistic content, a variety of methods have been proposed. Among these approaches, visual content generation has attracted significant attention from academia and industry due to its vast potential in various applications. This survey provides an overview of recent advances in visual content generation conditioned on other modalities, such as text, audio, speech, and music, with a focus on their key contributions to the community. In addition, we summarize the existing publicly available datasets that can be used for training and benchmarking cross-modal visual content generation models. We provide an in-depth exploration of the datasets used for audio-to-visual content generation, filling a gap in the existing literature. Various evaluation metrics are also introduced along with the datasets. Furthermore, we discuss the challenges and limitations encountered in the area, such as modality alignment and semantic coherence. Last, we outline possible future directions for synthesizing visual content from other modalities including the exploration of new modalities, and the development of multi-task multi-modal networks. This survey serves as a resource for researchers interested in quickly gaining insights into this burgeoning field.
Fatemeh Nazarieh, Zhenhua Feng 0001, Muhammad Awais 0001, Wenwu Wang 0001, Josef Kittler
IEEE Trans. Circuits Syst. Video Technol.2
2024 Toward Transferable Attack via Adversarial Diffusion in Face Recognition
abstract
Modern face recognition systems widely use deep convolutional neural networks (DCNNs). However, DCNNs are susceptible to adversarial examples, posing security risks to these systems. Transferable adversarial examples that can be transferred from surrogate to target models greatly undermine the robustness of DCNNs. Numerous attempts have been made to generate transferable adversarial examples, but the existing methods often suffer from limited transferability or produce adversarial examples with poor image perceptual quality. Recently, diffusion models have shown remarkable success in image generation and have excelled in various downstream tasks. However, their potential in adversarial attacks remains largely unexplored. To bridge this gap, we propose a novel approach, namely Adversarial Diffusion Attack (ADA), in generation of transferable adversarial facial examples. ADA employs a dynamic game-like strategy between injection and denoising that progressively reinforces the robustness of adversarial perturbation in the reverse process of diffusion model. Additionally, both adversarial perturbation and residual image are embedded to drift benign distribution towards adversarial distribution, crafting adversarial examples with high image quality. Extensive experimental results obtained on two benchmarking datasets, LFW and CelebA-HQ, demonstrate that ADA achieves higher attack success rates and produces adversarial examples with superior image quality compared to the state-of-the-art methods.
Zhenhua Feng 0001, Xiaojun Wu 0001
IEEE Trans. Inf. Forensics Secur.3
2024 Vision-Based UAV Self-Positioning in Low-Altitude Urban Environments
abstract
Unmanned Aerial Vehicles (UAVs) rely on satellite systems for stable positioning. However, due to limited satellite coverage or communication disruptions, UAVs may lose signals for positioning. In such situations, vision-based techniques can serve as an alternative, ensuring the self-positioning capability of UAVs. However, most of the existing datasets are developed for the geo-localization task of the objects captured by UAVs, rather than UAV self-positioning. Furthermore, the existing UAV datasets apply discrete sampling to synthetic data, such as Google Maps, neglecting the crucial aspects of dense sampling and the uncertainties commonly experienced in practical scenarios. To address these issues, this paper presents a new dataset, DenseUAV, that is the first publicly available dataset tailored for the UAV self-positioning task. DenseUAV adopts dense sampling on UAV images obtained in low-altitude urban areas. In total, over 27K UAV- and satellite-view images of 14 university campuses are collected and annotated. In terms of methodology, we first verify the superiority of Transformers over CNNs for the proposed task. Then we incorporate metric learning into representation learning to enhance the model's discriminative capacity and to reduce the modality discrepancy. Besides, to facilitate joint learning from both the satellite and UAV views, we introduce a mutually supervised learning approach. Last, we enhance the Recall@K metric and introduce a new measurement, SDM@K, to evaluate both the retrieval and localization performance for the proposed task. As a result, the proposed baseline method achieves a remarkable Recall@1 score of 83.01% and an SDM@1 score of 86.50% on DenseUAV. The dataset and code have been made publicly available on https://github.com/Dmmm1997/DenseUAV.
Enhui Zheng, Zhenhua Feng 0001, Lei Qi 0001, Jiedong Zhuang, Wankou Yang
IEEE Trans. Image Process.3
2024 Reinforcement Learning for Online Dispatching Policy in Real-Time Train Timetable Rescheduling
abstract
Train Timetable Rescheduling (TTR) is a crucial task in the daily operation of high-speed railways to maintain punctuality and efficiency in the presence of unexpected disturbances. However, it is challenging to promptly create a rescheduled timetable in real time. In this study, we propose a reinforcement-learning-based method for real-time rescheduling of high-speed trains. The key innovation of the proposed method is to learn a well-generalized dispatching policy from a large amount of samples, which can be applied to the TTR task directly. At first, the problem is transformed into a multi-stage decision process, and the decision agent is designed to predict dispatching rules. To enhance the training efficiency, we generate a small yet good-quality action set to reduce invalid explorations. Besides, we propose an action sampling strategy for action selection, which implements forward planning with consideration of evaluation uncertainty, thus improving search efficiency. Extensive experimental results demonstrate the effectiveness and competitiveness of the proposed method. It has been proven that the local policies trained by the proposed method can be applied to numerous problem instances directly, rendering it unnecessary to use human-designed rules.
Peng Yue 0005, Yaochu Jin, Xuewu Dai, Zhenhua Feng 0001, Dongliang Cui
IEEE Trans. Intell. Transp. Syst.4
2024 Reinforcement Learning for Scalable Train Timetable Rescheduling With Graph Representation
abstract
Train timetable rescheduling (TTR) aims to promptly restore the original operation of trains after unexpected disturbances or disruptions. Currently, this work is still done manually by train dispatchers, which is challenging to maintain performance under various problem instances. To mitigate this issue, this study proposes a reinforcement learning-based approach to TTR, which makes the following contributions compared to existing work. First, we design a simple directed graph to represent the TTR problem, enabling the automatic extraction of informative states through graph neural networks. Second, we reformulate the construction process of TTR’s solution, not only decoupling the decision model from the problem size but also ensuring the generated scheme’s feasibility. Third, we design a learning curriculum for our model to handle the scenarios with different levels of delay. Finally, a simple local search method is proposed to assist the learned decision model, which can significantly improve solution quality with little additional computation cost, further enhancing the practical value of our method. Extensive experimental results demonstrate the effectiveness of our method. The learned decision model can achieve better performance for various problems with varying degrees of train delay and different scales when compared to handcrafted rules and state-of-the-art solvers.
Peng Yue 0005, Yaochu Jin, Xuewu Dai, Zhenhua Feng 0001, Dongliang Cui
IEEE Trans. Intell. Transp. Syst.4
2024 One-pass View-unaligned Clustering
abstract
Given a set of multi-view instances, the prevailing assumption in most existing clustering approaches is that they are complete and exhibit cross-view alignment. However, this assumption is often unrealistic. In such scenarios, it could be satisfied at the cost of data pre-processing, but this would be complex and inconsistent with practical applications. Therefore, developing more effective solutions for the View-unaligned Problem (VuP) is highly desirable. Several pioneering works have tackled the partially VuP, yet handling fully VuP remains a challenge due to the reliance on partially pre-aligned instances. In this paper, we propose One-pass View-unaligned Clustering (OpVuC) that simultaneously aligns and clusters instances in a unified framework. Specifically, we alig shuffled instances with a selected template using an innovative global-local alignment scheme based on the notion of geometric invariance and separate the fully aligned instances using a relaxed$k$-means algorithm. The proposed OpVuC method can handle VuP at any alignment level without requiring any pre-aligned instances. Extensive experiments conducted on several benchmark datasets demonstrate the effectiveness and merits of the proposed OpVuC method.
Wenhua Dong, Xiaojun Wu 0001, Zhenhua Feng 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler
IEEE Trans. Multim.3
2023 Ada2NPT: An Adaptive Nearest Proxies Triplet Loss for Attribute-Aware Face Recognition with Adaptively Compacted Feature Learning
Lei Ju 0005, Zhenhua Feng 0001, Muhammad Awais 0001, Josef Kittler
ACML2
2023 Variational Autoencoders with Decremental Information Bottleneck for Disentanglement
Jiantao Wu, Shentong Mo, Xingshen Zhang, Muhammad Awais 0001, Zhenhua Feng 0001, Lin Wang 0004
BMVC6
2023 MFR-DTA: a multi-functional and robust model for predicting drug-target binding affinity and region
abstract
MOTIVATION: Recently, deep learning has become the mainstream methodology for drug-target binding affinity prediction. However, two deficiencies of the existing methods restrict their practical applications. On the one hand, most existing methods ignore the individual information of sequence elements, resulting in poor sequence feature representations. On the other hand, without prior biological knowledge, the prediction of drug-target binding regions based on attention weights of a deep neural network could be difficult to verify, which may bring adverse interference to biological researchers. RESULTS: We propose a novel Multi-Functional and Robust Drug-Target binding Affinity prediction (MFR-DTA) method to address the above issues. Specifically, we design a new biological sequence feature extraction block, namely BioMLP, that assists the model in extracting individual features of sequence elements. Then, we propose a new Elem-feature fusion block to refine the extracted features. After that, we construct a Mix-Decoder block that extracts drug-target interaction information and predicts their binding regions simultaneously. Last, we evaluate MFR-DTA on two benchmarks consistently with the existing methods and propose a new dataset, sc-PDB, to better measure the accuracy of binding region prediction. We also visualize some samples to demonstrate the locations of their binding sites and the predicted multi-scale interaction regions. The proposed method achieves excellent performance on these datasets, demonstrating its merits and superiority over the state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: https://github.com/JU-HuaY/MFR.
Yang Hua 0002, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001
Bioinform.3
2023 ETM-face: effective training sample selection and multi-scale feature learning for face detection
Junyuan He, Xiaoning Song, Zhenhua Feng 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
Multim. Tools Appl.3
2023 Global Context-Aware Feature Extraction and Visible Feature Enhancement for Occlusion-Invariant Pedestrian Detection in Crowded Scenes
Zhen Liu 0015, Xiaoning Song, Zhenhua Feng 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
Neural Process. Lett.3
2023 NPT-Loss: Demystifying Face Recognition Losses With Nearest Proxies Triplet
abstract
Face recognition (FR) using deep convolutional neural networks (DCNNs) has seen remarkable success in recent years. One key ingredient of DCNN-based FR is the design of a loss function that ensures discrimination between various identities. The state-of-the-art (SOTA) solutions utilise normalised Softmax loss with additive and/or multiplicative margins. Despite being popular and effective, these losses are justified only intuitively with little theoretical explanations. In this work, we show that under the LogSumExp (LSE) approximation, the SOTA Softmax losses become equivalent to a proxy-triplet loss that focuses on nearest-neighbour negative proxies only. This motivates us to propose a variant of the proxy-triplet loss, entitled Nearest Proxies Triplet (NPT) loss, which unlike SOTA solutions, converges for a wider range of hyper-parameters and offers flexibility in proxy selection and thus outperforms SOTA techniques. We generalise many SOTA losses into a single framework and give theoretical justifications for the assertion that minimising the proposed loss ensures a minimum separability between all identities. We also show that the proposed loss has an implicit mechanism of hard-sample mining. We conduct extensive experiments using various DCNN architectures on a number of FR benchmarks to demonstrate the efficacy of the proposed scheme over SOTA methods.
Syed Safwan Khalid, Muhammad Awais 0001, Zhenhua Feng 0001, Chi-Ho Chan, Ammarah Farooq, Ali Akbari 0003, Josef Kittler
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Attention-guided evolutionary attack with elastic-net regularization on face recognition
Zhenhua Feng 0001, Xiaojun Wu 0001
Pattern Recognit.3
2023 Keep an eye on faces: Robust face detection with heatmap-Assisted spatial attention and scale-Aware layer attention
abstract
Modern anchor-based face detectors learn discriminative features using large-capacity networks and extensive anchor settings. In spite of their promising results, they are not without problems. First, most anchors extract redundant features from the background. As a consequence, the performance improvements are achieved at the expense of a disproportionate computational complexity. Second, the predicted face boxes are only distinguished by a classifier supervised by pre-defined positive, negative and ignored anchors. This strategy may ignore potential contributions from cohorts of anchors labeled negative/ignored during inference simply because of their inferior initialisation, although they can regress well to a target. In other words, true positives and representative features may get filtered out by unreliable confidence scores. To deal with the first concern and achieve more efficient face detection, we propose a Heatmap-assisted Spatial Attention (HSA) module and a Scale-aware Layer Attention (SLA) module to extract informative features using lower computational costs. To be specific, SLA incorporates the information from all the feature pyramid layers, weighted adaptively to remove redundant layers. HSA predicts a reshaped Gaussian heatmap and employs it to facilitate a spatial feature selection by better highlighting facial areas. For more reliable decision-making, we merge the predicted heatmap scores and classification results by voting. Since our heatmap scores are based on the distance to the face centres, they are able to retain all the well-regressed anchors. The experiments obtained on several well-known benchmarks demonstrate the merits of the proposed method.
Lei Ju 0005, Josef Kittler, Muhammad Awais Rana, Wankou Yang, Zhenhua Feng 0001
Pattern Recognit.5
2023 CPInformer for Efficient and Robust Compound-Protein Interaction Prediction
abstract
Recently, deep learning has become the mainstream methodology for Compound-Protein Interaction (CPI) prediction. However, the existing compound-protein feature extraction methods have some issues that limit their performance. First, graph networks are widely used for structural compound feature extraction, but the chemical properties of a compound depend on functional groups rather than graphic structure. Besides, the existing methods lack capabilities in extracting rich and discriminative protein features. Last, the compound-protein features are usually simply combined for CPI prediction, without considering information redundancy and effective feature mining. To address the above issues, we propose a novel CPInformer method. Specifically, we extract heterogeneous compound features, including structural graph features and functional class fingerprints, to reduce prediction errors caused by similar structural compounds. Then, we combine local and global features using dense connections to obtain multi-scale protein features. Last, we apply ProbSparse self-attention to protein features, under the guidance of compound features, to eliminate information redundancy, and to improve the accuracy of CPInformer. More importantly, the proposed method identifies the activated local regions that link a CPI, providing a good visualisation for the CPI state. The results obtained on five benchmarks demonstrate the merits and superiority of CPInformer over the state-of-the-art approaches.
Yang Hua 0002, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler, Dongjun Yu
IEEE ACM Trans. Comput. Biol. Bioinform.3
2023 Motion-Driven Spatial and Temporal Adaptive High-Resolution Graph Convolutional Networks for Skeleton-Based Action Recognition
abstract
Graph convolutional networks (GCN) have attracted increasing interest in action recognition in recent years. GCN models human skeleton sequences as spatio-temporal graphs. Also, attention mechanisms are often jointly used with GCNs to highlight important frames or body joints in a sequence. However, attention modules learn parameters offline and are fixed, so may not adapt well to unseen samples. In this paper, we propose a simple but effective motion-driven spatial and temporal adaptation strategy to dynamically strengthen the features of important frames and joints for skeleton-based action recognition. The rationale is that the joints and frames with dramatic motions are generally more informative and discriminative. We combine the spatial and temporal refinements by using a two-branch structure, in which the joint and frame-wise feature refinements perform in parallel. Such a structure can lead to learn more complementary feature representations. Moreover, we propose to use the fully connected graph convolution to learn the long-range spatial dependencies. Besides, we investigate two high-resolution skeleton graphs by creating virtual joints, aiming to improve the representation of skeleton features. By combining the above proposals, we develop a novel motion-driven spatial and temporal adaptive high-resolution GCN. Experimental results demonstrate that the proposed model achieves state-of-the-art (SOTA) results on the challenging large-scale Kinetics-Skeleton and UAV-Human datasets, and it is on par with the SOTA methods on the two NTU-RGB+D 60&120 datasets. Additionally, our motion-driven adaptation method shows encouraging performance when compared with the attention mechanisms.
Zengxi Huang, Yusong Qin, Xiaobing Lin, Tianlin Liu, Zhenhua Feng 0001, Yiguang Liu
IEEE Trans. Circuits Syst. Video Technol.5
2023 Toward Robust Visual Object Tracking With Independent Target-Agnostic Detection and Effective Siamese Cross-Task Interaction
abstract
Advanced Siamese visual object tracking architectures are jointly trained using pair-wise input images to perform target classification and bounding box regression. They have achieved promising results in recent benchmarks and competitions. However, the existing methods suffer from two limitations: First, though the Siamese structure can estimate the target state in an instance frame, provided the target appearance does not deviate too much from the template, the detection of the target in an image cannot be guaranteed in the presence of severe appearance variations. Second, despite the classification and regression tasks sharing the same output from the backbone network, their specific modules and loss functions are invariably designed independently, without promoting any interaction. Yet, in a general tracking task, the centre classification and bounding box regression tasks are collaboratively working to estimate the final target location. To address the above issues, it is essential to perform target-agnostic detection so as to promote cross-task interactions in a Siamese-based tracking framework. In this work, we endow a novel network with a target-agnostic object detection module to complement the direct target inference, and to avoid or minimise the misalignment of the key cues of potential template-instance matches. To unify the multi-task learning formulation, we develop a cross-task interaction module to ensure consistent supervision of the classification and regression branches, improving the synergy of different branches. To eliminate potential inconsistencies that may arise within a multi-task architecture, we assign adaptive labels, rather than fixed hard labels, to supervise the network training more effectively. The experimental results obtained on several benchmarks, i.e., OTB100, UAV123, VOT2018, VOT2019, and LaSOT, demonstrate the effectiveness of the advanced target detection module, as well as the cross-task interaction, exhibiting superior tracking performance as compared with the state-of-the-art tracking methods.
Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
IEEE Trans. Image Process.2
2022 Memory-Token Transformer for Unsupervised Video Anomaly Detection
abstract
Video anomaly detection is crucial for behavior analysis, which has witnessed continuous progress in recent years with the auto-encoder based reconstruction framework. However, in some cases, abnormal frames may also be reconstructed well due to the strong representation ability of deep networks, increasing missed detection. To mitigate this issue, the existing methods usually the memory bank method. This method records normal patterns and assigns high errors for the reconstruction of abnormal frames into normal frames. In this paper, to better use the semantic information of normal videos recorded in the memory module, we introduce the Memory-Token Transformer (MTT) to boost the reconstruction performance on normal frames. We assume that the anomalies in a video mainly concentrate on the regions containing people and relevant objects. Therefore, during the decoding stage, we first extract the semantic concepts of a feature map and generate the corresponding semantic tokens. Then the tokens are combined with the proposed memory module. Last, we introduce a transformer to fuse the complex relationship among different tokens, and use 3D convolution with the pooling operator in our encoder to enhance spatio-temporal feature extraction as compared with 2D models. The experimental results obtained on various benchmarks demonstrate the effectiveness of the proposed method.
Youyu Li, Xiaoning Song, Tianyang Xu 0001, Zhenhua Feng 0001
ICPR4
2022 KITPose: Keypoint-Interactive Transformer for Animal Pose Estimation
Jiyong Rao, Tianyang Xu 0001, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001
PRCV (1)4
2022 Distribution Cognisant Loss for Cross-Database Facial Age Estimation With Sensitivity Analysis
abstract
Existing facial age estimation studies have mostly focused on intra-database protocols that assume training and test images are captured under similar conditions. This is rarely valid in practical applications, where we typically encounter training and test sets with different characteristics. In this article, we deal with such situations, namely subjective-exclusive cross-database age estimation. We formulate the age estimation problem as the distribution learning framework, where the age labels are encoded as a probability distribution. To improve the cross-database age estimation performance, we propose a new loss function which provides a more robust measure of the difference between ground-truth and predicted distributions. The desirable properties of the proposed loss function are theoretically analysed and compared with the state-of-the-art approaches. In addition, we compile a new balanced large-scale age estimation database. Last, we introduce a novel evaluation protocol, called subject-exclusive cross-database age estimation protocol, which provides meaningful information of a method in terms of the generalisation capability. The experimental results demonstrate that the proposed approach outperforms the state-of-the-art age estimation methods under both intra-database and subject-exclusive cross-database evaluation protocols. In addition, in this article, we provide a comparative sensitivity analysis of various algorithms to identify trends and issues inherent to their performance. This analysis introduces some open problems to the community which might be considered when designing a robust age estimation system.
Ali Akbari 0003, Muhammad Awais 0001, Zhenhua Feng 0001, Ammarah Farooq, Josef Kittler
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Target-Cognisant Siamese Network for Robust Visual Object Tracking
Yingjie Jiang, Xiaoning Song, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
Pattern Recognit. Lett.4
2022 Feature Alignment for Robust Acoustic Scene Classification Across Devices
abstract
This letter presents a feature alignment method for domain adaptive Acoustic Scene Classification (ASC) across recording devices. First, we design a two-stream network, in which each stream processes two features,i.e., Log-Mel spectrogram and delta-deltas, using two sub-networks. Second, we investigate different loss functions for feature alignment between the feature maps obtained by the source and target domains. Last, we present an alternate training strategy to deal with the data imbalance problem between paired and unpaired samples. The experimental results obtained on the DCASE benchmarks demonstrate the effectiveness and superiority of the proposed method. The source code of the proposed method is available athttps://github.com/Jingqiao-Zhao/FAASC.
Jingqiao Zhao, Qiuqiang Kong, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001
IEEE Signal Process. Lett.4
2022 Robust Visual Object Tracking Via Adaptive Attribute-Aware Discriminative Correlation Filters
abstract
In recent years, attention mechanisms have been widely studied in Discriminative Correlation Filter (DCF) based visual object tracking. To realise spatial attention and discriminative feature mining, existing approaches usually apply regularisation terms to the spatial dimension of multi-channel features. However, these spatial regularisation approaches construct a shared spatial attention pattern for all multi-channel features, without considering the diversity across channels. As each feature map (channel) focuses on a specific visual attribute, a shared spatial attention pattern limits the capability for mining important information from different channels. To address this issue, we advocate channel-specific spatial attention for DCF-based trackers. The key ingredient of the proposed method is an Adaptive Attribute-Aware spatial attention mechanism for constructing a novel DCF-based tracker (A$^3$DCF). To highlight the discriminative elements in each feature map, spatial sparsity is imposed in the filter learning stage, moderated by the prior knowledge regarding the expected concentration of signal energy. In addition, we perform a post processing of the identified spatial patterns to alleviate the impact of less significant channels. The net effect is that the irrelevant and inconsistent channels are removed by the proposed method. The results obtained on a number of well-known benchmarking datasets, including OTB2015, DTB70, UAV123, VOT2018, LaSOT, GOT-10 K and TrackingNet, demonstrate the merits of the proposed A$^3$DCF tracker, with improved performance compared to the state-of-the-art methods.
Xuefeng Zhu 0003, Xiaojun Wu 0001, Tianyang Xu 0001, Zhenhua Feng 0001, Josef Kittler
IEEE Trans. Multim.4
2021 Separable Batch Normalization for Robust Facial Landmark Localization
Shuangping Jin, Zhenhua Feng 0001, Wankou Yang, Josef Kittler
BMVC2
2021 Adaptive Channel Selection for Robust Visual Object Tracking with Discriminative Correlation Filters
abstract
Abstract Discriminative Correlation Filters (DCF) have been shown to achieve impressive performance in visual object tracking. However, existing DCF-based trackers rely heavily on learning regularised appearance models from invariant image feature representations. To further improve the performance of DCF in accuracy and provide a parsimonious model from the attribute perspective, we propose to gauge the relevance of multi-channel features for the purpose of channel selection. This is achieved by assessing the information conveyed by the features of each channel as a group, using an adaptive group elastic net inducing independent sparsity and temporal smoothness on the DCF solution. The robustness and stability of the learned appearance model are significantly enhanced by the proposed method as the process of channel selection performs implicit spatial regularisation. We use the augmented Lagrangian method to optimise the discriminative filters efficiently. The experimental results obtained on a number of well-known benchmarking datasets demonstrate the effectiveness and stability of the proposed method. A superior performance over the state-of-the-art trackers is achieved using less than $$10\%$$ 10 % deep feature channels.
Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
Int. J. Comput. Vis.2
2021 Complementary Discriminative Correlation Filters Based on Collaborative Representation for Visual Object Tracking
abstract
In recent years, discriminative correlation filter (DCF) based algorithms have significantly advanced the state of the art in visual object tracking. The key to the success of DCF is an efficient discriminative regression model trained with powerful multi-cue features, including both hand-crafted and deep neural network features. However, the tracking performance is hindered by their inability to respond adequately to abrupt target appearance variations. This issue is posed by the limited representation capability of fixed image features. In this work, we set out to rectify this shortcoming by proposing a complementary representation of a visual content. Specifically, we propose the use of a collaborative representation between successive frames to extract the dynamic appearance information from a target with rapid appearance changes, which results in suppressing the undesirable impact of the background. The resulting collaborative representation coefficients are combined with the original feature maps using a spatially regularised DCF framework for performance boosting. The experimental results on several benchmarking datasets demonstrate the effectiveness and robustness of the proposed method, as compared with a number of state-of-the-art tracking algorithms.
Xuefeng Zhu 0003, Xiaojun Wu 0001, Tianyang Xu 0001, Zhenhua Feng 0001, Josef Kittler
IEEE Trans. Circuits Syst. Video Technol.4
2021 From RGB to Depth: Domain Transfer Network for Face Anti-Spoofing
Yahang Wang, Xiaoning Song, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001
IEEE Trans. Inf. Forensics Secur.4
2021 SP-GAN: Self-Growing and Pruning Generative Adversarial Networks
abstract
This article presents a new Self-growing and Pruning Generative Adversarial Network (SP-GAN) for realistic image generation. In contrast to traditional GAN models, our SP-GAN is able to dynamically adjust the size and architecture of a network in the training stage by using the proposed self-growing and pruning mechanisms. To be more specific, we first train two seed networks as the generator and discriminator; each contains a small number of convolution kernels. Such small-scale networks are much easier and faster to train than large-capacity networks. Second, in the self-growing step, we replicate the convolution kernels of each seed network to augment the scale of the network, followed by fine-tuning the augmented/expanded network. More importantly, to prevent the excessive growth of each seed network in the self-growing stage, we propose a pruning strategy that reduces the redundancy of an augmented network, yielding the optimal scale of the network. Finally, we design a new adaptive loss function that is treated as a variable loss computational process for the training of the proposed SP-GAN model. By design, the hyperparameters of the loss function can dynamically adapt to different training stages. Experimental results obtained on a set of data sets demonstrate the merits of the proposed method, especially in terms of the stability and efficiency of network training. The source code of the proposed SP-GAN method is publicly available at https://github.com/Lambert-chen/SPGAN.git.
Xiaoning Song, Zhenhua Feng 0001, Guosheng Hu, Dongjun Yu, Xiaojun Wu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2020 A Flatter Loss for Bias Mitigation in Cross-dataset Facial Age Estimation
abstract
The most existing studies in the facial age estimation assume training and test images are captured under similar shooting conditions. However, this is rarely valid in real-worlds applications, where training and test sets usually have different characteristics. In this paper, we advocate a cross-dataset protocol for age estimation benchmarking. In order to improve the cross-dataset age estimation performance, we mitigate the inherent bias caused by the learning algorithm itself. To this end, we propose a novel loss function that is more effective for neural network training. The relative smoothness of the proposed loss function is its advantage with regards to the optimisation process performed by stochastic gradient descent (SGD). Compared with existing loss functions, the lower gradient of the proposed loss function leads to the convergence of SGD to a better optimum point, and consequently a better generalisation. The cross-dataset experimental results demonstrate the superiority of the proposed method over the state-of-the-art algorithms in terms of accuracy and generalisation capability.
Ali Akbari 0003, Muhammad Awais 0001, Zhenhua Feng 0001, Ammarah Farooq, Josef Kittler
ICPR3
2020 Subspace Clustering via Joint Unsupervised Feature Selection
abstract
Any high-dimensional data arising from practical applications usually contains irrelevant features that may impact on the performance of existing subspace clustering methods. This paper proposes a novel subspace clustering method which reconstructs the feature matrix by the means of unsupervised feature selection (UFS) to achieve a better dictionary for subspace clustering (SC). Different from most existing clustering methods, the proposed approach uses the reconstructed feature matrix as the dictionary rather than the original data matrix. As the feature matrix reconstructed by representative features is more discriminative and closer to the ground-truth, it results in improved performance. The corresponding non-convex optimization problem is effectively solved using the half-quadratic and augmented Lagrange multiplier methods. Extensive experiments on four real datasets demonstrate the effectiveness of the proposed method.
Wenhua Dong, Xiaojun Wu 0001, Hui Li 0037, Zhenhua Feng 0001, Josef Kittler
ICPR4
2020 Adaptive Context-Aware Discriminative Correlation Filters for Robust Visual Object Tracking
abstract
In recent years, Discriminative Correlation Filters (DCFs) have gained popularity due to their superior performance in visual object tracking. However, existing DCF trackers usually learn filters using fixed attention mechanisms that focus on the centre of an image and suppresses filter amplitudes in surroundings. In this paper, we propose an Adaptive Context-Aware Discriminative Correlation Filter (ACA-DCF) that is able to improve the existing DCF formulation with complementary attention mechanisms. Our ACA-DCF integrates foreground attention and background attention for complementary context-aware filter learning. More importantly, we ameliorate the design using an adaptive weighting strategy that takes complex appearance variations into account. The experimental results obtained on several well-known benchmarks demonstrate the effectiveness and superiority of the proposed method over the state-of-the-art approaches.
Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
ICPR2
2020 Biased Feature Learning for Occlusion Invariant Face Recognition
abstract
To address the challenges posed by unknown occlusions, we propose a Biased Feature Learning (BFL) framework for occlusion-invariant face recognition. We first construct an extended dataset using a multi-scale data augmentation method. For model training, we modify the label loss to adjust the impact of normal and occluded samples. Further, we propose a biased guidance strategy to manipulate the optimization of a network so that the feature embedding space is dominated by non-occluded faces. BFL not only enhances the robustness of a network to unknown occlusions but also maintains or even improves its performance for normal faces. Experimental results demonstrate its superiority as well as the generalization capability with different network architectures and loss functions.
Chang-Bin Shao, Jing Huo, Lei Qi 0001, Zhenhua Feng 0001, Wenbin Li 0006, Chuanqi Dong, Yang Gao 0001
IJCAI4
2020 Rectified Wing Loss for Efficient and Robust Facial Landmark Localisation with Convolutional Neural Networks
abstract
Abstract Efficient and robust facial landmark localisation is crucial for the deployment of real-time face analysis systems. This paper presents a new loss function, namely Rectified Wing (RWing) loss, for regression-based facial landmark localisation with Convolutional Neural Networks (CNNs). We first systemically analyse different loss functions, including L2, L1 and smooth L1. The analysis suggests that the training of a network should pay more attention to small-medium errors. Motivated by this finding, we design a piece-wise loss that amplifies the impact of the samples with small-medium errors. Besides, we rectify the loss function for very small errors to mitigate the impact of inaccuracy of manual annotation. The use of our RWing loss boosts the performance significantly for regression-based CNNs in facial landmarking, especially for lightweight network architectures. To address the problem of under-representation of samples with large pose variations, we propose a simple but effective boosting strategy, referred to as pose-based data balancing. In particular, we deal with the data imbalance problem by duplicating the minority training samples and perturbing them by injecting random image rotation, bounding box translation and other data augmentation strategies. Last, the proposed approach is extended to create a coarse-to-fine framework for robust and efficient landmark localisation. Moreover, the proposed coarse-to-fine framework is able to deal with the small sample size problem effectively. The experimental results obtained on several well-known benchmarking datasets demonstrate the merits of our RWing loss and prove the superiority of the proposed method over the state-of-the-art approaches.
Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001, Xiaojun Wu 0001
Int. J. Comput. Vis.1
2020 Learning a representation with the block-diagonal structure for pattern classification
He-Feng Yin, Xiaojun Wu 0001, Josef Kittler, Zhenhua Feng 0001
Pattern Anal. Appl.4
2020 An accelerated correlation filter tracker
Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
Pattern Recognit.2
2020 Learning Low-Rank and Sparse Discriminative Correlation Filters for Coarse-to-Fine Visual Object Tracking
abstract
Discriminative correlation filter (DCF) has achieved advanced performance in visual object tracking with remarkable efficiency guaranteed by its implementation in the frequency domain. However, the effect of the structural relationship of DCF and object features has not been adequately explored in the context of the filter design. To remedy this deficiency, this paper proposes a Low-rank and Sparse DCF (LSDCF) that improves the relevance of features used by discriminative filters. To be more specific, we extend the classical DCF paradigm from ridge regression to lasso regression, and constrain the estimate to be of low-rank across frames, thus identifying and retaining the informative filters distributed on a low-dimensional manifold. To this end, specific temporal-spatial-channel configurations are adaptively learned to achieve enhanced discrimination and interpretability. In addition, we analyse the complementary characteristics between hand-crafted features and deep features, and propose a coarse-to-fine heuristic tracking strategy to further improve the performance of our LSDCF. Last, the augmented Lagrange multiplier optimisation method is used to achieve efficient optimisation. The experimental results obtained on a number of well-known benchmarking datasets, including OTB2013, OTB50, OTB100, TC128, UAV123, VOT2016 and VOT2018, demonstrate the effectiveness and robustness of the proposed method, delivering outstanding performance compared to the state-of-the-art trackers.
Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
IEEE Trans. Circuits Syst. Video Technol.2
2019 Joint Group Feature Selection and Discriminative Filter Learning for Robust Visual Object Tracking
abstract
We propose a new Group Feature Selection method for Discriminative Correlation Filters (GFS-DCF) based visual object tracking. The key innovation of the proposed method is to perform group feature selection across both channel and spatial dimensions, thus to pinpoint the structural relevance of multi-channel features to the filtering system. In contrast to the widely used spatial regularisation or feature selection methods, to the best of our knowledge, this is the first time that channel selection has been advocated for DCF-based tracking. We demonstrate that our GFS-DCF method is able to significantly improve the performance of a DCF tracker equipped with deep neural network features. In addition, our GFS-DCF enables joint feature selection and filter learning, achieving enhanced discrimination and interpretability of the learned filters. To further improve the performance, we adaptively integrate historical information by constraining filters to be smooth across temporal frames, using an efficient low-rank approximation. By design, specific temporal-spatial-channel configurations are dynamically learned in the tracking process, highlighting the relevant features, and alleviating the performance degrading impact of less discriminative representations and reducing information redundancy. The experimental results obtained on OTB2013, OTB2015, VOT2017, VOT2018 and TrackingNet demonstrate the merits of our GFS-DCF and its superiority over the state-of-the-art trackers. The code is publicly available at \url{https://github.com/XU-TIANYANG/GFS-DCF}.
Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
ICCV2
2019 Collaborative representation based face classification exploiting block weighted LBP and analysis dictionary learning
Xiaoning Song, Youming Chen, Zhenhua Feng 0001, Guosheng Hu, Tao Zhang 0010, Xiaojun Wu 0001
Pattern Recognit.3
2019 Fast SRC using quadratic optimisation in downsized coefficient solution subspace
Xiaoning Song, Guosheng Hu, Jian-Hao Luo, Zhenhua Feng 0001, Dongjun Yu, Xiaojun Wu 0001
Signal Process.4
2019 Mining Hard Augmented Samples for Robust Facial Landmark Localization With CNNs
abstract
Effective data augmentation is crucial for facial landmark localization with convolutional neural networks (CNNs). In this letter, we investigate different data augmentation techniques that can be used to generate sufficient data for training CNN-based facial landmark localization systems. To the best of our knowledge, this is the first study that provides a systematic analysis of different data augmentation techniques in the area. In addition, an online hard augmented example mining (HAEM) strategy is advocated for further performance boosting. We examine the effectiveness of those techniques using a regression-based CNN architecture. The experimental results obtained on the AFLW and COFW datasets demonstrate the importance of data augmentation and the effectiveness of HAEM. The performance achieved using these techniques is superior to the state-of-the-art algorithms.
Zhenhua Feng 0001, Josef Kittler, Xiaojun Wu 0001
IEEE Signal Process. Lett.1
2019 Learning Adaptive Discriminative Correlation Filters via Temporal Consistency Preserving Spatial Feature Selection for Robust Visual Object Tracking
abstract
With efficient appearance learning models, discriminative correlation filter (DCF) has been proven to be very successful in recent video object tracking benchmarks and competitions. However, the existing DCF paradigm suffers from two major issues, i.e., spatial boundary effect and temporal filter degradation. To mitigate these challenges, we propose a new DCF-based tracking method. The key innovations of the proposed method include adaptive spatial feature selection and temporal consistent constraints, with which the new tracker enables joint spatial-temporal filter learning in a lower dimensional discriminative manifold. More specifically, we apply structured spatial sparsity constraints to multi-channel filters. Consequently, the process of learning spatial filters can be approximated by the lasso regularization. To encourage temporal consistency, the filter model is restricted to lie around its historical value and updated locally to preserve the global structure in the manifold. Last, a unified optimization framework is proposed to jointly select temporal consistency preserving spatial features and learn discriminative filters with the augmented Lagrangian method. Qualitative and quantitative evaluations have been conducted on a number of well-known benchmarking datasets such as OTB2013, OTB50, OTB100, Temple-Colour, UAV123, and VOT2018. The experimental results demonstrate the superiority of the proposed method over the state-of-the-art approaches.
Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
IEEE Trans. Image Process.2
2018 Wing Loss for Robust Facial Landmark Localisation With Convolutional Neural Networks
abstract
We present a new loss function, namely Wing loss, for robust facial landmark localisation with Convolutional Neural Networks (CNNs). We first compare and analyse different loss functions including L2, L1 and smooth L1. The analysis of these loss functions suggests that, for the training of a CNN-based localisation model, more attention should be paid to small and medium range errors. To this end, we design a piece-wise loss function. The new loss amplifies the impact of errors from the interval (-w, w) by switching from L1 loss to a modified logarithm function. To address the problem of under-representation of samples with large out-of-plane head rotations in the training set, we propose a simple but effective boosting strategy, referred to as pose-based data balancing. In particular, we deal with the data imbalance problem by duplicating the minority training samples and perturbing them by injecting random image rotation, bounding box translation and other data augmentation approaches. Last, the proposed approach is extended to create a two-stage framework for robust facial landmark localisation. The experimental results obtained on AFLW and 300W demonstrate the merits of the Wing loss function, and prove the superiority of the proposed method over the state-of-the-art approaches.
Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001, Patrik Huber 0001, Xiaojun Wu 0001
CVPR1
2018 Evaluation of Dense 3D Reconstruction from 2D Face Images in the Wild
abstract
This paper investigates the evaluation of dense 3D face reconstruction from a single 2D image in the wild. To this end, we organise a competition that provides a new benchmark dataset that contains 2000 2D facial images of 135 subjects as well as their 3D ground truth face scans. In contrast to previous competitions or challenges, the aim of this new benchmark dataset is to evaluate the accuracy of a 3D dense face reconstruction algorithm using real, accurate and high-resolution 3D ground truth face scans. In addition to the dataset, we provide a standard protocol as well as a Python script for the evaluation. Last, we report the results obtained by three state-of-the-art 3D face reconstruction systems on the new benchmark dataset. The competition is organised along with the 2018 13th IEEE Conference on Automatic Face & Gesture Recognition.
Zhenhua Feng 0001, Patrik Huber 0001, Josef Kittler, Peter J. B. Hancock, Xiaojun Wu 0001, Qijun Zhao, Willem P. Koppen, Matthias Rätsch
FG1
2018 Improve the Spoofing Resistance of Multimodal Verification with Representation-Based Measures
Zengxi Huang, Zhenhua Feng 0001, Josef Kittler, Yiguang Liu
PRCV (3)2
2018 Gaussian mixture 3D morphable face model
abstract
3D Morphable Face Models (3DMM) have been used in pattern recognition for some time now. They have been applied as a basis for 3D face recognition, as well as in an assistive role for 2D face recognition to perform geometric and photometric normalisation of the input image, or in 2D face recognition system training. The statistical distribution underlying 3DMM is Gaussian. However, the single-Gaussian model seems at odds with reality when we consider different cohorts of data, e.g. Black and Chinese faces. Their means are clearly different. This paper introduces the Gaussian Mixture 3DMM (GM-3DMM) which models the global population as a mixture of Gaussian subpopulations, each with its own mean. The proposed GM-3DMM extends the traditional 3DMM naturally, by adopting a shared covariance structure to mitigate small sample estimation problems associated with data in high dimensional spaces. We construct a GM-3DMM, the training of which involves a multiple cohort dataset, SURREY-JNU, comprising 942 3D face scans of people with mixed backgrounds. Experiments in fitting the GM-3DMM to 2D face images to facilitate their geometric and photometric normalisation for pose and illumination invariant face recognition demonstrate the merits of the proposed mixture of Gaussians 3D face model.
Willem P. Koppen, Zhenhua Feng 0001, Josef Kittler, Muhammad Awais 0001, William J. Christmas, Xiaojun Wu 0001, He-Feng Yin
Pattern Recognit.2
2018 Dictionary Integration Using 3D Morphable Face Models for Pose-Invariant Collaborative-Representation-Based Classification
abstract
The paper presents a dictionary integration algorithm using 3D morphable face models (3DMM) for pose-invariant collaborative-representation-based face classification. To this end, we first fit a 3DMM to the 2D face images of a dictionary to reconstruct the 3D shape and texture of each image. The 3D faces are used to render a number of virtual 2D face images with arbitrary pose variations to augment the training data, by merging the original and rendered virtual samples to create an extended dictionary. Second, to reduce the information redundancy of the extended dictionary and improve the sparsity of reconstruction coefficient vectors using collaborative-representation-based classification (CRC), we exploit an on-line class elimination scheme to optimise the extended dictionary by identifying the training samples of the most representative classes for a given query. The final goal is to perform pose-invariant face classification using the proposed dictionary integration method and the on-line pruning strategy under the CRC framework. Experimental results obtained for a set of well-known face data sets demonstrate the merits of the proposed method, especially its robustness to pose variations.
Xiaoning Song, Zhenhua Feng 0001, Guosheng Hu, Josef Kittler, Xiaojun Wu 0001
IEEE Trans. Inf. Forensics Secur.2
2017 Dynamic Attention-Controlled Cascaded Shape Regression Exploiting Training Data Augmentation and Fuzzy-Set Sample Weighting
abstract
We present a new Cascaded Shape Regression (CSR) architecture, namely Dynamic Attention-Controlled CSR (DAC-CSR), for robust facial landmark detection on unconstrained faces. Our DAC-CSR divides facial landmark detection into three cascaded sub-tasks: face bounding box refinement, general CSR and attention-controlled CSR. The first two stages refine initial face bounding boxes and output intermediate facial landmarks. Then, an online dynamic model selection method is used to choose appropriate domain-specific CSRs for further landmark refinement. The key innovation of our DAC-CSR is the fault-tolerant mechanism, using fuzzy set sample weighting, for attention-controlled domain-specific model training. Moreover, we advocate data augmentation with a simple but effective 2D profile face generator, and context-aware feature extraction for better facial feature representation. Experimental results obtained on challenging datasets demonstrate the merits of our DAC-CSR over the state-of-the-art methods.
Zhenhua Feng 0001, Josef Kittler, William J. Christmas, Patrik Huber 0001, Xiaojun Wu 0001
CVPR1
2017 Dynamic dictionary optimization for sparse-representation-based face classification using local difference images
Chang-Bin Shao, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001, Yuhui Zheng
Inf. Sci.3
2017 Efficient 3D morphable face model fitting
Guosheng Hu, Fei Yan 0001, Josef Kittler, William J. Christmas, Chi-Ho Chan, Zhenhua Feng 0001, Patrik Huber 0001
Pattern Recognit.6
2017 Half-Face Dictionary Integration for Representation-Based Classification
abstract
This paper presents a half-face dictionary integration (HFDI) algorithm for representation-based classification. The proposed HFDI algorithm measures residuals between an input signal and the reconstructed one, using both the original and the synthesized dual-column (row) half-face training samples. More specifically, we first generate a set of virtual half-face samples for the purpose of training data augmentation. The aim is to obtain high-fidelity collaborative representation of a test sample. In this half-face integrated dictionary, each original training vector is replaced by an integrated dual-column (row) half-face matrix. Second, to reduce the redundancy between the original dictionary and the extended half-face dictionary, we propose an elimination strategy to gain the most robust training atoms. The last contribution of the proposed HFDI method is the use of a competitive fusion method weighting the reconstruction residuals from different dictionaries for robust face classification. Experimental results obtained from the Facial Recognition Technology, Aleix and Robert, Georgia Tech, ORL, and Carnegie Mellon University-pose, illumination and expression data sets demonstrate the effectiveness of the proposed method, especially in the case of the small sample size problem.
Xiaoning Song, Zhenhua Feng 0001, Guosheng Hu, Xiaojun Wu 0001
IEEE Trans. Cybern.2
2016 Towards multi-scale fuzzy sparse discriminant analysis using local third-order tensor model of face images
Xiaoning Song, Zhenhua Feng 0001, Xibei Yang, Xiaojun Wu 0001, Jing-Yu Yang 0001
Neurocomputing2
2015 Fitting 3D Morphable Face Models using local features
abstract
In this paper, we propose a novel fitting method that uses local image features to fit a 3D Morphable Face Model to 2D images. To overcome the obstacle of optimising a cost function that contains a non-differentiable feature extraction operator, we use a learning-based cascaded regression method that learns the gradient direction from data. The method allows to simultaneously solve for shape and pose parameters. Our method is thoroughly evaluated on Morphable Model generated data and first results on real data are presented. Compared to traditional fitting methods, which use simple raw features like pixel colour or edge maps, local features have been shown to be much more robust against variations in imaging conditions. Our approach is unique in that we are the first to use local features to fit a 3D Morphable Model. Because of the speed of our method, it is applicable for realtime applications. Our cascaded regression framework is available as an open source library at github.com/patrikhuber/superviseddescent.
Patrik Huber 0001, Zhenhua Feng 0001, William J. Christmas, Josef Kittler, Matthias Rätsch
ICIP2
2015 Random Cascaded-Regression Copse for Robust Facial Landmark Detection
abstract
In this letter, we present a random cascaded-regression copse (R-CR-C) for robust facial landmark detection. Its key innovations include a new parallel cascade structure design, and an adaptive scheme for scale-invariant shape update and local feature extraction. Evaluation on two challenging benchmarks shows the superiority of the proposed algorithm to state-of-the-art methods.
Zhenhua Feng 0001, Patrik Huber 0001, Josef Kittler, William J. Christmas, Xiaojun Wu 0001
IEEE Signal Process. Lett.1
2015 Cascaded Collaborative Regression for Robust Facial Landmark Detection Trained Using a Mixture of Synthetic and Real Images With Dynamic Weighting
abstract
A large amount of training data is usually crucial for successful supervised learning. However, the task of providing training samples is often time-consuming, involving a considerable amount of tedious manual work. In addition, the amount of training data available is often limited. As an alternative, in this paper, we discuss how best to augment the available data for the application of automatic facial landmark detection. We propose the use of a 3D morphable face model to generate synthesized faces for a regression-based detector training. Benefiting from the large synthetic training data, the learned detector is shown to exhibit a better capability to detect the landmarks of a face with pose variations. Furthermore, the synthesized training data set provides accurate and consistent landmarks automatically as compared to the landmarks annotated manually, especially for occluded facial parts. The synthetic data and real data are from different domains; hence the detector trained using only synthesized faces does not generalize well to real faces. To deal with this problem, we propose a cascaded collaborative regression algorithm, which generates a cascaded shape updater that has the ability to overcome the difficulties caused by pose variations, as well as achieving better accuracy when applied to real faces. The training is based on a mix of synthetic and real image data with the mixing controlled by a dynamic mixture weighting schedule. Initially, the training uses heavily the synthetic data, as this can model the gross variations between the various poses. As the training proceeds, progressively more of the natural images are incorporated, as these can model finer detail. To improve the performance of the proposed algorithm further, we designed a dynamic multi-scale local feature extraction method, which captures more informative local features for detector training. An extensive evaluation on both controlled and uncontrolled face data sets demonstrates the merit of the proposed algorithm.
Zhenhua Feng 0001, Guosheng Hu, Josef Kittler, William J. Christmas, Xiaojun Wu 0001
IEEE Trans. Image Process.1
2012 Automatic face annotation by multilinear AAM with Missing Values
Zhenhua Feng 0001, Josef Kittler, William J. Christmas, Xiaojun Wu 0001, Sebastian Pfeiffer
ICPR1