Tianyang Xu 0001

dblp:169/4627 · DBLP profile ↗
← Back
140ranked-venue papers
13as first author
134since 2021 · last 2026
0000-0002-9015-3128ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 84 · 9 first-author · 80 since 2021Graphics, computer vision, multimedia, augmented reality and games · 69 · 6 first-author · 64 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Bidirectional Channel-selective Semantic Interaction for Semi-Supervised Medical Segmentation
abstract
Semi-supervised medical image segmentation is an effective method for addressing scenarios with limited labeled data. Existing methods mainly rely on frameworks such as mean teacher and dual-stream consistency learning. These approaches often face issues like error accumulation and model structural complexity, while also neglecting the interaction between labeled and unlabeled data streams. To overcome these challenges, we propose a Bidirectional Channel-selective Semantic Interaction (BCSI) framework for semi-supervised medical image segmentation. First, we propose a Semantic-Spatial Perturbation (SSP) mechanism, which disturbs the data using two strong augmentation operations and leverages unsupervised learning with pseudo-labels from weak augmentations. Additionally, we employ consistency on the predictions from the two strong augmentations to further improve model stability and robustness. Second, to reduce noise during the interaction between labeled and unlabeled data, we propose a Channel-selective Router (CR) component, which dynamically selects the most relevant channels for information exchange. This mechanism ensures that only highly relevant features are activated, minimizing unnecessary interference. Finally, the Bidirectional Channel-wise Interaction (BCI) strategy is employed to supplement additional semantic information and enhance the representation of important channels. Experimental results on multiple benchmarking 3D medical datasets demonstrate that the proposed method outperforms existing semi-supervised approaches.
Kaiwen Huang 0002, Yizhe Zhang 0001, Yi Zhou 0007, Tianyang Xu 0001, Tao Zhou 0002
AAAI4
2026 Learning Topology-Driven Multi-Subspace Fusion for Grassmannian Deep Networks
abstract
Grassmannian manifolds offer a powerful carrier for geometric representation learning by modelling high-dimensional data as low-dimensional subspaces. However, existing approaches predominantly rely on static single-subspace representations, neglecting the dynamic interplay between multiple subspaces critical for capturing complex geometric structures. To address this limitation, we propose a topology-driven multi-subspace fusion framework that enables adaptive subspace collaboration on the Grassmannian. Our solution introduces two key innovations: (1) an adaptive multi-subspace construction mechanism that dynamically selects and weights task-relevant subspaces via topological convergence analysis, and (2) a multi-subspace interaction block that fuses heterogeneous geometric representations through Fréchet mean optimisation on the manifold. Theoretically, we establish the convergence guarantees of adaptive subspaces under a projection metric topology, ensuring stable gradient-based optimisation. Practically, we integrate Riemannian batch normalisation and mutual information regularisation to enhance discriminability and robustness. Extensive experiments on 3D action recognition (HDM05, FPHA), EEG classification (MAMEM-SSVEPII), and graph tasks demonstrate state-of-the-art performance. Our work not only advances geometric deep learning but also successfully adapts the proven multi-channel interaction philosophy of Euclidean networks to non-Euclidean domains, achieving superior discriminability and interpretability.
Tianyang Xu 0001
AAAI2
2026 Coarse-to-fine dual flexible-competition hybrid collaborative-nonnegative representation method for image classification
Zi-Qi Li, Xiaoning Song, Tianyang Xu 0001
Expert Syst. Appl.6
2026 A Color Information Driven Collaborative Training of Dual Task Parallel Network for Visible and Thermal Infrared Image Fusion and Saliency Object Detection
Zeyang Zhang 0002, Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Muhammad Awais 0001, Josef Kittler
Int. J. Comput. Vis.3
2026 Attack Intensity is Target-Related: Exploration of Sparse Adversarial Attack against Visual Object Trackers
Shao-Chuan Zhao, Tianyang Xu 0001, Zhangyong Tang, Xiaojun Wu 0001, Josef Kittler
Int. J. Comput. Vis.2
2026 Positive and negative neighbor dual-flexible nonnegative representation method for image classification
Xiaoning Song, Tianyang Xu 0001
Inf. Process. Manag.6
2026 Hausdorff-weighted contrastive fusion for multi-view clustering
Jun Sun 0008, Tianyang Xu 0001
Knowl. Based Syst.5
2026 Interactive image-to-video transfer learning
Cong Wu 0006, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
Neural Networks2
2026 LoongTrack: Exploring long-sequence modeling for visual tracking
Tianyang Xu 0001, Mu Nie, Wankou Yang
Neural Networks2
2026 TATrack: Target-oriented adaptive vision transformer for UAV tracking
Tianyang Xu 0001, Wankou Yang
Neural Networks2
2026 SPD-Updater: Symmetric positive definite manifold geometry based temporal updating for visual object tracking
abstract
Visual object tracking has witnessed continuous advances in recent years along with the exciting developments in backbone networks. In general, all advanced solutions adhere to the template-based tracking framework, which exhibits powerful representative capacity gained via offline training. However, when the target undergoes appearance changes or occlusion, the tracker, which relies on a fixed template defined in the initial frame, struggles to locate it accurately in such complex situations. To achieve online adaptation, recent studies have introduced dynamic templates. Typically, the adopted solution is to compute reliability scores in the traditional Euclidean space to assess the confidence of the dynamic template. However, the Euclidean metric is unreliable to some extent in high-dimensional feature spaces, potentially resulting in a negative impact by involving incorrect dynamic templates. To overcome this problem, we exploit the compact geometric representation capacity of the Symmetric Positive Definite (SPD) manifold to design a novel score prediction module for the tracker update (SPD-Updater). By switching to an SPD manifold metric, we obtain a more accurate and stable dynamic template, thereby enhancing the model capacity to handle complex situations. To validate the reliability of manifold metric in tracking models, we conduct experiments with trackers using different backbones. The experimental results on LaSOT, GOT-10k, TrackingNet, and UAV123 demonstrate the effectiveness of our approach, reflecting the merit of the SPD metric in online tracking adaptation.
Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001, Josef Kittler
Neural Networks2
2026 EvaNet: Toward More Efficient and Consistent Infrared and Visible Image Fusion Assessment
abstract
Evaluation is essential in image fusion research, yet most existing metrics are directly borrowed from other vision tasks without proper adaptation. These traditional metrics, often based on complex image transformations, not only fail to capture the true quality of the fusion results but also are computationally demanding. To address these issues, we propose a unified evaluation framework specifically tailored for image fusion. At its core is a lightweight network designed efficiently to approximate widely used metrics, following a divide-and-conquer strategy. Unlike conventional approaches that directly assess similarity between fused and source images, we first decompose the fusion result into infrared and visible components. The evaluation model is then used to measure the degree of information preservation in these separated components, effectively disentangling the fusion evaluation process. During training, we incorporate a contrastive learning strategy and inform our evaluation model by perceptual scene assessment provided by a large language model. Last, we propose the first consistency evaluation framework, which measures the alignment between image fusion metrics and human visual perception, using both independent no-reference scores and downstream tasks performance as objective references. Extensive experiments show that our learning-based evaluation paradigm delivers both superior efficiency (up to 1,000 times faster) and greater consistency across a range of standard image fusion benchmarks.
Chunyang Cheng, Tianyang Xu 0001, Xiaojun Wu 0001, Tao Zhou 0002, Hui Li 0037, Zhangyong Tang, Josef Kittler
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Vision Mamba-enhanced Multi-level Context Aggregation Network for polyp segmentation
Jiaye Chen, Tianyang Xu 0001, Rui Wang 0050, Xiaoning Song, Tao Zhou 0002
Pattern Recognit.3
2026 Unfolded ISTA for deep sparse subspace clustering
Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
Pattern Recognit.2
2026 Full-combination contrastive learning for multi-view clustering
Zhe Chen 0018, Heng Liu 0003, Hui Li 0037, Tianyang Xu 0001
Pattern Recognit.5
2026 Switcher: Adaptive framework for unified and customized multi-modal object tracking
He Wang 0028, Tianyang Xu 0001, Zhangyong Tang, Xiaojun Wu 0001, Josef Kittler
Pattern Recognit.2
2026 Harmonising adaptive receptive field with multi-scale and multi-task perception for real-time face detection
He-Feng Yin, Tianyang Xu 0001, Xiaojun Wu 0001
Pattern Recognit.2
2026 Visual complexity guided diffusion defender for video object tracking and recognition
Shao-Chuan Zhao, Tianyang Xu 0001, Hui Li 0037, Xiaojun Wu 0001, Josef Kittler
Pattern Recognit.2
2026 Dual-Reliable Contrastive Fusion for Multi-View Clustering
abstract
Multi-view clustering (MVC) has garnered significant attention in recent years due to its ability to leverage shared information across heterogeneous data sources. However, most existing methods focus on improving clustering performance while often neglecting potential conflicts between views and the uncertainty in view distributions. To address this limitation, we propose a novel framework named dual-reliable contrastive fusion multi-view clustering (DRCFMVC). This framework organically integrates uncertainty and conflict within a multi-view contrastive clustering paradigm for the first time. Through an innovative dual reliability weighting mechanism, uncertainty and conflict are systematically incorporated into contrastive learning. Specifically, high-dimensional features are first mapped to cluster distributions via a clustering network. Subsequently, Dempster–Shafer Theory (DST) estimates prediction uncertainty within each view, while Jensen–Shannon (JS) divergence quantifies conflict levels between views. These two complementary types of information are further integrated through the proposed dual-reliable weighting (DRW) fusion strategy, effectively guiding the contrastive optimization process to learn more stable and discriminative representations. Extensive experiments on eleven public benchmark datasets demonstrate that the proposed method achieves superior performance across multiple clustering evaluation metrics compared to existing state-of-the-art approaches. The source code can be made accessible at https://github.com/li-zi-qi/DRCFMVC.
Tianyang Xu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Adaptive Continual Learning for Online Visual Object Tracking via Dynamic Grassmannian Appearance Modeling
abstract
Online visual object tracking fundamentally constitutes a continual learning challenge, demanding persistent adaptation to target variations within dynamic video streams, while preserving critical features to prevent forgetting. Existing template-based trackers are particularly vulnerable to severe target deformation and drastic background changes in real-world scenarios. Recent explorations enhance adaptability through local update strategies—such as dynamic templates, appearance tokens, and parameter fine-tuning—to address rapid appearance variations. However, these methods inherently propagate target appearance changes without explicit modelling and prioritise short-term adaptation over long-term global representation. Consequently, they fail to balance initial target features with current observations, leading to tracking failure in long-term scenarios. To overcome these limitations, we propose DG-Track, an adaptive continual learning framework leveraging Grassmannian manifold geometry. Specifically, we represent target appearance within a Grassmannian affine subspace and perform continual adaptation via incremental learning. Compared to Euclidean geometry, the Grassmannian manifold captures nonlinear appearance representations, yielding more compact and geometrically consistent temporal dynamics modelling. Furthermore, an adaptive forgetting module dynamically regulates the interplay between the current observed subspace and the initial template, ensuring stable long-term tracking. DG-Track is a plug-and-play solution for online tracking, adding no learnable parameters. Comprehensive experiments with diverse baseline trackers on LaSOT, GOT-10k, TrackingNet, and UAV123 validate the efficacy of our continual learning framework and Grassmannian manifold geometry in enhancing visual object tracking performance. Code is available at https://github.com/xiaoqing0825/DGTrack.
Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001, Josef Kittler
IEEE Trans. Circuits Syst. Video Technol.2
2026 BusReF: Infrared-Visible Images Registration and Fusion Focus on Reconstructible Area Using One Set of Features
abstract
In multi-modal imaging scenarios, the misalignment of images presents a persistent challenge. Conventional image fusion algorithms, aiming to enhance the performance of downstream vision tasks, presuppose strictly registered inputs to achieve satisfactory results. To relax this assumption, a common approach is to register the images first; however, existing multi-modal registration methods are often hindered by complex architectures and a heavy reliance on semantic information. This article proposes BusRef, a unified framework that jointly addresses image registration and fusion, with a specific focus on the Infrared-Visible Image Registration and Fusion (IVRF) task. Within this framework, unaligned image pairs are processed through three sequential stages: coarse registration, fine registration, and fusion. We demonstrate that this integrated approach enables more robust and accurate IVRF. Key to our framework is a novel training and evaluation strategy that employs masks to mitigate the influence of non-reconstructible regions on the loss function, thereby significantly improving the model’s accuracy and robustness. Furthermore, we introduce a gradient-aware fusion network designed to effectively preserve complementary information from both modalities. Comprehensive experiments demonstrate that BusRef achieves superior performance when compared against various state-of-the-art registration and fusion algorithms. Our code is available at https://github.com/Yukarizz/BusReF .
Zeyang Zhang 0002, Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Congcong Bian, Josef Kittler
ACM Trans. Multim. Comput. Commun. Appl.3
2025 R-DTI: Drug Target Interaction Prediction Based on Second-Order Relevance Exploration
abstract
Drug Target Interaction (DTI) prediction has witnessed promising performance boosts accompanied by advanced multimodal feature extraction. However, existing approaches suffer from two main difficulties. First, the complex protein structures cannot be well represented by current protein-sequence-based feature extractors. Second, the gap between protein and drug features increases the vulnerability of the obtained classifier thus degrading the prediction robustness. To address these issues, we propose a novel R-DTI method by exploring the second-order relevance in both protein structural feature extraction and DTI prediction phases. Specifically, we construct a pre-trained structural feature extractor that mines the atomic relevance of each amino acid. Then, an inter-feature structure-preserved Riemannian network is designed to expand the existing protein extraction patterns. To improve the prediction robustness, we also develop a Riemannian classifier that uses the second-order protein-drug relevance with a unified feature space. Extensive experimental results demonstrate the merits and superiority of our R-DTI against the state-of-the-art, achieving 1.4% and 1.9% higher AUC-ROC on the BindingDB and DrugBank datasets, respectively.
Yang Hua 0002, Tianyang Xu 0001, Xiaoning Song, Zhenhua Feng 0001, Rui Wang 0050, Wenjie Zhang 0009, Xiaojun Wu 0001
AAAI2
2025 PETS2025: Multi-Authority Multi-Sensor Maritime Surveillance Challenge and Evaluation
abstract
This paper presents the outcomes of the PETS2025 challenge, held in conjunction with AVSS 2025 and sponsored by the EU-funded EURMARS project. The challenge introduces a novel maritime surveillance dataset comprising image sequences captured by diverse multi-altitude, multimodal sensors, reflecting the real-world multi-authority environment. The key tasks include: (1) object detection using various sensors across different platforms (ground-based and low-altitude aerial) and spectral ranges (visible, thermal, ultraviolet (UV), and short-wave infrared (SWIR)); (2) long-term tracking of targets in maritime environments spanning both sea and land; and (3) approximating target geolocations by using sensor imagery and telemetry data. Performance evaluations of results submitted by 12 international participants are discussed. The results show the effectiveness of these submissions and highlight ongoing challenges posed by heterogeneous sensors and complex environments. These challenges emphasise the need to further improve detection, tracking, and geolocation approximation for maritime and coastal surveillance.
Thanet Markchom, Jonathan N. Boyle, Lulu Chen, James M. Ferryman, Matteo Marturini, Stephan Veigl, Andreas Opitz, Andreas Kriechbaum-Zabini, Romaios Bratskas, Anastasios Gkamaris, Dimitris Papachristos, George Leventakis, Wenjun Fan, Hsiang-Wei Huang, Jeng-Neng Hwang, Pyong-Kun Kim, Kwangju Kim, Chung-I Huang, Kenta Saito, Shunta Kaneko, Kyoko Sudo, Nguyen Thanh Thien, Meng-Yu Kao, Jun-Wei Hsieh, Teepakorn Lilek, Tossapol Pomsuwan, Jinjie Gu, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001, Josef Kittler, Stephanie Stacy, Alfredo Gabaldon, Peter Tu, Dongyoung Kim, Kyoungoh Lee
AVSS28
2025 One Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image Fusion
abstract
Advanced image fusion methods mostly prioritise high-level missions, where task interaction struggles with semantic gaps, requiring complex bridging mechanisms. In contrast, we propose to leverage low-level vision tasks from digital photography fusion, allowing for effective feature interaction through pixel-level supervision. This new paradigm provides strong guidance for unsupervised multimodal fusion without relying on abstract semantics, enhancing task-shared feature learning for broader applicability. Owning to the hybrid image features and enhanced universal representations, the proposed GIFNet supports diverse fusion tasks, achieving high performance across both seen and unseen scenarios with a single model. Uniquely, experimental results reveal that our framework also supports single-modality enhancement, offering superior flexibility for practical applications. Our code will be available at https://github.com/AWCXV/GIFNet.
Chunyang Cheng, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Zhangyong Tang, Hui Li 0037, Zeyang Zhang 0002, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler
CVPR2
2025 Adaptive Hyper-Graph Convolution Network for Skeleton-Based Human Action Recognition with Virtual Connections
abstract
The shared topology of human skeletons motivated the recent investigation of graph convolutional network (GCN) solutions for action recognition. However, most of the existing GCNs rely on the binary connection of two neighboring vertices (joints) formed by an edge (bone), overlooking the potential of constructing multi-vertex convolution structures. Although some studies have attempted to utilize hyper-graphs to represent the topology, they rely on a fixed construction strategy, which limits their adaptivity in uncovering the intricate latent relationships within the action. In this paper, we address this oversight and explore the merits of an adaptive hyper-graph convolutional network (Hyper-GCN) to achieve the aggregation of rich semantic information conveyed by skeleton vertices. In particular, our Hyper-GCN adaptively optimises the hyper-graphs during training, revealing the action-driven multi-vertex relations. Besides, virtual connections are often designed to support efficient feature aggregation, implicitly extending the spectrum of dependencies within the skeleton. By injecting virtual connections into hyper-graphs, the semantic clues of diverse action categories can be highlighted. The results of experiments conducted on the NTU-60, NTU-120, and NW-UCLA datasets demonstrate the merits of our Hyper-GCN, compared to the state-of-the-art methods. The code is available at https://github.com/6UOOON9/Hyper-GCN.
Youwei Zhou, Tianyang Xu 0001, Cong Wu 0006, Xiaojun Wu 0001, Josef Kittler
ICCV2
2025 Hybrid Batch Normalisation: Resolving the Dilemma of Batch Normalisation in Federated Learning
abstract
Batch Normalisation (BN) is widely used in conventional deep neural network training to harmonise the input-output distributions for each batch of data. However, federated learning, a distributed learning paradigm, faces the challenge of dealing with non-independent and identically distributed data among the client nodes. Due to the lack of a coherent methodology for updating BN statistical parameters, standard BN degrades the federated learning performance. To this end, it is urgent to explore an alternative normalisation solution for federated learning. In this work, we resolve the dilemma of the BN layer in federated learning by developing a customised normalisation approach, Hybrid Batch Normalisation (HBN). HBN separates the update of statistical parameters (*i.e.*, means and variances used for evaluation) from that of learnable parameters (*i.e.*, parameters that require gradient updates), obtaining unbiased estimates of global statistical parameters in distributed scenarios. In contrast with the existing solutions, we emphasise the supportive power of global statistics for federated learning. The HBN layer introduces a learnable hybrid distribution factor, allowing each computing node to adaptively mix the statistical parameters of the current batch with the global statistics. Our HBN can serve as a powerful plugin to advance federated learning performance. It reflects promising merits across a wide range of federated learning settings, especially for small batch sizes and heterogeneous data. Code is available at https://github.com/Hongyao-Chen/HybridBN.
Hongyao Chen, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
ICML2
2025 SymCL: Riemannian Contrastive Learning on the Symmetric Positive Definite Manifold for Visual Classification
abstract
Symmetric Positive Definite (SPD) matric has been proven to be an effective feature descriptor in the realm of artificial intelligence, as it can encode spatiotemporal statistical information of data on a curved Riemannian manifold, i.e., SPD manifold. Although existing Riemannian neural networks have demonstrated superiority in many scientific fields, the inherent reliance on labels within supervised learning renders them susceptible to label errors. Besides, it is insufficient to depend solely on labels to learn effective feature distributions in some complicated data scenarios. Drawing inspiration from the considerable achievements of contrastive learning (CL) across diverse tasks, we extend the conventional CL paradigm to the context of SPD manifolds, which we denote SymCL, paving the way for a novel approach in SPD matrix-based visual classification. Furthermore, we inject a Riemannian triplet loss-based Riemannian metric learning (RML) into the designed SPD manifold CL framework for the sake of improving the discrimination of the learned geometric representations. Extensive experimental results on four datasets verify the effectiveness of the proposed algorithm.
Yusheng Bao, Rui Wang 0050, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
IJCNN3
2025 Learning a Discriminative Grassmannian Neural Network for Visual Classification
abstract
Learning representations on the Grassmannian manifold is popular in quite a few visual classification tasks. With the development of deep learning techniques, several neural networks have recently emerged for processing subspace data. However, the diversely changed appearance of the signal data (video clips and image sets), makes it impossible for the existing Grassmannian networks (GrasNets) that rely on a single cross-entropy loss for end-to-end training to learn effective geometric representations, especially for complicated visual scenarios. To solve this problem, a Riemannian triplet loss-based Riemannian metric learning mechanism is introduced to the original GrasNet, which can explicitly encode and learn the characteristics of the intra- and inter-class data distributions conveyed by the input data during network training. Additionally, given the existence of intra-class diversity and inter-class ambiguity of the input data, we propose a hard sample reward strategy (HSR) to further improve the discriminability of the learned network embedding. Extensive experimental results obtained on four benchmarking datasets demonstrate the effectiveness of the proposed method.
Yusheng Bao, Rui Wang 0050, Tianyang Xu 0001, Xiaojun Wu 0001, Umapada Pal 0001, Josef Kittler
IJCNN3
2025 Serial Over Parallel: Learning Continual Unification for Multi-Modal Visual Object Tracking and Benchmarking
abstract
Unifying multiple multi-modal visual object tracking (MMVOT) tasks draws increasing attention due to the complementary nature of different modalities in building robust tracking systems. Existing practices mix all data sensor types in a single training procedure, structuring a parallel paradigm from the data-centric perspective and aiming for a global optimum on the joint distribution of the involved tasks. However, the absence of a unified benchmark where all types of data coexist forces evaluations on separated benchmarks, causing inconsistency between training and testing, thus leading to performance degradation. To address these issues, this work advances in two aspects: A unified benchmark, coined as UniBench300, is introduced to bridge the inconsistency by incorporating multiple task data, reducing inference passes from three to one and cutting time consumption by 27%. The unification process is reformulated in a serial format, progressively integrating new tasks. In this way, the performance degradation can be specified as knowledge forgetting of previous tasks, which naturally aligns with the philosophy of continual learning (CL), motivating further exploration of injecting CL into the unification process. Extensive experiments conducted on two baselines and four benchmarks demonstrate the significance of UniBench300 and the superiority of CL in supporting a stable unification process. Moreover, while conducting dedicated analyses, the performance degradation is found to be negatively correlated with network capacity. Additionally, modality discrepancies contribute to varying degradation levels across tasks (RGBT > RGBD > RGBE in MMVOT), offering valuable insights for future multi-modal vision research. Source codes and the proposed benchmark is available at https://github.com/Zhangyong-Tang/UniBench300.
Zhangyong Tang, Tianyang Xu 0001, Xuefeng Zhu 0003, Chunyang Cheng, Tao Zhou 0002, Xiaojun Wu 0001, Josef Kittler
ACM Multimedia2
2025 Collaborating Vision, Depth, and Thermal Signals for Multi-Modal Tracking: Dataset and Algorithm
abstract
Existing multi-modal object tracking approaches primarily focus on dual-modal paradigms, such as RGB-Depth or RGB-Thermal, yet remain challenged in complex scenarios due to limited input modalities. To address this gap, this work introduces a novel multi-modal tracking task that leverages three complementary modalities, including visible RGB, Depth (D), and Thermal Infrared (TIR), aiming to enhance robustness in complex scenarios. To support this task, we construct a new multi-modal tracking dataset, coined RGBDT500, which consists of 500 videos with synchronised frames across the three modalities. Each frame provides spatially aligned RGB, depth, and thermal infrared images with precise object bounding box annotations.Furthermore, we propose a novel multi-modal tracker, dubbed RDTTrack.RDTTrack integrates tri-modal information for robust tracking by leveraging a pretrained RGB-only tracking model and prompt learning techniques.In specific, RDTTrack fuses thermal infrared and depth modalities under a proposed orthogonal projection constraint, then integrates them with RGB signals as prompts for the pre-trained foundation tracking model, effectively harmonising tri-modal complementary cues.The experimental results demonstrate the effectiveness and advantages of the proposed method, showing significant improvements over existing dual-modal approaches in terms of tracking accuracy and robustness in complex scenarios. The dataset and source code are publicly available at https://xuefeng-zhu5.github.io/RGBDT500.
Xuefeng Zhu 0003, Tianyang Xu 0001, Yifan Pan, Jinjie Gu, Xi Li 0001, Jiwen Lu, Xiaojun Wu 0001, Josef Kittler
NeurIPS2
2025 Cross-Modal Supervised Contrastive Learning for RGB-T Semantic Segmentation
Chuanjiang Zhang, Tianyang Xu 0001, Zhangyong Tang, Xiaojun Wu 0001
PRCV (5)2
2025 Fastere: a fast framework for entity relation extractions
Wenjie Zhang 0009, Tianyang Xu 0001, Yang Hua 0002, Zhenhua Feng 0001, Xiaoning Song
Data Min. Knowl. Discov.2
2025 Learning adaptive detection and tracking collaborations with augmented UAV synthesis for accurate anti-UAV system
Shihan Liu, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001, Josef Kittler
Expert Syst. Appl.2
2025 Towards fine-grained adaptive video captioning via Quality-Aware Recurrent Feedback Network
Tianyang Xu 0001, Xiaoning Song, Xiaojun Wu 0001
Expert Syst. Appl.1
2025 FusionBooster: A Unified Image Fusion Boosting Paradigm
Chunyang Cheng, Tianyang Xu 0001, Xiaojun Wu 0001, Hui Li 0037, Xi Li 0001, Josef Kittler
Int. J. Comput. Vis.2
2025 SMLNet: A SPD Manifold Learning Network for Infrared and Visible Image Fusion
Huan Kang, Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Rui Wang 0050, Chunyang Cheng, Josef Kittler
Int. J. Comput. Vis.3
2025 Learning Structure-Supporting Dependencies via Keypoint Interactive Transformer for General Mammal Pose Estimation
Tianyang Xu 0001, Jiyong Rao, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001
Int. J. Comput. Vis.1
2025 Dual attention focus network for few-shot skeleton-based action recognition
Chongben Tao, Cong Wu 0006, Tianyang Xu 0001, Xizhao Luo, Zufeng Zhang, Sai Xu
Knowl. Based Syst.5
2025 Anchor Graph Learning with Double Noise Removal for Multi-View Clustering
Zhe Chen 0018, Mingzhi Zhu, Hui Li 0037, Tianyang Xu 0001
Neural Networks4
2025 TENet: Targetness entanglement incorporating with multi-scale pooling and mutually-guided fusion for RGB-E object tracking
Pengcheng Shao, Tianyang Xu 0001, Zhangyong Tang, Linze Li 0002, Xiaojun Wu 0001, Josef Kittler
Neural Networks2
2025 Adaptive pooling with dual-stage fusion for skeleton-based action recognition
Cong Wu 0006, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler
Neural Networks3
2025 Temporal aggregation for real-time RGBT tracking via fast decision-level fusion
Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
Pattern Recognit. Lett.2
2025 BCN: Bidirectional Contrastive Learning Net for Multi-View Clustering
abstract
Contrastive learning for deep multi-view clustering aims to learn discriminative representations across multiple views. However, prevailing cluster-level alignment approaches fail to fully leverage cross-view consistency and complementarity, as they neglect instance-level semantic coherence. To address this limitation, we propose a novel bidirectional contrastive learning network for multi-view clustering. By simultaneously contrasting the inter-view semantic label matrix along the row and column directions (i.e., at instance-level and cluster-level), the labels of the same instance in different views are consistent, and instances assigned to the same cluster across different views remain consistent. Moreover, we use dual-channel MLPs to avoid information conflicts caused by bidirectional contrastive learning. The proposed framework also demonstrates strong generalization capability, serving as a plug-and-play module that can be seamlessly integrated with existing methods to improve their clustering performance. Extensive experiments on publicly available datasets demonstrate the superiority of our method over several state-of-the-art techniques.
Zhe Chen 0018, Jun Huang 0003, Tianyang Xu 0001, Xiaojun Wu 0001
IEEE Signal Process. Lett.4
2025 M3Track: Meta-Prompt for Multi-Modal Tracking
abstract
Prompt-tuning has shown remarkable success in multi-modal visual tracking, which enhances RGB tracking by incorporating an additional modality,e.g., thermal infrared (T), Depth (D), or Event (E), forming an RGB+X tracking paradigm. However, in the testing phase, current approaches utilise a frozen prompt for the entire benchmark, failing to account for the diversity and unique characteristics of individual videos. To address this issue, we inject meta-learning solutions into current prompt-based tracking technique, thereby emphasising sequence-level adaptation. Unlike traditional prompt-based trackers, which keep the parameters of prompt blocks fixed during testing, our approach updates these parameters in the first frame of each video using a meta-learning solution. This allows for enhanced discriminative tracking capabilities tailored to each video. Building on this advancement, we extend our methodology beyond separate implementations for RGBT, RGBD, and RGBE tasks. A unified multi-modal tracker is further derived, resulting in the first unified tracker without any task priors (notification of task type) employed in both training and testing phases. Extensive experimental results on LasHeR, DepthTrack, VisEvent, GTOT, RGBT234, and RGBD1K consistently demonstrate the superiority of the proposed method against the existing prompt-tuning paradigm. Source codes are available athttps://github.com/Zhangyong-Tang/M3Track.
Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
IEEE Signal Process. Lett.2
2025 Adaptive Colour-Depth Aware Attention for RGB-D Object Tracking
abstract
Recent advances in RGB-D tracking have been driven by the synergistic combination of high-performing RGB-only trackers and auxiliary depth information. However, most existing methods rely on visual feature descriptors to extract depth features, which are then fused with vision features. This pipeline may lead to performance degradation due to the incongruence between the RGB and depth modalities. In this letter, we propose an efficient and effective transformer-based framework, that explicitly models colour and depth information for RGB-D tracking. Specifically, we first statistically code the colour and depth information of the foreground and background for the template. Then, the spatial attention maps of the search region are obtained using these colour-depth statistical models, enhancing the visual features of the search region for improved object localisation accuracy. The comprehensive experimental results obtained on multiple benchmarks demonstrate the effectiveness and merits of the proposed approach in explicit colour-depth coding for RGB-D tracking. The code and the models are publicly accessible athttps://github.com/xuefeng-zhu5/CDAAT.
Xuefeng Zhu 0003, Tianyang Xu 0001, Xiaojun Wu 0001, Zhenhua Feng 0001, Josef Kittler
IEEE Signal Process. Lett.2
2025 Deep Discriminative Multi-View Clustering
abstract
Multi-view clustering based on deep auto-encoder networks has garnered increasing attention and made significant progress in recent years. However, we argue that most existing methods inadequately explore the discriminability while learning clustering assignments, resulting in models struggling to accurately cluster data, particularly those with ambiguous semantics. To address this problem, we propose a novel framework termed deep discriminative multi-view clustering (DDMvC). This framework is designed to further increase the inter-cluster distances by learning a discriminative projection dictionary with global prior information. To begin with, we enhance the reliability of the dictionary atoms by initializing them with class-specific prototypes derived from concatenated global features across multiple views. Subsequently, we iteratively refine the atoms to guarantee their independence from any specific cluster. Simultaneously, we incorporate contrastive learning for the cluster assignments projected by these atoms, striving for inter-view consistent clustering results. Experimental results on benchmark multi-view datasets demonstrate that our framework achieves the state-of-the-art clustering performance.
Zhe Chen 0018, Xiaojun Wu 0001, Tianyang Xu 0001, Hui Li 0037, Josef Kittler
IEEE Trans. Circuits Syst. Video Technol.3
2025 Revisiting RGBT Tracking Benchmarks From the Perspective of Modality Validity: A New Benchmark, Problem, and Solution
abstract
RGBT tracking draws increasing attention because of its robustness in multi-modal warranting (MMW) scenarios, such as nighttime and adverse weather conditions, where relying on a single sensing modality fails to ensure stable tracking results. However, existing benchmarks predominantly contain videos collected in common scenarios where both RGB and thermal infrared (TIR) information are of sufficient quality. This weakens the representativeness of existing benchmarks in severe imaging conditions, leading to tracking failures in MMW scenarios. To bridge this gap, we present a new benchmark considering the modality validity, MV-RGBT, captured specifically from MMW scenarios where either RGB (extreme illumination) or TIR (thermal truncation) modality is invalid. Hence, it is further divided into two subsets according to the valid modality, offering a new compositional perspective for evaluation and providing valuable insights for future designs. Moreover, MV-RGBT is the most diverse benchmark of its kind, featuring 36 different object categories captured across 19 distinct scenes. Furthermore, considering severe imaging conditions in MMW scenarios, a new problem is posed in RGBT tracking, named 'when to fuse', to stimulate the development of fusion strategies for such scenarios. To facilitate its discussion, we propose a new solution with a mixture of experts, named MoETrack, where each expert generates independent tracking results along with a confidence score. Extensive results demonstrate the significant potential of MV-RGBT in advancing RGBT tracking and elicit the conclusion that fusion is not always beneficial, especially in MMW scenarios. Besides, MoETrack achieves state-of-the-art results on several benchmarks, including MV-RGBT, GTOT, and LasHeR. Source codes and benchmarks are available at https://github.com/Zhangyong-Tang/MVRGBT.
Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Xuefeng Zhu 0003, Chunyang Cheng, Zhenhua Feng 0001, Josef Kittler
IEEE Trans. Image Process.2
2025 DFL-Net: Disentangled Feature Learning Network for Multi-View Clustering
abstract
Multi-view clustering aims at partitioning data into their underlying categories by mining shared and complementary information conveyed by different views. Although the integration of deep learning and disentanglement learning has markedly improved clustering performance, our analysis reveals two fundamental limitations in existing approaches: inadequate separation between view-shared and view-exclusive features; and the negative effects of clustering-irrelevant information on feature decoupling. To tackle these issues, we present a novel Disentangled Feature Learning Network (DFL-Net), which utilizes a progressive learning framework to systematically disentangle features. DFL-Net initially establishes view-shared representations through semantic disparity minimization, followed by the construction of orthogonal feature subspaces using cross-view and intra-view independence constraints to isolate view-specific features. Subsequently, DFL-Net enforces clustering consistency across views to adaptively eliminate irrelevant information, thus enhancing the overall effectiveness of disentanglement learning. The framework introduces two significant innovations: a comprehensive feature independence criterion that concurrently reduces intra-view and cross-view feature dependencies, and an irrelevance filtering mechanism that ensures cross-view clustering consistency. Extensive experiments on benchmark datasets demonstrate the superior performance of DFL-Net compared to state-of-the-art methods.
Zhe Chen 0018, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler
IEEE Trans. Knowl. Data Eng.3
2025 I Know How You Move: Explicit Motion Estimation for Human Action Recognition
abstract
Enabled by hierarchical convolutions and nonlinear mappings, recent action recognition studies have continuously boosted performance with spatiotemporal modelling. In general, motion clues are essential in video-oriented tasks, while existing approaches aggregate the spatial and temporal signatures via specially designed modules in the middle or output stages. To highlight the privilege provided by temporal motions, in this paper, we propose a simple but effectiveMOTion Estimator(MOTE) to generate the motion patterns from every single frame, avoiding complex dense-frame input. In particular, MOTE follows an encoder-decoder structure, which takes the short-term motion features generated by the pretrained dense-frame network as the learning target. The spatial information of a single frame is utilized to estimate the instantaneous motion appearance. It can support the expression of vulnerable regions, such as the ‘hand’ in ‘waving hands’, which would otherwise be suppressed in the feature maps as the ‘hand’ suffers from motion blur. The training process of MOTE is independent of the action recognition system. Therefore, the trained MOTE can be transplanted to the input-end of existing action recognition methods to provide instantaneous motion estimation as feature enhancement according to practical requirements. Our experiments performed on Something-Something V1, V2, Kinetics-400, and Diving48 verify the effectiveness of the proposed method.
Xiaojun Wu 0001, Hui Li 0037, Tianyang Xu 0001, Cong Wu 0006
IEEE Trans. Multim.4
2025 ATMNet: Adaptive Two-Stage Modular Network for Accurate Video Captioning
abstract
In recent years, pretrained language-image models (PLIMs) have delivered advances in video captioning. However, existing PLIMs primarily focus on extracting global feature representations from still images and text sequences, while neglecting fine-grained semantic alignment and temporal variations between vision and text pairs. To this end, we propose a global-local alignment module and a temporal parsing module to reflect the detailed correspondence and temporal perception between the two modalities, respectively. In particular, the global-local alignment module enables cross-modal registration at two levels, i.e., the sentence-video level and the word-frame level, to obtain mixed-granularity semantic video features. The temporal parsing module is a dedicated self-attention structure that highlights temporal order cues along video frames, compensating for the limited temporal capacity of PLIMs. In addition, an adaptive two-stage gating structure is designed to leverage the linguistic predictions further. The linguistic information derived from the first stage prediction is dynamically routed through an adaptive decision gate, allowing for quality assessment of whether the information should proceed to the second stage. This structure can effectively reduce the computational burden for easy samples and further improve the accuracy of the prediction results. The experimental results obtained on several benchmark datasets demonstrate the effectiveness of the proposed solution, with improved performance compared to state-of-the-art methods.
Tianyang Xu 0001, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2024 Generative-Based Fusion Mechanism for Multi-Modal Tracking
abstract
Generative models (GMs) have received increasing research interest for their remarkable capacity to achieve comprehensive understanding. However, their potential application in the domain of multi-modal tracking has remained unexplored. In this context, we seek to uncover the potential of harnessing generative techniques to address the critical challenge, information fusion, in multi-modal tracking. In this paper, we delve into two prominent GM techniques, namely, Conditional Generative Adversarial Networks (CGANs) and Diffusion Models (DMs). Different from the standard fusion process where the features from each modality are directly fed into the fusion block, we combine these multi-modal features with random noise in the GM framework, effectively transforming the original training samples into harder instances. This design excels at extracting discriminative clues from the features, enhancing the ultimate tracking performance. Based on this, we conduct extensive experiments across two multi-modal tracking tasks, three baseline methods, and four challenging benchmarks. The experimental results demonstrate that the proposed generative-based fusion mechanism achieves state-of-the-art performance by setting new records on GTOT, LasHeR and RGBD1K. Code will be available at https://github.com/Zhangyong-Tang/GMMT.
Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Xuefeng Zhu 0003, Josef Kittler
AAAI2
2024 SCD-Net: Spatiotemporal Clues Disentanglement Network for Self-Supervised Skeleton-Based Action Recognition
abstract
Contrastive learning has achieved great success in skeleton-based action recognition. However, most existing approaches encode the skeleton sequences as entangled spatiotemporal representations and confine the contrasts to the same level of representation. Instead, this paper introduces a novel contrastive learning framework, namely Spatiotemporal Clues Disentanglement Network (SCD-Net). Specifically, we integrate the decoupling module with a feature extractor to derive explicit clues from spatial and temporal domains respectively. As for the training of SCD-Net, with a constructed global anchor, we encourage the interaction between the anchor and extracted clues. Further, we propose a new masking strategy with structural constraints to strengthen the contextual associations, leveraging the latest development from masked image modelling into the proposed SCD-Net. We conduct extensive evaluations on the NTU-RGB+D (60&120) and PKU-MMD (I&II) datasets, covering various downstream tasks such as action recognition, action retrieval, transfer learning, and semi-supervised learning. The experimental results demonstrate the effectiveness of our method, which outperforms the existing state-of-the-art (SOTA) approaches significantly. Our code and supplementary material can be found at https://github.com/cong-wu/SCD-Net.
Cong Wu 0006, Xiaojun Wu 0001, Josef Kittler, Tianyang Xu 0001, Muhammad Awais 0001, Zhenhua Feng 0001
AAAI4
2024 LabelPrompt: Effective prompt-based learning for relation classification
Wenjie Zhang 0009, Xiaoning Song, Zhenhua Feng 0001, Tianyang Xu 0001, Xiaojun Wu 0001
ACML4
2024 C2C: Component-to-Composition Learning for Zero-Shot Compositional Action Recognition
Rongchang Li 0001, Zhenhua Feng 0001, Tianyang Xu 0001, Linze Li 0002, Xiaojun Wu 0001, Muhammad Awais 0001, Sara Atito Ali Ahmed, Josef Kittler
ECCV (38)3
2024 Efficient Few-Shot Action Recognition via Multi-level Post-reasoning
Cong Wu 0006, Xiaojun Wu 0001, Linze Li 0002, Tianyang Xu 0001, Zhenhua Feng 0001, Josef Kittler
ECCV (3)4
2024 Spatio-Temporal Domain-Aware Network for Skeleton-Based Action Representation Learning
Jiannan Hu, Cong Wu 0006, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
ICPR (29)3
2024 Learning Explicit Modulation Vectors for Disentangled Transformer Attention-Based RGB-D Visual Tracking
Yifan Pan, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaoqing Luo, Xiaojun Wu 0001, Josef Kittler
ICPR (16)2
2024 Infrared and Visible Image Fusion Method Based on Learnable Joint Sparse Low-Rank Decomposition
Wenfeng Song, Naiyun Huang, Xiaoqing Luo, Zhancheng Zhang, Tianyang Xu 0001, Xiaojun Wu 0001
ICPR (5)5
2024 IFFusion: Illumination-Free Fusion Network for Infrared and Visible Images
Hui Li 0037, Tianyang Xu 0001, Zeyang Zhang 0002, Xiaojun Wu 0001
ICPR (5)3
2024 Harmonizing Regression-Classification Inconsistency for Task-Specific Decoupling in Underwater Object Detection
Minrui Xiang, Tianyang Xu 0001, Xiaojun Wu 0001
ICPR (16)2
2024 Attention-Based Patch Matching and Motion-Driven Point Association for Accurate Point Tracking
Han Zang, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaoning Song, Xiaojun Wu 0001, Josef Kittler
ICPR (16)2
2024 Infrared and Visible Image Fusion Based on CNN and Transformer Cross-Interaction with Semantic Modulations
Yusu Zhang, Xiaojun Wu 0001, Tianyang Xu 0001
ICPR (4)3
2024 Novel Clustering Aggregation and Multi-grained Alignment for Image-Text Matching
Xiaojun Wu 0001, Tianyang Xu 0001, Donglin Zhang 0001
ICPR (21)3
2024 Multi-frequency Fine-Grained Matching for Audio-Visual Segmentation
Yinhao Zhang, Tianyang Xu 0001, Xiaojun Wu 0001, Shao-Chuan Zhao, Josef Kittler
ICPR (29)2
2024 A Riemannian Residual Learning Mechanism for SPD Network
abstract
The generalization of Euclidean network paradigm to the Riemannian manifolds has attracted much attention for offering useful geometric representations in processing manifold-valued data in recent years. However, the information degradation during data compression mapping hinders Riemannian networks from going deeper, and there are very few solutions specifically designed for this problem. Given the remarkable success of deep Residual learning in Euclidean networks, a novel Riemannian residual learning mechanism (RRLM) is proposed in the context of Symmetric Positive Definite (SPD) manifolds, enabling the characterization of deep spatiotemporal features while preserving the manifold properties. Based on RRLM, a stack of SPD manifold-constrained residual-like blocks is designed on the tail of the original SPDNet(backbone) for the sake of conducting deep Riemannian residual learning. For simplicity, we refer to the network architecture introduced above as Riemannian residual SPD network (ResSPDNet). The experimental results achieved on three types of visual classification tasks, i.e., facial emotion recognition, drone recognition, and action recognition, demonstrate that our method can achieve improved accuracy with a deepened network structure.
Zhenyu Cai, Rui Wang 0050, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
IJCNN3
2024 MMDRFuse: Distilled Mini-Model with Dynamic Refresh for Multi-Modality Image Fusion
abstract
In recent years, Multi-Modality Image Fusion (MMIF) has been applied to many fields, which has attracted many scholars to endeavour to improve the fusion performance. However, the prevailing focus has predominantly been on the architecture design, rather than the training strategies. As a low-level vision task, image fusion is supposed to quickly deliver output images for observation and supporting downstream tasks. Thus, superfluous computational and storage overheads should be avoided. In this work, a lightweight Distilled Mini-Model with a Dynamic Refresh strategy (MMDRFuse) is proposed to achieve this objective. To pursue model parsimony, an extremely small convolutional network with a total of 113 trainable parameters (0.44 KB) is obtained by three carefully designed supervisions. First, digestible distillation is constructed by emphasising external spatial feature consistency, delivering soft supervision with balanced details and saliency for the target network. Second, we develop a comprehensive loss to balance the pixel, gradient, and perception clues from the source images. Third, an innovative dynamic refresh training strategy is used to collaborate history parameters and current supervision during training, together with an adaptive adjust function to optimise the fusion network. Extensive experiments on several public datasets demonstrate that our method exhibits promising advantages in terms of model efficiency and complexity, with superior performance in multiple image fusion tasks and downstream pedestrian detection application. The code of this work is publicly available at https://github.com/yanglinDeng/MMDRFuse.
Yanglin Deng, Tianyang Xu 0001, Chunyang Cheng, Xiaojun Wu 0001, Josef Kittler
ACM Multimedia2
2024 Dynamic Subframe Splitting and Spatio-Temporal Motion Entangled Sparse Attention for RGB-E Tracking
Pengcheng Shao, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001, Josef Kittler
PRCV (13)2
2024 Local Point Matching for Collaborative Image Registration and RGBT Anti-UAV Tracking
Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001
PRCV (12)2
2024 M-adapter: Multi-level image-to-video adaptation for video action recognition
Rongchang Li 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Linze Li 0002, Josef Kittler
Comput. Vis. Image Underst.2
2024 Scene adaptive mechanism for action recognition
Cong Wu 0006, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler
Comput. Vis. Image Underst.3
2024 Learning Adaptive Spatio-Temporal Inference Transformer for Coarse-to-Fine Animal Visual Tracking: Algorithm and Benchmark
Tianyang Xu 0001, Ze Kang, Xuefeng Zhu 0003, Xiaojun Wu 0001
Int. J. Comput. Vis.1
2024 Learning Feature Restoration Transformer for Robust Dehazing Visual Object Tracking
Tianyang Xu 0001, Yifan Pan, Zhenhua Feng 0001, Xuefeng Zhu 0003, Chunyang Cheng, Xiaojun Wu 0001, Josef Kittler
Int. J. Comput. Vis.1
2024 A Spatio-Temporal Robust Tracker with Spatial-Channel Transformer and Jitter Suppression
Shao-Chuan Zhao, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
Int. J. Comput. Vis.2
2024 UniMod1K: Towards a More Universal Large-Scale Dataset and Benchmark for Multi-modal Learning
Xuefeng Zhu 0003, Tianyang Xu 0001, Zongtao Liu, Zhangyong Tang, Xiaojun Wu 0001, Josef Kittler
Int. J. Comput. Vis.2
2024 View-shuffled clustering via the modified Hungarian algorithm
Wenhua Dong, Xiaojun Wu 0001, Tianyang Xu 0001, Zhenhua Feng 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler
Neural Networks3
2024 CRTrack: Learning Correlation-Refine network for visual object tracking
Tianyang Xu 0001, Jiang Zhai, Wankou Yang
Pattern Recognit.3
2024 Self-supervised learning for RGB-D object tracking
Xuefeng Zhu 0003, Tianyang Xu 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Xiaojun Wu 0001, Zhenhua Feng 0001, Josef Kittler
Pattern Recognit.2
2024 Towards accurate unsupervised video captioning with implicit visual feature injection and explicit
Tianyang Xu 0001, Xiaoning Song, Xuefeng Zhu 0003, Zhenhua Feng 0001, Xiaojun Wu 0001
Pattern Recognit. Lett.2
2024 Feature enhancement and coarse-to-fine detection for RGB-D tracking
Xuefeng Zhu 0003, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
Pattern Recognit. Lett.2
2024 Unified Referring Expression Generation for Bounding Boxes and Segmentations
abstract
Referring expression generation (REG) is a challenging task at the intersection of computer vision and natural language processing, which aims at generating natural language descriptions that uniquely refer to a specific object within an image. Existing REG approaches solely utilize bounding boxes in a rather primitive manner to specify target objects, and employ the classical Convolutional Neural Networks (CNNs) for image encoding, followed by recurrent layers for text generation. In this letter, we propose a novel end-to-end REG model. Our model highlights the target using bounding boxes and segmentations in a unified fashion. Specifically, we propose two settings for utilizing these signals: employing them as inputs to the model and as supervision signals for pre-training tasks. Additionally, we harness the power of the recently prevailed self-attention architecture to bridge targeted visual clues and text correspondence. During inference, our method achieves state-of-the-art performance in a one-stage manner, reflecting the potential of both bounding boxes and segmentation references in constructing REG solutions.
Zongtao Liu, Tianyang Xu 0001, Xiaoning Song, Xiaojun Wu 0001
IEEE Signal Process. Lett.2
2024 APMG: 3D Molecule Generation Driven by Atomic Chemical Properties
abstract
Recently, mask-fill-based 3D Molecular Generation (MG) methods have become very popular in virtual drug design. However, the existing MG methods ignore the chemical properties of atoms and contain inappropriate atomic position training data, which limits their generation capability. To mitigate the above issues, this paper presents a novel mask-fill-based 3D molecule generation model driven by atomic chemical properties (APMG). Specifically, we construct a new attention-MPNN-based encoder and introduce the electronic information into atom representations to enrich chemical properties. Also, a multi-functional classifier is designed to predict the electronic information of each generated atom, guiding the type prediction of elements and bonds. By design, the proposed method uses the chemical properties of atoms and their correlations for high-quality molecule generation. Second, to optimize the atomic position training data, we propose a novel atomic training position generation approach using the Chi-Square distribution. We evaluate our APMG method on the CrossDocked dataset and visualize the docking states of the pockets and generated molecules. The obtained results demonstrate the superiority and merits of APMG over the state-of-the-art approaches.
Yang Hua 0002, Zhenhua Feng 0001, Xiaoning Song, Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Dongjun Yu
IEEE ACM Trans. Comput. Biol. Bioinform.5
2024 Deep Metric Learning on the SPD Manifold for Image Set Classification
abstract
Thanks to the efficacy of Symmetric Positive Definite (SPD) manifold in characterizing video sequences (image sets), image set-based visual classification has made remarkable progress. However, the issue of large intra-class diversity and inter-class similarity is still an open challenge for the research community. Although several recent studies have alleviated the above issue by constructing Riemannian neural networks for SPD matrix nonlinear processing, the degradation of structural information during multi-stage feature transformation impedes them from going deeper. Besides, a single cross-entropy loss is insufficient for discriminative learning as it neglects the peculiarities of data distribution. To this end, this paper develops a novel framework for image set classification. Specifically, we first choose a mainstream neural network built on the SPD manifold (SPDNet)[25]as the backbone with a stacked SPD manifold autoencoder (SSMAE) built on the tail to enrich the structured representations. Due to the associated reconstruction error terms, the embedding mechanism of both SSMAE and each SPD manifold autoencoder (SMAE) forms an approximate identity mapping, simplifying the training of the suggested deeper network. Then, the ReCov layer is introduced with a nonlinear function for the constructed architecture to narrow the discrepancy of the intra-class distributions from the perspective of regularizing the local statistical information of the SPD data. Afterward, two progressive metric learning stages are coupled with the proposed SSMAE to explicitly capture, encode, and analyze the geometric distributions of the generated deep representations during training. In consequence, not only a more powerful Riemannian network embedding but also effective classifiers can be obtained. Finally, a simple maximum voting strategy is applied to the outputs of the learned multiple classifiers for classification. The proposed model is evaluated on three typical visual classification tasks using widely adopted benchmarking datasets. Extensive experiments show its superiority over the state of the arts.
Rui Wang 0050, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler
IEEE Trans. Circuits Syst. Video Technol.3
2024 Motion Complement and Temporal Multifocusing for Skeleton-Based Action Recognition
abstract
Modeling sequences with spatial-temporal graph convolutional networks has become a mainstream paradigm in skeleton-based action recognition. However, many existing methods adopt redundant or cluttered structures to mine the key action features, thus making it difficult to achieve a balanced or leading performance in accuracy and efficiency. In this paper, we propose a novel framework, referred to as Motion Complement and Temporal Multifocusing Network (MCTM-Net), to capture the relationships within skeleton sequences by means of an efficient decomposition of the spatiotemporal graph model. Specifically, for spatial modeling, we introduce a motion-related relational descriptor that extends the channel dimension so as to enhance the modeling of motion salient regions as a complement to the conventional physical adjacency relationships. An improved parameterized physical relationship model is also proposed to better fit the data characteristics. As for temporal modeling, we propose an efficient multi-focus temporal information acquisition strategy that aggregates the information from multiple temporal spans and adjacent regions. We conduct extensive experiments on multiple representative datasets, including NTU-RGB+D (60&120), Northwestern-UCLA, and UWA3D Multiview Activity II, to validate our innovations. The experimental results show the effectiveness of our method. The code will be available athttps://github.com/cong-wu/MCMT-Net.
Cong Wu 0006, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler
IEEE Trans. Circuits Syst. Video Technol.3
2024 Distillation, Ensemble and Selection for Building a Better and Faster Siamese Based Tracker
abstract
Visual object tracking has witnessed continuous improvements in performance, thanks to deep CNN learning that recently emerged. More complex CNN models invariably offer better accuracy. However, there is a conflict between the tracking efficiency and model complexity, which poses a challenge in balancing speed against accuracy. To optimize the trade-off between these two performance criteria, a distillation-ensemble-selection framework is proposed in this paper. Without any modification to the baseline network architecture, the proposed approach enables the construction of a Siamese-based tracker with improved capacity and efficiency. Specifically, multiple student trackers are designed by means of knowledge distillation from a given teacher tracking model. To manage the varying granularity of unknown targets, an ensemble module combines the outputs of the student trackers with the help of a learnable fine-grained attention module. Besides, in the online tracking stage, a selection module adaptively controls the complexity of the tracker by identifying an appropriate subset of the candidate tracker models. We verify the effectiveness of the proposed method in both anchor-based and anchor-free paradigms. The experimental results obtained on standard benchmarking datasets demonstrate the effectiveness of the proposed method, with an outstanding and balanced performance in both accuracy and speed.
Shao-Chuan Zhao, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
IEEE Trans. Circuits Syst. Video Technol.2
2024 HabLSTM: A Nonstationary Feature Focusing LSTM for Spatiotemporal Prediction of Harmful Algal Bloom
abstract
Harmful algal bloom (HAB) has long been one of the most formidable environmental problems in the world. HAB is influenced by multifactors, and its dynamic is highly nonstationary, making its prediction challenging. The existing machine learning (ML)-based HAB prediction methods mainly use time-series data, which ignore the intrinsic relationship between spatial and temporal variations in HAB. To achieve more accurate HAB spatiotemporal prediction, a novel long short-term memory (LSTM)-based nonstationary focusing prediction model (HabLSTM) is proposed in this article. The HabLSTM network is constructed by stacking HabLSTM units consisting of the hidden states spatial differential block (HSSD) and the combined states temporal differential (CSTD) block. The HSSD block uses the gating mechanism and the difference in hidden states to generate differential features between adjacent frames and guides the network to learn short-term nonstationary features by controlling the feature update of the hidden state in the HabLSTM unit. The CSTD block uses the gating mechanism and the difference in combined states to generate the differential features of the current input sequence and guides the network to learn long-term nonstationary features by controlling the feature update of the memory state in the HabLSTM unit. These two differential features guide the HabLSTM network to focus on learning nonstationary spatiotemporal features and boost HAB spatiotemporal prediction accuracy. In addition, two new spatiotemporal datasets of HAB named as Taihu HAB A and Taihu HAB B are established using the year-A and year-B normalized difference vegetation index (NDVI) images collected by Himawari-8 satellite, respectively. The experimental results on the two HAB datasets and spatiotemporal predictive learning (ST-PL) benchmark dataset MovingMNIST++ validate the outstanding HAB prediction and nonstationary spatiotemporal features’ learning capability of HabLSTM. The source code is available athttps://github.com/lxq-jnu/HabLSTM.
Xiaoqing Luo, Peirui Wang, Zhancheng Zhang, Zhengming Zhou, Shuyang Chen, Tianyang Xu 0001, Xiaojun Wu 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Pluggable Attack for Visual Object Tracking
abstract
Performing adversarial attacks on a visual tracker aims to drift the apparent target to the background by adding malicious perturbations to the source images. Demonstrating convincingly their ability to decrease accuracy, existing tracking attackers mislead the target predictions at the decision level, but this is tracker design specific, narrowing their applicability to other tracking approaches. In contrast, we advocate that attacks be performed by corrupting the feature-level clues, i.e., the feature representations extracted by deep networks. The proposed approach provides a general attacking framework for backbone-head tracking architectures. Motivated by the knowledge that the quality of intermediate-level features strongly influences the decision making, four intermediate-level attack methods are proposed to maximise the difference between the feature distributions of natural and adversarial samples, thus decoupling the attack strategies from the form of the output of specific victim trackers. Interestingly, our intermediate-level attacks are compatible with existing decision-level attacks, thus a joint optimisation of these two kinds of adversarial objective functions has the potential to achieve better attacking performance. Hence, the proposed adversarial attack methodology can be used in conjunction with several mainstream tracking paradigms (Discriminative correlation filters, Siamese networks, and Transformer trackers), demonstrating its pluggability. The experimental results on four popular benchmarks, e.g., OTB100, UAV123, LaSOT, and TLP, verify that our method can produce impressive and consistent accuracy degeneration.
Shao-Chuan Zhao, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
IEEE Trans. Inf. Forensics Secur.2
2024 Adaptive Log-Euclidean Metrics for SPD Matrix Learning
abstract
Symmetric Positive Definite (SPD) matrices have received wide attention in machine learning due to their intrinsic capacity to encode underlying structural correlation in data. Many successful Riemannian metrics have been proposed to reflect the non-Euclidean geometry of SPD manifolds. However, most existing metric tensors are fixed, which might lead to sub-optimal performance for SPD matrix learning, especially for deep SPD neural networks. To remedy this limitation, we leverage the commonly encountered pullback techniques and propose Adaptive Log-Euclidean Metrics (ALEMs), which extend the widely used Log-Euclidean Metric (LEM). Compared with the previous Riemannian metrics, our metrics contain learnable parameters, which can better adapt to the complex dynamics of Riemannian neural networks with minor extra computations. We also present a complete theoretical analysis to support our ALEMs, including algebraic and Riemannian properties. The experimental and theoretical results demonstrate the merit of the proposed metrics in improving the performance of SPD neural networks. The efficacy of our metrics is further showcased on a set of recently developed Riemannian building blocks, including Riemannian batch normalization, Riemannian Residual blocks, and Riemannian classifiers.
Ziheng Chen 0001, Yue Song 0002, Tianyang Xu 0001, Zhiwu Huang, Xiaojun Wu 0001, Nicu Sebe
IEEE Trans. Image Process.3
2024 Perceiving Actions via Temporal Video Frame Pairs
abstract
Video action recognition aims at classifying the action category in given videos. In general, semantic-relevant video frame pairs reflect significant action patterns such as object appearance variation and abstract temporal concepts like speed, rhythm, and so on. However, existing action recognition approaches tend to holistically extract spatiotemporal features. Though effective, there is still a risk of neglecting the crucial action features occurring across frames with a long-term temporal span. Motivated by this, in this article, we propose to perceive actions via frame pairs directly and devise a novel Nest Structure with frame pairs as basic units. Specifically, we decompose a video sequence into all possible frame pairs and hierarchically organize them according to temporal frequency and order, thus transforming the original video sequence into a Nest Structure. Through naturally decomposing actions, the proposed structure can flexibly adapt to diverse action variations such as speed or rhythm changes. Next, we devise a Temporal Pair Analysis module (TPA) to extract discriminative action patterns based on the proposed Nest Structure. The designed TPA module consists of a pair calculation part to calculate the pair features and a pair fusion part to hierarchically fuse the pair features for recognizing actions. The proposed TPA can be flexibly integrated into existing backbones, serving as a side branch to capture various action patterns from multi-level features. Extensive experiments show that the proposed TPA module can achieve consistent improvements over several typical backbones, reaching or updating CNN-based SOTA results on several challenging action recognition benchmarks.
Rongchang Li 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
ACM Trans. Intell. Syst. Technol.2
2024 Multi-Level Fusion for Robust RGBT Tracking via Enhanced Thermal Representation
abstract
Due to the limitations of visible (RGB) sensors in challenging scenarios, such as nighttime and foggy environments, the thermal infrared (TIR) modality draws increasing attention as an auxiliary source for robust tracking systems. Currently, the existing methods extract both the RGB and TIR (RGBT) clues in a similar approach, i.e., utilising RGB-pretrained models with or without finetuning, and then aggregate the multi-modal information through a fusion block embedded in a single level. However, the different imaging principles of RGB and TIR data raise questions about the suitability of RGB-pretrained models for thermal data. In this article, it is argued that the modality gap is overlooked, and an alternative training paradigm is proposed for TIR data to ensure consistency between the training and test data, which is achieved by optimising the TIR feature extractor with only TIR data involved. Furthermore, with the goal of making better use of the enhanced thermal representations, a multi-level fusion strategy is inspired by the observation that various fusion strategies at different levels can contribute to a better performance. Specifically, fusion modules at both the feature and decision levels are derived for a comprehensive fusion procedure while the pixel-level fusion strategy is not considered due to the misalignment of multi-modal image pairs. The effectiveness of our method is demonstrated by extensive qualitative and quantitative experiments conducted on several challenging benchmarks. Code will be released at https://github.com/Zhangyong-Tang/MELT .
Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Riemannian Local Mechanism for SPD Neural Networks
abstract
The Symmetric Positive Definite (SPD) matrices have received wide attention for data representation in many scientific areas. Although there are many different attempts to develop effective deep architectures for data processing on the Riemannian manifold of SPD matrices, very few solutions explicitly mine the local geometrical information in deep SPD feature representations. Given the great success of local mechanisms in Euclidean methods, we argue that it is of utmost importance to ensure the preservation of local geometric information in the SPD networks. We first analyse the convolution operator commonly used for capturing local information in Euclidean deep networks from the perspective of a higher level of abstraction afforded by category theory. Based on this analysis, we define the local information in the SPD manifold and design a multi-scale submanifold block for mining local geometry. Experiments involving multiple visual tasks validate the effectiveness of our approach.
Ziheng Chen 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Rui Wang 0050, Zhiwu Huang, Josef Kittler
AAAI2
2023 RGBD1K: A Large-Scale Dataset and Benchmark for RGB-D Object Tracking
abstract
RGB-D object tracking has attracted considerable attention recently, achieving promising performance thanks to the symbiosis between visual and depth channels. However, given a limited amount of annotated RGB-D tracking data, most state-of-the-art RGB-D trackers are simple extensions of high-performance RGB-only trackers, without fully exploiting the underlying potential of the depth channel in the offline training stage. To address the dataset deficiency issue, a new RGB-D dataset named RGBD1K is released in this paper. The RGBD1K contains 1,050 sequences with about 2.5M frames in total. To demonstrate the benefits of training on a larger RGB-D data set in general, and RGBD1K in particular, we develop a transformer-based RGB-D tracker, named SPT, as a baseline for future visual object tracking studies using the new dataset. The results, of extensive experiments using the SPT tracker demonstrate the potential of the RGBD1K dataset to improve the performance of RGB-D tracking, inspiring future developments of effective tracker designs. The dataset and codes will be available on the project homepage: https://github.com/xuefeng-zhu5/RGBD1K.
Xuefeng Zhu 0003, Tianyang Xu 0001, Zhangyong Tang, Zucheng Wu, Xiaojun Wu 0001, Josef Kittler
AAAI2
2023 Mixture-of-Domain-Adapters: Decoupling and Injecting Domain Knowledge to Pre-trained Language Models' Memories
abstract
Pre-trained language models (PLMs) demonstrate excellent abilities to understand texts in the generic domain while struggling in a specific domain.Although continued pre-training on a large domain-specific corpus is effective, it is costly to tune all the parameters on the domain.In this paper, we investigate whether we can adapt PLMs both effectively and efficiently by only tuning a few parameters.Specifically, we decouple the feed-forward networks (FFNs) of the Transformer architecture into two parts: the original pre-trained FFNs to maintain the old-domain knowledge and our novel domain-specific adapters to inject domainspecific knowledge in parallel.Then we adopt a mixture-of-adapters gate to fuse the knowledge from different domain adapters dynamically.Our proposed Mixture-of-Domain-Adapters (MixDA) employs a two-stage adapter-tuning strategy that leverages both unlabeled data and labeled data to help the domain adaptation: i) domain-specific adapter on unlabeled data; followed by ii) the task-specific adapter on labeled data.MixDA can be seamlessly plugged into the pretraining-finetuning paradigm and our experiments demonstrate that MixDA achieves superior performance on in-domain tasks (GLUE), out-of-domain tasks (ChemProt, RCT, IMDB, Amazon), and knowledge-intensive tasks (KILT).Further analyses demonstrate the reliability, scalability, and efficiency of our method.1 * Equal Contribution. 1 The code is available at https://github.com/ Amano-Aki/Mixture-of-Domain-Adapters.
Shizhe Diao, Tianyang Xu 0001, Ruijia Xu, Tong Zhang 0001
ACL (1)2
2023 VLNet: A Multi-task Network for Joint Vehicle and Lane Detection
Aiqi Feng, Tianyang Xu 0001, Donglin Zhang 0001, Xiaojun Wu 0001
ICIG (2)3
2023 When Diffusion Model Meets with Adversarial Attack: Generating Transferable Adversarial Examples on Face Recognition
Tianyang Xu 0001, Xiaojun Wu 0001
ICIG (5)3
2023 Multi-modal Stream Fusion for Skeleton-Based Action Recognition
Ruixuan Pang, Rongchang Li 0001, Tianyang Xu 0001, Xiaoning Song, Xiaojun Wu 0001
ICIG (3)3
2023 Semantic-Guided Multi-feature Fusion for Accurate Video Captioning
Tianyang Xu 0001, Xiaoning Song, Zhenghua Feng, Xiaojun Wu 0001
ICIG (4)2
2023 ACMA-GAN: Adaptive Cross-Modal Attention for Text-to-Image Generation
Longlong Zhou, Xiaojun Wu 0001, Tianyang Xu 0001
ICIG (4)3
2023 COMIM-GAN: Improved Text-to-Image Generation via Condition Optimization and Mutual Information Maximization
Longlong Zhou, Xiaojun Wu 0001, Tianyang Xu 0001
MMM (1)3
2023 Asymmetric Attention Fusion for Unsupervised Video Object Segmentation
Hongfan Jiang, Xiaojun Wu 0001, Tianyang Xu 0001
PRCV (6)3
2023 ETM-face: effective training sample selection and multi-scale feature learning for face detection
Junyuan He, Xiaoning Song, Zhenhua Feng 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
Multim. Tools Appl.4
2023 U-SPDNet: An SPD manifold learning-based neural network for visual classification
Rui Wang 0050, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler
Neural Networks3
2023 Enhanced robust spatial feature selection and correlation filter learning for UAV tracking
Jiajun Wen 0001, Hong-Lin Chu, Zhihui Lai 0001, Tianyang Xu 0001, LinLin Shen
Neural Networks4
2023 Global Context-Aware Feature Extraction and Visible Feature Enhancement for Occlusion-Invariant Pedestrian Detection in Crowded Scenes
Zhen Liu 0015, Xiaoning Song, Zhenhua Feng 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
Neural Process. Lett.4
2023 Scalable Affine Multi-view Subspace Clustering
Wanrong Yu, Xiaojun Wu 0001, Tianyang Xu 0001, Ziheng Chen 0001, Josef Kittler
Neural Process. Lett.3
2023 LRRNet: A Novel Representation Learning Guided Fusion Network for Infrared and Visible Images
abstract
Deep learning based fusion methods have been achieving promising performance in image fusion tasks. This is attributed to the network architecture that plays a very important role in the fusion process. However, in general, it is hard to specify a good fusion architecture, and consequently, the design of fusion networks is still a black art, rather than science. To address this problem, we formulate the fusion task mathematically, and establish a connection between its optimal solution and the network architecture that can implement it. This approach leads to a novel method proposed in the paper of constructing a lightweight fusion network. It avoids the time-consuming empirical network design by a trial-and-test strategy. In particular we adopt a learnable representation approach to the fusion task, in which the construction of the fusion network architecture is guided by the optimisation algorithm producing the learnable model. The low-rank representation (LRR) objective is the foundation of our learnable model. The matrix multiplications, which are at the heart of the solution are transformed into convolutional operations, and the iterative process of optimisation is replaced by a special feed-forward network. Based on this novel network architecture, an end-to-end lightweight fusion network is constructed to fuse infrared and visible light images. Its successful training is facilitated by a detail-to-semantic information loss function proposed to preserve the image details and to enhance the salient features of the source images. Our experiments show that the proposed fusion network exhibits better fusion performance than the state-of-the-art fusion methods on public datasets. Interestingly, our network requires a fewer training parameters than other existing methods.
Hui Li 0037, Tianyang Xu 0001, Xiaojun Wu 0001, Jiwen Lu, Josef Kittler
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Learning Motion-Perceive Siamese network for robust visual object tracking
Ze Kang, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001
Pattern Recognit. Lett.2
2023 Hybrid Riemannian Graph-Embedding Metric Learning for Image Set Classification
abstract
With the continuously increasing amount of video data, image set classification has recently received widespread attention in the CV&PR community. However, the intra-class diversity and inter-class ambiguity of representations remain an open challenge. To tackle this issue, several methods have been put forward to perform multiple geometry-aware image set modelling and learning. Although the extracted complementary geometric information is beneficial for decision making, the sophisticated computational paradigm (e.g., scatter matrices computation and iterative optimisation) of such algorithms is counterproductive. As a countermeasure, we propose an effective hybrid Riemannian metric learning framework in this paper. Specifically, we design a multiple graph embedding-guided metric learning framework for the sake of fusing these complementary kernel features, obtained via the explicit RKHS embeddings of the Grassmannian manifold, SPD manifold, and Gaussian embedded Riemannian manifold, into a unified subspace for classification. Furthermore, the involved optimisation problem of the developed model can be solved in terms of a series of sub-problems, achieving improved efficiency theoretically and experimentally. Substantial experiments are carried out to evaluate the efficacy of our approach. The experimental results suggest the superiority of it over the state-of-the-art methods.
Ziheng Chen 0001, Tianyang Xu 0001, Xiaojun Wu 0001, Rui Wang 0050, Josef Kittler
IEEE Trans. Big Data2
2023 Fast Self-Guided Multi-View Subspace Clustering
abstract
Multi-view subspace clustering is an important topic in cluster analysis. Its aim is to utilize the complementary information conveyed by multiple views of objects to be clustered. Recently, view-shared anchor learning based multi-view clustering methods have been developed to speed up the learning of common data representation. Although widely applied to large-scale scenarios, most of the existing approaches are still faced with two limitations. First, they do not pay sufficient consideration on the negative impact caused by certain noisy views with unclear clustering structures. Second, many of them only focus on the multi-view consistency, yet are incapable of capturing the cross-view diversity. As a result, the learned complementary features may be inaccurate and adversely affect clustering performance. To solve these two challenging issues, we propose a Fast Self-guided Multi-view Subspace Clustering (FSMSC) algorithm which skillfully integrates the view-shared anchor learning and global-guided-local self-guidance learning into a unified model. Such an integration is inspired by the observation that the view with clean clustering structures will play a more crucial role in grouping the clusters when the features of all views are concatenated. Specifically, we first learn a locally-consistent data representation shared by all views in the local learning module, then we learn a globally-discriminative data representation from multi-view concatenated features in the global learning module. Afterwards, a feature selection matrix constrained by the ℓ2,1-norm is designed to construct a guidance from global learning to local learning. In this way, the multi-view consistent and diverse information can be simultaneously utilized and the negative impact caused by noisy views can be overcame to some extent. Extensive experiments on different datasets demonstrate the effectiveness of our proposed fast self-guided learning model, and its promising performance compared to both, the state-of-the-art non-deep and deep multi-view clustering algorithms. The code of this paper is available at https://github.com/chenzhe207/FSMSC.
Zhe Chen 0018, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler
IEEE Trans. Image Process.3
2023 Toward Robust Visual Object Tracking With Independent Target-Agnostic Detection and Effective Siamese Cross-Task Interaction
abstract
Advanced Siamese visual object tracking architectures are jointly trained using pair-wise input images to perform target classification and bounding box regression. They have achieved promising results in recent benchmarks and competitions. However, the existing methods suffer from two limitations: First, though the Siamese structure can estimate the target state in an instance frame, provided the target appearance does not deviate too much from the template, the detection of the target in an image cannot be guaranteed in the presence of severe appearance variations. Second, despite the classification and regression tasks sharing the same output from the backbone network, their specific modules and loss functions are invariably designed independently, without promoting any interaction. Yet, in a general tracking task, the centre classification and bounding box regression tasks are collaboratively working to estimate the final target location. To address the above issues, it is essential to perform target-agnostic detection so as to promote cross-task interactions in a Siamese-based tracking framework. In this work, we endow a novel network with a target-agnostic object detection module to complement the direct target inference, and to avoid or minimise the misalignment of the key cues of potential template-instance matches. To unify the multi-task learning formulation, we develop a cross-task interaction module to ensure consistent supervision of the classification and regression branches, improving the synergy of different branches. To eliminate potential inconsistencies that may arise within a multi-task architecture, we assign adaptive labels, rather than fixed hard labels, to supervise the network training more effectively. The experimental results obtained on several benchmarks, i.e., OTB100, UAV123, VOT2018, VOT2019, and LaSOT, demonstrate the effectiveness of the advanced target detection module, as well as the cross-task interaction, exhibiting superior tracking performance as compared with the state-of-the-art tracking methods.
Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
IEEE Trans. Image Process.1
2023 WATCH: Two-Stage Discrete Cross-Media Hashing
abstract
Due to the explosive growth of multimedia data in recent years, cross-media hashing (CMH) approaches have recently received increasing attention. To learn the hash codes, most existing supervised CMH algorithms employ the strict binary label information, which has small margins between the incorrect labels (0) and the true labels (1), increasing the risk of classification error. Besides, most existing CMH approaches are one-stage algorithms, in which the hash functions and binary codes can be learned simultaneously, complicating the optimization. To avoid NP-hard optimization, many approaches utilize a relaxation strategy. However, this optimisation trick may cause large quantization errors. To address this, we present a novel tWo-stAge discreTe Cross-media Hashing method based on smooth matrix factorization and label relaxation, named WATCH. The proposed WATCH controls the margins adaptively by the novel label relaxation strategy. This innovation reduces the quantization error significantly. Besides, WATCH is a two-stage model. In stage 1, we employ a discrete smooth matrix factorization model. Then, the hash codes can be generated discretely, reducing the large quantization loss greatly. In stage 2, we adopt an effective hash function learning strategy, which produces more effective hash functions. Comprehensive experiments on several datasets demonstrate that WATCH outperforms some state-of-the-art methods.
Donglin Zhang 0001, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler
IEEE Trans. Knowl. Data Eng.3
2023 DAH: Discrete Asymmetric Hashing for Efficient Cross-Media Retrieval
abstract
Given the merits in high computational efficiency and low storage cost, hashing techniques have been widely studied in cross-media retrieval. Existing methods usually adopt the equal length encoding scheme to represent the multimedia data. However, the strictly equal length scheme maybe not optimal because the dimension of different modalities is often various. Besides, there exists other challenges in designing a cross-media retrieval system, e.g., how to address the discrete constraints, how to avoid using the n*n similarity matrix, and how to effectively exploit the discriminative label information. To conquer the above challenges, we propose a novel method, i.e., discrete asymmetric hashing (DAH). Specifically, DAH exploits a flexible model, which can seamlessly deal with equal or unequal length encoding scenarios. Moreover, DAH constructs a supervised semantic embedding framework by jointly minimizing the distance-distance difference and label reconstructing error, significantly reducing the computational complexity. An asymmetric strategy is employed to establish the connection between hash codes and the latent subspace. Furthermore, the hash codes can be learned discretely by the designed optimization algorithm. In the training stage2, a semantic intersection scheme is proposed to learn more powerful hash functions. Experiments show that our DAH is effective in equal and unequal scenarios.
Donglin Zhang 0001, Xiaojun Wu 0001, Tianyang Xu 0001, He-Feng Yin
IEEE Trans. Knowl. Data Eng.3
2023 Discriminative Dictionary Pair Learning With Scale-Constrained Structured Representation for Image Classification
abstract
The dictionary pair learning (DPL) model aims to design a synthesis dictionary and an analysis dictionary to accomplish the goal of rapid sample encoding. In this article, we propose a novel structured representation learning algorithm based on the DPL for image classification. It is referred to as discriminative DPL with scale-constrained structured representation (DPL-SCSR). The proposed DPL-SCSR utilizes the binary label matrix of dictionary atoms to project the representation into the corresponding label space of the training samples. By imposing a non-negative constraint, the learned representation adaptively approximates a block-diagonal structure. This innovative transformation is also capable of controlling the scale of the block-diagonal representation by enforcing the sum of within-class coefficients of each sample to 1, which means that the dictionary atoms of each class compete to represent the samples from the same class. This implies that the requirement of similarity preservation is considered from the perspective of the constraint on the sum of coefficients. More importantly, the DPL-SCSR does not need to design a classifier in the representation space as the label matrix of the dictionary can also be used as an efficient linear classifier. Finally, the DPL-SCSR imposes the$l_{2,p}$-norm on the analysis dictionary to make the process of feature extraction more interpretable. The DPL-SCSR seamlessly incorporates the scale-constrained structured representation learning, within-class similarity preservation of representation, and the linear classifier into one regularization term, which dramatically reduces the complexity of training and parameter tuning. The experimental results on several popular image classification datasets show that our DPL-SCSR can deliver superior performance compared with the state-of-the-art (SOTA) dictionary learning methods. The MATLAB code of this article is available athttps://github.com/chenzhe207/DPL-SCSR.
Zhe Chen 0018, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler
IEEE Trans. Neural Networks Learn. Syst.3
2022 DreamNet: A Deep Riemannian Manifold Network for SPD Matrix Learning
Rui Wang 0050, Xiaojun Wu 0001, Ziheng Chen 0001, Tianyang Xu 0001, Josef Kittler
ACCV (6)4
2022 FoGMesh: 3D Human Mesh Recovery in Videos with Focal Transformer and GRU
Yihao He, Xiaoning Song, Tianyang Xu 0001, Yang Hua 0002, Xiaojun Wu 0001
BMVC3
2022 Memory-Token Transformer for Unsupervised Video Anomaly Detection
abstract
Video anomaly detection is crucial for behavior analysis, which has witnessed continuous progress in recent years with the auto-encoder based reconstruction framework. However, in some cases, abnormal frames may also be reconstructed well due to the strong representation ability of deep networks, increasing missed detection. To mitigate this issue, the existing methods usually the memory bank method. This method records normal patterns and assigns high errors for the reconstruction of abnormal frames into normal frames. In this paper, to better use the semantic information of normal videos recorded in the memory module, we introduce the Memory-Token Transformer (MTT) to boost the reconstruction performance on normal frames. We assume that the anomalies in a video mainly concentrate on the regions containing people and relevant objects. Therefore, during the decoding stage, we first extract the semantic concepts of a feature map and generate the corresponding semantic tokens. Then the tokens are combined with the proposed memory module. Last, we introduce a transformer to fuse the complex relationship among different tokens, and use 3D convolution with the pooling operator in our encoder to enhance spatio-temporal feature extraction as compared with 2D models. The experimental results obtained on various benchmarks demonstrate the effectiveness of the proposed method.
Youyu Li, Xiaoning Song, Tianyang Xu 0001, Zhenhua Feng 0001
ICPR3
2022 KITPose: Keypoint-Interactive Transformer for Animal Pose Estimation
Jiyong Rao, Tianyang Xu 0001, Xiaoning Song, Zhenhua Feng 0001, Xiaojun Wu 0001
PRCV (1)2
2022 FPNFuse: A lightweight feature pyramid network for infrared and visible image fusion
abstract
Abstract A novel deep learning structure of infrared and visible image fusion is proposed. In particular, feature pyramid networks are developed for enhanced feature extraction across multiple convolutional layers. Besides, a fusion strategy is improved based on the channel attention mechanism to highlight the relevant attributes in the fusion stage. The fusion method consists of four parts: encoder, feature pyramid networks, fusion strategy and decoder, respectively. First, the multi‐scale deep features are extracted from the source images by encoder with embedded feature pyramid networks, realizing cross‐layer interaction. Second, these features are fused by the improved fusion strategy with channel attention for each scale. Finally, the fused features are reconstructed by the designed decoder to produce the informative fused image. The experimental results show that the proposed fusion method achieves state‐of‐the‐art results in both qualitative and quantitative evaluation with a lightweight architecture.
Zi-Han Zhang, Xiaojun Wu 0001, Tianyang Xu 0001
IET Image Process.3
2022 One-step kernelized sparse clustering on grassmann manifolds
Wenbo Hu 0008, Xiaojun Wu 0001, Tianyang Xu 0001
Multim. Tools Appl.3
2022 Learning a discriminative SPD manifold neural network for image set classification
Rui Wang 0050, Xiaojun Wu 0001, Ziheng Chen 0001, Tianyang Xu 0001, Josef Kittler
Neural Networks4
2022 Multi-view Subspace Clustering via Joint Latent Representations
Wenhua Dong, Xiaojun Wu 0001, Tianyang Xu 0001
Neural Process. Lett.3
2022 Structured classifier-based dictionary pair learning for pattern classification
Yu-Hong Cai, Xiaojun Wu 0001, Zhe Chen 0018, Tianyang Xu 0001
Pattern Anal. Appl.4
2022 Target-Cognisant Siamese Network for Robust Visual Object Tracking
Yingjie Jiang, Xiaoning Song, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
Pattern Recognit. Lett.3
2022 FEXNet: Foreground Extraction Network for Human Action Recognition
abstract
As most human actions in video sequences embody the continuous interactions between foregrounds rather than the background scene, it is significant to disentangle these foregrounds from the background for advanced action recognition systems. In this paper, therefore, we propose a Foreground EXtraction (FEX) block to explicitly model the foreground clues to achieve effective management of action subjects. In particular, the designed FEX block contains two components. The first part is a Foreground Enhancement (FE) module, which highlights the potential feature channels related to the action attributes, providing channel-level refinement for the following spatiotemporal modeling. The second phase is a Scene Segregation (SS) module, which splits feature maps into foreground and background. Specifically, a temporal model with dynamic enhancement is constructed for the foreground part, reflecting the essential nature of the action category. While the background is modeled using simple spatial convolutions, mapping the inputs to the consistent feature space. The FEX blocks can be inserted into existing 2D CNNs (denoted as FEXNet) for spatiotemporal modeling, concentrating on the foreground clues for effective action inference. Our experiments performed on Something-Something V1, V2 and Kinetics400 verify the effectiveness of the proposed method.
Xiaojun Wu 0001, Tianyang Xu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Robust Visual Object Tracking Via Adaptive Attribute-Aware Discriminative Correlation Filters
abstract
In recent years, attention mechanisms have been widely studied in Discriminative Correlation Filter (DCF) based visual object tracking. To realise spatial attention and discriminative feature mining, existing approaches usually apply regularisation terms to the spatial dimension of multi-channel features. However, these spatial regularisation approaches construct a shared spatial attention pattern for all multi-channel features, without considering the diversity across channels. As each feature map (channel) focuses on a specific visual attribute, a shared spatial attention pattern limits the capability for mining important information from different channels. To address this issue, we advocate channel-specific spatial attention for DCF-based trackers. The key ingredient of the proposed method is an Adaptive Attribute-Aware spatial attention mechanism for constructing a novel DCF-based tracker (A$^3$DCF). To highlight the discriminative elements in each feature map, spatial sparsity is imposed in the filter learning stage, moderated by the prior knowledge regarding the expected concentration of signal energy. In addition, we perform a post processing of the identified spatial patterns to alleviate the impact of less significant channels. The net effect is that the irrelevant and inconsistent channels are removed by the proposed method. The results obtained on a number of well-known benchmarking datasets, including OTB2015, DTB70, UAV123, VOT2018, LaSOT, GOT-10 K and TrackingNet, demonstrate the merits of the proposed A$^3$DCF tracker, with improved performance compared to the state-of-the-art methods.
Xuefeng Zhu 0003, Xiaojun Wu 0001, Tianyang Xu 0001, Zhenhua Feng 0001, Josef Kittler
IEEE Trans. Multim.3
2022 Two-Stage Supervised Discrete Hashing for Cross-Modal Retrieval
abstract
Recently, hashing-based multimodal learning systems have received increasing attention due to their query efficiency and parsimonious storage costs. However, impeded by the quantization loss caused by numerical optimization, the existing cross-media hashing approaches are unable to capture all the discriminative information present in the original multimodal data. Besides, most cross-modal methods belong to the one-step paradigm, which learn the binary codes and hash function simultaneously, increasing the complexity of optimization. To address these issues, we propose a novel two-stage approach, named the two-stage supervised discrete hashing (TSDH) method. In particular, in the first phase, TSDH generates a latent representation for each modality. These representations are then mapped to a common Hamming space to generate the binary codes. In addition, TSDH directly endows the hash codes with the semantic labels, enhancing the discriminatory power of the learned binary codes. A discrete hash optimization approach is developed to learn the binary codes without relaxation, avoiding the large quantization loss. The proposed hash function learning scheme reuses the semantic information contained by the embeddings, endowing the hash functions with enhanced discriminability. Extensive experiments on several databases demonstrate the effectiveness of the developed TSDH, outperforming several recent competitive cross-media algorithms.
Donglin Zhang 0001, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler
IEEE Trans. Syst. Man Cybern. Syst.3
2021 MSC-Fuse: An Unsupervised Multi-scale Convolutional Fusion Framework for Infrared and Visible Image
Guo-Yang Chen, Xiaojun Wu 0001, Hui Li 0037, Tianyang Xu 0001
ICIG (1)4
2021 Locality-Constrained Collaborative Representation with Multi-resolution Dictionary for Face Recognition
Zhen Liu 0015, Xiaojun Wu 0001, He-Feng Yin, Tianyang Xu 0001, Zhenqiu Shu
PRCV (1)4
2021 Adaptive Channel Selection for Robust Visual Object Tracking with Discriminative Correlation Filters
abstract
Abstract Discriminative Correlation Filters (DCF) have been shown to achieve impressive performance in visual object tracking. However, existing DCF-based trackers rely heavily on learning regularised appearance models from invariant image feature representations. To further improve the performance of DCF in accuracy and provide a parsimonious model from the attribute perspective, we propose to gauge the relevance of multi-channel features for the purpose of channel selection. This is achieved by assessing the information conveyed by the features of each channel as a group, using an adaptive group elastic net inducing independent sparsity and temporal smoothness on the DCF solution. The robustness and stability of the learned appearance model are significantly enhanced by the proposed method as the process of channel selection performs implicit spatial regularisation. We use the augmented Lagrangian method to optimise the discriminative filters efficiently. The experimental results obtained on a number of well-known benchmarking datasets demonstrate the effectiveness and stability of the proposed method. A superior performance over the state-of-the-art trackers is achieved using less than $$10\%$$ 10 % deep feature channels.
Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
Int. J. Comput. Vis.1
2021 Dynamic information enhancement for video classification
Rongchang Li 0001, Xiaojun Wu 0001, Cong Wu 0006, Tianyang Xu 0001, Josef Kittler
Image Vis. Comput.4
2021 Adaptive feature fusion for visual object tracking
Shao-Chuan Zhao, Tianyang Xu 0001, Xiaojun Wu 0001, Xuefeng Zhu 0003
Pattern Recognit.2
2021 Learning Alternating Deep-Layer Cascaded Representation
abstract
We propose an alternating deep-layer cascade (A-DLC) architecture for representation learning in the context of image classification. The merits of the proposed model are threefold. First, A-DLC is the first-ever method that alternatively cascades the sparse and collaborative representations using the class-discriminant softmax vector representation at the interface of each cascade section so that the sparsity and collaborativity can simultaneously be considered. Second, A-DLC inherits the hierarchy learning capability that effectively extends the traditional shallow sparse coding to a multi-layer learning model, thus enabling a full exploitation of the inherent latent discriminative information. Third, the simulation results show a significant amelioration in the classification accuracy, compared to earlier one-step single-layer classification algorithms. The Matlab code of this paper is available at https://github.com/chenzhe207/A-DLC.
Zhe Chen 0018, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler
IEEE Signal Process. Lett.3
2021 Complementary Discriminative Correlation Filters Based on Collaborative Representation for Visual Object Tracking
abstract
In recent years, discriminative correlation filter (DCF) based algorithms have significantly advanced the state of the art in visual object tracking. The key to the success of DCF is an efficient discriminative regression model trained with powerful multi-cue features, including both hand-crafted and deep neural network features. However, the tracking performance is hindered by their inability to respond adequately to abrupt target appearance variations. This issue is posed by the limited representation capability of fixed image features. In this work, we set out to rectify this shortcoming by proposing a complementary representation of a visual content. Specifically, we propose the use of a collaborative representation between successive frames to extract the dynamic appearance information from a target with rapid appearance changes, which results in suppressing the undesirable impact of the background. The resulting collaborative representation coefficients are combined with the original feature maps using a spatially regularised DCF framework for performance boosting. The experimental results on several benchmarking datasets demonstrate the effectiveness and robustness of the proposed method, as compared with a number of state-of-the-art tracking algorithms.
Xuefeng Zhu 0003, Xiaojun Wu 0001, Tianyang Xu 0001, Zhenhua Feng 0001, Josef Kittler
IEEE Trans. Circuits Syst. Video Technol.3
2021 From RGB to Depth: Domain Transfer Network for Face Anti-Spoofing
Yahang Wang, Xiaoning Song, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001
IEEE Trans. Inf. Forensics Secur.3
2021 Correlation tracking with implicitly extending search region
Xiaojun Wu 0001, Josef Kittler, Tianyang Xu 0001
Vis. Comput.4
2020 Adaptive Context-Aware Discriminative Correlation Filters for Robust Visual Object Tracking
abstract
In recent years, Discriminative Correlation Filters (DCFs) have gained popularity due to their superior performance in visual object tracking. However, existing DCF trackers usually learn filters using fixed attention mechanisms that focus on the centre of an image and suppresses filter amplitudes in surroundings. In this paper, we propose an Adaptive Context-Aware Discriminative Correlation Filter (ACA-DCF) that is able to improve the existing DCF formulation with complementary attention mechanisms. Our ACA-DCF integrates foreground attention and background attention for complementary context-aware filter learning. More importantly, we ameliorate the design using an adaptive weighting strategy that takes complex appearance variations into account. The experimental results obtained on several well-known benchmarks demonstrate the effectiveness and superiority of the proposed method over the state-of-the-art approaches.
Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
ICPR1
2020 An accelerated correlation filter tracker
Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
Pattern Recognit.1
2020 Learning Low-Rank and Sparse Discriminative Correlation Filters for Coarse-to-Fine Visual Object Tracking
abstract
Discriminative correlation filter (DCF) has achieved advanced performance in visual object tracking with remarkable efficiency guaranteed by its implementation in the frequency domain. However, the effect of the structural relationship of DCF and object features has not been adequately explored in the context of the filter design. To remedy this deficiency, this paper proposes a Low-rank and Sparse DCF (LSDCF) that improves the relevance of features used by discriminative filters. To be more specific, we extend the classical DCF paradigm from ridge regression to lasso regression, and constrain the estimate to be of low-rank across frames, thus identifying and retaining the informative filters distributed on a low-dimensional manifold. To this end, specific temporal-spatial-channel configurations are adaptively learned to achieve enhanced discrimination and interpretability. In addition, we analyse the complementary characteristics between hand-crafted features and deep features, and propose a coarse-to-fine heuristic tracking strategy to further improve the performance of our LSDCF. Last, the augmented Lagrange multiplier optimisation method is used to achieve efficient optimisation. The experimental results obtained on a number of well-known benchmarking datasets, including OTB2013, OTB50, OTB100, TC128, UAV123, VOT2016 and VOT2018, demonstrate the effectiveness and robustness of the proposed method, delivering outstanding performance compared to the state-of-the-art trackers.
Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
IEEE Trans. Circuits Syst. Video Technol.1
2019 Joint Group Feature Selection and Discriminative Filter Learning for Robust Visual Object Tracking
abstract
We propose a new Group Feature Selection method for Discriminative Correlation Filters (GFS-DCF) based visual object tracking. The key innovation of the proposed method is to perform group feature selection across both channel and spatial dimensions, thus to pinpoint the structural relevance of multi-channel features to the filtering system. In contrast to the widely used spatial regularisation or feature selection methods, to the best of our knowledge, this is the first time that channel selection has been advocated for DCF-based tracking. We demonstrate that our GFS-DCF method is able to significantly improve the performance of a DCF tracker equipped with deep neural network features. In addition, our GFS-DCF enables joint feature selection and filter learning, achieving enhanced discrimination and interpretability of the learned filters. To further improve the performance, we adaptively integrate historical information by constraining filters to be smooth across temporal frames, using an efficient low-rank approximation. By design, specific temporal-spatial-channel configurations are dynamically learned in the tracking process, highlighting the relevant features, and alleviating the performance degrading impact of less discriminative representations and reducing information redundancy. The experimental results obtained on OTB2013, OTB2015, VOT2017, VOT2018 and TrackingNet demonstrate the merits of our GFS-DCF and its superiority over the state-of-the-art trackers. The code is publicly available at \url{https://github.com/XU-TIANYANG/GFS-DCF}.
Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
ICCV1
2019 Learning Adaptive Discriminative Correlation Filters via Temporal Consistency Preserving Spatial Feature Selection for Robust Visual Object Tracking
abstract
With efficient appearance learning models, discriminative correlation filter (DCF) has been proven to be very successful in recent video object tracking benchmarks and competitions. However, the existing DCF paradigm suffers from two major issues, i.e., spatial boundary effect and temporal filter degradation. To mitigate these challenges, we propose a new DCF-based tracking method. The key innovations of the proposed method include adaptive spatial feature selection and temporal consistent constraints, with which the new tracker enables joint spatial-temporal filter learning in a lower dimensional discriminative manifold. More specifically, we apply structured spatial sparsity constraints to multi-channel filters. Consequently, the process of learning spatial filters can be approximated by the lasso regularization. To encourage temporal consistency, the filter model is restricted to lie around its historical value and updated locally to preserve the global structure in the manifold. Last, a unified optimization framework is proposed to jointly select temporal consistency preserving spatial features and learn discriminative filters with the augmented Lagrangian method. Qualitative and quantitative evaluations have been conducted on a number of well-known benchmarking datasets such as OTB2013, OTB50, OTB100, Temple-Colour, UAV123, and VOT2018. The experimental results demonstrate the superiority of the proposed method over the state-of-the-art approaches.
Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Josef Kittler
IEEE Trans. Image Process.1
2018 Non-negative Subspace Representation Learning Scheme for Correlation Filter Based Tracking
abstract
Discriminative correlation filter (DCF) based tracking methods have achieved great success recently. However, the temporal learning scheme in the current paradigm is of a linear recursion form determined by a fixed learning rate which can not adaptively feedback appearance variations. In this paper, we propose a unified non-negative subspace representation constrained leaning scheme for DCF. The subspace is constructed by several templates with auxiliary memory mechanisms. Then the current template is projected onto the subspace to find the non-negative representation and to determine the corresponding template weights. Our learning scheme enables efficient combination of correlation filter and subspace structure. The experimental results on OTB50 demonstrate the effectiveness of our learning formulation.
Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler
ICPR1