EDBT 2026 Demo / reviewers in the wild / expert
Xuefeng Zhu 0003
dblp:70/6628-3
· DBLP profile ↗
24ranked-venue papers
8as first author
24since 2021 · last 2026
0000-0003-0262-5891ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPD-Updater: Symmetric positive definite manifold geometry based temporal updating for visual object trackingabstractVisual object tracking has witnessed continuous advances in recent years along with the exciting developments in backbone networks. In general, all advanced solutions adhere to the template-based tracking framework, which exhibits powerful representative capacity gained via offline training. However, when the target undergoes appearance changes or occlusion, the tracker, which relies on a fixed template defined in the initial frame, struggles to locate it accurately in such complex situations. To achieve online adaptation, recent studies have introduced dynamic templates. Typically, the adopted solution is to compute reliability scores in the traditional Euclidean space to assess the confidence of the dynamic template. However, the Euclidean metric is unreliable to some extent in high-dimensional feature spaces, potentially resulting in a negative impact by involving incorrect dynamic templates. To overcome this problem, we exploit the compact geometric representation capacity of the Symmetric Positive Definite (SPD) manifold to design a novel score prediction module for the tracker update (SPD-Updater). By switching to an SPD manifold metric, we obtain a more accurate and stable dynamic template, thereby enhancing the model capacity to handle complex situations. To validate the reliability of manifold metric in tracking models, we conduct experiments with trackers using different backbones. The experimental results on LaSOT, GOT-10k, TrackingNet, and UAV123 demonstrate the effectiveness of our approach, reflecting the merit of the SPD metric in online tracking adaptation. Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001, Josef Kittler |
Neural Networks | 3 |
| 2026 | Adaptive Continual Learning for Online Visual Object Tracking via Dynamic Grassmannian Appearance ModelingabstractOnline visual object tracking fundamentally constitutes a continual learning challenge, demanding persistent adaptation to target variations within dynamic video streams, while preserving critical features to prevent forgetting. Existing template-based trackers are particularly vulnerable to severe target deformation and drastic background changes in real-world scenarios. Recent explorations enhance adaptability through local update strategies—such as dynamic templates, appearance tokens, and parameter fine-tuning—to address rapid appearance variations. However, these methods inherently propagate target appearance changes without explicit modelling and prioritise short-term adaptation over long-term global representation. Consequently, they fail to balance initial target features with current observations, leading to tracking failure in long-term scenarios. To overcome these limitations, we propose DG-Track, an adaptive continual learning framework leveraging Grassmannian manifold geometry. Specifically, we represent target appearance within a Grassmannian affine subspace and perform continual adaptation via incremental learning. Compared to Euclidean geometry, the Grassmannian manifold captures nonlinear appearance representations, yielding more compact and geometrically consistent temporal dynamics modelling. Furthermore, an adaptive forgetting module dynamically regulates the interplay between the current observed subspace and the initial template, ensuring stable long-term tracking. DG-Track is a plug-and-play solution for online tracking, adding no learnable parameters. Comprehensive experiments with diverse baseline trackers on LaSOT, GOT-10k, TrackingNet, and UAV123 validate the efficacy of our continual learning framework and Grassmannian manifold geometry in enhancing visual object tracking performance. Code is available at https://github.com/xiaoqing0825/DGTrack. Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | PETS2025: Multi-Authority Multi-Sensor Maritime Surveillance Challenge and EvaluationabstractThis paper presents the outcomes of the PETS2025 challenge, held in conjunction with AVSS 2025 and sponsored by the EU-funded EURMARS project. The challenge introduces a novel maritime surveillance dataset comprising image sequences captured by diverse multi-altitude, multimodal sensors, reflecting the real-world multi-authority environment. The key tasks include: (1) object detection using various sensors across different platforms (ground-based and low-altitude aerial) and spectral ranges (visible, thermal, ultraviolet (UV), and short-wave infrared (SWIR)); (2) long-term tracking of targets in maritime environments spanning both sea and land; and (3) approximating target geolocations by using sensor imagery and telemetry data. Performance evaluations of results submitted by 12 international participants are discussed. The results show the effectiveness of these submissions and highlight ongoing challenges posed by heterogeneous sensors and complex environments. These challenges emphasise the need to further improve detection, tracking, and geolocation approximation for maritime and coastal surveillance. Thanet Markchom, Jonathan N. Boyle, Lulu Chen, James M. Ferryman, Matteo Marturini, Stephan Veigl, Andreas Opitz, Andreas Kriechbaum-Zabini, Romaios Bratskas, Anastasios Gkamaris, Dimitris Papachristos, George Leventakis, Wenjun Fan, Hsiang-Wei Huang, Jeng-Neng Hwang, Pyong-Kun Kim, Kwangju Kim, Chung-I Huang, Kenta Saito, Shunta Kaneko, Kyoko Sudo, Nguyen Thanh Thien, Meng-Yu Kao, Jun-Wei Hsieh, Teepakorn Lilek, Tossapol Pomsuwan, Jinjie Gu, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001, Josef Kittler, Stephanie Stacy, Alfredo Gabaldon, Peter Tu, Dongyoung Kim, Kyoungoh Lee |
AVSS | 29 |
| 2025 | Serial Over Parallel: Learning Continual Unification for Multi-Modal Visual Object Tracking and BenchmarkingabstractUnifying multiple multi-modal visual object tracking (MMVOT) tasks draws increasing attention due to the complementary nature of different modalities in building robust tracking systems. Existing practices mix all data sensor types in a single training procedure, structuring a parallel paradigm from the data-centric perspective and aiming for a global optimum on the joint distribution of the involved tasks. However, the absence of a unified benchmark where all types of data coexist forces evaluations on separated benchmarks, causing inconsistency between training and testing, thus leading to performance degradation. To address these issues, this work advances in two aspects: A unified benchmark, coined as UniBench300, is introduced to bridge the inconsistency by incorporating multiple task data, reducing inference passes from three to one and cutting time consumption by 27%. The unification process is reformulated in a serial format, progressively integrating new tasks. In this way, the performance degradation can be specified as knowledge forgetting of previous tasks, which naturally aligns with the philosophy of continual learning (CL), motivating further exploration of injecting CL into the unification process. Extensive experiments conducted on two baselines and four benchmarks demonstrate the significance of UniBench300 and the superiority of CL in supporting a stable unification process. Moreover, while conducting dedicated analyses, the performance degradation is found to be negatively correlated with network capacity. Additionally, modality discrepancies contribute to varying degradation levels across tasks (RGBT > RGBD > RGBE in MMVOT), offering valuable insights for future multi-modal vision research. Source codes and the proposed benchmark is available at https://github.com/Zhangyong-Tang/UniBench300. Zhangyong Tang, Tianyang Xu 0001, Xuefeng Zhu 0003, Chunyang Cheng, Tao Zhou 0002, Xiaojun Wu 0001, Josef Kittler |
ACM Multimedia | 3 |
| 2025 | Collaborating Vision, Depth, and Thermal Signals for Multi-Modal Tracking: Dataset and AlgorithmabstractExisting multi-modal object tracking approaches primarily focus on dual-modal paradigms, such as RGB-Depth or RGB-Thermal, yet remain challenged in complex scenarios due to limited input modalities. To address this gap, this work introduces a novel multi-modal tracking task that leverages three complementary modalities, including visible RGB, Depth (D), and Thermal Infrared (TIR), aiming to enhance robustness in complex scenarios. To support this task, we construct a new multi-modal tracking dataset, coined RGBDT500, which consists of 500 videos with synchronised frames across the three modalities. Each frame provides spatially aligned RGB, depth, and thermal infrared images with precise object bounding box annotations.Furthermore, we propose a novel multi-modal tracker, dubbed RDTTrack.RDTTrack integrates tri-modal information for robust tracking by leveraging a pretrained RGB-only tracking model and prompt learning techniques.In specific, RDTTrack fuses thermal infrared and depth modalities under a proposed orthogonal projection constraint, then integrates them with RGB signals as prompts for the pre-trained foundation tracking model, effectively harmonising tri-modal complementary cues.The experimental results demonstrate the effectiveness and advantages of the proposed method, showing significant improvements over existing dual-modal approaches in terms of tracking accuracy and robustness in complex scenarios. The dataset and source code are publicly available at https://xuefeng-zhu5.github.io/RGBDT500. Xuefeng Zhu 0003, Tianyang Xu 0001, Yifan Pan, Jinjie Gu, Xi Li 0001, Jiwen Lu, Xiaojun Wu 0001, Josef Kittler |
NeurIPS | 1 |
| 2025 | Learning adaptive detection and tracking collaborations with augmented UAV synthesis for accurate anti-UAV system
Shihan Liu, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001, Josef Kittler |
Expert Syst. Appl. | 3 |
| 2025 | Adaptive Colour-Depth Aware Attention for RGB-D Object TrackingabstractRecent advances in RGB-D tracking have been driven by the synergistic combination of high-performing RGB-only trackers and auxiliary depth information. However, most existing methods rely on visual feature descriptors to extract depth features, which are then fused with vision features. This pipeline may lead to performance degradation due to the incongruence between the RGB and depth modalities. In this letter, we propose an efficient and effective transformer-based framework, that explicitly models colour and depth information for RGB-D tracking. Specifically, we first statistically code the colour and depth information of the foreground and background for the template. Then, the spatial attention maps of the search region are obtained using these colour-depth statistical models, enhancing the visual features of the search region for improved object localisation accuracy. The comprehensive experimental results obtained on multiple benchmarks demonstrate the effectiveness and merits of the proposed approach in explicit colour-depth coding for RGB-D tracking. The code and the models are publicly accessible athttps://github.com/xuefeng-zhu5/CDAAT. Xuefeng Zhu 0003, Tianyang Xu 0001, Xiaojun Wu 0001, Zhenhua Feng 0001, Josef Kittler |
IEEE Signal Process. Lett. | 1 |
| 2025 | Revisiting RGBT Tracking Benchmarks From the Perspective of Modality Validity: A New Benchmark, Problem, and SolutionabstractRGBT tracking draws increasing attention because of its robustness in multi-modal warranting (MMW) scenarios, such as nighttime and adverse weather conditions, where relying on a single sensing modality fails to ensure stable tracking results. However, existing benchmarks predominantly contain videos collected in common scenarios where both RGB and thermal infrared (TIR) information are of sufficient quality. This weakens the representativeness of existing benchmarks in severe imaging conditions, leading to tracking failures in MMW scenarios. To bridge this gap, we present a new benchmark considering the modality validity, MV-RGBT, captured specifically from MMW scenarios where either RGB (extreme illumination) or TIR (thermal truncation) modality is invalid. Hence, it is further divided into two subsets according to the valid modality, offering a new compositional perspective for evaluation and providing valuable insights for future designs. Moreover, MV-RGBT is the most diverse benchmark of its kind, featuring 36 different object categories captured across 19 distinct scenes. Furthermore, considering severe imaging conditions in MMW scenarios, a new problem is posed in RGBT tracking, named 'when to fuse', to stimulate the development of fusion strategies for such scenarios. To facilitate its discussion, we propose a new solution with a mixture of experts, named MoETrack, where each expert generates independent tracking results along with a confidence score. Extensive results demonstrate the significant potential of MV-RGBT in advancing RGBT tracking and elicit the conclusion that fusion is not always beneficial, especially in MMW scenarios. Besides, MoETrack achieves state-of-the-art results on several benchmarks, including MV-RGBT, GTOT, and LasHeR. Source codes and benchmarks are available at https://github.com/Zhangyong-Tang/MVRGBT. Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Xuefeng Zhu 0003, Chunyang Cheng, Zhenhua Feng 0001, Josef Kittler |
IEEE Trans. Image Process. | 4 |
| 2024 | Generative-Based Fusion Mechanism for Multi-Modal TrackingabstractGenerative models (GMs) have received increasing research interest for their remarkable capacity to achieve comprehensive understanding. However, their potential application in the domain of multi-modal tracking has remained unexplored. In this context, we seek to uncover the potential of harnessing generative techniques to address the critical challenge, information fusion, in multi-modal tracking. In this paper, we delve into two prominent GM techniques, namely, Conditional Generative Adversarial Networks (CGANs) and Diffusion Models (DMs). Different from the standard fusion process where the features from each modality are directly fed into the fusion block, we combine these multi-modal features with random noise in the GM framework, effectively transforming the original training samples into harder instances. This design excels at extracting discriminative clues from the features, enhancing the ultimate tracking performance. Based on this, we conduct extensive experiments across two multi-modal tracking tasks, three baseline methods, and four challenging benchmarks. The experimental results demonstrate that the proposed generative-based fusion mechanism achieves state-of-the-art performance by setting new records on GTOT, LasHeR and RGBD1K. Code will be available at https://github.com/Zhangyong-Tang/GMMT. Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Xuefeng Zhu 0003, Josef Kittler |
AAAI | 4 |
| 2024 | Learning Explicit Modulation Vectors for Disentangled Transformer Attention-Based RGB-D Visual Tracking
Yifan Pan, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaoqing Luo, Xiaojun Wu 0001, Josef Kittler |
ICPR (16) | 3 |
| 2024 | Attention-Based Patch Matching and Motion-Driven Point Association for Accurate Point Tracking
Han Zang, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaoning Song, Xiaojun Wu 0001, Josef Kittler |
ICPR (16) | 3 |
| 2024 | Dynamic Subframe Splitting and Spatio-Temporal Motion Entangled Sparse Attention for RGB-E Tracking
Pengcheng Shao, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001, Josef Kittler |
PRCV (13) | 3 |
| 2024 | Local Point Matching for Collaborative Image Registration and RGBT Anti-UAV Tracking
Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001 |
PRCV (12) | 3 |
| 2024 | Learning Adaptive Spatio-Temporal Inference Transformer for Coarse-to-Fine Animal Visual Tracking: Algorithm and Benchmark
Tianyang Xu 0001, Ze Kang, Xuefeng Zhu 0003, Xiaojun Wu 0001 |
Int. J. Comput. Vis. | 3 |
| 2024 | Learning Feature Restoration Transformer for Robust Dehazing Visual Object Tracking
Tianyang Xu 0001, Yifan Pan, Zhenhua Feng 0001, Xuefeng Zhu 0003, Chunyang Cheng, Xiaojun Wu 0001, Josef Kittler |
Int. J. Comput. Vis. | 4 |
| 2024 | UniMod1K: Towards a More Universal Large-Scale Dataset and Benchmark for Multi-modal Learning
Xuefeng Zhu 0003, Tianyang Xu 0001, Zongtao Liu, Zhangyong Tang, Xiaojun Wu 0001, Josef Kittler |
Int. J. Comput. Vis. | 1 |
| 2024 | Self-supervised learning for RGB-D object tracking
Xuefeng Zhu 0003, Tianyang Xu 0001, Sara Atito Ali Ahmed, Muhammad Awais 0001, Xiaojun Wu 0001, Zhenhua Feng 0001, Josef Kittler |
Pattern Recognit. | 1 |
| 2024 | Towards accurate unsupervised video captioning with implicit visual feature injection and explicit
Tianyang Xu 0001, Xiaoning Song, Xuefeng Zhu 0003, Zhenhua Feng 0001, Xiaojun Wu 0001 |
Pattern Recognit. Lett. | 4 |
| 2024 | Feature enhancement and coarse-to-fine detection for RGB-D tracking
Xuefeng Zhu 0003, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. Lett. | 1 |
| 2023 | RGBD1K: A Large-Scale Dataset and Benchmark for RGB-D Object TrackingabstractRGB-D object tracking has attracted considerable attention recently, achieving promising performance thanks to the symbiosis between visual and depth channels. However, given a limited amount of annotated RGB-D tracking data, most state-of-the-art RGB-D trackers are simple extensions of high-performance RGB-only trackers, without fully exploiting the underlying potential of the depth channel in the offline training stage. To address the dataset deficiency issue, a new RGB-D dataset named RGBD1K is released in this paper. The RGBD1K contains 1,050 sequences with about 2.5M frames in total. To demonstrate the benefits of training on a larger RGB-D data set in general, and RGBD1K in particular, we develop a transformer-based RGB-D tracker, named SPT, as a baseline for future visual object tracking studies using the new dataset. The results, of extensive experiments using the SPT tracker demonstrate the potential of the RGBD1K dataset to improve the performance of RGB-D tracking, inspiring future developments of effective tracker designs. The dataset and codes will be available on the project homepage: https://github.com/xuefeng-zhu5/RGBD1K. Xuefeng Zhu 0003, Tianyang Xu 0001, Zhangyong Tang, Zucheng Wu, Xiaojun Wu 0001, Josef Kittler |
AAAI | 1 |
| 2023 | Learning Motion-Perceive Siamese network for robust visual object tracking
Ze Kang, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001 |
Pattern Recognit. Lett. | 3 |
| 2022 | Robust Visual Object Tracking Via Adaptive Attribute-Aware Discriminative Correlation FiltersabstractIn recent years, attention mechanisms have been widely studied in Discriminative Correlation Filter (DCF) based visual object tracking. To realise spatial attention and discriminative feature mining, existing approaches usually apply regularisation terms to the spatial dimension of multi-channel features. However, these spatial regularisation approaches construct a shared spatial attention pattern for all multi-channel features, without considering the diversity across channels. As each feature map (channel) focuses on a specific visual attribute, a shared spatial attention pattern limits the capability for mining important information from different channels. To address this issue, we advocate channel-specific spatial attention for DCF-based trackers. The key ingredient of the proposed method is an Adaptive Attribute-Aware spatial attention mechanism for constructing a novel DCF-based tracker (A$^3$DCF). To highlight the discriminative elements in each feature map, spatial sparsity is imposed in the filter learning stage, moderated by the prior knowledge regarding the expected concentration of signal energy. In addition, we perform a post processing of the identified spatial patterns to alleviate the impact of less significant channels. The net effect is that the irrelevant and inconsistent channels are removed by the proposed method. The results obtained on a number of well-known benchmarking datasets, including OTB2015, DTB70, UAV123, VOT2018, LaSOT, GOT-10 K and TrackingNet, demonstrate the merits of the proposed A$^3$DCF tracker, with improved performance compared to the state-of-the-art methods. Xuefeng Zhu 0003, Xiaojun Wu 0001, Tianyang Xu 0001, Zhenhua Feng 0001, Josef Kittler |
IEEE Trans. Multim. | 1 |
| 2021 | Adaptive feature fusion for visual object tracking
Shao-Chuan Zhao, Tianyang Xu 0001, Xiaojun Wu 0001, Xuefeng Zhu 0003 |
Pattern Recognit. | 4 |
| 2021 | Complementary Discriminative Correlation Filters Based on Collaborative Representation for Visual Object TrackingabstractIn recent years, discriminative correlation filter (DCF) based algorithms have significantly advanced the state of the art in visual object tracking. The key to the success of DCF is an efficient discriminative regression model trained with powerful multi-cue features, including both hand-crafted and deep neural network features. However, the tracking performance is hindered by their inability to respond adequately to abrupt target appearance variations. This issue is posed by the limited representation capability of fixed image features. In this work, we set out to rectify this shortcoming by proposing a complementary representation of a visual content. Specifically, we propose the use of a collaborative representation between successive frames to extract the dynamic appearance information from a target with rapid appearance changes, which results in suppressing the undesirable impact of the background. The resulting collaborative representation coefficients are combined with the original feature maps using a spatially regularised DCF framework for performance boosting. The experimental results on several benchmarking datasets demonstrate the effectiveness and robustness of the proposed method, as compared with a number of state-of-the-art tracking algorithms. Xuefeng Zhu 0003, Xiaojun Wu 0001, Tianyang Xu 0001, Zhenhua Feng 0001, Josef Kittler |
IEEE Trans. Circuits Syst. Video Technol. | 1 |