EDBT 2026 Demo / reviewers in the wild / expert
Zhangyong Tang
dblp:236/3195
· DBLP profile ↗
16ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0001-8187-9384ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Computer networks · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Attack Intensity is Target-Related: Exploration of Sparse Adversarial Attack against Visual Object Trackers
Shao-Chuan Zhao, Tianyang Xu 0001, Zhangyong Tang, Xiaojun Wu 0001, Josef Kittler |
Int. J. Comput. Vis. | 4 |
| 2026 | EvaNet: Toward More Efficient and Consistent Infrared and Visible Image Fusion AssessmentabstractEvaluation is essential in image fusion research, yet most existing metrics are directly borrowed from other vision tasks without proper adaptation. These traditional metrics, often based on complex image transformations, not only fail to capture the true quality of the fusion results but also are computationally demanding. To address these issues, we propose a unified evaluation framework specifically tailored for image fusion. At its core is a lightweight network designed efficiently to approximate widely used metrics, following a divide-and-conquer strategy. Unlike conventional approaches that directly assess similarity between fused and source images, we first decompose the fusion result into infrared and visible components. The evaluation model is then used to measure the degree of information preservation in these separated components, effectively disentangling the fusion evaluation process. During training, we incorporate a contrastive learning strategy and inform our evaluation model by perceptual scene assessment provided by a large language model. Last, we propose the first consistency evaluation framework, which measures the alignment between image fusion metrics and human visual perception, using both independent no-reference scores and downstream tasks performance as objective references. Extensive experiments show that our learning-based evaluation paradigm delivers both superior efficiency (up to 1,000 times faster) and greater consistency across a range of standard image fusion benchmarks. Chunyang Cheng, Tianyang Xu 0001, Xiaojun Wu 0001, Tao Zhou 0002, Hui Li 0037, Zhangyong Tang, Josef Kittler |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Curriculum adaptation for one-stream RGB-T tracking
Xiantao Hu, Fansheng Zeng, Bineng Zhong 0001, Zhangyong Tang, Wenxuan Fang 0001, Jun Li 0027, Ying Tai, Jian Yang 0003 |
Pattern Recognit. | 4 |
| 2026 | Switcher: Adaptive framework for unified and customized multi-modal object tracking
He Wang 0028, Tianyang Xu 0001, Zhangyong Tang, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. | 3 |
| 2025 | One Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image FusionabstractAdvanced image fusion methods mostly prioritise high-level missions, where task interaction struggles with semantic gaps, requiring complex bridging mechanisms. In contrast, we propose to leverage low-level vision tasks from digital photography fusion, allowing for effective feature interaction through pixel-level supervision. This new paradigm provides strong guidance for unsupervised multimodal fusion without relying on abstract semantics, enhancing task-shared feature learning for broader applicability. Owning to the hybrid image features and enhanced universal representations, the proposed GIFNet supports diverse fusion tasks, achieving high performance across both seen and unseen scenarios with a single model. Uniquely, experimental results reveal that our framework also supports single-modality enhancement, offering superior flexibility for practical applications. Our code will be available at https://github.com/AWCXV/GIFNet. Chunyang Cheng, Tianyang Xu 0001, Zhenhua Feng 0001, Xiaojun Wu 0001, Zhangyong Tang, Hui Li 0037, Zeyang Zhang 0002, Sara Atito Ali Ahmed, Muhammad Awais 0001, Josef Kittler |
CVPR | 5 |
| 2025 | Serial Over Parallel: Learning Continual Unification for Multi-Modal Visual Object Tracking and BenchmarkingabstractUnifying multiple multi-modal visual object tracking (MMVOT) tasks draws increasing attention due to the complementary nature of different modalities in building robust tracking systems. Existing practices mix all data sensor types in a single training procedure, structuring a parallel paradigm from the data-centric perspective and aiming for a global optimum on the joint distribution of the involved tasks. However, the absence of a unified benchmark where all types of data coexist forces evaluations on separated benchmarks, causing inconsistency between training and testing, thus leading to performance degradation. To address these issues, this work advances in two aspects: A unified benchmark, coined as UniBench300, is introduced to bridge the inconsistency by incorporating multiple task data, reducing inference passes from three to one and cutting time consumption by 27%. The unification process is reformulated in a serial format, progressively integrating new tasks. In this way, the performance degradation can be specified as knowledge forgetting of previous tasks, which naturally aligns with the philosophy of continual learning (CL), motivating further exploration of injecting CL into the unification process. Extensive experiments conducted on two baselines and four benchmarks demonstrate the significance of UniBench300 and the superiority of CL in supporting a stable unification process. Moreover, while conducting dedicated analyses, the performance degradation is found to be negatively correlated with network capacity. Additionally, modality discrepancies contribute to varying degradation levels across tasks (RGBT > RGBD > RGBE in MMVOT), offering valuable insights for future multi-modal vision research. Source codes and the proposed benchmark is available at https://github.com/Zhangyong-Tang/UniBench300. Zhangyong Tang, Tianyang Xu 0001, Xuefeng Zhu 0003, Chunyang Cheng, Tao Zhou 0002, Xiaojun Wu 0001, Josef Kittler |
ACM Multimedia | 1 |
| 2025 | Cross-Modal Supervised Contrastive Learning for RGB-T Semantic Segmentation
Chuanjiang Zhang, Tianyang Xu 0001, Zhangyong Tang, Xiaojun Wu 0001 |
PRCV (5) | 3 |
| 2025 | TENet: Targetness entanglement incorporating with multi-scale pooling and mutually-guided fusion for RGB-E object tracking
Pengcheng Shao, Tianyang Xu 0001, Zhangyong Tang, Linze Li 0002, Xiaojun Wu 0001, Josef Kittler |
Neural Networks | 3 |
| 2025 | Temporal aggregation for real-time RGBT tracking via fast decision-level fusion
Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. Lett. | 1 |
| 2025 | M3Track: Meta-Prompt for Multi-Modal TrackingabstractPrompt-tuning has shown remarkable success in multi-modal visual tracking, which enhances RGB tracking by incorporating an additional modality,e.g., thermal infrared (T), Depth (D), or Event (E), forming an RGB+X tracking paradigm. However, in the testing phase, current approaches utilise a frozen prompt for the entire benchmark, failing to account for the diversity and unique characteristics of individual videos. To address this issue, we inject meta-learning solutions into current prompt-based tracking technique, thereby emphasising sequence-level adaptation. Unlike traditional prompt-based trackers, which keep the parameters of prompt blocks fixed during testing, our approach updates these parameters in the first frame of each video using a meta-learning solution. This allows for enhanced discriminative tracking capabilities tailored to each video. Building on this advancement, we extend our methodology beyond separate implementations for RGBT, RGBD, and RGBE tasks. A unified multi-modal tracker is further derived, resulting in the first unified tracker without any task priors (notification of task type) employed in both training and testing phases. Extensive experimental results on LasHeR, DepthTrack, VisEvent, GTOT, RGBT234, and RGBD1K consistently demonstrate the superiority of the proposed method against the existing prompt-tuning paradigm. Source codes are available athttps://github.com/Zhangyong-Tang/M3Track. Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
IEEE Signal Process. Lett. | 1 |
| 2025 | Revisiting RGBT Tracking Benchmarks From the Perspective of Modality Validity: A New Benchmark, Problem, and SolutionabstractRGBT tracking draws increasing attention because of its robustness in multi-modal warranting (MMW) scenarios, such as nighttime and adverse weather conditions, where relying on a single sensing modality fails to ensure stable tracking results. However, existing benchmarks predominantly contain videos collected in common scenarios where both RGB and thermal infrared (TIR) information are of sufficient quality. This weakens the representativeness of existing benchmarks in severe imaging conditions, leading to tracking failures in MMW scenarios. To bridge this gap, we present a new benchmark considering the modality validity, MV-RGBT, captured specifically from MMW scenarios where either RGB (extreme illumination) or TIR (thermal truncation) modality is invalid. Hence, it is further divided into two subsets according to the valid modality, offering a new compositional perspective for evaluation and providing valuable insights for future designs. Moreover, MV-RGBT is the most diverse benchmark of its kind, featuring 36 different object categories captured across 19 distinct scenes. Furthermore, considering severe imaging conditions in MMW scenarios, a new problem is posed in RGBT tracking, named 'when to fuse', to stimulate the development of fusion strategies for such scenarios. To facilitate its discussion, we propose a new solution with a mixture of experts, named MoETrack, where each expert generates independent tracking results along with a confidence score. Extensive results demonstrate the significant potential of MV-RGBT in advancing RGBT tracking and elicit the conclusion that fusion is not always beneficial, especially in MMW scenarios. Besides, MoETrack achieves state-of-the-art results on several benchmarks, including MV-RGBT, GTOT, and LasHeR. Source codes and benchmarks are available at https://github.com/Zhangyong-Tang/MVRGBT. Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Xuefeng Zhu 0003, Chunyang Cheng, Zhenhua Feng 0001, Josef Kittler |
IEEE Trans. Image Process. | 1 |
| 2024 | Generative-Based Fusion Mechanism for Multi-Modal TrackingabstractGenerative models (GMs) have received increasing research interest for their remarkable capacity to achieve comprehensive understanding. However, their potential application in the domain of multi-modal tracking has remained unexplored. In this context, we seek to uncover the potential of harnessing generative techniques to address the critical challenge, information fusion, in multi-modal tracking. In this paper, we delve into two prominent GM techniques, namely, Conditional Generative Adversarial Networks (CGANs) and Diffusion Models (DMs). Different from the standard fusion process where the features from each modality are directly fed into the fusion block, we combine these multi-modal features with random noise in the GM framework, effectively transforming the original training samples into harder instances. This design excels at extracting discriminative clues from the features, enhancing the ultimate tracking performance. Based on this, we conduct extensive experiments across two multi-modal tracking tasks, three baseline methods, and four challenging benchmarks. The experimental results demonstrate that the proposed generative-based fusion mechanism achieves state-of-the-art performance by setting new records on GTOT, LasHeR and RGBD1K. Code will be available at https://github.com/Zhangyong-Tang/GMMT. Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Xuefeng Zhu 0003, Josef Kittler |
AAAI | 1 |
| 2024 | UniMod1K: Towards a More Universal Large-Scale Dataset and Benchmark for Multi-modal Learning
Xuefeng Zhu 0003, Tianyang Xu 0001, Zongtao Liu, Zhangyong Tang, Xiaojun Wu 0001, Josef Kittler |
Int. J. Comput. Vis. | 4 |
| 2024 | Multi-Level Fusion for Robust RGBT Tracking via Enhanced Thermal RepresentationabstractDue to the limitations of visible (RGB) sensors in challenging scenarios, such as nighttime and foggy environments, the thermal infrared (TIR) modality draws increasing attention as an auxiliary source for robust tracking systems. Currently, the existing methods extract both the RGB and TIR (RGBT) clues in a similar approach, i.e., utilising RGB-pretrained models with or without finetuning, and then aggregate the multi-modal information through a fusion block embedded in a single level. However, the different imaging principles of RGB and TIR data raise questions about the suitability of RGB-pretrained models for thermal data. In this article, it is argued that the modality gap is overlooked, and an alternative training paradigm is proposed for TIR data to ensure consistency between the training and test data, which is achieved by optimising the TIR feature extractor with only TIR data involved. Furthermore, with the goal of making better use of the enhanced thermal representations, a multi-level fusion strategy is inspired by the observation that various fusion strategies at different levels can contribute to a better performance. Specifically, fusion modules at both the feature and decision levels are derived for a comprehensive fusion procedure while the pixel-level fusion strategy is not considered due to the misalignment of multi-modal image pairs. The effectiveness of our method is demonstrated by extensive qualitative and quantitative experiments conducted on several challenging benchmarks. Code will be released at https://github.com/Zhangyong-Tang/MELT . Zhangyong Tang, Tianyang Xu 0001, Xiaojun Wu 0001, Josef Kittler |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | RGBD1K: A Large-Scale Dataset and Benchmark for RGB-D Object TrackingabstractRGB-D object tracking has attracted considerable attention recently, achieving promising performance thanks to the symbiosis between visual and depth channels. However, given a limited amount of annotated RGB-D tracking data, most state-of-the-art RGB-D trackers are simple extensions of high-performance RGB-only trackers, without fully exploiting the underlying potential of the depth channel in the offline training stage. To address the dataset deficiency issue, a new RGB-D dataset named RGBD1K is released in this paper. The RGBD1K contains 1,050 sequences with about 2.5M frames in total. To demonstrate the benefits of training on a larger RGB-D data set in general, and RGBD1K in particular, we develop a transformer-based RGB-D tracker, named SPT, as a baseline for future visual object tracking studies using the new dataset. The results, of extensive experiments using the SPT tracker demonstrate the potential of the RGBD1K dataset to improve the performance of RGB-D tracking, inspiring future developments of effective tracker designs. The dataset and codes will be available on the project homepage: https://github.com/xuefeng-zhu5/RGBD1K. Xuefeng Zhu 0003, Tianyang Xu 0001, Zhangyong Tang, Zucheng Wu, Xiaojun Wu 0001, Josef Kittler |
AAAI | 3 |
| 2018 | Time Context-Aware IPTV Program Recommendation Based on Tensor LearningabstractIPTV provides a variety of services to users, e.g., Live TV, Video on Demand (VOD), and Catch-up TV, which allows users to freely select from a huge pool of program genres. A high-quality program recommendation system that can well predict users' dynamic preferences is desirable to improve user satisfaction. In this paper, we make the first attempt to design a personalized TV program recommendation system based on tensor learning that leverages temporal context information and implicit feedback. We establish the user preference model by carefully analyzing the influence of program genres and time context on users' viewing behavior in IPTV viewing logs. To obtain the optimal program recommendation, we adopt tensor decomposition to mine the latent relationship among users, program genres and the time contexts, while address issues of data sparsity and missing information. Our extensive experimental results on a real dataset from a major IPTV service provider in China have confirmed the superiority of our proposed recommendation algorithm over baseline algorithms in achieving a higher recommendation accuracy and covering a wider range of programs. Xiaoyan Yin 0001, Yanjiao Chen, Xiaoqian Mi, Zhangyong Tang, Dingyi Fang |
GLOBECOM | 5 |