EDBT 2026 Demo / reviewers in the wild / expert
Yaozong Zheng
dblp:344/6907
· DBLP profile ↗
17ranked-venue papers
5as first author
17since 2021 · last 2026
0009-0007-2664-0574ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unified Multi-Modal Tracking via Proxy Visual PromptsabstractRecent unified multi-modal tracking frameworks often encounter high computational overhead due to complex fusion operations. In this paper, we propose PoATrack, a proxy-based bidirectional fusion framework designed to facilitate efficient multi-modal tracking via interactive proxy prompts. Specifically, the framework treats features from all modalities equally and introduces an adaptive feature enhancement module to improve spatial representations within the search region. To further reduce the fusion cost, a lightweight proxy prompt fusion module is developed following a two-stage strategy: (1) representing modality-specific features through proxy visual prompts, and (2) dynamically learning hierarchical cross-modal relationships via proxy-guided fusion for enriched contextual modeling. Extensive experiments on six public benchmarks (RGB-T, RGB-E, and RGB-D) demonstrate that PoATrack achieves competitive accuracy while operating 1.8× faster than SDSTrack [17] under identical hardware settings. Bineng Zhong 0001, Yaozong Zheng, Qihua Liang, Zhiruo Zhu, Shuxiang Song 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2026 | FMTrack: Frequency-Aware Interaction and Multi-Expert Fusion for RGB-T TrackingabstractRecently, RGB-T tracking has received increasing attention due to its robustness. However, existing RGB-T trackers mainly use cross-attention for modal feature interaction, limiting the utilization of complementary information. In addition, these trackers employ fixed dominant-auxiliary paradigms for feature fusion, ignoring modal quality fluctuations. To address these issues, we propose FMTrack, an effective framework for fully capturing complementary information. FMTrack consists of two key components, a frequency-aware interaction network (FIN) and a multi-expert fusion module (MEFM). To emphasize the valuable information in each modality, FIN utilizes frequency masks to perform high-pass and low-pass filtering on RGB and TIR data. FIN explicitly establishes cross-modal interactions via frequency domain learning, which facilitates the sharing of complementary information. Besides, MEFM extracts diverse features via the differentiated expert network and then adjusts feature combinations according to modal reliability, achieving deep understanding and flexible fusion of multimodal data. With FIN and MEFM, FMTrack makes full use of the advantageous information of each modality to highlight target representations, thus improving performance in complex scenes. Extensive experiments on four popular RGBT tracking datasets (LasHeR, VTUAV, RGBT234, and RGBT210) show that our FMTrack achieves leading performance. The code is available at https://github.com/xyl-507/FMTrack. Yuanliang Xue, Guodong Jin, Bineng Zhong 0001, Lining Tan, Chaocan Xue, Yaozong Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Robust RGB-T Tracking via Learnable Visual Fourier Prompt Fine-Tuning and Modality Fusion Prompt GenerationabstractRecently, visual prompt tuning is introduced to RGB-Thermal (RGB-T) tracking as a parameter-efficient finetuning (PEFT) method. However, these PEFT-based RGB-T tracking methods typically rely solely on spatial domain information as prompts for feature extraction. As a result, they often fail to achieve optimal performance by overlooking the crucial role of frequency-domain information in prompt learning. To address this issue, we propose an efficient Visual Fourier Prompt Tracking (named VFPTrack) method to learn modality-related prompts via Fast Fourier Transform (FFT). Our method consists of symmetric feature extraction encoder with shared parameters, visual-fourier prompts, and Modality Fusion Prompt Generator that generates bidirectional interaction prompts through multi-modal feature fusion. Specifically, we first use a frozen feature extraction encoder to extract RGB and thermal infrared (TIR) modality features. Then, we combine the visual prompts in the spatial domain with the frequency domain prompts obtained from the FFT, which allows for the full extraction and understanding of modality features from different domain information. Finally, unlike previous fusion methods, the modality fusion prompt generation module we use combines features from different modalities to generate a fused modality prompt. This modality prompt is interacted with each individual modality to fully enable feature interaction across different modalities. Extensive experiments conducted on three popular RGB-T tracking benchmarks show that our method demonstrates outstanding performance. Bineng Zhong 0001, Qihua Liang, Zhiruo Zhu, Yaozong Zheng, Ning Li 0044 |
IEEE Trans. Multim. | 5 |
| 2025 | Less Is More: Token Context-Aware Learning for Object TrackingabstractRecently, several studies have shown that utilizing contextual information to perceive target states is crucial for object tracking. They typically capture context by incorporating multiple video frames. However, these naive frame-context methods fail to consider the importance of each patch within a reference frame, making them susceptible to noise and redundant tokens, which deteriorates tracking performance. To address this challenge, we propose a new token context-aware tracking pipeline named LMTrack, designed to automatically learn high-quality reference tokens for efficient visual tracking. Embracing the principle of Less is More, the core idea of LMTrack is to analyze the importance distribution of all reference tokens, where important tokens are collected, continually attended to, and updated. Specifically, a novel Token Context Memory module is designed to dynamically collect high-quality spatio-temporal information of a target in an autoregressive manner, eliminating redundant background tokens from the reference frames. Furthermore, an effective Unidirectional Token Attention mechanism is designed to establish dependencies between reference tokens and search frame, enabling robust cross-frame association and target localization. Extensive experiments demonstrate the superiority of our tracker, achieving state-of-the-art results on tracking benchmarks such as GOT-10K, TrackingNet, and LaSOT. Chenlong Xu, Bineng Zhong 0001, Qihua Liang, Yaozong Zheng, Guorong Li, Shuxiang Song 0001 |
AAAI | 4 |
| 2025 | Decoupled Spatio-Temporal Consistency Learning for Self-Supervised TrackingabstractThe success of visual tracking has been largely driven by datasets with manual box annotations. However, these box annotations require tremendous human effort, limiting the scale and diversity of existing tracking datasets. In this work, we present a novel Self-Supervised Tracking framework, named SSTrack, designed to eliminate the need of box annotations. Specifically, a decoupled spatio-temporal consistency training framework is proposed to learn rich target information across timestamps through global spatial localization and local temporal association. This allows for the simulation of appearance and motion variations of instances in real-world scenarios. Furthermore, an instance contrastive loss is designed to learn instance-level correspondences from a multi-view perspective, offering robust instance supervision without additional labels. This new design paradigm enables SSTrack to effectively learn generic tracking representations in a self-supervised manner, while reducing reliance on extensive box annotations. Extensive experiments on nine benchmark datasets demonstrate that SSTrack surpasses SOTA self-supervised tracking methods, achieving an improvement of more than 25.3%, 20.4%, and 14.8% in AUC (AO) score on the GOT10K, LaSOT, TrackingNet datasets, respectively. Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Ning Li 0044, Shuxiang Song 0001 |
AAAI | 1 |
| 2025 | Similarity-Guided Layer-Adaptive Vision Transformer for UAV TrackingabstractVision transformers (ViTs) have emerged as a popular backbone for visual tracking. However, complete ViT architectures are too cumbersome to deploy for unmanned aerial vehicle (UAV) tracking which extremely emphasizes efficiency. In this study, we discover that many layers within lightweight ViT-based trackers tend to learn relatively redundant and repetitive target representations. Based on this observation, we propose a similarity-guided layer adaptation approach to optimize the structure of ViTs. Our approach dynamically disables a large number of representation-similar layers and selectively retains only a single optimal layer among them, aiming to achieve a better accuracy-speed trade-off. By incorporating this approach into existing ViTs, we tailor previously complete ViT architectures into an efficient similarity-guided layer-adaptive framework, namely SGLATrack, for real-time UAV tracking. Extensive experiments on six tracking benchmarks verify the effectiveness of the proposed approach, and show that our SGLATrack achieves a state-of-the-art real-time speed while maintaining competitive tracking precision. Codes and models are available at https://github.com/GXNU-ZhongLab/SGLATrack. Chaocan Xue, Bineng Zhong 0001, Qihua Liang, Yaozong Zheng, Ning Li 0044, Yuanliang Xue, Shuxiang Song 0001 |
CVPR | 4 |
| 2025 | Towards Universal Modal Tracking With Online Dense Temporal Token LearningabstractWe propose a universal video-level modality-awareness tracking model with online dense temporal token learning (called UM-ODTrack). It is designed to support various tracking tasks, including RGB, RGB+Thermal, RGB+Depth, and RGB+Event, utilizing the same model architecture and parameters. Specifically, our model is designed with three core goals: Video-level Sampling. We expand the model's inputs to a video sequence level, aiming to see a richer video context from an near-global perspective. Video-level Association. Furthermore, we introduce two simple yet effective online dense temporal token association mechanisms to propagate the appearance and motion trajectory information of target via a video stream manner. Modality Scalable. We propose two novel gated perceivers that adaptively learn cross-modal representations via a gated attention mechanism, and subsequently compress them into the same set of model parameters via a one-shot training manner for multi-task inference. This new solution brings the following benefits: (i) The purified token sequences can serve as temporal prompts for the inference in the next video frames, whereby previous information is leveraged to guide future inference. (ii) Unlike multi-modal trackers that require independent training, our one-shot training scheme not only alleviates the training burden, but also improves model representation. Extensive experiments on visible and multi-modal benchmarks show that our UM-ODTrack achieves a new SOTA performance. Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Shengping Zhang, Guorong Li, Xianxian Li, Rongrong Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | AVLTrack: Dynamic Sparse Learning for Aerial Vision-Language TrackingabstractThe introduction of natural language for vision-language (VL) tracking has been proven to improve performance. However, natural language remains under-explored in existing aerial trackers. Moreover, existing VL trackers ignore the misalignment of language with dynamic target states, which is prominent in complex UAV scenarios. In this work, we present AVLTrack, a flexible framework for aerial vision-language tracking. It consists of three key components, a dynamic sparse learning (DSL) module, an efficient Transformer backbone, and a multi-level language perception (MLP) strategy. First, DSL sparsely connects language and images via dynamic sparse attention, providing accurate multi-modal prompts. To adapt to target state variations, the sparsity in DSL is dynamically adjusted based on semantic information, flexibly highlighting target-specific tokens. Next, the Transformer backbone follows highly parallelized one-stream architectures, allowing efficient multi-modal feature extraction and interaction. Finally, MLP enables the iterative interaction of language and visual information, aiming to utilize language priori to guide the generation of discriminative visual features. Moreover, we construct the DTB70-NLP dataset to facilitate UAV vision-language tracking. Extensive experiments on WebUAV-3M and DTB70-NLP demonstrate the leading performance of AVLTrack compared to existing outstanding trackers while maintaining a high running speed of 80.5 FPS. The dataset and codes are available athttps://github.com/xyl-507/AVLTrack. Yuanliang Xue, Bineng Zhong 0001, Guodong Jin, Lining Tan, Ning Li 0044, Yaozong Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Adaptive Expert Decision for RGB-T TrackingabstractThe features provided by RGB and Thermal Infrared (TIR) images have their own characteristics. Therefore, how to adaptively fuse multi-modal features according to different tracking scenarios is crucial for RGB-T tracking. However, current mainstream RGB-T tracking algorithms often use fixed fusion operations for modal interaction in different scenarios. Consequently, their tracking permanence is deteriorated due to they are unable to dynamically adjust the fused multi-modal features based on the current scenes. To address this issue, we propose a novel RGB-T tracking algorithm called AETrack, which can dynamically extract effective modal features in different scenarios for adaptive fusion. Firstly, we design an adaptive expert decision mechanism that employs multiple experts to process the input features. Each expert focuses on and learns different relevant features. Based on this mechanism, we then propose a feature-guided method that leverages the correlations between modalities to provide cross-modal information. This guidance enables the adaptive expert mechanism to adaptively select the most suitable expert to output effective features based on different scenarios, ensuring that our proposed AETrack prioritizes effective features and thus alleviates interference from irrelevant information. Finally, we design a Progressive Cross-modal Fusion operation to achieve multi-level adaptive fusion of effective features across different modalities. Benefiting from this adaptive fusion process, we can effectively achieve multi-modal interaction in different scenarios to guide robust tracking. Extensive experiments on three popular benchmarks (i.e., LasHeR, RGBT210, RGBT234) show that our proposed AETrack can significantly improve tracking performance. Zhiruo Zhu, Bineng Zhong 0001, Qihua Liang, Yaozong Zheng, Ning Li 0044 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | ODTrack: Online Dense Temporal Token Learning for Visual TrackingabstractOnline contextual reasoning and association across consecutive video frames are critical to perceive instances in visual tracking. However, most current top-performing trackers persistently lean on sparse temporal relationships between reference and search frames via an offline mode. Consequently, they can only interact independently within each image-pair and establish limited temporal correlations. To alleviate the above problem, we propose a simple, flexible and effective video-level tracking pipeline, named ODTrack, which densely associates the contextual relationships of video frames in an online token propagation manner. ODTrack receives video frames of arbitrary length to capture the spatio-temporal trajectory relationships of an instance, and compresses the discrimination features (localization information) of a target into a token sequence to achieve frame-to-frame association. This new solution brings the following benefits: 1) the purified token sequences can serve as prompts for the inference in the next video frame, whereby past information is leveraged to guide future inference; 2) the complex online update strategies are effectively avoided by the iterative propagation of token sequences, and thus we can achieve more efficient model representation and computation. ODTrack achieves a new SOTA performance on seven benchmarks, while running at real-time speed. Code and models are available at https://github.com/GXNU-ZhongLab/ODTrack. Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shengping Zhang, Xianxian Li |
AAAI | 1 |
| 2024 | Personalized federated learning based on feature fusionabstractFederated learning (FL) enables distributed clients to collaborate on training while storing their data locally to protect client privacy. However, due to data heterogeneity, including issues related to label distributions skew in heterogeneous scenarios, the resulting global model may not be suitable for all clients. In this work, we introduce a personalized federated learning method called pFedPM, which focuses on addressing this challenge of label distributions skew in heterogeneous scenarios. We replace traditional gradient uploading with feature uploading, and introduce a novel feature fusion scheme to learn personalized local model for clients. Specifically, the server receives feature information from clients, aggregates global features, and sends them back to the clients. Clients achieve personalization by fusing local and global features. Furthermore, we introduce a relation network as an additional decision layer, providing a non-linear learnable classifier to predict labels. Through the novel modeling techniques, our proposed method reduces communication costs and supports heterogeneous client models. Experimental results demonstrate that our approach outperforms recent FL methods on the MNIST, FEMNIST, and CIFAR-10 datasets while requiring less communication. Wolong Xing, Zhenkui Shi, Hongyan Peng, Xiantao Hu, Yaozong Zheng, Xianxian Li |
CSCWD | 5 |
| 2024 | Exploration of Railway Signal Unloading Task Based on Deep Reinforcement Learning Method
Ting Ke, Yaozong Zheng, Zhanshuo Liu, Jianan Shen |
ICIC (2) | 2 |
| 2024 | A general maximal margin hyper-sphere SVM for multi-class classification
Ting Ke, Xuechun Ge, Feifei Yin, Yaozong Zheng, Chuanlei Zhang, Jianrong Li |
Expert Syst. Appl. | 5 |
| 2024 | Top-Down Cross-Modal Guidance for Robust RGB-T TrackingabstractMost RGB-T trackers heavily rely on bottom-up attention and thus overlook top-down cross-modal guidance for learning target features. Consequently, the discriminative power of the learnt target features is weak. To address this issue, we propose a novel RGB-T tracker (called TGTrack) that designs a Top-down Cross-modal Guidance mechanism to learn target features in two stages. In the first stage, our TGTrack effectively generates top-down cross-modal guidance signals with multi-modal encoders-decoders and prior vectors. In the second stage, these signals are transmitted and integrated to improve the discriminative power of our target features by the attention layers of the cross-modal encoders. Moreover, we introduce an Attention-Driven Spatio-Temporal Updater for updating discriminative target features. Through cross-frame attention guidance, it can effectively eliminates irrelevant features within the search region. As a result, our TGTrack can effectively avoid the complex multi-modal fusion modules and thus achieve robust RGB-T tracking. Extensive experiments on three popular RGB-T tracking benchmarks (i.e., LasHeR, RGBT234, and RGBT210) demonstrate that our TGTrack achieves new state-of-the-art performances. Bineng Zhong 0001, Qihua Liang, Yaozong Zheng, Zhiyi Mo, Shuxiang Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Robust Tracking via Combing Top-Down and Bottom-Up AttentionabstractTransformer attention plays an important role in current top-performing trackers. However, it is bottom-up, driven by stimulus and lacks intrinsic prior guidance. This bottom-up attention mechanism leads to an emphasis on all objects in the input images, rather than the task related objects. As a result, the performance of the bottom-up attention based trackers is deteriorated in complicated scenes. To address this issue, we propose a robust tracker that combines bottom-up attention with top-down attention to comply with the existing ViT framework, named TBTrack. TBTrack can not only utilize the existing bottom-up attention mechanisms to model the long-range relationship of input tokens, but also utilize a newly added top-down attention mechanism to pay more attention to task related object and further eliminate interference from similar objects and backgrounds. Specifically, we firstly design a top-down prior generation module using an adaptive learning parameter combined with the template inputs to obtain top-down task guided signals. Then, we inject the prior signals into a bottom-up attention module to obtain a top-down and bottom-up attention combination block (TB-Block). Finally, we stack these TB-Blocks to construct our tracker (TBTrack) with top-down prior guidance capability, which focuses more on the task related object. Through extensive experiments, our TBTrack achieves impressive performance on multiple tracking benchmarks, including GOT-10k, LaSOT, LaSOText, TNL2K, TrackingNet, UAV123 and so on. The code and trained models will be publicly available. Ning Li 0044, Bineng Zhong 0001, Yaozong Zheng, Qihua Liang, Zhiyi Mo, Shuxiang Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Toward Unified Token Learning for Vision-Language TrackingabstractIn this paper, we present a simple, flexible and effective vision-language (VL) tracking pipeline, termed MMTrack, which casts VL tracking as a token generation task. Traditional paradigms address VL tracking task indirectly with sophisticated prior designs, making them over-specialize on the features of specific architectures or mechanisms. In contrast, our proposed framework serializes language description and bounding box into a sequence of discrete tokens. In this new design paradigm, all token queries are required to perceive the desired target and directly predict spatial coordinates of the target in an auto-regressive manner. The design without other prior modules avoids multiple sub-tasks learning and hand-designed loss functions, significantly reducing the complexity of VL tracking modeling and allowing our tracker to use a simple cross-entropy loss as unified optimization objective for VL tracking task. Extensive experiments on TNL2K, LaSOT, LaSOT$_{\mathrm{ext}}$and OTB99-Lang benchmarks show that our approach achieves promising results, compared to other state-of-the-arts. Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Guorong Li, Rongrong Ji, Xianxian Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Leveraging Local and Global Cues for Visual Tracking via Parallel Interaction NetworkabstractDespite that both local and context information are crucial for robust tracking, existing CNN-based and transformer-based methods mainly focus on one of these aspects. Consequently, the former fails to exploit rich global context information due to the limited receptive field, while the latter suffers from the deficiencies in constructing the local relationship among neighboring regions. To address this issue, we propose the SiamPIN tracker, based on our Parallel Interaction Network. It consists of two effective modules, namely Global Aggregation Block (GAB) and Local Process Block (LPB). GAB perceives the global context to capture the long-range spatial dependency through a transformer-based architecture. Meanwhile, LPB performs local information extraction using a CNN model to retain the detailed appearance information of the target. These two modules are connected consecutively to compose a Trans-Conv unit block, which transmits the global context information to the local feature extraction procedure, hence enables the interaction of global-local information flow. Several such blocks are cascaded so that our model can learn to aggregate local and context information interactively. The proposed tracker achieves state-of-the-art performance on six benchmark datasets, while maintaining a real time running speed. Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Zhenjun Tang, Rongrong Ji, Xianxian Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |