VLDB 2026 Research / reviewers in the wild / expert
Guangtong Zhang
dblp:363/8765
· DBLP profile ↗
8ranked-venue papers
6as first author
8since 2021 · last 2026
0009-0001-1513-0313ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Aware Distillation for Robust Vision-Language Tracking Under Linguistic SparsityabstractVision-language object tracking overcomes the limitations of relying solely on visual features by leveraging language descriptions of objects to provide cross-modal semantic information, thereby enhancing model robustness in complex scenarios. However, most existing high-performance vision-language trackers are trained jointly on pure visual data and vision-language multimodal data. Due to the relative sparsity of language annotations in the data, the trackers tend to prioritize the localization role of visual features, diminishing the model's attention to language information. To mitigate this issue, we propose a novel vision-language tracker: Aware Distillation for Robust Vision-Language Tracking under Linguistic Sparsity (ADTrack). We introduce a knowledge distillation framework employing a knowledge-rich teacher model and a lightweight student model to establish modality correlations between vision and language, enabling efficient modeling between visual information and language descriptions. Specifically, our lightweight student module simultaneously distills language encoding capabilities from large language models through teacher-guided learning on input language, while performing target-aware perception on template images using language descriptions to generate more effective template features for subsequent visual extraction. Furthermore, to ensure perceptual robustness in linguistically sparse scenarios, we simulate language-deficient conditions during training and employ contrastive learning to enhance model adaptability. Extensive experiments demonstrate that ADTrack reduces parameters by over 50% while achieving state-of-the-art (SOTA) performance and speed on vision-language tracking benchmarks, including LaSOT, LaSOText, TNL2K, OTB-Lang and MGIT. Guangtong Zhang, Bineng Zhong 0001, Shirui Yang, Tian Bai 0002 |
AAAI | 1 |
| 2026 | Selective distillation of language tokens for redundancy suppression in vision-language tracking
Tian Bai 0002, Shirui Yang, Guangtong Zhang |
Expert Syst. Appl. | 4 |
| 2026 | RWKV-Inspired Multi-Modal Relation Modeling for Vision-Language TrackingabstractVision-language object tracking can provide more state representations for targets by introducing the language modality, achieving more robust tracking and localization. Therefore, designing multi-modal interactions to achieve feature alignment between vision and language has been one of the research hotspots. However, existing multi-modal interaction methods face two key issues: on the one hand, they lack effective exploration of modeling the relationship between the contextual information of language sequences and visual features; on the other hand, the introduction of modalities leads to increased computational time costs in multi-modal interactions, which severely affects the real-time performance of vision-language tracking algorithms. To address these challenges, we propose a vision-language tracking framework called RWKV-Inspired Multi-modal Relation Modeling for Vision-Language Tracking (RrmTrack). We introduce a novel modality interaction method specific to vision-language object tracking based on RWKV, providing customized interaction for different modalities in vision-language tracking and effectively reducing the computational time cost of cross-modal interaction. Specifically, this method uses a time mixing module to model the relationship between language information and image features, and a channel mixing module to facilitate information interaction between images. By combining parallelized training with a linear attention mechanism and efficient RNN inference, it enables accurate and fast target localization in vision-language tracking. Additionally, we propose a novel feature extraction structure that integrates Siamese and One-stream architectures. An information restoration module is designed to reduce the information interference introduced by the search image to the template image during interaction. RrmTrack achieves state-of-the-art results and speed on multiple vision-language object tracking benchmarks, including TNL2k, LaSOT, OTB-Lang, LaSOText, and MGIT. Guangtong Zhang, Bineng Zhong 0001, Yuhao Mu, Tian Bai 0002 |
IEEE Trans. Multim. | 1 |
| 2024 | Visual Adapt for RGBD TrackingabstractRecent RGBD trackers have employed cueing techniques by overlaying Depth modality images as cues onto RGB modality images, which are then fed into the RGB-based model for tracking. However, the direct overlaying interaction method between modalities not only introduces more noise into the feature space but also exhibits the inadaptability of the RGB-based model to mixed-modality inputs. To address these issues, we introduce Visual Adapt for RGBD Tracking (VADT). Specifically, we maintain the input of the RGB-based model as the RGB modality. Additionally, we have devised a fusion module to enable modality interaction between depth and RGB features. Subsequently, a Depth Adapt module has been formulated to facilitate image interaction with the fused features. This module involves cross-attending to the obtained depth-assisted features and the RGB search frame features produced by the RGB-based model’s output. Experimental results indicate that our proposed tracker achieves state-of-the-art results on various RGBD benchmark tests. Guangtong Zhang, Qihua Liang, Zhiyi Mo, Ning Li 0044, Bineng Zhong 0001 |
ICASSP | 1 |
| 2024 | Diffusion Mask-Driven Visual-language Tracking
Guangtong Zhang, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shuxiang Song 0001 |
IJCAI | 1 |
| 2024 | Dual-stream Multi-modal Interactive Vision-language Tracking
Zhiyi Mo, Guangtong Zhang, Jian Nong, Bineng Zhong 0001, Zhi Li 0017 |
MMAsia | 2 |
| 2024 | One-Stream Stepwise Decreasing for Vision-Language TrackingabstractBased on the fixed language descriptions in the initial frames, a vision-language tracker typically adopts a two-stream model structure to align vision and language features at the feature fusion stages. However, this paradigm may degrade the tracking performance due to inaccurate language descriptions and lacks further modal interaction. To address these issues, we propose a one-stream vision-language model called One-stream Stepwise Decreasing for Vision-Language Tracking (OSDT). Specifically, we first encode the language description using a language encoder. The obtained language features are then combined with visual images and entered jointly into a visual encoder, in which the encoder’s self-attention mechanism is utilized to facilitate more interactions between language and visual features. Moreover, to mitigate the problems caused by inaccurate language descriptions, we design a stepwise decreasing multi-modal interaction framework, in which a Feature Filter Module (FFM) is introduced to select language features that are more relevant to visual information to provide semantic guidance for visual feature extraction. Furthermore, without additional feature fusion modules, our one-stream model framework can efficiently utilize the proposed feature filtering module for feature selection. Consequently, our tracker can achieve fast tracking speed in the vision-language tracking domain compared to existing state-of-the-art methods. We extensively evaluate our tracker on three benchmarks, i.e. TNL2K, LaSOT, and OTB99, demonstrating competing performance compared to state-of-the-art vision-language tracking methods. Guangtong Zhang, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Ning Li 0044, Shuxiang Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Robust Tracking via Unifying Pretrain-Finetuning and Visual Prompt TuningabstractThe finetuning paradigm has been a widely used methodology for the supervised training of top-performing trackers. However, the finetuning paradigm faces one key issue: it is unclear how best to perform the finetuning method to adapt a pretrained model to tracking tasks while alleviating the catastrophic forgetting problem. To address this problem, we propose a novel partial finetuning paradigm for visual tracking via unifying pretrain-finetuning and visual prompt tuning (named UPVPT), which can not only efficiently learn knowledge from the tracking task but also reuse the prior knowledge learned by the pre-trained model for effectively handling various challenges in tracking task. Firstly, to maintain the pre-trained prior knowledge, we design a Prompt-style method to freeze some parameters of the pretrained network. Then, to learn knowledge from the tracking task, we update the parameters of the prompt and MLP layers. As a result, we cannot only retain useful prior knowledge of the pre-trained model by freezing the backbone network but also effectively learn target domain knowledge by updating the Prompt and MLP layer. Furthermore, the proposed UPVPT can easily be embedded into existing Transformer trackers (e.g., OSTracker and SwinTracker) by adding only a small number of model parameters (less than 1% of a Backbone network). Extensive experiments on five tracking benchmarks (i.e., UAV123, GOT-10k, LaSOT, TNL2K, and TrackingNet) demonstrate that the proposed UPVPT can improve the robustness and effectiveness of the model, especially in complex scenarios. Guangtong Zhang, Qihua Liang, Ning Li 0044, Zhiyi Mo, Bineng Zhong 0001 |
MMAsia | 1 |