VLDB 2026 Research / reviewers in the wild / expert
Qihua Liang
dblp:304/1822
· DBLP profile ↗
35ranked-venue papers
1as first author
35since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 31 since 2021Artificial intelligence and machine learning · 13 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MUTrack: A Memory-Aware Unified Representation Framework for Visual TrackingabstractBuilding a unified target representation that simultaneously achieves short-term adaptability and long-term stability is crucial for robust visual tracking. However, existing trackers typically face an inherent trade-off. Methods primarily relying on short-term appearance and motion cues achieve rapid adaptation, but they often struggle with long-term identity consistency. Conversely, trackers that emphasize extensive temporal context provide strong robustness, yet this approach can compromise their short-term adaptability. To bridge this gap, we propose a novel tracker, MUTrack, which comprehensively integrates both long-term and short-term memories into a unified target representation for more robust tracking. Specifically, we design a unified memory bank that stores and manages long-term memory for maintaining long-term identity consistency, and short-term memory for adapting to instantaneous appearance changes. To fully leverage the complementary nature of both long-term and short-term temporal information, we introduce a perception interaction module that dynamically fuses these memory types through deep and bidirectional interactions, enabling mutual refinement where one guides the other. This ultimately generates a highly adaptive target representation, which effectively balances adaptability to instantaneous changes with robustness against long-term identity drift. Extensive experiments on GOT10k, TrackingNet, LaSOT, LaSOT_ext, NfS, and OTB100 consistently demonstrate that MUTrack achieves SOTA performance. Weijing Wu, Qihua Liang, Bineng Zhong 0001, Yufei Tan, Ning Li 0044, Yuanliang Xue |
AAAI | 2 |
| 2026 | Motion-Aware Object Tracking via Motion and Geometry-Aware CuesabstractUnderstanding motion is essential for visual object tracking, especially in complex and dynamic scenarios. Yet, many existing methods rely on simplistic strategies such as template updates or temporal feature propagation, often overlooking the deeper modeling of motion information. To mitigate this limitation, we introduce a motion-aware spatio-temporal framework that enhances motion perception by explicitly matching motion patterns and modeling inter-frame motion relationships. Central to our design is a motion pattern dictionary, which encodes a diverse set of representative motion cues as learnable features. During tracking, features from the search region interact with the dictionary to retrieve the most relevant motion patterns, allowing the model to adapt to the current motion state. A dedicated decoder further incorporates temporal correlations to refine motion awareness. To complement motion modeling, we embed geometric cues into the search region features, which strengthens spatial perception, reduces ambiguity under occlusion, and improves foreground-background separation. Extensive evaluations on seven challenging benchmarks demonstrate the effectiveness of our design. In particular, MoDTrack_384 surpasses recent SOTA trackers on LaSOT by 1.2% in AUC, highlighting the benefits of motion pattern modeling and geometry-guided enhancement in mitigating tracking drift. Bineng Zhong 0001, Qihua Liang, Xiantao Hu, Yufei Tan, Haiying Xia, Shuxiang Song 0001 |
AAAI | 3 |
| 2026 | Unified Multi-Modal Tracking via Proxy Visual PromptsabstractRecent unified multi-modal tracking frameworks often encounter high computational overhead due to complex fusion operations. In this paper, we propose PoATrack, a proxy-based bidirectional fusion framework designed to facilitate efficient multi-modal tracking via interactive proxy prompts. Specifically, the framework treats features from all modalities equally and introduces an adaptive feature enhancement module to improve spatial representations within the search region. To further reduce the fusion cost, a lightweight proxy prompt fusion module is developed following a two-stage strategy: (1) representing modality-specific features through proxy visual prompts, and (2) dynamically learning hierarchical cross-modal relationships via proxy-guided fusion for enriched contextual modeling. Extensive experiments on six public benchmarks (RGB-T, RGB-E, and RGB-D) demonstrate that PoATrack achieves competitive accuracy while operating 1.8× faster than SDSTrack [17] under identical hardware settings. Bineng Zhong 0001, Yaozong Zheng, Qihua Liang, Zhiruo Zhu, Shuxiang Song 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2026 | Mamba-Driven Diffusion Model for Salient Object Detection in Optical Remote Sensing ImagesabstractExisting Optical Remote Sensing Image Salient Object Detection (ORSI-SOD) methods mainly rely on a semantic segmentation paradigm, which relies on pixel-wise probabilities, leading to overconfident mispredictions. In contrast, the random sampling process of the diffusion model allows multiple possible predictions to be drawn from the mask distribution, effectively alleviating this problem. However, existing diffusion models mainly use Transformers as conditional feature extraction networks. Although they are good at global modeling, they have limited ability to handle long-range dependencies due to computational complexity. To overcome these challenges, we introduce MambaDif, an innovative diffusion model architecture based on Mamba. Specifically, we regard ORSI-SOD as a conditional mask generation task leveraging the diffusion model and achieving target distribution matching by adding noise to the mask and iteratively denoising it to match the target distribution. Then, we adopt Mamba to extract global features, efficiently process long sequences, and capture global contextual information with linear complexity. In addition, we introduce the global-local feature collaborative completion module (GLM), which combines the ability of convolutional layers to extract local features with the advantage of Mamba in capturing long-range dependencies, thereby achieving excellent denoising performance. Extensive experiments show that MambaDif outperforms SOTA methods in eight evaluation metrics on two standard datasets (EORSSD and ORSSD). We also report the generalization performance of the model on the challenging ORSI-4199 to evaluate its robustness. Bineng Zhong 0001, Qihua Liang, Yufei Tan, Haiying Xia, Shuxiang Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Robust RGB-T Tracking via Learnable Visual Fourier Prompt Fine-Tuning and Modality Fusion Prompt GenerationabstractRecently, visual prompt tuning is introduced to RGB-Thermal (RGB-T) tracking as a parameter-efficient finetuning (PEFT) method. However, these PEFT-based RGB-T tracking methods typically rely solely on spatial domain information as prompts for feature extraction. As a result, they often fail to achieve optimal performance by overlooking the crucial role of frequency-domain information in prompt learning. To address this issue, we propose an efficient Visual Fourier Prompt Tracking (named VFPTrack) method to learn modality-related prompts via Fast Fourier Transform (FFT). Our method consists of symmetric feature extraction encoder with shared parameters, visual-fourier prompts, and Modality Fusion Prompt Generator that generates bidirectional interaction prompts through multi-modal feature fusion. Specifically, we first use a frozen feature extraction encoder to extract RGB and thermal infrared (TIR) modality features. Then, we combine the visual prompts in the spatial domain with the frequency domain prompts obtained from the FFT, which allows for the full extraction and understanding of modality features from different domain information. Finally, unlike previous fusion methods, the modality fusion prompt generation module we use combines features from different modalities to generate a fused modality prompt. This modality prompt is interacted with each individual modality to fully enable feature interaction across different modalities. Extensive experiments conducted on three popular RGB-T tracking benchmarks show that our method demonstrates outstanding performance. Bineng Zhong 0001, Qihua Liang, Zhiruo Zhu, Yaozong Zheng, Ning Li 0044 |
IEEE Trans. Multim. | 3 |
| 2025 | MambaLCT: Boosting Tracking via Long-term Context State Space ModelabstractEffectively constructing context information with long-term dependencies from video sequences is crucial for object tracking. However, the context length constructed by existing work is limited, only considering object information from adjacent frames or video clips, leading to insufficient utilization of contextual information. To address this issue, we propose MambaLCT, which constructs and utilizes target variation cues from the first frame to the current frame for robust tracking. First, a novel unidirectional Context Mamba module is designed to scan frame features along the temporal dimension, gathering target change cues throughout the entire sequence. Specifically, target-related information in frame features is compressed into a hidden state space through a selective scanning mechanism. The target information across the entire video is continuously aggregated into target variation cues. Next, we inject the target change cues into the attention mechanism, providing temporal information for modeling the relationship between the template and search frames. The advantage of MambaLCT is its ability to continuously extend the length of the context, capturing complete target change cues, which enhances the stability and robustness of the tracker. Extensive experiments show that long-term context information enhances the model's ability to perceive targets in complex scenarios. MambaLCT achieves new SOTA performance on six benchmarks while maintaining real-time runing speeds. Xiaohai Li, Bineng Zhong 0001, Qihua Liang, Guorong Li, Zhiyi Mo, Shuxiang Song 0001 |
AAAI | 3 |
| 2025 | Robust Tracking via Mamba-based Context-aware Token LearningabstractHow to make a good trade-off between performance and computational cost is crucial for a tracker. However, current famous methods typically focus on complicated and time-consuming learning that combining temporal and appearance information by input more and more images (or features). Consequently, these methods not only increase the model's computational source and learning burden but also introduce much useless and potentially interfering information. To alleviate the above issues, we propose a simple yet robust tracker that separates temporal information learning from appearance modeling and extracts temporal relations from a set of representative tokens rather than several images (or features). Specifically, we introduce one track token for each frame to collect the target's appearance information in the backbone. Then, we design a mamba-based Temporal Module for track tokens to be aware of context by interacting with other track tokens within a sliding window. This module consists of a mamba layer with autoregressive characteristic and a cross-attention layer with strong global perception ability, ensuring sufficient interaction for track tokens to perceive the appearance changes and movement trends of the target. Finally, track tokens serve as a guidance to adjust the appearance feature for the final prediction in the head. Experiments show our method is effective and achieves competitive performance on multiple benchmarks at a real-time speed. Jinxia Xie, Bineng Zhong 0001, Qihua Liang, Ning Li 0044, Zhiyi Mo, Shuxiang Song 0001 |
AAAI | 3 |
| 2025 | Less Is More: Token Context-Aware Learning for Object TrackingabstractRecently, several studies have shown that utilizing contextual information to perceive target states is crucial for object tracking. They typically capture context by incorporating multiple video frames. However, these naive frame-context methods fail to consider the importance of each patch within a reference frame, making them susceptible to noise and redundant tokens, which deteriorates tracking performance. To address this challenge, we propose a new token context-aware tracking pipeline named LMTrack, designed to automatically learn high-quality reference tokens for efficient visual tracking. Embracing the principle of Less is More, the core idea of LMTrack is to analyze the importance distribution of all reference tokens, where important tokens are collected, continually attended to, and updated. Specifically, a novel Token Context Memory module is designed to dynamically collect high-quality spatio-temporal information of a target in an autoregressive manner, eliminating redundant background tokens from the reference frames. Furthermore, an effective Unidirectional Token Attention mechanism is designed to establish dependencies between reference tokens and search frame, enabling robust cross-frame association and target localization. Extensive experiments demonstrate the superiority of our tracker, achieving state-of-the-art results on tracking benchmarks such as GOT-10K, TrackingNet, and LaSOT. Chenlong Xu, Bineng Zhong 0001, Qihua Liang, Yaozong Zheng, Guorong Li, Shuxiang Song 0001 |
AAAI | 3 |
| 2025 | Decoupled Spatio-Temporal Consistency Learning for Self-Supervised TrackingabstractThe success of visual tracking has been largely driven by datasets with manual box annotations. However, these box annotations require tremendous human effort, limiting the scale and diversity of existing tracking datasets. In this work, we present a novel Self-Supervised Tracking framework, named SSTrack, designed to eliminate the need of box annotations. Specifically, a decoupled spatio-temporal consistency training framework is proposed to learn rich target information across timestamps through global spatial localization and local temporal association. This allows for the simulation of appearance and motion variations of instances in real-world scenarios. Furthermore, an instance contrastive loss is designed to learn instance-level correspondences from a multi-view perspective, offering robust instance supervision without additional labels. This new design paradigm enables SSTrack to effectively learn generic tracking representations in a self-supervised manner, while reducing reliance on extensive box annotations. Extensive experiments on nine benchmark datasets demonstrate that SSTrack surpasses SOTA self-supervised tracking methods, achieving an improvement of more than 25.3%, 20.4%, and 14.8% in AUC (AO) score on the GOT10K, LaSOT, TrackingNet datasets, respectively. Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Ning Li 0044, Shuxiang Song 0001 |
AAAI | 3 |
| 2025 | Dynamic Updates for Language Adaptation in Visual-Language TrackingabstractThe consistency between the semantic information provided by the multi-modal reference and the tracked object is crucial for visual-language (VL) tracking. However, existing VL tracking frameworks rely on static multi-modal references to locate dynamic objects, which can lead to semantic discrepancies and reduce the robustness of the tracker. To address this issue, we propose a novel vision-language tracking framework, named DUTrack, which captures the latest state of the target by dynamically updating multimodal references to maintain consistency. Specifically, we introduce a Dynamic Language Update Module, which leverages a large language model to generate dynamic language descriptions for the object based on visual features and object category information. Then, we design a Dynamic Template Capture Module, which captures the regions in the image that highly match the dynamic language descriptions. Furthermore, to ensure the efficiency of description generation, we design an update strategy that assesses changes in target displacement, scale, and other factors to decide on updates. Finally, the dynamic template and language descriptions that record the latest state of the target are used to update the multi-modal references, providing more accurate reference information for subsequent inference and enhancing the robustness of the tracker. DUTrack achieves new state-of-the-art performance on five mainstream vision-language and two vision-only tracking benchmarks, including LaSOT, LaSOText, TNL2K, OTB99-Lang, MGIT, GOT-10K, and UAV123. Code and models are available at https://github.com/GXNU-ZhongLab/DUTrack. Xiaohai Li, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Jian Nong, Shuxiang Song 0001 |
CVPR | 3 |
| 2025 | Similarity-Guided Layer-Adaptive Vision Transformer for UAV TrackingabstractVision transformers (ViTs) have emerged as a popular backbone for visual tracking. However, complete ViT architectures are too cumbersome to deploy for unmanned aerial vehicle (UAV) tracking which extremely emphasizes efficiency. In this study, we discover that many layers within lightweight ViT-based trackers tend to learn relatively redundant and repetitive target representations. Based on this observation, we propose a similarity-guided layer adaptation approach to optimize the structure of ViTs. Our approach dynamically disables a large number of representation-similar layers and selectively retains only a single optimal layer among them, aiming to achieve a better accuracy-speed trade-off. By incorporating this approach into existing ViTs, we tailor previously complete ViT architectures into an efficient similarity-guided layer-adaptive framework, namely SGLATrack, for real-time UAV tracking. Extensive experiments on six tracking benchmarks verify the effectiveness of the proposed approach, and show that our SGLATrack achieves a state-of-the-art real-time speed while maintaining competitive tracking precision. Codes and models are available at https://github.com/GXNU-ZhongLab/SGLATrack. Chaocan Xue, Bineng Zhong 0001, Qihua Liang, Yaozong Zheng, Ning Li 0044, Yuanliang Xue, Shuxiang Song 0001 |
CVPR | 3 |
| 2025 | Towards Universal Modal Tracking With Online Dense Temporal Token LearningabstractWe propose a universal video-level modality-awareness tracking model with online dense temporal token learning (called UM-ODTrack). It is designed to support various tracking tasks, including RGB, RGB+Thermal, RGB+Depth, and RGB+Event, utilizing the same model architecture and parameters. Specifically, our model is designed with three core goals: Video-level Sampling. We expand the model's inputs to a video sequence level, aiming to see a richer video context from an near-global perspective. Video-level Association. Furthermore, we introduce two simple yet effective online dense temporal token association mechanisms to propagate the appearance and motion trajectory information of target via a video stream manner. Modality Scalable. We propose two novel gated perceivers that adaptively learn cross-modal representations via a gated attention mechanism, and subsequently compress them into the same set of model parameters via a one-shot training manner for multi-task inference. This new solution brings the following benefits: (i) The purified token sequences can serve as temporal prompts for the inference in the next video frames, whereby previous information is leveraged to guide future inference. (ii) Unlike multi-modal trackers that require independent training, our one-shot training scheme not only alleviates the training burden, but also improves model representation. Extensive experiments on visible and multi-modal benchmarks show that our UM-ODTrack achieves a new SOTA performance. Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Shengping Zhang, Guorong Li, Xianxian Li, Rongrong Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | SIEVL-Track: Exploring Semantic Information Enhancement for Visual-Language Object TrackingabstractWith the assistance of language descriptions, Visual-Language (VL) object tracking can obtain more accurate semantic information compared to traditional Visual-Only object tracking. However, the ability of current VL trackers to obtain target semantic information has not been fully developed due to limitations such as wasted modeling capabilities and insufficient utilization of historical temporal information. On the one hand, the modeling output from Transformer shallow encoders often does not directly participate in the prediction of tracking results, resulting in a certain degree of model capability waste. On the other hand, the semantic information of historical tracking results has also not been fully utilized in the tracking process, resulting in a certain degree of lack of semantic assistance capability. Therefore, we propose a novel hierarchical multi-stage VL tracker called SIEVL-Track to enhance target semantic information. Specifically, we first design a multi-stage visual language tracking framework for modeling multi-scale semantic information in Visual-Language tracking pipeline. Secondly, we propose a selective deep and shallow semantic information fusion module (S-DSFM) that explicitly integrates shallow output features into deep output features, so to reduce the waste of modeling capabilities and obtain more high-frequency semantic information related to the target. Finally, we design a temporal cue modeling module based on linguistic classification and multi-frame historical information(MHLS-TCM), with the aim of more comprehensive utilization of historical temporal semantic information. Benefit from the above designs, our VL tracker can obtain stronger target semantic information. Competitive performance from extensive experimental results on five popular vision-language tracking benchmarks, including LaSOT, OTB99-Lang, WebUAV-3M, LaSOText and TNL2K, have demonstrated the superiority and effectiveness of our SIEVL-Track. Ning Li 0044, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Jian Nong, Shuxiang Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Mamba Adapter: Efficient Multi-Modal Fusion for Vision-Language TrackingabstractUtilizing the high-level semantic information of language to compensate for the limitations of vision information is a highly regarded approach in single-object tracking. However, most existing vision-language (VL) trackers employ full-parameter fine-tuning, which can easily lead to catastrophic forgetting. Therefore, they fail to fully exploit the prior knowledge of pre-trained models from upstream tasks, resulting in unsatisfactory tracking performance. To alleviate the above problem, we propose a simple yet effective Vision-Language Tracking pipeline based on Mamba Adapter, named MAVLT, which adopts the idea of parameter-efficient fine-tuning (PEFT) to realize the interaction between vision-language modalities. This novel approach offers the following advantages: (1)The knowledge of the upstream pre-trained model is efficiently inherited by freezing its parameters. This ensures that the VL tracking framework only learns the modules for vision and language interaction, with a focus on the fusion between modalities. (2)The modal interaction between language and vision encoders is flexibly bridged in each encoder layer via proposed mamba adapter, enabling efficient interaction of visual and language information at multiple levels. Extensive experiments on five popular vision-language tracking benchmarks validate the effectiveness of the proposed MAVLT. Particularly, the MAVLT achieves 73.4% AUC score on the LaSOT benchmarks with only 0.18%(0.32M) of the total parameters updates. Code and models are available at https://github.com/GXNU-ZhongLab/MAVLT. Liangtao Shi, Bineng Zhong 0001, Qihua Liang, Xiantao Hu, Zhiyi Mo, Shuxiang Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Unifying Motion and Appearance Cues for Visual Tracking via Shared QueriesabstractThe rich motion and appearance cues between consecutive frames are crucial for robust visual tracking. However, most existing tracking methods are still limited in designing different components to separately employ corresponding cues and even ignore one of them. This makes them difficult to maintain effective interaction between different cues, thus hindering the models from fostering a comprehensive understanding of the target objects. To address these issues, we propose a unified spatio-temporal cues learning framework (named USCLTrack) that comprehensively mines the variation patterns of targets between consecutive frames in complex video streams. Specifically, USCLTrack firstly aggregates motion and appearance cues into shared queries to provide the bridge of interaction between both cues. Then, it directly generates object locations on the condition of these shared queries in an autoregressive manner, unifying different cues to guide future inferences. To effectively learn multiple spatio-temporal cues aggregated in the shared queries, we develop a spatio-temporal attention mechanism. This mechanism integrates motion cues with appearance cues according to the time steps for ensuring temporal consistency. Moreover, it concurrently captures motion trends and appearance changes to facilitate the understanding of the target objects. Extensive experiments on eight popular tracking benchmarks validate the effectiveness of the proposed USCLTrack. Chaocan Xue, Bineng Zhong 0001, Qihua Liang, Haiying Xia, Shuxiang Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Adaptive Expert Decision for RGB-T TrackingabstractThe features provided by RGB and Thermal Infrared (TIR) images have their own characteristics. Therefore, how to adaptively fuse multi-modal features according to different tracking scenarios is crucial for RGB-T tracking. However, current mainstream RGB-T tracking algorithms often use fixed fusion operations for modal interaction in different scenarios. Consequently, their tracking permanence is deteriorated due to they are unable to dynamically adjust the fused multi-modal features based on the current scenes. To address this issue, we propose a novel RGB-T tracking algorithm called AETrack, which can dynamically extract effective modal features in different scenarios for adaptive fusion. Firstly, we design an adaptive expert decision mechanism that employs multiple experts to process the input features. Each expert focuses on and learns different relevant features. Based on this mechanism, we then propose a feature-guided method that leverages the correlations between modalities to provide cross-modal information. This guidance enables the adaptive expert mechanism to adaptively select the most suitable expert to output effective features based on different scenarios, ensuring that our proposed AETrack prioritizes effective features and thus alleviates interference from irrelevant information. Finally, we design a Progressive Cross-modal Fusion operation to achieve multi-level adaptive fusion of effective features across different modalities. Benefiting from this adaptive fusion process, we can effectively achieve multi-modal interaction in different scenarios to guide robust tracking. Extensive experiments on three popular benchmarks (i.e., LasHeR, RGBT210, RGBT234) show that our proposed AETrack can significantly improve tracking performance. Zhiruo Zhu, Bineng Zhong 0001, Qihua Liang, Yaozong Zheng, Ning Li 0044 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Robust Multi-Stage Tracking via Multi-Scale and Multi-Level Representation LearningabstractHow to learn multi-scale and multi-level representations is crucial for robust tracking. However, most current one-stream structure based trackers with visual transformers (dubbed ViTs) cannot effectively capture multi-scale representations due to the structure of their adopted ViTs is non-hierarchical. Meanwhile, they often only use the output features from the final layer for predicting results (i.e., ignoring the utilization of low-level features from the shallow layers) which may result in a certain degree of lacking multi-level representation learning ability. To address these issues, we propose a robust multi-stage tracker that effectively combines the advantages of both hierarchical and one-stream structured ViT as a tracking backbone to improve the multi-scale and multi-level representation learning abilities. Specifically, first of all, we design a hierarchical tracker with a three-stage backbone. In the first two stages of our tracker, we utilize a dual-branch structure to obtain multi-scale features of the template and search region separately. Especially, We design the local scale awareness modules based on simple MLP layers to capture multi-scale features. These modules remove complex operations such as convolutions or shifted window attentions, thus avoiding the performance degradation caused by traditional hierarchical ViTs. In the third stage (i.e. the main stage), we construct a global encoder based on the one-stream ViT to achieve efficient feature extraction and feature interaction for our tracker. Then, we design a multi-level feature integration module in the main stage to explicitly utilize the representation information learned from the shallow layers and fuse them with the features of the final layer to obtain multi-level representation information. Lastly, benefit from the these designs, our tracker can effectively capture more multi-scale and multi-level representations for robust tracking. Comprehensive experiments on GOT-10k, LaSOT, LaSOT$_{ext}$, TNL2K, UAV123, TrackingNet and VOT2020 benchmarks validate the effectiveness and robustness of our method. Ning Li 0044, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shuxiang Song 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Uncertainty-Guided Diffusion Model for Camouflaged Object DetectionabstractRecently, diffusion models have significantly improved the performance of Camouflaged Object Detection (COD) by adding noise to a mask and iteratively denoising it to match the target distributions. Due to the direct extraction of features from noisy masks and the lack of conditional constraints on a prediction area, the diffusion model may deviate from a correct prediction range and produces mispredictions in regions with high uncertainty. To address this issue, we propose an uncertainty-guided diffusion model (UGDNet) for COD, which explicitly quantifies uncertainty and integrates it as an anchor condition into the diffusion models to provide an initialization of the diffusion regions. The core idea is first to utilize a probability representation and transformer to explicitly model uncertainty, aiming to identify areas where a model may generate overconfident mispredictions. Then, we use the uncertainty as an anchor condition to provide a reference prediction range for the diffusion model, guiding each step of the diffusion process. Furthermore, we use uncertainty to guide feature aggregation, prompting the model to pay extra attention to the semantic features of regions with high uncertainty to refine the segmentation results further. The experimental results indicate that our proposed UGDNet achieves higher accuracy than existing state-of-the-art models on five COD benchmarks, including COD10K, NC4K, CAMO, CHAMELEON, and CDS2K. Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shengping Zhang, Shuxiang Song 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Explicit Visual Prompts for Visual Object TrackingabstractHow to effectively exploit spatio-temporal information is crucial to capture target appearance changes in visual tracking. However, most deep learning-based trackers mainly focus on designing a complicated appearance model or template updating strategy, while lacking the exploitation of context between consecutive frames and thus entailing the when-and-how-to-update dilemma. To address these issues, we propose a novel explicit visual prompts framework for visual tracking, dubbed EVPTrack. Specifically, we utilize spatio-temporal tokens to propagate information between consecutive frames without focusing on updating templates. As a result, we cannot only alleviate the challenge of when-to-update, but also avoid the hyper-parameters associated with updating strategies. Then, we utilize the spatio-temporal tokens to generate explicit visual prompts that facilitate inference in the current frame. The prompts are fed into a transformer encoder together with the image tokens without additional processing. Consequently, the efficiency of our model is improved by avoiding how-to-update. In addition, we consider multi-scale information as explicit visual prompts, providing multiscale template features to enhance the EVPTrack's ability to handle target scale changes. Extensive experimental results on six benchmarks (i.e., LaSOT, LaSOText, GOT-10k, UAV123, TrackingNet, and TNL2K.) validate that our EVPTrack can achieve competitive performance at a real-time speed by effectively exploiting both spatio-temporal and multi-scale information. Code and models are available at https://github.com/GXNU-ZhongLab/EVPTrack. Liangtao Shi, Bineng Zhong 0001, Qihua Liang, Ning Li 0044, Shengping Zhang, Xianxian Li |
AAAI | 3 |
| 2024 | ODTrack: Online Dense Temporal Token Learning for Visual TrackingabstractOnline contextual reasoning and association across consecutive video frames are critical to perceive instances in visual tracking. However, most current top-performing trackers persistently lean on sparse temporal relationships between reference and search frames via an offline mode. Consequently, they can only interact independently within each image-pair and establish limited temporal correlations. To alleviate the above problem, we propose a simple, flexible and effective video-level tracking pipeline, named ODTrack, which densely associates the contextual relationships of video frames in an online token propagation manner. ODTrack receives video frames of arbitrary length to capture the spatio-temporal trajectory relationships of an instance, and compresses the discrimination features (localization information) of a target into a token sequence to achieve frame-to-frame association. This new solution brings the following benefits: 1) the purified token sequences can serve as prompts for the inference in the next video frame, whereby past information is leveraged to guide future inference; 2) the complex online update strategies are effectively avoided by the iterative propagation of token sequences, and thus we can achieve more efficient model representation and computation. ODTrack achieves a new SOTA performance on seven benchmarks, while running at real-time speed. Code and models are available at https://github.com/GXNU-ZhongLab/ODTrack. Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shengping Zhang, Xianxian Li |
AAAI | 3 |
| 2024 | Visual Adapt for RGBD TrackingabstractRecent RGBD trackers have employed cueing techniques by overlaying Depth modality images as cues onto RGB modality images, which are then fed into the RGB-based model for tracking. However, the direct overlaying interaction method between modalities not only introduces more noise into the feature space but also exhibits the inadaptability of the RGB-based model to mixed-modality inputs. To address these issues, we introduce Visual Adapt for RGBD Tracking (VADT). Specifically, we maintain the input of the RGB-based model as the RGB modality. Additionally, we have devised a fusion module to enable modality interaction between depth and RGB features. Subsequently, a Depth Adapt module has been formulated to facilitate image interaction with the fused features. This module involves cross-attending to the obtained depth-assisted features and the RGB search frame features produced by the RGB-based model’s output. Experimental results indicate that our proposed tracker achieves state-of-the-art results on various RGBD benchmark tests. Guangtong Zhang, Qihua Liang, Zhiyi Mo, Ning Li 0044, Bineng Zhong 0001 |
ICASSP | 2 |
| 2024 | Diffusion Mask-Driven Visual-language Tracking
Guangtong Zhang, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Shuxiang Song 0001 |
IJCAI | 3 |
| 2024 | Top-Down Cross-Modal Guidance for Robust RGB-T TrackingabstractMost RGB-T trackers heavily rely on bottom-up attention and thus overlook top-down cross-modal guidance for learning target features. Consequently, the discriminative power of the learnt target features is weak. To address this issue, we propose a novel RGB-T tracker (called TGTrack) that designs a Top-down Cross-modal Guidance mechanism to learn target features in two stages. In the first stage, our TGTrack effectively generates top-down cross-modal guidance signals with multi-modal encoders-decoders and prior vectors. In the second stage, these signals are transmitted and integrated to improve the discriminative power of our target features by the attention layers of the cross-modal encoders. Moreover, we introduce an Attention-Driven Spatio-Temporal Updater for updating discriminative target features. Through cross-frame attention guidance, it can effectively eliminates irrelevant features within the search region. As a result, our TGTrack can effectively avoid the complex multi-modal fusion modules and thus achieve robust RGB-T tracking. Extensive experiments on three popular RGB-T tracking benchmarks (i.e., LasHeR, RGBT234, and RGBT210) demonstrate that our TGTrack achieves new state-of-the-art performances. Bineng Zhong 0001, Qihua Liang, Yaozong Zheng, Zhiyi Mo, Shuxiang Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Toward Modalities Correlation for RGB-T TrackingabstractRecently, RGB-T tracking methods have made significant progress, demonstrating remarkable capabilities in addressing the complexities of tracking tasks within demanding environments. However, these methods overlook instability of modal validity in real-world scenarios. This limits the model’s ability to understand the correlation between modalities, thereby hindering the model’s ability to fully leverage the synergistic effects of RGB and TIR. To address this challenge, we propose a novel RGB-T tracking model named MCTrack, from the perspective of leveraging correlation among modalities. First, during the feature extraction stage, we design a novel module based on channel matching modeling to construct bidirectional channel context information flow for two modalities. By leveraging information flow, specific modalities correlation information can be transmitted to two modes, augmenting the correlation between the two modes adaptively. Subsequently, after the feature extraction network, the features of each modality are decoded and transformed to generate more correlated feature representations. During this stage, we extract distinctive and collective features by leveraging the correlation among modalities. Then fusing these features and generated search region features specifically for localization. This aids the model in comprehending the correlation between RGB and TIR under complex scenarios, thereby enhancing its ability to capture and utilize key features. Based on extensive experiments conducted on four popular RGB-T tracking benchmarks, our model demonstrates superior performance, particularly showcasing impressive results on the LasHeR dataset with an achieved Precision of 71.6%. Xiantao Hu, Bineng Zhong 0001, Qihua Liang, Shengping Zhang, Ning Li 0044, Xianxian Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Transformer Tracking via Frequency FusionabstractTransformer has achieved impressive progress in visual tracking due to their capability of global modeling, which enables them to learn low-frequency features(i.e., high-level semantic information). However, it seems to overlook the high-frequency features(i.e., low-level texture and edge information) which are crucial to identify different intra-class object instances in the tracking task. To address this issue, we propose a transformer based tracker via frequency fusion perspective that investigated whether high-frequency and low-frequency features can be effectively combined to achieve robust tracking. Specifically, we design a simple yet effective two-stage fusion strategy and use an appropriate frequency fusion strategy in tracking process of each stage so as to make full use of frequency domain information. In the feature extraction stage, we use wavelet decomposition of high-frequency subbands to solve the performance loss caused by the transformer’s catastrophic forgetting of high-frequency information. In the prediction head stage, we use a variety of wavelet decomposition subbands to model the multi-frequency information. The two-stage fusion strategy makes our model extract more balanced and beneficial multi-frequency information, enabling it to effectively capture target texture information and local edge information while also being sensitive to global information. Extensive experiments on six challenging benchmarks (i.e., LaSOT$_{ext}$, UAV123, TNL2K, LaSOT, TrackingNet, and GOT-10k) demonstrates the superior performance of our tracker. Xiantao Hu, Bineng Zhong 0001, Qihua Liang, Shengping Zhang, Ning Li 0044, Xianxian Li, Rongrong Ji |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Robust Tracking via Combing Top-Down and Bottom-Up AttentionabstractTransformer attention plays an important role in current top-performing trackers. However, it is bottom-up, driven by stimulus and lacks intrinsic prior guidance. This bottom-up attention mechanism leads to an emphasis on all objects in the input images, rather than the task related objects. As a result, the performance of the bottom-up attention based trackers is deteriorated in complicated scenes. To address this issue, we propose a robust tracker that combines bottom-up attention with top-down attention to comply with the existing ViT framework, named TBTrack. TBTrack can not only utilize the existing bottom-up attention mechanisms to model the long-range relationship of input tokens, but also utilize a newly added top-down attention mechanism to pay more attention to task related object and further eliminate interference from similar objects and backgrounds. Specifically, we firstly design a top-down prior generation module using an adaptive learning parameter combined with the template inputs to obtain top-down task guided signals. Then, we inject the prior signals into a bottom-up attention module to obtain a top-down and bottom-up attention combination block (TB-Block). Finally, we stack these TB-Blocks to construct our tracker (TBTrack) with top-down prior guidance capability, which focuses more on the task related object. Through extensive experiments, our TBTrack achieves impressive performance on multiple tracking benchmarks, including GOT-10k, LaSOT, LaSOText, TNL2K, TrackingNet, UAV123 and so on. The code and trained models will be publicly available. Ning Li 0044, Bineng Zhong 0001, Yaozong Zheng, Qihua Liang, Zhiyi Mo, Shuxiang Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Positive-Sample-Free Object Tracking via a Soft ConstraintabstractMost of the existing bounding box-based trackers rely on a classification subnetwork and a regression subnetwork to predict the location and scale of the bounding box. They learn the classification subnetwork by processing each sample individually and applying the suggested classification confidence to produce the final prediction. They typically involve heuristic positive sample configurations, which inevitably introduce mislabelled training samples and therefore deteriorate their tracking performance. Moreover, the parallel prediction of the bounding box position and scale may lead to misalignment of classification and regression. To address these issues,we propose a simple yet effective soft constraint-based tracking framework without positive samples (named SoftCT). SoftCT adaptively senses the target’s pixel position through a soft constraint mechanism, which eliminates potential performance gaps caused by artificially marking the target’s pixel position. In addition, SoftCT computes the state of the bounding box by aggregating such positional information, thereby allowing the tracker to avoid misalignment in classification and regression due to uninformed communication. Specifically, SoftCT directly senses the position of the target pixel and fuses this information into the bounding box prediction, rather than requiring explicit annotation or regression of the target pixel. Extensive experiments on six tracking benchmarks including GOT-10k, TrackingNet, LaSOT, UAV123, LaSOText and TNL2K demonstrate that our tracker achieves state-of-the-art performance, confirming its effectiveness and efficiency. Jiaxin Ye, Bineng Zhong 0001, Qihua Liang, Shengping Zhang, Xianxian Li, Rongrong Ji |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | One-Stream Stepwise Decreasing for Vision-Language TrackingabstractBased on the fixed language descriptions in the initial frames, a vision-language tracker typically adopts a two-stream model structure to align vision and language features at the feature fusion stages. However, this paradigm may degrade the tracking performance due to inaccurate language descriptions and lacks further modal interaction. To address these issues, we propose a one-stream vision-language model called One-stream Stepwise Decreasing for Vision-Language Tracking (OSDT). Specifically, we first encode the language description using a language encoder. The obtained language features are then combined with visual images and entered jointly into a visual encoder, in which the encoder’s self-attention mechanism is utilized to facilitate more interactions between language and visual features. Moreover, to mitigate the problems caused by inaccurate language descriptions, we design a stepwise decreasing multi-modal interaction framework, in which a Feature Filter Module (FFM) is introduced to select language features that are more relevant to visual information to provide semantic guidance for visual feature extraction. Furthermore, without additional feature fusion modules, our one-stream model framework can efficiently utilize the proposed feature filtering module for feature selection. Consequently, our tracker can achieve fast tracking speed in the vision-language tracking domain compared to existing state-of-the-art methods. We extensively evaluate our tracker on three benchmarks, i.e. TNL2K, LaSOT, and OTB99, demonstrating competing performance compared to state-of-the-art vision-language tracking methods. Guangtong Zhang, Bineng Zhong 0001, Qihua Liang, Zhiyi Mo, Ning Li 0044, Shuxiang Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Toward Unified Token Learning for Vision-Language TrackingabstractIn this paper, we present a simple, flexible and effective vision-language (VL) tracking pipeline, termed MMTrack, which casts VL tracking as a token generation task. Traditional paradigms address VL tracking task indirectly with sophisticated prior designs, making them over-specialize on the features of specific architectures or mechanisms. In contrast, our proposed framework serializes language description and bounding box into a sequence of discrete tokens. In this new design paradigm, all token queries are required to perceive the desired target and directly predict spatial coordinates of the target in an auto-regressive manner. The design without other prior modules avoids multiple sub-tasks learning and hand-designed loss functions, significantly reducing the complexity of VL tracking modeling and allowing our tracker to use a simple cross-entropy loss as unified optimization objective for VL tracking task. Extensive experiments on TNL2K, LaSOT, LaSOT$_{\mathrm{ext}}$and OTB99-Lang benchmarks show that our approach achieves promising results, compared to other state-of-the-arts. Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Guorong Li, Rongrong Ji, Xianxian Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Robust Tracking via Bidirectional Transduction With Mask InformationabstractIn the tracking literature, foreground and background information have been extensively investigated to discriminate a target from its surrounding background. However, both foreground and background possess their own spatial-temporal correlation relationship that provide significant information to separate the target from its surrounding background, which has been usually ignored by existing work. To address this issue, we propose a bidirectional transductive network based tracker, which incorporates long-range spatial-temporal and bidirectional constraints. Specifically, our tracker consists of two modules, namely the mask generation module (MGM) and the transduction attention module (TAM). MGM aggregates long-range interdependencies of a target along the history frames for generating accurate target masks. TAM retrieves back to the history frames to find patches similar to the current frame, which are then forwarded along with the target masks generated by MGM. In this manner, each position in the current frame can determine its own identity, whether belonging to either the background or the foreground, hence accurately distinguishing the target from its distractors. We conduct systematically experiments and achieve state-of-the-art performance on several benchmarks, obtaining 69.2% AO on GOT-10k and 82.1% on TrackingNet. TianYu Ning, Bineng Zhong 0001, Qihua Liang, Zhenjun Tang, Xianxian Li |
IEEE Trans. Multim. | 3 |
| 2023 | Robust Tracking via Unifying Pretrain-Finetuning and Visual Prompt TuningabstractThe finetuning paradigm has been a widely used methodology for the supervised training of top-performing trackers. However, the finetuning paradigm faces one key issue: it is unclear how best to perform the finetuning method to adapt a pretrained model to tracking tasks while alleviating the catastrophic forgetting problem. To address this problem, we propose a novel partial finetuning paradigm for visual tracking via unifying pretrain-finetuning and visual prompt tuning (named UPVPT), which can not only efficiently learn knowledge from the tracking task but also reuse the prior knowledge learned by the pre-trained model for effectively handling various challenges in tracking task. Firstly, to maintain the pre-trained prior knowledge, we design a Prompt-style method to freeze some parameters of the pretrained network. Then, to learn knowledge from the tracking task, we update the parameters of the prompt and MLP layers. As a result, we cannot only retain useful prior knowledge of the pre-trained model by freezing the backbone network but also effectively learn target domain knowledge by updating the Prompt and MLP layer. Furthermore, the proposed UPVPT can easily be embedded into existing Transformer trackers (e.g., OSTracker and SwinTracker) by adding only a small number of model parameters (less than 1% of a Backbone network). Extensive experiments on five tracking benchmarks (i.e., UAV123, GOT-10k, LaSOT, TNL2K, and TrackingNet) demonstrate that the proposed UPVPT can improve the robustness and effectiveness of the model, especially in complex scenarios. Guangtong Zhang, Qihua Liang, Ning Li 0044, Zhiyi Mo, Bineng Zhong 0001 |
MMAsia | 2 |
| 2023 | SpectralTracker: Jointly High and Low-Frequency Modeling for Tracking
Yimin Rong, Qihua Liang, Ning Li 0044, Zhiyi Mo, Bineng Zhong 0001 |
PRCV (12) | 2 |
| 2023 | Leveraging Local and Global Cues for Visual Tracking via Parallel Interaction NetworkabstractDespite that both local and context information are crucial for robust tracking, existing CNN-based and transformer-based methods mainly focus on one of these aspects. Consequently, the former fails to exploit rich global context information due to the limited receptive field, while the latter suffers from the deficiencies in constructing the local relationship among neighboring regions. To address this issue, we propose the SiamPIN tracker, based on our Parallel Interaction Network. It consists of two effective modules, namely Global Aggregation Block (GAB) and Local Process Block (LPB). GAB perceives the global context to capture the long-range spatial dependency through a transformer-based architecture. Meanwhile, LPB performs local information extraction using a CNN model to retain the detailed appearance information of the target. These two modules are connected consecutively to compose a Trans-Conv unit block, which transmits the global context information to the local feature extraction procedure, hence enables the interaction of global-local information flow. Several such blocks are cascaded so that our model can learn to aggregate local and context information interactively. The proposed tracker achieves state-of-the-art performance on six benchmark datasets, while maintaining a real time running speed. Yaozong Zheng, Bineng Zhong 0001, Qihua Liang, Zhenjun Tang, Rongrong Ji, Xianxian Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Teacher-student knowledge distillation for real-time correlation tracking
Qihuang Chen, Bineng Zhong 0001, Qihua Liang, Qingyong Deng, Xianxian Li |
Neurocomputing | 3 |
| 2021 | Reference-agnostic representation and visualization of pan-genomesabstractBACKGROUND: The pan-genome of a species is the union of the genes and non-coding sequences present in all individuals (cultivar, accessions, or strains) within that species. RESULTS: Here we introduce PGV, a reference-agnostic representation of the pan-genome of a species based on the notion of consensus ordering. Our experimental results demonstrate that PGV enables an intuitive, effective and interactive visualization of a pan-genome by providing a genome browser that can elucidate complex structural genomic variations. CONCLUSIONS: The PGV software can be installed via conda or downloaded from https://github.com/ucrbioinfo/PGV . The companion PGV browser at http://pgv.cs.ucr.edu can be tested using example bed tracks available from the GitHub page. Qihua Liang, Stefano Lonardi |
BMC Bioinform. | 1 |