VLDB 2026 Research / reviewers in the wild / expert
Yuanliang Xue
dblp:329/0407
· DBLP profile ↗
11ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-8753-4990ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MUTrack: A Memory-Aware Unified Representation Framework for Visual TrackingabstractBuilding a unified target representation that simultaneously achieves short-term adaptability and long-term stability is crucial for robust visual tracking. However, existing trackers typically face an inherent trade-off. Methods primarily relying on short-term appearance and motion cues achieve rapid adaptation, but they often struggle with long-term identity consistency. Conversely, trackers that emphasize extensive temporal context provide strong robustness, yet this approach can compromise their short-term adaptability. To bridge this gap, we propose a novel tracker, MUTrack, which comprehensively integrates both long-term and short-term memories into a unified target representation for more robust tracking. Specifically, we design a unified memory bank that stores and manages long-term memory for maintaining long-term identity consistency, and short-term memory for adapting to instantaneous appearance changes. To fully leverage the complementary nature of both long-term and short-term temporal information, we introduce a perception interaction module that dynamically fuses these memory types through deep and bidirectional interactions, enabling mutual refinement where one guides the other. This ultimately generates a highly adaptive target representation, which effectively balances adaptability to instantaneous changes with robustness against long-term identity drift. Extensive experiments on GOT10k, TrackingNet, LaSOT, LaSOT_ext, NfS, and OTB100 consistently demonstrate that MUTrack achieves SOTA performance. Weijing Wu, Qihua Liang, Bineng Zhong 0001, Yufei Tan, Ning Li 0044, Yuanliang Xue |
AAAI | 7 |
| 2026 | Dynamic and Consistent Doubly Stochastic similarity learning for multi-view and multi-order clustering
Nian Wang 0001, Zhigao Cui, Yanzhao Su, Aihua Li, Yuanliang Xue, Wenqi Ren |
Pattern Recognit. | 5 |
| 2026 | Weakly Supervised Image Dehazing via Physics-Based DecompositionabstractRecent weakly supervised image dehazing (WSID) works have succeeded to improve models’ generalization ability to real scene dehazing by using generative adversarial network (GAN) for unpaired image training. However, it is still difficult for current WSID methods to train one effective dehazing model for various scenes since 1) they always result in residual haze due to insufficient generalization to the feature distribution of real scenes, and 2) they are prone to cause distortions like color shifts, artifacts or halos etc, owing to embedding manual prior or threshold hypothesis for image reconstruction. To solve above problems, in this paper, we propose a novel WSID model via physics-based decomposition (PBD), which estimates atmospheric light, scattering coefficient and scene depth of real haze input to effectively capture the illumination information and haze distribution to recover a preliminary dehazed image by minimizing reconstruction loss. With this constraint, we subtly design a discrete wavelet discriminator (DWD) to effectively improve the generalization to real scene from both spatial and frequency aspect under the supervision of unpaired real clear image. Our PBD is a purely data-driven model freeing from any manual setting or partially correct prior, thus simultaneously ensuring the realness and visibility of dehazed images. Experiments on seven benchmarks verified the strong generalization ability of our PBD, which achieves SOTA dehazing performance with realistic details. Code will be published at https://github.com/NianWang-HJJGCDX/PBD. Nian Wang 0001, Zhigao Cui, Yanzhao Su, Yunwei Lan, Yuanliang Xue, Aihua Li |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | FMTrack: Frequency-Aware Interaction and Multi-Expert Fusion for RGB-T TrackingabstractRecently, RGB-T tracking has received increasing attention due to its robustness. However, existing RGB-T trackers mainly use cross-attention for modal feature interaction, limiting the utilization of complementary information. In addition, these trackers employ fixed dominant-auxiliary paradigms for feature fusion, ignoring modal quality fluctuations. To address these issues, we propose FMTrack, an effective framework for fully capturing complementary information. FMTrack consists of two key components, a frequency-aware interaction network (FIN) and a multi-expert fusion module (MEFM). To emphasize the valuable information in each modality, FIN utilizes frequency masks to perform high-pass and low-pass filtering on RGB and TIR data. FIN explicitly establishes cross-modal interactions via frequency domain learning, which facilitates the sharing of complementary information. Besides, MEFM extracts diverse features via the differentiated expert network and then adjusts feature combinations according to modal reliability, achieving deep understanding and flexible fusion of multimodal data. With FIN and MEFM, FMTrack makes full use of the advantageous information of each modality to highlight target representations, thus improving performance in complex scenes. Extensive experiments on four popular RGBT tracking datasets (LasHeR, VTUAV, RGBT234, and RGBT210) show that our FMTrack achieves leading performance. The code is available at https://github.com/xyl-507/FMTrack. Yuanliang Xue, Guodong Jin, Bineng Zhong 0001, Lining Tan, Chaocan Xue, Yaozong Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Similarity-Guided Layer-Adaptive Vision Transformer for UAV TrackingabstractVision transformers (ViTs) have emerged as a popular backbone for visual tracking. However, complete ViT architectures are too cumbersome to deploy for unmanned aerial vehicle (UAV) tracking which extremely emphasizes efficiency. In this study, we discover that many layers within lightweight ViT-based trackers tend to learn relatively redundant and repetitive target representations. Based on this observation, we propose a similarity-guided layer adaptation approach to optimize the structure of ViTs. Our approach dynamically disables a large number of representation-similar layers and selectively retains only a single optimal layer among them, aiming to achieve a better accuracy-speed trade-off. By incorporating this approach into existing ViTs, we tailor previously complete ViT architectures into an efficient similarity-guided layer-adaptive framework, namely SGLATrack, for real-time UAV tracking. Extensive experiments on six tracking benchmarks verify the effectiveness of the proposed approach, and show that our SGLATrack achieves a state-of-the-art real-time speed while maintaining competitive tracking precision. Codes and models are available at https://github.com/GXNU-ZhongLab/SGLATrack. Chaocan Xue, Bineng Zhong 0001, Qihua Liang, Yaozong Zheng, Ning Li 0044, Yuanliang Xue, Shuxiang Song 0001 |
CVPR | 6 |
| 2025 | Multi-order graph based clustering via dynamical low rank tensor approximation
Nian Wang 0001, Zhigao Cui, Aihua Li, Yuanliang Xue, Rong Wang 0001, Feiping Nie 0001 |
Neurocomputing | 4 |
| 2025 | AVLTrack: Dynamic Sparse Learning for Aerial Vision-Language TrackingabstractThe introduction of natural language for vision-language (VL) tracking has been proven to improve performance. However, natural language remains under-explored in existing aerial trackers. Moreover, existing VL trackers ignore the misalignment of language with dynamic target states, which is prominent in complex UAV scenarios. In this work, we present AVLTrack, a flexible framework for aerial vision-language tracking. It consists of three key components, a dynamic sparse learning (DSL) module, an efficient Transformer backbone, and a multi-level language perception (MLP) strategy. First, DSL sparsely connects language and images via dynamic sparse attention, providing accurate multi-modal prompts. To adapt to target state variations, the sparsity in DSL is dynamically adjusted based on semantic information, flexibly highlighting target-specific tokens. Next, the Transformer backbone follows highly parallelized one-stream architectures, allowing efficient multi-modal feature extraction and interaction. Finally, MLP enables the iterative interaction of language and visual information, aiming to utilize language priori to guide the generation of discriminative visual features. Moreover, we construct the DTB70-NLP dataset to facilitate UAV vision-language tracking. Extensive experiments on WebUAV-3M and DTB70-NLP demonstrate the leading performance of AVLTrack compared to existing outstanding trackers while maintaining a high running speed of 80.5 FPS. The dataset and codes are available athttps://github.com/xyl-507/AVLTrack. Yuanliang Xue, Bineng Zhong 0001, Guodong Jin, Lining Tan, Ning Li 0044, Yaozong Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Target-Distractor Aware UAV Tracking via Global AgentabstractObject tracking is a basic task of the uncrewed aerial vehicle (UAV)-based intelligent visual perception system. The presence of similar targets and complex backgrounds in the airborne perspective poses significant challenges to aerial trackers. However, existing target-aware or distractor-aware trackers fail to capture discriminative cues from both target and background information in a balanced manner, resulting in limited improvement. To address these issues, this paper proposes a global agent-based Target-Distractor Aware Tracker (TDAT) to enhance the discrimination of the target. TDAT comprises two effective modules: a global agent generator and an interactor. First, the generator aggregates the target and background regions into representative agents and then performs self-attention on these agents to explicitly model the global relationships between the target and backgrounds. Next, the interactor realizes the bidirectional information interaction between global agents and local regions via self-attention. Based on the global dependencies encoded in global agents, the interactor extracts target-oriented features and enhances the understanding of the target. TDAT embedded with target-distractor awareness effectively widens the gap between target and background distractors. Experimental results on multiple UAV benchmarks show that TDAT achieves outstanding performance with a speed of 34.5 frames/s. The code is available at https://github.com/xyl-507/TDAT Yuanliang Xue, Guodong Jin, Lining Tan, Nian Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Consistent Representation Mining for Multi-Drone Single Object TrackingabstractAerial tracking has received growing attention due to its broad practical applications. However, single-view aerial trackers are still limited by challenges such as severe appearance variations and occlusions. Existing multi-view trackers utilize cross-drone information to address these issues but struggle to overcome heterogenous differences. In this paper, we propose a novel Transformer-based consistent representation mining (CRM) module to capture invariant target information and suppress the heterogenous differences in cross-drone information. First, CRM divides the heterogenous input into regions and measures semantic relevance by modeling the relations between these regions. Then reliable target regions are roughly localized by selecting the top k most relevant regions. Next, the global perception is performed on these reliable regions via multi-head sparse self-attention, further enhancing the understanding of the target and suppressing background regions. In particular, CRM, as a plug-and-play module, can be flexibly embedded into different tracking frameworks (CRM-Siam and CRM-DiMP). Besides, the multi-view correction strategy is designed to ensure timely correction of multi-view information and full utilization of its own information. Extensive experiments on the multi-drone dataset, MDOT, demonstrate that CRM-assisted trackers effectively improve the accuracy and robustness of the multi-drone tracking system, outperforming other outstanding trackers. The code and models are available athttps://github.com/xyl-507/CRM. Yuanliang Xue, Guodong Jin, Lining Tan, Nian Wang 0001, Lianfeng Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | SmallTrack: Wavelet Pooling and Graph Enhanced Classification for UAV Small Object TrackingabstractAerial object tracking has recently shown great potential in the field of remote sensing. However, small objects with limited feature information pose a huge challenge to aerial trackers. Despite significant improvements, most trackers still struggle to capture enough discriminative features and to overcome background disturbances. In this work, we propose an efficient aerial tracker (SmallTrack) based on the Siamese network to improve the discrimination of small objects. It consists of two effective modules, namely Wavelet Pooling Layer (WPL) and Graph Enhanced Module (GEM). First, WPL decomposes the input into four subbands via wavelet domain learning, and fully utilizes the high- and low-frequency information in the subbands to preserve the discriminative features of small objects. Second, GEM embeds the pixels on the classification responses as nodes in graph learning through graph neural networks, which naturally mines the similarity between pixels. Based on the pixel-level modulation constructed from graph theory, GEM enhances the understanding of small objects and highlights them in the classification responses. The proposed tracker achieves leading performance on five aerial benchmarks, while maintaining a high running speed of 72.5 frames/s. Besides, real-world tests on an aerial platform have proven the effectiveness of SmallTrack. The code and models are available at https://github.com/xyl-507/SmallTrack. Yuanliang Xue, Guodong Jin, Lining Tan, Nian Wang 0001, Lianfeng Wang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | MobileTrack: Siamese efficient mobile network for high-speed UAV trackingabstractAbstract Recently, Siamese‐based trackers have drawn amounts of attention in visual tracking field because of their excellent performance. However, visual object tracking on Unmanned Aerial Vehicles platform encounters difficulties under circumstances such as small objects and similar objects interference. Most existing tracking methods for aerial tracking adopt deeper networks or inefficient policies to promote performance, but most trackers can hardly meet real‐time requirements on mobile platforms with limited computing resources. Thus, in this work, an efficient and lightweight siamese tracker (MobileTrack) is proposed for high‐time Unmanned Aerial Vehicles tracking, realising the balance between performance and speed. Firstly, a lightweight convolutional network (D‐MobileNet) is designed to enhance the characterisation ability of small objects. Secondly, an efficient object‐aware module is proposed for local cross‐channel information exchange, enhancing the feature information of the tracking object. Besides, an anchor‐free region proposal network is introduced to predict the object pixel by pixel. Finally, deep and shallow feature information is fully utilised by cascading multiple anchor‐free region proposal networks for accurate locating and robust tracking. Extensive experiments on the three Unmanned Aerial Vehicles benchmarks show that the proposed tracker achieves outstanding performance while keeping a beyond‐real‐time speed. Yuanliang Xue, Guodong Jin, Lining Tan, Xiaohan Hou |
IET Image Process. | 1 |