EDBT 2026 Demo / reviewers in the wild / expert
Xucheng Wang
dblp:323/8501
· DBLP profile ↗
13ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning motion blur robust vision transformers for real-time UAV tracking
You Wu 0009, Xucheng Wang, Dan Zeng 0002, Hengzhou Ye, Xiaolan Xie 0002, Qijun Zhao, Shuiwang Li |
Expert Syst. Appl. | 2 |
| 2026 | Learning an Adaptive and View-Invariant Vision Transformer for Real-Time UAV TrackingabstractTransformer-based models have improved visual tracking, but most still cannot run in real time on resource-limited devices, especially for unmanned aerial vehicle (UAV) tracking. To achieve a better balance between performance and efficiency, we propose AVTrack, an adaptive computation tracking framework that adaptively activates transformer blocks through an Activation Module (AM), which dynamically optimizes the ViT architecture by selectively engaging relevant components. To address extreme viewpoint variations, we propose to learn view-invariant representations via mutual information (MI) maximization. In addition, we propose AVTrack-MD, an enhanced tracker incorporating a novel MI maximization-based multi-teacher knowledge distillation framework. Leveraging multiple off-the-shelf AVTrack models as teachers, we maximize the MI between their aggregated softened features and the corresponding softened feature of the student model, improving the generalization and performance of the student, especially under noisy conditions. Extensive experiments show that AVTrack-MD achieves performance comparable to AVTrack’s performance while reducing model complexity and boosting average tracking speed by over 17%. Codes is available at https://github.com/wuyou3474/AVTrack. You Wu 0009, Xucheng Wang, Xiangyang Yang 0001, Hengzhou Ye, Dan Zeng 0002, Qijun Zhao, Shuiwang Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Toward Real-Time UAV Tracking With Adaptive and Background-Aware Vision TransformersabstractIn Unmanned Aerial Vehicle (UAV) tracking, discriminative correlation filters (DCF) are popular for their speed, making them ideal for real-time use with limited resources. Recently, lightweight convolutional neural networks (CNNs) have offered a new approach. Through filter pruning, these CNNs maintain high accuracy and efficiency, making them a strong alternative, especially for greater precision. Despite these advancements, the potential of pure vision transformers (ViTs) in UAV tracking remains largely untapped, especially based on the paradigm of conditional computation. In this work, we introduce an adaptive and background-aware Vision Transformer (Aba-ViT) and leverage it to develop a real-time UAV tracking framework called Aba-ViTrack. The proposed Aba-ViT exploits an adaptive and background-aware token computation method to reduce inference time. This approach adaptively discards tokens based on learned halting probabilities, which a priori are higher for background tokens than target ones. To further improve efficiency, this paper proposes a novel classroom-style learning (CSL) approach, where robustness knowledge is transmitted vertically from teacher to students, and generalization capability is enhanced through horizontal mutual learning among students. This method is used to compress Aba-ViTrack, resulting in Aba-ViTrack++. The upgraded version achieves a better balance between accuracy and efficiency in real-time UAV tracking. This version achieves a better balance between accuracy and efficiency for real-time UAV tracking. Extensive experiments on six UAV tracking benchmarks demonstrate that the proposed method achieves state-of-the-art performance in UAV tracking. The code is available at https://github.com/xyyang317/Aba-ViTrack. Xiangyang Yang 0001, Dan Zeng 0002, Xucheng Wang, Hengzhou Ye, Xiaolan Xie 0002, Qijun Zhao, Jihua Zhu, Shuiwang Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Learning Occlusion-Robust Vision Transformers for Real-Time UAV TrackingabstractSingle-stream architectures using Vision Transformer (ViT) backbones show great potential for real-time UAV tracking recently. However, frequent occlusions from obstacles like buildings and trees expose a major drawback: these models often lack strategies to handle occlusions effectively. New methods are needed to enhance the occlusion resilience of single-stream ViT models in aerial tracking. In this work, we propose to learn Occlusion-Robust Representations (ORR) based on ViTs for UAV tracking by enforcing an invariance of the feature representation of a target with respect to random masking operations modeled by a spatial Cox process. Hopefully, this random masking approximately simulates target occlusions, thereby enabling us to learn ViTs that are robust to target occlusion for UAV tracking. This framework is termed ORTrack. Additionally, to facilitate real-time applications, we propose an Adaptive Feature-Based Knowledge Distillation (AFKD) method to create a more compact tracker, which adaptively mimics the behavior of the teacher model ORTrack according to the task’s difficulty. This student model, dubbed ORTrack-D, retains much of ORTrack’s performance while offering higher efficiency. Extensive experiments on multiple benchmarks validate the effectiveness of our method, demonstrating its state-of-the-art performance. Codes is available at https://github.com/wuyou3474/ORTrack. You Wu 0009, Xucheng Wang, Xiangyang Yang 0001, Dan Zeng 0002, Hengzhou Ye, Shuiwang Li |
CVPR | 2 |
| 2025 | MambaNUT: Nighttime UAV Tracking via Mamba-based Adaptive Curriculum LearningabstractHarnessing low-light enhancement and domain adaptation, nighttime UAV tracking has made substantial strides. However, over-reliance on image enhancement, limited high-quality nighttime data, and a lack of integration between daytime and nighttime trackers hinder the development of an end-to-end trainable framework. Additionally, current ViT-based trackers demand heavy computational resources due to their reliance on the self-attention mechanism. In this paper, we propose a novel pure Mamba-based tracking framework (MambaNUT) that employs a state space model with linear complexity as its backbone, incorporating a single-stream architecture that integrates feature learning and template-search coupling within Vision Mamba. We introduce an adaptive curriculum learning (ACL) approach that dynamically adjusts sampling strategies and loss weights, thereby improving the model’s ability of generalization. Our ACL is composed of two levels of curriculum schedulers: (1) sampling scheduler that transforms the data distribution from imbalanced to balanced, as well as from easier (daytime) to harder (nighttime) samples; (2) loss scheduler that dynamically assigns weights based on the size of the training set and IoU of individual instances. Exhaustive experiments on multiple nighttime UAV tracking benchmarks demonstrate that the proposed MambaNUT achieves state-of-the-art performance while requiring lower computational costs. The code will be available at https://github.com/wuyou3474/MambaNUT. You Wu 0009, Xiangyang Yang 0001, Xucheng Wang, Hengzhou Ye, Dan Zeng 0002, Shuiwang Li |
IROS | 3 |
| 2025 | Adaptively bypassing vision transformer blocks for efficient visual tracking
Xiangyang Yang 0001, Dan Zeng 0002, Xucheng Wang, You Wu 0009, Hengzhou Ye, Qijun Zhao, Shuiwang Li |
Pattern Recognit. | 3 |
| 2025 | Exploiting rank-based filter pruning for real-time UAV tracking
Xucheng Wang, Dan Zeng 0002, Qijun Zhao, Shuiwang Li |
Signal Process. Image Commun. | 1 |
| 2024 | Learning Adaptive and View-Invariant Vision Transformer for Real-Time UAV TrackingabstractHarnessing transformer-based models, visual tracking has made substantial strides. However, the sluggish performance of current trackers limits their practicality on devices with constrained computational capabilities, especially for real-time unmanned aerial vehicle (UAV) tracking. Addressing this challenge, we introduce AVTrack, an adaptive computation framework tailored to selectively activate transformer blocks for real-time UAV tracking in this work. Our novel Activation Module (AM) dynamically optimizes ViT architecture, selectively engaging relevant components and enhancing inference efficiency without compromising much tracking performance. Moreover, we bolster the effectiveness of ViTs, particularly in addressing challenges arising from extreme changes in viewing angles commonly encountered in UAV tracking, by learning view-invariant representations through mutual information maximization. Extensive experiments on five tracking benchmarks affirm the effectiveness and versatility of our approach, positioning it as a state-of-the-art solution in visual tracking. Code is released at: https://github.com/wuyou3474/AVTrack. You Wu 0009, Xucheng Wang, Xiangyang Yang 0001, Shuiwang Li |
ICML | 4 |
| 2024 | Learning Target-Aware Vision Transformers for Real-Time UAV TrackingabstractIn recent years, the field of unmanned aerial vehicle (UAV) tracking has grown rapidly, finding numerous applications across various industries. While the discriminative correlation filters (DCF)-based trackers remain the most efficient and widely used in the UAV tracking, recently lightweight convolutional neural network (CNN)-based trackers using filter pruning have also demonstrated impressive efficiency and precision. However, the performance of these lightweight CNN-based trackers is still far from satisfactory. In the generic visual tracking, emerging vision transformer (ViT)-based trackers have shown great success by using cross-attention instead of correlation operation, enabling more effective capturing of relationships between the target and the search image. But to best of the authors’ knowledge, the UAV tracking community has not yet well explored the potential of ViTs for more effective and efficient template-search coupling for UAV tracking. In this article, we propose an efficient ViT-based tracking framework for real-time UAV tracking. Our framework integrates feature learning and template-search coupling into an efficient one-stream ViT to avoid an extra heavy relation modeling module. However, we observe that it tends to weaken the target information through transformer blocks due to the significantly more background tokens. To address this problem, we propose to maximize the mutual information (MI) between the template image and its feature representation produced by the ViT. The proposed method is dubbed TATrack. In addition, to further enhance efficiency, we introduce a novel MI maximization-based knowledge distillation, which strikes a better trade-off between accuracy and efficiency. Exhaustive experiments on five benchmarks show that the proposed tracker achieves state-of-the-art performance in UAV tracking. Code is released at:https://github.com/xyyang317/TATrack. Shuiwang Li, Xiangyang Yang 0001, Xucheng Wang, Dan Zeng 0002, Hengzhou Ye, Qijun Zhao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Adaptive and Background-Aware Vision Transformer for Real-Time UAV TrackingabstractWhile discriminative correlation filters (DCF)-based trackers prevail in UAV tracking for their favorable efficiency, lightweight convolutional neural network (CNN)-based trackers using filter pruning have also demonstrated remarkable efficiency and precision. However, the use of pure vision transformer models (ViTs) for UAV tracking remains unexplored, which is a surprising finding given that ViTs have been shown to produce better performance and greater efficiency than CNNs in image classification. In this paper, we propose an efficient ViT-based tracking framework, Aba-ViTrack, for UAV tracking. In our framework, feature learning and template-search coupling are integrated into an efficient one-stream ViT to avoid an extra heavy relation modeling module. The proposed Aba-ViT exploits an adaptive and background-aware token computation method to reduce inference time. This approach adaptively discards tokens based on learned halting probabilities, which a priori are higher for background tokens than target ones. Extensive experiments on six UAV tracking benchmarks demonstrate that the proposed Aba-ViTrack achieves state-of-the-art performance in UAV tracking. Code is available at https://github.com/xyyang317/Aba-ViTrack. Shuiwang Li, Xiangxyang Yang, Dan Zeng 0002, Xucheng Wang |
ICCV | 4 |
| 2023 | Learning Disentangled Representation with Mutual Information Maximization for Real-Time UAV TrackingabstractEfficiency has been a critical problem in UAV tracking due to limitations in computation resources, battery capacity, and unmanned aerial vehicle maximum load. Although discriminative correlation filters (DCF)-based trackers prevail in this field for their favorable efficiency, some recently proposed lightweight deep learning (DL)-based trackers using model compression demonstrated quite remarkable CPU efficiency as well as precision. Unfortunately, the model compression methods utilized by these works, though simple, are still unable to achieve satisfying tracking precision with higher compression rates. This paper aims to exploit disentangled representation learning with mutual information maximization (DR-MIM) to further improve DL-based trackers’ precision and efficiency for UAV tracking. The proposed disentangled representation separates the feature into an identity-related and an identity-unrelated features. Only the latter is used, which enhances the effectiveness of the feature representation for subsequent classification and regression tasks. Extensive experiments on four UAV benchmarks, including UAV123@10fps, DTB70, UAVDT and VisDrone2018, show that our DR-MIM tracker significantly outperforms state-of-the-art UAV tracking methods. Xucheng Wang, Xiangyang Yang 0001, Hengzhou Ye, Shuiwang Li |
ICME | 1 |
| 2023 | Towards Discriminative Representations with Contrastive Instances for Real-Time UAV TrackingabstractMaintaining high efficiency and high precision are two fundamental challenges in UAV tracking due to the constraints of computing resources, battery capacity, and UAV maximum load. Discriminative correlation filters (DCF)-based trackers can yield high efficiency on a single CPU but with inferior precision. Lightweight Deep learning (DL)-based trackers can achieve a good balance between efficiency and precision but performance gains are limited by the compression rate. High compression rate often leads to poor discriminative representations. To this end, this paper aims to enhance the discriminative power of feature representations from a new feature-learning perspective. Specifically, we attempt to learn more disciminative representations with contrastive instances for UAV tracking in a simple yet effective manner, which not only requires no manual annotations but also allows for developing and deploying a lightweight model. We are the first to explore contrastive learning for UAV tracking. Extensive experiments on four UAV benchmarks, including UAV123@10fps, DTB70, UAVDT and VisDrone2018, show that the proposed DRCI tracker significantly outperforms state-of-the-art UAV tracking methods. Dan Zeng 0002, Mingliang Zou, Xucheng Wang, Shuiwang Li |
ICME | 3 |
| 2022 | Rank-Based Filter Pruning for Real-Time UAV TrackingabstractUnmanned aerial vehicle (UAV) tracking has wide poten-tial applications in such as agriculture, navigation, and public security. However, the limitations of computing resources, battery capacity, and maximum load of UAV hinder the de-ployment of deep learning-based tracking algorithms on UAV. Consequently, discriminative correlation filters (DCF) track-ers stand out in the UAV tracking community because of their high efficiency. However, their precision is usually much lower than trackers based on deep learning. Model compression is a promising way to narrow the gap (i.e., effciency, precision) between DCF- and deep learning- based trackers, which has not caught much attention in UAV tracking. In this paper, we propose the P-SiamFC++ tracker, which is the first to use rank-based filter pruning to compress the SiamFC++ model, achieving a remarkable balance between efficiency and precision. Our method is general and may encourage further studies on UAV tracking with model compression. Extensive experiments on four UAV benchmarks, including UAV123@10fps, DTB70, UAVDT and Vistrone2018, show that P-SiamFC++ tracker significantly outperforms state-of-the-art UAV tracking methods. Xucheng Wang, Dan Zeng 0002, Qijun Zhao, Shuiwang Li |
ICME | 1 |