Hengzhou Ye

dblp:82/181 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0001-6646-8747ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning motion blur robust vision transformers for real-time UAV tracking
You Wu 0009, Xucheng Wang, Dan Zeng 0002, Hengzhou Ye, Xiaolan Xie 0002, Qijun Zhao, Shuiwang Li
Expert Syst. Appl.4
2026 Learning an Adaptive and View-Invariant Vision Transformer for Real-Time UAV Tracking
abstract
Transformer-based models have improved visual tracking, but most still cannot run in real time on resource-limited devices, especially for unmanned aerial vehicle (UAV) tracking. To achieve a better balance between performance and efficiency, we propose AVTrack, an adaptive computation tracking framework that adaptively activates transformer blocks through an Activation Module (AM), which dynamically optimizes the ViT architecture by selectively engaging relevant components. To address extreme viewpoint variations, we propose to learn view-invariant representations via mutual information (MI) maximization. In addition, we propose AVTrack-MD, an enhanced tracker incorporating a novel MI maximization-based multi-teacher knowledge distillation framework. Leveraging multiple off-the-shelf AVTrack models as teachers, we maximize the MI between their aggregated softened features and the corresponding softened feature of the student model, improving the generalization and performance of the student, especially under noisy conditions. Extensive experiments show that AVTrack-MD achieves performance comparable to AVTrack’s performance while reducing model complexity and boosting average tracking speed by over 17%. Codes is available at https://github.com/wuyou3474/AVTrack.
You Wu 0009, Xucheng Wang, Xiangyang Yang 0001, Hengzhou Ye, Dan Zeng 0002, Qijun Zhao, Shuiwang Li
IEEE Trans. Circuits Syst. Video Technol.6
2026 Toward Real-Time UAV Tracking With Adaptive and Background-Aware Vision Transformers
abstract
In Unmanned Aerial Vehicle (UAV) tracking, discriminative correlation filters (DCF) are popular for their speed, making them ideal for real-time use with limited resources. Recently, lightweight convolutional neural networks (CNNs) have offered a new approach. Through filter pruning, these CNNs maintain high accuracy and efficiency, making them a strong alternative, especially for greater precision. Despite these advancements, the potential of pure vision transformers (ViTs) in UAV tracking remains largely untapped, especially based on the paradigm of conditional computation. In this work, we introduce an adaptive and background-aware Vision Transformer (Aba-ViT) and leverage it to develop a real-time UAV tracking framework called Aba-ViTrack. The proposed Aba-ViT exploits an adaptive and background-aware token computation method to reduce inference time. This approach adaptively discards tokens based on learned halting probabilities, which a priori are higher for background tokens than target ones. To further improve efficiency, this paper proposes a novel classroom-style learning (CSL) approach, where robustness knowledge is transmitted vertically from teacher to students, and generalization capability is enhanced through horizontal mutual learning among students. This method is used to compress Aba-ViTrack, resulting in Aba-ViTrack++. The upgraded version achieves a better balance between accuracy and efficiency in real-time UAV tracking. This version achieves a better balance between accuracy and efficiency for real-time UAV tracking. Extensive experiments on six UAV tracking benchmarks demonstrate that the proposed method achieves state-of-the-art performance in UAV tracking. The code is available at https://github.com/xyyang317/Aba-ViTrack.
Xiangyang Yang 0001, Dan Zeng 0002, Xucheng Wang, Hengzhou Ye, Xiaolan Xie 0002, Qijun Zhao, Jihua Zhu, Shuiwang Li
IEEE Trans. Circuits Syst. Video Technol.4
2025 Learning Occlusion-Robust Vision Transformers for Real-Time UAV Tracking
abstract
Single-stream architectures using Vision Transformer (ViT) backbones show great potential for real-time UAV tracking recently. However, frequent occlusions from obstacles like buildings and trees expose a major drawback: these models often lack strategies to handle occlusions effectively. New methods are needed to enhance the occlusion resilience of single-stream ViT models in aerial tracking. In this work, we propose to learn Occlusion-Robust Representations (ORR) based on ViTs for UAV tracking by enforcing an invariance of the feature representation of a target with respect to random masking operations modeled by a spatial Cox process. Hopefully, this random masking approximately simulates target occlusions, thereby enabling us to learn ViTs that are robust to target occlusion for UAV tracking. This framework is termed ORTrack. Additionally, to facilitate real-time applications, we propose an Adaptive Feature-Based Knowledge Distillation (AFKD) method to create a more compact tracker, which adaptively mimics the behavior of the teacher model ORTrack according to the task’s difficulty. This student model, dubbed ORTrack-D, retains much of ORTrack’s performance while offering higher efficiency. Extensive experiments on multiple benchmarks validate the effectiveness of our method, demonstrating its state-of-the-art performance. Codes is available at https://github.com/wuyou3474/ORTrack.
You Wu 0009, Xucheng Wang, Xiangyang Yang 0001, Dan Zeng 0002, Hengzhou Ye, Shuiwang Li
CVPR6
2025 MambaNUT: Nighttime UAV Tracking via Mamba-based Adaptive Curriculum Learning
abstract
Harnessing low-light enhancement and domain adaptation, nighttime UAV tracking has made substantial strides. However, over-reliance on image enhancement, limited high-quality nighttime data, and a lack of integration between daytime and nighttime trackers hinder the development of an end-to-end trainable framework. Additionally, current ViT-based trackers demand heavy computational resources due to their reliance on the self-attention mechanism. In this paper, we propose a novel pure Mamba-based tracking framework (MambaNUT) that employs a state space model with linear complexity as its backbone, incorporating a single-stream architecture that integrates feature learning and template-search coupling within Vision Mamba. We introduce an adaptive curriculum learning (ACL) approach that dynamically adjusts sampling strategies and loss weights, thereby improving the model’s ability of generalization. Our ACL is composed of two levels of curriculum schedulers: (1) sampling scheduler that transforms the data distribution from imbalanced to balanced, as well as from easier (daytime) to harder (nighttime) samples; (2) loss scheduler that dynamically assigns weights based on the size of the training set and IoU of individual instances. Exhaustive experiments on multiple nighttime UAV tracking benchmarks demonstrate that the proposed MambaNUT achieves state-of-the-art performance while requiring lower computational costs. The code will be available at https://github.com/wuyou3474/MambaNUT.
You Wu 0009, Xiangyang Yang 0001, Xucheng Wang, Hengzhou Ye, Dan Zeng 0002, Shuiwang Li
IROS4
2025 Adaptive Dynamic Service Placement Approach for Edge-Enabled Vehicular Networks Based on SAC and RF
abstract
ABSTRACT Edge computing offers crucial computational and storage support to vehicles by providing various services within the framework of the Internet of Vehicles in intelligent transportation systems. Service placement (SP) becomes particularly challenging when edge resources are limited and vehicles exhibit high‐mobility. Many current dynamic placement methods rely on real‐time placement, often leading to increased costs, instability, and frequent changes. This paper proposes SACRF‐SP, an adaptive dynamic service placement algorithm based on Soft Actor‐Critic (SAC) and Random Forest (RF), for dynamic urban traffic scenarios. This algorithm utilizes the SAC method to identify optimal placement nodes and integrates an RF model to predict service request trends. A decision network is constructed to assess the necessity of redeployment. Extensive simulation experiments demonstrate that SACRF‐SP significantly reduces latency, resource usage, and the frequency of redeployment.
Hengzhou Ye, Gaoxing Li
Concurr. Comput. Pract. Exp.2
2025 Adaptively bypassing vision transformer blocks for efficient visual tracking
Xiangyang Yang 0001, Dan Zeng 0002, Xucheng Wang, You Wu 0009, Hengzhou Ye, Qijun Zhao, Shuiwang Li
Pattern Recognit.5
2024 Tracking Transforming Objects: A Benchmark
You Wu 0009, Yuelong Wang, Yaxin Liao, Fuliang Wu, Hengzhou Ye, Shuiwang Li
PRCV (13)5
2024 Three-partition coevolutionary differential evolution algorithm for mixed-variable optimization problems
abstract
Both in industrial and scientific fields, many optimization problems involve continuous and discrete decision variables. Such problems are called mixed-variable optimization problems (MVOPs). MVOPs remain challenging due to the different spatial distribution characteristics of continuous, ordinal, and categorical variables. In this study, a new variant of differential evolution (DE), called the three-partition coevolutionary DE algorithm for MVOPs (TCDEmv) is proposed. First, a mixed-variable three-partition coevolutionary scheme that can simultaneously handle MVOPs comprising continuous, ordinal, and categorical variables with the same evolution operator is proposed. Additionally, the TCDEmv adopts a dynamic adaptive (DA) mechanism to maintain the balance between ordinal and categorical variables, avoiding the quantity dominance issue. Furthermore, to enhance the efficiency of the TCDEmv, a statistical probability-based two-layer optimization strategy (SPT) was employed for ordinal and categorical variables. The experimental results on 34 artificial MVOPs show that the TCDEmv obtained better solutions and convergence than seven representative algorithms. Compared with similar algorithms in three real-world MVOPs, the TCDEmv also shows competitive performance.
Guojun Gan, Hengzhou Ye, Minggang Dong
Eng. Appl. Artif. Intell.2
2024 Tripartite Game Theory-Based Edge Resource Pricing Approach for Edge Federation
Hengzhou Ye, Bochao Feng, Qiu Lu
J. Grid Comput.1
2024 EdgeSim++: A Realistic, Versatile, and Easily Customizable Edge Computing Simulator
abstract
In the edge computing environment, due to factors such as expensive devices, complex scenarios, and fine control requirements, it is necessary to use simulators to construct heterogeneous devices, simulate task offloading scenarios, and evaluate and optimize decision-making algorithms. However, existing edge computing simulators overly simplify the simulation process, neglecting device interactions and task complexities, making it difficult to finely control the task offloading process and customize complex scenarios. We propose EdgeSim++, a realistic, versatile, and easily customizable edge computing simulator. EdgeSim++ has the capability to generate a batch of heterogeneous devices with various features, supporting the simulation of multi-layer cloud-edge-end scenarios and various MEC architectures. It dynamically sets task and device resource attributes, enables devices to interact and schedule behaviors through the network, and facilitates easy customization of both machine learning and non-machine learning offloading strategies. Additionally, EdgeSim++ supports network topology visualization and result visualization, providing intuitive displays of task transmission processes and offloading results, and real-time monitoring of device resource changes. To test the effectiveness of EdgeSim++, we provided default machine learning and non-machine learning offloading strategies, detailed simulation steps, demonstrated scenario construction, showcased process control, monitored resource changes, and compared algorithm effects. The results illustrated that our simulator could realistically simulate complex edge computing scenarios, flexibly control decision processes, and customize complex offloading algorithms, exhibiting good scalability and reusability.
Qiu Lu, Gaoxing Li, Hengzhou Ye
IEEE Internet Things J.3
2024 Learning Target-Aware Vision Transformers for Real-Time UAV Tracking
abstract
In recent years, the field of unmanned aerial vehicle (UAV) tracking has grown rapidly, finding numerous applications across various industries. While the discriminative correlation filters (DCF)-based trackers remain the most efficient and widely used in the UAV tracking, recently lightweight convolutional neural network (CNN)-based trackers using filter pruning have also demonstrated impressive efficiency and precision. However, the performance of these lightweight CNN-based trackers is still far from satisfactory. In the generic visual tracking, emerging vision transformer (ViT)-based trackers have shown great success by using cross-attention instead of correlation operation, enabling more effective capturing of relationships between the target and the search image. But to best of the authors’ knowledge, the UAV tracking community has not yet well explored the potential of ViTs for more effective and efficient template-search coupling for UAV tracking. In this article, we propose an efficient ViT-based tracking framework for real-time UAV tracking. Our framework integrates feature learning and template-search coupling into an efficient one-stream ViT to avoid an extra heavy relation modeling module. However, we observe that it tends to weaken the target information through transformer blocks due to the significantly more background tokens. To address this problem, we propose to maximize the mutual information (MI) between the template image and its feature representation produced by the ViT. The proposed method is dubbed TATrack. In addition, to further enhance efficiency, we introduce a novel MI maximization-based knowledge distillation, which strikes a better trade-off between accuracy and efficiency. Exhaustive experiments on five benchmarks show that the proposed tracker achieves state-of-the-art performance in UAV tracking. Code is released at:https://github.com/xyyang317/TATrack.
Shuiwang Li, Xiangyang Yang 0001, Xucheng Wang, Dan Zeng 0002, Hengzhou Ye, Qijun Zhao
IEEE Trans. Geosci. Remote. Sens.5
2023 Learning Disentangled Representation with Mutual Information Maximization for Real-Time UAV Tracking
abstract
Efficiency has been a critical problem in UAV tracking due to limitations in computation resources, battery capacity, and unmanned aerial vehicle maximum load. Although discriminative correlation filters (DCF)-based trackers prevail in this field for their favorable efficiency, some recently proposed lightweight deep learning (DL)-based trackers using model compression demonstrated quite remarkable CPU efficiency as well as precision. Unfortunately, the model compression methods utilized by these works, though simple, are still unable to achieve satisfying tracking precision with higher compression rates. This paper aims to exploit disentangled representation learning with mutual information maximization (DR-MIM) to further improve DL-based trackers’ precision and efficiency for UAV tracking. The proposed disentangled representation separates the feature into an identity-related and an identity-unrelated features. Only the latter is used, which enhances the effectiveness of the feature representation for subsequent classification and regression tasks. Extensive experiments on four UAV benchmarks, including UAV123@10fps, DTB70, UAVDT and VisDrone2018, show that our DR-MIM tracker significantly outperforms state-of-the-art UAV tracking methods.
Xucheng Wang, Xiangyang Yang 0001, Hengzhou Ye, Shuiwang Li
ICME3
2023 Cost-aware resource management based on market pricing mechanisms in edge federation environments
Hengzhou Ye, Wei Hao 0004
J. Supercomput.2
2023 A game-based approach for cloudlet resource pricing for cloudlet federation
Hengzhou Ye, Bochao Feng, Xinxiao Li
J. Supercomput.1
2006 Software Leasing: A Virtual Enterprise Oriented ASP Solution to Small and Middle Enterprise Information Service
abstract
The main challenges to information technology applying in small and middle enterprises in China can be summarized to the three aspects: the financial problem, the lack of professional knowledge in IT, and the shortage of resource sharing and communication between the enterprises. Based on the analysis of small and middle enterprises in China, this paper presents a virtual enterprise oriented ASP solution using software leasing technology. The two innovative ideas in software leasing have been described in details, which are terminal service and network authentication. An application architecture using the software leasing has been also included in this paper
Qingzhou Niu, Hengzhou Ye, Xinxiao Li, Xiaomei Tao
CSCWD2