Ye Wang 0020

dblp:44/6292-20 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2025
0009-0009-5689-8271ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Hyperspectral Tracker With Constrained Object Adaptive Learning and Trajectory Construction
abstract
Hyperspectral imaging offers significant potential for precise object tracking, yet the scarcity of dataset volumes specifically tailored for hyperspectral tracking algorithms hinders progress, particularly for deep models with complex structures. Additionally, current deep learning-based hyperspectral trackers typically enhance model accuracy via online or adversarial learning, adversely affecting tracking speed. To address these challenges, this paper introduces the Constrained Object Adaptive Learning hyperspectral Tracker (COALT), an effective parameter-efficient fine-tuning tracker tailored for hyperspectral tracking. COALT integrates Pixel-level Object Constrained Spectral Prompt (POCSP) and Temporal Sequence Trajectory Prompt (TSTP) through Adaptive Learning with Parameter-efficient Fine-tuning (ALPEFT), enabling a transformer-based tracker to capture detailed spectral features and relationships in hyperspectral image sequences through trainable rank decomposition matrices. Specifically, POCSP is designed to retain optimal spectral information with low internal correlation and high object representativeness, enabling rapid image reconstruction. Then, the most representative spectral template and search are fused into a single stream as spectral prompts for the Encoder and Decoder layers. Concurrently, the previous coordinates within the same sequence are tokenized and utilized as temporal prompts by TSTP in the decoder layers. The model is trained with ALPEFT to optimize spectral information learning, which substantially reduces the number of training parameters, alleviating overfitting issues arising from limited data. Meanwhile, the proposed tracker not only retains the ability of pre-trained model to estimate object trajectories in an autoregressive manner but also effectively utilizes spectral information and enhances target location perception during the fine-tuning process. Extensive experiments and evaluations are conducted on two public hyperspectral tracking datasets. The results demonstrate that the proposed COALT tracker achieves satisfactory performance with leading processing speed. The code will be available at https://github.com/PING-CHUANG/COALT.
Ye Wang 0020, Mingyang Ma 0004, Ge Zhang 0006, Tao Gao 0001, Shaohui Mei
IEEE Trans. Circuits Syst. Video Technol.1
2025 Hyperspectral Object Tracking With Context-Aware Learning and Category Consistency
abstract
Hyperspectral imaging technology is of crucial importance to improve the performance of object tracking in many remote sensing surveillance areas. Previous methods primarily focused on feature fusion strategies by employing additional enhancement modules. However, these methods commonly lack contextual understanding to distinguish the target from the background and totally ignore the category information of the targets. To address these limitations, a novel hyperspectral object tracker is proposed to incorporate context-aware learning and category consistency tracker (CCTrack), which can adaptively learn context-aware representations in hyperspectral scenarios to obtain global target information with memory storage, while constructing an interframe category consistency constraint to enhance tracking process. Specifically, CCTrack integrates an adaptive context-aware learning (ACL) mechanism, which includes a feature decoupling module (FDM) to extract specific representations from decoupled features, and a Mamba layer to retain and update long-range dependencies. To align with prior knowledge of target recognition and motion patterns, an alignment transformation module (ATM) is employed with the ACL mechanism, fully leveraging spatial-spectral representations. In addition, category consistency constraint modules (C3Ms) are introduced to enforce category consistency across frames by computing the similarities between the target features and the corresponding category name, serving as the constraint to improve tracking performance. Extensive experiments over the hyperspectral object tracking (HOT) benchmark covering various remote sensing scenarios demonstrate that CCTrack outperforms state-of-the-art methods by a significant margin.
Ye Wang 0020, Shaohui Mei, Mingyang Ma 0004, Tao Gao 0001, Huiyang Han
IEEE Trans. Geosci. Remote. Sens.1
2025 HTACPE: A Hybrid Transformer With Adaptive Content and Position Embedding for Sample Learning Efficiency of Hyperspectral Tracker
abstract
Transformer architecture has demonstrated significant potential in hyperspectral object tracking by leveraging global correlation learning to accurately represent the data distribution. However, existing hyperspectral object trackers based on transformer models typically rely on costly pre-trained models, making them prone to crashing due to overfitting when tuned on small-scale hyperspectral videos, greatly limiting their performance. To address this challenge, in this paper, a Hybrid Transformer with Adaptive Content and Position Embedding (HTACPE) tracker is proposed to improve the learning efficiency of the tracking model, and fully explore the spectral-spatial information. Specifically, an Adaptive Content and Position Embedding Module (ACPEM) is designed to dynamically learn the balance between focusing on positional and content-based information, which allows the model to effectively handle datasets of various sizes. To enhance the spectral-spatial information, a Spectral Grouping Module (SGM) is designed to learn the highfrequency information in complex scenarios, thereby enhancing diversified features. It operates in parallel with the ACPEM feature learning module. Furthermore, a Dynamic Reliability Refinement Module (DRRM) is incorporated to address challenges related to accurate object position perception, iteratively refining prediction parameters to enhance the reliability of the model. Extensive experiments demonstrate that the proposed HTACPE achieves satisfactory tracking performance both qualitatively and quantitatively, especially with insufficient training data.
Ye Wang 0020, Shaohui Mei, Mingyang Ma 0004, Yuru Su
IEEE Trans. Multim.1
2023 Rethinking Transformers for Semantic Segmentation of Remote Sensing Images
abstract
Transformer has been widely applied in image processing tasks as a substitute for Convolutional Neural Networks (CNNs) for feature extraction due to its superiority in global context modeling and flexibility in model generalization. However, the existing transformer-based methods for semantic segmentation of Remote Sensing (RS) images are still with several limitations, which can be summarized into two main aspects: 1) the transformer encoder is generally combined with CNN-based decoder, leading to inconsistency in feature representations; 2) the strategies for global and local context information utilization are not sufficiently effective. Therefore, in this paper, a Global-Local Transformer Segmentor (GLOTS) framework is proposed for semantic segmentation of RS images to acquire consistent feature representations by adopting transformers for both encoding and decoding, in which a Masked Image Modeling (MIM) pretrained transformer encoder is adopted to learn semantic-rich representations of input images, and a multi-scale global-local transformer decoder is designed to fully exploit the global and local features. Specifically, the transformer decoder uses a feature separation-aggregation module (FSAM) to utilize the feature adequately at different scales and adopts a global-local attention module (GLAM) containing Global Attention Block (GAB) and Local Attention Block (LAB) to capture the global and local context information respectively. Furthermore, a Learnable Progressive Upsampling Strategy (LPUS) is proposed to restore the resolution progressively, which can flexibly recover the fine-grained details in the upsampling process. Experimental results on the three benchmark RS datasets demonstrate that the proposed GLOTS is capable of achieving better performance with some state-of-the-art methods, and the superiority of the proposed framework is also verified by ablation studies. The code will be available at https://github.com/lyhnsn/GLOTS.
Yifan Zhang 0006, Ye Wang 0020, Shaohui Mei
IEEE Trans. Geosci. Remote. Sens.3
2023 Reconstruction-Assisted and Distance-Optimized Adversarial Training: A Defense Framework for Remote Sensing Scene Classification
abstract
Despite deep neural networks (DNNs) have been widely applied in remote sensing (RS) scene classification and achieved satisfying performance, the vulnerability of DNNs towards adversarial examples significantly degrades their performance. Moreover, the relatively limited labeled samples of RS scene classification make DNNs more likely to overfit, leading to weak generalizability and noise sensitivity. This may result in DNNs being more vulnerable to adversarial examples. Consequently, the defense of adversarial examples is of crucial importance to improve both the generalizability and robustness of DNNs in the RS scene classification task. However, few studies have been conducted on defense for RS scene classification, especially ignoring the intrinsic characteristics of RS images. In this paper, an effective defense framework for RS scene classification, named reconstruction-assisted and distance-optimized adversarial training (RDAT), is proposed to defend adversarial examples. In order to solve the problems caused by high interclass similarity, a distance-optimized (DO) strategy is designed for adversarial training to strengthen the learning of underfitting content, increase the interclass distance, and improve the robustness of the networks. Furthermore, in order to generate high quality samples for adversarial training, a reconstruction-assisted (RA) block is proposed to eliminate adversarial perturbations in adversarial examples. Specifically, in this block, by swin transformer (SwinT) block and multi-scale convolution (MSC) block, SwinT-MSC-UNet (SMUNet) is constructed to fully extract global and multi-scale local features to adapt to the characteristics of RS images with large variance of ground object scales. Extensive experiments on the benchmark datasets, i.e., UC Merced (UCM) and Aerial Image Dataset (AID), have demonstrate that the proposed RDAT can effectively resist multiple adversarial attacks and yield superior results than other defense methods for RS scene classification.
Yuru Su, Ge Zhang 0006, Shaohui Mei, Jiawei Lian, Ye Wang 0020, Shuai Wan
IEEE Trans. Geosci. Remote. Sens.5
2022 Semantic Segmentation of High-Resolution Remote Sensing Images Using an Improved Transformer
abstract
Semantic segmentation has been widely researched for high level analysis of High Spatial Resolution (HSR) remote sensing images, where Convolutional Neural Network (CNN) is the mainstream method. However, the transformer with attention mechanism has its unique capacity of extracting global information which is generally ignored by CNN models. In this paper, a Swin Transformer with UPer head (STUP) is proposed to tackle with semantic segmentation problem on a challenging remote sensing land-cover dataset called LoveDA, which owns complex background samples and inconsistent classes distributions. The proposed STUP combines the Swin Transformer with Uper Head in the form of an encoder-decoder structure, to extract features of HSR images for segmentation. Furthermore, Focal Loss is adopted to handle the unbalanced distribution problem in the training step. Experimental results demonstrate that the proposed STUP clearly outperforms several state-of-the-art models.
Shaohui Mei, Ye Wang 0020, Mingyi He, Qian Du 0001
IGARSS4
2022 Gaussian Information Entropy based band Reduction for Unsupervised Hyperspectral Video Tracking
abstract
Hyperspectral videos, which provide extra spectral characteristics besides spatial and temporal information, can improve the performance of object tracking using spectral signatures. However, there is a lack of labeled hyperspectral videos to support deep learning based model design. On the contrary, object tracking in the color space has been well developed in the past decade with many benchmark tracking models, e.g., SiamBAN. Therefore, how to transfer models designed in the color space to the hyperspectral space is of great importance. In this paper, hyperspectral videos are reduced into 3 bands using a band reduction algorithm, by which the existing well-trained trackers can be directly used. Specifically, Gaussian Information Entropy (GIE) is used to transform a hyperspectral video into a 3-band pseudo-color video, by which hyperspectral object tracking is conducted in an unsupervised mode. Experimental results demonstrate that object trackers designed in the color space can be transferred to hyperspectral videos using band reduction algorithms and the GIE based reduction is more effective than several well-known band reduction algorithms when using SiamBAN.
Yuru Su, Shaohui Mei, Ge Zhang 0006, Ye Wang 0020, Mingyi He, Qian Du 0001
IGARSS4
2021 Siammraan: Siamese Multi-Level Residual Attention Adaptive Network for Hyperspectral Videos Tracking
abstract
The deep learning based techniques have been widely applied to object tracking in color videos. When these techniques are applied to hyperspectral videos, how to fully explore unique spectral signatures of tracking objects is of crucial importance as well as simultaneously utilizing spatial and temporal information. Different with color videos, hyperspectral videos record continuous spectral reflectance of targets in light wavelength indexed band images and it is more difficult to explore unique spectral feature of tracking objects. Aiming to take advantage of existing object tracking techniques in color videos, a Siamese Multi-level Residual Attention Adaptive Network (SiamMRAAN) is designed to handle 3-band images by using the well-trained ResNet50 as backbone. By grouping hyperspectral videos into several 3-band-image subsets, the proposed SiamMRAAN can be used to explore high-dimensional spectral information. We design a loss function to fuse the tracking results over these subsets to improve the tracking performance. Finally, experiments over 75 hyperspectral videos confirmed that using spectral information is critical to improve the performance of object tracking in color videos, and also demonstrated that the proposed SiamMRAAN based strategy outperforms several compared networks for hyperspectral videos.
Ye Wang 0020, Shaohui Mei, Qian Du 0001
IGARSS1