Chen Wu 0003

dblp:78/3213-3 · DBLP profile ↗
← Back
52ranked-venue papers
7as first author
37since 2021 · last 2026
0000-0001-6461-8377ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 34 · 2 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2026 From segmentation to change: Releasing segment anything model for remote sensing change detection
Kaixuan Jiang, Chen Wu 0003, Zhenghui Zhao, Bo Du 0001, Liangpei Zhang
Pattern Recognit.2
2025 Spatial-Spectral Graph Convolutional Network for Hyperspectral Target Detection
abstract
Deep learning-based hyperspectral target detection methods commonly face challenges such as insufficient target samples and inadequate use of spatial context. To address these limitations, we propose a novel hyperspectral target detection approach leveraging spatial-spectral graph convolutional networks. First, we introduce an innovative sample augmentation strategy utilizing pre-detection and target implantation, which can simulate different backgrounds around the target and expand the target sample, thereby enhancing the representation ability of the model. Next, a graph-based representation strategy is proposed to integrate spatial and spectral information. Finally, we develop four specialized graph network detectors (GCND, GATD, GCN-GAT1, and GCN-GAT2) with structural optimizations involving multi-scale feature fusion and dynamic neighborhood adjustments. Extensive experiments demonstrate the superior detection performance of our method compared to existing techniques, highlighting its practical significance.
Chenxing Li, Dehui Zhu, Chen Wu 0003
IEEE Geosci. Remote. Sens. Lett.3
2025 HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation Model
abstract
Accurate hyperspectral image (HSI) interpretation is critical for providing valuable insights into various earth observation-related applications such as urban planning, precision agriculture, and environmental monitoring. However, existing HSI processing methods are predominantly task-specific and scene-dependent, which severely limits their ability to transfer knowledge across tasks and scenes, thereby reducing the practicality in real-world applications. To address these challenges, we present HyperSIGMA, a vision transformer-based foundation model that unifies HSI interpretation across tasks and scenes, scalable to over one billion parameters. To overcome the spectral and spatial redundancy inherent in HSIs, we introduce a novel sparse sampling attention (SSA) mechanism, which effectively promotes the learning of diverse contextual features and serves as the basic block of HyperSIGMA. HyperSIGMA integrates spatial and spectral features using a specially designed spectral enhancement module. In addition, we construct a large-scale hyperspectral dataset, HyperGlobal-450K, for pre-training, which contains about 450 K hyperspectral images, significantly surpassing existing datasets in scale. Extensive experiments on various high-level and low-level HSI tasks demonstrate HyperSIGMA's versatility and superior representational capability compared to current state-of-the-art methods. Moreover, HyperSIGMA shows significant advantages in scalability, robustness, cross-modal transferring capability, real-world applicability, and computational efficiency.
Di Wang 0023, Meiqi Hu, Yuchun Miao, Jiaqi Yang 0005, Yichu Xu, Xiaolei Qin, Jiaqi Ma 0002, Chenxing Li, Chuan Fu, Hongruixuan Chen, Chengxi Han, Naoto Yokoya, Jing Zhang 0037, Minqiang Xu, Lefei Zhang, Chen Wu 0003, Bo Du 0001, Dacheng Tao, Liangpei Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.19
2025 ACR-Net: Adaptive Correlation Refined Hyperspectral Unmixing
abstract
Hyperspectral unmixing aims to resolve the prevalent issue of mixed pixels in hyperspectral imagery and serves as an effective technique for sub-pixel level image interpretation. Recent years have seen the emergence of advanced unmixing algorithms that integrate both spatial and spectral information. However, existing methods mainly focus on spatial context and lack depth in modeling spatial correlations. Both relevant and irrelevant spatial information is introduced into the spectral mixing model for local pixels, with the irrelevant information acting as noise that impacts the unmixing accuracy. To address these challenges, we propose an advanced spectral mixing model, Adaptive Correlation Refined Network (ACR-Net) which integrates refined spatial correlation based on self-attention. A key component of ACR-Net is the Adaptive Correlation Aggregated Decoder (ACAD), which extracts affinity information from the encoder’s feature map and adaptively amplifies the influence of highly correlated regions in the unmixing process. Additionally, the Composite Active Spatial Attention (CASA) Module emphasizes unique spectral characteristics across bands, improving spatial distribution representation and enabling more accurate abundance estimation. We conducted abundant evaluations on six datasets, including three real-world and two synthetic hyperspectral unmixing datasets as well as a large benchmark classification dataset. Extensive experiments have demonstrated that the proposed algorithm can achieve exceptional or comparable unmixing results among various state-of-the-art algorithms. Moreover, the proposed method achieved optimal classification performance on the classic PaviaU dataset, indicating its strong potential for hyperspectral classification task.
Meiqi Hu, Chen Wu 0003, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 LGCANet: Local-Global and Change-Aware Network via Segment Anything Model for Remote Sensing Images Change Detection
abstract
Change detection (CD) is a very fundamental and challenging task in remote sensing. Many deep learning-based CD methods generally utilize Siamese networks to extract image features. However, the semantic features extracted by these methods are still not fine-grained. In addition, these CD methods ignore the object scale diverse in remote sensing images and the interaction information between bi-temporal images, which leads to the problem that the network is unable to capture more efficient feature embeddings, with ambiguous or erroneous detection results. To alleviate the above issues, we propose Local-Global and Change Aware Network via Fast Segment Anything Model (LGCANet). The Segment Everything Model (SAM) can accurately segment objects in various scene images. In this work, we intend to utilize the powerful recognition capabilities of SAM to refine the CD task. Therefore, LGCANet employs more efficient FastSAM and ResNet as encoders to extract potential feature representations in remote sensing images. FastSAM can effectively extract global contextual information, combined with ResNet’s powerful deep feature extraction capability, which enables the network to comprehensively model features. LGCANet contains three modules: content aware attention module (CAAM), fore-background aware module (FAM), and edge-reinforce hybrid-selection module (EHM). CAAM delivers feature extraction from local to global perception, realizing dynamic attention to various scales of objects. FAM can effectively learn foreground and background representations through feature interaction, which significantly enhances the model’s capability of recognizing changed regions. EHM can utilize direction-awareness to extract edge information and generate fine-grained detection maps by adaptively selecting discriminative features through designed attention mechanisms. Experiments on publicly available CD datasets show that LGCANet achieves superior detection performance compared to other state-of-the-art methods. The code is available at https://github.com/Jscript10/LGCANet.
Kaixuan Jiang, Chen Wu 0003
IEEE Trans. Geosci. Remote. Sens.2
2025 Essential Hierarchical Information That Warrants Attention: A Semi-Supervised Hyperspectral Hierarchical Classification Network
abstract
The hierarchical classification method can utilize the hierarchical structure to achieve fine classification from coarse-grained to fine-grained. Hierarchical classification has demonstrated outstanding performance across various fields, benefiting from the continuous expansion of dataset sizes. Nevertheless, the majority of current hyperspectral image (HSI) deep learning classification methods employ flat classification strategy, disregarding the hierarchical structure present in HSI data. Can hierarchical structure information assist deep learning in completing HSI classification tasks? In this paper, we validated the feasibility of using hierarchical structure information to assist deep learning in completing HSI classification tasks through intuitive experiments and then proposed a semi-supervised hierarchical classification network (Semi-HCN) to achieve hierarchical classification of HSI. Semi-HCN first generates a large number of pseudo-labels through a fast self-training module for subsequent network training, then extracts hierarchical features through a multi-branch network and progressively consolidates them across layers. Finally, hierarchical constraints are applied to the network output through hierarchical cross entropy loss. Experiments on three real datasets have demonstrated that Semi-HCN can achieve competitive classification performance with limited training samples. The code will be released at https://github.com/jinyaoWHU/Semi-HCN.
Yanni Dong, Chen Wu 0003
IEEE Trans. Geosci. Remote. Sens.3
2025 Embracing Hierarchical Classification: A Multifeature Space Hierarchical Network for Hyperspectral Image Classification
abstract
Classes in hyperspectral images (HSI) often exhibit inherent hierarchical structures, such as the family–genus–species hierarchy in tree species. Previous studies have shown that modeling these hierarchical structures can improve classification performance. However, existing deep learning methods often overlook such structures, limiting their ability to distinguish fine-grained categories. To address this issue, we propose a Multi-Feature Space Hierarchical Network (MFS-HiNet), which enhances fine-grained discrimination by modeling hierarchical relationships and guiding classification top-down. The framework consists of three key components: Hierarchical Structure Mining (HSM), Multi-Feature Space Classification Network (MFSCN), and Parameter Inheritance Strategy (PIS). Specifically, HSM automatically mines latent hierarchical relationships in HSI and constructs the hierarchy; MFSCN integrates multi-feature space information for node classification to improve the discrimination of subtle inter-class differences; and PIS leverages parent node parameters to guide the learning of child nodes, fully exploiting inter-level correlations. Experimental results on five HSI datasets demonstrate that MFS-HiNet outperforms existing methods in classification accuracy, validating the effectiveness of the framework and highlighting the potential of hierarchical classification. The source code will be made available at https://github.com/jin-yaoWHU/ MFS-HiNet.
Chen Wu 0003
IEEE Trans. Geosci. Remote. Sens.2
2025 TransWCD: Scene-Adaptive Joint Constrained Framework for Weakly Supervised Change Detection
abstract
Change detection (CD) based on deep learning typically requires costly pixel-level change labels. Recently, weakly supervised CD (WSCD) has emerged as a more label-efficient approach, using scene-level (i.e., image-level) labels to identify pixel-level changes in bitemporal images. With only scene-level labels, existing WSCD methods are typically trained as scene-level change classification models. However, these methods often suffer from label-prediction inconsistency, with false changes frequently predicted in unchanged scenes. To address this issue, we propose TransWCD-SA, an end-to-end classifier-predictor framework. TransWCD-SA consists of a hierarchical transformer-based TransWCD classifier and a scene-adaptive (SA) predictor. This classifier-predictor framework is trained with two-stage joint constraints in an end-to-end learning manner. Specifically, the TransWCD classifier integrates hierarchical transformer blocks and multiscale class activation maps (CAMs), capturing pixel-level changes across various scales under weak supervision. The SA predictor dynamically introduces different pixel-level information for scenes labeled as changed and unchanged. Furthermore, a scene gated constraint is proposed as a penalty for label-prediction inconsistency, which is activated by the Dirac delta function and rectify features of mispredicted pixels in the embedding space. We validate the effectiveness of TransWCD-SA on three datasets: Wuhan University building CD (WHU-CD), learning, vision, and remote sensing CD (LEVIR-CD), and DSIFN-CD, demonstrating significant improvement. The code is available athttps://github.com/zhenghuizhao/TransWCD.
Zhenghui Zhao, Lixiang Ru, Chen Wu 0003, Di Wang 0023
IEEE Trans. Geosci. Remote. Sens.3
2025 Advancing Weakly-Supervised Change Detection in Satellite Images via Adversarial Class Prompting
abstract
Weakly-Supervised Change Detection (WSCD) aims to distinguish specific object changes (e.g., objects appearing or disappearing) from background variations (e.g., environmental changes due to light, weather, or seasonal shifts) in paired satellite images, relying only on paired image (i.e., image-level) classification labels. This technique significantly reduces the need for dense annotations required in fully-supervised change detection. However, as image-level supervision only indicates whether objects have changed in a scene, WSCD methods often misclassify background variations as object changes, especially in complex remote-sensing scenarios. In this work, we propose an Adversarial Class Prompting (AdvCP) method to address this co-occurring noise problem, including two phases: a) Adversarial Prompt Mining: After each training iteration, we introduce adversarial prompting perturbations, using incorrect one-hot image-level labels to activate erroneous feature mappings. This process reveals co-occurring adversarial samples under weak supervision, namely background variation features that are likely to be misclassified as object changes. b) Adversarial Sample Rectification: We integrate these adversarially prompt-activated pixel samples into training by constructing an online global prototype. This prototype is built from an exponentially weighted moving average of the current batch and all historical training data. Serving as an unbiased anchor, the global prototype guides the rectification of adversarial pixel samples. Our AdvCP can be seamlessly integrated into current WSCD methods without adding additional inference cost. Experiments on ConvNet, Transformer, and Segment Anything Model (SAM)-based baselines demonstrate significant performance enhancements, achieving up to 7.37%, 7.46%, and 6.56% IoU improvements on the WHU-CD, LEVIR-CD, and DSIFN-CD datasets. Furthermore, we demonstrate the generalizability of AdvCP to other multi-class weakly-supervised dense prediction scenarios. Code is available at https://github.com/zhenghuizhao/AdvCP.
Zhenghui Zhao, Chen Wu 0003, Di Wang 0023, Hongruixuan Chen, Cuiqun Chen, Zhuo Zheng, Bo Du 0001, Liangpei Zhang 0001
IEEE Trans. Image Process.2
2024 Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models
abstract
Recently, there has been a surge in interest in Large Language Models (LLMs), with ChatGPT standing out for its exceptional capabilities in language comprehension, reasoning, and interactive communication. These models have garnered attention from a diverse array of users and researchers across various disciplines. While LLMs have demonstrated remarkable proficiency in mimicking human task execution through natural language, their application in remote sensing interpretation remains largely uncharted. Furthermore, the current lack of automation in remote sensing task planning limits the accessibility of these sophisticated interpretation techniques, especially for non-specialists in the field. To bridge this gap, we introduce Remote Sensing ChatGPT, an innovative LLM-driven agent that integrates ChatGPT with a suite of AI-powered remote sensing models to tackle complex interpretation challenges. This system is designed to interpret user requests, delineate task planning based on the functionalities required, execute each subtask sequentially, and compile the final output by synthesizing the results from each stage. Given that LLMs, trained predominantly on natural language, do not inherently comprehend visual elements present in remote sensing imagery, we have devised a method to incorporate visual cues, effectively embedding the visual context of remote sensing images into the ChatGPT framework. With Remote Sensing ChatGPT, users can effortlessly submit a remote sensing image alongside their query and promptly receive detailed interpretation outcomes along with comprehensive linguistic feedback. Experiments and case studies demonstrate that our method is adept at handling a diverse range of remote sensing tasks and has the potential to be expanded to encompass an even wider array of applications with the integration of more advanced models, such as remote sensing foundation model. The code and demo of Remote Sensing ChatGPT is publicly available at https://github.com/HaonanGuo/Remote-Sensing-ChatGPT.
Xin Su 0003, Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001, DeRen Li
IGARSS3
2024 OFA-NET: Optical Flow Aligning Network for Time Series Change Detection
abstract
With the rapid advancement of remote sensing and Earth observation technologies, a plethora of Time Series remote sensing Images (TSIs) from platforms like Landsat and Sentinel-2 have become accessible, providing valuable data for Time Series remote sensing images Change Detection (TSCD). In TSCD, TSIs captured at the same geographic location but at different times often exhibit misalignment issues due to variations in radiation incidence angles, satellite orbit deviations, and other factors. To tackle these challenges, an Optical Flow Aligning Network (OFA-NET) for TSCD is proposed in this paper. In detail, a simple and lightweight Encoder-Decoder is introduced to extract multiscale spatial features from TSIs and restore these features to the size of the original input of Encoder-Decoder. Moreover, Encoder-Decoder employs the Siamese structure to effectively preserve the original features of each input image as much as possible. Subsequently, the outputs of Encoder-Decoder are utilized for optical flow prediction, which is further employed to align TSIs and extract differences between them for TSCD. Experiments conducted on the UTRNet datasets illustrate that the proposed OFA-NET achieves superior accuracy and produces more distinct change detection results compared to other methods.
Chen Wu 0003
IGARSS2
2024 Game-Theoretic Internal Competitive Learning for Scene-to-Pixel Weakly-Supervised Change Detection
abstract
Weakly supervised change detection (WSCD), which leverages scene-level labels as supervision signals to identify pixel-level changes in multi-temporal images, has recently garnered increasing attention. Existing weakly-supervised change detection tasks typically rely only on scene-level classification, using cross-entropy loss. However, this approach fails to account for correlations between pixels, leading to a lack of clear distinction between changed and unchanged pixel-level features. In this paper, we propose Internal Competitive (IC) learning, an approach that incorporates a game-theoretic framework. This method implements a dynamic competition mechanism between changed and unchanged regions. It does so by utilizing the similarity between pixel-wise changed and unchanged features within the differential feature embedding space as a metric for decision-making. Experiments conducted on the WHU-CD and LEVIR-CD datasets demonstrate the effectiveness of our IC learning approach in both ConvNet-based and Transformer-based models.
Zhenghui Zhao, Chen Wu 0003
IGARSS2
2024 Adversarial pair-wise distribution matching for remote sensing image cross-scene classification
Sihan Zhu, Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001
Neural Networks2
2024 Building-Road Collaborative Extraction From Remote Sensing Images via Cross-Task and Cross-Scale Interaction
abstract
Buildings and roads are the two most basic man-made environments that carry and interconnect human society. Building and road information has important application value in the frontier fields of regional coordinated development, disaster prevention, auto-driving, etc. Mapping buildings and roads from very high-resolution (VHR) remote sensing images has become a hot research topic. However, the existing methods often extract buildings and roads with separate models, ignoring their strong spatial correlation. To fully utilize their complementary relation, we propose a method that simultaneously extracts buildings and roads from remote sensing images. The accuracy of both tasks can be improved using our proposed multi-task feature interaction and cross-scale feature interaction modules. To be specific, a multi-task interaction module is proposed to interact information across building extraction and road extraction tasks while preserving the unique information of each task. Furthermore, a cross-scale interaction module is designed to automatically learn the optimal reception field for buildings and roads under varied appearances and structures. Compared with existing methods that train individual models for each task separately, the proposed collaborative extraction method can utilize the complementary advantages between buildings and roads and reduce the inference time by half using a single model. Experiments on a wide range of urban and rural scenarios show that the proposed algorithm can achieve building-road extraction with outstanding performance and efficiency.
Xin Su 0003, Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 C2F-SemiCD: A Coarse-to-Fine Semi-Supervised Change Detection Method Based on Consistency Regularization in High-Resolution Remote Sensing Images
abstract
A high-precision feature extraction model is crucial for change detection. In the past, many deep learning-based supervised change detection methods learned to recognize change feature patterns from a large number of labelled bi-temporal images, whereas labelling bi-temporal remote sensing images is very expensive and often time-consuming. Therefore, we propose a coarse-to-fine semi-supervised change detection method based on consistency regularization (C2F-SemiCD), which includes a coarse-to-fine change detection network with a multi-scale attention mechanism(C2FNet) and a semi-supervised update method. Among them, the C2FNet network "gradually" completes the extraction of change features from coarse-grained to fine-grained through multi-scale feature fusion, channel attention mechanism, spatial attention mechanism, global context module, feature refine module, initial aggregation module, and final aggregation module. The semi-supervised update method uses the mean teacher method. The parameters of the student model are updated to the parameters of the teacher Model by using the exponential moving average (EMA) method. Through extensive experiments on three datasets and meticulous ablation studies, including crossover experiments across datasets, we verify the significant effectiveness and efficiency of the proposed C2F-SemiCD method. The code will be open at: https://github.com/ChengxiHAN/C2F-SemiCD-and-C2FNet.
Chengxi Han, Chen Wu 0003, Meiqi Hu, Jiepan Li, Hongruixuan Chen
IEEE Trans. Geosci. Remote. Sens.2
2024 Global Overcomplete Dictionary-Based Sparse and Nonnegative Collaborative Representation for Hyperspectral Target Detection
abstract
The combined sparse and collaborative representation-based algorithm is one of the most effective methods among hyperspectral target detection methods based on representation and dictionary learning. It encourages target atoms to compete with each other and background atoms to collaborate in the representation. However, this method suffers from several drawbacks. In sparse representation, an overcomplete dictionary is necessary, whereas, in collaborative representation, non-negative coefficients are required. Besides, the local dual window approach may result in impure background dictionaries obtained from the outer window. To address these issues, we propose a novel approach for hyperspectral target detection, referred to as the global overcomplete dictionary-based sparse and nonnegative collaborative representation (GODSNCR) detector. First, a hierarchical density clustering algorithm is used to complete the dictionary atom extraction to construct a joint overcomplete dictionary to satisfy the dictionary overcompleteness problem required for sparse representation. Second, a nonnegative constraint on the coefficient matrix and a “sum to one” constraint for the joint representation are incorporated to make it more consistent with the physical meaning. Finally, the limitation of the local dual window approach is overcome by substituting the local background dictionary with a global background dictionary. Through the aforementioned strategies, we can use a joint overcomplete dictionary for achieving the sparse representation of targets and utilize a global background dictionary for the collaborative representation of background, the final detection results are obtained by calculating the residuals. The experimental results clearly demonstrate that the proposed algorithm has significant improvement in detection accuracy and strong robustness compared to other typical representation-based hyperspectral target detection methods. Our model will be available at https://github.com/Chenxing-Li/GODSNCR.
Chenxing Li, Dehui Zhu, Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Robust Remote Sensing Image Cross-Scene Classification Under Noisy Environment
abstract
In recent years, great progress has been made in the field of cross-scene classification. However, most existing cross-scene methods assume that the source domain has massive and carefully annotated data, which is time-consuming and labor-intensive in practice. Datasets in real-life applications usually contain a large number of noisy labels, which will significantly affect cross-scene classification performance. How to perform more discriminative and generalized cross-scene classification in the presence of noisy samples needs to be urgently addressed. Apart from that, existing methods tend to implement global matching between domains, causing problems such as unbalanced adaptation and negative transfer, limiting the cross-scene performance of the model. For more effective and reliable cross-scene classification under noisy environment, robust adaptation with noise (RAN) is proposed in this article. RAN explores which samples are noiseless and transferable to enable positive and robust cross-scene transfer. The curriculum learning strategy is used to filter out noisy samples for better source supervised learning and cross-domain matching. To further improve the stability and effectiveness of cross-scene adaptation, the class weighting factor and the public weighting factor are introduced to consider the class information of the source and target domains. RAN is an efficient plug-and-play adaptation framework, which is easily implemented and can be embedded in existing methods. Experimental results demonstrate that the proposed RAN can achieve remarkable performance on cross-scene classification tasks in noisy environments.
Sihan Zhu, Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 SAAN: Similarity-Aware Attention Flow Network for Change Detection With VHR Remote Sensing Images
abstract
Change detection (CD) is a fundamental and important task for monitoring the land surface dynamics in the earth observation field. Existing deep learning-based CD methods typically extract bi-temporal image features using a weight-sharing Siamese encoder network and identify change regions using a decoder network. These CD methods, however, still perform far from satisfactorily as we observe that 1) deep encoder layers focus on irrelevant background regions; and 2) the models' confidence in the change regions is inconsistent at different decoder stages. The first problem is because deep encoder layers cannot effectively learn from imbalanced change categories using the sole output supervision, while the second problem is attributed to the lack of explicit semantic consistency preservation. To address these issues, we design a novel similarity-aware attention flow network (SAAN). SAAN incorporates a similarity-guided attention flow module with deeply supervised similarity optimization to achieve effective change detection. Specifically, we counter the first issue by explicitly guiding deep encoder layers to discover semantic relations from bi-temporal input images using deeply supervised similarity optimization. The extracted features are optimized to be semantically similar in the unchanged regions and dissimilar in the changing regions. The second drawback can be alleviated by the proposed similarity-guided attention flow module, which incorporates similarity-guided attention modules and attention flow mechanisms to guide the model to focus on discriminative channels and regions. We evaluated the effectiveness and generalization ability of the proposed method by conducting experiments on a wide range of CD tasks. The experimental results demonstrate that our method achieves excellent performance on several CD tasks, with discriminative features and semantic consistency preserved.
Xin Su 0003, Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001
IEEE Trans. Image Process.3
2023 HCGMNet: A Hierarchical Change Guiding Map Network for Change Detection
abstract
Very-high-resolution (VHR) remote sensing (RS) image change detection (CD) has been a challenging task for its very rich spatial information and sample imbalance problem. In this paper, we have proposed a hierarchical change guiding map network (HCGMNet) for VHR RS change detection. The model uses hierarchical convolution operations to extract multi-scale features, continuously merges multi-scale features layer by layer to improve the expression of global and local information, and guides the model to gradually refine edge features and comprehensive performance by a change guide module (CGM), which is a self-attention with changing guide map. Extensive experiments on two CD datasets show that the proposed HCGMNet architecture achieves better CD performance than existing state-of-the-art (SOTA) CD methods.
Chengxi Han, Chen Wu 0003, Bo Du 0001
IGARSS2
2023 EMS-NET: Efficient Multi-Temporal Self-Attention for Hyperspectral Change Detection
abstract
Hyperspectral change detection plays an essential role of monitoring the dynamic urban development and detecting precise fine object evolution and alteration. In this paper, we have proposed an original Efficient Multi-temporal Self-attention Network (EMS-Net) for hyperspectral change detection. The designed EMS module cuts redundancy of those similar and containing-no-changes feature maps, computing efficient multi-temporal change information for precise binary change map. Besides, to explore the clustering characteristics of the change detection, a novel supervised contrastive loss is provided to enhance the compactness of the unchanged. Experiments implemented on two hyperspectral change detection datasets manifests the out-standing performance and validity of proposed method.
Meiqi Hu, Chen Wu 0003, Bo Du 0001
IGARSS2
2023 LGFormer: Local-to-Global Transformer for Hyperspectral Image Classification
abstract
Recently, many transformer-based approaches have emerged in the field of hyperspectral image (HSI) classification. However, existing transformer-based works either model the information of spectral vectors using a transformer or introduce a vision transformer (ViT) for feature expression of the spatial patch after principal component analysis (PCA), neglecting the spectral-spatial correlation of HSI. Besides, local details and global distributions are always challenging to simultaneously extract in these methods. To address the above issues, a local-to-global transformer (LGFormer) is proposed in this paper. In detail, the proposed approach directly extracts inherent features on the originally spectral-spatial patch, which not only takes the spatial distribution and spectral continuity into account but also preserves the intrinsically spectral-spatial correlation of the HSI cube. Moreover, a local-to-global self-attention (LGSA) including 3-D convolutional neural network (CNN) and ViT is designed in the presented LGFormer. With the above HSI-tailored structure, both fine- and coarse-grained features can be captured by progressively spectral-spatial feature learning. Experimental results on benchmark HSI datasets demonstrate that the proposed LGFormer can outperform other methods in terms of higher accuracy and finer classification maps.
Jiaqi Yang 0005, Bo Du 0001, Chen Wu 0003
IGARSS3
2023 Fully Convolutional Change Detection Framework With Generative Adversarial Network for Unsupervised, Weakly Supervised and Regional Supervised Change Detection
abstract
Deep learning for change detection is one of the current hot topics in the field of remote sensing. However, most end-to-end networks are proposed for supervised change detection, and unsupervised change detection models depend on traditional pre-detection methods. Therefore, we proposed a fully convolutional change detection framework with generative adversarial network, to unify unsupervised, weakly supervised, regional supervised, and fully supervised change detection tasks into one end-to-end framework. A basic Unet segmentor is used to obtain change detection map, an image-to-image generator is implemented to model the spectral and spatial variation between multi-temporal images, and a discriminator for changed and unchanged is proposed for modeling the semantic changes in weakly and regional supervised change detection task. The iterative optimization of segmentor and generator can build an end-to-end network for unsupervised change detection, the adversarial process between segmentor and discriminator can provide the solutions for weakly and regional supervised change detection, the segmentor itself can be trained for fully supervised task. The experiments indicate the effectiveness of the propsed framework in unsupervised, weakly supervised and regional supervised change detection. This article provides new theorical definitions for unsupervised, weakly supervised and regional supervised change detection tasks with the proposed framework, and shows great potentials in exploring end-to-end network for remote sensing change detection (https://github.com/Cwuwhu/FCD-GAN-pytorch).
Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Decoupling Semantic and Edge Representations for Building Footprint Extraction From Remote Sensing Images
abstract
Very high-resolution (VHR) earth observation systems provide an ideal data source for man-made structure detection such as building footprint extraction. Manually delineating building footprints from the remotely sensed VHR images, however, is laborious and time-intensive; thus, automation is needed in the building extraction process to increase productivity. Recently, many researchers have focused on developing building extraction algorithms based on the encoder-decoder architecture of convolutional neural networks. However, we observe that this widely adopted architecture cannot well preserve the precise boundaries and integrity of the extracted buildings. Moreover, features obtained by shallow convolutional layers contain irrelevant background noises that degrade building feature representations. This paper addresses these problems by presenting a feature decoupling network (FD-Net) that exploits two essential building information from the input image, including semantic information that concerns building integrity and edge information that improves building boundaries. The proposed FD-Net improves the existing encoder-decoder framework by decoupling image features into the edge subspace and the semantic subspace; the decoupled features are then integrated by a supervision-guided fusion process considering the heterogeneity between edge and semantic features. Furthermore, a lightweight and effective global context attention module is introduced to capture contextual building information and thus enhance feature representations. Comprehensive experimental results on three real-world datasets confirm the effectiveness of FD-Net in large-scale building mapping. We applied the proposed method to various encoder-decoder variants to verify the generalizability of the proposed framework. Experimental results show remarkable accuracy improvements with less computational cost.
Xin Su 0003, Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 Adversarial Divergence Training for Universal Cross-Scene Classification
abstract
Cross-scene classification has recently gained increasing interest, which improves the classification performance on label-scarce domains by transferring knowledge learned from label-rich domains. Domain adaptation (DA) attempts to solve the domain gap problem and it is widely used in the cross-scene applications. The priori knowledge that the label space between the source and target domains is identical is a key requirement for existing cross-scene approaches to work successfully. However, label sets of different domains can always be different in real applications, that is to say, there will be common as well as private categories for different domains. Universal DA (UniDA) has been proposed to deal with the above difficulty by relaxing all constraints on the label sets. In order to complete more general remote sensing cross-scene classification tasks regardless of label sets, we propose a UniDA cross-scene classification approach, adversarial divergence training (ADT), to simultaneously classify the target common categories and detect the target private categories based on the divergence of different classifiers. ADT attempts to train the classifier and feature extractor against each other (in adversarial) in order to extract more domain-invariant and discriminative features. At the same time, divergence optimization of different classifiers is used to distinguish the target private class. The former makes it capable of cross-scene tasks, while the latter weakens the effect of the label set on the performance of the algorithm. Experiments show that ADT outperforms baselines in the UniDA setting and even in other settings.
Sihan Zhu, Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Binary Change Guided Hyperspectral Multiclass Change Detection
abstract
Characterized by tremendous spectral information, hyperspectral image is able to detect subtle changes and discriminate various change classes for change detection. The recent research works dominated by hyperspectral binary change detection, however, cannot provide fine change classes information. And most methods incorporating spectral unmixing for hyperspectral multiclass change detection (HMCD), yet suffer from the neglection of temporal correlation and error accumulation. In this study, we proposed an unsupervised Binary Change Guided hyperspectral multiclass change detection Network (BCG-Net) for HMCD, which aims at boosting the multiclass change detection result and unmixing result with the mature binary change detection approaches. In BCG-Net, a novel partial-siamese united-unmixing module is designed for multi-temporal spectral unmixing, and a groundbreaking temporal correlation constraint directed by the pseudo-labels of binary change detection result is developed to guide the unmixing process from the perspective of change detection, encouraging the abundance of the unchanged pixels more coherent and that of the changed pixels more accurate. Moreover, an innovative binary change detection rule is put forward to deal with the problem that traditional rule is susceptible to numerical values. The iterative optimization of the spectral unmixing process and the change detection process is proposed to eliminate the accumulated errors and bias from unmixing result to change detection result. The experimental results demonstrate that our proposed BCG-Net could achieve comparative or even outstanding performance of multiclass change detection among the state-of-the-art approaches and gain better spectral unmixing results at the same time.
Meiqi Hu, Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001
IEEE Trans. Image Process.2
2022 Dual-Tasks Siamese Transformer Framework for Building Damage Assessment
abstract
Accurate and fine-grained information about the extent of damage to buildings is essential for humanitarian relief and disaster response. However, as the most commonly used architecture in remote sensing interpretation tasks, Convolutional Neural Networks (CNNs) have limited ability to model the non-local relationship between pixels. Recently, Transformer architecture first proposed for modeling long-range dependency in natural language processing has shown promising results in computer vision tasks. Considering the frontier advances of Transformer architecture in the computer vision field, in this paper, we present a Transformer-based damage assessment architecture (DamFormer). In DamFormer, a siamese Transformer encoder is first constructed to extract non-local and representative deep features from input multitemporal image-pairs. Then, a multitemporal fusion module is designed to fuse information for downstream tasks. Finally, a lightweight dual-tasks decoder aggregates multi-level features for final prediction. To the best of our knowledge, it is the first time that such a deep Transformer-based network is proposed for multitemporal remote sensing interpretation tasks. The experimental results on the large-scale damage assessment dataset xBD demonstrate the potential of the Transformer-based architecture.
Hongruixuan Chen, Edoardo Nemni, Sofia Vallecorsa, Chen Wu 0003, Lars Bromley
IGARSS5
2022 Multi-Temporal Spatial-Spectral Comparison Network For Hyperspectral Anomalous Change Detection
abstract
Hyperspectral anomalous change detection has been a challenging task for its emphasis on the dynamics of small and rare objects against the prevalent changes. In this paper, we have proposed a Multi-Temporal spatial-spectral Comparison Network for hyperspectral anomalous change detection (MTC-NET). The whole model is a deep siamese network, aiming at learning the prevalent spectral difference resulting from the complex imaging conditions from the hyperspectral images by contrastive learning. A three-dimensional spatial spectral attention module is designed to effectively extract the spatial semantic information and the key spectral differences. Then the gaps between the multi-temporal features are minimized, boosting the alignment of the semantic and spectral features and the suppression of the multi-temporal background spectral difference. The experiments on the “Viareggio 2013” datasets demonstrate the effectiveness of proposed MTC-NET.
Meiqi Hu, Chen Wu 0003, Bo Du 0001
IGARSS2
2022 Hybrid Vision Transformer Model for Hyperspectral Image Classification
abstract
Due to the local connectivity property, convolutional neural network (CNN) can effectively extract contextual detailed information. Therefore, a large number of CNN-based methods are introduced to hyperspectral image (HSI) classification. However, receptive fields of these methods are greatly limited, and information extraction process is usually inadequate. Recently, transformer structure has attracted extensive attention owing to its ability to capture global dependency. With a self-attention mechanism, transformer can extract long-tail distribution and model global features to enhance the representation of data. Consequently, it is a natural idea to combine CNN and transformer to obtain both local detail and global distribution. In this paper, we propose a hybrid vision transformer model (Hybrid ViT) to jointly learn global and local information of HSI, including a convolution block and a vision transformer block. With the unified architecture, Hybrid ViT model can not only access detailed features of narrow targets but also extract the global distribution of large objects. Experimental results on benchmark HSI datasets demonstrate that the proposed Hybrid ViT can outperform other methods with higher classification accuracy and finer classification maps.
Jiaqi Yang 0005, Bo Du 0001, Chen Wu 0003
IGARSS3
2022 Weakly-Supervised Semantic Segmentation with Visual Words Learning and Hybrid Pooling
Lixiang Ru, Bo Du 0001, Yibing Zhan, Chen Wu 0003
Int. J. Comput. Vis.4
2022 Unsupervised Change Detection in Multitemporal VHR Images Based on Deep Kernel PCA Convolutional Mapping Network
abstract
With the development of Earth observation technology, a very-high-resolution (VHR) image has become an important data source of change detection (CD). These days, deep learning (DL) methods have achieved conspicuous performance in the CD of VHR images. Nonetheless, most of the existing CD models based on DL require annotated training samples. In this article, a novel unsupervised model, called kernel principal component analysis (KPCA) convolution, is proposed for extracting representative features from multitemporal VHR images. Based on the KPCA convolution, an unsupervised deep siamese KPCA convolutional mapping network (KPCA-MNet) is designed for binary and multiclass CD. In the KPCA-MNet, the high-level spatial-spectral feature maps are extracted by a deep siamese network consisting of weight-shared KPCA convolutional layers. Then, the change information in the feature difference map is mapped into a 2-D polar domain. Finally, the CD results are generated by threshold segmentation and clustering algorithms. All procedures of KPCA-MNet do not require labeled data. The theoretical analysis and experimental results in two binary CD datasets and one multiclass CD datasets demonstrate the validity, robustness, and potential of the proposed method.
Chen Wu 0003, Hongruixuan Chen, Bo Du 0001, Liangpei Zhang 0001
IEEE Trans. Cybern.1
2022 Unsupervised Multimodal Change Detection Based on Structural Relationship Graph Representation Learning
abstract
Unsupervised multimodal change detection is a practical and challenging topic that can play an important role in time-sensitive emergency applications. To address the challenge that multimodal remote sensing images cannot be directly compared due to their modal heterogeneity, we take advantage of two types of modality-independent structural relationships in multimodal images. In particular, we present a structural relationship graph representation learning framework for measuring the similarity of the two structural relationships. First, structural graphs are generated from preprocessed multimodal image pairs by means of an object-based image analysis approach. Then, a structural relationship graph convolutional autoencoder (SR-GCAE) is proposed to learn robust and representative features from graphs. Two loss functions aiming at reconstructing vertex information and edge information are presented to make the learned representations applicable for structural relationship similarity measurement. Subsequently, the similarity levels of two structural relationships are calculated from learned graph representations, and two difference images are generated based on the similarity levels. After obtaining the difference images, an adaptive fusion strategy is presented to fuse the two difference images. Finally, a morphological filtering-based postprocessing approach is employed to refine the detection results. Experimental results on six datasets with different modal combinations demonstrate the effectiveness of the proposed method.
Hongruixuan Chen, Naoto Yokoya, Chen Wu 0003, Bo Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 HyperNet: Self-Supervised Hyperspectral Spatial-Spectral Feature Understanding Network for Hyperspectral Change Detection
abstract
The fast development of self-supervised learning lowers the bar learning feature representation from massive unlabeled data and has triggered a series of researches on change detection of remote sensing images. Challenges in adapting self-supervised learning from natural images classification to remote sensing images change detection arise from difference between the two tasks. The learned patch-level feature representations are not satisfying for the pixel-level precise change detection. In this paper, we proposed a novel pixel-level self-supervised hyperspectral spatial-spectral understanding network (HyperNet) to accomplish pixel-wise feature representation for effective hyperspectral change detection. Concretely, not patches but the whole images are fed into the network and the multi-temporal spatial-spectral features are compared pixel by pixel. Instead of processing the two-dimensional imaging space and spectral response dimension in hybrid style, a powerful spatial-spectral attention module is put forward to explore the spatial correlation and discriminative spectral features of multi-temporal hyperspectral images (HSIs), separately. Only the positive samples at the same location of bi-temporal HSIs are created and forced to be aligned, aiming at learning the spectral difference-invariant features. Moreover, a new similarity loss function named focal cosine is proposed to solve the problem of imbalanced easy and hard positive samples comparison, where the weights of those hard samples are enlarged and highlighted to promote the network training. Six hyperspectral datasets have been adopted to test the validity and generalization of proposed HyperNet. The extensive experiments demonstrate the superiority of HyperNet over the state-of-the-art algorithms on downstream hyperspectral change detection tasks.
Meiqi Hu, Chen Wu 0003, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2021 Automatically Adjustable Multi-Scale Feature Extraction Framework for Hyperspectral Image Classification
abstract
Recently, deep learning-based methods have shown the great potential in hyperspectral image (HSI) classification. Nevertheless, feature extraction by convolutional neural network (CNN) is often performed on only one scale, resulting in multi-scale information loss. To address this problem, in this paper, we propose an automatically adjustable multi-scale feature extraction framework (A2MFE-Framework) for hyperspectral classification, including a scale reference network and two scale transformation networks. With the well-designed architecture, A2MFE-Framework can not only extract multiscale features, but also automatically change the network structure to match input features of different scales. Experimental results on two benchmark HSI datasets demonstrate that the A2MFE-Framework can better capture multi-scale features of different objects via an automatically adjustable feature extraction framework with higher classification accuracy compared with previous methods.
Jiaqi Yang 0005, Bo Du 0001, Chen Wu 0003, Liangpei Zhang 0001
IGARSS3
2021 Learning Visual Words for Weakly-Supervised Semantic Segmentation
abstract
Current weakly-supervised semantic segmentation (WSSS) methods with image-level labels mainly adopt class activation maps (CAM) to generate the initial pseudo labels. However, CAM usually only identifies the most discriminative object extents, which is attributed to the fact that the network doesn't need to discover the integral object to recognize image-level labels. In this work, to tackle this problem, we proposed to simultaneously learn the image-level labels and local visual word labels. Specifically, in each forward propagation, the feature maps of the input image will be encoded to visual words with a learnable codebook. By enforcing the network to classify the encoded fine-grained visual words, the generated CAM could cover more semantic regions. Besides, we also proposed a hybrid spatial pyramid pooling module that could preserve local maximum and global average values of feature maps, so that more object details and less background were considered. Based on the proposed methods, we conducted experiments on the PASCAL VOC 2012 dataset. Our proposed method achieved 67.2% mIoU on the val set and 67.3% mIoU on the test set, which outperformed recent state-of-the-art methods.
Lixiang Ru, Bo Du 0001, Chen Wu 0003
IJCAI3
2021 Enhanced Multiscale Feature Fusion Network for HSI Classification
abstract
Deep learning-based hyperspectral image (HSI) classification methods have recently attracted significant attention. However, features captured by convolutional neural network (CNN) are always partial due to the restrictions of the respective fields and the loss of multiscale information, which lead to features being discontinuous when extracted. In a departure from existing approaches, in this article, we propose a novel Enhanced Multiscale Feature Fusion Network (EMFFN). As a deeper and wider network, EMFFN can extract sufficiently multiscale features from the parallel multipath of three stages for HSI classification purposes. There are two subnetworks for multiscale spectral and spatial information in EMFFN, respectively. First, we propose a spectral Cascaded Dilated Convolutional Network (CDCN) designed to obtain a larger respective field for long-ranged information and extract multiscale features. Subsequently, a Parallel Multipath Network (PMN) is proposed to capture large-scale, middle-scale, and small-scale spatial features in parallel during all three stages. In the next step, hierarchical features are fused successively, and shallower feature maps can achieve better learning performance when guided by deeper semantic information. As PMN deepens in different stages, more multiscale information flows into the network, enabling finer classification results. To incorporate abundant spectral and spatial features, moreover, we combine features collected from two subnetworks into EMFFN using the designed consolidated loss function. As a result, the network facilitates the learning of not only localization-preserved features, but also high-level semantic features. In our experiments, three benchmark HSIs are utilized to evaluate the performance of the proposed method. Our results demonstrate that the proposed EMFFN can outperform state-of-the-art methods.
Jiaqi Yang 0005, Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2021 Multi-Temporal Scene Classification and Scene Change Detection With Correlation Based Fusion
abstract
Classifying multi-temporal scene land-use categories and detecting their semantic scene-level changes for remote sensing imagery covering urban regions could straightly reflect the land-use transitions. Existing methods for scene change detection rarely focus on the temporal correlation of bi-temporal features, and are mainly evaluated on small scale scene change detection datasets. In this work, we proposed a CorrFusion module that fuses the highly correlated components in bi-temporal feature embeddings. We first extract the deep representations of the bi-temporal inputs with deep convolutional networks. Then the extracted features will be projected into a lower-dimensional space to extract the most correlated components and compute the instance-level correlation. The cross-temporal fusion will be performed based on the computed correlation in CorrFusion module. The final scene classification results are obtained with softmax layers. In the objective function, we introduced a new formulation to calculate the temporal correlation more efficiently and stably. The detailed derivation of backpropagation gradients for the proposed module is also given. Besides, we presented a much larger scale scene change detection dataset with more semantic categories and conducted extensive experiments on this dataset. The experimental results demonstrated that our proposed CorrFusion module could remarkably improve the multi-temporal scene classification and scene change detection results.
Lixiang Ru, Bo Du 0001, Chen Wu 0003
IEEE Trans. Image Process.3
2021 HRSiam: High-Resolution Siamese Network, Towards Space-Borne Satellite Video Tracking
abstract
Tracking moving objects from space-borne satellite videos is a new and challenging task. The main difficulty stems from the extremely small size of the target of interest. First, because the target usually occupies only a few pixels, it is hard to obtain discriminative appearance features. Second, the small object can easily suffer from occlusion and illumination variation, making the features of objects less distinguishable from features in surrounding regions. Current state-of-the-art tracking approaches mainly consider high-level deep features of a single frame with low spatial resolution, and hardly benefit from inter-frame motion information inherent in videos. Thus, they fail to accurately locate such small objects and handle challenging scenarios in satellite videos. In this article, we successfully design a lightweight parallel network with a high spatial resolution to locate the small objects in satellite videos. This architecture guarantees real-time and precise localization when applied to the Siamese Trackers. Moreover, a pixel-level refining model based on online moving object detection and adaptive fusion is proposed to enhance the tracking robustness in satellite videos. It models the video sequence in time to detect the moving targets in pixels and has ability to take full advantage of tracking and detecting. We conduct quantitative experiments on real satellite video datasets, and the results show the proposed HIGH-RESOLUTION SIAMESE NETWORK (HRSiam) achieves state-of-the-art tracking performance while running at over 30 FPS.
Jia Shao, Bo Du 0001, Chen Wu 0003, Mingming Gong, Tongliang Liu
IEEE Trans. Image Process.3
2020 Change Detection in Multisource VHR Images via Deep Siamese Convolutional Multiple-Layers Recurrent Neural Network
abstract
With the rapid development of Earth observation technology, very-high-resolution (VHR) images from various satellite sensors are more available, which greatly enrich the data source of change detection (CD). Multisource multitemporal images can provide abundant information on observed landscapes with various physical and material views, and it is exigent to develop efficient techniques to utilize these multisource data for CD. In this article, we propose a novel and general deep siamese convolutional multiple-layers recurrent neural network (RNN) (SiamCRNN) for CD in multitemporal VHR images. Superior to most VHR image CD methods, SiamCRNN can be used for both homogeneous and heterogeneous images. Integrating the merits of both convolutional neural network (CNN) and RNN, SiamCRNN consists of three subnetworks: deep siamese convolutional neural network (DSCNN), multiple-layers RNN (MRNN), and fully connected (FC) layers. The DSCNN has a flexible structure for multisource image and is able to extract spatial-spectral features from homogeneous or heterogeneous VHR image patches. The MRNN stacked by long-short term memory (LSTM) units is responsible for mapping the spatial-spectral features extracted by DSCNN into a new latent feature space and mining the change information between them. In addition, FC, the last part of SiamCRNN, is adopted to predict change probability. The experimental results in two homogeneous data sets and one challenging heterogeneous VHR images data set demonstrate that the promising performances of the proposed network outperform several state-of-the-art approaches.
Hongruixuan Chen, Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2019 PASiam: Predicting Attention Inspired Siamese Network, for Space-Borne Satellite Video Tracking
abstract
Tracking a moving target of interests from a space-borne satellite video is really challenging. The difficulty lies in that the target usually occupies only several pixels, so that its features are very difficult to obtain. Besides, appearance features of the target would be unobvious when it is occluded, suffers from illumination variation influence or moves to similarity surroundings. In this paper, we propose a PREDICTING ATTENTION Inspired SIAMESE NETWORK (PASiam) for space-borne satellite video tracking, which constructs a fully convolutional Siamese network with shallow-layer features to obtain fine-grained appearance features. Moreover, a predicting attention is proposed to deal with occlusion and obscure. It employs Gaussian mixture models (GMM) to detect the target's motion status, and Kalman filter to predict and correct the target's location. Quantitative evaluations are performed on three real satellite video datasets. The results show our approach outperforms the state-of-the-art tracking methods while running at 54.83 FPS.
Jia Shao, Bo Du 0001, Chen Wu 0003, Pingkun Yan
ICME3
2019 Scene Change Detection VIA Deep Convolution Canonical Correlation Analysis Neural Network
abstract
Scene change detection is the process of identifying the differences between the multi-temporal image scenes at the semantic level, which has significant potential in the application of urban development and land management. In this paper, we propose a novel deep convolution canonical correlation analysis neural network (DCCANet) architecture, which could consider the spectral-spatial-temporal correlation feature for scene change detection in remote sensing images. For this purpose, we put together the convolutional neural networks (CNNs) and the deep canonical correlation analysis (DCCA) into the end-to-end network. The CNN could get the spectral-spatial feature information for scene representation, while the later could enhance the temporal correlation by nonlinear high- dimensional transformations between the multi-temporal image scenes for scene change detection. Experiments with high-resolution remote sensing image scene datasets demonstrated that our proposed approach can get a better performance in scene classification and change detection.
Bo Du 0001, Lixiang Ru, Chen Wu 0003, Hui Luo 0012
IGARSS4
2019 Unsupervised Deep Slow Feature Analysis for Change Detection in Multi-Temporal Remote Sensing Images
abstract
Change detection has been a hotspot in the remote sensing technology for a long time. With the increasing availability of multi-temporal remote sensing images, numerous change detection algorithms have been proposed. Among these methods, image transformation methods with feature extraction and mapping could effectively highlight the changed information and thus has a better change detection performance. However, the changes of multi-temporal images are usually complex, and the existing methods are not effective enough. In recent years, the deep network has shown its brilliant performance in many fields, including feature extraction and projection. Therefore, in this paper, based on the deep network and slow feature analysis (SFA) theory, we proposed a new change detection algorithm for multi-temporal remotes sensing images called deep SFA (DSFA). In the DSFA model, two symmetric deep networks are utilized for projecting the input data of bi-temporal imagery. Then, the SFA module is deployed to suppress the unchanged components and highlight the changed components of the transformed features. The change vector analysis pre-detection is employed to find unchanged pixels with high confidence as training samples. Finally, the change intensity is calculated with chi-square distance and the changes are determined by threshold algorithms. The experiments are performed on two real-world data sets and a public hyperspectral data set. The visual comparison and the quantitative evaluation have shown that DSFA could outperform the other state-of-the-art algorithms, including other SFA-based and deep learning methods.
Bo Du 0001, Lixiang Ru, Chen Wu 0003, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.3
2019 Tracking Objects From Satellite Videos: A Velocity Feature Based Correlation Filter
abstract
Satellite video target tracking is a new topic in the remote sensing field, which refers to tracking moving objects of interest from satellite video in real time. The target of interest usually occupies only a few pixels in a satellite video image, even when the train is long. Thus, satellite video target tracking still faces new challenges compared with traditional visual tracking, including the detection of low-resolution targets, features with less representation, and targets with an extremely similar background. Little research has been done on satellite video target tracking, and little is known about whether or not the existing tracking algorithms can still work on the satellite video data. This paper, for the first time, intensively investigated 13 typical trackers in traditional visual tracking. The experimental results suggest that most of the state-of-the-art tracking algorithms mainly rely on luminance, color features, or convolutional features, and they fail to track satellite video targets due to their inadequate representation features. To overcome this difficulty, we propose a velocity correlation filter (VCF) algorithm, which employs both a velocity feature and an inertia mechanism (IM) to construct a specific kernel correlation filter for the satellite video target tracking. The velocity feature has a high discriminative ability to detect moving targets in satellite videos, and the IM can prevent model drift adaptively. Experimental results on three real satellite video data sets show that the VCF outperforms state-of-the-art tracking methods with regard to precision and success plots while running at over 100 frames per second.
Jia Shao, Bo Du 0001, Chen Wu 0003, Lefei Zhang
IEEE Trans. Geosci. Remote. Sens.3
2019 Can We Track Targets From Space? A Hybrid Kernel Correlation Filter Tracker for Satellite Video
abstract
Despite the great success of correlation filter-based trackers in visual tracking, it is questionable whether they can still perform on the satellite video data, acquired by a satellite or space station very high above the earth. The difficulty lies in that the targets usually occupy only a few pixels compared with the image size of over one million pixels and almost melt into the similar background. Since correlation filter models strongly depend on the quality of features and the spatial layout of the tracked object, they would probably fail on satellite video tracking tasks. In this paper, we propose a hybrid kernel correlation filter (HKCF) tracker employing two complementary features adaptively in a ridge regression framework. One feature is the optical flow that can detect variation pixels of the target. The other one is the histogram of oriented gradient that can capture the contour and texture information in the target, and an adaptive fusion strategy is proposed to employ the strengths of both features in different satellite videos. Quantitative evaluations are performed on six real satellite video data sets. The results show that our approach outperforms state-of-the-art tracking methods while running at more than 100 frames/s.
Jia Shao, Bo Du 0001, Chen Wu 0003, Lefei Zhang
IEEE Trans. Geosci. Remote. Sens.3
2018 VCF: Velocity Correlation Filter, Towards Space-Borne Satellite Video Tracking
abstract
Tracking a moving target of interests from a space-borne satellite video is really a difficulty, since the target usually occupies only a few pixels in each frame of satellite video. Even it is a long train. Most state-of-the-art tracking algorithms mainly rely on luminance or color features, failing to handle this video tracking problem due to the extremely inadequate quality of targets features. To overcome this difficulty, we propose a velocity correlation filter (VCF), employing velocity feature and inertia mechanism to construct a kernel correlation filter for satellite video targets tracking. The velocity feature has a high discriminative ability to detect moving targets in satellite videos and the inertia mechanism can prevent model drift adaptively. Experimental results on three real satellite video datasets show our proposed approach outperforms state-of-the-art tracking methods with a more than 100 frames per second.
Jia Shao, Bo Du 0001, Chen Wu 0003, Jia Wu 0001, Ruimin Hu, Xuelong Li 0001
ICME3
2018 Object Tracking in Satellite Videos by Fusing the Kernel Correlation Filter and the Three-Frame-Difference Algorithm
abstract
Object tracking is a popular topic in the field of computer vision. The detailed spatial information provided by a very high resolution remote sensing sensor makes it possible to track targets of interest in satellite videos. In recent years, correlation filters have yielded promising results. However, in terms of dealing with object tracking in satellite videos, the kernel correlation filter (KCF) tracker achieves poor results due to the fact that the size of each target is too small compared with the entire image, and the target and the background are very similar. Therefore, in this letter, we propose a new object tracking method for satellite videos by fusing the KCF tracker and a three-frame-difference algorithm. A specific strategy is proposed herein for taking advantage of the KCF tracker and the three-frame-difference algorithm to build a strong tracker. We evaluate the proposed method in three satellite videos and show its superiority to other state-of-the-art tracking methods.
Bo Du 0001, Shihan Cai, Chen Wu 0003, Qian Du 0001
IEEE Geosci. Remote. Sens. Lett.4
2017 Real-time tracking based on weighted compressive tracking and a cognitive memory model
Bo Du 0001, Chen Wu 0003, Lefei Zhang, Liangpei Zhang 0001
Signal Process.3
2017 Kernel Slow Feature Analysis for Scene Change Detection
abstract
Scene change detection between multitemporal image scenes can be used to interpret the variation of regional land use, and has significant potential in the application of urban development monitoring at the semantic level. The traditional methods directly comparing the independent semantic classes neglect the temporal correlation, and thus suffer from accumulated classification errors. In this paper, we propose a novel scene change detection method via kernel slow feature analysis (KSFA) and postclassification fusion, which integrates independent scene classification with scene change detection to accurately determine scene changes and identify the “from-to” transition type. After representation with the bag-of-visual-words model, KSFA is proposed to extract the nonlinear temporally invariant features, to better measure the change probability between corresponding multitemporal image scenes. Two postclassification fusion methods, which are based on Bayesian theory and predefined rules, respectively, are then employed to identify the optimal coupled class combinations of multitemporal scene pairs. Furthermore, in addition to identifying semantic changes, the proposed method can also improve the performance of scene classification, since the unchanged scenes are more likely to belong to the same class. Two experiments with high-resolution remote sensing image scene data sets confirm that the proposed method can increase the accuracy of scene change detection, scene transition identification, and scene classification.
Chen Wu 0003, Liangpei Zhang 0001, Bo Du 0001
IEEE Trans. Geosci. Remote. Sens.1
2016 A scene change detection framework for multi-temporal very high resolution remote sensing images
Chen Wu 0003, Lefei Zhang, Liangpei Zhang 0001
Signal Process.1
2015 Hyperspectral anomaly change detection with slow feature analysis
Chen Wu 0003, Liangpei Zhang 0001, Bo Du 0001
Neurocomputing1
2014 Slow Feature Analysis for Change Detection in Multispectral Imagery
abstract
Change detection was one of the earliest and is also one of the most important applications of remote sensing technology. For multispectral images, an effective solution for the change detection problem is to exploit all the available spectral bands to detect the spectral changes. However, in practice, the temporal spectral variance makes it difficult to separate changes and nonchanges. In this paper, we propose a novel slow feature analysis (SFA) algorithm for change detection. Compared with changed pixels, the unchanged ones should be spectrally invariant and varying slowly across the multitemporal images. SFA extracts the most temporally invariant component from the multitemporal images to transform the data into a new feature space. In this feature space, the differences in the unchanged pixels are suppressed so that the changed pixels can be better separated. Three SFA change detection approaches, comprising unsupervised SFA, supervised SFA, and iterative SFA, are constructed. Experiments on two groups of real Enhanced Thematic Mapper data sets show that our proposed method performs better in detecting changes than the other state-of-the-art change detection methods.
Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.1
2014 Automatic Radiometric Normalization for Multitemporal Remote Sensing Imagery With Iterative Slow Feature Analysis
abstract
Multitemporal imagery analysis has attracted widespread interest in recent years due to the large number of applications. Multitemporal remote sensing imagery analysis is very important for Earth observation, in order to allow an understanding of the relationships and interactions between human and natural phenomena. Radiometric variance of the same targets due to differences in environmental conditions is one of the most important issues. In this paper, we propose an automatic radiometric normalization method with iterative slow feature analysis (ISFA) to reduce the radiometric variance. Slow feature analysis extracts invariant features from the quickly varying input signals. It is first reformulated for the multitemporal imagery problem and then improved to an iterative version. In the iteration, high weights are assigned to unchanged pixels. After convergence, the linear function of the radiometric normalization is directly obtained with all the pixels and their weights. If the ISFA is negatively affected by the changed pixels in some special cases and cannot find the correct regression line, initial seeds are selected as the initial weights in the iteration, to improve the performance, which is called S-ISFA. Two pairs of multitemporal ETM images from different seasons and years were used to test the effectiveness of our proposed method. The quantitative evaluation showed that our proposed method performs better, with smaller differences in the statistical distributions and radiometric values than the other state-of-the-art methods. The robustness with regard to the selection of initial seeds was also proved in the experiment.
Liangpei Zhang 0001, Chen Wu 0003, Bo Du 0001
IEEE Trans. Geosci. Remote. Sens.2
2013 An Automatic Relative Radiometric Correction Method Based on Slow Feature Analysis
abstract
Radiometric correction is very important for temporal remote sensing images analysis. The key of relative radiometric correction is to accurately select pseudo-invariant features (PIFs). This process should be automatic. Slow feature analysis is a new learning algorithm to extract invariant feature from input signals. It is appreciate to separate the unchanged pixels. We apply iteration process to assign high weights to unchanged pixels. After convergence, the linear function is calculated directly with all the pixels and their weights. The experiment demonstrates that our automatic relative radiometric correction method can get a good performance.
Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001
ICIG1