EDBT 2026 Demo / reviewers in the wild / expert
Bing Tu
dblp:124/3257
· DBLP profile ↗
66ranked-venue papers
26as first author
42since 2021 · last 2026
0000-0001-5802-9496ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 40 · 15 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ERCAD: An embedding replay method for continual anomaly detection and segmentation
Zhipeng Deng, Bing Tu, Junfeng Man |
Pattern Recognit. | 3 |
| 2026 | Spectral State Fusion Tree Mamba for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) data possess complex spatial structures and high-dimensional spectral information. Mamba has been applied to address the limitations of general methods in HSI classification, including restricted receptive fields and high computational complexity. However, the scan mechanism of traditional Mamba unreasonably constructs spatial distance relationships between neighboring row pixels and fails to adaptively construct the optimal scanning path based on the spectral similarity of pixels. Additionally, the characteristic of traditional Mamba scanning each channel independently overlooks the feature extraction from high-dimensional spectral information. This work proposes a Spectral State Fusion Tree Mamba (SSFTM) architecture to resolve these limitations. The Tree Scan (TS) mechanism computes cosine distances among spatial neighboring pixels and spectral channels to construct adaptive minimum spanning trees in both the spatial and spectral domains, thereby establishing reasonable spatial-spectral relationships and enabling efficient joint feature extraction. The Spectral State Fusion (SSF) mechanism applies multi-layer one-dimensional dilated convolutions along the spectral dimension to the state space vectors, enabling inter-channel interaction and promoting multi-scale spectral feature extraction. The proposed SSFTM demonstrates superior classification accuracy across multiple datasets compared to SOTA methods and exhibits acceptable computational complexity. The code is available at https://github.com/copawloroous/SSFTM. Bing Tu, Zhenghao Hu, Bo Liu 0020 |
IEEE Trans. Image Process. | 1 |
| 2025 | Multiscale Spectral-Morpho Fusion-Based Unsupervised Cross-Domain Learning for Hyperspectral ClassificationabstractHyperspectral image classification (HSIC) plays a key role in remote sensing, but the interpretation of scene features with limited samples remains challenging. In this letter, we propose a multiscale spectral morphological fusion network (MSMFNet) for unsupervised domain adaptation (UDA) in HSIC. This method consists of several key steps. First, the network uses principal component analysis (PCA) to extract spectral features, which are then combined with extended morphological profiles (EMPs) to enhance spatial structure representation. Then, a multiscale heterogeneous feature aggregation (MHFA) module is introduced to capture heterogeneous features across different scales and directions. Next, the multiscale global attention (MGA) module generates multilevel responses, exploring correlations between local and global information, thereby improving the representation of fine-grained features and boosting cross-domain feature fusion and classification performance. Finally, contrastive learning (CL) is applied to efficiently extract domain-invariant features. The experimental results demonstrate that MSMFNet achieves superior performance in cross-domain adaptation and fine-grained feature discrimination, achieving accuracies of 77.48% on the Houston dataset and 93.66% on the Pavia dataset, with Kappa coefficients that exceed the state of the art by 2.46 and 2.48, respectively. Yishu Peng, Bing Tu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | AF2DN: Attention-Guided Frequency Feature Decomposition Network for Hyperspectral and LiDAR Data ClassificationabstractTransformers have gained significant attention in multimodal remote sensing fusion due to their strong global context modeling capability. Although Transformer-based methods excel at processing high-dimensional spectral sequences and joint spatial-spectral information, most current research remains focused on the spatial domain. Consequently, the exploration of frequency-domain features—particularly implicit frequency representations—is often neglected. Moreover, efficiently fusing multimodal data features while emphasizing more discriminative information remains a challenging task. To address these challenges, this paper proposes an Attention-guided Frequency Feature Decomposition Network (AF2DN) for Hyperspectral and LiDAR Data Classification. First, a Transformer-based Frequency Feature Decomposition(TFFD) method is proposed, employing window attention to capture distinct directional frequency components from multimodal remote sensing data. Through this approach, low-frequency components are utilized to characterize global structural information, while various high-frequency components are employed to extract diverse texture and edge features. Second, an Attention Frequency Modulation(AFM) module is developed, incorporating a weight learning matrix in the frequency domain. This matrix is designed to selectively amplify and suppress different frequency components, thereby reducing data redundancy resulting from frequency feature decomposition. Finally, an adaptive Multimodal Same-Frequency Feature Fusion (AMSF3) module is designed to achieve cross-modal feature integration at identical frequency bands. Extensive experiments are conducted on three benchmark datasets, and the results demonstrate that the proposed framework outperforms existing state-of-the-art methods while exhibiting stronger adaptability in complex environments. Zhuoyu Chen, Bing Tu, Bo Liu 0020, Jun Li 0009, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Classification of Multisource Remote Sensing Data Using Slice MambaabstractThe multisource remote sensing (RS) data have yielded promising results in target detection and classification tasks. However, most existing methods primarily focus on the spatial features inherent in spectral information, while the continuous spectral characteristics are generally neglected. This oversight leads to insufficient extraction of spectral information, thereby limiting detection performance. Recently, the Mamba architecture, based on state space models (SSMs), integrates the advantages of long-range sequence modeling and linear computational efficiency, demonstrating significant potential in low-dimensional scenarios. Inspired by this, we propose Slice Mamba for multisource RS data fusion classification. Specifically, we design two scanning methods: lateral slice scanning (LatSS) and longitudinal slice scanning (LonSS), which construct sequences from lateral and longitudinal perspectives to facilitate information interaction between pixels. In conjunction with the Mamba architecture, we develop the lateral slice Mamba block (LatSMB) and the longitudinal slice Mamba block (LonSMB) to capture continuous spatial-spectral features. Based on this, we establish the slice feature extraction (SFE) module for extracting spatial-spectral feature information and design the cross-information fusion (CIF) module to form a complementary structure for effectively modeling spatial-spectral features, thereby achieving the fusion and classification of multisource heterogeneous features. Experimental results on three benchmark datasets demonstrate that Slice Mamba outperforms existing advanced methods in fusion classification performance and exhibits greater robustness when applied to multispectral datasets. Bing Tu, Puzhao Jiang, Bo Liu 0020, Jun Li 0009, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | HSI-MFormer: Integrating Mamba and Transformer Experts for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is fundamental to numerous remote sensing applications, enabling detailed analysis of material properties and environmental conditions. Recent Mamba built upon selective state space models (S6) have demonstrated exceptional advantages in long-range sequence modeling with linear computational efficiency, while Transformer based on self-attention mechanisms is particularly adept at capturing short-range dependencies. To leverage the complementary strengths of these models, this paper introduces a novel hybrid Mamba-Transformer framework (HSI-MFormer), effectively exploring the multiscale properties of hyperspectral data for HSI classification. Initially, a Multiscale Token Generation module (MTG) is developed, which converts the HSI cube into multiple spatial-spectral token groups across different scales. To adequately capture fine-grained multiscale spatial-spectral patterns, an Inner-scale Transformer Expert (ITE) is designed, which incorporates grouped self-attention operations to perform short-range sequence modeling within token groups at each scale. Meanwhile, a Cross-scale Mamba Expert (CME) is introduced, which integrates a cross-scale serialization mechanism and bidirectional Mamba block for long-range sequence modeling, further exploring the interactions and complementarity between token groups across different scales. Several hybrid strategies for integrating the ITE and CME are investigated to maximize their complementarity, including parallel, interval, and serial structures. Extensive experiments demonstrate that the propsed HSI-MFormer significantly out-performs the state-of-the-art Transformer-based and Mamba-based HSI classification methods. The code is available at https://github.com/tubingnuist/HSI-MFormer. Bing Tu, Bo Liu 0020, Jun Li 0009, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Self-Supervised Graph Masked Autoencoders for Hyperspectral Image ClassificationabstractTraditional supervised deep learning (DL) methods for hyperspectral image (HSI) classification are severely limited by the quality and quantity of labels. Furthermore, existing feature extraction methods generally lack the fusion of multiscale feature information, struggling to handle complex scenarios. To counter these problems, this work investigates a feature extraction module based on self-supervised graph masked autoencoders (SGMAEs). It innovatively employs graph masked autoencoders to achieve self-supervised label-free feature extraction for the complete set of samples, utilizing a multiscale graph convolutional network encoder (MGCNE) and cross correlation decoder (CCD) to extract and fuse multiscale spatial-spectral features of HSI data, respectively. Specifically, the HSI data is first converted into an edge-masked perturbed graph to label-freely extract multiscale feature representations of all pixel samples, and then fed into the MGCNE to obtain multilayer feature vectors for the pixel nodes. To reconstruct the masked edges for the fusion of multiscale features, the CCD applies cross correlation calculations to the nodes of the true edges at the masked positions and the fake edges at the random positions. The contrastive learning loss function is proposed for training of the autoencoder, which calculates the loss for the existence estimates of edges generated by cross correlation calculations. The pretrained MGCNE possesses an efficient self-supervised multiscale spatial-spectral feature extraction capability, along with strong generalizability, which significantly improves the accuracy of various mainstream models in downstream classification tasks. Extensive experiments and analyses on multiple HSI datasets demonstrate that our proposed SGMAE significantly enhances the model performance of various supervised classifiers and achieves superior performance in comparison to mainstream models. Zhenghao Hu, Bing Tu, Bo Liu 0020, Jun Li 0009, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Adaptive Feature Self-Attention in Spiking Neural Networks for Hyperspectral ClassificationabstractHyperspectral image (HSI) classification is crucial for remote sensing research, while its high-dimensional features make traditional algorithms difficult to cope with. Despite the breakthroughs in deep learning, the high computational complexity and energy consumption limit its application in resource-limited environments. Spiking neural networks (SNNs), mimicking the brain’s information processing with low power consumption, have emerged as a promising alternative for edge computing. However, SNNs struggle with complex tasks due to the nondifferentiability of spike signals, which complicates training and exhibits limitations in extracting deep features and modeling long-range dependencies. In this article, we propose a novel SNN framework that addresses these challenges by enhancing feature extraction and efficiently capturing dependencies in hyperspectral data. Our framework integrates an adaptive refocusing convolutional layer with a spike self-attention (SSA) mechanism. The adaptive refocusing convolutional layer employs learnable parameters to dynamically adjust the convolutional kernel’s response to input spike data, improving feature representation. The adaptive refocusing convolutional layer uses learnable parameters to dynamically adjust kernel responses to input spike data, enhancing feature representation. Experimental results show that this model achieves over 96% classification accuracy in a single time step, significantly surpassing current methods and effectively solving the problem of low accuracy at short time steps in SNNs. Additionally, this framework reduces computational energy consumption by approximately$12.5\times $compared to similar, offering new potential for edge intelligence applications. Bing Tu, Bo Liu 0020, Jun Li 0009, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Spatial-Frequency Domain Transformation for Infrared Small Target DetectionabstractWith the development of infrared technology, infrared small target detection (IRSTD) is widely applied in fields such as environmental monitoring, marine rescue, and forest fire prevention. Existing IRSTD methods are often based on spatial domain approaches, which preserve target features in the spatial domain but overlook the characteristics of infrared small targets in the frequency domain. In frequency domain methods, infrared small targets are typically considered as high-frequency components, while the continuous background is regarded as low-frequency components. However, infrared small targets often have complex backgrounds, strong edges, and noise generated during imaging, all of which are also reflected as high-frequency components, leading to false detections. To overcome this issue and fully explore the potential of IRSTD in the frequency domain, we propose a novel network, SFDTNet, which integrates frequency-domain attention and U-Structure for IRSTD. In the encoding phase, spatial feature extraction is applied to the infrared small target. In the decoding stage, global-scale spatial features are modeled in the frequency domain to achieve more precise reconstruction of small targets while reducing the interference of background high-frequency clutter. Frequency domain self-attention (FDSA) introduces an attention mechanism to model global information in the frequency domain and capture the importance of different frequency components. Adaptive frequency selection network (AFSN) incorporates learnable masks to adaptively modulate high- and low-frequency components in the frequency domain. Finally, a deep supervision strategy is employed to help the network learn features more effectively. Experimental results demonstrate that it effectively retains the shape and contours of small targets while achieving a very low false detection rate. Compared with existing state-of-the-art methods, our approach shows superior performance and better robustness. Bing Tu, Bo Liu 0020, Jun Li 0009, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Superpixel-Integrated Dual-Stage Mamba for Hyperspectral Image Classification
Qinghua Song, Bing Tu, Bo Liu 0020, Jun Li 0009, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | High-Resolution Aerosol Retrieval Algorithm via Convolutional Kolmogorov-Arnold NetworkabstractAccurately obtaining the optical/microphysical characteristics of aerosols from satellite data is important for environmental protection and air quality monitoring. In this study, using Sentinel-2 satellite data, we propose a high-resolution satellite aerosol retrieval algorithm (CKAN) via convolutional neural network (CNN) and Kolmogorov-Arnold network (KAN). Unlike traditional retrieval algorithms that require the construction of physical models, the CKAN algorithm relies entirely on deep learning. This algorithm focuses on extracting high-dimensional information from the data through CNN and learning the potential nonlinear relationships between the data through the powerful fitting ability of KAN. Compared with the existing algorithms, the CKAN algorithm is characterized by simplicity and accuracy, and does not require a large amount of auxiliary meteorological data (e.g., relative humidity and ground air pressure) enables retrieval of various aerosol parameters, including Aerosol Optical Depth (AOD), Fine-mode AOD (FAOD), Coarse-mode AOD (CAOD), and Single Scattering Albedo (SSA). To demonstrate the effectiveness of the algorithm, we retrieved aerosol optical/microphysical characteristics from Sentinel-2 imagery for four study areas. Results indicate that both the AOD and FAOD retrieved by the CKAN algorithm exhibit a high degree of correlation ( R > 0.90) with AERONET products. For CAOD and SSA, although a few poor retrievals resulted in a low overall correlation, CKAN’s retrievals show good agreement with AERONET products. The CKAN algorithm has good accuracy in high-resolution satellite aerosol retrievals, and is expected to be widely used in urban-scale aerosol monitoring. Bing Tu, Chengxin Hu, Bo Liu 0020, Jun Li 0009, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Multimodal Data Fusion Classification via Adaptive Frequency Domain Sparse Enhancement
Bing Tu, Bo Liu 0020, Jun Li 0009, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | From Weak Textures to Dense Arrangements: Leveraging Prior Knowledge for Small-Object Detection in Remote Sensing ImagesabstractDetecting small objects in remote sensing images is a significant challenge due to their weak texture, scale variations, and dense spatial arrangements. Existing approaches often overlook the importance of prior contextual information and the aggregation of features in densely packed small objects, both of which are crucial for improving the performance of remote sensing small object detection (RSSOD). In this work, we propose the Prior Guided Context Fusion Network (PGCFNet), which enhances small object detection by decoupling scene contextual information through three novel components: the Prior Guided Context Fusion Module (PGCFM), the DepthWise Aggregator (DWA), and the Prior Guided Small Object Detector (PGSOD). This architecture facilitates a deeper exploration of the relationships between small objects and their surrounding environment. Specifically, PGCFM improves feature representation by integrating multi-scale features and applying prior-guided dynamic channel weighting, addressing the challenge of weak textures. Additionally, DWA refines feature aggregation using dilated convolutions and dynamic feature adjustment, enabling precise multi-scale detection in environments with dense small objects. Furthermore, PGSOD leverages prior knowledge to reduce background interference, enhancing small object detection across varying scales and orientations. Collectively, these modules work synergistically to advance small object detection in remote sensing images, overcoming key challenges in complex environments. Extensive experiments on three public datasets demonstrate that the performance of the proposed method outperforms several state-of-the-art detectors, especially for tiny object detection. Specifically, PGCFNet achieves 86.0% mAP on the DIOR dataset, 95.59% mAP on the NWPU VHR-10 dataset, and 58.5% mAP on the AI-TOD dataset. Additionally, we conducted generalization experiments for PGCFM, DWA and PGSOD, demonstrating its effectiveness across different datasets and detection networks with varying model sizes. Wei He 0021, Guoyun Zhang, Jianhui Wu 0002, Bing Tu |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Hyperspectral Image Classification via Neighborhood Adaptive Graph Isomorphism NetworkabstractGraph convolutional network (GCN) has garnered significant attention in hyperspectral image (HSI) classification due to their ability to model non-Euclidean structured data. Compared with convolutional neural network (CNN), GCN can perform convolutions over irregular image regions and learn global dependencies among pixels in the whole image. Most existing GCN-based methods in the HSI community rely on average aggregation or weighted average aggregation strategies to aggregate neighboring node features. This process tends to obscure the differences between the nodes. However, for HSI classification tasks with obvious intra-class variability, average aggregation is a suboptimal choice. Moreover, the quality of the initial graph structure plays a crucial role in the model’s capacity to represent spectral relationships effectively. To mitigate these issues, we propose a neighborhood adaptive graph isomorphism network (NAGIN) for HSI classification to ensure that the diversified spectra representation of land-cover can be effectively captured. The neighborhood adaptive block (NAB) enhances spectral discriminability between land-cover classes via spectral reconstruction, enabling more precise removal of anomalous pixels in neighboring nodes. The graph isomorphism network (GIN) aggregates the features of neighboring nodes in an isomorphic manner to obtain multiple spectral expressions of the same type of land-cover, ensuring that the spectral features of different land-cover classes can be accurately distinguished. The Kolmogorov-Arnold network (KAN) leverages its ability to learn adaptive activation functions to better extract and refine the spectral features aggregated by GIN. Experimental results demonstrate that NAB can effectively improve the quality of the graph structure, the GIN aggregation method is competitive in HSI classification, and the proposed NAGIN outperforms the state-of-the-art methods on several public HSI datasets. Bing Tu, Bo Liu 0020, Jun Li 0009, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Self-Supervised Masked Graph Autoencoder for Hyperspectral Anomaly DetectionabstractHyperspectral image anomaly detection faces the challenge of difficulty in annotating anomalous targets. Autoencoder(AE)-based methods are widely used due to their excellent image reconstruction capability. However, traditional grid-based image representation methods struggle to capture long-range dependencies and model non-Euclidean structures. To address these issues, this paper proposes a self-supervised Masked Graph AutoEncoder (MGAE) for hyperspectral anomaly detection. MGAE utilizes a Graph Attention Network (GAT) autoencoder to reconstruct the background of hyperspectral images and identifies anomalies by comparing the reconstructed features with the original features. Specifically, we constructs a topological graph structure of the hyperspectral image, which is then input into the GAT autoencoder for reconstruction, leveraging the multi-head attention mechanism to learn spatial and spectral features. To prevent the decoder from learning trivial solutions, we introduce a re-masking strategy that randomly masks both the input features and hidden representations during training, forcing the model to learn and reconstruct features under limited information, thereby improving detection performance. Additionally, the proposed loss function with graph Laplacian regularization (Twice Loss) minimizes variations in feature representations, leading to more consistent background reconstruction. Experimental results on several real-world hyperspectral datasets demonstrate that MGAE outperforms existing methods. Bing Tu, Baoliang He, Bo Liu 0020, Jun Li 0009, Antonio Plaza |
IEEE Trans. Image Process. | 1 |
| 2025 | Multi-Scale Autoencoder Suppression Strategy for Hyperspectral Image Anomaly DetectionabstractAutoencoders (AEs) have received extensive attention in hyperspectral anomaly detection (HAD) due to their capability to separate the background from the anomaly based on the reconstruction error. However, the existing AE methods routinely fail to adequately exploit spatial information and may precisely reconstruct anomalies, thereby affecting the detection accuracy. To address these issues, this study proposes a novel Multi-scale Autoencoder Suppression Strategy (MASS). The underlying principle of MASS is to prioritize the reconstruction of background information over anomalies. In the encoding stage, the Local Feature Extractor, which integrates Convolution and Omni-Dimensional Dynamic Convolution (ODConv), is combined with the Global Feature Extractor based on Transformer to effectively extract multi-scale features. Furthermore, a Self-Attention Suppression module (SAS) is devised to diminish the influence of anomalous pixels, enabling the network to focus more intently on the precise reconstruction of the background. During the process of network learning, a mask derived from the test outcomes of each iteration is integrated into the loss function computation, encompassing only the positions with low anomaly scores from the preceding detection round. Experiments on eight datasets demonstrate that the proposed method is significantly superior to several traditional methods and deep learning methods in terms of performance. Bing Tu, Bo Liu 0020, Jun Li 0009, Antonio Plaza |
IEEE Trans. Image Process. | 1 |
| 2025 | Anomaly Detection in Hyperspectral Images Using Adaptive Graph Frequency LocationabstractGraph theory-based techniques have recently been adopted for anomaly detection in hyperspectral images (HSIs). However, these methods rely excessively on the relational structure within the constructed graphs and tend to downplay the importance of spectral features in the original HSI. To address this issue, we introduce graph frequency analysis to hyperspectral anomaly detection (HAD), which can serve as a natural tool for integrating graph structure and spectral features. We treat anomaly detection as a problem of graph frequency location, achieved by constructing a beta distribution-based graph wavelet space, where the optimal wavelet can be identified adaptively for anomaly detection. Initially, a high-dimensional, undirected, unweighted graph is built using the pixels in the HSI as vertices. By leveraging the observation of energy shifting to higher frequencies caused by anomalies, we can dynamically pinpoint the specific Beta wavelet associated with the anomalies' high-frequency content to accurately extract anomalies in the context of HSIs. Furthermore, we introduce a novel entropy definition to address the frequency location problem in an adaptive manner. Experimental results from seven real HSIs validate the remarkable detection performance of our newly proposed approach when compared to various state-of-the-art anomaly detection methods. Bing Tu, Xianchang Yang, Baoliang He, Yunyun Chen, Jun Li 0009, Antonio Plaza |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Infrared Small Target Detection via Nested U-Structure With Attention and Multiscale Feature PyramidabstractWith the advancement of infrared technology, infrared small target detection plays a crucial role in precise guidance and warning systems, attracting widespread attention from researchers. Due to their characteristics of low contrast and low signal-to-noise ratio, infrared small targets are easily affected by noise and background interference, making accurate detection challenging. To address the problem of infrared small target detection and shape-preserving segmentation in complex backgrounds, we propose a novel network named U2AMFP-Net. Specifically, we improve the residual U-attention-blocks (RUAB) structure, as the original residual U-blocks (RSU) structure tends to retain excessive invalid features of non-target regions when extracting intra-layer and inter-layer information at different scales. To overcome this issue, we employ attention mechanisms to focus the model more on small target features while suppressing irrelevant information. Additionally, we design the multiscale feature pyramid network (MFPN) on the network to avoid the problem of edge information loss caused by excessive skip connections, thereby further improving detection rates. Experimental evaluations demonstrate that our method preserves more complete shapes of weak infrared small targets while ensuring accurate detection. Compared with other state-of-the-art methods, our approach exhibits superior performance across various metrics. Furthermore, we construct a new training dataset containing 10 000 images of infrared small targets using existing datasets. Longyuan Guo, Bing Tu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | IGroupSS-Mamba: Interval Group Spatial-Spectral Mamba for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification has garnered substantial attention in remote sensing fields. Recent mamba architectures built upon the selective state-space models (S6) have demonstrated enormous potential in long-range sequence modeling. However, the high dimensionality of hyperspectral data and information redundancy pose challenges to the application of S6 in HSI classification, suffering from suboptimal performance and computational efficiency. In light of this, this article investigates a lightweight interval group spatial-spectral mamba framework (IGroupSS-Mamba) for HSI classification, which allows for multidirectional and multiscale global spatial-spectral information extraction in a grouping and hierarchical manner. Technically, an interval group S6 mechanism (IGSM) is developed as the core component, which partitions high-dimensional features into multiple nonoverlapping groups at intervals, and then integrates a unidirectional S6 for each group with a specific scanning direction to achieve nonredundant sequence modeling. Compared with conventional applying multidirectional scanning to all bands, this grouping strategy leverages the complementary strengths of different scanning directions while decreasing computational costs. To adequately capture the spatial-spectral contextual information, an interval group spatial-spectral block (IGSSB) is introduced, in which two IGSM-based spatial and spectral operators are cascaded to characterize the global spatial-spectral relationship along the spatial and spectral dimensions, respectively. IGroupSS-Mamba is constructed as a hierarchical structure stacked by multiple IGSSB blocks, integrating a pixel aggregation-based downsampling strategy for multiscale spatial-spectral semantic learning from shallow to deep stages. Extensive experiments demonstrate that IGroupSS-Mamba significantly outperforms the state-of-the-art methods in classification accuracy and achieves lower model parameters and floating point operations (FLOPs). The code is available athttps://github.com/IIP-Team/IGroupSS-Mamba. Bing Tu, Puzhao Jiang, Bo Liu 0020, Jun Li 0009, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Hybrid Multiscale Spatial-Spectral Transformer for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification constitutes a significant foundation for remote sensing analysis. Transformer architecture establishes long-range dependencies with a self-attention mechanism (SA), which exhibits advantages in HSI classification. However, most existing transformer-based methods are inadequate in exploring the multiscale properties of hybrid spatial and spectral information inherent in HSI data. To countermeasure this problem, this work investigates a hybrid multiscale spatial–spectral framework (HMSSF). It innovatively models global dependencies across multiple scales from both spatial and spectral domains, which allows for cooperatively capturing hybrid multiscale spatial and spectral characteristics for HSI classification. Technically, a spatial–spectral token generation (SSTG) module is first designed to generate the spatial tokens and spectral tokens. Then, a multiscale SA (MSSA) is developed to achieve multiscale attention modeling by constructing different dimensional attention heads per attention layer. This mechanism is adaptively integrated into both spatial and spectral branches for hybrid multiscale feature extraction. Furthermore, a spatial–spectral attention aggregation (SSAA) module is introduced to dynamically fuse the multiscale spatial and spectral features to enhance the classification robustness. Experimental results and analysis demonstrate that the proposed method outperforms the state-of-the-art methods on several public HSI datasets. Bing Tu, Bo Liu 0020, Yunyun Chen, Jun Li 0009, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | 3DSS-Mamba: 3D-Spectral-Spatial Mamba for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification constitutes the fundamental research in remote sensing fields. Convolutional neural networks (CNNs) and Transformers have demonstrated impressive capability in capturing spectral-spatial contextual dependencies. However, these architectures suffer from limited receptive fields and quadratic computational complexity, respectively. Fortunately, recent Mamba architectures built upon the state space models (SSMs) integrate the advantages of long-range sequence modeling and linear computational efficiency, exhibiting substantial potential in low-dimensional scenarios. Motivated by this, we propose a novel 3D-spectral-spatial mamba (3DSS-Mamba) framework for HSI classification, allowing for global spectral-spatial relationship modeling with greater computational efficiency. Technically, a spectral-spatial token generation (SSTG) module is designed to convert the HSI cube into a set of 3-D spectral-spatial tokens. To overcome the limitations of traditional Mamba, which is confined to modeling causal sequences and inadaptable to high-dimensional scenarios, a 3D-spectral-spatial selective scanning (3DSS) mechanism is introduced, which performs pixel-wise selective scanning on 3-D hyperspectral tokens along the spectral and spatial dimensions. Five scanning routes are constructed to investigate the impact of dimension prioritization. The 3DSS scanning mechanism combined with conventional mapping operations forms the 3D-spectral-spatial mamba block (3DMB), enabling the extraction of global spectral-spatial semantic representations. Experimental results and analysis demonstrate that the proposed method outperforms the state-of-the-art methods on HSI classification benchmarks. The code is available athttps://github.com/IIP-Team/3DSS-Mamba. Bing Tu, Bo Liu 0020, Jun Li 0009, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Hyperspectral Image Classification via Multiscale Multiangle Attention NetworkabstractHyperspectral images (HSIs) provide a large amount of spatial and spectral information to characterize ground objects. However, they also contain a lot of redundant information, which makes it difficult to extract complex local and global spatial-spectral features. Considering that HSIs present multi-scale similarity and anisotropic image features, multi-scale and multi-angle information can be used to effectively model local and global features and reduce the complexity of self-attention. This paper proposes a new multi-scale multi-angle attention network (MMAN) for HSI classification that models the internal relationship between image features at local and global scales. Firstly, three spectral-spatial feature extraction modules (at different scales) are constructed to extract the low-level features of the image. These modules are first used by a 3D convolutional layer for spectral feature extraction, and then input to a 2D convolutional layer for spatial feature extraction. Next, the serialized tokens are input to the multi-angle attention module. Finally, the learnable labels are identified through a linear layer, and the features of different scales are fused through a fully connected layer to realize the classification of samples. Experimental results on four standard datasets show that the proposed exhibits comparable or superior classification performance than other state-of-the-art methods. Jianghong Hu, Bing Tu, Qi Ren, Xiaolong Liao, Zhaolou Cao, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Cross-Resolution Perceptual Knowledge Propagation for Low-Resolution Aerial Photograph CategorizationabstractAccurately recognizing the semantic categories of LR (low-resolution) aerial photos is an indispensable technique in remote sensing. In practice, however, this task is nontrivial due to: 1) the difficulty to encode human visual perception in the recognition process, 2) the intolerable human resources to label sufficient training LR aerial photos, and 3) the challenge to select high quality features for categorization. To handle these problems, a novel cross-resolution perceptual knowledge propagation (CPKP) is proposed, focusing on leveraging the visual perceptual experiences deeply learned from HR (high-resolution) aerial photos to enhance categorizing LR ones. Specifically, by mimicking human vision system, a novel low-rank model is proposed to decompose each LR aerial photo into multiple visually/semantically salient foreground regions coupled with the background non-salient ones. This model can 1) produce the gaze shifting path (GSP) simulating human gaze shifting sequence, and 2) calculate the deep feature for each GSP. Afterward, a kernel-induced feature selection (FS) algorithm is formulated to obtain a succinct set of deep GSP features discriminative across LR and HR aerial photos. Based on these, the labels from LR and HR aerial photos are collaboratively utilized to train a linear classifier for categorizing LR ones. Noticeably, our CPKP framework can effectively optimize the linear classifier training. This is because labels of HR aerial photos can be acquired more conveniently than LR ones practically. Comprehensive visualization results and comparative study have validated the superiority of our method. Bing Tu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Retargeting HR Aerial Photos Under Contaminated Labels With Application in Smart NavigationabstractRetargeting aims to shrink a photo wherein the perceptually prominent regions are appropriately kept. In practice, optimally shrinking a high resolution (HR) aerial photo is a useful tool for smart navigation. Nowadays, vehicle drivers’ path planning is generally guided by an HR aerial photo recommended by a navigation App like Google Maps. Owing to the limited and various resolution of vehicle displays, we have to retarget each original HR aerial photo accordingly, wherein the navigation-aware regions can be well preserved. In practice, HR aerial photo retargeting is non-trivial due to three challenges: 1) the rich number of internal objects and their complex spatial layouts, 2) deriving the region-level semantics from potentially contaminated image labels, and 3) the inefficiency of retargeting each HR aerial photo with millions of pixels. To handle these problems, we propose a novel HR aerial photo retargeting pipeline that can intelligently avoid the negative effects from incorrect image labels. The key is a noise-tolerant hashing algorithm that converts image-level semantics into the hash codes corresponding to different regions, which guides the HR aerial photo shrinking. More specifically, for each HR aerial photo, we extract visually/semantically salient object patches inside it. To explicitly encode their spatial layout, we construct a graphlet by linking the spatially adjacent object patches into a small graph. Subsequently, a binary matrix factorization (MF) is designed to exploit the underlying semantics of these graphlets, wherein three attributes: i) binary hash codes learning, ii) noisy labels refinement, iii) deep image-level semantics, are collaboratively encoded. Such binary MF can be solved iteratively and each graphlet is subsequently converted into the binary hash codes. Finally, the hash codes corresponding to graphlets within each HR aerial photo are utilized to learn a Gaussian mixture model (GMM) that optimizes the HR aerial photo retargeting. During the experimental validation, we compiled a smart navigation dataset including 132743 planned paths annotated from 10132 HR aerial photos, based on which comparative study has demonstrated the superiority of our method. Bing Tu, Yingjie Xia |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Hyperspectral Anomaly Detection Using Reconstruction Fusion of Quaternion Frequency Domain AnalysisabstractMost existing techniques consider hyperspectral anomaly detection (HAD) as background modeling and anomaly search problems in the spatial domain. In this article, we model the background in the frequency domain and treat anomaly detection as a frequency-domain analysis problem. We illustrate that spikes in the amplitude spectrum correspond to the background, and a Gaussian low-pass filter performing on the amplitude spectrum is equivalent to an anomaly detector. The initial anomaly detection map is obtained by the reconstruction with the filtered amplitude and the raw phase spectrum. To further suppress the nonanomaly high-frequency detailed information, we illustrate that the phase spectrum is critical information to perceive the spatial saliency of anomalies. The saliency-aware map obtained by phase-only reconstruction (POR) is used to enhance the initial anomaly map, which realizes a significant improvement in background suppression. In addition to the standard Fourier transform (FT), we adopt the quaternion FT (QFT) for conducting multiscale and multifeature processing in a parallel way, to obtain the frequency domain representation of the hyperspectral images (HSIs). This helps with robust detection performance. Experimental results on four real HSIs validate the remarkable detection performance and excellent time efficiency of our proposed approach when compared to some state-of-the-art anomaly detection methods. Bing Tu, Xianchang Yang, Wei He 0021, Jun Li 0009, Antonio Plaza |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Graph Evolution-Based Vertex Extraction for Hyperspectral Anomaly DetectionabstractAnomaly detection is a fundamental task in hyperspectral image (HSI) processing. However, most existing methods rely on pixel feature vectors and overlook the relational structure information between pixels, limiting the detection performance. In this article, we propose a novel approach to hyperspectral anomaly detection that characterizes the HSI data using a vertex- and edge-weighted graph with the pixels as vertices. The constructed graph encodes rich structural information in an affinity matrix. A crucial innovation of our method is the ability to obtain internal relations between pixels at multiple topological scales by processing different powers of the affinity matrix. This power processing is viewed as a graph evolution, which enables anomaly detection using vertex extraction formulated as a quadratic programming problem on graphs of varying topological scales. We also design a hierarchical guided filtering architecture to fuse multiscale detection results derived from graph evolution, which significantly reduces the false alarm rate. Our approach effectively characterizes the topological properties of HSIs, leveraging the structural information between pixels to improve anomaly detection accuracy. Experimental results on four real HSIs demonstrate the superior detection performance of our proposed approach compared to some state-of-the-art hyperspectral anomaly detection methods. Xianchang Yang, Bing Tu, Qianming Li, Jun Li 0009, Antonio Plaza |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Background subtraction via regional multi-feature-frequency model in complex scenes
Ping Lei, Wei He 0021, Guoyun Zhang, Jianhui Wu 0002, Bing Tu |
Soft Comput. | 7 |
| 2023 | Class-wise Graph Embedding-Based Active Learning for Hyperspectral Image ClassificationabstractDeep learning (DL) techniques have shown remarkable progress in remotely sensed hyperspectral image (HSI) classification tasks. The performance of DL-based models highly relies on the quality and quantity of labeled data. However, manual labeling is a laborious and expensive process that requires substantial efforts from human experts. Active learning (AL) techniques have been developed to alleviate the burden of manual annotation by selecting the most informative and uncertain samples for labeling. In this paper, we propose a new class-wise graph embedding-based AL (CGE-AL) framework implemented by a class-wise graph convolutional network (CGCN). First, we train a classifier with labeled data and infer latent features from labeled and unlabeled samples with the trained parameter. Then, we group the labeled data into multiple one-label sets by category. In a class-wise manner, we initialize the nodes of the graph with one-label and unlabeled features, which are then fed into CGCN. By updating the graph parameters with binary loss, CGCNs measure the uncertainty between labeled nodes and unlabeled nodes. To select the most valuable sample for labeling, we adopt the class minimum uncertainty to query the unlabeled nodes with higher overall uncertainty. We repeat this process with the updated labeled set to retrain our classification model and CGCNs. Extensive experiments demonstrate the outstanding performance of our method compared to other state-of-the-art AL-based approaches. Xiaolong Liao, Bing Tu, Jun Li 0009, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | CBW-MSSANet: A CNN Framework With Compact Band Weighting and Multiscale Spatial Attention for Hyperspectral Image Change DetectionabstractChange detection (CD), aims to detect the changing area of the same scene at different times, which is an important application of remote sensing images. As the key data source of CD, hyperspectral image (HSI) is widely used in CD technology because of its rich spectral-spatial information. However, how to mine the multi-level spatial information of dual-temporal hyperspectral images (HSIs) and focus on the features of the pixels to be classified individually remains a problem in the spatial attention mechanism (SAM). To make full use of the spectral-spatial information of HSIs, in this paper we propose a CNN framework with compact band weighting and multi-scale spatial attention (CBW-MSSANet) for HSI pixel-level CD. The main contributions of this article are as follows: 1) a new method of pseudo-label training sample selection based on k-means (KM) centroid distance is designed; 2) apply the compact band weighting (CBW) module to HSI CD to take full advantage of the spectral information of HSIs; 3) a multi-scale spatial attention (MSSA) module is developed for pixel-level CD, which can mine multi-level spatial information and pay more attention to the features of the pixels to be classified, and combine the spatial information of adjacent pixels to make it more conducive to pixel-level CD. Experimental results on four real HSI datasets demonstrated that the performance of MSSA surpasses the classical single-scale SAM, and CBW-MSSANet is superior to some representative CD methods. Xianfeng Ou, Liangzhen Liu, Bing Tu, Linbo Qing, Guoyun Zhang, Zifei Liang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | A New Context-Aware Framework for Defending Against Adversarial Attacks in Hyperspectral Image ClassificationabstractDeep neural networks play a significant role in hyperspectral image (HSI) processing, yet they can be easily fooled when trained with adversarial samples (generated by adding tiny perturbations to clean samples). These perturbations are invisible to the human eye, but can easily lead to misclassification by the deep learning model. Recent research on defense against adversarial samples in HSI classification has improved the robustness of deep networks by exploiting global contextual information. However, available methods do not distinguish between different classes of contextual information, which makes the global context unreliable and increases the success rate of attacks. To solve this problem, we propose a robust context-aware network able to defend against adversarial samples in HSI classification. The proposed model generates a global contextual representation by aggregating the features learned via dilated convolution, and then explicitly models intraclass and interclass contextual information by constructing a class context-aware learning module (including affinity loss) to further refine the global context. The module helps pixels obtain more reliable long-range dependencies and improves the overall robustness of the model against adversarial attacks. Experiments on several benchmark HSI datasets demonstrate that the proposed method is more robust and exhibits better generalization than other advanced techniques. Bing Tu, Wangquan He, Qianming Li, Yishu Peng, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | UHD Aerial Photograph Categorization by Leveraging Deep Multiattribute Matrix FactorizationabstractThere are thousands of observation satellites orbiting the earth, each of which captures massive-scale photographs covering millions of square kilometers everyday. In practice, these aerial photos are with ultra-high-definitions (UHD) and may contain tens to hundreds of ground objects (e.g., vehicles and rooftops). Understanding the multiple categories of a rich variety of UHD aerial photos is an indispensable technique for many applications, such as intelligent transportation, natural disaster prediction, and smart agriculture. In this work, we propose a novel multi-label UHD aerial photo categorization pipeline, wherein the key is to topologically represent the spatial layouts of the ground objects and further deeply encode them using a deep multi-clue matrix factorization (DMCMF) that robustly handles noisy labels at image-level. More specifically, for each UHD aerial photo, we extract visually/semantically salient object patches inside it. To explicitly encode their spatial layout, we construct a graphlet by linking the spatially adjacent object patches into a small graph. Subsequently, a binary MF is designed to intelligently exploit the semantics of these graphlets, wherein four clues: i) binary hash codes learning, ii) noisy labels refinement, iii) deep image-level semantics, and iv) adaptive data graph updating are incorporated. Such DMCMF can be solved iteratively and each graphlet is then converted into the discrete hash codes. Finally, the hash codes corresponding to graphlets within each UHD aerial photo are quantized into a feature vector by a kernel machine for multi-label categorization. Toward a comprehensive comparative study, we complied a million-scale UHD aerial photo set collected from 100 top-ranking cities worldwide. Experiments have shown that 1) our method is highly competitive in learning categorization model from imperfect labels at image-level, and 2) the four clues are elaborately designed and seamlessly combined to learn hash codes for representing UHD aerial photos. Yinfu Feng, Bing Tu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Multi-Objective Unsupervised Band Selection Method for Hyperspectral Images ClassificationabstractWith the increasing spectral dimension of hyperspectral images (HSI), how correctly choose bands based on band correlation and information has become more significant, but also complicated. Band selection is a combinatorial optimization problem, and intelligent optimization algorithms have been shown to be crucial in solving combinatorial optimization problems. However, major of them only use a single objective as the selection index, while neglecting the overall features of hyperspectral images, which may lead to inaccuracy in object detection. To tackle this, we propose a band selection method based on a multi-objective cuckoo search algorithm (MOCS) when constructing a multi-objective unsupervised band selection model based on the amount of information and correlation of the bands (MOCS-BS). Specifically, an adaptive strategy based on population crowding degree is first proposed to assist Lévy flight in overcoming the influence of the parameter constancy. Then, an information-sharing strategy based on grouping and crossover is designed to balance the search ability between global exploration and local exploitation, which can overcome the shortcomings caused by the lack of information interaction between individuals. Finally, the HSI classification experiments are performed by Random Forest and KNN classifiers based on the subset of bands selected by the proposed MOCS-BS method. The proposed method is compared with state-of-the-art algorithms including neighborhood grouping normalized matched filter (NGNMF) and multi-objective artificial bee colony with band selection (MABC-BS) on four HSI datasets. The experimental results demonstrate that MOCS-BS is more effective and robust than other methods. Xianfeng Ou, Meng Wu 0007, Bing Tu, Guoyun Zhang, Wujing Li |
IEEE Trans. Image Process. | 3 |
| 2022 | SIM-MFR: Spatial interactions mechanisms based multi-feature representation for background modeling
Wei He 0021, Jiexin Li, Bing Tu, Xianfeng Ou, Longyuan Guo |
J. Vis. Commun. Image Represent. | 4 |
| 2022 | Spatial Peak-Aware Collaborative Representation for Hyperspectral Imagery ClassificationabstractIn this letter, a novel spatial peak-aware collaborative representation (SPaCR) method is proposed for hyperspectral imagery (HSI) classification, which introduces spectral–spatial information among superpixel clusters into regularization terms to construct a new collaborative representation (CR)-based closed-form solution. The proposed method is composed of the following key steps. First, the raw HSI is clustered into many superpixels according to an oversegmentation strategy. Then, cluster pixels are determined based on spectral–spatial correlation between pixels within each superpixel. Next, spectral distance and spatial coherence of superpixel clusters corresponding to training samples and testing pixels are fused to define differences between pixels. Finally, the difference information between clusters as a spectral–spatial feature-induced regularization term is incorporated into the objective function. Experimental results on the Indian Pines and the University of Pavia HSIs indicated that the proposed SPaCR method, without any preprocessing and postprocessing, outperforms well-known and state-of-the-art classifiers on the limited labeled samples. Chengle Zhou, Bing Tu, Qi Ren |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | A CNN Framework With Slow-Fast Band Selection and Feature Fusion Grouping for Hyperspectral Image Change DetectionabstractChange detection approaches can detect changed areas of the same scene at different times. Hyperspectral remote-sensing images contain large amounts of spectral information at high resolution. As hyperspectral datasets become abundant, more and more change detection technologies use hyperspectral images as raw data. Hyperspectral images suffer from band redundancy. There is an urgent need to improve the directionality of change of information features. To solve these problems, in this article, we propose a CNN framework involving slow-fast band selection (SFBS) and feature fusion grouping (SFBS-FFGNET) for hyperspectral image change detection. The main contributions of this article are as follows: 1) based on slow feature analysis (SFA), an SFBS method is proposed, which selects slow and fast feature bands to better extract changed and unchanged features, to more effectively separate changed and unchanged pixels; 2) we used a difference matrix to enrich the level of change information to provide more change characteristics for change detection; and 3) an FFG method was used to generate a more discriminative feature group, and the related loss function was designed. Experimental results on multiple real hyperspectral datasets showed that SFBS can reduce the operating load on the computer and improve the accuracy of change detection, and FFG can also improve the accuracy of change detection. In summary, SFBS-FFGNET is superior to most existing change detection methods. Xianfeng Ou, Liangzhen Liu, Bing Tu, Guoyun Zhang, Zhi Xu 0005 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Spatial-Spectral Transformer With Cross-Attention for Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) have been widely used in hyperspectral image (HSI) classification tasks because of their excellent local spatial feature extraction capabilities. However, because it is difficult to establish dependencies between long sequences of data for CNNs, there are limitations in the process of processing hyperspectral spectral sequence features. To overcome these limitations, inspired by the Transformer model, a spatial–spectral transformer with cross-attention (CASST) method is proposed. Overall, the method consists of a dual-branch structures, i.e., spatial and spectral sequence branches. The former is used to capture fine-grained spatial information of HSI, and the latter is adopted to extract the spectral features and establish interdependencies between spectral sequences. Specifically, to enhance the consistency among features and relieve computational burden, we design a spatial–spectral cross-attention module with weighted sharing to extract the interactive spatial–spectral fusion feature intra Transformer block, while also developing a spatial–spectral weighted sharing mechanism to capture the robust semantic feature inter Transformer block. Performance evaluation experiments are conducted on three hyperspectral classification datasets, demonstrating that the CASST method achieves better accuracy than the state-of-the-art Transformer classification models and mainstream classification networks. Yishu Peng, Bing Tu, Qianming Li, Wujing Li |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Local Semantic Feature Aggregation-Based Transformer for Hyperspectral Image ClassificationabstractHyperspectral images (HSIs) contain abundant information in the spatial and spectral domains, allowing for a precise characterization of categories of materials. Convolutional neural networks (CNNs) have achieved great success in HSI classification, owing to their excellent ability in local contextual modeling. However, CNNs suffer from fixed filter weights and deep convolutional layers, which lead to a limited receptive field and high computational burden. The recent Vision Transformer (ViT) models long-range dependencies with a self-attention mechanism and has been an alternative backbone to the CNNs traditionally used in HSI classification. However, such transformer-based architectures designate all input pixels of the receptive field as feature tokens in terms of feature embedding and self-attention, which inevitably limits the ability for learning multi-scale features and increases the computational cost. To overcome this issue, we propose a local semantic feature aggregation-based transformer (LSFAT) architecture which allows transformers to represent long-range dependencies of multi-scale features more efficiently. We introduce the concept of the homogeneous region into the transformer by considering a pixel aggregation strategy and further propose neighborhood aggregation-based embedding (NAE) and attention (NAA) modules, which are able to adaptively form multi-scale features and capture locally spatial semantics among them in a hierarchical transformer architecture. A reusable classification token is included together with the feature tokens in the attention calculation. In the last stage, a fully connected layer is employed to perform classification on the reusable token after transformer encoding. We verify the effectiveness of the NAE and NAA modules compared with the traditional ViT through extensive experiments. Our results demonstrate the excellent classification performance of the proposed method in comparison with other state-of-the-art approaches on several public HSIs. Bing Tu, Xiaolong Liao, Qianming Li, Yishu Peng, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Hyperspectral Anomaly Detection Using the Spectral-Spatial GraphabstractAnomaly detection is an important technique for hyperspectral image processing. It aims to find pixels that are markedly different from the background when the target spectrum is unavailable. Many anomaly detection methods have been proposed over the past years, among which graph-based ones have attracted extensive attention. And they usually just consider the spectral information to build the adjacency matrix of the graph, which does not think over the effect of spatial information in this process. This paper proposes a new anomaly detection method using the Spectral-Spatial Graph (SSG) that considers both the spatial and spectral information. Thus, the spatial adjacency matrix and spectral adjacency matrix are constructed from the spatial and spectral dimensions, respectively. To obtain a spectral-spatial graph with more discriminant characteristics, and two different local neighborhood detection strategies are used to measure the similarity of the SSG. Furthermore, global anomaly detection results on hyperspectral images were obtained by the graph Laplacian anomaly detection method and the global and local anomaly detection results were optimized by the differential fusion method. Compared with other anomaly detection algorithms on several synthetic and real data sets, the proposed algorithm shows superior detection performance. Bing Tu, Zhi Wang 0022, Huiting Ouyang, Xianchang Yang, Jun Li 0009, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Ensemble Entropy Metric for Hyperspectral Anomaly DetectionabstractIn hyperspectral anomaly detection, anomalies are rare targets that exhibit distinct spectral signatures from the background. Thus, anomalies are with low probabilities of occurrence in hyperspectral images. In this article, we develop a new technique for hyperspectral anomaly detection that adopts a new information theory perspective, to fully utilize the aforementioned concepts. Our goal is to transform system entropy into quantitative metrics of anomaly conspicuousness of pixels. To do so, two tasks are first completed: first, the construction of occurrence probability of pixels based on the density peak clustering algorithm, and second, the valid system definitions for pixels in specific anomaly detection problems with multiviews. Specifically, three types of systems are separately established by pixel pairs to conform to the definitions of three entropy definitions in information theory, i.e., Shannon entropy, joint entropy, and relative entropy. Then, three individual entropy-based metrics that assess the anomaly conspicuousness are defined. In addition, we design a standard deviation-based ensemble strategy for the integrated representation of the three individual metrics, which considers both logic “OR” and “AND” operations to simultaneously improve the detection rate and reduce the false alarm rate. Our experimental results obtained on two publicly available datasets with anomalies of different sizes and shapes demonstrate the superiority of our newly proposed anomaly detection method. Bing Tu, Xianchang Yang, Xianfeng Ou, Guoyun Zhang, Jun Li 0009, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Detection of moving objects using adaptive multi-feature histograms
Wei He 0021, Wujing Li, Guoyun Zhang, Bing Tu, Yong Kwan Kim, Jianhui Wu 0002 |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | Spectral-Spatial Hyperspectral Classification via Structural-Kernel Collaborative RepresentationabstractThis letter introduces a novel spatial-spectral classification method for hyperspectral images (HSIs) based on a structural-kernel collaborative representation (SKCR), which considers one weak assumption of spatial neighborhood that of the pixels in a superpixel belong to the same class when exploiting contextual information in HSI. The proposed method consists of the following steps. First, a superpixel segmentation strategy is used to construct self-adaptive regions for the HSI. Then, the structural information within each superpixel block is extracted based on the density peak and K nearest neighbors. Next, dual kernels are separately utilized for the exploitation of the spectral and the spatial information. Finally, the dual kernels are combined and incorporated into a support-vector-machine classifier. Since the weak assumption of spatial neighborhood is well considered in the collaborative representation, the proposed method showed excellent classification performance for two widely used real hyperspectral data sets even when the number of training samples was relatively small. Bing Tu, Chengle Zhou, Xiaolong Liao, Guoyun Zhang, Yishu Peng |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2021 | Feature Extraction via 3-D Block Characteristics Sharing for Hyperspectral Image ClassificationabstractSpectral–spatial information plays an essential role in hyperspectral image (HSI) classification compared to pure spectral information. However, the neighbor spectral–spatial information of a pixel tends to be mixed into other ground coverings due to various external factors such as the weather and sensor jitter, and mainstream HSI classification methods present low sensitivity for spatial information in this situation. This article proposes a novel feature extraction method via 3-D block characteristics sharing (3-D-BCS) for HSI classification that redefines spatial–spectral information of a local region based on a superpixel perspective to overcome the spectral–spatial weak assumptions in feature extraction that consists of the following steps. First, 3-D blocks are obtained by performing an oversegmentation method on the raw HSI. Then, instead of global operation, a 3-D block-based Gabor filter is applied to the principal components of an HSI to extract the textural features. Next, an average operation is conducted on each shape adaptive region to address the spatial weak assumption and Gaussian weight is introduced into each superpixel block to overcome the spectral weak assumption. Thus, 3-D characteristics sharing blocks can be constructed by reshaping the above three kinds of spectral–spatial feature. Finally, the majority-based support vector machine (SVM) classifier is utilized to determine the final class labels of HSI at the decision fusion level. Experiments performed on several real hyperspectral data sets with limited training samples show that the proposed 3-D-BCS method outperforms the other types of the classification method. Bing Tu, Chengle Zhou, Xiaolong Liao, Qianming Li, Yishu Peng |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Dual-Stage Construction of Probability for Hyperspectral Image ClassificationabstractRecently, feature extraction-based methods have received increasing attention in the hyperspectral image. In this letter, to ensure a more powerful discriminative ability of extracted features, a dual-stage construction of probability (DSCP) method is proposed for hyperspectral image classification. Specifically, the extended multi-attribute profiles (EMAP) method is applied to extract the shape feature of hyperspectral remote sensing image (HSI) to obtain a more accurate initial probability map. Considering that there are still some noises in the boundaries of the initial probability map, an effective edge-preserving filter-based approach named rolling guidance filter is used for probability post-optimization. Consequently, the class label of each pixel can be determined according to the optimized probability maps. Experiments demonstrate significantly the efficiency of the proposed method in comparison with other advanced methods. Bing Tu, Guangzhe Zhao, Guoyun Zhang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2020 | Density Peak Covariance Matrix for Feature Extraction of Hyperspectral ImageabstractThe clustering methods have a good application in many aspects, in which the density peak (DP) clustering can effectively cluster similar neighboring pixels so that the features can be extracted well for hyperspectral images (HSIs) classification. In this work, a DP based covariance matrix (DPCM) method is proposed for the feature extraction of HSIs, which not only can effectively extract features but also can reduce the within-class variations and the between-class interference. The proposed method consists of the following steps: First, maximum noise fraction is employed on the original HSI to reduce the computational complexity and eliminate noise. Second, the local densities of the sample are calculated by the DP clustering. Therefore, a reconstructed image can be obtained in which each pixel has a density feature vector. Then, the covariance matrix between each density pixel in the density map is calculated. Last, the extracted covariance matrices are fed back to the support vector machine based on the logarithm Euclidean kernel for label assignment. Experiments are conducted on the Indian pine data set, in which each of the five randomly selected marker data are selected as the training sample. The experimental results show that the method can effectively improve the classification accuracy and is superior to other classification methods. Guangzhe Zhao, Nanying Li, Bing Tu, Guoyun Zhang, Wei He 0021 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | Hyperspectral anomaly detection via density peak clustering
Bing Tu, Xianchang Yang, Nanying Li, Chengle Zhou, Danbing He |
Pattern Recognit. Lett. | 1 |
| 2020 | Hyperspectral Anomaly Detection Using Dual Window DensityabstractHyperspectral anomaly detection is one of the most active topics in hyperspectral image (HSI) analysis. The fine spectral information of HSIs allows us to uncover anomalies with very high accuracy. Recently, an intrinsic image decomposition (IID) model has been introduced for low-rank IID (LRIID) in multispectral images. Inspired by the LRIID, which is able to effectively recover the reflectance and shading components of the multispectral image, this article adapts the LRIID for obtaining the reflectance component of HSIs (which is the key feature for the discrimination of different objects). In order to exploit the reflectance component, we also propose a new dual window density (DWD)-based detector for anomaly detection, which is based on the idea that anomalies are usually rare pixels and, thus, exhibit low density in the image. The density analysis of DWD is intended not only to circumvent the Gaussian assumption regarding the distribution of HSI data, but also to mitigate the contamination of background statistics caused by anomalies. The dual window operation of our DWD is specifically designed to adaptively calculate the density of each pixel under test, so as to identify anomalies with nonspecific sizes. Our experimental results, obtained on a database of real HSIs including Airport, Beach, and Urban scenes, demonstrate the superiority of the proposed method in terms of detection performance when compared to other widely used anomaly detection methods. Bing Tu, Xianchang Yang, Chengle Zhou, Danbing He, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Hyperspectral Classification With Noisy Label Detection via Superpixel-to-Pixel Weighting DistanceabstractClassification is an important technique for remotely sensed hyperspectral image (HSI) exploitation. Often, the presence of wrong (noisy) labels presents a drawback for accurate supervised classification. In this article, we introduce a new framework for noisy label detection that combines a superpixel-to-pixel weighting distance (SPWD) and density peak clustering. The proposed method is able to accurately detect and remove noisy labels in the training set before HSI classification. It considers two weak assumptions when exploiting the spectral-spatial information contained in the HSI: 1) all the pixels in a superpixel belong to the same class and 2) close pixels in spectral space have the same label. The proposed method consists of the following steps. First, a superpixel segmentation step is used to obtain self-adaptive spatial information for each training sample. Then, a metric is utilized to measure the spectral distance information between each superpixel and pixel. Meanwhile, in order to overcome the first weak assumption, we use K nearest neighbors to obtain the closest neighborhoods of pixels around each superpixel, and a Gaussian weight is employed to mitigate the second weak assumption by adapting the original distance information. Next, the noisy labels in the original training set are removed by a density threshold-based decision function. Finally, the support vector machine (SVM) classifier is employed to evaluate the effectiveness of the proposed SPWD detection method in terms of classification accuracy. Experiments performed on several real HSI data sets demonstrate that the method can effectively improve the performance of classifiers trained with noisy training sets in terms of classification accuracy. Bing Tu, Chengle Zhou, Danbing He, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Density Peak Based Covariance Matrix for Hyperspectral Images ClassificationabstractThe clustering methods have a good application in many aspects, in which the density peak (DP) clustering can effectively cluster similar neighboring pixels, so that the features can be extracted well for hyperspectral images (HSIs) classification. In this work, a density peak based covariance matrix (DPCM) method is proposed for HSIs classification, which not only can effectively extracts features, but also can reduce the within-class variations and the between-class interference. The proposed method consists of the following steps: first, maximum noise fraction (MNF) is employed on the original HSI to reduce the computational complexity and eliminate noise. Second, the local densities of sample is calculated by the DP clustering. Therefore, the density map can be obtained in which each pixel has a density value in the original image. Then, the covariance matrix between each density pixel in the density map is calculated. Last, the extracted covariance matrices are fed back to the support vector machine (SVM) based on the logarithm Euclidean kernel for label assignment. Experiments on the Indian pine data set show that this method is superior to other classification methods. Bing Tu, Nanying Li, Wenlan Kuang, Chengle Zhou |
IGARSS | 1 |
| 2019 | Multiple convolutional layers fusion framework for hyperspectral image classification
Guangzhe Zhao, Guangyun Liu, Leyuan Fang, Bing Tu, Pedram Ghamisi |
Neurocomputing | 4 |
| 2019 | Deep feature representation for anti-fraud system
Bing Tu, Danbing He, Yongheng Shang, Chengle Zhou, Wujing Li |
J. Vis. Commun. Image Represent. | 1 |
| 2019 | A multi-view object tracking using triplet model
Bing Tu, Wenlan Kuang, Yongheng Shang, Danbing He |
J. Vis. Commun. Image Represent. | 1 |
| 2019 | Spatiotemporal local compact binary pattern for background subtraction in complex scenes
Wei He 0021, Hak-Lim Ko, Yong Kwan Kim, Jianhui Wu 0002, Guoyun Zhang, Bing Tu, Xianfeng Ou |
Multim. Tools Appl. | 7 |
| 2019 | Study of multiple moving targets' detection in fisheye video based on the moving blob model
Jianhui Wu 0002, Wenjing Hu, Wei He 0021, Bing Tu, Longyuan Guo, Xianfeng Ou, Guoyun Zhang |
Multim. Tools Appl. | 5 |
| 2019 | Hyperspectral image classification with a class-dependent spatial-spectral mixed metric
Bing Tu, Nanying Li, Leyuan Fang, Xianchang Yang, Jianhui Wu 0002 |
Pattern Recognit. Lett. | 1 |
| 2019 | Spatial Density Peak Clustering for Hyperspectral Image Classification With Noisy LabelsabstractThe “noisy label” problem is one of the major challenges in hyperspectral image (HSI) classification. In order to address this problem, a spatial density peak (SDP) clustering-based method is proposed to detect mislabeled samples in the training set. Specifically, the proposed methods consist of the following steps: first, the correlation coefficients among the training samples in each class are estimated. In this step, instead of measuring the correlation coefficients by considering individual samples, all neighbor samples or K representative neighbor samples in a local window surrounding each training sample are considered. By this way, the spatial contextual information could be used, and two versions of the proposed method, i.e., measuring the correlation coefficients using all neighbor samples or K representative samples, are referred as SDP and K-SDP, respectively. Second, with the correlation coefficients calculated above, the local density of each training sample can be obtained by the DP clustering algorithm. Finally, those mislabeled samples which usually have lower local densities in each class are able to be identified by a defined decision function. The effectiveness of the proposed detection method is evaluated using a series of spectral and spectral-spatial classification methods on several real hyperspectral data sets. Bing Tu, Xudong Kang, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Density Peak-Based Noisy Label Detection for Hyperspectral Image ClassificationabstractMislabeled training samples may have a negative effect on the performance of hyperspectral image classification. In order to solve this problem, a new density peak (DP) clustering-based noisy label detection method is proposed, which consists of the following steps. First, the distances among the training samples of each class are calculated using four representative distance metrics, i.e., the Euclidean distance (ED), orthogonal projection divergence (OPD), spectral information divergence (SID), and correlation coefficient (CC). Then, the local density of each training sample can be obtained using the DP clustering algorithm. Finally, a local density-based decision function is used to detect the noisy labels. The effectiveness of the proposed method is evaluated using the support vector machines on several real hyperspectral data sets. Experimental results demonstrate that the proposed noisy label detection method indeed helps in improving the classification performance. Bing Tu, Xudong Kang, Guoyun Zhang, Shutao Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Local Compact Binary Patterns for Background Subtraction in Complex ScenesabstractBackground modeling in complex scenes is a challenging problem. In this paper, a novel background subtraction method is proposed to address it. First, the textures are modeled with local compact binary patterns (LCBP), which have excellent robustness, strong discriminative power, and fast computation speed. To make LCBP more effective to appearance changes in complex scenarios, spatiotemporal local compact binary patterns (STLCBP) are then considered in which spatial texture information and temporal motion information are combined together. Multiple color spaces are also presented to separate foreground pixels more accurately from the background. To our knowledge, this is the first time that LCBP have been used for background modeling. Extensive experimental results on a widely used dataset clearly show that the proposed method outperforms other state-of-the-art methods and works effectively in complex scenes. Wei He 0021, Yongkwan Kim, Jianhui Wu 0002, Guoyun Zhang, Longyuan Guo, Bing Tu |
ICPR | 7 |
| 2018 | Spatial-spectral classification of hyperspectral image via group tensor decomposition
Guangzhe Zhao, Bing Tu, Hongyan Fei, Nanying Li, Xianchang Yang |
Neurocomputing | 2 |
| 2018 | Sub-Pixel Level Defect Detection Based on Notch Filter and Image RegistrationabstractGeneral machine vision algorithms are difficult to detect LCD sub-pixel level defects. By studying the LCD screen images, we found that the pixels in the LCD screen are regularly arranged. The spectrum distribution of LCD images, which is obtained by the Fourier transform, is relatively consistent. According to this feature, a method of sub-pixel defect detection based on notch filter and image registration is proposed. First, we take a defect-free template image to establish registration template and notch-filtering template; then we take the defect images for image registration with registration template, and solve the offset problem. After the notch-filter template filtering the background texture, the defect is more obvious; Finally the defects are obtained by the threshold segmentation method. The experiment results show that the proposed method can detect sub-pixel defects accurately and quickly. Longyuan Guo, Shinan Li, Wenjing Hu, Jianhui Wu 0002, Bing Tu, Wei He 0021, Xianfeng Ou, Guoyun Zhang |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2018 | An overview of face-related technologies
Hongyan Fei, Bing Tu, Ququ Chen, Danbing He, Chengle Zhou, Yishu Peng |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Classification of hyperspectral images via weighted spatial correlation representation
Bing Tu, Nanying Li, Leyuan Fang, Hongyan Fei, Danbing He |
J. Vis. Commun. Image Represent. | 1 |
| 2018 | Adaptive total variation-based spectral-spatial feature extraction of hyperspectral image
Guoyun Zhang, Hongyan Fei, Bing Tu |
J. Vis. Commun. Image Represent. | 5 |
| 2018 | Hyperspectral Imagery Noisy Label Detection by Spectral Angle Local Outlier FactorabstractThis letter presents the hyperspectral imagery (HSI) noisy label detection using a spectral angle and the local outlier factor (SALOF) algorithm. The noisy label is caused by a mislabeled training pixel, and thus, noisy training samples mixed with correct and incorrect labels are formed in the supervised classification. The LOF algorithm is first used in the noisy label detection of the HSI to improve the supervised classification accuracy. The proposed method SALOF mainly includes the following steps. First, k nearest neighbors of different training samples of each class are calculated based on the spectral angle mapper. Second, the reachability distance and local reachability density of all training samples are obtained. Third, the LOF is determined among different classes of training samples. Then, a segmentation threshold of the LOF is established to achieve an abnormal probability of these training samples. Finally, the support vector machines are applied to measure the detection efficiency of the proposed method. The experiments performed on the Kennedy Space Center data set demonstrate that the proposed method can effectively detect noisy labels. Bing Tu, Chengle Zhou, Wenlan Kuang, Longyuan Guo, Xianfeng Ou |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2018 | Hyperspectral Image Classification via Fusing Correlation Coefficient and Joint Sparse RepresentationabstractThe joint sparse representation (JSR)-based classifier assumes that pixels in a local window can be jointly and sparsely represented by a dictionary constructed by the training samples. The class label of each pixel can be decided according to the representation residual. However, once the local window of each pixel includes pixels from different classes, the performance of the JSR classifier may be seriously decreased. Since correlation coefficient (CC) is able to measure the spectral similarity among different pixels efficiently, this letter proposes a new classification method via fusing CC and JSR, which attempts to use the within-class similarity between training and test samples while decreasing the between-class interference. First, the CCs among the training and test samples are calculated. Then, the JSR-based classifier is used to obtain the representation residuals of different pixels. Finally, a regularization parameter λ is introduced to achieve the balance between the JSR and the CC. Experimental results obtained on the Indian Pines data set demonstrate the competitive performance of the proposed approach with respect to other widely used classifiers. Bing Tu, Xudong Kang, Guoyun Zhang, Jianhui Wu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2018 | Hyperspectral Image Classification via Superpixel Spectral Metrics RepresentationabstractThis letter proposes a new hyperspectral classification method that fuses superpixel spectral metrics and joint sparse representation (JSR), which is termed as superpixel spectral metrics representation (SSMR). Recently, superpixel segmentation has proven to be a powerful tool to exploit the spatial information of hyperspectral images (HSIs), since the size and shape of each superpixel can be adaptively changed in different structural textures. Moreover, spectral information divergence (SID) has superiority compared to other distance-based similarity measures, particularly when using with a JSR classifier. Taking the aforemementioned advantages into account, superpixel segmentation, SID, and JSR are availably combined to effectively utilize the spectralspatial information of the HSI. The proposed SSMR method includes the following main steps. First, superpixel segmentation is utilized to divide the original map into several superpixels. Second, similarity metric SID among test samples in all superpixels and training samples are calculated. Next, the JSR model is employed to obtain the reconstruction residuals of each class. Then, a regularization parameter λ is introduced to attain balance between JSR and SID. Finally, pixel's label is determined by the minimal total residual. Experimental results on the Indian Pines dataset show better performance than several well-known classification methods. Bing Tu, Wenlan Kuang, Guangzhe Zhao, Hongyan Fei |
IEEE Signal Process. Lett. | 1 |
| 2017 | Temporal-Spatial Symmetric Distributed Multi-View Video Coding SchemeabstractTo improve the rate stability and make a balance for different viewpoints in distributed multi-view video coding (DMVC) system, a novel symmetric DMVC (SDMVC) scheme is proposed in this paper. In the proposed scheme, every frame from all views adopts the same encoding mode and stable output rates are achieved, which are significant to improve the transmission efficiency in the channel. Both temporal and spatial correlations are exploited, in addition, a novel side information (SI) generation algorithm aiming at better exploring the correlations of proposed scheme has been proposed to obtain better performance. The simulation results show that the proposed SDMVC scheme gets a much more stable rate than the asymmetric scheme, only with neglectable bit-rate increasing. Meanwhile, the proposed SI generation algorithm significantly improves the coding performance. Guoyun Zhang, Canqun Xiang, Xianfeng Ou, Hong Yue, Longyuan Guo, Jianhui Wu 0002, Bing Tu, Wei He 0021 |
Int. J. Pattern Recognit. Artif. Intell. | 7 |