EDBT 2026 Demo / reviewers in the wild / expert
Yanfeng Gu
dblp:48/6204
· DBLP profile ↗
158ranked-venue papers
36as first author
80since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 134 · 25 first-author · 72 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 3 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pillar-voxel fusion network for 3D object detection in airborne hyperspectral point clouds
Yanze Jiang, Yanfeng Gu, Xian Li 0001 |
Sci. China Inf. Sci. | 2 |
| 2026 | TV Subgradient-Guided Multi-Source Fusion for Spectral Imaging in Dual-Camera CASSI SystemsabstractBalancing spectral, spatial, and temporal resolutions is a key challenge in spectral imaging. The Dual-Camera Coded Aperture Snapshot Spectral Imaging (DC-CASSI) system alleviates this trade-off but suffers from severely ill-posed reconstruction problems due to its high compression ratio. Existing methods are constrained by scene-specific tuning or excessive reliance on paired training data. To address these issues, we propose a Total Variation (TV) subgradient-guided multi-source fusion framework for DC-CASSI reconstruction, comprising three core components: (1) An end-to-end Single-Disperser CASSI (SD-CASSI) observation model based on the tensor-form Kronecker δ, which establishes a rigorous mathematical foundation for physical constraints while enabling efficient adjoint operator implementation; (2) An adaptive spatial reference generator that integrates SD-CASSI’s physical model and RGB subspace constraint, generating the reference image as reliable spatial prior; (3) A TV subgradient-guided regularization term that encodes local structural directions from the reference image into spectral reconstruction, achieving high-quality fused results. The framework is validated on simulated datasets and real-world datasets. Experimental results demonstrate that it achieves state-of-the-art reconstruction performance and robust noise resilience. This work not only establishes an interpretable theoretical foundation for subgradient-guided fusion but also provides a practical fusion-based paradigm for high-fidelity spectral image reconstruction in DC-CASSI systems. Source code: https://github.com/bestwishes43/ADMM-TVDS. Weiqiang Zhao, Tianzhu Liu, Yuzhe Gui, Wei Bian 0001, Yanfeng Gu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | An enhanced classification method based on adaptive multi-scale fusion for long-tailed multispectral point clouds
Tianzhu Liu, Bangyan Hu, Yanfeng Gu, Xian Li 0001, Aleksandra Pizurica |
Sci. China Inf. Sci. | 3 |
| 2025 | Hypergraph Contrastive Learning for Large-Scale Hyperspectral Image ClusteringabstractLarge-scale hyperspectral image (HSI) clustering has become an important research task owing to its promising applications in various fields. Recently, beneficial from the correlation modeling capability of graphs, graph contrastive learning methods have received increasing attention in the clustering task. However, these methods usually have limited ability to explore the high-order correlation as well as beneficial clustering information of large-scale HSI, thus limiting the clustering performance on large-scale HSI. To this end, a novel hypergraph contrastive learning network (HCL-Net) for large-scale HSI clustering is proposed in this paper. Specifically, a diffusion hypergraph-based contrastive clustering mechanism is presented, in which a diffusion hypergraph is constructed to model the high-order correlation in large-scale HSI, thus guiding contrastive learning for obtaining more discriminative representations. Besides, by mining the confident clustering information, a confidence-guided positive-negative updating strategy is designed to dynamically update positives and negatives for contrastive learning, thereby obtaining a more compact clustering structure. The proposed method is evaluated on three public large-scale HSI datasets. The experimental results have demonstrated the superior performance of the proposed HCL-Net over state-of-the-art methods. Bo Peng 0007, Tianyi Qin, Yanfeng Gu, Nam Ling, Jianjun Lei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Unsupervised Occluded Target Detection Based on Spherical Shell With Multispectral Point CloudsabstractMultispectral Point Clouds (MPCs) acquired from Unmanned Aerial Vehicles (UAVs), leveraging LiDAR’s canopy-penetrating capacity, provide distinct advantages for detecting occluded targets beneath vegetation canopy. However, limited samples, missing target spatial morphology, and unstructured data format have led to low occluded detection accuracy. Given these constraints, an unsupervised stereo detection method for occluded targets based on the spherical shell model with MPCs from UAVs has been proposed for the first time. The method exploits the spatial-spectral differences between targets and the background, treating targets as anomalies within the background, and enables occluded target detection without requiring training. The spherical shell model has been constructed in MPCs to avoid contaminating the global background with targets, leveraging its local separation characteristics for unsupervised detection of occluded targets. An adaptive radius and multi-scale spatial feature extraction have been designed to enhance the method’s robustness. The collaborative representation model has been utilized to achieve unsupervised detection by suppressing expressible points in local background features. To bridge the gap between algorithmic metrics and practical requirements, we have further proposed a target-level evaluation method that overcomes the limitations of conventional methods, which are susceptible to density-induced false alarm distortions in unstructured MPCs. Experiments on three real-world MPCs and a public airborne hyperspectral and LiDAR dataset show that our method achieves higher detection accuracy than existing spectral-based detectors and performs better on exposed targets. Likun Chen, Yanfeng Gu, Xian Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Digital Surface Model-Embedded Intrinsic Hyperspectral UnmixingabstractHigh-precision spectral unmixing (SU) of hyperspectral image (HSI) faces a challenging problem in that it is difficult to distinguish the different objective classes presenting similar spectra without elevation information. A digital surface model (DSM) with the same spatial resolution as HSI can provide additional geometric information that can be useful for the HSI SU task. Existing SU methods that incorporate the DSM data miss the consideration that DSM data and hyperspectral data are actually different dimensions of information from different sensors observing the same scene. The intrinsic hyperspectral unmixing model can obtain the shading component and the reflectance component, which can be further decomposed into the endmember and the abundance. From a physical modeling perspective, the shadow component can be considered as the interaction of illumination and the geometric component from the DSM data; thus, the quality of hyperspectral unmixing under invariant illumination conditions can be significantly enhanced. Furthermore, the elevation information within the DSM data contributes to the unmixing process by distinguishing different objective classes with similar physiochemical properties at varying altitudes. Experimental validation is conducted using three HSI datasets. The results can indicate the robustness and superiority of the proposed unmixing method. Yanyuan Huang, Tianzhu Liu, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | TG-ADet: Terrain-Guided Network for 3-D Object Detection in ALS Point CloudsabstractAirborne Laser Scanning (ALS) offers significant potential for three-dimensional (3D) object detection due to its ability to penetrate the canopy and acquire high-precision 3D spatial information. However, complex terrain distribution and backgrounds similar to objects hinder effective object detection in airborne scenes. To address these challenges, we propose TG-ADet, the first 3D object detection network explicitly designed for ALS point clouds. Our approach introduces three key components and integrates them into a unified framework. A multi-stage terrain guidance module predicts the terrain distribution and guides multiple detection stages based on prediction results, focusing on objects under various terrain conditions. A sparse feature enhancement module that aggregates voxel features and leverages auxiliary tasks to improve the backbone’s feature representation and suppress background interference. Additionally, an integrated data augmentation method generates training samples that align with ALS data distributions during network training, while increasing terrain complexity during testing. Experiments on two ALS point cloud datasets demonstrate that TG-ADet significantly outperforms state-of-the-art methods and achieves robust detection performance in challenging scenarios. Yanze Jiang, Xian Li 0001, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | A Sub-Pixel Coupled Dictionary Learning Method for Large-Scale and High-Spatial-Resolution Satellite Hyperspectral Image Reconstruction
Tianzhu Liu, Zitong Liu, Xianhao Zhang, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Real-Time Vehicle Detection in Satellite Videos: Transitioning From Large Scenes to ClustersabstractDetecting vehicles from satellite videos presents several significant challenges: 1) Satellite video frames typically possess extremely high resolutions, often containing millions or even hundreds of millions of pixels, whereas onboard computational resources remain constrained; 2) The non-uniform spatial distribution of vehicles results in inefficient allocation of computational resources; 3) Vehicles are typically small with limited distinguishing features, further complicating the detection task. In this article, we propose LSCNet (Large Scenes to Clusters), an efficient and lightweight satellite video vehicle detection network specifically designed to address these challenges. To alleviate the difficulties associated with large-scale imagery and uneven vehicle distributions, we introduce a plug-and-play Object Cluster Module (OCM). The OCM leverages inter-frame information from satellite video to adaptively identify and prioritize clustered regions, thereby enhancing detection precision. Furthermore, to improve the extraction of discriminative features from small-sized vehicles, we propose a lightweight Multi-frame Feature Aggregation Module (MFAG), which effectively captures the spatiotemporal characteristics of vehicles while maintaining computationally efficiency. Additionally, we refine the regression loss function by integrating the Kullback-Leibler Divergence (KLD), enabling the generation of higher-quality bounding boxes and significantly boosting the detection performance for small objects. Experimental evaluations on the Jilin-1 satellite video dataset demonstrate that the proposed method achieves improved detection accuracy while maintaining real-time performance, thereby validating its robustness and practical effectiveness. Jialei Pan, Yanfeng Gu, Guoming Gao, Shaochuan Wu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Masking Graph Cross-Convolution Network for Multispectral Point Cloud ClassificationabstractAchieving accurate 3-D environment perception is a key task in the field of remote sensing. Multispectral point cloud has rich integrated 3-D spatial–spectral information, which provides a data basis for realizing more detailed scene understanding and perception. However, the diversity of land covers and the complexity of its features pose challenges to classification. In addition, the current methods mechanically pool and fuse local features to obtain global information, which has limited the utility for multispectral point cloud classification. In this article, we propose a masking graph cross-convolution network (MGC2N), which aims to address these problems by utilizing spectral features to construct point-to-point relationships independent of spatial distance. A self-attention masking (SAM) module and a spatial–spectral cross-convolution (S2C2) module are innovatively designed into the proposed MGC2N. The former is used to adaptively adjust the nodes and edges of the adjacency matrix to dynamically extract effective features for different land covers; the latter is used to extract spatial distribution features and local spectral features of the land covers in the scene to enhance the discriminative ability of the learned features. Our method achieves the best-in-class performance on two real multispectral point cloud datasets, demonstrating its effectiveness in improving classification accuracy and robustness. Qingwang Wang, Xueqian Chen, Yuanqin Meng, Tao Shen 0004, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Satellite Video Event of Interest Detection Using Deep Spatiotemporal MetricabstractThe rapid advancements in video satellite technology have enabled dynamic Earth observation. However, the unique characteristics of satellite video data, such as extensive spatial coverage, sparse motion information, and significant foreground–background imbalance, introduce numerous challenges for practical applications. To address these difficulties, the task of event of interest (EOI) detection has emerged as a critical solution by extracting meaningful regions containing dynamic information from satellite videos. In this article, a novel deep learning framework for EOI detection, integrating spatiotemporal slice analysis and metric learning to tackle these challenges is proposed. The framework employs spatiotemporal slice analysis to convert 3-D motion information into 2-D motion trajectories, preserving essential dynamic information while minimizing background redundancy. A ResNet-based network is adopted for robust feature extraction, and a custom Earth mover’s distance (EMD) metric is introduced to enhance detection precision, enabling accurate differentiation between EOIs and non-EOIs in satellite video data. Experimental results demonstrate the superior performance of the proposed method. Guoming Gao, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | LPRnet: A Self-Supervised Registration Network for LiDAR and Photogrammetric Point CloudsabstractLiDAR and photogrammetry are active and passive remote sensing techniques for point cloud acquisition, respectively, offering complementary advantages and heterogeneous. Due to the fundamental differences in sensing mechanisms, spatial distributions, and coordinate systems, their point clouds exhibit significant discrepancies in density, precision, noise, and overlap. Coupled with the lack of ground truth for large-scale scenes, integrating the heterogeneous point clouds is a highly challenging task. This article proposes a self-supervised registration network based on a masked autoencoder, focusing on heterogeneous LiDAR and photogrammetric point clouds. At its core, the method introduces a multiscale masked training strategy to extract robust features from heterogeneous point clouds under self-supervision. To further enhance registration performance, a rotation-translation embedding module is designed to effectively capture the key features essential for accurate rigid transformations. Building upon robust representations, a transformer-based architecture seamlessly integrates local and global features, fostering precise alignment across diverse point cloud datasets. The proposed method demonstrates strong feature extraction capabilities for both LiDAR and photogrammetric point clouds, addressing the challenges of acquiring ground truth at the scene level. Experiments conducted on two real-world datasets validate the effectiveness of the proposed method in solving heterogeneous point cloud registration problems. Chen Wang 0060, Yanfeng Gu, Xian Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Community Structure Guided Network for Hyperspectral Image ClassificationabstractRecently, the hypergraph convolutional network (HGCN) has attracted increasing attention in hyperspectral image (HSI) classification. Compared to graph convolutional networks, HGCN has a stronger ability to mine nonlinear high-order correlations. However, the problems of intraclass variability and interclass similarity exist due to the effects of light, environment, and sensor bias, resulting in insufficient reliability of hypergraphs constructed by directly utilizing the original spectral features. Motivated by the observation that the land cover in HSI contains the spatial distribution semantic information of community structures, which can be used to extract deeper contextual semantic features, we propose a novel community structure guided network (CSGNet) for HSI classification. Specifically, CSGNet adopts a dual-branch architecture: the HGCN branch focuses on superpixel-level high-order feature extraction, while the convolutional neural network (CNN) branch enhances pixel-level local features. In HGCN branch, a novel reliable hypergraph construction approach is introduced, which strikes a balance between depth-first search (DFS) and breadth-first search (BFS), effectively representing different community structure features and improving the ability of edge detection. Meanwhile, kernel function mapping is used to achieve more accurate node connections and enhances classification within classes. Finally, to achieve balanced training of the HGCN and CNN branches, we add their cross-entropy loss as an auxiliary component in the backpropagation process. Experimental results demonstrate that CSGNet outperforms the state-of-the-art methods. The code will be released athttps://github.com/KustTeamWQW/CSGNet. Qingwang Wang, Jiangbo Huang, Shunyuan Wang, Zhen Zhang 0035, Tao Shen 0004, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | S4DR-Net: Self-Supervised Spatial-Spectral Distance Reconstruction Network for Multispectral Point Cloud ClassificationabstractMultispectral LiDAR point clouds are valuable in remote sensing for their spatial-spectral consistency, yet their high acquisition and annotation costs pose significant challenges. To mitigate this, self-supervised learning has emerged as a promising solution, reducing reliance on annotated data while improving model generalization. However, existing self-supervised frameworks for point clouds often overlook the complexity of ground object distribution in large-scale remote sensing scenarios and fail to leverage the spectral information inherent in multispectral point clouds. In this paper, we introduce the Self-Supervised Spatial-Spectral Distance Reconstruction Network (S4DR-Net), a novel self-supervised pre-training network designed for multispectral point cloud classification. Serving as the key component of the network, the Spatial-Spectral Distance Prediction module (S-SDP) effectively addresses these limitations by reconstructing the distance relationships between voxel blocks in three-dimensional Euclidean as well as spectral spaces. By jointly considering spatial and spectral distances, S-SDP enables the network to learn a unified representation that captures the intrinsic spatial-spectral consistency of multispectral point clouds. This design allows S4DR-Net to generate low-dimensional feature representations in a self-supervised manner, without reliance on manual annotations. We conducted experiments and evaluated on two real-world multispectral point cloud datasets. The results demonstrate that S4DR-Net consistently outperforms existing self-supervised pre-training methods, achieving superior accuracy and generalization compared with current state-of-the-art approaches. The code will be released at https://github.com/KustTeamWQW/S4DR-Net. Qingwang Wang, Jianling Kuang, Tao Shen 0004, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | PESAT: A Parameter-Efficient Spatiotemporal Adapter Tuning Framework for Satellite Video Scene ClassificationabstractSatellite video scene classification (SVSC) is a critical task for dynamic earth observation. However, it remains challenging due to the distinct spatiotemporal characteristics of satellite videos and the scarcity of annotated data, both of which differ significantly from general video datasets. Although Vision Transformers (ViTs) have shown strong performance in general video classification, their direct application to SVSC often results in suboptimal performance and overfitting. To address these challenges, we propose PESAT, a novel parameter-efficient spatiotemporal adapter tuning framework specifically tailored for SVSC tasks. PESAT enables the effective adaptation of pre-trained ViTs for SVSC by keeping the backbone model largely frozen and fine-tuning only a small number of strategically inserted adapter modules. Our framework incorporates three key innovations: an efficient temporal attention modeling (TAM) mechanism that reuses pre-trained self-attention weights for temporal feature extraction without adding new parameters; a sensitivity-guided adapter insertion strategy that identifies optimal locations within the ViT to place adapters, maximizing their impact; and a hybrid gated adapter (HGA) module, which combines depthwise convolution and a dynamic gating mechanism to capture complex spatiotemporal contexts specific to satellite video data. Experimental results demonstrate the superior performance of the proposed method. Guoming Gao, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Heterogeneous Open-Set Cross-Domain Manifold Embedding Aligned for HSI-MSI Collaborative ClassificationabstractHyperspectral images (HSI) have higher spectral resolution than multispectral images (MSI), but due to limitations of imaging equipment, their width is narrower than MSI. When using partially overlapping HSI-MSI to improve the classification capabilities of MSI, there may be unknown classes that do not exist in HSI-MSI overlapping regions. To solve this problem, this paper proposes a heterogeneous open-set cross-domain manifold embedding aligned method for HSI-MSI collaborative classification. The method designs manifold embedding to align HSI-MSI features to map into subspaces, and gradually selects target domain samples for pseudo-labeling through the designed strategy while rejecting unknown class samples. The feature alignment and pseudo-labeled sample selection are continuously iterated to promote each other, reducing the intra-class distance while pushing the rejected target data away from known classes. The experimental results verify the superiority of our method. Bin Guo 0015, Xiangrong Zhang, Tianzhu Liu, Yanfeng Gu |
IGARSS | 4 |
| 2024 | A Multimodal Hyperspectral Unmixing Method Under Spectral VariabilityabstractVariation in illumination conditions can give rise to divergences in the reflectance profiles corresponding to identical feature categories. This phenomenon has always been a very important challenge that cannot be avoided in spectral unmixing (SU) algorithms. Current methods primarily focus on modeling spectral errors or using spectral libraries for optimization. This paper presents a new multimodal hyperspectral unmixing method under spectral variability that incorporates DSM data to model external shading variations in complicated illumination conditions. The proposed model aims to reduce spectral variability caused by external imaging changes by fitting the shading information with the digital surface model (DSM) data and the illumination information. The reflectance information, which represents the properties of the features themselves, is used for spectral unmixing after introducing the intrinsic decomposition. The method proposed in this paper can effectively attenuate the effect of spectral variability, as demonstrated by experimental validation on real MUFFL multimodal datasets. Yanyuan Huang, Yanfeng Gu, Tianzhu Liu |
IGARSS | 2 |
| 2024 | An Automated Workflow for Pixel-Level BRDF Extraction Using UAV-Based Multispectral ImagesabstractBidirectional Reflectance Distribution Function (BRDF) plays a vital role in quantitative remote sensing. Recently, UAV has gradually emerged as the leading choice for BRDF acquirement. However, challenges remain in extracting usable BRDF data from UAV multispectral images (MSIs), including labor-intensive and limited accuracy. To tackle these challenges, an automated workflow for pixel-level BRDF extraction is proposed in this paper, comprising two main stages: 3D reconstruction and back projection. Pixel-level accuracy is achieved through 3D reconstruction, enhanced by the integration of commercial software for streamlined automation. Back projection is critical for precise location. Experiments were carried out for both selected ROIs and all pixel areas, validating the efficacy of our proposed approach. Zhenqiang Qin, Xian Li 0001, Yanfeng Gu, Xiangrong Zhang |
IGARSS | 3 |
| 2024 | Application of Landweber with Optimization for Small Footprint Waveform Lidar DecompositionabstractSmall-footprint waveform LiDAR requires waveform decomposition for accurate target structure characterization. To improve the ability for identifying close targets, this paper first introduces the Landweber (LW) deconvolution method to decompose the small-footprint LiDAR waveforms. Our study emphasizes the advantages of the deconvolution methods in capturing more targets of waveforms. Generally, the LW approach introduced with optimization excels in detecting more targets after false target removal. Experiments were conducted on datasets collected under various conditions using small-footprint waveform LiDAR system. The findings highlight an average target distance error of 0.083m, showcasing superior performance compared to direct decomposition methods. When compared with the GOLD and RL methods, the decomposition accuracy is nearly indistinguishable, but the success rates are higher. Our research establishes the LW method as a viable waveform decomposition method, contributing to the diversity of choices for waveform data processing. Yanfeng Gu, Xian Li 0001, Xiangrong Zhang |
IGARSS | 2 |
| 2024 | Lurking in the Shadows: Imperceptible Shadow Black-Box Attacks Against Lane Detection Models
Xiaoshu Cui, Yalun Wu, Yanfeng Gu, Endong Tong, Jiqiang Liu, Wenjia Niu |
KSEM (3) | 3 |
| 2024 | Navigating Data in UAV Networks: Harmonic Function-Based Potential Field for Interference-Aware Multi-Hop RoutingabstractMulti-hop packet routing is critical for unmanned aerial vehicle (UAV) networks to enable efficient communication between terminals in diverse environments. However, the complexity of routing design exacerbates due to interference from the environment and link instability caused by high-speed mobility. To address this challenge, we propose a harmonic function-based potential field to assess the impact of interference on UAV networks quantitatively. The proposed field maps the communication quality in terms of interference and mobility onto a virtual three-dimensional plane, providing a metric to establish routing paths. Based on this, two routing algorithms are designed to address two distinct routing requirements of UAV networks, timeliness and losslessness. Leveraging the natural adaptation to the potential field, the two proposed routing algorithms can effectively avoid interference while meeting different requirements. Simulation results demonstrate the effectiveness of the proposed potential field in representing the influences of interference and mobility. Additionally, the results validate the ability of the two routing algorithms to fulfill data communication requirements in terms of delay and accuracy while effectively mitigating interference. Hanze Liu, Zhutian Yang, Nan Zhao 0001, Yanfeng Gu, Chau Yuen |
WCNC | 5 |
| 2024 | Multi-sensor multispectral reconstruction framework based on projection and reconstruction
Tianshuai Li, Tianzhu Liu, Xian Li 0001, Yanfeng Gu, Yushi Chen 0002 |
Sci. China Inf. Sci. | 4 |
| 2024 | An adaptive 3D reconstruction method for asymmetric dual-angle multispectral stereo imaging system on UAV platform
Chen Wang 0060, Xian Li 0001, Yanfeng Gu |
Sci. China Inf. Sci. | 3 |
| 2024 | Interference-Aware Multihop Routing in UAV Networks: A Harmonic-Function-Based Potential Field ApproachabstractMulti-hop packet routing is critical for unmanned aerial vehicle (UAV) networks to enable efficient communication between terminals in diverse environments. However, the complexity of routing design exacerbates due to interference from the environment and link instability caused by high-speed mobility. To address this challenge, we propose a harmonic function-based potential field to assess the impact of interference on UAV networks quantitatively. The proposed field maps the communication quality in terms of interference and mobility onto a virtual three-dimensional plane, providing a metric to establish routing paths. Based on this, two routing algorithms are designed to address two distinct routing requirements of UAV networks, timeliness and losslessness. Leveraging the natural adaptation to the potential field, the two proposed routing algorithms can effectively avoid interference while meeting different requirements. Simulation results demonstrate the effectiveness of the proposed potential field in representing the influences of interference and mobility. Additionally, the results validate the ability of the two routing algorithms to fulfill data communication requirements in terms of delay and accuracy while effectively mitigating interference. Hanze Liu, Zhutian Yang, Nan Zhao 0001, Yanfeng Gu, Chau Yuen |
IEEE Internet Things J. | 4 |
| 2024 | EHGNN: Enhanced Hypergraph Neural Network for Hyperspectral Image ClassificationabstractRecently, the hypergraph neural network (HGNN) has drawn increasing attention in modeling complex high-order correlations. Compared to simple graph neural networks, HGNNs exhibit more powerful representational ability. There are two limitations in the application of hypergraph theory to hyperspectral image (HSI) classification. One is the inadequate explicit representation of semantic information contained in HSI. Another is the loss of pixel-level spectral-spatial information. Thus, an enhanced hypergraph neural network (EHGNN) is proposed to promote the application of hypergraph theory to HSI classification. Specifically, two important enhancements are introduced: 1) the concept of key hypergraph, providing more rich semantic information and improving the interpretability for complex distribution structures, and 2) the integration of convolutional neural network (CNN) and HGNN architectures into an end-to-end framework, the loss of spectral-spatial information at the pixel-level is effectively reduced. Through these two enhancements, EHGNN exhibits a 4% improvement in overall accuracy (OA) on the Pavia University dataset and a 2% improvement in OA on the Xuzhou dataset compared to HGNN. Furthermore, the test results on two HSI datasets demonstrate that our EHGNN achieves competitive performance compared to other state-of-the-art methods. Qingwang Wang, Jiangbo Huang, Tao Shen 0004, Yanfeng Gu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Generative ConvNet Foundation Model With Sparse Modeling and Low-Frequency Reconstruction for Remote Sensing Image InterpretationabstractFoundation models offer a highly versatile and precise solution for intelligent interpretation of remote sensing images, thus greatly facilitating various remote sensing applications. Nevertheless, conventional remote sensing foundational models based on generative transformers neglect the consideration of multiscale features and frequency information, limiting their potential for dense prediction tasks in remote sensing scenarios. In this article, we make the first attempt to propose a generative convolutional neural network (ConvNet) foundation model tailored for remote sensing scenarios, which comprises two key components: First, a large dataset named GeoSense, containing approximately nine million diverse remote sensing images, is constructed to enhance the robustness and generalization of the foundation model during the pretraining phase. Second, a sparse modeling and low-frequency reconstruction (SMLFR) framework is designed for self-supervised representation learning of the ConvNet foundation model. Specifically, a sparse modeling strategy is proposed in masked image modeling (MIM), which allows ConvNet to process variable-length sequences by treating unmasked patches as voxels and sparsifying the encoder. In addition, a low-frequency reconstruction target is designed to guide the model’s attention toward essential ground object features in remote sensing images, while mitigating unnecessary detail interference. To evaluate the general performance of our proposed foundation model, comprehensive experiments have been carried out on five datasets across three downstream tasks. Experimental results demonstrate that our method consistently achieves state-of-the-art performance across all the benchmark datasets and downstream tasks. The code and pretrained models will be available athttps://github.com/HIT-SIRS/SMLFR. Yanfeng Gu, Tianzhu Liu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | UPetu: A Unified Parameter-Efficient Fine-Tuning Framework for Remote Sensing Foundation ModelabstractRecent advancements in remote sensing foundation models have unveiled their tremendous potential in addressing earth observation tasks. Presently, when large-scale foundation models are transferred to downstream tasks, the prevalent approach is to adopt the full-tuning strategy, resulting in significant increases in storage demands and computational costs. Although the introduction of parameter-efficient fine-tuning (PEFT) has mitigated this issue to some extent, mainstream PEFT methods are primarily designed for classification tasks and often prove insufficient to meet the demands of dense prediction tasks. In order to overcome the aforementioned limitations, we propose a unified PEFT framework UPetu, encompassing two essential and complementary modules: the efficient quantization adapter module (EQAM) and the context-aware prompt module (CAPM). EQAM is specifically designed to enhance the correlation between fine-grained feature information and task-specific knowledge through the introduction of quantization linear layers and non-linear activation functions. Additionally, CAPM is introduced to acquire rich contextual features by incorporating trainable prompts into multi-scale features. The synergistic integration of both modules enhances the representation learning capability and generalization transferability of the foundation model. Extensive experiments on three remote sensing scene classification datasets demonstrate the superiority of UPetu over other fine-tuning methods. With the update of only 0.73% of ConvNeXt-B parameters, our UPetu achieves superior performance compared to full-tuning on the UCM-55, AID-28, and AID-55 datasets. Furthermore, experiments conducted on semantic segmentation and change detection tasks provide additional evidence of the effectiveness and generalization capabilities of the proposed UPetu. Yanfeng Gu, Tianzhu Liu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | A Multimodal Unified Representation Learning Framework With Masked Image Modeling for Remote Sensing ImagesabstractThe coordinated utilization of diverse types of satellite sensors provides a more comprehensive view of the Earth’s surface. However, due to the significant heterogeneity across modalities and the scarcity of high-quality labels, most existing methods face bottlenecks in the underutilization of massive unlabeled multimodal satellite data, making it challenging to understand the scene comprehensively. To this end, we propose a multimodal unified representation learning framework (MURLF) based on masked image modeling (MIM) for remote sensing (RS) images, aiming to make better use of massive unlabeled multimodal RS data. MURLF leverages the consistency and complementarity relationships among modalities to extract both common and distinctive features, mitigating the challenges faced by encoders due to significant heterogeneity across various data types. In addition, MURLF uses multilevel masking independently across different modalities, using visual tokens both within the same modality and across modalities to jointly recover masked pixels as the pretext task, facilitating comprehensive cross-modal information interaction. Furthermore, we design a preselected sensor-specific feature extractor (PSFE) to exploit the heterogeneous characteristics of various data sources, thereby extracting discriminative features. By integrating the multistage PSFE with the ViT backbone, MURLF can naturally extract multimodal hierarchical representations for downstream tasks, fully preserving valuable information from each modality. The proposed MURLF is not restricted to multimodal inputs but also supports single-modal inputs during the fine-tuning stage, significantly broadening the framework’s application. Extensive experiments across multiple tasks demonstrate the superiority of the proposed MURLF compared with several advanced multimodal models. The code will be released soon. Dakuan Du, Tianzhu Liu, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Few-Shot Multispectral-Hyperspectral Image Collaborative Classification With Feature Distribution Enhancement and Subdomain AlignmentabstractWith the development of observation technology, multispectral (MS) images of large scenes are easy to obtain, but the low spectral resolution limits their classification ability. Moreover, the collection of training samples is difficult and time-consuming, and limited labeled samples are a challenge for the precise classification of large-scene MS images. This article attempts to use hyperspectral (HS) images with limited labels to help classify MS images of large scenes, so as to achieve better classification results. To solve this problem, a few-shot MS-HS image collaborative classification method combining feature distribution enhancement (FDE) and subdomain alignment is proposed. Specifically, a residual 3-D convolution network embedded with a 3-D FDE module is designed to improve the diversity of the feature distribution extracted by the network and increase the generalization ability of the model under the few-shot condition. Furthermore, the local domain alignment between the source and target domains is achieved by subdomain alignment, which better aligns the categories in the source domain and the target domain, and achieves the distribution alignment of the subdomains. In addition, the feature bias adjustment (FBA) module is introduced in the test phase to correct the bias of the MS image feature representation, and to alleviate the cross-domain problem to some extent. The few-shot learning (FSL) is applied in the source and target domains to learn better feature mapping. The results of comparative experiments on three datasets show that the proposed method is superior to the most advanced method in the case of limited labeled samples. Bin Guo 0015, Tianzhu Liu, Xiangrong Zhang, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Few-Shot Open-Set Collaborative Classification of Multispectral and Hyperspectral Images With Adaptive Joint Similarity MetricabstractHyperspectral images (HSIs) have higher spectral resolution than multispectral (MS) images, but they have a narrower swath than MS images. The limited spectral resolution of MS images constrains their classification capabilities, and annotating remote sensing data is time-consuming and laborious. In addition, large-scale MS images may contain unknown classes not present in the training data. This article attempts to use partially overlapping HS images with limited labels to assist in the classification of large-scene MS images. It can correctly distinguish known classes and simultaneously identify unknown classes, thereby achieving better classification results for MS images. To address this challenge, a few-shot open-set HS–MS image collaborative classification method is proposed. Specifically, a spectral–spatial feature interactive enhancement (SSFIE) module is designed for richer feature extraction and enhanced classification capabilities in the feature extraction stage. In the few-shot learning (FSL) stage, an adaptive joint similarity metric criterion is proposed to improve feature mapping between the source and target domains. Discriminative joint probability adaptation (DJPA) is used for domain adaptation and to enhance feature discriminability, while batch nuclear-norm maximization (BNM) is employed to increase the feature diversity. In the testing phase, the open-set classification module is designed to correctly classify samples of known classes while simultaneously distinguishing unknown classes. The experimental results on four cross-domain HS–MS data pairs demonstrate that our proposed method outperforms state-of-the-art methods. Bin Guo 0015, Xiangrong Zhang, Tianzhu Liu, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Spectral Reconstruction for Paired Images Based on Semi-Supervised Deep LearningabstractSpectral reconstruction (SR) techniques can generate hyperspectral images (HSIs) from multispectral images (MSIs) with the same spatial resolution, thus alleviating the problem of limited availability and low spatial resolution of satellite HSIs. However, in scenarios where both HSIs and MSIs can be acquired simultaneously, spectral mapping relationship (SMR) among real images may not align with the sensor’s spectral response function (SRF), due to factors such as sensor noise and calibration errors. This mismatch can result in discrepancies in reflectivity between the reconstructed HSIs and the real HSIs. To solve the above problems, this article proposes a semi-supervised transfer learning SR (SSTSR) model based on gradient direction constraints. Through semi-supervised learning, SSTSR acquires precise SMRs in overlapping regions and extracts spectral trend information of ground objects from historical models in nonoverlapping regions. Experiments on two datasets demonstrate that the reconstructed HSIs closely resemble real HSIs, leading to impressive classification performance when employing a real HSI classifier. Tianshuai Li, Tianzhu Liu, Yanfeng Gu, Yushi Chen 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | SATDark: A Satellite Video Low-Light Tracking Benchmark for Dark and Weak VehiclesabstractSatellite video single object tracking (SVSOT) stands as a pivotal research area. However, it faces significant challenges in low-light environments, particularly when dealing with dark and weak vehicles. Previous studies have predominantly focused on tracking methods under favorable lighting conditions, neglecting the complexities introduced by inadequate illumination. The difficulty in extracting features from targets in low-light environments, coupled with the susceptibility of dark and weak targets to background noise, exacerbates these challenges. In low-light environments, dark and weak vehicles exhibit less distinctive features and are more susceptible to background interference due to the reduced contrast. To tackle the above challenges, this work proposes an innovative correlation filter (CF)-based tracker (RETrack) that incorporates a retinex-inspired target enhancement. This enhancer integrates an effective low-light enhancement within the CF-based tracker, enhancing target visibility by reallocating target energy based on the characteristics of target motion. Moreover, to mitigate background interference and leverage background information efficiently, an adaptive label update mechanism is developed to suppress background disturbance. Furthermore, this work constructs a satellite video low-light tracking benchmark SATDark, which comprises 120 sequences of dark and weak vehicles. Comprehensive experiments show that RETrack surpasses current leading trackers on the SATDark, providing innovative insights and advancing the field of satellite video object tracking. Additionally, RETrack achieves real-time processing speeds exceeding 30 frames/s on a single CPU, underscoring its practical applicability and efficiency. Jialei Pan, Yanfeng Gu, Guoming Gao, Qiang Wang 0001, Shaochuan Wu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Hemisphere Harmonics Basis: A Universal Approach to Remote Sensing BRDF ApproximationabstractBidirectional Reflectance Distribution Function (BRDF) is an important quantity in remote sensing, describing the variations of reflectance factors with viewing geometries. Current empirical or semi-empirical BRDF models are often constrained to limited types of landcovers due to the assumptions regarding surface cavity distribution. In this paper, a universal approach to remote sensing BRDF approximation based on theHemiSphere Harmonics (HSH)basis function is proposed. We derived the HSH, which match the BRDF definition domain, as basis functions of an infinite series to achieve high-precision BRDF representation. Besides, the proposed approach is universal across various landcovers since HSH basis functions are complete and orthogonal, enabling to approximate arbitrary BRDF data. To our knowledge, it is the first time to utilize the basis function for the BRDF approximation in remote sensing. The proposed approach was validated on both the satellite dataset with 16 landcovers and UAV dataset with 6 landcovers. The results demonstrate that the proposed approach outperforms current BRDF models, and show universality across different landcovers, spectral bands and platforms. Zhenqiang Qin, Xian Li 0001, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | MPS2L: Mutual Prediction Self-Supervised Learning for Remote Sensing Image Change DetectionabstractIn this article, we propose a novel mutual prediction self-supervised learning (MPS2L) method for remote sensing (RS) image change detection (CD). Compared with the previous self-supervised CD methods based on contrastive learning (CL), MPS2L employing a pixel-level training strategy based on masked image modeling (MIM) can effectively train the model to interpret the local scene of RS images. Utilizing global and local scenes and temporal change features extracted from masked bitemporal images to achieve cross-temporal mutual prediction makes the model have the ability to understand the overall observation scene and capture the change information. The training of the two abilities is carried out simultaneously, avoiding the problem of multiobjective conflict or mutual inhibition. To better focus on the changing regions in RS scenes, we further introduce a change feature interaction module (CFIM), comprising spatial and channel feature interaction. The channel interaction module (CIM) can facilitate the cross-temporal transmission of global scene information by channel attention, and the spatial interaction module (SIM) can promote the network to capture information on changing regions by spatial attention. The experimental results on three benchmark RS CD datasets demonstrate the effectiveness and priority of our proposed MPS2L compared to some existing state-of-the-art (SOTA) methods. The source code of the proposed MPS2L will be made available publicly athttps://github.com/KustTeamWQW/MPS2L. Qingwang Wang, Yujie Qiu, Pengcheng Jin, Tao Shen 0004, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Unsupervised Domain Adaptation for Cross-Scene Multispectral Point Cloud ClassificationabstractRemote sensing cross-scene classification has always been an important research field, especially in the field of 3-D classification, which is of great significance. Considering the diversity of collection conditions, seasons, and regional styles, deep learning networks well-trained on one source domain dataset tend to suffer from severe performance degradation when applied to other target domain datasets. To tackle the issue, in this article, we propose a new cross-scene classification method, which combines pre-alignment and Shannon entropy constraint to accomplish unsupervised domain adaptive classification (PS-UDA). On the one hand, the pre-alignment employs$L_{2}$-paradigm constraint and Laplace matrix to pre-align the features. With the$L_{2}$-paradigm constraint, the originally distant features of the source and target domain are constrained to the same sphere surface, and it is easier to make the distribution alignment on the sphere surface. Further, the Laplace matrix is used to map the source and target domain. In this way, similar features of the source and target domain are further aligned, and dissimilar features become discrete from each other. On the other hand, this article employs the Shannon entropy constraint to motivate the network to obtain more high-confidence target domain pseudo-labels. In addition, to fully utilize the unlabeled target domain information, the target domain features are augmented using the adjacency matrix. Experimental results of two cross-scene multispectral point cloud classifications demonstrate that the proposed PS-UDA can effectively mitigate the spectral shift issue in cross-scene multispectral point clouds, achieving state-of-the-art performance. Qingwang Wang, Mingye Wang, Jiangbo Huang, Tianzhu Liu, Tao Shen 0004, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | A High-Resolution and Efficient Waveform Decomposition Method for Small-Footprint LiDARabstractSmall-footprint waveform LiDAR necessitates waveform decomposition for accurate target structure characterization. However, the limited range resolution and heavy computations lead to this being hindered. Addressing these challenges, we propose a high-resolution and efficient waveform decomposition method on fundamentals of the LiDAR physics model. To enhance LiDAR ranging resolution, we introduce a novel technique that separates the transmitted pulses and received multi-target waveforms to simulate narrow transmitted pulse conditions. Then, the separated pulses and waveforms are input into a deconvolution algorithm, which incorporates an automatic stopping criterion for iteration to ensure accurate results. For efficient processing of waveforms, we design a lightweight classification method that categorizes waveforms into single-target and multi-target waveforms before waveform decomposition, with only the latter undergoing complex downstream processing. Indoor and airborne experiments are conducted on datasets collected using small-footprint full waveform LiDAR. The indoor results demonstrate that the average target distance error is reduced to 0.064mwith a significant improvement in efficiency, surpassing mainstream methods. The airborne results reveal that our method is able to decompose faster to get more points for better structural characterization. Yanfeng Gu, Xian Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Intrinsic Hyperspectral Image Recovery for UAV Strips Stitching
Wen Xie 0003, Tianzhu Liu, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | MP2Net: Mask Propagation and Motion Prediction Network for Multiobject Tracking in Satellite VideosabstractMainstream multi-object tracking (MOT) algorithms employ global object detection and association methods. However, when dealing with scenarios involving crowded tiny objects in satellite videos, existing global trackers often yield numerous missed detections and unstable trajectories. To address this issue, we propose a novel joint-detection-and-tracking framework, MP2Net, which integrates local detection enhancements for tiny targets and bridges the gap between detection and association. Specifically, our approach incorporates a mask propagation network that enhances feature representation for tiny targets by matching frame-by-frame to capture local details. Additionally, we utilize an implicit and explicit motion prediction strategy that merges tracking information into detection at both feature and instance levels, thereby improving tracking robustness. Experimental results on two large-scale datasets demonstrate the effectiveness and robustness of MP2Net, achieving state-of-the-art performance on typical moving objects in satellite videos, such as 66.7% MOTA and 75.9% IDF1 on the SatVideoDT challenge dataset. The code will be available at https://github.com/DonDominic/MP2Net. Manqi Zhao, Shengyang Li, Han Wang 0049, Yuhan Sun 0004, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Satellite Video Multi-Label Scene Classification With Spatial and Temporal Feature Cooperative Encoding: A Benchmark Dataset and MethodabstractSatellite video multi-label scene classification predicts semantic labels of multiple ground contents to describe a given satellite observation video, which plays an important role in applications like ocean observation, smart cities, et al. However, the lack of a high-quality and large-scale dataset prevents further improvement of the task. And existing methods on general videos have the difficulty to represent the local details of ground contents when directly applied to the satellite videos. In this paper, our contributions include (1) we develop the first publicly available and large-scale satellite video multi-label scene classification dataset. It consists of 18 classes of static and dynamic ground contents, 3549 videos, and 141960 frames. (2) we propose a baseline method with the novel Spatial and Temporal Feature Cooperative Encoding (STFCE). It exploits the relations between local spatial and temporal features, and models long-term motion information hidden in inter-frame variations. In this way, it can enhance features of local details and obtain the powerful video-scene-level feature representation, which raises the classification performance effectively. Experimental results show that our proposed STFCE outperforms 13 state-of-the-art methods with a global average precision (GAP) of 0.8106 and the careful fusion and joint learning of the spatial, temporal, and motion features are beneficial to achieve a more robust and accurate model. Moreover, benchmarking results show that the proposed dataset is very challenging and we hope it could promote further development of the satellite video multi-label scene classification task. Weilong Guo, Shengyang Li, Feixiang Chen, Yuhan Sun 0004, Yanfeng Gu |
IEEE Trans. Image Process. | 5 |
| 2023 | A multi-frame sparse self-learning PWC-Net for motion estimation in satellite video scenes
Tengfei Wang 0001, Yanfeng Gu, Shengyang Li |
Sci. China Inf. Sci. | 2 |
| 2023 | UTFNet: Uncertainty-Guided Trustworthy Fusion Network for RGB-Thermal Semantic SegmentationabstractIn real-world scenarios, the information quality provided by RGB and thermal (RGB-T) sensors often varies across samples. This variation will negatively impact the performance of semantic segmentation models in utilizing complementary information from RGB-T modalities, resulting in a decrease in accuracy and fusion credibility. Dynamically estimating the uncertainty of each modality for different samples could help the model perceive such information quality variation and then provide guidance for a reliable fusion. With this in mind, we propose a novel uncertainty-guided trustworthy fusion network (UTFNet) for RGB-T semantic segmentation. Specifically, we design an uncertainty estimation and evidential fusion (UEEF) module to quantify the uncertainty of each modality and then utilize the uncertainty to guide the information fusion. In the UEEF module, we introduce the Dirichlet distribution to model the distribution of the predicted probabilities, parameterized with evidence from each modality and then integrate them with the Dempster-Shafer theory (DST). Moreover, illumination evidence gathering (IEG) and multi-scale evidence gathering (MEG) modules by considering illumination and target multi-scale information respectively are designed to gather more reliable evidence. In the IEG module, we calculate the illumination probability and model it as the illumination evidence. The MEG module can collect evidence for each modality across multiple scales. Both qualitative and quantitative results demonstrate the effectiveness of our proposed model in accuracy, robustness and trustworthiness. The code will be accessible at https://github.com/KustTeamWQW/UTFNet. Qingwang Wang, Haochen Song, Tao Shen 0004, Yanfeng Gu |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | A Normalized Spatial-Spectral Supervoxel Segmentation Method for Multispectral Point Cloud DataabstractAirborne LiDAR point cloud segmentation (PCS) is often employed as a preprocessing step for the subsequent object recognition for scene interpretation. Current segmentation methods often aim at single-wavelength LiDAR data by fully exploiting the spatial information, which makes them unsuitable for multispectral point cloud (MPC) data due to ignoring the use of spectral signatures. In this article, a normalized spatial–spectral supervoxel segmentation method is proposed for MPC data. Specifically, a normalized spectral–spatial metric is developed to construct the${k}$-dimensional tree (KD tree) for MPC data clustering. Considering the uneven density distribution of MPC, an adaptive energy minimization principle based on the sum of the distance is devised to accurately select the seed points of voxels, solving the problem of undersegmentation. To reduce the cross-boundary points, the normalized spectral–spatial metric with the concave–convex judgment is extended to further optimize the edges between adjacent voxels. An important asset of our method is to segment MPC without the need for any manual annotation. Experiments on two MPC datasets show that the proposed method yields better performance compared to several comparative methods. Likun Chen, Yanfeng Gu, Xian Li 0001, Xiangrong Zhang, Baisen Liu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Distilling Segmenters From CNNs and Transformers for Remote Sensing Images' Semantic SegmentationabstractSemantic segmentation is a crucial task in remote sensing and has been predominantly performed using convolutional neural networks (CNNs) for the past decade. Recently, transformers with self-attention mechanisms have demonstrated superior performance compared to CNNs. However, due to the locality of CNN and the high computational complexity and massive data resource requirements of transformer, neither of them can be well applied in resource-constrained practical remote sensing scenarios. Motivated by the limitations of using either convolutional neural networks (CNNs) or transformers alone in the task of semantic segmentation of remote sensing images, a novel cross-model knowledge distillation framework, named distilling segmenters from CNNs and transformers (DSCT), is proposed in this paper to harness the complementary advantages of both models. The framework utilizes a channel-weighted attention-guided feature distillation (CAFD) module to condense the feature from the teacher model and enhance the student model’s focus on the teacher-focused regions. Additionally, a target-nontarget knowledge distillation (TNKD) module is proposed that decouples logit distillation into target and nontarget knowledge distillation to guide the student model in learning the underlying representations and decision boundaries from the teacher model. By learning the complementary knowledge from the teacher, our proposed DSCT framework improves the student’s segmentation performance without adding trainable parameters. Experiments on four available remote sensing datasets (ISPRS Potsdam, Vaihingen, GID and LoveDA) indicate that the proposed DSCT outperforms the state-of-the-art knowledge distillation methods and demonstrates its effectiveness and robustness. Guoming Gao, Tianzhu Liu, Yanfeng Gu, Xiangrong Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Spatial and Semantic Consistency Contrastive Learning for Self-Supervised Semantic Segmentation of Remote Sensing ImagesabstractA critical requirement for the success of supervised deep learning lies in having numerous annotated images, which is often challenging to fulfill in remote sensing semantic segmentation tasks. Self-supervised contrastive learning (CL) offers a strategy for learning general feature representations by pre-training neural networks on vast amounts of unlabeled data and subsequently fine-tuning them on downstream tasks with limited annotations. However, the vast majority of CL methods are designed based on instance discriminative pretext tasks, focusing solely on learning the global representation of the entire image while disregarding the essential spatial and semantic correlations crucial for semantic segmentation tasks. To address the above issues, in this paper, we propose a spatial and semantic consistency contrastive learning (SSCCL) framework for the semantic segmentation task of remote sensing images. Specifically, a consistency branch in SSCCL is designed to learn feature representations with spatial and semantic consistency by maximizing the similarity of the overlapping regions of the two augmented views. Additionally, an instance branch is introduced to learn global representations by enforcing the similarity of two augmented views from one image. Through the integration of the consistency branch and instance branch, the proposed SSCCL framework can learn robust and informative feature representations for semantic segmentation in remote sensing scenarios. The proposed method was evaluated on three publicly available remote sensing semantic segmentation datasets, and the experimental results show that our method achieves superior segmentation performance with limited annotations compared to state-of-the-art CL methods as well as ImageNet pre-training method. Tianzhu Liu, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Spectral Reconstruction From Satellite Multispectral Imagery Using Convolution and Transformer Joint NetworkabstractSpectral reconstruction based on satellite multispectral (MS) images can produce high spatial resolution hyperspectral (HS) images at a reasonable cost, significantly expanding the application of satellite-based HS remote sensing. As a challenging ill-posed problem, existing methods have difficulty making full use of local and global information of space and spectra to guide the reconstruction, resulting in limited accuracy in large-scale scenes with complex ground features and severe spectral mixing. In this article, we propose a novel convolution and Transformer joint network (CTJN) to address the challenge of high-accuracy spectral reconstruction in complex scenes. The CTJN is cascaded with shallow feature extraction modules (SFEMs) and deep feature extraction modules (DFEMs), which can explore local spatial features and global spectral features. Besides, a high-frequency Transformer block (HF-TB) is designed to highlight the detailed features of the images to prevent significant high-frequency information loss, which could improve the reconstruction results in regions with drastic feature changes. Moreover, a spatial–spectral recalibration block (SSRB) is proposed to perform explicit constraints on the reconstructed points by exploiting the correlation among neighboring pixels and adjacent spectra. Extensive experimental results on four HS–MS datasets and one MS dataset demonstrate that the proposed CTJN outperforms the state-of-the-art methods in large-scale and small-scale scenes. Dakuan Du, Yanfeng Gu, Tianzhu Liu, Xian Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Intrinsic Decomposition Embedded Spectral Unmixing for Satellite Hyperspectral Images With Endmembers From UAV PlatformabstractTraditional spectral unmixing (SU) of satellite hyperspectral images (HSIs) faces two main challenges: One is that limited by the low resolution of satellite HSIs, it is difficult to guarantee the accuracy of endmember extraction due to severe spectral mixing; the other is that the spectral variability is unavoidable due to external factors such as atmospheric, illumination, and topographic variations, as well as internal factors such as physical changes of the features themselves. Unmanned aerial vehicle (UAV) HSIs of high spatial resolution can provide a highly accurate reflectance curves from regions of interest (ROIs), and the intrinsic image decomposition (IID) technique can reduce the spectral variability caused by external factors. Based on this, a novel IID embedded UAV-satellite spectral unmixing model is proposed. On the one hand, the spectral variability is solved by an embedded IID framework in the inverse problem of SU. The proposed method replaces the input,i.e., the original HSI, with the reflectance component, which is independent of the spectral variability caused by external factors. On the other hand, a UAV spectral library constructed from the UAV HSI is introduced to guarantee the accuracy of the endmember. Thus, by IID embedded in the framework of UAV-satellite collaborative spectral unmixing, the proposed method is able to address the aforementioned problems. Experimental validation is conducted using UAV HSI and three sets of satellite HSI from the Yellow River Delta region. The results indicate that the proposed method can effectively improve the robustness and superiority of the unmixing results. Yanfeng Gu, Yanyuan Huang, Tianzhu Liu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Unsupervised Satellite Video Deep Intrinsic Decomposition Using Physical Prior ConstraintsabstractSatellite video intrinsic decomposition has emerged as a promising area of research with significant application potential. However, existing methods still have certain limitations that hinder their effectiveness in extracting high-quality intrinsic information from complex scenes, ensuring temporal stability of the reflectance component, and achieving computational efficiency. In this article, an unsupervised satellite video intrinsic decomposition network (USVIDNet) is proposed, which overcomes the limitations encountered by existing methods. The USVIDNet incorporates three loss functions based on physical priors: reconstruction loss, chromaticity consistency loss, and spatiotemporal reflectance similarity loss, which provide constraints to guide the intrinsic decomposition process, eliminating the dependence on ground truth intrinsic images required for supervised learning. The network is based on a U-Net architecture variant, consisting of an encoder and two decoders. The encoder captures essential features of the input satellite video, while the decoders focus on predicting two components: reflectance and shading. To enhance the processing efficiency of intrinsic decomposition, a novel initialization-decomposition mode is proposed by leveraging the invariant background characteristics of staring satellites. Experiments are conducted on six Jilin-1 satellite videos to assess the performance of the proposed method in terms of intrinsic component extraction and improved ability of satellite video applications. The experimental results demonstrate the superiority of the proposed method. Yanfeng Gu, Guoming Gao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | A Spatial Alignment Method for UAV LiDAR Strip Adjustment in Nonurban ScenesabstractLiDAR strip adjustment is a key prerequisite for subsequent applications based on point cloud data since it inevitably suffers from spatial discrepancies caused by laser ranging errors, mounting errors, etc. Most current LiDAR strip adjustment methods rely on the extraction of structural features which are often unsuitable for non-urban scenes. Alternative strip adjustment methods based on correspondence distance minimalization ignore spatial alignment. To overcome these limitations, this paper presents an accurate spatial alignment method for UAV LiDAR strip adjustment in non-urban scenes. Firstly, we construct a novel point cloud feature descriptor called Spherical Shell Point Feature (SSPF) to extract multi-dimensional non-structural features that are robust to non-urban point clouds. The constructed SSPF is then combined with point coordinates to generate embedded features, which simultaneously consider the point coordinates and spatial alignment. Finally, the embedded features are utilized by a two-stage matching method to match pair-wise points of two adjacent strips. The proposed method is validated on two non-urban datasets collected by two types of LiDARs, which reduces the digital surface model discrepancies by 0.252mand 0.221m, respectively, and proves its superiority compared to mainstream strip adjustment methods as well. Yanfeng Gu, Xian Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Structure Preserved Discriminative Distribution Adaptation for Multihyperspectral Image Collaborative ClassificationabstractThe fine spectra of the hyperspectral (HS) images can fully reflect the subtle features of the spectra of different objects. However, due to the limitation of the imaging equipment, its swath is not as large as that of multispectral (MS) images. The acquisition of MS images is more convenient, but the discrimination of spectral features is relatively poor. This paper aims to investigate how partially overlapping HS images can be utilized to improve the classification accuracy of large-scene MS images. Due to the spectral mismatch existing between MS and HS features, traditional transfer learning methods cannot solve the problem of classification with heterogeneous features. To address this issue, a novel structure-preserving discriminative distribution adaptive MS-HS image collaborative classification method is proposed in this paper, which aims to improve the classification accuracy of large-scene MS images by discriminative features. Specifically, this method combines statistical properties and geometric constraints in transfer learning, and jointly maximizes the distance between different classes by discriminative least squares to maximize classification accuracy. Moreover, the source and target domains are probabilistically adaptive while maintaining the local structure of MS-HS features, so that the data distribution is fully aligned and the distance between different classes is increased. The learned mapping matrix enables the mapping of multi-scale spectral-spatial features of MS-HS images to subspaces for classification. Compared with related advanced methods, three sets of MS-HS data sets show that the proposed method can effectively reduce the differences between MS-HS data and achieve better classification results. Bin Guo 0015, Tianzhu Liu, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | A Multitask Benchmark Dataset for Satellite Video: Object Detection, Tracking, and SegmentationabstractVideo satellites can continuously image large areas and provide dynamic, real-time monitoring of hotspots and objects. The intelligent processing and analysis of satellite video have become a research hotspot in the field of remote sensing. However, the lack of high-quality satellite video datasets limits the development of relevant object detection, object tracking, and object segmentation. In this paper, we build the largest scale satellite video dataset with the most task types supported and object categories, named Satellite Video Multi-Mission Benchmark (SAT-MTB). First, multi-task annotation of aircraft, ships, cars, trains, and their corresponding 14 categories of fine-grained objects in 249 satellite videos is performed based on horizontal bounding boxes (HBB), oriented bounding boxes (OBB), masks, which cover more than 50,000 frames and 1,033,511 annotated object instances. Then, we review the tasks of object detection, object tracking, and object segmentation based on satellite videos, providing a comprehensive overview of progress in related datasets and algorithm research. Finally, we establish the first public benchmark of multi-task algorithms for satellite video object detection, object tracking, and object segmentation, evaluating and analyzing the performance of a total of 47 representative algorithms under different tasks on the constructed dataset. The proposed SAT-MTB will significantly advance research in intelligent processing and analysis of satellite video and related applications. Shengyang Li, Manqi Zhao, Weilong Guo, Yixuan Lv, Longxuan Kou, Han Wang 0049, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2023 | A Robust Multispectral Point Cloud Generation Method Based on 3-D Reconstruction From Multispectral ImagesabstractMultispectral point cloud is a novel type of data rich in spectral and spatial information. 3D reconstruction is a low-cost solution for acquiring multispectral point cloud. However, most of the existing methods have been developed for RGB images, which are inapplicable to multispectral images due to the special structure of multispectral sensors and the nonlinear intensity differences. In this paper, a robust 3D reconstruction method for multispectral images is proposed to generate multispectral point cloud by harnessing their spatial and spectral information. Considering the characteristics of multispectral image acquisition, reflectance correction and band alignment steps are introduced into the proposed method, aiming to reduce the impact of band differences and spatial errors on 3D reconstruction. Subsequently, a fused multispectral feature extraction is employed to provide more potential reconstruction feature points. To reduce the mismatched feature points induced by the spectra of vegetation regions, an NDVI-guided feature matching algorithm is proposed that provides accurate correspondence calculation for multispectral images reconstruction. The experiments compared with several well-known methods and a commercial software on two datasets have shown superior reconstruction performance. Chen Wang 0060, Yanfeng Gu, Xian Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Quantitative Inversion of Oil Film Thickness Based on Airborne Hyperspectral Data Using the 1DCNN_GRU ModelabstractOil film thickness (OFT) is an important indicator for estimating the amount of oil spill, and accurately quantifying the OFT is of great significance for loss assessment. In this paper, hyperspectral images (HSIs) of different OFTs (0.01-3.04 mm) through a ground experiment were obtained, and the spectral characteristics were analyzed. To address the issue of poor spectral separability for different OFTs, the 1DConvolutional Neural Network_Gate Recurrent Unit (1DCNN_GRU) model was developed for the quantitative inversion of OFT. It was validated through experiments on airborne Cubert-S185 HSI. The experimental results indicated that: (1) The proposed 1DCNN_GRU model effectively addressed the issue of reduced quantitative inversion accuracy resulting from poor spectral separability. The inversion results of it outperformed those of the SVR, CNN, and GRU models. Moreover, the optimal time for hyperspectral sensor to monitor OFT was at noon. (2) The proposed model using airborne hyperspectral data exhibited excellent inversion performance for OFT greater than 0.07 mm, especially with the best performance in 0.60-0.90mm. (3) The accuracy of HSI based OFT inversion assisted by brightness temperature (BT) data was superior to that of OFT inversion using single-source data. In particular, the proposed model had advantages in the feature level and decision level inversion of OFT in the ranges of 0.01-0.30mm and 1.00-3.04mm, respectively. This research provides technical support for the detection of OFT. Junfang Yang, Shanwei Liu, Yanfeng Gu, Mingming Xu 0001, Yi Ma 0004, Jie Zhang 0019, Jianhua Wan |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Hyperspectral Intrinsic Image Decomposition Based on Physical Prior-Driven Unsupervised LearningabstractDeep learning-based intrinsic image decomposition (IID) has gained significant attention in computer vision due to the high efficiency and accuracy of learning-based methods. However, the development of deep learning-based IID methods in the remote sensing field has been limited by the lack of experimental datasets. This article proposes a two-stream encoder-decoder network for the single hyperspectral (HS) image IID task. The proposed network comprises one reflectance estimation subnetwork and one shading estimation subnetwork, which predict intrinsic properties separately. The proposed model introduces three physical losses to enhance performance: 1) In the reflectance estimation subnetwork, the self-similarity loss on the reflectance component is added to satisfy the basic assumption that pixels with similar intensity tend to have a similar reflectance property. 2) In the shading estimation subnetwork, the shading structure loss is added to ensure that the structure of the shading component conforms to physical observation. 3) Reconstruction loss connecting two subnetworks is required to ensure the estimated intrinsic components are physically correct. Finally, to avoid an unreasonable composition, the entire network is initialized by reflectance estimated by the physical model. The quantitative experimental results of intraclass consistency and classification metrics demonstrate that the proposed physical prior-driven unsupervised learning-based IID network outperforms the current available learning or optimization-based approaches. Wen Xie 0003, Yanfeng Gu, Tianzhu Liu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Shadow-Less Intrinsic Hyperspectral Point Cloud Generation From HSIs and LiDARabstractGenerating hyperspectral point cloud from hyperspectral images (HSIs) and light detection and ranging (LiDAR) has become more and more common in the remote sensing field and supported various applications. One challenge here is that hyperspectral imaging is a passive imaging method and is suffering from shadows in a natural scene. Intrinsic information recovery can effectively eliminate the spectral variation caused by illumination changes; however, it assumes a uniform light and neglects the shadows in the scene. In this article, we provide a novel hyperspectral point cloud intrinsic model that can detect the shaded regions and recover reflectance information in them. We first estimate the global illumination of the scene using an intrinsic information recovery method. Then, we perform supervoxel segmentation on hyperspectral point cloud to calculate the blocking relation of supervoxels and therefore accurately detect shaded regions. Finally, we estimate the illumination and reflectance of shaded regions based on an illumination-invariant spectral prior. The experimental results show that the proposed method can effectively detect shaded areas and robustly generate shadow-less intrinsic hyperspectral point cloud. Wen Xie 0003, Xudong Jin, Yanfeng Gu, Tianzhu Liu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Superpixel Consistency Saliency Map Generation for Weakly Supervised Semantic Segmentation of Remote Sensing ImagesabstractThe weakly supervised semantic segmentation (WSSS) method aims to assign semantic labels to each image pixel from weak (image-level) instead of strong (pixel-level) labels, which can greatly reduce human labor costs. However, there are some problems in WSSS of remote sensing images such as how to locate labels accurately, and how to get precise segmentation edges. To address these issues, we propose a novel framework directly transferring the scene classification model to perform semantic segmentation. We first train a multi-label scene classification network as the encoder to obtain the pre-trained model, then the feature learned by the model is transferred to the decoder. Different from other methods, we propose a saliency map generator instead of the Class Activation Map for more accurate location information by making pixels belonging to the same class lie close together while different classes are separated in feature space. Meanwhile, we take the superpixel patch as processing unit to provide precise boundary inhibition for the saliency map. To assign semantic labels for each patch, combined with extracted salient region, we propose a module responsible for exploiting the consistency of spatial and semantic similarity between different patches. Finally, we incorporate the above two modules to supervise the training process of the decoder without generating pseudo labels as most methods do, thus simplifying the training process. Experimental results show that our method outperforms other weakly supervised approaches on DLRSD and WHDLD datasets with at least a 3% improvement on mean intersection over union. Xiaopeng Zeng, Tengfei Wang 0001, Xiangrong Zhang, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | An End-to-End Framework for Joint Denoising and Classification of Hyperspectral ImagesabstractImage denoising and classification are typically conducted separately and sequentially according to their respective objectives. In such a setup, where the two tasks are decoupled, the denoising operation does not optimally serve the classification task and sometimes even deteriorates it. We introduce here a unified deep learning framework for joint denoising and classification of high-dimensional images, and we particularly apply it in the framework of hyperspectral imaging. Earlier works on joint image denoising and classification are very scarce, and to the best of our knowledge, no deep learning models were proposed or studied yet for this type of multitask image processing. A key component in our joint learning model is a compound loss function, designed in such a way that the denoising and classification operations benefit each other iteratively during the learning process. Hyperspectral images (HSIs) are particularly challenging for both denoising and classification due to their high dimensionality and varying noise statistics across the bands. We argue that a well-designed end-to-end deep learning framework for joint denoising and classification is superior to current deep learning approaches for processing HSI data, and we substantiate this by results on real HSI images in remote sensing. We experimentally show that the proposed joint learning framework substantially improves the classification performance compared to the common deep learning approaches in HSI processing, and as a by-product, the denoising results are enhanced as well, especially in terms of the semantic content, benefiting from the classification. Xian Li 0001, Mingli Ding, Yanfeng Gu, Aleksandra Pizurica |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Key Region Extraction Via Scene Classification ModelabstractKey regions are the similar regions of the scenes of the same category and they are the crucial and explicable areas. In this paper, we extract the key regions from the satellite images. The proposed method transfers the model of scene classification and obtains the result without the segmentation labels. First, the features of each layer of current scene classification models are extracted to find the relationships of key regions and scene classification labels. Second, we reconstruct the images from the relative features to obtain end-to-end results. There are areas of similar features between the images of the same categories and no areas of different categories. Last, we design the masks to obtain the key regions from the end-to-end results. The experiments were conducted on a typical scene classification dataset. The experimental result demonstrates the feasibility of extracting key regions by scene classification without relabeling the dataset. Tengfei Wang 0001, Yanfeng Gu, Xiaopeng Zeng |
IGARSS | 2 |
| 2022 | A New Radiometric Correction Method for Multiple UAV Multispectral Images Under Varying Illumination ConditionabstractConsidering the spectral accuracy of UAV multispectral images is mainly influenced by the sensor factor and illumination factor, we separate radiometric correction into radiometric calibration and illumination and reflectance spectra separation (IRSS). Then focusing on the influence of varying illumination, we proposed the multiple-image IRSS (MI-IRSS) model for relative radiometric correction for multiple images. MI-IRSS can not only do a radiometric correction process for multiple images without irradiance measurement, but also can give a clear physical meaning which is illumination function to relative correction coefficients. Using reflectance images corrected by precise irradiance as validation data, the experimental results show that MI-IRSS can eliminate changes in radiometric level due to the varying illumination, with the mean root mean square error (MRMSE) of reflectance reaching 0.048, and the spectral angle mapping (SAM) between the estimated and measured illumination function reaching 0.154. Zhenqiang Qin, Yanfeng Gu |
IGARSS | 2 |
| 2022 | A Method for Generating True Digital Orthophoto Map of UAV Platform Push-Broom Hyperspectral Scanners Assisted by LidarabstractUAV platforms equipped with hyperspectral sensors are widely welcomed in various fields for rich spectral and spatial information. This paper presents a method for generating true digital orthophoto map of push-broom hyperspectral scanners assisted by Lidar, which is suitable for UAV platforms with low flight altitude. Different from the traditional digital orthophoto map generation method, this method utilizes Lidar to obtain the elevation information lost in the hyperspectral imaging process. GNSS/INS unit is used for direct georeferencing to remove geometric distortions caused by platform instability. Furthermore, we reduce the projection errors caused by sudden changes in elevation assisted by Lidar data. The UAV flight test in urban scenes verifies the effectiveness and accuracy of the method. Chen Wang 0060, Yanfeng Gu |
IGARSS | 2 |
| 2022 | Satellite Video Intrinsic DecompositionabstractExisting satellite video processing methods are mainly based on original video, ignoring the use of invariant background characteristics of staring satellites, and easy to be disturbed by rapid light changes. In order to improve application capability of satellite video, this paper establishes the satellite video intrinsic decomposition (SVID) model, including satellite video signal composition model, decomposition constraint with time-spatial unity similarity constraint, static and dynamic components separation by improving TRPCA, and decomposition acceleration based on reflectance transfer. With SVID, intrinsic decomposition and dynamic and static component separation are realized. Five Jilin-1 satellite videos are used to verify the validity, superiority and the potential applications of the proposed algorithm. By comparing with state-of-the-art intrinsic image decomposition method and intrinsic video decomposition method, the experimental results prove the superiority of the SVID method in extracting reflectance component. In addition, the experimental results also prove SVID has excellent application ability in scene background analysis and moving target tracking. Guoming Gao, Yanfeng Gu, Shengyang Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Multitemporal Intrinsic Image Decomposition With Temporal-Spatial Energy Constraints for Remote Sensing Image AnalysisabstractDue to interference with remote imaging by some natural factors, the multitemporal analysis ability is limited by the spectral drift between images. In this article, a new approach to optimize the existing multitemporal analysis system is proposed: multitemporal intrinsic image decomposition (MIID). The MIID method is designed to extract common spectral reflectance from multitemporal images. With MIID, multitemporal classification, changing detection, and index extracting will become extremely easy and more accurate. Firstly, without considering land cover change, the general MIID framework is proposed by adding local temporal–spatial energy constraints in traditional intrinsic images decomposition. On this basis, an improved MIID method with change detection (CD) (CD-MIID) capability is proposed to make the model adapt to the land cover change situation. Finally, specific steps of how to use MIID methods in the multitemporal analysis are given. Multitemporal multispectral/hyperspectral remote sensing images from GF-1, GF-2, GF-5, Landsat TM, and two groups of captured datasets with reflectance truth map are used to evaluate the performance. The experimental results show the following two points: first, the MIID methods achieve better extraction results of spectral reflectance. Second, the proposed MIID methods have better performance both on multitemporal classification and CD. Guoming Gao, Baisen Liu, Xiangrong Zhang, Xudong Jin, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | An Intensity-Independent Stereo Registration Method of Push-Broom Hyperspectral Scanner and LiDAR on UAV PlatformsabstractUnmanned aerial vehicles (UAVs) equipped with hyperspectral scanners and LiDARs can flexibly acquire rich spectral and geometric information about the observation scene. To combine the complementary advantages of multi-source data, the stereo registration of hyperspectral images and LiDAR data has become one of the hot topics in remote sensing community. However, existing research works are more focused on exploiting intensity information from multi-source data, which is applicable to data acquired on manned vehicle platforms or satellite platforms. For UAV platforms with poor stability and limited load, the low signal-to-noise ratio of LiDAR data and the complex distortion of push-broom images bring great challenges to stereo registration. Under this circumstance, an intensity-independent stereo registration method is proposed in this paper, which is based on the physical model of the integrated system and the sensor detection principles. Specifically, the proposed method utilizes the position and orientation system (POS) to reduce the impact of UAV platform motion on hyperspectral imaging, and projection errors are eliminated by the ray tracing model with aid of LiDAR data. Finally, a virtual ray decomposition model based on geometric features is constructed to realize the stereo registration of hyperspectral images and LiDAR data. Compared with an advanced solution and professional processing software, the proposed method has shown better registration performance on two data of different scenarios. Yanfeng Gu, Chen Wang 0060, Xian Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Hyperspectral Intrinsic Image Decomposition With Enhanced Spatial InformationabstractHyperspectral intrinsic image decomposition (HyperIID) has been proven to be a very useful approach to reduce the spectral uncertainty in the remote sensing imaging process and improve the classification. In this article, a new HyperIID with enhanced spatial information, called ESI-IID, is proposed to overcome the deficiency of low spatial resolution in the existing HyperIID methods. With the aid of high-resolution (HR) panchromatic (PAN) image, the proposed method embeds the HR spatial information into the intrinsic decomposition model and enhances spatial details of the intrinsic component. The proposed ESI-IID introduces three constraints: 1) we make the constraint on spectral information to protect it from distortion during the spatial resolution enhancement process; 2) we add the constraint on spatial information to make sure that the details of edges will be well kept; and 3) based on the assumption that the reflectance component has a strong correlation in the local neighborhood, we add the self-constraint on reflectance component, in which the similarity matrix consists of two parts extracted from hyperspectral images and PAN image, respectively. Finally, we build a matrix energy function according to the aforementioned constraints and solve it by finding the minimum Frobenius norm iteratively. Both visual and quantitative experiments on simulated and real datasets demonstrate that the proposed method outperforms other alternative methods with high reliability. Yanfeng Gu, Wen Xie 0003, Xian Li 0001, Xudong Jin |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Integrating Coupled Dictionary Learning and Distance Preserved Probability Distribution Adaptation for Multispectral-Hyperspectral Image Collaborative ClassificationabstractWith the development of observation technology in remote sensing (RS), large-area multispectral (MS) images can be easily obtained. However, due to the limitation of imaging devices, only a limited range of hyperspectral (HS) images with higher spectral resolution can be obtained. This article mainly focuses on how to use limited HS images to improve the classification performance of MS images. In order to solve this problem, this article proposes an MS–HS image collaborative classification method, which integrates coupled dictionary learning and distance preserved probability distribution. First, image reconstruction based on coupled dictionary learning is performed, in which sparse representation and dictionary learning are used to generate HS images from MS images through spectral superresolution, so that the spectral features of the MS data and HS data are converted to the same feature space for feature space alignment. Second, the probability distribution is adapted, in which the marginal and conditional probabilities are adapted to further narrow the difference between the real HS data and the generated HS data. At the same time, the consistency of the data structure of the source domain before and after the mapping is maintained, so that the same class of data is more compact after the mapping and reduces the spacing within the same class. Compared with the state-of-the-art methods, this article conducts the experiments on three MS–HS RS datasets, which demonstrate the superiority of the proposed method. Bin Guo 0015, Tianzhu Liu, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Supervoxel-Based Intrinsic Scene Properties From Hyperspectral Images and LiDARabstractThe combination of spectral and 3-D elevation information provided by hyperspectral images (HSIs) and Light Detection and Ranging (LiDAR) has gained increased attention in the remote sensing field and enabled numerous applications. While various methods have been proposed to fuse these two data streams in pixel, feature, or decision level, a deeper view into the intrinsic relation of surface geometry, material reflectance, and environment illumination is still lacking. In this article, we present a novel supervoxel-based joint intrinsic decomposition framework for HSIs and LiDAR. First, we proposed a novel intrinsic scene model for HSIs and LiDAR point cloud, which tells how we can map LiDAR point cloud into HSI pixels with point-cloud-level normals, reflectance, and incident light direction. Then, we extract supervoxels from the LiDAR point cloud using a graph-based supervoxel method. Finally, we formulate the intrinsic decomposition problem within a supervoxel-based framework which can be optimized effectively and efficiently. The outputs of the proposed model are intrinsic scene properties like incident light direction and point-cloud-level hyperspectral reflectance, with which we can then generate intrinsic hyperspectral point cloud (IHSPC) where each point possesses not only 3-D coordinates and normals but also the reflectance over each wavelength. The performance of our approach is demonstrated with both synthetic and real data. Xudong Jin, Yanfeng Gu, Tianzhu Liu, Wen Xie 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Intrinsic Hyperspectral Image Decomposition With DSM CuesabstractIntrinsic hyperspectral image decomposition (IHID) aims to recover physical scene properties such as reflectance and illumination from a given hyperspectral image (HSI), which directly respects the physical imaging process and can benefit many HSI processing tasks. It is a severely ill-posed problem and is challenging to solve using HSI alone. Additional geometric information provided by digital surface models (DSMs) can otherwise help immensely. While intrinsic image decomposition for RGB images and RGB-D images has been studied extensively during the past few decades and has seen significant progress, studies of the problem for other types of data, such as HSIs and DSMs, are still needed. It is much more challenging to handle an HSI with hundreds of channels than an RGB image with only three channels. Moreover, compared with RGB-D data, HSIs and DSM data usually have much lower spatial resolutions and more complicated land covers, making it difficult to extend the RGB-D intrinsic image method directly. In this article, we present a novel IHID framework for HSIs with DSM cues. Utilizing spherical-harmonic illumination, we first propose a convenient HSI rendering model with DSM, which describes the interplay of material reflectance, geometric distribution, and environment illumination. Then, we introduce local and nonlocal priors on reflectance that ensure the local smooth and global consistency of recovered reflectance. Experiments on synthetic and real data demonstrate that the proposed method outperforms the state-of-the-art methods and is robust to illumination changes. Xudong Jin, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Progressive Spatial-Spectral Joint Network for Hyperspectral Image ReconstructionabstractHyperspectral (HS) images are widely used to identify and characterize objects in a scene of interest with high acquisition costs and low spatial resolution. It is an inexpensive way to obtain high-spatial-resolution HS images (HSIs) by spectral reconstruction from high-spatial-resolution multispectral (MS) images. In this article, we proposed a progressive spatial–spectral joint network (PSJN) to reconstruct HSIs from MS images. PSJN is composed of a 2-D spatial feature extraction module, a 3-D progressive spatial–spectral feature construction module, and a spectral postprocessing module. PSJN makes full use of the shallow spatial features extracted by the 2-D spatial feature extraction module with the spatial–spectral features extracted by the 3-D progressive spatial–spectral feature construction module. The 3-D progressive spatial–spectral feature construction module is designed to extract spatial–spectral information from local spectra in local space and construct spectral information from a few bands to a lot of bands in a pyramidal structure. Besides, a network updating mechanism is proposed to improve the spectral reconstruction effect of the images with poor original spectral reconstruction effect. The experimental results on three HS–MS datasets and one MS dataset demonstrate the efficacy of the proposed methods. Compared with the most advanced spectral reconstruction methods based on dictionary learning and deep learning, our method achieves the best performance of the latest methods in similarity evaluation and classification performance evaluation. Tianshuai Li, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Unified Multiview Spectral Feature Learning Framework for Hyperspectral Image ClassificationabstractRecent progress in spectral classification is dominated by the use of deep learning models. While various learning architectures have been developed, they all extract spectral features from a single view input. In this paper, we investigate a different perspective and develop a unified multiview spectral feature learning framework, which extracts discriminative spectral features from multiple views of inputs. To our knowledge, this is the first reported multiview spectral feature learning method based on deep learning. In this framework, we introduce a multiview spectrum construction method by transforming the input spectral vector into multiple 3D image patches with different sizes, termed as multiview spectrum. This multiview spectrum is fed to a well-designed triple-stream architecture, where a global and two local spectral feature learning networks operate in parallel, capturing thus both global and local spectral contextual features simultaneously. Another important contribution of this work is a novel interactive attention mechanism to identify the most informative spectral contextual features. The model is trained in an end-to-end fashion from scratch with a joint loss. Experimental results on four data sets demonstrate excellent performance compared to the current state-of-the-art. Xian Li 0001, Yanfeng Gu, Aleksandra Pizurica |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Spectral Reconstruction Network From Multispectral Images to Hyperspectral Images: A Multitemporal CaseabstractHyperspectral satellite data has been widely applied in many fields due to its numerous bands. Along with the advantages of high spectral resolution, hyperspectral satellite data are still limited by some disadvantages of high acquisition cost, low revisiting capability, and low spatial resolution. Compared with hyperspectral satellites, multispectral satellites have a large number, large width, strong coverage and high spatial resolution. Therefore, multispectral data can be used as the input to the spectral reconstruction to obtain hyperspectral data with high temporal resolution. Better hyperspectral data can be obtained by spectral reconstructing with these continuous multi-temporal data than with single-temporal data. A multi-temporal spectral reconstruction network (MTSRN) is proposed in this paper, which is used to reconstruct hyperspectral images from multi-temporal multispectral images. The proposed MTSRN comprises multiple single-temporal spectral reconstruction networks (STSRN) for extracting temporal features and a multi-temporal fusion network (MTFN). The parallel component alternative (PA) post-processing method enhances the physical plausibility of reconstructed hyperspectral data. To demonstrate performance of the proposed method in aspects of multi-temporal reconstruction, experiments are conducted on four multi-temporal hyperspectral and multispectral satellite datasets. The experimental results prove that the proposed MTSRN obtains better spectral reconstruction results compared with the spectral reconstruction method based on single-temporal information. Tianshuai Li, Tianzhu Liu, Xian Li 0001, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Deep Joint Estimation Network for Satellite Video Super-Resolution With Multiple DegradationsabstractSuper-resolution (SR) for satellite video data has been a hot research topic in the field of remote sensing video analysis. The existing satellite video SR methods assume that the blur kernel in the imaging degradation model is known. However, the blur kernel in real satellite videos is usually unknown, which inevitably results in poor performance when the true blur kernel is not consistent with a predefined blur kernel. To address this issue, this article proposes a deep joint estimation network for satellite video SR (JENSVSR), which jointly estimates blur kernels and SR frames. Specifically, JENSVSR is composed of a video SR subnetwork and a blur kernel estimation subnetwork. On one hand, the video SR subnetwork makes use of multiple video frames to generate super-resolved satellite frames. To effectively fuse information from adjacent frames, an alignment and fusion module is proposed in the feature space. On the other hand, the blur estimation subnetwork is also proposed to predict blur kernels. The two subnetworks are coupled by cross-task feature fusion modules (CTFFMs) to achieve joint estimation rather than two-step independent estimation. The performance of our proposed method is evaluated on synthetic and real satellite videos. The experimental results show that our proposed method is superior to the current state-of-the-art SR methods. Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Separable Coupled Dictionary Learning for Large-Scene Precise Classification of Multispectral ImagesabstractLarge-scene precise classification of multispectral images (MSIs) has become one of the hot topics in remote sensing field. MSIs usually have wide swath and a meter or even submeter level of spatial resolution, which make large-scene observation possible. However, the limited number of spectral bands leads to the confusion of land covers in classification, especially for the large-scene conditions with abundant land cover types. Therefore, overlapped hyperspectral images (HSIs) can be used to improve the precision degree of classification. To achieve this purpose, coupled dictionary learning has been proposed as a major means. Aiming at separating the class-specific characteristics and mutual patterns among different land covers, this paper proposed a separable coupled dictionary learning (SCDL) method, which converts the separation of mutual features into the construction of separable coupled dictionaries and learns both class-specific coupled dictionaries and mutual coupled dictionaries simultaneously with the aid of label information. More specifically, the proposed method uses the labels of training samples to construct class-specific reconstruction error constraint, class-specificity constraint and separable dictionary incoherence constraint as regularization terms, to make sure that the learned coupled dictionaries to be both compact and discriminative. The learned separable coupled dictionaries facilitate pixels belong to the same category to be represented by the mutual dictionary and the class-specific sub-dictionary of corresponding class. The experiments compared with several state-of-the-art methods on three pairs of HSI and MSI have shown better classification performance. Tianzhu Liu, Yanfeng Gu, Wenyong Yu, Xiuping Jia, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Intrinsic Satellite Video Decomposition With Motion Target Energy ConstraintabstractSatellite videos dynamically monitor the Earth’s surface by using staring imaging, which has gained increased attention and enabled target tracking applications. While various tracking methods are processed on the original video, the rapid light changes due to staring imaging are not considered and have a negative effect on tracking. To reduce the effects of illumination and improve the performance of satellite video target tracking, an intrinsic satellite video decomposition model with motion target energy constraint, called MTE-ISVD, is proposed in this paper. The proposed algorithm introduces two main constraints: The first is a temporal constraint of reflectance, which can solve the flicker problem by preserving reflectance coherence in the time domain with the property that the background pixels in satellite videos are nearly consistent between adjacent frames. The second is a motion target energy constraint, which can concentrate the signal energy of the motion targets in the reflectance by representing them with the surrounding background in the shading. The decomposition problem is reformulated as a quadratic function minimization, which can be addressed using the standard conjugate gradient in closed form. For visual and quantitative comparisons, we perform experiments on five Jilin-1 satellite videos and analyze the results in terms of visual comparison, target tracking improvement, stability evaluation and processing time comparison. The experimental results demonstrate that our proposed method outperforms the other representative intrinsic decomposition methods in terms of processing speed, stability, and motion target representation. Jialei Pan, Yanfeng Gu, Shengyang Li, Guoming Gao, Shaochuan Wu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | An Illumination Estimation and Compensation Method for Radiometric Correction of UAV Multispectral ImagesabstractThe multispectral imaging of Unmanned Aerial Vehicle (UAV) is often affected by the variation of illumination, resulting in serious spectral radiation distortion. Precise illumination estimation and compensation is a key step to carry out the radiometric correction on UAV multispectral images (MSIs), especially for the case without irradiance sensors. To accurately estimate the illumination for the radiometric correction, a physics-based illumination estimation and compensation method is proposed in this paper. In the proposed method, an illumination estimation model is built based on the intra-image hypothesis on illumination consistency and the inter-image hypothesis on reflectance consistency. This model is used to obtain the illumination irradiance of each one from numerous MSIs simultaneously. Then the influence of varying illumination can be alleviated with the estimated irradiance based on the physical imaging principle. To validate the effectiveness of the proposed method, numerical experiments are conducted on three UAV datasets acquired under cloudy weather. The experimental results demonstrate that the proposed method outperforms current methods and the Normalized Root Mean Square Error (NRMSE) on the three datasets are noticeably reduced. Zhenqiang Qin, Xian Li 0001, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Satellite Video Scene Classification Using Low-Rank Sparse Representation Two-Stream NetworksabstractSatellite video scene classification (SVSC) is a challenging work in remote sensing. The main procedure of SVSC is spatial–temporal feature extraction. Unfortunately, massive numbers of dim small moving targets and the low signal-to-noise ratio (SNR) of satellite video bring great challenges to feature extraction. It is difficult to apply traditional feature extraction methods to SVSC because they are used to classify the actions of high-quality video. According to the theory of low-rank sparse decomposition, a Low-rank Sparse Representation Two-stream Network (LSRTN) is designed to increase the classification accuracy of two-stream networks. First, we propose a Low-rank Sparse Component Analysis Network (LSCAN) to decompose satellite videos into low-rank background images and sparse moving target sequences. The LSCAN possesses the advantage of low-rank sparse decomposition to solve small targets and has the capability to adjust the features using the data. Moreover, the LSCAN can efficiently improve the feature extraction of low SNR video. Second, a two-stream structure that was proven to be effective for multiclass video classification was applied to obtain the spatial features and temporal features in each stream. Finally, a fully connected layer integrates the features to classify the satellite video scenes. To utilize the label information, we refine the loss function to adjust the degree of low-rank sparse characteristics and ensure the classification accuracy of training. The experimental results demonstrate that the proposed method achieves better performance than the baseline methods for the SVSC task. Tengfei Wang 0001, Yanfeng Gu, Guoming Gao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Efficient Convolutional Neural Architecture Search for LiDAR DSM ClassificationabstractLight detection and ranging (LiDAR) data provide rich elevation information, so it plays an irreplaceable role in ground object classification. Recently, convolutional neural networks (CNNs) have shown excellent performance in LiDAR digital surface models (DSMs) classification. However, the architecture of CNN model relies heavily on manual design, so it has great limitations. In addition, different sensors capture LiDAR datasets with different properties, so the model should be designed to suit for different datasets, which further increases the workload of architecture design. Therefore, this article proposes a method of automatic design of LiDAR DSM classification model. First, attention mechanism is introduced into search space to improve the feature extraction capability of the model. Then, a gradient-based search strategy is used to obtain the optimal architecture from this search space. Second, a learning rate adjustment strategy is proposed to reduce the time spent in the search stage and evaluation stage to improve the classification accuracy of the model. Finally, a regularization scheme is introduced to enhance the robustness of the model and avoid overfitting. Experimental results on three public LiDAR datasets (Bayview Park, Recology, and Houston) obtained from different sensors show that the proposed neural architecture search method achieves the impressive classification performance compared to several state-of-the-art classification methods and improves the classification accuracy under the condition of limited training samples. Aili Wang 0001, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | A Novel Low Rank Smooth Flat-Field Correction Algorithm for Hyperspectral Microscopy ImagingabstractA flat-field correction method is proposed for multiple measured hyperspectral microscopy imaging in this paper. As the most crucial preprocessing process in quantitative microscopic analysis, flat-field correction solves the uneven illumination caused by vignetting in microscopic images, and guarantees the precision of spatial and spectral information in hyperspectral microscopic imaging. In order to carry out flat-field correction and extract uneven illumination among groups of hyperspectral microscopic data containing hundreds of bands simultaneously, two properties of vignetting have been exploited: i) low-rank property is reflected by little information contained in vignetting; ii) local smoothness can be observed as a gradual change in brightness of vignetting, which is typically equivalent to the sparseness in spatial frequency domain. Combining the two properties above, a novel Low Rank Smooth Flat-field Correction (LRSFC) model modified from common orthogonal basis extraction is proposed, while an optimization is solved based on alternating direction multiplier method (ADMM), obtaining a unique flat-field term with low-rank and smooth properties. Qualitative and quantitative experimental assessments indicate that LRSFC does not add extra cell texture to the extracted flat-field term, whose performance appears prior to other state-of-the-art flat-field correction methods. Yanfeng Gu |
IEEE Trans. Medical Imaging | 2 |
| 2021 | Vehicle Detection Using Deep Learning with Deformable ConvolutionabstractAiming at accurately detect vehicles in high-resolution remote sensing images, this paper proposes a target detection framework combining region-based fully convolutional networks (R-FCN) and deformable convolution (DCN). The difficulty of vehicle detection is that its pixel range is small and difficult to detect, R-FCN calculates confidence scores pixel by pixel, and uses a confidence scoring map related to the number of categories and local parts of the target as the output of the network, which can make full use of the limited feature information of vehicles. As to the precision reduction caused by geometric deformation of vehicle images, the fixed structure of the convolution kernel is improved, and the convolution kernel of part of the convolution layers and region of interest (RoI) pooling layers in the network are deformable to make it adapt to the deformation of targets. Experiments show that the R-FCN equipped with deformable convolution and deformable RoI pooling has advantages in detection precision and detection time. Shujia Ye, Guoming Gao, Yanfeng Gu |
IGARSS | 5 |
| 2021 | Multimodal hyperspectral remote sensing: an overview and perspective
Yanfeng Gu, Tianzhu Liu, Guoming Gao, Guangbo Ren, Jocelyn Chanussot, Xiuping Jia |
Sci. China Inf. Sci. | 1 |
| 2021 | Rotation adaptive correlation filter for moving object tracking in satellite videos
Shiyu Xuan, Shengyang Li, Zifei Zhao, Wanfeng Zhang, Hong Tan, Gui-Song Xia, Yanfeng Gu |
Neurocomputing | 8 |
| 2021 | An Online Distributed Satellite Cooperative Observation Scheduling Algorithm Based on Multiagent Deep Reinforcement LearningabstractThe provision of real-time information services is one of the crucial functions of satellites. In comparison with the centralized scheduling, the distributed scheduling can provide better robustness and extendibility. However, the existing distributed satellite scheduling algorithms require a large amount of communication between satellites to coordinate tasks, which makes it difficult to support scheduling in real-time. This letter proposes a multiagent deep reinforcement learning (MADRL)-based method to solve the problem of scheduling real-time multisatellite cooperative observation. The method enables satellites to share their decision policy, but it is not necessary to share data on the decisions they make or data on their current internal state. The satellites can use the decision policy to infer the decisions of other satellites to decide whether to accept a task when they receive a new request for observations. In this way, our method can significantly reduce the communication overhead and improve the response time. The pillar of the architecture is a multiagent deep deterministic policy gradient network. Our simulation results show that the proposed method is stable and effective. In comparison with the Contract Net Protocol method, our algorithm can reduce the communication overhead and achieve better use of satellite resources. Li Dalin, Wang Haijiao, Yanfeng Gu, Shi Shen |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2020 | Weak Target Detection in High-Resolution Remote Sensing Images by Combining Super-Resolution and Deformable FPNabstractWeak target detection plays an important role in military and civilian fields. However, due to the limitation of the target size and the influence of complex background, the detection of weak target is a huge challenge. Therefore, based on high-resolution remote sensing image, this paper proposes a weak target detection network which combines super-resolution and deformable convolution. Firstly, the high-resolution remote sensing image is expanded and enhanced to eliminate the influence of complex background. Secondly, a detection network based on the deformable convolution and feature pyramid network (FPN) is used to solve the problem of less information caused by the fewer target pixels. In addition, this paper establishes a detection dataset only containing weak vehicles. The experimental results show that the proposed method achieves better detection results in the weak target detection problem. Tongyuan Zou, Shujia Ye, Zhenqiang Qin, Guoming Gao, Yanfeng Gu |
IGARSS | 6 |
| 2020 | Spatial-Spectral Smooth Graph Convolutional Network for Multispectral Point Cloud ClassificationabstractMultispectral point cloud, as a new type of data containing both spectrum and spatial geometry, opens the door to three-dimensional (3D) land cover classification at a finer scale. In this paper, we model the multispectral point cloud as a spatial-spectral graph and propose a smooth graph convolutional network for multispectral point cloud classification, abbreviated 3SGCN. We construct the spectral graph and spatial graph respectively to mine patterns in spectral and spatial geometric domains. Then, the multispectral point cloud graph is generated by combining the spatial and spectral graphs. For remote sensing scene classification tasks, it is usually desirable to make the classification map relatively smooth and avoid salt and pepper noise. Heat operator is introduced to enhance the low- frequency filters and enforce the smoothness in the graph signal. Further, a graph -based smoothness prior is deployed in our loss function. Experiments are conducted on real multispectral point cloud. The experimental results demonstrate that 3 SGCN can achieve significant improvements in comparison with several state-of-the art algori thms. Qingwang Wang, Xiangrong Zhang, Yanfeng Gu |
IGARSS | 3 |
| 2020 | Deep feature extraction and motion representation for satellite video scene classification
Yanfeng Gu, Tengfei Wang 0001, Shengyang Li, Guoming Gao |
Sci. China Inf. Sci. | 1 |
| 2020 | Detection of Event of Interest for Satellite Video UnderstandingabstractSatellite videos provide rich dynamic information of observed scenes at a large spatial and temporal scale and will play an important role in the future space information network. This work devotes to revealing events of interest (EOI) from satellite video scenes by using a two-stream method. In satellite videos, individual frames reflect the static information like the basic scenes where the event was happening, while a sequence of frames determines the motion information. Considering these facts, a novel two-stream EOI detection framework is proposed, where one stream extracts static spatial information of satellite videos by AlexNet, whereas the other stream extracts the motion information using a local trajectories analysis method. First, the whole video scene is segmented into small spatial-temporal patches, where labeling EOI and non-EOI is completed. Next, the trajectories are extracted from 3-D satellite video cubes that are generated from event scene patches. Finally, this trajectory classification process is treated as a weak supervision learning problem and solved by sparse dictionary learning. The experimental results demonstrate that the proposed two-stream method is effective for EOI detection and has a huge potential for satellite video scenes analysis and understanding. The proposed method also outperforms the existing competitive models for video analysis. Yanfeng Gu, Tengfei Wang 0001, Xudong Jin, Guoming Gao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Hyperspectral Image Recovery Using Nonconvex Sparsity and Low-Rank RegularizationsabstractHyperspectral image (HSI) restoration is an important preprocessing step in HSI data analysis to improve the image quality for subsequent applications of HSI. In this article, we introduce a spatial-spectral patch-based nonconvex sparsity and low-rank regularization method for HSI restoration. In contrast to traditional approaches based on convex penalties or nonconvex spectral penalty alone, we consider the sparsity of HSI in the spatial-spectral domain and combine the nonconvex low-rank penalty and the nonconvex 3-D total variation (TV)-like sparsity regularization to fully exploit the correlations in both spatial-spectral dimensions of the HSI data set. In addition, we propose a fast iterative variable splitting-based algorithm to effectively solve the corresponding optimization problem. Numerical experiments on both simulated and real HSI data sets demonstrate that the proposed nonconvex low-rank and TV (NonLRTV) method significantly improves the recovered image quality compared with the state-of-the-art algorithms. Yue Hu 0003, Xiaodi Li 0003, Yanfeng Gu, Mathews Jacob |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Region-Enhanced Convolutional Neural Network for Object Detection in Remote Sensing ImagesabstractThe convolutional neural networks (CNNs) have recently demonstrated to be a powerful tool for object detection. However, with the complex scenes in remote sensing images, feature extraction of the object in the CNN will be seriously affected by background information. To address this issue, in this article, a region-enhanced CNN (RECNN) is proposed for the object detection of remote sensing images. The RECNN introduces the saliency constraint and multilayer fusion strategy into the CNN model, which can effectively enhance the object regions for better detection. Specifically, the saliency map is extracted and utilized to guide the training of the proposed model to strengthen saliency regions in feature maps. In addition, since different layers can reflect the object regions in varied resolutions, a multilayer fusion strategy is introduced to connect different convolutional layers and explore the context, where the feature maps of object regions are further enhanced. Experimental results on a publicly available ten-class object detection data set demonstrate the superiority of the RECNN over several competitive object detection methods. Jianjun Lei 0001, Leyuan Fang, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2020 | Satellite Video Super-Resolution Based on Adaptively Spatiotemporal Neighbors and Nonlocal Similarity RegularizationabstractRecently, super-resolution (SR) of satellite videos has received increasing attention as it can overcome the limitation of spatial resolution in applications of satellite videos to dynamic analysis. The low quality of satellite videos presents big challenges to the development of the spatial SR techniques, e.g., accurate motion estimation and motion compensation for multiframe SR. Therefore, reasonable image priors in maximum a posteriori (MAP) framework, where motion information among adjacent frames is involved, are needed to regularize the solution space and generate the corresponding high-resolution frames. In this article, an effective satellite video SR framework based on locally spatiotemporal neighbors and nonlocal similarity modeling is proposed. Firstly, local prior knowledge is represented by means of adaptively exploiting spatiotemporal neighbors. In this way, implicitly local motion information can be captured without explicit motion estimation. Secondly, the nonlocal spatial similarity is integrated into the proposed SR framework to enhance texture details. Finally, the locally spatiotemporal regularization and nonlocal similarity modeling bring out a complex optimization problem, which is solved via the iterated reweighted least squares in the proposed SR framework. The videos from the Jilin-1 satellite and the OVS-1A satellite are used for evaluating the proposed method. Experimental results show that the proposed method demonstrates better SR performance in preserving edges and texture details compared with the-state-of-art video SR methods. Yanfeng Gu, Tengfei Wang 0001, Shengyang Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | A Discriminative Tensor Representation Model for Feature Extraction and Classification of Multispectral LiDAR DataabstractMultispectral light detection and ranging (MS-LiDAR) systems open the door to the possibility in the 3-D land cover classification at a finer scale using only point cloud data. This article proposes a model based on the tensor representation for multispectral point cloud classification. The proposed method combines the 3-D local spatial structure of each multispectral point by characterizing the point with a second-order tensor. The first mode of the tensor indicates the spatial location and spectral information of each point (i.e., the row of the second-order tensor) and the second mode denotes the neighborhood geometric and spectral structures (i.e., the column of the second-order tensor). Then we develop a novel tensor manifold discriminant embedding (TMDE) algorithm to extract the geometric-spectral features for multispectral point clouds classification. TMDE solves the mapping matrices of each mode by preserving the intraclass samples' distribution further making it more compact and maximizing the distance of different classes. Finally, the support vector machine classifier with the extracted features as input is used to implement the classification of multispectral point clouds. Experiments are conducted on two real multispectral point cloud data sets. The experimental results demonstrate that the proposed method can achieve significant improvements in classification accuracies in comparison with several state-of-the-art algorithms. Qingwang Wang, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Very High Resolution Image Scene Classification with Capsule NetworkabstractConvolutional Neural Network (CNN) has boosted the performance of Very High Resolution (VHR) remote sensing data classification. Moreover, the continuous development of CNN techniques for image scenes description has entered a new challenge. The deep neural network models require a huge number of training samples, which is the main limitation of processing remote sensing data. To overcome this issue, a new method, based on the Capsule neural network for VHR image scenes recognition is proposed in this work. Experiments on the public Aerial Image Dataset (AID) benchmark, containing different areal categories with sub-meter spatial resolution are conducted. The obtained results demonstrate the effectiveness of the proposed method, as compared with the classical CNN model. Souleyman Chaib, Mohammed El Amin Larabi, Yanfeng Gu, Khadidja Bakhti, Moussa Sofiane Karoui |
IGARSS | 3 |
| 2019 | Unsupervised Temporal-Adaptation with Multiple Geodesic Flow Kernels for Hyperspectral Image ClassificationabstractThe miniaturization of hyperspectral sensors and the popularity of the unmanned aerial vehicle (UAV) make it possible to obtain a series of hyperspectral images (HSIs) in the same geographical area at different time-points by same or different sensors. When classifying these multi-temporal HSIs, temporal-adaptation is required to deal with the spectral drift and band inconsistency problems. Since most studies focus on semi-supervised domain adaptation (DA) strategy, and spatial features are usually absent during most of the DA procedure, an unsupervised temporal-adaptation method is realized by spatial-spectral multiple Geodesic Flow Kernels (S2-GFKs) to classify bi-temporal HSIs. Experiments conducted on two real HSI datasets and compared with several well-known methods demonstrate the availability of the proposed model. Tianzhu Liu, Yanfeng Gu |
IGARSS | 2 |
| 2019 | Unsupervised Multitemporal Domain Adaptation With Source Labels LearningabstractMultitemporal domain adaptation (DA) is very useful for solving the spectral drift problem between different images and is a basis step of multitemporal classification. However, for high-resolution images, they always have a few spectral bands. A few spectral bands are difficult to establish accurate alignment model. In order to achieving accurate multitemporal alignment on a few spectral bands' high-resolution images, source label learning step is proposed in this letter and used to optimize traditional manifold alignment (MA). The core of this method is to improve the erroneous manifold structure by combining majority voting and weighting coefficients. Besides, this method is a universal step and can be used for optimizing all MA methods. Two groups of data sets captured by Chinese GF1 and GF2 satellites are used for performance evaluation. The experimental results demonstrate the effectiveness of our method and indicate our method significantly outperforms the traditional DA methods. Baisen Liu, Guoming Gao, Yanfeng Gu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2019 | Tensorized Principal Component Alignment: A Unified Framework for Multimodal High-Resolution Images ClassificationabstractHigh-resolution (HR) remote sensing (RS) imaging opens the door to very accurate geometrical analysis for objects. However, it is difficult to simultaneous use massive HR RS images in practical applications, because these HR images are often collected in different multimodal conditions (multisource, multiarea, multitemporal, multiresolution, and multiangular) and learning method trained for one situation is difficult to use for others. The key problem is how to simultaneously tackle three main problems: spectral drift, spatial deformation, and band inconsistency. To deal with these problems, we propose an unsupervised tensorized principal component alignment framework in this paper. In this framework, local spatial-spectral patch data are used as basic units in order to achieve simultaneously multidimensional alignment. This framework seeks a domain-invariant tensor feature space by learning multilinear mapping functions which align the source tensor subspace with the target tensor subspace on different dimensions. In addition, an approach based on the Mahalanobis distance for dimensionality estimation of tensor subspace is proposed to determine best sizes of the aligned tensor subspace for reducing computational complexity. HR images from GF-1, GF-2, DEIMOS-2, WorldView-2, and WorldView-3 satellites are used to evaluate the performance. The experimental results show the following two points: first, the proposed alignment framework for multimodal HR images not only can align the different multimodal data more accurately than existing state-of-the-art domain adaptation methods, but also has a fast and simple procedure for large-scale data situation which is caused by HR imaging. Second, the proposed tensor dimensionality estimation method is an efficient technology for seeking the intrinsic dimensions of high-order data. Guoming Gao, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Superpixel Tensor Model for Spatial-Spectral Classification of Remote Sensing ImagesabstractNowadays, many methods of spatial-spectral classification have been developed and achieved good results for classification with high-resolution remotely sensed images, especially superpixel-based methods. However, these methods generally consider a superpixel as a group of pixels instead of one entity, ignoring the spectral-spatial entirety in the third-order RSI data cube. In order to fully exploit the third-order spectral-spatial information, in this paper, we propose a superpixel-based tensor model for RSI classification, where a multiattribute superpixel tensor (MAST) model is constructed on the top of multiattribute superpixel maps based on the concept of extended morphological profiles (EMAPs). In order to manage the adaptive spatial nature of superpixels, we develop an increment strategy to augment all superpixels with filling up their own envelop rectangles including three different ways, i.e., 0 vector, mean vector of all the pixels within the superpixel, or original pixels. Then, we use CANDECOMP/PARAFAC (CP) decomposition to obtain the features of the unified dimension from the MASTs of various sizes. Especially, CP decomposition can deal with missing data, so we also got a fourth means of constructing the MAST. Finally, base kernels calculated, respectively, from the original spectral feature, EMAP features and MAST features are learned by multiple kernel learning methods, with the optimal kernel fed to a support vector machine to complete the classification task. The experiments conducted on four real RSIs and compared with several well-known methods demonstrate the effectiveness of the proposed model. Yanfeng Gu, Tianzhu Liu, Jun Li 0009 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Intrinsic Image Recovery From Remote Sensing Hyperspectral ImagesabstractIn this paper, a novel reflectance model is proposed to recover intrinsic images from remote sensing hyperspectral images (HSIs). Intrinsic image recovery is a well-known challenging and underconstrained problem in computer vision, and it becomes even more severely illposed for HSIs. To reduce the uncertainties and improve the recovery accuracy, two kinds of priors are introduced: 1) shading prior which describes the geometric relation between illuminate and object surface and 2) reflectance prior based on L1-graph coding, which describes the relation between pigment density with reflectance. These priors can effectively eliminate the reflectance inhomogeneity caused by surface normal changes or pigment density variations other than material changes. Then, a noniterative optimization method is proposed to combine the shading prior and reflectance prior, with which closed-form solutions can be derived and thus avoided falling into local optimums. The experimental results demonstrate that the proposed method can efficiently improve the spectral reflectance homogeneity within a class while preserving the image boundaries; it also produces a competitive performance with the state of the art when utilizing the extracted intrinsic hyperspectral reflectance feature in the task of HSI classification. Xudong Jin, Yanfeng Gu, Tianzhu Liu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Unsupervised Cross-Temporal Classification of Hyperspectral Images With Multiple Geodesic Flow Kernel LearningabstractWith the increasing acquisition ability of hyperspectral remote sensing images, unsupervised cross-temporal classification (UCTC) of hyperspectral images (HSIs) has attracted more and more attention. In this paper, we focus on cross-temporal HSI classification, i.e., using one labeled HSI to classify the other unlabeled HSI. A multiple geodesic flow kernel learning (MGFKL) framework is proposed to exploit both spatial and spectral features for UCTC with bitemporal HSIs and called S2-MGFKL. The proposed S2-MGFKL method first extracts extended multi-attribute profiles (EMAPs) from the original bitemporal HSIs. The spatial features of the bitemporal HSIs obtained by the same attribute filter are paired up, so are the original spectral features. Second, each pair of features from both source and target domains are used to construct multiple geodesic flows. According to the original definition of GFK, we can obtain the construction of Gaussian base GFKs. The base kernels consist of two parts, the spectral part is obtained base on the same geodesic flow (which is constructed on the bitemporal spectral features) by tuning the kernel scale, while the spatial part is obtained under the same kernel scale but different geodesic flows constructed on different spatial feature pairs. After that, the mean rule is adopted to acquire the combined kernel, which is fed into the supervised vector machine (SVM) to implement the cross-temporal classification task. Experiments are conducted on two real HSI data sets, and the results compared with several well-known methods demonstrate the effectiveness of the proposed method. Tianzhu Liu, Xiangrong Zhang, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | A Fast Intra Coding Algorithm for Spatial Scalability in SHVCabstractScalable High Efficiency Video Coding (SHVC) provides high compression efficiency at the expense of considerable computational complexity. In this paper, a fast algorithm is proposed to reduce the computational complexity of SHVC intra coding. The coding depth information, texture complexity, and spatio-temporal correlation are jointly used to achieve a faster depth decision process. The mode dependency between the base layer and the enhancement layer is combined with the temporal correlation to simplify the mode decision process. Experimental results demonstrate that the proposed scheme saves encoding time by up to 59% compared with the SHM12.0 encoder with negligible degradation in Rate Distortion (RD). Xin Lu 0001, Yanfeng Gu, Graham R. Martin |
ICIP | 3 |
| 2018 | Very High Resolution Image Scene Classification with Semantic Fisher VectorsabstractVery high resolution (VHR) image scene classification is the most challenging of remote sensing data analysis, that has attracted researchers' attention. To improve the precision of VHR image scene classification, we propose a new method based on convolutional features extracted by convolutional neural network (CNN). First, Visual Geometry Group Network (VGG-Net) model is introduced as a feature extractor form the original VHR images. Second, we select the fifth convolutional layer constructed by VGG-Net, which is supposed as convolutional features descriptors. Third, based on Improved Fisher Vector (IFV) coding method, we compute the visual word corresponding to the convolutional features of the image scene. We conduct experiments on the public AID benchmark dataset, which contains 30 different areal categories with sub-meter resolution. Experimental results demonstrate the effectiveness of the proposed method, as compared with several state-of-the-art methods. Souleyman Chaib, Yanfeng Gu, Hongxun Yao, Khaled Belkadi |
IGARSS | 2 |
| 2018 | Combine Reflectance with Shading Component for Hyperspectral Image ClassificationabstractIntrinsic image decomposition (IID) of hyperspectral images (HSIs) aims to separate the reflectance cube and shading component from the original image data. The reflectance cube contains the spectral information reflecting the intrinsic properties of the material, whereas the shading component contains the spatial information reflecting geometric structure of the object like the surface orientation changes. From the perspective of hyperspectral image classification, combining spectral information with spatial information can be useful for improving the classification performance. In this paper, a new optimization algorithm is proposed for intrinsic image decomposition of hyperspectral images, and composite kernel learning (CKL) method is further utilized to combine reflectance with shading component. Xudong Jin, Yanfeng Gu |
IGARSS | 2 |
| 2018 | Multi-Attribute Super-Tensor Model for Remote Sensing Image Classification with High Spatial ResolutionabstractWith the development of remote sensors, it is much easier to acquire large amount of remote sensing images (RSIs) with very high spatial resolution, which has made the spatial characteristics play an important role in classification task. Many work of spatial-spectral classification have been done and achieved good results, especially superpixel-based methods. However, these methods didn't take each superpixel as an entirety, which had ignored the relationship between spatial and spectral signature. It is well known that RSI can be treated as a third-order data cube, thus it can also be represented by a third-order tensor. This paper proposed a Multi-Attribute Superpixel Tensor (MAST) model to address the aforementioned problem. Experiments conducted on two real RSIs and compared with several well-known methods demonstrate the effectiveness of the proposed model. Tianzhu Liu, Yanfeng Gu |
IGARSS | 2 |
| 2018 | Cnn Based Renormalization Method for Ship Detection in Vhr Remote Sensing ImagesabstractShip detection with very high resolution (VHR) remote sensing image has recently been an attractive topic due to rapid development of deep learning. Current researches on ship detection are generally confronted with a big challenge that existing methods failed to get high quality of object proposal with good intersection-over-union (IOU) before detection. In this paper, a Convolutional Neural Network (CNN) based renormalization method is proposed to improve the quality of object proposal. First, CNN is used to predict shape information of candidate ships' which are involved with rotation, location and scale in patches. Then, a renormalization net is designed to adjust the candidate ships in patches by correcting the shape information and renormalizing it to uniform patch. In this way, good candidate objects in patches could be generated and will be helpful with improving following ship detection. The proposed renormalization net was tested on a Google-Earth handcraft dataset. The experimental result demonstrates the proposed renormalization net greatly improve the ship detection with both of good detection accuracy and high IOU. Tengfei Wang 0001, Yanfeng Gu |
IGARSS | 2 |
| 2018 | Robust cost function for optimizing chamfer masks
Baraka Jacob Maiseli, Lifei Bai, Xianqiang Yang 0001, Yanfeng Gu, Huijun Gao |
Vis. Comput. | 4 |
| 2017 | Multi-temporal images classification with evidential fusion of manifold alignmentabstractMulti-temporal remote sensing images classification have attracted more and more attention in the last decade because of a wide range of applications of multi-temporal images in long-term environmental monitoring and land cover change detection and increasing multi-temporal data available. At present, most papers investigated two temporal remote-sensing images classification. In fact, there is lots of distinctive information to be unexploited between two or more temporal images which can enhance classification effect and improve ability of detecting change area. In this paper, we present an evidential fusion framework of manifold alignment to combine more than two multi-temporal remote sensing images. Embedding of multi-groups two temporal images pairs after MA can be intergraded based a layered structure of D-S theory. The proposed method was evaluated using five Landsat 8 images. Results confirmed that the proposed algorithm performed better than those with only two temporal images. Tianzhu Liu, Guoming Gao, Yanfeng Gu |
IGARSS | 4 |
| 2017 | Recent developments and trends in point set registration methods
Baraka Jacob Maiseli, Yanfeng Gu, Huijun Gao |
J. Vis. Commun. Image Represent. | 2 |
| 2017 | Deep Fusion of Remote Sensing Data for Accurate ClassificationabstractThe multisensory fusion of remote sensing data has obtained a great attention in recent years. In this letter, we propose a new feature fusion framework based on deep neural networks (DNNs). The proposed framework employs deep convolutional neural networks (CNNs) to effectively extract features of multi-/hyperspectral and light detection and ranging data. Then, a fully connected DNN is designed to fuse the heterogeneous features obtained by the previous CNNs. Through the aforementioned deep networks, one can extract the discriminant and invariant features of remote sensing data, which are useful for further processing. At last, logistic regression is used to produce the final classification results. Dropout and batch normalization strategies are adopted in the deep fusion framework to further improve classification accuracy. The obtained results reveal that the proposed deep fusion model provides competitive results in terms of classification accuracy. Furthermore, the proposed deep learning idea opens a new window for future remote sensing data fusion. Yushi Chen 0002, Pedram Ghamisi, Xiuping Jia, Yanfeng Gu |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2017 | Discriminative Graph-Based Fusion of HSI and LiDAR Data for Urban Area ClassificationabstractA novel discriminative graph-based fusion (DGF) method is proposed for urban area classification to fuse heterogeneous features from two data sources, i.e., hyperspectral image (HSI) and light detecting and ranging (LiDAR) data. The features include spectral characteristics in HSI, height in LiDAR data, and geometry in image processing technologies like morphological profiles (MPs). Our proposed DGF method couples dimension reduction and heterogeneous feature fusion. The core idea of the proposed method is to search for a projection matrix by minimizing the similarity term that preserves the local geometry of each class and maximizing the dissimilarity term that contains the relation of between-class distance. As a result, the proposed method can pull close together samples of the same class while pushing those of different classes apart in the projected space by fusing graphs constructed by different groups of heterogeneous features. The edges of the graphs are measured by kernel. Furthermore, the multiscale DGF (MS-DGF) is introduced to utilize the capability of similarity measure of different scales of kernel and avoid finding the optimal scale simultaneously. Experiments are conducted on real HSI along with LiDAR data. The corresponding results demonstrate that the proposed method can make an effective fusion of heterogeneous features to make full use of the complementary information of HSI and LiDAR, which facilitates fine classification task of urban area, compared with several state-of-the-art algorithms. Yanfeng Gu, Qingwang Wang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | A hierarchical energy minimization method for building roof segmentation from airborne LiDAR data
Yanfeng Gu, Zhimin Cao, Limin Dong |
Multim. Tools Appl. | 1 |
| 2017 | Deep Feature Fusion for VHR Remote Sensing Scene ClassificationabstractThe rapid development of remote sensing technology allows us to get images with high and very high resolution (VHR). VHR imagery scene classification has become an important and challenging problem. In this paper, we introduce a framework for VHR scene understanding. First, the pretrained visual geometry group network (VGG-Net) model is proposed as deep feature extractors to extract informative features from the original VHR images. Second, we select the fully connected layers constructed by VGG-Net in which each layer is regarded as separated feature descriptors. And then we combine between them to construct final representation of the VHR image scenes. Third, discriminant correlation analysis (DCA) is adopted as feature fusion strategy to further refine the original features extracting from VGG-Net, which allows a more efficient fusion approach with small cost than the traditional feature fusion strategies. We apply our approach to three challenging data sets: 1) UC MERCED data set that contains 21 different areal scene categories with submeter resolution; 2) WHU-RS data set that contains 19 challenging scene categories with various resolutions; and 3) the Aerial Image data set that has a number of 10 000 images within 30 challenging scene categories with various resolutions. The experimental results demonstrate that our proposed method outperforms the state-of-the-art approaches. Using feature fusion technique achieves a higher accuracy than solely using the raw deep features. Moreover, the proposed method based on DCA fusion produces good informative features to describe the images scene with much lower dimension. Souleyman Chaib, Yanfeng Gu, Hongxun Yao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2017 | Supervised Multiview Feature Selection Exploring Homogeneity and Heterogeneity With ℓ1, 2-Norm and Automatic View GenerationabstractIt is useful and challenging to analyze and select object features of very high resolution (VHR) remote sensing imagery. The overwhelming majority of existing feature selection methods always concatenate all of the features into a long feature vector and then select features from the vector, ignoring the homogeneity and heterogeneity of underlying feature subspaces. In this paper, we propose a supervised multiview feature selection (SMFS) method. Unlike the existing multiview methods, SMFS requires no prior knowledge of the number of views, and is independent of a prefixed classifier. By utilizing homogeneity and heterogeneity of the data, SMFS employs affinity propagation to automatically decompose features into multiple disjoint and meaningful feature groups or views without any prior knowledge. A group or view consists of homogeneous features, describing a unique data characteristic. Different views represent heterogeneous data characteristics. Then, features are evaluated and selected based on joint ℓ1,2-norm minimization of a loss function and a regularization term. Different from the popular ℓ2,1-norm, joint ℓ1,2-norm enforces the intraview sparsity, instead of interview sparsity. Consequently, a view can be represented by a few representative features in each view, and the information of heterogeneous views can be well kept by the remaining representative features. The experimental results on four VHR satellite images attest to the effectiveness and practicability of SMFS in comparison with single-view algorithms. Furthermore, some discussions are conducted to give insights into homogeneity and heterogeneity of features. Xi Chen 0004, Gongjian Zhou, Yushi Chen 0002, Guofan Shao, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2017 | Multitemporal Landsat Missing Data Recovery Based on Tempo-Spectral Angle ModelabstractMultitemporal Landsat images play an important role in remote sensing applications. Unfortunately, missing data caused by cloud cover and sensor-specific problems have seriously limited its application. To improve the usability of Landsat data, several recovery methods have been proposed to fill the missing values. But, current studies mostly focus on spatial dimension and ignore the continuity of data in time dimension. More importantly, multitemporal images have more potential than single image in selecting similar pixels for recovering the missing pixels. In this paper, to recover missing pixels by jointly utilizing multispectral and multitemporal information, tempo-spectral angle mapping (TSAM) is proposed at first to measure tempo-spectral similarity between pixels described in spectral dimension and temporal dimension. Then, a multitemporal replacement method is used to recover missing data with the pixel selected by TSAM. Two new indices are also proposed to evaluate the effectiveness of TSAM. Simulated and actual multitemporal scan-line corrector-off and cloud cover-Enhanced Thematic Mapper Plus images were used to assess the performance of our filling method. The quantitative evaluations suggest that the proposed method can predict the missing values accurately. The recovered results show that our method can keep the continuity of the boundary and is robust for the data with high percentage of missing. Guoming Gao, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Multiple Kernel Learning for Hyperspectral Image Classification: A ReviewabstractWith the rapid development of spectral imaging techniques, classification of hyperspectral images (HSIs) has attracted great attention in various applications such as land survey and resource monitoring in the field of remote sensing. A key challenge in HSI classification is how to explore effective approaches to fully use the spatial-spectral information provided by the data cube. Multiple kernel learning (MKL) has been successfully applied to HSI classification due to its capacity to handle heterogeneous fusion of both spectral and spatial features. This approach can generate an adaptive kernel as an optimally weighted sum of a few fixed kernels to model a nonlinear data structure. In this way, the difficulty of kernel selection and the limitation of a fixed kernel can be alleviated. Various MKL algorithms have been developed in recent years, such as the general MKL, the subspace MKL, the nonlinear MKL, the sparse MKL, and the ensemble MKL. The goal of this paper is to provide a systematic review of MKL methods, which have been applied to HSI classification. We also analyze and evaluate different MKL algorithms and their respective characteristics in different cases of HSI classification cases. Finally, we discuss the future direction and trends of research in this area. Yanfeng Gu, Jocelyn Chanussot, Xiuping Jia, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | Multiple Kernel Sparse Representation for Airborne LiDAR Data ClassificationabstractTo effectively learn heterogeneous features extracted from raw LiDAR point cloud data for landcover classification, a multiple kernel sparse representation classification (MKSRC) framework is proposed in this paper. In the MKSRC, multiple kernel learning (MKL) is embedded into sparse representation classification (SRC). The heterogeneous features are first extracted from the raw LiDAR point cloud data before classification. These features contain useful information from different dimensions, including single point features and neighbor features. Based on feature extraction, on the one hand, MKL is reasonably integrated into the SRC, namely, different base kernels that are constructed with each heterogeneous feature separately are utilized in the process of sparse representation. Furthermore, joint sparsity model is also introduced into the MKSRC framework and multiple kernel joint SRC (MKJSRC) is then proposed. On the other hand, improved kernel alignment (IKA) methods are proposed to more effectively determine the weights of base kernels in both of MKSRC and MKJSRC. Experiments are conducted on three real airborne LiDAR data sets. The experimental results demonstrate that MKSRC and MKJSRC frameworks can effectively learn the heterogeneous features for LiDAR point cloud classification and outperforms the other state-of-the-art sparse representation-based classifiers and the recent MKL algorithm. Moreover, the proposed IKA is helpful to better determine the “optimal” weights of the base kernels in both MKSRC and MKJSRC than in the existing kernel alignment method. Yanfeng Gu, Qingwang Wang, Bingqian Xie |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2017 | Superpixel-Based Intrinsic Image Decomposition of Hyperspectral ImagesabstractIn this paper, we propose a novel superpixel-based intrinsic image decomposition (SIID) framework for hyperspectral images. Intrinsic images are usually referred to the separation of shading and reflectance components from an input image. Considering the high dimensionality of hyperspectral images, we further decompose the shading component into the product of environment illumination and surface orientation changes, thus modeling the problem more properly. The proposed method consists of the following steps. First, we build two superpixel segmentation maps of different scales, i.e., a finer one that is oversegmented and a coarser one that is undersegmented. Based on the observation that the finer superpixel map achieves a higher segmentation accuracy, whereas the coarser superpixel map tends to reserve the objectness of the original image, we model the SIID decomposition problem in a matrix form based on the finer superpixel map and define a constraint matrix by integrating the information in the coarser superpixel map. The constraint matrix is introduced as a secondary constraint in order to make the ill-posed IID problem solvable. Finally, we transform the original decomposition problem into minimizing the Frobenius norm of the proposed matrix energy function and iteratively derive the solution. Our experimental results demonstrate that the proposed method is able to achieve a performance outperforming the state-of-the-art while making a great improvement in efficiency. Xudong Jin, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Multimorphological Superpixel Model for Hyperspectral Image ClassificationabstractWith the development of hyperspectral sensors, nowadays, we can easily acquire large amount of hyperspectral images (HSIs) with very high spatial resolution, which has led to a better identification of relatively small structures. Owing to the high spatial resolution, there are much less mixed pixels in the HSIs, and the boundaries between these categories are much clearer. However, the high spatial resolution also leads to complex and fine geometrical structures and high inner-class variability, which make the classification results very “noisy.” In this paper, we propose a multimorphological superpixel (MMSP) method to extract the spectral and spatial features and address the aforementioned problems. To reduce the difference within the same class and obtain multilevel spatial information, morphological features (multistructuring element extended morphological profile or multiattribute filter extended multi-attribute profiles) are first obtained from the original HSI. After that, simple linear iterative clustering segmentation method is performed on each morphological feature to acquire the MMSPs. Then, uniformity constraint is used to merge the MMSPs belonging to the same class which can avoid introducing the information from different classes and acquire spatial structures at object level. Subsequently, mean filtering is utilized to extract the spatial features within and among MMSPs. At last, base kernels are obtained from the spatial features and original HSI, and several multiple kernel learning methods are used to obtain the optimal kernel to incorporate into the support vector machine. Experiments conducted on three widely used real HSIs and compared with several well-known methods demonstrate the effectiveness of the proposed model. Tianzhu Liu, Yanfeng Gu, Jocelyn Chanussot, Mauro Dalla Mura |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Tensor Matched Subspace Detector for Hyperspectral Target DetectionabstractIn this paper, a new framework for tensor hyperspectral target detection is proposed. In this new framework, tensor is well integrated into the conventional target detection algorithm. As a result, a tensor matched subspace detector (MSD) for hyperspectral target detection is proposed. The proposed method is mainly applied to detect multipixel targets rather than subpixel targets. In this new method, the hyperspectral data are considered as a form of third-order tensor in order to jointly utilize the information of multidimensional data. In conventional detection methods, the spatial-spectral information has not been taken into account, even some algorithms have been presented for improving the utilization efficiency of the spatial-spectral structural feature, but the overall structural characteristic of the extracted feature is still ignored. In our algorithm, the tensor subspace projection is defined for the first time, which is easily calculated by three predetermined orthogonal direction mapping matrices without any iteration. Then, the test tensor blocks are projected into the tensor subspace and finally measured by the ratio of residual energy, just like the general likelihood ratio test. The proposed method can be regarded as an extension of conventional MSD. The reliability and superiority are demonstrated by the experiments on real hyperspectral imaging data sets. The experimental results indicate that our approach compares favorably to some classical and novel methods by jointly processing multidimensional data with tensorial form. Yongjian Liu, Guoming Gao, Yanfeng Gu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2016 | A VHR scene classification method integrating sparse PCA and saliency computingabstractUnderstanding a scene provided by very high resolution (VHR) satellite imagery has become a more and more challenging problem. In this paper, we propose a new method for scene classification based on saliency computing of patches sampling from the VHR images. Sparse principal component analysis (sPCA) is then adopted to select the corresponding informative salient patches for image scene representation. The proposed technique for selecting informative salient patches is efficient and robust for scene understanding. We conduct experiments on the public UC Merced benchmark dataset, which contains 21 different areal categories with sub-meter resolution. Experimental results demonstrate the effectiveness of the proposed method, as compared with several state-of-the-art methods. Souleyman Chaib, Yanfeng Gu, Hongxun Yao, Sicheng Zhao |
IGARSS | 2 |
| 2016 | Deep fusion of hyperspectral and LiDAR data for thematic classificationabstractRecently, the fusion of hyperspectral and light detection and ranging (LiDAR) data has obtained a great attention in the remote sensing community. In this paper, we propose a new feature fusion framework using deep neural network (DNN). The proposed framework employs a novel 3D convolutional neural network (CNN) to extract the spectral-spatial features of hyperspectral data, a deep 2D CNN to extract the elevation features of LiDAR data, and then a fully connected deep neural network to fuse the extracted features in the previous CNNs. Through the aforementioned three deep networks, one can extract the discriminant and invariant features of hyperspectral and LiDAR data. At last, logistic regression is used to produce the final classification results. The experimental results reveal that the proposed deep fusion model provides competitive results. Furthermore, the proposed deep fusion idea opens a new window for future research. Pedram Ghamisi, Chunyu Shi, Yanfeng Gu |
IGARSS | 5 |
| 2016 | Improved neighborhood similar pixel interpolator for filling unsacn multi-temporal Landsat ETM+ data without referenceabstractSince the scan line corrector (SLC) of the Landsat ETM+ sensor failed permanently in 2003, about 22% of the pixels in an SLC-off image are missed. Traditional gap filling methods always need a SLC-on image for reference, but the most similar sensor (Landsat TM) closed at 2011. And the potential of multi-temporal was also neglected in traditional filling methods. In this paper, a multi-temporal Landsat ETM+ gap filling method is proposed without using SLC-on reference which has ability to increase the utilization efficiency of multi-temporal images. The proposed method is mainly based on neighborhood similar pixel interpolator (NSPI) and the major contribution are find an effective way to select valid temporal and conjunctive use the temporal advantage to calculate of the target pixel value. Similarity both in spatial and temporal can be obtained in our method. Real multi-temporal Landsat data and missing gap location are used to assess the efficacy of the proposed method. Both qualitative and quantitative evaluations results suggest that our proposed method can predict the missing values very accurately and improve the utilization efficiency of multi-temporal. Guoming Gao, Tianzhu Liu, Yanfeng Gu |
IGARSS | 3 |
| 2016 | LiDAR point classification based on joint sparse representation in kernel spaceabstractResent years, sparse representation theory has been widely used in signal processing field. Researchers introduce this theory into the application of pattern recognition and classification and get the sparse representation classifier (SRC). In this paper, we use the SRC to achieve the classification of LiDAR (Light Detection and Ranging) points. To get a better performance, we introduce the kernel method into SRC, for the advancement of kernel in solving nonlinear problem. Also, a joint sparse representation is used for the category similarity of neighboring LiDAR points. Bingqian Xie, Yanfeng Gu, Qingwang Wang |
IGARSS | 2 |
| 2016 | Sample-screening MKL method via boosting strategy for hyperspectral image classification
Yanfeng Gu |
Neurocomputing | 1 |
| 2016 | An Informative Feature Selection Method Based on Sparse PCA for VHR Scene ClassificationabstractUnderstanding the scenes provided by very high resolution satellite (VHR) imagery has become a critical task. In this letter, we propose a new informative feature selection method for VHR scene classification. First, scale-invariant feature transform and speeded up robust feature operators are used to extract local features from the original VHR images to construct a visual dictionary. A sparse principal component analysis (sPCA) is then adopted to learn a set of informative features from the visual dictionary for each category. Finally, the scenes are represented by sparse informative low-level features. We conducted experiments on the University of California at Merced data set containing 21 different areal scene categories with submeter resolution and the Sydney data set containing seven land-use categories with 0.5-m spatial resolution. The experimental results demonstrate that the proposed method outperforms the state-of-the-art methods even without saliency detection. Souleyman Chaib, Yanfeng Gu, Hongxun Yao |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | Nonlinear Multiple Kernel Learning With Multiple-Structure-Element Extended Morphological Profiles for Hyperspectral Image ClassificationabstractIn this paper, we propose a novel multiple kernel learning (MKL) framework to incorporate both spectral and spatial features for hyperspectral image classification, which is called multiple-structure-element nonlinear MKL (MultiSE-NMKL). In the proposed framework, multiple structure elements (MultiSEs) are employed to generate extended morphological profiles (EMPs) to present spatial-spectral information. In order to better mine interscale and interstructure similarity among EMPs, a nonlinear MKL (NMKL) is introduced to learn an optimal combined kernel from the predefined linear base kernels. We integrate this NMKL with support vector machines (SVMs) and reduce the min-max problem to a simple minimization problem. The optimal weight for each kernel matrix is then solved by a projection-based gradient descent algorithm. The advantages of using nonlinear combination of base kernels and multiSE-based EMP are that similarity information generated from the nonlinear interaction of different kernels is fully exploited, and the discriminability of the classes of interest is deeply enhanced. Experiments are conducted on three real hyperspectral data sets. The experimental results show that the proposed method achieves better performance for hyperspectral image classification, compared with several state-of-the-art algorithms. The MultiSE EMPs can provide much higher classification accuracy than using a single-SE EMP. Yanfeng Gu, Tianzhu Liu, Xiuping Jia, Jón Atli Benediktsson, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | Class-Specific Sparse Multiple Kernel Learning for Spectral-Spatial Hyperspectral Image ClassificationabstractIn recent years, many studies on hyperspectral image classification have shown that using multiple features can effectively improve the classification accuracy. As a very powerful means of learning, multiple kernel learning (MKL) can conveniently be embedded in a variety of characteristics. This paper proposes a class-specific sparse MKL (CS-SMKL) framework to improve the capability of hyperspectral image classification. In terms of the features, extended multiattribute profiles are adopted because it can effectively represent the spatial and spectral information of hyperspectral images. CS-SMKL classifies the hyperspectral images, simultaneously learns class-specific significant features, and selects class-specific weights. Using an $L_{1}$-norm constraint (i.e., group lasso) as the regularizer, we can enforce the sparsity at the group/feature level and automatically learn a compact feature set for the classification of any two classes. More precisely, our CS-SMKL determines the associated weights of optimal base kernels for any two classes and results in improved classification performances. The advantage of the proposed method is that only the features useful for the classification of any two classes can be retained, which leads to greatly enhanced discriminability. Experiments are conducted on three hyperspectral data sets. The experimental results show that the proposed method achieves better performances for hyperspectral image classification compared with several state-of-the-art algorithms, and the results confirm the capability of the method in selecting the useful features. Tianzhu Liu, Yanfeng Gu, Xiuping Jia, Jón Atli Benediktsson, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2016 | Discriminative Multiple Kernel Learning for Hyperspectral Image ClassificationabstractIn this paper, we propose a discriminative multiple kernel learning (DMKL) method for spectral image classification. The core idea of the proposed method is to learn an optimal combined kernel from predefined basic kernels by maximizing separability in reproduction kernel Hilbert space. DMKL achieves the maximum separability via finding an optimal projective direction according to statistical significance, which leads to the minimum within-class scatter and maximum between-class scatter instead of a time-consuming search for the optimal kernel combination. Fisher criterion (FC) and maximum margin criterion (MMC) are used to find the optimal projective direction, thus leading to two variants of the proposed method, DMKL-FC and DMKL-MMC, respectively. After learning the projective direction, all basic kernels are projected to generate a discriminative combined kernel. Three merits are realized by DMKL. First, DMKL can achieve a substantial improvement in classification performance without strict limitation for selection of basic kernels. Second, the discriminating scales of a Gaussian kernel, the useful bands for classification, and the competitive sizes of spatial filters can be selected by ranking the corresponding weights, where the large weights correspond to the most relevant. Third, DMKL reduces the computational burden by requiring fewer support vectors. Experiments are conducted on two hyperspectral data sets and one multispectral data set. The corresponding experimental results demonstrate that the proposed algorithms can achieve the best performance with satisfactory computational efficiency for spectral image classification, compared with several state-of-the-art algorithms. Qingwang Wang, Yanfeng Gu, Devis Tuia |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2015 | L2, 0-norm regularization based feature selection for very high resolution remote sensing imagesabstractThis paper presents a ℓ2,0-norm regularization based feature selection method to analyze very high resolution remote sensing imagery. The method tackles the feature selection problem based on a ℓ2,1-norm based objective function and a ℓ2, 0-norm equality constraint. The constrained optimization problem is solved by an efficient algorithm based on augmented Lagrangian method to figure out a stable local solution. Though the ℓ2, 0-norm regularization based feature selection method should handle a non-convex and non-smooth problem, it outperforms the ℓ2,1-norm regularization based approximate convex counterparts and state-of-art feature selection methods in light of classification accuracies by 1-NN and SVM classifiers. The experimental results demonstrate the effectiveness of the presented method in selecting features with great generalization capabilities. Xi Chen 0004, Yanfeng Gu, Ye Zhang 0008 |
IGARSS | 2 |
| 2015 | A novel multiple kernel boosting method for hyperspectral image classificationabstractMultiple kernel learning (MKL) combines multiple base kernels and is becoming more and more popular in machine learning. The choice of kernels is crucial importance for classification performance. In this paper, we propose a new RMKL (RMKBoost) framework for classification in hyperspectral images. The classification is performed in separate two steps. The key boosting strategy is embedded in the first step, which aims to learn an optimally or suboptimally linear combined kernel from the predefined base kernels. Then, the proposed boosting framework generates weak multiple kernel classifiers using a part of the base kernels randomly selected rather than using all base kernels with randomly training samples. Experiments are conducted on the real hyperspectral data set, and the corresponding experimental result shows that RMKBoost algorithm provides the best performances compared with the state-of-the-art kernel methods. Tianzhu Liu, Yanfeng Gu |
IGARSS | 3 |
| 2015 | Building LiDAR point cloud denoising processing through sparse representationabstractNowdays, airborne LiDAR comes into a popular way to survey the ground scene, particularly for the application of building reconstruction. However, the LiDAR point cloud acquired is usually polluted by noise for the existence of LiDAR system's inherent error and aircraft's shock. Thus, before LiDAR data is used, a preprocessing such as denoising is needed. This paper focus on the denoising of building LiDAR data. First, the building LiDAR point cloud is rasterized into a two- dimensional image. Then, a dictionary learned from training samples is used to denoise the image according to signal's sparse representation theory. Last, we can get the building's raster image with little noise. Bingqian Xie, Yanfeng Gu, Zhimin Cao |
IGARSS | 2 |
| 2015 | Hyperspectral target detection via exploiting spatial-spectral joint sparsity
Yanfeng Gu, He Zheng, Yue Hu 0003 |
Neurocomputing | 1 |
| 2015 | Class-Specific Feature Selection With Local Geometric Structure and Discriminative Information Based on Sparse Similar SamplesabstractIt is necessary while quite challenging to select features strongly relevant to a thematic class, i.e., class-specific features, from very high resolution (VHR) remote sensing images. To meet this challenge, a class-specific feature selection method based on sparse similar samples (CFS4) is proposed. Specifically, CFS4 incorporates the local geometrical structure and discriminative information of the data into a sparsity regularization problem. The experimental results on VHR satellite images well validate the effectiveness and practicability of the proposed method. Xi Chen 0004, Yanfeng Gu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2015 | A Novel MKL Model of Integrating LiDAR Data and MSI for Urban Area ClassificationabstractA novel multiple-kernel learning (MKL) model is proposed for urban classification to integrate heterogeneous features (HF-MKL) from two data sources, i.e., spectral images and LiDAR data. The features include spectral, spatial, and elevation attributes of urban objects from the two data sources. With these heterogeneous features (HFs), the new MKL model is designed to carry out feature fusion that is embedded in classification. First, Gaussian kernels with different bandwidths are used to measure the similarity of samples on each feature at different scales. Then, these multiscale kernels with different features are integrated using a linear combination. In the combination, the weights of the kernels with different features are determined by finding a projection based on the maximum variance. This way, the discriminative ability of the HFs is exploited at different scales and is also integrated to generate an optimal combined kernel. Finally, the optimization of the conventional support vector machine with this kernel is performed to construct a more effective classifier. Experiments are conducted on two real data sets, and the experimental results show that the HF-MKL model achieves the best performance in terms of classification accuracies in integrating the HFs for classification when compared with several state-of-the-art algorithms. Yanfeng Gu, Qingwang Wang, Xiuping Jia, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | Linear discriminant multiple kernel learning for multispectral image classificationabstractIn the past decade, with the development of kernel-based machine learning, many different multiple kernel learning (MKL) methods were proposed which focus on selecting the pivotal kernel to be preserved and confirming the optimal kernel combination. In this paper, we address the question mentioned above by using subspace projection method and put forward a linear discriminant based MKL (LDMKL) algorithm. LDMKL algorithm reduces the computational burden and keeps the excellent property of MKL in terms of good classification accuracy by finding the optimal projective direction which makes the intraclass scatter minimum and interclass scatter maximum instead of the time-consuming search for optimal kernel combination. Experimental results indicate that LDMKL algorithm provides the best performances among several the state-of-the-art algorithms while demonstrating satisfactory computational efficiency. Yanfeng Gu, Qingwang Wang, Pigang Liu, Deshan Zuo |
ICIP | 1 |
| 2014 | Hyperspectral image classification with multiple kernel Boosting algorithmabstractMultiple kernel learning (MKL) is becoming more and more popular in machine learning. Traditional MKL methods usually learn the optimal combinations of both kernels and classifiers as the optimization task which is difficult to be solved. In this paper, we study a Boosting framework of MKL for classification in hyperspectral images. The multiple kernel Boosting (MKBoost) is proposed to solve the MKL problem, which apply the idea of Boosting to the multiple kernel classifiers based on the SVM. Experiments are conducted on different real hyperspectral data sets, and the corresponding experimental results show that MKBoost algorithm provides the best performances compared with the state-of-the-art kernel methods. Yanfeng Gu, Guoming Gao, Qingwang Wang |
ICIP | 2 |
| 2014 | Mapping of cloud thickness with MODIS and CloudSat data through multiple kernel learningabstractIn this paper, we present an efficient approach based on multiple kernel learning (MKL) for mapping cloud thickness with MODIS and CloudSat data. In order to adapt the characteristics of radar data, we generalize a signal model from the gas imaging model, and the signal model provides a way for transforming the mapping of cloud thickness into a linear estimate problem. Then, considering the disadvantage of complexity and nonlinearity of the MODIS data, the MKL method which has been shown to improve the performance of many learning tasks is qualified for the mapping of cloud thickness. The MODIS data in a real scenarios is used to test the performance of the develop method and the experimental results indicate that the proposed MKL method outperforms single kernel method for the research of mapping cloud thickness. Yanfeng Gu, Pigang Liu, Qingwang Wang, Shizhe Wang |
IGARSS | 1 |
| 2014 | Orientation estimation of building using DSM and optical images based on Zernike momentsabstractDigital Surface Models (DSMs) generated from airborne laser-scanning or stereo satellite images provide a very useful source of information for building orientation estimation. However, due to the high complexity of the building structure, it is hardly to get the details of the building boundary information. In this paper, a new orientation estimation method is proposed, which merge the edge information extracted from optical images into DSM, and take advantage of the ZM phase information in the orientation estimation process. The classical way of extracting Zernike features only takes into account the magnitude of the moments and loses the phase information. The novelty of our approach is to take advantage of the phase information of Zernike moments to capture these variabilities in a way that makes it robust in the context of rotation angle estimation. Shu Tian, Ye Zhang 0008, Yanfeng Gu |
IGARSS | 4 |
| 2014 | Ultrasound echocardiography despeckling with non-local means time series filter
Yanfeng Gu, Zhaoyu Cui, Chunhong Xiu, Lanfeng Wang |
Neurocomputing | 1 |
| 2014 | Three-Dimensional Reconstruction of Multiplatform Stereo Data With Variance Component EstimationabstractIn this paper, we address a problem of 3-D reconstruction with generalized stereo data from multiple platforms of remote sensing. Nowadays, rational function model (RFM)-based 3-D reconstruction with stereo images obtained from a single platform of remote sensing like a satellite or an airborne platform has been widely investigated, but there are little attentions to be paid to the problem of 3-D reconstruction with stereo images from multiple platforms in the existing literature. In order to make full use of the generalized stereo images from different platforms with different rigorous sensor models for 3-D reconstruction, we need to form the least squares estimation model of the corresponding RFM-based forward-intersection task after collecting observations from different platforms. However, resolutions of the stereo images from different platforms are greatly different so that the observations in the corresponding least squares problem are mathematically seriously unbalanced. To solve this problem for achieving precise reconstruction, we first model how the spatial resolution of the observation images of different platforms changes pixel by pixel and then embed the variance-component-estimation technique into the RFM-based 3-D reconstruction procedure to adaptively adjust weights for different observations. Experiments are conducted on simulated and real data sets. Experimental results show that the proposed algorithm can efficiently fulfill the 3-D reconstruction task for multiplatform stereo images with noticeable improvement over the classical RFM-based 3-D reconstruction method in terms of precision. Yanfeng Gu, Zhimin Cao, Ye Zhang 0008 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2013 | A novel model for building information acquisition optimization technology of remote sensing observationabstractIt is an important problem in remote sensing that using limited observing points acquire the maximum quantity of building information. In this paper, a building information acquisition (BIA) model based on Support Vector Machine (SVM) is proposed for quantitative description of the mathematical relationship between the information quantity acquisition and the observing angles, which is optimized to obtain the maximum information quantity in the multi-temporal remote sensing observation. The main idea of the BIA model is that, to calculate information quantity at different observing angles, the target is decomposed into multiple faces whose information is described by the combined vector. Further, the modified bee colony algorithm is utilized to optimize the model to achieve the ideal maximum information quantity. The corresponding combined vector is optimal observing angles combination. The proposed model method performs well in our imaging simulation system data. Experiment results demonstrate that the proposed BIA model optimized will provide much more information quantity than observing randomly. Nan Su 0001, Ye Zhang 0008, Yanfeng Gu |
IGARSS | 4 |
| 2013 | Rare signal component extraction based on kernel methods for anomaly detection in hyperspectral imagery
Yanfeng Gu, Lin Zhang 0011 |
Neurocomputing | 1 |
| 2013 | Spectral Unmixing in Multiple-Kernel Hilbert Space for Hyperspectral ImageryabstractIn this paper, we address a spectral unmixing problem for hyperspectral images by introducing multiple-kernel learning (MKL) coupled with support vector machines. To effectively solve issues of spectral unmixing, an MKL method is explored to build new boundaries and distances between classes in multiple-kernel Hilbert space (MKHS). Integrating reproducing kernel Hilbert spaces (RKHSs) spanned by a series of different basis kernels in MKHS is able to provide increased power in handling general nonlinear problems than traditional single-kernel learning in RKHS. The proposed method is developed to solve multiclass unmixing problems. To validate the proposed MKL-based algorithm, both synthetic data and real hyperspectral image data were used in our experiments. The experimental results demonstrate that the proposed algorithm has a strong ability to capture interclass spectral differences and improve unmixing accuracy, compared to the state-of-the-art algorithms tested. Yanfeng Gu, Shizhe Wang, Xiuping Jia |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2012 | L1-graph semisupervised learning for hyperspectral image classificationabstractRecently, research in semisupervised learning (SSL) based on sparse representation has shown huge potential for many classification tasks. In this paper, we address a hyperspectral image classification by integrating L1-graph and SSL. We propose a semisupervised classification method with L1-graph which has more attractive merits than traditional graph method, such as parameter free, sparsity and robustness. Our method firstly obtains the graph weights by solving a L1 optimization problem, and then generates a way of SSL with the L1-graph weights to deal with classification of hyperspectral images. The experiments are designed to cope with challenging real hyperspectral image classification task with a few labeled samples. The experimental results demonstrate the effectiveness of the L1-graph semisupervised method. Yanfeng Gu |
IGARSS | 1 |
| 2012 | Multiple-kernel learning-based unmixing algorithm for estimation of cloud fractions with MODIS and CloudSat dataabstractDetection of clouds in satellite-generated radiance images, including those from MODIS, is an important first step in many applications of these data. In this paper we apply spectral unmixing to this problem with the aim of estimating subpixel cloud fractions, as opposed to identification only of whether or not a pixel radiance contains cloud contributions. We formulate the spectral unmixing approach in terms of multiple-kernel learning (MKL). To this end we propose a MKL-based unmixing algorithm that drives a multiple-kernel description of cloud, enabling estimation of sub-pixel cloud fractions. This approach is based on supervised learning. We generate training and testing samples by using CloudSat and CALIPSO data to compute cloud fractions within individual MODIS pixels. Results of our study on limited data (1875 training and testing MODIS pixels along with their CloudSat and CALIPSO based sub-pixel cloud fractions) show that the proposed algorithm can effectively estimate sub-pixel MODIS cloud fraction and outperforms support vector machine (SVM) in terms of estimation performance. Yanfeng Gu, Shizhe Wang, Yinghui Lu, Eugene E. Clothiaux, Bin Yu 0001 |
IGARSS | 1 |
| 2012 | Representative Multiple Kernel Learning for Classification in Hyperspectral ImageryabstractRecently, multiple kernel learning (MKL) methods have been developed to improve the flexibility of kernel-based learning machine. The MKL methods generally focus on determining key kernels to be preserved and their significance in optimal kernel combination. Unfortunately, computational demand of finding the optimal combination is prohibitive when the number of training samples and kernels increase rapidly, particularly for hyperspectral remote sensing data. In this paper, we address the MKL for classification in hyperspectral images by extracting the most variation from the space spanned by multiple kernels and propose a representative MKL (RMKL) algorithm. The core idea embedded in the algorithm is to determine the kernels to be preserved and their weights according to statistical significance instead of time-consuming search for optimal kernel combination. The noticeable merits of RMKL consist that it greatly reduces the computational load for searching optimal combination of basis kernels and has no limitation from strict selection of basis kernels like most MKL algorithms do; meanwhile, RMKL keeps excellent properties of MKL in terms of both good classification accuracy and interpretability. Experiments are conducted on different real hyperspectral data, and the corresponding experimental results show that RMKL algorithm provides the best performances to date among several the state-of-the-art algorithms while demonstrating satisfactory computational efficiency. Yanfeng Gu, Di You, Yuhang Zhang 0002, Shizhe Wang, Ye Zhang 0008 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2011 | MUlti information based Ground Control Points selection methodabstractGround Control Points (GCPs) are one of the most important data used in many fields of Remote Sensing. The number and distribution of GCPs are always the key factors for the success of some researches. A GCPs selection method by integrating the three dimensional spatial information (i.e. the earth coordinates (X, Y, Z) ) and the corresponding feature information underlying the data itself was proposed. To testify the performance of this method, a Rational Function Model Resolving experiment is conducted. Experiment results show that the accuracy and time-consuming performance are both improved using GCPs selected by the proposed method. Yanfeng Gu, Zhimin Cao, Ye Zhang 0008, Xiangrong Zhang |
IGARSS | 1 |
| 2011 | A self-adjustive geometric correction method for seriously oblique aero imageabstractThe projection errors caused by curvature of the earth and relief are two important problems in geometric correction, especially in the case of imaging with large view angles. General polynomial correction model is only effective to flat area and ineffective to correct projection errors caused by both curvature of the earth and relief. According to the generated characteristic of the projection errors, this paper proposes a self-adjustive polynomial model which adds an adjustive factor in the large view angle direction to make effective correction on distorted images with both kinds of projection errors in the absence of precise attitude parameters. This paper first analyzes the principle of projection errors caused by curvature of the earth and relief, and then proposes the improved polynomial model. Experiments show that the proposed model has greatly improved accuracy for correcting the projection errors in the large angle direction compared with general polynomial models. Ye Zhang 0008, Pigang Liu, Yanfeng Gu |
IGARSS | 5 |
| 2011 | Enhanced Self-Training Superresolution Mapping Technique for Hyperspectral ImageryabstractAn efficient superresolution technique through spatial-spectral data fusion for hyperspectral (HS) imagery is proposed in this letter. The spatial and spectral contents of an HS image are extracted using a linear mixture model and a fully constrained least squares unmixing technique. These data are then combined using a spatial correlation model through a learning-based superresolution mapping (SRM) algorithm. The proposed spatial correlation model realistically simulates a mapping model between the low-resolution (LR) HS image and its subsampled version ( LR2HS image) to train the designed SRM algorithm for mapping from the LR to high resolution. The experiments on real HS images validate the accuracy and low complexity of the proposed autonomous technique for key information detection in HS imagery. Fereidoun A. Mianji, Yanfeng Gu, Ye Zhang 0008, Junping Zhang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2011 | Kernel-based regularized-angle spectral matching for target detection in hyperspectral imagery
Yanfeng Gu, Shizhe Wang, Ye Zhang 0008 |
Pattern Recognit. Lett. | 1 |
| 2010 | A Kernel-Based Nonparametric Regression Method for Clutter Removal in Infrared Small-Target Detection ApplicationsabstractSmall-target detection in infrared imagery with a complex background is always an important task in remote-sensing fields. Complex clutter background usually results in serious false alarm in target detection for low contrast of infrared imagery. In this letter, a kernel-based nonparametric regression method is proposed for background prediction and clutter removal, furthermore applied in target detection. First, a linear mixture model is used to represent each pixel of the observed infrared imagery. Second, adaptive detection is performed on local regions in the infrared image by means of kernel-based nonparametric regression and two-parameter constant false alarm rate (CFAR) detector. Kernel regression, which is one of the nonparametric regression approaches, is adopted to estimate complex clutter background. Then, CFAR detection is performed on “pure” target-like region after estimation and removal of clutter background. Experimental results prove that the proposed algorithm is effective and adaptable to small-target detection under a complex background. Yanfeng Gu, BaoXue Liu, Ye Zhang 0008 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2009 | Kernel Regression-based Background Predicting Method for Target Detection in SAR ImageabstractTarget detection with SAR image is one of important research topics in remote sensing. In this paper, a kernel regression-based predicting method is proposed for target detection in SAR image. Badly speckle noise and background clutter are two main factors which make the target detection with SAR image difficult. In the proposed method, the kernel regression on local image is used to exactly predict the background interferences and make Gaussian assumption in conventional detector better followed after kernel regression-based prediction and suppression of background clutter. Thus, final CFAR detection is performed on the background clutter-removed SAR image. Experiments conducted on real SAR image show that the proposed algorithm can effectively predict and suppress background clutters, and greatly improve the performance of the conventional CFAR detector. Yanfeng Gu, Jinglong Han, Ye Zhang 0008 |
IGARSS (4) | 1 |
| 2009 | Resolution Enhancement of Hyperspectral Images using a Learning-based Super-resolution Mapping TechniqueabstractA fast and efficient spatial-spectral fusion method for resolution enhancement of hyperspectral imagery is proposed in this paper. A linear mixture model and fully constrained least squares based unmixing algorithm are applied for spectral unmixing of the hyperspectral imagery and the resulted fractional images are processed using a spatial-spectral information correlation model through a learning-based super-resolution mapping technique. To validate the performance of the method, experiments are carried out on real images. The obtained results validate the reliability of the technique. The main advantages of the proposed method include its autonomous nature so that it doesn't need any high resolution secondary source of data, its acceptable performance, and its low computational cost which makes it favorable for realtime target recognition and tracking applications. Fereidoun A. Mianji, Ye Zhang 0008, Yanfeng Gu |
IGARSS (3) | 3 |
| 2009 | Spatial-spectral Data Fusion for Resolution Enhancement of Hyperspectral ImageryabstractA new spatial-spectral data fusion technique based on spectral mixture analysis and super-resolution mapping for spatial resolution enhancement of hyperspectral imagery is proposed in this paper. To this end, a linear mixture model and a constrained least squares based unmixing algorithm are applied for spectral unmixing of the hyperspectral imagery and the resulted fractional images are processed based on a spatial-spectral information correlation model through a super-resolution mapping technique. The obtained results validate the effectiveness of the method. It doesn't need any a priori information of the scene or secondary high resolution source of data, and is fast. Fereidoun A. Mianji, Ye Zhang 0008, Yanfeng Gu, Asad Babakhani |
IGARSS (3) | 3 |
| 2009 | Robust Feature Matching and Selection Methods for Multisensor Image RegistrationabstractThe crucial problem of multisensor image registration is how to establish the correspondences between the features extracted from reference and input images. Generally, most existing methods only consider how to extract features, the quality of the features is ignored. In this paper, we combine scale invariant feature transform (SIFT) and maximally stable extremal region (MSER) to initialize the process of extracting plenty of control points(CPs) pairs. A concept of distribution quality(DQ) is introduced to quantify the distribution of CPs pairs, experimental analysis is illustrated to analyze the effects of CPs pairs number and DQ on the registration root mean square error(RMSE). An automatic feature matching and selection algorithm is then proposed, extensive experiments demonstrate the effectiveness of the proposed algorithm by aligning real images. Ye Zhang 0008, Yanfeng Gu |
IGARSS (3) | 3 |
| 2008 | A Selective KPCA Algorithm Based on High-Order Statistics for Anomaly Detection in Hyperspectral ImageryabstractIn this letter, a selective kernel principal component analysis (KPCA) algorithm based on high-order statistics is proposed for anomaly detection in hyperspectral imagery. First, KPCA is performed on the original hyperspectral data to fully mine the high-order correlation between spectral bands. Then, the average local singularity (LS) is defined based on the high-order statistics in the local sliding window, which is used as a measure for selecting the most informative nonlinear component for anomaly detection. By the selective KPCA, information on anomalous targets is extracted to maximum extent, and background clutters are well suppressed in the selected component. Finally, the selected component with maximum average LS is used as input for anomaly detectors. Numerical experiments are conducted on real hyperspectral images collected by the airborne visible/infrared imaging spectrometer. The results strongly prove the effectiveness of the proposed algorithm. Yanfeng Gu, Ye Zhang 0008 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2008 | Integration of Spatial-Spectral Information for Resolution Enhancement in Hyperspectral ImagesabstractIn this paper, a new algorithm is proposed for resolution enhancement in hyperspectral images (HSIs). The key techniques are included: spectral unmixing and superresolution mapping, by which spatial and spectral information of HSIs is substantially fused. The proposed algorithm first represents each pixel in scene as a linear combination of landcover spectra and noise. Then, a fully constrained least squares algorithm is used to obtain the proportion of each landcover in each pixel, i.e., abundance, subjecting to two constraints: nonnegativity and sum-to-one. After that, superresolution mapping is performed on high-resolution grids according to spectral unmixing abundances of each landcover and following spatial correlation of clutters. Thus, by reasonably integrating spatial and spectral information of landcovers in HSIs, the proposed algorithm realizes resolution enhancement of the HSIs based on a back-propagation neural network. The proposed algorithm is independent from thea prioriinformation associated with original HSIs, i.e., a main merit of the algorithm. In order to evaluate the performance of the new algorithm, numerical experiments are conducted on both simulated images and real HSIs collected by the Airborne Visible/Infrared Imaging Spectrometer. The proposed algorithm is compared with the traditional method in the experiments. The experimental results prove that the proposed algorithm effectively enhances the resolution of HSIs and indicate its applicability. Yanfeng Gu, Ye Zhang 0008, Junping Zhang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2006 | A Selective Kernel PCA Algorithm for Anomaly Detection in Hyperspectral ImageryabstractIn this paper, a selective kernel principal component analysis algorithm is proposed for anomaly detection in hyperspectral imagery. The proposed algorithm tries to solve the problem brought by high dimensionality of hyperspectral images in anomaly detection. This algorithm firstly performs kernel principal component analysis (KPCA) on the original data to fully mine high-order correlation between spectral bands. Then, high-order statistics in local scene are exploited to define local average singularity (LAS), which is used to measure the singularity of each nonlinear principal component transformed. Based on LAS, one component transformed with maximum singularity is selected after KPCA. Finally, with RX detector, anomaly detection is performed on the component selected. Numerical experiments are conducted on real hyperspectral images collected by AVIRIS. The results prove that the proposed algorithm outperforms the conventional RX algorithm Yanfeng Gu, Ye Zhang 0008 |
ICASSP (2) | 1 |
| 2006 | Unmixing Component Analysis for Anomaly Detection in Hyperspectral ImageryabstractAnomaly detection is one of the most important applications for hyperspectral images. In this paper, a new algorithm called unmixing component analysis (UCA) is proposed for anomaly detection in hyperspectral imagery. The proposed algorithm firstly performs spectral unmixing only with background endmembers on original hyperspectral images, and the unmixing error data are retained. Secondly, kernel principal analysis (KPCA) is performed on the error data to concentrate and extract useful information about anormalous targets. After that, non-linear principal component that includes the most information about anomalous targets is selected based on non-gaussianity measures. Finally, anomaly detection is conducted on the selected non-linear principal component using RX detector. Numerical experiments are performed on AVIRIS data with 126 bands. The experimental results show the proposed algorithm greatly modifies performance of the conventional RX algorithm and has good detection performance with low false alarms. Yanfeng Gu, Ye Zhang 0008 |
ICIP | 1 |
| 2006 | Target Detection For Hyperspectral Images Using ICA-Based Feature ExtractionabstractIn this paper we present a target detection method for hyperspectral images using feature extraction based on independent component analysis (ICA). This method makes good use of the high order statistic of image data and greatly overcome the spectral signature variability. ICA aims to find a linear representation of the observed data in order that the components are statistically independent, or as independent as possible. Such an independent component can capture the intrinsic structure of data and extract image features, including target feature that will be used in detection. First each pixel, which is assumed to be a linear mixture of target and background spectra, is projected onto the orthogonal background subspace to remove the background spectral portion from the corresponding pixel spectrum. Then the targets in the background-removed image are estimated through matched filtering with the feature of target component extracted by ICA. The method has been testified on airborne visible and infrared imaging spectrometer (AVIRIS) data. The experimental results show that targets are successfully separated from the background, demonstrating the good performance of this method to detect targets in hyperspectral images. Junping Zhang, Yanfeng Gu |
IGARSS | 3 |
| 2004 | Kernel-based invariant subspace method for hyperspectral target detectionabstractIn this paper, a kernel-based invariant subspace detection method is proposed for small target detection of hyperspectral images. The method combines kernel principal component analysis (KPCA) and the linear mixture model (LMM). The LMM is used to describe each pixel in the hyper-spectral image as a mixture of target, background and noise. The KPCA is used to build subspaces of the target and background. A generalized likelihood ratio test is used to detect whether each pixel in the hyperspectral image includes the target. Numerical experiments are performed on AVIRIS hyperspectral data with 126 bands. The experimental results show the effectiveness of the proposed method and prove that this method can commendably overcome spectral variability in hyperspectral target detection, and it has good ability to separate target from background. Ye Zhang 0008, Yanfeng Gu |
ICASSP (5) | 2 |
| 2003 | Unsupervised subspace linear spectral mixture analysis for hyperspectral imagesabstractIn this paper, an unsupervised subspace linear spectral unmixing algorithm for hyperspectral data is investigated, which includes two key techniques: subspace minimum noise fraction transformation (SMNFT) and independent component analysis (ICA). The SMNFT is used to reduce noise, remove correlation between neighboring bands and determine intrinsic dimensionality of hyperspectral data. Then the ICA is applied to unmix hyperspectral images and obtain independent endmembers. The main merits of the proposed algorithm are that it can fast unsupervisedly separate useful and independent endmembers resident in hyperspectral images. The experimental results demonstrate that this algorithm can effectively identify independent endmembers. Meanwhile, the results show high computational efficiency of the algorithm. The time consumed by the SMNFT is merely one fifth of the traditional minimum noise fraction transformation. Yanfeng Gu, Ye Zhang 0008 |
ICIP (1) | 1 |
| 2002 | A kernel based nonlinear subspace projection method for reduction of hyperspectral image dimensionalityabstractA kernel based nonlinear subspace projection (KNSP) method is proposed for reduction of hyperspectral image dimensionality. This method involves three steps: subspace partition of full data space, feature extraction based on kernel principal component analysis (KPCA) in subspace and feature selection based on class separability criterion. The main merit of the proposed method is that it is more suitable for feature extraction than linear principal component analysis (PCA) and segmented principal component: transformation (SPCT), in particular, when hyperspectral data have nonlinear characteristics. In order to testify the effectiveness of the KNSP method for reduction of hyperspectral image dimensionality, hyperspectral image classification is performed on AVIRIS data. Experimental results show that when the hyperspectral dimensionality is reduced to a few features, the average classification accuracy of the new method is higher than those of PCA and SPCT methods. Yanfeng Gu, Ye Zhang 0008, Junping Zhang |
ICIP (2) | 1 |