Ping Zhong 0001

dblp:75/4407-1 · DBLP profile ↗
← Back
66ranked-venue papers
11as first author
39since 2021 · last 2026
0000-0002-8686-3928ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 37 · 7 first-author · 20 since 2021Artificial intelligence and machine learning · 19 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 12 since 2021
YearPublicationVenuePosition
2026 Separate to generalization: Two-branch feature separation framework for generalized underwater image restoration
Jiahao Qi, Chen Chen 0152, Kangcheng Bin, Ping Zhong 0001
Pattern Recognit.5
2026 Adversarial spectral perturbation for single-domain generalized object detection
Xiangsheng Wang, Kangcheng Bin, Ping Zhong 0001
Pattern Recognit.6
2026 Robust Fine-Grained Oriented Ship Detection for Remote Sensing Imagery via Controllable Generative Pretraining
abstract
Fine-grained ship recognition in remote sensing imagery is essential for maritime applications. However, its development is hindered by two challenges: 1) the limited granularity of existing ship detection datasets, and 2) the disturbance of complex maritime conditions as well as the arbitrary ship orientations and distributions. To address the first issue, we annotated a large-scale fine-grained ship instance detection dataset (LAFI), comprising 48,717 ship instances worldwide with 49 categories. To tackle the challenges of marine disturbance and diverse ship status, we proposed a controllable generative knowledge-driven ship detection framework (COSD). It employs a controllable diffusion model guided by ship-marine textual prompt to generate millions of synthetic images that not only preserve ship structures but also cover diverse sea and weather conditions for robust pretraining. The pretraining stage then utilizes masked reconstruction to learn component-level cues under occlusion, clutter, fog, and illumination changes. Furthermore, a heterogeneous feature alignment decoder is designed to align multi-modal metrics of orientation and distribution features in the latent space, allowing for accurate representation of diverse ship status. Extensive experiments on two benchmark datasets showed that our method respectively increased 0.011 and 0.030 mean average precision (mAP@50) over SOTA methods, particularly in scenarios involving small, densely packed and arbitrary oriented ships.
Da He, Xikun Hu, Ping Zhong 0001, Qian Shi 0001, Xiaoping Liu 0001, Yanfei Zhong, Liangpei Zhang 0001
IEEE Trans. Image Process.6
2026 Fusion of Infrared and Visible Images Based on Iterative Dual-Branch Attention and Modality Discrepancy Guidance
abstract
The fusion of infrared and visible images aims to generate images that provide a more comprehensive description of the scene. Convolutional neural network and transformer are two commonly used methods in image fusion. The former focuses on extracting local features but lacks a global receptive field, and the latter can extract global information but ignores the discrepancy information between modalities. In view of this, we propose an efficient fusion network based on iterative dual-branch attention and modality discrepancy guidance (IDMDNet), which consists of three components. In the first component, we use shallow and deep feature extraction modules to extract features. In the second component, we first design an iterative dual-branch attention module to capture important features of each modality, thereby preserving important information. Secondly, to model the global information of the source image, we introduce a transformer module to achieve the fusion of shallow global information. Further, we design a modality discrepancy guided fusion module to promote the fusion of modality discrepancy information and ultimately achieve modality information complementation. In the third component, we introduce an invertible neural network to reduce the loss of feature information, thus achieving high quality image reconstruction. Finally, we construct a loss function for IDMDNet that includes intensity, gradient, and multiscale structure to motivate the network to preserve texture and target details. Experiments on infrared and visible image fusion benchmark datasets show that the proposed IDMDNet has competitive fusion performance. Remarkably, it can be effectively applied to other infrared and visible datasets without the need for fine-tuning, highlighting its good generalization ability. The source code of IDMDNet has been released athttps://github.com/AHUT-MILAGroup/IDMDNet.
Shuting Zhu, Yazhou Yao, Ping Zhong 0001
IEEE Trans. Multim.5
2025 Dive into Aerial Remote Sensing Underwater Depth Estimation with Hyperspectral Imagery
abstract
Visible spectrum images capture limited information from just three discrete bands, often resulting in suboptimal performance in underwater depth estimation (UDE) due to significant information loss from water absorption. In contrast, HSIs, which include hundreds of continuous bands, provide abundant spectral information that offers greater resilience against the adverse effects of water absorption. In this paper, we conduct a comprehensive study to investigate how spectral information can enhance remote sensing UDE through two key aspects: the benchmark dataset and the general framework. For the benchmark dataset, we construct a real-world hyperspectral UDE (HUDE) dataset ATR-HUDE, comprising approximately 500 synchronized hyperspectral and LiDAR data pairs collected from diverse coastal scenes and flight altitudes. Regarding the general framework, we integrate recent advances in state space models and physical imaging models to design a novel HUDE framework named HUDEMamba that estimates underwater depth using both model-driven and data-driven approaches. Experimental results on the constructed benchmark dataset validate the potential of HUDE and the effectiveness of HUDEMamba.
Jiahao Qi, Chen Chen 0152, Dehui Zhu, Kangcheng Bin, Ping Zhong 0001
AAAI6
2025 UCM-VeID V2: A Richer Dataset and A Pre-training Method for UAV Cross-Modality Vehicle Re-Identification
abstract
Cross-Modality Re-Identification (VT-ReID) aims to achieve around-the-clock target matching, benefiting from the strengths of both RGB and infrared (IR) modalities. However, the field is hindered by limited datasets, particularly for vehicle VT-ReID, and by challenges such as modality bias training (MBT), stemming from biased pre-training on ImageNet. To tackle the above issues, this paper introduces an dataset benchmark, named UCM-VeID V2, for vehicle VT-ReID, and proposes a new self-supervised pre-training method, Cross-Modality Patch-Mixed Self-Supervised Learning (PMSL). UCM-VeID V2 dataset features a significant increase in data volume, along with enhancements in multiple aspects. PMSL addresses MBT by learning modality-invariant features through Patch-Mixed Image Reconstruction (PMIR) and Modality Discrimination Adversarial Learning (MDAL), and enhances discriminability with Modality-Augmented Contrasting Cluster (MACC). Comprehensive experiments are carried out to validate the proposed method.
Jiahao Qi, Chen Chen 0152, Kangcheng Bin, Ping Zhong 0001
CVPR5
2025 Fusion Meets Diverse Conditions: A High-Diversity Benchmark and Baseline for UAV-Based Multimodal Object Detection with Condition Cues
abstract
Unmanned aerial vehicles (UAV)-based object detection with visible (RGB) and infrared (IR) images facilitates robust around-the-clock detection, driven by advancements in deep learning techniques and the availability of high-quality dataset. However, the existing dataset struggles to fully capture real-world complexity for limited imaging conditions. To this end, we introduce a high-diversity dataset ATR-UMOD covering varying scenarios, spanning altitudes from 80m to 300m, angles from 0° to 75°, and all-day, all-year time variations in rich weather and illumination conditions. Moreover, each RGB-IR image pair is annotated with 6 condition attributes, offering valuable high-level contextual information. To meet the challenge raised by such diverse conditions, we propose a novel prompt-guided condition-aware dynamic fusion (PCDF) to adaptively reassign multimodal contributions by leveraging annotated condition cues. By encoding imaging conditions as text prompts, PCDF effectively models the relationship between conditions and multimodal contributions through a task-specific soft-gating transformation. A prompt-guided condition-decoupling module further ensures the availability in practice without condition annotations. Experiments on ATR-UMOD dataset reveal the effectiveness of PCDF.
Chen Chen 0152, Kangcheng Bin, Jiahao Qi, Tianpeng Liu, Zhen Liu 0004, Yongxiang Liu, Ping Zhong 0001
ICCV9
2025 Spectral-Feature-Guided Controllable Diffusion for SAR-to-Optical Satellite Imagery Generation in Wildfire Mapping
Yushan Zou, Xikun Hu, Ping Zhong 0001
PRCV (15)3
2025 Partial Attention Feature Aggregation Network for Lightweight Remote Sensing Image Super-Resolution
abstract
Most lightweight super-resolution networks are designed to improve performance by introducing an attention mechanism and to reduce model parameters by designing lightweight convolutional layers. However, the introduction of the attention mechanism often leads to an increase in the number of parameters. In addition, the lightweight convolutional layer has a limited receptive field and cannot effectively capture long-range dependencies. In this letter, we design a novel lightweight base module called partial attention convolution (PAConv) and develop three variants of PAConv with different receptive fields to collaboratively exploit non-local information. Based on PAConv, we further propose a lightweight super-resolution network called partial attention feature aggregation network (PAFAN). Specifically, we arrange the PAConv variants in a progressive iterative manner to form the attention progressive feature distillation block (APFDB), which aims to gradually optimize the distilled features. Furthermore, we construct a multi-level aggregation spatial attention (MASA) via a stacking of the PAConv variants to systematically coordinate multi-scale structural information. Extensive experiments conducted on benchmark datasets show that PAFAN achieves an optimal balance between reconstruction quality and computational efficiency. In particular, with only 123K parameters and 0.49G FLOPs, PAFAN can maintain a performance comparable to that of SOTA methods.
Tiancheng Shao, Mingyang Du, Ping Zhong 0001
IEEE Geosci. Remote. Sens. Lett.5
2025 Discriminative Latent-Space Learning for Fine-Grained Object Detection in Remote Sensing Images
abstract
Abstract—Fine-grained object detection (FOD) is essential in many remote sensing image interpretation tasks. Existing FOD methods have achieved remarkable progress in modeling discriminative features for FOD in remote sensing images. However, they receive unsatisfactory recognition accuracy due to the curse of dimensionality (CoD) problem. In this paper, we propose an orthogonal constraint-based discriminative latent-space learning (DLL) method to address the CoD problem. We first optimize a sparse optimization paradigm with convex relaxation to extract shared features between fine-grained objects into a low-dimensional latent-space. Then, we solve the orthogonal space of the latent-space for extracting irrelevant features related to common features from input features, i.e., discriminative features. The sparse optimization paradigm reprojects the underlying trends hidden in high-dimensional feature space into a low-dimensional latent space and hence addresses the CoD problem. We use neural parameters with latent-space and orthogonal constraints to approximately solve the proposed DLL, which can be efficiently optimized under convex programming tools. We theoretically prove the effectiveness of the proposed DLL. By adding our DLL to the existing deep learning-based object detection method, extensive experiments conducted on two datasets demonstrate that our method achieves superior performance compared with other state-of-the-art methods.
Xikun Hu, Wenlin Liu, Xiangsheng Wang, Chen Chen 0152, Ya Jiang, Ping Zhong 0001
IEEE Trans. Geosci. Remote. Sens.8
2025 MSP-O: A Two-Stage Mask Pretraining Method With Object Focus for Remote Sensing Object Detection
abstract
Pretraining is showing fundamental importance in computer vision, to initialize the backbones of deep learning models with large natural image datasets, e.g., ImageNet. However, there is seldom a special pretraining technique focused on remote sensing object detection due to the distinct characteristics of remote sensing images, such as specific imaging angles, complex backgrounds, and various target categories. This article aims to bridge the gap between natural image data and remote sensing image data and to facilitate pretraining for the specific object detection task. This article introduces masked remote sensing pretraining with object (MSP-O), an innovative two-stage masked pretraining method. Traditional single-stage pretraining frameworks aim to extract intrinsic features from large-scale image datasets to learn object discrimination and localization ability. MSP-O builds upon this by adding a second-stage pretraining phase to learn specific features related to remote sensing objects of interest themselves, facilitating the downstream object detection task. Specifically, MSP-O incorporates a pixel-level restoration task focused on the objects of interest, guiding the model to prioritize the objects to be detected. Compared with traditional pretraining methods, MSP-O is able to learn feature representations that are more relevant to the detection task, especially when the data scale is limited. Extensive experiments on the DOTA and DIOR datasets demonstrate the effectiveness of MSP-O, showing substantial improvements in detection performance compared with conventional methods.
Wenlin Liu, Xikun Hu, Ping Zhong 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Physics-Informed Curriculum Learning Framework for Hyperspectral Underwater Target Characterization
abstract
Hyperspectral imaging (HSI) provides fine-grained spectral information essential for material identification and target detection, particularly in complex environments such as underwater scenarios. However, hyperspectral underwater target detection (HUTD) remains challenging due to severe spectral distortions and variability introduced by wavelength-dependent absorption and the dynamic nature of aquatic environments. Existing separation-based and characterization-based methods are often constrained by weak signal responses or a heavy reliance on accurate environmental parameter estimation, which is difficult to achieve in practice. To overcome these limitations, we propose PCL-HUTD, a novel physics-informed curriculum learning framework for robust underwater target characterization without requiring explicit environmental modeling. PCL-HUTD integrates a physics-guided target construction module with a hard-sample aware contrastive learning strategy, enhanced by unsupervised clustering and a perturbation-consistency based sample selection mechanism. Furthermore, a closed-loop curriculum learning paradigm is introduced to progressively refine target representations throughout training. Extensive experiments on three real-world HUTD datasets demonstrate that PCL-HUTD achieves state-of-the-art performance in both detection accuracy and robustness, particularly under challenging conditions with strong background interference. These results validate the effectiveness of our parameter-free, physics-informed approach for underwater hyperspectral target detection.
Jiahao Qi, Chen Chen 0152, Dehui Zhu, Kangcheng Bin, Ping Zhong 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 Empowering Hybrid-Level Contrastive Learning for Hyperspectral Underwater Target Detection in Nearshore Environment
abstract
UAV-borne hyperspectral remote sensing has emerged as a promising technique for underwater target detection (UTD). Existing hyperspectral UTD (HUTD) methods in deep-sea environments typically rely on bathymetric models to characterize spectral attenuation caused by the water column, often coupled with restoration- or prediction-based pipelines for target detection. However, in nearshore environments, high turbidity and complex seabed topography introduce nonlinear and heterogeneous attenuation effects that cannot be accurately modeled by conventional bathymetric models. Moreover, the highly dynamic nature of nearshore waters induces significant spectral variability in target signatures, rendering restoration- and prediction-based pipelines ineffective due to their limited capacity to model such variability. To address these challenges, we propose the Hyperspectral Underwater Contrastive Learning Network (HUCLNet), a data-driven framework for HUTD that eliminates dependence on bathymetric models. Rather than modeling spectral attenuation, HUCLNet establishes a semantically meaningful latent space with enhanced target-background separability through a hybrid-level contrastive learning framework. A reliability-guided clustering strategy is introduced to refine the contrastive learning inputs and improve representation robustness. Furthermore, a self-paced learning mechanism flexibly integrates the clustering and contrastive modules to stabilize training and accelerate convergence. Extensive experiments demonstrate that HUCLNet outperforms state-of-the-art methods across various evaluation metrics.
Jiahao Qi, Ping Zhong 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Subtle Spectral Difference Discriminative Deep Metric Learning With Spectral Center Construction for Hyperspectral Target Detection
abstract
Detection of targets for hyperspectral images (HSIs) persists as a fundamental task in remote sensing image processing. Exploring the discriminative ability of deep Siamese networks to distinguish targets from backgrounds is the mainstream method for target detection in HSIs. Nevertheless, these methods enhance the discriminative ability of networks by learning the separability distance between the target and the overall backgrounds, where the backgrounds are considered as a single category. As a result, they may struggle to effectively suppress backgrounds with solely subtle spectral differences from the target, resulting in a limited separability performance, and the inability to accurately detect the targets. To alleviate this problem, we propose a novel subtle spectral difference discriminative deep metric learning-based target detector for HSIs (denoted as S2D3ML) in this work. The proposed S2D3ML constructs a deep metric learning framework embedded with a discriminative constraint to learn a deep metric feature space for addressing limited separability, in which the subtle feature differences between targets and different ground objects can be distinguished. In addition, we investigate a new multi-block sparse representation score-based strategy to obtain sufficient samples and spectral centers of backgrounds for training the S2D3ML framework. Finally, the detection of targets is executed within the learned metric space. A comprehensive suite of experiments is rigorously conducted on four benchmark datasets, and the results indicate that the S2D3ML achieves superior performance in HSIs target detection.
Dehui Zhu, Yuetian Lu, Ping Zhong 0001, Bo Du 0001, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Active Object Detection for UAV Remote Sensing via Behavior Cloning and Enhanced Q-Network With Shallow Features
abstract
Object detection in Unmanned Aerial Vehicle (UAV) remote sensing imagery faces two critical challenges: the inability to effectively handle multi-scale targets and the difficulty in addressing partial occlusions, which significantly impact detection accuracy in real-world applications. To overcome these limitations, we propose an active object detection (AOD) method that dynamically adjusts viewing angles and scales during the detection process. Our key innovation lies in a novel architecture that uniquely combines shallow Feature Pyramid Network features with detector output bounding boxes, combining with a self-supervised learning framework. The technical originality of our approach is further enhanced by our two-stage learning methodology, which initially employs behavior cloning to establish robust foundational performance, followed by Q-learning fine-tuning to optimize detection strategies. To facilitate comprehensive evaluation, we introduce CARLA-AOD, a new benchmark dataset that encompasses 18 diverse scenarios across hemispheric space above targets, specifically designed for UAV remote sensing AOD applications. Extensive experimental validation demonstrates the effectiveness of our approach, achieving substantial improvements over baseline detectors across multiple datasets: 7.9% on Small Airport, 14.0% on Virtual Park, and 10.2% on our CARLA-AOD dataset. The two-stage learning process proves particularly effective, with Q-learning fine-tuning providing an additional performance boost of up to 7.3% beyond the behavior cloning baseline. Moreover, our method achieves a 60.2% reduction in inference time compared to the state-of-the-art DCCL algorithm, making it particularly suitable for time-critical applications such as emergency response and real-time surveillance. The dataset is available at IEEE Dataport. IEEE Dataport.
Zhuocheng Zou, Xikun Hu, Ping Zhong 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Interpreting Hidden Semantics in the Intermediate Layers of 3D Point Cloud Classification Neural Network
abstract
Although 3D point cloud classification neural network models have been widely used, the in-depth interpretation of the activation of the neurons and layers is still a challenge. We propose a novel approach, named Relevance Flow, to interpret the hidden semantics of 3D point cloud classification neural networks. It delivers the class Relevance to the activated neurons in the intermediate layers in a back-propagation manner, and associates the activation of neurons with the input points to visualize the hidden semantics of each layer. Specially, we reveal that the 3D point cloud classification neural network has learned the plane-level and part-level hidden semantics in the intermediate layers, and utilize the normal and IoU to evaluate the consistency of both levels' hidden semantics. Besides, by using the hidden semantics, we generate the adversarial attack samples to attack 3D point cloud classifiers. Experiments show that our proposed method reveals the hidden semantics of the 3D point cloud classification neural network on ModelNet40 and ShapeNet, which can be used for the unsupervised point cloud part segmentation without labels and attacking the 3D point cloud classifiers.
Weiquan Liu, Minghao Liu 0007, Shijun Zheng, Xuesheng Bian, Ping Zhong 0001, Cheng Wang 0003
IEEE Trans. Multim.7
2025 Masked Spatial-Spectral Autoencoders Are Excellent Hyperspectral Defenders
abstract
Deep learning (DL) methodology contributes a lot to the development of hyperspectral image (HSI) analysis community. However, it also makes HSI analysis systems vulnerable to adversarial attacks. To this end, we propose a masked spatial-spectral autoencoder (MSSA) in this article under self-supervised learning theory, for enhancing the robustness of HSI analysis systems. First, a masked sequence attention learning (MSAL) module is conducted to promote the inherent robustness of HSI analysis systems along spectral channel. Then, we develop a graph convolutional network (GCN) with learnable graph structure to establish global pixel-wise combinations. In this way, the attack effect would be dispersed by all the related pixels among each combination, and a better defense performance is achievable in spatial aspect. Finally, to improve the defense transferability and address the problem of limited labeled samples, MSSA employs spectra reconstruction as a pretext task and fits the datasets in a self-supervised manner. Comprehensive experiments over three benchmarks verify the effectiveness of MSSA in comparison with the state-of-the-art hyperspectral classification methods and representative adversarial defense strategies.
Jiahao Qi, Zhiqiang Gong, Chen Chen 0152, Ping Zhong 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 Weakly Misalignment-Free Adaptive Feature Alignment for UAVs-Based Multimodal Object Detection
abstract
Visible-infrared (RGB-IR) image fusion has shown great potentials in object detection based on unmanned aerial ve-hicles (UAVs). However, the weakly misalignment problem between multimodal image pairs limits its performance in object detection. Most existing methods often ignore the modality gap and emphasize a strict alignment, resulting in an upper bound of alignment quality and an increase of implementation costs. To address these challenges, we propose a novel method named Offset-guided Adaptive Feature Alignment (OAFA), which could adaptively adjust the relative positions between multimodal features. Considering the impact of modality gap on the cross-modality spa-tial matching, a Cross-modality Spatial Offset Modeling (CSOM) module is designed to establish a common sub-space to estimate the precise feature-level offsets. Then, an Offset-guided Deformable Alignment and Fusion (ODAF) module is utilized to implicitly capture optimal fusion po-sitions for detection task rather than conducting a strict alignment. Comprehensive experiments demonstrate that our method not only achieves state-of-the-art performance in the UAVs-based object detection task but also shows strong robustness to the weakly misalignment problem.
Chen Chen 0152, Jiahao Qi, Kangcheng Bin, Ruigang Fu, Xikun Hu, Ping Zhong 0001
CVPR7
2024 Lighten CARAFE: Dynamic Lightweight Upsampling with Guided Reassemble Kernels
Ruigang Fu, Qingyong Hu, Xiaohu Dong, Yinghui Gao, Ping Zhong 0001
ICPR (4)6
2024 Content-Aware Feature Upsampling for Voxel-Based 3D Semantic Segmentation
Ruigang Fu, Qingyong Hu, Ping Zhong 0001
ICPR (30)5
2024 Near Real-Time Burned Area Progression Mapping With Multispectral Data Using Ensemble Learning
abstract
Monitoring the wildfire progression is essential to quantify the fire-disturbance areas for emergency responses. To combine the advantages of pixelwise machine learning (ML) method and region-based deep learning (DL) segmentation model, this study proposes a two-phase hybrid framework for near real-time burned area progression mapping: the first one intends to depict burned area delimitation using a contextual algorithm HRNet to exclude the unburned areas outside the perimeter and minimize omission errors, which partially remain unburned patches within the delimitation as commission errors. The second phase refines the burned area spatially using ensemble fusion based on an updating support vector machine (SVM) model under the voting scheme as new imagery arrives to reduce the commission errors consecutively. The validation results showed that the accuracy of perimeter prediction using the HRNet can reach 96.77% in Kappa. The iterative optimization can improve the average Kappa value from 62.55% to 70.75% for burned area pixel classification using pixelwise SVM alone. The proposed ensemble learning framework can further refine the burned area progression results, reaching an average Kappa up to 85.19%, at four acquisition dates with Sentinel-2 and Landsat-8 available during the Sand fire event that occurred in California.
Xikun Hu, Puzhao Zhang, Ka-Veng Yuen, Ping Zhong 0001
IEEE Geosci. Remote. Sens. Lett.5
2024 Attention-based Sparse and Collaborative Spectral Abundance Learning for Hyperspectral Subpixel Target Detection
Dehui Zhu, Ping Zhong 0001, Bo Du 0001, Liangpei Zhang 0001
Neural Networks2
2024 Deep Intrinsic Decomposition With Adversarial Learning for Hyperspectral Image Classification
abstract
Convolutional neural networks (CNNs) have shown their potential ability to extract discriminative features for hyperspectral image classification. However, traditional deep learning methods using CNNs tend to overlook the influence of complex environmental factors. These factors contribute to an increase in intraclass variance and a decrease in interclass variance, making it considerably more challenging to extract meaningful features. To overcome this problem, this work develops a novel deep intrinsic decomposition with adversarial learning, namely AdverDecom, for hyperspectral image classification to mitigate the negative impact of environmental factors on classification performance. First, we develop a generative network for hyperspectral images (HyperNet) to extract the environment-related features and category-related features from the image. Then, a discriminative network is constructed to distinguish different environmental categories. Finally, an environment-category joint learning loss is developed for adversarial learning to make the deep model learn discriminative features. Experiments are conducted over four commonly used real-world datasets and the comparison results show the superiority of the proposed method. The implementation of the proposed method could be accessed athttps://github.com/shendu-sw/Adversarial_Learning_Intrinsic_Decompositionfor the sake of reproducibility.
Zhiqiang Gong, Jiahao Qi, Ping Zhong 0001, Xian Zhou 0003, Wen Yao 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 HyperDID: Hyperspectral Intrinsic Image Decomposition With Deep Feature Embedding
abstract
The dissection of hyperspectral images into intrinsic components through hyperspectral intrinsic image decomposition (HIID) enhances the interpretability of hyperspectral data, providing a foundation for more accurate classification outcomes. However, the classification performance of HIID is constrained by the model’s representational ability. To address this limitation, this study rethinks hyperspectral intrinsic image decomposition for classification tasks by introducing deep feature embedding. The proposed framework, HyperDID, incorporates the Environmental Feature Module (EFM) and Categorical Feature Module (CFM) to extract intrinsic features. Additionally, a Feature Discrimination Module (FDM) is introduced to separate environment-related and category-related features. Experimental results across three commonly used datasets validate the effectiveness of HyperDID in improving hyperspectral image classification performance. This novel approach holds promise for advancing the capabilities of hyperspectral image analysis by leveraging deep feature embedding principles. The implementation of the proposed method could be accessed soon at https://github.com/shendu-sw/HyperDID for the sake of reproducibility.
Zhiqiang Gong, Xian Zhou 0003, Wen Yao 0001, Xiaohu Zheng, Ping Zhong 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 Detecting Nearshore Underwater Targets With Hyperspectral Nonlinear Unmixing Autoencoder
abstract
Hyperspectral underwater target detection (HUTD) is a promising and challenging task in remote sensing image processing. Existing methods face significant challenges when adapting to nearshore environments, where cluttered backgrounds hinder the extraction of target signatures and exacerbate signal distortion. Hyperspectral unmixing (HU) demonstrates potential effectiveness for nearshore underwater target detection (UTD) by simultaneously extracting water background endmembers and separating target signals. To this end, this article investigates a novel nonlinear unmixing network for hyperspectral UTD, denoted as nonlinear unmixing network for hyperspectral-UTD (NUN-UTD), in which a well-designed autoencoder-based unmixing network is used to obtain the abundance map as the detection result. To address the weak underwater target signals, a target prior spectral preservation scheme is employed to guide the unmixing network in learning the accurate target abundance. Besides, to address the complexity of the nearshore environment, a pseudomixed data classification constraint is incorporated into the objective function to enhance the discriminative capability between the background and the target. Moreover, we adopt an additive postnonlinear model in the decoder to deal with the interactions between underwater spectra to account for the nonlinear effects between spectra of underwater substances. To validate the effectiveness of the proposed method, we constructed a hyperspectral dataset for nearshore UTD. Extensive experiments conducted on three real-world datasets and one simulated dataset demonstrate that our method achieves outstanding performance in HUTD.
Jiahao Qi, Dehui Zhu, Hejun Jiang, Ping Zhong 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 DSRNet: Diagonal Subsampling Reconstruction Network for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection seeks to locate pixels in a scene exhibiting substantial spectral discrepancies from the surrounding background pixels, holding essential applications in both civilian and military domains. However, in real-world scenarios, the inherent properties of hyperspectral imaging, the irregular forms of anomalous targets, and the lack of prior information pose significant challenges for anomaly detection methodologies. To address these issues, we first explore a novel reconstruction-based modeling approach for hyperspectral anomaly detection, offering a rational motivation for the modeling approach and a detailed exposition of its effective implementation. Furthermore, we propose a diagonal subsampling reconstruction network (DSRNet) for anomaly detection of hyperspectral data. Specifically, DSRNet consists of a paired training data generation algorithm using subsampling and a self-supervision training process enforced with a reconstruction consistency constraint. The training input is derived by a randomly diagonal averaging subsampler, where training pairs are derived from the same original hyperspectral data. Extensive experiments on four public datasets demonstrate the superiority of our DSRNet compared over several state-of-the-art baselines, with an average AUC score increase of 0.0156.
Xikun Hu, Ping Zhong 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Empowering Physical Attacks With Jacobian Matrix Regularization Against ViT-Based Detectors in UAV Remote Sensing Images
abstract
Vision transformers (ViTs) have achieved great success in unmanned aerial vehicle (UAV) target detection tasks. However, little attention has been paid to the adversarial attack against ViT-based detectors, and the generated adversarial examples cannot take physical realizability and attack transferability into account at the same time. To overcome the limitation, we focus on transferable attacks toward ViT-based detectors in optical UAV-based remote sensing images and generate adversarial examples in the physical world. Concretely, we design unique perturbation patches deployed within and beyond the target object rather than requiring the patches to be aligned with image tokens. To narrow the gap between limited digital samples and complex physical scenarios, we conduct data augmentation on training images at global and local levels. In addition, we propose a novel transferable attack method named Jacobian matrix regularization (JMR), which consists of feature variance regularization (FVR) and attention weight regularization (AWR). Specifically, FVR calculates feature variances of different channels within specific layers and then sets the features as zeros for channels with top variances. AWR is achieved by masking the largest self-attention weights. We conduct extensive transferable experiments with typical detectors in both digital and physical UAV-based remote sensing scenarios. The results indicate that our method could achieve competitive transferability compared with state-of-the-art methods.
Yu Zhang 0221, Zhiqiang Gong, Wenlin Liu, Jiahao Qi, Xikun Hu, Ping Zhong 0001
IEEE Trans. Geosci. Remote. Sens.8
2024 SR-Adv: Salient Region Adversarial Attacks on 3D Point Clouds for Autonomous Driving
abstract
Autonomous driving safety based on LiDAR perception is increasingly becoming a hot spot. Specifically, 3D adversarial examples always make the prediction results of deep neural network models unpredictable, which poses a major security risk to autonomous driving systems. However, the vulnerability of 3D neural network models to adversarial examples is less explored. At present, existing adversarial attack methods often obtain 3D adversarial examples by perturbing the entire point cloud, which requires a large number of perturbed points, i.e. requires a large perturbation budget. In this paper, we propose a Salient Region Adversarial attack method (SR-Adv) to generate adversarial point clouds, by perturbing fewer regions and fewer points. To our knowledge, we are the first to propose region-based attacks for 3D point clouds. First, the proposed SR-Adv employs game theory to extract salient regions of point clouds. This mechanism assigns a value to each region to measure its importance to the 3D neural network model prediction results and realizes the vulnerability analysis of the 3D model. Second, we propose a novel optimization-based gradient attack algorithm to achieve adversarial attacks on salient regions. We evaluate the proposed SR-Adv attack method on the synthetic datasets ModelNet40 and ShapeNetPart as well as the real-world dataset KITTI and NuScenes. Experimental results show that the proposed SR-Adv achieves a state-of-the-art attack success rate and better imperceptibility by perturbing fewer points on 3D point clouds.
Shijun Zheng, Weiquan Liu, Chenglu Wen, Ming Cheng 0002, Ping Zhong 0001, Cheng Wang 0003
IEEE Trans. Intell. Transp. Syst.8
2024 Relation-Aware Weight Sharing in Decoupling Feature Learning Network for UAV RGB-Infrared Vehicle Re-Identification
abstract
Owing to the capacity of performing full-time target searches, cross-modality vehicle re-identification based on unmanned aerial vehicles (UAV) is gaining more attention in both video surveillance and public security. However, this promising and innovative research has not been studied sufficiently due to the issue of data inadequacy. Meanwhile, the cross-modality discrepancy and orientation discrepancy challenges further aggravate the difficulty of this task. To this end, we pioneer a cross-modality vehicle Re-ID benchmark named UAV Cross-Modality Vehicle Re-ID (UCM-VeID), containing 753 identities with16015RGB and13913infrared images. Moreover, to meet cross-modality discrepancy and orientation discrepancy challenges, we present a hybrid weights decoupling network (HWDNet) to learn the shared discriminative orientation-invariant features. For the first challenge, we proposed a hybrid weights siamese network with a well-designed weight restrainer and its corresponding objective function to learn both modality-specific and modality shared information. In terms of the second challenge, three effective decoupling structures with two pretext tasks are investigated to flexibly conduct orientation-invariant feature separation task. Comprehensive experiments are carried out to validate the effectiveness of the proposed method.
Jiahao Qi, Chen Chen 0152, Kangcheng Bin, Ping Zhong 0001
IEEE Trans. Multim.5
2023 A CNN with noise inclined module and denoise framework for hyperspectral image classification
abstract
Abstract Deep Neural Networks have been successfully applied in hyperspectral image classification. However, most of prior works adopt general deep architectures while ignore the intrinsic structure of the hyperspectral image, such as the physical noise generation. This would make these deep models unable to generate discriminative features and provide impressive classification performance. To leverage such intrinsic information, this work develops a novel deep learning framework with the noise inclined module and denoise framework for hyperspectral image classification. First, the spectral signature of hyperspectral image is modeled with the physical noise model to describe the high intra‐class variance of each class and great overlapping between different classes in the image. Then, a noise inclined module is developed to capture the physical noise within each object and a denoise framework is then followed to remove such noise from the object. Finally, the CNN with noise inclined module and the denoise framework is developed to obtain discriminative features and provides good classification performance of hyperspectral image. Experiments are conducted over two commonly used real‐world datasets and the experimental results show the effectiveness of the proposed method. The implementation of the proposed method and other compared methods could be accessed at https://github.com/shendu‐sw/noise‐physical‐framework .
Zhiqiang Gong, Ping Zhong 0001, Wen Yao 0001, Weien Zhou, Jiahao Qi, Panhe Hu
IET Image Process.2
2023 Self-Aligned Spatial Feature Extraction Network for UAV Vehicle Reidentification
abstract
Compared with existing vehicle reidentification (VeID) tasks conducted with datasets collected by fixed surveillance cameras, VeID for an unmanned aerial vehicle (UAV) is still under-explored and could be more challenging. Vehicles with the same color and type show extremely similar appearances from the UAV’s perspective so that mining fine-grained characteristics becomes necessary. Recent works tend to extract distinguishing information by regional features and component features. The former requires input images to be aligned and the latter entails detailed annotations, both of which are difficult to meet in UAV application. To extract efficient fine-grained features and avoid tedious annotating work, this letter develops an unsupervised self-aligned network consisting of three branches. The network introduced a self-alignment module to convert the input images with variable orientations to a uniform orientation, which is implemented under the constraint of a triple loss function designed with spatial features. On this basis, spatial features, obtained by vertical and horizontal segmentation methods, and global features are integrated to improve the representation ability in embedded space. Extensive experiments are conducted on UAV-VeID dataset, and our method achieves the best performance compared with recent reidentification (ReID) works.
Aihuan Yao, Jiahao Qi, Ping Zhong 0001
IEEE Geosci. Remote. Sens. Lett.3
2023 Boosting transferability of physical attack against detectors by redistributing separable attention
abstract
The research on attack transferability is of great importance as it can guide how to conduct an adversarial attack without knowing any information about target models. However, it remains challenging for adversarial examples to maintain a good attack transferability performance, especially for the black-box attack implemented in the physical world. To enhance black-box transferability of physical attacks on object detectors, we present a novel adversarial learning method to produce adversarial patches by redistributing separable attention maps. Concretely, we first develop smoothed multilayer attention maps by introducing serial composite transformations, which could suppress model-specific noise on the one hand, and cover objects to be concealed at various resolutions on the other hand. Besides, our method resorts to a scalable mask to separate object attention from the background and adjust their distribution with a novel loss function. Extensive experiments show that our approach outperforms state-of-the-art methods in both the digital space and the physical world. Our code is available at https://github.com/zhangyu13a/transPhyAtt .
Yu Zhang 0221, Zhiqiang Gong, Yichuang Zhang, Kangcheng Bin, Yongqian Li, Jiahao Qi, Ping Zhong 0001
Pattern Recognit.8
2023 Triplet Spectralwise Transformer Network for Hyperspectral Target Detection
abstract
Recently, deep learning methods have demonstrated their potentials in extracting spectral information for hyperspectral images and have been widely applied in hyperspectral target detection (HTD). However, prior deep learning methods, represented by the convolutional neural networks, mainly focus on the local information and representation, which cannot well capture the long-range dependence. Besides, limited target references cannot meet the need of massive labelled samples for the training process. This work develops a triplet spectral-wise transformer-based target detector (TSTTD) to deal with these problems. First, this work explores a novel triplet spectral-wise transformer network for HTD task, and a data augmentation method is utilized to construct sufficient and balanced training samples for balanced learning. The proposed network shows advantages in learning local features from multiple adjacent bands and global features with long-range dependence. Second, for improving the separability between targets and backgrounds, a novel inter-category separation and intra-category aggregation (ISIA) loss function is proposed, which joints the hard-negative-mining triplet loss and the binary cross entropy loss. Third, experimental results on six data sets show that our proposed method is effective in leading to excellent detection performance when compared with other state-of-the-art methods.
Jinyue Jiao, Zhiqiang Gong, Ping Zhong 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Attention Mask-Based Network With Simple Color Annotation for UAV Vehicle Re-Identification
abstract
Vehicle re-identification (VeID) has attracted a growing research interest in recent years, and excellent performance has been shown with fixed traffic cameras. However, vehicle ReID in aerial images taken by unmanned aerial vehicles (UAVs), possessing both variable locations and special viewpoints, is still under-explored. Recent works tend to extract meaningful local features by careful annotation, which are effective but time-consuming. In order to extract discriminative features and avoid tedious annotating work, this letter develops an attention mask (AM)-based network with simple color annotation for object enhancement and background reduction. The network makes full use of deep features obtained by a pretrained color classification network and then utilizes principal component analysis (PCA) as a mapping function to achieve AMs without partial annotation. Besides, we introduce weighted triplet loss (WTL) function to deal with the problem of great similarity between classes caused by overlook views of UAVs. The loss function concentrates more on negative pairs to facilitate the identification ability of network. Rich experiments are conducted on both UAV dataset and surveillance dataset, and our method achieves competitive performance compared with recent ReID works.
Aihuan Yao, Mengmeng Huang, Jiahao Qi, Ping Zhong 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Deep Manifold Embedding for Hyperspectral Image Classification
abstract
Deep learning methods have played a more important role in hyperspectral image classification. However, general deep learning methods mainly take advantage of the samplewise information to formulate the training loss while ignoring the intrinsic data structure of each class. Due to the high spectral dimension and great redundancy between different spectral channels in the hyperspectral image, these former training losses usually cannot work so well for the deep representation of the image. To tackle this problem, this work develops a novel deep manifold embedding method (DMEM) for deep learning in hyperspectral image classification. First, each class in the image is modeled as a specific nonlinear manifold, and the geodesic distance is used to measure the correlation between the samples. Then, based on the hierarchical clustering, the manifold structure of the data can be captured and each nonlinear data manifold can be divided into several subclasses. Finally, considering the distribution of each subclass and the correlation between different subclasses under data manifold, DMEM is constructed as the novel training loss to incorporate the special classwise information in the training process and obtain discriminative representation for the hyperspectral image. Experiments over four real-world hyperspectral image datasets have demonstrated the effectiveness of the proposed method when compared with general sample-based losses and showed superiority when compared with state-of-the-art methods.
Zhiqiang Gong, Weidong Hu, Xiaoyong Du 0002, Ping Zhong 0001, Panhe Hu
IEEE Trans. Cybern.4
2022 Dual Interactive Graph Convolutional Networks for Hyperspectral Image Classification
abstract
Recently, graph convolutional network (GCN) has progressed significantly and gained increasing attention in hyperspectral image (HSI) classification due to its impressive representation power. However, existing GCN-based methods do not give full consideration to the multiscale spatial information, since the convolution operations are governed by fixed neighborhood. As a result, their performances can be limited, particularly in the regions with diverse land cover appearances. In this article, we develop a new dual interactive GCN (DIGCN) which introduces the dual GCN branches to capture spatial information at different scales. More significantly, the dual interactive module is embedded across the GCN branches, so that the correlation of multiscale spatial information can be leveraged to refine the graph information. To be concrete, the edge information contained in one GCN branch can be refined by incorporating the feature representations from the other branch. Analogously, improved feature representations can be generated in one GCN branch by fusing the edge information from the other branch. As such, the refined graph information can help enhance the representation power of the model. Furthermore, to avoid the negative effects of the manually constructed graph, our proposed model adaptively learns a discriminative region-induced graph, which also accelerates the convolution operation. We comprehensively evaluate the proposed method on four commonly used HSI benchmark data sets, and the state-of-the-art results can be achieved when compared with several typical HSI classification methods.
Sheng Wan, Shirui Pan, Ping Zhong 0001, Xiaojun Chang, Jian Yang 0003, Chen Gong 0002
IEEE Trans. Geosci. Remote. Sens.3
2021 Unmixing-Based Underwater Target Detection for Hyperspectral Imagery
abstract
In this paper, we proposed a novel underwater target detection method based on hyperspectral unmixing methodology which composes of three different modules. The first module employed a classical anomaly detector to pick out the target-water mixed pixels from background, which can remove the influence of background pixels at the same time. Then, a bathymetric model-based autoencoder is designed in the second module to unmix target-water mixed pixels for attaining underwater target spectra without any prior information about water environment. Finally, we desgin a multi-criterion spectrum distance metric for target spectrum recognition and get the final detection result. Experimental results on simulated data sets demonstrate the effectiveness and efficiency of our method in comparison with the state-of-the-art underwater target detection methods.
Jiahao Qi, Aihuan Yao, Ping Zhong 0001
IGARSS4
2021 Hyperspectral Image Classification With Context-Aware Dynamic Graph Convolutional Network
abstract
In hyperspectral image (HSI) classification, spatial context has demonstrated its significance in achieving promising performance. However, conventional spatial context-based methods simply assume that spatially neighboring pixels should correspond to the same land-cover class, so they often fail to correctly discover the contextual relations among pixels in complex situations, and thus leading to imperfect classification results on some irregular or inhomogeneous regions such as class boundaries. To address this deficiency, we develop a new HSI classification method based on the recently proposed graph convolutional network (GCN), as it can flexibly encode the relations among arbitrarily structured non-Euclidean data. Different from traditional GCN, there are two novel strategies adopted by our method to further exploit the contextual relations for accurate HSI classification. First, since the receptive field of traditional GCN is often limited to fairly small neighborhood, we proposed to capture long-range contextual relations in HSI by performing successive graph convolutions on a learned region-induced graph which is transformed from the original 2-D image grids. Second, we refine the graph edge weight and the connective relationships among image regions simultaneously by learning the improved similarity measurement and the “edge filter,” so that the graph can be gradually refined to adapt to the representations generated by each graph convolutional layer. Such updated graph will in turn result in faithful region representations, and vice versa. The experiments carried out on four real-world benchmark data sets demonstrate the effectiveness of the proposed method.
Sheng Wan, Chen Gong 0002, Ping Zhong 0001, Shirui Pan, Jian Yang 0003
IEEE Trans. Geosci. Remote. Sens.3
2021 Statistical Loss and Analysis for Deep Learning in Hyperspectral Image Classification
abstract
Nowadays, deep learning methods, especially the convolutional neural networks (CNNs), have shown impressive performance on extracting abstract and high-level features from the hyperspectral image. However, the general training process of CNNs mainly considers the pixelwise information or the samples' correlation to formulate the penalization while ignores the statistical properties especially the spectral variability of each class in the hyperspectral image. These sample-based penalizations would lead to the uncertainty of the training process due to the imbalanced and limited number of training samples. To overcome this problem, this article characterizes each class from the hyperspectral image as a statistical distribution and further develops a novel statistical loss with the distributions, not directly with samples for deep learning. Based on the Fisher discrimination criterion, the loss penalizes the sample variance of each class distribution to decrease the intraclass variance of the training samples. Moreover, an additional diversity-promoting condition is added to enlarge the interclass variance between different class distributions, and this could better discriminate samples from different classes in the hyperspectral image. Finally, the statistical estimation form of the statistical loss is developed with the training samples through multivariant statistical analysis. Experiments over the real-world hyperspectral images show the effectiveness of the developed statistical loss for deep learning.
Zhiqiang Gong, Ping Zhong 0001, Weidong Hu
IEEE Trans. Neural Networks Learn. Syst.2
2020 Multiscale Dynamic Graph Convolutional Network for Hyperspectral Image Classification
abstract
Convolutional neural network (CNN) has demonstrated impressive ability to represent hyperspectral images and to achieve promising results in hyperspectral image classification. However, traditional CNN models can only operate convolution on regular square image regions with fixed size and weights, and thus, they cannot universally adapt to the distinct local regions with various object distributions and geometric appearances. Therefore, their classification performances are still to be improved, especially in class boundaries. To alleviate this shortcoming, we consider employing the recently proposed graph convolutional network (GCN) for hyperspectral image classification, as it can conduct the convolution on arbitrarily structured non-Euclidean data and is applicable to the irregular image regions represented by graph topological information. Different from the commonly used GCN models that work on a fixed graph, we enable the graph to be dynamically updated along with the graph convolution process so that these two steps can be benefited from each other to gradually produce the discriminative embedded features as well as a refined graph. Moreover, to comprehensively deploy the multiscale information inherited by hyperspectral images, we establish multiple input graphs with different neighborhood scales to extensively exploit the diversified spectral-spatial correlations at multiple scales. Therefore, our method is termed multiscale dynamic GCN (MDGCN). The experimental results on three typical benchmark data sets firmly demonstrate the superiority of the proposed MDGCN to other state-of-the-art methods in both qualitative and quantitative aspects.
Sheng Wan, Chen Gong 0002, Ping Zhong 0001, Bo Du 0001, Lefei Zhang, Jian Yang 0003
IEEE Trans. Geosci. Remote. Sens.3
2020 Multiple Instance Learning for Multiple Diverse Hyperspectral Target Characterizations
abstract
A practical hyperspectral target characterization task estimates a target signature from imprecisely labeled training data. The imprecisions arise from the characteristics of the real-world tasks. First, accurate pixel-level labels on training data are often unavailable. Second, the subpixel targets and occluded targets cause the training samples to contain mixed data and multiple target types. To address these imprecisions, this paper proposes a new hyperspectral target characterization method to produce diverse multiple hyperspectral target signatures under a multiple instance learning (MIL) framework. The proposed method uses only bag-level training samples and labels, which solves the problems arising from the mixed data and lack of pixel-level labels. Moreover, by formulating a multiple characterization MIL and including a diversity-promoting term, the proposed method can learn a set of diverse target signatures, which solves the problems arising from multiple target types in training samples. The experiments on hyperspectral target detections using the learned multiple target signatures over synthetic and real-world data show the effectiveness of the proposed method.
Ping Zhong 0001, Zhiqiang Gong, Jiaxin Shan
IEEE Trans. Neural Networks Learn. Syst.1
2019 An End-to-End Joint Unsupervised Learning of Deep Model and Pseudo-Classes for Remote Sensing Scene Representation
abstract
This work develops a novel end-to-end deep unsupervised learning method based on convolutional neural network (CNN) with pseudo-classes for remote sensing scene representation. First, we introduce center points as the centers of the pseudo classes and the training samples can be allocated with pseudo labels based on the center points. Therefore, the CNN model, which is used to extract features from the scenes, can be trained supervised with the pseudo labels. Moreover, a pseudo-center loss is developed to decrease the variance between the samples and the corresponding pseudo center point. The pseudo-center loss is important since it can update both the center points with the training samples and the CNN model with the center points in the training process simultaneously. Finally, joint learning of the pseudo-center loss and the pseudo softmax loss which is formulated with the samples and the pseudo labels is developed for unsupervised remote sensing scene representation to obtain discriminative representations from the scenes. Experiments are conducted over two commonly used remote sensing scene datasets to validate the effectiveness of the proposed method and the experimental results show the superiority of the proposed method when compared with other state-of-the-art methods.
Zhiqiang Gong, Ping Zhong 0001, Weidong Hu, BingWei Hui
IJCNN2
2019 A CNN With Multiscale Convolution and Diversified Metric for Hyperspectral Image Classification
abstract
Recently, researchers have shown the powerful ability of deep methods with multilayers to extract high-level features and to obtain better performance for hyperspectral image classification. However, a common problem of traditional deep models is that the learned deep models might be suboptimal because of the limited number of training samples, especially for the image with large intraclass variance and low interclass variance. In this paper, novel convolutional neural networks (CNNs) with multiscale convolution (MS-CNNs) are proposed to address this problem by extracting deep multiscale features from the hyperspectral image. Moreover, deep metrics usually accompany with MS-CNNs to improve the representational ability for the hyperspectral image. However, the usual metric learning would make the metric parameters in the learned model tend to behave similarly. This similarity leads to obvious model's redundancy and, thus, shows negative effects on the description ability of the deep metrics. Traditionally, determinantal point process (DPP) priors, which encourage the learned factors to repulse from one another, can be imposed over these factors to diversify them. Taking advantage of both the MS-CNNs and DPP-based diversity-promoting deep metrics, this paper develops a CNN with multiscale convolution and diversified metric to obtain discriminative features for hyperspectral image classification. Experiments are conducted over four real-world hyperspectral image data sets to show the effectiveness and applicability of the proposed method. Experimental results show that our method is better than original deep models and can produce comparable or even better classification performance in different hyperspectral image data sets with respect to spectral and spectral-spatial features.
Zhiqiang Gong, Ping Zhong 0001, Yang Yu 0006, Weidong Hu, Shutao Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2018 Diversifying Deep Multiple Choices for Remote Sensing Scene Classification
abstract
Recently, deep models have shown powerful ability for remote sensing scene representation. However, the training process of these deep methods requires large amount of labelled samples while usual remote sensing image datasets cannot provide enough training samples. Therefore, the learned model is usually suboptimal. To solve the problem, this work focuses on obtaining multiple choices by training multiple models simultaneously, and then the human oracle can choose a proper one from these choices. However, training several models separately usually makes the obtained results similar. This paper tries to diversify the obtained choices by encouraging the obtained choices to repulse from each other. Experiments are conducted on Ucmerced Land Use dataset to validate the effectiveness of the proposed method to provide multiple diversified choices.
Zhiqiang Gong, Ping Zhong 0001, Jiaxin Shan, Weidong Hu
IGARSS2
2018 An Unsupervised Convolutional Feature Fusion Network for Deep Representation of Remote Sensing Images
abstract
Unsupervised learning of a convolutional neural network (CNN) is a feasible method to represent and classify remote sensing images, where labeling the observed data to prepare training samples is a highly expensive and time-consuming task. In this letter, we propose an unsupervised convolutional feature fusion network to formulate an easy-to-train but effective CNN representation of remote sensing images. The efficiency and effectiveness are derived from the following two aspects. First, the proposed method trains a deep CNN through unsupervised learning of each CNN layer in a greedy layer-wise manner, which makes the training relatively easy and efficient. Second, the feature fusion strategy in the proposed network can effectively use both the information from individual layers and the important interactions between different layers. As a result, the proposed network requires only several layers to obtain comparable or even better results than very deep networks. The experiments on unsupervised deep representations and the classification of remote sensing images demonstrate the efficiency and effectiveness of the proposed method.
Yang Yu 0006, Zhiqiang Gong, Cheng Wang 0003, Ping Zhong 0001
IEEE Geosci. Remote. Sens. Lett.4
2018 Diversity-Promoting Deep Structural Metric Learning for Remote Sensing Scene Classification
abstract
Deep models with multiple layers have demonstrated their potential in learning abstract and invariant features for better representation and classification of remote sensing images. Moreover, metric learning (ML) is usually introduced into the deep models to further increase the discrimination of deep representations. However, the usual deep ML methods treat the training samples in each training batch in the stochastic gradient descent-based learning procedure independently, and thus, they neglect the important contextual (structural) information in the training samples. In this paper, we first introduce deep structural ML (DSML) into the literature of remote sensing scene classification and specifically capture and use the structural information during the training on the remote sensing images. Further analysis demonstrates that DSML usually makes many learned metric parameters similar. This similarity leads to obvious model redundancy and thus decreases the representational ability of the model. To address this problem, this paper proposes a new diversity-promoting DSML (D-DSML) method by regularizing the learning procedure by a diversity-promoting prior over the parameter factors. The proposed D-DSML encourages the parameter factors to be uncorrelated, such that each factor can model unique information, and thus, the model's description ability and classification performance would be significantly improved. Experiments over six real-world remote sensing scene data sets demonstrate that the proposed method obtains much better results than those obtained by the original deep models and has comparable or even better performances when compared with state-of-the-art methods.
Zhiqiang Gong, Ping Zhong 0001, Yang Yu 0006, Weidong Hu
IEEE Trans. Geosci. Remote. Sens.2
2017 Unsupervised Representation Learning with Deep Convolutional Neural Network for Remote Sensing Images
Yang Yu 0006, Zhiqiang Gong, Ping Zhong 0001, Jiaxin Shan
ICIG (2)3
2017 Diversified deep structural metric learning for land use classification in remote sensing images
abstract
In this work, a diversified deep structural metric learning is proposed for remote sensing image classification. Firstly, a deep structural metric learning is introduced to take full advantage of structural information of training batches. Secondly, we impose a diversity regularization over the factors of deep structural metric learning to encourage them to be uncorrelated, such that each factor tends to model unique information during the training phase and all factors sums up to capture a larger proportion of information. The diversified model could benefit the classification of remote sensing images. Experiments are conducted on two real-world remote sensing image datasets to evaluate the effectiveness and wide applicability of the proposed approach. The results show that our proposed method can obtain comparable or even better results on remote sensing image classification when compared with the recent results.
Zhiqiang Gong, Ping Zhong 0001, Yang Yu 0006, Weidong Hu
IGARSS2
2017 Balanced data driven sparsity for unsupervised deep feature learning in remote sensing images classification
abstract
There are many attempts that utilize deep learning methods to solve the problem of classification in remote sensing images. Convolutional Neural Networks (CNN) have made very good performance for various visual tasks, and marked their important place in all deep learning models. However, for some classification tasks of remote sensing images, CNN could not demonstrate their full potential because of lacking large amounts of labeled training data. Some efforts have been made to combine CNN with unlabeled data to tackle the problem by performing unsupervised learning. In this work we propose the balanced data driven sparsity to help train CNN in an unsupervised way. The experiments over the real world remote sensing images demonstrate that the proposed method improves the performance of the recent methods.
Yang Yu 0006, Ping Zhong 0001, Zhiqiang Gong
IGARSS2
2017 Combining depth-skeleton feature with sparse coding for action recognition
Hanling Zhang, Ping Zhong 0001, Jiale He, Chenxing Xia
Neurocomputing2
2017 Learning to Diversify Deep Belief Networks for Hyperspectral Image Classification
abstract
In the literature of remote sensing, deep models with multiple layers have demonstrated their potentials in learning the abstract and invariant features for better representation and classification of hyperspectral images. The usual supervised deep models, such as convolutional neural networks, need a large number of labeled training samples to learn their model parameters. However, the real-world hyperspectral image classification task provides only a limited number of training samples. This paper adopts another popular deep model, i.e., deep belief networks (DBNs), to deal with this problem. The DBNs allow unsupervised pretraining over unlabeled samples at first and then a supervised fine-tuning over labeled samples. But the usual pretraining and fine-tuning method would make many hidden units in the learned DBNs tend to behave very similarly or perform as “dead” (never responding) or “potential over-tolerant” (always responding) latent factors. These results could negatively affect description ability and thus classification performance of DBNs. To further improve DBN's performance, this paper develops a new diversified DBN through regularizing pretraining and fine-tuning procedures by a diversity promoting prior over latent factors. Moreover, the regularized pretraining and fine-tuning can be efficiently implemented through usual recursive greedy and back-propagation learning framework. The experiments over real-world hyperspectral images demonstrated that the diversity promoting prior in both pretraining and fine-tuning procedure lead to the learned DBNs with more diverse latent factors, which directly make the diversified DBNs obtain much better results than original DBNs and comparable or even better performances compared with other recent hyperspectral image classification methods.
Ping Zhong 0001, Zhiqiang Gong, Shutao Li 0001, Carola-Bibiane Schönlieb
IEEE Trans. Geosci. Remote. Sens.1
2016 A DBN-crf for spectral-spatial classification of hyperspectral data
abstract
This work shows how to improve hyperspectral image classification through using both a deep representation and contextual information. To implement this objective, this work proposes a new Conditional Random Field (CRF) model (named DBN-CRF) with potentials defined over deep features produced by the Deep Belief Networks (DBNs). The newly formulated DBN-CRF model takes advantage of strength of the DBNs in learning a good representation and the ability of CRFs to model contextual (spatial) information in both observations and labels. Within a piecewise training framework, an efficient training method is proposed to train the whole DBN-CRF model end-to-end. This means that parameters in DBN and CRF can be jointly trained and thus the proposed method can fully use the strength of both DBN and CRF. Moreover, in the proposed training method, the end-to-end training can be implemented with a standard back-propagation algorithm, avoiding the repeated inference usually involved in CRF training and thus is computationally efficient. Experiments on real-world hyperspectral data show that our method outperforms the most recent approaches in hyperspectral image classification.
Ping Zhong 0001, Zhiqiang Gong, Carola-Bibiane Schönlieb
ICPR1
2016 Spectral-spatial classification of hyperspectral images with Gaussian process
abstract
In this paper, a spectral-spatial classification method with Gaussian process was proposed for hyperspectral image classification. This method exploits the relationship among adjacent pixels and integrates it into spectral information to obtain spectral-spatial classification. In the proposed approach, the spatial information of a single pixel is weighted by the cosine similarity value between the adjacent pixels in the neighborhood. Experiments were conducted on the AVIRIS Indian Pines data set to evaluate the performance of the proposed approach. And the results demonstrated the effectiveness of the proposed methods to improve the classification performance by consideration of the spatial relationship between adjacent pixels in the hyperspectral image.
Shujin Sun, Ping Zhong 0001, Huaitie Xiao, Zhiqiang Gong, Runsheng Wang
IGARSS2
2016 Integrating global and local visual features with semantic hierarchies for two-level image annotation
Zhiming Qian, Ping Zhong 0001
Neurocomputing2
2015 Personalized image annotation via class-specific cross-domain learning
Zhiming Qian, Ping Zhong 0001, Runsheng Wang
Signal Process. Image Commun.2
2015 Tag Refinement for User-Contributed Images via Graph Learning and Nonnegative Tensor Factorization
abstract
Social image tagging systems mostly suffer from poor performance for image retrieval due to the noisy and incomplete correspondences between user-contributed images and their associated tags. In this letter, we aim to refine tag allocations in the social tagging data provided by these systems. In particular, we propose to harness the tagged and untagged data with a two-stage strategy according to different types of data relations, i.e. item similarity defined by prior knowledge and item co-occurrence learned from data statistics. To solve the sparsity problem, we first introduce a new graph learning (GL) method for enriching the tagging data according to item similarities. Then, we develop a method of nonnegative tensor factorization (NTF) for learning more coherent ternary relations among users, images and tags coupled by the manifold constraints learned from item co-occurrences. Experimental results with the tagging data from the NUS-WIDE dataset have been reported to validate the effectiveness of the proposed method.
Zhiming Qian, Ping Zhong 0001, Runsheng Wang
IEEE Signal Process. Lett.2
2015 Active Learning With Gaussian Process Classifier for Hyperspectral Image Classification
abstract
Gaussian process (GP) classifiers represent a powerful and interesting theoretical framework for the Bayesian classification of hyperspectral images. However, the collection of labeled samples is time consuming and costly for hyperspectral data, and the training samples available are often not enough for an adequate learning of the GP classifier. Moreover, the computational cost of performing inference using GP classifiers scales cubically with the size of the training set. To address the limitations of GP classifiers for hyperspectral image classification, reducing the label cost and keeping the training set in a moderate size, this paper introduces an active learning (AL) strategy to collect the most informative training samples for manual labeling. First, we propose three new AL heuristics based on the probabilistic output of GP classifiers aimed at actively selecting the most uncertain and confusing candidate samples from the unlabeled data. Moreover, we develop an incremental model updating scheme to avoid the repeated training of the GP classifiers during the AL process. The proposed approaches are tested on the classification of two realworld hyperspectral data. Comparison with random sampling method reveals a better accuracy gain and faster convergence with the number of queries, and comparison with recent active learning approaches shows a competitive performance. Experimental results also verified the efficiency of the incremental model updating scheme.
Shujin Sun, Ping Zhong 0001, Huaitie Xiao, Runsheng Wang
IEEE Trans. Geosci. Remote. Sens.2
2014 $L_{1/2}$-Regularized Deconvolution Network for the Representation and Restoration of Optical Remote Sensing Images
abstract
Optical remote sensing images of land cover are composed of many natural and man-made objects and thus exhibit rich image features ranging from low to high levels. Extracting such a wide range of features to represent the optical remote sensing images, beyond edge primitives, is a long-standing goal in the remote sensing and vision research community. The recently proposed deconvolution network (DN) can effectively learn and capture features in a variety of forms: low-level edges, midlevel edge junctions, high-level object parts, and complete objects. The approach is based on the convolutional decomposition of images under an L1sparsity constraint. Unfortunately, the L1regularizer cannot enforce further sparsity, hence limiting the practical efficacy of the DN in optical remote sensing representation and processing. In this paper, we extend the DN by incorporating the L1/2sparsity constraint, which we name the L1/2-DN. The L1/2regularizer not only induces sparsity but is also a better choice among Lq, (01/2-DN algorithm is more efficient, provides a sparser representation, and results in more accurate recovery than the DN. We illustrate the utility of our method on a wide range of optical remote sensing images and compare our results to those yielded by other state-of-the-art methods.
Jun Zhang 0067, Ping Zhong 0001, Yangtai Chen, Shuohao Li
IEEE Trans. Geosci. Remote. Sens.2
2014 Jointly Learning the Hybrid CRF and MLR Model for Simultaneous Denoising and Classification of Hyperspectral Imagery
abstract
Despite much advance obtained in hyperspectral image sensors, they are still very sensitive to the noise, and thus cause the captured data to carry enough noise to degrade the classification results. The traditional approach first resorts to image denoising and then feeds the denoised image into a classifier. However, such a straightforward approach, treating denoising and classification separately, suffers greatly from neglecting their impacts on each other. This paper presents a new simultaneous denoising and classification method in the pursuit of cleanest image for optimal classification in the sense of given task evaluation measures. To obtain this objective, we develop a hybrid conditional random field (CRF) (for denoising) and multinomial logistic regression (MLR) (for classification) model at first, and then to train the proposed hybrid model, we propose a new joint learning method, which can effectively capture the impacts of denoising on classification, or vice versa, the effects of classification on denoising. Through the proposed joint learning method, the CRF and MLR, and thus the denoising and classification procedure, can be tightly combined. Moreover, the proposed joint learning method can directly optimize a large class of application specific performance measures including both the linear measures, such as the overall accuracy, and the nonlinear measures, such as kappa statistics. Meanwhile, the consistency between the criteria of model learning and model application has the potential to obtain the denoised image, which is at its best for optimal classification in the sense of the given measure. The extensive experiments of simultaneous denoising and classification tasks are conducted in both simulated and real noisy conditions to test our jointly learned model, which are shown to outperform the conventional methods of treating the two tasks independently.
Ping Zhong 0001, Runsheng Wang
IEEE Trans. Neural Networks Learn. Syst.1
2013 Multiple-Spectral-Band CRFs for Denoising Junk Bands of Hyperspectral Imagery
abstract
Denoising of hyperspectral imagery in the domain of imaging spectroscopy by conditional random fields (CRFs) is addressed in this work. For denoising of hyperspectral imagery, the strong dependencies across spatial and spectral neighbors have been proved to be very useful. Many available hyperspectral image denoising algorithms adopt multidimensional tools to deal with the problems and thus naturally focus on the use of the spectral dependencies. However, few of them were specifically designed to use the spatial dependencies. In this paper, we propose a multiple-spectral-band CRF (MSB-CRF) to simultaneously model and use the spatial and spectral dependencies in a unified probabilistic framework. Furthermore, under the proposed MSB-CRF framework, we develop two hyperspectral image denoising algorithms, which, thanks to the incorporated spatial and spectral dependencies, can significantly remove the noise, while maintaining the important image details. The experiments are conducted in both simulated and real noisy conditions to test the proposed denoising algorithms, which are shown to outperform the popular denoising methods described in the previous literatures.
Ping Zhong 0001, Runsheng Wang
IEEE Trans. Geosci. Remote. Sens.1
2011 Modeling and Classifying Hyperspectral Imagery by CRFs With Sparse Higher Order Potentials
abstract
Hyperspectral images exhibit strong dependencies across spatial and spectral neighbors, which have been proved to be very useful for hyperspectral image classification. The recently defined conditional random field (CRF) can effectively model and use the dependencies for classification of hyperspectral images in a unified probabilistic framework. However, in order to be computationally tractable, the usual CRFs are limited to incorporate only pairwise potentials. Thus, the usual CRFs can capture only pairwise interactions and neglect higher order dependencies, which are potentially useful high-level properties particularly for the classification of hyperspectral image consisting of complex components. This paper overcomes this limitation by developing hyperspectral image classification algorithm based on a CRF with sparse higher order potentials, which are specially designed to incorporate complex characteristics of hyperspectral images. To efficiently implement the CRF model at training step, this paper develops an efficient local method under the piecewise training framework, while at inference step, this proposes a simple strategy to combine the piecewisely trained model to overcome the possible over-counting problems. Moreover, the combined model with the specially defined potentials can be efficiently inferred by graph cut method. Experiments on the real-world data attest to the accuracy, effectiveness, and efficiency of the proposed model on modeling and classifying hyperspectral images.
Ping Zhong 0001, Runsheng Wang
IEEE Trans. Geosci. Remote. Sens.1
2010 Learning Conditional Random Fields for Classification of Hyperspectral Images
abstract
Hyperspectral images exhibit strong dependencies across spatial and spectral neighbors, which have been proved to be very useful for hyperspectral image classification. State-of-the-art hyperspectral image classification algorithms use the dependencies in a heuristic way or in probabilistic frameworks but impose unreasonable assumptions on observed data. In this paper, we formulate a conditional random field (CRF) to replace such heuristics and unreasonable assumptions for the classification of hyperspectral images. Moreover, because of avoiding explicit modeling of the observed data, the proposed method can incorporate the classification of hyperspectral images with different statistics characteristics into a unified probabilistic framework. Since the usual classification task for hyperspectral images needs the proposed CRF to be trained on local samples, available global training methods cannot be directly used. Under piecewise training framework, this paper develops an efficient local method to train the CRF. It is efficiently implemented through separated training of simple classifiers defined by corresponding potentials. However, the independent classifier training may lead to over-counting problems during inference. So we further propose a strategy to combine the independently trained models to obtain final CRF model. Experiments on real-world hyperspectral data show that our algorithm is competitive with the most recent results in hyperspectral image classification.
Ping Zhong 0001, Runsheng Wang
IEEE Trans. Image Process.1
2008 Dynamic Learning of SMLR for Feature Selection and Classification of Hyperspectral Data
abstract
Feature selection is an important task in the analysis of hyperspectral data. Recently developed methods for learning sparse classifiers, which combine the automatic feature selection and classifier design, established themselves among the state of the art in the literature of machine learning. In this letter, the sparse multinomial logistic regression (SMLR) is introduced into the community of remote sensing and is utilized for the feature selection in the classification of hyperspectral data. To relieve the heavy degeneration of classification performance caused by the characteristics of the hyperspectral data and the oversparsity when the SMLR selects a small feature subset, we develop a dynamic learning framework to train the SMLR. Experimental results attest to the effectiveness of the proposed method.
Ping Zhong 0001, Peng Zhang 0079, Runsheng Wang
IEEE Geosci. Remote. Sens. Lett.1
2008 Learning Sparse CRFs for Feature Selection and Classification of Hyperspectral Imagery
abstract
Feature selection is an important task in hyperspectral data analysis. This paper presents a sparse conditional random field (SCRF) model to select relevant features for the classification of hyperspectral images and, meanwhile, to exploit the contextual information in the form of spatial dependences in the images. The sparsity arises from the use of a Laplacian prior on the CRF parameters, which encourages the parameter estimates to be either significantly large or exactly zero. To joint the feature selection and classifier design, this paper develops an efficient sparse training method, which divides the training of SCRF into the sparse trainings of two simpler classifiers. Experiments on the real-world hyperspectral image attest to the accuracy, sparsity, and efficiency of the proposed model.
Ping Zhong 0001, Runsheng Wang
IEEE Trans. Geosci. Remote. Sens.1
2007 Using Combination of Statistical Models and Multilevel Structural Information for Detecting Urban Areas From a Single Gray-Level Image
abstract
With the complex building composition and imaging condition, urban areas show versatile characteristics in remote sensing images. In the literature of land-cover analysis, many algorithms utilize the features with structural information to characterize urban areas. Typically, these are more successful on some types of imagery than others, since they usually use only one kind or a few kinds of structural information. On the other hand, since levels of development in neighboring areas are not statistically independent, the multiple features (encoding the multilevel structural information) of each site in urban area depend on that of neighboring sites. In this paper, a new-come discriminative model, i.e., conditional random field (CRF), is introduced to learn the dependencies and fuse the multilevel structural information to obtain the essential detection. To meet the higher needs of some users, we introduce a two-component-based Markov random field model and show how to integrate it tightly with CRF model to refine the results from essential detection. Experiments on a wide range of images show that our algorithms are competitive with recent results in urban area detection
Ping Zhong 0001, Runsheng Wang
IEEE Trans. Geosci. Remote. Sens.1
2007 A Multiple Conditional Random Fields Ensemble Model for Urban Area Detection in Remote Sensing Optical Images
abstract
With complex building composition and imaging condition, urban areas show versatile characteristics in remote sensing optical images. It demonstrates that multiple features should be utilized to characterize urban areas. On the other hand, since levels of development in neighboring areas are not statistically independent, the features of each urban area site depend on those of neighboring sites. In this paper, we present a multiple conditional random fields (CRFs) ensemble model to incorporate multiple features and learn their contextual information. This model involves two aspects: one is to use a CRF as the base classifier to automatically generate a set of CRFs by changing input features, and the other is to integrate the set of CRFs by defining a conditional distribution. The model has some distinct merits: each CRF component models a kind of feature, so that the ensemble model can learn different aspects of training data. Moreover, it lets the ensemble model search in a wide solution space. The ensemble model can also avoid the well-known overfitting problem of a single CRF, i.e., the many features may cause the redundancy of irrelevant information and result in counter-effect. Experiments on a wide range of images show that our ensemble model produces higher detection accuracy than single CRF and is also competitive with recent results in urban area detection.
Ping Zhong 0001, Runsheng Wang
IEEE Trans. Geosci. Remote. Sens.1