EDBT 2026 Demo / reviewers in the wild / expert
Jie Feng 0003
dblp:24/7003-3
· DBLP profile ↗
83ranked-venue papers
27as first author
61since 2021 · last 2026
0000-0002-8032-7542ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 58 · 20 first-author · 42 since 2021Artificial intelligence and machine learning · 20 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hyperspectral image classification based on multi-scale equivariant feature extraction and geometric equivariant self-attention
Jinhong Ren, Ronghua Shang, Kun Xie 0011, Jie Feng 0003, Dongzhu Feng |
Expert Syst. Appl. | 5 |
| 2026 | Cross-scene hyperspectral image classification based on cross-domain feature extraction and category decision collaborative optimizationabstractCross-scene hyperspectral image classification aims to enable the model to complete the classification of unlabeled target domain data by learning from labeled source domain data. Aiming at the problem that most current cross-scene hyperspectral image classification algorithms do not fully consider the cross-domain feature representation and category decision boundary optimization, a cross-domain Feature Extraction and Category Decision collaborative optimization (FECD) network is proposed. First, an adaptive feature discovery based on dynamic masks is designed. In this mechanism, the dynamically scaled masks are applied to the 3D representation of source and target domain data to generate an informative feature space and enhance the cross-scene discrimination potential of the model. Second, a dual-stream convolutional cross-domain feature extraction based on Mamba stream and ViT stream is constructed. Long sequence modeling and convolutional attention mechanisms are used to capture cross-domain spectral features between pixel, and self-attention mechanisms and multi-scale convolution are used to excavate cross-domain space patterns of pixel. Finally, a category decision based on the co-optimization of dual-stream classifiers is implemented. The spectral and spatial boundaries learned by the dual streams are fused to optimize the category decision. Therefore, the risk of false labeling is avoided while obtaining more accurate category boundaries. Compared with seven state-of-the-art algorithms on three widely used datasets, FECD obtains better categorization results on three categorization metrics: OA, AA, and Kappa. Ronghua Shang, Yangyang Li 0001, Jie Feng 0003, Songhua Xu |
Expert Syst. Appl. | 4 |
| 2026 | Domain-consistent networks for cross-scene hyperspectral image classification
Ronghua Shang, Yangyang Li 0001, Jie Feng 0003, Songhua Xu |
Neurocomputing | 4 |
| 2026 | Noise correction and distribution fine-tuning for long-tailed partial multi-label learning
Jingyu Zhong, Ronghua Shang, Jie Feng 0003 |
Pattern Recognit. | 5 |
| 2026 | Correlation-Induced Negative Suppression Disambiguation Loss for Partial Multi-Label Image ClassificationabstractPartial multi-label image classification (PMLIC) learns from typical weak supervision, where each image is labeled with a set of candidate labels, only some of which are correct. We find that noisy labels generate conflicting gradient signals that disrupt the learning of latent true labels, causing the model to prefer learning clean negative labels that provide consistent supervisory signals, thereby hindering disambiguation. Meanwhile, noisy labels cause the model to activate misattributed pixel regions, which interfere with feature pattern extraction, leading to inaccurate label correlation. In this paper, we propose a PMLIC framework that constructs a correlation-induced negative suppression disambiguation loss (CoNeS). First, we exploit the property that networks tend to learn clean labels first by extracting class activation maps to identify and screen misattributed pixel regions. Meanwhile, we aggregate noise-disturbed feature patterns into more expressive representations via k-means clustering and construct accurate label correlations to aid disambiguation. In addition, we design the negative suppression disambiguation loss to focus the model on disambiguation by introducing a weight distribution to suppress the contribution of negative labels. This weighting distribution can be adaptively inferred by a closed-form solution. Extensive experiments demonstrate that the CoNeS framework achieves significant advantages over current state-of-the-art methods. Specifically, it achieves average mAP improvements of 1.26%, 2.74%, 0.85%, and 0.33% on the VOC 2007, MS-COCO, VG-256, and CUB-200 datasets at different resolutions and noise rates. Code has been made available at https://github.com/zhongjingyu1/CoNeS. Jingyu Zhong, Ronghua Shang, Shasha Mao, Jinhong Ren, Jie Feng 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | MCIB: Multi-Modal Complementary Information Bottleneck for Hyperspectral and LiDAR ClassificationabstractThe effective fusion of multi-modal remote sensing images, particularly hyperspectral imagery (HSI) and light detection and ranging (LiDAR) data, is pivotal for accurate land use and land cover (LULC) classification. However, this process is hindered by two inherent challenges: pervasive data redundancy and the underutilization of cross-modal complementarity, largely due to the lack of a unifying theoretical framework. To address these limitations, we propose the multi-modal complementary information bottleneck (MCIB) framework, which extends the IB principle to learn compact, sufficient, and complementary representations for multi-modal scenes. From a theoretical perspective, we formalize the MCIB objective and introduce structured priors to derive tractable information-theoretic bounds, providing a principled and computationally feasible approach to reduce redundancy and enhance complementarity simultaneously. Building on the obtained theoretical insights, we design an end-to-end variational optimization strategy with a novel supervised conditional InfoNCE (SCInfoNCE). Efficiently reusing existing model components, this new supervised contrastive method optimizes the conditional mutual information terms crucial for synergy. Extensive experiments on benchmark HSI-LiDAR datasets demonstrate superior classification performance of MCIB. This work not only fills a theoretical gap in multi-modal representation learning, but offers a robust and principled solution for LULC classification using complex heterogeneous remote sensing images. Hao Zhu 0009, Bo Yang 0047, Changzhe Jiao, Jie Feng 0003, Jinjian Wu |
IEEE Trans. Image Process. | 5 |
| 2025 | Partially Matching Submap Helps: Uncertainty Modeling and Propagation for Text to Point Cloud Localization
Mingtao Feng, Longlong Mei, Jianqiao Luo, Fenghao Tian, Jie Feng 0003, Weisheng Dong, Yaonan Wang 0001 |
ICCV | 6 |
| 2025 | Group-spectral superposition and position self-attention transformer for hyperspectral image classification
Mingwei Hu, Sihan Hou, Ronghua Shang, Jie Feng 0003, Songhua Xu |
Expert Syst. Appl. | 5 |
| 2025 | Adaptive Homophily Clustering: Structure Homophily Graph Learning With Adaptive Filter for Hyperspectral ImageabstractHyperspectral image (HSI) clustering is a fundamental yet challenging task that typically operates without training labels. Recent advancements in deep graph clustering methods have shown promise for HSI due to their ability to effectively encode spatial structural information. However, limitations such as inadequate utilization of structural information, poor feature representation, and weak graph update capabilities hinder their performance. In this article, we propose an adaptive homophily structure graph clustering (AHSGC) method for HSI. Our approach begins with the generation of homogeneous regions to process HSI and construct the initial graph. Next, we design an adaptive filter graph encoder that captures both high and low-frequency features for subsequent processing. We then develop a graph embedding clustering self-training decoder using KL Divergence to generate pseudo-labels for network training. To enhance graph learning, we introduce homophily-enhanced structure learning, which updates the graph based on the clustering task. This involves estimating node connections through orient correlation estimation and dynamically adjusting graph edges via graph edge sparsification. Finally, we implement joint network optimization to facilitate self-training and graph updates, with K-means used to express latent features. The clustering accuracy on three datasets is 83.60%, 63.65%, and 86.03%, the FLOPs are 3.57G, 30.62G, and 2.95G. The source code will be available athttps://github.com/DY-HYX. Yao Ding 0010, Weijie Kang, Aitao Yang, Junyang Zhao, Jie Feng 0003, Danfeng Hong, Qinghe Zheng |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | SA-MixNet: Structure-Aware Mixup and Invariance Learning for Scribble-Supervised Road Extraction in Remote Sensing ImagesabstractMainstreamed weakly supervised road extractors rely on highly confident pseudo-labels propagated from scribbles, and their performance often degrades gradually as the image scenes tend to vary. We argue that such degradation is due to the poor model’s invariance to scenes with different complexities, whereas existing solutions to this problem are commonly based on crafted priors that cannot be derived from scribbles. To eliminate the reliance on such priors, we propose a novel structure-aware mixup and invariance learning framework (SA-MixNet) for weakly supervised road extraction that improves the model invariance in a data-driven manner. Specifically, we design a structure-aware mixup (SA-Mix) scheme to paste road regions from one image onto another to create an image scene with increased complexity while preserving the road’s structural integrity. Then, an invariance regularization is imposed on the predictions of constructed and origin images to minimize their conflicts, which thus forces the model to behave consistently in various scenes. Moreover, a discriminator-based regularization is designed to enhance connectivity while preserving the structure of roads. Combining these designs, our framework demonstrates superior performance on the DeepGlobe, Wuhan, and Massachusetts datasets, outperforming the state-of-the-art techniques by 1.47%, 2.12%, and 4.09%, respectively, in IoU metrics, and showing its potential as a plug-and-play solution. Our source code is available athttps://github.com/xdu-jjgs. Jie Feng 0003, Junpeng Zhang 0002, Weisheng Dong, Dingwen Zhang, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Few-Shot Learning Based on Embedded Self-Distillation and Adaptive Wasserstein Distance for Hyperspectral Image ClassificationabstractDue to the domain shift, it is challenging to achieve ideal experimental results for cross-domain few-shot learning (FSL) in hyperspectral image (HSI) classification. Most existing FSL algorithms are impacted by the limited samples, and they do not effectively leverage the representations from different layers of the network. Therefore, this article proposes an FSL based on embedded self-distillation and adaptive Wasserstein (ESAW-FSL) distance for HSI classification. First, the embedding self-distillation network is proposed in the feature extraction process of the source domain (SD) and the target domain (TD). The embedding self-distillation network utilizes self-distillation from different perspectives to get discriminative features. In the SD, the mask evaluation of embedded features is employed to guarantee the learning of guiding features. Second, a domain adaptation based on adaptive Wasserstein distance is designed to alleviate the domain shift problem between the domains. A lightweight feature correlation network learns the comprehensive cost matrix in the Wasserstein distance adaptively, and the obtained cost matrix helps achieve domain adaptation by an iterative algorithm. Finally, a focal loss based on double softening is adopted in the process of FSL. The probability is double softened to improve the ratio of correctly classifying hard samples. Experiments are conducted on three widely used hyperspectral datasets and compared with six state-of-the-art algorithms. The overall accuracy (OA) and average accuracy (AA) are achieved in multiple experiments, demonstrating the effectiveness of ESAW-FSL. Shizhe Shang, Ronghua Shang, Dongzhu Feng, Chao Wang 0099, Jie Feng 0003, Songhua Xu |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | Oriented Vehicle Joint Detection and Tracking in Satellite Video via Identifier-Free Point SupervisionabstractOriented vehicle detection and tracking play a crucial role in various real-world applications. Yet, existing advanced models heavily rely on abundant and accurate oriented bounding box and tracking identifier annotations, which are extremely labor-intensive for satellite videos. In this paper, we endeavor to employ identifier-free point annotation to achieve the competitive performance while minimizing annotation costs. Specifically, each instance across video frames are labeled by single points, without providing its instance identifier. Building upon this setting, we introduce an oriented vehicle joint detection and tracking framework for satellite video, focusing on enhancing model performance by carefully-designed sample acquisition and robust learning processes. Firstly, we leverage temporal and visual information to generate sequence-aligned pseudo-labels and visually-aligned synthetic objects, which complement each other during training by providing both exact appearance and annotation information. Secondly, a novel spatio-temporal consistency metric is developed to assess sample quality, which is then incorporated into a curriculum learning schedule. This strategy facilitates a gradual learning progression from high-quality data to low-quality or noisy examples, thereyby minimizing interference from potentially misleading samples. Finally, an end-to-end oriented object joint detection and tracking network is constructed to enable effective oriented vehicle dynamic analysis. Extensive ablation and experimental results on two satellite video datasets demonstrate the superiority of our proposed method. Yuping Liang, Jinjian Wu, Junpeng Zhang 0002, Yuxuan Chang, Jie Feng 0003, Guangming Shi |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Edge-Enhanced Cascaded MRF for SAR Image SegmentationabstractMarkov Random Fields (MRF) effectively capture local contextual information by modeling the spatial dependencies between pixels, which helps highlight details and enhances segmentation smoothness. To fully exploit MRF for synthetic aperture radar (SAR) image segmentation, we propose a novel edge-enhanced cascaded MRF (ECMRF) approach. Specifically, we introduce multiple edge-constrained filters to emphasize SAR image boundaries and provide relatively clean features. Building on this, we present a cascaded MRF framework that sequentially integrates region-level and pixel-level segmentation with feature perturbation and fusion to generate the final segmentation output. The framework comprises four key components: (1) a region-level MRF, regulated by edge features, to achieve precise region segmentation; (2) a pixel-level MRF with selective label smoothing to refine edges and reduce noise clusters; (3) equal-channel feature perturbation to increase feature diversity; and (4) a random probability-based feature fusion scheme to merge the input features. Experimental results demonstrate that our ECMRF outperforms six state-of-the-art comparable methods, underscoring its competitive performance. Ronghua Shang, Kang Liu 0025, Jie Feng 0003, Chao Wang 0099, Songhua Xu, Yangyang Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Knowledge Distillation Based on Adaptive Learning and Channel Amplification Features for PolSAR Image ClassificationabstractThe models currently used for Polarimetric Synthetic Aperture Radar (PolSAR) image classification tasks have problems such as complex network structures, poor distinction of detailed features, and fixed loss weights during the training process. In response to these problems, this paper proposes a PolSAR image classification method based on knowledge distillation using adaptive learning and channel amplification features. Firstly, this paper builds a knowledge distillation framework for PolSAR. Using a teacher network trained in advance that can acquire global knowledge to guide the student. This framework reduces the computational complexity and improves the classification accuracy of the student. Then, an adaptive loss weight learning mechanism is designed, which sets the weight of the Kullback-Leibler divergence loss during training into a learnable mode. The weight can be automatically adjusted according to the actual training situation of the student. Finally, a scheme for channel amplification to enhance features is proposed. This scheme obtains channel weights based on the student’s feature map information. These weights are amplified, strengthening the network’s ability to obtain feature information. Compared with the five PolSAR image classification algorithms, the method proposed in this paper uses lower computational complexity to obtain higher classification accuracy on the Flevoland, San Francisco, and Xi’an datasets. Ronghua Shang, Mingwei Hu, Lei Liu 0014, Jie Feng 0003, Songhua Xu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | GTCFN: A Graph-Based Transformer and Convolution Fusion Network for Hyperspectral Image ClassificationabstractGraph Neural Networks (GNN) are capable of modeling complex non-Euclidean structures through information transfer, and thus have been party widely used in the field of Hyperspectral Image (HSI) classification. However, conventional GNNs often have difficulty in handling regular grid data, which in turn loses positional information or spatial coherence, as well as in capturing long-range dependencies, which affects their performance in heterogeneous and limited-sample condition. To address these limitations, this paper proposes a novel Graph-based Transformer and Convolution Fusion Network (GTCFN) that integrates the local representation power of Convolutional Neural Networks (CNNs) with the global reasoning capability of graph-based Transformers. GTCFN consists of two synergistic branches: a Graph Transformer sub-network (GTsN) that models high-level semantic structures among superpixels via attention-based topology learning, and a Spectral–Spatial Convolutional sub-network (S2CsN) that extracts multi-scale fine-grained features using 5×5, 7×7, and 9×9 convolutional kernels. To enhance efficiency and generalization, GTCFN incorporates kernelized attention with random feature mapping, reducing the complexity fromO(M2) toO(M). At the same time, attention oversmoothing is avoided by introducing a Gumbel-based multi-head random aggregation mechanism. Experiments conducted on four benchmark datasets, namely Indian Pines, Pavia University, Salinas and WHU-Hi-HongHu, show that GTCFN achieves state-of-the-art performance with OA of 95.62%, 98.34%, 97.88% and 96.69%, which is significantly better than 12 other algorithms, such as CNNs, graph-based models and hybrid network models. The core code for GTCFN is posted on https://github.com/ Majunyi310321/GTCFN. Junyi Ma, Yao Ding 0010, Jie Feng 0003 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | S4DL: Shift-Sensitive Spatial-Spectral Disentangling Learning for Hyperspectral Image Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) techniques, extensively studied in hyperspectral image (HSI) classification, aim to use labeled source domain data and unlabeled target domain data to learn domain invariant features for cross-scene classification. Compared to natural images, numerous spectral bands of HSIs provide abundant semantic information, but they also increase the domain shift significantly. In most existing methods, both explicit alignment and implicit alignment simply align feature distribution, ignoring domain information in the spectrum. We noted that when the spectral channel between source and target domains is distinguished obviously, the transfer performance of these methods tends to deteriorate. Additionally, their performance fluctuates greatly owing to the varying domain shifts across various datasets. To address these problems, a novel shift-sensitive spatial-spectral disentangling learning (S4DL) approach is proposed. In S4DL, gradient-guided spatial-spectral decomposition (GSSD) is designed to separate domain-specific and domain-invariant representations by generating tailored masks under the guidance of the gradient from domain classification. A shift-sensitive adaptive monitor is defined to adjust the intensity of disentangling according to the magnitude of domain shift. Furthermore, a reversible neural network is constructed to retain domain information that lies not only in semantic but also the shallow-level detailed information. Extensive experimental results on several cross-scene HSI datasets consistently verified that S4DL is better than the state-of-the-art UDA methods. Our source code will be available athttps://github.com/xdu-jjgs/IEEE_TNNLS_S4DL. Jie Feng 0003, Junpeng Zhang 0002, Ronghua Shang, Weisheng Dong, Guangming Shi, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Learnable Prompts-Based Transformers for Domain Generalization of Hyperspectral Image ClassificationabstractExtensive pre-trained visual-language alignment models, such as Contrastive Language-Image Pre-training (CLIP), have demonstrated significant potential for learning representations transferable to domain generation tasks. In hyperspectral image (HSI) classification, a major challenge in deploying such models lies in prompt engineering, which requires particular expertise and substantial time investment. Moreover, existing methods ignore correlation information cross spectral bands. To address these issues, a novel method named learnable prompts-based Transformer (LPFormer) is proposed in this paper. In LPFormer, cross-band correlation information is extracted by self-attention of the transformer, which converted into positional embedding within the transformer framework to obtain the visual features. Subsequently, prompt words are modeled using learnable parameters that turn into efficient expertise. Finally, contrast learning method is used to align visual and textual features. Experimental results on two HSI datasets shows that the proposed LPFormer outperforms other domain adaptation methods. Baofa He, Jie Feng 0003, Ronghua Shang, Jinjian Wu, Licheng Jiao |
IGARSS | 2 |
| 2024 | A Dual-Branch Network for End-to-End Point-Supervised Object Detection on Remote Sensing ImagesabstractLearning object detectors for remote sensing images commonly requires for a huge number of annotated boundary boxes, which are not available without enormous manual efforts in annotating. Alternatively, points can indicate the existence of the objects of interests with reduced labeling cost. Existing Point-supervised object detection (PSOD) methods predominantly employ a two-stage training strategy, which involves propagating point annotations to pseudo boxes at the first stage then training an object detector with these pseudo boxes in a fully supervised manner. However, such paradigm substantially impedes the end-to-end flow of training gradients. In this work, we propose a novel dual-branch network (DBNet) for end-to-end weakly supervised object detection on remote sensing images. Firstly, a pseudo box generation network is attached to the object detector as a sibling branch, which produces semantic response maps for the objects of interest then extracts pseudo boxes by examining their spatial connectivity. Then, instead of training this pseudo box generation network separately, we jointly adjust the pseudo box generation network and the detection network through a multi-task loss. Experimental results on the DOTA-v1.0 dataset demonstrate the effectiveness of our proposed method, achieving an average precision (mAP50) of 32.3%. Jie Feng 0003, Junpeng Zhang 0002, Ronghua Shang, Xiangrong Zhang, Licheng Jiao |
IGARSS | 2 |
| 2024 | Multi-agent deep reinforcement learning for hyperspectral band selection with hybrid teacher guide
Jie Feng 0003, Qiyang Gao, Ronghua Shang, Xianghai Cao, Gaiqin Bai, Xiangrong Zhang, Licheng Jiao |
Knowl. Based Syst. | 1 |
| 2024 | Ellipse IoU Loss: Better Learning for Rotated Bounding Box RegressionabstractRotated object detection is an important research content in the field of remote-sensing images. However, in the rotated object detection, the inconsistency between the loss function and the final detection metric has become an important factor restricting the improvement of detection accuracy. So, in this letter, an ellipse intersection over union (IoU) loss (EPIoU loss) is proposed to solve these problems. The EPIoU loss uses IoU between the bounding boxes’ inscribed ellipses, which is approximate to the original bounding box IoU. This loss function can jointly optimize the prediction box parameters and promote the model to locate the object better. Compared to the complex intersection of rotated rectangles, the intersection calculation of two rotated ellipses is simple. A unified and differentiable process is also designed to calculate EPIoU, which avoids the complexity of the original bounding box IoU calculation. The experiments on DOTA, DIOR, and HRSC datasets verify that the proposed loss function can effectively improve the accuracy of the model. Ronghua Shang, Zihan Ju, Jie Feng 0003, Songhua Xu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Multiscale Spatial-Channel Transformer Architecture Search for Remote Sensing Image Change DetectionabstractDeep learning-based approaches play important roles and achieve impressive performances in remote sensing image change detection. Most of the networks are designed by the researchers with rich experiences. It is difficult to design the fixed networks with universal and good performance on various datasets. Regarding this issue, this study presents a transformer architecture search for remote sensing image change detection. The transformer architectures can be designed automatically by a two-stage transformer architecture search. The first stage can search for the effective combinations of attention mechanisms, while the second stage can identify the suitable combinations of multiscale modules. A gradient-based optimization method is employed for enabling the stable and efficient transformer architecture search. The experiments on various change detection datasets can verify the effectiveness of this work. Mengxuan Zhang 0003, Long Liu 0004, Zhikun Lei, Kun Ma 0003, Jie Feng 0003, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | CFDRM: Coarse-to-Fine Dynamic Refinement Model for Weakly Supervised Moving Vehicle Detection in Satellite VideosabstractDeep learning methods have gradually developed into the mainstream methods of moving vehicle detection in satellite videos. However, these methods require labor-intensive and time-consuming box-level annotations to predict accurate locations and sizes, which is challenging for large-scale satellite video datasets with hundreds of vehicles. To address this problem, a novel coarse-to-fine dynamic refinement framework (CFDRM) is proposed for moving vehicle detection in satellite videos only under the supervision of point-level annotations. CFDRM generates initial proposal boxes and performs spatio-guided matching with point annotations to obtain coarse box-level pseudo annotations. The initial priority of these coarse annotations is calculated by leveraging locally-consistent prior tailored to satellite videos. Then, a dynamic refinement detector is constructed to transfer coarse annotations to fine annotations with prior and predictive collaborative curriculum refinement. During the curriculum learning process, the coarse annotations are sequentially learned with a certain priority, where the priority is inferred by considering the prior knowledge from the locally-consistent prior and the knowledge itself from the predicted detector. Ultimately, a novel ambiguity-aware loss is designed to optimize the dynamic refinement detector from coarse annotations to fine annotations in an adaptively-weighted fashion. Extensive experiments have been conducted on the Jilin-1 and SkySat satellite video datasets demonstrate the superiority of CFDRM. Jie Feng 0003, Quanpeng Jiang, Junpeng Zhang 0002, Yuping Liang, Ronghua Shang, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Class-Aligned and Class-Balancing Generative Domain Adaptation for Hyperspectral Image ClassificationabstractThe task of hyperspectral image (HSI) classification is fundamental and crucial in HSI processing. Currently, domain adaptive methods have become a research hotspot in HSI classification. However, most domain adaptive methods ignore the class alignment in different domains. Additionally, HSIs have the characteristics of category imbalance and complex spatial-spectral distribution, which restricts the adaptation performance in HSIs. To address these problems, a class-aligned and class-balancing generative domain adaptation (CCGDA) method is proposed for HSI classification. The architecture of CCGDA is designed by using the classifier, domain discriminator, sampler and two weight-sharing generators. In the classifier, split-level capsule network is constructed by extracting rich spatial information of shallow layer and spectral features of deep layer with equivariant characteristic. Then, the classifier provides the pseudo label of samples in the target domain. To prevent the generators from mode collapse caused by category imbalance, the sampler is designed. It samples and re-samples the samples of the target domain in an adaptive proportion according to the statistical calculation through confidence and distribution of pseudo labels. Finally, a novel class-aligned domain adversarial loss is defined to jointly optimize the generators and discriminator. It incorporates the class shift adjusting and adaptive sampling for the samples of the target domain to better adapt the discriminant boundary of the classifier to the target domain. Experiments on benchmark HSI datasets verify the superiority of the proposed method for domain adaptive classification. Jie Feng 0003, Ziyu Zhou 0009, Ronghua Shang, Jinjian Wu, Xiangrong Zhang, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | SAR Image Segmentation Based on Complicated Region-Sensitive Adaptive Superpixel Generation and Hybrid Edge CorrectionabstractSuperpixel segmentation algorithms are predominently based on simple linear iterative clustering (SLIC), and treat homogeneous and complex regions equally. This can lead to suboptimal segmentation results, especially in complex images with multiple objects. We address this problem by proposing an SAR image segmentation algorithm based on complicated region-sensitive adaptive superpixel generation and hybrid edge correction (RSASGEC). First, a dynamic initialization algorithm for superpixel seeds based on region complexity is designed. Specifically, a new superpixel representation structure for superpixel seeds is constructed by combining superpixel complexity and the number of contained pixels. The algorithm gives priority to regions with high complexity, dynamically selecting the region with the highest complexity for further partitioning. This results in a dense distribution of superpixel seeds in complex regions, and sparse distributions in homogeneous regions with low complexity. Second, an iterative superpixel segmentation process based on an adaptive energy function is proposed. The Lagrange multiplier mathematical strategy is employed to optimize the adaptive energy function within an adjustable search window, resulting in more compact superpixel segmentation. Finally, a label correction method, based on edge mixture model constraints, is proposed for postprocessing. By integrating edge information from the Gaussian edge detector and the Canny algorithm as constraints, this method leverages majority voting and region growth methods to mitigate edge noise and outliers, refining the superpixel labels. The RSASGEC algorithm is verified in experiments, using one simulated image and six real SAR images. The results indicate that RSASGEC outperforms six representative algorithms, achieving more satisfactory segmentation performance. Jinhong Ren, Ronghua Shang, Jiansheng Chen 0004, Jie Feng 0003, Chao Wang 0099, Songhua Xu, Rustam Stolkin |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Joint Adversarial Network With Semantic and Topology Fusion for Cross-Scene Hyperspectral Image ClassificationabstractHyperspectral image cross-scene classification (HSICC) poses a significant challenge due to distribution variations between source and target domains. Existing unsupervised domain adaptation methods primarily focus on local knowledge transfer, often neglecting the critical semantic information and sample topological structure inherent in hyperspectral images (HSIs). To address these limitations, this article introduces an end-to-end joint adversarial network with semantic and topology fusion (JAN-STF). This network liberates from the constraints of local perception by integrating semantic and topological information into both domain- and class-level adversarial learning processes. First, the network constructs a semantic-guided cross-domain graph structure to obtain cross-domain features. Subsequently, domain-level adversarial learning is conducted using these features to achieve domain-invariant representation with robust transferability. Moreover, to bolster stability in the ensuing class-level adversarial procedure, the network dynamically computes cross-domain category center distance loss utilizing an intra-domain topological semantic attention mechanism, thereby mapping features to proximate spaces. Finally, class-level adversarial learning is performed by leveraging the prediction discrepancy between the local classifier and the topological classifier, thus enhancing the discriminative performance of the domain-invariant representation. Extensive experiments on three broadly utilized HSICC datasets demonstrate JAN-STF’s superiority in accuracy and Kappa coefficient (KC) metrics over nine leading algorithms. Ronghua Shang, Jie Feng 0003, Songhua Xu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | A Multitask Framework for Hyperspectral Change Detection and Band Reweighting With Unbalanced Contrastive LearningabstractMultitask learning has been widely applied in visual learning to significantly enhance the performance. The combination of hyperspectral change detection (HCD) and band reweighting can achieve discriminative feature enhancement for improving detection performance. However, existing multitask models for these two tasks are unidirectional, with band reweighting unable to learn from task guidance. To address this challenge, a multitask HCD (MHCD) framework with differential band reweighting and unbalanced contrastive learning is proposed. MHCD consists of a differential band reweighting network (DBRN) and a Siamese detection network. DBRN extracts discriminative information for HCD by analyzing the differential spatial-spectral information across time states, whose optimization is under the guidance of HCD. Furthermore, a multitemporal interaction module and multidomain fusion module are inserted into the Siamese detection network. They hierarchically connect cross-temporal features and fuse features from spatial, spectral, and temporal domains, providing complementary clues in these different domains. Considering the sample imbalance and enormous variation within a class in binary HCD, an unbalanced contrastive learning method based on multiple prototypes (UCLM) tailored has been considered. It estimates multiple prototypes to flexibly adjust the contribution of different classes of samples to the loss. The proposed method has been validated using three public benchmark datasets, demonstrating improvements in multiple metrics for change detection. The code of our paper is available at:https://github.com/jiefeng0109/MHCD. Xiande Wu, Paolo Gamba, Jie Feng 0003, Ronghua Shang, Xiangrong Zhang, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Domain Adversarial Debiased Self-Training for Hyperspectral Image ClassificationabstractUnsupervised domain adaptation (UDA) has been widely used in hyperspectral image (HSI) classification. Domain adversarial learning methods and self-training methods are two major UDA methods. Most existing methods use these two methods independently, which limit their capacity to knowledge transferability and robustness of classification. Therefore, a novel domain adversarial debiased self-training (DADST) is proposed to combine these two methods, where domain adversarial learning reduces the domain discrepancy and debiased self-training updates model via a self-paced curriculum policy. To this end, we first apply debiased self-training into HSI classification. Then domain adversarial learning module is combined to align joint distribution between the source domain and the target domain. Experimental results on two cross-scene HSI datasets demonstrate that the proposed DADST method outperforms other domain adaptation approaches. Jie Feng 0003, Ziyu Zhou 0009, Xiangrong Zhang, Licheng Jiao |
IGARSS | 2 |
| 2023 | 3D-Mglnet: Moving Vehicle Detection in Satellite Videos with 3D Motion-Guided Lightweight NetworkabstractObject detectors based on convolutional neural networks have been widely-applied to detect moving vehicles in satellite videos. However, many detectors render superior detection accuracy at the expense of increased computational complexity and decreased inference speed. This prevents these detectors from being deployed into mobile devices. In this paper, an efficient 3D motion-guided lightweight network (3D-MGLNet) is proposed. Specifically, 3D-MGLNet constructs a motion-guided module based on 3D convolution to extract motion cues from spatial-temporal information. This module uses model compression strategies to detect moving vehicles in real-time while following the principle of "fewer channels, smaller convolution kernels," significantly reducing the number of parameters and computational complexity. Extensive experiments are conducted on the Jilin-1 and SkySat satellite video datasets. The results demonstrate that 3D-MGLNet gains strong performance by striking an excellent tradeoff between resource and accuracy, resulting in the fewest parameters (0.35M) and fastest speed (66.84 fps) compared to other popular models. Jie Li 0001, Jie Feng 0003, Quanpeng Jiang, Xiangrong Zhang, Licheng Jiao |
IGARSS | 3 |
| 2023 | Semi-supervised feature learning for disjoint hyperspectral imagery classification
Xianghai Cao, Jie Feng 0003, Licheng Jiao |
Neurocomputing | 3 |
| 2023 | Multi-Angle Models and Lightweight Unbiased Decoding-Based Algorithm for Human Pose EstimationabstractWhen a top-down method is taken to the task of human pose estimation, the accuracy of joint point localization is often limited by the accuracy of human detection. In addition, conventional algorithms commonly encode the image to generate a heat map before processing, but the systematic error in decoding the heat map back to the original image has an impact on the positioning. Therefore, to address the two problems, we propose an algorithm that uses multiple angle models to generate the human boxes and then performs lightweight decoding to recover the image. The new boxes can better fit humans and the recovery error can be reduced. First, we split the backbone network into three sub-networks, the first sub-network is responsible for generating the original human box, the second sub-network is responsible for generating a coarse pose estimation in the boxes, and the third sub-network is responsible for a high-precision pose estimation. In order to make the human box fit the human body better, with only a small number of interfering pixels inside the box, models of the human boxes with multiple rotation angles are generated. The results from the second sub-network are used to select the best human box. Using this human box as input to the third sub-network can significantly improve the accuracy of the pose estimation. Then to reduce the errors arising from image decoding, we propose a lightweight unbiased decoding strategy that differs from traditional methods by combining multiple possible offsets to select the direction and size of the final offset. On the MPII dataset and the COCO dataset, we compare the proposed algorithm with 11 state-of-the-art algorithms. The experimental results show that the algorithm achieves a large improvement in accuracy for a wide range of image sizes and different metrics. Jianghai He, Ronghua Shang, Jie Feng 0003, Licheng Jiao |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2023 | MR-Selection: A Meta-Reinforcement Learning Approach for Zero-Shot Hyperspectral Band SelectionabstractBand selection is an effective method to deal with the difficulties in image transmission, storage, and processing caused by redundant and noisy bands in hyperspectral images (HSIs). Existing band selection methods usually need to learn a specific model for each HSI dataset, which ignores the inherent correlation and common knowledge among different band selection tasks. Meanwhile, these methods lead to a huge waste of computation. In this article, a novel zero-shot band selection method, called MR-Selection, is proposed for HSI classification. It formalizes zero-shot band selection as a metalearning problem, where advantage actor–critic algorithm-based reinforcement learning (A2C-RL) is designed to extract the metaknowledge in the band selection tasks of various seen hyperspectral datasets through a shared agent. To learn a consistent representation among different tasks, a dynamic structure-aware graph convolutional network is constructed to build a shared agent in A2C-RL. In A2C-RL, the state is tailored in a feasible way and easy to adapt to various tasks. Meanwhile, the reward is defined according to an efficient evaluation network, which can evaluate each state effectively without any fine-tuning. Furthermore, a two-stage optimization strategy is designed to coordinate optimization directions of a shared agent from different tasks effectively. Once the shared agent is optimized, it can be directly applied to unseen HSI band selection tasks without any available samples. Experimental results demonstrate the effectiveness and efficiency of the MR-Selection on the band selection of unseen HSI datasets. Jie Feng 0003, Gaiqin Bai, Xiangrong Zhang, Ronghua Shang, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Multi-Complementary Generative Adversarial Networks With Contrastive Learning for Hyperspectral Image ClassificationabstractIn the last decade, generative adversarial network (GAN) and its variants provide a powerful training mechanism for hyperspectral image (HSI) classification. In HSIs, the distribution of samples is more complicated due to the existence of abundant spatial-spectral information and multi-scale information. The single generation pattern of GANs is prone to modal collapse for the sample generation of HSIs. Moreover, the promotion of the generator only relies on adversarial learning with the discriminator, which limits the generator’s performance. To address these problems, a multi-complementary GANs with contrastive learning (CMC-GAN) is proposed. CMC-GAN consists of two groups of GANs, where coarse-grained GAN adopts the structure in encoder-decoder form for hidden fine-scale and coarse-scale generation, and another fine-grained GAN is responsible for fine-scale generation. In fine-grained GAN, the discriminator is constructed to distinguish the fine-scale samples from different generators, which enforces the joint optimization of these two groups of GANs and makes GANs generate diverse multi-scale samples. Furthermore, a novel contrastive learning constraint is added into GANs, where a unidirectional contrastive loss guarantees the generators to extract intra-class invariant representation and a class-specific contrastive loss urges the discriminators to learn more discriminative features for classification. Finally, both discriminators are adaptively-fused to extract complementary multi-scale spatial-spectral features for classification under the guidance of diverse generated samples. The experimental results demonstrate CMC-GAN has superior classification performance, especially for small sample classification. Jie Feng 0003, Zizhuo Gao, Ronghua Shang, Xiangrong Zhang, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | MidNet: An Anchor-and-Angle-Free Detector for Oriented Ship Detection in Aerial ImagesabstractShip detection in aerial images remains an active yet challenging task due to its arbitrary object orientation and various aspect ratios from the bird’s-eye perspective. Most existing oriented objection detection methods rely on angular prediction or predefined anchor boxes, making these methods highly sensitive to unstable angular regression and excessive hyper-parameter setting. To address these issues, we replace the angular-based object encoding with an anchor-and-angle-free paradigm, and propose a novel detector deploying a center and four midpoints for encoding each oriented object, namely MidNet. Moreover, MidNet designs a novel symmetrical deformable convolution for enhanceing the features of midpoints, then the center and midpoints for an identical ship are adaptively matched by predicting corresponding centripetal shift and matching radius. Finally, a concise analytical geometry algorithm is proposed to calculate the ship orientation and refine the keypoints step-wisely for building precise oriented bounding boxes. On two public ship detection datasets, HRSC2016 and FGSD2021, MidNet outperforms the state-of-the-art detectors by achieving APs of 90.52% and 86.50%. Yuping Liang, Jie Feng 0003, Xiangrong Zhang, Junpeng Zhang 0002, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | CAST: A Cascade Spectral-Aware Transformer for Hyperspectral Image Change DetectionabstractHyperspectral image change detection (HSI-CD) aims to detect subtle changes on the Earth’s surface through approximately continuous spectral information, which has gradually become a very important research hotspot in the field of remote sensing (RS). In recent years, convolutional neural networks (CNNs) based HSI-CD methods have shown strong feature extraction capabilities. However, due to the simple fusion of spectral information in the channel dimension by CNN, the medium and long-term sequence properties of spectral features cannot be well mined and represented. Most previous studies mainly extract semantic features from images at different times, ignoring the temporal correlation between features, which cannot fully extract and effectively utilize temporal-spatial-spectral features. To this end, this paper proposes a cascade spectral aware transformer (CAST) for HSI-CD. First, we propose a temporal-spatial transformer (TS-Former) to enhance the temporal correlation and spatial global relationship of extracted features, thereby addressing the insufficient consideration of temporal correlation. Second, a spectral awareness transformer (SA-Former) is designed to better mine and represent the sequence properties of spectral features, especially the medium and long-term dependencies. Finally, we observe a spectral distortion in the process of extracting temporal-spatial features and based on this present a spectral constraint module (SCM) to preserve the sequence properties of spectral features and reduce the distortion of the spectrum. Extensive experiments on three challenging hyperspectral datasets demonstrate that our method achieves state-of-the-art results. The code is available at: https://github.com/tianshunli/CAST. Xiangrong Zhang, Shunli Tian, Guanchun Wang, Xu Tang 0004, Jie Feng 0003, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Bidirectional Multiple Object Tracking Based on Trajectory Criteria in Satellite VideosabstractMultiple object tracking (MOT) in satellite videos requires to detect all objects belonging to specified categories and identify each object, which plays a basic and necessary role in automatic driving, traffic surveillance, and smart city. The traditional MOT methods in satellite videos mostly follow the detection–association framework. However, the detection–association framework works under a strict assumption that all objects are correctly localized by the detector. In practice, MOT in satellite videos faces challenges such as low resolution, tiny objects, and the wide field of view, which leads to the degradation of detector performance. In order to reduce the impact of detector degradation, we propose a bidirectional MOT framework based on trajectory criteria (BMTC) in satellite videos. In BMTC, the single object tracking (SOT) tracker carries out locating the objects between consecutive frames and the detector is just used for finding new objects. Therefore, it is less dependent on the detector performance. According to the characteristics of satellite videos, the trajectory criteria are designed to control the state of the tracker, which includes trajectory density, the limit of consecutive virtual motion predictions, and trajectory similarity measurement. Invalid fragment trajectory backtracking is implemented to alleviate the misalignment caused by the above subsection trajectory criteria. The method is validated on the VISO benchmark and SkySat-1 dataset. The experimental results show the improvement of completeness and accuracy, and the proposed tracker achieves the state-of-the-art performance. Xiangrong Zhang, Zhongjian Huang, Xina Cheng, Jie Feng 0003, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | SDANet: Semantic-Embedded Density Adaptive Network for Moving Vehicle Detection in Satellite VideosabstractIn satellite videos, moving vehicles are extremely small-sized and densely clustered in vast scenes. Anchor-free detectors offer great potential by predicting the keypoints and boundaries of objects directly. However, for dense small-sized vehicles, most anchor-free detectors miss the dense objects without considering the density distribution. Furthermore, weak appearance features and massive interference in the satellite videos limit the application of anchor-free detectors. To address these problems, a novel semantic-embedded density adaptive network (SDANet) is proposed. In SDANet, the cluster-proposals, including a variable number of objects, and centers are generated parallelly through pixel-wise prediction. Then, a novel density matching algorithm is designed to obtain each object via partitioning the cluster-proposals and matching the corresponding centers hierarchically and recursively. Meanwhile, the isolated cluster-proposals and centers are suppressed. In SDANet, the road is segmented in vast scenes and its semantic features are embedded into the network by weakly supervised learning, which guides the detector to emphasize the regions of interest. By this way, SDANet reduces the false detection caused by massive interference. To alleviate the lack of appearance information on small-sized vehicles, a customized bi-directional conv-RNN module extracts the temporal information from consecutive input frames by aligning the disturbed background. The experimental results on Jilin-1 and SkySat satellite videos demonstrate the effectiveness of SDANet, especially for dense objects. Jie Feng 0003, Yuping Liang, Xiangrong Zhang, Junpeng Zhang 0002, Licheng Jiao |
IEEE Trans. Image Process. | 1 |
| 2022 | Hyperspectral Image Classification Based on Spectrally-Enhanced and Densely Connected Transformer ModelabstractDeep learning methods have been widely used in hyperspectral image classification. Most existing CNN-based deep learning methods can extract spatial features by the local receptive field. However, these methods do not pay enough attention to the long-term dependence on HSI data and cannot capture sequence attributes. To solve this problem, we design spectrally-enhanced and densely connected transformer model (SEDT). Due to the ability of transformers in obtaining global dependencies, dense connections are added on this basis to fuse the features from shallow layers to deep layers. Furthermore, a spectrally-enhanced module is constructed to enhance the network's ability to capture local contextual and semantic features. Extensive experiments show that the proposed method exhibits highly competitive classification performance on hyperspectral image datasets. Yongen Wu, Jie Feng 0003, Gaiqin Bai, Qiyang Gao, Xiangrong Zhang |
IGARSS | 2 |
| 2022 | Spectral Constrained Residual Attention Network for Hyperspectral PansharpeningabstractDeep learning methods have been widely used in the task of hyperspectral pansharpening. However, most of these methods regard the Panchromatic (PAN) image as a kind of auxiliary information, which is mainly used as spatial details to add on the hyperspectral image (HSI) after processing. Obviously, this kind of methods utilize the PAN image insufficiently, resulting in the imbalance of spatial preservation and spatial preservation. In this paper, a spectral constrained residual attention network (SCRAN) is proposed by using the PAN image as the foundation of the pansharpening task and concerning on the spectral and spatial learning. The proposed SCRAN method consists of three parts: a spectral feature extraction net, an attention spatial residual net and a spectral reconstruction net. A spectral constrained loss function is designed to enhance the spectral learning ability of SCRAN. Additionally, in SCRAN, a deep back-projection network (DBPN) is operated to upsample the HSI, and the histogram matching is applied to the PAN image to make it closer to the HSI in terms of spectral bands. Ziyu Zhou 0009, Jie Feng 0003, Xiande Wu, Jiao Shi, Xiangrong Zhang |
IGARSS | 2 |
| 2022 | Feature selection via Non-convex constraint and latent representation learning with Laplacian embedding
Ronghua Shang, Jiarui Kong, Jie Feng 0003, Licheng Jiao |
Expert Syst. Appl. | 3 |
| 2022 | Sparse and low-dimensional representation with maximum entropy adaptive graph for feature selection
Ronghua Shang, Jie Feng 0003, Yangyang Li 0001, Licheng Jiao |
Neurocomputing | 3 |
| 2022 | CMNet: Classification-oriented multi-task network for hyperspectral pansharpening
Xiande Wu, Jie Feng 0003, Ronghua Shang, Xiangrong Zhang, Licheng Jiao |
Knowl. Based Syst. | 2 |
| 2022 | Deep Mutual-Teaching for Hyperspectral Imagery ClassificationabstractHyperspectral imagery (HSI) classification is a widely used method in remote sensing, which can provide accurate label information for each pixel. Though the classification accuracy is very high for many publicly available data sets in many research articles, they often exhibit much worse performance in practical applications. Because most of the articles adopt a random sampling strategy to select training and test samples from the same image, the high correlation between training and test samples will bring optimistic results. However, this strategy is not suitable for practical application. Because the training and test samples are collected from different locations in most situations, in this letter, the nonoverlapped sampling is adopted to reduce the correlation between training and test samples. Four key factors are presented to analyze the HSI classification; then, a new deep mutual-teaching method is proposed to classify the HSI. The experimental results show that the performance of the proposed method outperforms comparison methods. Jin Zhao 0002, Zixuan Ba, Xianghai Cao, Jie Feng 0003, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Uncorrelated feature selection via sparse latent representation and extended OLSDA
Ronghua Shang, Jiarui Kong, Jie Feng 0003, Licheng Jiao, Rustam Stolkin |
Pattern Recognit. | 4 |
| 2022 | Nonoverlapped Sampling for Hyperspectral Imagery: Performance Evaluation and a Cotraining-Based Classification StrategyabstractFor hyperspectral imagery (HSI) classification, most of the studies focus on how to improve the classification accuracy, while the influence of sampling strategy for classification performance attracts little attention. For now, random sampling (RS) is the most adopted strategy. That is, for a hyperspectral image, a certain number of labeled samples are randomly selected as the training set, and the remaining labeled samples are taken as the test set. However, the RS strategy will produce over optimistic results when used for performance evaluation because of the overlap between training set and test set. Though spectral-spatial classification methods benefit most from the RS strategy, the pixel-wise classification methods can also benefit from it because of the high spectral correlation between training and test samples. However, in practical applications, the RS strategy is not feasible. Because the training and test samples are often collected from different locations. In this situation, the correlation between training and test samples will decrease dramatically and the performance of HSI classification methods will be affected. In this article, a nonoverlapped sampling method is adopted to reduce the correlation between training and test samples and different classic classification methods are evaluated. Experimental results show that the classification performance of all methods drops a lot when nonoverlapped sampling strategy is adopted. After the analysis of some important factors for HSI classification, we also propose a cotraining-based classification method to relief the influence of sampling strategy and obtains much better performance compared with those classic spectral-spatial classification methods. Xianghai Cao, Zuji Liu, Jie Feng 0003, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Deep Reinforcement Learning for Semisupervised Hyperspectral Band SelectionabstractBand selection is an important step in efficient processing of hyperspectral images (HSIs), which can be seen as the combination of powerful band search technique and effective evaluation criterion. The existing deep-learning-based methods make the network parameters sparse to search the spectral bands using threshold-based functions or regularization terms. These methods may lead to an intractable optimization problem. Furthermore, these methods need to repeatedly train deep networks for evaluating candidate band subsets. In this article, we formalize hyperspectral band selection as a reinforcement learning (RL) problem. Band search is regarded as a sequential decision-making process, where each state in the search space is a feasible band subset. To evaluate each state, a semisupervised convolutional neural network (CNN), called EvaluateNet, is constructed by adding the intraclass compactness constraint of both limited labeled and sufficient unlabeled samples. A simple stochastic band sampling method is designed to train EvaluateNet, making it possible to efficiently evaluate without any fine-tuning. In RL, new reward functions are defined by taking the EvaluateNet and the penalty of repeated selection into account. Finally, advantage actor–critic algorithms are designed to explore in the state space and select the band subset according to the expected accumulated reward. The experimental results on HSI data sets demonstrate the effectiveness and efficiency of the proposed algorithms for hyperspectral band selection. Jie Feng 0003, Xianghai Cao, Ronghua Shang, Xiangrong Zhang, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Self-Supervised Divide-and-Conquer Generative Adversarial Network for Classification of Hyperspectral ImagesabstractGenerative adversarial network (GAN) has been rapidly developed because of its powerful generating ability. However, imbalanced class distribution of hyperspectral images (HSIs) easily causes mode collapse in GAN. Moreover, limited training samples in HSIs restrict the generating ability of GAN. These issues may further deteriorate the classification performance of the discriminator. To conquer these issues, a novel self-supervised divide-and-conquer GAN (SDC-GAN) is proposed for HSI classification. In SDC-GAN, a pretext cluster task with an encoder-decoder architecture is designed by leveraging abundant unlabeled samples. By transferring the learned cluster representation from the cluster task, limited labeled samples are divided effectively in the downstream classification. According to the division of clustering, SDC-GAN constructs a generic and several specific branches for both the generator and discriminator. The generator generates all-class and specific-class samples by using the generic and specific branches separately and combines them adaptively. It can weaken the generation preference for the classes with large sample sizes and alleviate the mode collapse problem. Meanwhile, the classification ability of the discriminator is improved by integrating the judgment of specific branches into the generic branch. Experimental results show that SDC-GAN achieves competitive results for HSI classification compared with several state-of-the-art methods. Jie Feng 0003, Ronghua Shang, Xiangrong Zhang, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Hyperspectral Image Classification Based on Multiscale Cross-Branch Response and Second-Order Channel AttentionabstractRecently, most convolutional neural network-based methods use convolutional kernels of fixed size to extract features, which ignore the inherent spatial structure information of ground objects and lose spatial details. In addition, rough first-order statistics is not enough to capture subtle differences between different categories and extract non local context information. To address these issues, a hyperspectral image (HSI) classification method based on multi-scale cross-branch response and second-order channel attention (MCRSCA) is proposed in this paper. Firstly, a multi-scale cross-branch response module (MCBR) is proposed, which uses convolution kernels of different sizes for feature extraction. It adds and concatenates the features of different scales respectively to obtain rich and complementary spatial context information. Then, element multiplication and element addition are performed on the fused multi-scale features to promote the propagation of the multi-scale information and enhance the nonlinear expression ability. Next, the second-order channel attention module (SOCA) is designed to interact the channel information through the feature covariance matrix to obtain the long-term dependence between channels. This module pays more attention to the significant channels and suppresses the redundant channels. Finally, the residual connection is used to embed MCBR and SOCA into the residual block to improve the gradient back propagation and accelerate the training process. Experiments on four commonly used HSI benchmark datasets show that the results of MCRSCA is competitive compared with other state-of-the-art methods. Ronghua Shang, Huidong Chang, Jie Feng 0003, Yangyang Li 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Region-Level SAR Image Segmentation Based on Edge Feature and Label AssistanceabstractThis paper proposes a novel segmentation algorithm for synthetic aperture radar (SAR) images. The algorithm performs region-level segmentation based on edge feature and label assistance (REFLA). It demonstrates improved performance in terms of segmentation accuracy while better preserving image edges. Firstly, an edge detection scheme is implemented, which fuses information from two advanced edge detection methods, thereby obtaining a more precise edge strength map (ESM). Secondly, a Canny algorithm is performed to divide the SAR image into edge regions and homogeneous regions, and different smoothing templates are selected according to pixel positions. Therefore, an anisotropic smoothing on the SAR image can be achieved, aiming at suppressing the noise within targets while also accurately maintaining the target boundaries. Thirdly, K-means clustering is applied on the smoothed result, to generate an initial set of labels. Using ESM and the initial labels as inputs, a watershed transformation and a majority voting strategy are employed to realize an initial segmentation at the region level. Finally, a label-aided region merging (LaRM) strategy is used to correctly segment the wrongly labeled regions, to give the final segmentation result. The LaRM, with merging rules based on label rather than gray characteristics, can avoid the need for calculating a large number of complex formulae, thus accelerating the region merging. Results are presented of experiments, on both simulated and real SAR images, in which the proposed REFLA method is compared against six state-of-the-art algorithms from the literature. REFLA achieves higher accuracy, while better retaining the image edges. Ronghua Shang, Licheng Jiao, Jie Feng 0003, Yangyang Li 0001, Rustam Stolkin |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | SAR Image Segmentation Based on Constrained Smoothing and Hierarchical Label CorrectionabstractSynthetic aperture radar (SAR) is widely used in the field of modern remote sensing due to its high resolution for a comparatively small antenna. However, there are still some difficulties in the processing of SAR images. In particular, accurate segmentation of small targets and image corners remains an important challenge, as these can easily be lost during conventional image smoothing and denoising methods. To address this, we propose an SAR image segmentation algorithm based on constrained smoothing and hierarchical label correction (CSHLC). First, a Canny algorithm is used to extract the edges of SAR images, and the Gaussian smoothing is performed on SAR images under edge constraints to achieve noise reduction so that the edges of small and big targets are well preserved. Second, a preliminary K-means clustering is conducted on the smoothing results, and then, a Markov random field (MRF) model is used on the clustering results (“original label” results), iteratively calculating a maximum likelihood set of pixel labels. Finally, through two label correction methods, pixel group counting comparison (PGCC) and gray similarity comparison (GSC), the labels of the MRF output are further checked and corrected to obtain final segmentation results. Compared with seven state-of-the-art algorithms, simulation results on both simulated SAR images and real SAR images show that the proposed CSHLC delivers higher accuracy while better retaining corners and small targets. Ronghua Shang, Junkai Lin, Jie Feng 0003, Yangyang Li 0001, Rustam Stolkin, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Simplified Nonlocal Network Based on Adaptive Projection Attention Method for Hyperspectral Image ClassificationabstractNonlocal convolutional neural networks have difficulties in dealing with the imbalanced samples in hyperspectral images effectively, so the networks cannot achieve ideal experiment results. Therefore, this paper proposes an adaptive projection attention-based simplified nonlocal neural network for hyperspectral image classification. Firstly, the local information is calculated in horizontal and vertical directions. Then the information is passed to a simplified nonlocal network to learn the global semantic information. The simplified nonlocal network can reduce information redundancy and improve classification accuracy at the same time. Secondly, the global semantic information is adaptively projected according to the spatial features and compressed using multi-scale pooling layers. After that, the pooled results are reassigned channel weights through two fully connected layers and extended using multi-scale pooling layers. Then the extended features are concatenated with the global semantic information, which can alleviate the imbalanced sample existing in the dataset. Then a simplified nonlocal approach is used to fuse shallow and deep information to improve the robustness and classification performance of the network. In this paper, experiments of the proposed method are conducted on three widely used hyperspectral datasets compared with those of seven state-of-the-art algorithms, and satisfactory overall and average accuracies are achieved, demonstrating the effectiveness of the proposed algorithm. Ronghua Shang, Jie Feng 0003, Yangyang Li 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Hyperspectral Image Classification Based on Pyramid Coordinate Attention and Weighted Self-DistillationabstractAttention mechanism-based Hyperspectral Image (HSI) classification algorithms typically extract spectral and spatial features by spectral attention and spatial attention network respectively. However, these algorithms lack joint attention and ignore imbalanced samples, leading to insufficient information extraction. To address this problem, this paper proposes a novel HSI classification algorithm based on the pyramidal coordinate attention and weighted self-distillation (PCA-WSD). To perform the joint attention of spectral and spatial features, the proposed PCA mechanism uses spectral attention to cope with the diverse spatial features. The PCA mechanism consists of two components. First, the spatial pyramid coordinate squeeze (SPCS) is designed to aggregate spatial features with local and global information. Then, the tailored spatial pyramid coordinate excitation (SPCE) adaptively enhances their informative spectral features for the obtained spatial features, realizing the joint attention to spectral-spatial features. Further, considering the imbalance of samples, WSD is proposed. Specifically, weighted cross-entropy is integrated into WSD. Extensive experiments are evaluated on the four HSI benchmark datasets: Indian Pine (IP), Pavia University (UP), Kennedy Space Center (KSC), and Pavia Center (PC). Compared with the seven advanced algorithms, experimental results of the proposed algorithm1. reveal superior classification performance, especially for the imbalanced samples. Ronghua Shang, Jinhong Ren, Songling Zhu, Jie Feng 0003, Yangyang Li 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Unsupervised Hyperspectral Band Selection With Multigraph Integrated Embedding and Robust Self-Contained RegressionabstractBand selection is an effective means to alleviate the curse of dimensionality in hyperspectral data. Many methods select a compact and low redundant band subset, which is inadequate as it may degrade the classification performance. Instead, more emphasis shall be put on selecting representative bands. In this article, we propose a robust unsupervised band selection method to address this issue. Our method reveals bandwise representativeness based on the comprehensive interband neighborhood structure. It incorporates an interband neighborhood graph into a sparse self-contained regression model in order to provide a reasonable measure for bandwise representativeness. The derived coefficient matrix not only uncovers bandwise importance values but also is coherent to the generalized interband local neighborhood structure. For constructing the interband neighboring structural graph, an integrated multigraph model is employed to achieve better generalization performance. It combines the benefit of multiple graphs but is insusceptible to the defects of a single one. To enhance the reliability of this model, a joint trace minimum and nonnegative constraint is imposed on the coefficient matrix. Accordingly, a multigraph integrated embedding and robust self-contained regression model (MGRSR) is formulated. In addition, an iterative update algorithm is developed to solve the problem. Comparative experiments on three hyperspectral data sets illustrate that MGRSR is robust to various data and has superior performance compared with several state-of-the-art methods. Chenhong Sui, Jun Zhou 0001, Chang Li 0001, Jie Feng 0003, Xiaoguang Mei, Jing Wang 0062 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Multiobjective Guided Divide-and-Conquer Network for Hyperspectral PansharpeningabstractDeep learning methods have gained rapid development in hyperspectral pansharpening (HP) due to powerful spatial–spectral feature extraction ability. However, most of these methods are optimized using a single reconstruction objective. It is difficult for these methods to find a balance between spectral preservation and spatial preservation. Furthermore, these methods adopt interpolation or convolution to upsample the hyperspectral images (HSIs), which tends to cause noticeable spectral distortion. To conquer these issues, a novel multiobjective guided divide-and-conquer network (MO-DCN) is proposed for HP. It consists of a deconvolution long short-term memories (LSTMs) network (DLSTM) and a divide-and-conquer network (DCN). DLSTM leverages bi-direction learning to upsample HSIs by considering 3-D spatiotemporal dependencies. Then, DCN designs a two-branch architecture to reconstruct spatial and spectral information from upsampled HSIs and panchromatic images (PANIs), respectively, where the spatial branch designs an attention-in-attention module (AIAM) to emphasize complementary attention in a coarse-to-fine way. Finally, co-improvement of spatial and spectral information is formulated as an Epsilon-constraint-based multiobjective optimization. The Epsilon constraint method transforms one objective into a constraint and regards it as a penalty bound to make an excellent tradeoff between different objectives. Experimental results demonstrated that the proposed method markedly improves pansharpening performance in both the spatial and spectral domains and has superior fusion performance than state-of-the-art methods. Xiande Wu, Jie Feng 0003, Ronghua Shang, Xiangrong Zhang, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Spatial Pooling Graph Convolutional Network for Hyperspectral Image ClassificationabstractGraph convolution networks (GCNs) have been applied in a variety of fields due to their powerful ability in processing graph-like data. However, the massive number of hyperspectral pixels makes it challenging to define general graph structures on hyperspectral images (HSIs). On the other hand, convolutional neural networks (CNNs) take in regular image regions with fixed square size, and have demonstrated impressive accuracy while being efficient in computation. Inspired by the classification framework of CNNs, we develop a GCN-based model that generates effective local spectral–spatial features for HSI classification. Specifically, graph convolutions are performed separately on every local region, which significantly limits the graph’s size. While graph convolution extracts features of every pixel, it does not reduce the number of them. To fuse suitable representations for the classification task, we develop a graph pooling operation to preserve classification-specific features and reduce redundant pixels. Based on local regions of HSIs, pooling in the graph domain is equivalent to spatial pooling in the spatial domain. The proposed method is thus named the spatial pooling graph convolutional network (SPGCN). Experimental results on several typical datasets demonstrated that the proposed SPGCN provides competitive results compared with other state-of-the-art CNN-based methods. Xiangrong Zhang, Peng Zhu 0004, Xu Tang 0004, Jie Feng 0003, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Spectral Partitioning Residual Network With Spatial Attention Mechanism for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is one of the most important tasks in hyperspectral data analysis. Convolutional neural networks (CNN) have been introduced to HSI classification and achieved good performance. In this article, an effective and efficient CNN-based spectral partitioning residual network (SPRN) is proposed for HSI classification. The SPRN splits the input spectral bands into several nonoverlapping continuous subbands and uses cascaded parallel improved residual blocks to extract spectral–spatial features from these subbands, respectively. Finally, the features are fused and fed into a classifier. By equivalently using grouped convolutions, the spectral partition and feature extraction are embedded into an end-to-end network. Experimental results show that the proposed SPRN achieves state-of-the-art performance, meanwhile, with relatively fewer parameters and computational costs. Usually, the CNN takes a patch that contains continuous spatial information as the input and results in a class label of the center pixel. The large size of the input patch includes more spatial information, whereas also introduces interfering pixels that may lead to a degradation of classification accuracies. For that reason, we propose a novel spatial attention module named homogeneous pixel detection module (HPDM). The module alleviates the degradation of performance as the input patch size increases by capturing the homogeneous pixels in the input patch. The module can be integrated into any CNN-based HSI classification framework. Xiangrong Zhang, Shouwang Shang, Xu Tang 0004, Jie Feng 0003, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Automatic Design Recurrent Neural Network for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is a hot research direction in remote sensing community. In HSIs, there are hundreds of narrow and continuous spectral bands. Due to powerful sequence data processing ability, recurrent neural network (RNN) has shown great potential in HSI classification in recent years. In order to ensure a good classification performance, how to design a proper RNN structure is of great importance. In this paper, an automatic RNN (Auto-RNN) for HSI classification is proposed. Firstly, a number of candidate modules, including ReLU, tanh, sigmoid, and identity, are provided. Then, a policy gradient-based reinforcement learning is us ed to search the feasible deep architecture. Finally, the best RNN structure is selected through evaluating on the validation set. The experimental results demonstrate that the proposed algorithm can yield promising classification performance compared with some existing methods. Jie Feng 0003, Gaiqin Bai, Zizhuo Gao, Xiangrong Zhang, Xu Tang 0004 |
IGARSS | 1 |
| 2021 | Improved SiamRPN++ with Clustering-Based Frame Differencing for Object Tracking of Remote Sensing VideosabstractDeep learning (DL) based object tracking methods have achieved encouraging results on natural videos. However, directly applying these DL-based methods to the vehicle tracking of optical remote sensing videos (ORSV) still faces many challenges. Different with the vehicles in nature videos, most vehicles in ORSV are blurry, small in size, and highly similar to other vehicles. Furthermore, blurred vehicles easily blend into the background and are difficult to distinguish. To solve these problems, an improved Siamrpn++ with clustering-based frame differencing (CFD-SiamRPN++) is proposed. In CFD-SiamRPN++, a clustering method is used to divide the differencing map among adjacent frames into two clusters. By using the statistics of the clustering results, each cluster is judged to the target or not. According to the judgment of each cluster, the differencing map is refined by reducing the background noises and retaining moving information of vehicles. Then, the refined differencing map is fused with the original frame as the input of the tracking network to enhance the discriminative ability of small-sized and blurred vehicles. In the tracking phase, SiamRPN++ is selected as the tracking network due to its feature extraction capability with multi-scale feature fusion and efficient tracking performance. Experiment results based on Jilin-1 ORSV dataset show that the proposed method provides a competitive tracking performance over state-of art deep learning methods. Jie Feng 0003, Bingyu Hui, Yuping Liang, Quanhe Yao, Xiangrong Zhang |
IGARSS | 1 |
| 2021 | Dual-graph convolutional network based on band attention and sparse constraint for hyperspectral band selection
Jie Feng 0003, Zhanwei Ye, Shuai Liu 0016, Xiangrong Zhang, Jiantong Chen, Ronghua Shang, Licheng Jiao |
Knowl. Based Syst. | 1 |
| 2021 | Convolutional Neural Network Based on Bandwise-Independent Convolution and Hard Thresholding for Hyperspectral Band SelectionabstractBand selection has been widely utilized in hyperspectral image (HSI) classification to reduce the dimensionality of HSIs. Recently, deep-learning-based band selection has become of great interest. However, existing deep-learning-based methods usually implement band selection and classification in isolation, or evaluate selected spectral bands by training the deep network repeatedly, which may lead to the loss of discriminative bands and increased computational cost. In this article, a novel convolutional neural network (CNN) based on bandwise-independent convolution and hard thresholding (BHCNN) is proposed, which combines band selection, feature extraction, and classification into an end-to-end trainable network. In BHCNN, a band selection layer is constructed by designing bandwise 1×1 convolutions, which perform for each spectral band of input HSIs independently. Then, hard thresholding is utilized to constrain the weights of convolution kernels with unselected spectral bands to zero. In this case, these weights are difficult to update. To optimize these weights, the straight-through estimator (STE) is devised by approximating the gradient. Furthermore, a novel coarse-to-fine loss calculated by full and selected spectral bands is defined to improve the interpretability of STE. In the subsequent layers of BHCNN, multiscale 3-D dilated convolutions are constructed to extract joint spatial-spectral features from HSIs with selected spectral bands. The experimental results on several HSI datasets demonstrate that the proposed method uses selected spectral bands to achieve more encouraging classification performance than current state-of-the-art band selection methods. Jie Feng 0003, Jiantong Chen, Qigong Sun, Ronghua Shang, Xianghai Cao, Xiangrong Zhang, Licheng Jiao |
IEEE Trans. Cybern. | 1 |
| 2021 | Attention Multibranch Convolutional Neural Network for Hyperspectral Image Classification Based on Adaptive Region SearchabstractConvolutional neural networks (CNNs) have demonstrated outstanding performance on image classification. To classify the hyperspectral images (HSIs), existing CNN-based approaches commonly adopt the architecture using single or several fixed spatial windows as inputs. This kind of architecture may lose contextual information or incorporate heterogeneous information due to the neglect of various land-cover distributions in HSIs. To deal with this problem, a novel attention multibranch CNN method based on adaptive region search (RS-AMCNN) is proposed for HSI classification. In RS-AMCNN, sizes and locations of spatial windows are searched in the nonlocal candidate region adaptively according to sample-specific distribution. These flexible spatial windows are input into several branches of RS-AMCNN. In each branch, convolutional long short-term memories (ConvLSTMs) are merged into CNN from shallow to deep layers, which not only extracts joint spatial-spectral features, but also exploits complementary information among different layers. Then, a branch attention mechanism is devised to emphasize more discriminative branches and suppress less useful ones. It forces RS-AMCNN to extract multiscale and multicontextual attention features for classification. Finally, RS-AMCNN is optimized end-to-end by combining the losses from the ramose classifiers of different branches and the main classifier. Experiments carried on several benchmark HSI data sets demonstrate that RS-AMCNN provides promising classification performance, especially in edge preservation and region uniformity. Jie Feng 0003, Xiande Wu, Ronghua Shang, Chenhong Sui, Jie Li 0001, Licheng Jiao, Xiangrong Zhang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | PSMD-Net: A Novel Pan-Sharpening Method Based on a Multiscale Dense NetworkabstractPan sharpening is used to fuse a low-resolution multispectral (MS) image and a high-resolution panchromatic (PAN) image to obtain a high-resolution MS image. This article proposes PSMD-Net, an end-to-end pan-sharpening method based on a multi-scale dense network. A shallow feature extraction layer (SFEL) extracts the shallow features from the original images, and these are used as an input to a global dense feature fusion (GDFF) network to learn the global features for image reconstruction. A multiscale dense block (MDB) is designed to fully extract the spatial and spectral information from the shallow features in the GDFF network. In the proposed network, multiple MDBs are stacked to extract rich, multi-scale dense hierarchical features, and a global dense connection (GDC) is designed to allow direct connections from the state of the current MDB to all subsequent MDBs to extract more advanced features. The extracted hierarchical features are sent to the global feature fusion layer (GFFL) to adaptively learn the global features for image reconstruction. Finally, global residual learning (GRL) is adopted to force the network to pay more attention to the changing part of the image. We perform experiments on simulated and real data from WorldView-2 and WorldView-3 satellites. Visual and quantitative assessment results demonstrate that PSMD-Net yields higher-resolution fusion images than the state-of-the-art methods. Jinye Peng 0001, Lu Liu 0025, Jun Wang 0078, Erlei Zhang, Xuan Zhu 0003, Yongqin Zhang, Jie Feng 0003, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2020 | Small Object Detection in Optical Remote Sensing Video with Motion Guided R-CNNabstractDeep learning (DL) based object detection methods have been making great achievements for natural images, which guides the vehicle detection of optical remote sensing videos (ORSV). Compared with natural images, objects in ORSV are smaller and blurrier, and most of vehicles are crowded. Thus, it is difficult for DL to detect these small objects only using the single-frame image. To address this problem, a motion guided R-CNN (MG-RCNN) is proposed. In MG-RCNN, motion information from consecutive frames is extracted by the mean differencing method and merged into apparent information to obtain motion-related discriminative features. Then, high-quality proposals are generated on the feature maps by mini-region proposal network (MRPN). For small targets, an improved loss function is defined by incorporating smooth factor, which makes the regression of shapes more stable. Experiments on ORSV demonstrate the proposed method shows superior detection performance over state-of-the-art deep learning methods. Jie Feng 0003, Yuping Liang, Zhanwei Ye, Xiande Wu, Dening Zeng, Xiangrong Zhang, Xu Tang 0004 |
IGARSS | 1 |
| 2020 | Hyperspectral Image Classification Based on Semi-Supervised Dual-Branch Convolutional Autoencoder with Self-AttentionabstractDeep learning method shows its powerful classification performance with sufficient available data. However, the labeled data is limited in hyperspectral images (HSIs). Semi-supervised algorithms have unique advantages on dealing with this problem. Therefore, a semi-supervised convolutional neural network is proposed in this paper. It consists of two branches, which use limited labeled samples and a large number of unlabeled samples, respectively. The first branch includes an encoder-decoder model to extract contextual information of unlabeled samples. The other one uses the similar construction except extra classification layers to extract discriminative features of labeled samples. In order to fuse contextual and discriminative information, we cascade the features of low-level layers from different branches. Furthermore, self-attention is added to the first branch, which focuses more on the global information for classification. The experiment results show that the proposed model provides a competitive result compared with state-of-the-art methods. Jie Feng 0003, Zhanwei Ye, Yuping Liang, Xu Tang 0004, Xiangrong Zhang |
IGARSS | 1 |
| 2020 | Unsupervised Manifold-Preserving and Weakly Redundant Band Selection Method for Hyperspectral ImageryabstractHyperspectral band selection is of great value to alleviate the curse of dimensionality. For many band selection methods, however, the neglect of bandwise usefulness tends to result in the loss of valuable bands, but the retention of useless ones; consequently, this causes deterioration of the classification performance. In this sense, bandwise significance should be emphasized. To address this issue, this article proposes a manifold-preserving and weakly redundant (MPWR) unsupervised band selection method. In the method, a manifold-preserving band-importance metric is put forward to measure the bandwise essentiality. This ensures the retention of bands involving abundant intrinsic structures conductive to classification. Specifically, aimed at obtaining the presented band-importance metric, an attainment algorithm is presented, which mainly relies on the embedding learning and linear regression, followed by the introduction of multi-normalization combination. In addition, concerning the massive redundancy caused by the highly correlated bands, MPWR further establishes a constrained band-weight optimization model. Then, both bandwise manifold-preserving capability and intraband correlation are fully integrated into the band selection process. To solve the problem, a corresponding algorithm within the framework of the alternating direction method of multipliers (ADMM) is also developed. Regarding evaluating the effectiveness of the proposed method, comparative experiments with the state-of-the-art methods are conducted on three public hyperspectral data sets. Experimental results demonstrate the superiority and robustness of MPWR. Chenhong Sui, Chang Li 0001, Jie Feng 0003, Xiaoguang Mei |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Hyperspectral Band Selection Based On Ternary Weight Convolutional Neural NetworkabstractIn this paper, a novel ternary weight convolution neural network (TWCNN) is proposed for band selection of hyperspectral images. TWCNN constructs deep-wise convolution layer with 1×1 filters as the first layer of the network, which is used for band selection. In the deep-wise convolution layer, weights are constrained to -1, 0, or 1. -1 and 1 represent that the corresponding band is selected, while 0 indicates it's not. TWCNN constructs subsequent layers to extract features and classify for selected spectral bands. It combines band selection, feature extraction and classification into a unified optimization procedure, which makes it to achieve end-to-end band selection and classification. Furthermore, the constraints of the number of spectral bands is added to the cost function of TWCNN. The specific number of spectral bands can be selected. The experiment results show that the proposed model provides a competitive result to state-of-the-art methods. Jie Feng 0003, Jiantong Chen, Xiangrong Zhang, Xu Tang 0004, Xiande Wu |
IGARSS | 1 |
| 2019 | Joint Multilayer Spatial-Spectral Classification of Hyperspectral Images Based on CNN and ConvlstmabstractThe following topics are dealt with: remote sensing; geophysical image processing; synthetic aperture radar; radar imaging; remote sensing by radar; image classification; learning (artificial intelligence); feature extraction; vegetation; image resolution. Jie Feng 0003, Xiande Wu, Jiantong Chen, Xiangrong Zhang, Xu Tang 0004 |
IGARSS | 1 |
| 2019 | Classification of Hyperspectral Images Based on Multiclass Spatial-Spectral Generative Adversarial NetworksabstractGenerative adversarial networks (GANs) are famous for generating samples by training a generator and a discriminator via an adversarial procedure. For hyperspectral image classification, the collection of samples is always difficult. However, directly applying GAN to hyperspectral image classification exists two problems. One is that the generated samples lack discriminative information. Meanwhile, the discriminator has no discriminative ability for multiclassification. Another is that spatial and spectral information requires to be considered in hyperspectral image classification simultaneously. To address these problems, a novel multiclass spatial-spectral GAN (MSGAN) method is proposed. In MSGAN, two generators are devised to generate the samples containing spatial and spectral information, respectively, and the discriminator is devised to extract joint spatial-spectral features and output multiclass probabilities. Moreover, novel adversarial objectives for multiclass are defined. The discriminator is devised to predict training samples belonging to true classes and generated samples belonging to all the classes with the same probability. The generators are devised to make the discriminator mistake. By adversarial learning between the discriminator and generators, the classification performance of the discriminator is promoted with the assistance of discriminative generated samples. Experimental results on hyperspectral images demonstrate that the proposed method achieves encouraging classification performance compared with several state-of-the-art methods, especially with the limited training samples. Jie Feng 0003, Haipeng Yu, Xianghai Cao, Xiangrong Zhang, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Spatial-Spectral Graph-Based Nonlinear Embedding Dimensionality Reduction for Hyperspectral Image ClassificaitonabstractDimensionality reduction (DR) is one of the most important tasks to improve the performance of hyperspectral images classification. Recently, a sparse and low-rank graph embedding based method (SLGE) has been proposed to describe the intrinsic structure of data combined with the local and global constraint simultaneously, which is effective to reduce the dimension of hyperspectral data and obtain a better classification accuracy. However, SLGE is based on an assumption that low-dimensional feature can be obtained utilizing a linear projection. Its performance may degrade under nonlinearly distributed data. Moreover, spatial prior of HSI is not considered in the framework. In this paper, we proposed a novel dimensionality reduction method named spatial-spectral graph-based non-linear embedding (SSGNE). To generate a new graph-trained data, the segmentation strategy based on superpixel is adopted. The spatial-spectral graph is constructed by constraining the sparsity and low-rankness simultaneously on graph-trained data set. Finally, the kernel trick is adopted to extend the general graph embedding framework to nonlinearly space, which fully considers the complexity of real data. Experimental results show that the proposed method outperforms the state-of-the-art methods in terms of the classification accuracy. Xiangrong Zhang, Yaru Han, Ning Huyan, Chen Li 0011, Jie Feng 0003, Xiaoxiao Ma 0003 |
IGARSS | 5 |
| 2017 | Hyperspectral image classification based on stacked marginal discriminative autoencoderabstractIn this paper, a novel stacked marginal discriminative autoencoder (SMDAE) method is proposed for hyperspectral image classification. It uses a deep neural network to learn discriminative features from hyperspectral images automatically. In hyperspectral images, the collection of training samples is difficult. When the number of training samples is not enough, these training samples are difficult to estimate the statistical distribution of hyperspectral images accurately. In order to solve the small sample problem and improve the classification performance of the autoencoder, the marginal samples are selected through the distribution characteristics of samples. The marginal samples are searched based on k nearest neighbors between different classes. These samples are used to fine-tune the SMDAE network. The experimental results show that the proposed SMDAE method can achieve satisfying performance under small training set. Jie Feng 0003, Liguo Liu, Xiangrong Zhang, Rongfang Wang, Hongying Liu 0001 |
IGARSS | 1 |
| 2016 | Terrain classification with Polarimetric SAR based on Deep Sparse Filtering NetworkabstractA new method for Polarimetric Synthetic Aperture Radar (PolSAR) terrain classification based on Deep Sparse Filtering Network (DSFN) is proposed in this paper. It uses a novel deep learning network to learn features from the input raw data automatically. And the spatial information between pixels on PolSAR image is combined into the input data. Moreover, unlike the conventional deep networks, the DSFN only needs to tune very few parameters during pre-training and fine-tuning. A real PolSAR data is used to verify the proposed method. Experimental results show that the proposed DSFN is efficient with less parameters and effectively improves the classification accuracy compared with conventional deep networks. Hongying Liu 0001, Qiang Min, Jin Zhao 0002, Shuyuan Yang 0001, Biao Hou, Jie Feng 0003, Licheng Jiao |
IGARSS | 7 |
| 2016 | Unsupervised feature selection based on maximum information and minimum redundancy for hyperspectral images
Jie Feng 0003, Licheng Jiao, Fang Liu 0001, Tao Sun 0007, Xiangrong Zhang |
Pattern Recognit. | 1 |
| 2016 | Multiple Kernel Learning Based on Discriminative Kernel Clustering for Hyperspectral Band SelectionabstractIn hyperspectral images, band selection plays a crucial role for land-cover classification. Multiple kernel learning (MKL) is a popular feature selection method by selecting the relevant features and classifying the images simultaneously. Unfortunately, a large number of spectral bands in hyperspectral images result in excessive kernels, which limit the application of MKL. To address this problem, a novel MKL method based on discriminative kernel clustering (DKC) is proposed. In the proposed method, a discriminative kernel alignment (KA) (DKA) is defined. Traditional KA measures kernel similarity independently of the current classification task. Compared with KA, DKA measures the similarity of discriminative information by introducing the comparison of intraclass and interclass similarities. It can evaluate both kernel redundancy and kernel synergy for classification. Then, DKA-based affinity-propagation clustering is devised to reduce the kernel scale and retain the kernels having high discrimination and low redundancy for classification. Additionally, an analysis of necessity for DKC in hyperspectral band selection is provided by empirical Rademacher complexity. Experimental results on several hyperspectral images demonstrate the effectiveness of the proposed band selection method in terms of classification performance and computation efficiency. Jie Feng 0003, Licheng Jiao, Tao Sun 0007, Hongying Liu 0001, Xiangrong Zhang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Scaling cut criterion-based discriminant analysis for supervised dimension reduction
Xiangrong Zhang, Yudi He, Licheng Jiao, Ruochen Liu 0006, Jie Feng 0003 |
Knowl. Inf. Syst. | 5 |
| 2015 | Imbalanced Hyperspectral Image Classification Based on Maximum MarginabstractHyperspectral remote sensing images own rich spectral information to distinguish different land-cover classes. Sometimes, it may encounter the case that some classes have much fewer pixels than other classes. In this case, traditional classification methods are not appropriate because they are prone to assign all the pixels to the classes with a large number of pixels. For such an imbalanced problem, ensemble learning is a good method by partitioning the majority classes into different groups with small sizes. However, the existing ensemble schemes are independent of classifiers, which will not get the best performance for a certain classifier. In this letter, the selected classifier, i.e., a support vector machine (SVM), is considered in an ensemble procedure to improve the classification accuracy. Specifically, the criterion of the SVM, i.e., the maximum margin, is adopted to guide the ensemble learning procedure for imbalanced hyperspectral image classification. Experiments state that our method obtains higher classification accuracy than the SVM and several representative imbalanced classification methods for hyperspectral images. Tao Sun 0007, Licheng Jiao, Jie Feng 0003, Fang Liu 0001, Xiangrong Zhang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2015 | Mutual-Information-Based Semi-Supervised Hyperspectral Band Selection With High Discrimination, High Information, and Low RedundancyabstractThe large number of spectral bands in hyperspectral images provides abundant information to distinguish different land covers. However, these spectral bands have much redundancy and bring an extra computational burden. Thus, band selection is important for hyperspectral images. Since the labeled samples are difficult to obtain, a semi-supervised criterion based on maximum discrimination and information (MDI) is defined by using both limited labeled samples and sufficient unlabeled samples. This MDI criterion aims to select the most highly discriminative and informative bands, but it is hard to accurately calculate. Therefore, a novel criterion based on high discrimination, high information, and low redundancy (DIR) is proposed as its low-order approximation. Moreover, from an information theory perspective, a theoretical proof is given that many traditional semi-supervised feature selection criteria are the low-order approximations of this MDI criterion. Compared with them, the proposed criterion needs more relaxed approximation conditions. To search and optimize the proposed criterion, a novel clonal selection algorithm is proposed, where the adaptive clone and mutation operators are devised to speed up the convergence. Experimental results on hyperspectral images demonstrate the effectiveness of the proposed semi-supervised band selection method. Jie Feng 0003, Licheng Jiao, Fang Liu 0001, Tao Sun 0007, Xiangrong Zhang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2014 | A simplified low rank and sparse graph for semi-supervised learning
Miaoyun Zhao, Licheng Jiao, Jie Feng 0003 |
Neurocomputing | 3 |
| 2014 | Hyperspectral Band Selection Based on Trivariate Mutual Information and Clonal SelectionabstractBand selection is an important preprocessing step for hyperspectral data processing. It involves two crucial problems, i.e., suitable measure criterion and effective search strategy. Mutual information (MI) has been widely used as the measure criterion for its nonlinear and nonparametric characteristics. For efficient calculation, traditional MI-based criteria commonly use bivariate MI (BMI) to approximate the ideal MI-based criterion. However, these BMI-based criteria may miss the bands having discriminative information and do not give the condition of the approximation. In this paper, a novel criterion based on trivariate MI (TMI) is proposed to measure the redundancy for classification. From the multivariate MI perspective, the proposed TMI-based and traditional BMI-based criteria are proved as the low-order approximations of the ideal criterion under some assumptions. Compared with the BMI-based criteria, a more relaxed assumption condition is required for the TMI-based criterion. To alleviate the problem of few labeled samples existing in hyperspectral images, the TMI-based criterion is extended to the semisupervised TMI-based (STMI) method by adding a graph regulation term. Additionally, to search an appropriate band subset by the TMI- and STMI-based criteria, a new clonal selection algorithm (CSA) is proposed. In CSA, integer encoding and adaptive operators are devised to reduce space and time cost. Experimental results demonstrate the effectiveness of the proposed algorithms for hyperspectral band selection. Jie Feng 0003, Licheng Jiao, Xiangrong Zhang, Tao Sun 0007 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2013 | Spatial-spectral classification based on group sparse coding for hyperspectral imageabstractIn this paper, a novel hyperspectral image classification method is proposed, based on group sparse coding. The method is based on this acknowledgement that larger spatial variation exists in high spatial resolution hyperspectral image, which degrades the separability of hyperspectral image. In order to obtain a smooth representation, each pixel and its spatial neighbors are coded together by group sparse coding. Although nothing about class information is included, the neighbor pixels in a small spatial window are inclined to belong to the same class. Thus, that will reduce the within-class scatter and be favorable to the classification task. Then, the obtained sparse representation vectors are used for hyperspectral image classification with SVM. Experimental results show that our method exceeds the classical classification algorithms in accuracy and regional consistency. Xiangrong Zhang, Peng Weng, Jie Feng 0003, Erlei Zhang, Biao Hou |
IGARSS | 3 |
| 2013 | Selective multiple kernel learning for classification with ensemble strategy
Tao Sun 0007, Licheng Jiao, Fang Liu 0001, Shuang Wang 0001, Jie Feng 0003 |
Pattern Recognit. | 5 |
| 2013 | Robust non-local fuzzy c-means algorithm with edge preservation for SAR image segmentation
Jie Feng 0003, Licheng Jiao, Xiangrong Zhang, Maoguo Gong, Tao Sun 0007 |
Signal Process. | 1 |
| 2012 | SAR image change detection based on low rank matrix decompositionabstractIn this paper we propose an unsupervised approach for SAR image change detection task. A new method based on compressed sensing is applied. First using the PPB method for the speckle reduction, and then the logarithm ratio method is applied to generate a simple change map, and then the compressed sensing-based method is used to part the change map into a low rank part and a sparse part, where the sparse part is correspond to the changed area, finally k-means algorithm is applied to cluster the sparse part into two clusters. Experiment results show the effectiveness and feasibility of the proposed method. Xiangrong Zhang, Yaoguo Zheng, Jie Feng 0003, Shuiping Gou |
IGARSS | 3 |
| 2011 | Bag-of-Visual-Words Based on Clonal Selection Algorithm for SAR Image ClassificationabstractSynthetic aperture radar (SAR) image classification involves two crucial issues: suitable feature representation technique and effective pattern classification methodology. Here, we concentrate on the first issue. By exploiting a famous image feature processing strategy, Bag-of-Visual-Words (BOV) in image semantic analysis and the artificial immune systems (AIS)'s abilities of learning and adaptability to solve complicated problems, we present a novel and effective image representation method for SAR image classification. In BOV, an effective fused feature sets for local feature representation are first formulated, which are viewed as the low-level features in it. After that, clonal selection algorithm (CSA) in AIS is introduced to optimize the prediction error of k-fold cross-validation for getting more suitable visual words from the low-level features. Finally, the BOV features are represented by the learned visual words for subsequent pattern classification. Compared with the other four algorithms, the proposed algorithm obtains more satisfactory and cogent classification experimental results. Jie Feng 0003, Licheng Jiao, Xiangrong Zhang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2010 | Adaptive ranks clone and k-nearest neighbor list-based immune multi-objective optimizationabstractArtificial immune systems (AIS) are computational systems inspired by the principles and processes of the vertebrate immune system. The AIS‐based algorithms typically exploit the immune system's characteristics of learning and adaptability to solve some complicated problems. Although, several AIS‐based algorithms have proposed to solve multi‐objective optimization problems (MOPs), little focus have been placed on the issues that adaptively use the online discovered solutions. Here, we proposed an adaptive selection scheme and an adaptive ranks clone scheme by the online discovered solutions in different ranks. Accordingly, the dynamic information of the online antibody population is efficiently exploited, which is beneficial to the search process. Furthermore, it has been widely approved that one‐off deletion could not obtain excellent diversity in the final population; therefore, ak‐nearest neighbor list (wherekis the number of objectives) is established and maintained to eliminate the solutions in the archive population. Thek‐nearest neighbors of each antibody are founded and stored in a list memory. Once an antibody with minimal product ofk‐nearest neighbors is deleted, the neighborhood relations of the remaining antibodies in the list memory are updated. Finally, the proposed algorithm is tested on 10 well‐known and frequently used multi‐objective problems and two many‐objective problems with 4, 6, and 8 objectives. Compared with five other state‐of‐the‐art multi‐objective algorithms, namely NSGA‐II, SPEA2, IBEA, HYPE, and NNIA, our method achieves comparable results in terms of convergence, diversity metrics, and computational time. Licheng Jiao, Maoguo Gong, Jie Feng 0003 |
Comput. Intell. | 4 |