EDBT 2026 Demo / reviewers in the wild / expert
Hao Zhu 0009
dblp:10/3520-9
· DBLP profile ↗
51ranked-venue papers
9as first author
46since 2021 · last 2026
0000-0003-4411-0933ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 28 · 6 first-author · 25 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Direction Perception via Atomic Dot-Product Operators for Rotation-Invariant Point Clouds LearningabstractPoint cloud processing has become a cornerstone technology in many 3D vision tasks. However, arbitrary rotations introduce variations in point cloud orientations, posing a long-standing challenge for effective representation learning. The core of this issue is the disruption of the point cloud's intrinsic directional characteristics caused by rotational perturbations. Recent methods attempt to implicitly model rotational equivariance and invariance, preserving directional information and propagating it into deep semantic spaces. Yet, they often fall short of fully exploiting the multiscale directional nature of point clouds to enhance feature representations. To address this, we propose the Direction-Perceptive Vector Network (DiPVNet). At its core is an atomic dot-product operator that simultaneously encodes directional selectivity and rotation invariance--endowing the network with both rotational symmetry modeling and adaptive directional perception. At the local level, we introduce a Learnable Local Dot-Product (L2DP) Operator, which enables interactions between a center point and its neighbors to adaptively capture the non-uniform local structures of point clouds. At the global level, we leverage generalized harmonic analysis to prove that the dot-product between point clouds and spherical sampling vectors is equivalent to a direction-aware spherical Fourier transform (DASFT). This leads to the construction of a global directional response spectrum for modeling holistic directional structures. We rigorously prove the rotation invariance of both operators. Extensive experiments on challenging scenarios involving noise and large-angle rotations demonstrate that DiPVNet achieves state-of-the-art performance on point cloud classification and segmentation tasks. Chenyu Hu, Hao Zhu 0009, Biao Hou |
AAAI | 3 |
| 2026 | A trust-aware singular fusion network for multimodal image classification
Wenping Ma 0001, Mengru Ma, Hekai Zhang, Hao Zhu 0009, Licheng Jiao |
Neurocomputing | 5 |
| 2026 | Recurrent progressive fusion-based learning for multi-source remote sensing image classification
Hao Zhu 0009, Biao Hou, Wenhao Zhao, Xiaoyu Yi 0002, Wenping Ma 0001, Licheng Jiao |
Pattern Recognit. | 2 |
| 2026 | A Progressive Semi-Distillation Model for Dual-Source Remote Sensing Image ClassificationabstractPanchromatic images (PANs) and multispectral (MS) images (MSs) are widely used for dual-source remote sensing image classification, gradually becoming a research hotspot. However, making the most of dual-source image information with insufficiently labeled samples is a significant challenge. This article proposes a progressive semi-distillation model (PSDM) to classify dual-source remote sensing images with insufficient samples. We design a framework of rookie teacher network (RTN)-teaching assistant system (TAS)-student grouping network (SGN) in the case of a traditional teacher network (TN) (i.e., rookie TN (RTN)) that does not provide excellent guidance to student network (SN) due to insufficient samples. The PSDM expands the samples and compresses the space through the RTN-SGN structure to cope with the dilemma of insufficient samples. To make RTN better guide the SGN, we design TAS, which can gradually guide SGN to learn the samples from easy to difficult. It can also further assist SGN training to improve the classification performance of SGN with insufficient samples. We design SGN and add cooperation and correction mechanism to better learn dual- source information. These strategies can eliminate SGN's over-dependence on the RTN, help SGN outperform the RTN, and achieve the effect of semi-distillation. Experimental results and theoretical analysis have sufficiently pointed out the proposed method's accuracy, efficiency, and robustness under insufficient sample situations. Our model is available at https://github.com/MarjordCpz/PSDM. Hao Zhu 0009, Peizhou Cao, Licheng Jiao, Biao Hou, Xiaoyu Yi 0002, Wenhao Zhao, Wenping Ma 0001 |
IEEE Trans. Cybern. | 1 |
| 2026 | MCIB: Multi-Modal Complementary Information Bottleneck for Hyperspectral and LiDAR ClassificationabstractThe effective fusion of multi-modal remote sensing images, particularly hyperspectral imagery (HSI) and light detection and ranging (LiDAR) data, is pivotal for accurate land use and land cover (LULC) classification. However, this process is hindered by two inherent challenges: pervasive data redundancy and the underutilization of cross-modal complementarity, largely due to the lack of a unifying theoretical framework. To address these limitations, we propose the multi-modal complementary information bottleneck (MCIB) framework, which extends the IB principle to learn compact, sufficient, and complementary representations for multi-modal scenes. From a theoretical perspective, we formalize the MCIB objective and introduce structured priors to derive tractable information-theoretic bounds, providing a principled and computationally feasible approach to reduce redundancy and enhance complementarity simultaneously. Building on the obtained theoretical insights, we design an end-to-end variational optimization strategy with a novel supervised conditional InfoNCE (SCInfoNCE). Efficiently reusing existing model components, this new supervised contrastive method optimizes the conditional mutual information terms crucial for synergy. Extensive experiments on benchmark HSI-LiDAR datasets demonstrate superior classification performance of MCIB. This work not only fills a theoretical gap in multi-modal representation learning, but offers a robust and principled solution for LULC classification using complex heterogeneous remote sensing images. Hao Zhu 0009, Bo Yang 0047, Changzhe Jiao, Jie Feng 0003, Jinjian Wu |
IEEE Trans. Image Process. | 2 |
| 2026 | Geo-SelfSSC: Integrating Dense Geometric Priors for Enhanced Self-Supervised Semantic Scene CompletionabstractAccurate 3D scene understanding is vital for applications like autonomous driving and robotics. However, existing voxel-based methods struggle with the reliance on large-scale labeled data, inherent voxel-pixel misalignments, and high computational costs. Recent self-supervised Semantic Scene Completion (SSC) methods using Neural Radiance Fields (NeRF) reduce 3D annotation needs but assume Lambertian surfaces and rely on photometric consistency, which fails in real-world scenes with non-Lambertian effects and sparse camera coverage. In this paper, we introduce Geo-SelfSSC, a self-supervised framework that leverages temporal information from consecutive frames and slight variations in camera poses for supervision, while integrating dense geometric cues to enhance the reconstruction quality and efficiency of neural implicit models. Specifically, Geo-SelfSSC leverages depth priors along with complementary semi-local and local geometric supervision to facilitate efficient sampling, while ensuring effective complementarity between photometric and geometric cues. Our method yields significant improvements in reconstructing reflective and under-observed regions where conventional photometric-based strategies struggle. Comprehensive experiments on challenging tasks demonstrate that Geo-SelfSSC not only achieves strong results in semantic scene completion but also establishes competitive performance in geometry-only 3D occupancy prediction and monocular depth estimation. Our code and models are available at: https://github.com/Xidian AIGroup190726/GeoSelfSSC. Hao Zhu 0009, Pute Guo, Longsheng Qu, Jinjian Wu |
IEEE Trans. Multim. | 1 |
| 2025 | Semi-SNN: Biological-inspired semi-supervised image classification with spiking neural networks
Biao Hou, Chuanfeng Ma, Leida Li, Hao Zhu 0009, Licheng Jiao |
Neurocomputing | 6 |
| 2025 | Dual-model spiking neural network for remote sensing image classification using mutual knowledge distillation
Biao Hou, Hao Zhu 0009, Yifan Ge, Licheng Jiao |
Neurocomputing | 5 |
| 2025 | A two-stage strategy for brain-inspired unsupervised learning in spiking neural networks
Chuanfeng Ma, Biao Hou, Leida Li, Hao Zhu 0009, Dou Quan, Licheng Jiao |
Neurocomputing | 6 |
| 2025 | Attribute-guided feature fusion network with knowledge-inspired attention mechanism for multi-source remote sensing classification
Changzhe Jiao, Bo Yang 0047, Hao Zhu 0009, Jinjian Wu |
Neural Networks | 4 |
| 2025 | An Adaptive Dual-Supervised Cross-Deep Dependency Network for Pixel-Wise ClassificationabstractWith the advancement of remote sensing (RS) technology and satellite observation, the task of fusing multisource data, such as multispectral (MS) and panchromatic (PAN) images, has become increasingly important. However, image fusion involving certain semantic differences can hinder the model’s ability to learn effective feature mappings. To reconstruct richer and more consistent features during fusion, we propose an adaptive dual-supervised cross-deep dependency network (ADCD-Net), which consists of two training stages. Stage I uses a semantic perceptual self-supervision strategy (SPS) to learn deep features across different modalities, thereby reducing semantic differences while mining its own non-singular features. Stage II uses the deep temporal Mamba module (DTM-Module) to interactively learn the output of each network layers, which are able to take part in the deep feature reinforcement and improve the classification performance of semantic information. Finally, to eliminate channel redundancy during the two-stage network training process while enhancing spatial location memory and feature discrimination in the 2-D features, we propose a deformable interactive attention module (DIA-Module) to further bolster feature representation capabilities. Additionally, we conduct comparative and transfer experiments on multiple RS datasets, achieving outstanding classification results. Our code is available athttps://github.com/ChenC1027/ADCD-Net. Wenping Ma 0001, Mengru Ma, Hekai Zhang, Hao Zhu 0009, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Dual-Path Prototype Feature Decoupling Alignment Network for Panchromatic and Multispectral ClassificationabstractIn recent years, with the rapid advancements and widespread application of satellite photography technology, it has become increasingly possible to obtain high-quality panchromatic (PAN) and multispectral (MS) data, which has provided new opportunities and challenges for multisource information fusion and classification research. Remote sensing data have the characteristics of small interclass differences and large intraclass differences, which easily leads to category confusion in network learning. In addition, how to fully tap the advantages of multisource data, better align multisource features, improve classification accuracy, and achieve collaborative classification are key issues that need to be solved urgently. In this article, a dual-path prototype feature decoupling alignment network (DPFDA-Net) is designed to solve the above issues. The network consists of two components: a prototype feature embedding (PFE) module and a feature alignment module (FAM) based on prototype decoupling. In the feature extraction stage, the PFE module uses the prototype concept to learn the discriminative prototype features of each category of the dual-source data separately, making the boundaries between categories more obvious. The FAM operates at the dual-source prototype feature level and achieves feature alignment by decoupling single-source prototype features and performing feature transformation to supplement the missing information of another data source. Finally, we use the aligned features for classification. The results of the experiment demonstrate that our approach has made significant progress in improving classification precision. The code is available athttps://github.com/Xidian-AIGroup190726/DPFDANet. Wenping Ma 0001, Yanshan Guo, Hao Zhu 0009, Wenhao Zhao, Mengru Ma, Yue Wu 0004, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | A Diff-Attention Aware State-Space Fusion Model for Remote Sensing ClassificationabstractMultispectral (MS) and panchromatic (PAN) images describe the same land surface, so these images not only have their own advantages, but also share a significant amount of redundant information. In order to separate similar information and each modality’s unique advantages, thereby reducing feature redundancy at the fusion stage, this paper introduces a diff-attention aware state space fusion model (DASF-Model) for multimodal remote sensing image classification. Based on the selective state space model, a cross-modal diff-attention module (CDAM) is designed to extract and separate the common features and their respective dominant features of MS and PAN images. Specifically, space preserving visual mamba (SPVM) retains image spatial features and captures local features by appropriately optimizing visual mamba’s input. Considering that features in the fusion stage will have large semantic differences after feature separation and traditional mean fusion method fails to effectively integrate these features with significant discrepancies, an attention-aware linear fusion module (ALFM) is proposed. It performs pixel-wise linear fusion by calculating influence coefficients. This mechanism can fuse features with large semantic differences while keeping the feature size unchanged. Empirical evaluations indicate that the presented method achieves better results than alternative approaches. The relevant code can be found at: https://github.com/AVKSKVL/DAS-F-Model. Wenping Ma 0001, Boyou Xue, Mengru Ma, Hekai Zhang, Hao Zhu 0009 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Dense-Weak Ship Detection Based on Foreground-Guided Background Generation Network in SAR ImagesabstractCurrently, ship detection based on Synthetic Aperture Radar (SAR) images still faces significant challenges, particularly in detecting weak and densely distributed ships within complex backgrounds. In areas such as ports and land, the complex background features often resemble those of densely distributed ships, leading to reduced detection accuracy. Additionally, the overlapping and mutual interference of features among dense ships can cause the network to miss detections or produce false positives. Therefore, this paper proposes a Foreground-Guided Background Generation Network (FGBG-Net), which includes a Gaussian Foreground Localization (GFL) model and a Background Feature Removal (BFR) module. The GFL module identifies the approximate high-probability regions of ship foregrounds on the feature map, guiding the network to focus on these regions. The BFR module then progressively removes background interference features based on the positions provided by the GFL module, generating feature maps that are more suitable for detecting weak and dense ships. Our network has been validated on multiple SAR ship datasets, and the experimental results demonstrate noticeable performance improvements, with a mean Average Precision (mAP) increase of 3.4% on the SSDD and HRSID datasets. The relevant code is available at the following link: https://github.com/Xidian-AIGroup190726/FBGBNet/tree/master. Wenping Ma 0001, Xiaoting Yang, Hao Zhu 0009, Xiaoteng Wang, Biao Hou, Mengru Ma, Yue Wu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | FAFormer: Frequency-Analysis-Based Transformer Focusing on Correlation and Specificity for PansharpeningabstractPan-sharpening refers to fusing remote sensing multispectral (MS) and panchromatic (PAN) images to generate high-resolution multispectral (HR-MS) images. Recent advancements in deep learning-based pan-sharpening techniques have shown promising results. However, they face the following two issues. On one hand, there is a modality gap between MS and PAN images. Directly fusing them can lead to spectral and spatial distortions. On the other hand, the fusion process is prone to information loss, which can lead to image blurriness. To tackle these issues, we develop a Transformer-based model: FAFormer, which incorporates frequency analysis and focuses on the correlation and specificity of the PAN and MS images. Focusing on correlation can reduce the spectral and spatial distortions while focusing on specificity can reflect the specific information from MS and PAN images in the fusion result. We utilize the Discrete Wavelet Transform (DWT) to obtain the correlate and specific features. We introduce bijective functions based on the Transformer to design an Integrated Attention Block (IAB). As a critical component of the model, it effectively utilizes the correlation and specificity of the two images. In designing the model’s overall framework, we employ a Correlative Feature Attention Module (CFAM) to leverage the correlation between MS and PAN. We utilize a Specific Feature Attention Module (SFAM) to integrate specific information into fused features gradually. Experimental results show that our method improves pan-sharpening performance and has practical value. Codes are available at https://github.com/Xidian-AIGroup190726/FAFormer. Yifan Meng, Hao Zhu 0009, Xiaoyu Yi 0002, Biao Hou, Shuang Wang 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Sample-Level Improved Cross-Source Contrastive Learning for PAN and MS Joint ClassificationabstractIn recent years, the number and ways of acquiring panchromatic images (PAN) and multispectral images (MS) have increased, and manual labeling costs have also increased. Processing these data efficiently has become a challenge. In this paper, we propose a sample-level improved cross-source contrastive learning method for PAN and MS joint classification (SLCL), which aims to provide a self-supervised pre-training model using unlabeled samples for downstream joint classification using a small quantity of labeled samples. First, we propose a sample weighting and screening (SWS) strategy, which enables the model to learn inter- and intra-source sample representations, while balancing the interference from false samples so that the model learns true samples. It solves the problems of homologous similar features embedded far away and false negative samples bringing the wrong learning direction, which exist in existing contrastive learning methods. In addition, we design a hard sample learning (HSL) module for the problem of mining and optimization of hard samples. The module efficiently mines hard samples and uses a new loss function to make the model more focused on hard sample optimization. It further improves the accuracy of pre-training models for downstream tasks. Our method performs best on multiple datasets, and it is experimentally validated and analyzed. The code is available at: https://github.com/Xidian-AIGroup190726/SLCL. Pengyu Tian, Hao Zhu 0009, Biao Hou, Pute Guo, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | EGPO: Enhanced Guidance and Pseudo-Label Optimization for Semi-Supervised Semantic Segmentation of Remote Sensing ImagesabstractOwing to complex remote sensing image features boosting manually labeled costs, semi-supervised semantic segmentation learns from limited labeled and abundant unlabeled data to alleviate this dilemma. However, there are still many challenges in the practical application of this technology, such as incorrect unsupervised information misleading the model and errors in pseudo-labels causing error accumulation. This paper proposes a semi-supervised semantic segmentation method for remote sensing images. The enhanced guided learning module we designed uses a label guided model to predict unlabeled data in the correct direction, enriching ground features and unsupervised information, and alleviating the problems of misprediction and consistency regularization failure caused by the lack of labels. At the same time, facing the noise generated during the data augmentation process, our designed unsupervised loss dynamic screening module aims to locate and suppress the noise adaptively. In addition, in the face of inevitable erroneous predictions in pseudo-labels, we design a pixel category selection module that produces high-quality and high-confidence pseudo-labels through multi-step filtering and dual model fusion. Ultimately, by conducting experiments on the DFC22, iSAID, MER, GID-15, and Vaihingen datasets, we successfully verified the effectiveness of our proposed method. The source code has been made public: https://github.com/Xidian-AIGroup190726/EGPO. Xifeng Xue, Hao Zhu 0009, Longsheng Qu, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Contour-Aware Dynamic Low-High Frequency Integration for Pan-SharpeningabstractPan-sharpening is the process of fusing panchromatic (PAN) and multispectral (MS) images. Its critical focus lies in accurately capturing the contour information from the PAN image during the fusion process and presenting it at a high resolution. However, existing deep learning methods lack the precise capture of delicate and smooth contour information, resulting in contour diffusion that affects the fusion results. Therefore, we introduce contourlet decomposition to capture multiscale directional delicate contour features and construct multiscale graph structures for semantic mining of dual-source contour features, continually updated through dynamic learning. By incorporating global features, we guide the multihead attention mechanism with directional decoding, enabling the network to pay more attention to high-resolution contour features, thereby gaining an advantage in image reconstruction. Cross-decoding between modalities provides strong representational capabilities for the advantageous features of both modalities, effectively enhancing the sharpening effect. Our algorithm achieves state-of-the-art results, and its effectiveness and advantages have been thoroughly validated across multiple datasets, including GaoFen-2, WorldView2, WorldView3, etc. Our code is available athttps://github.com/Xidian-AIGroup190726/CDFInet. Xiaoyu Yi 0002, Hao Zhu 0009, Pute Guo, Biao Hou, Bo Ren 0001, Xiaoteng Wang, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Adaptive Complex Wavelet Informed Transformer OperatorabstractVisual transformers have achieved great success in representation learning. This is mainly due to efficient token dependency modeling via self-attention. However, the computational burden increases sharply as the input pixels increase. Although recent Fourier-based global frequency-domain mixing methods attempt to improve the efficiency of transformers for high-resolution image inputs, the Fourier operator has limited ability to capture the local geometric structure. Complex wavelets can perform local attention in both the spatial domain and the frequency domain. Therefore, we propose the complex wavelet informed transformer operator that uses the real and imaginary wavelets of the dual-tree complex wavelet transform to simulate the interaction in the attention kernel. In order to further reduce the computational burden of operators, we introduce an adaptive local block shared attention mechanism in the channel domain for our wavelet informed operators. Further, we construct the deep multi-head operator network consisting of a hybrid stack of complex wavelet informed transformer operators and self-attention layers. This enables the Transformer to more sparsely capture multi-scale and multi-directional structured features in the process of learning dependencies. Extensive experimental results show that our adaptive complex wavelet informed transformer operator under the Transformer architecture achieves highly competitive accuracy performance on multiple image classification benchmark datasets. And the proposed operators can be flexibly and effectively migrated to vision tasks in dynamic video scenarios. Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Hao Zhu 0009, Xu Liu 0006, Lingling Li 0002, Wenping Ma 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Complex Dual-Tree Pyramid Scattering TransformerabstractAttention-based transformer networks have recently played an increasingly important role in computer vision tasks. However, since pixel-by-pixel attention multiplication does not involve constraint assumptions such as spatial invariance, the computational complexity grows quadratically with the increase of input pixels. Therefore, this article proposes a complex pyramid scattering Transformer in dense scale space, which introduces sparse scattering constraints with a small number of wavelet basis parameters. It enhances the Transformer's flexibility and sparsity in multiscale space and, to a certain extent, slows down the increase in computational complexity caused by multiresolution input. In addition, compared with the general single-tree real wavelet transform, the dual-tree complex scattering method improves the aliasing of the scattering attention layer and helps obtain a more robust feature representation. At the same time, the multihead stepwise pyramid scattering coupling mechanism helps increase the abundance of directional priors. We conduct experiments in image classification and video tracking scenarios and verify the reliability and superiority of our dual-tree complex pyramid scattering Transformer for visual tasks with different scale requirements. The performance is better than that of the baseline Transformer and other advanced wavelet scattering networks at the same parameter scale. The code is available at https://github.com/Dawn5786/CPSTFormer. Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Hao Zhu 0009, Xin Zhang 0167, Xu Liu 0006, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Pseudo Label Learning for Partial Point Cloud RegistrationabstractPartial point cloud registration plays a crucial role in computer vision and has widespread applications in 3D map construction, pose estimation, and high-precision localization. However, the collected point clouds often contain missing data due to hardware limitations and complex environments. Various partial registration algorithms have been proposed, most of which rely on estimating overlap regions. However, a significant proportion of these algorithms rely heavily on ground truth labels. Manual labeling is both time-consuming and labor-intensive, whereas algorithmic automatic labeling lacks sufficient accuracy. To tackle this issue, we present PSEudo Label learning for unsupervised partial point cloud registration (PSEL). This method utilizes complementary tasks to learn reliable pseudo labels for overlap regions and correspondences without depending on ground truth labels. The key idea is to use the complementarity between overlap estimation and registration to generate two types of pseudo labels based on the nearest points in pairs of aligned point clouds. These pseudo labels are then employed to supervise the learning of overlap regions and correspondences, gradually enhancing their accuracy throughout the learning process and ultimately establishing an unsupervised learning framework. PSEL consists of an overlap estimation module and a correspondence filtering module. The pseudo labels generated after registration are used to supervise both modules. Notably, the correspondence filtering module has two pipelines. The similarity and difference of the corresponding point features are used to eliminate false correspondences during the training and inference stages, respectively, with only the latter being optimized with pseudo labels. To validate the effectiveness of our registration method, we conducted experiments using the synthetic dataset ModelNet40, the indoor dataset 3DMatch, and the outdoor dataset KITTI. Wenping Ma 0001, Yue Wu 0004, Yue Zhang 0040, Hao Zhu 0009, Biao Hou, Licheng Jiao |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | MCDet: Multi-Content Collaboration Detector for Multiscale Remote Sensing ObjectabstractIn previous works, powerful CNN backbones are typically used for one- or two-stage detectors to facilitate multi-categories object classification. Unfortunately, continuous convolution and pooling operations tend to weaken the detailed information. We propose an end-to-end Multi-content Collaboration Detector (MCDet) to improve object recognition accuracy. First, we summarize the reasons for the disappearance of detailed features in traditional feature extraction backbone networks, and propose a Shallow Clue Refinement (SCR) module, which helps us to retain more critical local detail information in the downsampling process. Second, to receive more suitable contextual information, we design a Self-dilating Spatial Pooling (SSP) module, it adaptively learns a contextual reception field, thereby alleviating the mismatch between the theoretical receptive field of the network design and the practical requirements. Finally, extensive experiments on the NWPU VHR-10 and DIOR datasets have shown that the proposed MCDet significantly improves detection accuracy. Our code is available at https://github.com/Xidian-AIGroup190726/RS-objectdetection-MCDet. Wenping Ma 0001, Hao Zhu 0009, Yue Wu 0004, Biao Hou, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Intra- and Intersource Interactive Representation Learning Network for Remote Sensing Images ClassificationabstractRecently, remote sensing technology has developed faster and faster, and obtaining high-quality panchromatic (PAN) and multispectral (MS) images has become more accessible. The complementarity between them provides new opportunities in multisource remote sensing image classification. However, solving the problem of the semantic gap between multisource high-level features and, at the same time, utilizing the complementary properties between them to reduce intersource information redundancy is still a challenge. This article constructs an$I^{3}$RL-Net for the multisource remote sensing image classification task. Specifically, we design a cross-source interactive enhanced fusion module (CIEF-Module). For multilevel multisource features, by strengthening the dependencies of intrasource features and conducting intersource enhanced fusion, intrasource correlation features are refined, and the problem of the intersource semantic gap can be effectively alleviated. During the cross-source interaction process, we design a complementary representation supervised learning strategy (CRSL-Strategy). According to the similarities and differences of multisource features, it can adaptively promote complementary feature learning, thus generating a nonredundant multisource representation. The method has been verified to be effective on multiple RS datasets. The code is open source at:https://github.com/Xidian-AIGroup190726/Ping-Pie-I3RL-Net.git. Wenping Ma 0001, Yanshan Guo, Hao Zhu 0009, Xiaoyu Yi 0002, Wenhao Zhao, Yue Wu 0004, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Significant Feature Elimination and Sample Assessment for Remote Sensing Small Objects' DetectionabstractIn recent years, small object detection has remained challenging in remote sensing tasks. Firstly, small objects inherently have fewer pixels, making them susceptible to interference from prominently featured larger objects during feature extraction. Secondly, existing detection methods solely based on the Intersection over Union (IOU) loss are disadvantageous for small object detection and fail to leverage the rich prior information in remote sensing images. Based on these observations, we propose a significant feature elimination and sample assessment network for small object detection called SESA-Net, based on the Facet derivative model. SESA-Net introduces prior information to the network through the directional derivatives characteristic of remote sensing images. The overall network comprises the ADM module and SIA strategy. The ADM module eliminates significant responses from shallow large objects, directing the network’s focus towards the features of shallow small objects. The Sample Importance Assessment (SIA) strategy addresses the limitations of the IOU loss function by using high-quality positive samples generated by ADM to provide an evaluation strategy for different positive samples of small objects. This enables the network to focus more on high-quality positive samples, thereby improving the accuracy of small object detection. The effectiveness of the proposed algorithm has been validated on multiple datasets. Our code is available at https://github.com/Xidian-AIGroup190726/RS-objectdetection-SESANet. Wenping Ma 0001, Xiaoteng Wang, Hao Zhu 0009, Xiaoting Yang, Xiaoyu Yi 0002, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Adaptive Feature Separation Network for Remote Sensing Object DetectionabstractWith the development of remote sensing technology, remote sensing object detection has been widely applied in various fields, but it still faces some thorny challenges, such as the following: 1) the complexity of object scale changes in remote sensing images makes it difficult to improve the performance of small object detection and 2) remote sensing images have complex backgrounds and densely arranged small and weak objects, which pose a serious problem of feature interference. To alleviate these challenges, we propose an end-to-end adaptive feature separation network called AFSNet, which includes a scale-aware module (SAM) and a class-aware module (CAM). The SAM mainly enables feature maps of different resolutions to detect objects of different scales. Shallow feature maps mainly suppress the features of large objects they contain to focus on small object detection, while deep feature maps increase the detailed features of large objects they contain to focus on large object detection. The CAM is mainly used to distinguish the features in the feature map by category, separating the features of different categories into different channels, thus mitigating the problem of inter class feature interference, and blocking background interference. The effectiveness of this article has been proven on the NWPU VHR-10, IPIU-M, DIOR, and DOTA2.0 datasets. It can be widely applied in civilian, military, and other fields. Through experimental verification, our AFSNet achieved 97.70% mAP on the NWPU VHR-10 dataset, 78.9% mAP on the DIOR dataset, and 58.22% mAP on the DOTA2.0 dataset. Our code is available at:https://github.com/Xidian-AIGroup190726/AFSNet. Wenping Ma 0001, Yiting Wu, Hao Zhu 0009, Wenhao Zhao, Yue Wu 0004, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Self-Supervised Learning Guided by SAR Image Factors for Terrain ClassificationabstractEffective feature representation is the key to SAR image terrain classification. Limited by the abstract appearance and the scarcity of high-quality labeled data in this field, the features learned by current methods, especially deep learning models, do not have enough directivity and applicability, which hampers the performance. This paper proposes Multi-image Factor Self-Supervised Learning(MFSSL) to achieve directional feature learning and obtain generalized features with few patch-level labeled data. The framework consists of an upstream multi-factor image style transfer task and a downstream terrain classification task. In the upstream task, the goal of feature learning is first set up by multiple SAR image factors, including the observation region, the terrain category, and the imaging parameters. And then, different styles of SAR terrain images are generated and reconstructed under this goal. Through this bidirectional generative learning, the low-level external appearance of the terrain is removed, while the essential and discriminative feature representation is retained and shared across different factors. Finally, the downstream model inherits the general feature from the upstream model and implements the terrain classification task using a small amount of labeled data. Experiments conducted on three broad SAR scenes with different image factors demonstrate that the proposed framework can improve pixel-level terrain classification only with a few patch-level labeled data. Zhongle Ren, Zhe Du, Biao Hou, Weibin Li 0002, Hao Zhu 0009, Bo Ren 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Few-Shot MS and PAN Joint Classification With Improved Cross-Source Contrastive LearningabstractThe joint classification of multispectral (MS) and panchromatic (PAN) images aims to provide a more detailed and accurate interpretation of land features. Although deep-learning-based methods have achieved remarkable success in this task, the generalization performance of networks is compromised when labeled samples are insufficient. In this study, we explore the possibility of leveraging unlabeled remote sensing images (RSIs) through contrastive learning and demonstrate the challenges associated with directly applying contrastive learning to RSIs. To end this, we propose a cross-source contrastive learning method for few-shot MS and PAN joint classification (CrossCLMP), which aims to learn sufficient transferable representations in a self-supervised contrastive manner so as to provide a robust pretrained model for fine-tuning the downstream joint classification task. Specifically, we design: 1) intersource and intrasource alignment loss (ER-Align) to achieve self-supervised feature extraction and alignment; 2) the source-unique feature adaptive separation (SUAS) strategy to model source-unique information explicitly; and 3) the auxiliary contrastive learning (ACL) strategy to mitigate the adverse impact of numerous false-negative samples in the pretraining stage. The experimental results and the theoretical analyses on multiple popular datasets comprehensively demonstrate the effectiveness and robustness of the proposed method under few-shot. Our code is available at:https://github.com/Xidian-AIGroup190726/CrossCLMP. Hao Zhu 0009, Pute Guo, Biao Hou, Changzhe Jiao, Bo Ren 0001, Licheng Jiao, Shuang Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | High-Low-Frequency Progressive-Guided Diffusion Model for PAN and MS ClassificationabstractWith the rapid development of remote sensing technology, satellites can easily acquire multispectral (MS) and panchromatic (PAN) images. It is challenging to utilize their complementarity to effectively combine each other’s advantages and mitigate the differences between different modes. In this article, we propose a high-low-frequency progressive-guided diffusion model. It is used to generate an image with the advantages of both MS and PAN, which can be complementary to MS and PAN and, thus, can better reduce the modal differences between them. Therefore, we use the fusion image as an auxiliary mode and an intermediate bridge, which can better connect the characteristics between various sources. First, we design guidance information that contains the advantages of MS and PAN, and some operations can make this information better guide the generation stage. In addition, we design a high-low-frequency progressive guidance strategy; by using this strategy, we can first ensure the overall structure and layout of the image in the generation stage and then refine the local details and features of the image. This dramatically improves the quality of the generated image. Finally, we use mathematical knowledge to explain the rationality of the strategy. We validate our method on multiple datasets and achieve the best performance. Our code ishttps://github.com/Xidian-AIGroup190726/HLF-GDiffusion. Hao Zhu 0009, Fengtao Yan, Pute Guo, Biao Hou, Shuang Wang 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | ConvGRU-Based Multiscale Frequency Fusion Network for PAN-MS Joint ClassificationabstractAs a hot research topic in remote sensing, effectively integrating the advantageous features of multispectral and panchromatic images is the main challenge for fusing these two remote sensing images. This article proposes a multiscale frequency fusion network based on ConvGRU. To address the underutilization of texture features, we extract multiscale bandpass and low-pass sub-bands representing texture and content features through Contourlet decomposition. Multiscale bandpass sub-bands contain more comprehensive and concentrated texture details. Then, by proposing a multiscale frequency feature extractor based on ConvGRU, we effectively integrate and enhance sub-bands of different scales and frequencies, fully utilizing the characteristics of multispectral and panchromatic images and scale transmission. With these enhanced sub-band features, we obtain more comprehensive scale-enhanced texture features. Simultaneously, content features are also preserved as dual-source image features. Moreover, to reduce redundancy between fused features and make more efficient use of the obtained enhanced features, we designed an Inver-band integrator (IBI) module. It can fuse enhanced features at different scales, improve the complementarity between features, and thus achieve effective fusion. Experimental results demonstrate the effectiveness and robustness of our model on multiple datasets. Our codes are available athttps://github.com/Xidian-AIGroup190726/GMFnet. Hao Zhu 0009, Xiaoyu Yi 0002, Biao Hou, Changzhe Jiao, Wenping Ma 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | A Semantically Nonredundant Continuous-Scale Feature Network for Panchromatic and Multispectral ClassificationabstractIn recent years, panchromatic (PAN) images and multispectral (MS) images, as a type of multimodal remote sensing data, are attracting increasingly more attention to their classification problems. However, effectively representing size variations of targets in remote sensing images and reducing redundant representations of different modalities’ deep semantic features to enhance classification accuracy remains a challenge. In this article, we propose a semantically nonredundant continuous-scale feature network (SNCF-Net) for PAN and MS classification, consisting of two modules: the texture-enhanced continuous scale input generation module and the cross-modal feature Kernel interaction (CMKI) module. By simulating the human eye’s adjustment of distance to observe objects of different sizes, we employ 3-D convolution to extract continuous-scale images generated by the texture-enhanced continuous-scale input generation (TCIG) module, enabling optimal feature representation of objects in remote sensing images. Additionally, the texture enhancement (TE) strategy in the TCIG module alleviates texture diffusion in scale space, enhancing the network’s ability to represent texture features. Subsequently, the CMKI module utilizes the response differences between different features to generate convolution kernels from deep feature maps, enabling feature interaction between the PAN modal and MS modal. This reduces redundant representations of essential image content information in deep features of two modalities, facilitating a better mapping between dual-modal features and categories. Our results achieve state-of-the-art performance on multiple datasets. The code is available athttps://github.com/Xidian-AIGroup190726/SNCFNet. Hao Zhu 0009, Wenhao Zhao, Biao Hou, Changzhe Jiao, Zhongle Ren, Wenping Ma 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | A Complex-Former Tracker With Dynamic Polar Spatio-Temporal EncodingabstractRecently, the excellent performance of transformer has attracted the attention of the visual community. Visual transformer models usually reshape images into sequence format and encode them sequentially. However, it is difficult to explicitly represent the relative relationship in distance and direction of visual data with typical 2-D spatial structures. Also, the temporal motion properties of consecutive frames are hardly exploited when it comes to dynamic video tasks like tracking. Therefore, we propose a novel dynamic polar spatio-temporal encoding for video scenes. We use spiral functions in polar space to fully exploit the spatial dependences of distance and direction in real scenes. We then design a dynamic relative encoding mode for continuous frames to capture the continuous spatio-temporal motion characteristics among video frames. Finally, we construct a complex-former framework with the proposed encoding applied to video-tracking tasks, where the complex fusion mode (CFM) realizes the effective fusion of scenes and positions for consecutive frames. The theoretical analysis demonstrates the feasibility and effectiveness of our proposed method. The experimental results on multiple datasets validate that our method can improve tracker performance in various video scenarios. Licheng Jiao, Hao Zhu 0009, Zhongjian Huang, Fang Liu 0001, Lingling Li 0002, Puhua Chen, Shuyuan Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | An Adaptive Migration Collaborative Network for Multimodal Image ClassificationabstractThe multispectral (MS) and the panchromatic (PAN) images belong to different modalities with specific advantageous properties. Therefore, there is a large representation gap between them. Moreover, the features extracted independently by the two branches belong to different feature spaces, which is not conducive to the subsequent collaborative classification. At the same time, different layers also have different representation capabilities for objects with large size differences. In order to dynamically and adaptively transfer the dominant attributes, reduce the gap between them, find the best shared layer representation, and fuse the features of different representation capabilities, this article proposes an adaptive migration collaborative network (AMC-Net) for multimodal remote-sensing (RS) images classification. First, for the input of the network, we combine principal component analysis (PCA) and nonsubsampled contourlet transformation (NSCT) to migrate the advantageous attributes of the PAN and the MS images to each other. This not only improves the quality of images themselves, but also increases the similarity between the two images, thereby reducing the representational gap between them and the pressure on the subsequent classification network. Second, for the interaction on the feature migrate branch, we design a feature progressive migration fusion unit (FPMF-Unit) based on the adaptive cross-stitch unit of correlation coefficient analysis (CCA), which can make the network automatically learn the features that need to be shared and migrated, aiming to find the best shared-layer representation for multifeature learning. And we design an adaptive layer fusion mechanism module (ALFM-Module), which can adaptively fuse features of different layers, aiming to clearly model the dependencies among multiple layers for different sized objects. Finally, for the output of the network, we add the calculation of the correlation coefficient to the loss function, which can make the network converge to the global optimum as much as possible. The experimental results indicate that AMC-Net can achieve competitive performance. And the code for the network framework is available at: https://github.com/ru-willow/A-AFM-ResNet. Wenping Ma 0001, Mengru Ma, Licheng Jiao, Fang Liu 0001, Hao Zhu 0009, Xu Liu 0006, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Explore the Influence of Shallow Information on Point Cloud RegistrationabstractFeature extraction is a key step for deep-learning-based point cloud registration. In the correspondence-free point cloud registration task, the previous work commonly aggregates deep information for global feature extraction and numerous shallow information which is positive to point cloud registration will be ignored with the deepening of the neural network. Shallow information tends to represent the structural information of the point cloud, while deep information tends to represent the semantic information of the point cloud. In addition, fusing information of different dimensions is conducive to making full use of shallow information. Inspired by this, we verify shallow information in the middle layers can bring a positive impact on the point cloud registration task. We design various architectures to combine shallow information and deep information to extract global features for point cloud registration. Experimental results on the ModelNet40 dataset illustrate that feature extractors that incorporate shallow information will bring positive performance. Wenping Ma 0001, Mingyu Yue, Yue Wu 0004, Yongzhe Yuan, Hao Zhu 0009, Biao Hou, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | A Collaborative Learning Tracking Network for Remote Sensing VideosabstractWith the increasing accessibility of remote sensing videos, remote sensing tracking is gradually becoming a hot issue. However, accurately detecting and tracking in complex remote sensing scenes is still a challenge. In this article, we propose a collaborative learning tracking network for remote sensing videos, including a consistent receptive field parallel fusion module (CRFPF), dual-branch spatial-channel co-attention (DSCA) module, and geometric constraint retrack strategy (GCRT). Considering the small-size objects of remote sensing scenes are difficult for general forward networks to extract effective features, we propose a CRFPF-module to establish parallel branches with consistent receptive fields to separately extract from shallow to deep features and then fuse hierarchical features adaptively. Since the objects and their background are difficult to distinguish, the proposed DSCA-module uses the spatial-channel co-attention mechanism to collaboratively learn the relevant information, which enhances the saliency of the objects and regresses to precise bounding boxes. Considering the interference of similar objects, we designed a GCRT-strategy to judge whether there is a false detection through the estimated motion trajectory and then recover the correct object by weakening the feature response of interference. The experimental results and theoretical analysis on multiple datasets demonstrate our proposed method's feasibility and effectiveness. Code and net are available at https://github.com/Dawn5786/CoCRF-TrackNet. Licheng Jiao, Hao Zhu 0009, Fang Liu 0001, Shuyuan Yang 0001, Xiangrong Zhang, Shuang Wang 0001, Rong Qu |
IEEE Trans. Cybern. | 3 |
| 2023 | SSMU-Net: A Style Separation and Mode Unification Network for Multimodal Remote Sensing Image ClassificationabstractThe rapid progress in remote sensing technology has made it convenient for satellites to capture both multispectral (MS) and panchromatic (PAN) images. MS has more spectral information, and PAN has higher spatial resolution. How to exploit the complementarity between MS and PAN images, and effectively combine their respective advantageous features while alleviating mode differences, has become a crucial research task. This paper designs a Style Separation and Mode Unification network (SSMU-Net) for MS and PAN image classification from a novel and effective perspective. The network can be divided into two stages: style separation and mode unification. In the style separation stage, we use wavelet decomposition and techniques similar to generative adversarial networks to preliminarily separate the information of MS and PAN into different components. These components better preserve complete information from the original data and have their own advantages in style and content. Then we propose a Symmetrical Triplet Traction module to perform style traction on different components, making style features more unique and content features more unified, achieving feature separation and purification. In the mode unification stage, we design an encoder-decoder model to reduce the impact of mode differences. The experimental results from multiple datasets validate the effectiveness of our proposed method. Our overall accuracy improved by approximately 4% on the Shanghai and Beijing datasets, and it has exceeded 99.28% on the Hohhot and Vancouver datasets. Our code is available at: https://github.com/proudpie/SSMU-Net. Hao Zhu 0009, Licheng Jiao, Xiaoyu Yi 0002, Biao Hou, Wenping Ma 0001, Shuang Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | A Dual-Stream Transformer With Diff-Attention for Multispectral and Panchromatic ClassificationabstractTo minimize the feature redundancy of multispectral (MS) and panchromatic (PAN) images and maximize the complementary advantages of PAN and MS, a Dual-Stream Transformer with Diff-attention (DSTD)-Net is proposed for PAN and MS classification in this paper. Firstly, in terms of feature extraction, we use Self-attention and Co-attention (SCA) block to extract both specific advantageous features and common essential features. Based on that, a self-attention module strengthened by diff-attention (SSDA) that pays attention to the difference between two specific advantageous features is designed to reduce the essential redundancy in specific features. It can take advantage of the difference between two specific features and reduce the essential redundancy of the specific advantageous features, making them purer and better for classification. Finally, since the specific features and common features of multispectral (MS) and panchromatic (PAN) images make different contributions to classification, a Multi-stage Gated Fusion (MGF) strategy is used. The MGF strategy mainly uses Gated multisource units (GMU) to adapt the weight of different features and fuse them. So, our MGF strategy can strengthen the specific advantageous features beneficial for classification. Above all, the several experiment results verify our proposed networks’ effectiveness and robustness. Our code is available at: https://github.com/blackkiring/DSTD. Lin Xu 0012, Hao Zhu 0009, Licheng Jiao, Wenhao Zhao, Biao Hou, Zhongle Ren, Wenping Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | A Multi-Scale Progressive Collaborative Attention Network for Remote Sensing Fusion ClassificationabstractWith the development of remote sensing technology, panchromatic images (PANs) and multispectral images (MSs) can be easily obtained. PAN has higher spatial resolution, while MS has more spectral information. So how to use the two kinds of images' characteristics to design a network has become a hot research field. In this article, a multi-scale progressive collaborative attention network (MPCA-Net) is proposed for PAN and MS's fusion classification. Compared to the traditional multi-scale convolution operations, we adopt an adaptive dilation rate selection strategy (ADR-SS) to adaptively select the dilation rate to deal with the problem of category area's excessive scale differences. For the traditional pixel-by-pixel sliding window sampling strategy, the patches which are generated by adjacent pixels but belonging to different categories contain a considerable overlap of information. So we change original sampling strategy and propose a center pixel migration (CPM) strategy. It migrates the center pixel to the most similar position of the neighborhood information for classification, which reduces network confusion and increases its stability. Moreover, due to the different spatial and spectral characteristics of PAN and MS, the same network structure for the two branches ignores their respective advantages. For a certain branch, as the network deepens, characteristic has different representations in different stages, so using the same module in multiple feature extraction stages is inappropriate. Thus we carefully design different modules for each feature extraction stage of the two branches. Between the two branches, because the strong mapping methods of directly cascading their features are too rough, we design collaborative progressive fusion modules to eliminate the differences. The experimental results verify that our proposed method can achieve competitive performance. Wenping Ma 0001, Hao Zhu 0009, Licheng Jiao, Jianchao Shen, Biao Hou |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | RDLNet: A Regularized Descriptor Learning NetworkabstractLocal image descriptor learning has been instrumental in various computer vision tasks. Recent innovations lie with similarity measurement of descriptor vectors with metric learning for randomly selected Siamese or triplet patches. Local image descriptor learning focuses more on hard samples since easy samples do not contribute much to optimization. However, few studies focus on hard samples of image patches from the perspective of loss functions and design appropriate learning algorithms to obtain a more compact descriptor representation. This article proposes a regularized descriptor learning network (RDLNet) that makes the network focus on the learning of hard samples and compact descriptor with triplet networks. A novel hard sample mining strategy is designed to select the hardest negative samples in mini-batch. Then batch margin loss concerned with hard samples is adopted to optimize the distance of extreme cases. Finally, for a more stable network and preventing network collapsing, orthogonal regularization is designed to constrain convolutional kernels and obtain rich deep features. RDLNet provides a compact discriminative low-dimensional representation and can be embedded in other pipelines easily. This article gives extensive experimental results for large benchmarks in multiple scenarios and generalization in matching applications with significant improvements. Jun Zhang 0045, Licheng Jiao, Wenping Ma 0001, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Hao Zhu 0009 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2022 | Multitask Semantic Boundary Awareness Network for Remote Sensing Image SegmentationabstractIn remote sensing images, boundary information plays a crucial role in land-cover segmentation. However, it is a challenging problem that sufficiently extracts complete and sharp boundaries from complex very-high-resolution (VHR) remote sensing images. To tackle this problem, we propose a semantic boundary awareness network (SBANet). The SBANet captures refined boundary information of land covers in feature extraction and then supervises its learning with a designed boundary loss. The key of SBANet includes boundary attention module (BA-module) and adaptive weights of multitask learning (AWML). The BA-module is proposed to capture land-cover boundary information from hierarchical features aggregation in a bottom-up manner. It emphasizes useful boundary information and relieves noise information in low-level features with the guidance of high-level features. To directly learn the boundary information, AWML adds a boundary loss to the original semantic loss by an adaptive fusion manner. This multitask learning enables the semantic information and the boundary information to work collaboratively and promote each other. Note that the BA-module and AWML are plug-and-play. Experimental results demonstrate the effectiveness of the proposed SBANet on the available ISPRS 2-D semantic labeling Potsdam and Vaihingen data sets. The SBANet also achieves the state-of-the-art performance in terms of overall accuracy (OA) and mean$F_{1}$score (m-$F_{1}$). Aijin Li, Licheng Jiao, Hao Zhu 0009, Lingling Li 0002, Fang Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | A Two-Stage Mutual Fusion Network for Multispectral and Panchromatic Image ClassificationabstractWith the rapid development of remote sensing technology, satellites can easily obtain multispectral (MS) and panchromatic (PAN) images. How to mine the essence and peculiarity of the MS and PAN images and utilize their complementary to improve classification performance is still a challenge. This paper designs a two-stage mutual fusion network (TSMF-Net) for MS and PAN image classification. The network can be divided into two stages: data fusion and feature fusion. In the data fusion stage, we propose an adaptive twin intensity-hue-saturation (ATIHS) strategy. It not only aligns the size and channels of the MS and PAN images by a novel q-Split operation, but also introduces an adaptive soft-average mask to reduce the differences between replacement components, effectively mitigating spectral distortion and paving the way for the next stage. In the feature fusion stage, we propose a feature graft block (FG-Block) in which we introduce triplet loss and design an interlaced channel addition (ICA) module. Under the supervision of triplet loss, the FG-Block separates and hauls each branch’s essential and peculiar features. With the help of the ICA module, it can effectively graft the essential feature between branches and retain the peculiar feature of each branch, improving the utilization and discrimination of features. Finally, composed of the ATIHS, FG-Blocks, and output layers, our TSMF-Net is proven to improve the accuracy of the remote sensing classification task. The experimental results on multiple datasets verify the effectiveness of our proposed algorithms. Our code is available at: https://github.com/liaoyinuo/TSMF-Net. Yinuo Liao, Hao Zhu 0009, Licheng Jiao, Kenan Sun, Xu Tang 0004, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Feature Split-Merge-Enhancement Network for Remote Sensing Object DetectionabstractRecently, multicategory object detection in high-resolution remote sensing images is still a challenge. First, objects with significant scale differences exist in one scene simultaneously, so it is generally difficult for the detectors to balance the detection performance of large and small objects. Second, because of the complex background and the objects’ densely distributed characteristics in the remote sensing images, the extracted features usually have noise and blurred boundaries, which interfere with the detection performance of the object detectors. With this observation, we propose an end-to-end scale-aware network called feature split–merge–enhancement network (SME-Net) for remote sensing object detection, composed of the feature split-and-merge (FSM) module, the offset-error rectification (OER) module, and the object saliency enhancement (OSE) strategy. FSM eliminates salient information of large objects to highlight the features of small objects in the shallow feature maps. It also transmits the effective detailed features of large objects to the deep feature maps, alleviating feature confusion between multiscale objects. OER corrects the inconsistency of the features spatial layout among the multilayer feature maps by the proposed offset loss, so as to achieve supervised elimination and transmission in FSM. OSE enhances the features of interests and suppresses the background information by the proposed membership function, thus preventing false detection and missed detection caused by noise and blurred boundaries. The effectiveness of the proposed algorithm has been verified on multiple datasets. Our code is available at:https://github.com/Momuli/SMENet.git Wenping Ma 0001, Hao Zhu 0009, Licheng Jiao, Xu Tang 0004, Yuwei Guo 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | A Collaborative Correlation-Matching Network for Multimodality Remote Sensing Image ClassificationabstractRecently, with the increasing availability of the high-quality panchromatic (PAN) and multispectral (MS) remote sensing (RS) images, the inherent complementarity between PAN and MS images provides a wide development prospect for the multimodality RS image classification task. However, how to cleverly relieve the modal differences and effectively integrate the single-modality PAN and MS features is still a challenge. In this article, we design a collaborative correlation-matching network (CCM-Net) for multimodality RS image classification. Concretely, we first propose a bidirectional dominant feature supervision (Bi-DFS) learning, it utilizes single-modality dominant features as supplementary supervision information to establish the joint optimization loss function, thereby adaptively narrowing the differences between modalities before the feature extraction. In the feature extraction stage, the interactive correlation feature matching (ICFM) learning, composing the spatial feature matching (Spa-FM) and spectral feature matching (Spe-FM) strategies, is proposed to establish interactive matching and enhancement between multimodality strong correlation features from the perspective of spatial and spectral, respectively, thereby effectively alleviating the semantic deviation of multimodality features. Finally, we aggregate finer multilevel multimodality features to obtain top-level features with high discrimination. The effectiveness of the proposed algorithm has been verified on multiple datasets. Our code is available at:https://github.com/Momuli/CCM-Net.git. Wenping Ma 0001, Hao Zhu 0009, Kenan Sun, Zhongle Ren, Xu Tang 0004, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | A Novel Adaptive Hybrid Fusion Network for Multiresolution Remote Sensing Images ClassificationabstractWith the rapid development of earth observation technology, panchromatic (PAN) and multispectral (MS) images have also become easier to obtain. The multiresolution classification of PAN and MS images as a basic MS image analysis task has become a research hotspot. The main challenge in this field is how to process data and extract features to improve classification accuracy effectively. In this article, we design a novel adaptive hybrid fusion network (AHF-Net) for multiresolution remote sensing image classification. It includes two parts: data fusion and feature fusion. In the data fusion part, we propose an adaptive weighted intensity-hue-saturation (AWIHS) strategy, which can reduce the difference between MS and PAN images by adaptively adding each other’s unique information from the perspective of information sharing. In the feature fusion part, starting from the second-order correlation of features, we propose a correlation-based attention feature fusion (CAFF) module. It can improve the discrimination of fusion features by adaptively determining the fusion coefficient according to the importance of the input feature channel. Based on AWIHS and CAFF, inspired by the idea of feature pyramid, we combine the multilevel feature fusion and the dual-branch residual network as the backbone network of AHF-Net. By combining AWIHS and CAFF modules with the backbone network, our AHF-Net can effectively improve the classification accuracy of multiresolution remote sensing images. The effectiveness of the proposed algorithm has been verified on multiple data sets. Our code and model are available athttps://github.com/1826133674/AHF-Net. Wenping Ma 0001, Jianchao Shen, Hao Zhu 0009, Jun Zhang 0045, Jiliang Zhao, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Adaptive Dual-Path Collaborative Learning for PAN and MS ClassificationabstractDue to the limitation of sensor technology, researchers tend to obtain high-quality image information from panchromatic (PAN) images and multispectral (MS) images with different resolutions. Therefore, the classification of remote sensing images of PAN and MS have become a research hotspot. In this paper, we propose an adaptive dual-path collaborative learning method for PAN and MS classification. In the stage of sample generation and training, we propose an adaptive neighborhood sample grading (ANSG) strategy in the establishing sample stage so that each pixel to be classified can obtain neighborhood information suitable for itself. Further, to simulate biological cognitive mechanisms, we divide the samples into different levels, and design the self-paced progressive loss (SPL), thus allowing the network to do preference training in different stages. The network’s training can quickly reach the optimal of the current stage and the overall convergence is more thorough. In the network structure, we propose a dual-path module (DPM) to effectively alleviate the gradient degradation in theresidual path, while ensuring maximum gradient loss information flow between every two layers in thedensely connected path. This module can extract more robust features to cope with the complex characteristics of remote-sensing images. Moreover, using the characteristics of the dual path to better fuse the features by the gradual collaborative fusion (GCF) way. The experimental results and theoretical analysis have demonstrated the proposed approach’s effectiveness, feasibility, and robustness. Our model are available at https://github.com/AIpy-nan/DBFI-Net. Hao Zhu 0009, Kenan Sun, Licheng Jiao, Fang Liu 0001, Biao Hou, Shuang Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Hyperspectral image classification based on spatial and spectral kernels generation network
Wenping Ma 0001, Hao Zhu 0009, Licheng Jiao, Biao Hou |
Inf. Sci. | 3 |
| 2021 | A two-stage hybrid ant colony optimization for high-dimensional feature selection
Wenping Ma 0001, Hao Zhu 0009, Licheng Jiao |
Pattern Recognit. | 3 |
| 2020 | A Novel Object Re-Track Framework for 3D Point Cloudsabstract3D point cloud data is an important data source for autonomous vehicles to perceive the surroundings. Achieving accurate object tracking of 3D point clouds has become a challenging task. In this paper, we propose a 3D object two-stage re-track framework directly utilizing point clouds as the input, without using the ground truth as the reference box. The framework consists of a coarse stage and a fine stage. By tracking back the previous T frames and expanding the search space for each frame, we add the fine stage to re-track the lost objects of the coarse stage. Moreover, we design a dense AutoEncoder to enhance the discrimination in the latent space and improve shape completion performance, thus improving tracking performance. A Sample Update Strategy is also proposed to aggregate similar model shape samples in different frames, which improves the quality of the model shape. In terms of motion models for the proposed re-track framework, we further compare Kalman Filter with PointLSTM and do an extensive analysis. Finally, we test the re-track framework on the KITTI tracking dataset and outperform the public benchmark by 17.1%/15.5% in Success and Precision, respectively. Our code and model are available at https://github.com/FengZicai/Re-Track. Tuo Feng 0001, Licheng Jiao, Hao Zhu 0009 |
ACM Multimedia | 3 |
| 2019 | Adaptive Multiscale Deep Fusion Residual Network for Remote Sensing Image ClassificationabstractWith the development of remote sensing imaging technology, remote sensing images with high-resolution and complex structure can be acquired easily. The classification of remote sensing images is always a hot and challenging problem. In order to improve the performance of remote sensing image classification, we propose an adaptive multiscale deep fusion residual network (AMDF-ResNet). The AMDF-ResNet consists of a backbone network and a fusion network. The backbone network including several residual blocks generates multiscale hierarchy features, which contain semantic information from low to high levels. In the fusion network, the adaptive feature fusion module proposed can emphasize useful information and suppress useless information by learning the weights, which represent the importance of the features. The AMDF-ResNet can make full use of the multiscale hierarchy features and the extracted feature is discriminative. In addition, we propose a samples selection method named important samples selection strategy (ISSS). Based on superpixels segmentation result, gradient information and spatial distribution are used as two references to determine the selection numbers and select samples. Compared with the random selection strategy, training samples selected by ISSS are more representative and diverse. The experimental results on four data sets demonstrate that the AMDF-ResNet and ISSS are effective. Lingling Li 0002, Hao Zhu 0009, Xu Liu 0006, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | A Novel Two-Step Registration Method for Remote Sensing Images Based on Deep and Local FeaturesabstractAutomatic remote sensing image registration has achieved great accomplishment. However, it is still a vital challenging problem to develop a robust and accurate registration method due to the negative effects of noise and imaging differences between images. For these images, it is difficult to guarantee the accuracy and robustness at the same time for one-step registration methods. To address this issue, we introduce an effective coarse-to-fine strategy and develop a new two-step registration method based on deep and local features in this paper. The first step is to calculate the approximate spatial relationship, which is obtained by a convolutional neural network. This step makes full use of the deep features to match and can generate stable results. For the second step, a matching strategy considering spatial relationship is applied to the local feature-based method. In addition, this step adopts more accurate features in location to adjust the results of the previous step. A variety of homologous and multimodal remote sensing images, including optical, synthetic aperture radar, and general map images, are used to evaluate the proposed method. The comparison experiments demonstrate that our method can apparently increase the correct correspondences, can improve the ratio of correct correspondences, and is highly robust and accurate. Wenping Ma 0001, Jun Zhang 0045, Yue Wu 0004, Licheng Jiao, Hao Zhu 0009, Wei Zhao 0014 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2019 | A Novel Neural Network for Remote Sensing Image MatchingabstractRapid development of remote sensing (RS) imaging technology makes the acquired images have larger size, higher resolution, and more complex structure, which goes beyond the reach of classical hand-crafted feature-based matching. In this paper, we propose a feature learning approach based on two-branch networks to transform the image matching task into a two-class classification problem. To match two key points, two image patches centered at the key points are entered into the proposed network. The network aims to learn discriminative feature representations for patch matching, so that more matching pairs can be obtained on the premise of maintaining higher subpixel matching accuracy. The proposed network adopts a two-stage training mode to deal with the complex characteristics of RS images. An adaptive sample selection strategy is proposed to determine the size of each patch by the scale of its central key point. Thus, each patch can preserve the texture structure around its key point rather than all patches have a predetermined size. In the matching prediction stage, two strategies, namely, superpixel-based sample graded strategy and superpixel-based ordered spatial matching, are designed to improve the matching efficiency and matching accuracy, respectively. The experimental results and theoretical analysis demonstrate the feasibility, robustness, and effectiveness of the proposed method. Hao Zhu 0009, Licheng Jiao, Wenping Ma 0001, Fang Liu 0001, Wei Zhao 0014 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2016 | SAR Image Registration Based on Multifeature Detection and Arborescence Network MatchingabstractIn this letter, a novel synthetic aperture radar (SAR) image registration method, including two operators for feature detection and arborescence network matching (ANM) for feature matching, is proposed. The two operators, namely, SAR scale-invariant feature transform (SIFT) and R-SIFT, can detect corner points and texture points in SAR images, respectively. This process has an advantage of preserving two types of feature information in SAR images simultaneously. The ANM algorithm has a two-stage process for finding matching pairs. The backbone network and the branch network are successively built. This ANM algorithm combines feature constraints with spatial relations among feature points and possesses a larger number of matching pairs and higher subpixel matching precision than the original version. Experimental results on various SAR images show that the proposed method provides superior performance than other approaches investigated. Hao Zhu 0009, Wenping Ma 0001, Biao Hou, Licheng Jiao |
IEEE Geosci. Remote. Sens. Lett. | 1 |