Wenping Ma 0001

dblp:27/2987-1 · DBLP profile ↗
← Back
128ranked-venue papers
25as first author
94since 2021 · last 2026
0000-0001-8872-2195ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 56 · 8 first-author · 39 since 2021Applied, interdisciplinary, general and emerging computing · 43 · 15 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 1 first-author · 36 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Evolving Semantic Propagation for Aerial Semantic 3D Gaussian Splatting
abstract
Semantic understanding of large-scale aerial scenes represents a critical challenge in 3D computer vision, hindered by the prohibitive cost of dense annotation. This paper introduces EvoPropGS, a novel approach for the semantic segmentation of 3D Gaussian Splatting models that requires only minimal supervision. Our core insight is to leverage the inherent structural repetitions within aerial environments to propagate semantic information from a sparse set of annotations across the entire 3D scene. Our approach constructs a prompt library by pairing SAM-generated mask candidates with DINOv2 feature embeddings from annotated views. For unannotated regions, we generate pseudo-labels by matching region proposals with these featured prompts via cosine similarity. We then formulate optimal prompt selection as a discrete optimization problem solved via evolutionary search, guided by our novel fitness function that evaluates both 3D consistency and 2D semantic coherence. Extensive experiments demonstrate that EvoPropGS achieves accurate segmentation with only 2 percent annotated pixels.
Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Licheng Jiao, Puhua Chen, Wenping Ma 0001, Shuyuan Yang 0001
AAAI7
2026 Hybrid Vector-Occupancy Field for Robust Implicit 3D Surface Reconstruction
abstract
We introduce the Hybrid Vector-Occupancy Field (HVOF), a new implicit 3D representation for reconstructing both open and closed surfaces from sparse point clouds. Existing approaches, such as occupancy field and signed distance fields, face severe limitations. They struggle with open surfaces, while unsigned distance field and neural vector field exhibit directional instability in complex topologies and ridge regions. HVOF addresses these challenges by incorporating a smoothly decaying occupancy field around the surface, while capturing precise local geometry using truncated displacement vectors, naturally mitigating direction-field ambiguities near ridge regions. This unified design forms a robust hybrid representation that leverages both occupancy and vector fields. To fulfill it, we design a Hybrid Field variational autoencoder including a hierarchical cross-attention encoder and dual-branch decoder that jointly learn occupancy and vector fields through continuous weighting. Extensive experiments demonstrate that HVOF consistently outperforms state-of-the-art methods across ShapeNet, ABC, and MGN datasets, accurately reconstructing both open and closed surfaces while preserving fine geometric details in complex regions.
Yue Wu 0004, Tengfei Xiao, Can Qin, Yongzhe Yuan, Hao Li 0009, Kaiyuan Feng, Wenping Ma 0001
AAAI8
2026 A trust-aware singular fusion network for multimodal image classification
Wenping Ma 0001, Mengru Ma, Hekai Zhang, Hao Zhu 0009, Licheng Jiao
Neurocomputing1
2026 Like Human Rethinking: Contour Transformer AutoRegression for Referring Remote Sensing Interpretation
abstract
Referring remote sensing interpretation holds significant application value in various scenarios such as ecological protection, resource exploration, and emergency management. However, referring remote sensing expression comprehension and segmentation (RRSECS) faces critical challenges, including micro-target localization drift problem caused by insufficient extraction of boundary features in existing paradigms. Moreover, when transferred to remote sensing domains, polygon-based methods encounter issues such as contour-boundary misalignment and multi-task co-optimization conflicts problems. In this paper, we propose SeeFormer, a novel contour autoregressive paradigm specifically designed for RRSECS, which accurately locates and segments micro, irregular targets in remote sensing imagery. We first introduce a brain-inspired feature refocus learning (BIFRL) module that progressively attends to effective object features via a coarse-to-fine scheme, significantly boosting small-object localization and segmentation. Next, we present a language-contour enhancer (LCE) that injects shape-aware contour priors, and a corner-based contour sampler (CBCS) to improve mask-polygon reconstruction fidelity. Finally, we develop an autoregressive dual-decoder paradigm (ARDDP) that preserves sequence consistency while alleviating multi-task optimization conflicts. Extensive experiments on RefDIOR, RRSISD, and OPTRSVG datasets under varying scenarios, scales, and task paradigms demonstrate transformative performance gains: compared to the baseline PolyFormer, our proposed SeeFormer improves oIoU and mIoU by 27.58% and 39.37% for referring image segmentation and by 18.94% and 28.90% for visual grounding on the RefDIOR dataset.
Jinming Chai, Licheng Jiao, Xiaoqiang Lu, Lingling Li 0002, Fang Liu 0001, Xu Liu 0006, Wenping Ma 0001, Weibin Li 0002
IEEE Trans. Pattern Anal. Mach. Intell.8
2026 Physics-Informed Matrix Factorization Operator
abstract
Matrix factorization is a fundamental characterization model in machine learning and is usually solved using mathematical decomposition reconstruction loss. However, matrix factorization is a data-driven model whose results depend on data quality, making it susceptible to noise. Inspired by physics, the law of conservation of energy is used to introduce physical laws into matrix factorization, which is called Physics-informed Matrix Factorization operator (PiMF). The PiMF operator uses the heat conduction equation to construct the energy objective function for matrix factorization, thereby retaining the mathematical model's decomposition meaning and satisfying the interpretability of physics. The PiMF follows the physical laws, thereby suppressing irregular or sudden noise signals that violate these physical principles. The solutions of the PiMF operator include more comprehensive knowledge of mathematics and physics, which improves the ability to generalize complex data, especially for noisy data. We demonstrate the consistency of the energy objective function and the mathematical model, which verifies the feasibility of matrix factorization using physical energy laws. In addition, the physical interpretability of the PiMF operator is proved from the perspective of energy decline. This study proposes two practical algorithms for PiMF in classification and clustering tasks, enhancing the practicability of matrix factorization by incorporating task-specific prior information constraints. The experimental results of PiMF for classification and clustering demonstrate the advantages of the proposed operator. The importance of physics-informed matrix factorization is verified, especially for noisy data.
Chenxi Tian, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 VCGPrompt: Visual Concept Graph-Aware Prompt Learning for Vision-Language Models
Mengjia Wang, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001
Pattern Recognit.8
2026 Recurrent progressive fusion-based learning for multi-source remote sensing image classification
Hao Zhu 0009, Biao Hou, Wenhao Zhao, Xiaoyu Yi 0002, Wenping Ma 0001, Licheng Jiao
Pattern Recognit.7
2026 ERFC: Energy-Aware Reinforcement Feedback Calibration for Zero-Shot Captioning
abstract
Zero-shot captioning aims to generate descriptive captions for unseen image and video data by leveraging the potential of visual language models (VLMs) and language models (LMs) without requiring task-specific training. It has emerged as a critical task, but its performance is often hindered by the inherent gap between the training distribution and unseen test data. The fundamental challenge lies in the model’s strong dependence on the marginal distribution of the training data, which leads to biased predictions when handling test samples. To address this issue, we propose an Energy-aware Reinforcement Feedback Calibration (ERFC) framework to calibrate the distribution and predictions of caption models from a novel energy perspective. The calibration process of ERFC is divided into two key components: 1) We first construct an Energy Stabilizer (ES) based on the caption model, where energy is considered a measure of the affinity between the input sample and the model’s learned distribution. ES iteratively adjusts the embedding features of the input sample using Langevin Dynamics, reducing its energy to implicitly align the model’s distribution with the unseen target domain. 2) We deploy a Reinforcement Calibrator (RC) to refine and calibrate the generated captions through a reward-feedback mechanism. RC leverages the expert CLIP model as a reward signal to assess the quality of the generated captions and employs the policy gradient algorithm to reward or penalize the model, thereby improving its performance. By iteratively combining energy-based optimization and reward-driven calibration, ERFC achieves superior zero-shot generalization capabilities, as demonstrated on image benchmarks such as MSCOCO, Flickr30K, and NoCaps, as well as video benchmarks such as MSR-VTT and MSVD.
Qianyue Bao, Fang Liu 0001, Licheng Jiao, Yang Liu 0349, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Wenping Ma 0001
IEEE Trans. Circuits Syst. Video Technol.9
2026 A Progressive Semi-Distillation Model for Dual-Source Remote Sensing Image Classification
abstract
Panchromatic images (PANs) and multispectral (MS) images (MSs) are widely used for dual-source remote sensing image classification, gradually becoming a research hotspot. However, making the most of dual-source image information with insufficiently labeled samples is a significant challenge. This article proposes a progressive semi-distillation model (PSDM) to classify dual-source remote sensing images with insufficient samples. We design a framework of rookie teacher network (RTN)-teaching assistant system (TAS)-student grouping network (SGN) in the case of a traditional teacher network (TN) (i.e., rookie TN (RTN)) that does not provide excellent guidance to student network (SN) due to insufficient samples. The PSDM expands the samples and compresses the space through the RTN-SGN structure to cope with the dilemma of insufficient samples. To make RTN better guide the SGN, we design TAS, which can gradually guide SGN to learn the samples from easy to difficult. It can also further assist SGN training to improve the classification performance of SGN with insufficient samples. We design SGN and add cooperation and correction mechanism to better learn dual- source information. These strategies can eliminate SGN's over-dependence on the RTN, help SGN outperform the RTN, and achieve the effect of semi-distillation. Experimental results and theoretical analysis have sufficiently pointed out the proposed method's accuracy, efficiency, and robustness under insufficient sample situations. Our model is available at https://github.com/MarjordCpz/PSDM.
Hao Zhu 0009, Peizhou Cao, Licheng Jiao, Biao Hou, Xiaoyu Yi 0002, Wenhao Zhao, Wenping Ma 0001
IEEE Trans. Cybern.8
2026 Adaptive Multi-Modal Visual Tracking With Dynamic Semantic Prompts
abstract
RGB-based object tracking is a fundamental task in computer vision, aiming to identify, locate, and continuously track objects of interest across sequential video frames. Despite the significant advancements in the performance of traditional RGB trackers, they still face challenges in maintaining accuracy and robustness in the presence of complex backgrounds, occlusions, and rapid movements. To tackle these challenges, combining visual auxiliary modalities has gained significant attention. Beyond this, integrating natural language information offers additional advantages by providing high-level semantic context, enhancing robustness, and clarifying target priorities, further elevating tracker performance. This work proposes theAdaptiveMulti-modalVisual Tracking with Dynamic Semantic Prompts (AMVTrack) tracker, which efficiently incorporates image descriptions and avoids text dependency during tracking to improve flexibility and adaptability. AMVTrack significantly reduces computational resource consumption by freezing the parameters of the image encoder, text encoder, and Box Head and only optimizing a few learnable prompt parameters. Additionally, we introduce the Adaptive Dynamic Semantic Prompt Generator (ADSPG), which dynamically generates semantic prompts based on visual features, and theVisual-LanguageFusionAdaptation (V-L FA) method, which integrates multi-modal features to ensure consistency and complementarity of information. Additionally, we partition the Image Encoder to conduct an in-depth investigation into the relationship between the importance of features across different depth and width regions. Experimental results demonstrate that AMVTrack achieves significant performance improvements on multiple benchmark datasets, proving its effectiveness and robustness in complex scenarios.
Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001
IEEE Trans. Multim.9
2026 Regularized-Aware Discriminative Transformer Tracker for Satellite Videos
Xin Zhang 0167, Licheng Jiao, Lingling Li 0002, Zhongjian Huang, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001
IEEE Trans. Multim.8
2026 Multiscale Spatial-Frequency Learning for Degradation Decoupling in RS Image Restoration
abstract
Remote sensing (RS) images are prone to various degradations, which poses challenges to downstream tasks. Although existing single-task remote sensing image restoration methods are effective, they lack generalizability across tasks. All-in-one methods can handle multiple degradation tasks, but they usually focus on spatial information, ignoring the physical properties of the degradation information. To address the above limitations, we propose a Multiscale Spatial-Frequency Degradation Decoupling framework for All-in-One remote sensing image restoration (SFD$^{2}$IR), which decouples degradation features across different tasks to guide the model in performing task-specific image restoration. Specifically, a task-specific instruction generator (TIG) is proposed first to transform degradation features into task-specific prompts. Then, a multi-scale multi-frequency enhancement (MME) module is designed to decouple degradation effects from both spatial and frequency perspectives, thus enhancing the model's adaptability to various degradation types. Finally, a prompt feature refinement (PFR) module is developed to further refine the model's response to degraded tasks. Extensive experiments demonstrate that the proposed method achieves excellent performance on different RSIR tasks, including cloud removal, deblurring, dehazing, and super-resolution. The source code will be publicly available at SFD$^{2}$IR.
Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001
IEEE Trans. Multim.6
2026 Edit-Your-Motion: Space-Time Diffusion Decoupling Learning for Video Motion Editing
abstract
Existing diffusion-based methods have achieved impressive results in human motion editing. However, these methods often exhibit significant ghosting and body distortion in unseen in-the-wild cases. In this paper, we introduce Edit-Your-Motion, a video motion editing method that tackles these challenges through one-shot fine-tuning on unseen cases. Specifically, firstly, we utilized DDIM inversion to initialize the noise, preserving the appearance of the source video and designed a lightweight motion attention adapter module to enhance motion fidelity. DDIM inversion aims to obtain the implicit representations by estimating the prediction noise from the source video, which serves as a starting point for the sampling process, ensuring the appearance consistency between the source and edited videos. The Motion Attention Module (MA) enhances the model's motion editing ability by resolving the conflict between the skeleton features and the appearance features. Secondly, to effectively decouple motion and appearance of source video, we design a spatio-temporal two-stage learning strategy (STL). In the first stage, we focus on learning temporal features of human motion and propose recurrent causal attention (RCA) to ensure consistency between video frames. In the second stage, we shift focus on learning the appearance features of the source video. With Edit-Your-Motion, users can edit the motion of humans in the source video, creating more engaging and diverse content. Extensive qualitative and quantitative experiments, along with user preference studies, show that Edit-Your-Motion outperforms other methods.
Yi Zuo 0003, Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Wenping Ma 0001, Shuyuan Yang 0001, Yuwei Guo 0001
IEEE Trans. Multim.6
2026 DGNMF: Dynamic Diffusion Graph Nonnegative Matrix Factorization
abstract
In feature learning (FL), structural information shows advantages in retaining information and maintaining stability. Graph diffusion, a graph learning method that can focus on neighborhood structure and transmit information, has great research potential. In this study, a novel dynamic diffusion graph nonnegative matrix factorization (DGNMF) method is proposed, which uses a diffusion graph to improve the performance of FL and further enhances the effectiveness and stability of downstream classification tasks. DGNMF aims to mine and retain structural information more deeply in FL to build a more powerful and stable FL method. First, the model embeds graph learning into FL to obtain features containing structural information. Second, dynamic diffusion graph learning is used to mine deeper and more global structural information. Finally, we construct an updateable indicator matrix to enhance the discriminability of features. The classification experimental results of DGNMF on six databases demonstrate its advantages, verify its effectiveness and stability, and prove the importance of diffusion graph in improving FL.
Chenxi Tian, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001
IEEE Trans. Neural Networks Learn. Syst.7
2026 Spatial-Temporal Diffusion Model for Matrix Factorization
abstract
Matrix factorization (MF) is a fundamental problem in machine learning, which is usually used as a feature learning method in various fields. For complex data involving spatiotemporal interactions, MF that only handles 2-D data will disrupt spatial dependence or temporal dynamics, failing to effectively couple spatial information with temporal factors. According to Markov chain principle, the spatial information of the present time is related to the spatial state of the previous time. We propose a spatial-temporal diffusion model for MF (STDMF), which uses graph diffusion to couple spatial-temporal information. Then, MF is used to learn the joint feature of data and spatial-temporal diffusion graph. Specifically, STDMF utilizes the graph diffusion with physical laws to generate spatial-temporal structure information. It obtains the underlying core structure of complex systems from a global perspective, which enhances the generalization ability of MF in noisy time-series data. To learn the lowest rank subspace of MF in time-series data, STDMF uses structural learning to constrain the rank of the learned features. Finally, STDMF is applied to clustering and anomaly detection of dynamic graph. The effectiveness of this method is verified by sufficient experiments, especially for noisy data.
Chenxi Tian, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Licheng Jiao, Shuyuan Yang 0001
IEEE Trans. Neural Networks Learn. Syst.6
2025 MUCD: Unsupervised Point Cloud Change Detection via Masked Consistency
abstract
3D Change Detection (3DCD) has gradually become another research hotspot after image change detection. Recent works focus on using artificial labels for supervised or weakly-supervised training of siamese networks to segment changed points. However, labeling every points of multi-temporal point clouds is very expensive and time-consuming. In addition, these works lack effective self-supervised signals, and existing self-supervised signals often fail to capture sufficiently rich change information. To solve this problem, we assume that the powerful representation of 3D objects should model the consistency information of unchanged regions and distinguish different objects. Based on this assumption, we propose a new unsupervised framework called MUCD to learn change information of multi-temporal point clouds through bidirectional optimization of change segmentor and feature extractor. The training of network is divided into two stages. We first design a foreknowledge point contrastive loss based on the characteristics of the 3DCD task to initialize the feature extractor, and then propose a masked consistency loss to further learn the shared geometric information of unchanged regions in the multi-temporal point clouds, utilizing it as a free and powerful supervised signal to train a change segmentor. In the inference stage, only the segmentor is used to take multi-temporal point clouds as input and produce change segmentation result. Extensive experiments are conducted on SLPCCD and Urb3DCD, two real-world datasets of streets and urban buildings, to verify that our proposed unsupervised method is highly competitive and even outperforms supervised methods in scenes where semantic information changes occur, exhibiting better performance in generalization ability and robustness.
Yue Wu 0004, Yongzhe Yuan, Maoguo Gong, Hao Li 0009, Mingyang Zhang 0002, Wenping Ma 0001, Qiguang Miao
AAAI7
2025 Where Precision Meets Efficiency: Transformation Diffusion Model for Point Cloud Registration
abstract
We propose a transformation diffusion model for point cloud registration to balance precision and efficiency. Our method formulates point cloud registration as a denoising diffusion process from noisy transformation to object transformation, which is represented by quaternion and translation. Specifically, in training stage, object transformation diffuses from ground-truth transformation to random distribution, and the model learns to reverse this noising process. In sampling stage, the model refines randomly generated transformation to the optimal transformation in a progressive way. We derive the variational bound in closed form for training and provide instantiation of the model. Our diffusion model maps transformation into latent space, and splits the transformation into two components (rotation and translation) based on the fact that they belong to different solution spaces. In addition, our work provides the following crucial findings: (i) Point cloud registration, one of the representative discriminative tasks, can be solved by a generative way and mapped into latent space to obtain new unified probabilistic formulation. (ii) Our model, Transformation Diffusion Model (TDM) can be a plug-and-play agent for point cloud registration, making our method applicable to different deep registration networks. Experimental results on synthetic and real-world datasets demonstrate that, in correspondence-free and correspondence-based scenarios, TDM can both achieve exceeding 60% performance improvements and higher efficiency simultaneously.
Yongzhe Yuan, Yue Wu 0004, Xiaolong Fan, Maoguo Gong, Qiguang Miao, Wenping Ma 0001
AAAI6
2025 Partial Point Cloud Registration with Multi-view 2D Image Learning
abstract
Learning representations from numerous 2D image data has shown promising performance, yet very few works apply this representations to point cloud registration. In this paper, we explore how to leverage the 2D information to assist the point cloud registration, and propose IAPReg, an Image-Assisted Partial 3D point cloud Registration framework with the multi-view images generated by the input point cloud. It is expected to enrich 3D information with 2D knowledge, and leverage 2D knowledge to assist with point cloud registration. Specifically, we create multi-view depth maps by projecting the input point cloud from several specific views, and then extract 2D and 3D features using some well-established models. To fuse the information learned from 2D and 3D modalities, inter-modality multi-view learning module is proposed to enhance geometric information and complement semantic information. Weighted SVD is a common method to reduce the impact of inaccurate correspondences on registration. However, determining the correspondence weights is not trivial. Therefore, we design a 2D-weighted SVD method, where the 2D knowledge is employed to provide weight information of correspondences. Extensive experiments perform that our method outperform the state-of-the-art method without additional 2D training data.
Yue Zhang 0040, Yue Wu 0004, Wenping Ma 0001, Maoguo Gong, Hao Li 0009, Biao Hou
AAAI3
2025 Logits DeConfusion with CLIP for Few-Shot Learning
abstract
With its powerful visual-language alignment capability, CLIP performs well in zero-shot and few-shot learning tasks. However, we found in experiments that CLIP’s logits suffer from serious inter-class confusion problems in down-stream tasks, and the ambiguity between categories seriously affects the accuracy. To address this challenge, we propose a novel method called Logits DeConfusion, which effectively learns and eliminates inter-class confusion in logits by combining our Multi-level Adapter Fusion (MAF) module with our Inter-Class Deconfusion (ICD) module. Our MAF extracts features from different levels and fuses them uniformly to enhance feature representation. Our ICD learnably eliminates inter-class confusion in logits with a residual structure. Experimental results show that our method can significantly improve the classification performance and alleviate the inter-class confusion problem. The code is available at https://github.com/LiShuo1001/LDC.
Shuo Li 0010, Fang Liu 0001, Zehua Hao, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Wenping Ma 0001
CVPR8
2025 Disentangled Pose and Appearance Guidance for Multi-Pose Generation
abstract
Human pose generation is a complex task due to the non-rigid and highly variable nature of human body structures and appearances. However, existing methods often overlook the fundamental differences between spatial transformations of poses and texture generation for appearance, which makes them prone to overfitting. To address this issue, we propose a multi-pose generation framework driven by disentangled pose and appearance guidance. Our approach includes a Global-aware Pose Generation module that iteratively generates pose embeddings, enabling effective control over non-rigid body deformations. Additionally, we introduce the Global-aware Transformer Decoder, which leverages similarity queries and attention mechanisms to achieve spatial transformations and enhance pose consistency through a Global-aware block. In the appearance generation phase, we condition a diffusion model on pose embeddings produced in the initial stage and introduce an Appearance Adapter that extracts high-level contextual semantic information from multi-scale features, enabling further refinement of pose appearance textures and providing appearance guidance. Extensive experiments on the UBC Fashion and TikTok datasets demonstrate that our framework achieves state-of-the-art results in both quality and fidelity, establishing it as a powerful approach for complex pose generation tasks.
Tengfei Xiao, Yue Wu 0004, Can Qin, Maoguo Gong, Qiguang Miao, Wenping Ma 0001
CVPR7
2025 Knowledge-Guided Part Segmentation
Xuejian Gou, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Lingling Li 0002, Hao Wang 0211, Xu Liu 0006, Puhua Chen, Wenping Ma 0001
ICCV9
2025 Domain-Aware Category-Level Geometry Learning Segmentation for 3D Point Clouds
Pei He, Lingling Li 0002, Licheng Jiao, Ronghua Shang, Fang Liu 0001, Shuang Wang 0001, Xu Liu 0006, Wenping Ma 0001
ICCV8
2025 Hierarchical Variational Test-Time Prompt Generation for Zero-Shot Generalization
Zhaoyang Wu, Fang Liu 0001, Licheng Jiao, Shuo Li 0010, Lingling Li 0002, LiXu Liu, Puhua Chen, Wenping Ma 0001
ICCV8
2025 PointTruss: K-Truss for Point Cloud Registration
abstract
Point cloud registration is a fundamental task in 3D computer vision. Recent advances have shown that graph-based methods are effective for outlier rejection in this context. However, existing clique-based methods impose overly strict constraints and are NP-hard, making it difficult to achieve both robustness and efficiency. While the k-core reduces computational complexity, which only considers node degree and ignores higher-order topological structures such as triangles, limiting its effectiveness in complex scenarios. To overcome these limitations, we introduce the $k$-truss from graph theory into point cloud registration, leveraging triangle support as a constraint for inlier selection. We further propose a consensus voting-based low-scale sampling strategy to efficiently extract the structural skeleton of the point cloud prior to $k$-truss decomposition. Additionally, we design a spatial distribution score that balances coverage and uniformity of inliers, preventing selections that concentrate on sparse local clusters. Extensive experiments on KITTI, 3DMatch, and 3DLoMatch demonstrate that our method consistently outperforms both traditional and learning-based approaches in various indoor and outdoor scenarios, achieving state-of-the-art results.
Yue Wu 0004, Yongzhe Yuan, Maoguo Gong, Qiguang Miao, Hao Li 0009, Mingyang Zhang 0002, Wenping Ma 0001
NeurIPS8
2025 Unveiling and Mitigating Generalized Biases of DNNs Through the Intrinsic Dimensions of Perceptual Manifolds
abstract
Building fair deep neural networks (DNNs) is a crucial step towards achieving trustworthy artificial intelligence. Delving into deeper factors that affect the fairness of DNNs is paramount and serves as the foundation for mitigating model biases. However, current methods are limited in accurately predicting DNN biases, relying solely on the number of training samples and lacking more precise measurement tools. Here, we establish a geometric perspective for analyzing the fairness of DNNs, comprehensively exploring how DNNs internally shape the intrinsic geometric characteristics of datasets-the intrinsic dimensions (IDs) of perceptual manifolds, and the impact of IDs on the fairness of DNNs. Based on multiple findings, we propose Intrinsic Dimension Regularization (IDR), which enhances the fairness and performance of models by promoting the learning of concise and ID-balanced class perceptual manifolds. In various image recognition benchmark tests, IDR significantly mitigates model bias while improving its performance.
Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Lingling Li 0002, Wenping Ma 0001, Shuyuan Yang 0001, Xu Liu 0006, Puhua Chen
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Predicting and Enhancing the Fairness of DNNs With the Curvature of Perceptual Manifolds
abstract
To address the challenges of long-tailed classification, researchers have proposed several approaches to reduce model bias, most of which assume that classes with few samples are weak classes. However, recent studies have shown that tail classes are not always hard to learn, and model bias has been observed on sample-balanced datasets, suggesting the existence of other factors that affect model bias. In this work, we first establish a geometric perspective for analyzing model fairness and then systematically propose a series of geometric measurements for perceptual manifolds in deep neural networks. Subsequently, we comprehensively explore the effect of the geometric characteristics of perceptual manifolds on classification difficulty and how learning shapes the geometric characteristics of perceptual manifolds. An unanticipated finding is that the correlation between the class accuracy and the separation degree of perceptual manifolds gradually decreases during training, while the negative correlation with the curvature gradually increases, implying that curvature imbalance leads to model bias. We thoroughly validate this finding across multiple networks and datasets, providing a solid experimental foundation for future research. We also investigate the convergence consistency between the loss function and curvature imbalance, demonstrating the lack of curvature constraints in existing optimization objectives. Building upon these observations, we propose curvature regularization to facilitate the model to learn curvature-balanced and flatter perceptual manifolds. Evaluations on multiple long-tailed and non-long-tailed datasets show the excellent performance and exciting generality of our approach, especially in achieving significant performance improvements based on current state-of-the-art techniques. Our work opens up a geometric analysis perspective on model bias and reminds researchers to pay attention to model bias on non-long-tailed and even sample-balanced datasets.
Yanbiao Ma, Licheng Jiao, Fang Liu 0001, Maoji Wen, Lingling Li 0002, Wenping Ma 0001, Shuyuan Yang 0001, Xu Liu 0006, Puhua Chen
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 VLPA-CLIP: Video Language Prompting and Adapting CLIP for efficient video action recognition
Hao Wang 0211, Fang Liu 0001, Licheng Jiao, Jiahao Wang 0002, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001
Pattern Recognit.9
2025 Prompt-Based Concept Learning for Few-Shot Class-Incremental Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) faces a huge stability-plasticity challenge due to continuously learning knowledge from new classes with a small number of training samples without forgetting the knowledge of previously seen old classes. To alleviate this challenge, we propose a novel method called Prompt-based Concept Learning (PCL) for FSCIL, which generalizes conceptual knowledge learned from old classes to new classes by simulating human learning capabilities. In our PCL, in the base session, we simultaneously learn common basic concepts from the training data and the class-concept weight of each class in a prompt learning manner, and in each incremental session, class-concept weights between new classes and previously learned basic concepts are learned to achieve incremental learning. Furthermore, in order to avoid catastrophic forgetting, we propose a distribution estimation module to retain feature distributions of previously seen classes and a data replay module to randomly sample features of previously seen classes in incremental sessions. We verify the effectiveness of our PCL on widely used benchmarks, such as miniImageNet, CIFAR-100, and CUB-200. Experimental results show that our PCL achieves competitive results compared with other state-of-the-art methods, especially we achieve an average accuracy of 94.02% across all sessions on the miniImageNet benchmark.
Shuo Li 0010, Fang Liu 0001, Licheng Jiao, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001
IEEE Trans. Circuits Syst. Video Technol.7
2025 Triple Point Masking
abstract
Existing 3D mask learning methods encounter performance bottlenecks under limited data, and our objective is to overcome this limitation. In this paper, we introduce a triple point masking scheme, named TPM, which serves as a scalable plug-and-play framework for MAE pre-training to achieve multi-mask learning for 3D point clouds. Specifically, we augment the baseline methods with two additional mask choices (i.e., medium mask and low mask) as our core insight is that the recovery process of an object can manifest in diverse ways. Previous high-masking schemes focus on capturing the global representation information but lack fine-grained recovery capabilities, so that the generated pre-training weights tend to play a limited role in the fine-tuning process. With the support of the proposed TPM, current methods can exhibit more flexible and accurate completion capabilities, enabling the potential autoencoder in the pre-training stage to consider multiple representations of a single 3D point cloud object. In addition, during the fine-tuning stage, an SVM-guided weight selection module is proposed to fill the encoder parameters for downstream networks with the optimal weight, maximizing linear accuracy and facilitating the acquisition of intricate representations for new objects. Extensive experimental results and theoretical analysis show that five baselines equipped with the proposed TPM achieve comprehensive performance improvements on various downstream tasks. Our code and models are available athttps://github.com/liujia99/TPM.
Linghe Kong, Yue Wu 0004, Maoguo Gong, Hao Li 0009, Qiguang Miao, Wenping Ma 0001, Can Qin
IEEE Trans. Circuits Syst. Video Technol.7
2025 Equivariance-Based Markov Decision Process for Unsupervised Point Cloud Registration
abstract
Unsupervised point cloud registration is crucial in 3D computer vision. However, most unsupervised methods struggle to construct effective optimization objectives and reliable unsupervised signals to enhance the performance of the model. To address these issues, with the observation of the significant alignment between the registration process and the Markov Decision Process (MDP), we model point cloud registration as MDP, which can provide more reliable unsupervised signals through the reward. We propose a colored noise based cross-entropy method, which introduces colored noise into sampling process, regulating the power spectral density of the action sequence and expanding the search space, improving the registration effect. Particularly, to strengthen constraints on MDP and training in the transformation space, we utilize equivariance theory to construct transformation equivariant constraint as a new optimization objective and derive equivariant constraint solutions for optimization, providing more reliable unsupervised signals. Extensive experiments demonstrate the superior performance of our method on benchmark datasets.
Yue Wu 0004, Jiayi Lei, Yongzhe Yuan, Xiaolong Fan, Maoguo Gong, Wenping Ma 0001, Qiguang Miao, Mingyang Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.6
2025 Contour Knowledge-Aware Perception Learning for Semantic Segmentation
abstract
The diversity of contextual information is of great importance for accurate semantic segmentation. However, most methods focus on single spatial contextual information, which results in an overlap of the semantic content of categories and a loss of contour information of objects. In this article, we propose a novel contour knowledge-aware perception learning network (CKPL-Net) to capture diverse contextual information by space-category aggregation module (SCAM) and contour-aware calibration module (CACM). First, SCAM is introduced to enhance intraclass consistency and interclass differentiation of features. By integrating space-aware and category-aware attention, SCAM reduces the redundancy of features from a categorical perspective while maintaining spatial correlation of pixels, substantially avoiding the overlap of the semantic content in categories. Second, CACM is designed to maintain the integrity of objects by perceiving contour contextual information. It develops a novel contour-aware knowledge and adaptively transforms the grid structure of convolutions for boundary pixels, which effectively calibrates the representation of features near boundaries. Finally, the quantitative and qualitative analyses on the three public datasets: ISPRS Potsdam dataset, ISPRS Vaihingen dataset, and WHDLD dataset, demonstrate that the proposed CKPL-Net achieves superior performance compared with prevalent methods, which indicates diverse contextual information is beneficial for accurate segmentation.
Chao You, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Multitask Multiscale Feature Selection for Point Cloud Registration
abstract
3D point cloud registration is a process of solving the geometric transformation between two point clouds. This process is an important issue in computer vision and pattern recognition. The registration methods based on geometric features are highly sensitive to the scale of feature extraction. Changes in scale can introduce inaccuracies in feature descriptions, thereby compromising the reliability of the registration results. To mitigate the impact of feature scale on the outcomes and the high-dimensional issue arising from features of different scales, we propose a method for multi-scale point cloud feature selection. We solve the high-dimensional problem of feature selection by designing a multi-task framework. By designing a mutual information dimensionality reduction method, we decomposed the high-dimensional feature selection task of different descriptors with multi-scale features into multiple related low-dimensional feature selection tasks. Then, by means of the knowledge transfer among these low-dimensional feature selection tasks, we sought the best feature subset to obtain more robust feature information. We evaluate the effectiveness of our method by conducting extensive experiments on various datasets. The experimental results show that the method outperforms other feature descriptors in terms of descriptive power and robustness and improves the effectiveness of point cloud registration.
Yue Wu 0004, Chuang Luo, Maoguo Gong, Hangqi Ding, Jinlong Sheng, Qiguang Miao, Hao Li 0009, Wenping Ma 0001
IEEE Trans. Evol. Comput.8
2025 Evolutionary Multitasking Descriptor Optimization for Point Cloud Registration
abstract
Point cloud registration (PCR) is an important task for other point cloud tasks. Feature-based methods are widely adopted for their speed and efficiency in PCR. The descriptive capability of features extracted by a single geometric descriptor is limited. Descriptive capabilities can be improved by concatenating features extracted from multiple descriptors. However, due to the existence of redundant and irrelevant features, the correct corresponding points are difficult to match, which further affects the registration effect. We propose an evolutionary multitasking point cloud descriptor optimization method. Integrate existing descriptors to optimize descriptors with stronger description ability. Labeling features to calculate the feature importance for the registration and generating multitasks. In optimized processing, approximate evaluation which is calculated by prior correspondence saved in the database replaces the expensive searching correspondences process in the entire point cloud. Finally, a multiscale filter is developed to remove error correspondences by the geometric information from multiple scale descriptor features. Experimental demonstrate that the proposed approach can optimize a feature subset with higher-descriptive capability compared to other methods and show superior PCR performance on 14 point cloud models. This is the first paper on point cloud descriptor optimization, which provides a new idea for PCR research.
Yue Wu 0004, Jinlong Sheng, Hangqi Ding, Peiran Gong, Hao Li 0009, Maoguo Gong, Wenping Ma 0001, Qiguang Miao
IEEE Trans. Evol. Comput.7
2025 An Adaptive Dual-Supervised Cross-Deep Dependency Network for Pixel-Wise Classification
abstract
With the advancement of remote sensing (RS) technology and satellite observation, the task of fusing multisource data, such as multispectral (MS) and panchromatic (PAN) images, has become increasingly important. However, image fusion involving certain semantic differences can hinder the model’s ability to learn effective feature mappings. To reconstruct richer and more consistent features during fusion, we propose an adaptive dual-supervised cross-deep dependency network (ADCD-Net), which consists of two training stages. Stage I uses a semantic perceptual self-supervision strategy (SPS) to learn deep features across different modalities, thereby reducing semantic differences while mining its own non-singular features. Stage II uses the deep temporal Mamba module (DTM-Module) to interactively learn the output of each network layers, which are able to take part in the deep feature reinforcement and improve the classification performance of semantic information. Finally, to eliminate channel redundancy during the two-stage network training process while enhancing spatial location memory and feature discrimination in the 2-D features, we propose a deformable interactive attention module (DIA-Module) to further bolster feature representation capabilities. Additionally, we conduct comparative and transfer experiments on multiple RS datasets, achieving outstanding classification results. Our code is available athttps://github.com/ChenC1027/ADCD-Net.
Wenping Ma 0001, Mengru Ma, Hekai Zhang, Hao Zhu 0009, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2025 Dual-Path Prototype Feature Decoupling Alignment Network for Panchromatic and Multispectral Classification
abstract
In recent years, with the rapid advancements and widespread application of satellite photography technology, it has become increasingly possible to obtain high-quality panchromatic (PAN) and multispectral (MS) data, which has provided new opportunities and challenges for multisource information fusion and classification research. Remote sensing data have the characteristics of small interclass differences and large intraclass differences, which easily leads to category confusion in network learning. In addition, how to fully tap the advantages of multisource data, better align multisource features, improve classification accuracy, and achieve collaborative classification are key issues that need to be solved urgently. In this article, a dual-path prototype feature decoupling alignment network (DPFDA-Net) is designed to solve the above issues. The network consists of two components: a prototype feature embedding (PFE) module and a feature alignment module (FAM) based on prototype decoupling. In the feature extraction stage, the PFE module uses the prototype concept to learn the discriminative prototype features of each category of the dual-source data separately, making the boundaries between categories more obvious. The FAM operates at the dual-source prototype feature level and achieves feature alignment by decoupling single-source prototype features and performing feature transformation to supplement the missing information of another data source. Finally, we use the aligned features for classification. The results of the experiment demonstrate that our approach has made significant progress in improving classification precision. The code is available athttps://github.com/Xidian-AIGroup190726/DPFDANet.
Wenping Ma 0001, Yanshan Guo, Hao Zhu 0009, Wenhao Zhao, Mengru Ma, Yue Wu 0004, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2025 LSCF: Long-Term Semantic-Guidance ConvFormer for Referring Remote Sensing Image Segmentation
abstract
Referring Remote Sensing Image Segmentation (RRSIS) task aims to generate segmentation masks for target objects based on language descriptions. It requires precise localization while distinguishing between visually similar yet semantically distinct objects. Fusing vision-language features only during extraction causes information loss and semantic forgetting in the decoder, harming similar target distinction. Additionally, high-resolution remote sensing images present challenges, including complex backgrounds, diverse object scales, and intricate boundaries, limiting the effectiveness of previous methods. To address these issues, we propose the Long-term Semantic-guidance ConvFormer (LSCF) Network. First, we fuse multi-receptive-field local features extracted by the Multi-scale CoordConv (MCC) module with language-aware global features from the Cross-modal Attention (CA) module to obtain multi-modal representations. Second, the Sampling Attention (SA) module enables fine-grained vision context alignment under semantic guidance. Finally, the Global Language Fusion (GLF) module is incorporated in the decoder to maintain long-term vision-language alignment and mitigate semantic degradation. Experimental validation on the RefSegRS, RRSIS-D, and RISBench datasets demonstrates that LSCF achieves oIoU scores of 83.27%, 77.42%, and 74.88%, and mIoU scores of 77.44%, 64.25%, and 68.53%, respectively. On RefSegRS, LSCF surpasses the SOTA method FIANet by 5.53% (oIoU) and 9.58% (mIoU), while delivering competitive performance on RRSIS-D and RISBench. Code and experimental configurations will be released.
Lingling Li 0002, Xiaoqiang Lu, Licheng Jiao, Fang Liu 0001, Wenping Ma 0001, Xu Liu 0006
IEEE Trans. Geosci. Remote. Sens.6
2025 A Diff-Attention Aware State-Space Fusion Model for Remote Sensing Classification
abstract
Multispectral (MS) and panchromatic (PAN) images describe the same land surface, so these images not only have their own advantages, but also share a significant amount of redundant information. In order to separate similar information and each modality’s unique advantages, thereby reducing feature redundancy at the fusion stage, this paper introduces a diff-attention aware state space fusion model (DASF-Model) for multimodal remote sensing image classification. Based on the selective state space model, a cross-modal diff-attention module (CDAM) is designed to extract and separate the common features and their respective dominant features of MS and PAN images. Specifically, space preserving visual mamba (SPVM) retains image spatial features and captures local features by appropriately optimizing visual mamba’s input. Considering that features in the fusion stage will have large semantic differences after feature separation and traditional mean fusion method fails to effectively integrate these features with significant discrepancies, an attention-aware linear fusion module (ALFM) is proposed. It performs pixel-wise linear fusion by calculating influence coefficients. This mechanism can fuse features with large semantic differences while keeping the feature size unchanged. Empirical evaluations indicate that the presented method achieves better results than alternative approaches. The relevant code can be found at: https://github.com/AVKSKVL/DAS-F-Model.
Wenping Ma 0001, Boyou Xue, Mengru Ma, Hekai Zhang, Hao Zhu 0009
IEEE Trans. Geosci. Remote. Sens.1
2025 Dense-Weak Ship Detection Based on Foreground-Guided Background Generation Network in SAR Images
abstract
Currently, ship detection based on Synthetic Aperture Radar (SAR) images still faces significant challenges, particularly in detecting weak and densely distributed ships within complex backgrounds. In areas such as ports and land, the complex background features often resemble those of densely distributed ships, leading to reduced detection accuracy. Additionally, the overlapping and mutual interference of features among dense ships can cause the network to miss detections or produce false positives. Therefore, this paper proposes a Foreground-Guided Background Generation Network (FGBG-Net), which includes a Gaussian Foreground Localization (GFL) model and a Background Feature Removal (BFR) module. The GFL module identifies the approximate high-probability regions of ship foregrounds on the feature map, guiding the network to focus on these regions. The BFR module then progressively removes background interference features based on the positions provided by the GFL module, generating feature maps that are more suitable for detecting weak and dense ships. Our network has been validated on multiple SAR ship datasets, and the experimental results demonstrate noticeable performance improvements, with a mean Average Precision (mAP) increase of 3.4% on the SSDD and HRSID datasets. The relevant code is available at the following link: https://github.com/Xidian-AIGroup190726/FBGBNet/tree/master.
Wenping Ma 0001, Xiaoting Yang, Hao Zhu 0009, Xiaoteng Wang, Biao Hou, Mengru Ma, Yue Wu 0004
IEEE Trans. Geosci. Remote. Sens.1
2025 A Mamba-Aware Spatial-Spectral Cross-Modal Network for Remote Sensing Classification
abstract
This study introduces a novel cross-modal spatial-spectral interaction Mamba (CMS2I-Mamba) for remote sensing image fusion classification. Unlike convolution-based models focusing on local details and Transformer-based models with high computational complexity, CMS2I-Mamba efficiently models global long-range dependencies in a linear complexity manner. First, multispectral (MS) and panchromatic (PAN) images each have unique advantages in the spectral and spatial attributes. Given this, this paper innovatively designs the multi-path selective-scan mechanism (MPS2M), which applies different path scanning strategies to deeply capture the global features from both spectral and spatial dimensions, aiming to enhance the robustness and complementarity of spatial-spectral features. Secondly, to overcome the characterization differences between images acquired by different sensors, this paper further introduces the channel interaction alignment module (CIAM). This module employs efficient former-last and oddeven channel interaction strategies to achieve precise semantic alignment of deep features between modalities. Finally, to leverage the shared fusion features to guide the unique singular features, this paper proposes a semantic-aware calibration module (SACM), which accurately constraints and calibrates the same semantic information in deep features. This not only enhances the model’s ability to understand scene semantics, but also promotes the deep fusion and utilization of information between different modalities. Through experimental verification on multiple datasets, the CMS2I-Mamba proposed in this paper shows excellent recognition performance and computational efficiency (parameter quantity and running speed) in fusion classification tasks. The code for CMS2I-Mamba is available at: https://github.com/ru-willow/CMSI-Mamba.
Mengru Ma, Jiaxuan Zhao, Wenping Ma 0001, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Change Knowledge-Guided Vision-Language Remote Sensing Change Detection
abstract
Remote sensing image change detection plays a critical role in applications like video surveillance and geographic information systems. However, existing binary and semantic change detection methods often rely solely on visual information, neglecting language information, which limits interpretability and the ability to provide specific change details. This work proposes the Change Knowledge-Guided Vision-Language Remote Sensing Change Detection (CKCD) method to address these limitations. By introducing change knowledge as language information, CKCD enhances semantic understanding and change detail representation. A Cross-Modal Affinity (CMA) module is designed to effectively fuse visual and textual features, improving information complementarity and fusion coherence. CKCD further enhances data utilization efficiency by merging change area detection and change category information into a single output through endto- end learning. This design reduces redundant data representations and simplifies the detection process, leading to a more compact and efficient use of the input data without requiring additional branches or multiple output heads. Experimental results demonstrate consistent performance improvements over traditional methods across multiple change detection datasets.
Jiahao Wang 0002, Fang Liu 0001, Licheng Jiao, Hao Wang 0211, Shuo Li 0010, Lingling Li 0002, Puhua Chen, Xu Liu 0006, Wenping Ma 0001
IEEE Trans. Geosci. Remote. Sens.9
2025 Dual Causal-Aware Detection Transformer for Remote Sensing Images
abstract
Deep neural networks often inherit biases from training data, compromising generalization. In visual recognition, distinguishing foreground from background is insufficient, as models tend to rely on spurious correlations rather than learning essential causal patterns. To address this issue, this paper proposes a novel transformer architecture, termed dual causal-aware detection transformer (DCDT), specifically designed for object detection in optical remote sensing images from a causal perspective. Specifically, we begin by constructing a structural causal model to intuitively analyze the causal effects inherent in the overall visual patterns. Building on this foundation, DCDT introduces causal constraints at the attention level by embedding dynamic multi-scale causal prototypes into the attention mechanism. The derived causal priors are subsequently used to enhance features at the representation level, thereby enforcing feature-level causal modulation. This dual causal-aware strategy enables the precise extraction and reinforcement of causally relevant features, improving both robustness and discriminative capability in complex detection scenarios. In addition, a sparse kernel-region mask is incorporated to decouple local information from global representations, effectively strengthening the modeling of fine-grained structures. Extensive experiments conducted on two challenging public datasets, DIOR and HRRSD, demonstrate that DCDT consistently outperforms existing methods and baselines. These results validate the effectiveness of DCDT in capturing both global causal semantics and local fine-grained features, highlighting its practicality in complex remote sensing scenarios.
Yuhan Wang 0007, Lingling Li 0002, Xu Liu 0006, Licheng Jiao, Zhongjian Huang, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001
IEEE Trans. Geosci. Remote. Sens.7
2025 Local-Global Spectral Feature-Aware Learning for Hyperspectral Imagery Classification
abstract
Effective modeling of the relationship between local spectral details and global contextual information remains a core challenge in hyperspectral image (HSI) classification. In this paper, a local-global spectral feature-aware network (LGSFA-Net) is proposed, which achieves local and global spectral feature learning through synergistic integration of local convolutional inductive biases and global state-space models (SSMs). The architecture of LGSFA-Net comprises three sequential components, including an embedding stage using convolutions for fundamental feature extraction, an encoding stage that cascades standard Mamba blocks with specialized interactive Mamba (IMamba) blocks and an enhanced spatial-spectral feature fusion (ESSFF) module. The proposed IMamba blocks employ separable convolutions and feature interaction learning for explicitly modeling the cross-channel spectral correlations learning, which can be effective in awareness of the spectral feature. And then, the ESSFF module utilizes self-attention mechanisms to dynamically balance local and global spatial-spectral feature contributions. The final prediction stage incorporates a lightweight classification head for efficient inference. Experimental results validate the effectiveness of the proposed methods for HSI classification on four benchmark datasets, including the PaviaU, Houston, Honghu, and Hanchuan datasets. The proposed LGSFA-Net achieves approximately 1.48%-2.51% increased overall accuracy (OA), 1.34%-2.06% increased average accuracy(AA), and 1.37%-3.75% increased Kappa on the aforementioned four datasets, respectively, outperforming the contrasting methods. The code implementation will be available at https://github.com/yutinyang/LGSFA-Net.
Yuting Yang 0008, Lingling Li 0002, Xu Liu 0006, Licheng Jiao, Fang Liu 0001, Shuo Li 0010, Wenping Ma 0001
IEEE Trans. Geosci. Remote. Sens.7
2025 CSCT: Channel-Spatial Coherent Transformer for Remote Sensing Image Super-Resolution
abstract
Remote sensing image super-resolution (RSISR) techniques are crucial in practice as an economical approach to enhancing the resolution of remote sensing images (RSIs). The scale of structural information and the richness of texture details in RSIs far exceed those in natural images. Therefore, accurately restoring and preserving edge and detail information are a critical challenge in the super-resolution (SR) process. Currently, convolutional neural network (CNN)-based methods primarily rely on local feature extraction, which fails to effectively capture and integrate global contextual information. Generative adversarial network (GAN)-based methods, while improving the visual quality, often suffer from artifacts and training instability, adversely affecting image quality. Moreover, these approaches struggle to accurately represent high-frequency features, leading to blurriness or distortion when reconstructing fine details and edges. To address these limitations, we introduce the channel–spatial coherent transformer (CSCT). The core of CSCT includes the channel–spatial coherent attention (CSCA) and the frequency-gated feed-forward network (FGFN), which work synergistically to enhance edge and detail preservation while significantly improving overall image clarity. CSCA efficiently aggregates channel and spatial information, while FGFN adaptively adjusts frequency information to enhance high-frequency details and suppress low-frequency noise. Moreover, this article leverages advanced data augmentation methods that markedly boost RSISR performance, offering new avenues for further exploration. The empirical analysis across several remote sensing SR benchmark datasets reveals that our approach excels in detail restoration, effectively reduces artifacts and noise, and significantly enhances the quality of SR images.
Kexin Zhang 0003, Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Wenping Ma 0001, Fang Liu 0001, Shuyuan Yang 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Adaptive Complex Wavelet Informed Transformer Operator
abstract
Visual transformers have achieved great success in representation learning. This is mainly due to efficient token dependency modeling via self-attention. However, the computational burden increases sharply as the input pixels increase. Although recent Fourier-based global frequency-domain mixing methods attempt to improve the efficiency of transformers for high-resolution image inputs, the Fourier operator has limited ability to capture the local geometric structure. Complex wavelets can perform local attention in both the spatial domain and the frequency domain. Therefore, we propose the complex wavelet informed transformer operator that uses the real and imaginary wavelets of the dual-tree complex wavelet transform to simulate the interaction in the attention kernel. In order to further reduce the computational burden of operators, we introduce an adaptive local block shared attention mechanism in the channel domain for our wavelet informed operators. Further, we construct the deep multi-head operator network consisting of a hybrid stack of complex wavelet informed transformer operators and self-attention layers. This enables the Transformer to more sparsely capture multi-scale and multi-directional structured features in the process of learning dependencies. Extensive experimental results show that our adaptive complex wavelet informed transformer operator under the Transformer architecture achieves highly competitive accuracy performance on multiple image classification benchmark datasets. And the proposed operators can be flexibly and effectively migrated to vision tasks in dynamic video scenarios.
Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Hao Zhu 0009, Xu Liu 0006, Lingling Li 0002, Wenping Ma 0001
IEEE Trans. Multim.8
2025 Uncertainty-Aware Semi-Supervised Learning Segmentation for Remote Sensing Images
abstract
Deep learning based remote sensing (RS) image segmentation significantly impacts several real application scenarios. Behind its success, massive labeled data plays an important role. However, annotating high-resolution RS images requires time-consuming and relevant expertise efforts. To address it, many works dive into semi-supervised learning which utilizes raw information embedded in unlabeled data to improve the segmentation model. Nevertheless, previous studies ignore the integrity and effectiveness of the potential context information hidden in RS data. In this work, we propose an uncertainty-aware masked consistency learning (U-MCL) framework that contains an uncertainty-aware masked denoising (U-MD) module and an uncertainty-aware masked image consistency (U-MIC) module. U-MCL initially generates a patch-wise uncertainty map for each unlabeled image during each training iteration, which is then used to derive an adaptive mask ratio for pseudo-label denoising in U-MD. Simultaneously, the uncertainty map is adopted to model a masked unlabeled image for reasoning unseen areas in U-MIC. Consequently, U-MCL is capable of enhancing model performance by engaging in accurate and stable consistency learning while preserving the integrity of the context and employing the context to infer the predictions of the masked regions safely. Extensive experiments on six RS datasets, i.e., ISPRS Vaihingen, FloodNet, MiniFrance, LoveDA, MER, and MSL, demonstrate the superiority of our U-MCL over recent most advanced methods, achieving new state-of-the-art performance under all benchmarks.
Xiaoqiang Lu, Lingling Li 0002, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001
IEEE Trans. Multim.6
2025 A 3D Self-Awareness Diffusion Network for Multimodal Classification
abstract
As imaging sensor technology in remote sensing has advanced quickly, multimodal fusion classification has become an important research direction in land cover and urban planning classification tasks. While generative models and image classification have greatly benefited from diffusion models, the present ones primarily concentrate on single-modality-driven diffusion processes. Therefore, this paper presents a 3D self-awareness diffusion network (3DSA-DiffNet) for multispectral (MS) and panchromatic (PAN) image fusion classification, which would make it easier to classify heterogeneous data from various sensors. First, in order to model the relationship between multi-channel spectra and multi-pixel spatial distributions as well as samples, respectively, a spatial-spectral joint denoising network (S$^{2}$JD-Net) is proposed. It can incorporate the diffusion process into the neural network to enhance the quality of diffusion features. Secondly, to imitate the brain's spatial-spectral coexistence learning mechanism, this work offers a 3D self-awareness module (3DSA-Module) that can learn the weight of each pixel in 3D space, resulting in extraordinarily high feature representation capabilities. Finally, experimental verification demonstrates that the 3D self-awareness diffusion fusion network driven by brain inspiration outperforms more sophisticated approaches on the Xi'an, Huhhot, and Muufl datasets.
Mengru Ma, Wenping Ma 0001, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Yuwei Guo 0001
IEEE Trans. Multim.2
2025 Tracking Like Human: Dynamic Scene Learning Reasoning Tracker in Satellite Videos
abstract
In satellite video object tracking, the individual frame analysis method is usually used for target localization, ignoring informative cues of the dynamic scene. Temporal information could contribute to identifying the target from distractors. In this work, a novel dynamic scene learning reasoning tracker is proposed for satellite videos, which reasons over temporal dynamic information to derive the target location. It is inspired by the tracking pattern through human perception and reasoning. First, static-dynamic united analysis is designed to construct dynamic scenes by concatenating the static searching results along the temporal dimension. Second, the information of each response object is aggregated by wavelet transforms. Meanwhile, these scenes are projected into low-frequency and high-frequency subspaces, which could imitate different levels of perceptions of humans for scenes. Third, an object-aware reasoning transformer is proposed to utilize the temporal dynamics of input response objects. In each subspace, it models the mutual interactions between dynamic objects and further learns the intrinsic property of each object for target reasoning. Finally, to obtain the current reasoning result, inverse wavelet transforms are utilized to integrate the results of low-frequency and high-frequency subspaces. The effectiveness of the proposed method is validated on three public satellite video datasets, including SV248S, SkySat, and VISO. Qualitative and quantitative experimental results show that the proposed tracker outperforms 22 popular approaches in seven challenging tracking satellite scenarios.
Licheng Jiao, Yangyang Li 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001
IEEE Trans. Multim.8
2025 Heterogeneous Riemannian Few-Shot Learning Network
abstract
How to learn and accurately distinguish new concepts from few samples, as humans do, is a long-standing concern in artificial intelligence (AI). Studies in brain science and neuroscience have shown that human brain perception is based on nonlinear manifolds, and high-dimensional manifolds can facilitate concept learning in neural circuits. Based on this inspiration, in this paper, we propose a heterogeneous Riemannian few-shot learning network (HRFL-Net), which is the first few-shot learning method to perform end-to-end deep learning on heterogeneous Riemannian manifolds. Specifically, to enhance the geometric invariance of the image representation, the image features are projected into three heterogeneous Riemannian manifold spaces. Then, the implicit Riemannian kernel function maps the manifolds to the separable high-dimensional reproducing Hilbert space. It is assumed that the embedded kernel features of the complementary manifolds are mapped to the same common subspace. Thus, a novel neural network-based Riemannian metric learning method is designed to solve the subspace feature vectors by imposing orthogonal normalized projection, which overcomes the data extension limitation of the Riemannian metric. Finally, with the optimization objective of increasing the interclass distance and decreasing the intraclass distance in Hilbert space, the HRFL-Net is trained with end-to-end stochastic optimization, and the optimal aggregation subspace is learned during the gradient descent process. Thus, the proposed HRFL-Net can be easily generalized to challenging nonconvex data. The evaluation of four public datasets shows that the proposed HRFL-Net has significant superiority and also achieves competitive results compared with the state-of-the-art methods.
Jie Chen 0098, Lingling Li 0002, Licheng Jiao, Fang Liu 0001, Xu Liu 0006, Yuwei Guo 0001, Puhua Chen, Wenping Ma 0001
IEEE Trans. Neural Networks Learn. Syst.8
2025 Brain-Inspired Learning, Perception, and Cognition: A Comprehensive Review
abstract
The progress of brain cognition and learning mechanisms has provided new inspiration for the next generation of artificial intelligence (AI) and provided the biological basis for the establishment of new models and methods. Brain science can effectively improve the intelligence of existing models and systems. Compared with other reviews, this article provides a comprehensive review of brain-inspired deep learning algorithms for learning, perception, and cognition from microscopic, mesoscopic, macroscopic, and super-macroscopic perspectives. First, this article introduces the brain cognition mechanism. Then, it summarizes the existing studies on brain-inspired learning and modeling from the perspectives of neural structure, cognitive module, learning mechanism, and behavioral characteristics. Next, this article introduces the potential learning directions of brain-inspired learning from four aspects: perception, cognition, understanding, and decision-making. Finally, the top-ten open problems that brain-inspired learning, perception, and cognition currently face are summarized, and the next generation of AI technology has been prospected. This work intends to provide a quick overview of the research on brain-inspired AI algorithms and to motivate future research by illuminating the latest developments in brain science.
Licheng Jiao, Mengru Ma, Pei He, Xueli Geng, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001, Biao Hou, Xu Tang 0004
IEEE Trans. Neural Networks Learn. Syst.7
2025 Chain-of-Situation Aware Progressive Inference Learning
abstract
The grounded situation recognition (GSR) task aims to recognize the structured semantics of an image to achieve "human-like" event understanding. Most previous studies primarily focus on the visual features of the situation, overlooking the step-by-step cognitive reasoning process that humans employ in complex task settings. Recently, the emergence of multimodal large language models (MLLMs) has provided novel directions for addressing complex problems. However, directly deploying MLLMs on the GSR task is suboptimal due to their tendency to exhibit "hallucination" issues. Additionally, fine-tuning MLLMs for the GSR task incurs high training costs. To address these challenges, inspired by human cognitive theory and the chain-of-thought (CoT) strategy, we propose the chain-of-situation progressive inference learning (CoS-PIL) framework, a lightweight approach that progressively completes verb prediction, noun prediction, and role grounding. The prediction of each step depends on the historical information of the previous step. Specifically, we first design situation prompts tailored to the GSR task and utilize MLLMs to analyze the input image and language prompts, generating heuristic response text for the current situation in the image. Instead of fine-tuning the MLLM, we activate the reasoning capabilities of the frozen MLLM and adapt its generated responses into three lightweight modules: CoS-Verb, CoS-Noun, and CoS-Ground. Considering that MLLMs may generate redundant content, we carefully design the chain-of-interest predictor (CoI-Predictor) to extract key information from the extensive response text and inject it into the model as prompts to enhance the performance. Extensive experiments on the challenging SWiG benchmark demonstrate that CoS-PIL outperforms other state-of-the-art methods. The code is publically available at https://github.com/XDLiuyyy/CoS-PIL.
Yang Liu 0349, Fang Liu 0001, Licheng Jiao, Qianyue Bao, Shuo Li 0010, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Wenping Ma 0001
IEEE Trans. Neural Networks Learn. Syst.9
2025 Pseudo Label Learning for Partial Point Cloud Registration
abstract
Partial point cloud registration plays a crucial role in computer vision and has widespread applications in 3D map construction, pose estimation, and high-precision localization. However, the collected point clouds often contain missing data due to hardware limitations and complex environments. Various partial registration algorithms have been proposed, most of which rely on estimating overlap regions. However, a significant proportion of these algorithms rely heavily on ground truth labels. Manual labeling is both time-consuming and labor-intensive, whereas algorithmic automatic labeling lacks sufficient accuracy. To tackle this issue, we present PSEudo Label learning for unsupervised partial point cloud registration (PSEL). This method utilizes complementary tasks to learn reliable pseudo labels for overlap regions and correspondences without depending on ground truth labels. The key idea is to use the complementarity between overlap estimation and registration to generate two types of pseudo labels based on the nearest points in pairs of aligned point clouds. These pseudo labels are then employed to supervise the learning of overlap regions and correspondences, gradually enhancing their accuracy throughout the learning process and ultimately establishing an unsupervised learning framework. PSEL consists of an overlap estimation module and a correspondence filtering module. The pseudo labels generated after registration are used to supervise both modules. Notably, the correspondence filtering module has two pipelines. The similarity and difference of the corresponding point features are used to eliminate false correspondences during the training and inference stages, respectively, with only the latter being optimized with pseudo labels. To validate the effectiveness of our registration method, we conducted experiments using the synthetic dataset ModelNet40, the indoor dataset 3DMatch, and the outdoor dataset KITTI.
Wenping Ma 0001, Yue Wu 0004, Yue Zhang 0040, Hao Zhu 0009, Biao Hou, Licheng Jiao
IEEE Trans. Vis. Comput. Graph.1
2024 M3SOT: Multi-Frame, Multi-Field, Multi-Space 3D Single Object Tracking
abstract
3D Single Object Tracking (SOT) stands a forefront task of computer vision, proving essential for applications like autonomous driving. Sparse and occluded data in scene point clouds introduce variations in the appearance of tracked objects, adding complexity to the task. In this research, we unveil M3SOT, a novel 3D SOT framework, which synergizes multiple input frames (template sets), multiple receptive fields (continuous contexts), and multiple solution spaces (distinct tasks) in ONE model. Remarkably, M3SOT pioneers in modeling temporality, contexts, and tasks directly from point clouds, revisiting a perspective on the key factors influencing SOT. To this end, we design a transformer-based network centered on point cloud targets in the search area, aggregating diverse contextual representations and propagating target cues by employing historical frames. As M3SOT spans varied processing perspectives, we've streamlined the network—trimming its depth and optimizing its structure—to ensure a lightweight and efficient deployment for SOT applications. We posit that, backed by practical construction, M3SOT sidesteps the need for complex frameworks and auxiliary components to deliver sterling results. Extensive experiments on benchmarks such as KITTI, nuScenes, and Waymo Open Dataset demonstrate that M3SOT achieves state-of-the-art performance at 38 FPS. Our code and models are available at https://github.com/ywu0912/TeamCode.git.
Yue Wu 0004, Maoguo Gong, Qiguang Miao, Wenping Ma 0001, Can Qin
AAAI5
2024 Evolutionary Multitasking with Compatibility Graph for Point Cloud Registration
abstract
3D point cloud registration is a fundamental task in computer vision, aimed at estimating a transformation to align a pair of point clouds. For point cloud registration, where the popular methods are used to build a compatibility graph. Different methods of constructing compatibility graphs can result in different information regions to be searched in the graph, due to varying constraints. Evolutionary multitask optimization has gained attention in the field of evolutionary computation, as it enables knowledge transfer among multitasks to enhance the exploration of information. Inspired by this theory, this paper proposes a method to utilize compatibility graphs as tasks through evolutionary multitasking for solving the problem of point cloud registration. We first construct two compatibility graph tasks with different tightness constraints for point cloud registration. Due to the different constraints of the two proposed tasks, the local information of the graph of interest might be biased. This bias can be effectively utilized in the evolutionary multitasking framework to enhance the ability to discover meaningful consensus relationships in the search space of the graph. Then we map the two tasks to a unified search space by designing clique selection strategies, and carrying out knowledge transfer between the two tasks, aiming to emphasize more on the local consensus information in the graphs. Lastly, the efficacy of our proposed method is further validated on multiple registration datasets.
Hangqi Ding, Yue Wu 0004, Hao Li 0009, Maoguo Gong, Wenping Ma 0001, Qiguang Miao
CEC6
2024 Inlier Confidence Calibration for Point Cloud Registration
abstract
Inliers estimation constitutes a pivotal step in partially overlapping point cloud registration. Existing methods broadly obey coordinate-based scheme, where inlier con-fidence is scored through simply capturing coordinate differences in the context. However, this scheme results in massive inlier misinterpretation readily, consequently affecting the registration performance. In this paper, we explore to extend a new definition called inlier confidence calibration (ICC) to alleviate the above issues. Firstly, we provide finely initial correspondences for ICC in order to generate high quality reference point cloud copy corresponding to the source point cloud. In particular, we develop a soft assignment matrix optimization theorem that offers faster speed and greater precision compared to Sinkhorn. Benefiting from the high quality reference copy, we argue the neighborhood patch formed by inlier and its neighborhood should have consistency between source point cloud and its reference copy. Based on this insight, we construct transformation-invariant geometric constraints and capture geometric structure consistency to calibrate inlier confidence for estimated correspondences between source point cloud and its reference copy. Finally, transformation is further calculated by the weighted SVD algorithm with the calibrated inlier confidence. Our model is trained in an unsupervised manner, and extensive experiments on synthetic and real-world datasets illustrate the effectiveness of the proposed method.
Yongzhe Yuan, Yue Wu 0004, Xiaolong Fan, Maoguo Gong, Qiguang Miao, Wenping Ma 0001
CVPR6
2024 Evolutionary Multitasking with Two-level Knowledge Transfer for Multi-view Point Cloud Registration
abstract
Point cloud registration is a hot research topic in the field of computer vision. In recent years, the registration method based on evolutionary computation has attracted more and more attention because of its robustness to initial pose and flexibility of objective function design. However, most of the current evolutionary computation-based point cloud registration methods do not take into account the multi-view problem, that is, to capture the close relationship between point clouds from different perspectives. We fully realize that if these relations are used correctly, the registration performance can be improved. Therefore, this paper proposes an evolutionary multitasking multi-view point cloud registration method, which solves the problem of multi-view error accumulation. To ensure the unity of global and local, a two-level knowledge transfer strategy is proposed, which divides the multi-view cloud registration task into two levels. This strategy unifies the search space of two registration tasks, solves the negative transfer phenomenon, and avoids the problem of falling into the local optimum. Finally, the effectiveness of the method is verified by sufficient experiments. This method has strong robustness to noise and outliers, and can be effectively implemented in various registration scenarios.
Hangqi Ding, Haoran Xu 0005, Yue Wu 0004, Hao Li 0009, Maoguo Gong, Wenping Ma 0001, Qiguang Miao, Jiao Shi, Yu Lei 0002
GECCO6
2024 PointMC: Multi-instance Point Cloud Registration based on Maximal Cliques
abstract
Multi-instance point cloud registration is the problem of estimating multiple rigid transformations between two point clouds. Existing solutions rely on global spatial consistency of ambiguity and the time-consuming clustering of highdimensional correspondence features, making it difficult to handle registration scenarios where multiple instances overlap. To address these problems, we propose a maximal clique based multiinstance point cloud registration framework called PointMC. The key idea is to search for maximal cliques on the correspondence compatibility graph to estimate multiple transformations, and cluster these transformations into clusters corresponding to different instances to efficiently and accurately estimate all poses. PointMC leverages a correspondence embedding module that relies on local spatial consistency to effectively eliminate outliers, and the extracted discriminative features empower the network to circumvent missed pose detection in scenarios involving multiple overlapping instances. We conduct comprehensive experiments on both synthetic and real-world datasets, and the results show that the proposed PointMC yields remarkable performance improvements.
Yue Wu 0004, Xidao Hu, Yongzhe Yuan, Xiaolong Fan, Maoguo Gong, Hao Li 0009, Mingyang Zhang 0002, Qiguang Miao, Wenping Ma 0001
ICML9
2024 MCDet: Multi-Content Collaboration Detector for Multiscale Remote Sensing Object
abstract
In previous works, powerful CNN backbones are typically used for one- or two-stage detectors to facilitate multi-categories object classification. Unfortunately, continuous convolution and pooling operations tend to weaken the detailed information. We propose an end-to-end Multi-content Collaboration Detector (MCDet) to improve object recognition accuracy. First, we summarize the reasons for the disappearance of detailed features in traditional feature extraction backbone networks, and propose a Shallow Clue Refinement (SCR) module, which helps us to retain more critical local detail information in the downsampling process. Second, to receive more suitable contextual information, we design a Self-dilating Spatial Pooling (SSP) module, it adaptively learns a contextual reception field, thereby alleviating the mismatch between the theoretical receptive field of the network design and the practical requirements. Finally, extensive experiments on the NWPU VHR-10 and DIOR datasets have shown that the proposed MCDet significantly improves detection accuracy. Our code is available at https://github.com/Xidian-AIGroup190726/RS-objectdetection-MCDet.
Wenping Ma 0001, Hao Zhu 0009, Yue Wu 0004, Biao Hou, Licheng Jiao
IEEE Geosci. Remote. Sens. Lett.1
2024 MBSI-Net: Multimodal Balanced Self-Learning Interaction Network for Image Classification
abstract
A growing number of earth observation satellites are able to simultaneously gather multimodal images of the same area due to the expanding availability and resolution of satellite remote sensing data. This paper proposes a novel multimodal balanced self-learning interaction network (MBSI-Net) for the classification task. It involves a dual-branch teacher-student network that enables knowledge interaction and transfer between the multimodalities. Firstly, in order to introduce statistical information in addition to local and global structural information, a texture feature equalization module (TFE-Module) is proposed. This can enhance the texture information of features through histogram equalization and further improve the representation ability of features. Secondly, to enable the student network to provide timely feedback questions, the paper proposes a feature fusion module (F2-Module) that models and enhances teacher features through the student network. This helps to raise the classification’s accuracy by incorporating information from multimodal images. Finally, the paper proposes a loss function based on structural similarity analysis to ensure balanced self-learning between the student and the teacher networks. Taking the multispectral (MS) and the panchromatic (PAN) images of the same scene as examples, through experimental verification, the proposed method can achieve good results on multiple datasets compared with other methods. Therefore, it offers an effective method for classifying and fusing multimodal data.
Mengru Ma, Wenping Ma 0001, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Shuyuan Yang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Learning Discriminative Features via Multi-Hierarchical Mutual Information for Unsupervised Point Cloud Registration
abstract
Extracting discriminative representations is the key step for correspondence-free point cloud registration. The extracted representations require to be discriminative to transformation, which demands representations to reduce the influence of redundant information irrelevant to transformation. However, recently proposed methods ignore this crucial property, resulting in limited ability to represent point cloud. In addition, researching correspondence-free point cloud registration has stagnated in recent years. In this paper, we try to relieve features redundancy issue for correspondence-free point cloud registration from a new perspective. Specifically, our method comprises two stages: feature extraction stage and rigid body transformation stage. In feature extraction stage, we aim to maximize multi-hierarchical mutual information between different hierarchical features, which can provide discriminative and less redundancy representations to regress transformation parameters for next stage. In rigid body transformation stage, we utilize dual quaternion to estimate transformation parameters, which combines rotation and translation simultaneously within a unified framework and obtains a compact representations for rigid transformation. The proposed model is trained in an unsupervised manner on the ModelNet40 dataset. The experimental results illustrate that our method achieves higher accuracy and robustness compared with existing correspondence-free methods.
Yongzhe Yuan, Yue Wu 0004, Mingyu Yue, Maoguo Gong, Xiaolong Fan, Wenping Ma 0001, Qiguang Miao
IEEE Trans. Circuits Syst. Video Technol.6
2024 Evolutionary Multiform Optimization With Two-Stage Bidirectional Knowledge Transfer Strategy for Point Cloud Registration
abstract
Point cloud registration is an important task in computer vision, where the goal is to estimate a transformation to align a pair of point clouds. Most of the existing registration methods face the problems of poor robustness and getting stuck in local optima. Evolutionary multitasking is an effective paradigm to enhance global search capability and improve convergence characteristics through knowledge transfer across multiple related tasks. Inspired by evolutionary multitasking, this article proposes a multiform optimization approach through evolutionary multitasking for solving the point cloud registration problems. We first construct two related registration tasks with different functional landscapes to form a multiform optimization problem. Compared with methods that only focus on a single registration attribute, the two proposed tasks focus on robustness and precision of registration, respectively. Then, a new two-stage bidirectional knowledge transfer strategy is presented, which can implement efficient knowledge transfer among two related tasks. Finally, both simulations and real experiments show the power of our method. The proposed method is robust to noise, outliers, and partial overlaps and is effective in multiple real registration scenarios, such as object registration, scene reconstruction, and simultaneous localization and mapping.
Yue Wu 0004, Hangqi Ding, Maoguo Gong, A. K. Qin 0001, Wenping Ma 0001, Qiguang Miao, Kay Chen Tan
IEEE Trans. Evol. Comput.5
2024 Intra- and Intersource Interactive Representation Learning Network for Remote Sensing Images Classification
abstract
Recently, remote sensing technology has developed faster and faster, and obtaining high-quality panchromatic (PAN) and multispectral (MS) images has become more accessible. The complementarity between them provides new opportunities in multisource remote sensing image classification. However, solving the problem of the semantic gap between multisource high-level features and, at the same time, utilizing the complementary properties between them to reduce intersource information redundancy is still a challenge. This article constructs an$I^{3}$RL-Net for the multisource remote sensing image classification task. Specifically, we design a cross-source interactive enhanced fusion module (CIEF-Module). For multilevel multisource features, by strengthening the dependencies of intrasource features and conducting intersource enhanced fusion, intrasource correlation features are refined, and the problem of the intersource semantic gap can be effectively alleviated. During the cross-source interaction process, we design a complementary representation supervised learning strategy (CRSL-Strategy). According to the similarities and differences of multisource features, it can adaptively promote complementary feature learning, thus generating a nonredundant multisource representation. The method has been verified to be effective on multiple RS datasets. The code is open source at:https://github.com/Xidian-AIGroup190726/Ping-Pie-I3RL-Net.git.
Wenping Ma 0001, Yanshan Guo, Hao Zhu 0009, Xiaoyu Yi 0002, Wenhao Zhao, Yue Wu 0004, Biao Hou, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2024 Significant Feature Elimination and Sample Assessment for Remote Sensing Small Objects' Detection
abstract
In recent years, small object detection has remained challenging in remote sensing tasks. Firstly, small objects inherently have fewer pixels, making them susceptible to interference from prominently featured larger objects during feature extraction. Secondly, existing detection methods solely based on the Intersection over Union (IOU) loss are disadvantageous for small object detection and fail to leverage the rich prior information in remote sensing images. Based on these observations, we propose a significant feature elimination and sample assessment network for small object detection called SESA-Net, based on the Facet derivative model. SESA-Net introduces prior information to the network through the directional derivatives characteristic of remote sensing images. The overall network comprises the ADM module and SIA strategy. The ADM module eliminates significant responses from shallow large objects, directing the network’s focus towards the features of shallow small objects. The Sample Importance Assessment (SIA) strategy addresses the limitations of the IOU loss function by using high-quality positive samples generated by ADM to provide an evaluation strategy for different positive samples of small objects. This enables the network to focus more on high-quality positive samples, thereby improving the accuracy of small object detection. The effectiveness of the proposed algorithm has been validated on multiple datasets. Our code is available at https://github.com/Xidian-AIGroup190726/RS-objectdetection-SESANet.
Wenping Ma 0001, Xiaoteng Wang, Hao Zhu 0009, Xiaoting Yang, Xiaoyu Yi 0002, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2024 Adaptive Feature Separation Network for Remote Sensing Object Detection
abstract
With the development of remote sensing technology, remote sensing object detection has been widely applied in various fields, but it still faces some thorny challenges, such as the following: 1) the complexity of object scale changes in remote sensing images makes it difficult to improve the performance of small object detection and 2) remote sensing images have complex backgrounds and densely arranged small and weak objects, which pose a serious problem of feature interference. To alleviate these challenges, we propose an end-to-end adaptive feature separation network called AFSNet, which includes a scale-aware module (SAM) and a class-aware module (CAM). The SAM mainly enables feature maps of different resolutions to detect objects of different scales. Shallow feature maps mainly suppress the features of large objects they contain to focus on small object detection, while deep feature maps increase the detailed features of large objects they contain to focus on large object detection. The CAM is mainly used to distinguish the features in the feature map by category, separating the features of different categories into different channels, thus mitigating the problem of inter class feature interference, and blocking background interference. The effectiveness of this article has been proven on the NWPU VHR-10, IPIU-M, DIOR, and DOTA2.0 datasets. It can be widely applied in civilian, military, and other fields. Through experimental verification, our AFSNet achieved 97.70% mAP on the NWPU VHR-10 dataset, 78.9% mAP on the DIOR dataset, and 58.22% mAP on the DOTA2.0 dataset. Our code is available at:https://github.com/Xidian-AIGroup190726/AFSNet.
Wenping Ma 0001, Yiting Wu, Hao Zhu 0009, Wenhao Zhao, Yue Wu 0004, Biao Hou, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2024 ISSP-Net: An Interactive Spatial-Spectral Perception Network for Multimodal Classification
abstract
Coordinated and complementary spatial-spectral information is represented by the panchromatic (PAN) and multispectral (MS) images. The optimal utilization of the advantages of these images has become a subject of intense research interest. This article introduces the interactive spatial-spectral perception network (ISSP-Net) for multimodal remote sensing image classification, addressing the challenge of optimal utilization of complementary information from PAN and MS images. First, the pixel-guided spatial enhancement module (PGSE-Module) improves spatial location interaction using the spatial location enhancement learning strategy (SLEL-Strategy) and the cross-spatial aggregation learning strategy (CSAL-Strategy), integrating multiscale contextual information and emphasizing pixel-level features. Second, the time-frequency collaborative spectral enhancement module (TFCSE-Module) distinguishes useful frequency domain features through channel separation, lightweight convolutions, and adaptive Fourier transform learning. This approach enables comprehensive utilization of both primary and auxiliary information from multimodal data. Finally, experiments on four datasets demonstrate the ISSP-Net’s state-of-the-art performance in classifying MS and PAN images, with good generalization to hyperspectral (HS) and LiDAR data. The code is provided at:https://github.com/sun740936222/ISSP-Net.
Wenping Ma 0001, Hekai Zhang, Mengru Ma, Biao Hou
IEEE Trans. Geosci. Remote. Sens.1
2024 Adaptive Multi-Scale Transformer Tracker for Satellite Videos
abstract
Satellite video tracking tasks are often characterized by blurred foreground boundaries in vast scenes, a wide range of targets varying in scale, and irregular changes in appearance. These challenges significantly impact the optimization of robust tracker performance. Therefore, it is imperative to extract diverse features with dynamic adaptive learning capabilities for the target being tracked in each sequence. In this article, we explore a novel adaptive multi-scale Transformer (MT) tracker for satellite videos to explore the potential spatiotemporal information of the target effectively. Specifically, a multi-scale spatial Transformer (MSST) is designed to leverage stage-by-stage spatial reduction and channel doubling, thereby enhancing the representation capabilities for the tracked target. In dynamic feature learning, an adaptive temporal Transformer (ATT) is then introduced based on multiple cross attentions, which analyzes the adaptive learning capacity for the dynamic target. It analyzes the weight proportion of different attentions automatically in the specific sequence through the learnable parameters. Finally, a multi-scale feature (MSF) regression module is crafted to improve the positioning accuracy of targets with low pixel counts in satellite scenes. This module accomplishes precise annotation of target boxes by effectively fusing features from diverse stages. We evaluate the proposed tracker performance on several public satellite datasets, including SatSOT, SV248S, and VISO. Experimental results show that the performance of our model can be comparable to the state-of-the-art trackers.
Xin Zhang 0167, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Globally-Aware Continuous-Time Redistribution Learning for RS Image Change Detection
abstract
Change detection (CD) based on deep learning has achieved excellent performance in recent years. However, these models exhibit limited capability in complete temporal modeling or face problems with fine-grained spatial features being overshadowed by the temporal context. Pure CNN-based CD pipelines also struggle to establish long-range connections. In this article, a globally aware continuous-time redistribution network (GCRNet) is proposed for RSCD. First, a boundary extraction branch is designed to preserve the semantic invariance of objects within the same boundary. This is achieved by providing boundary attention to adaptively guide the integration of temporal and spatial information. Then, a globally aware operator (GAO) is developed to obtain global interaction features. GAO utilizes the convolution theorem, which combines the Fourier transform and inverse Fourier transform, achieving it with low computational costs. Finally, an adaptive feature redistribution (AFR) module is designed to increase the distance between positive and negative samples in the latent space with change perception. It alleviates the effects of the severe class imbalance issue. Experimental results demonstrate that our proposed GCRNet surpasses 13 state-of-the-art CD methods. It achieves F1-score 0.33%, 0.62%, 0.84%, 0.17%, and 1.54% higher than the second-best model on the LEVIR-CD, LEVIR-CD+, WHU, CDD, and DSIFN datasets. The code of GCRNet is available athttps://github.com/XiaowenZhang-kuku/GCRNet.
Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 ConvGRU-Based Multiscale Frequency Fusion Network for PAN-MS Joint Classification
abstract
As a hot research topic in remote sensing, effectively integrating the advantageous features of multispectral and panchromatic images is the main challenge for fusing these two remote sensing images. This article proposes a multiscale frequency fusion network based on ConvGRU. To address the underutilization of texture features, we extract multiscale bandpass and low-pass sub-bands representing texture and content features through Contourlet decomposition. Multiscale bandpass sub-bands contain more comprehensive and concentrated texture details. Then, by proposing a multiscale frequency feature extractor based on ConvGRU, we effectively integrate and enhance sub-bands of different scales and frequencies, fully utilizing the characteristics of multispectral and panchromatic images and scale transmission. With these enhanced sub-band features, we obtain more comprehensive scale-enhanced texture features. Simultaneously, content features are also preserved as dual-source image features. Moreover, to reduce redundancy between fused features and make more efficient use of the obtained enhanced features, we designed an Inver-band integrator (IBI) module. It can fuse enhanced features at different scales, improve the complementarity between features, and thus achieve effective fusion. Experimental results demonstrate the effectiveness and robustness of our model on multiple datasets. Our codes are available athttps://github.com/Xidian-AIGroup190726/GMFnet.
Hao Zhu 0009, Xiaoyu Yi 0002, Biao Hou, Changzhe Jiao, Wenping Ma 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.6
2024 A Semantically Nonredundant Continuous-Scale Feature Network for Panchromatic and Multispectral Classification
abstract
In recent years, panchromatic (PAN) images and multispectral (MS) images, as a type of multimodal remote sensing data, are attracting increasingly more attention to their classification problems. However, effectively representing size variations of targets in remote sensing images and reducing redundant representations of different modalities’ deep semantic features to enhance classification accuracy remains a challenge. In this article, we propose a semantically nonredundant continuous-scale feature network (SNCF-Net) for PAN and MS classification, consisting of two modules: the texture-enhanced continuous scale input generation module and the cross-modal feature Kernel interaction (CMKI) module. By simulating the human eye’s adjustment of distance to observe objects of different sizes, we employ 3-D convolution to extract continuous-scale images generated by the texture-enhanced continuous-scale input generation (TCIG) module, enabling optimal feature representation of objects in remote sensing images. Additionally, the texture enhancement (TE) strategy in the TCIG module alleviates texture diffusion in scale space, enhancing the network’s ability to represent texture features. Subsequently, the CMKI module utilizes the response differences between different features to generate convolution kernels from deep feature maps, enabling feature interaction between the PAN modal and MS modal. This reduces redundant representations of essential image content information in deep features of two modalities, facilitating a better mapping between dual-modal features and categories. Our results achieve state-of-the-art performance on multiple datasets. The code is available athttps://github.com/Xidian-AIGroup190726/SNCFNet.
Hao Zhu 0009, Wenhao Zhao, Biao Hou, Changzhe Jiao, Zhongle Ren, Wenping Ma 0001, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.7
2024 Inter-Modal Masked Autoencoder for Self-Supervised Learning on Point Clouds
abstract
Masked autoencoder (MAE) is a recently widely used self-supervised learning method that has achieved great success in NLP and computer vision. However, the potential advantages of masked pre-training for point cloud understanding have not been fully explored. There is preliminary work on MAE-based point clouds using the Transformer architecture to explore low-level geometric representations in 3D space, which is insufficient for fine-grained decoding completion and downstream tasks. Inspired by multimodality, we propose Inter-MAE, a inter-modal MAE method for self-supervised learning on point clouds. Specifically, we first use Point-MAE as a baseline to partition point clouds into random low percentage of visible and high percentage of masked point patches. Then, a standard Transformer-based autoencoder is built by asymmetric design and shifting mask operations, and latent features are learned from the visible point patches aiming to recover the masked point patches. In addition, we generate image features based on ViT after point cloud rendering to form inter-modal contrastive learning with the decoded features of the completed point patches. Extensive experiments show that the proposed Inter-MAE generates pre-trained models that are effective and exhibit superior results in various downstream tasks. For example, an accuracy of 85.4% is achieved on ScanObjectNN and 86.3% on ShapeNetPart, outperforming other state-of-the-art self-supervised learning methods. Notably, our work establishes for the first time the feasibility of applying image modality to masked point clouds.
Yue Wu 0004, Maoguo Gong, Zhixiao Liu, Qiguang Miao, Wenping Ma 0001
IEEE Trans. Multim.6
2024 Self-Supervised Intra-Modal and Cross-Modal Contrastive Learning for Point Cloud Understanding
abstract
Learning effective representations from unlabeled data is a challenging task for point cloud understanding. As the human visual system can map concepts learned from 2D images to the 3D world, and inspired by recent multimodal research, we introduce data from point cloud modality and image modality for joint learning. Based on the properties of point clouds and images, we propose CrossNet, a comprehensive intra- and cross-modal contrastive learning method that learns 3D point cloud representations. The proposed method achieves 3D-3D and 3D-2D correspondences of objectives by maximizing the consistency of point clouds and their augmented versions, and with the corresponding rendered images in invariant space. We further distinguish the rendered images into RGB and grayscale images to extract color and geometric features, respectively. These training objectives combine feature correspondences between modalities to combine rich learning signals from point clouds and images. Our CrossNet is simple: we add a feature extraction module and a projection head module to the point cloud and image branches, respectively, to train the backbone network in a self-supervised manner. After the network is pretrained, only the point cloud feature extraction module is required for fine-tuning and directly predicting results for downstream tasks. Our experiments on multiple benchmarks demonstrate improved point cloud classification and segmentation results, and the learned representations can be generalized across domains.
Yue Wu 0004, Maoguo Gong, Peiran Gong, Xiaolong Fan, A. K. Qin 0001, Qiguang Miao, Wenping Ma 0001
IEEE Trans. Multim.8
2024 MPCT: Multiscale Point Cloud Transformer With a Residual Network
abstract
The self-attention (SA) network revisits the essence of data and has achieved remarkable results in text processing and image analysis. SA is conceptualized as a set operator that is insensitive to the order and number of data, making it suitable for point sets embedded in 3D space. However, working with point clouds still poses challenges. To tackle the issue of exponential growth in complexity and singularity induced by the original SA network without position encoding, we modify the attention mechanism by incorporating position encoding to make it linear, thus reducing its computational cost and memory usage and making it more feasible for point clouds. This article presents a new framework called multiscale point cloud transformer (MPCT), which improves upon prior methods in cross-domain applications. The utilization of multiple embeddings enables the complete capture of the remote and local contextual connections within point clouds, as determined by our proposed attention mechanism. Additionally, we use a residual network to facilitate the fusion of multiscale features, allowing MPCT to better comprehend the representations of point clouds at each stage of attention. Experiments conducted on several datasets demonstrate that MPCT outperforms the existing methods, such as achieving accuracies of 94.2% and 84.9% in classification tasks implemented on ModelNet40 and ScanObjectNN, respectively.
Yue Wu 0004, Maoguo Gong, Zhixiao Liu, Qiguang Miao, Wenping Ma 0001
IEEE Trans. Multim.6
2024 An Adaptive Migration Collaborative Network for Multimodal Image Classification
abstract
The multispectral (MS) and the panchromatic (PAN) images belong to different modalities with specific advantageous properties. Therefore, there is a large representation gap between them. Moreover, the features extracted independently by the two branches belong to different feature spaces, which is not conducive to the subsequent collaborative classification. At the same time, different layers also have different representation capabilities for objects with large size differences. In order to dynamically and adaptively transfer the dominant attributes, reduce the gap between them, find the best shared layer representation, and fuse the features of different representation capabilities, this article proposes an adaptive migration collaborative network (AMC-Net) for multimodal remote-sensing (RS) images classification. First, for the input of the network, we combine principal component analysis (PCA) and nonsubsampled contourlet transformation (NSCT) to migrate the advantageous attributes of the PAN and the MS images to each other. This not only improves the quality of images themselves, but also increases the similarity between the two images, thereby reducing the representational gap between them and the pressure on the subsequent classification network. Second, for the interaction on the feature migrate branch, we design a feature progressive migration fusion unit (FPMF-Unit) based on the adaptive cross-stitch unit of correlation coefficient analysis (CCA), which can make the network automatically learn the features that need to be shared and migrated, aiming to find the best shared-layer representation for multifeature learning. And we design an adaptive layer fusion mechanism module (ALFM-Module), which can adaptively fuse features of different layers, aiming to clearly model the dependencies among multiple layers for different sized objects. Finally, for the output of the network, we add the calculation of the correlation coefficient to the loss function, which can make the network converge to the global optimum as much as possible. The experimental results indicate that AMC-Net can achieve competitive performance. And the code for the network framework is available at: https://github.com/ru-willow/A-AFM-ResNet.
Wenping Ma 0001, Mengru Ma, Licheng Jiao, Fang Liu 0001, Hao Zhu 0009, Xu Liu 0006, Shuyuan Yang 0001, Biao Hou
IEEE Trans. Neural Networks Learn. Syst.1
2024 Explore the Influence of Shallow Information on Point Cloud Registration
abstract
Feature extraction is a key step for deep-learning-based point cloud registration. In the correspondence-free point cloud registration task, the previous work commonly aggregates deep information for global feature extraction and numerous shallow information which is positive to point cloud registration will be ignored with the deepening of the neural network. Shallow information tends to represent the structural information of the point cloud, while deep information tends to represent the semantic information of the point cloud. In addition, fusing information of different dimensions is conducive to making full use of shallow information. Inspired by this, we verify shallow information in the middle layers can bring a positive impact on the point cloud registration task. We design various architectures to combine shallow information and deep information to extract global features for point cloud registration. Experimental results on the ModelNet40 dataset illustrate that feature extractors that incorporate shallow information will bring positive performance.
Wenping Ma 0001, Mingyu Yue, Yue Wu 0004, Yongzhe Yuan, Hao Zhu 0009, Biao Hou, Licheng Jiao
IEEE Trans. Neural Networks Learn. Syst.1
2024 RORNet: Partial-to-Partial Registration Network With Reliable Overlapping Representations
abstract
Three-dimensional point cloud registration is an important field in computer vision. Recently, due to the increasingly complex scenes and incomplete observations, many partial-overlap registration methods based on overlap estimation have been proposed. These methods heavily rely on the extracted overlapping regions with their performances greatly degraded when the overlapping region extraction underperforms. To solve this problem, we propose a partial-to-partial registration network (RORNet) to find reliable overlapping representations from the partially overlapping point clouds and use these representations for registration. The idea is to select a small number of key points called reliable overlapping representations from the estimated overlapping points, reducing the side effect of overlap estimation errors on registration. Although it may filter out some inliers, the inclusion of outliers has a much bigger influence than the omission of inliers on the registration task. The RORNet is composed of overlapping points' estimation module and representations' generation module. Different from the previous methods of direct registration after extraction of overlapping areas, RORNet adds the step of extracting reliable representations before registration, where the proposed similarity matrix downsampling method is used to filter out the points with low similarity and retain reliable representations, and thus reduce the side effects of overlap estimation errors on the registration. Besides, compared with previous similarity-based and score-based overlap estimation methods, we use the dual-branch structure to combine the benefits of both, which is less sensitive to noise. We perform overlap estimation experiments and registration experiments on the ModelNet40 dataset, outdoor large scene dataset KITTI, and natural data Stanford Bunny dataset. The experimental results demonstrate that our method is superior to other partial registration methods. Our code is available at https://github.com/superYuezhang/RORNet.
Yue Wu 0004, Yue Zhang 0040, Wenping Ma 0001, Maoguo Gong, Xiaolong Fan, Mingyang Zhang 0002, A. K. Qin 0001, Qiguang Miao
IEEE Trans. Neural Networks Learn. Syst.3
2024 EGST: Enhanced Geometric Structure Transformer for Point Cloud Registration
abstract
We explore the effect of geometric structure descriptors on extracting reliable correspondences and obtaining accurate registration for point cloud registration. The point cloud registration task involves the estimation of rigid transformation motion in unorganized point cloud, hence it is crucial to capture the contextual features of the geometric structure in point cloud. Recent coordinates-only methods ignore numerous geometric information in the point cloud which weaken ability to express the global context. We propose Enhanced Geometric Structure Transformer to learn enhanced contextual features of the geometric structure in point cloud and model the structure consistency between point clouds for extracting reliable correspondences, which encodes three explicit enhanced geometric structures and provides significant cues for point cloud registration. More importantly, we report empirical results that Enhanced Geometric Structure Transformer can learn meaningful geometric structure features using none of the following: (i) explicit positional embeddings, (ii) additional feature exchange module such as cross-attention, which can simplify network structure compared with plain Transformer. Extensive experiments on the synthetic dataset and real-world datasets illustrate that our method can achieve competitive results.
Yongzhe Yuan, Yue Wu 0004, Xiaolong Fan, Maoguo Gong, Wenping Ma 0001, Qiguang Miao
IEEE Trans. Vis. Comput. Graph.5
2023 Dense Cross-Scale Transformer with Channel Learning for Remote Sensing Scene Classification
abstract
The recent explosive on Transformer has suggested its potential to become the mainstream model for feature representation and classification of remote sensing. Although Transformer has excellent global modeling capabilities, it lacks inductive bias. In contrast, CNNs show excellent performance in computer vision due to their strong inductive bias. To solve the above issue, and benefit from combining the Transformer model with CNNs, a dense cross-scale Transformer is proposed for remote sensing scene classification. Firstly, the attention aggregation based feature pyramid network (A2-FPN) is adopted to obtain the multi-scale features. And then, the multi-scale features are input into the multi-scale global features learning with channel attention (MSCA) module to obtain the global features. Besides, the multi-scale features are input into the dense cross-scale attention (DCSA) module to learn the multi-level cross-scale features. Finally, the outputs of these two modules are considered for computing the final class score. Experimental results obtained from the three public datasets indicate that the proposed method surpasses other remote sensing classification methods.
Yuting Yang 0008, Xu Liu 0006, Wenping Ma 0001, Licheng Jiao
IGARSS4
2023 Exploring Dual Representations in Large-Scale Point Clouds: A Simple Weakly Supervised Semantic Segmentation Framework
abstract
Existing work shows that 3D point clouds produce only about a 4% drop in semantic segmentation even at 1% random point annotation, which inspires us to further explore how to achieve better results at lower cost. As scene point clouds provide position and color information and often used in tandem as the only input, with little work going into segmentation by fusing information from dual spaces. To optimize point cloud representations, we propose a novel framework for the dual representation query network (DRQNet). The proposed framework partitions the input point cloud into position and color spaces, using the separately extracted geometric structure and semantic context to create an internal supervisory mechanism that bridges the dual spaces and fuses the information. Adopting sparsely annotated points as the query set, DRQNet provide guidance and perceptual information for multi-stage point clouds through random sampling. More, to differentiate and enhance the features generated by local neighbourhoods within multiple perceptual fields, we design a representation selection module to identify the contributions made by the position and color of each query point, and weight them adaptively according to reliability. The proposed DRQNet is robust to point cloud analysis and eliminates the effects of irregularities and disorder. Our method achieves significant performance gains on three mainstream benchmarks.
Yue Wu 0004, Maoguo Gong, Qiguang Miao, Wenping Ma 0001
ACM Multimedia5
2023 SACF-Net: Skip-Attention Based Correspondence Filtering Network for Point Cloud Registration
abstract
Rigid registration is a transformation estimation problem between two point clouds. The two point clouds captured may partially overlap owing to different viewpoints and acquisition times. Some previous correspondence matching based methods utilize an encoder-decoder network to carry out partial-to-partial registration task and adopt a skip-connection structure to convey information between the encoder and decoder. However, equally revisiting them with skip-connection may introduce the information redundancy, and limit the feature learning ability of the entire network. To address these problems, we propose a skip-attention based correspondence filtering network (SACF-Net) for point cloud registration. A novel feature interaction mechanism is designed to utilize both low-level geometric information and high-level context-aware information to enhance the original pointwise matching map. Additionally, a skip-attention based correspondence filtering method is proposed to selectively revisits features in the encoder at different resolutions, allowing the decoder to extract high-quality correspondences within overlapping regions. We conduct comprehensive experiments on indoor and outdoor scene datasets, and the results show that the proposed SACF-Net yields unprecedented performance improvements.
Yue Wu 0004, Xidao Hu, Yue Zhang 0040, Maoguo Gong, Wenping Ma 0001, Qiguang Miao
IEEE Trans. Circuits Syst. Video Technol.5
2023 INENet: Inliers Estimation Network With Similarity Learning for Partial Overlapping Registration
abstract
Point cloud registration is a key problem in the application of computer vision to robotics, autopilot and other fields. However, because the object is partially covered up or the resolution of 3D scanners is different, point clouds collected by the same sense may be inconsistent and even incomplete. Inspired by the recently proposed learning-based approaches, we propose Inliers Estimation Network (INENet) which includes a self-designed threshold prediction network and a probability estimation network with adaptive similarity mutual attention to help to find the overlapping area of the point clouds. In order to solve the above problems, we divide the partially overlapping point cloud registration task into two sub-tasks: overlapping areas detection and registration. The threshold prediction network can automatically calculate the threshold according to the input point clouds, and then the probability estimation network estimates the overlapping points by using threshold. The advantages of the proposed approach include: (1) threshold prediction network avoids bias and the complexity of manually adjusting the threshold. (2) Probability estimation network with similarity matrix can deeply fuse the information between a pair of point clouds, which is helpful to improve the accuracy. (3) INENet can be easily integrated into other overlapping region sensitive algorithms and without adjusting parameters. We conduct experiments on the ModelNet40, S3DIS and 3DMatch data sets. Specifically, the rotation error of the registration algorithm integrated with INENet is improved by at least 25% compared with direct partial overlap registration, our method improves the$F_{1} $score by 5% and has better anti-noise ability compared with the existing overlap detection methods, showing the effectiveness of the proposed method.
Yue Wu 0004, Yue Zhang 0040, Xiaolong Fan, Maoguo Gong, Qiguang Miao, Wenping Ma 0001
IEEE Trans. Circuits Syst. Video Technol.6
2023 SSMU-Net: A Style Separation and Mode Unification Network for Multimodal Remote Sensing Image Classification
abstract
The rapid progress in remote sensing technology has made it convenient for satellites to capture both multispectral (MS) and panchromatic (PAN) images. MS has more spectral information, and PAN has higher spatial resolution. How to exploit the complementarity between MS and PAN images, and effectively combine their respective advantageous features while alleviating mode differences, has become a crucial research task. This paper designs a Style Separation and Mode Unification network (SSMU-Net) for MS and PAN image classification from a novel and effective perspective. The network can be divided into two stages: style separation and mode unification. In the style separation stage, we use wavelet decomposition and techniques similar to generative adversarial networks to preliminarily separate the information of MS and PAN into different components. These components better preserve complete information from the original data and have their own advantages in style and content. Then we propose a Symmetrical Triplet Traction module to perform style traction on different components, making style features more unique and content features more unified, achieving feature separation and purification. In the mode unification stage, we design an encoder-decoder model to reduce the impact of mode differences. The experimental results from multiple datasets validate the effectiveness of our proposed method. Our overall accuracy improved by approximately 4% on the Shanghai and Beijing datasets, and it has exceeded 99.28% on the Hohhot and Vancouver datasets. Our code is available at: https://github.com/proudpie/SSMU-Net.
Hao Zhu 0009, Licheng Jiao, Xiaoyu Yi 0002, Biao Hou, Wenping Ma 0001, Shuang Wang 0001
IEEE Trans. Geosci. Remote. Sens.7
2023 A Dual-Stream Transformer With Diff-Attention for Multispectral and Panchromatic Classification
abstract
To minimize the feature redundancy of multispectral (MS) and panchromatic (PAN) images and maximize the complementary advantages of PAN and MS, a Dual-Stream Transformer with Diff-attention (DSTD)-Net is proposed for PAN and MS classification in this paper. Firstly, in terms of feature extraction, we use Self-attention and Co-attention (SCA) block to extract both specific advantageous features and common essential features. Based on that, a self-attention module strengthened by diff-attention (SSDA) that pays attention to the difference between two specific advantageous features is designed to reduce the essential redundancy in specific features. It can take advantage of the difference between two specific features and reduce the essential redundancy of the specific advantageous features, making them purer and better for classification. Finally, since the specific features and common features of multispectral (MS) and panchromatic (PAN) images make different contributions to classification, a Multi-stage Gated Fusion (MGF) strategy is used. The MGF strategy mainly uses Gated multisource units (GMU) to adapt the weight of different features and fuse them. So, our MGF strategy can strengthen the specific advantageous features beneficial for classification. Above all, the several experiment results verify our proposed networks’ effectiveness and robustness. Our code is available at: https://github.com/blackkiring/DSTD.
Lin Xu 0012, Hao Zhu 0009, Licheng Jiao, Wenhao Zhao, Biao Hou, Zhongle Ren, Wenping Ma 0001
IEEE Trans. Geosci. Remote. Sens.8
2023 Boundary-Aware Multiscale Learning Perception for Remote Sensing Image Segmentation
abstract
For remote sensing image segmentation, the boundaries of objects are difficult to distinguish, which is ignored by most methods. Therefore, it is challenging how to excavate and recover the boundaries of objects accurately. In this article, we propose a boundary-aware multi-scale network (BMNet) to solve this problem. The key components of BMNet include the scale attention module (SA-module) and boundary guidance module (BG-module). Specifically, SA-module is proposed to guide the refinement of multi-scale features in a context-aware way. It enhances the discriminability of multi-scale features by establishing contextual dependencies, which enables the refinement of the prediction of objects. Then, BG-module is proposed to enable networks to distinguish the boundary of objects. It utilizes manifold information of features to generate boundary guidance maps and forces the network to focus more on the boundary of objects. The effectiveness of the proposed BMNet is demonstrated on two public remote sensing datasets: ISPRS 2-D semantic labeling Potsdam dataset and Vaihingen dataset, where BMNet achieves better segmentation than prevalent methods. Finally, the experimental results indicate that BMNet can produce sharper boundaries of objects to reconstruct more detailed segmentation results.
Chao You, Licheng Jiao, Xu Liu 0006, Lingling Li 0002, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001
IEEE Trans. Geosci. Remote. Sens.6
2023 Transformer Based Conditional GAN for Multimodal Image Fusion
abstract
Multimodal Image fusion is becoming urgent in multi-sensor information utilization. However, existing end-to-end image fusion frameworks ignore a priori knowledge integration and long-distance dependencies across domains, which brings challenges to the network convergence and global image perception in complex scenes. In this paper, a conditional generative adversarial network with transformer (TCGAN) is proposed for multimodal image fusion. The generator is to generate a fused image with the source images content. The discriminators are adopted to distinguish the differences between the fused image and the source images. Adversarial training makes the final fused image to maintain the structural and textural details in the cross-modal images simultaneously. In particular, a wavelet fusion module makes the inputs contain image content from different domains as much as possible. The extracted convolutional features interact in the multiscale cross-modal transformer fusion module to fully complement the associated information. It makes the generator to focus on both local and global context. TCGAN fully considers the training efficiency of the adversarial process and the integrated retention of redundant information. Various experimental results of TCGAN have highlighted targets, rich details, and fast convergence properties on public datasets.
Jun Zhang 0045, Licheng Jiao, Wenping Ma 0001, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Puhua Chen, Shuyuan Yang 0001
IEEE Trans. Multim.3
2023 A Multi-Scale Progressive Collaborative Attention Network for Remote Sensing Fusion Classification
abstract
With the development of remote sensing technology, panchromatic images (PANs) and multispectral images (MSs) can be easily obtained. PAN has higher spatial resolution, while MS has more spectral information. So how to use the two kinds of images' characteristics to design a network has become a hot research field. In this article, a multi-scale progressive collaborative attention network (MPCA-Net) is proposed for PAN and MS's fusion classification. Compared to the traditional multi-scale convolution operations, we adopt an adaptive dilation rate selection strategy (ADR-SS) to adaptively select the dilation rate to deal with the problem of category area's excessive scale differences. For the traditional pixel-by-pixel sliding window sampling strategy, the patches which are generated by adjacent pixels but belonging to different categories contain a considerable overlap of information. So we change original sampling strategy and propose a center pixel migration (CPM) strategy. It migrates the center pixel to the most similar position of the neighborhood information for classification, which reduces network confusion and increases its stability. Moreover, due to the different spatial and spectral characteristics of PAN and MS, the same network structure for the two branches ignores their respective advantages. For a certain branch, as the network deepens, characteristic has different representations in different stages, so using the same module in multiple feature extraction stages is inappropriate. Thus we carefully design different modules for each feature extraction stage of the two branches. Between the two branches, because the strong mapping methods of directly cascading their features are too rough, we design collaborative progressive fusion modules to eliminate the differences. The experimental results verify that our proposed method can achieve competitive performance.
Wenping Ma 0001, Hao Zhu 0009, Licheng Jiao, Jianchao Shen, Biao Hou
IEEE Trans. Neural Networks Learn. Syst.1
2023 RDLNet: A Regularized Descriptor Learning Network
abstract
Local image descriptor learning has been instrumental in various computer vision tasks. Recent innovations lie with similarity measurement of descriptor vectors with metric learning for randomly selected Siamese or triplet patches. Local image descriptor learning focuses more on hard samples since easy samples do not contribute much to optimization. However, few studies focus on hard samples of image patches from the perspective of loss functions and design appropriate learning algorithms to obtain a more compact descriptor representation. This article proposes a regularized descriptor learning network (RDLNet) that makes the network focus on the learning of hard samples and compact descriptor with triplet networks. A novel hard sample mining strategy is designed to select the hardest negative samples in mini-batch. Then batch margin loss concerned with hard samples is adopted to optimize the distance of extreme cases. Finally, for a more stable network and preventing network collapsing, orthogonal regularization is designed to constrain convolutional kernels and obtain rich deep features. RDLNet provides a compact discriminative low-dimensional representation and can be embedded in other pipelines easily. This article gives extensive experimental results for large benchmarks in multiple scenarios and generalization in matching applications with significant improvements.
Jun Zhang 0045, Licheng Jiao, Wenping Ma 0001, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Hao Zhu 0009
IEEE Trans. Neural Networks Learn. Syst.3
2022 Deep Shearlet Network for Change Detection in SAR Images
abstract
Convolutional neural networks (CNN) can extract shift-invariant features, and have been widely applied in change detection task. However, common CNN lacks noise robustness and needs supervised data, to alleviate these problems, in this paper, we propose a novel deep shearlet network (ShearNet) for change detection in SAR images. In the network, a shearlet denoising layer (SDL) is designed to enhance the representation ability of common CNN. In SDL, feature maps are decomposed into subband coefficients by shearlet transform (ST). Due to optimal sparse representation property and highly direction sensitivity of ST, the network can capture important geometric information. Then, hard-threshold shrinkage is applied to high frequency subbands to drop small coefficients that are most likely to be noise, so that reduce the effect of noise. Finally, ShearNet is trained by introducing a noise-robust loss with noisy labels. The noisy labels are obtained by deep clustering that shows more robustness than existing preclassification methods. This fine-tuning process novelly follows the paradigm of learning from noisy labels to aside the difficulty of precisely labeling samples. Our experimental results on multiple real SAR datasets show that ShearNet can boost accuracy, and have better applicability for change detection in SAR images. The source code is available at https://github.com/yizhilanmaodhh/ShearNet.
Huihui Dong, Licheng Jiao, Wenping Ma 0001, Fang Liu 0001, Xu Liu 0006, Lingling Li 0002, Shuyuan Yang 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 A Multiscale Self-Attention Deep Clustering for Change Detection in SAR Images
abstract
Synthetic aperture radar (SAR) image change detection (CD) is an important application in the field of remote sensing. Due to the lack of labeled data especially in the pixelwise task, it is urgent to develop unsupervised techniques to effectively detect changes. In this article, we propose a novel unsupervised representation learning framework for CD in SAR images, called multiscale self-attention (SA) deep clustering based on octave convolution. The main motivation is that a convolutional neural network (CNN) has the ability to extract significant feature hidden in input images, but it relies heavily on annotated data. Clustering is typically free from supervision; however, SAR images always suffer from speckle noise, which is unfriendly for clustering. Thus, we integrate unsupervised clustering with CNN to learn clustering-friendly feature representations. In the unified framework, CNN feature learning and clustering can be optimized end-to-end without supervision. To better suppress speckle noise and boost the joint optimization for distinguishing changes and unchanges, we use the K-means++ algorithm that is robust to noise as the clustering algorithm. In the meanwhile, we introduce the octave convolution and SA mechanism into the network to fully mine important spatial structure information for enhancing noise resistance of the network. Moreover, multiscale fusion modules are proposed to fuse multiscale input into a complementary feature representation that contains more context and semantic information around each pixel so that it refines the difference feature extraction while reducing speckle noise. Experiments on challenging SAR data sets demonstrate the effectiveness and potential of the proposed model compared with the current state-of-the-art algorithms.
Huihui Dong, Wenping Ma 0001, Licheng Jiao, Fang Liu 0001, Lingling Li 0002
IEEE Trans. Geosci. Remote. Sens.2
2022 Feature Split-Merge-Enhancement Network for Remote Sensing Object Detection
abstract
Recently, multicategory object detection in high-resolution remote sensing images is still a challenge. First, objects with significant scale differences exist in one scene simultaneously, so it is generally difficult for the detectors to balance the detection performance of large and small objects. Second, because of the complex background and the objects’ densely distributed characteristics in the remote sensing images, the extracted features usually have noise and blurred boundaries, which interfere with the detection performance of the object detectors. With this observation, we propose an end-to-end scale-aware network called feature split–merge–enhancement network (SME-Net) for remote sensing object detection, composed of the feature split-and-merge (FSM) module, the offset-error rectification (OER) module, and the object saliency enhancement (OSE) strategy. FSM eliminates salient information of large objects to highlight the features of small objects in the shallow feature maps. It also transmits the effective detailed features of large objects to the deep feature maps, alleviating feature confusion between multiscale objects. OER corrects the inconsistency of the features spatial layout among the multilayer feature maps by the proposed offset loss, so as to achieve supervised elimination and transmission in FSM. OSE enhances the features of interests and suppresses the background information by the proposed membership function, thus preventing false detection and missed detection caused by noise and blurred boundaries. The effectiveness of the proposed algorithm has been verified on multiple datasets. Our code is available at:https://github.com/Momuli/SMENet.git
Wenping Ma 0001, Hao Zhu 0009, Licheng Jiao, Xu Tang 0004, Yuwei Guo 0001, Biao Hou
IEEE Trans. Geosci. Remote. Sens.1
2022 A Collaborative Correlation-Matching Network for Multimodality Remote Sensing Image Classification
abstract
Recently, with the increasing availability of the high-quality panchromatic (PAN) and multispectral (MS) remote sensing (RS) images, the inherent complementarity between PAN and MS images provides a wide development prospect for the multimodality RS image classification task. However, how to cleverly relieve the modal differences and effectively integrate the single-modality PAN and MS features is still a challenge. In this article, we design a collaborative correlation-matching network (CCM-Net) for multimodality RS image classification. Concretely, we first propose a bidirectional dominant feature supervision (Bi-DFS) learning, it utilizes single-modality dominant features as supplementary supervision information to establish the joint optimization loss function, thereby adaptively narrowing the differences between modalities before the feature extraction. In the feature extraction stage, the interactive correlation feature matching (ICFM) learning, composing the spatial feature matching (Spa-FM) and spectral feature matching (Spe-FM) strategies, is proposed to establish interactive matching and enhancement between multimodality strong correlation features from the perspective of spatial and spectral, respectively, thereby effectively alleviating the semantic deviation of multimodality features. Finally, we aggregate finer multilevel multimodality features to obtain top-level features with high discrimination. The effectiveness of the proposed algorithm has been verified on multiple datasets. Our code is available at:https://github.com/Momuli/CCM-Net.git.
Wenping Ma 0001, Hao Zhu 0009, Kenan Sun, Zhongle Ren, Xu Tang 0004, Biao Hou, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2022 Transfer Representation Learning Meets Multimodal Fusion Classification for Remote Sensing Images
abstract
To maximize the complementary advantages of synergistic multimodal, a transfer representation learning fusion network (TRLF-Net) is proposed for multisource remote sensing images collaborative classification in this article. First, with respect to the feature encoding, we design a dual-branch attention sparse transfer module (DAST-Module), which combines the spatial and channel attention (CA) masks to migrate the advantage attributes of the panchromatic (PAN) and the MS images mutually. This not only enhances their respective image advantages but also facilitates the sparse fusion of low-level features. Second, for the separation of multiscale information, a deep dual-scale decomposition module (DDSD-Module) is designed, which allows the decompose of high-frequency and low-frequency components. Then it uses the decomposed information to make the essential difference as small as possible, and the surrounding contour difference is as large as possible of the complementary multimodal image through the design of the loss function. Finally, to address the problem of large intraclass and small interclass differences, we develop a representation fusion of the global and local features’ module (RFGAL-Module). It mainly adopts global features to sort local features within classes, and then outputs them in a cascade. Thus, the characterization ability of features is improved, and the global and local features are used in a coordinated manner to accomplish the sample classification tasks. In particular, the experimental results demonstrate that TRLF-Net can obtain much improved accuracy and efficiency. The code is accessible in:https://github.com/ru-willow/SRLF-Net.
Mengru Ma, Wenping Ma 0001, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Shuyuan Yang 0001, Biao Hou
IEEE Trans. Geosci. Remote. Sens.2
2022 A Novel Adaptive Hybrid Fusion Network for Multiresolution Remote Sensing Images Classification
abstract
With the rapid development of earth observation technology, panchromatic (PAN) and multispectral (MS) images have also become easier to obtain. The multiresolution classification of PAN and MS images as a basic MS image analysis task has become a research hotspot. The main challenge in this field is how to process data and extract features to improve classification accuracy effectively. In this article, we design a novel adaptive hybrid fusion network (AHF-Net) for multiresolution remote sensing image classification. It includes two parts: data fusion and feature fusion. In the data fusion part, we propose an adaptive weighted intensity-hue-saturation (AWIHS) strategy, which can reduce the difference between MS and PAN images by adaptively adding each other’s unique information from the perspective of information sharing. In the feature fusion part, starting from the second-order correlation of features, we propose a correlation-based attention feature fusion (CAFF) module. It can improve the discrimination of fusion features by adaptively determining the fusion coefficient according to the importance of the input feature channel. Based on AWIHS and CAFF, inspired by the idea of feature pyramid, we combine the multilevel feature fusion and the dual-branch residual network as the backbone network of AHF-Net. By combining AWIHS and CAFF modules with the backbone network, our AHF-Net can effectively improve the classification accuracy of multiresolution remote sensing images. The effectiveness of the proposed algorithm has been verified on multiple data sets. Our code and model are available athttps://github.com/1826133674/AHF-Net.
Wenping Ma 0001, Jianchao Shen, Hao Zhu 0009, Jun Zhang 0045, Jiliang Zhao, Biao Hou, Licheng Jiao
IEEE Trans. Geosci. Remote. Sens.1
2021 Hyperspectral image classification based on spatial and spectral kernels generation network
Wenping Ma 0001, Hao Zhu 0009, Licheng Jiao, Biao Hou
Inf. Sci.1
2021 Dynamic threshold ECDSA signature and application to asset custody in blockchain
Huili Wang 0002, Wenping Ma 0001, Fuyang Deng, Haibin Zheng, Qianhong Wu
J. Inf. Secur. Appl.2
2021 A two-stage hybrid ant colony optimization for high-dimensional feature selection
Wenping Ma 0001, Hao Zhu 0009, Licheng Jiao
Pattern Recognit.1
2019 High-order graph matching based on ant colony optimization
Yue Wu 0004, Maoguo Gong, Wenping Ma 0001, Shanfeng Wang
Neurocomputing3
2019 PolSAR image classification based on multi-scale stacked sparse autoencoder
Lu Zhang 0028, Licheng Jiao, Wenping Ma 0001, Yiping Duan
Neurocomputing3
2019 Multimodal Remote Sensing Image Registration Based on Image Transfer and Local Features
abstract
Automatic registration is still a challenging problem for multimodal remote sensing images including optical, light detection and ranging, synthetic aperture radar images, and so on. Due to the differences in imaging principles, the gray value, texture, and landscape characteristic of these images are different in the local area. This also makes it difficult to obtain satisfactory results for the conventional image registration methods. In order to achieve registration of multimodal images to obtain complementary information, we apply the transfer algorithm based on a deep image analogy to the preprocessing of image registration. It eliminates the differences in multimodal remote sensing images by blending the original image structure and texture. The conventional local feature-based method is applied to match the original and generated images. Correspondences are increased and the registration error is reduced. The experiments demonstrate that our method can effectively deal with multimodal data and produce more accurate results. The algorithm is based on the joint of image deep semantic features and indirectly achieves matching of the original image pair. It provides a new solution to the problem of multimodal images registration.
Jun Zhang 0045, Wenping Ma 0001, Yue Wu 0004, Licheng Jiao
IEEE Geosci. Remote. Sens. Lett.2
2019 A Pareto-Based Sparse Subspace Learning Framework
abstract
High-dimensionality is a common characteristic of real-world data, which often results in high time and space complexity or poor performance of ensuing methods. Subspace learning, as one kind of dimension reduction method, provides a way to overcome the aforementioned problem. In this paper, we introduce multiobjective evolutionary optimization into subspace learning, and propose a Pareto-based sparse subspace learning algorithm for classification tasks. The proposed algorithm aims at minimizing two conflicting objective functions, the reconstruction error and the sparsity. A kernel trick derived from Gaussian kernel is implemented to the sparse subspace learning for the nonlinear phenomena of nature. In order to speed up the convergence, an entropy-driven initialization scheme and a gradient-descent mutation scheme are designed specifically. At last, a knee point is selected from the Pareto front to guarantee that we can obtain a solution with good classification performance, and yet as sparse as possible. The experiments and detailed analysis on real-life datasets and the hyperspectral images demonstrated that the proposed model achieves comparable results with the existing conventional subspace learning and evolutionary feature selection algorithms. Hence, this paper provides a more flexible and efficient approach for sparse subspace learning.
Juanjuan Luo, Licheng Jiao, Fang Liu 0001, Shuyuan Yang 0001, Wenping Ma 0001
IEEE Trans. Cybern.5
2019 A Novel Two-Step Registration Method for Remote Sensing Images Based on Deep and Local Features
abstract
Automatic remote sensing image registration has achieved great accomplishment. However, it is still a vital challenging problem to develop a robust and accurate registration method due to the negative effects of noise and imaging differences between images. For these images, it is difficult to guarantee the accuracy and robustness at the same time for one-step registration methods. To address this issue, we introduce an effective coarse-to-fine strategy and develop a new two-step registration method based on deep and local features in this paper. The first step is to calculate the approximate spatial relationship, which is obtained by a convolutional neural network. This step makes full use of the deep features to match and can generate stable results. For the second step, a matching strategy considering spatial relationship is applied to the local feature-based method. In addition, this step adopts more accurate features in location to adjust the results of the previous step. A variety of homologous and multimodal remote sensing images, including optical, synthetic aperture radar, and general map images, are used to evaluate the proposed method. The comparison experiments demonstrate that our method can apparently increase the correct correspondences, can improve the ratio of correct correspondences, and is highly robust and accurate.
Wenping Ma 0001, Jun Zhang 0045, Yue Wu 0004, Licheng Jiao, Hao Zhu 0009, Wei Zhao 0014
IEEE Trans. Geosci. Remote. Sens.1
2019 Local Restricted Convolutional Neural Network for Change Detection in Polarimetric SAR Images
abstract
To detect changed areas in multitemporal polarimetric synthetic aperture radar (SAR) images, this paper presents a novel version of convolutional neural network (CNN), which is named local restricted CNN (LRCNN). CNN with only convolutional layers is employed for change detection first, and then LRCNN is formed by imposing a spatial constraint called local restriction on the output layer of CNN. In the training of CNN/LRCNN, the polarimetric property of SAR image is fully used instead of manual labeled pixels. As a preparation, a similarity measure for polarimetric SAR data is proposed, and several layered difference images (LDIs) of polarimetric SAR images are produced. Next, the LDIs are transformed into discriminative enhanced LDIs (DELDIs). CNN/LRCNN is trained to model these DELDIs by a regression pretraining, and then a classification fine-tuning is conducted with some pseudolabeled pixels obtained from DELDIs. Finally, the change detection result showing changed areas is directly generated from the output of the trained CNN/LRCNN. The relation of LRCNN to the traditional way for change detection is also discussed to illustrate our method from an overall point of view. Tested on one simulated data set and two real data sets, the effectiveness of LRCNN is certified and it outperforms various traditional algorithms. In fact, the experimental results demonstrate that the proposed LRCNN for change detection not only recognizes different types of changed/unchanged data, but also ensures noise insensitivity without losing details in changed areas.
Fang Liu 0034, Licheng Jiao, Xu Tang 0004, Shuyuan Yang 0001, Wenping Ma 0001, Biao Hou
IEEE Trans. Neural Networks Learn. Syst.5
2019 A Novel Neural Network for Remote Sensing Image Matching
abstract
Rapid development of remote sensing (RS) imaging technology makes the acquired images have larger size, higher resolution, and more complex structure, which goes beyond the reach of classical hand-crafted feature-based matching. In this paper, we propose a feature learning approach based on two-branch networks to transform the image matching task into a two-class classification problem. To match two key points, two image patches centered at the key points are entered into the proposed network. The network aims to learn discriminative feature representations for patch matching, so that more matching pairs can be obtained on the premise of maintaining higher subpixel matching accuracy. The proposed network adopts a two-stage training mode to deal with the complex characteristics of RS images. An adaptive sample selection strategy is proposed to determine the size of each patch by the scale of its central key point. Thus, each patch can preserve the texture structure around its key point rather than all patches have a predetermined size. In the matching prediction stage, two strategies, namely, superpixel-based sample graded strategy and superpixel-based ordered spatial matching, are designed to improve the matching efficiency and matching accuracy, respectively. The experimental results and theoretical analysis demonstrate the feasibility, robustness, and effectiveness of the proposed method.
Hao Zhu 0009, Licheng Jiao, Wenping Ma 0001, Fang Liu 0001, Wei Zhao 0014
IEEE Trans. Neural Networks Learn. Syst.3
2018 PSOSAC: Particle Swarm Optimization Sample Consensus Algorithm for Remote Sensing Image Registration
abstract
Image registration is an important preprocessing step for many remote sensing image processing applications, and its result will affect the performance of the follow-up procedures. Establishing reliable matches is a key issue in point matching-based image registration. Due to the significant intensity mapping difference between remote sensing images, it may be difficult to find enough correct matches from the tentative matches. In this letter, particle swarm optimization (PSO) sample consensus algorithm is proposed for remote sensing image registration. Different from random sample consensus (RANSAC) algorithm, the proposed method directly samples the modal transformation parameter rather than randomly selecting tentative matches. Thus, the proposed method is less sensitive to the correct rate than RANSAC, and it has the ability to handle lower correct rate and more matches. Meanwhile, PSO is utilized to optimize parameter as its efficiency. The proposed method is tested on several multisensor remote sensing image pairs. The experimental results indicate that the proposed method yields a better registration performance in terms of both the number of correct matches and aligning accuracy.
Yue Wu 0004, Qiguang Miao, Wenping Ma 0001, Maoguo Gong, Shanfeng Wang
IEEE Geosci. Remote. Sens. Lett.3
2018 Deep Self-Paced Residual Network for Multispectral Images Classification Based on Feature-Level Fusion
abstract
The classification methods based on fusion techniques of multisource multispectral (MS) images have been studied for a long time. However, it may be difficult to classify these data based on a feature level while avoiding the inconsistency of data caused by multisource and multiple regions or cities. In this letter, we propose a deep learning structure called 2-branch SPL-ResNet which combines the self-paced learning with deep residual network to classify multisource MS data based on the feature-level fusion. First, a 2-D discrete wavelet is used to obtain the multiscale features and sparse representation of MS data. Then, a 2-branch SPL-ResNet is established to extract respective characteristics of the two satellites. Finally, we implement the feature-level fusion by cascading the two feature vectors and then classify the integrated feature vector. We conduct the experiments on Landsat_8 and Sentinel_2 MS images. Compared with the commonly used classification methods such as support vector machine and convolutional neural networks, our proposed 2-branch SPL-ResNet framework has higher accuracy and more robustness.
Wenping Ma 0001, Licheng Jiao
IEEE Geosci. Remote. Sens. Lett.3
2018 A modified convolutional neural network for face sketch synthesis
Licheng Jiao, Lingling Li 0002, Fang Liu 0001, Wenping Ma 0001
Pattern Recognit.5
2017 Remote Sensing Image Registration With Modified SIFT and Enhanced Feature Matching
abstract
The scale-invariant feature transform algorithm and its many variants are widely used in feature-based remote sensing image registration. However, it may be difficult to find enough correct correspondences for remote image pairs in some cases that exhibit a significant difference in intensity mapping. In this letter, a new gradient definition is introduced to overcome the difference of image intensity between the remote image pairs. Then, an enhanced feature matching method by combining the position, scale, and orientation of each keypoint is introduced to increase the number of correct correspondences. The proposed algorithm is tested on multispectral and multisensor remote sensing images. The experimental results show that the proposed method improves the matching performance compared with several state-of-the-art methods in terms of the number of correct correspondences and aligning accuracy.
Wenping Ma 0001, Zelian Wen, Yue Wu 0004, Licheng Jiao, Maoguo Gong, Yafei Zheng, Liang Liu 0001
IEEE Geosci. Remote. Sens. Lett.1
2017 Remote Sensing Image Registration Based on Multifeature and Region Division
abstract
Although many feature-based registration methods have been proposed, automatic image registration is still a challenging task due to the influence of various conditions and uncertain difficulties for remote sensing images. In this letter, a novel image registration method, including two types of feature detectors and a region boundary constraint strategy for matching, is proposed. Two types of features detected by scale-invariant feature transform and Harris operators have advantages of keeping different structural information in the image and increasing the number of keypoints for later matching. Afterward, a region boundary constraint strategy based on the image sketch map is utilized in matching step. This strategy restricts the detected two types of features in their respective structural region and nonstructural region to reduce the incorrect correspondences. Experimental results demonstrate the superiority of our proposed registration algorithm compared with other research works in terms of correct matching number and aligning accuracy.
Wenping Ma 0001, Yue Wu 0004, Yafei Zheng, Zelian Wen, Liang Liu 0001
IEEE Geosci. Remote. Sens. Lett.1
2017 Superpixel-Based Multiple Local CNN for Panchromatic and Multispectral Image Classification
abstract
Recently, very high resolution (VHR) panchromatic and multispectral (MS) remote-sensing images can be acquired easily. However, it is still a challenging task to fuse and classify these VHR images. Generally, there are two ways for the fusion and classification of panchromatic and MS images. One way is to use a panchromatic image to sharpen an MS image, and then classify a pan-sharpened MS image. Another way is to extract features from panchromatic and MS images, respectively, and then combine these features for classification. In this paper, we propose a superpixel-based multiple local convolution neural network (SML-CNN) model for panchromatic and MS images classification. In order to reduce the amount of input data for the CNN, we extend simple linear iterative clustering algorithm for segmenting MS images and generating superpixels. Superpixels are taken as the basic analysis unit instead of pixels. To make full advantage of the spatial-spectral and environment information of superpixels, a superpixel-based multiple local regions joint representation method is proposed. Then, an SML-CNN model is established to extract an efficient joint feature representation. A softmax layer is used to classify these features learned by multiple local CNN into different categories. Finally, in order to eliminate the adverse effects on the classification results within and between superpixels, we propose a multi-information modification strategy that combines the detailed information and semantic information to improve the classification performance. Experiments on the classification of Vancouver and Xi’an panchromatic and MS image data sets have demonstrated the effectiveness of the proposed approach.
Wei Zhao 0014, Licheng Jiao, Wenping Ma 0001, Jiaqi Zhao 0001, Jin Zhao 0002, Hongying Liu 0001, Xianghai Cao, Shuyuan Yang 0001
IEEE Trans. Geosci. Remote. Sens.3
2016 Multi-objective population-based incremental learning for community structure detection
abstract
The community structure detection of complex networks has become a hot topic in the past several years. In this paper, a new discrete framework of population-based incremental learning for complex networks problem is proposed. Based on the proposed discrete framework, a novel multi-objective population-based incremental learning algorithm is proposed to solve community structure detection problem. The proposed algorithm combines population-based incremental learning with the multi-objective evolutionary algorithm based on decomposition, this makes the evolution get directionality and converge fast. In order to discourage premature convergence, a random perturbation operator is adopted. The proposed algorithm has two contradictory objective functions termed as negative ratio association and ratio cut, respectively. The community structure detection results are a set of tradeoff solutions by simultaneous optimizing these two contradictory objective functions. Each of these solutions corresponds to a network community structure at one hierarchical level. Experiments on both real-world and synthetic networks prove the effectiveness of the proposed algorithm.
Wenping Ma 0001, Yue Wu 0004, Jie Yun
CEC1
2016 Discrete particle swarm optimization for high-order graph matching
Maoguo Gong, Yue Wu 0004, Wenping Ma 0001, A. K. Qin 0001, Zhenkun Wang 0001, Licheng Jiao
Inf. Sci.4
2016 A new quantum-behaved particle swarm optimization based on cultural evolution mechanism for multiobjective problems
Licheng Jiao, Wenping Ma 0001, Jingjing Ma 0001, Ronghua Shang
Knowl. Based Syst.3
2016 Unsupervised High-Level Feature Extraction of SAR Imagery With Structured Sparsity Priors and Incremental Dictionary Learning
abstract
Sparse representation is an effective model for high-level feature extraction, and the dictionary is critical, since it can provide a sparse and discriminative feature for image classification. However, the traditional sparse model with ℓ1- norm is unstable and ignores spatial context dependence. Furthermore, the traditional off-line dictionary learning is less efficient. In this letter, a high-level feature extraction approach is proposed, in which structured sparsity priors are imposed on the sparse representation to exploit the context dependence and an incremental structured dictionary learning method is proposed to exploit the inherent structures of a dictionary. The experiment results on unsupervised synthetic aperture radar imagery classification show that the structured priors improve classification performance and the proposed algorithm is more efficient in dictionary learning compared with existing works.
Jiawei Chen 0001, Licheng Jiao, Wenping Ma 0001, Hongying Liu 0001
IEEE Geosci. Remote. Sens. Lett.3
2016 Stacked Sparse Autoencoder in PolSAR Data Classification Using Local Spatial Information
abstract
Terrain classification is an important topic in polarimetric synthetic aperture radar (PolSAR) image processing. Among various classification techniques, the stacked sparse autoencoder (SSAE) is a kind of deep learning method that can automatically learn useful features layer by layer in an unsupervised manner. However, the scattering measurements of individual pixels in PolSAR images are affected by the speckle; hence, the performance of pixel-based classification approaches would be poor. In this situation, a novel framework is proposed to learn robust features of PolSAR data. The local spatial information is introduced into SSAE to learn the deep spatial sparse features automatically for the first time. Furthermore, the influences of the neighbor pixels on the central pixel are controlled depending on the spatial distances from the neighbor pixels to the central pixel. Experimental results with fully PolSAR data indicate that the proposed method provides a competitive solution.
Lu Zhang 0028, Wenping Ma 0001
IEEE Geosci. Remote. Sens. Lett.2
2016 SAR Image Registration Based on Multifeature Detection and Arborescence Network Matching
abstract
In this letter, a novel synthetic aperture radar (SAR) image registration method, including two operators for feature detection and arborescence network matching (ANM) for feature matching, is proposed. The two operators, namely, SAR scale-invariant feature transform (SIFT) and R-SIFT, can detect corner points and texture points in SAR images, respectively. This process has an advantage of preserving two types of feature information in SAR images simultaneously. The ANM algorithm has a two-stage process for finding matching pairs. The backbone network and the branch network are successively built. This ANM algorithm combines feature constraints with spatial relations among feature points and possesses a larger number of matching pairs and higher subpixel matching precision than the original version. Experimental results on various SAR images show that the proposed method provides superior performance than other approaches investigated.
Hao Zhu 0009, Wenping Ma 0001, Biao Hou, Licheng Jiao
IEEE Geosci. Remote. Sens. Lett.2
2016 Classification and saliency detection by semi-supervised low-rank representation
Miaoyun Zhao, Licheng Jiao, Wenping Ma 0001, Hongying Liu 0001, Shuyuan Yang 0001
Pattern Recognit.3
2015 A Novel Point-Matching Algorithm Based on Fast Sample Consensus for Image Registration
abstract
Robustness and accuracy are the two main challenging problems in feature-based remote sensing image registration. In this letter, a novel point-matching algorithm is proposed. An improved random sample consensus (RANSAC) algorithm called fast sample consensus (FSC) is proposed. It divides the data set in RANSAC into two parts: the sample set and the consensus set. Sample set has high correct rate and consensus set has a large number of correct matches. An iterative method is put forward to increase the number of correct correspondences. A set of measures has been used to evaluate the registration result. The performance of the proposed method is validated on the evaluation of these measures and the mosaic images. FSC can get more correct matches than RANSAC in less number of iterations, iterative selection of correct matches algorithm and removal of the imprecise points algorithm effectively increase the accuracy of the result. Extensive experimental studies compared with three state-of-the-art methods prove that the proposed algorithm is robust and accurate.
Yue Wu 0004, Wenping Ma 0001, Maoguo Gong, Linzhi Su, Licheng Jiao
IEEE Geosci. Remote. Sens. Lett.2
2014 A memetic algorithm for solving flexible Job-Shop Scheduling Problems
abstract
The flexible Job-shop Scheduling Problem (FJSP) is an extension of the classical job-shop scheduling problem (JSP). In this paper, a memetic algorithm (MA) for the FJSP is presented. This MA is a hybrid genetic algorithm which explores the search space and two efficient local searchers to exploit information in the search region. An extensive computational study on 49 benchmark problems shows that the algorithm is effective and robust, with respect to other well-known effective algorithms.
Wenping Ma 0001, Yi Zuo 0003, Jiulin Zeng, Licheng Jiao
IEEE Congress on Evolutionary Computation1
2014 Reference direction based immune clone algorithm for many-objective optimization
Ruochen Liu 0006, Chenlin Ma, Wenping Ma 0001, Licheng Jiao
Frontiers Comput. Sci.4
2014 Quadratic interpolation based orthogonal learning particle swarm optimization algorithm
Ruochen Liu 0006, Wenping Ma 0001, Caihong Mu, Licheng Jiao
Nat. Comput.3
2014 A novel cooperative coevolutionary dynamic multi-objective optimization algorithm using a new predictive model
Ruochen Liu 0006, Wenping Ma 0001, Caihong Mu, Licheng Jiao
Soft Comput.3
2013 Fuzzy C-Means Clustering With Local Information and Kernel Metric for Image Segmentation
abstract
In this paper, we present an improved fuzzy C-means (FCM) algorithm for image segmentation by introducing a tradeoff weighted fuzzy factor and a kernel metric. The tradeoff weighted fuzzy factor depends on the space distance of all neighboring pixels and their gray-level difference simultaneously. By using this factor, the new algorithm can accurately estimate the damping extent of neighboring pixels. In order to further enhance its robustness to noise and outliers, we introduce a kernel distance measure to its objective function. The new algorithm adaptively determines the kernel parameter by using a fast bandwidth selection rule based on the distance variance of all data points in the collection. Furthermore, the tradeoff weighted fuzzy factor and the kernel distance measure are both parameter free. Experimental results on synthetic and real images show that the new algorithm is effective and efficient, and is relatively independent of this type of noise.
Maoguo Gong, Jiao Shi, Wenping Ma 0001, Jingjing Ma 0001
IEEE Trans. Image Process.4
2012 A Novel Immune Clonal Algorithm for MO Problems
abstract
Research on multiobjective optimization (MO) becomes one of the hot points of intelligent computation. Compared with evolutionary algorithm, the artificial immune system used for solving MO problems (MOPs) has shown many good performances in improving the convergence speed and maintaining the diversity of the antibody population. However, the simple clonal selection computation has some difficulties in handling some more complex MOPs. In this paper, the simple clonal selection strategy is improved and a novel immune clonal algorithm (NICA) is proposed. The improvements in NICA are mainly focus on four aspects. 1) Antibodies in the antibody population are divided into dominated ones and nondominated ones, which is suitable for the characteristic of one multiobjective optimization problem has a series Pareto-optimal solutions. 2) The entire cloning is adopted instead of different antibodies having different clonal rate. 3) The clonal selection is based on the Pareto-dominance and one antibody is selected or not depending on whether it is a nondominated one, which is different from the traditional clonal selection manner. 4) The antibody population updating operation after the clonal selection is adopted, which makes antibody population under a certain size and guarantees the convergence of the algorithm. The influences of the main parameters are analyzed empirically. Compared with the existed algorithms, simulation results on MOPs and constrained MOPs show that NICA in most problems is able to And much better spread of solutions and better convergence near the true Pareto-optimal front.
Ronghua Shang, Licheng Jiao, Fang Liu 0001, Wenping Ma 0001
IEEE Trans. Evol. Comput.4
2010 Immune algorithm with orthogonal design based initialization, cloning, and selection for global optimization
Maoguo Gong, Licheng Jiao, Fang Liu 0001, Wenping Ma 0001
Knowl. Inf. Syst.4
2009 Hybrid Immune Algorithm with intelligent recombination
abstract
In this study, we introduce a hybrid immune algorithm based on the intelligent recombination operator and clonal selection algorithm. The intelligent recombination operator uses orthogonal experimental design for factor analysis which identifies the potential gene segments from two individuals to improve their antigenic affinities. The new algorithm, termed as Hybrid Immune Algorithm with Recombination (HIAR), can avoid the decrease of gene diversity in evolutionary process. It evaluates the hamming distance before recombination and uses the two individuals which have the largest hamming distance between each other to implement intelligent recombination operator. It is shown empirically that HIAR has better performance in solving benchmark functions as compared with Intelligent Evolutionary Algorithm and Clonal Selection Algorithm.
Maoguo Gong, Licheng Jiao, Wenping Ma 0001, Ronghua Shang
IEEE Congress on Evolutionary Computation3
2009 Intelligent multi-user detection using an artificial immune system
Maoguo Gong, Licheng Jiao, Wenping Ma 0001, Jingjing Ma 0001
Sci. China Ser. F Inf. Sci.3
2008 Unsupervised texture image segmentation using multiobjective evolutionary clustering ensemble algorithm
abstract
Multiobjective evolutionary clustering approach has been successfully utilized in data clustering. In this paper, we propose a novel unsupervised machine learning algorithm namely multiobjective evolutionary clustering ensemble algorithm (MECEA) to perform the texture image segmentation. MECEA comprises two main phases. In the first phase, MECEA uses a multiobjective evolutionary clustering algorithm to optimize two complementary clustering objectives: one based on compactness in the same cluster, and the other based on connectedness of different clusters. The output of the first phase is a set of Pareto solutions, which correspond to different tradeoffs between two clustering objectives, and different numbers of clusters. In the second phase, we make use of the meta-clustering algorithm (MCLA) to combine all the Pareto solutions to get the final segmentation. The segmentation results are evaluated by comparing with three known algorithms: K-means, fuzzy K-means (FCM), and evolutionary clustering algorithm (ECA). It is shown that MECEA is an adaptive clustering algorithm, which outperforms the three algorithms in the experiments we carried out.
Xiaoxue Qian, Xiangrong Zhang, Licheng Jiao, Wenping Ma 0001
IEEE Congress on Evolutionary Computation4
2008 Multiobjective optimization using an immunodominance and clonal selection inspired algorithm
Maoguo Gong, Licheng Jiao, Wenping Ma 0001, Haifeng Du
Sci. China Ser. F Inf. Sci.3
2007 Solving multidimensional knapsack problems by an immune-inspired algorithm
abstract
This paper introduces a computational model simulating the dynamic process of human immune response to solve multidimensional knapsack problems. The new model is a quaternion (G, I, R, Al), where G denotes exterior stimulus or antigen, I denotes the set of valid antibodies, R denotes the set of reaction rules describing the interactions between antibodies, and Al denotes the dynamic algorithm describing how the reaction rules are applied to antibody population. The set of antibody-adjusting rules, the set of clonal selection rules, and a dynamic algorithm, named M P-PAISA, are designed for solving multidimensional knapsack problems. The efficiency of the proposed algorithm was validated by testing on 57 benchmark problems and comparing with three genetic algorithms. The results indicated that the proposed algorithm was suitable for solving multidimensional knapsack problems.
Maoguo Gong, Licheng Jiao, Wenping Ma 0001, Shuiping Gou
IEEE Congress on Evolutionary Computation3
2006 Immune Clonal Selection Evolutionary Strategy for Constrained Optimization
Wenping Ma 0001, Licheng Jiao, Maoguo Gong, Ronghua Shang
PRICAI1