EDBT 2026 Demo / reviewers in the wild / expert
Mengru Ma
dblp:180/9155
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A trust-aware singular fusion network for multimodal image classification
Wenping Ma 0001, Mengru Ma, Hekai Zhang, Hao Zhu 0009, Licheng Jiao |
Neurocomputing | 3 |
| 2025 | An Adaptive Dual-Supervised Cross-Deep Dependency Network for Pixel-Wise ClassificationabstractWith the advancement of remote sensing (RS) technology and satellite observation, the task of fusing multisource data, such as multispectral (MS) and panchromatic (PAN) images, has become increasingly important. However, image fusion involving certain semantic differences can hinder the model’s ability to learn effective feature mappings. To reconstruct richer and more consistent features during fusion, we propose an adaptive dual-supervised cross-deep dependency network (ADCD-Net), which consists of two training stages. Stage I uses a semantic perceptual self-supervision strategy (SPS) to learn deep features across different modalities, thereby reducing semantic differences while mining its own non-singular features. Stage II uses the deep temporal Mamba module (DTM-Module) to interactively learn the output of each network layers, which are able to take part in the deep feature reinforcement and improve the classification performance of semantic information. Finally, to eliminate channel redundancy during the two-stage network training process while enhancing spatial location memory and feature discrimination in the 2-D features, we propose a deformable interactive attention module (DIA-Module) to further bolster feature representation capabilities. Additionally, we conduct comparative and transfer experiments on multiple RS datasets, achieving outstanding classification results. Our code is available athttps://github.com/ChenC1027/ADCD-Net. Wenping Ma 0001, Mengru Ma, Hekai Zhang, Hao Zhu 0009, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Dual-Path Prototype Feature Decoupling Alignment Network for Panchromatic and Multispectral ClassificationabstractIn recent years, with the rapid advancements and widespread application of satellite photography technology, it has become increasingly possible to obtain high-quality panchromatic (PAN) and multispectral (MS) data, which has provided new opportunities and challenges for multisource information fusion and classification research. Remote sensing data have the characteristics of small interclass differences and large intraclass differences, which easily leads to category confusion in network learning. In addition, how to fully tap the advantages of multisource data, better align multisource features, improve classification accuracy, and achieve collaborative classification are key issues that need to be solved urgently. In this article, a dual-path prototype feature decoupling alignment network (DPFDA-Net) is designed to solve the above issues. The network consists of two components: a prototype feature embedding (PFE) module and a feature alignment module (FAM) based on prototype decoupling. In the feature extraction stage, the PFE module uses the prototype concept to learn the discriminative prototype features of each category of the dual-source data separately, making the boundaries between categories more obvious. The FAM operates at the dual-source prototype feature level and achieves feature alignment by decoupling single-source prototype features and performing feature transformation to supplement the missing information of another data source. Finally, we use the aligned features for classification. The results of the experiment demonstrate that our approach has made significant progress in improving classification precision. The code is available athttps://github.com/Xidian-AIGroup190726/DPFDANet. Wenping Ma 0001, Yanshan Guo, Hao Zhu 0009, Wenhao Zhao, Mengru Ma, Yue Wu 0004, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | A Diff-Attention Aware State-Space Fusion Model for Remote Sensing ClassificationabstractMultispectral (MS) and panchromatic (PAN) images describe the same land surface, so these images not only have their own advantages, but also share a significant amount of redundant information. In order to separate similar information and each modality’s unique advantages, thereby reducing feature redundancy at the fusion stage, this paper introduces a diff-attention aware state space fusion model (DASF-Model) for multimodal remote sensing image classification. Based on the selective state space model, a cross-modal diff-attention module (CDAM) is designed to extract and separate the common features and their respective dominant features of MS and PAN images. Specifically, space preserving visual mamba (SPVM) retains image spatial features and captures local features by appropriately optimizing visual mamba’s input. Considering that features in the fusion stage will have large semantic differences after feature separation and traditional mean fusion method fails to effectively integrate these features with significant discrepancies, an attention-aware linear fusion module (ALFM) is proposed. It performs pixel-wise linear fusion by calculating influence coefficients. This mechanism can fuse features with large semantic differences while keeping the feature size unchanged. Empirical evaluations indicate that the presented method achieves better results than alternative approaches. The relevant code can be found at: https://github.com/AVKSKVL/DAS-F-Model. Wenping Ma 0001, Boyou Xue, Mengru Ma, Hekai Zhang, Hao Zhu 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Dense-Weak Ship Detection Based on Foreground-Guided Background Generation Network in SAR ImagesabstractCurrently, ship detection based on Synthetic Aperture Radar (SAR) images still faces significant challenges, particularly in detecting weak and densely distributed ships within complex backgrounds. In areas such as ports and land, the complex background features often resemble those of densely distributed ships, leading to reduced detection accuracy. Additionally, the overlapping and mutual interference of features among dense ships can cause the network to miss detections or produce false positives. Therefore, this paper proposes a Foreground-Guided Background Generation Network (FGBG-Net), which includes a Gaussian Foreground Localization (GFL) model and a Background Feature Removal (BFR) module. The GFL module identifies the approximate high-probability regions of ship foregrounds on the feature map, guiding the network to focus on these regions. The BFR module then progressively removes background interference features based on the positions provided by the GFL module, generating feature maps that are more suitable for detecting weak and dense ships. Our network has been validated on multiple SAR ship datasets, and the experimental results demonstrate noticeable performance improvements, with a mean Average Precision (mAP) increase of 3.4% on the SSDD and HRSID datasets. The relevant code is available at the following link: https://github.com/Xidian-AIGroup190726/FBGBNet/tree/master. Wenping Ma 0001, Xiaoting Yang, Hao Zhu 0009, Xiaoteng Wang, Biao Hou, Mengru Ma, Yue Wu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | A Mamba-Aware Spatial-Spectral Cross-Modal Network for Remote Sensing ClassificationabstractThis study introduces a novel cross-modal spatial-spectral interaction Mamba (CMS2I-Mamba) for remote sensing image fusion classification. Unlike convolution-based models focusing on local details and Transformer-based models with high computational complexity, CMS2I-Mamba efficiently models global long-range dependencies in a linear complexity manner. First, multispectral (MS) and panchromatic (PAN) images each have unique advantages in the spectral and spatial attributes. Given this, this paper innovatively designs the multi-path selective-scan mechanism (MPS2M), which applies different path scanning strategies to deeply capture the global features from both spectral and spatial dimensions, aiming to enhance the robustness and complementarity of spatial-spectral features. Secondly, to overcome the characterization differences between images acquired by different sensors, this paper further introduces the channel interaction alignment module (CIAM). This module employs efficient former-last and oddeven channel interaction strategies to achieve precise semantic alignment of deep features between modalities. Finally, to leverage the shared fusion features to guide the unique singular features, this paper proposes a semantic-aware calibration module (SACM), which accurately constraints and calibrates the same semantic information in deep features. This not only enhances the model’s ability to understand scene semantics, but also promotes the deep fusion and utilization of information between different modalities. Through experimental verification on multiple datasets, the CMS2I-Mamba proposed in this paper shows excellent recognition performance and computational efficiency (parameter quantity and running speed) in fusion classification tasks. The code for CMS2I-Mamba is available at: https://github.com/ru-willow/CMSI-Mamba. Mengru Ma, Jiaxuan Zhao, Wenping Ma 0001, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | A 3D Self-Awareness Diffusion Network for Multimodal ClassificationabstractAs imaging sensor technology in remote sensing has advanced quickly, multimodal fusion classification has become an important research direction in land cover and urban planning classification tasks. While generative models and image classification have greatly benefited from diffusion models, the present ones primarily concentrate on single-modality-driven diffusion processes. Therefore, this paper presents a 3D self-awareness diffusion network (3DSA-DiffNet) for multispectral (MS) and panchromatic (PAN) image fusion classification, which would make it easier to classify heterogeneous data from various sensors. First, in order to model the relationship between multi-channel spectra and multi-pixel spatial distributions as well as samples, respectively, a spatial-spectral joint denoising network (S$^{2}$JD-Net) is proposed. It can incorporate the diffusion process into the neural network to enhance the quality of diffusion features. Secondly, to imitate the brain's spatial-spectral coexistence learning mechanism, this work offers a 3D self-awareness module (3DSA-Module) that can learn the weight of each pixel in 3D space, resulting in extraordinarily high feature representation capabilities. Finally, experimental verification demonstrates that the 3D self-awareness diffusion fusion network driven by brain inspiration outperforms more sophisticated approaches on the Xi'an, Huhhot, and Muufl datasets. Mengru Ma, Wenping Ma 0001, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Fang Liu 0001, Shuyuan Yang 0001, Yuwei Guo 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Brain-Inspired Learning, Perception, and Cognition: A Comprehensive ReviewabstractThe progress of brain cognition and learning mechanisms has provided new inspiration for the next generation of artificial intelligence (AI) and provided the biological basis for the establishment of new models and methods. Brain science can effectively improve the intelligence of existing models and systems. Compared with other reviews, this article provides a comprehensive review of brain-inspired deep learning algorithms for learning, perception, and cognition from microscopic, mesoscopic, macroscopic, and super-macroscopic perspectives. First, this article introduces the brain cognition mechanism. Then, it summarizes the existing studies on brain-inspired learning and modeling from the perspectives of neural structure, cognitive module, learning mechanism, and behavioral characteristics. Next, this article introduces the potential learning directions of brain-inspired learning from four aspects: perception, cognition, understanding, and decision-making. Finally, the top-ten open problems that brain-inspired learning, perception, and cognition currently face are summarized, and the next generation of AI technology has been prospected. This work intends to provide a quick overview of the research on brain-inspired AI algorithms and to motivate future research by illuminating the latest developments in brain science. Licheng Jiao, Mengru Ma, Pei He, Xueli Geng, Xu Liu 0006, Fang Liu 0001, Wenping Ma 0001, Shuyuan Yang 0001, Biao Hou, Xu Tang 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | MBSI-Net: Multimodal Balanced Self-Learning Interaction Network for Image ClassificationabstractA growing number of earth observation satellites are able to simultaneously gather multimodal images of the same area due to the expanding availability and resolution of satellite remote sensing data. This paper proposes a novel multimodal balanced self-learning interaction network (MBSI-Net) for the classification task. It involves a dual-branch teacher-student network that enables knowledge interaction and transfer between the multimodalities. Firstly, in order to introduce statistical information in addition to local and global structural information, a texture feature equalization module (TFE-Module) is proposed. This can enhance the texture information of features through histogram equalization and further improve the representation ability of features. Secondly, to enable the student network to provide timely feedback questions, the paper proposes a feature fusion module (F2-Module) that models and enhances teacher features through the student network. This helps to raise the classification’s accuracy by incorporating information from multimodal images. Finally, the paper proposes a loss function based on structural similarity analysis to ensure balanced self-learning between the student and the teacher networks. Taking the multispectral (MS) and the panchromatic (PAN) images of the same scene as examples, through experimental verification, the proposed method can achieve good results on multiple datasets compared with other methods. Therefore, it offers an effective method for classifying and fusing multimodal data. Mengru Ma, Wenping Ma 0001, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Shuyuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Knowledge Guided Evolutionary Transformer for Remote Sensing Scene ClassificationabstractSolving the complex challenges of sophisticated terrain and multi-scale targets in remote sensing (RS) images requires a synergistic combination of Transformer and convolutional neural network (CNN). However, crafting effective CNN architectures remains a major challenge. To address these difficulties, this study introduces the knowledge guided evolutionary Transformer for RS scene classification (Evo RSFormer). It amalgamates adaptive evolutionary CNN (Evo CNN) with Transformers in a hybrid strategy synergistically, which combines fine-grained local feature extraction of CNNs with long-range contextual dependency modeling of Transformers. Furthermore, for the development of Evo CNN blocks, this paper presents a knowledge-guided adaptive efficient multi-objective evolutionary neural architecture search (MOE2-NAS) strategy. This approach markedly diminishes the labor-intensive characteristics associated with traditional CNN design, striking a balance for both accuracy and compactness. Additionally, by leveraging domain knowledge from natural scene analysis into the RS field, MOE2-NAS facilitates the efficiency of classical NAS. It utilizes a priori knowledge to generate promising initial solutions and constructs a surrogate model for efficient search. The effectiveness of the proposed Evo RSFormer has been rigorously tested on various benchmark RS datasets, including UC Merced, NWPU45, and AID. Empirical results strongly support the superiority of Evo RSFormer over existing methods. Furthermore, experiments on MOE2-NAS have been studied to confirm the important role of knowledge guidance in improving the efficiency of NAS. Jiaxuan Zhao, Licheng Jiao, Chao Wang 0099, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Mengru Ma, Shuyuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | ISSP-Net: An Interactive Spatial-Spectral Perception Network for Multimodal ClassificationabstractCoordinated and complementary spatial-spectral information is represented by the panchromatic (PAN) and multispectral (MS) images. The optimal utilization of the advantages of these images has become a subject of intense research interest. This article introduces the interactive spatial-spectral perception network (ISSP-Net) for multimodal remote sensing image classification, addressing the challenge of optimal utilization of complementary information from PAN and MS images. First, the pixel-guided spatial enhancement module (PGSE-Module) improves spatial location interaction using the spatial location enhancement learning strategy (SLEL-Strategy) and the cross-spatial aggregation learning strategy (CSAL-Strategy), integrating multiscale contextual information and emphasizing pixel-level features. Second, the time-frequency collaborative spectral enhancement module (TFCSE-Module) distinguishes useful frequency domain features through channel separation, lightweight convolutions, and adaptive Fourier transform learning. This approach enables comprehensive utilization of both primary and auxiliary information from multimodal data. Finally, experiments on four datasets demonstrate the ISSP-Net’s state-of-the-art performance in classifying MS and PAN images, with good generalization to hyperspectral (HS) and LiDAR data. The code is provided at:https://github.com/sun740936222/ISSP-Net. Wenping Ma 0001, Hekai Zhang, Mengru Ma, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | An Adaptive Migration Collaborative Network for Multimodal Image ClassificationabstractThe multispectral (MS) and the panchromatic (PAN) images belong to different modalities with specific advantageous properties. Therefore, there is a large representation gap between them. Moreover, the features extracted independently by the two branches belong to different feature spaces, which is not conducive to the subsequent collaborative classification. At the same time, different layers also have different representation capabilities for objects with large size differences. In order to dynamically and adaptively transfer the dominant attributes, reduce the gap between them, find the best shared layer representation, and fuse the features of different representation capabilities, this article proposes an adaptive migration collaborative network (AMC-Net) for multimodal remote-sensing (RS) images classification. First, for the input of the network, we combine principal component analysis (PCA) and nonsubsampled contourlet transformation (NSCT) to migrate the advantageous attributes of the PAN and the MS images to each other. This not only improves the quality of images themselves, but also increases the similarity between the two images, thereby reducing the representational gap between them and the pressure on the subsequent classification network. Second, for the interaction on the feature migrate branch, we design a feature progressive migration fusion unit (FPMF-Unit) based on the adaptive cross-stitch unit of correlation coefficient analysis (CCA), which can make the network automatically learn the features that need to be shared and migrated, aiming to find the best shared-layer representation for multifeature learning. And we design an adaptive layer fusion mechanism module (ALFM-Module), which can adaptively fuse features of different layers, aiming to clearly model the dependencies among multiple layers for different sized objects. Finally, for the output of the network, we add the calculation of the correlation coefficient to the loss function, which can make the network converge to the global optimum as much as possible. The experimental results indicate that AMC-Net can achieve competitive performance. And the code for the network framework is available at: https://github.com/ru-willow/A-AFM-ResNet. Wenping Ma 0001, Mengru Ma, Licheng Jiao, Fang Liu 0001, Hao Zhu 0009, Xu Liu 0006, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Transfer Representation Learning Meets Multimodal Fusion Classification for Remote Sensing ImagesabstractTo maximize the complementary advantages of synergistic multimodal, a transfer representation learning fusion network (TRLF-Net) is proposed for multisource remote sensing images collaborative classification in this article. First, with respect to the feature encoding, we design a dual-branch attention sparse transfer module (DAST-Module), which combines the spatial and channel attention (CA) masks to migrate the advantage attributes of the panchromatic (PAN) and the MS images mutually. This not only enhances their respective image advantages but also facilitates the sparse fusion of low-level features. Second, for the separation of multiscale information, a deep dual-scale decomposition module (DDSD-Module) is designed, which allows the decompose of high-frequency and low-frequency components. Then it uses the decomposed information to make the essential difference as small as possible, and the surrounding contour difference is as large as possible of the complementary multimodal image through the design of the loss function. Finally, to address the problem of large intraclass and small interclass differences, we develop a representation fusion of the global and local features’ module (RFGAL-Module). It mainly adopts global features to sort local features within classes, and then outputs them in a cascade. Thus, the characterization ability of features is improved, and the global and local features are used in a coordinated manner to accomplish the sample classification tasks. In particular, the experimental results demonstrate that TRLF-Net can obtain much improved accuracy and efficiency. The code is accessible in:https://github.com/ru-willow/SRLF-Net. Mengru Ma, Wenping Ma 0001, Licheng Jiao, Xu Liu 0006, Fang Liu 0001, Lingling Li 0002, Shuyuan Yang 0001, Biao Hou |
IEEE Trans. Geosci. Remote. Sens. | 1 |