Chenzhong Gao

dblp:304/0235 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0003-4674-0613ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 HIMO: Cross-Arbitrary-Modality Image Invariant Feature Transform With Hierarchical Intrinsic Major Orientation
abstract
Invariant feature extraction is a critical challenge in intelligent image processing, particularly with the rapid advancement of multi-source/modal imaging. Cross-modal matching has attracted considerable attention, yet current studies primarily focus on targeted modalities rather than realizing a general approach. In this paper, cross-arbitrary-modal image invariant feature extraction and matching is studied. Inspired by human vision, a purely handcrafted invariant feature transform is proposed for universal cross-modal image matching, named Hierarchical Intrinsic Major Orientation (HIMO). Based on orientation information, a full-chain non-data-driven algorithm is designed that hinges on an Intrinsic Major Orientation (IMO) extraction. The HIMO incorporates a novel keypoint detector utilizing Difference-of-Feature Suppression (DoFS), a Polar-Pyramid descriptor (PolarP), and a Cascaded Dynamic Multi-scale Strategy (CDMS) to effectively address common challenges such as intensity distortion, rotation, scale differences, geometric deformation, and image noise. To validate the proposed method, two massive cross-modal datasets-General Cross-modal Zone (GCZ) and Wide-area Diverse Sources (WDS)-are introduced, alongside two practical evaluation metrics. Comprehensive experiments compared with 10 traditional and 15 deep-learning state-of-the-art algorithms on 5 datasets fully demonstrate that the proposed HIMO achieves superior performance in terms of robustness, stability, and generalization across diverse imaging conditions.
Chenzhong Gao, Wei Li 0032, Desheng Weng, Ran Tao 0003, Xiang-Gen Xia 0001, Qian Du 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 HOMO-Feature: Cross-Arbitrary-Modal Image Matching with Homomorphism of Organized Major Orientation
Chenzhong Gao, Wei Li 0032, Desheng Weng
ICCV1
2025 SwinMatcher: Universal Cross-Modal Remote Sensing Image Matching With Interactive Swin Transformer
abstract
Cross-modal remote sensing image matching serves as a key technique for collaborative utilization of multi-source information. However, modal differences and geometric distortions between multi-source images pose challenges to existing methods in terms of robustness and generalization. To achieve feature interaction in cross-modal scenes, this paper proposes SwinMatcher, an end-to-end matching model based on the Transformer architecture. Innovatively proposing the window/shifted-window cross-attention based on the window/shifted-window self-attention mechanisms of Swin Transformer, SwinMatcher enables efficient cross-modal feature interaction and multi-scale contextual modeling. It also incorporates a learnable matching module to directly generate semi-dense correspondences. Moreover, a cross-modal remote sensing image matching dataset is generated, which encompasses four modalities: visible light, synthetic aperture radar (SAR), light detection and ranging (LiDAR), and map, distributed across four representative scenes. The dataset includes 400 samples produced via random homography transformations, designed to enhance modal diversity and scene complexity. Experiments demonstrate that SwinMatcher outperforms state-of-the-art methods on this new dataset as well as public benchmarks, exhibiting superior robustness under complex scenes involving coupled modal and geometric distortions. The proposed method and dataset provide novel solutions and evaluation benchmarks for cross-modal remote sensing image matching. The code and testing dataset will be made publicly available at https://github.com/LotrL/SwinMatcher.
Wei Li 0032, Desheng Weng, Chenzhong Gao, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 DGIM: Cascaded Dynamic Data Generation for Robust Cross-Modal Image Matching
abstract
To address the challenges posed by extensive modality differences in image matching tasks, this paper proposes a cascaded learning framework. It guides the optimization of an end-to-end matching model via a dynamic data engine, which can provide sufficient cross-modal training data to support the model’s full adaptation to cross-modal features. The data engine integrates a random homography transformation module and a lightweight image generation model, enabling the online synthesis of cross-modal image pairs with geometric variations and diverse styles. This provides the matching model with rich cross-modal stimulation. The matching model adopts a hybrid architecture combining a CNN backbone and Transformer attention mechanisms, which integrates multi-scale local feature extraction with global context modeling. By adopting the proposed stepwise aggregation strategy, the efficiency of feature extraction is well ensured. Subsequently, a coarse-to-fine matching strategy is employed to achieve high accuracy and robustness of feature alignment. Comprehensive experiments on both self-collected and public cross-modal image matching datasets demonstrate that the proposed DGIM outperforms existing state-of-the-art approaches in cross-modal matching performance while achieving a good balance between efficiency and effectiveness. It also exhibits broad practical potential across multiple fields and scenes. This work provides novel solutions and evaluation benchmarks for cross-modal image matching tasks. The code and testing dataset will be made publicly available at https://github.com/LotrL/DGIM.
Desheng Weng, Wei Li 0032, Chenzhong Gao, Xiang-Gen Xia 0001, Zhicheng Shi, Bolun Cui
IEEE Trans. Geosci. Remote. Sens.3
2024 A New Invariant Feature for Multi-Modal Images Matching
abstract
This paper aims at providing an effective multi-modal images invariant feature extraction and matching algorithm for the application of multi-source data analysis. Focusing on the differences and correlation of multi-modal images, a feature-based matching algorithm is implemented. The key technologies include phase congruency (PC) and Shi-Tomasi feature point for keypoints detection, LogGabor filter and a weighted partial main orientation map (WPMOM) for feature extraction, and a multi-scale process to deal with scale differences and optimize matching results. The experimental results on practical data from multiple sources prove that the algorithm has effective performances on multi-modal images, which achieves accurate spatial alignment, showing practical application value and good generalization. The codes are available at https://github.com/MrPingQi.
Chenzhong Gao, Wei Li 0032, Yute Li
IGARSS1
2024 Multi-Temporal Images Generation for Building Change Detection Performance Promotion
abstract
The changes in building are important basis for urban monitoring. However, due to the rarity and sparsity of the occurrence of changes in buildings, collecting effective bitemporal image pairs is challenging, as it requires long-term observation over several months or even years. Additionally, annotating large-scale change detection datasets is time-consuming and labor-intensive. Consequently, data scarcity issues lead to insufficient training of building change detection models. To address this, we propose a data generation method Building Generation GAN (BG-GAN). Different from other GANs, the BG-GAN is trained based on adversarial consistency loss, enabling the model to generate new bi-temporal image pairs with various types of building changes. To verify the effectiveness of the proposed methods on change detection task, BG-GAN is utilized to perform building change samples generation on two building change detection datasets (LEVIR-CD and WHU-CD). The experimental results demonstrate that the proposed method can improve the robustness and generalization of change detection model to detect pseudo changes.
Yute Li, Wei Li 0032, Nan Wang 0038, Chenzhong Gao, Yin Zhuang, He Chen 0004
IGARSS4
2023 Hyperspectral pathology image classification using dimension-driven multi-path attention residual network
Xueyu Zhang, Wei Li 0032, Chenzhong Gao, Yue Yang 0041, Kan Chang
Expert Syst. Appl.3
2023 Central Attention Network for Hyperspectral Imagery Classification
abstract
In this article, the intrinsic properties of hyperspectral imagery (HSI) are analyzed, and two principles for spectral-spatial feature extraction of HSI are built, including the foundation of pixel-level HSI classification and the definition of spatial information. Based on the two principles, scaled dot-product central attention (SDPCA) tailored for HSI is designed to extract spectral-spatial information from a central pixel (i.e., a query pixel to be classified) and pixels that are similar to the central pixel on an HSI patch. Then, employed with the HSI-tailored SDPCA module, a central attention network (CAN) is proposed by combining HSI-tailored dense connections of the features of the hidden layers and the spectral information of the query pixel. MiniCAN as a simplified version of CAN is also investigated. Superior classification performance of CAN and miniCAN on three datasets of different scenarios demonstrates their effectiveness and benefits compared with state-of-the-art methods.
Huan Liu 0015, Wei Li 0032, Xiang-Gen Xia 0001, Mengmeng Zhang 0005, Chenzhong Gao, Ran Tao 0003
IEEE Trans. Neural Networks Learn. Syst.5
2022 MS-HLMO: Multiscale Histogram of Local Main Orientation for Remote Sensing Image Registration
abstract
Multi-source image registration is challenging due to intensity, rotation, and scale differences among the images. Considering the characteristics and differences of multi-source remote sensing images, a feature-based registration algorithm named Multi-scale Histogram of Local Main Orientation (MS-HLMO) is proposed. Harris corner detection is first adopted to generate feature points. The HLMO feature of each Harris feature point is extracted on a Partial Main Orientation Map (PMOM) with a Generalized Gradient Location and Orientation Histogram-like (GGLOH) feature descriptor, which provides high intensity, rotation, and scale invariance. The feature points are matched through a multi-scale matching strategy. Comprehensive experiments on 17 multi-source remote sensing scenes demonstrate that the proposed MS-HLMO and its simplified version MS-HLMO+outperform other competitive registration algorithms in terms of effectiveness and generalization.
Chenzhong Gao, Wei Li 0032, Ran Tao 0003, Qian Du 0001
IEEE Trans. Geosci. Remote. Sens.1
2021 Multi-Scale HARRIS-PIIFD Features for Registration of Visible and Infrared Images
abstract
This paper aims at providing visible and infrared images registered in geometric space for image fusion. Focusing on the characteristics and differences of visible and infrared images, a feature-based registration algorithm is implemented. The key technologies include image scale-space for implementing multi -scale properties, Harris corner detection for keypoints extraction, and partial intensity invariant feature descriptor (PIIFD) for keypoints description. Eventually, a multi-scale Harris- Piifd image registration algorithm framework is proposed. The experimental results of four sets of representative real data show that the algorithm has excellent, stable performance in visible and infrared image registration, and can achieve accurate spatial alignment, which has strong practical application value and certain generalization ability.
Chenzhong Gao, Wei Li 0032
IGARSS1