Guixin Zhao

dblp:287/8407 · DBLP profile ↗
← Back
20ranked-venue papers
0as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Multimodal Multi-Graph Fusion Learning for Alzheimer's Disease Diagnosis
abstract
Alzheimer's Disease (AD) is a prevalent and severe neurodegenerative disorder, and early diagnosis is essential for managing disease progression. Recently, multimodal graph learning has demonstrated significant potential in integrating both medical imaging and non-imaging data, as well as uncovering relationships between patients. However, the high-dimensional nature of multimodal medical data poses significant challenges for constructing and learning modality graph structures. Moreover, existing methods are often imprecise in modeling graph structures for continuous data. To address these issues, this paper introduces a novel multimodal multi-graph fusion learning method for Alzheimer's disease diagnosis. Specifically, multimodal state space networks (multimodal SSNs) are proposed to capture the dependencies between multimodal and high-dimensional features. Furthermore, a novel graph structure learning (KGSL) based on an initial K-nearest neighbors graph is proposed to separately construct graph structures for each modality. This method is particularly suitable for modeling the graph structures of Euclidean data. Finally, multimodal graph fusion integrates various modal graph structures into a single graph, leading to enhanced multimodal integration. In addition, this paper uses a learnable Chebyshev Graph Convolutional Network for the classification network, which enables end-to-end optimization. Experimental results demonstrate that our approach achieves excellent performance on public datasets.
Aimei Dong, Yongxing Cai, Guohua Lv, Guixin Zhao
IEEE Trans. Multim.6
2025 TDMF: Text-Guided Denoising and Interactive Medical Image Fusion
abstract
Multimodal image fusion aims to merge features from different modalities to create a comprehensively representative image. However, existing medical image fusion methods often struggle to handle noise generated during image acquisition, significantly diminishing their impact on visual quality. To address these challenges, we propose a semantically text-guided medical image fusion model, named TDMF. Specifically, TDMF guides classical image fusion through textual semantics and effectively coordinates the resolution of degradation and interaction issues during the fusion process. By integrating text encoders and interactive fusion modules, TDMF establishes a unified framework for denoising and interactive fusion of medical images. Extensive experiments have demonstrated that our proposed text-guided image fusion strategy offers significant advantages over state-of-the-art methods in medical image fusion performance.
Aimei Dong, Guohua Lv, Guixin Zhao, Jinyong Cheng
ICASSP5
2025 SSFSL: Self-Supervised and Few-Shot Learning for Cross-Domain Hyperspectral Image Classification
abstract
Few-shot learning (FSL) has gained increasing attention in hyperspectral image (HSI) classification due to its ability to perform cross-domain classification with minimal labeled samples. However, existing FSL methods overlook the continuity of HSI spectral sequences and fail to utilize the large amount of unlabeled samples in the target domain. To address these issues, we introduce a novel cross-domain HSI classification method that combines self-supervised learning with FSL (SSFSL). This approach uses self-supervised learning and FSL to extract transferable knowledge from the source domain and introduces an adaptive soft label generation algorithm to leverage unlabeled samples in the target domain. Compared to existing cross-domain FSL classification methods, the proposed approach considers the spectral sequence continuity of HSI and effectively extracts useful information from unlabeled samples in the target domain. Extensive experiments conducted on three datasets demonstrate that SSFSL outperforms state-of-the-art methods in both quantitative and qualitative aspects.
Guohua Lv, Qiang Chi, Guixin Zhao, Aimei Dong, Wei Li 0032
ICASSP4
2025 CGNet: Classification-Guided Multi-Task Interactive Network for Hyperspectral and Multispectral Image Fusion
abstract
The goal of fusing hyperspectral images (HSI) and multispectral images (MSI) is to generate high-resolution hyperspectral images for downstream tasks. However, most existing methods overlook the specific requirements of these tasks, leading to a gap between the fusion process and its subsequent applications due to insufficient guidance from downstream tasks. To address this issue, we propose a classification-guided multitask interactive network (CGNet) that integrates both fusion and classification tasks into a unified framework, with two branches producing the fused image and classification results, respectively. In the fusion branch, we design a multi-level residual refinement module to efficiently integrate spatial and spectral information. Additionally, an attention-based multi-scale fusion module, incorporating both spatial and channel attention, is carefully crafted to enhance representation learning. In the classification branch, both 2-D and 3-D convolutions are employed to improve classification performance. Moreover, an information interaction module is proposed to guide the fusion task based on classification outcomes. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches on the Pavia Centre and Pavia University datasets.
Guohua Lv, Yanlong Xu, Yongbiao Gao, Guixin Zhao, Xiangcheng Sun
ICASSP4
2025 A Grouping Strategy-Based Progressive Fusion Network for Hyperspectral Image Super-Resolution
abstract
Hyperspectral super-resolution involves combining low-resolution hyperspectral images with high-resolution multispectral images to produce a high-resolution hyperspectral image. Recently, although many methods for hyperspectral image super-resolution have been proposed, they often fail to fully utilize the high similarity among adjacent bands to enhance fusion performance. Therefore, we propose a grouping strategy-based progressive fusion network (GPFNet) for hyperspectral super-resolution. The core of GPFNet is the grouping strategy fusion block (GPF block), in which grouping-based spatial-spectral information fusion and spatial information refinement are performed. We design the spatial-spectral information fusion module (SSIFM) based on grouped convolutions to capture the feature differences from adjacent bands. To refine spatial details, we develop the spatial information enhancement module (SpaEM), which leverages the hierarchical features extracted by the multi-scale feature extraction module (MIEM). Additionally, a progressive fusion strategy, which involves using multiple upsampled hyperspectral images and downsampled multispectral images, further preserves spectral integrity and spatial details. Extensive experiments show that GPFNet outperforms state-of-the-art methods both qualitatively and quantitatively.
Guohua Lv, Baodong Zhang, Yongbiao Gao, Guixin Zhao, Juncan Wang
ICASSP4
2025 A Mutually Enhancement Network for Superpixel Segmentation and Classification of Hyperspectral Image
abstract
Most existing hyperspectral image (HSI) classification methods primarily focus on capturing subtle spectral variations by leveraging local spectral-spatial cues derived from patch-level representations. However, limited attention has been given to exploring the global spatial contextual correlations among pixels of HSI. In this study, we propose the Superpixel Segmentation and Classification Mutual Enhancement Network (S2CMEN), a novel framework that integrates global spatial correlations with spectral information through the mutual enhancement of superpixel segmentation and classification. Specifically, a global spatial adaptive module (GSAM) is designed to obtain the direct correlation of the global classes in HSI. It consists of an Adaptive Spectral-Superpixel Network (ASSN) and a Graph Convolutional Network (GCN), forming a synergistic architecture that effectively captures global spatial relationships by adaptively deriving superpixel results from HSIs. Notably, GSAM offers a transferable global spatial representation for HSI tasks, enabling integration with other spectral feature extraction models. Furthermore, we develop a Spatial-Spectral Fusion Module (SSFM) to obtain comprehensive spectral features and fuse them with the extracted global spatial features. Finally, under the constraint of a unit loss, the Mutual Enhancement Strategy (MES) can make the superpixel segmentation loss and the classification loss mutually enhance each other for better performance. We conducted extensive experiments on three public datasets. The proposed S2CMEN achieves overall classification accuracies of 97.38%, 92.33%, and 91.38% on Indian Pines, Pavia University, and Houston, respectively, consistently surpassing existing state-of-the-art methods.
Mengxin Cao, Yongmin Li 0001, Xu Zhang 0039, Guixin Zhao, Guohua Lv, Aimei Dong, Jinyong Cheng, Wei Li 0032, Xiangjun Dong 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Cross-Domain Hyperspectral Image Classification via Mamba-CNN and Knowledge Distillation
abstract
Domain adaptation (DA)-based cross-domain hyperspectral image (HSI) classification methods have garnered significant attention. The majority of DA techniques utilize models based on convolutional neural networks (CNNs) and Transformers for feature extraction. However, Transformers may struggle to capture local details in HSIs, while CNNs often underperform in handling long-range dependencies. Furthermore, many methods focus only on aligning marginal distributions while ignoring the consistency of inter-class features, which may lead to feature confusion and degraded classification accuracy. To overcome the challenges mentioned, we propose a Mamba-CNN and knowledge distillation network (MKDnet). Firstly, the network employs a feature extractor that integrates Mamba and CNN frameworks for cross-domain HSI classification, enabling the capture of both global and local features while effectively capturing long-range dependencies. Secondly, domain alignment is achieved through distribution alignment and graph alignment. In the distribution alignment phase, we design a knowledge distillation architecture that utilizes soft labels to enhance the understanding of relationships between classes, thereby improving the consistency of inter-class features. In the graph alignment phase, we use graph convolution to capture connections between nodes and edges and transfer class-level topological relationships across domains. Finally, the classifier is used to obtain classification results, with consistency constraints applied to balance features between classes more effectively. Extensive experiments have demonstrated that MKDnet outperforms other state-of-the-art methods on three public cross-domain HSI datasets.
Aoyan Du, Guixin Zhao, Mengxin Cao, Aimei Dong, Guohua Lv, Yongbiao Gao, Xiangjun Dong 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Hop-Gated Graph Attention Network for ASD Diagnosis via PC-Based Graph Regularization Sparse Representation
Aimei Dong, Xuening Zhang, Guixin Zhao
ICANN (8)3
2024 Multi-scale Convolutional Attention Fuzzy Broad Network for Few-Shot Hyperspectral Image Classification
Xiaopei Hu, Guixin Zhao, Xiangjun Dong 0001, Aimei Dong
ICANN (2)2
2024 Rafmnet: Reinforced Attention Fusion and Multiscale Network For Noisy Infrared and Visible Image Fusion
abstract
The purpose of infrared and visible image fusion is to combine the advantages of different types of images to produce more robust and informative images. However, if the source images are noisy, existing fusion methods may not produce clear results. To address this issue, we propose a novel method for infrared and visible image fusion with noise reduction. This method enhances the visual perception of fused images by integrating features of different scales extracted by the denoising network into the fusion network. By using deformable convolutional denoising networks, noise in images can be removed and features can be enhanced. Then, a set of reinforced attention fusion modules (RAFM) are designed to fuse the features extracted by the denoising network. Experimental results demonstrate the effectiveness of our proposed method, which outperforms existing state-of-the-art methods in terms of fusion accuracy and visual perception.
Guohua Lv, Xiyan Wang, Yongbiao Gao, Yi Zhai 0003, Guixin Zhao, Guangxiao Ma
ICIP5
2024 TLLFusion: An End-to-End Transformer-Based Method for Low-Light Infrared and Visible Image Fusion
Guohua Lv, Xinyue Fu, Yi Zhai 0003, Guixin Zhao, Yongbiao Gao
PRCV (3)4
2024 MFIFusion: An infrared and visible image enhanced fusion network based on multi-level feature injection
Aimei Dong, Guohua Lv, Guixin Zhao, Jinyong Cheng
Pattern Recognit.5
2024 Co-Enhancement of Multi-Modality Image Fusion and Object Detection via Feature Adaptation
abstract
The integration of multi-modality images significantly enhances the clarity of critical details for object detection. Valuable semantic data from object detection enriches the fusion process of these images. However, the potential reciprocal relationship that could enhance their mutual performance remains largely unexplored and underutilized, despite some semantic-driven fusion methodologies catering to specific application needs. To address these limitations, this study proposes a mutually reinforcing, dual-task-driven fusion architecture. Specifically, our design integrates a feature-adaptive interlinking module into both image fusion and object detection components, effectively managing the inherent feature discrepancies. The core idea is to channel distinct features from both tasks into a unified feature space after feature transformation. We then design a feature-adaptive selection module to generate features rich in target semantic information and compatible with the fusion network. Finally, effective combination and mutual enhancement of the two tasks are achieved through an alternating training process. A diverse range of swift evaluations is performed across various datasets to corroborate the potential efficiency of our framework, actualizing visible advancements in both fusion effectiveness and detection accuracy.
Aimei Dong, Guixin Zhao, Yi Zhai 0003, Guohua Lv, Jinyong Cheng
IEEE Trans. Circuits Syst. Video Technol.5
2024 Spatial-Spectral-Semantic Cross-Domain Few-Shot Learning for Hyperspectral Image Classification
abstract
Preprocessing procedures are commonly employed to reduce water-absorption bands and noise in hyperspectral images (HSIs). Nevertheless, they typically do not entirely eradicate noise. This is especially evident in scenarios that necessitate data of exceptional quality, such as cross-domain few-shot classification tasks. Within these specific conditions, the influence of remaining background noise on the ultimate results of classification is substantial. Furthermore, the presence of sample selection biases in the few-shot task might lead to the emergence of false statistical correlations between data from distinct domains, resulting in a decrease in the model’s ability to generalize. We propose a new method called spatial-spectral–semantic cross-domain few-shot learning (S3CFSL) to address the challenge. This method promotes the learning of transferable information by incorporating feature denoising operations in the feature extraction process to restore essential information. Concurrently, it enhances cross-domain distributional consistency by introducing a semantic-aware strategy to strengthen the association between cross-domain data and semantic information. Specifically, the spatial and spectral dual channels (SSDCs), in conjunction with the cross-spatial-spectral transformer (CSST), are designed as a feature extractor to acquire interactive spatial-spectral features. The feature-denoising operations can further acquiring transferable information from cross-domain features, thus facilitating meta-learning in both the source domain (SD) and the target domain (TD). Meanwhile, a semantic-enhanced domain alignment (SEDA) is designed to promote domain adaptation by using a semantic-aware strategy, which significantly enhances distributional consistency for cross-domain tasks. Our results exhibit exceptional classification efficacy in comparison to other state-of-the-art approaches on three public HSI datasets.
Mengxin Cao, Xu Zhang 0039, Jinyong Cheng, Guixin Zhao, Wei Li 0032, Xiangjun Dong 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Spectral-Spatial-Language Fusion Network for Hyperspectral, LiDAR, and Text Data Classification
abstract
The fusion classification of hyperspectral image (HSI) and light detection and ranging (LiDAR) data has gained widespread attention because of its ability to obtain more comprehensive spatial and spectral information. However, the heterogeneous gap between HSI and LiDAR data also adversely affects the classification performance. Despite the excellent performance of traditional multimodal fusion classification models, language information containing much linguistic priori knowledge to enrich visual representations needs to be addressed. Therefore, we design a Spectral-Spatial-Language fusion network (S2LFNet), which can fuse visual and language features to broaden the semantic space using linguistic priori knowledge commonly shared between spectral features and spatial features. First, we propose a dual-channel cascaded image fusion encoder (DCIFencoder) for visual feature extraction and progressive feature fusion of different levels for HSI and LiDAR data. Then, three aspects of Text data are designed to extract linguistic priori knowledge using the Text encoder. Finally, contrastive learning is utilized to construct a unified semantic space, and Spectral-Spatial-Language fusion features are obtained for classification tasks. We evaluate the classification performance of the proposed S2LFNet on three datasets through extensive experiments, and the results show that it outperforms the state-of-the-art fusion classification methods.
Mengxin Cao, Guixin Zhao, Guohua Lv, Aimei Dong, Ying Guo 0030, Xiangjun Dong 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Few-Shot Hyperspectral Image Classification Based on Cross-Domain Spectral Semantic Relation Transformer
abstract
In practical hyperspectral image (HSI) classification tasks, we often encounter the problems of few-shot classification and domain misalignment between source domains and target domains. To solve this classification paradigm, a meta-learning method of few-shot learning (FSL) is usually used. However, most existing FSL methods address the problem for domain alignment and neglect the exploration of semantic relationships of objects across domains. In this paper, we propose the cross-domain spectral semantic transformer FSL (SSTFSL), which can fully extract semantic features and spectral detail features for the cross-domain few-shot HSI classification task. Specifically, the multi-head self-attention (MSA) mechanism with enhancement process (EP) of the transformer is used to map out semantically relevant local regions and can enhance the ability of the model to distinguish subtle feature differences in the spectrum. In addition, the matching degree of different branches is computed by relational network learning, which ultimately enables cross-domain few-shot HSI classification. Through extensive experiments, we evaluate the classification performance of SSTFSL on HSI datasets. The results demonstrate that SSTFSL outperforms existing FSL methods and deep learning methods on HSI classification.
Mengxin Cao, Guixin Zhao, Aimei Dong, Guohua Lv, Ying Guo 0030, Xiangjun Dong 0001
ICIP2
2023 Mix-Net: Automatic Segmentation of Covid-19 ct Images Based on Parallel Design
abstract
Since the discovery of COVID-19 in late 2019, the viral pneumonia crisis has begun to spread rapidly around the world. Lesion segmentation can remove unnecessary background areas and help doctors diagnose the condition. However, the infected areas showed differences at different stages, and the border between the infected areas and the surrounding tissue was blurred. To solve this problem, a novel COVID-19 lung infection segmentation network (Mix-Net) is designed for the automatic identification of infected areas from chest CT slices. Specifically, first, the local and global features of the infected areas are extracted and interacted with using the mixing block. Then, the features extracted from multiple layers of the encoder are fused and connected to the decoder. Experiments show that Mix-Net outperforms most cutting-edge segmentation models and achieves good segmentation results.
Aimei Dong, Guohua Lv, Guixin Zhao, Yi Zhai 0003
ICIP4
2023 Few-Shot Hyperspectral Image Classification with Spectral-Spatial Feature Fusion Based on Fuzzy Broad Learning System
abstract
In the few-shot hyperspectral image (HSI) classification, most current models don't fully utilize the advantage of spectral-spatial feature fusion, resulting in low classification accuracy. Therefore, we propose a few-shot HSI classification model with spectral-spatial feature fusion based on fuzzy broad learning system (FBLS) (FSFBLS). Firstly, we use a Gaussian filter to suppress noise while smoothing spectral features based on spatial information to achieve the first fusion of spectral-spatial features. Secondly, we use FBLS with fuzzy rules to fully model the complex mapping relationship between spectral-spatial features and HSI labels to complete HSI classification. The fuzzy processing can extract rich discriminative features to enhance the recognition of different categories. Finally, the guided filter corrects the misclassified samples of FBLS based on the guided image to achieve the second fusion of spectral-spatial features. Extensive experimental results on three public datasets demonstrate that FSFBLS achieves state-of-the-art classification performance compared to nine popular models.
Xiaopei Hu, Guixin Zhao, Aimei Dong, Guohua Lv, Yi Zhai 0003, Ying Guo 0030, Xiangjun Dong 0001
ICIP2
2023 BS-YOLOv5s: Insulator Defect Detection with Attention Mechanism and Multi-Scale Fusion
abstract
With the rapid development of deep learning, the use of object detection algorithms for aerial insulator image defect detection has become the main way. To address the problems of low detection accuracy for small targets, weak representation ability of feature maps, insufficient extracted key information, and the shortage of aerial insulator defect datasets, this paper proposes an improved insulator defect detection method named BS-YOLOv5s based on 3-D attention mechanism and Bi-Slim-neck using YOLOv5s as the base network. Additionally, to solve the problem of the shortage of aerial insulator datasets, this paper proposes a new aerial insulator dataset Weather-Insulator (WI) containing a variety of defect scenarios. The experimental results demonstrate that the proposed method not only greatly improves the detection accuracy, but also maintains a high detection speed, satisfying the engineering requirements for insulator defect detection. The dataset and code for this paper are publicly available at https://github.com/jspron/insulator-defect.
Zengbin Zhang, Guohua Lv, Guixin Zhao, Yi Zhai 0003, Jinyong Cheng
ICIP3
2023 Autism Spectrum Disorder Diagnosis Using Graph Neural Network Based on Graph Pooling and Self-adjust Filter
Aimei Dong, Xuening Zhang, Guohua Lv, Guixin Zhao, Yi Zhai 0003
PRCV (13)4