VLDB 2026 Research / reviewers in the wild / expert
Guohua Lv
dblp:02/10772
· DBLP profile ↗
48ranked-venue papers
17as first author
42since 2021 · last 2027
0000-0003-3550-8026ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 12 first-author · 28 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Harmonizing micro textures and macro semantics: A multi-scale perception framework for aerial small object detection
Chongmiao Sun, Guohua Lv, Shaohua Wan 0001, Songtao Ding |
Expert Syst. Appl. | 4 |
| 2026 | Adaptive Momentum and EMA-weighted Modeling for Imbalanced Label Distribution LearningabstractLabel Distribution Learning (LDL) is a groundbreaking paradigm for addressing the task with label ambiguity. Subjectivity in annotating label description degrees often leads to imbalanced label distribution. Existing approaches either adopt representation alignment or decoupling strategies to solve the imbalanced label distribution learning (ILDL). However, representation alignment-based methods overlook the issue of gradient vanishing for non-dominant branches within imbalanced label distributions, while decoupling-based approaches fail to achieve adaptive weight optimization. To address these issues, we propose Adaptive Momentum and Exponential Moving Average weighted modeling (AMEMA). AMEMA combines EMA-based loss weighting with momentum allocation to mitigate gradient attenuation in non-dominant label learning and adaptively balance the optimization signals between dominant and non-dominant branches. It computes and updates Kullback-Leibler divergence losses for each branch using EMA, and applies different initial momenta to facilitate branch-specific optimization dynamics. Dynamic weighting coefficients, derived from EMA-smoothed losses, allow the model to adjust its learning direction adaptively and improve the learning of non-dominant labels. Extensive experiments on benchmark datasets show that AMEMA consistently outperforms state-of-the-art ILDL methods across various evaluation metrics. Yongbiao Gao, Xiangcheng Sun, Guohua Lv |
AAAI | 5 |
| 2026 | Edge-Cloud Collaborated Prototype Graph Network for Efficient Few-Shot Object DetectionabstractWith the rapid development of industrial automation, few-shot object detection has emerged as a promising solution for recognizing novel categories using only limited annotated data. However, existing approaches often suffer from high computational complexity and limited adaptability when deployed in resource-constrained industrial environments. To achieve precise detection, efficiency, and security, this paper proposes a collaborative computing framework based on an Edge-Cloud Dual-Prototype Graph Convolutional Network (EC-DP-GCN) for few-shot object detection with hierarchical knowledge embedding. The framework comprises three key components: a device–edge–cloud architecture, a Positive-Negative Prototype (PNP) module, and a Class-Prototype-Sample Hierarchical Graph (CPS-HG) module. Specifically, the PNP module explicitly models intra-class diversity by constructing discriminative positive and negative prototypes from limited support samples, thereby enhancing prototype representativeness. In addition, we further introduce the CPS-HG module, which treats the dual prototypes as class-based prior knowledge and models the relationships among samples through a hierarchical graph structure encompassing class, prototype, and sample levels. This design effectively expands the semantic margins in the embedding space to improve knowledge-guided detection. Extensive experiments on the PASCAL VOC and MS COCO benchmarks demonstrate that EC-DP-GCN significantly outperforms strong baselines and previous state-of-the-art methods, achieving an average improvement of 1.1% in 10-shot detection scenarios. Yirui Wu, Xinfu Liu 0001, Shaohua Wan 0001, Guohua Lv, Jiehan Zhou, Joel J. P. C. Rodrigues |
IEEE Internet Things J. | 5 |
| 2026 | Bridging the semantic gap in text-video retrieval: A diffusion generation network for enhanced visual representation
Songtao Ding, Chun Mao, Chun Geng, Hongyu Wang 0007, Yaqiong Xing, Guohua Lv, Shaohua Wan 0001 |
Pattern Recognit. | 6 |
| 2026 | Multimodal Multi-Graph Fusion Learning for Alzheimer's Disease DiagnosisabstractAlzheimer's Disease (AD) is a prevalent and severe neurodegenerative disorder, and early diagnosis is essential for managing disease progression. Recently, multimodal graph learning has demonstrated significant potential in integrating both medical imaging and non-imaging data, as well as uncovering relationships between patients. However, the high-dimensional nature of multimodal medical data poses significant challenges for constructing and learning modality graph structures. Moreover, existing methods are often imprecise in modeling graph structures for continuous data. To address these issues, this paper introduces a novel multimodal multi-graph fusion learning method for Alzheimer's disease diagnosis. Specifically, multimodal state space networks (multimodal SSNs) are proposed to capture the dependencies between multimodal and high-dimensional features. Furthermore, a novel graph structure learning (KGSL) based on an initial K-nearest neighbors graph is proposed to separately construct graph structures for each modality. This method is particularly suitable for modeling the graph structures of Euclidean data. Finally, multimodal graph fusion integrates various modal graph structures into a single graph, leading to enhanced multimodal integration. In addition, this paper uses a learnable Chebyshev Graph Convolutional Network for the classification network, which enables end-to-end optimization. Experimental results demonstrate that our approach achieves excellent performance on public datasets. Aimei Dong, Yongxing Cai, Guohua Lv, Guixin Zhao |
IEEE Trans. Multim. | 5 |
| 2025 | Reinforced Label Denoising for Weakly-Supervised Audio-Visual Video Parsing
Yongbiao Gao, Xiangcheng Sun, Guohua Lv, Deng Yu, Sijiu Niu |
CVM (3) | 3 |
| 2025 | TDMF: Text-Guided Denoising and Interactive Medical Image FusionabstractMultimodal image fusion aims to merge features from different modalities to create a comprehensively representative image. However, existing medical image fusion methods often struggle to handle noise generated during image acquisition, significantly diminishing their impact on visual quality. To address these challenges, we propose a semantically text-guided medical image fusion model, named TDMF. Specifically, TDMF guides classical image fusion through textual semantics and effectively coordinates the resolution of degradation and interaction issues during the fusion process. By integrating text encoders and interactive fusion modules, TDMF establishes a unified framework for denoising and interactive fusion of medical images. Extensive experiments have demonstrated that our proposed text-guided image fusion strategy offers significant advantages over state-of-the-art methods in medical image fusion performance. Aimei Dong, Guohua Lv, Guixin Zhao, Jinyong Cheng |
ICASSP | 4 |
| 2025 | SSFSL: Self-Supervised and Few-Shot Learning for Cross-Domain Hyperspectral Image ClassificationabstractFew-shot learning (FSL) has gained increasing attention in hyperspectral image (HSI) classification due to its ability to perform cross-domain classification with minimal labeled samples. However, existing FSL methods overlook the continuity of HSI spectral sequences and fail to utilize the large amount of unlabeled samples in the target domain. To address these issues, we introduce a novel cross-domain HSI classification method that combines self-supervised learning with FSL (SSFSL). This approach uses self-supervised learning and FSL to extract transferable knowledge from the source domain and introduces an adaptive soft label generation algorithm to leverage unlabeled samples in the target domain. Compared to existing cross-domain FSL classification methods, the proposed approach considers the spectral sequence continuity of HSI and effectively extracts useful information from unlabeled samples in the target domain. Extensive experiments conducted on three datasets demonstrate that SSFSL outperforms state-of-the-art methods in both quantitative and qualitative aspects. Guohua Lv, Qiang Chi, Guixin Zhao, Aimei Dong, Wei Li 0032 |
ICASSP | 1 |
| 2025 | CGNet: Classification-Guided Multi-Task Interactive Network for Hyperspectral and Multispectral Image FusionabstractThe goal of fusing hyperspectral images (HSI) and multispectral images (MSI) is to generate high-resolution hyperspectral images for downstream tasks. However, most existing methods overlook the specific requirements of these tasks, leading to a gap between the fusion process and its subsequent applications due to insufficient guidance from downstream tasks. To address this issue, we propose a classification-guided multitask interactive network (CGNet) that integrates both fusion and classification tasks into a unified framework, with two branches producing the fused image and classification results, respectively. In the fusion branch, we design a multi-level residual refinement module to efficiently integrate spatial and spectral information. Additionally, an attention-based multi-scale fusion module, incorporating both spatial and channel attention, is carefully crafted to enhance representation learning. In the classification branch, both 2-D and 3-D convolutions are employed to improve classification performance. Moreover, an information interaction module is proposed to guide the fusion task based on classification outcomes. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches on the Pavia Centre and Pavia University datasets. Guohua Lv, Yanlong Xu, Yongbiao Gao, Guixin Zhao, Xiangcheng Sun |
ICASSP | 1 |
| 2025 | Comprehensive Feature Processing Based on Attention Mechanism for Co-Salient Object DetectionabstractCo-salient object detection (CoSOD) aims to detect common salient objects across multiple related images. However, existing methods often struggle with limited attention coverage, missing some co-salient objects. To address this, we propose a two-stage feature processing module (FPM) comprising comprehensive feature extraction module (CFE) and feature enhancement module (FEM). CFE extracts comprehensive cosalient features while reducing background noise, and FEM enhances feature representation and adjusts attention weights for full object coverage. Additionally, we introduce an adversarial learning module (ALM) to improve prediction quality by reducing noise in the co-salient regions. Extensive experiments on three benchmark datasets—CoCA, CoSOD3k, and CoSal2015—demonstrate that our model significantly outperforms state-of-the-art methods. The source code is available at https://github.com/yaobaimiao/CFPAM. Guohua Lv, Mao Yuan, Zengbin Zhang, Zhengyang Zhang, Zhenhui Ding, Guangxiao Ma |
ICASSP | 1 |
| 2025 | A Grouping Strategy-Based Progressive Fusion Network for Hyperspectral Image Super-ResolutionabstractHyperspectral super-resolution involves combining low-resolution hyperspectral images with high-resolution multispectral images to produce a high-resolution hyperspectral image. Recently, although many methods for hyperspectral image super-resolution have been proposed, they often fail to fully utilize the high similarity among adjacent bands to enhance fusion performance. Therefore, we propose a grouping strategy-based progressive fusion network (GPFNet) for hyperspectral super-resolution. The core of GPFNet is the grouping strategy fusion block (GPF block), in which grouping-based spatial-spectral information fusion and spatial information refinement are performed. We design the spatial-spectral information fusion module (SSIFM) based on grouped convolutions to capture the feature differences from adjacent bands. To refine spatial details, we develop the spatial information enhancement module (SpaEM), which leverages the hierarchical features extracted by the multi-scale feature extraction module (MIEM). Additionally, a progressive fusion strategy, which involves using multiple upsampled hyperspectral images and downsampled multispectral images, further preserves spectral integrity and spatial details. Extensive experiments show that GPFNet outperforms state-of-the-art methods both qualitatively and quantitatively. Guohua Lv, Baodong Zhang, Yongbiao Gao, Guixin Zhao, Juncan Wang |
ICASSP | 1 |
| 2025 | Decoupled Imbalanced Label Distribution LearningabstractLabel Distribution Learning (LDL) has been successfully implemented in numerous practical applications. However, the imbalance in label distributions presents a significant challenge due to the substantial variation in annotation information. To tackle this issue, we introduce Decoupled Imbalance Label Distribution Learning (DILDL), which decomposes the imbalanced label distribution into a dominant label distribution and a non-dominant label distribution. Our empirical findings reveal that an excessively high description degree of dominant labels can result in substantial gradient information attenuation for non-dominant labels during the learning process. Therefore, we employ the decoupling approach to balance the description degrees of both dominant and non-dominant labels independently. Furthermore, we align the feature representations with the representations of dominant and non-dominant labels separately, aiming to effectively mitigate the distribution shift problem. Experimental results demonstrate that our proposed DILDL outperforms other state-of-the-art methods for imbalance label distribution learning. Yongbiao Gao, Xiangcheng Sun, Miaogen Ling, Yi Zhai 0003, Guohua Lv |
IJCAI | 6 |
| 2025 | OmniNet: Towards Unified Hyperspectral Image Super-Resolution
Yanlong Xu, Guohua Lv, Baodong Zhang, Yongbiao Gao |
PRCV (15) | 2 |
| 2025 | Self-supervised Co-salient Object Detection via Unified Multi-granularity Feature Learning
Mao Yuan, Guohua Lv, Guangxiao Ma, Zhengyang Zhang |
PRCV (16) | 2 |
| 2025 | BSMEF: Optimized multi-exposure image fusion using B-splines and Mamba
Jinyong Cheng, Qinghao Cui, Guohua Lv |
Image Vis. Comput. | 3 |
| 2025 | GLMR-Net: Global-to-local mutually reinforcing network for pneumonia segmentation and classification
Aimei Dong, Guohua Lv, Jinyong Cheng |
Pattern Recognit. | 3 |
| 2025 | SLFusion: A Structure-Aware Infrared and Visible Image Fusion Network for Low-Light ScenesabstractInfrared and visible image fusion is an image enhancement technique that generates a single image with rich textures and significant objectives in a variety of scenarios, providing great convenience for human discrimination and computer recognition. However, in low-light environments, low-intensity visible images tend to blur valuable information, and these details are often ignored during image fusion, resulting in the loss of important information. Although existing methods take into account the damage of low illumination and highlight the illumination in the fusion process, a large amount of structural information is lost in the process of adjusting illumination, resulting in the lack of texture details and poor performance in high-level vision tasks. To address the above challenges, this paper proposes a structure-aware image fusion method for low illumination scenes, called SLFusion, which enhances the illumination while reducing the loss of structural information, leading to a fused image with richer texture details. We first design an illumination enhancement module to separate the degraded illumination from the scene information in the visible image, and mine more details from the low-intensity regions. Based on the fact that image edge information has a good capability of modeling structures, we design an edge extraction network for low-light visible images to model the structural information, which can accurately highlight important structural information and inject it into the fusion image. The proposed method produces fusion results that not only have good visual perception, but also minimize the loss of structural information. Extensive experiments on benchmark datasets demonstrate that the proposed method outperforms state-of-the-art (SOTA) methods in terms of visual quality, quantitative metrics as well as advanced vision tasks. Guohua Lv, Aimei Dong, Zhonghe Wei, Jinyong Cheng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | A Mutually Enhancement Network for Superpixel Segmentation and Classification of Hyperspectral ImageabstractMost existing hyperspectral image (HSI) classification methods primarily focus on capturing subtle spectral variations by leveraging local spectral-spatial cues derived from patch-level representations. However, limited attention has been given to exploring the global spatial contextual correlations among pixels of HSI. In this study, we propose the Superpixel Segmentation and Classification Mutual Enhancement Network (S2CMEN), a novel framework that integrates global spatial correlations with spectral information through the mutual enhancement of superpixel segmentation and classification. Specifically, a global spatial adaptive module (GSAM) is designed to obtain the direct correlation of the global classes in HSI. It consists of an Adaptive Spectral-Superpixel Network (ASSN) and a Graph Convolutional Network (GCN), forming a synergistic architecture that effectively captures global spatial relationships by adaptively deriving superpixel results from HSIs. Notably, GSAM offers a transferable global spatial representation for HSI tasks, enabling integration with other spectral feature extraction models. Furthermore, we develop a Spatial-Spectral Fusion Module (SSFM) to obtain comprehensive spectral features and fuse them with the extracted global spatial features. Finally, under the constraint of a unit loss, the Mutual Enhancement Strategy (MES) can make the superpixel segmentation loss and the classification loss mutually enhance each other for better performance. We conducted extensive experiments on three public datasets. The proposed S2CMEN achieves overall classification accuracies of 97.38%, 92.33%, and 91.38% on Indian Pines, Pavia University, and Houston, respectively, consistently surpassing existing state-of-the-art methods. Mengxin Cao, Yongmin Li 0001, Xu Zhang 0039, Guixin Zhao, Guohua Lv, Aimei Dong, Jinyong Cheng, Wei Li 0032, Xiangjun Dong 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Cross-Domain Hyperspectral Image Classification via Mamba-CNN and Knowledge DistillationabstractDomain adaptation (DA)-based cross-domain hyperspectral image (HSI) classification methods have garnered significant attention. The majority of DA techniques utilize models based on convolutional neural networks (CNNs) and Transformers for feature extraction. However, Transformers may struggle to capture local details in HSIs, while CNNs often underperform in handling long-range dependencies. Furthermore, many methods focus only on aligning marginal distributions while ignoring the consistency of inter-class features, which may lead to feature confusion and degraded classification accuracy. To overcome the challenges mentioned, we propose a Mamba-CNN and knowledge distillation network (MKDnet). Firstly, the network employs a feature extractor that integrates Mamba and CNN frameworks for cross-domain HSI classification, enabling the capture of both global and local features while effectively capturing long-range dependencies. Secondly, domain alignment is achieved through distribution alignment and graph alignment. In the distribution alignment phase, we design a knowledge distillation architecture that utilizes soft labels to enhance the understanding of relationships between classes, thereby improving the consistency of inter-class features. In the graph alignment phase, we use graph convolution to capture connections between nodes and edges and transfer class-level topological relationships across domains. Finally, the classifier is used to obtain classification results, with consistency constraints applied to balance features between classes more effectively. Extensive experiments have demonstrated that MKDnet outperforms other state-of-the-art methods on three public cross-domain HSI datasets. Aoyan Du, Guixin Zhao, Mengxin Cao, Aimei Dong, Guohua Lv, Yongbiao Gao, Xiangjun Dong 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Long and Recent Preference Learning With Recent-K Items Distribution for Recommender SystemabstractReinforcement learning (RL) aims to formulate the recommendation task as a Markov decision process (MDP) and trains an agent to automatically learn the optimal recommendation policy from interaction trajectories through trial-and-error and reward mechanisms. However, most existing RL-based approaches overlook the correlation between items and the dynamics of user interests implied in temporally close interactions. Therefore, in this paper, we propose a reinforcement learning method that incorporates a “recent-k items” distribution to capture users' local preferences. Specifically, we model the output layer as two distinct branches. The “recent-k items” branch, formulated with a Kullback-Leibler divergence loss, learns the recent interests of users, whereas the other branch utilizes a one-step temporal difference error to capture long-term preferences. The proposed structure is integrated into deep Q-learning and actor-critics, resulting in two enhanced methods named R$k$Q and R$k$AC, respectively. Furthermore, a novel soft inter-reward is carefully designed to enhance the proposed method, and we theoretically prove the convergence of the proposed algorithm. We perform extensive experiments on two large real-world datasets and conduct further analysis of the influences of different action sequences, time intervals, and enhancement capabilities for state-of-the-art models. The experimental results demonstrate the efficacy of our proposed methods. Yongbiao Gao, Sijie Niu, Guohua Lv, Miaogen Ling, Xin Geng 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | IDFusion: An Infrared and Visible Image Fusion Network for Illuminating DarknessabstractThe purpose of infrared and visible image fusion is to combine the background information of the visible images and the thermal target information of the infrared images. Current fusion methods often neglect the challenges of low-light conditions. In the night scene, the existing methods fail to capture texture information from the visible images that is hidden in the darkness, producing suboptimal fusion results that can hinder subsequent visual applications. Therefore, we propose a method for infrared and visible image fusion under night scenes, termed as IDFusion. Specifically, our method is divided into three parts, first, a dense auto-encoder is designed to obtain more useful features from the source images. Then we design a brightness enhancement network that removes visible degraded illumination maps to obtain brightness-enhanced features. Finally, a texture fusion network is designed so that the infrared features and enhanced visible features avoid texture loss during the fusion process. Experimental results show that our network can obtain better fusion results, outperforming state-of-the-art methods in terms of subjective visual effects and quantitative metrics. Guohua Lv, Xiyan Wang, Zhonghe Wei, Jinyong Cheng, Guangxiao Ma, Hanju Bao |
CSCWD | 1 |
| 2024 | Illumination-Enhanced Infrared and Low-Light Visible Image FusionabstractInfrared and visible image fusion aims to generate fused images with rich texture details and salient targets. However, existing methods often ignore illumination conditions, and the fusion results under low-light conditions lack texture details. To solve this problem, we propose an illumination-enhanced infrared and visible image fusion method, named IEFusion. Specifically, we first design a feature decomposition network based on retinex theory to obtain enhanced features of infrared and visible images. Then, the cross-modal feature fusion module is used to fuse important features. Finally, a gradient-enhanced image reconstructor is designed to enhance the gradient and generate a fused image. In addition, we design a contrastive learning module to guide the network to focus on deep features and maximize the mutual information of fused and enhanced images. Extensive experiments demonstrate the superiority of our IEFusion over the state-of-the-art methods, in both qualitative and quantitative metrics. Guohua Lv, Xinyue Fu, Chaoqun Sima, Yanlong Xu, Baodong Zhang, Hanju Bao |
ICIP | 1 |
| 2024 | Rafmnet: Reinforced Attention Fusion and Multiscale Network For Noisy Infrared and Visible Image FusionabstractThe purpose of infrared and visible image fusion is to combine the advantages of different types of images to produce more robust and informative images. However, if the source images are noisy, existing fusion methods may not produce clear results. To address this issue, we propose a novel method for infrared and visible image fusion with noise reduction. This method enhances the visual perception of fused images by integrating features of different scales extracted by the denoising network into the fusion network. By using deformable convolutional denoising networks, noise in images can be removed and features can be enhanced. Then, a set of reinforced attention fusion modules (RAFM) are designed to fuse the features extracted by the denoising network. Experimental results demonstrate the effectiveness of our proposed method, which outperforms existing state-of-the-art methods in terms of fusion accuracy and visual perception. Guohua Lv, Xiyan Wang, Yongbiao Gao, Yi Zhai 0003, Guixin Zhao, Guangxiao Ma |
ICIP | 1 |
| 2024 | GLEGNet: Infrared and Visible Image Fusion via Global-Local Feature Extraction and Edge-Gradient Preservation
Guohua Lv, Wenkuo Song, Zhonghe Wei, Aimei Dong, Jinyong Cheng, Guangxiao Ma |
ICONIP (7) | 1 |
| 2024 | CEDP-YOLO: UAV Object Detection Based on Context Enhancement and Dynamic Perception
Zhenhui Ding, Zengbin Zhang, Mao Yuan, Guangxiao Ma, Guohua Lv |
PRCV (3) | 5 |
| 2024 | TLLFusion: An End-to-End Transformer-Based Method for Low-Light Infrared and Visible Image Fusion
Guohua Lv, Xinyue Fu, Yi Zhai 0003, Guixin Zhao, Yongbiao Gao |
PRCV (3) | 1 |
| 2024 | Scd-yolo: a novel object detection method for efficient road crack detection
Kuiye Ding, Zhenhui Ding, Zengbin Zhang, Mao Yuan, Guangxiao Ma, Guohua Lv |
Multim. Syst. | 6 |
| 2024 | MFIFusion: An infrared and visible image enhanced fusion network based on multi-level feature injection
Aimei Dong, Guohua Lv, Guixin Zhao, Jinyong Cheng |
Pattern Recognit. | 4 |
| 2024 | Co-Enhancement of Multi-Modality Image Fusion and Object Detection via Feature AdaptationabstractThe integration of multi-modality images significantly enhances the clarity of critical details for object detection. Valuable semantic data from object detection enriches the fusion process of these images. However, the potential reciprocal relationship that could enhance their mutual performance remains largely unexplored and underutilized, despite some semantic-driven fusion methodologies catering to specific application needs. To address these limitations, this study proposes a mutually reinforcing, dual-task-driven fusion architecture. Specifically, our design integrates a feature-adaptive interlinking module into both image fusion and object detection components, effectively managing the inherent feature discrepancies. The core idea is to channel distinct features from both tasks into a unified feature space after feature transformation. We then design a feature-adaptive selection module to generate features rich in target semantic information and compatible with the fusion network. Finally, effective combination and mutual enhancement of the two tasks are achieved through an alternating training process. A diverse range of swift evaluations is performed across various datasets to corroborate the potential efficiency of our framework, actualizing visible advancements in both fusion effectiveness and detection accuracy. Aimei Dong, Guixin Zhao, Yi Zhai 0003, Guohua Lv, Jinyong Cheng |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Spectral-Spatial-Language Fusion Network for Hyperspectral, LiDAR, and Text Data ClassificationabstractThe fusion classification of hyperspectral image (HSI) and light detection and ranging (LiDAR) data has gained widespread attention because of its ability to obtain more comprehensive spatial and spectral information. However, the heterogeneous gap between HSI and LiDAR data also adversely affects the classification performance. Despite the excellent performance of traditional multimodal fusion classification models, language information containing much linguistic priori knowledge to enrich visual representations needs to be addressed. Therefore, we design a Spectral-Spatial-Language fusion network (S2LFNet), which can fuse visual and language features to broaden the semantic space using linguistic priori knowledge commonly shared between spectral features and spatial features. First, we propose a dual-channel cascaded image fusion encoder (DCIFencoder) for visual feature extraction and progressive feature fusion of different levels for HSI and LiDAR data. Then, three aspects of Text data are designed to extract linguistic priori knowledge using the Text encoder. Finally, contrastive learning is utilized to construct a unified semantic space, and Spectral-Spatial-Language fusion features are obtained for classification tasks. We evaluate the classification performance of the proposed S2LFNet on three datasets through extensive experiments, and the results show that it outperforms the state-of-the-art fusion classification methods. Mengxin Cao, Guixin Zhao, Guohua Lv, Aimei Dong, Ying Guo 0030, Xiangjun Dong 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Multi-Source Domain Transfer Learning on Epilepsy DiagnosisabstractEpilepsy is a neurological disease that occurs in all ages and seriously threatens physical and mental health. There are two problems in the present study. One is the limitation of the amount of publicly available medical data. And the other is that the distributions of the data are different but correlated. Conventional machine learning methods are not applicable. But transfer learning method has shown promising performance in solving both problems. In this paper, a multi-source domain transfer learning method called MDTL for epilepsy diagnosis is proposed. In order to fully exploit the specific features and common features of the dataset, we propose a domain specific feature extractor and a common feature extractor. For enhancing data, we transform the signals into time-frequency diagrams to rotate and crop. The three types of electrocardiogram (ECG) time-frequency diagram are put to train model, and the model is transferred to electroencephalogram (EEG) time-frequency diagrams. The results confirm that MDTL is effective in epilepsy diagnosis. Aimei Dong, Zhiyun Qi, Yi Zhai 0003, Guohua Lv |
CSCWD | 4 |
| 2023 | A Deep Fusion Rule for Infrared and Visible Image Fusion: Feature Communication for Importance AssessmentabstractThe purpose of infrared and visible image fusion is to extract and combine information from source images to produce results that contain vital and complementary information. Existing fusion rules may not extract the most useful information and cannot effectively retain important information. To solve this problem, we propose a novel deep learning-based fusion rule. We perform deep feature communication and quantify the effect of feature substitution on image characteristics, including gradients and contrast. Based on the influence, the importance of feature maps can be objectively assessed. This designed fusion rule can preserve the thermal information of infrared images and the texture details of visible light images in a targeted manner, so as to obtain better fusion results. Qualitative and quantitative experiments have shown that our method can perform better than other advanced methods. Xuran Lv, Jinyong Cheng, Guohua Lv, Zhonghe Wei |
ICASSP | 3 |
| 2023 | Few-Shot Hyperspectral Image Classification Based on Cross-Domain Spectral Semantic Relation TransformerabstractIn practical hyperspectral image (HSI) classification tasks, we often encounter the problems of few-shot classification and domain misalignment between source domains and target domains. To solve this classification paradigm, a meta-learning method of few-shot learning (FSL) is usually used. However, most existing FSL methods address the problem for domain alignment and neglect the exploration of semantic relationships of objects across domains. In this paper, we propose the cross-domain spectral semantic transformer FSL (SSTFSL), which can fully extract semantic features and spectral detail features for the cross-domain few-shot HSI classification task. Specifically, the multi-head self-attention (MSA) mechanism with enhancement process (EP) of the transformer is used to map out semantically relevant local regions and can enhance the ability of the model to distinguish subtle feature differences in the spectrum. In addition, the matching degree of different branches is computed by relational network learning, which ultimately enables cross-domain few-shot HSI classification. Through extensive experiments, we evaluate the classification performance of SSTFSL on HSI datasets. The results demonstrate that SSTFSL outperforms existing FSL methods and deep learning methods on HSI classification. Mengxin Cao, Guixin Zhao, Aimei Dong, Guohua Lv, Ying Guo 0030, Xiangjun Dong 0001 |
ICIP | 4 |
| 2023 | Mix-Net: Automatic Segmentation of Covid-19 ct Images Based on Parallel DesignabstractSince the discovery of COVID-19 in late 2019, the viral pneumonia crisis has begun to spread rapidly around the world. Lesion segmentation can remove unnecessary background areas and help doctors diagnose the condition. However, the infected areas showed differences at different stages, and the border between the infected areas and the surrounding tissue was blurred. To solve this problem, a novel COVID-19 lung infection segmentation network (Mix-Net) is designed for the automatic identification of infected areas from chest CT slices. Specifically, first, the local and global features of the infected areas are extracted and interacted with using the mixing block. Then, the features extracted from multiple layers of the encoder are fused and connected to the decoder. Experiments show that Mix-Net outperforms most cutting-edge segmentation models and achieves good segmentation results. Aimei Dong, Guohua Lv, Guixin Zhao, Yi Zhai 0003 |
ICIP | 3 |
| 2023 | L2fusion: Low-Light Oriented Infrared and Visible Image FusionabstractInfrared and visible image fusion aims to integrate salient targets and abundant texture information into a single fused image. Existing methods typically ignore the issue of illumination, so that there are problems of weak texture details and poor visual perception in case of low illumination. To address this issue, we propose a low-light oriented infrared and visible image fusion network, named L2Fusion. In particular, we first design a decomposition network according to Retinex theory to obtain the reflectance features of a visible image with low-light. Then, these features are integrated with the features extracted from the corresponding infrared image by a residual network. The finally fused image largely eliminates the negative impact caused by low illumination, and contains both salient targets and abundant texture information. Extensive experiments demonstrate the superiority of our L2Fusion over the state-of-the-art methods, in terms of both visual effect and quantitative metrics. Guohua Lv, Aimei Dong, Zhonghe Wei, Jinyong Cheng |
ICIP | 2 |
| 2023 | Few-Shot Hyperspectral Image Classification with Spectral-Spatial Feature Fusion Based on Fuzzy Broad Learning SystemabstractIn the few-shot hyperspectral image (HSI) classification, most current models don't fully utilize the advantage of spectral-spatial feature fusion, resulting in low classification accuracy. Therefore, we propose a few-shot HSI classification model with spectral-spatial feature fusion based on fuzzy broad learning system (FBLS) (FSFBLS). Firstly, we use a Gaussian filter to suppress noise while smoothing spectral features based on spatial information to achieve the first fusion of spectral-spatial features. Secondly, we use FBLS with fuzzy rules to fully model the complex mapping relationship between spectral-spatial features and HSI labels to complete HSI classification. The fuzzy processing can extract rich discriminative features to enhance the recognition of different categories. Finally, the guided filter corrects the misclassified samples of FBLS based on the guided image to achieve the second fusion of spectral-spatial features. Extensive experimental results on three public datasets demonstrate that FSFBLS achieves state-of-the-art classification performance compared to nine popular models. Xiaopei Hu, Guixin Zhao, Aimei Dong, Guohua Lv, Yi Zhai 0003, Ying Guo 0030, Xiangjun Dong 0001 |
ICIP | 4 |
| 2023 | BS-YOLOv5s: Insulator Defect Detection with Attention Mechanism and Multi-Scale FusionabstractWith the rapid development of deep learning, the use of object detection algorithms for aerial insulator image defect detection has become the main way. To address the problems of low detection accuracy for small targets, weak representation ability of feature maps, insufficient extracted key information, and the shortage of aerial insulator defect datasets, this paper proposes an improved insulator defect detection method named BS-YOLOv5s based on 3-D attention mechanism and Bi-Slim-neck using YOLOv5s as the base network. Additionally, to solve the problem of the shortage of aerial insulator datasets, this paper proposes a new aerial insulator dataset Weather-Insulator (WI) containing a variety of defect scenarios. The experimental results demonstrate that the proposed method not only greatly improves the detection accuracy, but also maintains a high detection speed, satisfying the engineering requirements for insulator defect detection. The dataset and code for this paper are publicly available at https://github.com/jspron/insulator-defect. Zengbin Zhang, Guohua Lv, Guixin Zhao, Yi Zhai 0003, Jinyong Cheng |
ICIP | 2 |
| 2023 | Autism Spectrum Disorder Diagnosis Using Graph Neural Network Based on Graph Pooling and Self-adjust Filter
Aimei Dong, Xuening Zhang, Guohua Lv, Guixin Zhao, Yi Zhai 0003 |
PRCV (13) | 3 |
| 2023 | Online Class-Incremental Learning in Image Classification Based on Attention
Baoyu Du, Zhonghe Wei, Jinyong Cheng, Guohua Lv, Xiaoyu Dai |
PRCV (7) | 4 |
| 2023 | SIEFusion: Infrared and Visible Image Fusion via Semantic Information Enhancement
Guohua Lv, Wenkuo Song, Zhonghe Wei, Jinyong Cheng, Aimei Dong |
PRCV (3) | 1 |
| 2023 | Momentum contrast transformer for COVID-19 diagnosis with knowledge distillation
Aimei Dong, Zhonghe Wei, Yi Zhai 0003, Guohua Lv |
Pattern Recognit. | 6 |
| 2022 | Robust Registration of Multispectral Satellite Images Based on Structural and Geometrical SimilarityabstractAccurate registration of multispectral satellite images is a challenging task due to the significant and nonlinear radiometric differences between these data. To address this problem, this letter explores the strategy of geometrical similarity between triplets of feature points, and it is combined with the structural similarity between images in a feature-based image registration framework. The underlying principle is that the structural and geometrical similarities generally preserve across the images being registered. In this feature-based image registration framework, a set of control points (CPs) are first detected. Then, the geometric similarity between triplets of CPs is defined, followed by a ranking operation of these triplets of CPs. The highly ranked triplets are used to estimate a spatial transformation between images. Finally, initial matches obtained by a benchmark registration technique are refined by the estimated transformation. The experimental results demonstrate the great effectiveness of the proposed technique for registering multispectral satellite images. Guohua Lv, Qiang Chi, Mohammad Awrangjeb, Jian Li 0034 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2018 | A detector of structural similarity for multi-modal microscopic image registration
Guohua Lv, Shyh Wei Teng, Guojun Lu |
Multim. Tools Appl. | 1 |
| 2018 | COREG: a corner based registration technique for multimodal images
Guohua Lv, Shyh Wei Teng, Guojun Lu |
Multim. Tools Appl. | 1 |
| 2018 | Enhancing image registration performance by incorporating distribution and spatial distance of local descriptors
Guohua Lv, Shyh Wei Teng, Guojun Lu |
Pattern Recognit. Lett. | 1 |
| 2017 | e-NSPFI: Efficient Mining Negative Sequential Pattern from Both Frequent and Infrequent Positive Sequential PatternsabstractNegative sequential patterns (NSPs), which focus on nonoccurring but interesting behaviors (e.g. missing consumption records), provide a special perspective of analyzing sequential patterns. So far, very few methods have been proposed to solve for NSP mining problem, and these methods only mine NSP from positive sequential patterns (PSPs). However, as many useful negative association rules are mined from infrequent itemsets, many meaningful NSPs can also be found from infrequent positive sequences (IPSs). The challenge of mining NSP from IPS is how to constrain which IPS could be available used during NSP process because, if without constraints, the number of IPS would be too large to be handled. So in this study, we first propose a strategy to constrain which IPS could be available and utilized for mining NSP. Then we give a storage optimization method to hold this IPS information. Finally, an efficient algorithm called Efficient mining Negative Sequential Pattern from both Frequent and Infrequent positive sequential patterns (e-NSPFI) is proposed for mining NSP. The experimental results show that e-NSPFI can efficiently find much more interesting negative patterns than e-NSP. Yongshun Gong, Tiantian Xu 0002, Xiangjun Dong 0001, Guohua Lv |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2016 | Enhancing SIFT-based image registration performance by building and selecting highly discriminating descriptors
Guohua Lv, Shyh Wei Teng, Guojun Lu |
Pattern Recognit. Lett. | 1 |
| 2013 | Maximizing structural similarity in multimodal biomedical microscopic images for effective registrationabstractMultimodal image registration (MMIR) is the alignment of contents in images captured from different sensors or instruments. MMIR is important in medical applications as it enables the visualization of the complementary contents in biomedical microscopic images. The registration for such images can be challenging as the structures of their contents are usually only partially similar. Thus in this paper, we propose a new method to maximize the structural similarity of the contents in such images by utilizing intensity relationships among Red-Green-Blue color channels. Our experimental results will demonstrate that our proposed method substantially improves the accuracy of registering such images as compared to the state-of-the-art methods. Guohua Lv, Shyh Wei Teng, Guojun Lu, Martin Lackmann |
ICME | 1 |