Genji Yuan

dblp:236/2921 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
21since 2021 · last 2026
0000-0002-8710-2266ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Computer networks · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Empowering Semantic-Sensitive Underwater Image Enhancement with VLM
abstract
In recent years, learning-based underwater image enhancement (UIE) techniques have rapidly evolved. However, distribution shifts between high-quality enhanced outputs and natural images can hinder semantic cue extraction for downstream vision tasks, thereby limiting the adaptability of existing enhancement models. To address this challenge, this work proposes a new learning mechanism that leverages Vision-Language Models (VLMs) to empower UIE models with semantic-sensitive capabilities. To be concrete, our strategy first generates textual descriptions of key objects from a degraded image via a VLM. Subsequently, a text-image alignment model remaps these relevant descriptions back onto the image to produce a spatial semantic guidance map. This map then steers the UIE network through a dual-guidance mechanism, which combines cross-attention and an explicit alignment loss. This forces the network to focus its restorative power on semantic-sensitive regions during image reconstruction, rather than pursuing a globally uniform improvement, thereby ensuring the faithful restoration of key object features. Experiments confirm that when our strategy is applied to different UIE baselines, significantly boosts their performance on perceptual quality metrics as well as enhances their performance on detection and segmentation tasks, validating its effectiveness and adaptability.
Shengning Zhou, Genji Yuan, Jingchun Zhou, Jinjiang Li 0001
AAAI3
2026 MIGDUN: Multi-stage interactive guidance deep unfolding network for pansharpening remote sensing images
Hailin Tao, Genji Yuan, Jinjiang Li 0001
Neurocomputing2
2026 PFSC-Net: Physically-guided frequency-spatial-color fusion network for underwater image enhancement
Yongle Kan, Genji Yuan, Jinjiang Li 0001
Inf. Sci.3
2026 From unstable conditioning to disentangled prompts: Variance-stable diffusion model for harsh underwater image enhancement
Shengning Zhou, Genji Yuan
Pattern Recognit.3
2025 AGTCNet: Hybrid Network Based on AGT and Curvature Information for Skin Lesion Detection
Zhiwei Dong, Genji Yuan, Jinjiang Li 0001
CVM (1)2
2025 Mamba-enhanced spectral-attentive wavelet network for underwater image restoration
Baocai Chang, Genji Yuan, Jinjiang Li 0001
Eng. Appl. Artif. Intell.2
2025 CDME: Convolutional Dictionary Iterative Model for Pansharpening With Mixture of Experts
abstract
In this letter, we propose a convolutional dictionary iterative model for pansharpening with a mixture of experts. First, we define an observation model to model the common and unique feature information between multispectral (MS) and panchromatic (PAN) images. During this process, a proximal gradient algorithm is used to iteratively update the network parameters. The adaptive expert module (AEM) is designed to handle the unique and common features separately by using PAN mixture of experts (PMOE), multispectral mixture-of-experts (MMOE), and common mixture-of-experts (CMOE) modules, to achieve effective information reconstruction. Finally, the expert mixture fusion module (EMFM) adaptively integrates the information from the three mixture-of-experts (MOE) components by dynamically adjusting their respective weights, resulting in the final fused image. We conducted full-resolution and reduce-resolution experiments on GF2 and WV3 datasets with current state-of-the-art methods, and the experimental results show that our method performs best. The code is released onhttps://github.com/who15/CDME.
Genji Yuan, Zhen Hua, Jinjiang Li 0001
IEEE Geosci. Remote. Sens. Lett.3
2025 GSSR-Net: Geo-Spatial Structural Refinement Network for Remote Sensing Change Detection
abstract
Remote sensing (RS) change detection is a technique for identifying changes in the ground surface by comparing RS images from different periods. Although such tasks have been developed for a long time and some methods have been proposed to enhance the features of real changes in the ground objects, they still face the perception ambiguity caused by the heterogeneity of the ground objects in complex RS environments: 1) insufficient processing of nonstationary changes between dual-temporal image features and 2) high spatial heterogeneity leads to difficulties in structural identification. In order to solve the interference of these problems on the downstream tasks of change detection, this article proposes geo-spatial structural refinement network (GSSR-Net) for RS change detection. First, we introduce a DualTime Mamba structure with an omnidirectional scanning path, adjust its input matrix between dual-temporal image features in the deep scale, and allow the model to fully consider the spatiotemporal dependence and image structure information of the previous and next time points. In addition, this article designs a land-cover feature extraction (LCFE) method to improve the perception ability of the ground object target structure. Specifically, this method refines the image edge by separating high frequencies, adjusts the contour structure information of the dual-phase image by separating low frequencies, and then further models the relationship between pixels by combining feature structures and spatial offset mechanisms. Our experimental results on three datasets demonstrate the superiority of GSSR-Net. The network code address ishttps://github.com/SparrowTought/GSSR-Net.
Shuo Wang 0019, Genji Yuan, Jinjiang Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 Dual-Branch Network for Spatial-Channel Stream Modeling Based on the State-Space Model for Remote Sensing Image Segmentation
abstract
To quickly and effectively address the color similarity issue in remote sensing image segmentation, traditional channel attention methods typically use channel modeling based on global channel statistics mapped to weights. However, this approach either suffers from limitations in feature selection due to a lack of dynamic interactions and the loss of significant spatial information, resulting in poor performance in complex scenarios, or has high computational complexity, making it difficult to apply in high-resolution remote sensing images. To overcome these challenges, this article proposes an innovative streaming channel modeling method based on state-space models (SSMs), aimed at rapidly and efficiently tackling the color similarity problem in remote sensing image segmentation. Specifically, we designed the channel-position state-space model network (CPSSNet) framework, where the decoder comprises the spatial Mamba block (SMB) for spatial modeling and the channel Mamba block (CMB) for streaming channel modeling. The core component of SMB, position-selective-scan-2D, achieves multidirectional global modeling in the spatial domain through a combination of the spatial scanning algorithm and SSM, with linear complexity. The core component of CMB, channel-2D selective-scan (C-SS2D), fuses channel and spatial information into patches for streaming modeling using a combination of the channel scanning (CS) algorithm and SSM. We have further improved SSM within C-SS2D to enhance dynamic interactions between channels, allowing for more refined modeling while maintaining linear complexity. Experimental results demonstrate that CPSSNet exhibits outstanding performance in addressing color similarity challenges in remote sensing image segmentation. The code is available athttps://github.com/yysdck/CPSSNet.
Yunsong Yang, Genji Yuan, Jinjiang Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Disparity Map-Crack Detection: Combining Disparity Map Feature into Binary Segmentation for Accurate Crack Detection
abstract
To address the limitations of crack detection methods relying solely on RGB images, we propose an in-novative approach that incorporates disparity maps as an additional data source. This integration with RGB images aims to enhance crack detection performance. However, the generation of disparity maps is susceptible to image noise and matching errors, leading to inaccurate or mismatched disparity values that may impede precise ground crack detection. To mitigate this challenge, we apply a disparity transformation technique to refine the estimated disparity map, improving the differentiation of crack regions. Additionally, we employ a feature fusion method based on a connected low-loss subspace. This approach adaptively assigns feature weights to facilitate the complementary fusion of disparity and RGB features. Fur-thermore, the decoder includes a multi-scale feature alignment module that uses the fused encoder features to align each layer in the decoding process. This preserves image details and local features, enhancing the overall detection accuracy. Extensive testing experiments demonstrate a significant breakthrough in crack detection performance, achieving an Intersection over Union (IoU) of 80.48%. Our approach sets the benchmark in crack detection, effectively leveraging multi-source information, mitigating disparity map noise, and enhancing feature fusion.
Genji Yuan
SMC2
2024 Diffusion model-based text-guided enhancement network for medical image segmentation
Zhiwei Dong, Genji Yuan, Zhen Hua, Jinjiang Li 0001
Expert Syst. Appl.2
2024 DBEF-Net: Diffusion-Based Boundary-Enhanced Fusion Network for medical image segmentation
Zhenyang Huang, Jianjun Li 0011, Genji Yuan, Jinjiang Li 0001
Expert Syst. Appl.4
2024 DUCD: Deep Unfolding Convolutional-Dictionary network for pansharpening remote sensing image
Genji Yuan, Jinjiang Li 0001
Expert Syst. Appl.2
2024 Independence and Unity: Unseen Domain Segmentation Based on Federated Learning
abstract
The distinct attributes of Internet of Things (IoTs) devices, including the disparity between training and testing data distributions and limited availability of training data, pose challenges for deep learning models in effectively addressing unseen domain segmentation tasks. Federated Learning (FL) can increase the participation of various data contributors, thus has great potential to develop a unified framework to shed light on the relationship between unseen domains and generalized domains. In this paper, we proposed an FL-based unseen domain segmentation model. The architecture includes (1) an external memory module as an object feature guide to reduce the feature ambiguity of unseen domain objects. (2) A re-attention activation mechanism for better completing localization of unseen domain objects, enhancing the features of potential targets and suppressing interference features. (3) A self-supervised learning paradigm for achieving specific object feature exploration. Based on flexible splitting and combining, our model is able to capture both personalization and generalization capabilities, the client side retains a strong personalization ability, while the server side has a strong generalization ability. Moreover, taking into account the inherent limitations in computing and storage resources commonly associated with IoT devices, the introduced model leverages the concept of optional dependencies to enable efficient inference within resource-constrained client environments. Our proposed model is validated through extensive experiments. The approach proposed in this paper outperforms the generalization capabilities of state-of-the-art work on several benchmarks.
Genji Yuan, Yan Huang 0032, Zhenzhen Xie 0002, Junjie Pang, Zhipeng Cai 0001
IEEE Internet Things J.1
2024 MFDS-Net: Multiscale Feature Depth-Supervised Network for Remote Sensing Change Detection With Global Semantic and Detail Information
abstract
Change detection as an interdisciplinary discipline in the field of computer vision and remote sensing at present has been receiving extensive attention and research. Due to the rapid development of society, the geographic information captured by remote sensing satellites is changing faster and more complex, which undoubtedly poses a higher challenge and highlights the value of change detection tasks. We propose a multiscale feature depth-supervised network (MFDS-Net) for remote sensing change detection with global semantic and detail information (MFDS-Net) with the aim of achieving a more refined description of changing buildings as well as geographic information, enhancing the localization of changing targets and the acquisition of weak features. To achieve the research objectives, we use a modified$\text {ResNet}_{34}$as a backbone network to perform feature extraction. We propose the global semantic enhancement module (GSEM) to enhance the processing of high-level semantic information from a global perspective. The differential feature integration module (DFIM) is proposed to strengthen the fusion of different depth feature information, achieving learning and extraction of differential features. The entire network is trained and optimized using a deep supervision mechanism. The experimental outcomes of MFDS-Net surpass those of current mainstream change detection networks. On the LEVIR dataset, it achieved an F1 score of 91.589 and an IoU of 84.483. The code is available athttps://github.com/AOZAKIiii/MFDS-Net.
Zhenyang Huang, Zhaojin Fu, Jintao Song, Genji Yuan, Jinjiang Li 0001
IEEE Geosci. Remote. Sens. Lett.4
2024 ConMamba: CNN and SSM High-Performance Hybrid Network for Remote Sensing Change Detection
abstract
Accurate remote sensing change detection (RSCD) tasks rely on comprehensively processing multiscale information from local details to effectively integrate global dependencies. Hybrid models based on convolutional neural networks (CNNs) and Transformers have become mainstream approaches in RSCD due to their complementary advantages in local feature extraction and long-term dependency modeling. However, the Transformer faces application bottlenecks due to the high secondary complexity of its attention mechanism. In recent years, state-space models (SSMs) with efficient hardware-aware design, represented by Mamba, have gained widespread attention for their excellent performance in long-series modeling and have demonstrated significant advantages in terms of improved accuracy, reduced memory consumption, and reduced computational cost. Based on the high match between the efficiency of SSM in long sequence data processing and the requirements of the RSCD task, this study explores the potential of its application in the RSCD task. However, relying on SSM alone is insufficient in recognizing fine-grained features in remote sensing images. To this end, we propose a novel hybrid architecture, ConMamba, which constructs a high-performance hybrid encoder (CS-Hybridizer) by realizing the deep integration of the CNN and SSM through the feature interaction module (FIM). In addition, we introduce the spatial integration module (SIM) in the feature reconstruction stage to further enhance the model’s ability to integrate complex contextual information. Extensive experimental results on three publicly available RSCD datasets show that ConMamba significantly outperforms existing techniques in several performance metrics, validating the effectiveness and foresight of the hybrid architecture based on the CNN and SSM in RSCD.
Zhiwei Dong, Genji Yuan, Zhen Hua, Jinjiang Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 SFFNet: A Wavelet-Based Spatial and Frequency Domain Fusion Network for Remote Sensing Segmentation
abstract
To fully utilize spatial information for segmentation and address the challenge of handling areas with significant grayscale variations in remote sensing segmentation, we propose the spatial and frequency domain fusion network (SFFNet) framework. This framework employs a two-stage network design: the first stage extracts features using spatial methods to obtain features with sufficient spatial details and semantic information; the second stage maps these features in both spatial and frequency domains. In the frequency domain mapping, we introduce the wavelet transform feature decomposer (WTFD) structure, which decomposes features into low-frequency and high-frequency components using the Haar wavelet transform and integrates them with spatial features. To bridge the semantic gap between frequency and spatial features, facilitating significant feature selection to promote the combination of features from different representation domains, we design the multiscale dual-representation alignment filter (MDAF). This structure utilizes multiscale convolutions and dual-cross attentions. Comprehensive experimental results demonstrate that, compared to existing methods, SFFNet achieves superior performance in terms of mean intersection over union (mIoU), reaching 84.80% and 87.73%, respectively. The code is located athttps://github.com/yysdck/SFFNet.
Yunsong Yang, Genji Yuan, Jinjiang Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Edge-assisted Object Segmentation Using Multimodal Feature Aggregation and Learning
abstract
Object segmentation aims to perfectly identify objects embedded in the surrounding environment and has a wide range of applications. Most previous methods of object segmentation only use RGB images and ignore geometric information from disparity images. Making full use of heterogeneous data from different devices has proved to be a very effective strategy for improving segmentation performance. The key challenge of the multimodal fusion-based object segmentation task lies in the learning, transformation, and fusion of multimodal information. In this article, we focus on the transformation of disparity images and the fusion of multimodal features. We develop a multimodal fusion object segmentation framework, termed the Hybrid Fusion Segmentation Network (HFSNet). Specifically, HFSNet contains three key components, i.e., disparity convolutional sparse coding (DCSC), asymmetric dense projection feature aggregation (ADPFA), and multimodal feature fusion (MFF). The DCSC is designed based on convolutional sparse coding. It not only has better interpretability but also preserves the key geometric information of the object. ADPFA is designed to enhance texture and geometric information to fully exploit nonadjacent features. MFF is used to perform multimodal feature fusion. Extensive experiments show that our HFSNet outperforms existing state-of-the-art models on two challenging datasets.
Genji Yuan, Zheng Yang 0002
ACM Trans. Sens. Networks2
2022 Incorporating Self Attention Mechanism into Semantic Segmentation for Lane Detection
Genji Yuan
WASA (2)1
2021 DDCAttNet: Road Segmentation Network for Remote Sensing Images
Genji Yuan, Zhiqiang Lv, Yinong Li, Zhihao Xu 0002
WASA (2)1
2021 Image matting trimap optimization by ant colony algorithm
Genji Yuan, Jinjiang Li 0001, Zhen Hua
Multim. Tools Appl.1
2020 Robust trimap generation based on manifold ranking
Jinjiang Li 0001, Genji Yuan
Inf. Sci.2