Jinjiang Li 0001

dblp:97/7026-1 · DBLP profile ↗
← Back
106ranked-venue papers
8as first author
96since 2021 · last 2026
0000-0002-2080-8678ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 44 · 4 first-author · 38 since 2021Applied, interdisciplinary, general and emerging computing · 28 · 3 first-author · 25 since 2021Artificial intelligence and machine learning · 26 · 26 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Empowering Semantic-Sensitive Underwater Image Enhancement with VLM
abstract
In recent years, learning-based underwater image enhancement (UIE) techniques have rapidly evolved. However, distribution shifts between high-quality enhanced outputs and natural images can hinder semantic cue extraction for downstream vision tasks, thereby limiting the adaptability of existing enhancement models. To address this challenge, this work proposes a new learning mechanism that leverages Vision-Language Models (VLMs) to empower UIE models with semantic-sensitive capabilities. To be concrete, our strategy first generates textual descriptions of key objects from a degraded image via a VLM. Subsequently, a text-image alignment model remaps these relevant descriptions back onto the image to produce a spatial semantic guidance map. This map then steers the UIE network through a dual-guidance mechanism, which combines cross-attention and an explicit alignment loss. This forces the network to focus its restorative power on semantic-sensitive regions during image reconstruction, rather than pursuing a globally uniform improvement, thereby ensuring the faithful restoration of key object features. Experiments confirm that when our strategy is applied to different UIE baselines, significantly boosts their performance on perceptual quality metrics as well as enhances their performance on detection and segmentation tasks, validating its effectiveness and adaptability.
Shengning Zhou, Genji Yuan, Jingchun Zhou, Jinjiang Li 0001
AAAI6
2026 Diffusion prior guided deep model driven network for infrared and visible image fusion
abstract
Deep learning-based methods have achieved significant success in the field of image fusion, where the design of network architectures plays a crucial role in the fusion task. However, most deep learning fusion architectures still operate as black boxes and lack awareness of frequency domain information. To enhance the interpretability of fusion tasks and effectively utilize the frequency domain information from source images while ensuring the generation of high-quality fused images, we propose a novel deep model-driven network for infrared and visible image fusion guided by diffusion priors. The proposed algorithm generates a distribution of multi-channel input data through a diffusion process. Features produced by the denoiser serve as knowledge priors, guiding a custom model that leverages frequency domain knowledge to reconstruct high-quality fused images. Specifically, unlike some traditional fusion networks, our model can retain multi-channel data at the input stage rather than just single-channel spatial information. It uses the multi-channel features generated by the diffusion model as priors to guide the fusion process. Moreover, we innovatively design a frequency-domain-based objective function to guide the fusion process, constructing a frequency-domain learning module to simulate an interpretable deep model-driven network. Additionally, a task-driven loss function is developed to ensure the quality of the fused images. Extensive experimental evaluations across seven diverse datasets (e.g., MSRS, M3FD, RoadSence, TNO, Havard) and multiple scenarios demonstrate that the proposed algorithm significantly outperforms 9 state-of-the-art methods. Specifically, it delivers superior fusion results on eight metrics (e.g., EN, SF, VIF) with notable improvements in interpretability and robustness, as validated through comprehensive experiments on these seven benchmark datasets.
Shuo Wang 0019, Dapeng Cheng, Jinjiang Li 0001
Expert Syst. Appl.3
2026 MIGDUN: Multi-stage interactive guidance deep unfolding network for pansharpening remote sensing images
Hailin Tao, Genji Yuan, Jinjiang Li 0001
Neurocomputing4
2026 WaveGateNet: A wavelet-guided gated network for frequency-spatial collaborative underwater image enhancement
Shaotuo Zhang, Zhaolong Gao, Jingchun Zhou, Min Gan, Jinjiang Li 0001
Neurocomputing6
2026 PFSC-Net: Physically-guided frequency-spatial-color fusion network for underwater image enhancement
Yongle Kan, Genji Yuan, Jinjiang Li 0001
Inf. Sci.4
2026 WRSI: Wavelet scale reduction and spatiotemporal boundary interaction network for remote sensing change detection
Kaisheng Li, Shuo Wang 0019, Zhaolong Gao, Guangyong Chen, Jinjiang Li 0001
Inf. Sci.6
2026 WCI-Mamba: Overcoming intensity inhomogeneity in remote sensing image segmentation
Shengning Zhou, Yunsong Yang, Linwei Fan, Yihui Liu, Jinjiang Li 0001
Pattern Recognit.6
2026 DCD-UIE: Decoupled Chromatic Diffusion Model for Underwater Image Enhancement
abstract
Color distortion and structural degradation in underwater images are classic challenges in underwater image enhancement. The core goal is to restore degraded images to high-quality images with both color and structure that conform to visual perception. However, in the traditional RGB space, these two issues are highly coupled, resulting in existing enhancement methods often neglecting one over the other. To address this challenge, we propose a guided diffusion model based on the principle of decoupling. Our key insight is that in perceptual color spaces such as HSV, color (H, S) and structure (V) are naturally separated. To exploit this property, we first design an adaptive perceptual guidance module, which analyzes the degraded HSV image and generates two orthogonal guidance signals: a color guide and a structure guide, which guide the denoising process of the diffusion model. To ensure that this decoupled guidance is faithfully implemented, we propose a corresponding decoupled loss optimization module, which uses independent loss functions to supervise the final output color and structure. By combining the forward decoupled guidance with the backward decoupled supervision, we construct a closed-loop optimization framework. This framework enables the model to collaboratively optimize color and structure under various degradation scenarios. Extensive experiments demonstrate that our proposed method outperforms existing state-of-the-art approaches in a variety of underwater scenes, particularly those degraded by color casts and haze. Furthermore, it exhibits superior performance on no-reference image quality assessment metrics. The source code is available at https://github.com/zy-world/DCD-UIE.
Jingchun Zhou, Yakun Ju, Guang-Yong Chen, Jinjiang Li 0001, Alex Chichung Kot
IEEE Trans. Image Process.6
2026 BAMN: boundary-aware mamba network for skin lesion segmentation
Lutong Sun, Peng Duan 0004, Jinjiang Li 0001
J. Supercomput.3
2026 Dgenet: diffusion model-based graph convolution enhancement network for medical image segmentation
Yunfei Zhu, Jintao Song, Jinjiang Li 0001
J. Supercomput.3
2025 AGTCNet: Hybrid Network Based on AGT and Curvature Information for Skin Lesion Detection
Zhiwei Dong, Genji Yuan, Jinjiang Li 0001
CVM (1)3
2025 Mamba-enhanced spectral-attentive wavelet network for underwater image restoration
Baocai Chang, Genji Yuan, Jinjiang Li 0001
Eng. Appl. Artif. Intell.3
2025 SpectMamba: Remote sensing change detection network integrating frequency and visual state space model
Zhiwei Dong, Dapeng Cheng, Jinjiang Li 0001
Expert Syst. Appl.3
2025 Underwater variable zoom: Depth-guided perception network for underwater image enhancement
Zhixiong Huang, Xinying Wang 0005, Chengpei Xu, Jinjiang Li 0001, Lin Feng 0001
Expert Syst. Appl.4
2025 FSCMF: A Dual-Branch Frequency-Spatial Joint Perception Cross-Modality Network for visible and infrared image fusion
Chengpei Xu, Zhen Hua, Jinjiang Li 0001, Jingchun Zhou
Neurocomputing5
2025 PanRouter: Router-based dynamic decision interactive pansharpening network
Hailin Tao, Dapeng Cheng, Jinjiang Li 0001
Inf. Sci.3
2025 CDME: Convolutional Dictionary Iterative Model for Pansharpening With Mixture of Experts
abstract
In this letter, we propose a convolutional dictionary iterative model for pansharpening with a mixture of experts. First, we define an observation model to model the common and unique feature information between multispectral (MS) and panchromatic (PAN) images. During this process, a proximal gradient algorithm is used to iteratively update the network parameters. The adaptive expert module (AEM) is designed to handle the unique and common features separately by using PAN mixture of experts (PMOE), multispectral mixture-of-experts (MMOE), and common mixture-of-experts (CMOE) modules, to achieve effective information reconstruction. Finally, the expert mixture fusion module (EMFM) adaptively integrates the information from the three mixture-of-experts (MOE) components by dynamically adjusting their respective weights, resulting in the final fused image. We conducted full-resolution and reduce-resolution experiments on GF2 and WV3 datasets with current state-of-the-art methods, and the experimental results show that our method performs best. The code is released onhttps://github.com/who15/CDME.
Genji Yuan, Zhen Hua, Jinjiang Li 0001
IEEE Geosci. Remote. Sens. Lett.5
2025 Prototype-Driven Structure Synergy Network for Remote Sensing Images Segmentation
Jinjiang Li 0001, Yakun Ju, Alex Chichung Kot
IEEE Trans. Geosci. Remote. Sens.2
2025 GSSR-Net: Geo-Spatial Structural Refinement Network for Remote Sensing Change Detection
abstract
Remote sensing (RS) change detection is a technique for identifying changes in the ground surface by comparing RS images from different periods. Although such tasks have been developed for a long time and some methods have been proposed to enhance the features of real changes in the ground objects, they still face the perception ambiguity caused by the heterogeneity of the ground objects in complex RS environments: 1) insufficient processing of nonstationary changes between dual-temporal image features and 2) high spatial heterogeneity leads to difficulties in structural identification. In order to solve the interference of these problems on the downstream tasks of change detection, this article proposes geo-spatial structural refinement network (GSSR-Net) for RS change detection. First, we introduce a DualTime Mamba structure with an omnidirectional scanning path, adjust its input matrix between dual-temporal image features in the deep scale, and allow the model to fully consider the spatiotemporal dependence and image structure information of the previous and next time points. In addition, this article designs a land-cover feature extraction (LCFE) method to improve the perception ability of the ground object target structure. Specifically, this method refines the image edge by separating high frequencies, adjusts the contour structure information of the dual-phase image by separating low frequencies, and then further models the relationship between pixels by combining feature structures and spatial offset mechanisms. Our experimental results on three datasets demonstrate the superiority of GSSR-Net. The network code address ishttps://github.com/SparrowTought/GSSR-Net.
Shuo Wang 0019, Genji Yuan, Jinjiang Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Dual-Branch Network for Spatial-Channel Stream Modeling Based on the State-Space Model for Remote Sensing Image Segmentation
abstract
To quickly and effectively address the color similarity issue in remote sensing image segmentation, traditional channel attention methods typically use channel modeling based on global channel statistics mapped to weights. However, this approach either suffers from limitations in feature selection due to a lack of dynamic interactions and the loss of significant spatial information, resulting in poor performance in complex scenarios, or has high computational complexity, making it difficult to apply in high-resolution remote sensing images. To overcome these challenges, this article proposes an innovative streaming channel modeling method based on state-space models (SSMs), aimed at rapidly and efficiently tackling the color similarity problem in remote sensing image segmentation. Specifically, we designed the channel-position state-space model network (CPSSNet) framework, where the decoder comprises the spatial Mamba block (SMB) for spatial modeling and the channel Mamba block (CMB) for streaming channel modeling. The core component of SMB, position-selective-scan-2D, achieves multidirectional global modeling in the spatial domain through a combination of the spatial scanning algorithm and SSM, with linear complexity. The core component of CMB, channel-2D selective-scan (C-SS2D), fuses channel and spatial information into patches for streaming modeling using a combination of the channel scanning (CS) algorithm and SSM. We have further improved SSM within C-SS2D to enhance dynamic interactions between channels, allowing for more refined modeling while maintaining linear complexity. Experimental results demonstrate that CPSSNet exhibits outstanding performance in addressing color similarity challenges in remote sensing image segmentation. The code is available athttps://github.com/yysdck/CPSSNet.
Yunsong Yang, Genji Yuan, Jinjiang Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 SPRMamba: A Mamba-Based Saliency Proportion Reconciliatory Network With Squeezed Windows for Remote Sensing Change Detection
abstract
Remote sensing (RS) change detection (CD) faces challenges in effectively identifying non-salient change regions, such as subtle architectural modifications or changes that closely resemble the background. The primary difficulties stem from the weak feature representation of non-salient changes, which results in insufficient model response, and the high similarity between background and change regions, leading to misdetection or omission. To address this issue, we propose SPRMamba, a Mamba-based saliency proportion reconciliatory network with squeezed windows for remote sensing change detection. It introduces state space models with windowing operations and cross-window interaction mechanisms to improve the response to weak signals. To dynamically balance the representation of salient and non-salient features, we design the saliency proportion reconciler (SPR) to optimize the discrimination between background and change regions. In addition, we introduce a sparse saliency loss function, which imposes sparsity constraints on salient regions to enhance the feature representation of non-salient change regions. Experimental results show that SPRMamba significantly outperforms existing methods on several public datasets. Our code will be available at https://github.com/boomstarzzn/SPRMamba.
Shengning Zhou, Chengpei Xu, Jinjiang Li 0001, Zhen Hua, Jingchun Zhou
IEEE Trans. Geosci. Remote. Sens.4
2025 Transmission map-guided deep unfolding network for underwater image enhancement
Baocai Chang, Peng Duan 0004, Jinjiang Li 0001
J. Supercomput.3
2025 DFGIC-Net: diffusion feature-guided information complementary network for infrared and visible light fusion
Yekai Cui, Peng Duan 0004, Jinjiang Li 0001
J. Supercomput.3
2024 Frequency-aware robust multidimensional information fusion framework for remote sensing image segmentation
Junyu Fan, Jinjiang Li 0001, Yepeng Liu 0003, Fan Zhang 0045
Eng. Appl. Artif. Intell.2
2024 BADM: Boundary-Assisted Diffusion Model for Skin Lesion Segmentation
Zhenyang Huang, Jianjun Li 0011, Jinjiang Li 0001
Eng. Appl. Artif. Intell.4
2024 Transformer-based multi-attention hybrid networks for skin lesion segmentation
Zhiwei Dong, Jinjiang Li 0001, Zhen Hua
Expert Syst. Appl.2
2024 Diffusion model-based text-guided enhancement network for medical image segmentation
Zhiwei Dong, Genji Yuan, Zhen Hua, Jinjiang Li 0001
Expert Syst. Appl.4
2024 DBEF-Net: Diffusion-Based Boundary-Enhanced Fusion Network for medical image segmentation
Zhenyang Huang, Jianjun Li 0011, Genji Yuan, Jinjiang Li 0001
Expert Syst. Appl.5
2024 CrossWaveNet: A dual-channel network with deep cross-decomposition for Long-term Time Series Forecasting
Siyuan Huang 0006, Yepeng Liu 0003, Fan Zhang 0045, Jinjiang Li 0001, Caiming Zhang 0001
Expert Syst. Appl.5
2024 DUCD: Deep Unfolding Convolutional-Dictionary network for pansharpening remote sensing image
Genji Yuan, Jinjiang Li 0001
Expert Syst. Appl.3
2024 MEAformer: An all-MLP transformer with temporal external attention for long-term time series forecasting
Siyuan Huang 0006, Yepeng Liu 0003, Haoyi Cui, Fan Zhang 0045, Jinjiang Li 0001, Xiaofeng Zhang 0003, Caiming Zhang 0001
Inf. Sci.5
2024 Dual branch Transformer-CNN parametric filtering network for underwater image enhancement
Baocai Chang, Jinjiang Li 0001, Zheng Chen 0018
J. Vis. Commun. Image Represent.2
2024 Multi-Frequency Field Perception and Sparse Progressive Network for low-light image enhancement
Jinjiang Li 0001, Zheng Chen 0018
J. Vis. Commun. Image Represent.2
2024 Elevation Information-Guided Multimodal Fusion Robust Framework for Remote Sensing Image Segmentation
abstract
Currently, the task of remote sensing image segmentation still faces some challenges, such as variations in illumination, shadows, and occlusions present in remote sensing images. Additionally, there may be similarities and confusions between different types of terrain features. In this paper, we aim to explore how to utilize information exchange between multiple modalities to reduce the impact of interfering factors. To fully exploit the complementary information between different modalities, we establish an information exchange mechanism between optical images (visible light + infrared) features and Digital Surface Model (DSM) features. This allows them to interact and express themselves in a shared feature space, facilitating the acquisition of complementary information from different modalities. Furthermore, through a multimodal fusion encoder and decoder based on Transformer design, the optical features and DSM features are integrated, enabling the learning of high-level semantic representations in different dimensions. Extensive subjective, objective comparative experiments, and ablation experiments are conducted on the ISPRS Vaihingen and Potsdam datasets to evaluate the proposed method. The mIoU on the Vaihingen and Potsdam datasets reached 85.06% and 87.6% respectively, while the OA reached 92.01% and 91.92% respectively. The source code will be available at https://github.com/JunyuFan/MIEFNet.
Junyu Fan, Jinjiang Li 0001, Zhen Hua, Fan Zhang 0045, Caiming Zhang 0001
IEEE Geosci. Remote. Sens. Lett.2
2024 MFDS-Net: Multiscale Feature Depth-Supervised Network for Remote Sensing Change Detection With Global Semantic and Detail Information
abstract
Change detection as an interdisciplinary discipline in the field of computer vision and remote sensing at present has been receiving extensive attention and research. Due to the rapid development of society, the geographic information captured by remote sensing satellites is changing faster and more complex, which undoubtedly poses a higher challenge and highlights the value of change detection tasks. We propose a multiscale feature depth-supervised network (MFDS-Net) for remote sensing change detection with global semantic and detail information (MFDS-Net) with the aim of achieving a more refined description of changing buildings as well as geographic information, enhancing the localization of changing targets and the acquisition of weak features. To achieve the research objectives, we use a modified$\text {ResNet}_{34}$as a backbone network to perform feature extraction. We propose the global semantic enhancement module (GSEM) to enhance the processing of high-level semantic information from a global perspective. The differential feature integration module (DFIM) is proposed to strengthen the fusion of different depth feature information, achieving learning and extraction of differential features. The entire network is trained and optimized using a deep supervision mechanism. The experimental outcomes of MFDS-Net surpass those of current mainstream change detection networks. On the LEVIR dataset, it achieved an F1 score of 91.589 and an IoU of 84.483. The code is available athttps://github.com/AOZAKIiii/MFDS-Net.
Zhenyang Huang, Zhaojin Fu, Jintao Song, Genji Yuan, Jinjiang Li 0001
IEEE Geosci. Remote. Sens. Lett.5
2024 Attention based dual UNET network for infrared and visible image fusion
Zhen Hua, Jinjiang Li 0001
Multim. Tools Appl.3
2024 SFPN: segmentation-based feature pyramid network for multi-focus image fusion
Limai Jiang, Ying Li 0067, Jinjiang Li 0001
Multim. Tools Appl.5
2024 LBP-based multi-scale feature fusion enhanced dehazing networks
Ying Li 0067, Jinjiang Li 0001, Zhen Hua
Multim. Tools Appl.3
2024 ConMamba: CNN and SSM High-Performance Hybrid Network for Remote Sensing Change Detection
abstract
Accurate remote sensing change detection (RSCD) tasks rely on comprehensively processing multiscale information from local details to effectively integrate global dependencies. Hybrid models based on convolutional neural networks (CNNs) and Transformers have become mainstream approaches in RSCD due to their complementary advantages in local feature extraction and long-term dependency modeling. However, the Transformer faces application bottlenecks due to the high secondary complexity of its attention mechanism. In recent years, state-space models (SSMs) with efficient hardware-aware design, represented by Mamba, have gained widespread attention for their excellent performance in long-series modeling and have demonstrated significant advantages in terms of improved accuracy, reduced memory consumption, and reduced computational cost. Based on the high match between the efficiency of SSM in long sequence data processing and the requirements of the RSCD task, this study explores the potential of its application in the RSCD task. However, relying on SSM alone is insufficient in recognizing fine-grained features in remote sensing images. To this end, we propose a novel hybrid architecture, ConMamba, which constructs a high-performance hybrid encoder (CS-Hybridizer) by realizing the deep integration of the CNN and SSM through the feature interaction module (FIM). In addition, we introduce the spatial integration module (SIM) in the feature reconstruction stage to further enhance the model’s ability to integrate complex contextual information. Extensive experimental results on three publicly available RSCD datasets show that ConMamba significantly outperforms existing techniques in several performance metrics, validating the effectiveness and foresight of the hybrid architecture based on the CNN and SSM in RSCD.
Zhiwei Dong, Genji Yuan, Zhen Hua, Jinjiang Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 Multiscale Attention Fusion Graph Network for Remote Sensing Building Change Detection
abstract
With the development of imaging systems and satellite technology, higher quality high-resolution RS images are being applied in building change detection (BCD) techniques. Methods based on convolutional neural network (CNN) have achieved excellent success in BCD techniques due to their excellent feature discrimination ability. However, CNN relies heavily on the geometry of prior conditions and is limited by the size of the convolution kernel, making it easy to ignore global information. This makes it difficult to capture the long-range dependence of different building targets and handle complex spatial relationships in high-resolution satellite RS images. Considering that graph convolutional neural networks (GCN) have powerful internal relationship learning capabilities, we propose a multi-scale attention fusion graph network (MAFGNet) in this paper. MAFGNet uses a dual graph convolution module (DGM), which includes a spatial graph convolution network (SGCN) and a channel graph convolution network (CGCN), to effectively explore the long-range relationship between the detection target and the global at the spatial and channel levels. We also design a multi-scale attention fusion encoder that includes channel and spatial attention fusion modules to effectively combine valuable information from multi-scale features. In addition, an atrous context self-attention pyramid (ACSP) is designed to combine multi-scale context to enhance the feature representation of change information. We conducted qualitative and quantitative comparative experiments on different datasets to validate the effectiveness of our model. The experimental results show that our method performs better than advanced methods in terms of overall accuracy and visualization details. Our code is available at https://github.com/ShangGY805/MAFG.
Yu Shangguan, Jinjiang Li 0001, Zheng Chen 0018, Zhen Hua
IEEE Trans. Geosci. Remote. Sens.2
2024 Attention Filtering Network Based on Branch Transformer for Change Detection in Remote Sensing Images
abstract
The emergence of high-resolution (HR) remote sensing imagery showcases the continual advancements in remote sensing technology but also sets higher demands for related tasks in the field, including remote sensing image change detection. Due to their outstanding performance in extracting salient features, convolutional neural networks (CNNs) have played a significant role and become widely utilized in many computer vision tasks. The encoder–decoder structure has confirmed the effectiveness of integrating multilevel feature information, as it allows for the synthesis of both local and global information of features. The exploration of the potential relationships between multilevel features and their efficient integration remains of significant importance. Furthermore, thanks to the advent of the transformer, many modern approaches have seen great improvements in high-level semantic understanding of images. In this article, we propose an attention-filtering network based on a branch transformer for effective change detection in remote sensing images. A hybrid attention fusion module (HAFM) is used to efficiently fuse features of different granularities and perform progressive information filtering on the extracted multilevel features to obtain an effective change feature. We also propose a branch transformer block (BTB) to efficiently aggregate global long-range dependencies and spatial details from the change feature. Extensive comparative experiments conducted on three different HR remote sensing datasets have verified the effectiveness of our method.
Yu Shangguan, Jinjiang Li 0001, Yepeng Liu 0003, Fan Zhang 0045, Caiming Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Context Spatial Awareness Remote Sensing Image Change Detection Network Based on Graph and Convolution Interaction
abstract
Remote sensing images are characterized by high dimensionality, complex textures, and large scales. Traditional Convolutional Neural Network (CNN) methods may overlook spatial relationships and contextual information among pixels when dealing with remote sensing data. Therefore, Graph Convolutional Networks (GCN) have emerged as a promising solution. In this paper, we propose a Contextual Spatial Awareness Remote Sensing Image Change Detection Network Based on Graph and Convolution interaction (CSAGC). We aim to enhance the handling of contextual information by introducing multiple augmentation modules. In CSAGC, we propose a high-performance encoder called Congraph that integrates a CNN and a Graph Neural Network (GNN). By preserving the respective features of both branches, we effectively fuse local detailed features and global positional features, achieving superior feature extraction capabilities. Additionally, we design two modules to facilitate the integration of multiscale spatial information: Contextual Spatial Awareness Module (CSAM) and Spatial Integration Module (SIM). CSAM, a crucial module connecting the encoder and decoder, jointly explores contextual features using the current feature branch and high-low level feature branches, leveraging spatial positional information for better content acquisition. SIM, located in the decoder module, aims to integrate the multiscale information outputted by CSAM, complementing the contextual information and improving the overall network’s ability to capture spatial contextual information. We conducted extensive experiments on three datasets, namely LEVIR-CD, WHU-CD, and GZ-CD. The experimental results demonstrate that CSAGC exhibits excellent performance, achieving significant performance improvements compared to state-of-the-art (SOTA) methods.
Xinyang Song, Zhen Hua, Jinjiang Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 DUDB: Deep Unfolding-Based Dual-Branch Feature Fusion Network for Pan-Sharpening Remote Sensing Images
abstract
The proposed method aims to enhance the fusion of high-resolution multispectral (MS) images (HRMS) by extracting spatial and spectral features from panchromatic (PAN) images and MS images. However, existing pan-sharpening methods often suffer from the problem of missing spatial and spectral detail information. To better preserve these details, we introduce a dual-branch feature fusion pan-sharpening network based on deep unfolding. In this network, we utilize the algorithm unfolding iterative module (AUIF-Block) to continuously acquire detailed information from both MS and PAN images for image reconstruction. By leveraging the adaptive channel and spatial feature enhancement module (DEM-Block), the network can adjust spatial and channel features adaptively, leading to more accurate feature extraction and more complete image reconstruction. Finally, the detail-based fusion module (DBFM-Block) is employed to integrate and enrich the content of detailed information extracted from different channels, resulting in improved fusion performance. Experiments were conducted on QuickBird (QB) and WorldView-2 (WV2) datasets. Through qualitative analysis and quantitative comparisons, we demonstrate that this method outperforms existing approaches.
Hailin Tao, Jinjiang Li 0001, Zhen Hua, Fan Zhang 0045
IEEE Trans. Geosci. Remote. Sens.2
2024 SFFNet: A Wavelet-Based Spatial and Frequency Domain Fusion Network for Remote Sensing Segmentation
abstract
To fully utilize spatial information for segmentation and address the challenge of handling areas with significant grayscale variations in remote sensing segmentation, we propose the spatial and frequency domain fusion network (SFFNet) framework. This framework employs a two-stage network design: the first stage extracts features using spatial methods to obtain features with sufficient spatial details and semantic information; the second stage maps these features in both spatial and frequency domains. In the frequency domain mapping, we introduce the wavelet transform feature decomposer (WTFD) structure, which decomposes features into low-frequency and high-frequency components using the Haar wavelet transform and integrates them with spatial features. To bridge the semantic gap between frequency and spatial features, facilitating significant feature selection to promote the combination of features from different representation domains, we design the multiscale dual-representation alignment filter (MDAF). This structure utilizes multiscale convolutions and dual-cross attentions. Comprehensive experimental results demonstrate that, compared to existing methods, SFFNet achieves superior performance in terms of mean intersection over union (mIoU), reaching 84.80% and 87.73%, respectively. The code is located athttps://github.com/yysdck/SFFNet.
Yunsong Yang, Genji Yuan, Jinjiang Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Reference-based dual-task framework for motion deblurring
Cunzhe Liu, Zhen Hua, Jinjiang Li 0001
Vis. Comput.3
2023 Deep supervision feature refinement attention network for medical image segmentation
Zhaojin Fu, Jinjiang Li 0001, Zhen Hua, Linwei Fan
Eng. Appl. Artif. Intell.2
2023 MFBGR: Multi-scale feature boundary graph reasoning network for polyp segmentation
Fangjin Liu, Zhen Hua, Jinjiang Li 0001, Linwei Fan
Eng. Appl. Artif. Intell.3
2023 DPCTN: Dual path context-aware transformer network for medical image segmentation
Jinjiang Li 0001
Eng. Appl. Artif. Intell.3
2023 Joint transformer progressive self-calibration network for low light enhancement
abstract
Abstract When the lighting conditions are poor and the environmental light is weak, the image captured by the imaging device often has lower brightness and is accompanied by a lot of noise. The paper designs a progressive self‐calibration network model (PSCNet) for recovering high‐quality low‐light‐enhanced images. First, shallow features in low‐light images can be better focused and extracted with the help of attention mechanism. Next, the feature mapping is passed to the encoder and decoder modules, where the transformer and encoder‐decoder jump connection structures can be better combined with the semantic information of the context to learn rich deep feature information. Finally, the self‐calibration module can adaptively cascade the features decoded by the decoder and input them into the residual attention module quickly and accurately. Meanwhile, the LBP features of the image are also fused into the feature information of the residual attention module to enhance the detailed texture information of the image. Qualitative analysis and quantitative comparison of a large number of experimental results show that this method outperforms existing methods.
Junyu Fan, Jinjiang Li 0001, Zhen Hua, Linwei Fan
IET Image Process.2
2023 Multi-scale dynamic fusion for correcting uneven illumination images
Junyu Fan, Jinjiang Li 0001, Zheng Chen 0018
J. Vis. Commun. Image Represent.2
2023 Filter-cluster attention based recursive network for low-light enhancement
abstract
The poor quality of images recorded in low-light environments affects their further applications. To improve the visibility of low-light images, we propose a recurrent network based on filter-cluster attention (FCA), the main body of which consists of three units: difference concern, gate recurrent, and iterative residual. The network performs multi-stage recursive learning on low-light images, and then extracts deeper feature information. To compute more accurate dependence, we design a novel FCA that focuses on the saliency of feature channels. FCA and self-attention are used to highlight the low-light regions and important channels of the feature. We also design a dense connection pyramid (DenCP) to extract the color features of the low-light inversion image, to compensate for the loss of the image’s color information. Experimental results on six public datasets show that our method has outstanding performance in subjective and quantitative comparisons.
Zhixiong Huang, Jinjiang Li 0001, Zhen Hua, Linwei Fan
Frontiers Inf. Technol. Electron. Eng.2
2023 Enhanced Feature Interaction Network for Remote Sensing Change Detection
abstract
In the current research on remote sensing image change detection, the effective learning of mutual interactions between bi-temporal features has often been overlooked. To address this concern, we introduce a Patch Exchange Block aimed at capturing the interplay between bi-temporal channels by exchanging feature patches. This approach preserves feature structural information and prevents the introduction of unnecessary noise. Specifically, the feature maps of bi-temporal images are unfolded into multiple patches, followed by mutual patch exchanges and subsequent fusion operations. Additionally, we seek to leverage Transformers to tackle the model’s lack of effective global feature extraction capability. However, the standard Transformer aggregates features based on all query-key pairs, making the model susceptible to irrelevant features’ interference. Considering this, we introduce a Sparse Transformer in the decoder. It guides the model’s attention to areas of interest by selectively weighting the values produced by the Q and K operations, thus reducing interference from irrelevant information and focusing on the most valuable insights. Through experiments conducted on multiple datasets, we substantiate the effectiveness of our proposed EFIN approach.
Shike Liang, Zhen Hua, Jinjiang Li 0001
IEEE Geosci. Remote. Sens. Lett.3
2023 LHDACT: Lightweight Hybrid Dual Attention CNN and Transformer Network for Remote Sensing Image Change Detection
abstract
With the significant advancements of Deep Learning (DL) in the field of remote sensing imagery, a plethora of Change Detection (CD) methods based on CNNs, attention mechanisms, and transformers have emerged. Presently, a substantial amount of research has gradually relinquished control over parameter quantities in pursuit of enhanced outcomes, resulting in the inflation of networks with numerous stacked modules. This paper is dedicated to integrating lightweight approaches into the CD task.We introduce a Lightweight Hybrid Dual-Attention CNN and Transformer network (LHDACT) based on Depthwise Over-Parameterized Convolution (DO-Conv). In comparison to traditional convolution, DO-Conv combines both traditional and depthwise convolutions, achieving commendable performance enhancement with minimal additional cost. Furthermore, we leverage DO-Conv to enhance the Multi-Scale Average Pooling module (MSAP), ensuring global context with low computational overhead.To better discern regions of interest within complex images, we enhance the Dual Attention Module (DAM) by sharing weights across spatial and channel dimensions, thereby bolstering feature region identification. Lastly, we employ a compact transformer module to capture feature differences, enabling precise change detection CD. Our approach is evaluated on the LEVIR-CD, WHU-CD, and GZ-CD datasets, yielding F1 scores of 91.23%, 87.51%, and 85.32%, respectively. These results demonstrate high performance on a cost-effective scale.
Xinyang Song, Zhen Hua, Jinjiang Li 0001
IEEE Geosci. Remote. Sens. Lett.3
2023 Retinex low-light image enhancement network based on attention mechanism
Jinjiang Li 0001, Zhen Hua
Multim. Tools Appl.2
2023 Dual UNet low-light image enhancement network based on attention mechanism
Fangjin Liu, Zhen Hua, Jinjiang Li 0001, Linwei Fan
Multim. Tools Appl.3
2023 Multi-scale siamese networks for multi-focus image fusion
Zhen Hua, Jinjiang Li 0001
Multim. Tools Appl.3
2023 Attention-based dual-color space fusion network for low-light image enhancement
Zhixiong Huang, Jinjiang Li 0001, Zhen Hua, Linwei Fan
Signal Process. Image Commun.2
2023 Multi-Scale Video Super-Resolution Transformer With Polynomial Approximation
abstract
Video super-resolution techniques aim to obtain high-resolution equivalents of existing low-resolution videos through a series of operations. In recent research, transformers have been increasingly popular because of their remarkable abilities in parallel computing and efficient extraction of space-time sequence features from videos. Moreover, combining self-attention and multi-scale methods has yielded excellent results. However, the combination of the two methods has limitations, current up-sampling methods struggle to match the global modeling capacity of self-attention mechanisms. Therefore, this paper proposes three strategies to combine the two methods. Based on the approximation strategy, we first construct a new bilinear up-sampling method for multi-scale acquisition. Convolution and cross-attention techniques are then used to correct and align features at different scales to prevent large deviations in feature extraction at a specific scale, which can affect subsequent feature extraction. Finally, to effectively solve the common computational complexity,$ C^{0} $continuity, and neuron death problems of existing activation functions, a new method to construct the activation function is proposed. The cubic spline function is used to construct a new activation function approximating tanh. The new activation function is$ C^{2} $continuous, which is piecewise defined by cubic polynomial curves. In this study, better results were achieved on three public video super-resolution test sets: REDS4, Vid4, and Vimeo-90K-T. Experiments demonstrated that the proposed method could provide a new solution for video super-resolution tasks.
Fan Zhang 0045, Gongguan Chen, Hua Wang 0012, Jinjiang Li 0001, Caiming Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 SLDDNet: Stagewise Short and Long Distance Dependency Network for Remote Sensing Change Detection
abstract
With the rapid development of society, the pace of land change continues to accelerate. Consequently, remote sensing change detection has become a vital method for monitoring geographical information changes across various domains. However, the increasingly diverse and complex environments and structures where change targets exist pose significant challenges to change detection tasks. To address these challenges, we propose a novel approach called Stage-wise Short and Long Distance Dependency Network (SLDDNet). SLDDNet uses CNN-Transformer architecture and proposes the Transformer Semantic Selector to capture long-range feature dependencies and enhance semantic associations from a global perspective. Moreover, the Pyramid Structure Feature Stacking is proposed to capture short-range feature dependencies and emphasize local feature information. By integrating these two types of features at each layer, SLDDNet focuses on semantic information and improves its ability to attend to feature details. Furthermore, SLDDNet enhances target position information through Axial Semantic Enhancement and optimizes the network training process using a deep supervision mechanism. Through extensive experiments, SLDDNet outperforms mainstream and state-of-the-art methods on three datasets. Specifically, on the LEVIR-CD dataset, SLDDNet achieves an F1 score of 91.75% and an IoU of 84.76%. On the WHU-CD dataset, it achieves an F1 score of 92.76% and an IoU of 85.78%. Lastly, on the GZ-CD dataset, SLDDNet achieves an F1 score of 86.61% and an IoU of 76.38%. For those interested, the code for SLDDNet is available at https://github.com/Fuzhaojin/SLDDNet.
Zhaojin Fu, Jinjiang Li 0001, Zheng Chen 0018
IEEE Trans. Geosci. Remote. Sens.2
2023 CADUI: Cross-Attention-Based Depth Unfolding Iteration Network for Pansharpening Remote Sensing Images
abstract
Pansharpening is an important technology for remote sensing imaging systems to obtain high-resolution multispectral (HRMS) images. It mainly obtains high-resolution multi-spectral (HRMS) images with uniform spectral distribution and rich spatial details by fusing low-resolution multi-spectral (LRMS) images and high-spatial-resolution panchromatic (PAN) images. Therefore, how to extract features completely and reconstruct images with high quality is critical to obtain ideal fusion images. In this paper, we propose a new pansharpening method, called the Cross Attention-based Depth Unfolding Iteration Network for Pan-sharpening remote sensing images (CADUI), which achieves the desired fusion effect by iteratively optimizing the deep prior regularization and combining it with a cross-attention mechanism. The network consists of two parts: optimized iterations of deep prior regularization (DEIN-Block) and cross-attention mechanism (CAFM-Block). Among them, DEIN-Block introduces the depth prior as an implicit regularization and improves the adaptability and representation ability of the relevant data of the reconstructed image through iteration. CAFM-Block realizes dual-branch fusion through cross-attention fusion and channel-attention fusion to achieve better fusion results. Simulation experiments and real experiments are carried out on the standard datasets QuikBird (QB) and WorldView-2 (WV2). Through quantitative comparison and qualitative analysis, it is proved that the method is superior to the existing methods.
Jinjiang Li 0001, Fan Zhang 0045, Linwei Fan
IEEE Trans. Geosci. Remote. Sens.2
2023 CTMFNet: CNN and Transformer Multiscale Fusion Network of Remote Sensing Urban Scene Imagery
abstract
Semantic segmentation of remotely sensed urban scene images is widely demanded in areas such as land cover mapping, urban change detection, and environmental protection. With the development of deep learning, methods based on convolutional neural networks (CNNs) have been dominant due to their powerful ability to represent hierarchical feature information. However, the limitations of the convolution operation itself limit the network’s ability to extract global contextual information. With the successful use of transformer in computer vision in recent years, transformer has shown great potential for modeling global contextual information. However, transformer is not sufficiently capable of capturing local detailed information. In this article, to explore the potential of the joint CNN and transformer mechanism for semantic segmentation of remotely sensed urban scenes, we propose a CNN and transformer multiscale fusion network (CTMFNet) based on encoding–decoding for urban scene understanding. To couple local–global context information more efficiently, we designed a dual backbone attention fusion module (DAFM) to couple the local and global context information of the dual-branch encoder. In addition, to bridge the semantic gap between scales, we built a multi-layer dense connectivity network (MDCN) as our decoder. The MDCN enables the full flow of semantic information between multiple scales to be fused with each other through upsampling and residual connectivity. We conducted extensive subjective and objective comparison experiments and ablation experiments on both the International Society of Photogrammetry and Remote Sensing (ISPRS) Vaihingen and ISPRS Potsdam datasets. Numerous experimental results have proven the superiority of our method compared to currently popular methods.
Jinjiang Li 0001, Zhiyong An, Linwei Fan
IEEE Trans. Geosci. Remote. Sens.2
2023 AMCA: Attention-Guided Multiscale Context Aggregation Network for Remote Sensing Image Change Detection
abstract
Remote sensing image change detection is the key to understanding surface changes. Although the existing change detection methods have achieved good results, some structural details are missing and the detection accuracy needs to be improved. Therefore, we propose an attention-guided multi-scale context aggregation network (AMCA) for remote sensing image change detection. First, we use the fully attentional pyramid module (FAPM) to enhance the deep feature information of the original image. And we introduce the dense feature fusion module (DFFM) to fully fuse the bi-temporal features to obtain the change regions. Second, the introduction of channel-wise cross fusion transformer (CCT) and channel-wise cross attention (CCA) not only can effectively fuse channel features focusing on different semantic patterns, but also bridge the semantic gap between multi-scale features. Next, we use the transformer decoder to map the learned high-level semantic information into the pixel space to refine the original features. In addition, we use the context extraction module (CEM) to obtain the local and global associations of feature maps. Finally, the addition of attention aggregation module (AAM) can effectively combine the feature information at different scales. Extensive experiments on three public change detection datasets show that the proposed method has advantages over other methods in terms of both visual interpretation and quantitative analysis.
Xintao Xu, Jinjiang Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 Mutiscale Hybrid Attention Transformer for Remote Sensing Image Pansharpening
abstract
Pansharpening methods play a crucial role for remote sensing image processing. The existing pansharpening methods, in general, have the problems of spectral distortion and lack of spatial detail information. To mitigate these problems, we propose a multiscale hybrid attention Transformer pansharpening network (MHATP-Net). In the proposed network, the shallow feature (SF) is first acquired through an SF extraction module (SFEM), which contains the convolutional block attention module (CBAM) and dynamic convolution blocks. The CBAM in this module can filter initial information roughly, and the dynamic convolution blocks can enrich the SF information. Then, the multiscale Transformer module is used to obtain multiencoding feature images. We propose a hybrid attention module (HAM) in the multiscale feature recovery module to effectively address the balance between the spectral feature retention and the spatial feature recovery. In the training process, we use deep semantic statistics matching (D2SM) loss to optimize the output model. We have conducted extensive experiments on several known datasets, and the results show that this article has good performance compared with other state of the art (SOTA) methods.
Wengang Zhu, Jinjiang Li 0001, Zhiyong An, Zhen Hua
IEEE Trans. Geosci. Remote. Sens.2
2023 Attention-based adaptive feature selection for multi-stage image dehazing
Zhen Hua, Jinjiang Li 0001
Vis. Comput.3
2023 Cross-UNet: dual-branch infrared and visible image fusion framework based on cross-convolution and attention mechanism
Zhen Hua, Jinjiang Li 0001
Vis. Comput.3
2023 FLA-Net: multi-stage modular network for low-light image enhancement
Nana Yu, Jinjiang Li 0001, Zhen Hua
Vis. Comput.2
2022 Pyramid-attention based multi-scale feature fusion network for multispectral pan-sharpening
Yang Chi, Jinjiang Li 0001
Appl. Intell.2
2022 A multi-focus image fusion method based on attention mechanism and supervised learning
Limai Jiang, Jinjiang Li 0001
Appl. Intell.3
2022 Progressive edge-sensing dynamic scene deblurring
abstract
Deblurring images of dynamic scenes is a challenging task because blurring occurs due to a combination of many factors. In recent years, the use of multi-scale pyramid methods to recover high-resolution sharp images has been extensively studied. We have made improvements to the lack of detail recovery in the cascade structure through a network using progressive integration of data streams. Our new multi-scale structure and edge feature perception design deals with changes in blurring at different spatial scales and enhances the sensitivity of the network to blurred edges. The coarse-to-fine architecture restores the image structure, first performing global adjustments, and then performing local refinement. In this way, not only is global correlation considered, but also residual information is used to significantly improve image restoration and enhance texture details. Experimental results show quantitative and qualitative improvements over existing methods.
Jinjiang Li 0001
Comput. Vis. Media2
2022 TPET: Two-stage Perceptual Enhancement Transformer Network for Low-light Image Enhancement
Hengshuai Cui, Jinjiang Li 0001, Zhen Hua, Linwei Fan
Eng. Appl. Artif. Intell.2
2022 DPCFN: Dual path cross fusion network for medical image segmentation
Shen Jiang, Jinjiang Li 0001, Zhen Hua
Eng. Appl. Artif. Intell.2
2022 Underwater image enhancement via LBP-based attention residual network
abstract
Abstract Owing to the influence of light absorption and scattering in underwater environments, underwater images exhibit color deviation, low contrast and detail blur, and other degradations. This paper proposes an underwater image enhancement method combining a residual convolution network, local binary pattern (LBP), and self‐attention mechanism. The LBP operator processes the input underwater images. The LBP feature images and underwater images thus obtained constitute the network input. The network consists of three modules: a color correction module to remove the color deviation in underwater images, detail repair module to restore the integrity of details, and an LBP auxiliary enhancement module for global enhancement of image details. The correction and repair modules generate the correct color image and detailed supplement images, respectively. The final‐result image is obtained by superpositioning the two generated images. The experimental results confirm that our method can reproduce the bright colors and complete details of the visual effect, showing a significant improvement over other advanced methods in quantitative evaluation.
Zhixiong Huang, Jinjiang Li 0001, Zhen Hua
IET Image Process.2
2022 Two-stage single image dehazing network using swin-transformer
abstract
Abstract Hazy images often have color distortion, blur and other visible visual quality degradation, affecting the performance of some advanced visual tasks. Therefore, single image dehazing has always been a challenging and significant problem. Convolutional neural network has been widely used in image dehazing task, but the limitations of convolutional operation limit the development of dehazing task. Nowadays, Transformer offers a holistic approach to CV development and does not grow in location as the network deepens. For this reason, a hierarchical Transformer is introduced for use in the dehazing network. Specifically, the codec is improved and Transformer and CNN are combined to achieve basic feature extraction in the first stage. The encoder only models the global relationship at each layer, reducing the resolution of the feature map continuously and expanding the field of perception. In addition, an inter‐block supervision mechanism is added between encoder unit and decoder unit to refine features and supervise and select them, thus improving the efficiency of feature transmission. In the second stage, the original resolution block is used to extract the local features, and then feature fusion and interaction are carried out. In addition, to ensure the authenticity of the transmission of characteristic signals in the first stage and improve the transmission efficiency of the network, fusion attention mechanism is added between stages. It adds the residual image of the early input features to the image acquired in the first stage, then passes to the next stage. Ablation experiments show that the two‐stage network has significant benefits for image quality and visual effects. The experimental results on RESIDE, O‐Haze, and I‐Haze datasets show that the method is superior to advanced methods in dehazing effectiveness.
Zhen Hua, Jinjiang Li 0001
IET Image Process.3
2022 Two-stage progressive residual learning network for multi-focus image fusion
abstract
Abstract Boundary artifacts and color detail distortion are easily caused by the common multi‐focus image fusion methods. In order to solve this problem, we propose a two‐stage progressive residual learning network for multi‐focus image fusion. The proposed network can progressively learn color information and detail features through end‐to‐end mapping. The whole network is composed of two sub‐networks: the initial fusion block network and the enhanced fusion block network. First, the color information in the source image is fused by the initial fusion block network to generate the initial fusion image. Then on the basis of the initial fusion image, the detailed features of the source image are further fused by enhanced the fusion network to form the final fusion image. In order to solve the problem of lack of groundtruth when multi‐focus image fusion is carried out with supervised method, the multi‐focus image fusion problem is compared to the easy‐to‐solve image restoration problem. A synthetic dataset for network training is generated by ”degenerating” the VOC2012 dataset according to the set rules. After training, the method works well for fusion tasks without further processing. Experimental results show that the proposed method is superior to the existing methods in subjective visual perception and objective quantitative evaluation.
Zhen Hua, Jinjiang Li 0001
IET Image Process.3
2022 Attention-based multi-channel feature fusion enhancement network to process low-light images
abstract
Abstract In realistic low‐light environments, images captured by imaging devices often have problems such as low brightness and low contrast, serious loss of detail information, and a large amount of noise, posing major challenges to computer vision tasks. Low‐light image enhancement can effectively improve the overall quality of the image, which has important significance and application value. In this study, an attention‐based multi‐channel feature fusion enhancement network (M‐FFENet) is proposed to process low‐light images. In this network, a feature extraction model is first used to obtain the deep features of the downsampled low‐light images and fit them to an affine bilateral grid. Second, the addition of attention‐based residual dense blocks (ARDB) allows the network to focus on more details and spatial information. Meanwhile, all color channels are considered. The channel features and bilateral meshes are then linearly interpolated using the feature reconfiguration model (FRM) to obtain high‐quality features containing rich color and texture information. Next, the feature fusion module (FFM) is used to fuse features that contain different information. Enhancement model is used to further recover texture and detail in the image. Finally, the enhanced image is output. Numerous experimental results have shown that the method achieves better results in both quantitative and qualitative aspects compared to other methods.
Xintao Xu, Jinjiang Li 0001, Zhen Hua, Linwei Fan
IET Image Process.2
2022 LBP-based progressive feature aggregation network for low-light image enhancement
abstract
Abstract At night or in other low‐illumination environments, optical imaging devices cannot capture details and color information in images accurately because of the reduced number of photons captured and the low signal‐to‐noise ratio. Consequently, the image is very noisy with low contrast and inaccurate color information, which affects human visual perception and creates significant challenges in computer vision tasks. Low‐light image enhancement has great research value because it aims to reduce image noise and improve image quality. In this study, we propose an LBP‐based progressive feature aggregation network (P‐FANet) for low‐light image enhancement. The LBP feature has insensitivity to illumination, and it contains rich texture information. In the network, we input the LBP feature into each iteration of the network in an accompanying manner, which helps to restore some detailed information of the low‐light image. First, we input the low‐light image into the dual attention mechanism model to extract global features. Second, the extracted different features enter the feature aggregation module (FAM) for feature fusion. Third, we use the recurrent layer to share the features extracted at different stages, and use the residual layer to further extract deeper features. Finally, the enhanced image is output. The rationality of the method in this study has been verified through ablation experiments. Many experimental results show that the method in this study has greater advantages in subjective and objective evaluations compared with many other advanced methods.
Nana Yu, Jinjiang Li 0001, Zhen Hua
IET Image Process.2
2022 DDFN: a depth-differential fusion network for multi-focus image
Limai Jiang, Jinjiang Li 0001
Multim. Tools Appl.3
2022 Color layers -Based progressive network for Single image dehazing
Zhen Hua, Jinjiang Li 0001
Multim. Tools Appl.3
2022 Detail enhancement decolorization algorithm based on rolling guided filtering
Nana Yu, Jinjiang Li 0001, Zhen Hua
Multim. Tools Appl.2
2022 Attention based dual path fusion networks for multi-focus image
Nana Yu, Jinjiang Li 0001, Zhen Hua
Multim. Tools Appl.2
2022 Decolorization algorithm based on contrast pyramid transform fusion
Nana Yu, Jinjiang Li 0001, Zhen Hua
Multim. Tools Appl.2
2022 Near-infrared shadow detection based on HDR image
Wanwan Zhang, Jinjiang Li 0001, Zhen Hua
Multim. Tools Appl.2
2022 Multi-level receptive field feature reuse for multi-focus image fusion
Limai Jiang, Jinjiang Li 0001
Mach. Vis. Appl.3
2022 MFFE: Multi-scale Feature Fusion Enhanced Net for image dehazing
Jinjiang Li 0001, Zhen Hua
Signal Process. Image Commun.2
2022 Remote Sensing Image Change Detection Transformer Network Based on Dual-Feature Mixed Attention
abstract
Change detection (CD) of high-resolution remote sensing (RS) images is a basic task in RS image processing tasks. In recent years, CD tasks have made many attempts in pure convolutional networks, attention mechanism, and transformer, and have achieved good results. Based on the power of attention and transformers, we hope to find a method that can handle the details of the image better and has better generalization ability. In this article, we propose a dual-feature mixed attention-based transformer network (DMATNet). First, we adopt a dual-feature extraction method, using a simple convolutional neural network (CNN) to extract coarse features, and a CNN based on progressive sampling to extract fine features. Then, we fuse the fine and coarse features with dual-feature mixed attention (DFMA) module. It can not only extract more specific regions of interest, but also overcome the misjudgment caused by oversampling, and synchronize feature extraction and target information integration. Finally, we use transformer to optimize these extracted information and feedback into the original features in the encoder to help remodel the pixel space. We merged the DMAT network into a deep feature difference-based CD framework and conducted extensive experiments on four datasets, LEVIR-CD, DSIFN-CD, WHU-CD, and CLCD, respectively, with tested F1 and interconnection over union (IoU) results of 90.75%/84.13%, 71.23%/55.32%, 85.70%/74.98%, and 66.56%/59.87%. Experimental results show that our DMAT-based model performs significantly better than the existing state-of-the-art attention and transformer-based methods.
Xinyang Song, Zhen Hua, Jinjiang Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Transformer-Based Regression Network for Pansharpening Remote Sensing Images
abstract
The pansharpening entails obtaining images with uniform spectral distribution and rich spatial details by fusing multispectral images and panchromatic images, which has become a major image fusion problem in the field of remote sensing. Convolutional neural networks are widely used in image processing. We propose a transformer-based regression network (DR-NET) architecture. The first stage was feature extraction, which entailed extracting spectral information and spatial details from multispectral images and panchromatic images. The second stage was feature fusion, which entailed integrating the extracted feature information. In the third stage, image reconstruction, images with uniform distribution of spectral information, and sufficient spatial details were obtained. The fourth stage entailed optimizing the network performance and calculating the loss of shallow feature image and the image result after downsampling during image reconstruction. The performance of the DR-NET was optimized by optimizing the sum of all the loss values, which could be considered double regression. Simulated and real data experiments were conducted on the GF-2, QuickBird, and WorldView2 datasets to compare the proposed method with classical pansharpening methods. The qualitative and quantitative analyses proved that the spectral distribution of the image pansharpened using our method was uniform, the spatial details were completely retained, and the evaluation indicators were also optimal, which fully demonstrated the superior performance of the DR-NET.
Xunyang Su, Jinjiang Li 0001, Zhen Hua
IEEE Trans. Geosci. Remote. Sens.2
2022 Attention-Based and Staged Iterative Networks for Pansharpening of Remote Sensing Images
abstract
The pansharpening method combines complementary features from panchromatic (PAN) images and multispectral (MS) images to provide high-resolution MS images. Therefore, how to extract the features completely and reconstruct the image with high quality is the key link to obtain the ideal fusion image. We propose an attention-based and staged iterative network (ASIN) framework, which considers each subnetwork of the iterative network as a multistage process of pansharpening and carries out feature extraction and image reconstruction in each stage. Using the advantages of an iterative network for cross-stage depth features for the hierarchical extraction of refined features of MS images and PAN images for image reconstruction, we use the large kernel attention (LKA) module and the cascaded asymmetric coupling representation module to build the framework for feature extraction and the attention fusion module (AFM) to fuse the features of PAN and MS in the image reconstruction stage. LKA has channel and spatial adaptability, as well as strong long-range dependency establishment capability, which makes the feature extraction more complete. Asymmetric coupled representation module (ACRM) outputs refined spectral and spatial features by learning the hybrid correlation of MS and PAN images. AFM effectively utilizes the spectral and spatial features of the input, enabling the network to reduce information loss and retain important information. On the QuickBird (QB), WorldView-2 (WV2), and Gaofen-2 (GF-2) datasets, the superior performance of our method over the contrasting methods is demonstrated by quantitative comparison and qualitative analysis.
Xunyang Su, Jinjiang Li 0001, Zhen Hua
IEEE Trans. Geosci. Remote. Sens.2
2022 Attention-Based Multistage Fusion Network for Remote Sensing Image Pansharpening
abstract
Pansharpening is a significant branch in the field of remote sensing image processing, the goal of which is to fuse panchromatic (PAN) and multispectral (MS) images through certain rules to generate high-resolution MS (HRMS) images. Therefore, how to improve the spatial and spectral resolutions of the fused image is the problem that we need to solve urgently. In this article, a multistage remote sensing image fusion network (MRFNet) is proposed on the basis of in-depth research and exploration on the fusion of the PAN and MS images to obtain a clear fused image that can reflect the ground features more comprehensively and completely. The proposed network consists of three stages that are connected by cross-stage fusion. The first two stages are used to extract the features of the PAN and MS images. The structure of the encoder–decoder and the channel attention module are used to extract the features of the remote sensing image in the channel domain. The third stage is the image reconstruction stage fusing the extracted features with the original image to improve the spatial and spectral resolutions of the fused result. A series of experiments are conducted on the benchmark datasets WorldView II, GF-2, and QuickBird. Qualitative analysis and quantitative comparison show the superiority of MRFNet in visual effects and the values of evaluation indicators.
Wanwan Zhang, Jinjiang Li 0001, Zhen Hua
IEEE Trans. Geosci. Remote. Sens.2
2021 Multi-scale depth information fusion network for image dehazing
Zhen Hua, Jinjiang Li 0001
Appl. Intell.3
2021 Low-light image enhancement based on multi-illumination estimation
Xiaomei Feng, Jinjiang Li 0001, Zhen Hua, Fan Zhang 0045
Appl. Intell.2
2021 Low-light image enhancement based on exponential Retinex variational model
abstract
Abstract Aiming at the problems of residual noise, low contrast, and limited detail information caused by low‐light images, this paper proposes a new Retinex variational model. According to Retinex theory, it is necessary to estimate the illumination and reflectance components decomposed from the original image. In order to better maintain the edge information, texture richness, and prevent artefacts, the exponential forms of local variation deviation and total variation are used as illumination prior and reflectance prior, respectively, and mixed norms are used to constrain them, so as to deal with the illumination information and texture details of the image more effectively, and then use the bright channel prior to improve the colour reproduction sense of the original image, thereby constructing the objective function, and finally using the alternating iterative optimization method to find the optimal solution to the proposed model. Experiments show that compared with other existing image enhancement methods, the method proposed here improves the contrast of the image, overcomes the phenomenon of halo artefacts and colour distortion, is more consistent with human vision, and produces better results in terms of quantitative performance.
Jinjiang Li 0001, Zhen Hua
IET Image Process.2
2021 Hierarchical guided network for low-light image enhancement
abstract
Abstract Due to insufficient illumination in low‐light conditions, the brightness and contrast of the captured images are low, which affect the processing of other computer vision tasks. Low‐light enhancement is a challenging task that requires simultaneous processing of colour, brightness, contrast, artefacts and noise. To solve this problem, the authors apply the deep residual network to the low‐light enhancement task, and propose a hierarchical guided low‐light enhancement network. The key of this method is recombined hierarchical guided features through the feature aggregation module to realize low‐light enhancement. The network is based on the U‐Net network, and then hierarchically guided with the input pyramid branch in the encoding and decoding network. The input pyramid structure realizes multi‐level receptive fields and generates a hierarchical representation. The encoding and decoding structure concatenates the hierarchical features of the input pyramid and generates a set of hierarchical features. Finally, the feature aggregation module is used to fuse different features to achieve low‐light enhancement tasks. The effectiveness of the components is proved through ablation experiments. In addition, the authors are also evaluating on different data sets, and the experimental results show that the method proposed is superior to other methods in subjective and objective evaluation.
Xiaomei Feng, Jinjiang Li 0001
IET Image Process.2
2021 Pseudo-Siamese residual atrous pyramid network for multi-focus image fusion
abstract
Abstract Depth of field is one of the critical reasons to limit the richness of image information. Usually, in a scene with multiple targets, when the distance between each target and the lens is different, the clear scene image can be get within a certain distance range. This situation restricts the further image processing, such as semantic segmentation, object recognition and 3D reconstruction. Multi‐focus image fusion uses two or more images focused on different targets to fuse scene information, which can solve this problem to a great extent. In general, two or more multi‐focus images can cover almost all near/far targets. The fusion of more than two multi‐focus images can be accomplished by cascading the fusion results of the previous two images and the next image to be processed many times. Therefore, the paper focus on the fusion of two multi‐focus images. Inspired by this, new Pseudo‐Siamese neural network with several residual atrous convolution pyramids with multi‐level perception ability to perceive the multi‐level features and consistency relations of multi‐focus image pairs is proposed, and multi‐layer residual blocks are used to fuse the extracted features. In this process, the residual of the groundtruth and the generated image will be learned. Finally, a fully focused image without blur will be generated. After several ablation experiments and comparison experiments with other methods, the results show that the performance of the method proposed in this paper is state‐of‐the‐art, and overall better than other methods, which are advanced.
Limai Jiang, Jinjiang Li 0001, Changhe Tu
IET Image Process.3
2021 Iterative multi-scale residual network for deblurring
abstract
Abstract In dynamic scene deblurring, recent neural network–based methods have been very successful. But with the improvement of deep deblurring performance, network structure and learning become more complicated. Compared with large‐scale network parameters and complex network structures, an iterative multi‐scale residual network to achieve a more effective parameter sharing scheme is proposed. In each iterative unit, fast multi‐scale residual blocks to replace superimposed convolutional layers or classic residual blocks are used. On the basis of preventing model overfitting, the receptive field of the network is increased. At the same time, the gated recurrent unit is introduced to connect modules of different stages. The model does not rely on the estimation of the blur kernel and directly generates sharp images in an end‐to‐end manner. The experimental structure on the benchmark dataset and real‐world images showed that this method has better quality than the existing methods in terms of large‐scale blur and subjective perception effects, both in quantitative and qualitative terms.
Jinjiang Li 0001, Zhen Hua
IET Image Process.2
2021 Image matting trimap optimization by ant colony algorithm
Genji Yuan, Jinjiang Li 0001, Zhen Hua
Multim. Tools Appl.2
2021 Low-Light Image Enhancement via Progressive-Recursive Network
abstract
Low-light images have low brightness and contrast, which presents a huge obstacle to computer vision tasks. Low-light image enhancement is challenging because multiple factors (such as brightness, contrast, artifacts, and noise) must be considered simultaneously. In this study, we propose a neural network—a progressive-recursive image enhancement network (PRIEN)—to enhance low-light images. The main idea is to use a recursive unit, composed of a recursive layer and a residual block, to repeatedly unfold the input image for feature extraction. Unlike in previous methods, in the proposed study, we directly input low-light images into the dual attention model for global feature extraction. Next, we use a combination of recurrent layers and residual blocks for local feature extraction. Finally, we output the enhanced image. Furthermore, we input the global feature map of dual attention into each stage in a progressive way. In the local feature extraction module, a recurrent layer shares depth features across stages. In addition, we perform recursive operations on a single residual block, significantly reducing the number of parameters while ensuring good network performance. Although the network structure is simple, it can produce good results for a range of low-light conditions. We conducted experiments on widely adopted datasets. The results demonstrate the advantages of our method compared with other methods, from both qualitative and quantitative perspectives.
Jinjiang Li 0001, Xiaomei Feng, Zhen Hua
IEEE Trans. Circuits Syst. Video Technol.1
2020 Saliency-based image correction for colorblind patients
abstract
Improper functioning, or lack, of human cone cells leads to vision defects, making it impossible for affected persons to distinguish certain colors. Colorblind persons have color perception, but their ability to capture color information differs from that of normal people: colorblind and normal people perceive the same image differently. It is necessary to devise solutions to help persons with color blindness understand images and distinguish different colors. Most research on this subject is aimed at adjusting insensitive colors, enabling colorblind persons to better capture color information, but ignores the attention paid by colorblind persons to the salient areas of images. The areas of the image seen as salient by normal people generally differ from those seen by the colorblind. To provide the same saliency for colorblind persons and normal people, we propose a saliency-based image correction algorithm for color blindness. Adjusted colors in the adjusted image are harmonious and realistic, and the method is practical. Our experimental results show that this method effectively improves images, enabling the colorblind to see the same salient areas as normal people.
Jinjiang Li 0001, Xiaomei Feng
Comput. Vis. Media1
2020 Computing knots by quadratic and cubic polynomial curves
abstract
A new method is presented to determine parameter values (knot) for data points for curve and surface generation. With four adjacent data points, a quadratic polynomial curve can be determined uniquely if the four points form a convex polygon. When the four data points do not form a convex polygon, a cubic polynomial curve with one degree of freedom is used to interpolate the four points, so that the interpolant has better shape, approximating the polygon formed by the four data points. The degree of freedom is determined by minimizing the cubic coefficient of the cubic polynomial curve. The advantages of the new method are, firstly, the knots computed have quadratic polynomial precision, i.e., if the data points are sampled from a quadratic polynomial curve, and the knots are used to construct a quadratic polynomial, it reproduces the original quadratic curve. Secondly, the new method is affine invariant, which is significant, as most parameterization methods do not have this property. Thirdly, it computes knots using a local method. Experiments show that curves constructed using knots computed by the new method have better interpolation precision than for existing methods.
Fan Zhang 0045, Jinjiang Li 0001, Peiqiang Liu
Comput. Vis. Media2
2020 Image reflection removal using end-to-end convolutional neural network
abstract
Single image reflection removal is an ill‐posed problem. To solve this problem, this study develops a network structure based on a deep encoder–decoder RRnet. Unlike most deep learning strategies applied in this context, the authors find that redundant information increases the difficulty of predicting images on the network; thus, the proposed method uses mixed reflection image cascaded edges as input to the network. The proposed network structure is divided into two parts: the first part is a deep convolutional encoder–decoder network. Its function uses the mixed reflection image and the target edge as input to predict the target layer. The second part is an identical encoder–decoder network structure. Its function uses the mixed reflection image and the reflection edge as input to predict the image reflection layer. In addition, the authors use joint loss to optimise the network model. To train the neural network, they also create an image dataset for reflection removal, which includes a true mixed reflection image and a synthetic mixed reflection image. They use four evaluation indicators to evaluate the proposed method and the other six methods. The experimental results indicate that the proposed method is superior to previous methods.
Jinjiang Li 0001, Guihui Li
IET Image Process.1
2020 Robust trimap generation based on manifold ranking
Jinjiang Li 0001, Genji Yuan
Inf. Sci.1
2020 Low-light image enhancement algorithm based on an atmospheric physical model
Xiaomei Feng, Jinjiang Li 0001, Zhen Hua
Multim. Tools Appl.2
2020 Robust Visual Tracking Using Kernel Sparse Coding on Multiple Covariance Descriptors
abstract
In this article, we aim to improve the performance of visual tracking by combing different features of multiple modalities. The core idea is to use covariance matrices as feature descriptors and then use sparse coding to encode different features. The notion of sparsity has been successfully used in visual tracking. In this context, sparsity is used along appearance models often obtained from intensity/color information. In this work, we step outside this trend and propose to model the target appearance by local covariance descriptors (CovDs) in a pyramid structure. The proposed pyramid structure not only enables us to encode local and spatial information of the target appearance but also inherits useful properties of CovDs such as invariance to affine transforms. Since CovDs lie on a Riemannian manifold, we further propose to perform tracking through sparse coding by embedding the Riemannian manifold into an infinite-dimensional Hilbert space. Embedding the manifold into a Hilbert space allows us to perform sparse coding efficiently using the kernel trick. Our empirical study shows that the proposed tracking framework outperforms the existing state-of-the-art methods in challenging scenarios.
Changyong Guo, Jinjiang Li 0001, Xuesong Jiang, Jun Zhang 0017, Lei Zhang 0036
ACM Trans. Multim. Comput. Commun. Appl.3
2013 Curvature-direction measures for 3D feature detection
Jinjiang Li 0001
Sci. China Inf. Sci.1
2012 Robust Feature Extraction Based on Principal Curvature Direction
Jinjiang Li 0001
CVM1
2012 New progress in geometric computing for image and video processing
Jinjiang Li 0001, Hanyi Ge
Frontiers Comput. Sci.1
2010 Image Contour Extraction Based on Ant Colony Algorithm and B-snake
Jinjiang Li 0001
ICIC (1)1