Jinyong Cheng

dblp:35/3765 · DBLP profile ↗
← Back
29ranked-venue papers
2as first author
27since 2021 · last 2026
0000-0003-3432-4831ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Global Relation and Semantic Parts: A RaSPNet for Visible-Infrared Person Re-identification
Shanlei Gao, Jinyong Cheng, Zhe Zhuang
ICIC (10)2
2025 SD-CIL: A Structurally Decoupled Framework for Balancing Stability and Plasticity in Online Class-Incremental Learning
Mengyun Chen, Jinyong Cheng, Zhe Zhuang, Zhonghe Wei
IEEE Big Data2
2025 TDMF: Text-Guided Denoising and Interactive Medical Image Fusion
abstract
Multimodal image fusion aims to merge features from different modalities to create a comprehensively representative image. However, existing medical image fusion methods often struggle to handle noise generated during image acquisition, significantly diminishing their impact on visual quality. To address these challenges, we propose a semantically text-guided medical image fusion model, named TDMF. Specifically, TDMF guides classical image fusion through textual semantics and effectively coordinates the resolution of degradation and interaction issues during the fusion process. By integrating text encoders and interactive fusion modules, TDMF establishes a unified framework for denoising and interactive fusion of medical images. Extensive experiments have demonstrated that our proposed text-guided image fusion strategy offers significant advantages over state-of-the-art methods in medical image fusion performance.
Aimei Dong, Guohua Lv, Guixin Zhao, Jinyong Cheng
ICASSP6
2025 FDAFusion: A Joint Frequency Decomposition and Adaptive Fusion Network for Nighttime Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion aims to integrate multimodal complementary information from source images. However, image fusion faces dual challenges in dim environments and frequency-domain decomposition. Current methods perform well only within a single domain, overlooking the inherent relationships between them. Furthermore, feature fusion only takes a single feature representation from different modalities, making it unable to capture the interactions and impacts between the two modalities. Direct concatenation and attention-based weighting used to merge the modalities result in an incorrect feature weight distribution. To solve the above problems, in this paper, we propose a novel nighttime image fusion network (FDAFusion). It introduces the Frequency Domain Information Decomposition (FIDM) to achieve cyclic feedback among joint low-light enhancement, frequency-domain decomposition, and fusion. Moreover, an Adaptive Cross-Enhancement Fusion Module (ACFM) is developed to utilize the consistency matrix to capture the correlation between different modalities, achieving more efficient feature weight allocation and dynamic fusion. Experimental results demonstrate that our approach outperforms others in terms of contrast, texture detail, and visual fidelity.
Kening Cui, Jinyong Cheng
IJCNN2
2025 Un-CNL: An uncertainty-based continual noisy learning framework
Guangrui Guo, Jinyong Cheng
Neurocomputing2
2025 BSMEF: Optimized multi-exposure image fusion using B-splines and Mamba
Jinyong Cheng, Qinghao Cui, Guohua Lv
Image Vis. Comput.1
2025 GLMR-Net: Global-to-local mutually reinforcing network for pneumonia segmentation and classification
Aimei Dong, Guohua Lv, Jinyong Cheng
Pattern Recognit.4
2025 SLFusion: A Structure-Aware Infrared and Visible Image Fusion Network for Low-Light Scenes
abstract
Infrared and visible image fusion is an image enhancement technique that generates a single image with rich textures and significant objectives in a variety of scenarios, providing great convenience for human discrimination and computer recognition. However, in low-light environments, low-intensity visible images tend to blur valuable information, and these details are often ignored during image fusion, resulting in the loss of important information. Although existing methods take into account the damage of low illumination and highlight the illumination in the fusion process, a large amount of structural information is lost in the process of adjusting illumination, resulting in the lack of texture details and poor performance in high-level vision tasks. To address the above challenges, this paper proposes a structure-aware image fusion method for low illumination scenes, called SLFusion, which enhances the illumination while reducing the loss of structural information, leading to a fused image with richer texture details. We first design an illumination enhancement module to separate the degraded illumination from the scene information in the visible image, and mine more details from the low-intensity regions. Based on the fact that image edge information has a good capability of modeling structures, we design an edge extraction network for low-light visible images to model the structural information, which can accurately highlight important structural information and inject it into the fusion image. The proposed method produces fusion results that not only have good visual perception, but also minimize the loss of structural information. Extensive experiments on benchmark datasets demonstrate that the proposed method outperforms state-of-the-art (SOTA) methods in terms of visual quality, quantitative metrics as well as advanced vision tasks.
Guohua Lv, Aimei Dong, Zhonghe Wei, Jinyong Cheng
IEEE Trans. Circuits Syst. Video Technol.5
2025 A Mutually Enhancement Network for Superpixel Segmentation and Classification of Hyperspectral Image
abstract
Most existing hyperspectral image (HSI) classification methods primarily focus on capturing subtle spectral variations by leveraging local spectral-spatial cues derived from patch-level representations. However, limited attention has been given to exploring the global spatial contextual correlations among pixels of HSI. In this study, we propose the Superpixel Segmentation and Classification Mutual Enhancement Network (S2CMEN), a novel framework that integrates global spatial correlations with spectral information through the mutual enhancement of superpixel segmentation and classification. Specifically, a global spatial adaptive module (GSAM) is designed to obtain the direct correlation of the global classes in HSI. It consists of an Adaptive Spectral-Superpixel Network (ASSN) and a Graph Convolutional Network (GCN), forming a synergistic architecture that effectively captures global spatial relationships by adaptively deriving superpixel results from HSIs. Notably, GSAM offers a transferable global spatial representation for HSI tasks, enabling integration with other spectral feature extraction models. Furthermore, we develop a Spatial-Spectral Fusion Module (SSFM) to obtain comprehensive spectral features and fuse them with the extracted global spatial features. Finally, under the constraint of a unit loss, the Mutual Enhancement Strategy (MES) can make the superpixel segmentation loss and the classification loss mutually enhance each other for better performance. We conducted extensive experiments on three public datasets. The proposed S2CMEN achieves overall classification accuracies of 97.38%, 92.33%, and 91.38% on Indian Pines, Pavia University, and Houston, respectively, consistently surpassing existing state-of-the-art methods.
Mengxin Cao, Yongmin Li 0001, Xu Zhang 0039, Guixin Zhao, Guohua Lv, Aimei Dong, Jinyong Cheng, Wei Li 0032, Xiangjun Dong 0001
IEEE Trans. Geosci. Remote. Sens.7
2024 IDFusion: An Infrared and Visible Image Fusion Network for Illuminating Darkness
abstract
The purpose of infrared and visible image fusion is to combine the background information of the visible images and the thermal target information of the infrared images. Current fusion methods often neglect the challenges of low-light conditions. In the night scene, the existing methods fail to capture texture information from the visible images that is hidden in the darkness, producing suboptimal fusion results that can hinder subsequent visual applications. Therefore, we propose a method for infrared and visible image fusion under night scenes, termed as IDFusion. Specifically, our method is divided into three parts, first, a dense auto-encoder is designed to obtain more useful features from the source images. Then we design a brightness enhancement network that removes visible degraded illumination maps to obtain brightness-enhanced features. Finally, a texture fusion network is designed so that the infrared features and enhanced visible features avoid texture loss during the fusion process. Experimental results show that our network can obtain better fusion results, outperforming state-of-the-art methods in terms of subjective visual effects and quantitative metrics.
Guohua Lv, Xiyan Wang, Zhonghe Wei, Jinyong Cheng, Guangxiao Ma, Hanju Bao
CSCWD4
2024 A Multi-Scale Infrared and Visible Image Fusion Network Based on Context Perception
abstract
Existing image fusion methods often fail to fully utilize the contextual information of the images, resulting in blurry and distorted fusion results. In addition, although the existing multi-scale feature extraction methods can consider the information of the source image from multiple perspectives, they tend to ignore the local information. These limitations degrade the performance and effectiveness of image fusion techniques in practical applications. To address above issues, we propose a fusion network architecture called MIA-Nest, consisting of encoder network, fusion network, and decoder network. Specifically, we introduce the Squeeze-and-Excitation Blocks(SEBs) into the encoder network, so that it can take local information into account while extracting multi-scale features. At the same time, in order to allow the fusion network to integrate the context information of the extracted features, we introduce Global Context Blocks (GCBs), so as to enrich the details in the fusion result and enable it to contain more information. Finally, the decoder network reconstructs the fusion image from the fused features obtained by the fusion network. Experimental results demonstrate that our proposed method significantly improves the quality and detail preservation capability of the fused images.
Huixuan Zhao, Jinyong Cheng, Rundong Du
CSCWD2
2024 GLEGNet: Infrared and Visible Image Fusion via Global-Local Feature Extraction and Edge-Gradient Preservation
Guohua Lv, Wenkuo Song, Zhonghe Wei, Aimei Dong, Jinyong Cheng, Guangxiao Ma
ICONIP (7)5
2024 DTBNet: medical image segmentation model based on dual transformer bridge
abstract
Accurate medical image segmentation [1] techniques help doctors make better diagnoses. The traditional U-Net [2] model usually uses skip connections to directly connect the corresponding layers of the encoder and decoder, which can induce the model to retain more spatial information. Therefore, a series of ways to optimise skip connections, such as the introduction of residual connections, have appeared to optimise the convergence speed of the model and to improve the model’s performance. Although skip connections have achieved remarkable success in the field of medical image segmentation, since the structure of medical images is more complex than that of general images, skip connections may cause the model to learn some irrelevant features when performing image segmentation, and the transfer of the bottom layer information to the top layer will also cause some interference, in order to better achieve the multi-scale fusion of the bottom layer features and the top layer features and to reduce the interference between them, we proposed DTBNet for medical image segmentation, which is a U-shaped network with a bridge structure over the skip connections, and is characterised by the fact that the features output from the different layers of the encoder are first processed by the bridge structure before being transmitted to the corresponding layers of the decoder. Next, we design a SE_ASPP module in the encoder that can capture information using different scales of receptive fields and adaptively adjusts the weights for each channel. We also propose a neighbour fusion module(NFM) in the encoder for fusing features from adjacent layers of the encoder to enhance the representation of features at different scales. In addition, we construct a perceptual loss function [3] to comprehensively supervise the task-aware features from the bottom to the top layer, which helps the model to better capture the structures and textures in medical images. Our method achieves sota performance compared to previous work under different evaluation metrics on two medical image segmentation datasets including abdomen and heart.
Kequan Chen, Yuli Wang, Jinyong Cheng
IJCNN3
2024 MMFI-Net: Multilevel Multimodal Feature Injection Network for Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion aims at a better transfer of feature information from both modalities into the fused image. However, almost all existing methods currently focus on the performance of only one part of a single-modal feature extraction or image reconstruction, and do not link multimodal and multilevel features well. This leads to the problem of insufficient feature information in the image reconstruction process, and the resulting fused image lacks texture details and edge information. In this paper, we propose a multilevel multimodal feature injection network (MMFI-Net). First, the DMII module is used to complement the information of the difference modals. Then, the complemented feature information is decomposed more specifically into basic and detail information by fuzzy filtering and Laplace filtering. Next, in the final image reconstruction section, we concatenate the multi-level features from the feature extraction section with the low-level feature information from the image reconstruction section. In this way, the network is better able to capture information in the image and mitigate the problem of feature information loss. Our method demonstrates significant superiority in terms of contrast, edge information, and texture details, as evidenced by experimental results on publicly available datasets.
Rundong Du, Jinyong Cheng, Huixuan Zhao
IJCNN2
2024 EBcGAN: An Edge-Based Conditional Generative Adversarial Network for Image Fusion
Mengshu Li, Zheyuan Yang, Yuai Hua, Jinyong Cheng
PKAW4
2024 ULNet: A Lightweight Segmentation Network for Lane Detection
Xudong Meng, Jinyong Cheng
PRCV (9)3
2024 MFIFusion: An infrared and visible image enhanced fusion network based on multi-level feature injection
Aimei Dong, Guohua Lv, Guixin Zhao, Jinyong Cheng
Pattern Recognit.6
2024 Co-Enhancement of Multi-Modality Image Fusion and Object Detection via Feature Adaptation
abstract
The integration of multi-modality images significantly enhances the clarity of critical details for object detection. Valuable semantic data from object detection enriches the fusion process of these images. However, the potential reciprocal relationship that could enhance their mutual performance remains largely unexplored and underutilized, despite some semantic-driven fusion methodologies catering to specific application needs. To address these limitations, this study proposes a mutually reinforcing, dual-task-driven fusion architecture. Specifically, our design integrates a feature-adaptive interlinking module into both image fusion and object detection components, effectively managing the inherent feature discrepancies. The core idea is to channel distinct features from both tasks into a unified feature space after feature transformation. We then design a feature-adaptive selection module to generate features rich in target semantic information and compatible with the fusion network. Finally, effective combination and mutual enhancement of the two tasks are achieved through an alternating training process. A diverse range of swift evaluations is performed across various datasets to corroborate the potential efficiency of our framework, actualizing visible advancements in both fusion effectiveness and detection accuracy.
Aimei Dong, Guixin Zhao, Yi Zhai 0003, Guohua Lv, Jinyong Cheng
IEEE Trans. Circuits Syst. Video Technol.8
2024 Spatial-Spectral-Semantic Cross-Domain Few-Shot Learning for Hyperspectral Image Classification
abstract
Preprocessing procedures are commonly employed to reduce water-absorption bands and noise in hyperspectral images (HSIs). Nevertheless, they typically do not entirely eradicate noise. This is especially evident in scenarios that necessitate data of exceptional quality, such as cross-domain few-shot classification tasks. Within these specific conditions, the influence of remaining background noise on the ultimate results of classification is substantial. Furthermore, the presence of sample selection biases in the few-shot task might lead to the emergence of false statistical correlations between data from distinct domains, resulting in a decrease in the model’s ability to generalize. We propose a new method called spatial-spectral–semantic cross-domain few-shot learning (S3CFSL) to address the challenge. This method promotes the learning of transferable information by incorporating feature denoising operations in the feature extraction process to restore essential information. Concurrently, it enhances cross-domain distributional consistency by introducing a semantic-aware strategy to strengthen the association between cross-domain data and semantic information. Specifically, the spatial and spectral dual channels (SSDCs), in conjunction with the cross-spatial-spectral transformer (CSST), are designed as a feature extractor to acquire interactive spatial-spectral features. The feature-denoising operations can further acquiring transferable information from cross-domain features, thus facilitating meta-learning in both the source domain (SD) and the target domain (TD). Meanwhile, a semantic-enhanced domain alignment (SEDA) is designed to promote domain adaptation by using a semantic-aware strategy, which significantly enhances distributional consistency for cross-domain tasks. Our results exhibit exceptional classification efficacy in comparison to other state-of-the-art approaches on three public HSI datasets.
Mengxin Cao, Xu Zhang 0039, Jinyong Cheng, Guixin Zhao, Wei Li 0032, Xiangjun Dong 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 A Deep Fusion Rule for Infrared and Visible Image Fusion: Feature Communication for Importance Assessment
abstract
The purpose of infrared and visible image fusion is to extract and combine information from source images to produce results that contain vital and complementary information. Existing fusion rules may not extract the most useful information and cannot effectively retain important information. To solve this problem, we propose a novel deep learning-based fusion rule. We perform deep feature communication and quantify the effect of feature substitution on image characteristics, including gradients and contrast. Based on the influence, the importance of feature maps can be objectively assessed. This designed fusion rule can preserve the thermal information of infrared images and the texture details of visible light images in a targeted manner, so as to obtain better fusion results. Qualitative and quantitative experiments have shown that our method can perform better than other advanced methods.
Xuran Lv, Jinyong Cheng, Guohua Lv, Zhonghe Wei
ICASSP2
2023 L2fusion: Low-Light Oriented Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion aims to integrate salient targets and abundant texture information into a single fused image. Existing methods typically ignore the issue of illumination, so that there are problems of weak texture details and poor visual perception in case of low illumination. To address this issue, we propose a low-light oriented infrared and visible image fusion network, named L2Fusion. In particular, we first design a decomposition network according to Retinex theory to obtain the reflectance features of a visible image with low-light. Then, these features are integrated with the features extracted from the corresponding infrared image by a residual network. The finally fused image largely eliminates the negative impact caused by low illumination, and contains both salient targets and abundant texture information. Extensive experiments demonstrate the superiority of our L2Fusion over the state-of-the-art methods, in terms of both visual effect and quantitative metrics.
Guohua Lv, Aimei Dong, Zhonghe Wei, Jinyong Cheng
ICIP5
2023 BS-YOLOv5s: Insulator Defect Detection with Attention Mechanism and Multi-Scale Fusion
abstract
With the rapid development of deep learning, the use of object detection algorithms for aerial insulator image defect detection has become the main way. To address the problems of low detection accuracy for small targets, weak representation ability of feature maps, insufficient extracted key information, and the shortage of aerial insulator defect datasets, this paper proposes an improved insulator defect detection method named BS-YOLOv5s based on 3-D attention mechanism and Bi-Slim-neck using YOLOv5s as the base network. Additionally, to solve the problem of the shortage of aerial insulator datasets, this paper proposes a new aerial insulator dataset Weather-Insulator (WI) containing a variety of defect scenarios. The experimental results demonstrate that the proposed method not only greatly improves the detection accuracy, but also maintains a high detection speed, satisfying the engineering requirements for insulator defect detection. The dataset and code for this paper are publicly available at https://github.com/jspron/insulator-defect.
Zengbin Zhang, Guohua Lv, Guixin Zhao, Yi Zhai 0003, Jinyong Cheng
ICIP5
2023 Self-supervised anomaly detection of medical images based on dual-module discrepancy
abstract
Medical images anomaly detection plays a very important role in modern health care, which helps to improve the quality and efficiency of medical services and promote the development of human health. Due to the high cost of annotation in anomaly images and the fact that most existing methods do not fully utilize information from unlabeled images. Therefore, we propose a new reconstruction network and loss function that can better utilize unlabelled and normal images for anomaly identification. The framework used in this paper consists of two modules, each consisting of three reconstruction networks with the same architecture but different inputs. One module is trained only on normal images and is called the normal module (NM). The other module is trained on both normal images and unlabeled images, and is called the unknown module (UM). Furthermore, the internal differences of the normal module and the differences between the two modules will be used as two powerful anomaly scores, and these two anomaly scores will be refined to indicate anomalies. Experiments on four medical datasets show the state-of-the-art performance by the proposed approach.
Yuqing Song 0006, Jinyong Cheng
MMAsia2
2023 Online Class-Incremental Learning in Image Classification Based on Attention
Baoyu Du, Zhonghe Wei, Jinyong Cheng, Guohua Lv, Xiaoyu Dai
PRCV (7)3
2023 Enhancing Continual Noisy Label Learning with Uncertainty-Based Sample Selection and Feature Enhancement
Guangrui Guo, Zhonghang Wei, Jinyong Cheng
PRCV (8)3
2023 SIEFusion: Infrared and Visible Image Fusion via Semantic Information Enhancement
Guohua Lv, Wenkuo Song, Zhonghe Wei, Jinyong Cheng, Aimei Dong
PRCV (3)4
2023 Exemplar-Free Continual Learning in Vision Transformers via Feature Attention Distillation
abstract
In this paper, we propose a new approach for continual learning based on the Visual Transformers (ViTs). The purpose of continual learning is to address the catastrophic forgetting problem. One method for preventing forgetting previous tasks is exemplar replay. However, exemplar replay has limitations such as limited memory capacity, large storage requirements, and difficulty adapting to new tasks, which restrict its practical application. Therefore, we adopted knowledge distillation to implement exemplar-free continual learning. This approach transfers the knowledge of old tasks to the model's prior knowledge, helping the model to learn new tasks better while not increasing the storage burden. Based on this, we propose a new method for training the ViT model called Feature Attention Distillation (FAD). Specifically, we introduced a feature attention mechanism into the knowledge distillation model to help the student model better retain meaningful feature representations and attention distributions during the learning process of new tasks by passing attention information between the teacher and student models, thereby improving the learning efficiency and performance of the model. In addition, we also introduced a new normalization method called Continual Normalization (CN) into the student model. This method dynamically calculates the sample mean and variance of the current and historical tasks, allowing the student network to adapt better to the feature distribution changes between different tasks, thus improving the model's generalization ability and robustness. Extensive experiments on the CIFAR100 and ImageNet-32 datasets show that our exemplar-free method is competitive in performance compared to rehearsal-based ViT methods.
Xiaoyu Dai, Jinyong Cheng, Zhonghe Wei, Baoyu Du
SMC2
2020 Protein secondary structure prediction based on integration of CNN and LSTM model
Jinyong Cheng, Yihui Liu, Yuming Ma
J. Vis. Commun. Image Represent.1
2020 Chinese medical question answer selection via hybrid models based on CNN and GRU
Wenpeng Lu, Weihua Ou, Guoqiang Zhang 0003, Xu Zhang 0053, Jinyong Cheng, Weiyu Zhang 0001
Multim. Tools Appl.6