Jing Li 0040

dblp:l/JingLi40 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0003-0716-5329ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 LSRNet: A Novel Interpretable Low-rank Sparse Representation Guided Fusion Network for Polarization and Intensity Images
abstract
Polarization and intensity images fusion (PIF) has extracted extensive attentions as it can generate images with clear scene information and salient texture details of the object surface that are important for downstream applications. However, existing deep learning-based PIF methods usually lack interpretability and ignore the interactions among multi-modal features. To this end, we propose a novel interpretable low-rank sparse representation guided fusion network for polarization and intensity images (termed LSRNet). Specifically, a low-rank sparse representation deep unfolding module is designed to acquire the base and detail features of the source images, with the ability of improving the interpretability of the network. In addition, a cross-modal connection complementary feature extraction module is proposed, which aims to establish dependency among features of multi-modalities to fully extract complementary features of the source images. In order to demonstrate the validity of our LSRNet and take into account shortcomings of existing datasets for PIF, a multi-scene polarization and intensity image dataset, named MSPI dataset, is constructed, which includes 1034 high-resolution aligned image pairs. According to the best of our knowledge, this is the most comprehensive dataset for PIF that with a large number of image pairs, high resolution and multiple scene types. Extensive experiments on our MSPI dataset and two publicly available datasets (i.e., 12CFC and HCP) demonstrate the superior fusion performance, generalization ability, and desirable running efficiency of our LSRNet. Our codes and dataset will be publicly available at https://github.com/thebinyang/LSRNet.
Bin Yang 0008, Licheng Liu, Yu Liu 0023, Jing Li 0040
IEEE Trans. Image Process.5
2025 HAQJSK: Hierarchical-Aligned Quantum Jensen-Shannon Kernels for Graph Classification (Extended Abstract)
abstract
This paper proposes a family of Hierarchical Aligned Quantum Jensen-Shannon Kernels (HAQJSK) for un-attributed graphs. The HAQJSK kernels can incorporate hierarchical correspondence information between graphs, and thus transform arbitrary sized graphs into fix-sized aligned structures, i.e., the hierarchical transitive aligned Adjacency Matrix of vertices or Density Matrix of Continuous-Time Quantum Walks (CTQWs). For pairwise graphs, the resulting HAQJSK kernels are defined by computing the Quantum Jensen-Shannon Divergence (QJSD) between their aligned structures. Unlike classical graph kernels, the HAQJSK kernels can either reflect global intrinsic structure characteristics through CTQWs, or address the drawback of neglecting structural correspondence information, theoretically explaining the effectiveness.
Lu Bai 0001, Lixin Cui, Yue Wang 0014, Ming Li 0065, Jing Li 0040, Philip S. Yu, Edwin R. Hancock
ICDE5
2025 A multi-level detection guided and co-encoding network for infrared and visible image fusion
Renhua Wang, Jing Li 0040
Pattern Recognit.4
2025 Graph Representation Learning for Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion aims to extract complementary features to synthesize a single fused image. In our method, we covert the regular image format into the graph space and conduct graph convolutional networks (GCNs) to extract NLss for the reliable infrared and visible image fusion. More specifically, GCNs are first performed on each intra-modal set to aggregate the features and propagate the inherent information, thereby extracting independent intra-modal NLss. Then, such intra-modal non-local self-similarity (NLss) features of infrared and visible images are concatenated to explore cross-domain NLss inter-modally and reconstruct the fused images. Extensive experiments show the superior performance of our method with the qualitative and quantitative analysis on the TNO, RoadScene and M3FD datasets, respectively, outperforming many state-of-the-art (SOTA) methods for the robust and effective infrared and visible image fusion.
Jing Li 0040, Lu Bai 0001, Bin Yang 0008, Chang Li 0001, Lingfei Ma
IEEE Trans Autom. Sci. Eng.1
2025 Dual-Modal Prior Semantic Guided Infrared and Visible Image Fusion for Intelligent Transportation System
abstract
Infrared and visible image fusion (IVF) plays an important role in intelligent transportation system (ITS). The early works predominantly focus on boosting visual appeal of the fused result, although several recent approaches have tried to combine high-level vision task with IVF, they prioritize the design of cascaded structure to seek unified suitable features and fit different tasks. Thus, they tend to bias toward reconstructing raw pixels without considering the significance of semantic features. Therefore, we propose a novel prior semantic guided image fusion method based on the dual-modality strategy, improving the performance of IVF in ITS. Specifically, to explore the independent significant semantic of each modality, we first design two parallel semantic segmentation branches with a refined feature adaptive-modulation (RFaM) mechanism. RFaM can perceive the features that are semantically distinct enough in each semantic segmentation branch. Then, two pilot experiments based on the two branches are conducted to capture the significant prior semantic of source images, which is then applied to guide the fusion task in the integration of semantic segmentation branches and fusion branch. In addition, to aggregate both high-level semantics and impressive visual effects, we further investigate the frequency response of the prior semantics, and propose a multi-level representation-adaptive fusion (MRaF) module to explicitly integrate low-frequency prior semantic with high-frequency details. Extensive experiments on two public datasets demonstrate the superiority of our method over state-of-the-art fusion approaches. Our method has better performance on four quantitative metrics in fusion task and achieves the highest mIoU in semantic segmentation task.
Jing Li 0040, Lu Bai 0001, Bin Yang 0008, Chang Li 0001, Lingfei Ma, Lixin Cui, Edwin R. Hancock
IEEE Trans. Intell. Transp. Syst.1
2024 QBER: Quantum-based Entropic Representations for un-attributed graphs
Lixin Cui, Ming Li 0065, Lu Bai 0001, Yue Wang 0014, Jing Li 0040, Zhao Li 0007, Yunwen Chen, Edwin R. Hancock
Pattern Recognit.5
2024 RdmkNet & Toronto-RDMK: Large-Scale Datasets for Road Marking Classification and Segmentation
abstract
Effective road marking classification and segmentation play a pivotal role in advancing vehicle-to-everything (V2X) applications and refining road inventory databases. However, the irregular data formats and unordered permutation modes of 3D point clouds, along with the limited availability of large-scale datasets with point-level annotations, remain significant obstacles to designing deep learning-based networks with superior performance. To address these challenges, this paper proposes a novel multi-level feature optimization network structure, named MFPNet, and introduces two point cloud benchmarks, RdmkNet and Toronto-Rdmk, for road marking classification and segmentation in intricate urban environments. MFPNet is composed of three integral modules. First, the M-transformer module, consisting of three transformers obtained from different channels, fully captures rich point cloud background information and long-distance dependencies between objects. Then, the feature pooling aggregation module uses parallel structured pooling attention mechanisms to aggregate features captured by the M-transformer module, while the prediction refinement module further enhances the acquisition of semantic features. Comparative studies indicate that MFPNet can be embedded into general deep learning networks without changing their original network structures, significantly improving the accuracy of multiple baseline networks. Furthermore, extensive experiments demonstrate that the two newly-developed point cloud datasets are meaningful for road marking classification and segmentation tasks, contributing to the development of autonomous driving.
Jing Du 0007, Lingfei Ma, Jing Li 0040, Nannan Qin, John S. Zelek, Haiyan Guan, Jonathan Li 0001
IEEE Trans. Intell. Transp. Syst.3
2024 CEFusion: An Infrared and Visible Image Fusion Network Based on Cross-Modal Multi-Granularity Information Interaction and Edge Guidance
abstract
Infrared and visible image fusion (IVF) aims to generate a fused image with abundant texture details and salient thermal radiation targets, which can not only preserve the necessary scene information for traffic vision tasks, but also highlight the imperceptible targets that are crucial in intelligent transportation system (ITS). However, the existing image fusion methods often lack the information interactions between cross-modal features and among cross-granularity features, and they usually ignore the importance of edge information to the image, which affects the quality of the fused image. To this end, this study proposes an IVF network based on cross-modal multi-granularity information interaction and edge guidance, termed as CEFusion. On the one hand, a triple-branch scene fidelity module is designed to fuse the different modal features extracted by the encoder. This module can adequately mine difference information and infrared salient information of the cross-modal features through cross-modal information interaction. On the other hand, a progressive cross-granularity interaction feature enhancement module is employed to achieve the information interaction among cross-granularity features, which can further enrich the texture and structure information in the fused features. In addition, a novel edge loss function is proposed to guide the network to retain the edge information from source images. Extensive comparative and generalization experiments demonstrate that our CEFusion superior to the state-of-the-art methods in preserving texture details and thermal radiation targets. More importantly, the performance of our method in the high-level vision task suggests that it can provide reliable assistance for the ITS applications.
Bin Yang 0008, Jing Li 0040
IEEE Trans. Intell. Transp. Syst.4
2024 HAQJSK: Hierarchical-Aligned Quantum Jensen-Shannon Kernels for Graph Classification
abstract
In this work, we propose two novel quantum walk kernels, namely the Hierarchical Aligned Quantum Jensen-Shannon Kernels (HAQJSK), between un-attributed graph structures. Different from most classical graph kernels, the proposed HAQJSK kernels can incorporate hierarchical aligned structure information between graphs and transform graphs of random sizes into fixed-size aligned graph structures, i.e., the Hierarchical Transitive Aligned Adjacency Matrix of vertices and the Hierarchical Transitive Aligned Density Matrix of the Continuous-Time Quantum Walks (CTQW). With pairwise graphs to hand, the resulting HAQJSK kernels are defined by computing the Quantum Jensen-Shannon Divergence (QJSD) between their transitive aligned graph structures. We show that the proposed HAQJSK kernels not only reflect richer intrinsic whole graph characteristics in terms of the CTQW, but also address the drawback of neglecting structural correspondence information that arises in most R-convolution graph kernels. Moreover, unlike the previous QJSD based graph kernels associated with the QJSD and the CTQW, the proposed HAQJSK kernels can simultaneously guarantee the properties of permutation invariant and positive definiteness, explaining the theoretical advantages of the HAQJSK kernels. The experiment indicates the effectiveness of the new proposed kernels.
Lu Bai 0001, Lixin Cui, Yue Wang 0014, Ming Li 0065, Jing Li 0040, Philip S. Yu, Edwin R. Hancock
IEEE Trans. Knowl. Data Eng.5
2023 Enhancing Spatial Resolution of Building Datasets Using Transformer-Based Single-Image Super-Resolution
abstract
The spatial resolution of Earth Observation (EO) images plays a key role in building footprint extraction. For the spatial resolution enhancement, deep learning-based image super-resolution methods have been widely used due to their remarkable performance. Transformer-based networks are effective and has drawn much attention in computer vision but underutilized in remote sensing, especially for super-resolving building datasets. Therefore, in this paper, we developed a novel transformer-based Single-Image Super-Resolution (SISR) method, named Pyramid Vision Transformer-Residual Feature Aggregation Network (PVT_RFANet), to improve the spatial resolution of building datasets. Specifically, the PVT v2 network was embedded into our Momentum Spatial-Channel Attention Residual Feature Aggregation Network (MSCA-RFANet). Moreover we conducted a comparative study to compare our method with Bicubic interpolation (BI), Super-resolution Convolutional Neural Network (SRCNN), Deep Recursive Residual Network (DRRN), SRResNet, and MSCA-RFANet. Using Peak Signal-Noise Ratio (PSNR) and Similarity Structure Index Measurement (SSIM) as the evaluation metrics, our method showed highest performance with the PSNR of 22.01 dB and the SSIM of 0.50 on the WHU Building Dataset, which demonstrated the superior performance of the proposed method.
Yuwei Cai, Hongjie He 0003, Zhimeng He, Michael A. Chapman, Jing Li 0040, Lingfei Ma, Jonathan Li 0001
IGARSS5
2023 From Trained to Untrained: A Novel Change Detection Framework Using Randomly Initialized Models With Spatial-Channel Augmentation for Hyperspectral Images
abstract
Deep learning approaches have been extensively applied to change detection in hyperspectral images (HSIs). However, the majority of them encounter scarcity of training samples or rely on complex structures and learning strategies. Although untrained change detection models have been proved to be effective in relief above problems, they were constructed using regular convolutions and treated spatial locations and channels equally, which are insufficient to extract discriminative features and lead to limited accuracy. Given this, a novel untrained framework using randomly initialized models with spatial-channel augmentation (RICD) is proposed for HSI change detection in this paper. It consists of two major modules: 1) an enhanced feature extraction network using successive dilation-deformable feature extraction blocks, which can extract multiscale spatial-spectral features over unfixed sampling locations. It enlarges the field of view of convolutions and takes arbitrary neighborhood into consideration, which helps to increase the discriminativeness of the extracted features; 2) a change sensitive feature augmentation and comparison module integrating feature selection and spatial-channel augmentation strategies, which can exploit spatial context and channel importance. It magnifies difference between changed pixels and unchanged ones and emphasizes contribution of significant channels of the selected change sensitive features. Despite that convolution operations are included in RICD, all the weights are untrained and fixed once they are randomly initialized, indicating that the RICD can work in an unsupervised manner. Its performance is tested over three widely used hyperspectral datasets. Quantitative and qualitative comparisons with several state-of-the-art unsupervised methods reveal the effectiveness of the RICD method.
Bin Yang 0008, Yin Mao, Licheng Liu, Xinxin Liu 0002, Yuzhong Ma, Jing Li 0040
IEEE Trans. Geosci. Remote. Sens.6
2022 DSG-Fusion: Infrared and visible image fusion via generative adversarial networks and guided filter
Hongtao Huo, Jing Li 0040, Chang Li 0001, Xun Chen 0001
Expert Syst. Appl.3
2021 AttentionFGAN: Infrared and Visible Image Fusion Using Attention-Based Generative Adversarial Networks
abstract
Infrared and visible image fusion aims to describe the same scene from different aspects by combining complementary information of multi-modality images. The existing Generative adversarial networks (GAN) based infrared and visible image fusion methods cannot perceive the most discriminative regions, and hence fail to highlight the typical parts existing in infrared and visible images. To this end, we integrate multi-scale attention mechanism into both generator and discriminator of GAN to fuse infrared and visible images (AttentionFGAN). The multi-scale attention mechanism aims to not only capture comprehensive spatial information to help generator focus on the foreground target information of infrared image and background detail information of visible image, but also constrain the discriminators focus more on the attention regions rather than the whole input image. The generator of AttentionFGAN consists of two multi-scale attention networks and an image fusion network. Two multi-scale attention networks capture the attention maps of infrared and visible images respectively, so that the fusion network can reconstruct the fused image by paying more attention to the typical regions of source images. Besides, two discriminators are adopted to force the fused result keep more intensity and texture information from infrared and visible image respectively. Moreover, to keep more information of attention region from source images, an attention loss function is designed. Finally, the ablation experiments illustrate the effectiveness of the key parts of our method, and extensive qualitative and quantitative experiments on three public datasets demonstrate the advantages and effectiveness of AttentionFGAN compared with the other state-of-the-art methods.
Jing Li 0040, Hongtao Huo, Chang Li 0001, Renhua Wang
IEEE Trans. Multim.1
2020 Infrared and visible image fusion using dual discriminators generative adversarial networks with Wasserstein distance
Jing Li 0040, Hongtao Huo, Kejian Liu, Chang Li 0001
Inf. Sci.1
2019 Infrared and Visible Image Fusion via Multi-discriminators Wasserstein Generative Adversarial Network
abstract
Generative adversarial network (GAN) has been widely applied to infrared and visible image fusion. However, the existing GAN-based image fusion methods only establish one discriminator in the network to make the fused image capture gradient information from the visible image, which may result in the loss of some infrared intensity information and texture information on the fused images. To solve this problem and improve the performance of GAN, we extend GAN to multiple discriminators and propose an end-to-end multi-discriminators Wasserstein generative adversarial network (MD-WGAN). In this framework, the fused image can preserve major infrared intensity and detail information from the first discriminator, and keep more texture information that existing in visible image from the second discriminator. We also design a texture loss function via local binary patterns to preserve more texture from visible image. The extensive qualitative and quantitative experiments show the advantages of our method compared with other state-of-the-art fusion methods.
Jing Li 0040, Hongtao Huo, Kejian Liu, Chang Li 0001
ICMLA1