Jiawei Li 0016

dblp:12/3242-16 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-4313-8099ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Multi-Scale Spatial Channel Joint Representation for General Multi-Modality Image Fusion With Self-Supervision
abstract
The rapid advancement of multi-modality image fusion technology enables researchers to simultaneously acquire information from different modalities within a single fused image. In existing methods, some general approaches can implement both infrared and visible image fusion (IVIF) and medical image fusion (MIF) in the same framework. Nevertheless, these methods often ignore the learning of specific features in different modalities, resulting in unsatisfactory performance in fused results. To overcome this issue, we propose a multi-scale joint framework with self-supervision for general multi-modality image fusion, abbreviated as SCSFusion. It enables more targeted and robust implementation of IVIF and MIF. Specifically, in the fusion network, a joint attention module is employed to parallelly capture self-attention features in spatial and channel domains, which can keep fused results accurate in visual representation. Meanwhile, we utilize source images of different modalities to generate visual-focused maps as pseudo labels for self-supervised training of the fusion results. It effectively preserves the salient details in each fused image from being disrupted by other extracted information. Moreover, a medical dataset with segmentation labels, termed M2DF, is reorganized for fusion and down-stream tasks in MIF. With the help of M2DF, a pre-trained segmentation model can be cascaded with the fusion network, aiming to obtain high-level semantic features from inputs and enhance the data generalization in our general framework. We have conducted extensive experiments and analyses on SCSFusion in M$\rm ^{3}$FD, FMB, and M2DF datasets, respectively. The results indicate that the fused images generated by SCSFusion can not only achieve visually appealing results and superior performance metrics in MIF and IVIF, but also exhibit satisfactory performance in down-stream tasks.
Jiawei Li 0016, Jiansheng Chen 0001, Jinyuan Liu 0001, Xinlong Ding, Huimin Ma 0001
IEEE Trans. Multim.1
2025 A²RNet: Adversarial Attack Resilient Network for Robust Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion (IVIF) is a crucial technique for enhancing visual performance by integrating unique information from different modalities into one fused image. Exiting methods pay more attention to conducting fusion with undisturbed data, while overlooking the impact of deliberate interference on the effectiveness of fusion results. To investigate the robustness of fusion models, in this paper, we propose a novel adversarial attack resilient network, called A2RNet. Specifically, we develop an adversarial paradigm with an anti-attack loss function to implement adversarial attacks and training. It is constructed based on the intrinsic nature of IVIF and provide a robust foundation for future research advancements. We adopt a Unet as the pipeline with a transformer-based defensive refinement module (DRM) under this paradigm, which guarantees fused image quality in a robust coarse-to-fine manner. Compared to previous works, our method mitigates the adverse effects of adversarial perturbations, consistently maintaining high-fidelity fusion results. Furthermore, the performance of downstream tasks can also be well maintained under adversarial attacks.
Jiawei Li 0016, Jiansheng Chen 0002, Xinlong Ding, Jinyuan Liu 0001, Bochao Zou, Huimin Ma 0001
AAAI1
2025 Kaleidoscopic Background Attack: Disrupting Pose Estimation With Multi-Fold Radial Symmetry Textures
abstract
Camera pose estimation is a fundamental computer vision task that is essential for applications like visual localization and multi-view stereo reconstruction. In the object-centric scenarios with sparse inputs, the accuracy of pose estimation can be significantly influenced by background textures that occupy major portions of the images across different viewpoints. In light of this, we introduce the Kaleidoscopic Background Attack (KBA), which uses identical segments to form discs with multi-fold radial symmetry. These discs maintain high similarity across different viewpoints, enabling effective attacks on pose estimation models even with natural texture segments. Additionally, a projected orientation consistency loss is proposed to optimize the kaleidoscopic segments, leading to significant enhancement in the attack effectiveness. Experimental results show that optimized adversarial kaleidoscopic backgrounds can effectively attack various camera pose estimation models.
Xinlong Ding, Jiawei Li 0016, Bochao Zou, Huimin Ma 0001
ICCV3
2025 DADet: Safeguarding Image Conditional Diffusion Models Against Adversarial and Backdoor Attacks via Diffusion Anomaly Detection
Xinlong Ding, Jiawei Li 0016, Yudong Zhang 0008, Rongquan Wang, Huimin Ma 0001, Jiansheng Chen 0001
ICCV3
2025 LarTap: A Luminance-Aware Framework With Text-Correlation Priors for Multi-Exposure Image Fusion
abstract
Conventional imaging devices often struggle to produce high-dynamic-range (HDR) images that accurately represent natural scenes. To overcome this limitation, multi-exposure image fusion (MEF) techniques have been introduced as a viable solution. Existing MEF approaches aim to enhance performance by optimizing or searching architectures. However, they face challenges in precise feature extraction and scene reconstruction, leading to distortion in the fused images. Additionally, most methods do not adequately address luminance variations across different image regions, which may result in the loss of essential details. To address these challenges, we present a novel luminance-aware MEF framework that integrates text-correlation priors (LarTap). By embedding textual information into fusion process, the proposed framework enhances content extraction and comprehension. Specifically, it consist of two key components: the text-image correlation network (N1) and the multi-exposure fusion network (N2). First, N1 performs correlation training to achieve a holistic alignment between text and image pairs. Its iterative vision encoders (VEs) generate text-correlated prior knowledge to facilitate the fusion process in N2. Second, N2 leverages these priors for scene reconstruction and dynamically adjusts luminance based on comparative perception. Extensive experiments on multiple datasets demonstrate that LarTap outperforms state-of-the-art methods. The source code is available at https://github.com/EnLong-wang/LarTap.
Enlong Wang, Jiawei Li 0016, Tiantian Yan, Jia Lei 0001, Shihua Zhou, Bin Wang 0005, Jinyuan Liu 0001, Nikola K. Kasabov
IEEE Trans. Circuits Syst. Video Technol.2
2025 MLFuse: Multi-Scenario Feature Joint Learning for Multi-Modality Image Fusion
abstract
Multi-modality image fusion (MMIF) entails synthesizing images with detailed textures and prominent objects. Existing methods tend to use general feature extraction to handle different fusion tasks. However, these methods have difficulty breaking fusion barriers across various modalities owing to the lack of targeted learning routes. In this work, we propose a multi-scenario feature joint learning architecture, MLFuse, that employs the commonalities of multi-modality images to deconstruct the fusion progress. Specifically, we construct a cross-modal knowledge reinforcing network that adopts a multipath calibration strategy to promote information communication between different images. In addition, two professional networks are developed to maintain the salient and textural information of fusion results. The spatial-spectral domain optimizing network can learn the vital relationship of the source image context with the help of spatial attention and spectral attention. The edge-guided learning network utilizes the convolution operations of various receptive fields to capture image texture information. The desired fusion results are obtained by aggregating the outputs from the three networks. Extensive experiments demonstrate the superiority of MLFuse for infrared-visible image fusion and medical image fusion. The excellent results of downstream tasks (i.e., object detection and semantic segmentation) further verify the high-quality fusion performance of our method.
Jia Lei 0001, Jiawei Li 0016, Jinyuan Liu 0001, Bin Wang 0005, Shihua Zhou, Qiang Zhang 0008, Xiaopeng Wei, Nikola K. Kasabov
IEEE Trans. Multim.2
2024 Center of Pressure Estimation by Analyzing Walking Videos
abstract
Center of pressure (COP) serves as a widely utilized indicator for evaluating balance-related issues, e.g., gait quality of neurological disorders, fall risk of the elderly, and recovery of the injured. Existing methods for acquiring COP mostly rely on expensive force platforms or wearable force-sensing sensors that are relatively difficult to deploy. We propose a novel low-cost strategy to estimate COP with only visual information. Specifically, we first collect a database with walking videos. The plantar pressure of the subjects during walking is synchronously collected for calculating COP ground truth. Then, we propose an encoder-decoder model for estimating COPs using human pose and body shape as inputs. A perceptual codebook is used to tolerate the error in human pose estimation to improve the accuracy of the COP estimation. Experiments demonstrate the effectiveness of our proposed method. Our proposal achieves a correlation coefficient of 0.962, which is 0.8% better than the baseline model using Multi-layer Perceptron (MLP). The normalized RMSEs are improved by 13.9% and 19.0% in the anterior-posterior and medial-lateral directions, respectively. Compared to existing methods using wearable sensors or plantar pressure plates, our method is cheaper and easier to deploy.
Jiansheng Chen 0002, Yining Qin, Poyu Lin, Jiawei Li 0016, Youze Xue, Huimin Ma 0001
ICASSP4
2024 SDFuse: Semantic-injected dual-flow learning for infrared and visible image fusion
Enlong Wang, Jiawei Li 0016, Jia Lei 0001, Jinyuan Liu 0001, Shihua Zhou, Bin Wang 0005, Nikola K. Kasabov
Expert Syst. Appl.2
2024 GeSeNet: A General Semantic-Guided Network With Couple Mask Ensemble for Medical Image Fusion
abstract
At present, multimodal medical image fusion technology has become an essential means for researchers and doctors to predict diseases and study pathology. Nevertheless, how to reserve more unique features from different modal source images on the premise of ensuring time efficiency is a tricky problem. To handle this issue, we propose a flexible semantic-guided architecture with a mask-optimized framework in an end-to-end manner, termed as GeSeNet. Specifically, a region mask module is devised to deepen the learning of important information while pruning redundant computation for reducing the runtime. An edge enhancement module and a global refinement module are presented to modify the extracted features for boosting the edge textures and adjusting overall visual performance. In addition, we introduce a semantic module that is cascaded with the proposed fusion network to deliver semantic information into our generated results. Sufficient qualitative and quantitative comparative experiments (i.e., MRI-CT, MRI-PET, and MRI-SPECT) are deployed between our proposed method and ten state-of-the-art methods, which shows our generated images lead the way. Moreover, we also conduct operational efficiency comparisons and ablation experiments to prove that our proposed method can perform excellently in the field of multimodal medical image fusion. The code is available at https://github.com/lok-18/GeSeNet.
Jiawei Li 0016, Jinyuan Liu 0001, Shihua Zhou, Qiang Zhang 0008, Nikola K. Kasabov
IEEE Trans. Neural Networks Learn. Syst.1
2023 Learning a Graph Neural Network with Cross Modality Interaction for Image Fusion
abstract
Infrared and visible image fusion has gradually proved to be a vital fork in the field of multi-modality imaging technologies. In recent developments, researchers not only focus on the quality of fused images but also evaluate their performance in downstream tasks. Nevertheless, the majority of methods seldom put their eyes on mutual learning from different modalities, resulting in fused images lacking significant details and textures. To overcome this issue, we propose an interactive graph neural network (GNN)-based architecture between cross modality for fusion, called IGNet. Specifically, we first apply a multi-scale extractor to achieve shallow features, which are employed as the necessary input to build graph structures. Then, the graph interaction module can construct the extracted intermediate features of the infrared/visible branch into graph structures. Meanwhile, the graph structures of two branches interact for cross-modality and semantic learning, so that fused images can maintain the important feature expressions and enhance the performance of downstream tasks. Besides, the proposed leader nodes can improve information propagation in the same modality. Finally, we merge all graph features to get the fusion result. Extensive experiments on different datasets (i.e. TNO, MFNet, and M3FD) demonstrate that our IGNet can generate visually appealing fused images while scoring averagely 2.59% [email protected] and 7.77% mIoU higher in detection and segmentation than the compared state-of-the-art methods. The source code of the proposed IGNet can be available at https://github.com/lok-18/IGNet.
Jiawei Li 0016, Jiansheng Chen 0002, Jinyuan Liu 0001, Huimin Ma 0001
ACM Multimedia1
2023 Learning a Coordinated Network for Detail-Refinement Multiexposure Image Fusion
abstract
Nowadays, deep learning has made rapid progress in the field of multi-exposure image fusion. However, it is still challenging to extract available features while retaining texture details and color. To address this difficult issue, in this paper, we propose a coordinated learning network for detail-refinement in an end-to-end manner. Firstly, we obtain shallow feature maps from extreme over/under-exposed source images by a collaborative extraction module. Secondly, smooth attention weight maps are generated under the guidance of a self-attention module, which can draw a global connection to correlate patches in different locations. With the cooperation of the two aforementioned used modules, our proposed network can obtain a coarse fused image. Moreover, by assisting with an edge revision module, edge details of fused results are refined and noise is suppressed effectively. We conduct subjective qualitative and objective quantitative comparisons between the proposed method and twelve state-of-the-art methods on two available public datasets, respectively. The results show that our fused images significantly outperform others in visual effects and evaluation metrics. In addition, we also perform ablation experiments to verify the function and effectiveness of each module in our proposed method. The source code can be achieved athttps://github.com/lok-18/LCNDR.
Jiawei Li 0016, Jinyuan Liu 0001, Shihua Zhou, Qiang Zhang 0008, Nikola K. Kasabov
IEEE Trans. Circuits Syst. Video Technol.1