EDBT 2026 Demo / reviewers in the wild / expert
Xiaosong Li 0004
dblp:53/469-4
· DBLP profile ↗
19ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0003-4672-1527ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A location-aware network for nighttime image deraining with a semi-real synthetic paired benchmark dataset
Huichun Liu, Xiaosong Li 0004, Yang Liu 0335, Xiaoqi Cheng, Haishu Tan |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | CAWM-Mamba: A unified model for infrared-visible image fusion and compound adverse weather restoration
Huichun Liu, Xiaosong Li 0004, Zhuangfan Huang, Tao Ye 0002, Yang Liu 0335, Haishu Tan |
Expert Syst. Appl. | 2 |
| 2026 | FlexiD-Fuse: Flexible number of inputs multi-modal medical image fusion based on diffusion model
Yushen Xu, Xiaosong Li 0004, Xiaoqi Cheng, Huafeng Li 0001, Haishu Tan |
Expert Syst. Appl. | 2 |
| 2026 | FlexiSR-Diff: Flexible diffusion for multi-modal medical image fusion & super-resolution
Yushen Xu, Xiaosong Li 0004, Yang Liu 0335, Tao Ye 0002, Huafeng Li 0001 |
Pattern Recognit. | 2 |
| 2026 | AWM-Fuse: Multi-Modality Image Fusion for Adverse Weather via Global and Local Text PerceptionabstractMulti-modality image fusion (MMIF) in adverse weather aims to address the loss of visual information caused by weather-related degradations, providing clearer scene representations. Although a few studies have attempted to incorporate textual information to improve semantic perception, they often lack effective categorization and thorough analysis of textual content. To address these limitations, we propose AWM-Fuse, a unified fusion framework that handles diverse weather degradations via global and local text perception with shared parameters. In particular, a global text perception module leverages BLIP-generated captions to extract overall scene features and identify primary degradation types, thus promoting generalization across various adverse weather conditions. Complementing this, the local module employs detailed scene descriptions produced by ChatGPT to concentrate on specific degradation effects through concrete textual cues, enabling the recovery of subtle details. Furthermore, textual descriptions are used to constrain the generation of fused images, effectively steering the network learning process toward better alignment with semantic labels, thereby promoting the learning of more meaningful visual features. To facilitate text-guided fusion under adverse weather, we construct AWMM-Text, a large-scale benchmark providing paired global and local annotations for multi-modality image pairs. Extensive experiments demonstrate that AWM-Fuse consistently outperforms state-of-the-art methods under complex weather conditions and on multiple downstream tasks. Our code is available at https://github.com/Feecuin/AWM-Fuse. Xilai Li, Huichun Liu, Xiaosong Li 0004, Tao Ye 0002, Zhenyu Kuang, Huafeng Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | JDPNet: A Network Based on Joint Degradation Processing for Underwater Image EnhancementabstractGiven the complexity of underwater environments and the variability of water as a medium, underwater images are inevitably subject to various types of degradation. The degradations present nonlinear coupling rather than simple superposition, which renders the effective processing of such coupled degradations particularly challenging. Most existing methods focus on designing specific branches, modules, or strategies for specific degradations, with little attention paid to the potential information embedded in their coupling. Consequently, they struggle to effectively capture and process the nonlinear interactions of multiple degradations from a bottom-up perspective. To address this issue, we propose JDPNet, a joint degradation processing network, that mines and unifies the potential information inherent in coupled degradations within a unified framework. Specifically, we introduce a joint feature-mining module, along with a probabilistic bootstrap distribution strategy, to facilitate effective mining and unified adjustment of coupled degradation features. Furthermore, to balance color, clarity, and contrast, we design a novel AquaBalanceLoss to guide the network in learning from multiple coupled degradation losses. Experiments on six publicly available underwater datasets, as well as two new datasets constructed in this study, show that JDPNet exhibits state-of-the-art performance while offering a better tradeoff between performance, parameter size, and computational cost. Tao Ye 0002, Hongbin Ren, Chongbing Zhang, Xiaosong Li 0004 |
IEEE Trans. Image Process. | 5 |
| 2025 | PORSCHE: Progressive Optimization and Robust Spatial Convolution for Hybrid Enhancement in Visible-Infrared Vehicle Re-IdentificationabstractVisible-infrared vehicle re-identification has become crucial for stable 24-h surveillance of Visual Internet of Things (VIoT). It aims to leverage the shared information between different modalities to retrieve specific objects. Previous works primarily focus on projecting images from two modalities into a common embedding space to measure their similarity scores. However, the inherent distribution discrepancies between different modalities often lead models to rely on spurious features that are unrelated to vehicle identity, making effective feature fusion challenging. To address this unique problem, we propose the Progressive Optimization and Robust Spatial Convolution for Hybrid Enhancement (PORSCHE) model, which reduces the negative effects of spurious correlations and biases toward training pairs. Specifically, we introduce the Patch-wise Matching (PAM) module, which performs initial coarse-grained alignment between different modalities. Building upon this foundation, we develop the Point-wise Matching (POM) module to achieve fine-grained discriminative alignment through precise point-level feature matching, thereby enhancing identity-specific representation. To optimize these complementary PAM and POM components effectively, we implement a progressive training strategy that hierarchically refines feature representations from local patches to global structures, ensuring stable learning of modality-invariant characteristics. This coarse-to-fine architecture enables gradual fusion and alignment across modalities at both patch and point levels, effectively capturing the essential discriminative features required for robust cross-modality retrieval. Extensive experimental results on MSVR310, WMVeID863, and RGBN300 benchmarks demonstrate the effectiveness of our proposed method. The code will be available at https://github.com/HowardLiu28/PORSCHE. Yinhao Liu, Zhenyu Kuang, Yige Ma, Xinghao Ding, Yue Huang 0001, Congbo Cai, Xiaosong Li 0004 |
IEEE Internet Things J. | 8 |
| 2025 | UMCFuse: A Unified Multiple Complex Scenes Infrared and Visible Image Fusion FrameworkabstractInfrared and visible image fusion has emerged as a prominent research area in computer vision. However, little attention has been paid to the fusion task in complex scenes, leading to sub-optimal results under interference. To fill this gap, we propose a unified framework for infrared and visible images fusion in complex scenes, termed UMCFuse. Specifically, we classify the pixels of visible images from the degree of scattering of light transmission, allowing us to separate fine details from overall intensity. Maintaining a balance between interference removal and detail preservation is essential for the generalization capacity of the proposed method. Therefore, we propose an adaptive denoising strategy for the fusion of detail layers. Meanwhile, we fuse the energy features from different modalities by analyzing them from multiple directions. Extensive fusion experiments on real and synthetic complex scenes datasets cover adverse weather conditions, noise, blur, overexposure, fire, as well as downstream tasks including semantic segmentation, object detection, salient object detection, and depth estimation, consistently indicate the superiority of the proposed method compared with the recent representative methods. Our code is available at https://github.com/ixilai/UMCFuse. Xilai Li, Xiaosong Li 0004, Tianshu Tan, Huafeng Li 0001, Tao Ye 0002 |
IEEE Trans. Image Process. | 2 |
| 2024 | SAMF: Small-Area-Aware Multi-Focus Image Fusion for Object DetectionabstractExisting multi-focus image fusion (MFIF) methods often fail to preserve the uncertain transition region and detect small focus areas within large defocused regions accurately. To address this issue, this study proposes a new small-area-aware MFIF algorithm for enhancing object detection capability. First, we enhance the pixel attributes within the small focus and boundary regions, which are subsequently combined with visual saliency detection to obtain the pre-fusion results used to discriminate the distribution of focused pixels. To accurately ensure pixel focus, we consider the source image as a combination of focused, defocused, and uncertain regions and propose a three-region segmentation strategy. Finally, we design an effective pixel selection rule to generate segmentation decision maps and obtain the final fusion results. Experiments demonstrated that the proposed method can accurately detect small and smooth focus areas while improving object detection performance, outperforming existing methods in both subjective and objective evaluations. The source code is available at https://github.com/ixilai/SAMF. Xilai Li, Xiaosong Li 0004, Haishu Tan |
ICASSP | 2 |
| 2024 | Simultaneous Tri-Modal Medical Image Fusion and Super-Resolution Using Conditional Diffusion Model
Yushen Xu, Xiaosong Li 0004, Yuchan Jie, Haishu Tan |
MICCAI (7) | 2 |
| 2024 | Bridging the Gap between Multi-focus and Multi-modal: A Focused Integration Framework for Multi-modal Image FusionabstractMulti-modal image fusion (MMIF) integrates valuable information from different modality images into a fused one. However, the fusion of multiple visible images with different focal regions and infrared images is a unprecedented challenge in real MMIF applications. This is because of the limited depth of the focus of visible optical lenses, which impedes the simultaneous capture of the focal information within the same scene. To address this issue, in this paper, we propose a MMIF framework for joint focused integration and modalities information extraction. Specifically, a semi-sparsity-based smoothing filter is introduced to decompose the images into structure and texture components. Subsequently, a novel multi-scale operator is proposed to fuse the texture components, capable of detecting significant information by considering the pixel focus attributes and relevant data from various modal images. Additionally, to achieve an effective capture of scene luminance and reasonable contrast maintenance, we consider the distribution of energy information in the structural components in terms of multi-directional frequency variance and information entropy. Extensive experiments on existing MMIF datasets, as well as the object detection and depth estimation tasks, consistently demonstrate that the proposed algorithm can surpass the state-of-the-art methods in visual perception and quantitative evaluation. The code is available at https://github.com/ixilai/MFIF-MMIF. Xilai Li, Xiaosong Li 0004, Tao Ye 0002, Xiaoqi Cheng, Wuyang Liu, Haishu Tan |
WACV | 2 |
| 2024 | Generative Adversarial Network for Trimodal Medical Image Fusion Using Primitive Relationship ReasoningabstractMedical image fusion has become a hot biomedical image processing technology in recent years. The technology coalesces useful information from different modal medical images onto an informative single fused image to provide reasonable and effective medical assistance. Currently, research has mainly focused on dual-modal medical image fusion, and little attention has been paid on trimodal medical image fusion, which has greater application requirements and clinical significance. For this, the study proposes an end-to-end generative adversarial network for trimodal medical image fusion. Utilizing a multi-scale squeeze and excitation reasoning attention network, the proposed method generates an energy map for each source image, facilitating efficient trimodal medical image fusion under the guidance of an energy ratio fusion strategy. To obtain the global semantic information, we introduced squeeze and excitation reasoning attention blocks and enhanced the global feature by primitive relationship reasoning. Through extensive fusion experiments, we demonstrate that our method yields superior visual results and objective evaluation metric scores compared to state-of-the-art fusion methods. Furthermore, the proposed method also obtained the best accuracy in the glioma segmentation experiment. Jingxue Huang, Xiaosong Li 0004, Haishu Tan, Xiaoqi Cheng |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Multimodal Medical Image Fusion Based on Multichannel Aggregated Network
Jingxue Huang, Xiaosong Li 0004, Haishu Tan, Xiaoqi Cheng |
ICIG (5) | 2 |
| 2023 | FUFusion: Fuzzy Sets Theory for Infrared and Visible Image Fusion
Yuchan Jie, Xiaosong Li 0004, Haishu Tan, Xiaoqi Cheng |
PRCV (2) | 3 |
| 2023 | Medical image fusion based on extended difference-of-Gaussians and edge-preserving
Yuchan Jie, Xiaosong Li 0004, Fuqiang Zhou, Haishu Tan |
Expert Syst. Appl. | 2 |
| 2021 | Multimodal medical image fusion based on joint bilateral filter and local gradient energy
Xiaosong Li 0004, Fuqiang Zhou, Haishu Tan, Wanning Zhang, Congyang Zhao |
Inf. Sci. | 1 |
| 2021 | Joint image fusion and denoising via three-layer decomposition and sparse representation
Xiaosong Li 0004, Fuqiang Zhou, Haishu Tan |
Knowl. Based Syst. | 1 |
| 2021 | Multi-focus image fusion based on nonsubsampled contourlet transform and residual removal
Xiaosong Li 0004, Fuqiang Zhou, Haishu Tan, Yuanze Chen, Wangxia Zuo |
Signal Process. | 1 |
| 2016 | Multifocus image fusion by combining with mixed-order structure tensors and multiscale neighborhood
Huafeng Li 0001, Xiaosong Li 0004, Zhengtao Yu 0001, Cunli Mao |
Inf. Sci. | 2 |