VLDB 2026 Research / reviewers in the wild / expert
Xiaoli Zhang 0001
dblp:67/6767-1
· DBLP profile ↗
64ranked-venue papers
8as first author
48since 2021 · last 2026
0000-0001-8412-4956ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 6 first-author · 26 since 2021Artificial intelligence and machine learning · 19 · 2 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ControlFuse: Instruction-guided Multi-Granularity Controllable Image FusionabstractInfrared and Visible Image Fusion (IVIF) produces enhanced images by fusing complementary visual information. However, most existing methods generate fixed outputs and cannot flexibly adapt to user-specific requirements. Recent text-guided approaches offer partial control but are limited to global or semantic levels, lacking instance-level control. This limitation arises from two challenges: first, the lack of datasets that directly link textual instructions with corresponding spatial annotations, and second, the use of coarse cross-modal alignment methods that struggle to precisely match textual instructions with visual features. To overcome these challenges, we propose ControlFuse, a controllable IVIF framework enabling multi-granularity fusion across global, semantic, and instance levels, guided by user instructions. First, we construct an automated multi-granularity dataset that provides explicit textual-mask correspondences at these three levels. Second, inspired by manifold geometry, we design a Multimodal Feature Interaction Module (MFIM) comprising Feature Manifold Converter (FMC) and Curvature-Guided Interaction (CGI). FMC projects textual and visual features into a unified manifold space, while CGI leverages manifold curvature as a geometric cue to refine cross-modal alignment. Extensive experiments validate ControlFuse, outperforming state-of-the-art methods in robustness and flexibility. Xiaoli Zhang 0001, Zeyu Wang 0009 |
AAAI | 2 |
| 2026 | Diff-MEF: An efficient frequency domain-based diffusion model for multi-exposure image fusion
Kai Diao, Lisi Wei, Xihang Hu, Xiaoli Zhang 0001 |
Expert Syst. Appl. | 4 |
| 2026 | RL-CoSeg: A reinforcement learning-based collaborative localization and segmentation framework for medical image
Feilong Xu, Feiyang Yang, Xiaoli Zhang 0001, Zhaojun Liu |
Expert Syst. Appl. | 3 |
| 2026 | ATDFusion: Adapter-Tuned Dual-Branch Network for multimodal medical image fusionabstractMultimodal medical image fusion aims to combine complementary visual information from different imaging modalities to assist diagnosis and clinical decision-making. Existing methods often struggle to balance global semantic representation and local detail preservation, leading to blurred or incomplete features in salient regions. This work proposes ATDFusion (Adapter-Tuned Dual-Branch Network), featuring a global semantic branch and an auxiliary detail branch to jointly capture high-level context and fine-grained details. A Low-Rank Dynamic Token Adapter (LR-DTA) adaptively fine-tunes intermediate layers of pretrained models based on token number and rank. Additionally, a Fusion Enhancement Guidance (FEG) module imposes explicit spatial supervision via a saliency-aware loss to strictly preserve diagnostically critical regions. On three typical tasks (MRI-CT, MRI-PET, MRI-SPECT), ATDFusion surpasses state-of-the-art methods, improving SSIM and CC by 11.93% and 5.07%, respectively. Other metrics (e.g., EN, PSNR) also achieve leading results, validating its effectiveness. Furthermore, the model demonstrates strong zero-shot generalization on the unseen HECKTOR 2025 dataset. Code is available at https://github.com/pluto628/ATDFusion . Rui Zhu 0015, Hang Zhao 0001, Xiaoli Zhang 0001 |
Pattern Recognit. | 5 |
| 2026 | A Semantic-Aware and Multi-Guided Network for Infrared-Visible Image FusionabstractMulti-modality image fusion aims at fusing modality-specific (complementarity) and modality-shared (correlation) information from multiple source images. To tackle the overlooking of inter-feature relationships, high-frequency information loss, and the limited attention to downstream tasks, this paper focuses on efficiently extracting complementary in formation and aggregating multi-guided features. We propose a three-branch encoder-decoder architecture along with corresponding fusion layers as the fusion strategy. Firstly, shallow features from individual modalities are extracted by a depthwise convolution layer combined with the transformer block. In the three parallel branches of the encoder, Cross Attention and Invertible Block (CAI) extracts local features and preserves high frequency texture details. Base Feature Extraction Module (BFE) captures long-range dependencies and enhances modality-shared information. Graph Reasoning Module (GR) is introduced to reason high-level cross-modality relations and simultaneously ex tract low-level detail features as CAI's modality-specific complementary information. Experiments demonstrate the competitive results compared with state-of-the-art methods in visible/infrared image fusion and medical image fusion tasks. Moreover, the proposed algorithm surpasses the state-of-the-art methods in terms of subsequent tasks, averagely scoring 8.27% [email protected] higher in object detection and 5.85% mIoU higher in semantic segmentation. Xiaoli Zhang 0001, Liying Wang 0006, Siwei Ma 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | A Unified Loss for Handling Inter-Class and Intra-Class Imbalance in Medical Image SegmentationabstractIn utilizing deep learning techniques for medical image segmentation, two types of imbalance issues are observed: inter-class imbalance between majority and minority classes and intra-class imbalance between easy and hard samples. However, existing loss functions typically confuse these issues, leading to enhancements that cater to only one aspect. Moreover, loss functions optimized for specific tasks often exhibit limited generalizability. To address these issues, we propose Inter-class and Intra-class Balance loss, as well as a unified loss termed Balance loss. The Inter-class Balance loss controls the extent of hard sample mining for majority class samples by considering the frequency of minority classes present in each input image. This approach requires no manual adjustment weights and adapts automatically to different datasets. The Intra-class Balance loss enhances the network's ability to learn from hard samples by performing mining on hard samples within each class. We evaluate our loss functions on five segmentation tasks with varying degrees of class imbalance. The experimental results show that our proposed Balance loss enhances segmentation performance compared with the current loss functions and exhibits superior robustness. Feilong Xu, Feiyang Yang, Xiaoli Zhang 0001 |
AAAI | 4 |
| 2025 | MFANet: Multi-Feature Aggregation Network for Multi-focus Image FusionabstractExisting deep learning-based Multi-focus Image Fusion (MFIF) methods often rely on loss functions derived from linear combinations of image quality metrics, leading to complexities in training and only marginal improvements in image quality. Recognizing this, our study identifies input space and scale information as pivotal in enhancing MFIF performance. By augmenting raw spatial images with other feature spaces, i.e., gradient and dense Scale-Invariant Feature Transform (DSIFT), we enhance the model’s ability to detect edges, textures, and local structures, facilitating more accurate differentiation between focused and defocused areas. Additionally, smaller scale variations improve focus detection, while multi-scale learning within neural networks effectively suppresses artifacts without affecting focus detection accuracy. To achieve the above enhancements, we introduce the Multi-Feature Aggregation Network (MFANet), which employs a three-branch architecture to perform focused detection process in spatial, gradient, and DSIFT feature spaces. Each branch is equipped with a Pyramid Attention Fusion (PAF) module that utilizes attention mechanisms and a novel Light Spatial Aggregation Pyramid Module (LSAPM) to capture global feature relationships and aggregate multi-scale information. Experimental results demonstrate that MFANet surpasses other state-of-the-art fusion methods in both qualitative and quantitative evaluations. Xiaoli Zhang 0001, Mingjie Tian, Zeyu Wang 0009 |
ICASSP | 2 |
| 2025 | Cross-Counter-Repeat Attention for Enhanced Understanding of Visual Semantics in Radiology Report GenerationabstractRadiology report generation (RRG), intended to automatically generate a coherent free-text report describing the clinical observations of a radiograph, has been attracting increasing attention from researchers. In recent years, the Transformer-based encoder-decoder architecture has been adopted by most existing methods. However, they neglect the structural rationality issue when applying this single-modal architecture to the multi-modal RRG task, where information can only flow from visual features to textual features, but not in the opposite direction. This information asymmetry results in visual features having no knowledge of the textual features, sending out all visual information, including a large amount of heterogeneous noise. Consequently, this introduces significant resistance to the downstream decoder, which substantially limits or even harms the generation process. To tackle this problem, we present a method where a cross-counter-repeat attention is developed to integrate useful information from two separate modalities, and a memory-driven visual semantics enhancing module is designed to reinforce the visual features with strong time-ordered semantic information. Experimental results on the widely-used IU-Xray dataset show that our approach achieves the state-of-the-art performance, with a remarkable 6.9% improvement in BLEU-4 score. Further analyses also demonstrate that our method can generate sufficiently comprehensive reports to assist radiologists in their clinical decision-making. Xiaolei Bo, Feiyang Yang, Feilong Xu, Xiaoli Zhang 0001 |
ACM Multimedia | 4 |
| 2025 | ST-SAM: SAM-Driven Self-Training Framework for Semi-Supervised Camouflaged Object DetectionabstractSemi-supervised Camouflaged Object Detection (SSCOD) aims to reduce reliance on costly pixel-level annotations by leveraging limited annotated data and abundant unlabeled data. However, existing SSCOD methods based on Teacher-Student frameworks suffer from severe prediction bias and error propagation under scarce supervision, while their multi-network architectures incur high computational overhead and limited scalability. To overcome these limitations, we propose ST-SAM, a highly annotation-efficient yet concise framework that breaks away from conventional SSCOD constraints. Specifically, ST-SAM employs Self-Training strategy that dynamically filters and expands high-confidence pseudo-labels to enhance a single-model architecture, thereby fundamentally circumventing inter-model prediction bias. Furthermore, by transforming pseudo-labels into hybrid prompts containing domain-specific knowledge, ST-SAM effectively harnesses the Segment Anything Model's potential for specialized tasks to mitigate error accumulation in self-training. Experiments on COD benchmark datasets demonstrate that ST-SAM achieves state-of-the-art performance with only 1% labeled data, outperforming existing SSCOD methods and even matching fully supervised methods. Remarkably, ST-SAM requires training only a single network, without relying on specific models or loss functions. This work establishes a new paradigm for annotation-efficient SSCOD. Codes will be available at https://github.com/hu-xh/ST-SAM. Xihang Hu, Fuming Sun, Jiazhe Liu, Feilong Xu, Xiaoli Zhang 0001 |
ACM Multimedia | 5 |
| 2025 | An interpretable bilateral detail optimization deep unfolding network for pansharpening
Yufei Ge, Xiaoli Zhang 0001, Siwei Ma 0001 |
Neurocomputing | 2 |
| 2025 | Self-supervised multi-blind network for real image denoising via multivariate Gaussian-poisson noise
Hang Zhao 0001, Xiaoli Zhang 0001, Zhaojun Liu |
Neurocomputing | 3 |
| 2025 | Focusing on neglected natural images: A self-supervised learning model for pan-sharpening
Xiaoli Zhang 0001, Zeyu Wang 0009 |
Inf. Process. Manag. | 2 |
| 2025 | A deep unfolding network based on intrinsic image decomposition for pansharpening
Yufei Ge, Xiaoli Zhang 0001, Siwei Ma 0001 |
Knowl. Based Syst. | 2 |
| 2025 | BPDUN: Bidirectional Progressive Deep Unfolding Network for PansharpeningabstractPansharpening integrates low-resolution multispectral (LRMS) images with high-resolution panchromatic (PAN) images to produce high-resolution multispectral (HRMS) images for subsequent tasks. Existing methods have improved the image quality but still suffer from spectral distortion and blurry effects due to the inherent problems of scale inconsistency and frequency mismatch between source images. Although some deep learning (DL)-based methods adopt the progressive fusion strategy to mitigate their impact, they remain opaque and fail to incorporate domain-specific knowledge, limiting generalization ability. To address these problems, we propose a novel bidirectional progressive deep unfolding network (BPDUN) for pansharpening by decoupling the fusion process into a multistage PAN-guided multispectral (MS) image restoration task. In the forward direction, each stage performs iterative restoration of MS images at a specific spatial scale. Each iteration includes a prior denoising block (PDB) and an image reconstruction block (IRB) derived from variational optimization (VO) formulas with clear physical interpretations. The designed PDB implements band-aware prior fusion and dual-domain denoising, while a content-guided pixel attention module (CGPAM) is proposed for detail refinement. In the reverse direction, PAN images undergo two levels of decomposition via wavelet transform. The low-frequency component serves as a spatial prior to guide detail recovery for MS images at the same scale, while the high-frequency components are modulated between stages via a detail-enhanced upsampling module (DEUM). The proposed DEUM performs adaptive high-frequency fine-tuning to ensure spectral fidelity while increasing spatial detail during upsampling. Extensive experiments demonstrate that our method outperforms other state-of-the-art (SOTA) algorithms in spectral and spatial fidelity. Yufei Ge, Xiaoli Zhang 0001, Rui Zhu 0015, Siwei Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | DINet: Depth-Guided and Iterative Refinement Network for Salient Object Detection in Optical Remote Sensing ImagesabstractOptical remote sensing images (ORSI) feature unique scenes and complex imaging conditions. Specifically, they exhibit substantial variations in object scale, quantity, structure, and distribution. Consequently, salient object detection in ORSI (ORSI-SOD) is pivotal in ORSI content perception and understanding. Additionally, the limitations of the single modality impede the advancement of ORSI-SOD. To tackle these issues, we propose a Depth-guided and Iterative Refinement Network (DINet) for ORSI-SOD. By incorporating depth information as auxiliary cues, we introduce a multi-modal strategy for ORSI-SOD, resulting in improved accuracy in the localization and segmentation of salient objects. To address the variability of salient objects, we design an Aggregation Perception Enhancement (APE) Module. This module integrates complementary cues from cross-modal features using multi-dimensional attention mechanisms. By fostering cross-modal interactions, the APE module effectively preserves both detail and spatial location information. Furthermore, we propose an Iterative Guidance Refinement Decoder to handle boundary uncertainty. The decoder uses initial predictions to guide the decoding phase and iteratively refine results. Simultaneously, it minimizes noise from depth cues, yielding predictions with more accurate boundaries. Experimental comparisons with 22 state-of-the-art methods show that DINet exhibits superior performance while maintaining lightweight (11.98M) and real-time (55FPS) capabilities. Xihang Hu, Fuming Sun, Xiaoli Zhang 0001, Chuanmin Jia, Siwei Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | MAFS: Masked Autoencoder for Infrared-Visible Image Fusion and Semantic SegmentationabstractInfrared-visible image fusion methods aim at generating fused images with good visual quality and also facilitate the performance of high-level tasks. Indeed, existing semantic-driven methods have considered semantic information injection for downstream applications. However, none of them investigates the potential for reciprocal promotion between pixel-wise image fusion and cross-modal feature fusion perception tasks from a macroscopic task-level perspective. To address this limitation, we propose a unified network for image fusion and semantic segmentation. MAFS is a parallel structure, containing a fusion sub-network and a segmentation sub-network. On the one hand, we devise a heterogeneous feature fusion strategy to enhance semantic-aware capabilities for image fusion. On the other hand, by cascading the fusion sub-network and a segmentation backbone, segmentation-related knowledge is transferred to promote feature-level fusion-based segmentation. Within the framework, we design a novel multi-stage Transformer decoder to aggregate fine-grained multi-scale fused features efficiently. Additionally, a dynamic factor based on the max-min fairness allocation principle is introduced to generate adaptive weights of two tasks and guarantee smooth training in a multi-task manner. Extensive experiments demonstrate that our approach achieves competitive results compared with state-of-the-art methods. The code is available at https://github.com/Abraham-Einstein/MAFS/. Liying Wang 0006, Xiaoli Zhang 0001, Chuanmin Jia, Siwei Ma 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | MAINet: Modality-Aware Interaction Network for Medical Image FusionabstractDue to the limitations of imaging sensors, obtaining a medical image that simultaneously captures both functional metabolic data and structural tissue details remains a significant challenge in clinical diagnosis. To address this, Multimodal Medical Image Fusion (MMIF) has emerged as an effective technique for integrating complementary information from multimodal source images, such as CT, PET, and SPECT, which is critical for providing a comprehensive understanding of both anatomical and functional aspects of the human body. One of the key challenges in MMIF is how to exchange and aggregate this multimodal information. This article rethinks MMIF by addressing the harmony of modality gaps and proposes a novel Modality-Aware Interaction Network (MAINet), which leverages cross-modal feature interaction and progressively fuses multiple features in graph space. Specifically, we introduce two key modules: the Cascade Modality Interaction (CMI) module and the Dual-Graph Learning (DGL) module. The CMI module, integrated within a multi-scale encoder with triple branches, facilitates complementary multimodal feature learning and provides beneficial feedback to enhance discriminative feature learning across modalities. In the decoding process, the DGL module aggregates hierarchical features in two distinct graph spaces, enabling global feature interactions. Moreover, the DGL module incorporates a bottom-up guidance mechanism, where deeper semantic features guide the learning of shallower detail features, thus improving the fusion process by enhancing both scale diversity and modality awareness for visual fidelity results. Experimental results on medical image datasets demonstrate the superiority of the proposed method over existing fusion approaches in both subjective and objective evaluations. We also validated the performance of the proposed method in applications such as infrared-visible image fusion and medical image segmentation. Lisi Wei, Xiaoli Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Corrigendum: MAINet: Modality-Aware Interaction Network for Medical Image FusionabstractThis is a corrigendum for the article “MAINet: Modality-Aware Interaction Network for Medical Image Fusion” published in ACM Trans. Multimedia Comput. Commun. Appl. 21(6): 166:1–166:23 (2025). Lisi Wei, Xiaoli Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Patch-Level Knowledge Distillation and Regularization for Missing Modality Medical Image SegmentationabstractIn the context of medical image segmentation, complementary information among multi-modality images can improve segmentation performance. However, acquiring the complete multi-modality data in clinical settings is difficult. To tackle this problem, we propose a novel multi-modality knowledge distillation segmentation framework, which allows the inference performance of single-modality closer to that of multi-modality. In order to facilitate the extraction of valuable information from the multi-modality teacher network, we first introduce a subtask named patch-selection to distill the patch-level knowledge and improve the generalization capacity of networks simultaneously. Moreover, we employ contrastive learning distillation by defining patch-level positive and negative pairs in embedding, which can encourage the student network to extract more potential information from single-modality input and better understand the similarities and differences with the teacher network in representations. The evaluation process on the BraTS 2018 dataset shows the state-of-the-art performance of our method. Mingjie Tian, Feiyang Yang, Xiaoli Zhang 0001 |
ICASSP | 5 |
| 2024 | DFANet: A Dual-Stream Deep Feature Aware Network for Multi-focus Image Fusion
Yuye Dong, Xiaoli Zhang 0001 |
PRCV (15) | 4 |
| 2024 | Triple-loss driven generative adversarial network for pansharpeningabstractAbstract Pansharpening aims at fusing a panchromatic (PAN) image and a low‐resolution multispectral (LRMS) image into a high‐resolution multispectral (HRMS) image. In recent years, GAN‐based pansharpening methods have achieved excellent results, but they suffer from inadequate feature preservation and unstable training. To address these issues, a novel GAN‐based model named TriLossGAN is proposed. This method constructs three loss components with the help of the generator and the dual‐discriminator, which are calculated in both the original spatial domain and the transform domain to better preserve high‐frequency and low‐frequency information in the fused image. Additionally, a new training strategy is designed to stabilize the training process. In extensive experiments, the proposed method achieved satisfactory results on three datasets with QNR values of 0.9584 on GaoFen‐2, 0.9601 on QuickBird, and 0.9138 on WorldView‐3. Qualitative and quantitative comparisons demonstrate that TriLossGAN outperforms other state‐of‐the‐art methods. Xiaoli Zhang 0001 |
IET Image Process. | 3 |
| 2024 | SWPanGAN: A hybrid generative adversarial network for pansharpeningabstractAbstract Pansharpening is a vital technique in remote sensing that combines a low‐resolution multi‐spectral image with its corresponding panchromatic image to obtain a high‐resolution multi‐spectral image. Despite its potential benefits, the challenge lies in extracting features from the source images and eliminating artefacts in the fused images. In response to the challenge, a hybrid generative adversarial network‐based model, termed SWPanGAN, is proposed. For better feature extraction, the conventional convolution neural network is replaced with a Swin transformer in the generator, which provides the generator with the ability to model long‐range dependencies. Additionally, to suppress artefacts, a wavelet‐based discriminator is proposed for effectively distinguishing the frequency discrepancy. With these modifications, both the generator and discriminator networks of SWPanGAN are enhanced. Extensive experiments illustrate that our SWPanGAN can generate high‐quality pansharpening images and surpass other state‐of‐the‐art methods. Xiaoli Zhang 0001 |
IET Image Process. | 3 |
| 2024 | GCFormer: Multi-scale feature plays a crucial role in medical images segmentation
Yuncong Feng, Yeming Cong, Shuaijie Xing, Zihang Ren, Xiaoli Zhang 0001 |
Knowl. Based Syst. | 6 |
| 2024 | Electronic explosives inspection: a fine-grained X-ray benchmark and few-shot prohibited phone detection model
Jianzhao Cui, Xiaoli Zhang 0001, Sa Huang, Yuncong Feng |
Multim. Tools Appl. | 3 |
| 2024 | An image segmentation fusion algorithm based on density peak clustering and Markov random field
Yuncong Feng, Wanru Liu, Xiaoli Zhang 0001 |
Multim. Tools Appl. | 3 |
| 2024 | Efficient Camouflaged Object Detection Network Based on Global Localization Perception and Local Guidance RefinementabstractCamouflaged Object Detection (COD) is a challenging visual task due to its complex contour, diverse scales, and high similarity to the background. Existing COD methods encounter two predicaments: One is that they are prone to falling into local perception, resulting in inaccurate object localization; Another issue is the difficulty in achieving precise object segmentation due to a lack of detailed information. In addition, most COD methods typically require larger parameter amounts and higher computational complexity in pursuit of better performance. To this end, we propose a global localization perception and local guidance refinement network (PRNet), that simultaneously addresses performance and computational costs. Through effective aggregation and use of semantic and details information, the PRNet can achieve accurate localization and refined segmentation of camouflaged objects. Specifically, with the help of a Cascaded Attention Perceptron (CAP) designed, we can effectively integrate and perceive multi-scale information to localize camouflaged objects. We also design a Guided Refinement Decoder (GRD) in a top-down manner to extract context information and aggregate details to further refine camouflaged prediction results. Extensive experimental results demonstrate that our PRNet outperforms 12 state-of-the-art models on 4 challenging datasets. Meanwhile, the PRNet has a smaller number of parameters (12.74M), lower computational complexity (10.24G), and real-time inference speed (105FPS). Source codes are available at https://github.com/hu-xh/PRNet. Xihang Hu, Xiaoli Zhang 0001, Fasheng Wang, Jing Sun 0012, Fuming Sun |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | MRL-Seg: Overcoming Imbalance in Medical Image Segmentation With Multi-Step Reinforcement LearningabstractMedical image segmentation is a critical task for clinical diagnosis and research. However, dealing with highly imbalanced data remains a significant challenge in this domain, where the region of interest (ROI) may exhibit substantial variations across different slices. This presents a significant hurdle to medical image segmentation, as conventional segmentation methods may either overlook the minority class or overly emphasize the majority class, ultimately leading to a decrease in the overall generalization ability of the segmentation results. To overcome this, we propose a novel approach based on multi-step reinforcement learning, which integrates prior knowledge of medical images and pixel-wise segmentation difficulty into the reward function. Our method treats each pixel as an individual agent, utilizing diverse actions to evaluate its relevance for segmentation. To validate the effectiveness of our approach, we conduct experiments on four imbalanced medical datasets, and the results show that our approach surpasses other state-of-the-art methods in highly imbalanced scenarios. These findings hold substantial implications for clinical diagnosis and research. Feiyang Yang, Haoran Duan 0001, Feilong Xu, Yawen Huang, Xiaoli Zhang 0001, Yang Long 0001, Yefeng Zheng 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2023 | Detecting sparse building change with ambiguous label using Siamese full-scale connected network and instance augmentation
Xinze Lin, Zeyu Wang 0009, Xiaoli Zhang 0001 |
Appl. Intell. | 4 |
| 2023 | When Multi-Focus Image Fusion Networks Meet Traditional Edge-Preservation Technology
Zeyu Wang 0009, Haoran Duan 0001, Xiaoli Zhang 0001 |
Int. J. Comput. Vis. | 7 |
| 2023 | Wide aspect ratio matching for robust face detection
Shi Luo, Xiaoli Zhang 0001 |
Multim. Tools Appl. | 3 |
| 2023 | ERMF: Edge refinement multi-feature for change detection in bitemporal remote sensing images
Zixuan Song, Rui Zhu 0015, Zeyu Wang 0009, Xiaoli Zhang 0001 |
Signal Process. Image Commun. | 6 |
| 2023 | VSP-Fuse: Multifocus Image Fusion Model Using the Knowledge Transferred From Visual Salience PriorsabstractMultifocus image fusion (MFIF), as an efficient way to improve the visual effect of images with partial focus defects, is of great significance in the field of image enhancement. According to the imaging principle of the lens, we summarize the visual salience priors (VSP) from the daily photo scene and two relationships from MFIF. Thereby, an edge-sensitive model for MFIF is presented in this study. Supported by VSP, we consider the correlation between salience object detection (SOD) and MFIF, and select the former as a pre-training task. SOD provides the network with realistic depth of field and bokeh effects to learn, and enhances the network’s ability to extract and express the edges of focused objects. Meanwhile, given the scarcity of real multifocus training sets, we propose a randomized approach to generate massive training sets and pseudo-labels based on limited unlabeled data. Besides, two attention modules are designed based on isometric domain transformation (IDT) in the traditional edge-preservation field. IDT removes interference information from feature maps in a low-cost manner, thereby facilitating channel-wise and spatial-wise weight assignments. Experimental results on four datasets show that the performance of our model is superior to that of many supervised models, without the need of any real MFIF training set. Zeyu Wang 0009, Haoran Duan 0001, Xiaoli Zhang 0001, Jizheng Zhang, Shiping Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Medical Image Fusion Based on Pixel-Level Nonlocal Self-similarity Prior and Optimization
Rui Zhu 0015, Yu Wang 0164, Xiaoli Zhang 0001 |
DASFAA (3) | 4 |
| 2022 | IBMvSVM: An instance-based multi-view SVM algorithm for classification
Siru Sun, Hancheng Wang, Xiaoli Zhang 0001, Shiping Chen 0001 |
Appl. Intell. | 5 |
| 2022 | Multimodal medical image fusion using adaptive co-occurrence filter-based decomposition optimization modelabstractMOTIVATION: Medical image fusion has developed into an important technology, which can effectively merge the significant information of multiple source images into one image. Fused images with abundant and complementary information are desirable, which contributes to clinical diagnosis and surgical planning. RESULTS: In this article, the concept of the skewness of pixel intensity (SPI) and a novel adaptive co-occurrence filter (ACOF)-based image decomposition optimization model are proposed to improve the quality of fused images. Experimental results demonstrate that the proposed method outperforms 22 state-of-the-art medical image fusion methods in terms of five objective indices and subjective evaluation, and it has higher computational efficiency. AVAILABILITY AND IMPLEMENTATION: First, the concept of SPI is applied to the co-occurrence filter to design ACOF. The initial base layers of source images are obtained using ACOF, which relies on the contents of images rather than fixed scale. Then, the widely used iterative filter framework is replaced with an optimization model to ensure that the base layer and detail layer are sufficiently separated and the image decomposition has higher computational efficiency. The optimization function is constructed based on the characteristics of the ideal base layer. Finally, the fused images are generated by designed fusion rules and linear addition. The code and data can be downloaded at https://github.com/zhunui/acof. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rui Zhu 0015, Sa Huang, Xiaoli Zhang 0001 |
Bioinform. | 4 |
| 2022 | Bounding-box deep calibration for high performance face detectionabstractAbstract Modern convolutional neural networks (CNNs)‐based face detectors have achieved tremendous strides due to large annotated datasets. However, misaligned results with high detection confidence but low localization accuracy restrict the further improvement of detection performance. In this paper, the authors first predict high confidence detection results on the training set itself. Surprisingly, a considerable part of them exist in the same misalignment problem. Then, the authors carefully examine these cases and point out that annotation misalignment is the main reason. Later, a comprehensive discussion is given for the replacement rationality between predicted and annotated bounding‐boxes. Finally, the authors propose a novel Bounding‐Box Deep Calibration (BDC) method to reasonably replace misaligned annotations with model predicted bounding‐boxes and offer calibrated annotations for the training set. Extensive experiments on multiple detectors and two popular benchmark datasets show the effectiveness of BDC on improving models' precision and recall rate, without adding extra inference time and memory consumption. Our simple and effective method provides a general strategy for improving face detection, especially for light‐weight detectors in real‐time situations. Shi Luo, Xiaoli Zhang 0001 |
IET Comput. Vis. | 3 |
| 2022 | A measure for the evaluation of multi-focus image fusion at feature level
Yuncong Feng, Rui Guo 0008, Xuanjing Shen, Xiaoli Zhang 0001 |
Multim. Tools Appl. | 4 |
| 2022 | A Self-Supervised Residual Feature Learning Model for Multifocus Image FusionabstractMulti-focus image fusion (MFIF) attempts to achieve an "all-focused" image from multiple source images with the same scene but different focused objects. Given the lack of multi-focus image sets for network training, we propose a self-supervised residual feature learning model in this paper. The model consists of a feature extraction network and a fusion module. We select image super-resolution as a pretext task in the MFIF field, which is supported by a new residual gradient prior discovered by our theoretical study for low- and high-resolution (LR-HR) image pairs, as well as for multi-focus images. In the pretext task, our network's training set is LR-HR image pairs generated from natural images, and HR images can be regarded as pseudo-labels of LR images. In the fusion task, the trained network extracts residual features of multi-focus images firstly. Secondly, the fusion module, consisting of an activity level measurement and a new boundary refinement method, is leveraged for the features to generated decision maps. Experimental results, both subjective evaluations and objective evaluations, demonstrate that our approach outperforms other state-of-the-art fusion algorithms. Zeyu Wang 0009, Haoran Duan 0001, Xiaoli Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | HID: The Hybrid Image Decomposition Model for MRI and CT FusionabstractMultimodal medical image fusion can combine salient information from different source images of the same part and reduce the redundancy of information. In this paper, an efficient hybrid image decomposition (HID) method is proposed. It combines the advantages of spatial domain and transform domain methods and breaks through the limitations of the algorithms based on single category features. The accurate separation of base layer and texture details is conducive to the better effect of the fusion rules. First, the source anatomical images are decomposed into a series of high frequencies and a low frequency via nonsubsampled shearlet transform (NSST). Second, the low frequency is further decomposed using the designed optimization model based on structural similarity and structure tensor to get an energy texture layer and a base layer. Then, the modified choosing maximum (MCM) is designed to fuse base layers. The sum of modified Laplacian (SML) is used to fuse high frequencies and energy texture layers. Finally, the fused low frequency can be obtained by adding fused energy texture layer and base layer. And the fused image is reconstructed by the inverse NSST. The superiority of the proposed method is verified by amounts of experiments on 50 pairs of magnetic resonance imaging (MRI) images and computed tomography (CT) images and others, and compared with 12 state-of-the-art medical image fusion methods. It is demonstrated that the proposed hybrid decomposition model has a better ability to extract texture information than conventional ones. Rui Zhu 0015, Xiaoli Zhang 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | Medical image fusion based on convolutional neural networks and non-subsampled contourlet transform
Zeyu Wang 0009, Haoran Duan 0001, Yanchi Su, Xiaoli Zhang 0001, Xinjiang Guan |
Expert Syst. Appl. | 5 |
| 2021 | BIDI: A classification algorithm with instance difficulty invariance
Hancheng Wang, Xiaoli Zhang 0001, Shiping Chen 0001 |
Expert Syst. Appl. | 4 |
| 2021 | C_CART: An instance confidence-based decision tree algorithm for classificationabstractIn classification, a decision tree is a common model due to its simple structure and easy understanding. Most of decision tree algorithms assume all instances in a dataset have the same degree of confidence, so they use the same generation and pruning strategies for all training instances. In fact, the instances with greater degree of confidence are more useful than the ones with lower degree of confidence in the same dataset. Therefore, the instances should be treated discriminately according to their corresponding confidence degrees when training classifiers. In this paper, we investigate the impact and significance of degree of confidence of instances on the classification performance of decision tree algorithms, taking the classification and regression tree (CART) algorithm as an example. First, the degree of confidence of instances is quantified from a statistical perspective. Then, a developed CART algorithm named C_CART is proposed by introducing the confidence of instances into the generation and pruning processes of CART algorithm. Finally, we conduct experiments to evaluate the performance of C_CART algorithm. The experimental results show that our C_CART algorithm can significantly improve the generalization performance as well as avoiding the over-fitting problem to a certain extend. Hancheng Wang, Xiaoli Zhang 0001, Shiping Chen 0001 |
Intell. Data Anal. | 4 |
| 2021 | An instance-oriented performance measure for classification
Yuncong Feng, Xiaoli Zhang 0001, Shiping Chen 0001 |
Inf. Sci. | 4 |
| 2021 | Multi-focus image fusion based on L1 image transform
Xiaoli Zhang 0001, Shiping Chen 0001 |
Multim. Tools Appl. | 4 |
| 2021 | Medical image fusion algorithm based on L0 gradient minimization for CT and MRI
Rui Zhu 0015, Xiaoli Zhang 0001, Zeyu Wang 0009 |
Multim. Tools Appl. | 4 |
| 2021 | MRI enhancement based on visual-attention by adaptive contrast adjustment and image fusion
Rui Zhu 0015, Xiaoli Zhang 0001, Xiaowei Xu 0001 |
Multim. Tools Appl. | 3 |
| 2021 | Correction to: MRI enhancement based on visual-attention by adaptive contrast adjustment and image fusion
Rui Zhu 0015, Xiaoli Zhang 0001, Xiaowei Xu 0001 |
Multim. Tools Appl. | 3 |
| 2021 | A multi-focus image fusion framework based on multi-scale sparse representation in gradient domain
Yu Wang 0164, Rui Zhu 0015, Zeyu Wang 0009, Yuncong Feng, Xiaoli Zhang 0001 |
Signal Process. | 6 |
| 2020 | FabricGene: A Higher-Level Feature Representation of Fabric Patterns for Nationality Classification
Hancheng Wang, Xiaoli Zhang 0001, Shiping Chen 0001 |
ADMA | 4 |
| 2020 | Multi-focus image fusion based on fully convolutional networksabstractWe propose a multi-focus image fusion method, in which a fully convolutional network for focus detection (FD-FCN) is constructed. To obtain more precise focus detection maps, we propose to add skip layers in the network to make both detailed and abstract visual information available when using FD-FCN to generate maps. A new training dataset for the proposed network is constructed based on dataset CIFAR-10. The image fusion algorithm using FD-FCN contains three steps: focus maps are obtained using FD-FCN, decision map generation occurs by applying a morphological process on the focus maps, and image fusion occurs using a decision map. We carry out several sets of experiments, and both subjective and objective assessments demonstrate the superiority of the proposed fusion method to state-of-the-art algorithms. Rui Guo 0008, Xuanjing Shen, Xiao-yu Dong, Xiaoli Zhang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2019 | A fusion algorithm for medical structural and functional images based on adaptive image decomposition
Xuanjing Shen, Haipeng Chen 0002, Yingda Lv, Xiaoli Zhang 0001 |
Multim. Tools Appl. | 5 |
| 2019 | Multifocus image fusion using convolutional neural networks in the discrete wavelet transform domain
Zeyu Wang 0009, Haoran Duan 0001, Xiaoli Zhang 0001, Hancheng Wang |
Multim. Tools Appl. | 4 |
| 2019 | FusionCNN: a remote sensing image fusion algorithm based on deep convolutional neural networks
Fajie Ye, Xiaoli Zhang 0001 |
Multim. Tools Appl. | 3 |
| 2018 | Adaptive decomposition method for multi-modal medical image fusionabstractIn traditional image fusion, source images are separated into a fixed space. The low‐frequency part and the high‐frequency part are not discriminated according to the nature of the image. Traditional fusion rules often use a fixed proportion, causing colour distortion. In this study, a new adaptive decomposition algorithm is proposed to distinguish high frequency and low frequency of structure image to obtain smoothing layer and texture layer. The smoothing layer of the structural image and the colour information of the function image are fused according to dynamic rules, and then the texture layer is added. On the basis of the objective evaluation metrics, the spectral information evaluation metrics are introduced to evaluate the retention of colour. In the experiments, the proposed method is compared with other six classical image fusion methods. The experiment results show that the proposed method can retain the colour information and structure information very well at the same time. Concerning subjective and objective evaluation, the proposed algorithm is superior to other algorithms. Xiaoli Zhang 0001 |
IET Image Process. | 4 |
| 2017 | Multi-focus image fusion algorithm based on multilevel morphological component analysis and support vector machineabstractIn this study, a novel algorithm is proposed for multi‐focus image fusion based on multilevel morphological decomposition and classifier. The attractive feature of the algorithm is that it decomposes images into several layers with different morphological components, which makes it preserve more detail information of source images. In the algorithm, source images are first decomposed by the multilevel morphological component analysis. Then, feature vectors are extracted from nature layers, and they are classified by a trained two‐class support vector machine. Then, consistency verification is employed to verify the decision matrix sets. Finally, coefficients are fused based on the decision matrix sets. Experimental results demonstrate the superiority of the proposed method in terms of subjective and objective evaluation. Xiaoli Zhang 0001 |
IET Image Process. | 4 |
| 2017 | Segmentation fusion based on neighboring information for MR brain images
Yuncong Feng, Xuanjing Shen, Haipeng Chen 0002, Xiaoli Zhang 0001 |
Multim. Tools Appl. | 4 |
| 2017 | Image fusion based on simultaneous empirical wavelet transform
Xiaoli Zhang 0001, Yuncong Feng |
Multim. Tools Appl. | 1 |
| 2016 | A semi-automatic brain tumor segmentation algorithmabstractIn this paper, a novel semi-automatic segmentation algorithm is proposed to segment brain tumors from magnetic resonance imaging (MRI) images. First, an edge-aware filter is used to get the smoothed version of the original image. Secondly, Otsu based multilevel thresholding is performed on the smoothed image and the original image, respectively. Then the two segmentation maps are fused by the rule of K Nearest Neighbors (KNN) to obtain the refined segmentation result. The combination of the three steps can be denoted as multi-scale Otsu based segmentation. Finally, a bi-directional region growing method is employed to segment the brain tumor region around seeds which are inserted by the user. The proposed algorithm is tested on MRI-T2 images and it produces promising result: the segmented tumor regions are more accurate compared to those obtained by other state-of-the-art methods. Xiaoli Zhang 0001, Yuncong Feng |
ICME | 1 |
| 2016 | A classification performance measure considering the degree of classification difficulty
Xiaoli Zhang 0001, Yuncong Feng |
Neurocomputing | 1 |
| 2016 | A weighted-ROC graph based metric for image segmentation evaluation
Yuncong Feng, Xuanjing Shen, Haipeng Chen 0002, Xiaoli Zhang 0001 |
Signal Process. | 4 |
| 2016 | A new multifocus image fusion based on spectrum comparison
Xiaoli Zhang 0001, Yuncong Feng |
Signal Process. | 1 |
| 2015 | Image fusion with Internal Generative Mechanism
Xiaoli Zhang 0001, Yuncong Feng, Zhaojun Liu |
Expert Syst. Appl. | 1 |
| 2015 | The use of ROC and AUC in the validation of objective image fusion evaluation metrics
Xiaoli Zhang 0001, Yuncong Feng, Zhaojun Liu |
Signal Process. | 1 |
| 2014 | Multi-focus image fusion using image-partition-based focus detection
Xiaoli Zhang 0001, Zhaojun Liu, Yuncong Feng |
Signal Process. | 1 |