EDBT 2026 Demo / reviewers in the wild / expert
Zeyu Wang 0009
dblp:132/7882-9
· DBLP profile ↗
24ranked-venue papers
11as first author
22since 2021 · last 2026
0000-0002-1223-1661ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 8 first-author · 14 since 2021Artificial intelligence and machine learning · 9 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SigFusion: Unified Signal-Level Self-Supervised Learning Paradigm for Image FusionabstractImage Fusion (IF) aims to integrate complementary features from multiple source images into a single image. However, a key challenge in this field is the lack of large-scale real-world training datasets. Existing models typically rely on either small datasets or synthetic, less realistic datasets. To address this, we propose SigFusion, a unified signal-level self-supervised learning paradigm for various IF tasks.The core idea is to use signal-level Pseudo-Label Generation Networks (PLGN) to automatically synthesize training sets and pseudo labels with real multi-source signal characteristics from vast unlabeled natural images.PLGN includes two critical components: learnable 1D Signal Modulators (SM) and SigFormer. SM learns implicit 1D signal patterns across various source images and embeds them into natural images, reducing the domain gap between synthetic and real datasets. SigFormer integrates Transformer with signal processing methods, establishing an appropriate signal representation space for SM. Its cascaded, multi-level design allows hierarchical feature learning from coarse to fine detail. Moreover, SigFormer can serve as a flexible backbone for IF, as its design adheres to the classic decomposition-reconstruction paradigm. Experimental results demonstrate that SigFusion achieves state-of-the-art performance across multiple IF tasks, including medical image fusion, infrared-visible image fusion, multi-focus image fusion, and multi-exposure image fusion. Zeyu Wang 0009, Pengjie Wang 0001, Haiyu Song 0002 |
AAAI | 1 |
| 2026 | Breaking Task Boundaries: A Unified Model for 3D Medical Image Fusion and Segmentation Guided by Manifold Perspectiveabstract3D medical image fusion (MIF) and segmentation (MIS) are critical and inherently synergistic tasks in medical image analysis. However, fundamentally integrating them remains highly challenging, since effective collaborative paradigms are still scarce and their optimization objectives fundamentally diverge. Moreover, existing continual learning techniques are unable to achieve truly advanced performance for both tasks using a shared weight. To address these challenges, we propose M²-CoFS, a unified model capable of jointly handling both tasks. Our core contribution is a “network-guided network learning” paradigm designed to break the task boundaries. We model the weight spaces of MIF and MIS as high-dimensional manifolds and innovatively use a lightweight neural network to implicitly construct a shared manifold. Interestingly, this network yields a unified weight for both tasks. To ensure the shared manifold retains the intrinsic geometry of both original manifolds, we embed manifold distances into the loss function of this network as a constraint. Additionally, we design a tailored three-stage training paradigm for our core contribution mentioned above. Stage I focuses on independent task optimization for high-quality weights; Stage II aims to reduce parameter-space distance between tasks via our cross-task weight adaptation strategy; Our core innovation serves as stage III. Experimental results show that M²-CoFS consistently outperforms state-of-the-art comparison models on both MlF and MIS. Zeyu Wang 0009, Haiyu Song 0002 |
AAAI | 1 |
| 2026 | ControlFuse: Instruction-guided Multi-Granularity Controllable Image FusionabstractInfrared and Visible Image Fusion (IVIF) produces enhanced images by fusing complementary visual information. However, most existing methods generate fixed outputs and cannot flexibly adapt to user-specific requirements. Recent text-guided approaches offer partial control but are limited to global or semantic levels, lacking instance-level control. This limitation arises from two challenges: first, the lack of datasets that directly link textual instructions with corresponding spatial annotations, and second, the use of coarse cross-modal alignment methods that struggle to precisely match textual instructions with visual features. To overcome these challenges, we propose ControlFuse, a controllable IVIF framework enabling multi-granularity fusion across global, semantic, and instance levels, guided by user instructions. First, we construct an automated multi-granularity dataset that provides explicit textual-mask correspondences at these three levels. Second, inspired by manifold geometry, we design a Multimodal Feature Interaction Module (MFIM) comprising Feature Manifold Converter (FMC) and Curvature-Guided Interaction (CGI). FMC projects textual and visual features into a unified manifold space, while CGI leverages manifold curvature as a geometric cue to refine cross-modal alignment. Extensive experiments validate ControlFuse, outperforming state-of-the-art methods in robustness and flexibility. Xiaoli Zhang 0001, Zeyu Wang 0009 |
AAAI | 3 |
| 2026 | Infrared and visible image fusion via iterative feature decomposition and deep balanced fusion
Wei Li 0150, Baojia Li 0003, Haiyu Song 0002, Pengjie Wang 0001, Zeyu Wang 0009 |
Pattern Recognit. | 5 |
| 2026 | Rethinking Multi-Focus Image Fusion: An Input Space Optimization ViewabstractMulti-focus image fusion (MFIF) addresses the challenge of partial focus by integrating multiple source images taken at different focal depths. Unlike most existing methods that rely on complex loss functions or large-scale synthetic datasets, this study approaches MFIF from a novel perspective: optimizing the input space. The core idea is to construct a high-quality MFIF input space in a cost-effective manner by using intermediate features from well-trained, non-MFIF networks. To this end, we propose a cascaded framework comprising two feature extractors, a Feature Distillation and Fusion Module (FDFM), and a focus segmentation network Y ${}^{U}$ Net. Based on our observation that discrepancy and edge features are essential for MFIF, we select a image deblurring network and a salient object detection network as feature extractors. To transform these extracted features into an MFIF-suitable input space, we propose FDFM as a training-free feature adapter. To make FDFM compatible with high-dimensional feature maps, we extend the manifold theory from the edge-preserving field and design a novel isometric domain transformation. Extensive experiments on six benchmark datasets show that 1) our model consistently outperforms 13 state-of-the-art methods in both qualitative and quantitative evaluations, and 2) the constructed input space can directly enhance the performance of many MFIF models without additional requirements. Zeyu Wang 0009, Haoran Duan 0001, Yang Long 0001, Ling Shao 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Seg4Fusion: A Tumor-Aware Framework for 3D Medical Image Fusionabstract3D medical image fusion integrates complementary information from multiple imaging modalities such as MRI, CT, and PET to generate fused images with enhanced diagnostic value and structural representation. This technique is widely applied in lesion identification, preoperative planning, and clinical decision support. Fusion methods can be broadly categorized into visuallydriven and task-driven approaches, both aiming to preserve detailed and accurate information of critical lesion regions like tumors. Existing mainstream methods primarily focus on global structural features and modality complementarity but lack mechanisms for local feature enhancement of tumor regions, resulting in insufficient preservation of crucial lesion details. Moreover, non-lesion areas are processed uniformly, leading to inadequate suppression of background noise and redundant information. The scarcity of precise tumor annotations in training data further limits models' ability to focus on key lesion regions due to the lack of effective supervision. To address these challenges, this work proposes the Seg4Fusion framework. It introduces a tumoraware training strategy by incorporating segmentation-derived priors to Extract features into lesion and background regions. A dual-branch fusion architecture based on feature enhancement and sparse representation is designed to perform targeted fusion for different regions, strengthening lesion information while suppressing background interference. Additionally, a tumorregion supervised loss is developed to guide the model's attention toward critical areas, improving the discriminability and quality of fused images. Experimental results demonstrate that the proposed method effectively enhances tumor detail preservation and diagnostic information fusion, showing strong potential for clinical applications. Haiyu Song 0002, Maoyu Wang, Aohua Ma, Zeyu Wang 0009 |
BIBM | 6 |
| 2025 | Multi-Modal Medical Image Fusion via 3D Manifold Fitting and Dual-Domain Cross-AttentionabstractMedical image fusion (MIF) aims to extract complementary features from multi-modal source images and fuse them into a single image to assist in clinical diagnostics. Despite its importance, MIF faces two primary challenges: the lack of tailored paradigms for CMSF extraction and insufficient dual exploration of multi-modality and multi-frequency domains. To address these challenges, we propose a novel MIF model in this study. From the perspective of image manifolds, we reformulate CMSF extraction as a 3D manifold fitting problem and introduce a paradigm that uses mathematical fitting methods to generate CMSF. This approach achieves accurate feature extraction without the need for carefully designed loss functions as constraints, significantly reducing the number of parameters. Additionally, we introduce Cross-Modality Co-Frequency (CM-CoF) and Cross-Frequency Co-Modality (CF-CoM) attention modules, which explore implicit relationships between modalities and frequency domains. Experimental results demonstrate that the proposed model outperforms many state-of-the-art MIF algorithms. Zeyu Wang 0009, Haiyu Song 0002, Haoran Duan 0001 |
ICASSP | 1 |
| 2025 | MFANet: Multi-Feature Aggregation Network for Multi-focus Image FusionabstractExisting deep learning-based Multi-focus Image Fusion (MFIF) methods often rely on loss functions derived from linear combinations of image quality metrics, leading to complexities in training and only marginal improvements in image quality. Recognizing this, our study identifies input space and scale information as pivotal in enhancing MFIF performance. By augmenting raw spatial images with other feature spaces, i.e., gradient and dense Scale-Invariant Feature Transform (DSIFT), we enhance the model’s ability to detect edges, textures, and local structures, facilitating more accurate differentiation between focused and defocused areas. Additionally, smaller scale variations improve focus detection, while multi-scale learning within neural networks effectively suppresses artifacts without affecting focus detection accuracy. To achieve the above enhancements, we introduce the Multi-Feature Aggregation Network (MFANet), which employs a three-branch architecture to perform focused detection process in spatial, gradient, and DSIFT feature spaces. Each branch is equipped with a Pyramid Attention Fusion (PAF) module that utilizes attention mechanisms and a novel Light Spatial Aggregation Pyramid Module (LSAPM) to capture global feature relationships and aggregate multi-scale information. Experimental results demonstrate that MFANet surpasses other state-of-the-art fusion methods in both qualitative and quantitative evaluations. Xiaoli Zhang 0001, Mingjie Tian, Zeyu Wang 0009 |
ICASSP | 5 |
| 2025 | Highlight What You Want: Weakly-Supervised Instance-Level Controllable Infrared-Visible Image Fusion
Zeyu Wang 0009, Jizheng Zhang, Haiyu Song 0002, Mingyu Ge, Haoran Duan 0001 |
ICCV | 1 |
| 2025 | TRACE: Temporally Reliable Anatomically-Conditioned 3D CT Generation with Enhanced Efficiency
Minye Shao, Xingyu Miao, Haoran Duan 0001, Zeyu Wang 0009, Jingkun Chen, Yawen Huang, Xian Wu 0001, Jingjing Deng 0001, Yang Long 0001, Yefeng Zheng 0001 |
MICCAI (4) | 4 |
| 2025 | KMMF-Net: Implicit Fusion with KAN-Guided Mamba Modeling
Haiyu Song 0002, Yun Mao, Mingyu Ge, Zhengchi Du, Zeyu Wang 0009 |
PRCV (5) | 8 |
| 2025 | Multi-Text Guidance Is Important: Multi-Modality Image Fusion via Large Generative Vision-Language Model
Zeyu Wang 0009, Jizheng Zhang, Haiyu Song 0002, Jiana Meng |
Int. J. Comput. Vis. | 1 |
| 2025 | Focusing on neglected natural images: A self-supervised learning model for pan-sharpening
Xiaoli Zhang 0001, Zeyu Wang 0009 |
Inf. Process. Manag. | 3 |
| 2025 | Rethinking Brain Tumor Segmentation From the Frequency Domain PerspectiveabstractPrecise segmentation of brain tumors, particularly contrast-enhancing regions visible in post-contrast MRI (areas highlighted by contrast agent injection), is crucial for accurate clinical diagnosis and treatment planning but remains challenging. However, current methods exhibit notable performance degradation in segmenting these enhancing brain tumor areas, largely due to insufficient consideration of MRI-specific tumor features such as complex textures and directional variations. To address this, we propose the Harmonized Frequency Fusion Network (HFF-Net), which rethinks brain tumor segmentation from a frequency-domain perspective. To comprehensively characterize tumor regions, we develop a Frequency Domain Decomposition (FDD) module that separates MRI images into low-frequency components, capturing smooth tumor contours and high-frequency components, highlighting detailed textures and directional edges. To further enhance sensitivity to tumor boundaries, we introduce an Adaptive Laplacian Convolution (ALC) module that adaptively emphasizes critical high-frequency details using dynamically updated convolution kernels. To effectively fuse tumor features across multiple scales, we design a Frequency Domain Cross-Attention (FDCA) integrating semantic, positional, and slice-specific information. We further validate and interpret frequency-domain improvements through visualization, theoretical reasoning, and experimental analyses. Extensive experiments on four public datasets demonstrate that HFF-Net achieves an average relative improvement of 4.48% (ranging from 2.39% to 7.72%) in the mean Dice scores across the three major subregions, and an average relative improvement of 7.33% (ranging from 5.96% to 8.64%) in the segmentation of contrast-enhancing tumor regions, while maintaining favorable computational efficiency and clinical applicability. Our code is available at: https://github.com/VinyehShaw/HFF. Minye Shao, Zeyu Wang 0009, Haoran Duan 0001, Yawen Huang, Bing Zhai, Shizheng Wang, Yang Long 0001, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Detecting sparse building change with ambiguous label using Siamese full-scale connected network and instance augmentation
Xinze Lin, Zeyu Wang 0009, Xiaoli Zhang 0001 |
Appl. Intell. | 3 |
| 2023 | When Multi-Focus Image Fusion Networks Meet Traditional Edge-Preservation Technology
Zeyu Wang 0009, Haoran Duan 0001, Xiaoli Zhang 0001 |
Int. J. Comput. Vis. | 1 |
| 2023 | ERMF: Edge refinement multi-feature for change detection in bitemporal remote sensing images
Zixuan Song, Rui Zhu 0015, Zeyu Wang 0009, Xiaoli Zhang 0001 |
Signal Process. Image Commun. | 4 |
| 2023 | VSP-Fuse: Multifocus Image Fusion Model Using the Knowledge Transferred From Visual Salience PriorsabstractMultifocus image fusion (MFIF), as an efficient way to improve the visual effect of images with partial focus defects, is of great significance in the field of image enhancement. According to the imaging principle of the lens, we summarize the visual salience priors (VSP) from the daily photo scene and two relationships from MFIF. Thereby, an edge-sensitive model for MFIF is presented in this study. Supported by VSP, we consider the correlation between salience object detection (SOD) and MFIF, and select the former as a pre-training task. SOD provides the network with realistic depth of field and bokeh effects to learn, and enhances the network’s ability to extract and express the edges of focused objects. Meanwhile, given the scarcity of real multifocus training sets, we propose a randomized approach to generate massive training sets and pseudo-labels based on limited unlabeled data. Besides, two attention modules are designed based on isometric domain transformation (IDT) in the traditional edge-preservation field. IDT removes interference information from feature maps in a low-cost manner, thereby facilitating channel-wise and spatial-wise weight assignments. Experimental results on four datasets show that the performance of our model is superior to that of many supervised models, without the need of any real MFIF training set. Zeyu Wang 0009, Haoran Duan 0001, Xiaoli Zhang 0001, Jizheng Zhang, Shiping Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | A Self-Supervised Residual Feature Learning Model for Multifocus Image FusionabstractMulti-focus image fusion (MFIF) attempts to achieve an "all-focused" image from multiple source images with the same scene but different focused objects. Given the lack of multi-focus image sets for network training, we propose a self-supervised residual feature learning model in this paper. The model consists of a feature extraction network and a fusion module. We select image super-resolution as a pretext task in the MFIF field, which is supported by a new residual gradient prior discovered by our theoretical study for low- and high-resolution (LR-HR) image pairs, as well as for multi-focus images. In the pretext task, our network's training set is LR-HR image pairs generated from natural images, and HR images can be regarded as pseudo-labels of LR images. In the fusion task, the trained network extracts residual features of multi-focus images firstly. Secondly, the fusion module, consisting of an activity level measurement and a new boundary refinement method, is leveraged for the features to generated decision maps. Experimental results, both subjective evaluations and objective evaluations, demonstrate that our approach outperforms other state-of-the-art fusion algorithms. Zeyu Wang 0009, Haoran Duan 0001, Xiaoli Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Medical image fusion based on convolutional neural networks and non-subsampled contourlet transform
Zeyu Wang 0009, Haoran Duan 0001, Yanchi Su, Xiaoli Zhang 0001, Xinjiang Guan |
Expert Syst. Appl. | 1 |
| 2021 | Medical image fusion algorithm based on L0 gradient minimization for CT and MRI
Rui Zhu 0015, Xiaoli Zhang 0001, Zeyu Wang 0009 |
Multim. Tools Appl. | 5 |
| 2021 | A multi-focus image fusion framework based on multi-scale sparse representation in gradient domain
Yu Wang 0164, Rui Zhu 0015, Zeyu Wang 0009, Yuncong Feng, Xiaoli Zhang 0001 |
Signal Process. | 4 |
| 2020 | Multi-band weighted lp norm minimization for image denoising
Yanchi Su, Zhanshan Li, Haihong Yu, Zeyu Wang 0009 |
Inf. Sci. | 4 |
| 2019 | Multifocus image fusion using convolutional neural networks in the discrete wavelet transform domain
Zeyu Wang 0009, Haoran Duan 0001, Xiaoli Zhang 0001, Hancheng Wang |
Multim. Tools Appl. | 1 |