Haiyu Song 0002

dblp:52/7814-2 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0002-2693-6291ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SigFusion: Unified Signal-Level Self-Supervised Learning Paradigm for Image Fusion
abstract
Image Fusion (IF) aims to integrate complementary features from multiple source images into a single image. However, a key challenge in this field is the lack of large-scale real-world training datasets. Existing models typically rely on either small datasets or synthetic, less realistic datasets. To address this, we propose SigFusion, a unified signal-level self-supervised learning paradigm for various IF tasks.The core idea is to use signal-level Pseudo-Label Generation Networks (PLGN) to automatically synthesize training sets and pseudo labels with real multi-source signal characteristics from vast unlabeled natural images.PLGN includes two critical components: learnable 1D Signal Modulators (SM) and SigFormer. SM learns implicit 1D signal patterns across various source images and embeds them into natural images, reducing the domain gap between synthetic and real datasets. SigFormer integrates Transformer with signal processing methods, establishing an appropriate signal representation space for SM. Its cascaded, multi-level design allows hierarchical feature learning from coarse to fine detail. Moreover, SigFormer can serve as a flexible backbone for IF, as its design adheres to the classic decomposition-reconstruction paradigm. Experimental results demonstrate that SigFusion achieves state-of-the-art performance across multiple IF tasks, including medical image fusion, infrared-visible image fusion, multi-focus image fusion, and multi-exposure image fusion.
Zeyu Wang 0009, Pengjie Wang 0001, Haiyu Song 0002
AAAI5
2026 Breaking Task Boundaries: A Unified Model for 3D Medical Image Fusion and Segmentation Guided by Manifold Perspective
abstract
3D medical image fusion (MIF) and segmentation (MIS) are critical and inherently synergistic tasks in medical image analysis. However, fundamentally integrating them remains highly challenging, since effective collaborative paradigms are still scarce and their optimization objectives fundamentally diverge. Moreover, existing continual learning techniques are unable to achieve truly advanced performance for both tasks using a shared weight. To address these challenges, we propose M²-CoFS, a unified model capable of jointly handling both tasks. Our core contribution is a “network-guided network learning” paradigm designed to break the task boundaries. We model the weight spaces of MIF and MIS as high-dimensional manifolds and innovatively use a lightweight neural network to implicitly construct a shared manifold. Interestingly, this network yields a unified weight for both tasks. To ensure the shared manifold retains the intrinsic geometry of both original manifolds, we embed manifold distances into the loss function of this network as a constraint. Additionally, we design a tailored three-stage training paradigm for our core contribution mentioned above. Stage I focuses on independent task optimization for high-quality weights; Stage II aims to reduce parameter-space distance between tasks via our cross-task weight adaptation strategy; Our core innovation serves as stage III. Experimental results show that M²-CoFS consistently outperforms state-of-the-art comparison models on both MlF and MIS.
Zeyu Wang 0009, Haiyu Song 0002
AAAI3
2026 Infrared and visible image fusion via iterative feature decomposition and deep balanced fusion
Wei Li 0150, Baojia Li 0003, Haiyu Song 0002, Pengjie Wang 0001, Zeyu Wang 0009
Pattern Recognit.3
2025 Seg4Fusion: A Tumor-Aware Framework for 3D Medical Image Fusion
abstract
3D medical image fusion integrates complementary information from multiple imaging modalities such as MRI, CT, and PET to generate fused images with enhanced diagnostic value and structural representation. This technique is widely applied in lesion identification, preoperative planning, and clinical decision support. Fusion methods can be broadly categorized into visuallydriven and task-driven approaches, both aiming to preserve detailed and accurate information of critical lesion regions like tumors. Existing mainstream methods primarily focus on global structural features and modality complementarity but lack mechanisms for local feature enhancement of tumor regions, resulting in insufficient preservation of crucial lesion details. Moreover, non-lesion areas are processed uniformly, leading to inadequate suppression of background noise and redundant information. The scarcity of precise tumor annotations in training data further limits models' ability to focus on key lesion regions due to the lack of effective supervision. To address these challenges, this work proposes the Seg4Fusion framework. It introduces a tumoraware training strategy by incorporating segmentation-derived priors to Extract features into lesion and background regions. A dual-branch fusion architecture based on feature enhancement and sparse representation is designed to perform targeted fusion for different regions, strengthening lesion information while suppressing background interference. Additionally, a tumorregion supervised loss is developed to guide the model's attention toward critical areas, improving the discriminability and quality of fused images. Experimental results demonstrate that the proposed method effectively enhances tumor detail preservation and diagnostic information fusion, showing strong potential for clinical applications.
Haiyu Song 0002, Maoyu Wang, Aohua Ma, Zeyu Wang 0009
BIBM1
2025 Multi-Modal Medical Image Fusion via 3D Manifold Fitting and Dual-Domain Cross-Attention
abstract
Medical image fusion (MIF) aims to extract complementary features from multi-modal source images and fuse them into a single image to assist in clinical diagnostics. Despite its importance, MIF faces two primary challenges: the lack of tailored paradigms for CMSF extraction and insufficient dual exploration of multi-modality and multi-frequency domains. To address these challenges, we propose a novel MIF model in this study. From the perspective of image manifolds, we reformulate CMSF extraction as a 3D manifold fitting problem and introduce a paradigm that uses mathematical fitting methods to generate CMSF. This approach achieves accurate feature extraction without the need for carefully designed loss functions as constraints, significantly reducing the number of parameters. Additionally, we introduce Cross-Modality Co-Frequency (CM-CoF) and Cross-Frequency Co-Modality (CF-CoM) attention modules, which explore implicit relationships between modalities and frequency domains. Experimental results demonstrate that the proposed model outperforms many state-of-the-art MIF algorithms.
Zeyu Wang 0009, Haiyu Song 0002, Haoran Duan 0001
ICASSP3
2025 Highlight What You Want: Weakly-Supervised Instance-Level Controllable Infrared-Visible Image Fusion
Zeyu Wang 0009, Jizheng Zhang, Haiyu Song 0002, Mingyu Ge, Haoran Duan 0001
ICCV3
2025 KMMF-Net: Implicit Fusion with KAN-Guided Mamba Modeling
Haiyu Song 0002, Yun Mao, Mingyu Ge, Zhengchi Du, Zeyu Wang 0009
PRCV (5)2
2025 Multi-Text Guidance Is Important: Multi-Modality Image Fusion via Large Generative Vision-Language Model
Zeyu Wang 0009, Jizheng Zhang, Haiyu Song 0002, Jiana Meng
Int. J. Comput. Vis.5
2024 OR2Net: Online Re-weighting Relation Network for kinship verification
abstract
Kinship verification aims to infer whether there is a kin relation between different individuals from facial images . However, popular kinship datasets are often small and suffer from data imbalance. Most existing methods build complex networks to extract features but ignore some implicit information, like family information. They use balanced datasets with fixed negative samples for training, which overlooks valuable information from multiple negative samples, leading to poor performance and robustness. To address these issues, we propose a novel end-to-end framework for kinship verification called Online Re-weighting Relation Network (OR 2 Net) based on an online re-weighting strategy of meta-learning and relation network. Our novel relation network aims to extract fine-grained features and reduce differences between generations by using multi-scale features and mining family information from kinship datasets. Additionally, we design a lightweight meta re-weighting network that uses a small, clean meta-set to guide the adaptive weighing of training examples. This is done using one-step stochastic gradient descent (SGD) based on an online re-weighting strategy from meta-learning. This helps find effective hard negative samples and reduces the imbalance problem. Extensive experiments on three public kinship verification datasets show that our proposed method is more effective compared to state-of-the-art methods. The code is publicly available on https://github.com/XinZhao-dlnu/OR2N .
Houjie Li, Mengyin Wang, Haiyu Song 0002, Fuming Sun
Expert Syst. Appl.4
2014 An Eigen-based motion retrieval method for real-time animation
Pengjie Wang 0001, Rynson W. H. Lau, Jiang Wang 0015, Haiyu Song 0002
Comput. Graph.5
2013 The alpha parallelogram predictor: A lossless compression method for motion capture data
Pengjie Wang 0001, Rynson W. H. Lau, Haiyu Song 0002
Inf. Sci.5
2011 A real-time database architecture for motion capture data
abstract
Due to the popularity of motion capture data in many applications, such as games, movies and virtual environments, huge collections of motion capture data are now available. It is becoming important to store these data in compressed form while being able to retrieve them without much overhead. However, there is little work that addresses both issues together. In this paper, we address these two issues by proposing a novel database architecture. First, we propose a lossless compression algorithm to compress the motion clips, which is based on a novel Alpha Parallelogram Predictor (APP) to estimate the degree of freedom (DOF) of each child joint from its immediate neighbors and parents that have already been processed. Second, we propose to store selected eigenvalues and eigenvectors of each motion clip, which only require a very small amount of memory overheads, for faster filtering of irrelevant motions. Based on this architecture, real-time queries become a three-step process. In the first two steps, we perform a quick filtering to identify relevant motion clips in the database through a two-level indexing structure. In the third step, only a small number of candidate clips are uncompressed and accurately matched with a Dynamic Time Warping algorithm. Our results show that users can efficiently search clips from this losslessly compressed motion database.
Pengjie Wang 0001, Rynson W. H. Lau, Jiang Wang 0015, Haiyu Song 0002
ACM Multimedia5