Sifan Song

dblp:235/0567 · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-7940-650XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Beyond Adapter Retrieval: Latent Geometry-Preserving Composition via Sparse Task Projection
abstract
Recent advances in parameter-efficient transfer learning have demonstrated the utility of composing LoRA adapters from libraries of pretrained modules. However, most existing approaches rely on simple retrieval heuristics or uniform averaging, which overlook the latent structure of task relationships in representation space. We propose a new framework for adapter reuse that moves beyond retrieval, formulating adapter composition as a geometry-aware sparse reconstruction problem. Specifically, we represent each task by a latent prototype vector derived from the base model’s encoder and aim to approximate the target task prototype as a sparse linear combination of retrieved reference prototypes, under an L1-regularized optimization objective. The resulting combination weights are then used to blend the corresponding LoRA adapters, yielding a composite adapter tailored to the target task. This formulation not only preserves the local geometric structure of the task representation manifold, but also promotes interpretability and efficient reuse by selecting a minimal set of relevant adapters. We demonstrate the effectiveness of our approach across multiple domains—including medical image segmentation, medical report generation and image synthesis. Our results highlight the benefit of coupling retrieval with latent geometry-aware optimization for improved zero-shot generalization.
Pengfei Jin, Peng Shu, Sifan Song, Sekeun Kim, Qing Xiao 0003, Cheng Chen 0013, Tianming Liu 0001, Xiang Li 0001, Quanzheng Li
AAAI3
2025 ECHOPulse: ECG Controlled Echocardio-gram Video Generation
abstract
Echocardiography (ECHO) is essential for cardiac assessments, but its video quality and interpretation heavily relies on manual expertise, leading to inconsistent results from clinical and portable devices. ECHO video generation offers a solution by improving automated monitoring through synthetic data and generating high-quality videos from routine health data. However, existing models often face high computational costs, slow inference, and rely on complex conditional prompts that require experts' annotations. To address these challenges, we propose ECHOPulse, an ECG-conditioned ECHO video generation model. ECHOPulse introduces two key advancements: (1) it accelerates ECHO video generation by leveraging VQ-VAE tokenization and masked visual token modeling for fast decoding, and (2) it conditions on readily accessible ECG signals, which are highly coherent with ECHO videos, bypassing complex conditional prompts. To the best of our knowledge, this is the first work to use time-series prompts like ECG signals for ECHO video generation. ECHOPulse not only enables controllable synthetic ECHO data generation but also provides updated cardiac function information for disease monitoring and prediction beyond ECG alone. Evaluations on three public and private datasets demonstrate state-of-the-art performance in ECHO video generation across both qualitative and quantitative measures. Additionally, ECHOPulse can be easily generalized to other modality generation tasks, such as cardiac MRI, fMRI, and 3D CT generation. We will make the synthetic ECHO dataset, along with the code and model, publicly available upon acceptance.
Yiwei Li 0002, Sekeun Kim, Zihao Wu 0001, Hanqi Jiang, Yi Pan 0001, Pengfei Jin, Sifan Song, Xiaowei Yu 0001, Tianze Yang, Tianming Liu 0001, Quanzheng Li, Xiang Li 0001
ICLR7
2025 MAST-Pro: Dynamic Mixture-of-Experts for Adaptive Segmentation of Pan-Tumors with Knowledge-Driven Prompts
Runqi Meng, Sifan Song, Pengfei Jin, Yiqun Sun, Yujin Oh, Xiang Li 0001, Quanzheng Li, Dinggang Shen
MICCAI (16)2
2025 A Novel Fourier Adjacency Transformer for Advanced EEG Emotion Recognition
Jinfeng Wang 0008, Yanhao Huang, Sifan Song, Boqian Wang, Jionglong Su, Jiaman Ding
MICCAI (12)3
2025 SAMed-2: Selective Memory Enhanced Medical Segment Anything Model
Zhiling Yan, Sifan Song, Dingjie Song, Yiwei Li 0002, Rong Zhou 0007, Weixiang Sun, Zhennong Chen, Sekeun Kim, Hui Ren 0001, Tianming Liu 0001, Quanzheng Li, Xiang Li 0001, Lifang He 0001, Lichao Sun 0001
MICCAI (13)2
2025 Cascaded 3D Diffusion Models for Whole-Body 3D 18-F FDG PET/CT Synthesis from Demographics
Siyeop Yoon, Sifan Song, Pengfei Jin, Matthew Tivnan, Yujin Oh, Sekeun Kim, Dufan Wu, Xiang Li 0001, Quanzheng Li
MICCAI (3)2
2025 Implicit Image-to-Image Schrödinger Bridge for image restoration
Yuang Wang, Siyeop Yoon, Pengfei Jin, Matthew Tivnan, Sifan Song, Zhennong Chen, Li Zhang 0047, Quanzheng Li, Zhiqiang Chen 0001, Dufan Wu
Pattern Recognit.5
2025 EchoFM: Foundation Model for Generalizable Echocardiogram Analysis
abstract
Echocardiography is the first-line non-invasive cardiac imaging modality, providing rich spatio-temporal information on cardiac anatomy and physiology. Recently, foundation model trained on extensive and diverse datasets has shown strong performance in various downstream tasks. However, translating foundation models into the medical imaging domain remains challenging due to domain differences between medical and natural images, the lack of diverse patient and disease datasets. In this paper, we introduce EchoFM, a general-purpose vision foundation model for echocardiography trained on a large-scale dataset of over 20 million echocardiographic images from 6,500 patients. To enable effective learning of rich spatio-temporal representations from periodic videos, we propose a novel self-supervised learning framework based on a masked autoencoder with a spatio-temporal consistent masking strategy and periodic-driven contrastive learning. The learned cardiac representations can be readily adapted and fine-tuned for a wide range of downstream tasks, serving as a strong and flexible backbone model. We validate EchoFM through experiments across key downstream tasks in the clinical echocardiography workflow, leveraging public and multi-center internal datasets. EchoFM consistently outperforms SOTA methods, demonstrating superior generalization capabilities and flexibility. The code and checkpoints are available at: https://github.com/SekeunKim/EchoFM.git.
Sekeun Kim, Pengfei Jin, Sifan Song, Cheng Chen 0013, Yiwei Li 0002, Hui Ren 0001, Xiang Li 0001, Tianming Liu 0001, Quanzheng Li
IEEE Trans. Medical Imaging3
2024 Distortion-Disentangled Contrastive Learning
abstract
Self-supervised learning is well known for its remarkable performance in representation learning and various downstream computer vision tasks. Recently, Positive-pair-Only Contrastive Learning (POCL) has achieved reliable performance without the need to construct positive-negative training sets. It reduces memory requirements by lessening the dependency on the batch size. The POCL method typically uses a single objective function to extract the distortion invariant representation (DIR) which describes the proximity of positive-pair representations affected by different distortions. This objective function implicitly enables the model to filter out or ignore the distortion variant representation (DVR) affected by different distortions. However, some recent studies have shown that proper use of DVR in contrastive can optimize the performance of models in some downstream domain-specific tasks. In addition, these POCL methods have been observed to be sensitive to augmentation strategies. To address these limitations, we propose a novel POCL framework named Distortion-Disentangled Contrastive Learning (DDCL) and a Distortion-Disentangled Loss (DDL). Our approach is the first to explicitly and adaptively disentangle and exploit the DVR inside the model and feature stream to improve the representation utilization efficiency, robustness and representation ability. Experiments demonstrate our framework’s superiority to Barlow Twins and Simsiam in terms of convergence, representation quality (including transferability and generalization), and robustness on several datasets.
Jinfeng Wang 0008, Sifan Song, Jionglong Su, Shaohua Kevin Zhou
WACV2
2024 DualStreamFoveaNet: A Dual Stream Fusion Architecture With Anatomical Awareness for Robust Fovea Localization
abstract
Accurate fovea localization is essential for analyzing retinal diseases to prevent irreversible vision loss. While current deep learning-based methods outperform traditional ones, they still face challenges such as the lack of local anatomical landmarks around the fovea, the inability to robustly handle diseased retinal images, and the variations in image conditions. In this paper, we propose a novel transformer-based architecture called DualStreamFoveaNet (DSFN) for multi-cue fusion. This architecture explicitly incorporates long-range connections and global features using retina and vessel distributions for robust fovea localization. We introduce a spatial attention mechanism in the dual-stream encoder to extract and fuse self-learned anatomical information, focusing more on features distributed along blood vessels and significantly reducing computational costs by decreasing token numbers. Our extensive experiments show that the proposed architecture achieves state-of-the-art performance on two public datasets and one large-scale private dataset. Furthermore, we demonstrate that the DSFN is more robust on both normal and diseased retina images and has better generalization capacity in cross-dataset experiments.
Sifan Song, Jinfeng Wang 0008, Zilong Wang 0006, Hongxing Wang 0001, Jionglong Su, Kang Dang
IEEE J. Biomed. Health Informatics1
2023 From deterministic to stochastic: an interpretable stochastic model-free reinforcement learning framework for portfolio optimization
Zitao Song, Pin Qian, Sifan Song, Frans Coenen, Zhengyong Jiang, Jionglong Su
Appl. Intell.4
2023 GAMMA challenge: Glaucoma grAding from Multi-Modality imAges
Huihui Fang, Fei Li 0021, Huazhu Fu, Fengbin Lin, Jiongcheng Li, Yue Huang 0001, Qinji Yu, Sifan Song, Xinxing Xu, Yanyu Xu 0001, Wensai Wang, Shuai Lu 0003, Huiqi Li, Shihua Huang, Zhichao Lu, Chubin Ou, Xifei Wei, Bingyuan Liu, Riadh Kobbi, Xiaoying Tang 0001, Li Lin 0006, Hrvoje Bogunovic, José Ignacio Orlando, Xiulan Zhang, Yanwu Xu 0001
Medical Image Anal.9
2022 Stepwise Feature Fusion: Local Guides Global
Jinfeng Wang 0008, Qiming Huang, Jia Meng 0001, Jionglong Su, Sifan Song
MICCAI (3)6
2022 A New Convolutional Neural Network Architecture for Automatic Segmentation of Overlapping Human Chromosomes
Sifan Song, Tianming Bai, Yanxin Zhao, Wenbo Zhang 0010, Chunxiao Yang, Jia Meng 0001, Fei Ma 0002, Jionglong Su
Neural Process. Lett.1
2021 A Framework of Hierarchical Deep Q-Network for Portfolio Management
Ziming Gao, Sifan Song, Zhengyong Jiang, Jionglong Su
ICAART (2)4