EDBT 2026 Demo / reviewers in the wild / expert
Kaicong Sun
dblp:232/8715
· DBLP profile ↗
31ranked-venue papers
4as first author
30since 2021 · last 2026
0000-0002-9999-2542ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 27 · 2 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 first-author · 14 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LMGDM: A Lesion-aware Mutual Guidance Diffusion Model with attenuation prior constraint for self-attenuation correction of whole-body PET
Shengjun Li, Kaicong Sun, Caiwen Jiang, Zaixin Ou, Ruilong Dan, Qianjin Feng 0003, Dinggang Shen |
Medical Image Anal. | 2 |
| 2026 | A hierarchical prompt and prototype learning framework for brain disorder classification
Kaicong Sun, Yaping Wu, Weilin Zhou, Haoyue Yuan, Xintong Wu, Yichu He, Qingxia Wu, Zeng-Yang Che, Yiqiang Zhan, Sean Zhou, Dijia Wu, Feng Shi 0001, Dinggang Shen |
Medical Image Anal. | 2 |
| 2026 | Learning dual-scale context with overlap awareness for keypoint-driven partial-overlap medical image registration
Caiwen Jiang, Xiaosong Xiong, Kaicong Sun, Xiaohuan Cao, Dinggang Shen |
Medical Image Anal. | 5 |
| 2026 | HALO: High-frequency enhanced dose-aware diffusion model for arbitrary low-dose PET reconstruction
Caiwen Jiang, Kaicong Sun, Zhiming Cui 0001, Dinggang Shen |
Medical Image Anal. | 3 |
| 2026 | Multi-organ guided diagnosis of mild cognitive impairment via hierarchical alignment and knowledge distillation
Shilun Zhao, Kaicong Sun, Shuwei Bai, Weilin Zhou, Jiangtao Liang, Zhongxiang Ding, Han Zhang 0002, Dinggang Shen |
Medical Image Anal. | 3 |
| 2026 | Uncertainty-Guided Iterative Contrastive Fusion for Reliable Survival Prediction in Rectal CancerabstractIntegrating multimodal radiological images and clinical data is critical for survival prediction in rectal cancer. However, existing methods often lack sufficient consideration of 1) modality heterogeneity (caused by rectal peristalsis, noise artifacts, and missing modalities) and 2) site heterogeneity (caused by different imaging protocols and patient populations). These factors hinder the model from capturing reliable cross-modal relationships and adapting to distribution shifts across clinical sites. In this work, we propose UICSurv, a novel multimodal Survival prediction framework highlighted by Uncertainty-guided Iterative Contrastive fusion, to capture robust cross-site multimodal interactions while leveraging sample-level uncertainty to enhance fusion reliability. Specifically, UICSurv initializes a shared multimodal embedding and iteratively refines it by fusing each heterogeneous modality via the cross-attention mechanism. In each iteration, a novel Survival Contrastive Learning (SCL) strategy is designed to progressively enhance both cross-site alignment and survival discriminability of the multimodal embedding space. Moreover, we design an EvidenceHit module, which employs temporally consistent evidential learning to jointly estimate survival probabilities and uncertainty. The estimated uncertainty further guides the embedding alignment by reducing the interference of unreliable samples. All components operate synergistically within UICSurv to reinforce reliable survival prediction in rectal cancer. Extensive experiments on multimodal datasets of rectal cancer (collected from three sites) demonstrate the superiority of our method both in survival prediction and uncertainty estimation. The code is available open-source: https://github.com/ScorpioBao/UICSurv. Qingsen Bao, Lei Chen 0011, Kaicong Sun, Yiqun Sun, Fu Xiao 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 4 |
| 2026 | Positional Prompts-Enhanced Brain-Heart-Gut Interactions for Mild Cognitive Impairment DiagnosisabstractMild cognitive impairment (MCI) is the prodromal stage of dementia involving complex interactions between the brain and peripheral organs. Emerging evidence indicates that heart dysfunction and gut microbiota dysbiosis can contribute to MCI pathogenesis. Yet, these discoveries of cross-organ interactions have not been applied to assist MCI diagnosis. In this work, we propose a novel diagnostic framework that exploits the interactions of brain, heart, and gut using whole-body PET images to guide MCI diagnosis for scenarios when only brain MRI, PET, or PET&MRI are available. Specifically, we collected a multi-cohort, multi-modal dataset comprising 1,545 whole-body PET images, 6,010 brain MR images, and 2,446 brain PET images from eight data centers. Organ-specific image encoders are first pretrained for the brain, heart, and gut individually. Then, to effectively align and integrate brain, heart, and gut features, we introduce positional prompts to act as anatomical-level attention to highlight disease-relevant spatial regions, and further develop hierarchical Transformers to model brain-heart, brain-gut, and brain-heart-gut interactions. Finally, to achieve MCI diagnosis using only brain images, we transfer the above brain-heart-gut model to a brain-only model via an introduced multi-level knowledge distillation scheme, including sample-level contrastive distillation, group-level distribution alignment, and response-level supervision. Extensive experiments on multi-center data demonstrate the superiority of our method over the state-of-the-art methods by resorting to effective integration of heart and gut interactions for MCI diagnosis. Shilun Zhao, Shuwei Bai, Dengqiang Jia, Jiangtao Liang, Han Zhang 0002, Ya Zhang 0002, Zhongxiang Ding, Yin Xu 0001, Kaicong Sun, Dinggang Shen |
IEEE Trans. Medical Imaging | 11 |
| 2025 | Tree-Diffusion: Octree-Based Conditional Diffusion Model for Small Bowel Skeleton Generation with Geometric Direction ModelingabstractAccurate 3D reconstruction of the small bowel skeleton is vital for understanding intestinal morphology, de-tecting structural abnormalities, and supporting diagnosis, yet limited resolution, organ adhesion, complex anatomy, and scarce annotations make continuous skeleton extraction from masks challenging. Voxel-based methods often struggle with the sparse topology and geometric directionality inherent in the small bowel skeleton, leading to inefficiency and high memory cost. To address these limitations, we propose a novel octree-based conditional diffusion model (i.e., Tree-Diffusion) that generates anatomically consistent small bowel skeletons guided by 3D segmentation masks. Specifically, we introduce two modules that captures structural priors from masks and topology characteristics from skeletons, ensuring cross-domain alignment and high-quality skeleton generation. Besides, we design a synthesis strategy to generate anatomically plausible skeleton-mask pairs, serving as topological priors to guide the diffusion model toward realis-tic structure predictions. To efficiently represent the elongated skeleton, we adopt an octree- based spatial encoding of hierarchical geometric features. Compared with baselines, our model achieves superior performance in anatomical fidelity, directional consistency, and inference efficiency. The code is available at: https://github.com/Small-Bowel-Skeleton-GenerationlCode Zhichao Liang, Dengqiang Jia, Yaofei Duan, Xinyu Xie, Kaicong Sun, Zhiming Cui 0001, Tao Tan 0002, Dinggang Shen |
BIBM | 6 |
| 2025 | A Semi-Supervised Knowledge Distillation Framework for Left Ventricle Segmentation and Landmark Detection in Echocardiograms
Yonghao Li, Han Wu 0007, Kaicong Sun, Dinggang Shen |
MICCAI (8) | 6 |
| 2025 | Wavelet-Driven Decoupling and Physics-Informed Mapping Network for Accelerated Multi-parametric MR Imaging
Ruilong Dan, Kaicong Sun, Minqiang Jia, Han Zhang 0002, Xiaopeng Zong, Dinggang Shen |
MICCAI (1) | 2 |
| 2025 | Brain-Heart-Gut Guided Multi-constraint Knowledge Distillation for Early Alzheimer's Disease Diagnosis
Shilun Zhao, Shuwei Bai, Kai Zhang 0039, Yin Xu 0001, Ya Zhang 0002, Kaicong Sun, Dinggang Shen |
MICCAI (15) | 8 |
| 2025 | Unisyn: A Generative Foundation Model for Universal Medical Image Synthesis Across MRI, CT and PET
Honglin Xiong, Kaicong Sun, Jiameng Liu, Yuanzhe He, Qian Wang 0001, Dinggang Shen |
MICCAI (3) | 3 |
| 2025 | Location-Guided Automated Lesion Captioning in Whole-Body PET/CT Images
Mingyang Yu 0009, Yaozong Gao, Yiran Shu, Yanbo Chen 0003, Jingyu Liu 0002, Caiwen Jiang, Kaicong Sun, Zhiming Cui 0001, Weifang Zhang, Yiqiang Zhan, Xiang Sean Zhou, Shaonan Zhong, Xinlu Wang, Meixin Zhao, Dinggang Shen |
MICCAI (5) | 7 |
| 2025 | Learning contrast and content representations for synthesizing magnetic resonance image of arbitrary contrast
Honglin Xiong, Zhenrong Shen 0001, Kaicong Sun, Yu Fang 0008, Dinggang Shen, Qian Wang 0001 |
Medical Image Anal. | 4 |
| 2025 | A Modality-Flexible Framework for Alzheimer's Disease Diagnosis Following Clinical RoutineabstractDementia has high incidence among the elderly, and Alzheimer's disease (AD) is the most common dementia. The procedure of AD diagnosis in clinics usually follows a standard routine consisting of different phases, from acquiring non-imaging tabular data in the screening phase to MR imaging and ultimately to PET imaging. Most of the existing AD diagnosis studies are dedicated to a specific phase using either single or multi-modal data. In this paper, we introduce a modality-flexible classification framework, which is applicable for different AD diagnosis phases following the clinical routine. Specifically, our framework consists of three branches corresponding to three diagnosis phases: 1) a tabular branch using only tabular data for screening phase, 2) an MRI branch using both MRI and tabular data for uncertain cases in screening phase, and 3) ultimately a PET branch for the challenging cases using all the modalities including PET, MRI, and tabular data. To achieve effective fusion of imaging and non-imaging modalities, we introduce an image-tabular transformer block to adaptively scale and shift the image and tabular features according to modality importance determined by the network. The proposed framework is extensively validated on four cohorts containing 6495 subjects. Experiments demonstrate that our framework achieves superior diagnostic performance than the other representative methods across various AD diagnosis tasks, and shows promising performance for all the diagnosis phases, which exhibits great potential for clinical application. Yuanwang Zhang, Kaicong Sun, Qihao Guo, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | MIP-Enhanced Uncertainty-Aware Network for Fast 7T Time-of-Flight MRA ReconstructionabstractTime-of-flight (TOF) magnetic resonance angiography (MRA) is the dominant non-contrast MR imaging method for visualizing intracranial vascular system. The employment of 7T MRI for TOF-MRA is of great interest due to its outstanding spatial resolution and vessel-tissue contrast. However, high-resolution 7T TOF-MRA is undesirably slow to acquire. Besides, due to complicated and thin structures of brain vessels, reliability of reconstructed vessels is of great importance. In this work, we propose an uncertainty-aware reconstruction model for accelerated 7T TOF-MRA, which combines the merits of deep unrolling and evidential deep learning, such that our model not only provides promising MRI reconstruction, but also supports uncertainty quantification within a single inference. Moreover, we propose a maximum intensity projection (MIP) loss for TOF-MRA reconstruction to improve the quality of MIP images. In the experiments, we have evaluated our model on a relatively large in-house multi-coil 7T TOF-MRA dataset extensively, showing promising superiority of our model compared to state-of-the-art models in terms of both TOF-MRA reconstruction and uncertainty quantification. Kaicong Sun, Caohui Duan, Dinggang Shen |
IEEE Trans. Medical Imaging | 1 |
| 2025 | 3D MedDiffusion: A 3D Medical Latent Diffusion Model for Controllable and High-Quality Medical Image GenerationabstractThe generation of medical images presents significant challenges due to their high-resolution and three-dimensional nature. Existing methods often yield suboptimal performance in generating high-quality 3D medical images, and there is currently no universal generative framework for medical imaging. In this paper, we introduce a 3D Medical Latent Diffusion (3D MedDiffusion) model for controllable, high-quality 3D medical image generation. 3D MedDiffusion incorporates a novel, highly efficient Patch-Volume Autoencoder that compresses medical images into latent space through patch-wise encoding and recovers back into image space through volume-wise decoding. Additionally, we design a new noise estimator to capture both local details and global structural information during diffusion denoising process. 3D MedDiffusion can generate fine-detailed, high-resolution images (up to ${512}\times {512}\times {512}$ ) and effectively adapt to various downstream tasks as it is trained on large-scale datasets covering CT and MRI modalities and different anatomical regions (from head to leg). Experimental results demonstrate that 3D MedDiffusion surpasses state-of-the-art methods in generative quality and exhibits strong generalizability across tasks such as sparse-view CT reconstruction, fast MRI reconstruction, and data augmentation for segmentationand classification. Source code and checkpoints are available at https://github.com/ShanghaiTech-IMPACT/3D-MedDiffusion. Haoshen Wang, Kaicong Sun, Dinggang Shen, Zhiming Cui 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2024 | NaMa: Neighbor-Aware Multi-Modal Adaptive Learning for Prostate Tumor Segmentation on Anisotropic MR ImagesabstractAccurate segmentation of prostate tumors from multi-modal magnetic resonance (MR) images is crucial for diagnosis and treatment of prostate cancer. However, the robustness of existing segmentation methods is limited, mainly because these methods 1) fail to adaptively assess subject-specific information of each MR modality for accurate tumor delineation, and 2) lack effective utilization of inter-slice information across thick slices in MR images to segment tumor as a whole 3D volume. In this work, we propose a two-stage neighbor-aware multi-modal adaptive learning network (NaMa) for accurate prostate tumor segmentation from multi-modal anisotropic MR images. In particular, in the first stage, we apply subject-specific multi-modal fusion in each slice by developing a novel modality-informativeness adaptive learning (MIAL) module for selecting and adaptively fusing informative representation of each modality based on inter-modality correlations. In the second stage, we exploit inter-slice feature correlations to derive volumetric tumor segmentation. Specifically, we first use a Unet variant with sequence layers to coarsely capture slice relationship at a global scale, and further generate an activation map for each slice. Then, we introduce an activation mapping guidance (AMG) module to refine slice-wise representation (via information from adjacent slices) for consistent tumor segmentation across neighboring slices. Besides, during the network training, we further apply a random mask strategy to each MR modality to improve feature representation efficiency. Experiments on both in-house and public (PICAI) multi-modal prostate tumor datasets show that our proposed NaMa performs better than state-of-the-art methods. Runqi Meng, Xiao Zhang 0028, Yuning Gu, Guiqin Liu, Nizhuan Wang 0001, Kaicong Sun, Dinggang Shen |
AAAI | 8 |
| 2024 | k-t Self-consistency Diffusion: A Physics-Informed Model for Dynamic MR Imaging
Zhuo-Xu Cui, Kaicong Sun, Yuliang Zhu, Dinggang Shen, Dong Liang 0001 |
MICCAI (7) | 3 |
| 2024 | LM-UNet: Whole-Body PET-CT Lesion Segmentation with Dual-Modality-Based Annotations Driven by Latent Mamba U-Net
Anglin Liu, Dengqiang Jia, Kaicong Sun, Runqi Meng, Meixin Zhao, Yongluo Jiang, Zhijian Dong, Yaozong Gao, Dinggang Shen |
MICCAI (9) | 3 |
| 2024 | UinTSeg: Unified Infant Brain Tissue Segmentation with Anatomy Delineation
Jiameng Liu, Feihong Liu, Kaicong Sun, Caiwen Jiang, Islem Rekik, Dinggang Shen |
MICCAI (2) | 3 |
| 2024 | Hierarchical Symmetric Normalization Registration Using Deformation-Inverse Network
Qingrui Sha, Kaicong Sun, Yonghao Li, Zhong Xue, Xiaohuan Cao, Dinggang Shen |
MICCAI (2) | 2 |
| 2024 | Contrast Representation Learning from Imaging Parameters for Magnetic Resonance Image Synthesis
Honglin Xiong, Yu Fang 0008, Kaicong Sun, Xiaopeng Zong, Qian Wang 0001 |
MICCAI (7) | 3 |
| 2024 | Detail-preserving image warping by enforcing smooth image sampling
Qingrui Sha, Kaicong Sun, Caiwen Jiang, Zhong Xue, Xiaohuan Cao, Dinggang Shen |
Neural Networks | 2 |
| 2024 | Multi-Modal Modality-Masked Diffusion Network for Brain MRI Synthesis With Random Modality MissingabstractSynthesis of unavailable imaging modalities from available ones can generate modality-specific complementary information and enable multi-modality based medical images diagnosis or treatment. Existing generative methods for medical image synthesis are usually based on cross-modal translation between acquired and missing modalities. These methods are usually dedicated to specific missing modality and perform synthesis in one shot, which cannot deal with varying number of missing modalities flexibly and construct the mapping across modalities effectively. To address the above issues, in this paper, we propose a unified Multi-modal Modality-masked Diffusion Network (M2DN), tackling multi-modal synthesis from the perspective of "progressive whole-modality inpainting", instead of "cross-modal translation". Specifically, our M2DN considers the missing modalities as random noise and takes all the modalities as a unity in each reverse diffusion step. The proposed joint synthesis scheme performs synthesis for the missing modalities and self-reconstruction for the available ones, which not only enables synthesis for arbitrary missing scenarios, but also facilitates the construction of common latent space and enhances the model representation ability. Besides, we introduce a modality-mask scheme to encode availability status of each incoming modality explicitly in a binary mask, which is adopted as condition for the diffusion model to further enhance the synthesis performance of our M2DN for arbitrary missing scenarios. We carry out experiments on two public brain MRI datasets for synthesis and downstream segmentation tasks. Experimental results demonstrate that our M2DN outperforms the state-of-the-art models significantly and shows great generalizability for arbitrary missing modalities. Kaicong Sun, Jun Xu 0019, Xuming He 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Joint Cross-Attention Network With Deep Modality Prior for Fast MRI ReconstructionabstractCurrent deep learning-based reconstruction models for accelerated multi-coil magnetic resonance imaging (MRI) mainly focus on subsampled k-space data of single modality using convolutional neural network (CNN). Although dual-domain information and data consistency constraint are commonly adopted in fast MRI reconstruction, the performance of existing models is still limited mainly by three factors: inaccurate estimation of coil sensitivity, inadequate utilization of structural prior, and inductive bias of CNN. To tackle these challenges, we propose an unrolling-based joint Cross-Attention Network, dubbed as jCAN, using deep guidance of the already acquired intra-subject data. Particularly, to improve the performance of coil sensitivity estimation, we simultaneously optimize the latent MR image and sensitivity map (SM). Besides, we introduce Gating layer and Gaussian layer into SM estimation to alleviate the "defocus" and "over-coupling" effects and further ameliorate the SM estimation. To enhance the representation ability of the proposed model, we deploy Vision Transformer (ViT) and CNN in the image and k-space domains, respectively. Moreover, we exploit pre-acquired intra-subject scan as reference modality to guide the reconstruction of subsampled target modality by resorting to the self- and cross-attention scheme. Experimental results on public knee and in-house brain datasets demonstrate that the proposed jCAN outperforms the state-of-the-art methods by a large margin in terms of SSIM and PSNR for different acceleration factors and sampling masks. Our code is publicly available at https://github.com/sunkg/jCAN. Kaicong Sun, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 1 |
| 2024 | An Anatomy- and Topology-Preserving Framework for Coronary Artery SegmentationabstractCoronary artery segmentation is critical for coronary artery disease diagnosis but challenging due to its tortuous course with numerous small branches and inter-subject variations. Most existing studies ignore important anatomical information and vascular topologies, leading to less desirable segmentation performance that usually cannot satisfy clinical demands. To deal with these challenges, in this paper we propose an anatomy- and topology-preserving two-stage framework for coronary artery segmentation. The proposed framework consists of an anatomical dependency encoding (ADE) module and a hierarchical topology learning (HTL) module for coarse-to-fine segmentation, respectively. Specifically, the ADE module segments four heart chambers and aorta, and thus five distance field maps are obtained to encode distance between chamber surfaces and coarsely segmented coronary artery. Meanwhile, ADE also performs coronary artery detection to crop region-of-interest and eliminate foreground-background imbalance. The follow-up HTL module performs fine segmentation by exploiting three hierarchical vascular topologies, i.e., key points, centerlines, and neighbor connectivity using a multi-task learning scheme. In addition, we adopt a bottom-up attention interaction (BAI) module to integrate the feature representations extracted across hierarchical topologies. Extensive experiments on public and in-house datasets show that the proposed framework achieves state-of-the-art performance for coronary artery segmentation. Xiao Zhang 0028, Kaicong Sun, Dijia Wu, Xiaosong Xiong, Jiameng Liu, Linlin Yao, Shufang Li, Jun Feng 0003, Dinggang Shen |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Developing Large Pre-trained Model for Breast Tumor Segmentation from Ultrasound Images
Meiyu Li, Kaicong Sun, Yuning Gu, Kai Zhang 0039, Yiqun Sun, Zhenhui Li, Dinggang Shen |
MICCAI (7) | 2 |
| 2023 | Adult-Like Phase and Multi-scale Assistance for Isointense Infant Brain Tissue Segmentation
Jiameng Liu, Feihong Liu, Kaicong Sun, Mianxin Liu, Yuyan Ge, Dinggang Shen |
MICCAI (4) | 3 |
| 2022 | An FPGA-Based Residual Recurrent Neural Network for Real-Time Video Super-ResolutionabstractIn this paper, we propose a hardware-efficient residual recurrent neural network for real-time video super-resolution (VSR) based on field programmable gate array (FPGA). Although recent learning-based VSR methods have achieved remarkable performance, the large computational complexity prohibits the deployment of the sophisticated VSR models on FPGA for real-time applications. Limited by the hardware resources, state-of-the-art FPGA-based VSR methods perform single-image super-resolution over the video sequence and suffer from temporal inconsistency. In order to exploit the inter-frame temporal correlation for real-time VSR on low-complexity hardware, we introduce a hardware-efficient recurrent neural network ERVSR. Specially, the proposed ERVSR leverages the input frame and the temporal information entailed in the hidden state to reconstruct the high-resolution counterpart. To reduce the network parameters, the low-resolution input branch and the hidden state branch are convolved individually and a channel modulation coefficient is proposed to explicitly guide the network to allocate the amount of output feature channels to each branch. Additionally, in order to reduce the memory consumption, we perform a dedicated lightweight compression of the hidden state by introducing a statistical normalization scheme followed by a fixed-point quantization. Besides, we adopt group convolution and depthwise separable convolution to further compact the network. We evaluated the proposed ERVSR on multiple public datasets from different aspects. Experimental results demonstrate that ERVSR performs better than the existing state-of-the-art FPGA-based VSR methods in both image quality and data throughput. Kaicong Sun, Maurice Koch, Zhe Wang 0008, Slavisa Jovanovic, Hassan Rabah, Sven Simon 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Multi-frame super-resolution reconstruction based on mixed Poisson-Gaussian noise
Kaicong Sun, Trung-Hieu Tran, Roman Krawtschenko, Sven Simon 0001 |
Signal Process. Image Commun. | 1 |