Shuo Han 0009

dblp:20/7794-9 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Cross-Distribution Diffusion Priors-Driven Iterative Reconstruction for Sparse-View CT
abstract
Sparse-View CT (SVCT) reconstruction improves temporal resolution and reduces radiation dose, yet its clinical use is hindered by artifacts due to view reduction and domain shifts from scanner, protocol, or anatomical variations, leading to performance degradation in out-of-distribution (OOD) scenarios. We propose a Cross-Distribution Diffusion Priors-Driven Iterative Reconstruction (CDPIR) framework to tackle the OOD problem in SVCT. CDPIR integrates cross-distribution diffusion priors, derived from a Scalable Interpolant Transformer (SiT), with model-based iterative reconstruction methods. Specifically, we train a SiT backbone, an extension of the Diffusion Transformer (DiT) architecture, to establish a unified stochastic interpolant framework, leveraging Classifier-Free Guidance (CFG) across multiple datasets. By randomly dropping the conditioning with a null embedding during training, the model learns a more transferable cross-distribution prior that encourages domain-invariant anatomical structures while allowing domain-specific appearance modulation. During sampling, the globally sensitive transformer-based diffusion model exploits the cross-distribution prior within the unified stochastic interpolant framework, enabling flexible and stable control over multi-distribution-to-noise interpolation paths and decoupled sampling strategies, thereby improving adaptation to OOD reconstruction. By alternating between data fidelity and sampling updates, our model achieves state-of-the-art performance with superior detail preservation in SVCT reconstructions. Extensive experimental results demonstrate that CDPIR significantly outperforms existing approaches, particularly under OOD conditions, highlighting its robustness and potential clinical value in challenging imaging scenarios. The code is available at https://github.com/Graeme-Lee/CDPIR.
Shuo Han 0009, Haiyang Mao, Changsheng Fang, Jianjia Zhang, Weiwen Wu, Hengyong Yu
IEEE Trans. Medical Imaging2
2026 Clinical Metadata-Guided Limited-Angle CT Image Reconstruction
abstract
Limited-angle computed tomography (LACT) improves temporal resolution and reduces radiation dose, but suffers from severe artifacts due to missing projections. Clinical workflows record abundant patient- and acquisition-level metadata, yet such information remains underutilized in image reconstruction. To tackle the ill-posed LACT inverse problem, we propose a metadata-guided two-stage diffusion framework that leverages structured clinical contexts as semantic priors for robust reconstruction. In Stage-I, we learn a metadata-to-anatomy generative prior by conditioning a transformer-based diffusion model on clinical metadata (acquisition parameters, patient demographics, and diagnostic impressions), and sampling a coarse anatomical estimate from Gaussian noise. In Stage-II, a second conditional diffusion model performs coarse-to-fine refinement, using the Stage-I estimate as an image prior while re-injecting the same metadata to recover full-resolution anatomy. To preserve anatomical fidelity and suppress hallucinations, projection-domain data consistency is enforced periodically after denoising update via an ADMM-based solver. Experiments on the public multimodal CTRATE dataset demonstrate that the proposed framework outperforms iterative, CNN-based, and diffusion-based baselines, with the greatest gains under severe truncation, e.g., up to 5.23%/11.21% higher SSIM/PSNR than the strongest metadata-free diffusion competitor at 90°. On real clinical cardiac CT, it yields coronary artery calcium scores closer to full-view references, indicating improved clinical utility. Furthermore, the proposed method is generalized to out-of-distribution angular ranges and projection geometries, and ablation results confirm complementary contributions from different metadata types under limited-angle conditions. Our results highlight clinical metadata as actionable semantic priors to synergize with physics-informed constraints to improve both reconstruction fidelity and clinical quantification in LACT.
Shuyi Fan, Changsheng Fang, Shuo Han 0009, Li Zhou 0014, Bahareh Morovati, Dayang Wang, Hengyong Yu
IEEE Trans. Medical Imaging4
2025 ρ-NeRF: Leveraging Attenuation Priors in Neural Radiance Field for 3d Computed Tomography Reconstruction
abstract
This paper introduces ρ-NeRF, a novel approach that integrates attenuation priors into neural radiance fields, setting a new benchmark in novel view synthesis (NVS) and computed tomography (CT) reconstruction. ρ-NeRF models a continuous volumetric radiance field enriched with physics-based attenuation priors, representing a three-dimensional (3D) volume through a fully connected neural network. The model takes a single four-dimensional (4D) coordinate—spatial location (x, y, z) and an initialized attenuation value (ρ)—and outputs the attenuation coefficient at that position. By querying these 4D coordinates along X-ray paths, the classic forward projection technique integrates attenuation data across the 3D space. Matching and refining pre-initialized attenuation values from traditional reconstruction algorithms, like Feldkamp-Davis-Kress algorithm (FDK) or conjugate gradient least squares (CGLS), ρ-NeRF enhances both projection synthesis and image reconstruction with minimal computational overhead. This paper details the optimization of ρ-NeRF for accurate NVS and high-quality CT reconstruction from limited projections, establishing a new standard for sparse-view CT applications.
Li Zhou 0014, Changsheng Fang, Bahareh Morovati, Yongtong Liu, Shuo Han 0009, Yongshun Xu, Hengyong Yu
ICIP5
2025 Physics-Informed Score-Based Diffusion Model for Limited-Angle Reconstruction of Cardiac Computed Tomography
abstract
Cardiac computed tomography (CT) has emerged as a major imaging modality for the diagnosis and monitoring of cardiovascular diseases. High temporal resolution is essential to ensure diagnostic accuracy. Limited-angle data acquisition can reduce scan time and improve temporal resolution, but typically leads to severe image degradation and motivates for improved reconstruction techniques. In this paper, we propose a novel physics-informed score-based diffusion model (PSDM) for limited-angle reconstruction of cardiac CT. At the sampling time, we combine a data prior from a diffusion model and a model prior obtained via an iterative algorithm and Fourier fusion to further enhance the image quality. Specifically, our approach integrates the primal-dual hybrid gradient (PDHG) algorithm with score-based diffusion models, thereby enabling us to reconstruct high-quality cardiac CT images from limited-angle data. The numerical simulations and real data experiments confirm the effectiveness of our proposed approach.
Shuo Han 0009, Yongshun Xu, Dayang Wang, Bahareh Morovati, Li Zhou 0014, Jonathan S. Maltz, Ge Wang 0001, Hengyong Yu
IEEE Trans. Medical Imaging1
2024 Gradient Guided Co-Retention Feature Pyramid Network for LDCT Image Denoising
Li Zhou 0014, Dayang Wang, Yongshun Xu, Shuo Han 0009, Bahareh Morovati, Shuyi Fan, Hengyong Yu
MICCAI (12)4
2024 LoMAE: Simple Streamlined Low-Level Masked Autoencoders for Robust, Generalized, and Interpretable Low-Dose CT Denoising
abstract
Low-dose computed tomography (LDCT) offers reduced X-ray radiation exposure but at the cost of compromised image quality, characterized by increased noise and artifacts. Recently, transformer models emerged as a promising avenue to enhance LDCT image quality. However, the success of such models relies on a large amount of paired noisy and clean images, which are often scarce in clinical settings. In computer vision and natural language processing, masked autoencoders (MAE) have been recognized as a powerful self-pretraining method for transformers, due to their exceptional capability to extract representative features. However, the original pretraining and fine-tuning design fails to work in low-level vision tasks like denoising. In response to this challenge, we redesign the classical encoder-decoder learning model and facilitate a simple yet effective streamlined low-level vision MAE, referred to as LoMAE, tailored to address the LDCT denoising problem. Moreover, we introduce an MAE-GradCAM method to shed light on the latent learning mechanisms of the MAE/LoMAE. Additionally, we explore the LoMAE's robustness and generability across a variety of noise levels. Experimental findings show that the proposed LoMAE enhances the denoising capabilities of the transformer and substantially reduce their dependency on high-quality, ground-truth data. It also demonstrates remarkable robustness and generalizability over a spectrum of noise levels. In summary, the proposed LoMAE provides promising solutions to the major issues in LDCT including interpretability, ground truth data dependency, and model robustness/generalizability.
Dayang Wang, Shuo Han 0009, Yongshun Xu, Zhan Wu, Li Zhou 0014, Bahareh Morovati, Hengyong Yu
IEEE J. Biomed. Health Informatics2