Zhanli Hu

dblp:46/8349 · DBLP profile ↗
← Back
32ranked-venue papers
1as first author
31since 2021 · last 2026
0000-0003-0618-6240ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 25 · 1 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021
YearPublicationVenuePosition
2026 FourierPET: Deep Fourier-based Unrolled Network for Low-count PET Reconstruction
abstract
Low-count positron emission tomography (PET) reconstruction is a challenging inverse problem due to severe degradations arising from Poisson noise, photon scarcity, and attenuation correction errors. Existing deep learning methods typically address these in the spatial domain with an undifferentiated optimization objective, making it difficult to disentangle overlapping artifacts and limiting correction effectiveness. In this work, we perform a Fourier-domain analysis and reveal that these degradations are spectrally separable: Poisson noise and photon scarcity cause high-frequency phase perturbations, while attenuation errors suppress low-frequency amplitude components. Leveraging this insight, we propose FourierPET, a Fourier-based unrolled reconstruction framework grounded in the Alternating Direction Method of Multipliers. It consists of three tailored modules: a spectral consistency module that enforces global frequency alignment to maintain data fidelity, an amplitude–phase correction module that decouples and compensates for high-frequency phase distortions and low-frequency amplitude suppression, and a dual adjustment module that accelerates convergence during iterative reconstruction. Extensive experiments demonstrate that FourierPET achieves state-of-the-art performance with significantly fewer parameters, while offering enhanced interpretability through frequency-aware correction.
Hao Tang 0007, Zhanli Hu, Harry Qin
AAAI4
2026 Unlocking 2D/3D+T myocardial mechanics from cine MRI: a mechanically regularized space-time finite element correlation framework
Haizhou Liu, Xueling Qin, Yuxi Jin, Jidong Han, Lingtao Mao, François Hild, Hairong Zheng, Dong Liang 0001, Na Zhang 0001, Jiuping Liang, Dehong Luo, Zhanli Hu
Medical Image Anal.17
2026 Prompt guiding multi-scale adaptive sparse representation-driven network for low-dose CT MAR
Baoshun Shi, Huazhu Fu, Zhanli Hu
Medical Image Anal.5
2026 Mamba-SUM: A Mamba-Based Framework With Wavelet Transformation for Total-Body Ultra-Low-Dose PET/CT Imaging
abstract
Long-axial PET/CT systems have enabled ultrahigh sensitivity and a longer axial field of view for clinical imaging and diagnosis. However, radiation risks from radiotracers and CT scans have remained a persistent concern within total-body PET/CT systems. Conventional approaches focus mainly on PET radiotracer-based dose reduction, ignoring the substantial radiation burden inherent in CT acquisition. Therefore, we proposed a hybrid ultra-low-dose imaging framework (Mamba-SUM) for total-body PET/CT systems to restore high-quality PET images from ultra-low-dose PET (ULD PET) and ultra-low-dose CT (ULD CT) images. Our method innovatively integrates the Mamba architecture with wavelet transformation, enabling effective modeling of long-range dependencies while reducing computational overhead. Specifically, ULD PET and ULD CT images are first subjected to domain decomposition. Afterward, a custom-designed Low-Frequency Enhancement Module and a High-Frequency Denoising Module work in concert to leverage cross-domain and multimodal information, enhancing structural details and suppressing noise across different frequency subbands. Finally, a Mamba-based decoder progressively reconstructs the refined features to produce high-quality PET images with improved fidelity and diagnostic value. Experimental results have illustrated that our method achieved superior performance (PSNR: 28.88 dB $\pm ~3.26$ , SSIM: $0.92~\pm ~0.16$ , p< 0.05) compared with other models (VMamba, Mamba-Swin, SwinTransformer, CycleGAN and UNet). Moreover, the statistical analysis also revealed that the data distribution of our generated PET images was consistent with that of the ground truth (Pearson Correlation Coefficient> 0.95, p< 0.05). Our Mamba-SUM has provided a computationally effective approach for total-body ultra-low-dose PET/CT imaging. The code is available at https://github.com/LEE12365/Mamba-SUM.
Hairong Zheng, Dong Liang 0001, Zhanli Hu, Na Zhang 0001
IEEE Trans. Image Process.7
2026 Latent Diffusion Model With Estimation Posterior Sampling: A Unified Framework for General Medical Image Restoration
abstract
Clinical imaging protocols designed to accelerate acquisition or reduce radiation dose often lead to degraded image quality, compromising diagnostic confidence. The heterogeneity in degradation types and severities across imaging modalities further challenges the development of generalized restoration solutions. In this work, we introduce a unified framework that formulates medical image restoration as posterior sampling from self-supervised Latent Diffusion Models (LDMs), pretrained on multi-modal high-quality images. At the core of our method is an Estimation Posterior Sampling (EPS) strategy, which enhances both data fidelity and anatomical detail retention. EPS incorporates two key components: (i) estimated diffusion initialization to constrain sampling within the measurement-consistent solution space, and (ii) gradient-balanced optimization to adaptively trade off denoising strength and detail preservation throughout the diffusion trajectory. Unlike traditional task-specific models, our approach enables Plug-and-Play (PnP) deployment, supporting diverse degradations without retraining. Extensive experiments conducted on deterministic degradations (e.g., under-sampled MRI, sparse-view CT) and blind degradations (e.g., low-dose PET) across multiple degradation levels demonstrate superior quantitative and qualitative performance compared to both supervised baselines and state-of-the-art posterior sampling methods. Notably, our method achieves PSNR improvements of up to +2.9 dB (MRI), +1.1 dB (CT), and +0.9 dB (PET) in PnP mode. These results highlight the robustness and broad applicability of our framework for clinical deployment.
Qianhao Chen, Hanzhong Wang, Yi An, Meiyuan Wen, Hairong Zheng, Dong Liang 0001, Zhanli Hu
IEEE J. Biomed. Health Informatics8
2026 Low-Count PET Image Reconstruction With Generalized Sparsity Priors via Unrolled Deep Networks
abstract
Deep learning has demonstrated remarkable efficacy in reconstructing low-count PET (Positron EmissionTomography) images, attracting considerable attention in the medical imaging community. However, most existing deep learning approaches have not fully exploited the unique physical characteristics of PET imaging in the design of fidelity and prior regularization terms, resulting in constrained model performance and interpretability. In light of these considerations, we introduce an unrolled deep network based on maximum likelihood estimation for the Poisson distribution and a Generalized domain transformation for Sparsity learning, dubbed GS-Net. To address this complex optimization challenge, we employ the Alternating Direction Method of Multipliers (ADMM) framework, integrating a modified Expectation Maximization (EM) approach to address the primary objective and utilize the shrinkage thresholding approach to optimize the L1 norm term. Additionally, within this unrolled deep network, all hyperparameters are adaptively adjusted through end-to-end learning to eliminate the need for manual parameter tuning. Through extensive experiments on simulated patient brain datasets and real patient whole-body clinical datasets with multiple count levels, our method has demonstrated advanced performance compared to traditional non-iterative and iterative reconstruction, deep learning-based direct reconstruction, and hybrid unrolled methods, as demonstrated by qualitative and quantitative evaluations.
Minghan Fu, Bo Liao 0001, Dong Liang 0001, Zhanli Hu, Fang-Xiang Wu
IEEE J. Biomed. Health Informatics5
2026 Clinically Generalizable Low-Dose CT Denoising for Pediatric Imaging via Enhanced Diffusion Posterior Sampling
abstract
In total-body positron emission tomography and computed tomography (PET/CT) imaging, reducing the radiation dose of diagnostic CT scans is essential for minimizing overall radiation exposure, particularly in pediatric patients. Although deep learning-based denoising methods have shown promise in restoring low-dose CT (LDCT) to normal-dose CT (NDCT) quality, most approaches rely on structurally aligned paired data, which are difficult to acquire in clinical practice. Models trained on synthetic pairs often exhibit limited generalizability to real LDCT data. Unconditional diffusion models demonstrate outstanding generalizability, but fail to preserve structural fidelity. To address these challenges, we propose an enhanced diffusion posterior sampling (E-DPS) framework that combines a one-step denoiser U-Net with an unconditional diffusion model. Specifically, the U-Net estimator, trained on simulated LDCT-NDCT pairs, provides preliminary denoised outputs as structural constraints, whereas the diffusion model captures the prior distribution of NDCT images to enhance realism and generalizability. During inference, the U-Net predictions are integrated as constraints with tunable weights, thereby guiding diffusion posterior sampling. In addition, an intermediate-stage initialization strategy is introduced, significantly reducing the number of required sampling steps. Extensive experiments on simulated LDCT datasets across three dose levels demonstrate the superiority of our method, yielding average PSNR gains of +5.2% and +4.3% at unseen dose levels compared with state-of-the-art approaches. Moreover, on real LDCT images, E-DPS exhibits strong zero-shot generalizability, achieving better noise suppression while preserving anatomical detail. These results highlight the robustness and clinical potential of E-DPS for LDCT denoising.
Hongmei Tang, Qianhao Chen, Qiyang Zhang 0002, Zhaoting Cheng, Hairong Zheng, Dong Liang 0001, Zhanli Hu, Na Zhang 0001
IEEE J. Biomed. Health Informatics9
2026 An Automatic 3D PET Tumor Segmentation Framework Assisted by Geodesic Sequences
abstract
Positron Emission Tomography (PET) images reflect the metabolic rate of tracers in different tissues of the human body, crucial for early cancer diagnosis and treatment. Accurate tumor segmentation is essential to aid clinicians in determining drug dosages. Due to the low resolution of PET images, prior information (such as CT, MRI or distance information) are often incorporated to assist PET segmentation. In this paper, we propose an automatic 3D PET tumor segmentation framework assisted by geodesic sequences. Specifically, considering the intrinsic characteristics of PET images, we first construct geodesic prior, which effectively enhances the contrast between the tumor and background while suppressing noise and the influence of other tissues. To address the need for seed points in the geodesic prior, an automatic marking strategy is designed that identifies all suspected lesion regions and uses their central points as a series of seeds to generate the corresponding geodesic sequences. Subsequently, we develop a three-branch network architecture to simultaneously process PET images, geodesic sequences, and background geodesic information. To enhance image features, a distance attention mechanism is introduced at the end of the network encoder to effectively measure the similarity between different geodesic features, refining the image features. Finally, the network incorporates spatial regularization and local PET intensity information into the activation function via the Soft Threshold Dynamics with Local Intensity Fitting (STDLIF) module, further improving segmentation accuracy. Experimental results demonstrate that, compared to existing state-of-the-art algorithms, the proposed method shows better segmentation performance on both clinical and public datasets.
Dan Shao, Chuanli Cheng, Chao Zou, Zhenxing Huang, Hairong Zheng, Dong Liang 0001, Zhi-Feng Pang, Xue-Cheng Tai, Zhanli Hu
IEEE J. Biomed. Health Informatics10
2025 AGT-Diff: Anatomy-Guided Multimodal Adaptive Diffusion for Temporally-Aware PET Denoising
abstract
Positron emission tomography (PET) is essential in clinical imaging but limited by long scan times and radiation exposure. Enhancing low-dose PET (LDPET) images is vital for reducing patient burden while maintaining diagnostic reliability. Existing denoising methods either ignore anatomical information from CT or over-integrate it, leading to poor generalization and PET signal distortion. We propose AGT-Diff—Anatomy-Guided Multimodal Adaptive Diffusion for temporally-aware PET denoising. AGT-Diff introduces two key modules. The Anatomy-Guided Multimodal Adaptive Fusion (AGMA) module selectively incorporates CT-derived structural priors through learnable soft fusion, preserving PET-specific metabolic patterns. The Progressive Time-Aware Supervision (PTA) module aligns intermediateduration PET scans with diffusion steps via signal-to-noise-ratio scheduling, enabling realistic and duration-adaptive denoising. Extensive experiments on whole-body PET/CT datasets demonstrate that AGT-Diff surpasses state-of-the-art methods in quantitative accuracy and structural fidelity while reducing dependence on paired high-dose data. Its fast inference and adaptive multimodal design make AGT-Diff a practical and generalizable framework for clinical PET enhancement.
Jingxi Hu, Xiangjian He, Zhanli Hu
BIBM5
2025 Asynchronous Multi-modal Learning for Dynamic Risk Monitoring of Acute Respiratory Distress Syndrome in Intensive Care Units
Yidan Feng, Zhanli Hu, Harry Qin
MICCAI (15)4
2025 STMDiff: Spatiotemporal Matching Diffusion Model for Dual-Time-Point Total-Body PET/CT Imaging via Contrastive Learning
Zhenxing Huang, Lianghua Li, Chunyan Yang, Wenjian Qin, Na Zhang 0001, Hairong Zheng, Dong Liang 0001, Zhanli Hu
MICCAI (11)11
2025 BenchReAD: A Systematic Benchmark for Retinal Anomaly Detection
Chenyu Lian, Zhanli Hu, Harry Qin
MICCAI (2)3
2025 MMR-Mamba: Multi-modal MRI reconstruction with Mamba and spatial-frequency information fusion
Lanqing Liu, Qi Chen 0014, Zhanli Hu, Xiaohan Xing, Harry Qin
Medical Image Anal.5
2025 Multistage Diffusion Model With Phase Error Correction for Fast PET Imaging
abstract
Fast PET imaging is clinically important for reducing motion artifacts and improving patient comfort. While recent diffusion-based deep learning methods have shown promise, they often fail to capture the true PET degradation process, suffer from accumulated inference errors, introduce artifacts, and require extensive reconstruction iterations. To address these challenges, we propose a novel multistage diffusion framework tailored for fast PET imaging. At the coarse level, we design a multistage structure to approximate the temporal non-linear PET degradation process in a data-driven manner, using paired PET images collected under different acquisition duration. A Phase Error Correction Network (PECNet) ensures consistency across stages by correcting accumulated deviations. At the fine level, we introduce a deterministic cold diffusion mechanism, which simulates intra-stage degradation through interpolation between known acquisition durations—significantly reducing reconstruction iterations to as few as 10. Evaluations on [68Ga]FAPI and [18F]FDG PET datasets demonstrate the superiority of our approach, achieving peak PSNRs of 36.2 dB and 39.0 dB, respectively, with average SSIMs over 0.97. Our framework offers high-fidelity PET imaging with fewer iterations, making it practical for accelerated clinical imaging.
Zhenxing Huang, Xingyu Xie, Qianyi Yang, Xinlan Yang, Yongfeng Yang, Hairong Zheng, Dong Liang 0001, Ruohua Chen, Zhanli Hu
IEEE J. Biomed. Health Informatics12
2025 FE-DIC-Based Motion and Intensity Correction for Enhanced CEST-MRI Registration
abstract
Physiological and external motion cause inter-frame misalignment in chemical exchange saturation transfer magnetic resonance imaging (CEST-MRI), thereby compromising quantitative accuracy. In CEST-MRI, saturation effects induce intensity variations, resulting in motion-intensity coupling that makes registration particularly challenging. To address this issue, we extend the finite element digital image correlation (FE-DIC) framework by introducing an alternating correction strategy that iteratively refines both motion and intensity estimation. Unlike conventional FE-DIC approaches that assume intensity constancy, the proposed method incorporates mechanical regularization to suppress non-physical deformations, alongside intensity correction to compensate for reference-target contrast discrepancies. This mutual reinforcement enables progressively improved registration across the CEST sequence. The robustness and effectiveness of the method were evaluated on three datasets. In simulated liver data, it maintains RMSE within 0.4 pixels, reducing error by 0.5 pixels compared to RPCA & PCA (a PCA-based synthetic reference generation method for CEST registration). On clinical brain and pig cardiac data, it achieves average SSIM of 0.83, outperforming RPCA & PCA by 0.03 and surpassing CNN-based registration (e.g., AirLab) by 0.10. The consistent results across datasets highlight its generalizability, making it a promising tool for metabolic quantification in clinical and research settings.
Haizhou Liu, Yuxi Jin, Jidong Han, Ziang Di, Hairong Zheng, Dong Liang 0001, Dehong Luo, Zhanli Hu
IEEE J. Biomed. Health Informatics12
2025 IM-Diff: Implicit Multi-Contrast Diffusion Model for Arbitrary Scale MRI Super-Resolution
abstract
Diffusion models have garnered significant attention for MRI Super-Resolution (SR) and have achieved promising results. However, existing diffusion-based SR models face two formidable challenges: 1) insufficient exploitation of complementary information from multi-contrast images, which hinders the faithful reconstruction of texture details and anatomical structures; and 2) reliance on fixed magnification factors, such as 2× or 4×, which is impractical for clinical scenarios that require arbitrary scale magnification. To circumvent these issues, this paper introduces IM-Diff, an implicit multi-contrast diffusion model for arbitrary-scale MRI SR, leveraging the merits of both multi-contrast information and the continuous nature of implicit neural representation (INR). Firstly, we propose an innovative hierarchical multi-contrast fusion (HMF) module with reference-aware cross Mamba (RCM) to effectively incorporate target-relevant information from the reference image into the target image, while ensuring a substantial receptive field with computational efficiency. Secondly, we introduce multiple wavelet INR magnification (WINRM) modules into the denoising process by integrating the wavelet implicit neural non-linearity, enabling effective learning of continuous representations of MR images. The involved wavelet activation enhances space-frequency concentration, further bolstering representation accuracy and robustness in INR. Extensive experiments on three public datasets demonstrate the superiority of our method over existing state-of-the-art SR models across various magnification factors.
Lanqing Liu, Kang Wang 0004, Xuemiao Xu, Zhanli Hu, Harry Qin
IEEE J. Biomed. Health Informatics7
2025 Automatic Brain Segmentation for PET/MR Dual-Modal Images Through a Cross-Fusion Mechanism
abstract
The precise segmentation of different brain regions and tissues is usually a prerequisite for the detection and diagnosis of various neurological disorders in neuroscience. Considering the abundance of functional and structural dual-modality information for positron emission tomography/magnetic resonance (PET/MR) images, we propose a novel 3D whole-brain segmentation network with a cross-fusion mechanism introduced to obtain 45 brain regions. Specifically, the network processes PET and MR images simultaneously, employing UX-Net and a cross-fusion block for feature extraction and fusion in the encoder. We test our method by comparing it with other deep learning-based methods, including 3DUXNET, SwinUNETR, UNETR, nnFormer, UNet3D, NestedUNet, ResUNet, and VNet. The experimental results demonstrate that the proposed method achieves better segmentation performance in terms of both visual and quantitative evaluation metrics and achieves more precise segmentation in three views while preserving fine details. In particular, the proposed method achieves superior quantitative results, with a Dice coefficient of 85.73% 0.01%, a Jaccard index of 76.68% 0.02%, a sensitivity of 85.00% 0.01%, a precision of 83.26% 0.03% and a Hausdorff distance (HD) of 4.4885 14.85%. Moreover, the distribution and correlation of the SUV in the volume of interest (VOI) are also evaluated (PCC > 0.9), indicating consistency with the ground truth and the superiority of the proposed method. In future work, we will utilize our whole-brain segmentation method in clinical practice to assist doctors in accurately diagnosing and treating brain diseases.
Hongyan Tang, Zhenxing Huang, Yaping Wu, Jianmin Yuan, Yang Yang 0186, Harry Qin, Hairong Zheng, Dong Liang 0001, Zhanli Hu
IEEE J. Biomed. Health Informatics12
2025 Deep-Learning-Based Partial Volume Correction in 99mTc-TRODAT-1 SPECT for Parkinson's Disease: A Preliminary Study on Clinical Translation
abstract
99mTc-TRODAT-1 SPECT is effective for the early detection of Parkinson's disease (PD). However, SPECT images suffer from severe partial volume effect, which impairs tissue boundary clarity and subsequent quantification accuracy. This work proposes an anatomical prior- and segmentation-free deep learning (DL)-based partial volume correction (PVC) method using an attentionbased conditional generative adversarial network (Att-cGAN) for99mTc-TRODAT-1 SPECT. A population of 454 digital brain phantoms modelling anatomical and99mTc-TRODAT activity variations in different PD categories are used to generate realistic SPECT projections using the SIMIND Monte Carlo code, and then reconstructed using ordered subset expectation maximization algorithm. The dataset is split into 320, 44 and 90 used for training, validation, and testing. Att-cGAN, cGAN and U-Net are implemented based on simulated data, then directly tested on 100 retrospectively collected clinical99mTc-TRODAT data, with same acquisition and reconstruction parameters as in simulations. Non-DL PVC methods of Van-Cittert and iterative Yang are implemented for comparison. Physical and clinical metrics, as well as a no-gold standard technique (NGST) are applied to evaluate different PVC methods in the absence of clinical ground truth. Att-cGAN yields superior PVC performance in simulations as compared to other methods in physical and clinical evaluations. NGST assessment is generally consistent with the clinical metric evaluation. For the clinical study, Att-cGAN also obtains better NGST result than others striatal compartments can be discriminated on DLbased processed images. DL-PVC method is feasible for clinical PD SPECT using highly realistic simulated data.
Guang-Uei Hung, Zhanli Hu, Greta S. P. Mok
IEEE J. Biomed. Health Informatics7
2025 Prompt-Agent-Driven Integration of Foundation Model Priors for Low-Count PET Reconstruction
abstract
Low-count Positron Emission Tomography reconstruction is critical for maintaining high imaging quality while minimizing tracer doses and radiation exposure. Although integrating structural information from CT and MR data has been shown to enhance PET reconstruction, this typically requires simultaneous PET and CT/MRI scans, complicating workflows and increasing radiation exposure. Recent advancements in foundation models offer a promising alternative to in-person CT/MRI imaging, potentially overcoming these limitations. However, the use of foundation models' segmentation masks as semantic guides has been observed to introduce erroneous structures in low-count PET reconstructions. To address this challenge, this work introduces an innovative prompting agent-based framework that dynamically interacts with the foundation model to retrieve and refine priors, minimizing undue influence on the reconstruction process. Specifically, a box agent is designed for single-instance local area information retrieval, while a point agent is introduced to progressively prompt broader semantic structures globally, utilizing history point prompts. Additionally, an MDP paradigm has been developed to address the challenges of utilizing historical point prompts while maintaining the independence required by MDPs. Evaluated on both simulated and real datasets, the proposed method demonstrates superior qualitative and quantitative performance compared to state-of-the-art methods, even those leveraging in-person CT/MRI priors.
Xingyu Xie, Mu Nan, Yaping Wu, Hairong Zheng, Dong Liang 0001, Zhanli Hu
IEEE Trans. Medical Imaging9
2024 SG-Fusion: A swin-transformer and graph convolution-based multi-modal deep neural network for glioma prognosis
abstract
The integration of morphological attributes extracted from histopathological images and genomic data holds significant importance in advancing tumor diagnosis, prognosis, and grading. Histopathological images are acquired through microscopic examination of tissue slices, providing valuable insights into cellular structures and pathological features. On the other hand, genomic data provides information about tumor gene expression and functionality. The fusion of these two distinct data types is crucial for gaining a more comprehensive understanding of tumor characteristics and progression. In the past, many studies relied on single-modal approaches for tumor diagnosis. However, these approaches had limitations as they were unable to fully harness the information from multiple data sources. To address these limitations, researchers have turned to multi-modal methods that concurrently leverage both histopathological images and genomic data. These methods better capture the multifaceted nature of tumors and enhance diagnostic accuracy. Nonetheless, existing multi-modal methods have, to some extent, oversimplified the extraction processes for both modalities and the fusion process. In this study, we presented a dual-branch neural network, namely SG-Fusion. Specifically, for the histopathological modality, we utilize the Swin-Transformer structure to capture both local and global features and incorporate contrastive learning to encourage the model to discern commonalities and differences in the representation space. For the genomic modality, we developed a graph convolutional network based on gene functional and expression level similarities. Additionally, our model integrates a cross-attention module to enhance information interaction and employs divergence-based regularization to enhance the model's generalization performance. Validation conducted on glioma datasets from the Cancer Genome Atlas unequivocally demonstrates that our SG-Fusion model outperforms both single-modal methods and existing multi-modal approaches in both survival analysis and tumor grading.
Minghan Fu, Rayyan Azam Khan, Bo Liao 0001, Zhanli Hu, Fang-Xiang Wu
Artif. Intell. Medicine5
2024 Accurate Whole-Brain Image Enhancement for Low-Dose Integrated PET/MR Imaging Through Spatial Brain Transformation
abstract
Positron emission tomography/magnetic resonance imaging (PET/MRI) systems can provide precise anatomical and functional information with exceptional sensitivity and accuracy for neurological disorder detection. Nevertheless, the radiation exposure risks and economic costs of radiopharmaceuticals may pose significant burdens on patients. To mitigate image quality degradation during low-dose PET imaging, we proposed a novel 3D network equipped with a spatial brain transform (SBF) module for low-dose whole-brain PET and MR images to synthesize high-quality PET images. The FreeSurfer toolkit was applied to derive the spatial brain anatomical alignment information, which was then fused with low-dose PET and MR features through the SBF module. Moreover, several deep learning methods were employed as comparison measures to evaluate the model performance, with the peak signal-to-noise ratio (PSNR), structural similarity (SSIM) and Pearson correlation coefficient (PCC) serving as quantitative metrics. Both the visual results and quantitative results illustrated the effectiveness of our approach. The obtained PSNR and SSIM were 41.96 ± 4.91 dB (p < 0.01) and 0.9654 ± 0.0215 (p < 0.01), which achieved a 19% and 20% improvement, respectively, compared to the original low-dose brain PET images. The volume of interest (VOI) analysis of brain regions such as the left thalamus (PCC = 0.959) also showed that the proposed method could achieve a more accurate standardized uptake value (SUV) distribution while preserving the details of brain structures. In future works, we hope to apply our method to other multimodal systems, such as PET/CT, to assist clinical brain disease diagnosis and treatment.
Zhenxing Huang, Yaping Wu, Yun Dong 0002, Yongfeng Yang, Hairong Zheng, Dong Liang 0001, Zhanli Hu
IEEE J. Biomed. Health Informatics10
2024 MMCA-NET: A Multimodal Cross Attention Transformer Network for Nasopharyngeal Carcinoma Tumor Segmentation Based on a Total-Body PET/CT System
abstract
Nasopharyngeal carcinoma (NPC) is a malignant tumor primarily treated by radiotherapy. Accurate delineation of the target tumor is essential for improving the effectiveness of radiotherapy. However, the segmentation performance of current models is unsatisfactory due to poor boundaries, large-scale tumor volume variation, and the labor-intensive nature of manual delineation for radiotherapy. In this paper, MMCA-Net, a novel segmentation network for NPC using PET/CT images that incorporates an innovative multimodal cross attention transformer (MCA-Transformer) and a modified U-Net architecture, is introduced to enhance modal fusion by leveraging cross-attention mechanisms between CT and PET data. Our method, tested against ten algorithms via fivefold cross-validation on samples from Sun Yat-sen University Cancer Center and the public HECKTOR dataset, consistently topped all four evaluation metrics with average Dice similarity coefficients of 0.815 and 0.7944, respectively. Furthermore, ablation experiments were conducted to demonstrate the superiority of our method over multiple baseline and variant techniques. The proposed method has promising potential for application in other tasks.
Zhenxing Huang, Si Tang, Chuanli Cheng, Yongfeng Yang, Hairong Zheng, Dong Liang 0001, Zhanli Hu
IEEE J. Biomed. Health Informatics12
2024 OIF-Net: An Optical Flow Registration-Based PET/MR Cross-Modal Interactive Fusion Network for Low-Count Brain PET Image Denoising
abstract
The short frames of low-count positron emission tomography (PET) images generally cause high levels of statistical noise. Thus, improving the quality of low-count images by using image postprocessing algorithms to achieve better clinical diagnoses has attracted widespread attention in the medical imaging community. Most existing deep learning-based low-count PET image enhancement methods have achieved satisfying results, however, few of them focus on denoising low-count PET images with the magnetic resonance (MR) image modality as guidance. The prior context features contained in MR images can provide abundant and complementary information for single low-count PET image denoising, especially in ultralow-count (2.5%) cases. To this end, we propose a novel two-stream dual PET/MR cross-modal interactive fusion network with an optical flow pre-alignment module, namely, OIF-Net. Specifically, the learnable optical flow registration module enables the spatial manipulation of MR imaging inputs within the network without any extra training supervision. Registered MR images fundamentally solve the problem of feature misalignment in the multimodal fusion stage, which greatly benefits the subsequent denoising process. In addition, we design a spatial-channel feature enhancement module (SC-FEM) that considers the interactive impacts of multiple modalities and provides additional information flexibility in both the spatial and channel dimensions. Furthermore, instead of simply concatenating two extracted features from these two modalities as an intermediate fusion method, the proposed cross-modal feature fusion module (CM-FFM) adopts cross-attention at multiple feature levels and greatly improves the two modalities' feature fusion procedure. Extensive experimental assessments conducted on real clinical datasets, as well as an independent clinical testing dataset, demonstrate that the proposed OIF-Net outperforms the state-of-the-art methods.
Minghan Fu, Na Zhang 0001, Zhenxing Huang, Jianmin Yuan, Yongfeng Yang, Hairong Zheng, Dong Liang 0001, Fang-Xiang Wu, Zhanli Hu
IEEE Trans. Medical Imaging13
2024 Deep Generalized Learning Model for PET Image Reconstruction
abstract
Low-count positron emission tomography (PET) imaging is challenging because of the ill-posedness of this inverse problem. Previous studies have demonstrated that deep learning (DL) holds promise for achieving improved low-count PET image quality. However, almost all data-driven DL methods suffer from fine structure degradation and blurring effects after denoising. Incorporating DL into the traditional iterative optimization model can effectively improve its image quality and recover fine structures, but little research has considered the full relaxation of the model, resulting in the performance of this hybrid model not being sufficiently exploited. In this paper, we propose a learning framework that deeply integrates DL and an alternating direction of multipliers method (ADMM)-based iterative optimization model. The innovative feature of this method is that we break the inherent forms of the fidelity operators and use neural networks to process them. The regularization term is deeply generalized. The proposed method is evaluated on simulated data and real data. Both the qualitative and quantitative results show that our proposed neural network method can outperform partial operator expansion-based neural network methods, neural network denoising methods and traditional methods.
Qiyang Zhang 0002, Yumo Zhao, Debin Hu, Fuxiao Shi, Shuangliang Cao, Yongfeng Yang, Xin Liu 0053, Hairong Zheng, Dong Liang 0001, Zhanli Hu
IEEE Trans. Medical Imaging14
2023 MLNAN: Multi-level noise-aware network for low-dose CT imaging implemented with constrained cycle Wasserstein generative adversarial networks
Zhenxing Huang, Yunling Wang, Qiyang Zhang 0002, Yuxi Jin, Ruodai Wu, Guotao Quan, Dong Liang 0001, Zhanli Hu, Na Zhang 0001
Artif. Intell. Medicine10
2023 Adaptive weighted curvature-based active contour for ultrasonic and 3T/5T MR image segmentation
Zhi-Feng Pang, Mengxiao Geng, Yanru Zhou, Tieyong Zeng, Liyun Zheng, Na Zhang 0001, Dong Liang 0001, Hairong Zheng, Yongming Dai, Zhenxing Huang, Zhanli Hu
Signal Process.12
2023 A Two-Branch Neural Network for Short-Axis PET Image Quality Enhancement
abstract
The axial field of view (FOV) is a key factor that affects the quality of PET images. Due to hardware FOV restrictions, conventional short-axis PET scanners with FOVs of 20 to 35 cm can acquire only low-quality PET (LQ-PET) images in fast scanning times (2-3 minutes). To overcome hardware restrictions and improve PET image quality for better clinical diagnoses, several deep learning-based algorithms have been proposed. However, these approaches use simple convolution layers with residual learning and local attention, which insufficiently extract and fuse long-range contextual information. To this end, we propose a novel two-branch network architecture with swin transformer units and graph convolution operation, namely SW-GCN. The proposed SW-GCN provides additional spatial- and channel-wise flexibility to handle different types of input information flow. Specifically, considering the high computational cost of calculating self-attention weights in full-size PET images, in our designed spatial adaptive branch, we take the self-attention mechanism within each local partition window and introduce global information interactions between nonoverlapping windows by shifting operations to prevent the aforementioned problem. In addition, the convolutional network structure considers the information in each channel equally during the feature extraction process. In our designed channel adaptive branch, we use a Watts Strogatz topology structure to connect each feature map to only its most relevant features in each graph convolutional layer, substantially reducing information redundancy. Moreover, ensemble learning is adopted in our SW-GCN for mapping distinct features from the two well-designed branches to the enhanced PET images. We carried out extensive experiments on three single-bed position scans for 386 patients. The test results demonstrate that our proposed SW-GCN approach outperforms state-of-the-art methods in both quantitative and qualitative evaluations.
Minghan Fu, Yaping Wu, Na Zhang 0001, Yongfeng Yang, Fang-Xiang Wu, Hairong Zheng, Dong Liang 0001, Zhanli Hu
IEEE J. Biomed. Health Informatics12
2023 Four-Dimensional Cone Beam CT Imaging Using a Single Routine Scan via Deep Learning
abstract
A novel method is proposed to obtain four-dimensional (4D) cone-beam computed tomography (CBCT) images from a routine scan in patients with upper abdominal cancer. The projections are sorted according to the location of the lung diaphragm before being reconstructed to phase-sorted data. A multiscale-discriminator generative adversarial network (MSD-GAN) is proposed to alleviate the severe streaking artifacts in the original images. The MSD-GAN is trained using simulated CBCT datasets from patient planning CT images. The enhanced images are further used to estimate the deformable vector field (DVF) among breathing phases using a deformable image registration method. The estimated DVF is then applied in the motion-compensated ordered-subset simultaneous algebraic reconstruction approach to generate 4D CBCT images. The proposed MSD-GAN is compared with U-Net on the performance of image enhancement. Results show that the proposed method significantly outperforms the total variation regularization-based iterative reconstruction approach and the method using only MSD-GAN to enhance original phase-sorted images in simulation and patient studies on 4D reconstruction quality. The MSD-GAN also shows higher accuracy than the U-Net. The proposed method enables a practical way for 4D-CBCT imaging from a single routine scan in upper abdominal cancer treatment including liver and pancreatic tumors.
Tiffany Tsui, Xiaokun Liang, Yaoqin Xie, Zhanli Hu, Tianye Niu
IEEE Trans. Medical Imaging6
2021 FaNet: fast assessment network for the novel coronavirus (COVID-19) pneumonia based on 3D CT imaging and clinical symptoms
Zhenxing Huang, Xinfeng Liu, Rongpin Wang, Mudan Zhang, Xianchun Zeng, Jun Liu 0080, Yongfeng Yang, Xin Liu 0053, Hairong Zheng, Dong Liang 0001, Zhanli Hu
Appl. Intell.11
2021 Considering anatomical prior information for low-dose CT image enhancement using attribute-augmented Wasserstein generative adversarial networks
Zhenxing Huang, Xinfeng Liu, Rongpin Wang, Jincai Chen, Ping Lu 0006, Qiyang Zhang 0002, Changhui Jiang, Yongfeng Yang, Xin Liu 0053, Hairong Zheng, Dong Liang 0001, Zhanli Hu
Neurocomputing12
2021 Learning a Deep CNN Denoising Approach Using Anatomical Prior Information Implemented With Attention Mechanism for Low-Dose CT Imaging on Clinical Patient Data From Multiple Anatomical Sites
abstract
Dose reduction in computed tomography (CT) has gained considerable attention in clinical applications because it decreases radiation risks. However, a lower dose generates noise in low-dose computed tomography (LDCT) images. Previous deep learning (DL)-based works have investigated ways to improve diagnostic performance to address this ill-posed problem. However, most of them disregard the anatomical differences among different human body sites in constructing the mapping function between LDCT images and their high-resolution normal-dose CT (NDCT) counterparts. In this article, we propose a novel deep convolutional neural network (CNN) denoising approach by introducing information of the anatomical prior. Instead of designing multiple networks for each independent human body anatomical site, a unified network framework is employed to process anatomical information. The anatomical prior is represented as a pattern of weights of the features extracted from the corresponding LDCT image in an anatomical prior fusion module. To promote diversity in the contextual information, a spatial attention fusion mechanism is introduced to capture many local regions of interest in the attention fusion module. Although many network parameters are saved, the experimental results demonstrate that our method, which incorporates anatomical prior information, is effective in denoising LDCT images. Furthermore, the anatomical prior fusion module could be conveniently integrated into other DL-based methods and avails the performance improvement on multiple anatomical data.
Zhenxing Huang, Xinfeng Liu, Rongpin Wang, Zixiang Chen, Yongfeng Yang, Xin Liu 0053, Hairong Zheng, Dong Liang 0001, Zhanli Hu
IEEE J. Biomed. Health Informatics9
2011 Geometric Calibration of a Micro-CT System and Performance for Insect Imaging
abstract
Micro-CT with a high spatial resolution in combination with computer-based-reconstruction techniques is considered a powerful tool for morphological study of insects. The quality of CT images crucially depends on the precise knowledge of the scan geometry of the micro-CT system. In this paper, we have proposed a method to calculate the deviation of rotating axis for compensating deficiency of existing methods. A practical application of this geometric calibration method of the micro-CT system for insect imaging is presented. We have performed the computer-simulation study and experimental study with our prototype micro-CT system. The results demonstrate that the proposed technique is accurate and robust. In addition, we have evaluated the imaging characteristics of the detector in terms of modulation-transfer function (MTF). Finally, insect imaging performance and image reconstruction from data acquired with different energies are presented.
Zhanli Hu, Jianbao Gui, Junyan Rong, Qiyang Zhang 0002, Hairong Zheng
IEEE Trans. Inf. Technol. Biomed.1