EDBT 2026 Demo / reviewers in the wild / expert
Hairong Zheng
dblp:72/7023
· DBLP profile ↗
58ranked-venue papers
0as first author
46since 2021 · last 2026
0000-0002-8558-5102ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 48 · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 9 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unlocking 2D/3D+T myocardial mechanics from cine MRI: a mechanically regularized space-time finite element correlation framework
Haizhou Liu, Xueling Qin, Yuxi Jin, Jidong Han, Lingtao Mao, François Hild, Hairong Zheng, Dong Liang 0001, Na Zhang 0001, Jiuping Liang, Dehong Luo, Zhanli Hu |
Medical Image Anal. | 12 |
| 2026 | Switch-UMamba: Dynamic scanning vision Mamba UNet for medical image segmentation
Ziyao Zhang 0003, Qiankun Ma, Tong Zhang 0017, Jie Chen 0001, Hairong Zheng, Wen Gao 0001 |
Medical Image Anal. | 5 |
| 2026 | InterTeach: A Novel Approach for Semi-Supervised Medical Image Segmentation Using Cooperative Teacher-Student NetworksabstractIn medical image segmentation, the reliance on extensive, high-quality labeled datasets poses a significant challenge, especially considering the associated costs and the requirement for specialized expertise. In response, the field has progressively embraced semi-supervised learning (SSL) methods that leverage both labeled and unlabeled data. Nonetheless, these methods frequently encounter issues related to inconsistent label quality and constrained generalizability of models. To surmount these obstacles, we present InterTeach, an innovative SSL framework that seamlessly integrates cross-supervision with the mean teacher model. This framework facilitates effective knowledge transfer and boosts model performance through the implementation of two unique teacher-student training configurations. Herein, knowledge is exchanged between models via their respective teacher counterparts, facilitating mutual learning and enhancement. This strategy diverges from traditional SSL approaches, which mainly depend on mutual learning between two models updated through gradient descent. Furthermore, the incorporation of Feature Divergence Loss (FDL) in InterTeach encourages the transfer of diverse and complementary knowledge between models, thereby enriching the overall learning dynamics. The evaluation results revealed that our method could approach or even match the performance of fully supervised learning methods on certain evaluation metrics. This finding further confirms the effectiveness and wide applicability of the IntraTeach method in handling multi-modal and multi-dimensional medical image segmentation tasks. Ziyao Zhang 0003, Qiankun Ma, Jie Chen 0001, Hairong Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Mamba-SUM: A Mamba-Based Framework With Wavelet Transformation for Total-Body Ultra-Low-Dose PET/CT ImagingabstractLong-axial PET/CT systems have enabled ultrahigh sensitivity and a longer axial field of view for clinical imaging and diagnosis. However, radiation risks from radiotracers and CT scans have remained a persistent concern within total-body PET/CT systems. Conventional approaches focus mainly on PET radiotracer-based dose reduction, ignoring the substantial radiation burden inherent in CT acquisition. Therefore, we proposed a hybrid ultra-low-dose imaging framework (Mamba-SUM) for total-body PET/CT systems to restore high-quality PET images from ultra-low-dose PET (ULD PET) and ultra-low-dose CT (ULD CT) images. Our method innovatively integrates the Mamba architecture with wavelet transformation, enabling effective modeling of long-range dependencies while reducing computational overhead. Specifically, ULD PET and ULD CT images are first subjected to domain decomposition. Afterward, a custom-designed Low-Frequency Enhancement Module and a High-Frequency Denoising Module work in concert to leverage cross-domain and multimodal information, enhancing structural details and suppressing noise across different frequency subbands. Finally, a Mamba-based decoder progressively reconstructs the refined features to produce high-quality PET images with improved fidelity and diagnostic value. Experimental results have illustrated that our method achieved superior performance (PSNR: 28.88 dB $\pm ~3.26$ , SSIM: $0.92~\pm ~0.16$ , p< 0.05) compared with other models (VMamba, Mamba-Swin, SwinTransformer, CycleGAN and UNet). Moreover, the statistical analysis also revealed that the data distribution of our generated PET images was consistent with that of the ground truth (Pearson Correlation Coefficient> 0.95, p< 0.05). Our Mamba-SUM has provided a computationally effective approach for total-body ultra-low-dose PET/CT imaging. The code is available at https://github.com/LEE12365/Mamba-SUM. Hairong Zheng, Dong Liang 0001, Zhanli Hu, Na Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2026 | Latent Diffusion Model With Estimation Posterior Sampling: A Unified Framework for General Medical Image RestorationabstractClinical imaging protocols designed to accelerate acquisition or reduce radiation dose often lead to degraded image quality, compromising diagnostic confidence. The heterogeneity in degradation types and severities across imaging modalities further challenges the development of generalized restoration solutions. In this work, we introduce a unified framework that formulates medical image restoration as posterior sampling from self-supervised Latent Diffusion Models (LDMs), pretrained on multi-modal high-quality images. At the core of our method is an Estimation Posterior Sampling (EPS) strategy, which enhances both data fidelity and anatomical detail retention. EPS incorporates two key components: (i) estimated diffusion initialization to constrain sampling within the measurement-consistent solution space, and (ii) gradient-balanced optimization to adaptively trade off denoising strength and detail preservation throughout the diffusion trajectory. Unlike traditional task-specific models, our approach enables Plug-and-Play (PnP) deployment, supporting diverse degradations without retraining. Extensive experiments conducted on deterministic degradations (e.g., under-sampled MRI, sparse-view CT) and blind degradations (e.g., low-dose PET) across multiple degradation levels demonstrate superior quantitative and qualitative performance compared to both supervised baselines and state-of-the-art posterior sampling methods. Notably, our method achieves PSNR improvements of up to +2.9 dB (MRI), +1.1 dB (CT), and +0.9 dB (PET) in PnP mode. These results highlight the robustness and broad applicability of our framework for clinical deployment. Qianhao Chen, Hanzhong Wang, Yi An, Meiyuan Wen, Hairong Zheng, Dong Liang 0001, Zhanli Hu |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | Clinically Generalizable Low-Dose CT Denoising for Pediatric Imaging via Enhanced Diffusion Posterior SamplingabstractIn total-body positron emission tomography and computed tomography (PET/CT) imaging, reducing the radiation dose of diagnostic CT scans is essential for minimizing overall radiation exposure, particularly in pediatric patients. Although deep learning-based denoising methods have shown promise in restoring low-dose CT (LDCT) to normal-dose CT (NDCT) quality, most approaches rely on structurally aligned paired data, which are difficult to acquire in clinical practice. Models trained on synthetic pairs often exhibit limited generalizability to real LDCT data. Unconditional diffusion models demonstrate outstanding generalizability, but fail to preserve structural fidelity. To address these challenges, we propose an enhanced diffusion posterior sampling (E-DPS) framework that combines a one-step denoiser U-Net with an unconditional diffusion model. Specifically, the U-Net estimator, trained on simulated LDCT-NDCT pairs, provides preliminary denoised outputs as structural constraints, whereas the diffusion model captures the prior distribution of NDCT images to enhance realism and generalizability. During inference, the U-Net predictions are integrated as constraints with tunable weights, thereby guiding diffusion posterior sampling. In addition, an intermediate-stage initialization strategy is introduced, significantly reducing the number of required sampling steps. Extensive experiments on simulated LDCT datasets across three dose levels demonstrate the superiority of our method, yielding average PSNR gains of +5.2% and +4.3% at unseen dose levels compared with state-of-the-art approaches. Moreover, on real LDCT images, E-DPS exhibits strong zero-shot generalizability, achieving better noise suppression while preserving anatomical detail. These results highlight the robustness and clinical potential of E-DPS for LDCT denoising. Hongmei Tang, Qianhao Chen, Qiyang Zhang 0002, Zhaoting Cheng, Hairong Zheng, Dong Liang 0001, Zhanli Hu, Na Zhang 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2026 | An Automatic 3D PET Tumor Segmentation Framework Assisted by Geodesic SequencesabstractPositron Emission Tomography (PET) images reflect the metabolic rate of tracers in different tissues of the human body, crucial for early cancer diagnosis and treatment. Accurate tumor segmentation is essential to aid clinicians in determining drug dosages. Due to the low resolution of PET images, prior information (such as CT, MRI or distance information) are often incorporated to assist PET segmentation. In this paper, we propose an automatic 3D PET tumor segmentation framework assisted by geodesic sequences. Specifically, considering the intrinsic characteristics of PET images, we first construct geodesic prior, which effectively enhances the contrast between the tumor and background while suppressing noise and the influence of other tissues. To address the need for seed points in the geodesic prior, an automatic marking strategy is designed that identifies all suspected lesion regions and uses their central points as a series of seeds to generate the corresponding geodesic sequences. Subsequently, we develop a three-branch network architecture to simultaneously process PET images, geodesic sequences, and background geodesic information. To enhance image features, a distance attention mechanism is introduced at the end of the network encoder to effectively measure the similarity between different geodesic features, refining the image features. Finally, the network incorporates spatial regularization and local PET intensity information into the activation function via the Soft Threshold Dynamics with Local Intensity Fitting (STDLIF) module, further improving segmentation accuracy. Experimental results demonstrate that, compared to existing state-of-the-art algorithms, the proposed method shows better segmentation performance on both clinical and public datasets. Dan Shao, Chuanli Cheng, Chao Zou, Zhenxing Huang, Hairong Zheng, Dong Liang 0001, Zhi-Feng Pang, Xue-Cheng Tai, Zhanli Hu |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | HAVIR: Hierarchical Vision to Image Reconstruction Using Clip-Guided Versatile Diffusion
Dong Liang 0001, Hairong Zheng, Yihang Zhou |
BIBM | 3 |
| 2025 | STMDiff: Spatiotemporal Matching Diffusion Model for Dual-Time-Point Total-Body PET/CT Imaging via Contrastive Learning
Zhenxing Huang, Lianghua Li, Chunyan Yang, Wenjian Qin, Na Zhang 0001, Hairong Zheng, Dong Liang 0001, Zhanli Hu |
MICCAI (11) | 8 |
| 2025 | Test-Time Adaptation of Medical Vision-Language Models with Mixture of Modality ExpertsabstractTest-time adaptation (TTA) has emerged as a promising solution to improve the robustness of deep learning models under domain shifts, particularly in real-world scenarios where source data or true labels are unavailable. In this paper, we propose a novel TTA framework tailored for medical vision-language models (VLMs) that leverages a Mixture-of-Experts (MoE) mechanism. Building upon a frozen, pre-trained BiomedCLIP backbone, our method integrates parallel MoE adapters of different medical imaging modalities in each vision MLP block, enabling expert-specific adaptation without disrupting the core model representation. At inference time, only the MoE router is optimized through an entropy-regularized objective, which is further augmented by pseudo-label guidance and adaptive scaling strategies. Additionally, we propose an entropy-aware MoE scaling policy that dynamically adjusts expert influence based on prediction uncertainty, improving model adaptability. Extensive experiments on multiple medical imaging benchmarks demonstrate that our approach achieves substantial performance improvements over existing TTA baselines, while maintaining high efficiency and parameter sparsity. Our results highlight the potential of MoE-enhanced TTA to achieve robust and generalizable medical VLMs in unseen domains without access to source data. Code is available at https://openi.pcl.ac.cn/OpenMedIA/MoME-TTA.git Hancong Wang, Yue Yu 0001, Hairong Zheng, Tong Zhang 0017 |
ACM Multimedia | 3 |
| 2025 | Optimized Vessel Segmentation: A Structure-Agnostic Approach With Small Vessel Enhancement and Morphological CorrectionabstractAccurate segmentation of blood vessels is essential for various clinical assessments and postoperative analyses. However, the inherent challenges of vascular imaging-such as sparsity, fine granularity, low contrast, data distribution variability, and the critical need for preserving topological integrity-make generalized vessel segmentation particularly complex. While specialized segmentation methods have been developed for specific anatomical regions, their over-reliance on tailored models hinders broader applicability and generalization. General-purpose segmentation models introduced in medical imaging often fail to address critical vascular characteristics, including the connectivity of segmentation results. In this study, we propose OVS-Net, an optimized vessel segmentation framework designed to generalize across diverse vessel structures and imaging modalities. It introduces a dual-branch architecture design for improving small vessel segmentation and a morphology-aware correction module to preserve vascular topology and connectivity. We compiled a comprehensive multi-modality dataset from 17 datasets to train and benchmark the proposed OVS-Net against 6 SAM-based methods and 17 expert models under various conditions. The results demonstrate that our approach achieves superior segmentation accuracy, generalization, and a 34.6% improvement in connectivity, underscoring its potential for clinical applications. The code and dataset information are available at https://github.com/Hk416mod2/OVS-Net. Dongning Song, Weijian Huang, Jiarun Liu, Md Jahidul Islam, Hao Yang 0026, Shuqiang Wang, Hairong Zheng, Shanshan Wang 0002 |
IEEE Trans. Image Process. | 7 |
| 2025 | Multistage Diffusion Model With Phase Error Correction for Fast PET ImagingabstractFast PET imaging is clinically important for reducing motion artifacts and improving patient comfort. While recent diffusion-based deep learning methods have shown promise, they often fail to capture the true PET degradation process, suffer from accumulated inference errors, introduce artifacts, and require extensive reconstruction iterations. To address these challenges, we propose a novel multistage diffusion framework tailored for fast PET imaging. At the coarse level, we design a multistage structure to approximate the temporal non-linear PET degradation process in a data-driven manner, using paired PET images collected under different acquisition duration. A Phase Error Correction Network (PECNet) ensures consistency across stages by correcting accumulated deviations. At the fine level, we introduce a deterministic cold diffusion mechanism, which simulates intra-stage degradation through interpolation between known acquisition durations—significantly reducing reconstruction iterations to as few as 10. Evaluations on [68Ga]FAPI and [18F]FDG PET datasets demonstrate the superiority of our approach, achieving peak PSNRs of 36.2 dB and 39.0 dB, respectively, with average SSIMs over 0.97. Our framework offers high-fidelity PET imaging with fewer iterations, making it practical for accelerated clinical imaging. Zhenxing Huang, Xingyu Xie, Qianyi Yang, Xinlan Yang, Yongfeng Yang, Hairong Zheng, Dong Liang 0001, Ruohua Chen, Zhanli Hu |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | FE-DIC-Based Motion and Intensity Correction for Enhanced CEST-MRI RegistrationabstractPhysiological and external motion cause inter-frame misalignment in chemical exchange saturation transfer magnetic resonance imaging (CEST-MRI), thereby compromising quantitative accuracy. In CEST-MRI, saturation effects induce intensity variations, resulting in motion-intensity coupling that makes registration particularly challenging. To address this issue, we extend the finite element digital image correlation (FE-DIC) framework by introducing an alternating correction strategy that iteratively refines both motion and intensity estimation. Unlike conventional FE-DIC approaches that assume intensity constancy, the proposed method incorporates mechanical regularization to suppress non-physical deformations, alongside intensity correction to compensate for reference-target contrast discrepancies. This mutual reinforcement enables progressively improved registration across the CEST sequence. The robustness and effectiveness of the method were evaluated on three datasets. In simulated liver data, it maintains RMSE within 0.4 pixels, reducing error by 0.5 pixels compared to RPCA & PCA (a PCA-based synthetic reference generation method for CEST registration). On clinical brain and pig cardiac data, it achieves average SSIM of 0.83, outperforming RPCA & PCA by 0.03 and surpassing CNN-based registration (e.g., AirLab) by 0.10. The consistent results across datasets highlight its generalizability, making it a promising tool for metabolic quantification in clinical and research settings. Haizhou Liu, Yuxi Jin, Jidong Han, Ziang Di, Hairong Zheng, Dong Liang 0001, Dehong Luo, Zhanli Hu |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | Automatic Brain Segmentation for PET/MR Dual-Modal Images Through a Cross-Fusion MechanismabstractThe precise segmentation of different brain regions and tissues is usually a prerequisite for the detection and diagnosis of various neurological disorders in neuroscience. Considering the abundance of functional and structural dual-modality information for positron emission tomography/magnetic resonance (PET/MR) images, we propose a novel 3D whole-brain segmentation network with a cross-fusion mechanism introduced to obtain 45 brain regions. Specifically, the network processes PET and MR images simultaneously, employing UX-Net and a cross-fusion block for feature extraction and fusion in the encoder. We test our method by comparing it with other deep learning-based methods, including 3DUXNET, SwinUNETR, UNETR, nnFormer, UNet3D, NestedUNet, ResUNet, and VNet. The experimental results demonstrate that the proposed method achieves better segmentation performance in terms of both visual and quantitative evaluation metrics and achieves more precise segmentation in three views while preserving fine details. In particular, the proposed method achieves superior quantitative results, with a Dice coefficient of 85.73% 0.01%, a Jaccard index of 76.68% 0.02%, a sensitivity of 85.00% 0.01%, a precision of 83.26% 0.03% and a Hausdorff distance (HD) of 4.4885 14.85%. Moreover, the distribution and correlation of the SUV in the volume of interest (VOI) are also evaluated (PCC > 0.9), indicating consistency with the ground truth and the superiority of the proposed method. In future work, we will utilize our whole-brain segmentation method in clinical practice to assist doctors in accurately diagnosing and treating brain diseases. Hongyan Tang, Zhenxing Huang, Yaping Wu, Jianmin Yuan, Yang Yang 0186, Harry Qin, Hairong Zheng, Dong Liang 0001, Zhanli Hu |
IEEE J. Biomed. Health Informatics | 9 |
| 2025 | SPIRiT-Diffusion: Self-Consistency Driven Diffusion Model for Accelerated MRIabstractDiffusion models have emerged as a leading methodology for image generation and have proven successful in the realm of magnetic resonance imaging (MRI) reconstruction. However, existing reconstruction methods based on diffusion models are primarily formulated in the image domain, making the reconstruction quality susceptible to inaccuracies in coil sensitivity maps (CSMs). k-space interpolation methods can effectively address this issue but conventional diffusion models are not readily applicable in k-space interpolation. To overcome this challenge, we introduce a novel approach called SPIRiT-Diffusion, which is a diffusion model for k-space interpolation inspired by the iterative self-consistent SPIRiT method. Specifically, we utilize the iterative solver of the self-consistent term (i.e., k-space physical prior) in SPIRiT to formulate a novel stochastic differential equation (SDE) governing the diffusion process. Subsequently, k-space data can be interpolated by executing the diffusion process. This innovative approach highlights the optimization model's role in designing the SDE in diffusion models, enabling the diffusion process to align closely with the physics inherent in the optimization model-a concept referred to as model-driven diffusion. We evaluated the proposed SPIRiT-Diffusion method using a 3D joint intracranial and carotid vessel wall imaging dataset. The results convincingly demonstrate its superiority over image-domain reconstruction methods, achieving high reconstruction quality even at a substantial acceleration rate of 10. Our code are available at https://github.com/zhyjSIAT/SPIRiT-Diffusion. Zhuo-Xu Cui, Chentao Cao, Yue Wang 0119, Sen Jia 0005, Xin Liu 0053, Hairong Zheng, Dong Liang 0001, Yanjie Zhu |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Score-Based Diffusion Models With Self-Supervised Learning for Accelerated 3D Multi-Contrast Cardiac MR ImagingabstractLong scan time significantly hinders the widespread applications of three-dimensional multi-contrast cardiac magnetic resonance (3D-MC-CMR) imaging. This study aims to accelerate 3D-MC-CMR acquisition by a novel method based on score-based diffusion models with self-supervised learning. Specifically, we first establish a mapping between the undersampled k-space measurements and the MR images, utilizing a self-supervised Bayesian reconstruction network. Secondly, we develop a joint score-based diffusion model on 3D-MC-CMR images to capture their inherent distribution. The 3D-MC-CMR images are finally reconstructed using the conditioned Langenvin Markov chain Monte Carlo sampling. This approach enables accurate reconstruction without fully sampled training data. Its performance was tested on the dataset acquired by a 3D joint myocardial $ \text {T}_{{1}}$ and $ \text {T}_{{1}\rho }$ mapping sequence. The $ \text {T}_{{1}}$ and $ \text {T}_{{1}\rho }$ maps were estimated via a dictionary matching method from the reconstructed images. Experimental results show that the proposed method outperforms traditional compressed sensing and existing self-supervised deep learning MRI reconstruction methods. It also achieves high quality $ \text {T}_{{1}}$ and $ \text {T}_{{1}\rho }$ parametric maps close to the reference maps, even at a high acceleration rate of 14. Zhuo-Xu Cui, Shucong Qin, Hairong Zheng, Haifeng Wang 0003, Yihang Zhou, Dong Liang 0001, Yanjie Zhu |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Swin-UMamba†: Adapting Mamba-Based Vision Foundation Models for Medical Image SegmentationabstractVision foundation models have shown great potential in improving generalizability and data efficiency, especially for medical image segmentation since medical image datasets are relatively small due to high annotation costs and privacy concerns. However, current research on foundation models predominantly relies on transformers. The high quadratic complexity and large parameter counts make these models computationally expensive, limiting their potential for clinical applications. In this work, we introduce Swin-UMamba†, a novel Mamba-based model for medical image segmentation that seamlessly leverages the power of the vision foundation model, which is also computationally efficient with the linear complexity of Mamba. Moreover, we investigated and verified the impact of the vision foundation model on medical image segmentation, in which a self-supervised model adaptation scheme was designed to bridge the gap between natural and medical data. Notably, Swin-UMamba† outperforms 7 state-of-the-art methods, including CNN-based, transformer-based, and Mamba-based approaches across AbdomenMRI, Encoscopy, and Microscopy datasets. The code and models are publicly available at: https://github.com/JiarunLiu/Swin-UMamba. Jiarun Liu, Hao Yang 0026, Lequan Yu, Yong Liang 0001, Yizhou Yu, Shaoting Zhang 0001, Hairong Zheng, Shanshan Wang 0002 |
IEEE Trans. Medical Imaging | 8 |
| 2025 | Generalizable Reconstruction for Accelerating MR Imaging via Federated Learning With Neural Architecture SearchabstractHeterogeneous data captured by different scanning devices and imaging protocols can affect the generalization performance of the deep learning magnetic resonance (MR) reconstruction model. While a centralized training model is effective in mitigating this problem, it raises concerns about privacy protection. Federated learning is a distributed training paradigm that can utilize multi-institutional data for collaborative training without sharing data. However, existing federated learning MR image reconstruction methods rely on models designed manually by experts, which are complex and computationally expensive, suffering from performance degradation when facing heterogeneous data distributions. In addition, these methods give inadequate consideration to fairness issues, namely ensuring that the model's training does not introduce bias towards any specific dataset's distribution. To this end, this paper proposes a generalizable federated neural architecture search framework for accelerating MR imaging (GAutoMRI). Specifically, automatic neural architecture search is investigated for effective and efficient neural network representation learning of MR images from different centers. Furthermore, we design a fairness adjustment approach that can enable the model to learn features fairly from inconsistent distributions of different devices and centers, and thus facilitate the model to generalize well to the unseen center. Extensive experiments show that our proposed GAutoMRI has better performances and generalization ability compared with seven state-of-the-art federated learning methods. Moreover, the GAutoMRI model is significantly more lightweight, making it an efficient choice for MR image reconstruction tasks. The code will be made available at https://github.com/ternencewu123/GAutoMRI. Ruoyou Wu, Cheng Li 0008, Xinfeng Liu, Hairong Zheng, Shanshan Wang 0002 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Prompt-Agent-Driven Integration of Foundation Model Priors for Low-Count PET ReconstructionabstractLow-count Positron Emission Tomography reconstruction is critical for maintaining high imaging quality while minimizing tracer doses and radiation exposure. Although integrating structural information from CT and MR data has been shown to enhance PET reconstruction, this typically requires simultaneous PET and CT/MRI scans, complicating workflows and increasing radiation exposure. Recent advancements in foundation models offer a promising alternative to in-person CT/MRI imaging, potentially overcoming these limitations. However, the use of foundation models' segmentation masks as semantic guides has been observed to introduce erroneous structures in low-count PET reconstructions. To address this challenge, this work introduces an innovative prompting agent-based framework that dynamically interacts with the foundation model to retrieve and refine priors, minimizing undue influence on the reconstruction process. Specifically, a box agent is designed for single-instance local area information retrieval, while a point agent is introduced to progressively prompt broader semantic structures globally, utilizing history point prompts. Additionally, an MDP paradigm has been developed to address the challenges of utilizing historical point prompts while maintaining the independence required by MDPs. Evaluated on both simulated and real datasets, the proposed method demonstrates superior qualitative and quantitative performance compared to state-of-the-art methods, even those leveraging in-person CT/MRI priors. Xingyu Xie, Mu Nan, Yaping Wu, Hairong Zheng, Dong Liang 0001, Zhanli Hu |
IEEE Trans. Medical Imaging | 6 |
| 2024 | C2RG: Parameter-efficient Adaptation of 3D Vision and Language Foundation Model for Coronary CTA Report GenerationabstractMedical report generation (MRG) is a challenging yet highly demanding task in the application of multi-modal artificial intelligence in medicine. Typically, training an MRG model requires tens of thousands of labelled radiology images and reports datasets, which could be impractical for most clinical research groups. In this study, we present C2RG, a novel 3D vision and language foundation model tailored for Coronary Computed Tomography Angiography (CTA) Report Generation. Inspired by BLIP-2’s architecture, our method integrates a self-supervised pre-trained 3D cardiac vision model (ViT-B) and a general-purpose bilingual foundation model (ChatGLM-6B), with a lightweight querying Transformer (Q-Former). We also introduce a parallel high-resolution feature extractor module and a coronary calcification evaluation loss to simultaneously encode fine-grained 3D features and constrain the accuracy of report generation. We compared our model with six state-of-the-art MRG methods on a clinical dataset with 118 subjects, comprising 453 paired 3D CTA images and radiology reports. Experimental results with extensive ablations show the efficacy of our C2RG. Codes will be open-sourced after the conference. Zhiyu Ye, Bang Yang, Shibin Wu, Hancong Wang, Hairong Zheng, Tong Zhang 0017 |
BIBM | 8 |
| 2024 | Blind Proximal Diffusion Model for Joint Image and Sensitivity Estimation in Parallel MRI
Xing Li 0027, Yan Yang 0007, Hairong Zheng, Zongben Xu |
MICCAI (7) | 3 |
| 2024 | Swin-UMamba: Mamba-Based UNet with ImageNet-Based Pretraining
Jiarun Liu, Hao Yang 0026, Yan Xi, Lequan Yu, Cheng Li 0008, Yong Liang 0001, Guangming Shi, Yizhou Yu, Shaoting Zhang 0001, Hairong Zheng, Shanshan Wang 0002 |
MICCAI (9) | 11 |
| 2024 | An improved medical image segmentation framework with Channel-Height-Width-Spatial attention moduleabstractThis paper presents an improved version of the U-Net segmentation framework for medical image segmentation, called CHWS-UNet. To build the proposed framework CHWS-UNet, we first develop a novel lightweight channel attention module called LCAM, based on which we further propose the Channel-Height-Width-Spatial (CHWS) attention module for channel, height, width, and spatial dimension-level feature refinement. Our CHWS-UNet is constructed by integrating the proposed CHWS attention modules into the shortcut paths between the encoder and the decoder stem. To justify the effectiveness of the proposed modules and networks, we then carried out extensive experiments on four public medical image datasets, including BUSI, ISIC2017, ISIC2018, PH and a proprietary uterus lesion ultrasound dataset from Shenzhen Maternity and Child Healthcare Hospital. The results show that the proposed attention module can significantly improve the performance of baseline models, even on small medical image datasets, without introducing noticeable parameters and computational costs. Further, the proposed segmentation framework can achieve promising performance compared to edge-cutting frameworks. The code can be found at CHWS-UNet. Wenjia Guo, Yanqing Kong, Yudong Zhang 0001, Hairong Zheng, Shengli Li 0001 |
Eng. Appl. Artif. Intell. | 9 |
| 2024 | Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement
Weijian Huang, Cheng Li 0008, Hao Yang 0026, Jiarun Liu, Yong Liang 0001, Hairong Zheng, Shanshan Wang 0010 |
Medical Image Anal. | 6 |
| 2024 | ISP-IRLNet: Joint optimization of interpretable sampler and implicit regularization learning network for accerlerated MRI
Xing Li 0027, Yan Yang 0007, Hairong Zheng, Zongben Xu |
Pattern Recognit. | 3 |
| 2024 | Accurate Whole-Brain Image Enhancement for Low-Dose Integrated PET/MR Imaging Through Spatial Brain TransformationabstractPositron emission tomography/magnetic resonance imaging (PET/MRI) systems can provide precise anatomical and functional information with exceptional sensitivity and accuracy for neurological disorder detection. Nevertheless, the radiation exposure risks and economic costs of radiopharmaceuticals may pose significant burdens on patients. To mitigate image quality degradation during low-dose PET imaging, we proposed a novel 3D network equipped with a spatial brain transform (SBF) module for low-dose whole-brain PET and MR images to synthesize high-quality PET images. The FreeSurfer toolkit was applied to derive the spatial brain anatomical alignment information, which was then fused with low-dose PET and MR features through the SBF module. Moreover, several deep learning methods were employed as comparison measures to evaluate the model performance, with the peak signal-to-noise ratio (PSNR), structural similarity (SSIM) and Pearson correlation coefficient (PCC) serving as quantitative metrics. Both the visual results and quantitative results illustrated the effectiveness of our approach. The obtained PSNR and SSIM were 41.96 ± 4.91 dB (p < 0.01) and 0.9654 ± 0.0215 (p < 0.01), which achieved a 19% and 20% improvement, respectively, compared to the original low-dose brain PET images. The volume of interest (VOI) analysis of brain regions such as the left thalamus (PCC = 0.959) also showed that the proposed method could achieve a more accurate standardized uptake value (SUV) distribution while preserving the details of brain structures. In future works, we hope to apply our method to other multimodal systems, such as PET/CT, to assist clinical brain disease diagnosis and treatment. Zhenxing Huang, Yaping Wu, Yun Dong 0002, Yongfeng Yang, Hairong Zheng, Dong Liang 0001, Zhanli Hu |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | MMCA-NET: A Multimodal Cross Attention Transformer Network for Nasopharyngeal Carcinoma Tumor Segmentation Based on a Total-Body PET/CT SystemabstractNasopharyngeal carcinoma (NPC) is a malignant tumor primarily treated by radiotherapy. Accurate delineation of the target tumor is essential for improving the effectiveness of radiotherapy. However, the segmentation performance of current models is unsatisfactory due to poor boundaries, large-scale tumor volume variation, and the labor-intensive nature of manual delineation for radiotherapy. In this paper, MMCA-Net, a novel segmentation network for NPC using PET/CT images that incorporates an innovative multimodal cross attention transformer (MCA-Transformer) and a modified U-Net architecture, is introduced to enhance modal fusion by leveraging cross-attention mechanisms between CT and PET data. Our method, tested against ten algorithms via fivefold cross-validation on samples from Sun Yat-sen University Cancer Center and the public HECKTOR dataset, consistently topped all four evaluation metrics with average Dice similarity coefficients of 0.815 and 0.7944, respectively. Furthermore, ablation experiments were conducted to demonstrate the superiority of our method over multiple baseline and variant techniques. The proposed method has promising potential for application in other tasks. Zhenxing Huang, Si Tang, Chuanli Cheng, Yongfeng Yang, Hairong Zheng, Dong Liang 0001, Zhanli Hu |
IEEE J. Biomed. Health Informatics | 10 |
| 2024 | High-Frequency Space Diffusion Model for Accelerated MRIabstractDiffusion models with continuous stochastic differential equations (SDEs) have shown superior performances in image generation. It can serve as a deep generative prior to solving the inverse problem in magnetic resonance (MR) reconstruction. However, low-frequency regions of k -space data are typically fully sampled in fast MR imaging, while existing diffusion models are performed throughout the entire image or k -space, inevitably introducing uncertainty in the reconstruction of low-frequency regions. Additionally, existing diffusion models often demand substantial iterations to converge, resulting in time-consuming reconstructions. To address these challenges, we propose a novel SDE tailored specifically for MR reconstruction with the diffusion process in high-frequency space (referred to as HFS-SDE). This approach ensures determinism in the fully sampled low-frequency regions and accelerates the sampling procedure of reverse diffusion. Experiments conducted on the publicly available fastMRI dataset demonstrate that the proposed HFS-SDE method outperforms traditional parallel imaging methods, supervised deep learning, and existing diffusion models in terms of reconstruction accuracy and stability. The fast convergence properties are also confirmed through theoretical and experimental validation. Our code and weights are available at https://github.com/Aboriginer/HFS-SDE. Chentao Cao, Zhuo-Xu Cui, Yue Wang 0119, Shaonan Liu, Taijin Chen, Hairong Zheng, Dong Liang 0001, Yanjie Zhu |
IEEE Trans. Medical Imaging | 6 |
| 2024 | OIF-Net: An Optical Flow Registration-Based PET/MR Cross-Modal Interactive Fusion Network for Low-Count Brain PET Image DenoisingabstractThe short frames of low-count positron emission tomography (PET) images generally cause high levels of statistical noise. Thus, improving the quality of low-count images by using image postprocessing algorithms to achieve better clinical diagnoses has attracted widespread attention in the medical imaging community. Most existing deep learning-based low-count PET image enhancement methods have achieved satisfying results, however, few of them focus on denoising low-count PET images with the magnetic resonance (MR) image modality as guidance. The prior context features contained in MR images can provide abundant and complementary information for single low-count PET image denoising, especially in ultralow-count (2.5%) cases. To this end, we propose a novel two-stream dual PET/MR cross-modal interactive fusion network with an optical flow pre-alignment module, namely, OIF-Net. Specifically, the learnable optical flow registration module enables the spatial manipulation of MR imaging inputs within the network without any extra training supervision. Registered MR images fundamentally solve the problem of feature misalignment in the multimodal fusion stage, which greatly benefits the subsequent denoising process. In addition, we design a spatial-channel feature enhancement module (SC-FEM) that considers the interactive impacts of multiple modalities and provides additional information flexibility in both the spatial and channel dimensions. Furthermore, instead of simply concatenating two extracted features from these two modalities as an intermediate fusion method, the proposed cross-modal feature fusion module (CM-FFM) adopts cross-attention at multiple feature levels and greatly improves the two modalities' feature fusion procedure. Extensive experimental assessments conducted on real clinical datasets, as well as an independent clinical testing dataset, demonstrate that the proposed OIF-Net outperforms the state-of-the-art methods. Minghan Fu, Na Zhang 0001, Zhenxing Huang, Jianmin Yuan, Yongfeng Yang, Hairong Zheng, Dong Liang 0001, Fang-Xiang Wu, Zhanli Hu |
IEEE Trans. Medical Imaging | 9 |
| 2024 | Noise-Generating and Imaging Mechanism Inspired Implicit Regularization Learning Network for Low Dose CT ReconstrutionabstractLow-dose computed tomography (LDCT) helps to reduce radiation risks in CT scanning while maintaining image quality, which involves a consistent pursuit of lower incident rays and higher reconstruction performance. Although deep learning approaches have achieved encouraging success in LDCT reconstruction, most of them treat the task as a general inverse problem in either the image domain or the dual (sinogram and image) domains. Such frameworks have not considered the original noise generation of the projection data and suffer from limited performance improvement for the LDCT task. In this paper, we propose a novel reconstruction model based on noise-generating and imaging mechanism in full-domain, which fully considers the statistical properties of intrinsic noises in LDCT and prior information in sinogram and image domains. To solve the model, we propose an optimization algorithm based on the proximal gradient technique. Specifically, we derive the approximate solutions of the integer programming problem on the projection data theoretically. Instead of hand-crafting the sinogram and image regularizers, we propose to unroll the optimization algorithm to be a deep network. The network implicitly learns the proximal operators of sinogram and image regularizers with two deep neural networks, providing a more interpretable and effective reconstruction procedure. Numerical results demonstrate our proposed method improvements of > 2.9 dB in peak signal to noise ratio, > 1.4% promotion in structural similarity metric, and > 9 HU decrements in root mean square error over current state-of-the-art LDCT methods. Xing Li 0027, Kaili Jing, Yan Yang 0007, Jianhua Ma 0001, Hairong Zheng, Zongben Xu |
IEEE Trans. Medical Imaging | 6 |
| 2024 | Super Resolution Dual-Energy Cone-Beam CT Imaging With Dual-Layer Flat-Panel DetectorabstractIn flat-panel detector (FPD) based cone-beam computed tomography (CBCT) imaging, the native receptor array is usually binned into a smaller matrix size. By doing so, the signal readout speed could be increased by 4-9 times at the expense of a spatial resolution loss of 50%-67%. Clearly, such manipulation poses a key bottleneck in generating high spatial and high temporal resolution CBCT images at the same time. In addition, the conventional FPD is also difficult in generating dual-energy CBCT images. In this paper, we propose an innovative super resolution dual-energy CBCT imaging method, named as suRi, based on dual-layer FPD (DL-FPD) to overcome these aforementioned difficulties at once. With suRi, specifically, a 1D or 2D sub-pixel (half pixel in this study) shifted binning is applied instead of the conventionally aligned binning to double the spatial sampling rate during the dual-energy data acquisition. As a result, the suRi approach provides a new strategy to enable high spatial resolution CBCT imaging while at high readout speed. Moreover, a penalized likelihood material decomposition algorithm is developed to directly reconstruct the high resolution bases from these dual-energy CBCT projections containing sub-pixel shifts. Numerical and physical experiments are performed to validate this newly developed suRi method with phantoms and biological specimen. Results demonstrate that suRi can significantly improve the spatial resolution of the CBCT image. We believe this developed suRi method would greatly enhance the imaging performance of the DL-FPD based dual-energy CBCT systems in future. Ting Su 0004, Jiongtao Zhu, Yuhang Tan, Dong Zeng, Jinchuan Guo, Hairong Zheng, Jianhua Ma 0001, Dong Liang 0001, Yongshuai Ge |
IEEE Trans. Medical Imaging | 8 |
| 2024 | Non-Invasive Quantification of the Brain [¹⁸F]FDG-PET Using Inferred Blood Input Function Learned From Total-Body Data With Physical ConstraintabstractFull quantification of brain PET requires the blood input function (IF), which is traditionally achieved through an invasive and time-consuming arterial catheter procedure, making it unfeasible for clinical routine. This study presents a deep learning based method to estimate the input function (DLIF) for a dynamic brain FDG scan. A long short-term memory combined with a fully connected network was used. The dataset for training was generated from 85 total-body dynamic scans obtained on a uEXPLORER scanner. Time-activity curves from 8 brain regions and the carotid served as the input of the model, and labelled IF was generated from the ascending aorta defined on CT image. We emphasize the goodness-of-fitting of kinetic modeling as an additional physical loss to reduce the bias and the need for large training samples. DLIF was evaluated together with existing methods in terms of RMSE, area under the curve, regional and parametric image quantifications. The results revealed that the proposed model can generate IFs that closer to the reference ones in terms of shape and amplitude compared with the IFs generated using existing methods. All regional kinetic parameters calculated using DLIF agreed with reference values, with the correlation coefficient being 0.961 (0.913) and relative bias being 1.68±8.74% (0.37±4.93%) for [Formula: see text] ( [Formula: see text]. In terms of the visual appearance and quantification, parametric images were also highly identical to the reference images. In conclusion, our experiments indicate that a trained model can infer an image-derived IF from dynamic brain PET data, which enables subsequent reliable kinetic modeling. Yaping Wu, Zeheng Xia, Dong Liang 0001, Hairong Zheng, Yongfeng Yang, Shanshan Wang 0002, Tao Sun 0025 |
IEEE Trans. Medical Imaging | 9 |
| 2024 | Deep Generalized Learning Model for PET Image ReconstructionabstractLow-count positron emission tomography (PET) imaging is challenging because of the ill-posedness of this inverse problem. Previous studies have demonstrated that deep learning (DL) holds promise for achieving improved low-count PET image quality. However, almost all data-driven DL methods suffer from fine structure degradation and blurring effects after denoising. Incorporating DL into the traditional iterative optimization model can effectively improve its image quality and recover fine structures, but little research has considered the full relaxation of the model, resulting in the performance of this hybrid model not being sufficiently exploited. In this paper, we propose a learning framework that deeply integrates DL and an alternating direction of multipliers method (ADMM)-based iterative optimization model. The innovative feature of this method is that we break the inherent forms of the fidelity operators and use neural networks to process them. The regularization term is deeply generalized. The proposed method is evaluated on simulated data and real data. Both the qualitative and quantitative results show that our proposed neural network method can outperform partial operator expansion-based neural network methods, neural network denoising methods and traditional methods. Qiyang Zhang 0002, Yumo Zhao, Debin Hu, Fuxiao Shi, Shuangliang Cao, Yongfeng Yang, Xin Liu 0053, Hairong Zheng, Dong Liang 0001, Zhanli Hu |
IEEE Trans. Medical Imaging | 12 |
| 2023 | Adaptive weighted curvature-based active contour for ultrasonic and 3T/5T MR image segmentation
Zhi-Feng Pang, Mengxiao Geng, Yanru Zhou, Tieyong Zeng, Liyun Zheng, Na Zhang 0001, Dong Liang 0001, Hairong Zheng, Yongming Dai, Zhenxing Huang, Zhanli Hu |
Signal Process. | 9 |
| 2023 | PARCEL: Physics-Based Unsupervised Contrastive Representation Learning for Multi-Coil MR ImagingabstractWith the successful application of deep learning to magnetic resonance (MR) imaging, parallel imaging techniques based on neural networks have attracted wide attention. However, in the absence of high-quality, fully sampled datasets for training, the performance of these methods is limited. And the interpretability of models is not strong enough. To tackle this issue, this paper proposes a Physics-bAsed unsupeRvised Contrastive rEpresentation Learning (PARCEL) method to speed up parallel MR imaging. Specifically, PARCEL has a parallel framework to contrastively learn two branches of model-based unrolling networks from augmented undersampled multi-coil k-space data. A sophisticated co-training loss with three essential components has been designed to guide the two networks in capturing the inherent features and representations for MR images. And the final MR image is reconstructed with the trained contrastive networks. PARCEL was evaluated on two vivo datasets and compared to five state-of-the-art methods. The results show that PARCEL is able to learn essential representations for accurate MR reconstruction without relying on fully sampled datasets. The code will be made available at https://github.com/ternencewu123/PARCEL. Shanshan Wang 0002, Ruoyou Wu, Cheng Li 0008, Ziyao Zhang 0003, Qiegen Liu, Yan Xi, Hairong Zheng |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2023 | A Two-Branch Neural Network for Short-Axis PET Image Quality EnhancementabstractThe axial field of view (FOV) is a key factor that affects the quality of PET images. Due to hardware FOV restrictions, conventional short-axis PET scanners with FOVs of 20 to 35 cm can acquire only low-quality PET (LQ-PET) images in fast scanning times (2-3 minutes). To overcome hardware restrictions and improve PET image quality for better clinical diagnoses, several deep learning-based algorithms have been proposed. However, these approaches use simple convolution layers with residual learning and local attention, which insufficiently extract and fuse long-range contextual information. To this end, we propose a novel two-branch network architecture with swin transformer units and graph convolution operation, namely SW-GCN. The proposed SW-GCN provides additional spatial- and channel-wise flexibility to handle different types of input information flow. Specifically, considering the high computational cost of calculating self-attention weights in full-size PET images, in our designed spatial adaptive branch, we take the self-attention mechanism within each local partition window and introduce global information interactions between nonoverlapping windows by shifting operations to prevent the aforementioned problem. In addition, the convolutional network structure considers the information in each channel equally during the feature extraction process. In our designed channel adaptive branch, we use a Watts Strogatz topology structure to connect each feature map to only its most relevant features in each graph convolutional layer, substantially reducing information redundancy. Moreover, ensemble learning is adopted in our SW-GCN for mapping distinct features from the two well-designed branches to the enhanced PET images. We carried out extensive experiments on three single-bed position scans for 386 patients. The test results demonstrate that our proposed SW-GCN approach outperforms state-of-the-art methods in both quantitative and qualitative evaluations. Minghan Fu, Yaping Wu, Na Zhang 0001, Yongfeng Yang, Fang-Xiang Wu, Hairong Zheng, Dong Liang 0001, Zhanli Hu |
IEEE J. Biomed. Health Informatics | 10 |
| 2021 | Self-supervised Learning for MRI Reconstruction with a Parallel Network Training Framework
Cheng Li 0008, Haifeng Wang 0003, Qiegen Liu, Hairong Zheng, Shanshan Wang 0002 |
MICCAI (6) | 5 |
| 2021 | FaNet: fast assessment network for the novel coronavirus (COVID-19) pneumonia based on 3D CT imaging and clinical symptoms
Zhenxing Huang, Xinfeng Liu, Rongpin Wang, Mudan Zhang, Xianchun Zeng, Jun Liu 0080, Yongfeng Yang, Xin Liu 0053, Hairong Zheng, Dong Liang 0001, Zhanli Hu |
Appl. Intell. | 9 |
| 2021 | Considering anatomical prior information for low-dose CT image enhancement using attribute-augmented Wasserstein generative adversarial networks
Zhenxing Huang, Xinfeng Liu, Rongpin Wang, Jincai Chen, Ping Lu 0006, Qiyang Zhang 0002, Changhui Jiang, Yongfeng Yang, Xin Liu 0053, Hairong Zheng, Dong Liang 0001, Zhanli Hu |
Neurocomputing | 10 |
| 2021 | Ultrasound image reconstruction from plane wave radio-frequency data by self-supervised deep neural networkabstractImage reconstruction from radio-frequency (RF) data is crucial for ultrafast plane wave ultrasound (PWUS) imaging. Compared with the traditional delay-and-sum (DAS) method based on relatively imprecise assumptions, sparse regularization (SR) method directly solves the inverse problem of image reconstruction and has presented significant improvement in the image quality when the frame rate remains high. However, the computational complexity of SR is too high for practical implementation, which is inherently associated with its iterative process. In this work, a deep neural network (DNN), which is trained with an incorporated loss function including sparse regularization terms, is proposed to reconstruct PWUS images from RF data with significantly reduced computational time. It is remarkable that, a self-supervised learning scheme, in which the RF data are utilized as both the inputs and the labels during the training process, is employed to overcome the lack of the "ideal" ultrasound images as the labels for DNN. In addition, it has been also verified that the trained network can be used on the RF data obtained with steered plane waves (PWs), and thus the image quality can be further improved with coherent compounding. Using simulation data, the proposed method has significantly shorter reconstruction time (∼10 ms) than the conventional SR method (∼1-5 mins), with comparable spatial resolution and 1.5-dB higher contrast-to-noise ratio (CNR). Besides, the proposed method with single PW can achieve higher CNR than DAS with 75 PWs in reconstruction of in-vivo images of human carotid arteries. Jingke Zhang, Qiong He, Yang Xiao 0012, Hairong Zheng, Congzhi Wang, Jianwen Luo 0001 |
Medical Image Anal. | 4 |
| 2021 | MRI Based Radiomics Approach With Deep Learning for Prediction of Vessel Invasion in Early-Stage Cervical CancerabstractThis article aims to build deep learning-based radiomic methods in differentiating vessel invasion from non-vessel invasion in cervical cancer with multi-parametric MRI data. A set of 1,070 dynamic T1 contrast-enhanced (DCE-T1) and 986 T2 weighted imaging (T2WI) MRI images from 167 early-stage cervical cancer patients (January 2014 - August 2018) were used to train and validate deep learning models. Predictive performances were evaluated using receiver operating characteristic (ROC) curve and confusion matrix analysis, with the DCE-T1 showing more discriminative results than T2WI MRI. By adopting an attention ensemble learning strategy that integrates both MRI sequences, the highest average area was obtained under the ROC curve (AUC) of 0.911 (Sensitivity = 0.881 and Specificity = 0.752). The superior performances in this article, when compared to existing radiomic methods, indicate that a wealth of deep learning-based radiomics could be developed to aid radiologists in preoperatively predicting vessel invasion in cervical cancer patients. Xiran Jiang, Yangyang Kan, Shijie Chang, Xianzheng Sha, Hairong Zheng, Yahong Luo, Shanshan Wang 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2021 | Multi-View Mammographic Density Classification by Dilated and Attention-Guided Residual LearningabstractBreast density is widely adopted to reflect the likelihood of early breast cancer development. Existing methods of mammographic density classification either require steps of manual operations or achieve only moderate classification accuracy due to the limited model capacity. In this study, we present a radiomics approach based on dilated and attention-guided residual learning for the task of mammographic density classification. The proposed method was instantiated with two datasets, one clinical dataset and one publicly available dataset, and classification accuracies of 88.7 and 70.0 percent were obtained, respectively. Although the classification accuracy of the public dataset was lower than the clinical dataset, which was very likely related to the dataset size, our proposed model still achieved a better performance than the naive residual networks and several recently published deep learning-based approaches. Furthermore, we designed a multi-stream network architecture specifically targeting at analyzing the multi-view mammograms. Utilizing the clinical dataset, we validated that multi-view inputs were beneficial to the breast density classification task with an increase of at least 2.0 percent in accuracy and the different views lead to different model classification capacities. Our method has a great potential to be further developed and applied in computer-aided diagnosis systems. Our code is available at https://github.com/lich0031/Mammographic_Density_Classification. Cheng Li 0008, Jingxu Xu, Qiegen Liu, Yongjin Zhou 0002, Lisha Mou, Zuhui Pu, Yong Xia 0001, Hairong Zheng, Shanshan Wang 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2021 | Learning a Deep CNN Denoising Approach Using Anatomical Prior Information Implemented With Attention Mechanism for Low-Dose CT Imaging on Clinical Patient Data From Multiple Anatomical SitesabstractDose reduction in computed tomography (CT) has gained considerable attention in clinical applications because it decreases radiation risks. However, a lower dose generates noise in low-dose computed tomography (LDCT) images. Previous deep learning (DL)-based works have investigated ways to improve diagnostic performance to address this ill-posed problem. However, most of them disregard the anatomical differences among different human body sites in constructing the mapping function between LDCT images and their high-resolution normal-dose CT (NDCT) counterparts. In this article, we propose a novel deep convolutional neural network (CNN) denoising approach by introducing information of the anatomical prior. Instead of designing multiple networks for each independent human body anatomical site, a unified network framework is employed to process anatomical information. The anatomical prior is represented as a pattern of weights of the features extracted from the corresponding LDCT image in an anatomical prior fusion module. To promote diversity in the contextual information, a spatial attention fusion mechanism is introduced to capture many local regions of interest in the attention fusion module. Although many network parameters are saved, the experimental results demonstrate that our method, which incorporates anatomical prior information, is effective in denoising LDCT images. Furthermore, the anatomical prior fusion module could be conveniently integrated into other DL-based methods and avails the performance improvement on multiple anatomical data. Zhenxing Huang, Xinfeng Liu, Rongpin Wang, Zixiang Chen, Yongfeng Yang, Xin Liu 0053, Hairong Zheng, Dong Liang 0001, Zhanli Hu |
IEEE J. Biomed. Health Informatics | 7 |
| 2021 | A Coarse-to-Fine Deformable Transformation Framework for Unsupervised Multi-Contrast MR Image Registration with Dual Consistency ConstraintabstractMulti-contrast magnetic resonance (MR) image registration is useful in the clinic to achieve fast and accurate imaging-based disease diagnosis and treatment planning. Nevertheless, the efficiency and performance of the existing registration algorithms can still be improved. In this paper, we propose a novel unsupervised learning-based framework to achieve accurate and efficient multi-contrast MR image registration. Specifically, an end-to-end coarse-to-fine network architecture consisting of affine and deformable transformations is designed to improve the robustness and achieve end-to-end registration. Furthermore, a dual consistency constraint and a new prior knowledge-based loss function are developed to enhance the registration performances. The proposed method has been evaluated on a clinical dataset containing 555 cases, and encouraging performances have been achieved. Compared to the commonly utilized registration methods, including VoxelMorph, SyN, and LT-Net, the proposed method achieves better registration performance with a Dice score of 0.8397± 0.0756 in identifying stroke lesions. With regards to the registration speed, our method is about 10 times faster than the most competitive method of SyN (Affine) when testing on a CPU. Moreover, we prove that our method can still perform well on more challenging tasks with lacking scanning information data, showing the high robustness for the clinical application. Weijian Huang, Hao Yang 0026, Xinfeng Liu, Cheng Li 0008, Ian Zhang 0002, Rongpin Wang, Hairong Zheng, Shanshan Wang 0002 |
IEEE Trans. Medical Imaging | 7 |
| 2021 | Learned Low-Rank Priors in Dynamic MR ImagingabstractDeep learning methods have achieved attractive performance in dynamic MR cine imaging. However, most of these methods are driven only by the sparse prior of MR images, while the important low-rank (LR) prior of dynamic MR cine images is not explored, which may limit further improvements in dynamic MR reconstruction. In this paper, a learned singular value thresholding (Learned-SVT) operator is proposed to explore low-rank priors in dynamic MR imaging to obtain improved reconstruction results. In particular, we put forward a model-based unrolling sparse and low-rank network for dynamic MR imaging, dubbed as SLR-Net. SLR-Net is defined over a deep network flow graph, which is unrolled from the iterative procedures in the iterative shrinkage-thresholding algorithm (ISTA) for optimizing a sparse and LR-based dynamic MRI model. Experimental results on a single-coil scenario show that the proposed SLR-Net can further improve the state-of-the-art compressed sensing (CS) methods and sparsity-driven deep learning-based methods with strong robustness to different undersampling patterns, both qualitatively and quantitatively. Besides, SLR-Net has been extended to a multi-coil scenario, and achieved excellent reconstruction results compared with a sparsity-driven multi-coil deep learning-based method under a high acceleration. Prospective reconstruction results on an open real-time dataset further demonstrate the capability and flexibility of the proposed method on real-time scenarios. Ziwen Ke, Wenqi Huang 0003, Zhuo-Xu Cui, Sen Jia 0005, Haifeng Wang 0003, Xin Liu 0053, Hairong Zheng, Leslie Ying, Yanjie Zhu, Dong Liang 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2021 | Accelerated 3D bSSFP Using a Modified Wave-CAIPI Technique With Truncated Wave GradientsabstractThe Wave Controlled Aliasing In Parallel Imaging (Wave-CAIPI) technique manifests great potential to highly accelerate three-dimensional (3D) balanced steady-state free precession (bSSFP) through substantially reducing the geometric factor (g-factor) and aliasing artifacts of image reconstruction. However, severe banding artifacts appear in bSSFP imaging due to unbalanced gradients with nonzero 0thmoment applied by the conventional Wave-CAIPI technique. In this study, we propose a 3D Wave-bSSFP scheme that adopts truncated wave gradients with zero 0thmoment to avoid introducing additional banding artifacts and to maintain the advantages of wave encoding. The simulation results indicate that the number of wave cycles that are truncated and different options of applying wave gradients affect both the g-factor reduction and image quality, but the influence is limited. In phantom experiments, the proposed technique shows similar acceleration performance as the conventional Wave-CAIPI technique and effectively eliminates its introduced banding artifacts. Additionally, Wave-bSSFP obtains up to $12\times $ retrospective acceleration at 0.8 mm isotropic resolution in in vivo 3D brain experiments and is superior to the state-of-the-art Controlled Aliasing In Parallel Imaging Results IN Higher Acceleration (CAIPIRINHA) technique, according to both visual validation and quantitative analysis. Moreover, in vivo 3D spine and abdomen imaging demonstrate the potential clinical applications of Wave-bSSFP with fast acquisition speed, improved isotropic resolution and fine image quality. Shi Su, Zhilang Qiu, Caiyun Shi, Liwen Wan, Yanjie Zhu, Ye Li 0011, Xin Liu 0053, Hairong Zheng, Dong Liang 0001, Haifeng Wang 0003 |
IEEE Trans. Medical Imaging | 9 |
| 2020 | A Magnetic Resonance-Guided Focused Ultrasound Neuromodulation System With a Whole Brain Coil Array for Nonhuman Primates at 3 TabstractThe phased-array radio frequency (RF) coil plays a vital role in magnetic resonance-guided focused ultrasound (MRgFUS) neuromodulation studies, where accurate brain functional stimulations and neural circuit observations are required. Although various designs of phased-array coils have been reported, few are suitable for ultrasound stimulations. In this study, an MRgFUS neuromodulation system comprised of a whole brain coverage non-human primate (NHP) RF coil and an MRI-compatible ultrasound device was developed. When compared to a single loop coil, the NHP coil provided up to a 50% increase in the signal-to-noise ratio (SNR) in the brain and acquired better anatomical image-quality. The NHP coil also demonstrated the ability to achieve higher spatial resolution and reduce distortion in echo-planer imaging (EPI). Ultrasound beam characteristics and transcranial magnetic resonance acoustic radiation force (MR-ARF) were measured for simulated positions, and calculated B0 maps were employed to establish MRI-compatibility. The differences between focused off and on ultrasound techniques were measured using SNR, g-factors, and temporal SNR (tSNR) analyses and all deviations were under 2.3%. The EPI images quality and stable tSNR demonstrated the suitability of the MRgFUS neuromodulation system to conduct functional MRI studies. Last, the time course of the blood oxygen level dependent (BOLD) signal of posterior cingulate cortex in a focused ultrasound neuromodulation study was detected and repeated with MR thermometry. Ye Li 0011, Jo Lee, Xiaojing Long, Yangzi Qiao, Teng Ma 0004, Xiaoliang Zhang 0001, Hairong Zheng |
IEEE Trans. Medical Imaging | 9 |
| 2020 | Ultrafast Endoscopic Ultrasonography With Circular ArrayabstractRapid development of ultrafast ultrasound imaging has led to novel medical ultrasound applications, including shear wave elastography and super-resolution vascular imaging. However, these have yet to incorporate endoscopic ultrasonography (EUS) with a circular array, which provides a wider view in the alimentary canal than traditional linear and convex arrays. A coherent diverging wave compounding (CDWC) imaging method was proposed for ultrafast EUS imaging and implemented on a custom circular array. In CDWC, virtual acoustic point sources are allocated and virtually insonified diverging waves from each source are achieved by adjusting all circular array elements' emission time delays. Diverging waves emitted from different virtual sources are coherently compounded, generating synthetic transmit focusing at every location in the image plane. As the field of view of the circular array is centrally symmetric, all virtual sources are equidistantly distributed on a concentric circle of radius r . To achieve the highest frame rate possible with image quality comparable to that obtained with the traditional multi-focus imaging method, the effects of various radii r and virtual source quantities on the compounded image quality were theoretically analyzed and experimentally verified. Simulation, phantom, and ex-vivo experiments were conducted with an 8 MHz, 124-element circular array, with a 5.35 mm radius. When 16 virtual sources were used with r=1.605 mm, image quality comparable to that obtained with the multi-focus approach was achieved at a frame rate of 1000 frames/s. This demonstrates the feasibility of the proposed ultrafast EUS imaging method and promotes further development of multi-functional EUS devices. Qingyuan Tan, Congzhi Wang, Jiamei Liu, Jiqing Huang, Yongchuan Li, Yang Xiao 0012, Gui-Song Xia, Teng Ma 0004, Hairong Zheng |
IEEE Trans. Medical Imaging | 9 |
| 2019 | Learning Cross-Modal Deep Representations for Multi-Modal MR Image Segmentation
Cheng Li 0008, Zaiyi Liu, Hairong Zheng, Shanshan Wang 0002 |
MICCAI (2) | 5 |
| 2019 | CLCI-Net: Cross-Level Fusion and Context Inference Networks for Lesion Segmentation of Chronic Stroke
Hao Yang 0026, Weijian Huang, Kehan Qi, Cheng Li 0008, Xinfeng Liu, Hairong Zheng, Shanshan Wang 0002 |
MICCAI (3) | 7 |
| 2019 | A Dual-Mode Imaging Catheter for Intravascular Ultrasound ApplicationabstractBoth the morphological anatomy and functional parameters such as flow speed of the artery provide valuable information for the evaluation of cardiovascular diseases. Direct measurement of the arterial wall can be achieved by intravascular optical/ultrasound imaging methods, and however, no functional data are acquired with these methods. Fractional flow reserve and Doppler wire have been used to assess the blood flow information, but do not provide cross-sectional images of the artery. This paper is the first to design and fabricate a dual-mode imaging catheter that contains a forward-looking ultrasonic transducer and a side-looking ultrasonic transducer together in one catheter. This dual-mode catheter not only provides morphological information about the artery, but also a precise measurement of functional flow. The data indicate that the proposed catheter can be used to acquire multiple parameters of the artery with a one-time procedure. This novel one-catheter approach could be used for the functional diagnosis of atherosclerotic arteries. Jiehan Hong, Yanyan Yu, Zhiqiang Zhang 0008, Yaocai Huang, Peitian Mu, Hairong Zheng, Weibao Qiu |
IEEE Trans. Medical Imaging | 8 |
| 2018 | Local and Non-local Deep Feature Fusion for Malignancy Characterization of Hepatocellular Carcinoma
Tianyou Dou, Hairong Zheng, Wu Zhou 0002 |
MICCAI (4) | 3 |
| 2018 | A Dedicated 36-Channel Receive Array for Fetal MRI at 3TabstractDue to a lack of fetal imaging coils, the standard commercial abdominal coil is often used for fetal imaging, the performance of which is limited by its insufficient coverage, element number, and Signal-to-noise ratio (SNR). In this paper, a dedicated 36-channel coil array, of which size can best fit the body sizes of pregnancy gestation from 20 to 37+ weeks, was designed for fetal imaging at 3T. SNR with full phase encoding and G-factor denoted as noise amplification for parallel imaging were quantitatively evaluated by phantom studies. Compared with a commercial abdominal coil array, the proposed 36-channel fetal array provides not only SNR improvements in full phase encoding (with 10% in the region where the whole fetal body was located, and up to 40% in the edge region where the fetal brain and heart may appear) but also an augmented parallel imaging capability and remarkable SNR improvements at high acceleration factors. Qiaoyan Chen, Guoxi Xie, Jo Lee, Shi Su, Dong Liang 0001, Xiaoliang Zhang 0001, Xin Liu 0053, Ye Li 0011, Hairong Zheng |
IEEE Trans. Medical Imaging | 12 |
| 2018 | Learning Joint-Sparse Codes for Calibration-Free Parallel MR ImagingabstractThe integration of compressed sensing and parallel imaging (CS-PI) has shown an increased popularity in recent years to accelerate magnetic resonance (MR) imaging. Among them, calibration-free techniques have presented encouraging performances due to its capability in robustly handling the sensitivity information. Unfortunately, existing calibration-free methods have only explored joint-sparsity with direct analysis transform projections. To further exploit joint-sparsity and improve reconstruction accuracy, this paper proposes to Learn joINt-sparse coDes for caliBration-free parallEl mR imaGing (LINDBERG) by modeling the parallel MR imaging problem as an - - minimization objective with an norm constraining data fidelity, Frobenius norm enforcing sparse representation error and the mixed norm triggering joint sparsity across multichannels. A corresponding algorithm has been developed to alternatively update the sparse representation, sensitivity encoded images and K-space data. Then, the final image is produced as the square root of sum of squares of all channel images. Experimental results on both physical phantom and in vivo data sets show that the proposed method is comparable and even superior to state-of-the-art CS-PI reconstruction approaches. Specifically, LINDBERG has presented strong capability in suppressing noise and artifacts while reconstructing MR images from highly undersampled multichannel measurements. Shanshan Wang 0002, Sha Tan, Qiegen Liu, Leslie Ying, Taohui Xiao, Xin Liu 0053, Hairong Zheng, Dong Liang 0001 |
IEEE Trans. Medical Imaging | 9 |
| 2017 | Malignancy characterization of hepatocellular carcinoma using hybrid texture and deep featuresabstractMalignancy of hepatocellular carcinoma (HCC) is significant to establish a therapeutic strategy preoperatively for liver cancer and is one of critical issues that influence recurrence and patient survival. Recently, quantitative texture feature of HCC in arterial phase of Contrast-enhanced MR has been shown to be promising for malignancy characterization of HCC. However, such texture feature is low-level, which is usually insufficient to capture the complicated characteristics of HCC. In this work, we propose a systematic method to automatically extract deep feature from the arterial phase of Contrast-enhanced MR using convolution neural network (CNN) in order to characterize malignancy of HCC. Specifically, we resample each 3D tumor in three orthogonal views (Axial, Coronal and Sagittal) independently to increase training sets, and train one CNN for each view to generate its corresponding deep feature. We investigate a multi-kernel feature fusion method that can fuse deep features derived from three views or fuse deep feature and texture feature in a kernel space. Our experimental results demonstrate several interesting conclusions: (1) deep feature significantly outperforms previous texture feature for malignancy characterization of HCC, (2) fusion of deep feature and texture feature yields best results for malignancy characterization of HCC. Qiyao Wang, Yaoqin Xie, Hairong Zheng, Wu Zhou 0002 |
ICIP | 4 |
| 2017 | Development of a Mechanical Scanning Device With High-Frequency Ultrasound Transducer for Ultrasonic Capsule EndoscopyabstractWireless capsule endoscopy has opened a new era by enabling remote diagnostic assessment of the gastrointestinal tract in a painless procedure. Video capsule endoscopy is currently commercially available worldwide. However, it is limited to visualization of superficial tissue. Ultrasound (US) imaging is a complementary solution as it is capable of acquiring transmural information from the tissue wall. This paper presents a mechanical scanning device incorporating a high-frequency transducer specifically as a proof of concept for US capsule endoscopy (USCE), providing information that may usefully assist future research. A rotary solenoid-coil-based motor was employed to rotate the US transducer with sectional electronic control. A set of gears was used to convert the sectional rotation to circular rotation. A single-element focused US transducer with 39-MHz center frequency was used for high-resolution US imaging, connected to an imaging platform for pulse generation and image processing. Key parameters of US imaging for USCE applications were evaluated. Wire phantom imaging and tissue phantom imaging have been conducted to evaluate the performance of the proposed method. A porcine small intestine specimen was also used for imaging evaluation in vitro. Test results demonstrate that the proposed device and rotation mechanism are able to offer good image resolution ( [Formula: see text]) of the lumen wall, and they, therefore, offer a viable basis for the fabrication of a USCE device. Xingying Wang, Vipin Seetohul, Zhiqiang Zhang 0008, Ming Qian, Zhehao Shi, Peitian Mu, Congzhi Wang, Qifa Zhou, Hairong Zheng, Sandy Cochran, Weibao Qiu |
IEEE Trans. Medical Imaging | 12 |
| 2011 | Geometric Calibration of a Micro-CT System and Performance for Insect ImagingabstractMicro-CT with a high spatial resolution in combination with computer-based-reconstruction techniques is considered a powerful tool for morphological study of insects. The quality of CT images crucially depends on the precise knowledge of the scan geometry of the micro-CT system. In this paper, we have proposed a method to calculate the deviation of rotating axis for compensating deficiency of existing methods. A practical application of this geometric calibration method of the micro-CT system for insect imaging is presented. We have performed the computer-simulation study and experimental study with our prototype micro-CT system. The results demonstrate that the proposed technique is accurate and robust. In addition, we have evaluated the imaging characteristics of the detector in terms of modulation-transfer function (MTF). Finally, insect imaging performance and image reconstruction from data acquired with different energies are presented. Zhanli Hu, Jianbao Gui, Junyan Rong, Qiyang Zhang 0002, Hairong Zheng |
IEEE Trans. Inf. Technol. Biomed. | 6 |
| 2010 | Real-Time Visualized Freehand 3D Ultrasound Reconstruction Based on GPUabstractVisualized freehand 3-D ultrasound reconstruction offers to image incremental reconstruction during acquisition and guide users to scan interactively for high-quality volumes. We originally used the graphics processing unit (GPU) to develop a visualized reconstruction algorithm that achieves real-time level. Each newly acquired image was transferred to the memory of the GPU and inserted into the reconstruction volume on the GPU. The partially reconstructed volume was then rendered using GPU-based incremental ray casting. After visualized reconstruction, hole-filling was performed on the GPU to fill remaining empty voxels in the reconstruction volume. We examine the real-time nature of the algorithm using in vitro and in vivo datasets. The algorithm can image incremental reconstruction at speed of 26-58 frames/s and complete 3-D imaging in the acquisition time for the conventional freehand 3-D ultrasound. Yakang Dai, Jie Tian 0001, Di Dong, Guorui Yan, Hairong Zheng |
IEEE Trans. Inf. Technol. Biomed. | 5 |