EDBT 2026 Demo / reviewers in the wild / expert
Dinggang Shen
dblp:14/4383 · also Ding-Gang Shen
· DBLP profile ↗
733ranked-venue papers
26as first author
276since 2021 · last 2026
0000-0002-7934-5698ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 581 · 12 first-author · 225 since 2021Graphics, computer vision, multimedia, augmented reality and games · 352 · 7 first-author · 89 since 2021Artificial intelligence and machine learning · 125 · 15 first-author · 42 since 2021Databases, data management, data science and information retrieval · 4Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semi-supervised Fetal Brain Parcellation via Hierarchical Learning Framework
Kai Zhang 0039, Fangmei Zhu, Zhongxiang Ding, Geng Chen 0001, Dinggang Shen |
Medical Image Anal. | 6 |
| 2026 | A content-aware variable-rate framework for pathology learned image compression (PathoLIC)
Yonghao Li, Zhenhui Li, Jing Ke, Dinggang Shen |
Medical Image Anal. | 8 |
| 2026 | LMGDM: A Lesion-aware Mutual Guidance Diffusion Model with attenuation prior constraint for self-attenuation correction of whole-body PET
Shengjun Li, Kaicong Sun, Caiwen Jiang, Zaixin Ou, Ruilong Dan, Qianjin Feng 0003, Dinggang Shen |
Medical Image Anal. | 7 |
| 2026 | UniSurf: Universal lifespan cortical surface reconstruction
Zifeng Lian, Jiameng Liu, Xiaoye Li, Han Zhang 0002, Zhiming Cui 0001, Feng Shi 0001, Dinggang Shen |
Medical Image Anal. | 9 |
| 2026 | A hierarchical prompt and prototype learning framework for brain disorder classification
Kaicong Sun, Yaping Wu, Weilin Zhou, Haoyue Yuan, Xintong Wu, Yichu He, Qingxia Wu, Zeng-Yang Che, Yiqiang Zhan, Sean Zhou, Dijia Wu, Feng Shi 0001, Dinggang Shen |
Medical Image Anal. | 18 |
| 2026 | 3D vessel reconstruction from sparse-view dynamic DSA images via vessel probability guided attenuation learning
Huangxuan Zhao, Wenhui Qin, Zhenghong Zhou, Xinggang Wang, Wenping Wang 0001, Xiaochun Lai, Dinggang Shen, Zhiming Cui 0001 |
Medical Image Anal. | 8 |
| 2026 | Learning dual-scale context with overlap awareness for keypoint-driven partial-overlap medical image registration
Caiwen Jiang, Xiaosong Xiong, Kaicong Sun, Xiaohuan Cao, Dinggang Shen |
Medical Image Anal. | 7 |
| 2026 | HALO: High-frequency enhanced dose-aware diffusion model for arbitrary low-dose PET reconstruction
Caiwen Jiang, Kaicong Sun, Zhiming Cui 0001, Dinggang Shen |
Medical Image Anal. | 5 |
| 2026 | Two-stage robust 3D CTA-2D DSA alignment via vascular-aware rigid and pyramid-based hierarchical non-rigid registration
Xiaosong Xiong, Caiwen Jiang, Han Wu 0007, Xiao Zhang 0028, Yanli Song, Jiayin Zhang, Dijia Wu, Dinggang Shen |
Medical Image Anal. | 11 |
| 2026 | Multi-organ guided diagnosis of mild cognitive impairment via hierarchical alignment and knowledge distillation
Shilun Zhao, Kaicong Sun, Shuwei Bai, Weilin Zhou, Jiangtao Liang, Zhongxiang Ding, Han Zhang 0002, Dinggang Shen |
Medical Image Anal. | 11 |
| 2026 | MGCM: Multi-modal graph convolutional mamba for cancer survival prediction
Yilun Li, Dinggang Shen, Yan Wang 0015 |
Pattern Recognit. | 3 |
| 2026 | Unlocking shared-specific features of multi-modal brain graphs for accurate psychiatric diagnosis
Geng Chen 0001, Xuyun Wen, Lifang Wei, Han Zhang 0002, Dinggang Shen |
Pattern Recognit. | 6 |
| 2026 | FLEX-MoCo: Flexible MRI motion correction using motion recognition and adaptive routing
Feng Li 0039, Zhenrong Shen 0001, Jiangdong Cai, Rongrong Xie, Han Zhang 0002, Dinggang Shen, Feng Shi 0001, Qian Wang 0001 |
Pattern Recognit. | 7 |
| 2026 | Transferring ultrahigh-field representations for intensity-guided brain segmentation of low-field magnetic resonance imagingabstractUltrahigh-field (UHF) magnetic resonance imaging (MRI), 7T MRI, provides superior anatomical details of internal brain structures thanks to its enhanced signal-to-noise ratio and susceptibility-induced contrast. However, the widespread use of 7T MRI is limited by its high cost and lower accessibility compared to low-field (LF) MRI. This study proposes a SegUHF that systematically fuses the input LF MRI feature representations with the inferred 7T-like feature representations for brain image segmentation tasks in a 7T-absent environment. Specifically, our proposed adaptive fusion module within the SegUHF aggregates 7T-like features derived from the LF image using a pre-trained network and then refines them to effectively assimilate UHF guidance into LF image features. Using intensity-guided features obtained from such aggregation and assimilation, segmentation models can recognize subtle structural representations that are usually difficult to identify when relying only on LF features. Beyond these advantages, this strategy can be seamlessly utilized by modulating the contrast of LF features in alignment with UHF guidance, even when employing arbitrary segmentation models. Extensive experiments demonstrated that our method outperformed all baselines in both brain tissue and whole-brain segmentation, while also showcasing adaptability and scalability across various models and tasks. Code is available at https://github.com/ku-milab/UHF-guided_segmentation Kwanseok Oh, Da-Woon Heo, Dinggang Shen, Heung-Il Suk |
Pattern Recognit. | 4 |
| 2026 | HCRT: Hybrid network with correlation-aware region transformer for breast tumor segmentation in DCE-MRI
Lei Zheng 0017, Yuzhong Zhang, Tao Zhou 0002, Lei Zhou 0003, Dinggang Shen |
Pattern Recognit. | 7 |
| 2026 | MGTP: Multi-Granularity Textual Prompts for Low-Dose Brain PET Image Denoising via Adversarial Diffusion ModelabstractPositron emission tomography (PET) is an advanced nuclear imaging technique and has been widely applied in clinic. However, radiation risks associated with standard-dose PET imaging raise health concerns, whereas the quality of low-dose PET images fails to meet clinical requirements. To reduce the tracer dose while maintaining image quality, it is of great interest to estimate high-quality PET images from low-dose images. However, existing low-dose PET image denoising methods primarily focus on image data, overlooking crucial information in non-image textual data such as patients' clinical tabular and textual descriptions of general image quality. This neglect can lead to subpar denoising quality with inaccurate contexts and poor details. To address these problems, in this paper, we propose Multi-Granularity Textual Prompts, namely MGTP, to denoise low-dose PET images via an adversarial diffusion model. Different from prior methods that rely solely on image conditioning, our MGTP innovatively introduces textual prompts spanning diverse granularities to capture both high-level semantic-related contexts and low-level degradation-related details. To harmonize multi-granularity textual prompts with low-dose PET images, we design a Cross-Modality Selective Conditioning (CMSC) module, which prioritizes semantic- and detail-relevant information while eliminating irrelevant components. The resulting features are fed into diffusion model as conditions, enforcing a more controlled diffusion process. In addition, we develop a Masked Prompt Reconstruction Network (MPR-Net) to enhance the preservation of semantics and details in denoised images, mitigating distortions brought by the random noise in the diffusion process. Experiments on clinical PET data show that our method achieves the state-of-the-art performance. Xinyi Zeng, Pinxian Zeng, Bo Liu 0113, Xi Wu 0004, Deng Xiong, Jiliu Zhou, Yan Wang 0015, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 9 |
| 2026 | Super-Resolution Reconstruction of Fetal Brain MRI With Multi-View Interpolation Weight LearningabstractSuper-resolution reconstruction (SRR) of isotropic fetal brain MR images is critical for prenatal examinations but is hindered by fetal motion and misalignment of thick-slice scans. To address these challenges comprehensively, we introduce an innovative deep learning model, namely 3D-WISE, a 3D Weighted Interpolation for Super-resolution Estimation of fetal brain MRI. The model generates high-quality isotropic fetal brain MR images by learning the interpolation weights to correct misalignments between slices and volumes. These misalignments are estimated by extracting deep features from multiple motion-corrupted stacks. Specifically, 3D-WISE incorporates two key components: (1) a weight learning module for multi-view interpolation and (2) a feature extraction module guided by multi-type attention mechanisms. The weight learning module first maps motion-corrupted thick-slice stacks into latent feature spaces. The resulting features are then fed to an implicit decoding block to estimate interpolation weights of the surrounding points for a given coordinate. We further enhance our approach by incorporating convolutional block attention and atlas-induced cross-attention mechanisms. Extensive experiments on two benchmark datasets show that our 3D-WISE achieves remarkably improved performance compared to the widely adopted registration-reconstruction framework. We also extend the experiments on anatomical structure reconstruction and achieve promising results, highlighting the significant potential of our 3D-WISE for fetal brain MR images SRR in clinical settings. DengQiang Jia, Kai Zhang 0039, Lingnan Kong, Fangmei Zhu, Zhongxiang Ding, Geng Chen 0001, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 8 |
| 2026 | Pancreas Segmentation With Multi-Phase Feature Aggregation and Modality Adaptive TransformerabstractAutomatic pancreas segmentation can facilitate diagnosis and treatment of pancreatic diseases. The combination of non-contrast, arterial, and venous phases of CT imaging can enhance differentiation of the pancreas from its surrounding structures. However, existing multimodal methods, which try to integrate the multimodal information in computer-aided pancreas segmentation, often overlook the inter-modal relationships and have a limited capability for information fusion. In this paper, we propose a multi-phase pancreas segmentation method for incorporating Feature Aggregation Module (FAM) and Modality Adaptive Transformer (MAT). Specifically, we use the venous phase as the primary modality, while the non-contrast and arterial phases serve as supplementary modalities, based on clinical prior knowledge. Our FAM integrates spatial information from the primary and supplementary modalities, while our MAT adaptively enhances feature representation and establishes long-range dependencies among modalities. Our method outperforms state-of-the-art techniques on a large scale dataset. Based on the segmented pancreas region, We further perform a downstream task focused on pancreatic volume calculation. The prediction accuracy is on par with manual segmentation, demonstrating effectiveness and potential application of our proposed method. Lulu Tan, Wenda Sheng, Wenbin Zou, Dengqiang Jia, Qianqian Chen 0002, Jing Sheng, Yangyang Qian, Qianjin Feng 0001, Zhuan Liao, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 11 |
| 2026 | Uncertainty-Guided Iterative Contrastive Fusion for Reliable Survival Prediction in Rectal CancerabstractIntegrating multimodal radiological images and clinical data is critical for survival prediction in rectal cancer. However, existing methods often lack sufficient consideration of 1) modality heterogeneity (caused by rectal peristalsis, noise artifacts, and missing modalities) and 2) site heterogeneity (caused by different imaging protocols and patient populations). These factors hinder the model from capturing reliable cross-modal relationships and adapting to distribution shifts across clinical sites. In this work, we propose UICSurv, a novel multimodal Survival prediction framework highlighted by Uncertainty-guided Iterative Contrastive fusion, to capture robust cross-site multimodal interactions while leveraging sample-level uncertainty to enhance fusion reliability. Specifically, UICSurv initializes a shared multimodal embedding and iteratively refines it by fusing each heterogeneous modality via the cross-attention mechanism. In each iteration, a novel Survival Contrastive Learning (SCL) strategy is designed to progressively enhance both cross-site alignment and survival discriminability of the multimodal embedding space. Moreover, we design an EvidenceHit module, which employs temporally consistent evidential learning to jointly estimate survival probabilities and uncertainty. The estimated uncertainty further guides the embedding alignment by reducing the interference of unreliable samples. All components operate synergistically within UICSurv to reinforce reliable survival prediction in rectal cancer. Extensive experiments on multimodal datasets of rectal cancer (collected from three sites) demonstrate the superiority of our method both in survival prediction and uncertainty estimation. The code is available open-source: https://github.com/ScorpioBao/UICSurv. Qingsen Bao, Lei Chen 0011, Kaicong Sun, Yiqun Sun, Fu Xiao 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 8 |
| 2026 | An Alignment and Imputation Network (AINet) for Breast Cancer Diagnosis With Multimodal Multi-View Ultrasound ImagesabstractRecently, numerous deep learning models have been proposed for breast cancer diagnosis using multimodal multi-view ultrasound images. However, their performance could be highly affected by overlooking interactions between different modalities and views. Moreover, existing methods struggle to handle cases where certain modalities or views are missing, which limits their clinical applications. To address these issues, we propose a novel Alignment and Imputation Network (AINet) by integrating 1) alignment and imputation pre-training, and 2) hierarchical fusion fine-tuning. Specifically, in the pre-training stage, cross-modal contrastive learning is employed to align features across different modalities, for effectively capturing inter-modal interactions. To simulate missing modality (view) scenarios, we randomly mask out features and then impute them by leveraging inter-modal and inter-view relationships. Following the clinical diagnosis procedure, the subsequent fine-tuning stage further incorporates modality-level and view-level fusion in a hierarchical manner. The proposed AINet is developed and evaluated on three datasets, comprising 15,223 subjects in total. Experimental results demonstrate that AINet significantly outperforms state-of-the-art methods, particularly in handling missing modalities (views). This highlights its robustness and potential for real-world clinical applications. Yonghao Li, Yiqun Sun, Yaling Chen, Shichong Zhou, Zhenhui Li, Xuejun Qian, Dinggang Shen |
IEEE Trans. Medical Imaging | 11 |
| 2026 | Multi-Contrast MRI Super-Resolution in Brain Tumors: Arbitrary-Scale Implicit Sampling and Unsupervised Fine-TuningabstractMulti-contrast magnetic resonance imaging (MRI) has important value in clinical applications because it can reflect comprehensive tissue characterization from anatomy and function to metabolism. Previous studies utilize abundant details in high-resolution (HR) reference (Ref) images to guide the super-resolution (SR) of low-resolution (LR) images, termed multi-contrast MRI SR. Yet, their clinical applications are hindered by: 1) discrepancies in MRI equipment and acquisition protocols across hospitals (which lead to gaps in data distribution), and 2) lack of paired LR and HR images in certain modalities for supervised training. Herein, we rethink multi-contrast MRI from a clinical perspective, and propose an implicit sampling and generation (ISG) network plus an unsupervised fine-tuning (FT) framework. Briefly, the ISG network possesses a powerful representation capability, enabling arbitrary-scale LR inputs and SR outputs. The fine-tuning framework, as a test-time training technique, allows models to be adapted to testing data. Experiments are conducted on two clinical datasets containing amide proton transfer weighted (APTw) images from tumor patients and fluid-attenuated inversion recovery (FLAIR) images from a 5T scanner, respectively. For tumor patients, our ISG+FT proves $4{\times }$ SR capacity in APTw metabolic images, receiving good recognition from radiologists. In both quantitative and qualitative evaluations, ISG+FT outperforms state-of-the-art baselines. The ablation and robustness study further demonstrate the rationality of ISG+FT. Overall, our proposed method shows considerable promise in clinical scenarios. Wenxuan Chen, Zhongsen Li, Shuai Wang 0048, Sirui Wu, Chuyu Liu, Yonghong Fan, Benqi Zhao, Zhuozhao Zheng, Dinggang Shen, Xiaolei Song |
IEEE Trans. Medical Imaging | 10 |
| 2026 | Hierarchical Contrastive Learning for Precise Whole-Body Anatomical Localization in PET/CT ImagingabstractAutomatic anatomical localization is critical for radiology report generation. While many studies focus on lesion detection and segmentation, anatomical localization-accurately describing lesion positions in radiology reports-has received less attention. Conventional segmentation-based methods are limited to organ-level localization and often fail in severe disease cases due to low segmentation accuracy. To address these limitations, we reformulate anatomical localization as an image-to-text retrieval task. Specifically, we propose a CLIP-based framework that aligns lesion image patches with anatomically descriptive text embeddings in a shared multimodal space. By projecting lesion features into the semantic space and retrieving the most relevant anatomical descriptions in a coarse-to-fine manner, our method achieves fine-grained lesion localization with high accuracy across the entire body. Our main contributions are as follows: (1) hierarchical anatomical retrieval, which organizes 387 locations into a two-level hierarchy, by retrieving from the first level of 124 coarse categories to narrow down the search space and reduce localization complexity; (2) augmented location descriptions, which integrate domain-specific anatomical knowledge for enhancing semantic representation and improving visual-text alignment; and (3) semi-hard negative sample mining, which improves training stability and discriminative learning by avoiding selecting the overly similar negative samples that may introduce label noise or semantic ambiguity. We validate our method on two whole-body PET/CT datasets, achieving an 84.13% localization accuracy on the internal test set and 80.42% on the external test set, with a per-lesion inference time of 34 ms. The proposed framework also demonstrated superior robustness in complex clinical cases compared to segmentation-based approaches. Yaozong Gao, Yiran Shu, Mingyang Yu 0009, Yanbo Chen 0003, Jingyu Liu 0002, Shaonan Zhong, Weifang Zhang, Yiqiang Zhan, Xiang Sean Zhou, Xinlu Wang, Meixin Zhao, Dinggang Shen |
IEEE Trans. Medical Imaging | 12 |
| 2026 | Positional Prompts-Enhanced Brain-Heart-Gut Interactions for Mild Cognitive Impairment DiagnosisabstractMild cognitive impairment (MCI) is the prodromal stage of dementia involving complex interactions between the brain and peripheral organs. Emerging evidence indicates that heart dysfunction and gut microbiota dysbiosis can contribute to MCI pathogenesis. Yet, these discoveries of cross-organ interactions have not been applied to assist MCI diagnosis. In this work, we propose a novel diagnostic framework that exploits the interactions of brain, heart, and gut using whole-body PET images to guide MCI diagnosis for scenarios when only brain MRI, PET, or PET&MRI are available. Specifically, we collected a multi-cohort, multi-modal dataset comprising 1,545 whole-body PET images, 6,010 brain MR images, and 2,446 brain PET images from eight data centers. Organ-specific image encoders are first pretrained for the brain, heart, and gut individually. Then, to effectively align and integrate brain, heart, and gut features, we introduce positional prompts to act as anatomical-level attention to highlight disease-relevant spatial regions, and further develop hierarchical Transformers to model brain-heart, brain-gut, and brain-heart-gut interactions. Finally, to achieve MCI diagnosis using only brain images, we transfer the above brain-heart-gut model to a brain-only model via an introduced multi-level knowledge distillation scheme, including sample-level contrastive distillation, group-level distribution alignment, and response-level supervision. Extensive experiments on multi-center data demonstrate the superiority of our method over the state-of-the-art methods by resorting to effective integration of heart and gut interactions for MCI diagnosis. Shilun Zhao, Shuwei Bai, Dengqiang Jia, Jiangtao Liang, Han Zhang 0002, Ya Zhang 0002, Zhongxiang Ding, Yin Xu 0001, Kaicong Sun, Dinggang Shen |
IEEE Trans. Medical Imaging | 12 |
| 2026 | BrainSMM: Lifespan Brain Segmentation Model With Metadata-Driven Prompt LearningabstractAccurate and automatic segmentation of lifespan brain MRI into regions of interest (ROIs) is crucial for studying brain development, aging, and early diagnosis of neurological diseases. Existing segmentation methods are often tailored to specific age groups, such as infants or adults, resulting in inconsistent performance when processing brain data from different age groups. To overcome this limitation, we introduce BrainSMM, a novel metadata-driven model that incorporates text-based prompts to guide representation learning in a segmentation backbone. These prompts, extracted via a pretrained image-text alignment model, encode valuable prior knowledge (e.g., age, scanner, gender) and are infused into the vision model to condition the features according to domain-specific contexts. We evaluate BrainSMM on a large-scale lifespan brain MRI dataset with 5,565 T1w MR images spanning multiple ages. Our approach achieves an average DSC of 94.59% for tissue segmentation (i.e., gray matter, white matter, and cerebrospinal fluid) and 86.34% for anatomical region segmentation (e.g., hippocampus, putamen, etc.) with corresponding average ASD of 0.20 mm and 0.75 mm, respectively. Notably, BrainSMM shows strong consistency in segmentation accuracy across all age groups and demonstrates improved anatomical detail preservation compared to baseline methods. Additionally, our metadata prompt technique is easily transferable and compatible with multiple backbone architectures, highlighting its adaptability. Overall, BrainSMM offers a robust, generalizable solution for lifespan brain MRI segmentation and lays the groundwork for enhanced clinical and developmental neuroimaging applications. Zihao Zhao 0002, Feng Shi 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 5 |
| 2026 | uBrainSurf: Unified Curvature-Aware Deformation Framework for Lifespan Brain Cortical Surface ReconstructionabstractAccurate and automated reconstruction of cortical surfaces across the human lifespan is essential for studying brain development, aging, and the early diagnosis of neurological disorders. However, traditional neuroimaging pipelines require hours per subject, limiting scalability. Existing deep learning methods typically target narrow age ranges, struggling to generalize due to substantial age-related anatomical variability. This leads to inaccurate quantification of cortical properties, such as curvature and cortical thickness, thereby undermining their potential as reliable biomarkers for routine clinical brain analysis. To address these challenges, we present uBrainSurf, a unified curvature-aware deformation framework for lifespan cortical surface reconstruction. Specifically, uBrainSurf learns a sequence of stationary velocity fields (SVFs) from volumetric MR images, gradually deforming a smooth template mesh to subject-specific white-matter and pial surfaces through a coarse-to-fine strategy. To enhance the reconstruction accuracy, we introduce an auxiliary curvature prediction branch that provides an anatomical prior, guiding the model to prioritize anatomically important regions. Furthermore, we propose a novel curvature-driven loss function that encourages consistency between the curvatures of corresponding points on predicted and target surfaces, ensuring the reconstructed surfaces are directly suitable for downstream analyses. The uBrainSurf is evaluated on a large-scale brain dataset comprising 2,132 subjects spanning 0-100 years. Experimental results demonstrate that uBrainSurf achieves superior performance and generalizability while being several orders of magnitude faster than traditional pipelines. Our code is available at https://github.com/TL9792/CCF. Feng Shi 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 5 |
| 2026 | Semi-Supervised Landmark Tracking in Echocardiography Video via Spatial-Temporal Co-Training and Perception-Aware AttentionabstractPrecise landmark annotation in cardiac ultrasound images is fundamental for quantitative cardiac health assessment. However, the time-intensive nature of manual annotation typically constrains clinicians to annotate only selected key frames, limiting comprehensive temporal analysis capabilities. While recent automated landmark detection methods have demonstrated success for key-frame analysis, they fail to effectively utilize the intrinsic temporal information across cardiac sequence. To bridge this gap, we present SemiEchoTracker, a novel semi-supervised framework that enables comprehensive landmark tracking throughout echocardiography sequences while requiring supervision only on key frames. Our framework introduces three key innovative strategies: 1) a co-training mechanism that enforces mutual consistency between spatial detection and temporal tracking, enabling accurate intermediate frame detection without additional annotations, 2) a guided DINOv2 pretraining strategy that is specially tailored for extracting fine-grained echocardiography-specific spatial features, and 3) a perception-aware spatial-temporal (PAST) attention module that efficiently captures inter- and intra-frame relationships in echocardiography videos. Extensive validation on three datasets across multiple cardiac views demonstrates that our method not only achieves state-of-the-art detection performance on the keyframes but also yields accurate frame-by-frame prediction, which is important for dynamic cardiac analysis in clinicians. Han Wu 0007, Zhiming Cui 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 6 |
| 2026 | Enhancing Knee Disease Diagnosis via Multi-View Graph Representation With Multi-Task Pre-TrainingabstractMagnetic resonance imaging (MRI) is an indispensable tool for clinical knee examination, which often scans 2D stacked slices from multiple views. Radiologists typically locate lesion regions in one view, and then refer to other views to formulate a comprehensive diagnosis. However, existing computer-aided diagnosis methods fall short of identifying and fusing local regions in multi-view scans, leading to a decline in diagnostic performance and a heavy reliance on extensively annotated data. This paper introduces a novel framework that represents multi-view MRI scans as a knee graph, and conducts diagnosis using the proposed Knee Graph Network (KGNet). Moreover, KGNet is greatly enhanced by multi-task pre-training, which requires KGNet to reconstruct masked knee local patches and segment unmasked ones working alongside corresponding decoders. Experimental evaluations on public and in-house clinical datasets confirm that our framework outperforms existing approaches in diagnosing cartilage defects, anterior cruciate ligament tears, and knee abnormalities. In conclusion, our framework demonstrates the potential of enhancing knee disease diagnosis by representing multi-view MRI scans as a graph and employing multi-task pre-training in the graph network. The code is publicly available at https://github.com/zixuzhuang/KGNet. Zixu Zhuang, Dongdong Chen 0003, Sheng Wang 0014, Kai Xuan, Xiangyu Zhao 0003, Zhong Xue, Dinggang Shen, Lichi Zhang, Weiwu Yao, Qian Wang 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Tree-Diffusion: Octree-Based Conditional Diffusion Model for Small Bowel Skeleton Generation with Geometric Direction ModelingabstractAccurate 3D reconstruction of the small bowel skeleton is vital for understanding intestinal morphology, de-tecting structural abnormalities, and supporting diagnosis, yet limited resolution, organ adhesion, complex anatomy, and scarce annotations make continuous skeleton extraction from masks challenging. Voxel-based methods often struggle with the sparse topology and geometric directionality inherent in the small bowel skeleton, leading to inefficiency and high memory cost. To address these limitations, we propose a novel octree-based conditional diffusion model (i.e., Tree-Diffusion) that generates anatomically consistent small bowel skeletons guided by 3D segmentation masks. Specifically, we introduce two modules that captures structural priors from masks and topology characteristics from skeletons, ensuring cross-domain alignment and high-quality skeleton generation. Besides, we design a synthesis strategy to generate anatomically plausible skeleton-mask pairs, serving as topological priors to guide the diffusion model toward realis-tic structure predictions. To efficiently represent the elongated skeleton, we adopt an octree- based spatial encoding of hierarchical geometric features. Compared with baselines, our model achieves superior performance in anatomical fidelity, directional consistency, and inference efficiency. The code is available at: https://github.com/Small-Bowel-Skeleton-GenerationlCode Zhichao Liang, Dengqiang Jia, Yaofei Duan, Xinyu Xie, Kaicong Sun, Zhiming Cui 0001, Tao Tan 0002, Dinggang Shen |
BIBM | 9 |
| 2025 | Revolutionizing Disease Diagnosis with simultaneous functional PET/MR and Deeply Integrated Brain Metabolic, Hemodynamic, and Perfusion NetworksabstractSimultaneous functional PET/MR (sf-PET/MR) presents a cutting-edge multimodal neuroimaging technique. It provides an unprecedented opportunity for concurrently monitoring and integrating multifaceted brain networks built by spatiotemporally covaried metabolic activity, neural activity, and cerebral blood flow (perfusion). Albeit high scientific/clinical values, short in hardware accessibility of PET/MR hinders its applications, let alone modern AI-based PET/MR fusion models. Our objective is to develop a clinically feasible AI-based disease diagnosis model trained on comprehensive sf-PET/MR data with the power of, during inferencing, allowing single modality input (e.g., PET only) as well as enforcing multimodal-based accuracy. To this end, we propose MX-ARM, a multimodal MiXture-of-experts Alignment and Reconstruction Model. It is modality detachable and exchangeable, allocating different multi-layer perceptrons dynamically ("mixture of experts") through learnable weights to learn respective representations from different modalities. Such design will not sacrifice model performance in uni-modal situation. To fully exploit the inherent complex and nonlinear relation among modalities while producing fine-grained representations for uni-modal inference, a modal alignment module is utilized to line up a dominant modality (e.g., PET) with representations of auxiliary modalities (MR). We further adopt multimodal reconstruction to promote the quality of learned features. Experiments on precious multimodal sf-PET/MR data for Mild Cognitive Impairment diagnosis showcase the efficacy of MX-ARM toward clinically feasible precision medicine. Luoyu Wang, Yitian Tao, Siwei Liu 0013, Hongcheng Shi, Dinggang Shen, Han Zhang 0002 |
ICASSP | 7 |
| 2025 | ThicknessVAE: Learning a Lateral Prior for Clothed Human Body ReconstructionabstractSandwich-like structures have shown remarkable efficacy in clothed human reconstruction. However, these approaches often generate unrealistic side geometries due to inadequate handling of lateral regions. This paper addresses this limitation by incorporating the side geometry of clothed humans as a prior. We propose ThicknessVAE, a novel two-stage method that makes two key contributions: (1) We learn a prototype from point clouds for the lateral regions of clothed humans to extract common and detailed geometric features. (2) We utilize this prototype as a prior to transform geometric features into a thickness map associated with clothed human images, enabling refined normal integration for sandwich-like reconstruction methods. By seamlessly integrating our model into the sandwich-like reconstruction pipeline, we achieve highly realistic side views. Both qualitative and quantitative experiments demonstrate that our approach is comparable to state-of-the-art methods in terms of side-view realism. Xiaotao Wu, Zhaoxin Fan, Huiguang He, Dinggang Shen |
ICASSP | 4 |
| 2025 | BME2: A Plug-and-Play Bridge-Based Module for Misalignment Estimation and Elimination in Multi-scan Image Restoration
Wenxuan Chen, Caiwen Jiang, Xiaolei Song, Dinggang Shen |
MICCAI (13) | 4 |
| 2025 | A Semi-Supervised Knowledge Distillation Framework for Left Ventricle Segmentation and Landmark Detection in Echocardiograms
Yonghao Li, Han Wu 0007, Kaicong Sun, Dinggang Shen |
MICCAI (8) | 7 |
| 2025 | HiLa: Hierarchical Vision-Language Collaboration for Cancer Survival Prediction
Lu Wen, Yuchen Fei, Bo Liu 0113, Luping Zhou, Dinggang Shen, Yan Wang 0015 |
MICCAI (5) | 6 |
| 2025 | Predicting Alzheimer's Disease Progression Using a Regression-Based Survival Model with Longitudinal Data
Yiqun Sun, Jincheng Gu, Qinsen Bao, Feihong Liu, Dinggang Shen |
MICCAI (15) | 6 |
| 2025 | Wavelet-Driven Decoupling and Physics-Informed Mapping Network for Accelerated Multi-parametric MR Imaging
Ruilong Dan, Kaicong Sun, Minqiang Jia, Han Zhang 0002, Xiaopeng Zong, Dinggang Shen |
MICCAI (1) | 8 |
| 2025 | Leveraging Visual Prompt with Diffusion Adversarial Network for Radiotherapy Dose Prediction
Zhenghao Feng, Lu Wen, Xi Wu 0004, Jianghong Xiao, Xingchen Peng, Dinggang Shen, Yan Wang 0015 |
MICCAI (15) | 7 |
| 2025 | Sparsely Labeled fMRI Data Denoising with Meta-learning-Based Semi-supervised Domain Adaptation
Keun-Soo Heo, Ji-Wung Han, Soyeon Bak, Minjoo Lim, Bogyeong Kang, Weili Lin, Han Zhang 0002, Dinggang Shen, Tae-Eui Kam |
MICCAI (7) | 9 |
| 2025 | Task-Aligned fMRI Generation Model for Brain Disorder Diagnosis
Xiaotong Wu, Xiaocai Zhang, Haiteng Jiang, Weiwen Wu, Dinggang Shen, Jianjia Zhang |
MICCAI (12) | 6 |
| 2025 | Brain-Heart-Gut Guided Multi-constraint Knowledge Distillation for Early Alzheimer's Disease Diagnosis
Shilun Zhao, Shuwei Bai, Kai Zhang 0039, Yin Xu 0001, Ya Zhang 0002, Kaicong Sun, Dinggang Shen |
MICCAI (15) | 9 |
| 2025 | Adapting Foundation Model for Dental Caries Detection with Dual-View Co-training
Tao Luo 0010, Han Wu 0007, Dinggang Shen, Zhiming Cui 0001 |
MICCAI (16) | 4 |
| 2025 | MAST-Pro: Dynamic Mixture-of-Experts for Adaptive Segmentation of Pan-Tumors with Knowledge-Driven Prompts
Runqi Meng, Sifan Song, Pengfei Jin, Yiqun Sun, Yujin Oh, Xiang Li 0001, Quanzheng Li, Dinggang Shen |
MICCAI (16) | 12 |
| 2025 | A New Paradigm for Low-Dose PET/CT Reconstruction with Mamba-Powered Progressive Network and Physics-Informed Consistency
Caiwen Jiang, Zhiming Cui 0001, Dinggang Shen |
MICCAI (11) | 4 |
| 2025 | A Curvature-Guided Diffeomorphic Mesh Deformation Framework for Lifespan Brain Cortical Surface Reconstruction
Feng Shi 0001, Dinggang Shen |
MICCAI (1) | 4 |
| 2025 | Unisyn: A Generative Foundation Model for Universal Medical Image Synthesis Across MRI, CT and PET
Honglin Xiong, Kaicong Sun, Jiameng Liu, Yuanzhe He, Qian Wang 0001, Dinggang Shen |
MICCAI (3) | 9 |
| 2025 | Graph-Based Neighbor-Aware Network for Gaze-Supervised Medical Image Segmentation
Shaoxuan Wu, Jingkun Chen, Zhuo Jin, Peilin Zhang, Zhizezhang Gao, Jun Feng 0003, Xiao Zhang 0028, Dinggang Shen |
MICCAI (4) | 8 |
| 2025 | Tumor Segmentation with Heterogeneity Clustering in Non-Contrast Breast MRI
Xinyu Xie, Luyi Han, Yonghao Li, Yaofei Duan, Yue Sun 0001, Muzhen He, Tao Tan 0002, Dinggang Shen |
MICCAI (2) | 8 |
| 2025 | Location-Guided Automated Lesion Captioning in Whole-Body PET/CT Images
Mingyang Yu 0009, Yaozong Gao, Yiran Shu, Yanbo Chen 0003, Jingyu Liu 0002, Caiwen Jiang, Kaicong Sun, Zhiming Cui 0001, Weifang Zhang, Yiqiang Zhan, Xiang Sean Zhou, Shaonan Zhong, Xinlu Wang, Meixin Zhao, Dinggang Shen |
MICCAI (5) | 15 |
| 2025 | A$\upbeta $-PET Pattern Prediction via Graph Reconstruction-Aware Fusion (GRAF) of Functional and Structural Networks
Haoyue Yuan, Feihong Liu, Dinggang Shen |
MICCAI (12) | 4 |
| 2025 | MAK-GAN: Multi-level Adaptive Convolutional Kernels for Asymmetric Multi-modal PET Reconstruction
Xinyi Zeng, Pinxian Zeng, Yan Wang 0015, Luping Zhou, Caiwen Jiang, Han Zhang 0002, Dinggang Shen |
MICCAI (2) | 8 |
| 2025 | Neuro-AMS: Neuro-Informed Age-Aware and Medical Knowledge-Integrated Strategy for Diagnosis of Multiple Brain Disorders
Zhaoyu Qiu, Zehao Weng, Jinwei Kong, Feng Shi 0001, Dinggang Shen |
MICCAI (15) | 9 |
| 2025 | Draw Sketch, Draw Flesh: Whole-Body Computed Tomography from Any X-Ray Views
Yongsheng Pan, Yiwen Ye, Yanning Zhang 0001, Yong Xia 0001, Dinggang Shen |
Int. J. Comput. Vis. | 5 |
| 2025 | Guiding fusion of dynamic functional and effective connectivity in spatio-temporal graph neural network for brain disorder classification
Dongdong Chen 0003, Mengjun Liu, Sheng Wang 0014, Zheren Li, Lu Bai 0001, Qian Wang 0001, Dinggang Shen, Lichi Zhang |
Knowl. Based Syst. | 7 |
| 2025 | CLIK-Diffusion: Clinical Knowledge-informed Diffusion Model for Tooth Alignment
Yulong Dou, Han Wu 0007, Changjian Li 0001, Chen Wang 0054, Dinggang Shen, Zhiming Cui 0001 |
Medical Image Anal. | 7 |
| 2025 | A lung structure and function information-guided residual diffusion model for predicting idiopathic pulmonary fibrosis progression
Caiwen Jiang, Xiaodan Xing, Yang Nan 0002, Yingying Fang, Sheng Zhang 0024, Simon Walsh, Guang Yang 0006, Dinggang Shen |
Medical Image Anal. | 8 |
| 2025 | Learnable color space conversion and fusion for stain normalization in pathology images
Jing Ke, Yijin Zhou, Yiqing Shen 0003, Yi Guo 0001, Xiaodan Han, Dinggang Shen |
Medical Image Anal. | 7 |
| 2025 | Clinical knowledge-guided hybrid classification network for automatic periodontal disease diagnosis in X-ray image
Lanzhuju Mei, Zhiming Cui 0001, Yu Fang 0008, Yuan Liu 0025, Hongchang Lai, Maurizio Tonetti, Dinggang Shen |
Medical Image Anal. | 8 |
| 2025 | Predicting infant brain connectivity with federated multi-trajectory GNNs using scarce dataabstractThe understanding of the convoluted evolution of infant brain networks during the first postnatal year is pivotal for identifying the dynamics of early brain connectivity development. Thanks to the valuable insights into the brain's anatomy, existing deep learning frameworks focused on forecasting the brain evolution trajectory from a single baseline observation. While yielding remarkable results, they suffer from three major limitations. First, they lack the ability to generalize to multi-trajectory prediction tasks, where each graph trajectory corresponds to a particular imaging modality or connectivity type (e.g., T1-w MRI). Second, existing models require extensive training datasets to achieve satisfactory performance which are often challenging to obtain. Third, they do not efficiently utilize incomplete time series data. To address these limitations, we introduce FedGmTE-Net++, a federated graph-based multi-trajectory evolution network. Using the power of federation, we aggregate local learnings among diverse hospitals with limited datasets. As a result, we enhance the performance of each hospital's local generative model, while preserving data privacy. The three key innovations of FedGmTE-Net++ are: (i) presenting the first federated learning framework specifically designed for brain multi-trajectory evolution prediction in a data-scarce environment, (ii) incorporating an auxiliary regularizer in the local objective function to exploit all the longitudinal brain connectivity within the evolution trajectory and maximize data utilization, (iii) introducing a two-step imputation process, comprising a preliminary K-Nearest Neighbours based precompletion followed by an imputation refinement step that employs regressors to improve similarity scores and refine imputations. Our comprehensive experimental results showed the outperformance of FedGmTE-Net++ in brain multi-trajectory prediction from a single baseline graph in comparison with benchmark methods. Our source code is available at https://github.com/basiralab/FedGmTE-Net-plus. Michalis Pistos, Gang Li 0001, Weili Lin, Dinggang Shen, Islem Rekik |
Medical Image Anal. | 4 |
| 2025 | A topology-preserving three-stage framework for fully-connected coronary artery extractionabstractCoronary artery extraction is a crucial prerequisite for computer-aided diagnosis of coronary artery disease. Accurately extracting the complete coronary tree remains challenging due to several factors, including presence of thin distal vessels, tortuous topological structures, and insufficient contrast. These issues often result in over-segmentation and under-segmentation in current segmentation methods. To address these challenges, we propose a topology-preserving three-stage framework for fully-connected coronary artery extraction. This framework includes vessel segmentation, centerline reconnection, and missing vessel reconstruction. First, we introduce a new centerline enhanced loss in the segmentation process. Second, for the broken vessel segments, we further propose a regularized walk algorithm to integrate distance, probabilities predicted by a centerline classifier, and directional cosine similarity, for reconnecting the centerlines. Third, we apply implicit neural representation and implicit modeling, to reconstruct the geometric model of the missing vessels. Experimental results show that our proposed framework outperforms existing methods, achieving Dice scores of 88.53% and 85.07%, with Hausdorff Distances (HD) of 1.07 mm and 1.63 mm on ASOCA and PDSCA datasets, respectively. Code will be available at https://github.com/YH-Qiu/CorSegRec. Yuehui Qiu, Dandan Shan, Pei Dong, Dijia Wu, Xinnian Yang, Qingqi Hong, Dinggang Shen |
Medical Image Anal. | 8 |
| 2025 | Learning contrast and content representations for synthesizing magnetic resonance image of arbitrary contrast
Honglin Xiong, Zhenrong Shen 0001, Kaicong Sun, Yu Fang 0008, Dinggang Shen, Qian Wang 0001 |
Medical Image Anal. | 7 |
| 2025 | Learning better contrastive view from radiologist's gaze
Sheng Wang 0014, Zihao Zhao 0002, Zixu Zhuang, Xi Ouyang, Lichi Zhang, Zheren Li, Chong Ma 0004, Tianming Liu 0001, Dinggang Shen, Qian Wang 0001 |
Pattern Recognit. | 9 |
| 2025 | Dual-Domain Classification-Aided High-Quality PET Synthesis With Shared Information MaximizationabstractPositron emission tomography (PET) is widely applied in clinic for providing crucial diagnosis information. However, its inherent radiation exposure inevitably brings potential health risk for patient. To reduce radiation risk while also obtaining high-quality PET image, we plan to synthesize standard-dose PET (SPET) from low-dose PET (LPET). Since PET images can be represented in both projection domain and image domain (dual domains) emphasizing different information, considering dual domains in PET synthesis could contribute to better performance. In this way, we propose a novel dual-domain model for high-quality PET synthesis, named DCBi-GAN, by introducing a denoising network for the projection domain and an enhancing network for the image domain to effectively exploit dual-domain information. Concretely, the denoising network takes the LPET sinogram converted from LPET image to suppress noise and artifacts in the projection domain. Then, the enhancing network in the image domain takes the denoised LPET image (transferred back from the denoised sinogram) to enhance image quality. Notably, as LPET and SPET images come from the same subject, the abundant shared information between LPET and SPET can be used for boosting synthesis performance. Specially, we design a bi-directional contrastive generative adversarial network (GAN) to encourage maximal preservation of the shared information. Besides, we introduce a mild cognitive impairment (MCI) classification task to enhance clinical applicability of the synthesized PET. Evaluation on both Real Human Brain dataset and Phantom Brain dataset demonstrates effectiveness and superiority of our proposed model.Note to Practitioners—Positron emission tomography (PET) is a primary nuclear imaging technique for tumor detection and brain disorder diagnosis in the early stage of diseases, while the inherent radiation exposure inevitably raises concerns about potential health risk. This article proposes a novel PET image synthesis model to obtain clinically accepted PET image at low dose, namely DCBi-GAN, by taking account of the complementary multi-domain information and the modality shared content information, with a mild cognitive impairment (MCI) classification task to further boost clinical applicability of synthesized PET images. We experimentally validate the effectiveness of proposed DCBi-GAN on two datasets. Our proposed method could facilitate diagnosis and treatment of disease, to be used in the existing computer-aided medical systems. Yuchen Fei, Chen Zu, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015, Dinggang Shen |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | AugGPT: Leveraging ChatGPT for Text Data AugmentationabstractText data augmentation is an effective strategy for overcoming the challenge of limited sample sizes in many natural language processing (NLP) tasks. This challenge is especially prominent in the few-shot learning (FSL) scenario, where the data in the target domain is generally much scarcer and of lowered quality. A natural and widely used strategy to mitigate such challenges is to perform data augmentation to better capture data invariance and increase the sample size. However, current text data augmentation methods either can’t ensure the correct labeling of the generated data (lacking faithfulness), or can’t ensure sufficient diversity in the generated data (lacking compactness), or both. Inspired by the recent success of large language models (LLM), especially the development of ChatGPT, we propose a text data augmentation approach based on ChatGPT (named ”AugGPT”). AugGPT rephrases each sentence in the training samples into multiple conceptually similar but semantically different samples. The augmented samples can then be used in downstream model training. Experiment results on multiple few-shot learning text classification tasks show the superior performance of the proposed AugGPT approach over state-of-the-art text data augmentation methods in terms of testing accuracy and distribution of the augmented samples. Haixing Dai, Zhengliang Liu, Wenxiong Liao, Zihao Wu 0001, Lin Zhao 0004, Shaochen Xu, Fang Zeng, Wei Liu 0146, Ninghao Liu 0001, Sheng Li 0001, Dajiang Zhu, Hongmin Cai, Lichao Sun 0001, Quanzheng Li, Dinggang Shen, Tianming Liu 0001, Xiang Li 0001 |
IEEE Trans. Big Data | 17 |
| 2025 | Exploring the Trade-Offs: Unified Large Language Models vs Local Fine-Tuned Models for Highly-Specific Radiology NLI TaskabstractRecently, ChatGPT and GPT-4 have emerged and gained immense global attention due to their unparalleled performance in language processing. Despite demonstrating impressive capability in various open-domain tasks, their adequacy in highly specific fields like radiology remains untested. Radiology presents unique linguistic phenomena distinct from open-domain data due to its specificity and complexity. Assessing the performance of large language models (LLMs) in such specific domains is crucial not only for a thorough evaluation of their overall performance but also for providing valuable insights into future model design directions: whether model design should be generic or domain-specific. To this end, in this study, we evaluate the performance of ChatGPT/GPT-4 on a radiology natural language inference (NLI) task and compare it to other models fine-tuned specifically on task-related data samples. We also conduct a comprehensive investigation on ChatGPT/GPT-4’s reasoning ability by introducing varying levels of inference difficulty. Our results show that 1) ChatGPT and GPT-4 outperform other LLMs in the radiology NLI task and 2) other specifically fine-tuned Bert-based models require significant amounts of data samples to achieve comparable performance to ChatGPT/GPT-4. These findings not only demonstrate the feasibility and promise of constructing a generic model capable of addressing various tasks across different domains, but also highlight several key factors crucial for developing a unified model, particularly in a medical context, paving the way for future artificial general intelligence (AGI) systems. We release our code and data to the research community. Zihao Wu 0001, Lu Zhang 0050, Xiaowei Yu 0001, Zhengliang Liu, Lin Zhao 0004, Yiwei Li 0002, Haixing Dai, Chong Ma 0004, Gang Li 0001, Wei Liu 0146, Quanzheng Li, Dinggang Shen, Xiang Li 0001, Dajiang Zhu, Tianming Liu 0001 |
IEEE Trans. Big Data | 13 |
| 2025 | Multi-Modal Long-Short Distance Attention-Based Transformer-GAN for PET Reconstruction With Auxiliary MRIabstractTo obtain high-quality PET scans while minimizing potential radiation hazards for patients, various GAN-based methods have been developed to reconstruct high-quality standard-count PET (SPET) images from low-count PET (LPET) ones. While recent efforts try to integrate MRI or CT to enhance reconstruction in a multi-modal way, current architectures mainly face two limitations: 1) CNN backbones or simple Transformer bottleneck layers are insufficient for robust semantic understanding; and 2) the identical strategies for multi-modal feature extraction and fusion overlook each modality’s respective importance for the reconstruction task. In this work, we propose the Multi-modal Long-Short Distance Attention-based Transformer-GAN (MLSDA-GAN), a novel network combining 3D transformer and CNN architecture for PET image reconstruction. Specifically, to extract fine-grained features with a small number of parameters, our MLSDA-GAN integrates multi-scale convolution into the embedding part of the transformer. As for our multi-modal design, given the strong correlation between LPET and SPET in structural characteristics, we treat MRI as an auxiliary modality to LPET and achieve effective multi-modal extraction and fusion strategies. These strategies include 1) a PET-specific Self-attention Extraction (PSE) block for comprehensive feature extraction of the primary LPET and 2) a Multi-modality Cross-attention Fusion (MCF) block for effective multi-modal interaction and fusion, enabling us to more efficiently model both long- and short-range relationships in the corresponding feature extraction and fusion processes. Experiments demonstrate superiority of our method quantitatively and qualitatively. Code is available athttps://github.com/Aru321/MLSDA-GAN. Pinxian Zeng, Xinyi Zeng, Yan Wang 0015, Luping Zhou, Chen Zu, Xi Wu 0004, Jiliu Zhou, Dinggang Shen |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | Structure-Aware Brain Tissue Segmentation for Isointense Infant MRI Data Using Multi-Phase Multi-Scale Assistance NetworkabstractAccurate and automatic brain tissue segmentation is crucial for tracking brain development and diagnosing brain disorders. However, due to inherently ongoing myelination and maturation during the first postnatal year, the intensity distributions of gray matter and white matter in the infant brain MRI at the age of around 6 months old (a.k.a. isointense phase) are highly overlapped, which makes tissue segmentation very challenging, even for experts. To address this issue, in this study, we propose a multi-phase multi-scale assistance segmentation framework, which comprises a structure-preserved generative adversarial network (SPGAN) and a multi-phase multi-scale assisted segmentation network (MASN). SPGAN bi-directionally synthesizes isointense and adult-like data. The synthetic isointense data essentially augment the training dataset, combined with high-quality annotations transferred from its adult-like counterpart. By contrast, the synthetic adult-like data offers clear tissue structures and is concatenated with isointense data to serve as the input of MASN. In particular, MASN is designed with two-branch networks, which simultaneously segment tissues with two phases (isointense and adult-like) and two scales by also preserving their correspondences. We further propose a boundary refinement module to extract maximum gradients from local feature maps to indicate tissue boundaries, prompting MASN to focus more on boundaries where segmentation errors are prone to occur. Extensive experiments on the National Database for Autism Research and Baby Connectome Project datasets quantitatively and qualitatively demonstrate the superiority of our proposed framework compared with seven state-of-the-art methods. Jiameng Liu, Feihong Liu, Dong Nie, Yuning Gu, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Image-and-Label Conditioning Latent Diffusion Model: Synthesizing A$\beta$-PET From MRI for Detecting Amyloid StatusabstractDeposition of $\beta$-amyloid (A$\beta$), which is generally observed by A$\beta$-PET, is an important biomarker to evaluate subjects with early-onset dementia. However, acquisition of A$\beta$-PET usually suffers from high expense and radiation hazards, making A$\beta$-PET not commonly used as MRI. As A$\beta$-PET scans are only used to determine whether A$\beta$ deposition is positive or not, it is highly valuable to capture the underlying relationship between A$\beta$ deposition and other neuroimages (i.e., MRI) and detect amyloid status based on other neuroimages to reduce necessity of acquiring A$\beta$-PET. To this end, we propose an image-and-label conditioning latent diffusion model (IL-CLDM) to synthesize A$\beta$-PET scans from MRI scans by enhancing critical shared information to finally achieve MRI-based A$\beta$ classification. Specifically, two conditioning modules are introduced to enable IL-CLDM to implicitly learn joint image synthesis and diagnosis: 1) an image conditioning module, to extract meaningful features from source MRI scans to provide structural information, and 2) a label conditioning module, to guide the alignment of generated scans to the diagnosed label. Experiments on a clinical dataset of 510 subjects demonstrate that our proposed IL-CLDM achieves image quality superior to five widely used models, and our synthesized A$\beta$-PET scans (by IL-CLDM) can significantly help classification of A$\beta$ as positive or negative. Zaixin Ou, Yongsheng Pan, Qihao Guo, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Aleatoric-Uncertainty-Aware Maximum Intensity Projection-Based GAN for 7T-Like Generation From 3T TOF-MRAabstractTime-of-flight magnetic resonance angiography (TOF-MRA) is a prevalent vascular imaging technique for assessing cerebrovascular diseases. Compared to routine 3T TOF-MRA, 7T TOF-MRA provides vascular structures with a higher signal-to-noise ratio (SNR) and better vessel contrast, revealing greater vascular details. However, the inaccessibility of 7T scanners and specific physiological and technical concerns limit its clinical application. Therefore, we aimed to generate high-quality 7T-like TOF-MRA from 3T TOF-MRA. Considering the spatial sparsity of vessel signals, the visibility discrepancy of distal and small vessels between 3T and 7T images, and the subtle spatial misalignment between paired data, we proposed a novel aleatoric-uncertainty-aware maximum intensity projection-based generative adversarial network (AU-MIPGAN). In our method, we employed a knowledge distillation (KD) framework to incorporate multi-directional MIP information into the 3T-to-7T learning process to strengthen the learning of vessels and provide three-dimensional (3D) vascular morphological knowledge for the student model, facilitating accurate generation of vascular structures. Furthermore, we exploited AU modeling to compensate for the spatial misalignment between paired 3T and 7T images during the training procedure, which helped the model concentrate more on learning the intrinsic gap between 3T and 7T images. Qualitative and quantitative results demonstrated that the proposed AU-MIPGAN can achieve promising performance for 7T-like TOF-MRA generation. Yuxiang Dai, Zhang Shi, Ying-Hua Chu, Peixian Zhuang, Dinggang Shen, Chengyan Wang, He Wang 0016 |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Unified Model for Children's Brain Image Segmentation With Co-Registration Framework Guided by Longitudinal MRIabstractAccurate segmentation of brain structures is crucial for analyzing longitudinal changes in children's brains. However, existing methods are mostly based on models established at a single time-point due to difficulty in obtaining annotated data and dynamic variation of tissue intensity. The main problem with such approaches is that, when conducting longitudinal analysis, images from different time points are segmented by different models, leading to significant variation in estimating development trends. In this paper, we propose a novel unified model with co-registration framework to segment children's brain images covering neonates to preschoolers, which is formulated as two stages. First, to overcome the shortage of annotated data, we propose building gold-standard segmentation with co-registration framework guided by longitudinal data. Second, we construct a unified segmentation model tailored to brain images at 0-6 years old through the introduction of a convolutional network (named SE-VB-Net), which combines our previously proposed VB-Net with Squeeze-and-Excitation (SE) block. Moreover, different from existing methods that only require both T1- and T2-weighted MR images as inputs, our designed model also allows a single T1-weighted MR image as input. The proposed method is evaluated on the main dataset (320 longitudinal subjects with average 2 time-points) and two external datasets (10 cases with 6-month-old and 40 cases with 20-45 weeks, respectively). Results demonstrate that our proposed method achieves a high performance (>92%), even over a single time-point. This means that it is suitable for brain image analysis with large appearance variation, and largely broadens the application scenarios. Yichu He, Zehong Cao, Qianjin Feng 0003, Feng Shi 0001, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 8 |
| 2025 | A Modality-Flexible Framework for Alzheimer's Disease Diagnosis Following Clinical RoutineabstractDementia has high incidence among the elderly, and Alzheimer's disease (AD) is the most common dementia. The procedure of AD diagnosis in clinics usually follows a standard routine consisting of different phases, from acquiring non-imaging tabular data in the screening phase to MR imaging and ultimately to PET imaging. Most of the existing AD diagnosis studies are dedicated to a specific phase using either single or multi-modal data. In this paper, we introduce a modality-flexible classification framework, which is applicable for different AD diagnosis phases following the clinical routine. Specifically, our framework consists of three branches corresponding to three diagnosis phases: 1) a tabular branch using only tabular data for screening phase, 2) an MRI branch using both MRI and tabular data for uncertain cases in screening phase, and 3) ultimately a PET branch for the challenging cases using all the modalities including PET, MRI, and tabular data. To achieve effective fusion of imaging and non-imaging modalities, we introduce an image-tabular transformer block to adaptively scale and shift the image and tabular features according to modality importance determined by the network. The proposed framework is extensively validated on four cohorts containing 6495 subjects. Experiments demonstrate that our framework achieves superior diagnostic performance than the other representative methods across various AD diagnosis tasks, and shows promising performance for all the diagnosis phases, which exhibits great potential for clinical application. Yuanwang Zhang, Kaicong Sun, Qihao Guo, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Ethics of Foundation Models in Computational Pathology: Overview of Contemporary Issues and Future ImplicationsabstractArtificial intelligence (AI) has profoundly transformed our lives, reshaping industries and impacting nearly every aspect of society over the past few decades. It has recently become even more influential, primarily due to the rise of foundation models representing a new paradigm in AI development. These models, characterized by their large-scale training on vast datasets, have unique capabilities such as emergence and transference, enabling them to generalize across diverse tasks. Since their introduction, foundation models have been increasingly applied in fields such as autonomous driving, computer vision, marketing, finance, industrial robotics, and healthcare. Pathologists worldwide use computational methods to analyze diseases that profoundly impact human well-being, including cancer diagnosis and staging, genetic mutation prediction, and treatment and prognosis forecasting. In this article, we discuss how, despite the promise of foundation models in various applications, their development and application in computational pathology remain challenging due to inherent characteristics such as emergence, homogenization, hallucination, transference, compositionality, and explainability. While powerful, these traits introduce numerous ethical concerns and challenges, impacting safety and reliability, patient privacy, accountability, and equity and fairness in healthcare access. We examine these ethical issues, focusing on key concerns like algorithmic discrimination and misuse, accuracy, privacy breaches, transparency, public accessibility, and accountability. Furthermore, potential solutions to these challenges are analyzed, offering future perspectives on promoting the development and application of more ethical AI and foundation models in computational pathology. These insights aim to guide foundation models toward responsible integration of AI in healthcare. Rui Fei Du, Eduard Lloret Carbonell, Jiaxuan Huang, Xiaohang Wang 0015, Dinggang Shen, Jing Ke |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Geometry-Aware Attenuation Learning for Sparse-View CBCT ReconstructionabstractCone Beam Computed Tomography (CBCT) plays a vital role in clinical imaging. Traditional methods typically require hundreds of 2D X-ray projections to reconstruct a high-quality 3D CBCT image, leading to considerable radiation exposure. This has led to a growing interest in sparse-view CBCT reconstruction to reduce radiation doses. While recent advances, including deep learning and neural rendering algorithms, have made strides in this area, these methods either produce unsatisfactory results or suffer from time inefficiency of individual optimization. In this paper, we introduce a novel geometry-aware encoder-decoder framework to solve this problem. Our framework starts by encoding multi-view 2D features from various 2D X-ray projections with a 2D CNN encoder. Leveraging the geometry of CBCT scanning, it then back-projects the multi-view 2D features into the 3D space to formulate a comprehensive volumetric feature map, followed by a 3D CNN decoder to recover 3D CBCT image. Importantly, our approach respects the geometric relationship between 3D CBCT image and its 2D X-ray projections during feature back projection stage, and enjoys the prior knowledge learned from the data population. This ensures its adaptability in dealing with extremely sparse view inputs without individual training, such as scenarios with only 5 or 10 X-ray projections. Extensive evaluations on two simulated datasets and one real-world dataset demonstrate exceptional reconstruction quality and time efficiency of our method. Yu Fang 0008, Changjian Li 0001, Han Wu 0007, Yuan Liu 0025, Dinggang Shen, Zhiming Cui 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Transferring Adult-Like Phase Images for Robust Multi-View Isointense Infant Brain SegmentationabstractAccurate tissue segmentation of infant brain in magnetic resonance (MR) images is crucial for charting early brain development and identifying biomarkers. Due to ongoing myelination and maturation, in the isointense phase (6-9 months of age), the gray and white matters of infant brain exhibit similar intensity levels in MR images, posing significant challenges for tissue segmentation. Meanwhile, in the adult-like phase around 12 months of age, the MR images show high tissue contrast and can be easily segmented. In this paper, we propose to effectively exploit adult-like phase images to achieve robust multi-view isointense infant brain segmentation. Specifically, in one way, we transfer adult-like phase images to the isointense view, which have similar tissue contrast as the isointense phase images, and use the transferred images to train an isointense-view segmentation network. On the other way, we transfer isointense phase images to the adult-like view, which have enhanced tissue contrast, for training a segmentation network in the adult-like view. The segmentation networks of different views form a multi-path architecture that performs multi-view learning to further boost the segmentation performance. Since anatomy-preserving style transfer is key to the downstream segmentation task, we develop a Disentangled Cycle-consistent Adversarial Network (DCAN) with strong regularization terms to accurately transfer realistic tissue contrast between isointense and adult-like phase images while still maintaining their structural consistency. Experiments on both NDAR and iSeg-2019 datasets demonstrate a significant superior performance of our method over the state-of-the-art methods. Huabing Liu, Dengqiang Jia, Qian Wang 0001, Jun Xu 0019, Dinggang Shen |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Amyloid-β Deposition Prediction With Large Language Model Driven and Task-Oriented Learning of Brain Functional NetworksabstractAmyloid- positron emission tomography can reflect the Amyloid- protein deposition in the brain and thus serves as one of the golden standards for Alzheimer's disease (AD) diagnosis. However, its practical cost and high radioactivity hinder its application in large-scale early AD screening. Recent neuroscience studies suggest a strong association between changes in functional connectivity network (FCN) derived from functional MRI (fMRI), and deposition patterns of Amyloid- protein in the brain. This enables an FCN-based approach to assess the Amyloid- protein deposition with less expense and radioactivity. However, an effective FCN-based Amyloid- assessment remains lacking for practice. In this paper, we introduce a novel deep learning framework tailored for this task. Our framework comprises three innovative components: 1) a pre-trained Large Language Model Nodal Embedding Encoder, designed to extract task-related features from fMRI signals; 2) a task-oriented Hierarchical-order FCN Learning module, used to enhance the representation of complex correlations among different brain regions for improved prediction of Amyloid- deposition; and 3) task-feature consistency losses for promoting similarity between predicted and real Amyloid- values and ensuring effectiveness of predicted Amyloid- in downstream classification task. Experimental results show superiority of our method over several state-of-the-art FCN-based methods. Additionally, we identify crucial functional sub-networks for predicting Amyloid- depositions. The proposed method is anticipated to contribute valuable insights into the understanding of mechanisms of AD and its prevention. Mianxin Liu, Yuanwang Zhang, Yihui Guan, Qihao Guo, Dinggang Shen |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Guest Editorial Special Issue on Advancements in Foundation Models for Medical ImagingabstractPretrained on massive datasets, Foundation Models (FMs) are revolutionizing medical imaging by offering scalable and generalizable solutions to longstanding challenges. This Special Issue on Advancements in Foundation Models for Medical Imaging presents FM-related works that explore the potential of FMs to address data scarcity, domain shifts, and multimodal integration across a wide range of medical imaging tasks, including segmentation, diagnosis, reconstruction, and prognosis. The included papers also examine critical concerns such as interpretability, efficiency, benchmarking, and ethics in the adoption of FMs for medical imaging. Collectively, the articles in this Special Issue mark a significant step toward establishing FMs as a cornerstone of next-generation medical imaging AI. Tianming Liu 0001, Dinggang Shen, Jong Chul Ye, Marleen de Bruijne |
IEEE Trans. Medical Imaging | 2 |
| 2025 | MIP-Enhanced Uncertainty-Aware Network for Fast 7T Time-of-Flight MRA ReconstructionabstractTime-of-flight (TOF) magnetic resonance angiography (MRA) is the dominant non-contrast MR imaging method for visualizing intracranial vascular system. The employment of 7T MRI for TOF-MRA is of great interest due to its outstanding spatial resolution and vessel-tissue contrast. However, high-resolution 7T TOF-MRA is undesirably slow to acquire. Besides, due to complicated and thin structures of brain vessels, reliability of reconstructed vessels is of great importance. In this work, we propose an uncertainty-aware reconstruction model for accelerated 7T TOF-MRA, which combines the merits of deep unrolling and evidential deep learning, such that our model not only provides promising MRI reconstruction, but also supports uncertainty quantification within a single inference. Moreover, we propose a maximum intensity projection (MIP) loss for TOF-MRA reconstruction to improve the quality of MIP images. In the experiments, we have evaluated our model on a relatively large in-house multi-coil 7T TOF-MRA dataset extensively, showing promising superiority of our model compared to state-of-the-art models in terms of both TOF-MRA reconstruction and uncertainty quantification. Kaicong Sun, Caohui Duan, Dinggang Shen |
IEEE Trans. Medical Imaging | 4 |
| 2025 | 3D MedDiffusion: A 3D Medical Latent Diffusion Model for Controllable and High-Quality Medical Image GenerationabstractThe generation of medical images presents significant challenges due to their high-resolution and three-dimensional nature. Existing methods often yield suboptimal performance in generating high-quality 3D medical images, and there is currently no universal generative framework for medical imaging. In this paper, we introduce a 3D Medical Latent Diffusion (3D MedDiffusion) model for controllable, high-quality 3D medical image generation. 3D MedDiffusion incorporates a novel, highly efficient Patch-Volume Autoencoder that compresses medical images into latent space through patch-wise encoding and recovers back into image space through volume-wise decoding. Additionally, we design a new noise estimator to capture both local details and global structural information during diffusion denoising process. 3D MedDiffusion can generate fine-detailed, high-resolution images (up to ${512}\times {512}\times {512}$ ) and effectively adapt to various downstream tasks as it is trained on large-scale datasets covering CT and MRI modalities and different anatomical regions (from head to leg). Experimental results demonstrate that 3D MedDiffusion surpasses state-of-the-art methods in generative quality and exhibits strong generalizability across tasks such as sparse-view CT reconstruction, fast MRI reconstruction, and data augmentation for segmentationand classification. Source code and checkpoints are available at https://github.com/ShanghaiTech-IMPACT/3D-MedDiffusion. Haoshen Wang, Kaicong Sun, Dinggang Shen, Zhiming Cui 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Improving Self-Supervised Medical Image Pre-Training by Early Alignment With Human Eye Gaze InformationabstractAlignment between human knowledge and machine learning models is crucial for achieving efficient and interpretable AI systems. However, conventional self-supervised pre-training methods often suffer from low efficiency, as they do not incorporate human knowledge during the pre-training process and instead rely mainly on post-hoc alignment techniques. We propose Gaze Pre-Training (GzPT), a novel approach that introduces early alignment with human eye gaze information during the pre-training process to enhance both the learning efficiency and performance of self-supervised models. By leveraging contrastive learning to pull together images with similar gaze patterns, GzPT can effectively align the model with human attention during the pre-training. We demonstrate the effectiveness of our approach on three diverse medical image datasets, showing that GzPT can consistently outperform baseline methods and learn more meaningful and interpretable representations. Our findings also highlight the potential of incorporating human eye gaze as a form of passive knowledge to bridge the gap between human and machine learning in the self-supervised pre-training. Our code is available at Github. Sheng Wang 0014, Zihao Zhao 0002, Zhenrong Shen 0001, Bin Wang 0068, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Integrating Eye Tracking With Grouped Fusion Networks for Semantic Segmentation on Mammogram ImagesabstractMedical image segmentation has seen great progress in recent years, largely due to the development of deep neural networks. However, unlike in computer vision, high-quality clinical data is relatively scarce, and the annotation process is often a burden for clinicians. As a result, the scarcity of medical data limits the performance of existing medical image segmentation models. In this paper, we propose a novel framework that integrates eye tracking information from experienced radiologists during the screening process to improve the performance of deep neural networks with limited data. Our approach, a grouped hierarchical network, guides the network to learn from its faults by using gaze information as weak supervision. We demonstrate the effectiveness of our framework on mammogram images, particularly for handling segmentation classes with large scale differences. We evaluate the impact of gaze information on medical image segmentation tasks and show that our method achieves better segmentation performance compared to state-of-the-art models. A robustness study is conducted to investigate the influence of distraction or inaccuracies in gaze collection. We also develop a convenient system for collecting gaze data without interrupting the normal clinical workflow. Our work offers novel insights into the potential benefits of integrating gaze information into medical image segmentation tasks. Jiaming Xie, Zhiming Cui 0001, Chong Ma 0004, Wenping Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 7 |
| 2025 | AASeg: Artery-Aware Global-to-Local Framework for Aneurysm Segmentation in Head and Neck CTA ImagesabstractAneurysm segmentation in computed tomography angiography (CTA) images is essential for medical intervention aimed at preventing subarachnoid hemorrhages. However, most existing studies tend to overlook the topological characteristics of arteries related to aneurysms, often resulting in suboptimal performance in aneurysm segmentation. To address this challenge, we propose an artery-aware global-to-local framework for aneurysm segmentation (AASeg) using CTA images of head and neck. This framework consists of two key components: 1) a centerline graph network (CG-Net) for aneurysm global localization, and 2) a point cloud network (PC-Net) for local aneurysm segmentation. The centerline graph is generated by extracting artery centerline structures from vessel masks obtained through a pre-trained model for head and neck vessel segmentation. This representation serves as a high-level representation of the artery structure, allowing for analysis of aneurysms along the entire arteries. It facilitates aneurysm localization via aneurysm-segment graph classification along the arteries. Then, local region of aneurysm segment can be sampled from the vessel mask according to the aneurysm-segment graph. Subsequently, aneurysm segmentation is performed on the point cloud constructed from the aneurysm segment through the PC-Net. Extensive experiments show that the proposed framework achieves state-of-the-art performance in aneurysm localization on a main dataset and an external testing dataset, with Recall of 84.1% and 80.7%, false positives per case of 1.72 and 1.69, and segmentation DSC of 66.1% and 60.2%, respectively. Linlin Yao, Dongdong Chen 0003, Xiangyu Zhao 0003, Manman Fei, Zhiyun Song, Zhong Xue, Yiqiang Zhan, Bin Song 0002, Feng Shi 0001, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 11 |
| 2025 | Development-Driven Diffusion Model for Longitudinal Prediction of Fetal Brain MRI With Unpaired DataabstractLongitudinal magnetic resonance imaging (MRI) is essential for studying the early development of the brain, as it allows to observe and analyze how the brain changes over time. Unfortunately, existing cohort research suffers from lacking sufficient MRI data for studying the development of fetal brains. Apart from research data, a viable alternative is the use of large-scale clinical fetal brain MRI data, which is currently the primary source for longitudinal studies. Although clinical data has several benefits, it is impeded by the inherent drawback of incomplete data. In the context of clinical practice, nearly all subjects undergo only one MRI scan throughout their entire pregnancy, resulting in a lack of longitudinal data for any fetus. To address this issue and obtain longitudinal clinical fetal brain MRI data, we propose to generate MR images for two adjacent gestational weeks (GWs) within one subject, thereby bridging the information gap between three consecutive GWs. This fetal MRI prediction task suffers from two significant challenges, including 1) heterogeneous generation and 2) the lack of paired training data at adjacent GWs. To tackle these two challenges, we propose a new approach, called the Development-driven Diffusion Model (DDM). Specifically, our approach first involves training a conditional diffusion model using population development information spanning all GWs. This allows the model to generate images at various GWs. Next, during the inference stage, we incorporate individual development information of a specific subject using a specially designed perception feature guidance module. The DDM enables the generated 3D MR images to encompass both the general characteristics representative of the targeted GWs, as well as the distinct feature specific to each individual. To assess the efficacy of our approach, extensive experiments were carried out on a large-scale clinical dataset obtained from three different medical centers. The experimental results unequivocally establish the effectiveness of DDM for generating longitudinal MR images of fetal brains. Kai Zhang 0039, Geng Chen 0001, Fangmei Zhu, Zhongxiang Ding, Dinggang Shen |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Asynchronous Functional Brain Network Construction With Spatiotemporal Transformer for MCI ClassificationabstractConstruction and analysis of functional brain networks (FBNs) with resting-state functional magnetic resonance imaging (rs-fMRI) is a promising method to diagnose functional brain diseases. Nevertheless, the existing methods suffer from several limitations. First, the functional connectivities (FCs) of the FBN are usually measured by the temporal co-activation level between rs-fMRI time series from regions of interest (ROIs). While enjoying simplicity, the existing approach implicitly assumes simultaneous co-activation of all the ROIs, and models only their synchronous dependencies. However, the FCs are not necessarily always synchronous due to the time lag of information flow and cross-time interactions between ROIs. Therefore, it is desirable to model asynchronous FCs. Second, the traditional methods usually construct FBNs at individual level, leading to large variability and degraded diagnosis accuracy when modeling asynchronous FBN. Third, the FBN construction and analysis are conducted in two independent steps without joint alignment for the target diagnosis task. To address the first limitation, this paper proposes an effective sliding-window-based method to model spatiotemporal FCs in Transformer. Regarding the second limitation, we propose to learn common and individual FBNs adaptively with the common FBN as prior knowledge, thus alleviating the variability and enabling the network to focus on the individual disease-specific asynchronous FCs. To address the third limitation, the common and individual asynchronous FBNs are built and analyzed by an integrated network, enabling end-to-end training and improving the flexibility and discriminability. The effectiveness of the proposed method is consistently demonstrated on three data sets for mild cognitive impairment (MCI) diagnosis. Jianjia Zhang, Xiaotong Wu, Xiang Tang, Luping Zhou, Lei Wang 0001, Weiwen Wu, Dinggang Shen |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Coupled Diffusion Models for Metal Artifact Reduction of Clinical Dental CBCT ImagesabstractMetal dental implants may introduce metal artifacts (MA) during the CBCT imaging process, causing significant interference in subsequent diagnosis. In recent years, many deep learning methods for metal artifact reduction (MAR) have been proposed. Due to the huge difference between synthetic and clinical MA, supervised learning MAR methods may perform poorly in clinical settings. Many existing unsupervised MAR methods trained on clinical data often suffer from incorrect dental morphology. To alleviate the above problems, in this paper, we propose a new MAR method of Coupled Diffusion Models (CDM) for clinical dental CBCT images. Specifically, we separately train two diffusion models on clinical MA-degraded images and clinical clean images to obtain prior information, respectively. During the denoising process, the variances of noise levels are calculated from MA images and the prior of diffusion models. Then we develop a noise transformation module between the two diffusion models to transform the MA noise image into a new initial value for the denoising process. Our designs effectively exploit the inherent transformation between the misaligned MA-degraded images and clean images. Additionally, we introduce an MA-adaptive inference technique to better accommodate the MA degradation in different areas of an MA-degraded image. Experiments on our clinical dataset demonstrate that our CDM outperforms the comparison methods on both objective metrics and visual quality, especially for severe MA degradation. We will publicly release our code. Zhouzhuo Zhang, Juncheng Yan, Zhiming Cui 0001, Jun Xu 0019, Dinggang Shen |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Prototype Learning Guided Hybrid Network for Breast Tumor Segmentation in DCE-MRIabstractAutomated breast tumor segmentation on the basis of dynamic contrast-enhancement magnetic resonance imaging (DCE-MRI) has shown great promise in clinical practice, particularly for identifying the presence of breast disease. However, accurate segmentation of breast tumor is a challenging task, often necessitating the development of complex networks. To strike an optimal trade-off between computational costs and segmentation performance, we propose a hybrid network via the combination of convolution neural network (CNN) and transformer layers. Specifically, the hybrid network consists of a encoder-decoder architecture by stacking convolution and deconvolution layers. Effective 3D transformer layers are then implemented after the encoder subnetworks, to capture global dependencies between the bottleneck features. To improve the efficiency of hybrid network, two parallel encoder subnetworks are designed for the decoder and the transformer layers, respectively. To further enhance the discriminative capability of hybrid network, a prototype learning guided prediction module is proposed, where the category-specified prototypical features are calculated through online clustering. All learned prototypical features are finally combined with the features from decoder for tumor mask prediction. The experimental results on private and public DCE-MRI datasets demonstrate that the proposed hybrid network achieves superior performance than the state-of-the-art (SOTA) methods, while maintaining balance between segmentation accuracy and computation cost. Moreover, we demonstrate that automatically generated tumor masks can be effectively applied to identify HER2-positive subtype from HER2-negative subtype with the similar accuracy to the analysis based on manual tumor segmentation. The source code is available at https://github.com/ZhouL-lab/PLHN. Lei Zhou 0003, Yuzhong Zhang, Xuejun Qian, Chen Gong 0002, Zhongxiang Ding, Zhenhui Li, Zaiyi Liu, Dinggang Shen |
IEEE Trans. Medical Imaging | 11 |
| 2025 | ChatABL: Abductive Learning via Natural Language Interaction With ChatGPTabstractLarge language models (LLMs) such as ChatGPT have recently demonstrated significant potential in mathematical abilities, providing a valuable reasoning paradigm consistent with human natural language. However, LLMs currently have difficulty in bridging perception, language understanding, and reasoning (PLR) capabilities due to incompatibility of the underlying information flow among them, making their reasoning ability not fully elicited and challenging to accomplish complicated reasoning tasks autonomously. To resolve the above problem, a novel method called ChatABL is proposed by integrating LLMs into an abductive learning (ABL) framework, capable of unifying the three abilities effectively in a more user-friendly and understandable manner. Initially, the proposed method uses LLMs to correct the incomplete logical facts for optimizing the perception module, by summarizing and reorganizing domain knowledge represented in natural language format. Then, the perception module also provides necessary logical reasoning materials for feeding LLMs. Finally, these parts are integrated into a dynamic closed-loop system by introducing the feedback form and automatic learning strategies to mutually promote their performance. As a testbed, the variable-length handwritten equation decipherment (HED), an abstract expression of the Mayan calendar decoding, is used to demonstrate that ChatABL has reasoning ability beyond most existing state-of-the-art methods, which has been well-supported by comparative studies. To the best of authors' knowledge, the proposed ChatABL is the first attempt to explore a possible and novel avenue to approaching human-level cognitive ability via natural language interaction by means of ChatGPT. Tianyang Zhong, Yi Pan 0001, Yutong Zhang 0019, Yaonai Wei, Zhengliang Liu, Xiaozheng Wei, Wenjun Li 0001, Chong Ma 0004, Xi Jiang 0001, Dinggang Shen, Junwei Han 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 12 |
| 2024 | NaMa: Neighbor-Aware Multi-Modal Adaptive Learning for Prostate Tumor Segmentation on Anisotropic MR ImagesabstractAccurate segmentation of prostate tumors from multi-modal magnetic resonance (MR) images is crucial for diagnosis and treatment of prostate cancer. However, the robustness of existing segmentation methods is limited, mainly because these methods 1) fail to adaptively assess subject-specific information of each MR modality for accurate tumor delineation, and 2) lack effective utilization of inter-slice information across thick slices in MR images to segment tumor as a whole 3D volume. In this work, we propose a two-stage neighbor-aware multi-modal adaptive learning network (NaMa) for accurate prostate tumor segmentation from multi-modal anisotropic MR images. In particular, in the first stage, we apply subject-specific multi-modal fusion in each slice by developing a novel modality-informativeness adaptive learning (MIAL) module for selecting and adaptively fusing informative representation of each modality based on inter-modality correlations. In the second stage, we exploit inter-slice feature correlations to derive volumetric tumor segmentation. Specifically, we first use a Unet variant with sequence layers to coarsely capture slice relationship at a global scale, and further generate an activation map for each slice. Then, we introduce an activation mapping guidance (AMG) module to refine slice-wise representation (via information from adjacent slices) for consistent tumor segmentation across neighboring slices. Besides, during the network training, we further apply a random mask strategy to each MR modality to improve feature representation efficiency. Experiments on both in-house and public (PICAI) multi-modal prostate tumor datasets show that our proposed NaMa performs better than state-of-the-art methods. Runqi Meng, Xiao Zhang 0028, Yuning Gu, Guiqin Liu, Nizhuan Wang 0001, Kaicong Sun, Dinggang Shen |
AAAI | 9 |
| 2024 | Mining Gaze for Contrastive Learning toward Computer-Assisted DiagnosisabstractObtaining large-scale radiology reports can be difficult for medical images due to ethical concerns, limiting the effectiveness of contrastive pre-training in the medical image domain and underscoring the need for alternative methods. In this paper, we propose eye-tracking as an alternative to text reports, as it allows for the passive collection of gaze signals without ethical issues. By tracking the gaze of radiologists as they read and diagnose medical images, we can understand their visual attention and clinical reasoning. When a radiologist has similar gazes for two medical images, it may indicate semantic similarity for diagnosis, and these images should be treated as positive pairs when pre-training a computer-assisted diagnosis (CAD) network through contrastive learning. Accordingly, we introduce the Medical contrastive Gaze Image Pre-training (McGIP) as a plug-and-play module for contrastive learning frameworks. McGIP uses radiologist gaze to guide contrastive pre-training. We evaluate our method using two representative types of medical images and two common types of gaze data. The experimental results demonstrate the practicality of McGIP, indicating its high potential for various clinical scenarios and applications. Zihao Zhao 0002, Sheng Wang 0014, Qian Wang 0001, Dinggang Shen |
AAAI | 4 |
| 2024 | Multi-Modal Brain Graph Learning of Shared-Specific Features for Schizophrenia Diagnosis
Geng Chen 0001, Xuyun Wen, Dinggang Shen |
BIBM | 4 |
| 2024 | Decoding White Matter Fiber ODFs: A Mixture Learning Framework in x-q SpaceabstractDiffusion magnetic resonance imaging (dMRI), as a powerful non-invasive white matter imaging technology, plays an important role in studying brain white matter. The fiber orientation distribution functions (fODFs) derived from dMRI data provide the key directional information of fiber tracts for revealing the 3D geometric structure of brain white matter. The estimation of fODFs faces two challenges, including (i) the demand for dMRI data densely sampled in q-space and (ii) the joint consideration of x-q space. To address these challenges, we propose a mixture learning framework with q-space sparely sampled dMRI data as input. Specifically, we propose an x-space learning module based on 3D U-Net to learn x-space features and a q-space learning module based on spherical convolutional neural networks to learn q-space features. Two kinds of features are then fused with a mixture learning fusion module for fODFs estimation. The whole framework is supervised with an x-q space loss function. Our framework makes full use of joint x-q space information for fODFs estimation with clinically available q-space sparsely sampled dMRI data. Extensive experiments on three public datasets show that our framework is effective in fODFs estimation and outperforms cutting-edge models. Jiquan Ma, Chengdong Deng, Geng Chen 0001, Jaeil Kim, Xuyun Wen, Dinggang Shen |
BIBM | 8 |
| 2024 | Image2Points: A 3D Point-Based Context Clusters GAN for High-Quality Pet Image ReconstructionabstractTo obtain high-quality Positron emission tomography (PET) images while minimizing radiation exposure, numerous methods have been proposed to reconstruct standard-dose PET (SPET) images from the corresponding low-dose PET (LPET) images. However, these methods heavily rely on voxel-based representations, which fall short of adequately accounting for the precise structure and fine-grained context, leading to compromised reconstruction. In this paper, we propose a 3D point-based context clusters GAN, namely PCC-GAN, to reconstruct high-quality SPET images from LPET. Specifically, inspired by the geometric representation power of points, we resort to a point-based representation to enhance the explicit expression of the image structure, thus facilitating the reconstruction with finer details. Moreover, a context clustering strategy is applied to explore the contextual relationships among points, which mitigates the ambiguities of small structures in the reconstructed images. Experiments on both clinical and phantom datasets demonstrate that our PCC-GAN outperforms the state-of-the-art reconstruction methods qualitatively and quantitatively. Code is available at https://github.com/gluucose/PCCGAN. Yan Wang 0015, Lu Wen, Pinxian Zeng, Xi Wu 0004, Jiliu Zhou, Dinggang Shen |
ICASSP | 7 |
| 2024 | Synthesizing Aβ-Pet Via An Image And Label Conditioning Latent Diffusion Model For Detecting Amyloid StatusabstractDeposition of β-amyloid is a crucial biomarker to evaluate subjects with early-onset dementia, often evaluated through Aβ-PET imaging. Aβ-PET is expensive and radiation-heavy; thus, it’s advisable to avoid it unless medically necessary. Therefore there is a compelling need to classify Aβ and detect amyloid status using other neuroimaging modalities, capitalizing on the underlying relationship between different modalities. Here, we propose an image and label conditioning latent diffusion model to synthesize Aβ-PET scans for improving Aβ classification based on MRI and FDG-PET scans. We introduce two conditioning modules: (1) an image conditioning module to extract a meaningful feature map from two source modalities to provide structural and metabolism information for guidance, and (2) a label conditioning module to provide the specific guidance direction on image generation. Experiments on the clinical dataset demonstrate that our proposed method’s synthetic Aβ-PET scans are reliable for classifying Aβ. Zaixin Ou, Yongsheng Pan, Yuanning Li, Qihao Guo, Dinggang Shen |
ICASSP | 6 |
| 2024 | A Prior-information-guided Residual Diffusion Model for Multi-modal PET Synthesis from MRI
Zaixin Ou, Caiwen Jiang, Yongsheng Pan, Yuanwang Zhang, Zhiming Cui 0001, Dinggang Shen |
IJCAI | 6 |
| 2024 | D2GAN: A Dual-Domain Generative Adversarial Network for High-Quality PET Image ReconstructionabstractPositron emission tomography (PET) is a widely adopted nuclear imaging technique for early tumor detection and brain disorder diagnosis, while its intrinsic tracer radiation inevitably poses health risks for patients. Recently, to achieve high-quality PET imaging while reducing radiation exposure, numerous methods have been proposed to reconstruct high-quality standard-dose PET (SPET) images from low-dose PET (LPET) images. However, these methods usually overlooked crucial regions and details during the reconstruction, leading to high-frequency distortions in the reconstructed images. To this end, we propose D2GAN, a dual-domain generative adversarial network that utilizes spatial and frequency domain information to mitigate high-frequency disparities, facilitating high-quality PET reconstruction. The core of our approach is the Dual-Domain Learning Block (DLB), comprising a Spatial Domain Learning Block (SDLB) for identifying key regions and details in PET images, and a Frequency Domain Learning Block (FDLB) to further refine these areas by amplifying the high-frequency signals of the image. In addition, we introduce a multi-scale residual block (MSRB) to efficiently extract features at various scales and incorporate a focal frequency loss to encourage the consistency between the reconstructed and the real SPET images in the frequency domain. The DLBs and MSRBs are embedded into a U-shaped structure to form our generator. Furthermore, we apply a patch-based discriminator to enforce the data distribution consistency of the reconstructed PET images. Extensive experiments on two public datasets and an in-house clinical dataset demonstrate that our approach outperforms the state-of-the-art PET reconstruction methods. Binyu Yan, Jiliu Zhou, Yan Wang 0015, Dinggang Shen |
IJCNN | 7 |
| 2024 | Dynamic Hybrid Unrolled Multi-scale Network for Accelerated MRI Reconstruction
Xiaoxin Li 0001, Fang-Zheng Zhu, Yong Chen 0026, Dinggang Shen |
MICCAI (7) | 5 |
| 2024 | k-t Self-consistency Diffusion: A Physics-Informed Model for Dynamic MR Imaging
Zhuo-Xu Cui, Kaicong Sun, Yuliang Zhu, Dinggang Shen, Dong Liang 0001 |
MICCAI (7) | 7 |
| 2024 | LM-UNet: Whole-Body PET-CT Lesion Segmentation with Dual-Modality-Based Annotations Driven by Latent Mamba U-Net
Anglin Liu, Dengqiang Jia, Kaicong Sun, Runqi Meng, Meixin Zhao, Yongluo Jiang, Zhijian Dong, Yaozong Gao, Dinggang Shen |
MICCAI (9) | 9 |
| 2024 | UinTSeg: Unified Infant Brain Tissue Segmentation with Anatomy Delineation
Jiameng Liu, Feihong Liu, Kaicong Sun, Caiwen Jiang, Islem Rekik, Dinggang Shen |
MICCAI (2) | 8 |
| 2024 | A Graph-Embedded Latent Space Learning and Clustering Framework for Incomplete Multimodal Multiclass Alzheimer's Disease Diagnosis
Zaixin Ou, Caiwen Jiang, Yuanwang Zhang, Zhiming Cui 0001, Dinggang Shen |
MICCAI (7) | 6 |
| 2024 | Prompt-Based Segmentation Model of Anatomical Structures and Lesions in CT Images
Xi Ouyang, Dongdong Gu, Qianqian Chen 0002, Yiqiang Zhan, Xiang Sean Zhou, Feng Shi 0001, Zhong Xue, Dinggang Shen |
MICCAI (8) | 10 |
| 2024 | Hierarchical Symmetric Normalization Registration Using Deformation-Inverse Network
Qingrui Sha, Kaicong Sun, Yonghao Li, Zhong Xue, Xiaohuan Cao, Dinggang Shen |
MICCAI (2) | 7 |
| 2024 | HF-ResDiff: High-Frequency-Guided Residual Diffusion for Multi-dose PET Reconstruction
Caiwen Jiang, Zhiming Cui 0001, Dinggang Shen |
MICCAI (7) | 4 |
| 2024 | Knowledge-Guided Prompt Learning for Lifespan Brain MR Image Segmentation
Zihao Zhao 0002, Zehong Cao, Runqi Meng, Feng Shi 0001, Dinggang Shen |
MICCAI (2) | 7 |
| 2024 | Cephalometric Landmark Detection Across Ages with Prototypical Network
Han Wu 0007, Chong Wang 0012, Lanzhuju Mei, Dinggang Shen, Zhiming Cui 0001 |
MICCAI (5) | 6 |
| 2024 | TeethDreamer: 3D Teeth Reconstruction from Five Intra-Oral Photographs
Chenfan Xu, Yuan Liu 0025, Yulong Dou, Jiepeng Wang 0001, Minjiao Wang, Dinggang Shen, Zhiming Cui 0001 |
MICCAI (7) | 8 |
| 2024 | Exploiting Latent Classes for Medical Image Segmentation from Partially Labeled Datasets
Xiangyu Zhao 0003, Xi Ouyang, Lichi Zhang, Zhong Xue, Dinggang Shen |
MICCAI (8) | 5 |
| 2024 | LoCI-DiffCom: Longitudinal Consistency-Informed Diffusion Model for 3D Infant Brain Image Completion
Tianli Tao, Yitian Tao, Haowen Deng, Xinyi Cai, Gaofeng Wu, Kaidong Wang, Haifeng Tang, Lixuan Zhu, Zhuoyang Gu, Dinggang Shen, Han Zhang 0002 |
MICCAI (2) | 11 |
| 2024 | Eye-gaze Guided Multi-modal Alignment for Medical Representation LearningabstractIn the medical multi-modal frameworks, the alignment of cross-modality features presents a significant challenge. However, existing works have learned features that are implicitly aligned from the data, without considering the explicit relationships in the medical context. This data-reliance may lead to low generalization of the learned alignment relationships. In this work, we propose the Eye-gaze Guided Multi-modal Alignment (EGMA) framework to harness eye-gaze data for better alignment of medical visual and textual features. We explore the natural auxiliary role of radiologists' eye-gaze data in aligning medical images and text, and introduce a novel approach by using eye-gaze data, collected synchronously by radiologists during diagnostic evaluations. We conduct downstream tasks of image classification and image-text retrieval on four medical datasets, where EGMA achieved state-of-the-art performance and stronger generalization across different datasets. Additionally, we explore the impact of varying amounts of eye-gaze data on model performance, highlighting the feasibility and utility of integrating this auxiliary data into multi-modal alignment framework. Chong Ma 0004, Hanqi Jiang, Wenting Chen, Yiwei Li 0002, Zihao Wu 0001, Xiaowei Yu 0001, Zhengliang Liu, Lei Guo 0002, Dajiang Zhu, Dinggang Shen, Tianming Liu 0001, Xiang Li 0001 |
NeurIPS | 11 |
| 2024 | Real-time diagnosis of intracerebral hemorrhage by generating dual-energy CT from single-energy CT
Caiwen Jiang, Tianyu Wang 0014, Yongsheng Pan, Zhongxiang Ding, Dinggang Shen |
Medical Image Anal. | 5 |
| 2024 | Beam-wise dose composition learning for head and neck cancer dose prediction in radiotherapy
Bin Wang 0068, Xuanang Xu, Lanzhuju Mei, Qianjin Feng 0003, Dinggang Shen |
Medical Image Anal. | 7 |
| 2024 | 3D multi-modality Transformer-GAN for high-quality PET reconstruction
Yan Wang 0015, Yanmei Luo, Chen Zu, Bo Zhan, Zhengyang Jiao, Xi Wu 0004, Jiliu Zhou, Dinggang Shen, Luping Zhou |
Medical Image Anal. | 8 |
| 2024 | Constructing hierarchical attentive functional brain networks for early AD diagnosis
Jianjia Zhang, Yunan Guo, Luping Zhou, Lei Wang 0001, Weiwen Wu, Dinggang Shen |
Medical Image Anal. | 6 |
| 2024 | Detail-preserving image warping by enforcing smooth image sampling
Qingrui Sha, Kaicong Sun, Caiwen Jiang, Zhong Xue, Xiaohuan Cao, Dinggang Shen |
Neural Networks | 7 |
| 2024 | Hierarchical Encoding and Fusion of Brain Functions for Depression Subtype ClassificationabstractDepression is a serious mental disorder with complex etiology, exhibiting strong heterogeneity in clinical manifestations such as various subtypes. Research on depression subtypes may deepen the understanding of the disease, contributing to the diagnosis and prognosis. While brain functional network and graph neural networks (GNNs) provide such a means, the task is still challenged by limited feature encoding from the informative fMRI data, ineffective information fusion of brain functional network, and small size of the recruited subjects. Therefore, we propose a hierarchical encoding and fusion framework of brain functions. First, we pre-train a model to extract the features from individual brain regions, which signify nodes in the brain functional network. Then, distinct graphs are constructed to link the nodes within each subject, resulting in multi-view graphs of the brain functional network. We further develop a graph fusion strategy to integrate the multi-view information, by referring to the local encoding of the nodes and their interactions across multiple graph instances. Finally, we attain the classification of depression subtypes based on the fused graph representation. The experimental results demonstrate that our method can superiorly distinguish major depression subtypes and outperform the state-of-the-art methods. Mengjun Liu, Huifeng Zhang, Mianxin Liu, Dongdong Chen 0003, Rubai Zhou, Wenxian Lu, Lichi Zhang, Dinggang Shen, Qian Wang 0001, Daihui Peng |
IEEE Trans. Affect. Comput. | 8 |
| 2024 | 3D Point-Based Multi-Modal Context Clusters GAN for Low-Dose PET Image DenoisingabstractTo obtain high-quality Positron emission tomography (PET) images while minimizing radiation hazards, various methods have been developed to acquire standard-dose PET (SPET) images from low-dose PET (LPET) images. Recent efforts mainly focus on improving the denoising quality by utilizing multi-modal inputs. However, these methods exhibit certain limitations. First, they neglect the varied significance of each modality in denoising. Second, they rely on inflexible voxel-based representations, failing to explicitly preserve intricate structures and contexts in images. To alleviate these problems, we propose a 3D Point-based Multi-modal Context Clusters GAN, namely PMC2-GAN, for obtaining high-quality SPET images from LPET and magnetic resonance imaging (MRI) images. Specifically, we transform the 3D image into unorganized points to flexibly and precisely express its complex structure. Moreover, a self-context clusters (Self-CC) block is devised to explore fine-grained contextual relationships of the image from the perspective of points. Additionally, considering the diverse importance of different modalities, we introduce a cross-context clusters (Cross-CC) block, which prioritizes PET as the primary modality while regarding MRI as the auxiliary one, to effectively integrate the knowledge from the two modalities. Overall, built on the smart integration of Self- and Cross-CC blocks, our PMC2-GAN follows GAN architecture. Extensive experiments validate our superiority. Yan Wang 0015, Luping Zhou, Yuchen Fei, Jiliu Zhou, Dinggang Shen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Multimodal Brain Tumor Segmentation Boosted by Monomodal Normal Brain ImagesabstractMany deep learning based methods have been proposed for brain tumor segmentation. Most studies focus on deep network internal structure to improve the segmentation accuracy, while valuable external information, such as normal brain appearance, is often ignored. Inspired by the fact that radiologists often screen lesion regions with normal appearance as reference in mind, in this paper, we propose a novel deep framework for brain tumor segmentation, where normal brain images are adopted as reference to compare with tumor brain images in a learned feature space. In this way, features at tumor regions, i.e., tumor-related features, can be highlighted and enhanced for accurate tumor segmentation. It is known that routine tumor brain images are multimodal, while normal brain images are often monomodal. This causes the feature comparison a big issue, i.e., multimodal vs. monomodal. To this end, we present a new feature alignment module (FAM) to make the feature distribution of monomodal normal brain images consistent/inconsistent with multimodal tumor brain images at normal/tumor regions, making the feature comparison effective. Both public (BraTS2022) and in-house tumor brain image datasets are used to evaluate our framework. Experimental results demonstrate that for both datasets, our framework can effectively improve the segmentation accuracy and outperforms the state-of-the-art segmentation methods. Codes are available at https://github.com/hb-liu/Normal-Brain-Boost-Tumor-Segmentation. Huabing Liu, Zhengze Ni, Dong Nie, Dinggang Shen, Jinda Wang, Zhenyu Tang 0002 |
IEEE Trans. Image Process. | 4 |
| 2024 | Image Recovery Matters: A Recovery-Extraction Framework for Robust Fetal Brain Extraction From MR ImagesabstractThe extraction of the fetal brain from magnetic resonance (MR) images is a challenging task. In particular, fetal MR images suffer from different kinds of artifacts introduced during the image acquisition. Among those artifacts, intensity inhomogeneity is a common one affecting brain extraction. In this work, we propose a deep learning-based recovery-extraction framework for fetal brain extraction, which is particularly effective in handling fetal MR images with intensity inhomogeneity. Our framework involves two stages. First, the artifact-corrupted images are recovered with the proposed generative adversarial learning-based image recovery network with a novel region-of-darkness discriminator that enforces the network focusing on artifacts of the images. Second, we propose a brain extraction network for more effective fetal brain segmentation by strengthening the association between lower- and higher-level features as well as suppressing task-irrelevant features. Thanks to the proposed recovery-extraction strategy, our framework is able to accurately segment fetal brains from artifact-corrupted MR images. The experiments show that our framework achieves promising performance in both quantitative and qualitative evaluations, and outperforms state-of-the-art methods in both image recovery and fetal brain extraction. Ranlin Lu, Shilin Ye, Mengting Guang, Tewodros Megabiaw Tassew, Bin Jing, Guofu Zhang, Geng Chen 0001, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 9 |
| 2024 | FEFA: Frequency Enhanced Multi-Modal MRI Reconstruction With Deep Feature AlignmentabstractIntegrating complementary information from multiple magnetic resonance imaging (MRI) modalities is often necessary to make accurate and reliable diagnostic decisions. However, the different acquisition speeds of these modalities mean that obtaining information can be time consuming and require significant effort. Reference-based MRI reconstruction aims to accelerate slower, under-sampled imaging modalities, such as T2-modality, by utilizing redundant information from faster, fully sampled modalities, such as T1-modality. Unfortunately, spatial misalignment between different modalities often negatively impacts the final results. To address this issue, we propose FEFA, which consists of cascading FEFA blocks. The FEFA block first aligns and fuses the two modalities at the feature level. The combined features are then filtered in the frequency domain to enhance the important features while simultaneously suppressing the less essential ones, thereby ensuring accurate reconstruction. Furthermore, we emphasize the advantages of combining the reconstruction results from multiple cascaded blocks, which also contributes to stabilizing the training process. Compared to existing registration-then-reconstruction and cross-attention-based approaches, our method is end-to-end trainable without requiring additional supervision, extensive parameters, or heavy computation. Experiments on the public fastMRI, IXI and in-house datasets demonstrate that our approach is effective across various under-sampling patterns and ratios. Xuanmin Chen, Liyan Ma, Shihui Ying, Dinggang Shen, Tieyong Zeng |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | A Unified Multi-Modality Fusion Framework for Deep Spatio-Spectral-Temporal Feature Learning in Resting-State fMRI DenoisingabstractResting-state functional magnetic resonance imaging (rs-fMRI) is a commonly used functional neuroimaging technique to investigate the functional brain networks. However, rs-fMRI data are often contaminated with noise and artifacts that adversely affect the results of rs-fMRI studies. Several machine/deep learning methods have achieved impressive performance to automatically regress the noise-related components decomposed from rs-fMRI data, which are expressed as the pairs of a spatial map and its associated time series. However, most of the previous methods individually analyze each modality of the noise-related components and simply aggregate the decision-level information (or knowledge) extracted from each modality to make a final decision. Moreover, these approaches consider only the limited modalities making it difficult to explore class-discriminative spectral information of noise-related components. To overcome these limitations, we propose a unified deep attentive spatio-spectral-temporal feature fusion framework. We first adopt a learnable wavelet transform module at the input-level of the framework to elaborately explore the spectral information in subsequent processes. We then construct a feature-level multi-modality fusion module to efficiently exchange the information from multi-modality inputs in the feature space. Finally, we design confidence-based voting strategies for decision-level fusion at the end of the framework to make a robust final decision. In our experiments, the proposed method achieved remarkable performance for noise-related component detection on various rs-fMRI datasets. Minjoo Lim, Keun-Soo Heo, Junmo Kim 0001, Bogyeong Kang, Weili Lin, Han Zhang 0002, Dinggang Shen, Tae-Eui Kam |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | Structure-Aware Registration Network for Liver DCE-CT ImagesabstractImage registration of liver dynamic contrast-enhanced computed tomography (DCE-CT) is crucial for diagnosis and image-guided surgical planning of liver cancer. However, intensity variations due to the flow of contrast agents combined with complex spatial motion induced by respiration brings great challenge to existing intensity-based registration methods. To address these problems, we propose a novel structure-aware registration method by incorporating structural information of related organs with segmentation-guided deep registration network. Existing segmentation-guided registration methods only focus on volumetric registration inside the paired organ segmentations, ignoring the inherent attributes of their anatomical structures. In addition, such paired organ segmentations are not always available in DCE-CT images due to the flow of contrast agents. Different from existing segmentation-guided registration methods, our proposed method extracts structural information in hierarchical geometric perspectives of line and surface. Then, according to the extracted structural information, structure-aware constraints are constructed and imposed on the forward and backward deformation field simultaneously. In this way, all available organ segmentations, including unpaired ones, can be fully utilized to avoid the side effect of contrast agent and preserve the topology of organs during registration. Extensive experiments on an in-house liver DCE-CT dataset and a public LiTS dataset show that our proposed method can achieve higher registration accuracy and preserve anatomical structure more effectively than state-of-the-art methods. Peng Xue 0005, Jingyang Zhang, Lei Ma 0006, Mianxin Liu, Yuning Gu, Feihong Liu, Yongsheng Pan, Xiaohuan Cao, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 10 |
| 2024 | Attention-Based MultiOffset Deep Learning Reconstruction of Chemical Exchange Saturation Transfer (AMO-CEST) MRIabstractOne challenge of chemical exchange saturation transfer (CEST) magnetic resonance imaging (MRI) is the long scan time due to multiple acquisitions of images at different saturation frequency offsets. k-space under-sampling strategy is commonly used to accelerate MRI acquisition, while this could introduce artifacts and reduce signal-to-noise ratio (SNR). To accelerate CEST-MRI acquisition while maintaining suitable image quality, we proposed an attention-based multioffset deep learning reconstruction network (AMO-CEST) with a multiple radial k-space sampling strategy for CEST-MRI. The AMO-CEST also contains dilated convolution to enlarge the receptive field and data consistency module to preserve the sampled k-space data. We evaluated the proposed method on a mouse brain dataset containing 5760 CEST images acquired at a pre-clinical 3 T MRI scanner. Quantitative results demonstrated that AMO-CEST showed obvious improvement over zero-filling method with a PSNR enhancement of 11 dB, a SSIM enhancement of 0.15, and a NMSE decrease of [Formula: see text] in three acquisition orientations. Compared with other deep learning-based models, AMO-CEST showed visual and quantitative improvements in images from three different orientations. We also extracted molecular contrast maps, including the amide proton transfer (APT) and the relayed nuclear Overhauser enhancement (rNOE). The results demonstrated that the CEST contrast maps derived from the CEST images of AMO-CEST were comparable to those derived from the original high-resolution CEST images. The proposed AMO-CEST can efficiently reconstruct high-quality CEST images from under-sampled k-space data and thus has the potential to accelerate CEST-MRI acquisition. Zhikai Yang, Dinggang Shen, Kannie W. Y. Chan, Jianpan Huang |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | Modality-Specific Information Disentanglement From Multi-Parametric MRI for Breast Tumor Segmentation and Computer-Aided DiagnosisabstractBreast cancer is becoming a significant global health challenge, with millions of fatalities annually. Magnetic Resonance Imaging (MRI) can provide various sequences for characterizing tumor morphology and internal patterns, and becomes an effective tool for detection and diagnosis of breast tumors. However, previous deep-learning based tumor segmentation methods from multi-parametric MRI still have limitations in exploring inter-modality information and focusing task-informative modality/modalities. To address these shortcomings, we propose a Modality-Specific Information Disentanglement (MoSID) framework to extract both inter- and intra-modality attention maps as prior knowledge for guiding tumor segmentation. Specifically, by disentangling modality-specific information, the MoSID framework provides complementary clues for the segmentation task, by generating modality-specific attention maps to guide modality selection and inter-modality evaluation. Our experiments on two 3D breast datasets and one 2D prostate dataset demonstrate that the MoSID framework outperforms other state-of-the-art multi-modality segmentation methods, even in the cases of missing modalities. Based on the segmented lesions, we further train a classifier to predict the patients' response to radiotherapy. The prediction accuracy is comparable to the case of using manually-segmented tumors for treatment outcome prediction, indicating the robustness and effectiveness of the proposed segmentation method. The code is available at https://github.com/Qianqian-Chen/MoSID. Qianqian Chen 0002, Runqi Meng, Lei Zhou 0003, Zhenhui Li, Qianjin Feng 0003, Dinggang Shen |
IEEE Trans. Medical Imaging | 7 |
| 2024 | Prior Knowledge-Guided Triple-Domain Transformer-GAN for Direct PET Reconstruction From Low-Count SinogramsabstractTo obtain high-quality positron emission tomography (PET) images while minimizing radiation exposure, numerous methods have been dedicated to acquiring standard-count PET (SPET) from low-count PET (LPET). However, current methods have failed to take full advantage of the different emphasized information from multiple domains, i.e., the sinogram, image, and frequency domains, resulting in the loss of crucial details. Meanwhile, they overlook the unique inner-structure of the sinograms, thereby failing to fully capture its structural characteristics and relationships. To alleviate these problems, in this paper, we proposed a prior knowledge-guided transformer-GAN that unites triple domains of sinogram, image, and frequency to directly reconstruct SPET images from LPET sinograms, namely PK-TriDo. Our PK-TriDo consists of a Sinogram Inner-Structure-based Denoising Transformer (SISD-Former) to denoise the input LPET sinogram, a Frequency-adapted Image Reconstruction Transformer (FaIR-Former) to reconstruct high-quality SPET images from the denoised sinograms guided by the image domain prior knowledge, and an Adversarial Network (AdvNet) to further enhance the reconstruction quality via adversarial training. Specifically tailored for the PET imaging mechanism, we injected a sinogram embedding module that partitions the sinograms by rows and columns to obtain 1D sequences of angles and distances to faithfully preserve the inner-structure of the sinograms. Moreover, to mitigate high-frequency distortions and enhance reconstruction details, we integrated global-local frequency parsers (GLFPs) into FaIR-Former to calibrate the distributions and proportions of different frequency bands, thus compelling the network to preserve high-frequency details. Evaluations on three datasets with different dose levels and imaging scenarios demonstrated that our PK-TriDo outperforms the state-of-the-art methods. Pinxian Zeng, Xinyi Zeng, Jiliu Zhou, Yan Wang 0015, Dinggang Shen |
IEEE Trans. Medical Imaging | 8 |
| 2024 | Carotid Vessel Wall Segmentation Through Domain Aligner, Topological Learning, and Segment Anything Model for Sparse Annotation in MR ImagesabstractMedical image analysis poses significant challenges due to limited availability of clinical data, which is crucial for training accurate models. This limitation is further compounded by the specialized and labor-intensive nature of the data annotation process. For example, despite the popularity of computed tomography angiography (CTA) in diagnosing atherosclerosis with an abundance of annotated datasets, magnetic resonance (MR) images stand out with better visualization for soft plaque and vessel wall characterization. However, the higher cost and limited accessibility of MR, as well as time-consuming nature of manual labeling, contribute to fewer annotated datasets. To address these issues, we formulate a multi-modal transfer learning network, named MT-Net, designed to learn from unpaired CTA and sparsely-annotated MR data. Additionally, we harness the Segment Anything Model (SAM) to synthesize additional MR annotations, enriching the training process. Specifically, our method first segments vessel lumen regions followed by precise characterization of carotid artery vessel walls, thereby ensuring both segmentation accuracy and clinical relevance. Validation of our method involved rigorous experimentation on publicly available datasets from COSMOS and CARE-II challenge, demonstrating its superior performance compared to existing state-of-the-art techniques. Xibao Li, Xi Ouyang, Zhongxiang Ding, Yuyao Zhang 0005, Zhong Xue, Feng Shi 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 8 |
| 2024 | ScribFormer: Transformer Makes CNN Work Better for Scribble-Based Medical Image SegmentationabstractMost recent scribble-supervised segmentation methods commonly adopt a CNN framework with an encoder-decoder architecture. Despite its multiple benefits, this framework generally can only capture small-range feature dependency for the convolutional layer with the local receptive field, which makes it difficult to learn global shape information from the limited information provided by scribble annotations. To address this issue, this paper proposes a new CNN-Transformer hybrid solution for scribble-supervised medical image segmentation called ScribFormer. The proposed ScribFormer model has a triple-branch structure, i.e., the hybrid of a CNN branch, a Transformer branch, and an attention-guided class activation map (ACAM) branch. Specifically, the CNN branch collaborates with the Transformer branch to fuse the local features learned from CNN with the global representations obtained from Transformer, which can effectively overcome limitations of existing scribble-supervised segmentation methods. Furthermore, the ACAM branch assists in unifying the shallow convolution features and the deep convolution features to improve model's performance further. Extensive experiments on two public datasets and one private dataset show that our ScribFormer has superior performance over the state-of-the-art scribble-supervised segmentation methods, and achieves even better results than the fully-supervised segmentation methods. The code is released at https://github.com/HUANGLIZI/ScribFormer. Dandan Shan, Shuzhou Yang, Qingde Li, Beizhan Wang, Yuan-Ting Zhang, Qingqi Hong, Dinggang Shen |
IEEE Trans. Medical Imaging | 9 |
| 2024 | DSMT-Net: Dual Self-Supervised Multi-Operator Transformation for Multi-Source Endoscopic Ultrasound DiagnosisabstractPancreatic cancer has the worst prognosis of all cancers. The clinical application of endoscopic ultrasound (EUS) for the assessment of pancreatic cancer risk and of deep learning for the classification of EUS images have been hindered by inter-grader variability and labeling capability. One of the key reasons for these difficulties is that EUS images are obtained from multiple sources with varying resolutions, effective regions, and interference signals, making the distribution of the data highly variable and negatively impacting the performance of deep learning models. Additionally, manual labeling of images is time-consuming and requires significant effort, leading to the desire to effectively utilize a large amount of unlabeled data for network training. To address these challenges, this study proposes the Dual Self-supervised Multi-Operator Transformation Network (DSMT-Net) for multi-source EUS diagnosis. The DSMT-Net includes a multi-operator transformation approach to standardize the extraction of regions of interest in EUS images and eliminate irrelevant pixels. Furthermore, a transformer-based dual self-supervised network is designed to integrate unlabeled EUS images for pre-training the representation model, which can be transferred to supervised tasks such as classification, detection, and segmentation. A large-scale EUS-based pancreas image dataset (LEPset) has been collected, including 3,500 pathologically proven labeled EUS images (from pancreatic and non-pancreatic cancers) and 8,000 unlabeled EUS images for model development. The self-supervised method has also been applied to breast cancer diagnosis and was compared to state-of-the-art deep learning models on both datasets. The results demonstrate that the DSMT-Net significantly improves the accuracy of pancreatic and breast cancer diagnosis. Jiajia Li 0004, Lei Zhu 0003, Ruhan Liu, Dinggang Shen, Bin Sheng 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2024 | DTR-Net: Dual-Space 3D Tooth Model Reconstruction From Panoramic X-Ray ImagesabstractIn digital dentistry, cone-beam computed tomography (CBCT) can provide complete 3D tooth models, yet suffers from a long concern of requiring excessive radiation dose and higher expense. Therefore, 3D tooth model reconstruction from 2D panoramic X-ray image is more cost-effective, and has attracted great interest in clinical applications. In this paper, we propose a novel dual-space framework, namely DTR-Net, to reconstruct 3D tooth model from 2D panoramic X-ray images in both image and geometric spaces. Specifically, in the image space, we apply a 2D-to-3D generative model to recover intensities of CBCT image, guided by a task-oriented tooth segmentation network in a collaborative training manner. Meanwhile, in the geometric space, we benefit from an implicit function network in the continuous space, learning using points to capture complicated tooth shapes with geometric properties. Experimental results demonstrate that our proposed DTR-Net achieves state-of-the-art performance both quantitatively and qualitatively in 3D tooth model reconstruction, indicating its potential application in dental practice. Lanzhuju Mei, Yu Fang 0008, Yue Zhao 0012, Xiang Sean Zhou, Zhiming Cui 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 7 |
| 2024 | Multi-Modal Modality-Masked Diffusion Network for Brain MRI Synthesis With Random Modality MissingabstractSynthesis of unavailable imaging modalities from available ones can generate modality-specific complementary information and enable multi-modality based medical images diagnosis or treatment. Existing generative methods for medical image synthesis are usually based on cross-modal translation between acquired and missing modalities. These methods are usually dedicated to specific missing modality and perform synthesis in one shot, which cannot deal with varying number of missing modalities flexibly and construct the mapping across modalities effectively. To address the above issues, in this paper, we propose a unified Multi-modal Modality-masked Diffusion Network (M2DN), tackling multi-modal synthesis from the perspective of "progressive whole-modality inpainting", instead of "cross-modal translation". Specifically, our M2DN considers the missing modalities as random noise and takes all the modalities as a unity in each reverse diffusion step. The proposed joint synthesis scheme performs synthesis for the missing modalities and self-reconstruction for the available ones, which not only enables synthesis for arbitrary missing scenarios, but also facilitates the construction of common latent space and enhances the model representation ability. Besides, we introduce a modality-mask scheme to encode availability status of each incoming modality explicitly in a binary mask, which is adopted as condition for the diffusion model to further enhance the synthesis performance of our M2DN for arbitrary missing scenarios. We carry out experiments on two public brain MRI datasets for synthesis and downstream segmentation tasks. Experimental results demonstrate that our M2DN outperforms the state-of-the-art models significantly and shows great generalizability for arbitrary missing modalities. Kaicong Sun, Jun Xu 0019, Xuming He 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Joint Cross-Attention Network With Deep Modality Prior for Fast MRI ReconstructionabstractCurrent deep learning-based reconstruction models for accelerated multi-coil magnetic resonance imaging (MRI) mainly focus on subsampled k-space data of single modality using convolutional neural network (CNN). Although dual-domain information and data consistency constraint are commonly adopted in fast MRI reconstruction, the performance of existing models is still limited mainly by three factors: inaccurate estimation of coil sensitivity, inadequate utilization of structural prior, and inductive bias of CNN. To tackle these challenges, we propose an unrolling-based joint Cross-Attention Network, dubbed as jCAN, using deep guidance of the already acquired intra-subject data. Particularly, to improve the performance of coil sensitivity estimation, we simultaneously optimize the latent MR image and sensitivity map (SM). Besides, we introduce Gating layer and Gaussian layer into SM estimation to alleviate the "defocus" and "over-coupling" effects and further ameliorate the SM estimation. To enhance the representation ability of the proposed model, we deploy Vision Transformer (ViT) and CNN in the image and k-space domains, respectively. Moreover, we exploit pre-acquired intra-subject scan as reference modality to guide the reconstruction of subsampled target modality by resorting to the self- and cross-attention scheme. Experimental results on public knee and in-house brain datasets demonstrate that the proposed jCAN outperforms the state-of-the-art methods by a large margin in terms of SSIM and PSNR for different acceleration factors and sampling masks. Our code is publicly available at https://github.com/sunkg/jCAN. Kaicong Sun, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 3 |
| 2024 | Spatial and Modal Optimal Transport for Fast Cross-Modal MRI ReconstructionabstractMulti-modal magnetic resonance imaging (MRI) plays a crucial role in comprehensive disease diagnosis in clinical medicine. However, acquiring certain modalities, such as T2-weighted images (T2WIs), is time-consuming and prone to be with motion artifacts. It negatively impacts subsequent multi-modal image analysis. To address this issue, we propose an end-to-end deep learning framework that utilizes T1-weighted images (T1WIs) as auxiliary modalities to expedite T2WIs' acquisitions. While image pre-processing is capable of mitigating misalignment, improper parameter selection leads to adverse pre-processing effects, requiring iterative experimentation and adjustment. To overcome this shortage, we employ Optimal Transport (OT) to synthesize T2WIs by aligning T1WIs and performing cross-modal synthesis, effectively mitigating spatial misalignment effects. Furthermore, we adopt an alternating iteration framework between the reconstruction task and the cross-modal synthesis task to optimize the final results. Then, we prove that the reconstructed T2WIs and the synthetic T2WIs become closer on the T2 image manifold with iterations increasing, and further illustrate that the improved reconstruction result enhances the synthesis process, whereas the enhanced synthesis result improves the reconstruction process. Finally, experimental results from FastMRI and internal datasets confirm the effectiveness of our method, demonstrating significant improvements in image reconstruction quality even at low sampling rates. Qi Wang 0128, Zhijie Wen, Jun Shi 0004, Qian Wang 0001, Dinggang Shen, Shihui Ying |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Adaptive and Iterative Learning With Multi-Perspective Regularizations for Metal Artifact ReductionabstractMetal artifact reduction (MAR) is important for clinical diagnosis with CT images. The existing state-of-the-art deep learning methods usually suppress metal artifacts in sinogram or image domains or both. However, their performance is limited by the inherent characteristics of the two domains, i.e., the errors introduced by local manipulations in the sinogram domain would propagate throughout the whole image during backprojection and lead to serious secondary artifacts, while it is difficult to distinguish artifacts from actual image features in the image domain. To alleviate these limitations, this study analyzes the desirable properties of wavelet transform in-depth and proposes to perform MAR in the wavelet domain. First, wavelet transform yields components that possess spatial correspondence with the image, thereby preventing the spread of local errors to avoid secondary artifacts. Second, using wavelet transform could facilitate identification of artifacts from image since metal artifacts are mainly high-frequency signals. Taking these advantages of the wavelet transform, this paper decomposes an image into multiple wavelet components and introduces multi-perspective regularizations into the proposed MAR model. To improve the transparency and validity of the model, all the modules in the proposed MAR model are designed to reflect their mathematical meanings. In addition, an adaptive wavelet module is also utilized to enhance the flexibility of the model. To optimize the model, an iterative algorithm is developed. The evaluation on both synthetic and real clinical datasets consistently confirms the superior performance of the proposed method over the competing methods. Jianjia Zhang, Haiyang Mao, Dingyue Chang, Hengyong Yu, Weiwen Wu, Dinggang Shen |
IEEE Trans. Medical Imaging | 6 |
| 2024 | An Anatomy- and Topology-Preserving Framework for Coronary Artery SegmentationabstractCoronary artery segmentation is critical for coronary artery disease diagnosis but challenging due to its tortuous course with numerous small branches and inter-subject variations. Most existing studies ignore important anatomical information and vascular topologies, leading to less desirable segmentation performance that usually cannot satisfy clinical demands. To deal with these challenges, in this paper we propose an anatomy- and topology-preserving two-stage framework for coronary artery segmentation. The proposed framework consists of an anatomical dependency encoding (ADE) module and a hierarchical topology learning (HTL) module for coarse-to-fine segmentation, respectively. Specifically, the ADE module segments four heart chambers and aorta, and thus five distance field maps are obtained to encode distance between chamber surfaces and coarsely segmented coronary artery. Meanwhile, ADE also performs coronary artery detection to crop region-of-interest and eliminate foreground-background imbalance. The follow-up HTL module performs fine segmentation by exploiting three hierarchical vascular topologies, i.e., key points, centerlines, and neighbor connectivity using a multi-task learning scheme. In addition, we adopt a bottom-up attention interaction (BAI) module to integrate the feature representations extracted across hierarchical topologies. Extensive experiments on public and in-house datasets show that the proposed framework achieves state-of-the-art performance for coronary artery segmentation. Xiao Zhang 0028, Kaicong Sun, Dijia Wu, Xiaosong Xiong, Jiameng Liu, Linlin Yao, Shufang Li, Jun Feng 0003, Dinggang Shen |
IEEE Trans. Medical Imaging | 10 |
| 2024 | ChatCAD+: Toward a Universal and Reliable Interactive CAD Using LLMsabstractThe integration of Computer-Aided Diagnosis (CAD) with Large Language Models (LLMs) presents a promising frontier in clinical applications, notably in automating diagnostic processes akin to those performed by radiologists and providing consultations similar to a virtual family doctor. Despite the promising potential of this integration, current works face at least two limitations: (1) From the perspective of a radiologist, existing studies typically have a restricted scope of applicable imaging domains, failing to meet the diagnostic needs of different patients. Also, the insufficient diagnostic capability of LLMs further undermine the quality and reliability of the generated medical reports. (2) Current LLMs lack the requisite depth in medical expertise, rendering them less effective as virtual family doctors due to the potential unreliability of the advice provided during patient consultations. To address these limitations, we introduce ChatCAD+, to be universal and reliable. Specifically, it is featured by two main modules: (1) Reliable Report Generation and (2) Reliable Interaction. The Reliable Report Generation module is capable of interpreting medical images from diverse domains and generate high-quality medical reports via our proposed hierarchical in-context learning. Concurrently, the interaction module leverages up-to-date information from reputable medical websites to provide reliable medical advice. Together, these designed modules synergize to closely align with the expertise of human medical professionals, offering enhanced consistency and reliability for interpretation and advice. The source code is available at GitHub. Zihao Zhao 0002, Sheng Wang 0014, Jinchen Gu, Yitao Zhu, Lanzhuju Mei, Zixu Zhuang, Zhiming Cui 0001, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 9 |
| 2024 | Hierarchical Graph Convolutional Network Built by Multiscale Atlases for Brain Disorder Diagnosis Using Functional ConnectivityabstractFunctional connectivity network (FCN) data from functional magnetic resonance imaging (fMRI) is increasingly used for the diagnosis of brain disorders. However, state-of-the-art studies used to build the FCN using a single brain parcellation atlas at a certain spatial scale, which largely neglected functional interactions across different spatial scales in hierarchical manners. In this study, we propose a novel framework to perform multiscale FCN analysis for brain disorder diagnosis. We first use a set of well-defined multiscale atlases to compute multiscale FCNs. Then, we utilize biologically meaningful brain hierarchical relationships among the regions in multiscale atlases to perform nodal pooling across multiple spatial scales, namely "Atlas-guided Pooling (AP)." Accordingly, we propose a multiscale-atlases-based hierarchical graph convolutional network (MAHGCN), built on the stacked layers of graph convolution and the AP, for a comprehensive extraction of diagnostic information from multiscale FCNs. Experiments on neuroimaging data from 1792 subjects demonstrate the effectiveness of our proposed method in the diagnoses of Alzheimer's disease (AD), the prodromal stage of AD [i.e., mild cognitive impairment (MCI)], as well as autism spectrum disorder (ASD), with the accuracy of 88.9%, 78.6%, and 72.7%, respectively. All results show significant advantages of our proposed method over other competing methods. This study not only demonstrates the feasibility of brain disorder diagnosis using resting-state fMRI empowered by deep learning but also highlights that the functional interactions in the multiscale brain hierarchy are worth being explored and integrated into deep learning network architectures for a better understanding of the neuropathology of brain disorders. The codes for MAHGCN are publicly available at "https://github.com/MianxinLiu/MAHGCN-code." Mianxin Liu, Han Zhang 0002, Feng Shi 0001, Dinggang Shen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Editorial Special Issue on Explainable and Generalizable Deep Learning for Medical ImagingabstractThe rapid advancements in deep learning technologies have profoundly influenced the field of medical image analysis, yet their full integration into clinical radiology practices has not progressed as quickly as expected. A significant hurdle to their widespread adoption among radiologists and clinicians is the prevailing lack of trust and confidence in the outcomes produced by these technologies. This concern primarily stems from concerns regarding the explainability and generalizability of deep learning models within the realm of medical imaging. As part of the responses from the Medical Image Analysis Community to address these critical issues, we organized the IEEE Transactions on Neural Networks and Learning Systems (TNNLS) Special Issue on explainable and generalizable deep learning for medical imaging. This IEEE TNNLS Special Issue calls for original and innovative methodological contributions that aim to address the key challenges on explainability and generalizability of deep learning for medical imaging. This IEEE TNNLS Special Issue emphasizes the research and advanced development of the technical aspects of new image analysis methodologies, and all the developed new methods should also be evaluated or validated on real and large-scale medical imaging data. Tianming Liu 0001, Dajiang Zhu, Fei Wang 0001, Islem Rekik, Xia Ben Hu, Dinggang Shen |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Rectify ViT Shortcut Learning by Visual SaliencyabstractShortcut learning in deep learning models occurs when unintended features are prioritized, resulting in degenerated feature representations and reduced generalizability and interpretability. However, shortcut learning in the widely used vision transformer (ViT) framework is largely unknown. Meanwhile, introducing domain-specific knowledge is a major approach to rectifying the shortcuts that are predominated by background-related factors. For example, eye-gaze data from radiologists are effective human visual prior knowledge that has the great potential to guide the deep learning models to focus on meaningful foreground regions. However, obtaining eye-gaze data can still sometimes be time-consuming, labor-intensive, and even impractical. In this work, we propose a novel and effective saliency-guided ViT (SGT) model to rectify shortcut learning in ViT with the absence of eye-gaze data. Specifically, a computational visual saliency model (either pretrained or fine-tuned) is adopted to predict saliency maps for input image samples. Then, the saliency maps are used to filter the most informative image patches. Considering that this filter operation may lead to global information loss, we further introduce a residual connection that calculates the self-attention across all the image patches. The experiment results on natural and medical image datasets show that our SGT framework can effectively learn and leverage human prior knowledge without eye-gaze data and achieves much better performance than baselines. Meanwhile, it successfully rectifies the harmful shortcut learning and significantly improves the interpretability of the ViT model, demonstrating the promise of transferring human prior knowledge derived visual saliency in rectifying shortcut learning. Chong Ma 0004, Lin Zhao 0004, Yuzhong Chen 0002, Lei Guo 0002, Xintao Hu, Dinggang Shen, Xi Jiang 0001, Tianming Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Hierarchical Organ-Aware Total-Body Standard-Dose PET Reconstruction From Low-Dose PET and CT ImagesabstractPositron emission tomography (PET) is an important functional imaging technology in early disease diagnosis. Generally, the gamma ray emitted by standard-dose tracer inevitably increases the exposure risk to patients. To reduce dosage, a lower dose tracer is often used and injected into patients. However, this often leads to low-quality PET images. In this article, we propose a learning-based method to reconstruct total-body standard-dose PET (SPET) images from low-dose PET (LPET) images and corresponding total-body computed tomography (CT) images. Different from previous works focusing only on a certain part of human body, our framework can hierarchically reconstruct total-body SPET images, considering varying shapes and intensity distributions of different body parts. Specifically, we first use one global total-body network to coarsely reconstruct total-body SPET images. Then, four local networks are designed to finely reconstruct head-neck, thorax, abdomen-pelvic, and leg parts of human body. Moreover, to enhance each local network learning for the respective local body part, we design an organ-aware network with a residual organ-aware dynamic convolution (RO-DC) module by dynamically adapting organ masks as additional inputs. Extensive experiments on 65 samples collected from uEXPLORER PET/CT system demonstrate that our hierarchical framework can consistently improve the performance of all body parts, especially for total-body PET images with PSNR of 30.6 dB, outperforming the state-of-the-art methods in SPET image reconstruction. Zhiming Cui 0001, Caiwen Jiang, Fei Gao 0010, Dinggang Shen |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | 3D Structure-guided Network for Tooth Alignment in 2D Photograph
Yulong Dou, Lanzhuju Mei, Dinggang Shen, Zhiming Cui 0001 |
BMVC | 3 |
| 2023 | TriDo-Former: A Triple-Domain Transformer for Direct PET Reconstruction from Low-Dose Sinograms
Pinxian Zeng, Xinyi Zeng, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015, Dinggang Shen |
MICCAI (10) | 8 |
| 2023 | Contrastive Diffusion Model with Auxiliary Guidance for Coarse-to-Fine PET Reconstruction
Zeyu Han, Luping Zhou, Binyu Yan, Jiliu Zhou, Yan Wang 0015, Dinggang Shen |
MICCAI (10) | 8 |
| 2023 | Mammo-Net: Integrating Gaze Supervision and Interactive Information in Multi-view Mammogram Classification
Changkai Ji, Changde Du, Sheng Wang 0014, Chong Ma 0004, Jiaming Xie, Huiguang He, Dinggang Shen |
MICCAI (7) | 9 |
| 2023 | PET-Diffusion: Unsupervised PET Enhancement Based on the Latent Diffusion Model
Caiwen Jiang, Yongsheng Pan, Mianxin Liu, Lei Ma 0006, Xiao Zhang 0028, Jiameng Liu, Xiaosong Xiong, Dinggang Shen |
MICCAI (1) | 8 |
| 2023 | Developing Large Pre-trained Model for Breast Tumor Segmentation from Ultrasound Images
Meiyu Li, Kaicong Sun, Yuning Gu, Kai Zhang 0039, Yiqun Sun, Zhenhui Li, Dinggang Shen |
MICCAI (7) | 7 |
| 2023 | Adult-Like Phase and Multi-scale Assistance for Isointense Infant Brain Tissue Segmentation
Jiameng Liu, Feihong Liu, Kaicong Sun, Mianxin Liu, Yuyan Ge, Dinggang Shen |
MICCAI (4) | 7 |
| 2023 | Development and Fast Transferring of General Connectivity-Based Diagnosis Model to New Brain Disorders with Adaptive Graph Meta-Learner
Mianxin Liu, Yuanwang Zhang, Dinggang Shen |
MICCAI (8) | 4 |
| 2023 | HC-Net: Hybrid Classification Network for Automatic Periodontal Disease Diagnosis
Lanzhuju Mei, Yu Fang 0008, Zhiming Cui 0001, Nizhuan Wang 0001, Xuming He 0001, Yiqiang Zhan, Xiang Sean Zhou, Maurizio Tonetti, Dinggang Shen |
MICCAI (6) | 10 |
| 2023 | Revealing Anatomical Structures in PET to Generate CT for Attenuation Correction
Yongsheng Pan, Feihong Liu, Caiwen Jiang, Yong Xia 0001, Dinggang Shen |
MICCAI (10) | 6 |
| 2023 | CorSegRec: A Topology-Preserving Scheme for Extracting Fully-Connected Coronary Arteries from CT Angiography
Yuehui Qiu, Pei Dong, Dijia Wu, Xinnian Yang, Qingqi Hong, Dinggang Shen |
MICCAI (3) | 8 |
| 2023 | Multi-view Vertebra Localization and Identification from CT Images
Han Wu 0007, Yu Fang 0008, Nizhuan Wang 0001, Zhiming Cui 0001, Dinggang Shen |
MICCAI (5) | 7 |
| 2023 | NeuroExplainer: Fine-Grained Attention Decoding to Uncover Cortical Development Patterns of Preterm Infants
Chenyu Xue 0004, Fan Wang 0023, Yuanzhuo Zhu, Deyu Meng, Dinggang Shen, Chunfeng Lian |
MICCAI (2) | 6 |
| 2023 | SPR-Net: Structural Points Based Registration for Coronary Arteries Across Systolic and Diastolic Phases
Xiao Zhang 0028, Feihong Liu, Yuning Gu, Xiaosong Xiong, Caiwen Jiang, Jun Feng 0003, Dinggang Shen |
MICCAI (7) | 7 |
| 2023 | HENet: Hierarchical Enhancement Network for Pulmonary Vessel Segmentation in Non-contrast CT Images
Xiao Zhang 0028, Dongdong Gu, Sheng Wang 0014, Jiayu Huo, Zhihao Jiang 0001, Feng Shi 0001, Zhong Xue, Yiqiang Zhan, Xi Ouyang, Dinggang Shen |
MICCAI (3) | 12 |
| 2023 | CAS-Net: Cross-View Aligned Segmentation by Graph Representation of Knees
Zixu Zhuang, Xin Wang 0125, Sheng Wang 0014, Zhenrong Shen 0001, Xiangyu Zhao 0003, Mengjun Liu, Zhong Xue, Dinggang Shen, Lichi Zhang, Qian Wang 0001 |
MICCAI (4) | 8 |
| 2023 | AutoEncoder-Driven Multimodal Collaborative Learning for Medical Image Synthesis
Bing Cao 0002, Zhiwei Bi, Qinghua Hu, Han Zhang 0002, Nannan Wang 0001, Xinbo Gao 0001, Dinggang Shen |
Int. J. Comput. Vis. | 7 |
| 2023 | TransDose: Transformer-based radiotherapy dose prediction from CT images guided by super-pixel-level GCN classification
Zhengyang Jiao, Xingchen Peng, Yan Wang 0015, Jianghong Xiao, Dong Nie, Xi Wu 0004, Xin Wang 0045, Jiliu Zhou, Dinggang Shen |
Medical Image Anal. | 9 |
| 2023 | ClusterSeg: A crowd cluster pinpointed nucleus segmentation framework with cross-modality datasets
Jing Ke, Yizhou Lu, Yiqing Shen 0003, Junchao Zhu, Yijin Zhou, Jinghan Huang 0002, Jieteng Yao, Xiaoyao Liang, Yi Guo 0001, Zhonghua Wei, Fusong Jiang, Dinggang Shen |
Medical Image Anal. | 14 |
| 2023 | Bidirectional prediction of facial and bony shapes for orthognathic surgical planning
Lei Ma 0006, Chunfeng Lian, Daeseung Kim, Deqiang Xiao, Dongming Wei, Tianshu Kuang, Maryam Ghanbari, Guoshi Li, Jaime Gateno, Steve G. Shen, Li Wang 0026, Dinggang Shen, James J. Xia, Pew-Thian Yap |
Medical Image Anal. | 13 |
| 2023 | Image synthesis with disentangled attributes for chest X-ray nodule augmentation and detection
Zhenrong Shen 0001, Xi Ouyang, Bin Xiao 0010, Jie-Zhi Cheng, Dinggang Shen, Qian Wang 0001 |
Medical Image Anal. | 5 |
| 2023 | TS-DSANN: Texture and shape focused dual-stream attention neural network for benign-malignant diagnosis of thyroid nodules in ultrasound images
Lu Tang 0001, Chuangeng Tian, Zhiming Cui 0001, Dinggang Shen |
Medical Image Anal. | 7 |
| 2023 | A-GCL: Adversarial graph contrastive learning for fMRI analysis to diagnose neurodevelopmental disorders
Xiang Chen 0031, Bohan Ren, Haibo Yang 0002, Xi Jiang 0001, Dinggang Shen, Yuan Zhou 0004, Xiao-Yong Zhang |
Medical Image Anal. | 8 |
| 2023 | ML-DSVM+: A meta-learning based deep SVM+ for computer-aided diagnosis
Xiangmin Han, Jun Wang 0024, Shihui Ying, Jun Shi 0004, Dinggang Shen |
Pattern Recognit. | 5 |
| 2023 | Cross-level Feature Aggregation Network for Polyp Segmentation
Tao Zhou 0002, Yi Zhou 0007, Kelei He, Chen Gong 0002, Jian Yang 0003, Huazhu Fu, Dinggang Shen |
Pattern Recognit. | 7 |
| 2023 | Semi-Cycled Generative Adversarial Networks for Real-World Face Super-ResolutionabstractReal-world face super-resolution (SR) is a highly ill-posed image restoration task. The fully-cycled Cycle-GAN architecture is widely employed to achieve promising performance on face SR, but is prone to produce artifacts upon challenging cases in real-world scenarios, since joint participation in the same degradation branch will impact final performance due to huge domain gap between real-world and synthetic LR ones obtained by generators. To better exploit the powerful generative capability of GAN for real-world face SR, in this paper, we establish two independent degradation branches in the forward and backward cycle-consistent reconstruction processes, respectively, while the two processes share the same restoration branch. Our Semi-Cycled Generative Adversarial Networks (SCGAN) is able to alleviate the adverse effects of the domain gap between the real-world LR face images and the synthetic LR ones, and to achieve accurate and robust face SR performance by the shared restoration branch regularized by both the forward and backward cycle-consistent learning processes. Experiments on two synthetic and two real-world datasets demonstrate that, our SCGAN outperforms the state-of-the-art methods on recovering the face structures/details and quantitative metrics for real-world face SR. The code will be publicly released at https://github.com/HaoHou-98/SCGAN. Hao Hou, Jun Xu 0019, Yingkun Hou, Xiaotao Hu, Benzheng Wei, Dinggang Shen |
IEEE Trans. Image Process. | 6 |
| 2023 | Brain Status Transferring Generative Adversarial Network for Decoding Individualized Atrophy in Alzheimer's DiseaseabstractDeep learning has been widely investigated in brain image computational analysis for diagnosing brain diseases such as Alzheimer's disease (AD). Most of the existing methods built end-to-end models to learn discriminative features by group-wise analysis. However, these methods cannot detect pathological changes in each subject, which is essential for the individualized interpretation of disease variances and precision medicine. In this article, we propose a brain status transferring generative adversarial network (BrainStatTrans-GAN) to generate corresponding healthy images of patients, which are further used to decode individualized brain atrophy. The BrainStatTrans-GAN consists of generator, discriminator, and status discriminator. First, a normative GAN is built to generate healthy brain images from normal controls. However, it cannot generate healthy images from diseased ones due to the lack of paired healthy and diseased images. To address this problem, a status discriminator with adversarial learning is designed in the training process to produce healthy brain images for patients. Then, the residual between the generated and input images can be computed to quantify pathological brain changes. Finally, a residual-based multi-level fusion network (RMFN) is built for more accurate disease diagnosis. Compared to the existing methods, our method can model individualized brain atrophy for facilitating disease diagnosis and interpretation. Experimental results on T1-weighted magnetic resonance imaging (MRI) data of 1,739 subjects from three datasets demonstrate the effectiveness of our method. Hongrui Liu 0001, Feng Shi 0001, Dinggang Shen, Manhua Liu |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Individualized Assessment of Brain Aβ Deposition With fMRI Using Deep LearningabstractPET-based Alzheimer's disease (AD) assessment has many limitations in large-scale screening. Non-invasive techniques such as resting-state functional magnetic resonance imaging (rs-fMRI) have been proven valuable in early AD diagnosis. This study investigated feasibility of using rs-fMRI, especially functional connectivity (FC), for individualized assessment of brain amyloid-β deposition derived from PET. We designed a graph convolutional networks (GCNs) and random forest (RF) based integrated framework for using rs-fMRI-derived multi-level FC networks to predict amyloid-β PET patterns with the OASIS-3 (N = 258) and ADNI-2 (N = 291) datasets. Our method achieved satisfactory accuracy not only in Aβ-PET grade classification (for negative, intermediate, and positive grades, with accuracy in the three-class classification as 62.8% and 64.3% on two datasets, respectively), but also in prediction of whole-brain region-level Aβ-PET standard uptake value ratios (SUVRs) (with the mean square errors as 0.039 and 0.074 for two datasets, respectively). Model interpretability examination also revealed the contributive role of the limbic network. This study demonstrated high feasibility and reproducibility of using low-cost, more accessible magnetic resonance imaging (MRI) to approximate PET-based diagnosis. Chaolin Li, Mianxin Liu, Lang Mei, Feng Shi 0001, Han Zhang 0002, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 8 |
| 2023 | Two-Stage Self-Supervised Cycle-Consistency Transformer Network for Reducing Slice Gap in MR ImagesabstractMagnetic resonance (MR) images are usually acquired with large slice gap in clinical practice, i.e., low resolution (LR) along the through-plane direction. It is feasible to reduce the slice gap and reconstruct high-resolution (HR) images with the deep learning (DL) methods. To this end, the paired LR and HR images are generally required to train a DL model in a popular fully supervised manner. However, since the HR images are hardly acquired in clinical routine, it is difficult to get sufficient paired samples to train a robust model. Moreover, the widely used convolutional Neural Network (CNN) still cannot capture long-range image dependencies to combine useful information of similar contents, which are often spatially far away from each other across neighboring slices. To this end, a Two-stage Self-supervised Cycle-consistency Transformer Network (TSCTNet) is proposed to reduce the slice gap for MR images in this work. A novel self-supervised learning (SSL) strategy is designed with two stages respectively for robust network pre-training and specialized network refinement based on a cycle-consistency constraint. A hybrid Transformer and CNN structure is utilized to build an interpolation model, which explores both local and global slice representations. The experimental results on two public MR image datasets indicate that TSCTNet achieves superior performance over other compared SSL-based algorithms. Zhiyang Lu, Jian Wang 0135, Shihui Ying, Jun Wang 0024, Jun Shi 0004, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 7 |
| 2023 | Multi-Graph Attention Networks With Bilinear Convolution for Diagnosis of SchizophreniaabstractThe explorations of brain functional connectivity (FC) network using resting-state functional magnetic resonance imaging (rs-fMRI) can provide crucial insights into discriminative analysis of neuropsychiatric disorders such as schizophrenia (SZ). Graph attention network (GAT), which could capture the local stationary on the network topology and aggregate the features of neighboring nodes, has advantages in learning the feature representation of brain regions. However, GAT only can obtain the node-level features that reflect local information, ignoring the spatial information within the connectivity-based features that proved to be important for SZ diagnosis. In addition, existing graph learning techniques usually rely on a single graph topology to represent neighborhood information, and only consider a single correlation measure for connectivity features. Comprehensive analysis of multiple graph topologies and multiple measures of FC can leverage their complementary information that may contribute to identifying patients. In this paper, we propose a multi-graph attention network (MGAT) with bilinear convolution (BC) neural network framework for SZ diagnosis and functional connectivity analysis. Besides multiple correlation measures to construct connectivity networks from different perspectives, we further propose two different graph construction methods to capture both the low- and high-level graph topologies, respectively. Especially, the MGAT module is developed to learn multiple node interaction features on each graph topology, and the BC module is utilized to learn the spatial connectivity features of the brain network for disease prediction. Importantly, the rationality and advantages of our proposed method can be validated by the experiments on SZ identification. Therefore, we speculate that this framework may also be potentially used as a diagnostic tool for other neuropsychiatric disorders. Renping Yu, Xuan Fei, Mingming Chen 0005, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | Outcome Prediction of Unconscious Patients Based on Weighted Sparse Brain Network ConstructionabstractIt is quite challenging to establish a prompt and reliable prognosis assessment for acquired brain injury (ABI) patients with persistent severe disorders of consciousness (DOC) like unconscious comatose and unresponsive wakefulness syndrome (a.k.a., vegetative state). Recent advances in brain functional imaging and functional net-work analysis have demonstrated its potential in determining the consciousness level and prognostic outcome for ABI patients with DOC. However, the diagnostic and prognostic usefulness of the whole-brain functional connectome based on advanced machine learning techniques has not been fully evaluated. The first aim of this study is to predict the outcome of individual unconscious ABI patients during a three-month follow-up. The second aim is to conduct precise individualized differentiation among different consciousness levels for exploring the neurobiological mechanisms underlying DOC. Based on resting-state fMRI, we construct large-scale functional networks by using a weighted sparse model, which ensures sparsity and interpretability by preserving strong functional connections. The functional connection strengths are exploited as features for outcome prediction and consciousness level differentiation. We achieve significantly improved consciousness level classification (accuracy: 84.78%) and recovery outcome prediction (accuracy: 89.74%) compared to other network construction methods. More importantly, we reveal the contributive connections across the entire brain in both tasks. These connections could serve as the potential biomarkers for better understanding of consciousness and further provide new insight into the development of diagnostic, prognostic, and effective therapeutic guidelines for ABI patients with DOC. Renping Yu, Han Zhang 0002, Xuehai Wu, Xuan Fei, Zengxin Qi, Di Zang, Weijun Tang, Ying Mao 0002, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 11 |
| 2023 | TW-Net: Transformer Weighted Network for Neonatal Brain MRI SegmentationabstractAccurate neonatal brain MRI segmentation is valuable for investigating brain growth patterns and tracking the progression of neurodevelopmental disorders. However, it is a challenging task to use intensity-based methods to segment neonatal brain structures because of small contrast differences between brain regions caused by the inherent myelination process. Although convolutional neural networks offer the potential to segment brain structures in an intensity-independent manner, they suffer from lack of in-plane long-range dependency which is essential for the segmentation. To solve this problem, we propose a novel Transformer-Weighted network (TW-Net) to incorporate in-plane long-range dependency information. TW-Net employs a conventional encoder-decoder architecture with a Transformer module in the middle. The Transformer module uses a rotate-and-flip layer to better calculate the similarity between two patches in a slice to leverage similar patterns of geometrical and texture features within brain structures. In addition, a deep supervision module and squeeze-and-excitation blocks are introduced to incorporate boundary information of brain structures. Compared with state-of-the-art deep learning algorithms, TW-Net outperforms these methods for multiple-label tasks in 2D and 2.5D configurations on two independent public datasets, demonstrating that TW-Net is a promising method for neonatal brain MRI segmentation. Bohan Ren, Haibo Yang 0002, Xiaoyang Han, Xiang Chen 0031, Yuan Zhou 0004, Dinggang Shen, Xiao-Yong Zhang |
IEEE J. Biomed. Health Informatics | 8 |
| 2023 | A Federated Learning System for Histopathology Image Analysis With an Orchestral Stain-Normalization GANabstractCurrently, data-driven based machine learning is considered one of the best choices in clinical pathology analysis, and its success is subject to the sufficiency of digitized slides, particularly those with deep annotations. Although centralized training on a large data set may be more reliable and more generalized, the slides to the examination are more often than not collected from many distributed medical institutes. This brings its own challenges, and the most important is the assurance of privacy and security of incoming data samples. In the discipline of histopathology image, the universal stain-variation issue adds to the difficulty of an automatic system as different clinical institutions provide distinct stain styles. To address these two important challenges in AI-based histopathology diagnoses, this work proposes a novel conditional Generative Adversarial Network (GAN) with one orchestration generator and multiple distributed discriminators, to cope with multiple-client based stain-style normalization. Implemented within a Federated Learning (FL) paradigm, this framework well preserves data privacy and security. Additionally, the training consistency and stability of the distributed system are further enhanced by a novel temporal self-distillation regularization scheme. Empirically, on large cohorts of histopathology datasets as a benchmark, the proposed model matches the performance of conventional centralized learning very closely. It also outperforms state-of-the-art stain-style transfer methods on the downstream Federated Learning image classification task, with an accuracy increase of over 20.0% in comparison to the baseline classification model. Yiqing Shen 0003, Arcot Sowmya, Yulin Luo, Xiaoyao Liang, Dinggang Shen, Jing Ke |
IEEE Trans. Medical Imaging | 5 |
| 2023 | Semi-Supervised Standard-Dose PET Image Generation via Region-Adaptive Normalization and Structural Consistency ConstraintabstractPositron Emission Tomography (PET) is an important nuclear medical imaging technique, and has been widely used in clinical applications, e.g., tumor detection and brain disease diagnosis. As PET imaging could put patients at risk of radiation, the acquisition of high-quality PET images with standard-dose tracers should be cautious. However, if dose is reduced in PET acquisition, the imaging quality could become worse and thus may not meet clinical requirement. To safely reduce the tracer dose and also maintain high quality of PET imaging, we propose a novel and effective approach to estimate high-quality Standard-dose PET (SPET) images from Low-dose PET (LPET) images. Specifically, to fully utilize both the rare paired and the abundant unpaired LPET and SPET images, we propose a semi-supervised framework for network training. Meanwhile, based on this framework, we further design a Region-adaptive Normalization (RN) and a structural consistency constraint to track the task-specific challenges. RN performs region-specific normalization in different regions of each PET image to suppress negative impact of large intensity variation across different regions, while the structural consistency constraint maintains structural details during the generation of SPET images from LPET images. Experiments on real human chest-abdomen PET images demonstrate that our proposed approach achieves state-of-the-art performance quantitatively and qualitatively. Caiwen Jiang, Yongsheng Pan, Zhiming Cui 0001, Dong Nie, Dinggang Shen |
IEEE Trans. Medical Imaging | 5 |
| 2023 | Artifact Detection and Restoration in Histology Images With Stain-Style and Structural PreservationabstractThe artifacts in histology images may encumber the accurate interpretation of medical information and cause misdiagnosis. Accordingly, prepending manual quality control of artifacts considerably decreases the degree of automation. To close this gap, we propose a methodical pre-processing framework to detect and restore artifacts, which minimizes their impact on downstream AI diagnostic tasks. First, the artifact recognition network AR-Classifier first differentiates common artifacts from normal tissues, e.g., tissue folds, marking dye, tattoo pigment, spot, and out-of-focus, and also catalogs artifact patches by their restorability. Then, the succeeding artifact restoration network AR-CycleGAN performs de-artifact processing where stain styles and tissue structures can be maximally retained. We construct a benchmark for performance evaluation, curated from both clinically collected WSIs and public datasets of colorectal and breast cancer. The functional structures are compared with state-of-the-art methods, and also comprehensively evaluated by multiple metrics across multiple tasks, including artifact classification, artifact restoration, downstream diagnostic tasks of tumor classification and nuclei segmentation. The proposed system allows full automation of deep learning based histology image analysis without human intervention. Moreover, the structure-independent characteristic enables its processing with various artifact subtypes. The source code and data in this research are available at https://github.com/yunboer/AR-classifier-and-AR-CycleGAN. Jing Ke, Kai Liu 0034, Yuxiang Sun 0004, Yuying Xue, Jiaxuan Huang, Yizhou Lu, Yaobing Chen, Xiaodan Han, Yiqing Shen 0003, Dinggang Shen |
IEEE Trans. Medical Imaging | 11 |
| 2023 | A Hierarchical Graph V-Net With Semi-Supervised Pre-Training for Histological Image Based Breast Cancer ClassificationabstractNumerous patch-based methods have recently been proposed for histological image based breast cancer classification. However, their performance could be highly affected by ignoring spatial contextual information in the whole slide image (WSI). To address this issue, we propose a novel hierarchical Graph V-Net by integrating 1) patch-level pre-training and 2) context-based fine-tuning, with a hierarchical graph network. Specifically, a semi-supervised framework based on knowledge distillation is first developed to pre-train a patch encoder for extracting disease-relevant features. Then, a hierarchical Graph V-Net is designed to construct a hierarchical graph representation from neighboring/similar individual patches for coarse-to-fine classification, where each graph node (corresponding to one patch) is attached with extracted disease-relevant features and its target label during training is the average label of all pixels in the corresponding patch. To evaluate the performance of our proposed hierarchical Graph V-Net, we collect a large WSI dataset of 560 WSIs, with 30 labeled WSIs from the BACH dataset (through our further refinement), 30 labeled WSIs and 500 unlabeled WSIs from Yunnan Cancer Hospital. Those 500 unlabeled WSIs are employed for patch-level pre-training to improve feature representation, while 60 labeled WSIs are used to train and test our proposed hierarchical Graph V-Net. Both comparative assessment and ablation studies demonstrate the superiority of our proposed hierarchical Graph V-Net over state-of-the-art methods in classifying breast cancer from WSIs. The source code and our annotations for the BACH dataset have been released at https://github.com/lyhkevin/Graph-V-Net. Yonghao Li, Yiqing Shen 0003, Shujie Song, Zhenhui Li, Jing Ke, Dinggang Shen |
IEEE Trans. Medical Imaging | 7 |
| 2023 | Multi-Scale Transformer Network With Edge-Aware Pre-Training for Cross-Modality MR Image SynthesisabstractCross-modality magnetic resonance (MR) image synthesis can be used to generate missing modalities from given ones. Existing (supervised learning) methods often require a large number of paired multi-modal data to train an effective synthesis model. However, it is often challenging to obtain sufficient paired data for supervised training. In reality, we often have a small number of paired data while a large number of unpaired data. To take advantage of both paired and unpaired data, in this paper, we propose a Multi-scale Transformer Network (MT-Net) with edge-aware pre-training for cross-modality MR image synthesis. Specifically, an Edge-preserving Masked AutoEncoder (Edge-MAE) is first pre-trained in a self-supervised manner to simultaneously perform 1) image imputation for randomly masked patches in each image and 2) whole edge map estimation, which effectively learns both contextual and structural information. Besides, a novel patch-wise loss is proposed to enhance the performance of Edge-MAE by treating different masked patches differently according to the difficulties of their respective imputations. Based on this proposed pre-training, in the subsequent fine-tuning stage, a Dual-scale Selective Fusion (DSF) module is designed (in our MT-Net) to synthesize missing-modality images by integrating multi-scale features extracted from the encoder of the pre-trained Edge-MAE. Furthermore, this pre-trained encoder is also employed to extract high-level features from the synthesized image and corresponding ground-truth image, which are required to be similar (consistent) in the training. Experimental results show that our MT-Net achieves comparable performance to the competing methods even using 70% of all available paired data. Our code will be released at https://github.com/lyhkevin/MT-Net. Yonghao Li, Tao Zhou 0002, Kelei He, Yi Zhou 0007, Dinggang Shen |
IEEE Trans. Medical Imaging | 5 |
| 2023 | Structural Attention Graph Neural Network for Diagnosis and Prediction of COVID-19 SeverityabstractWith rapid worldwide spread of Coronavirus Disease 2019 (COVID-19), jointly identifying severe COVID-19 cases from mild ones and predicting the conversion time (from mild to severe) is essential to optimize the workflow and reduce the clinician's workload. In this study, we propose a novel framework for COVID-19 diagnosis, termed as Structural Attention Graph Neural Network (SAGNN), which can combine the multi-source information including features extracted from chest CT, latent lung structural distribution, and non-imaging patient information to conduct diagnosis of COVID-19 severity and predict the conversion time from mild to severe. Specifically, we first construct a graph to incorporate structural information of the lung and adopt graph attention network to iteratively update representations of lung segments. To distinguish different infection degrees of left and right lungs, we further introduce a structural attention mechanism. Finally, we introduce demographic information and develop a multi-task learning framework to jointly perform both tasks of classification and regression. Experiments are conducted on a real dataset with 1687 chest CT scans, which includes 1328 mild cases and 359 severe cases. Experimental results show that our method achieves the best classification (e.g., 86.86% in terms of Area Under Curve) and regression (e.g., 0.58 in terms of Correlation Coefficient) performance, compared with other comparison methods. Yanbei Liu, Henan Li, Tao Luo 0010, Changqing Zhang 0002, Zhitao Xiao, Ying Wei 0009, Yaozong Gao, Feng Shi 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 10 |
| 2023 | TMM-Nets: Transferred Multi- to Mono-Modal Generation for Lupus Retinopathy DiagnosisabstractRare diseases, which are severely underrepresented in basic and clinical research, can particularly benefit from machine learning techniques. However, current learning-based approaches usually focus on either mono-modal image data or matched multi-modal data, whereas the diagnosis of rare diseases necessitates the aggregation of unstructured and unmatched multi-modal image data due to their rare and diverse nature. In this study, we therefore propose diagnosis-guided multi-to-mono modal generation networks (TMM-Nets) along with training and testing procedures. TMM-Nets can transfer data from multiple sources to a single modality for diagnostic data structurization. To demonstrate their potential in the context of rare diseases, TMM-Nets were deployed to diagnose the lupus retinopathy (LR-SLE), leveraging unmatched regular and ultra-wide-field fundus images for transfer learning. The TMM-Nets encoded the transfer learning from diabetic retinopathy to LR-SLE based on the similarity of the fundus lesions. In addition, a lesion-aware multi-scale attention mechanism was developed for clinical alerts, enabling TMM-Nets not only to inform patient care, but also to provide insights consistent with those of clinicians. An adversarial strategy was also developed to refine multi- to mono-modal image generation based on diagnostic results and the data distribution to enhance the data augmentation performance. Compared to the baseline model, the TMM-Nets showed 35.19% and 33.56% F1 score improvements on the test and external validation sets, respectively. In addition, the TMM-Nets can be used to develop diagnostic models for other rare diseases. Ruhan Liu, Tianqin Wang, Huating Li, Ping Zhang 0016, Xiaokang Yang 0001, Dinggang Shen, Bin Sheng 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2023 | Eye-Gaze-Guided Vision Transformer for Rectifying Shortcut LearningabstractLearning harmful shortcuts such as spurious correlations and biases prevents deep neural networks from learning meaningful and useful representations, thus jeopardizing the generalizability and interpretability of the learned representation. The situation becomes even more serious in medical image analysis, where the clinical data are limited and scarce while the reliability, generalizability and transparency of the learned model are highly required. To rectify the harmful shortcuts in medical imaging applications, in this paper, we propose a novel eye-gaze-guided vision transformer (EG-ViT) model which infuses the visual attention from radiologists to proactively guide the vision transformer (ViT) model to focus on regions with potential pathology rather than spurious correlations. To do so, the EG-ViT model takes the masked image patches that are within the radiologists' interest as input while has an additional residual connection to the last encoder layer to maintain the interactions of all patches. The experiments on two medical imaging datasets demonstrate that the proposed EG-ViT model can effectively rectify the harmful shortcut learning and improve the interpretability of the model. Meanwhile, infusing the experts' domain knowledge can also improve the large-scale ViT model's performance over all compared baseline methods with limited samples available. In general, EG-ViT takes the advantages of powerful deep neural networks while rectifies the harmful shortcut learning with human expert's prior knowledge. This work also opens new avenues for advancing current artificial intelligence paradigms by infusing human intelligence. Chong Ma 0004, Lin Zhao 0004, Yuzhong Chen 0002, Sheng Wang 0014, Lei Guo 0002, Dinggang Shen, Xi Jiang 0001, Tianming Liu 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2023 | BowelNet: Joint Semantic-Geometric Ensemble Learning for Bowel Segmentation From Both Partially and Fully Labeled CT ImagesabstractAccurate bowel segmentation is essential for diagnosis and treatment of bowel cancers. Unfortunately, segmenting the entire bowel in CT images is quite challenging due to unclear boundary, large shape, size, and appearance variations, as well as diverse filling status within the bowel. In this paper, we present a novel two-stage framework, named BowelNet, to handle the challenging task of bowel segmentation in CT images, with two stages of 1) jointly localizing all types of the bowel, and 2) finely segmenting each type of the bowel. Specifically, in the first stage, we learn a unified localization network from both partially- and fully-labeled CT images to robustly detect all types of the bowel. To better capture unclear bowel boundary and learn complex bowel shapes, in the second stage, we propose to jointly learn semantic information (i.e., bowel segmentation mask) and geometric representations (i.e., bowel boundary and bowel skeleton) for fine bowel segmentation in a multi-task learning scheme. Moreover, we further propose to learn a meta segmentation network via pseudo labels to improve segmentation accuracy. By evaluating on a large abdominal CT dataset, our proposed BowelNet method can achieve Dice scores of 0.764, 0.848, 0.835, 0.774, and 0.824 in segmenting the duodenum, jejunum-ileum, colon, sigmoid, and rectum, respectively. These results demonstrate the effectiveness of our proposed BowelNet framework in segmenting the entire bowel from CT images. Chong Wang 0012, Zhiming Cui 0001, Miaofei Han, Gustavo Carneiro 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 6 |
| 2023 | Fast Multi-Contrast MRI Acquisition by Optimal Sampling of Information Complementary to Pre-Acquired MRI ContrastabstractRecent studies on multi-contrast MRI reconstruction have demonstrated the potential of further accelerating MRI acquisition by exploiting correlation between contrasts. Most of the state-of-the-art approaches have achieved improvement through the development of network architectures for fixed under-sampling patterns, without considering inter-contrast correlation in the under-sampling pattern design. On the other hand, sampling pattern learning methods have shown better reconstruction performance than those with fixed under-sampling patterns. However, most under-sampling pattern learning algorithms are designed for single contrast MRI without exploiting complementary information between contrasts. To this end, we propose a framework to optimize the under-sampling pattern of a target MRI contrast which complements the acquired fully-sampled reference contrast. Specifically, a novel image synthesis network is introduced to extract the redundant information contained in the reference contrast, which is exploited in the subsequent joint pattern optimization and reconstruction network. We have demonstrated superior performance of our learned under-sampling patterns on both public and in-house datasets, compared to the commonly used under-sampling patterns and state-of-the-art methods that jointly optimize the reconstruction network and the under-sampling patterns, up to 8-fold under-sampling factor. Xiaoxin Li 0001, Feihong Liu, Dong Nie, Pietro Liò, Haikun Qi, Dinggang Shen |
IEEE Trans. Medical Imaging | 7 |
| 2023 | TaG-Net: Topology-Aware Graph Network for Centerline-Based Vessel LabelingabstractAnatomical labeling of head and neck vessels is a vital step for cerebrovascular disease diagnosis. However, it remains challenging to automatically and accurately label vessels in computed tomography angiography (CTA) since head and neck vessels are tortuous, branched, and often spatially close to nearby vasculature. To address these challenges, we propose a novel topology-aware graph network (TaG-Net) for vessel labeling. It combines the advantages of volumetric image segmentation in the voxel space and centerline labeling in the line space, wherein the voxel space provides detailed local appearance information, and line space offers high-level anatomical and topological information of vessels through the vascular graph constructed from centerlines. First, we extract centerlines from the initial vessel segmentation and construct a vascular graph from them. Then, we conduct vascular graph labeling using TaG-Net, in which techniques of topology-preserving sampling, topology-aware feature grouping, and multi-scale vascular graph are designed. After that, the labeled vascular graph is utilized to improve volumetric segmentation via vessel completion. Finally, the head and neck vessels of 18 segments are labeled by assigning centerline labels to the refined segmentation. We have conducted experiments on CTA images of 401 subjects, and experimental results show superior vessel segmentation and labeling of our method compared to other state-of-the-art methods. Linlin Yao, Feng Shi 0001, Sheng Wang 0014, Xiao Zhang 0028, Zhong Xue, Xiaohuan Cao, Yiqiang Zhan, Lizhou Chen, Yuntian Chen, Bin Song 0002, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 12 |
| 2023 | Breast Fibroglandular Tissue Segmentation for Automated BPE Quantification With Iterative Cycle-Consistent Semi-Supervised LearningabstractBackground Parenchymal Enhancement (BPE) quantification in Dynamic Contrast-Enhanced Magnetic Resonance Imaging (DCE-MRI) plays a pivotal role in clinical breast cancer diagnosis and prognosis. However, the emerging deep learning-based breast fibroglandular tissue segmentation, a crucial step in automated BPE quantification, often suffers from limited training samples with accurate annotations. To address this challenge, we propose a novel iterative cycle-consistent semi-supervised framework to leverage segmentation performance by using a large amount of paired pre-/post-contrast images without annotations. Specifically, we design the reconstruction network, cascaded with the segmentation network, to learn a mapping from the pre-contrast images and segmentation predictions to the post-contrast images. Thus, we can implicitly use the reconstruction task to explore the inter-relationship between these two-phase images, which in return guides the segmentation task. Moreover, the reconstructed post-contrast images across multiple auto-context modeling-based iterations can be viewed as new augmentations, facilitating cycle-consistent constraints across each segmentation output. Extensive experiments on two datasets with various data distributions show great segmentation and BPE quantification accuracy compared with other state-of-the-art semi-supervised methods. Importantly, our method achieves 11.80 times of quantification accuracy improvement along with 10 times faster, compared with clinical physicians, demonstrating its potential for automated BPE quantification. The code is available at https://github.com/ZhangJD-ong/Iterative-Cycle-consistent-Semi-supervised-Learning-for-fibroglandular-tissue-segmentation. Zhiming Cui 0001, Luping Zhou, Yiqun Sun, Zhenhui Li, Zaiyi Liu, Dinggang Shen |
IEEE Trans. Medical Imaging | 7 |
| 2023 | Knee Cartilage Defect Assessment by Graph Representation and Surface ConvolutionabstractKnee osteoarthritis (OA) is the most common osteoarthritis and a leading cause of disability. Cartilage defects are regarded as major manifestations of knee OA, which are visible by magnetic resonance imaging (MRI). Thus early detection and assessment for knee cartilage defects are important for protecting patients from knee OA. In this way, many attempts have been made on knee cartilage defect assessment by applying convolutional neural networks (CNNs) to knee MRI. However, the physiologic characteristics of the cartilage may hinder such efforts: the cartilage is a thin curved layer, implying that only a small portion of voxels in knee MRI can contribute to the cartilage defect assessment; heterogeneous scanning protocols further challenge the feasibility of the CNNs in clinical practice; the CNN-based knee cartilage evaluation results lack interpretability. To address these challenges, we model the cartilages structure and appearance from knee MRI into a graph representation, which is capable of handling highly diverse clinical data. Then, guided by the cartilage graph representation, we design a non-Euclidean deep learning network with the self-attention mechanism, to extract cartilage features in the local and global, and to derive the final assessment with a visualized result. Our comprehensive experiments show that the proposed method yields superior performance in knee cartilage defect assessment, plus its convenient 3D visualization for interpretability. Zixu Zhuang, Liping Si, Sheng Wang 0014, Kai Xuan, Xi Ouyang, Yiqiang Zhan, Zhong Xue, Lichi Zhang, Dinggang Shen, Weiwu Yao, Qian Wang 0001 |
IEEE Trans. Medical Imaging | 9 |
| 2023 | Exploiting Sparse Self-Representation and Particle Swarm Optimization for CNN CompressionabstractStructured pruning has received ever-increasing attention as a method for compressing convolutional neural networks. However, most existing methods directly prune the network structure according to the statistical information of the parameters. Besides, these methods differentiate the pruning rates only in each pruning stage or even use the same pruning rate across all layers, rather than using learnable parameters. In this article, we propose a network redundancy elimination approach guided by the pruned model. Our proposed method can easily tackle multiple architectures and is scalable to the deeper neural networks because of the use of joint optimization during the pruning procedure. More specifically, we first construct a sparse self-representation for the filters or neurons of the well-trained model, which is useful for analyzing the relationship among filters. Then, we employ particle swarm optimization to learn pruning rates in a layerwise manner according to the performance of the pruned model, which can determine optimal pruning rates with the best performance of the pruned model. Under this criterion, the proposed pruning approach can remove more parameters without undermining the performance of the model. Experimental results demonstrate the effectiveness of our proposed method on different datasets and different architectures. For example, it can reduce 58.1% FLOPs for ResNet50 on ImageNet with only a 1.6% top-five error increase and 44.1% FLOPs for FCN_ResNet50 on COCO2017 with a 3% error increase, outperforming most state-of-the-art methods. Sijie Niu, Kun Gao 0002, Xizhan Gao, Hui Zhao 0009, Jiwen Dong, Yuehui Chen, Dinggang Shen |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2023 | Breast Tumor Segmentation in DCE-MRI With Tumor Sensitive SynthesisabstractSegmenting breast tumors from dynamic contrast-enhanced magnetic resonance (DCE-MR) images is a critical step for early detection and diagnosis of breast cancer. However, variable shapes and sizes of breast tumors, as well as inhomogeneous background, make it challenging to accurately segment tumors in DCE-MR images. Therefore, in this article, we propose a novel tumor-sensitive synthesis module and demonstrate its usage after being integrated with tumor segmentation. To suppress false-positive segmentation with similar contrast enhancement characteristics to true breast tumors, our tumor-sensitive synthesis module can feedback differential loss of the true and false breast tumors. Thus, by following the tumor-sensitive synthesis module after the segmentation predictions, the false breast tumors with similar contrast enhancement characteristics to the true ones will be effectively reduced in the learned segmentation model. Moreover, the synthesis module also helps improve the boundary accuracy while inaccurate predictions near the boundary will lead to higher loss. For the evaluation, we build a very large-scale breast DCE-MR image dataset with 422 subjects from different patients, and conduct comprehensive experiments and comparisons with other algorithms to justify the effectiveness, adaptability, and robustness of our proposed method. Shuai Wang 0003, Li Wang 0026, Liangqiong Qu, Fuhua Yan, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2022 | Curvature-Enhanced Implicit Function Network for High-quality Tooth Model Generation from CBCT Images
Yu Fang 0008, Zhiming Cui 0001, Lei Ma 0006, Lanzhuju Mei, Yue Zhao 0012, Zhihao Jiang 0001, Yiqiang Zhan, Yongsheng Pan, Dinggang Shen |
MICCAI (5) | 11 |
| 2022 | Classification-Aided High-Quality PET Image Synthesis via Bidirectional Contrastive GAN with Shared Information Maximization
Yuchen Fei, Chen Zu, Zhengyang Jiao, Xi Wu 0004, Jiliu Zhou, Dinggang Shen, Yan Wang 0015 |
MICCAI (6) | 6 |
| 2022 | Deep-Learning Based T1 and T2 Quantification from Undersampled Magnetic Resonance Fingerprinting Data to Track Tracer Kinetics in Small Laboratory Animals
Yuning Gu, Yongsheng Pan, Zhenghan Fang, Jingyang Zhang, Peng Xue 0005, Mianxin Liu, Yuran Zhu, Lei Ma 0006, Charlie Androjna, Dinggang Shen |
MICCAI (6) | 11 |
| 2022 | A Novel Knowledge Keeper Network for 7T-Free but 7T-Guided Brain Tissue Segmentation
Kwanseok Oh, Dinggang Shen, Heung-Il Suk |
MICCAI (5) | 3 |
| 2022 | Multimodal Brain Tumor Segmentation Using Contrastive Learning Based Feature Comparison with Monomodal Normal Brain Images
Huabing Liu, Dong Ni 0001, Dinggang Shen, Jinda Wang, Zhenyu Tang 0002 |
MICCAI (5) | 3 |
| 2022 | RandStainNA: Learning Stain-Agnostic Features from Histology Slides by Bridging Stain Augmentation and Normalization
Yiqing Shen 0003, Yulin Luo, Dinggang Shen, Jing Ke |
MICCAI (2) | 3 |
| 2022 | Deep Learning-Based Head and Neck Radiotherapy Planning Dose Prediction via Beam-Wise Dose Decomposition
Bin Wang 0068, Lanzhuju Mei, Zhiming Cui 0001, Xuanang Xu, Qianjin Feng 0003, Dinggang Shen |
MICCAI (8) | 7 |
| 2022 | 3D CVT-GAN: A 3D Convolutional Vision Transformer-GAN for PET Reconstruction
Pinxian Zeng, Luping Zhou, Chen Zu, Xinyi Zeng, Zhengyang Jiao, Xi Wu 0004, Jiliu Zhou, Dinggang Shen, Yan Wang 0015 |
MICCAI (6) | 8 |
| 2022 | Mapping in Cycles: Dual-Domain PET-CT Synthesis Framework with Cycle-Consistent Constraints
Zhiming Cui 0001, Caiwen Jiang, Jingyang Zhang, Fei Gao 0010, Dinggang Shen |
MICCAI (6) | 6 |
| 2022 | Learning Towards Synchronous Network Memorizability and Generalizability for Continual Segmentation Across Multiple Sites
Jingyang Zhang, Peng Xue 0005, Ran Gu, Yuning Gu, Mianxin Liu, Yongsheng Pan, Zhiming Cui 0001, Lei Ma 0006, Dinggang Shen |
MICCAI (5) | 10 |
| 2022 | Progressive Deep Segmentation of Coronary Artery via Hierarchical Topology Learning
Xiao Zhang 0028, Jingyang Zhang, Lei Ma 0006, Peng Xue 0005, Dijia Wu, Yiqiang Zhan, Jun Feng 0003, Dinggang Shen |
MICCAI (5) | 9 |
| 2022 | Local Graph Fusion of Multi-view MR Images for Knee Osteoarthritis Diagnosis
Zixu Zhuang, Sheng Wang 0014, Liping Si, Kai Xuan, Zhong Xue, Dinggang Shen, Lichi Zhang, Weiwu Yao, Qian Wang 0001 |
MICCAI (3) | 6 |
| 2022 | Forecasting Human Trajectory from Scene HistoryabstractPredicting the future trajectory of a person remains a challenging problem, due to randomness and subjectivity. However, the moving patterns of human in constrained scenario typically conform to a limited number of regularities to a certain extent, because of the scenario restrictions (\eg, floor plan, roads and obstacles) and person-person or person-object interactivity. Thus, an individual person in this scenario should follow one of the regularities as well. In other words, a person's subsequent trajectory has likely been traveled by others. Based on this hypothesis, we propose to forecast a person's future trajectory by learning from the implicit scene regularities. We call the regularities, inherently derived from the past dynamics of the people and the environment in the scene, \emph{scene history}. We categorize scene history information into two types: historical group trajectories and individual-surroundings interaction. To exploit these information for trajectory prediction, we propose a novel framework Scene History Excavating Network (SHENet), where the scene history is leveraged in a simple yet effective approach. In particular, we design two components, the group trajectory bank module to extract representative group trajectories as the candidate for future path, and the cross-modal interaction module to model the interaction between individual past trajectory and its surroundings for trajectory refinement, respectively. In addition, to mitigate the uncertainty in the evaluation, caused by the aforementioned randomness and subjectivity, we propose to include smoothness into evaluation metrics. We conduct extensive evaluations to validate the efficacy of proposed framework on ETH, UCY, as well as a new, challenging benchmark dataset PAV, demonstrating superior performance compared to state-of-the-art methods. Mancheng Meng, Ziyan Wu 0001, Terrence Chen, Xiran Cai, Xiang Sean Zhou, Fan Yang 0054, Dinggang Shen |
NeurIPS | 7 |
| 2022 | COVIDSum: A linguistically enriched SciBERT-based summarization model for COVID-19 scientific papers
Xiaoyan Cai, Sen Liu 0004, Libin Yang, Jintao Zhao, Dinggang Shen, Tianming Liu 0001 |
J. Biomed. Informatics | 6 |
| 2022 | D2FE-GAN: Decoupled dual feature extraction based GAN for MRI image synthesis
Bo Zhan, Luping Zhou, Xi Wu 0004, Yi-Fei Pu, Jiliu Zhou, Yan Wang 0015, Dinggang Shen |
Knowl. Based Syst. | 8 |
| 2022 | Auto-DenseUNet: Searchable neural network architecture for mass segmentation in 3D automated breast ultrasound
Houjin Chen, Yanfeng Li 0001, Yahui Peng, Yue Zhou 0006, Tianming Liu 0001, Dinggang Shen |
Medical Image Anal. | 8 |
| 2022 | Common feature learning for brain tumor MRI synthesis by context-aware generative adversarial network
Pu Huang 0001, Dengwang Li, Zhicheng Jiao, Dongming Wei, Bing Cao 0002, Zhanhao Mo, Qian Wang 0001, Han Zhang 0002, Dinggang Shen |
Medical Image Anal. | 9 |
| 2022 | Automatic Grading Assessments for Knee MRI Cartilage Defects via Self-ensembling Semi-supervised Learning with Dual-Consistency
Jiayu Huo, Xi Ouyang, Liping Si, Kai Xuan, Sheng Wang 0014, Weiwu Yao, Dahong Qian, Zhong Xue, Qian Wang 0001, Dinggang Shen, Lichi Zhang |
Medical Image Anal. | 12 |
| 2022 | Assessing clinical progression from subjective cognitive decline to mild cognitive impairment with incomplete multi-modal neuroimages
Yunbi Liu, Ling Yue, Shifu Xiao, Wei Yang 0006, Dinggang Shen, Mingxia Liu 0001 |
Medical Image Anal. | 5 |
| 2022 | Adaptive rectification based adversarial network with spectrum constraint for high-quality PET image synthesis
Yanmei Luo, Luping Zhou, Bo Zhan, Fei-Yue Wang 0001, Jiliu Zhou, Yan Wang 0015, Dinggang Shen |
Medical Image Anal. | 7 |
| 2022 | Multi-Class ASD Classification via Label Distribution Learning with Class-Shared and Class-Specific Decomposition
Jun Wang 0024, Fengyexin Zhang, Xiuyi Jia, Xin Wang 0084, Han Zhang 0002, Shihui Ying, Qian Wang 0001, Jun Shi 0004, Dinggang Shen |
Medical Image Anal. | 9 |
| 2022 | Semantic instance segmentation with discriminative deep supervision for medical images
Sihang Zhou 0001, Dong Nie, Ehsan Adeli-Mosabbeb, Xuhua Ren, Xinwang Liu 0002, En Zhu, Jianping Yin, Qian Wang 0001, Dinggang Shen |
Medical Image Anal. | 10 |
| 2022 | Disease-Image-Specific Learning for Diagnosis-Oriented Neuroimage Synthesis With Incomplete Multi-Modality DataabstractIncomplete data problem is commonly existing in classification tasks with multi-source data, particularly the disease diagnosis with multi-modality neuroimages, to track which, some methods have been proposed to utilize all available subjects by imputing missing neuroimages. However, these methods usually treat image synthesis and disease diagnosis as two standalone tasks, thus ignoring the specificity conveyed in different modalities, i.e., different modalities may highlight different disease-relevant regions in the brain. To this end, we propose a disease-image-specific deep learning (DSDL) framework for joint neuroimage synthesis and disease diagnosis using incomplete multi-modality neuroimages. Specifically, with each whole-brain scan as input, we first design a Disease-image-Specific Network (DSNet) with a spatial cosine module to implicitly model the disease-image specificity. We then develop a Feature-consistency Generative Adversarial Network (FGAN) to impute missing neuroimages, where feature maps (generated by DSNet) of a synthetic image and its respective real image are encouraged to be consistent while preserving the disease-image-specific information. Since our FGAN is correlated with DSNet, missing neuroimages can be synthesized in a diagnosis-oriented manner. Experimental results on three datasets suggest that our method can not only generate reasonable neuroimages, but also achieve state-of-the-art performance in both tasks of Alzheimer's disease identification and mild cognitive impairment conversion prediction. Yongsheng Pan, Mingxia Liu 0001, Yong Xia 0001, Dinggang Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Weakly Supervised Segmentation of COVID19 Infection with Scribble Annotation on CT Images
Xiaoming Liu 0004, Yaozong Gao, Kelei He, Jinshan Tang, Dinggang Shen |
Pattern Recognit. | 8 |
| 2022 | A cascaded nested network for 3T brain MR image segmentation guided by 7T labeling
Zhengwang Wu, Li Wang 0026, Toan Duc Bui, Liangqiong Qu, Pew-Thian Yap, Yong Xia 0001, Gang Li 0001, Dinggang Shen |
Pattern Recognit. | 9 |
| 2022 | Three-dimensional affinity learning based multi-branch ensemble network for breast tumor segmentation in MRI
Lei Zhou 0003, Tao Zhou 0002, Fuhua Yan, Dinggang Shen |
Pattern Recognit. | 6 |
| 2022 | Multiview Feature Learning With Multiatlas-Based Functional Connectivity Networks for MCI DiagnosisabstractFunctional connectivity (FC) networks built from resting-state functional magnetic resonance imaging (rs-fMRI) has shown promising results for the diagnosis of Alzheimer's disease and its prodromal stage, that is, mild cognitive impairment (MCI). FC is usually estimated as a temporal correlation of regional mean rs-fMRI signals between any pair of brain regions, and these regions are traditionally parcellated with a particular brain atlas. Most existing studies have adopted a predefined brain atlas for all subjects. However, the constructed FC networks inevitably ignore the potentially important subject-specific information, particularly, the subject-specific brain parcellation. Similar to the drawback of the "single view" (versus the "multiview" learning) in medical image-based classification, FC networks constructed based on a single atlas may not be sufficient to reveal the underlying complicated differences between normal controls and disease-affected patients due to the potential bias from that particular atlas. In this study, we propose a multiview feature learning method with multiatlas-based FC networks to improve MCI diagnosis. Specifically, a three-step transformation is implemented to generate multiple individually specified atlases from the standard automated anatomical labeling template, from which a set of atlas exemplars is selected. Multiple FC networks are constructed based on these preselected atlas exemplars, providing multiple views of the FC network-based feature representations for each subject. We then devise a multitask learning algorithm for joint feature selection from the constructed multiple FC networks. The selected features are jointly fed into a support vector machine classifier for multiatlas-based MCI diagnosis. Extensive experimental comparisons are carried out between the proposed method and other competing approaches, including the traditional single-atlas-based method. The results indicate that our method significantly improves the MCI classification, demonstrating its promise in the brain connectome-based individualized diagnosis of brain diseases. Yu Zhang 0009, Han Zhang 0002, Ehsan Adeli-Mosabbeb, Xiaobo Chen 0001, Mingxia Liu 0001, Dinggang Shen |
IEEE Trans. Cybern. | 6 |
| 2022 | Attention-Guided Hybrid Network for Dementia Diagnosis With Structural MR ImagesabstractDeep-learning methods (especially convolutional neural networks) using structural magnetic resonance imaging (sMRI) data have been successfully applied to computer-aided diagnosis (CAD) of Alzheimer's disease (AD) and its prodromal stage [i.e., mild cognitive impairment (MCI)]. As it is practically challenging to capture local and subtle disease-associated abnormalities directly from the whole-brain sMRI, most of those deep-learning approaches empirically preselect disease-associated sMRI brain regions for model construction. Considering that such isolated selection of potentially informative brain locations might be suboptimal, very few methods have been proposed to perform disease-associated discriminative region localization and disease diagnosis in a unified deep-learning framework. However, those methods based on task-oriented discriminative localization still suffer from two common limitations, that is: 1) identified brain locations are strictly consistent across all subjects, which ignores the unique anatomical characteristics of each brain and 2) only limited local regions/patches are used for model training, which does not fully utilize the global structural information provided by the whole-brain sMRI. In this article, we propose an attention-guided deep-learning framework to extract multilevel discriminative sMRI features for dementia diagnosis. Specifically, we first design a backbone fully convolutional network to automatically localize the discriminative brain regions in a weakly supervised manner. Using the identified disease-related regions as spatial attention guidance, we further develop a hybrid network to jointly learn and fuse multilevel sMRI features for CAD model construction. Our proposed method was evaluated on three public datasets (i.e., ADNI-1, ADNI-2, and AIBL), showing superior performance compared with several state-of-the-art methods in both tasks of AD diagnosis and MCI conversion prediction. Chunfeng Lian, Mingxia Liu 0001, Yongsheng Pan, Dinggang Shen |
IEEE Trans. Cybern. | 4 |
| 2022 | Task-Induced Pyramid and Attention GAN for Multimodal Brain Image Imputation and Classification in Alzheimer's DiseaseabstractWith the advance of medical imaging technologies, multimodal images such as magnetic resonance images (MRI) and positron emission tomography (PET) can capture subtle structural and functional changes of brain, facilitating the diagnosis of brain diseases such as Alzheimer's disease (AD). In practice, multimodal images may be incomplete since PET is often missing due to high financial costs or availability. Most of the existing methods simply excluded subjects with missing data, which unfortunately reduced the sample size. In addition, how to extract and combine multimodal features is still challenging. To address these problems, we propose a deep learning framework to integrate a task-induced pyramid and attention generative adversarial network (TPA-GAN) with a pathwise transfer dense convolution network (PT-DCN) for imputation and classification of multimodal brain images. First, we propose a TPA-GAN to integrate pyramid convolution and attention module as well as disease classification task into GAN for generating the missing PET data with their MRI. Then, with the imputed multimodal images, we build a dense convolution network with pathwise transfer blocks to gradually learn and combine multimodal features for final disease classification. Experiments are performed on ADNI-1/2 datasets to evaluate our method, achieving superior performance in image imputation and brain disease diagnosis compared to state-of-the-art methods. Feng Shi 0001, Dinggang Shen, Manhua Liu |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | GAN-Guided Deformable Attention Network for Identifying Thyroid Nodules in Ultrasound ImagesabstractEarly detection and identification of malignant thyroid nodules, a vital precursory to the treatment, is a difficult task even for experienced clinicians. Many Computer-Aided Diagnose (CAD) systems have been developed to assist clinicians in performing this task on ultrasonic images. Learning-based CAD systems for thyroid nodules generally accommodate both nodule detection/ segmentation and fine-grained classification for its malignancy, and prior researches often treat aforementioned tasks in separate stages, leading to additional computational costs. In this paper, we utilize an online class activation mapping (CAM) mechanism to guide the network to learn discriminative features for identifying thyroid nodules in ultrasound images, called CAM attention network. It takes nodule masks as localization cues for direct spatial attention of the classification module, thereby avoiding isolated training for classification. Meanwhile, we propose a deformable convolution module to add offsets to the regular grid sampling locations in the standard convolution, guiding the network to capture more discriminative features of nodule areas. Furthermore, we use a generative adversarial network (GAN)to ensure reliable deformations of nodules from the deformable convolution module. Our proposed CAM attention network has already achieved the 2nd place in the classification task of TN-SCUI 2020, a MICCAI 2020 Challenge with the largest set of thyroid nodule ultrasound images according to our knowledge. The further inclusion of our proposed GAN-guided deformable module allows for capturing more fine-grained features between benign and malignant nodules, and further improves the classification accuracy to a new state-of-the-art level. Jintao Lu, Xi Ouyang, Xueda Shen, Zhiming Cui 0001, Qian Wang 0001, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 7 |
| 2022 | Multiple B-Value Model-Based Residual Network (MORN) for Accelerated High-Resolution Diffusion-Weighted ImagingabstractSingle-Shot Echo Planar Imaging (SSEPI) based Diffusion Weighted Imaging (DWI) has shortcomings such as low resolution and severe distortions. In contrast, Multi-Shot EPI (MSEPI) provides optimal spatial resolution but increases scan time. This study proposed a Multiple b-value mOdel-based Residual Network (MORN) model to reconstruct multiple b-value high-resolution DWI from undersampled k-space data simultaneously. We incorporated Parallel Imaging (PI) into a residual U-net to reconstruct multiple b-value multi-coil data with the supervision of MUltiplexed Sensitivity-Encoding (MUSE) reconstructed Multi-Shot DWI (MSDWI). Moreover, asymmetric concatenations among different b-values and the combined loss to back propagate helped the feature transfer. After training and validation of the MORN in a dataset of 32 healthy cases, additional assessments were performed on 6 patients with different tumor types. The experimental results demonstrated that the MORN model outperformed conventional PI reconstruction (i.e. SENSE) and two state-of-the-art deep learning methods (SENSE-GAN and VSNet) in terms of PSNR (Peak Signal-to-Noise Ratio), SSIM (Structual SIMilarity) and apparent diffusion coefficient maps. In addition, using the pre-trained model under DWI, the MORN achieved consistent fractional anisotrophy and mean diffusivity reconstructed from multiple diffusion directions. Hence, the proposed method shows potential in clinical application according to the observations on tumor patients as well as images of multiple diffusion directions. Fanwen Wang, Hui Zhang 0005, Weibo Chen, Zidong Yang, Dinggang Shen, Chengyan Wang, He Wang 0016 |
IEEE J. Biomed. Health Informatics | 7 |
| 2022 | Divergent and Convergent Imaging Markers Between Bipolar and Unipolar Depression Based on Machine LearningabstractDistinguishing bipolar depression (BD) from unipolar depression (UD) based on symptoms only is challenging. Brain functional connectivity (FC), especially dynamic FC, has emerged as a promising approach to identify possible imaging markers for differentiating BD from UD. However, most of such studies utilized conventional FC and group-level statistical comparisons, which may not be sensitive enough to quantify subtle changes in the FC dynamics between BD and UD. In this paper, we present a more effective individualized differentiation model based on machine learning and the whole-brain "high-order functional connectivity (HOFC)" network. The HOFC, capturing temporal synchronization among the dynamic FC time series, a more complex "chronnectome" metric compared to the conventional FC, was used to classify 52 BD, 73 UD, and 76 healthycontrols (HC). We achieved a satisfactory accuracy (70.40%) in BD vs. UD differentiation. The resultant contributing features revealed the involvement of the coordinated flexible interactions among sensory (e.g., olfaction, vision, and audition), motor, and cognitive systems. Despite sharing common chronnectome of cognitive and affective impairments, BD and UD also demonstrated unique dynamic FC synchronization patterns. UD is more associated with abnormal visual-somatomotor inter-network connections, while BD is more related to impaired ventral attention-frontoparietal inter-network connections. Moreover, we found that the illness duration modulated the BD vs. UD separation, with the differentiation performance hampered by the secondary disease effects. Our findings suggest that BD and UD may have divergent and convergent neural substrates, which further expand our knowledge of the two different mental disorders. Huifeng Zhang, Zhen Zhou 0004, Chuangxin Wu, Meihui Qiu, Yueqi Huang, Ting Shen, Li-Ming Hsu, Han Zhang 0002, Dinggang Shen, Daihui Peng |
IEEE J. Biomed. Health Informatics | 13 |
| 2022 | Cross-Model Attention-Guided Tumor Segmentation for 3D Automated Breast Ultrasound (ABUS) ImagesabstractTumor segmentation in 3D automated breast ultrasound (ABUS) plays an important role in breast disease diagnosis and surgical planning. However, automatic segmentation of tumors in 3D ABUS images is still challenging, due to the large tumor shape and size variations, and uncertain tumor locations among patients. In this paper, we develop a novel cross-model attention-guided tumor segmentation network with a hybrid loss for 3D ABUS images. Specifically, we incorporate the tumor location into a segmentation network by combining an improved 3D Mask R-CNN head into V-Net as an end-to-end architecture. Furthermore, we introduce a cross-model attention mechanism that is able to aggregate the segmentation probability map from the improved 3D Mask R-CNN to each feature extraction level in the V-Net. Then, we design a hybrid loss to balance the contribution of each part in the proposed cross-model segmentation network. We conduct extensive experiments on 170 3D ABUS from 107 patients. Experimental results show that our method outperforms other state-of-the-art methods, by achieving the Dice similarity coefficient (DSC) of 64.57%, Jaccard coefficient (JC) of 53.39%, recall (REC) of 64.43%, precision (PRE) of 74.51%, 95th Hausdorff distance (95HD) of 11.91 mm, and average surface distance (ASD) of 4.63 mm. Our code will be available online (https://github.com/zhouyuegithub/CMVNet). Yue Zhou 0006, Houjin Chen, Yanfeng Li 0001, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | Identify Representative Samples by Conditional Random Field of Cancer Histology ImagesabstractPathology analysis is crucial to precise cancer diagnoses and the succeeding treatment plan as well. To detect abnormality in histopathology images with prevailing patch-based convolutional neural networks (CNNs), contextual information often serves as a powerful cue. However, as whole-slide images (WSIs) are characterized by intense morphological heterogeneity and extensive tissue scale, a straightforward visual span to a larger context may not well capture the information closely associated with the focal patch. In this paper, we propose a novel pixel-offset based patch-location method to identify high-representative tissues, with a CNN backbone. Pathology Deformable Conditional Random Field (PDCRF) is proposed to learn the offsets and weights of neighboring contexts in a spatial-adaptive manner, to search for high-representative patches. A CNN structure with the localized patches as training input is then capable of consistently reaching superior classification outcomes for histology images. Overall, the proposed method has achieved state-of-the-art performance, in terms of the test classification accuracy improvement to the baseline by 1.15-2.60%, 0.78-1.78%, and 1.47-2.18% on TCGA public datasets of TCGA-STAD, TCGA-COAD, and TCGA-READ respectively. It also achieves 88.95% test accuracy and 0.920 test AUC on Camelyon 16. To show the effectiveness of the proposed framework on downstream tasks, we take a further step by incorporating an active learning model, which noticeably reduces the number of manual annotations by PDCRF to reach a parallel patch-based histology classifier. Yiqing Shen 0003, Dinggang Shen, Jing Ke |
IEEE Trans. Medical Imaging | 2 |
| 2022 | Doubly Supervised Transfer Classifier for Computer-Aided Diagnosis With Imbalanced ModalitiesabstractTransfer learning (TL) can effectively improve diagnosis accuracy of single-modal-imaging-based computer-aided diagnosis (CAD) by transferring knowledge from other related imaging modalities, which offers a way to alleviate the small-sample-size problem. However, medical imaging data generally have the following characteristics for the TL-based CAD: 1) The source domain generally has limited data, which increases the difficulty to explore transferable information for the target domain; 2) Samples in both domains often have been labeled for training the CAD model, but the existing TL methods cannot make full use of label information to improve knowledge transfer. In this work, we propose a novel doubly supervised transfer classifier (DSTC) algorithm. In particular, DSTC integrates the support vector machine plus (SVM+) classifier and the low-rank representation (LRR) into a unified framework. The former makes full use of the shared labels to guide the knowledge transfer between the paired data, while the latter adopts the block-diagonal low-rank (BLR) to perform supervised TL between the unpaired data. Furthermore, we introduce the Schatten-p norm for BLR to obtain a tighter approximation to the rank function. The proposed DSTC algorithm is evaluated on the Alzheimer's disease neuroimaging initiative (ADNI) dataset and the bimodal breast ultrasound image (BBUI) dataset. The experimental results verify the effectiveness of the proposed DSTC algorithm. Xiangmin Han, Xiaoyan Fei, Jun Wang 0024, Tao Zhou 0002, Shihui Ying, Jun Shi 0004, Dinggang Shen |
IEEE Trans. Medical Imaging | 7 |
| 2022 | Localization of Craniomaxillofacial Landmarks on CBCT Images Using 3D Mask R-CNN and Local Dependency LearningabstractCephalometric analysis relies on accurate detection of craniomaxillofacial (CMF) landmarks from cone-beam computed tomography (CBCT) images. However, due to the complexity of CMF bony structures, it is difficult to localize landmarks efficiently and accurately. In this paper, we propose a deep learning framework to tackle this challenge by jointly digitalizing 105 CMF landmarks on CBCT images. By explicitly learning the local geometrical relationships between the landmarks, our approach extends Mask R-CNN for end-to-end prediction of landmark locations. Specifically, we first apply a detection network on a down-sampled 3D image to leverage global contextual information to predict the approximate locations of the landmarks. We subsequently leverage local information provided by higher-resolution image patches to refine the landmark locations. On patients with varying non-syndromic jaw deformities, our method achieves an average detection accuracy of 1.38± 0.95mm, outperforming a related state-of-the-art method. Yankun Lang, Chunfeng Lian, Deqiang Xiao, Hannah H. Deng, Kim-Han Thung, Peng Yuan 0001, Jaime Gateno, Tianshu Kuang, David M. Alfi, Li Wang 0026, Dinggang Shen, James J. Xia, Pew-Thian Yap |
IEEE Trans. Medical Imaging | 11 |
| 2022 | Semi-Supervised Deep Transfer Learning for Benign-Malignant Diagnosis of Pulmonary Nodules in Chest CT ImagesabstractLung cancer is the leading cause of cancer deaths worldwide. Accurately diagnosing the malignancy of suspected lung nodules is of paramount clinical importance. However, to date, the pathologically-proven lung nodule dataset is largely limited and is highly imbalanced in benign and malignant distributions. In this study, we proposed a Semi-supervised Deep Transfer Learning (SDTL) framework for benign-malignant pulmonary nodule diagnosis. First, we utilize a transfer learning strategy by adopting a pre-trained classification network that is used to differentiate pulmonary nodules from nodule-like tissues. Second, since the size of samples with pathological-proven is small, an iterated feature-matching-based semi-supervised method is proposed to take advantage of a large available dataset with no pathological results. Specifically, a similarity metric function is adopted in the network semantic representation space for gradually including a small subset of samples with no pathological results to iteratively optimize the classification network. In this study, a total of 3,038 pulmonary nodules (from 2,853 subjects) with pathologically-proven benign or malignant labels and 14,735 unlabeled nodules (from 4,391 subjects) were retrospectively collected. Experimental results demonstrate that our proposed SDTL framework achieves superior diagnosis performance, with accuracy = 88.3%, AUC = 91.0% in the main dataset, and accuracy = 74.5%, AUC = 79.5% in the independent testing dataset. Furthermore, ablation study shows that the use of transfer learning provides 2% accuracy improvement, and the use of semi-supervised learning further contributes 2.9% accuracy improvement. Results implicate that our proposed classification network could provide an effective diagnostic tool for suspected lung nodules, and might have a promising application in clinical practice. Feng Shi 0001, Bojiang Chen, Qiqi Cao, Ying Wei 0009, Yaojie Zhou, Rongrong Fan, Fan Yang 0054, Yanbo Chen 0003, Weimin Li 0003, Yaozong Gao, Dinggang Shen |
IEEE Trans. Medical Imaging | 15 |
| 2022 | Follow My Eye: Using Gaze to Supervise Computer-Aided DiagnosisabstractWhen deep neural network (DNN) was first introduced to the medical image analysis community, researchers were impressed by its performance. However, it is evident now that a large number of manually labeled data is often a must to train a properly functioning DNN. This demand for supervision data and labels is a major bottleneck in current medical image analysis, since collecting a large number of annotations from experienced experts can be time-consuming and expensive. In this paper, we demonstrate that the eye movement of radiologists reading medical images can be a new form of supervision to train the DNN-based computer-aided diagnosis (CAD) system. Particularly, we record the tracks of the radiologists' gaze when they are reading images. The gaze information is processed and then used to supervise the DNN's attention via an Attention Consistency module. To the best of our knowledge, the above pipeline is among the earliest efforts to leverage expert eye movement for deep-learning-based CAD. We have conducted extensive experiments on knee X-ray images for osteoarthritis assessment. The results show that our method can achieve considerable improvement in diagnosis performance, with the help of gaze supervision. Sheng Wang 0014, Xi Ouyang, Tianming Liu 0001, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 5 |
| 2022 | Recurrent Tissue-Aware Network for Deformable Registration of Infant Brain MR ImagesabstractDeformable registration is fundamental to longitudinal and population-based image analyses. However, it is challenging to precisely align longitudinal infant brain MR images of the same subject, as well as cross-sectional infant brain MR images of different subjects, due to fast brain development during infancy. In this paper, we propose a recurrently usable deep neural network for the registration of infant brain MR images. There are three main highlights of our proposed method. (i) We use brain tissue segmentation maps for registration, instead of intensity images, to tackle the issue of rapid contrast changes of brain tissues during the first year of life. (ii) A single registration network is trained in a one-shot manner, and then recurrently applied in inference for multiple times, such that the complex deformation field can be recovered incrementally. (iii) We also propose both the adaptive smoothing layer and the tissue-aware anti-folding constraint into the registration network to ensure the physiological plausibility of estimated deformations without degrading the registration accuracy. Experimental results, in comparison to the state-of-the-art registration methods, indicate that our proposed method achieves the highest registration accuracy while still preserving the smoothness of the deformation field. The implementation of our proposed registration network is available onlinehttps://github.com/Barnonewdm/ACTA-Reg-Net. Dongming Wei, Sahar Ahmad, Yuyu Guo 0002, Liyun Chen, Yunzhi Huang, Lei Ma 0006, Zhengwang Wu, Gang Li 0001, Li Wang 0026, Weili Lin, Pew-Thian Yap, Dinggang Shen, Qian Wang 0001 |
IEEE Trans. Medical Imaging | 12 |
| 2022 | Two-Stage Mesh Deep Learning for Automated Tooth Segmentation and Landmark Localization on 3D Intraoral ScansabstractAccurately segmenting teeth and identifying the corresponding anatomical landmarks on dental mesh models are essential in computer-aided orthodontic treatment. Manually performing these two tasks is time-consuming, tedious, and, more importantly, highly dependent on orthodontists' experiences due to the abnormality and large-scale variance of patients' teeth. Some machine learning-based methods have been designed and applied in the orthodontic field to automatically segment dental meshes (e.g., intraoral scans). In contrast, the number of studies on tooth landmark localization is still limited. This paper proposes a two-stage framework based on mesh deep learning (called TS-MDL) for joint tooth labeling and landmark identification on raw intraoral scans. Our TS-MDL first adopts an end-to-end iMeshSegNet method (i.e., a variant of the existing MeshSegNet with both improved accuracy and efficiency) to label each tooth on the downsampled scan. Guided by the segmentation outputs, our TS-MDL further selects each tooth's region of interest (ROI) on the original mesh to construct a light-weight variant of the pioneering PointNet (i.e., PointNet-Reg) for regressing the corresponding landmark heatmaps. Our TS-MDL was evaluated on a real-clinical dataset, showing promising segmentation and localization performance. Specifically, iMeshSegNet in the first stage of TS-MDL reached an averaged Dice similarity coefficient (DSC) at 0.964±0.054 , significantly outperforming the original MeshSegNet. In the second stage, PointNet-Reg achieved a mean absolute error (MAE) of 0.597±0.761 mm in distances between the prediction and ground truth for 66 landmarks, which is superior compared with other networks for landmark detection. All these results suggest the potential usage of our TS-MDL in orthodontics. Tai-Hsien Wu, Chunfeng Lian, Matthew Pastewait, Christian Piers, Fan Wang 0023, Li Wang 0026, Chiung-Ying Chiu, Wenchi Wang, Christina Jackson, Wei-Lun Chao, Dinggang Shen, Ching-Chang Ko |
IEEE Trans. Medical Imaging | 13 |
| 2022 | Cross-Site Severity Assessment of COVID-19 From CT Images via Domain AdaptationabstractEarly and accurate severity assessment of Coronavirus disease 2019 (COVID-19) based on computed tomography (CT) images offers a great help to the estimation of intensive care unit event and the clinical decision of treatment planning. To augment the labeled data and improve the generalization ability of the classification model, it is necessary to aggregate data from multiple sites. This task faces several challenges including class imbalance between mild and severe infections, domain distribution discrepancy between sites, and presence of heterogeneous features. In this paper, we propose a novel domain adaptation (DA) method with two components to address these problems. The first component is a stochastic class-balanced boosting sampling strategy that overcomes the imbalanced learning problem and improves the classification performance on poorly-predicted classes. The second component is a representation learning that guarantees three properties: 1) domain-transferability by prototype triplet loss, 2) discriminant by conditional maximum mean discrepancy loss, and 3) completeness by multi-view reconstruction loss. Particularly, we propose a domain translator and align the heterogeneous data to the estimated class prototypes (i.e., class centers) in a hyper-sphere manifold. Experiments on cross-site severity assessment of COVID-19 from CT images show that the proposed method can effectively tackle the imbalanced learning problem and outperform recent DA approaches. Gengxin Xu, Chen Liu 0026, Jun Liu 0075, Zhongxiang Ding, Feng Shi 0001, Man Guo, Wei Zhao 0040, Ying Wei 0009, Yaozong Gao, Chuan-Xian Ren, Dinggang Shen |
IEEE Trans. Medical Imaging | 12 |
| 2022 | Multimodal MRI Reconstruction Assisted With Spatial Alignment NetworkabstractIn clinical practice, multi-modal magnetic resonance imaging (MRI) with different contrasts is usually acquired in a single study to assess different properties of the same region of interest in the human body. The whole acquisition process can be accelerated by having one or more modalities under-sampled in the k -space. Recent research has shown that, considering the redundancy between different modalities, a target MRI modality under-sampled in the k -space can be more efficiently reconstructed with a fully-sampled reference MRI modality. However, we find that the performance of the aforementioned multi-modal reconstruction can be negatively affected by subtle spatial misalignment between different modalities, which is actually common in clinical practice. In this paper, we improve the quality of multi-modal reconstruction by compensating for such spatial misalignment with a spatial alignment network. First, our spatial alignment network estimates the displacement between the fully-sampled reference and the under-sampled target images, and warps the reference image accordingly. Then, the aligned fully-sampled reference image joins the multi-modal reconstruction of the under-sampled target image. Also, considering the contrast difference between the target and reference images, we have designed a cross-modality-synthesis-based registration loss in combination with the reconstruction loss, to jointly train the spatial alignment network and the reconstruction network. The experiments on both clinical MRI and multi-coil k -space raw data demonstrate the superiority and robustness of the multi-modal MRI reconstruction empowered with our spatial alignment network. Our code is publicly available at https://github.com/woxuankai/SpatialAlignmentNetwork. Kai Xuan, Lei Xiang 0001, Xiaoqian Huang, Lichi Zhang, Shu Liao, Dinggang Shen, Qian Wang 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2022 | Diffusion Kernel Attention Network for Brain Disorder ClassificationabstractConstructing and analyzing functional brain networks (FBN) has become a promising approach to brain disorder classification. However, the conventional successive construct-and-analyze process would limit the performance due to the lack of interactions and adaptivity among the subtasks in the process. Recently, Transformer has demonstrated remarkable performance in various tasks, attributing to its effective attention mechanism in modeling complex feature relationships. In this paper, for the first time, we develop Transformer for integrated FBN modeling, analysis and brain disorder classification with rs-fMRI data by proposing a Diffusion Kernel Attention Network to address the specific challenges. Specifically, directly applying Transformer does not necessarily admit optimal performance in this task due to its extensive parameters in the attention module against the limited training samples usually available. Looking into this issue, we propose to use kernel attention to replace the original dot-product attention module in Transformer. This significantly reduces the number of parameters to train and thus alleviates the issue of small sample while introducing a non-linear attention mechanism to model complex functional connections. Another limit of Transformer for FBN applications is that it only considers pair-wise interactions between directly connected brain regions but ignores the important indirect connections. Therefore, we further explore diffusion process over the kernel attention to incorporate wider interactions among indirectly connected brain regions. Extensive experimental study is conducted on ADHD-200 data set for ADHD classification and on ADNI data set for Alzheimer's disease classification, and the results demonstrate the superior performance of the proposed method over the competing methods. Jianjia Zhang, Luping Zhou, Lei Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 5 |
| 2022 | Two-Stream Graph Convolutional Network for Intra-Oral Scanner Image SegmentationabstractPrecise segmentation of teeth from intra-oral scanner images is an essential task in computer-aided orthodontic surgical planning. The state-of-the-art deep learning-based methods often simply concatenate the raw geometric attributes (i.e., coordinates and normal vectors) of mesh cells to train a single-stream network for automatic intra-oral scanner image segmentation. However, since different raw attributes reveal completely different geometric information, the naive concatenation of different raw attributes at the (low-level) input stage may bring unnecessary confusion in describing and differentiating between mesh cells, thus hampering the learning of high-level geometric representations for the segmentation task. To address this issue, we design a two-stream graph convolutional network (i.e., TSGCN), which can effectively handle inter-view confusion between different raw attributes to more effectively fuse their complementary information and learn discriminative multi-view geometric representations. Specifically, our TSGCN adopts two input-specific graph-learning streams to extract complementary high-level geometric representations from coordinates and normal vectors, respectively. Then, these single-view representations are further fused by a self-attention module to adaptively balance the contributions of different views in learning more discriminative multi-view representations for accurate and fully automatic tooth segmentation. We have evaluated our TSGCN on a real-patient dataset of dental (mesh) models acquired by 3D intraoral scanners. Experimental results show that our TSGCN significantly outperforms state-of-the-art methods in 3D tooth (surface) segmentation. Yue Zhao 0012, Yang Liu 0157, Deyu Meng, Zhiming Cui 0001, Chenqiang Gao, Xinbo Gao 0001, Chunfeng Lian, Dinggang Shen |
IEEE Trans. Medical Imaging | 9 |
| 2022 | Multi-Task Weakly-Supervised Attention Network for Dementia Status Estimation With Structural MRIabstractAccurate prediction of clinical scores (of neuropsychological tests) based on noninvasive structural magnetic resonance imaging (MRI) helps understand the pathological stage of dementia (e.g., Alzheimer's disease (AD)) and forecast its progression. Existing machine/deep learning approaches typically preselect dementia-sensitive brain locations for MRI feature extraction and model construction, potentially leading to undesired heterogeneity between different stages and degraded prediction performance. Besides, these methods usually rely on prior anatomical knowledge (e.g., brain atlas) and time-consuming nonlinear registration for the preselection of brain locations, thereby ignoring individual-specific structural changes during dementia progression because all subjects share the same preselected brain regions. In this article, we propose a multi-task weakly-supervised attention network (MWAN) for the joint regression of multiple clinical scores from baseline MRI scans. Three sequential components are included in MWAN: 1) a backbone fully convolutional network for extracting MRI features; 2) a weakly supervised dementia attention block for automatically identifying subject-specific discriminative brain locations; and 3) an attention-aware multitask regression block for jointly predicting multiple clinical scores. The proposed MWAN is an end-to-end and fully trainable deep learning model in which dementia-aware holistic feature learning and multitask regression model construction are integrated into a unified framework. Our MWAN method was evaluated on two public AD data sets for estimating clinical scores of mini-mental state examination (MMSE), clinical dementia rating sum of boxes (CDRSB), and AD assessment scale cognitive subscale (ADAS-Cog). Quantitative experimental results demonstrate that our method produces superior regression performance compared with state-of-the-art methods. Importantly, qualitative results indicate that the dementia-sensitive brain locations automatically identified by our MWAN method well retain individual specificities and are biologically meaningful. Chunfeng Lian, Mingxia Liu 0001, Li Wang 0026, Dinggang Shen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | TSGCNet: Discriminative Geometric Feature Learning With Two-Stream Graph Convolutional Network for 3D Dental Model SegmentationabstractThe ability to segment teeth precisely from digitized 3D dental models is an essential task in computer-aided orthodontic surgical planning. To date, deep learning based methods have been popularly used to handle this task. State-of-the-art methods directly concatenate the raw attributes of 3D inputs, namely coordinates and normal vectors of mesh cells, to train a single-stream network for fully-automated tooth segmentation. This, however, has the drawback of ignoring the different geometric meanings provided by those raw attributes. This issue might possibly confuse the network in learning discriminative geometric features and result in many isolated false predictions on the dental model. Against this issue, we propose a two-stream graph convolutional network (TSGCNet) to learn multi-view geometric information from different geometric attributes. Our TSGCNet adopts two graph-learning streams, designed in an input-aware fashion, to extract more discriminative high-level geometric representations from coordinates and normal vectors, respectively. These feature representations learned from the designed two different streams are further fused to integrate the multi-view complementary information for the cell-wise dense prediction task. We evaluate our proposed TSGCNet on a real-patient dataset of dental models acquired by 3D intraoral scanners, and experimental results demonstrate that our method significantly outperforms state-of-the-art methods for 3D shape segmentation. Yue Zhao 0012, Deyu Meng, Zhiming Cui 0001, Chenqiang Gao, Xinbo Gao 0001, Chunfeng Lian, Dinggang Shen |
CVPR | 8 |
| 2021 | VertNet: Accurate Vertebra Localization and Identification Network from CT Images
Zhiming Cui 0001, Changjian Li 0001, Lei Yang 0048, Chunfeng Lian, Feng Shi 0001, Wenping Wang 0001, Dijia Wu, Dinggang Shen |
MICCAI (5) | 8 |
| 2021 | CA2.5-Net Nuclei Segmentation Framework with a Microscopy Cell Benchmark Collection
Jinghan Huang 0002, Yiqing Shen 0003, Dinggang Shen, Jing Ke |
MICCAI (8) | 3 |
| 2021 | Contrastive Learning Based Stain Normalization Across Multiple Tumor in Histopathology
Jing Ke, Yiqing Shen 0003, Xiaoyao Liang, Dinggang Shen |
MICCAI (8) | 4 |
| 2021 | Multimodal MRI Acceleration via Deep Cascading Networks with Peer-Layer-Wise Dense Connections
Xiaoxin Li 0001, Xin-Jie Lou, Yong Chen 0026, Dinggang Shen |
MICCAI (6) | 6 |
| 2021 | Domain Generalization for Mammography Detection via Multi-style and Multi-view Contrastive Learning
Zheren Li, Zhiming Cui 0001, Sheng Wang 0014, Yuji Qi, Xi Ouyang, Qitian Chen, Yuezhi Yang, Zhong Xue, Dinggang Shen, Jie-Zhi Cheng |
MICCAI (7) | 9 |
| 2021 | Building Dynamic Hierarchical Brain Networks and Capturing Transient Meta-states for Early Mild Cognitive Impairment Diagnosis
Mianxin Liu, Han Zhang 0002, Feng Shi 0001, Dinggang Shen |
MICCAI (7) | 4 |
| 2021 | Two-Stage Self-supervised Cycle-Consistency Network for Reconstruction of Thin-Slice MR Images
Zhiyang Lu, Jun Wang 0024, Jun Shi 0004, Dinggang Shen |
MICCAI (6) | 5 |
| 2021 | 3D Transformer-GAN for High-Quality PET Reconstruction
Yanmei Luo, Yan Wang 0015, Chen Zu, Bo Zhan, Xi Wu 0004, Jiliu Zhou, Dinggang Shen, Luping Zhou |
MICCAI (6) | 7 |
| 2021 | Self-adversarial Learning for Detection of Clustered Microcalcifications in Mammograms
Xi Ouyang, Jifei Che, Qitian Chen, Zheren Li, Yiqiang Zhan, Zhong Xue, Qian Wang 0001, Jie-Zhi Cheng, Dinggang Shen |
MICCAI (7) | 9 |
| 2021 | Collaborative Image Synthesis and Disease Diagnosis for Classification of Neurodegenerative Disorders with Incomplete Multi-modal Neuroimages
Yongsheng Pan, Yuanyuan Chen 0001, Dinggang Shen, Yong Xia 0001 |
MICCAI (5) | 3 |
| 2021 | Predicting Symptoms from Multiphasic MRI via Multi-instance Attention Learning for Hepatocellular Carcinoma Grading
Zelin Qiu, Yongsheng Pan, Dijia Wu, Yong Xia 0001, Dinggang Shen |
MICCAI (5) | 6 |
| 2021 | Motion Correction for Liver DCE-MRI with Time-Intensity Curve Constraint
Dongming Wei, Zhiming Cui 0001, Yujia Zhou 0001, Caiwen Jiang, Jiameng Liu, Qianjin Feng 0003, Dinggang Shen |
MICCAI (7) | 8 |
| 2021 | Consistent Segmentation of Longitudinal Brain MR Images with Spatio-Temporal Constrained Networks
Feng Shi 0001, Zhiming Cui 0001, Yongsheng Pan, Yong Xia 0001, Dinggang Shen |
MICCAI (1) | 6 |
| 2021 | A Self-supervised Deep Framework for Reference Bony Shape Estimation in Orthognathic Surgical Planning
Deqiang Xiao, Hannah H. Deng, Tianshu Kuang, Lei Ma 0006, Xu Chen 0020, Chunfeng Lian, Yankun Lang, Daeseung Kim, Jaime Gateno, Steve G. Shen, Dinggang Shen, Pew-Thian Yap, James J. Xia |
MICCAI (4) | 12 |
| 2021 | Confidence-Aware Cascaded Network for Fetal Brain Segmentation on MR Images
Xukun Zhang, Zhiming Cui 0001, Changan Chen, Jingjiao Lou, Wenxin Hu, He Zhang 0023, Tao Zhou 0002, Feng Shi 0001, Dinggang Shen |
MICCAI (3) | 10 |
| 2021 | Nodule Synthesis and Selection for Augmenting Chest X-ray Nodule Detection
Zhenrong Shen 0001, Xi Ouyang, Zhuochen Wang, Yiqiang Zhan, Zhong Xue, Qian Wang 0001, Jie-Zhi Cheng, Dinggang Shen |
PRCV (3) | 8 |
| 2021 | Interactive medical image segmentation via a point-based interaction
Jian Zhang 0090, Yinghuan Shi, Jinquan Sun, Lei Wang 0001, Luping Zhou, Yang Gao 0001, Dinggang Shen |
Artif. Intell. Medicine | 7 |
| 2021 | Edge-preserving MRI image synthesis via adversarial network with iterative multi-scale fusion
Yanmei Luo, Dong Nie, Bo Zhan, Xi Wu 0004, Jiliu Zhou, Yan Wang 0015, Dinggang Shen |
Neurocomputing | 8 |
| 2021 | Diverse data augmentation for learning image segmentation with cross-modality annotations
Xu Chen 0020, Chunfeng Lian, Li Wang 0026, Hannah H. Deng, Tianshu Kuang, Steve H. Fung, Jaime Gateno, Dinggang Shen, James J. Xia, Pew-Thian Yap |
Medical Image Anal. | 8 |
| 2021 | ABCnet: Adversarial bias correction network for infant brain MR images
Liangjun Chen, Zhengwang Wu, Dan Hu 0004, Fan Wang 0023, J. Keith Smith, Weili Lin, Li Wang 0026, Dinggang Shen, Gang Li 0001 |
Medical Image Anal. | 8 |
| 2021 | TSegNet: An efficient and accurate tooth segmentation network on 3D dental model
Zhiming Cui 0001, Changjian Li 0001, Nenglun Chen, Guodong Wei, Runnan Chen, Yuanfeng Zhou, Dinggang Shen, Wenping Wang 0001 |
Medical Image Anal. | 7 |
| 2021 | Hypergraph learning for identification of COVID-19 with CT imaging
Donglin Di, Feng Shi 0001, Fuhua Yan, Liming Xia, Zhanhao Mo, Zhongxiang Ding, Bin Song 0002, Shengrui Li, Ying Wei 0009, Ying Shao, Miaofei Han, Yaozong Gao, He Sui, Yue Gao 0002, Dinggang Shen |
Medical Image Anal. | 16 |
| 2021 | Multi-Regression based supervised sample selection for predicting baby connectome evolution trajectory from neonatal timepoint
Olfa Ghribi, Gang Li 0001, Weili Lin, Dinggang Shen, Islem Rekik |
Medical Image Anal. | 4 |
| 2021 | Multi-site MRI harmonization via attention-guided deep domain adaptation for brain disorder identification
Yunbi Liu, Erkun Yang, Pew-Thian Yap, Dinggang Shen, Mingxia Liu 0001 |
Medical Image Anal. | 5 |
| 2021 | MetricUNet: Synergistic image- and voxel-level learning for precise prostate segmentation via online sampling
Kelei He, Chunfeng Lian, Ehsan Adeli-Mosabbeb, Jing Huo, Yang Gao 0001, Bing Zhang 0012, Dinggang Shen |
Medical Image Anal. | 8 |
| 2021 | Difficulty-aware hierarchical convolutional neural networks for deformable registration of brain MR images
Yunzhi Huang, Sahar Ahmad, Jingfan Fan, Dinggang Shen, Pew-Thian Yap |
Medical Image Anal. | 4 |
| 2021 | A novel multiple instance learning framework for COVID-19 severity assessment via data augmentation and self-supervised learning
Zekun Li 0010, Wei Zhao 0040, Feng Shi 0001, Lei Qi 0001, Xingzhi Xie, Ying Wei 0009, Zhongxiang Ding, Yang Gao 0001, Shangjie Wu, Jun Liu 0075, Yinghuan Shi, Dinggang Shen |
Medical Image Anal. | 12 |
| 2021 | Gaussianization of Diffusion MRI Data Using Spatially Adaptive Filtering
Feihong Liu, Jun Feng 0003, Geng Chen 0001, Dinggang Shen, Pew-Thian Yap |
Medical Image Anal. | 4 |
| 2021 | Incomplete multi-modal representation learning for Alzheimer's disease diagnosis
Yanbei Liu, Lianxi Fan, Changqing Zhang 0002, Tao Zhou 0002, Zhitao Xiao, Lei Geng, Dinggang Shen |
Medical Image Anal. | 7 |
| 2021 | Asymmetric multi-task attention network for prostate bed segmentation in computed tomography images
Xuanang Xu, Chunfeng Lian, Shuai Wang 0003, Ronald C. Chen, Andrew Z. Wang, Trevor J. Royce, Pew-Thian Yap, Dinggang Shen, Jun Lian |
Medical Image Anal. | 9 |
| 2021 | Multi-task learning for segmentation and classification of tumors in 3D automated breast ultrasound images
Yue Zhou 0006, Houjin Chen, Yanfeng Li 0001, Xuanang Xu, Pew-Thian Yap, Dinggang Shen |
Medical Image Anal. | 8 |
| 2021 | Joint prediction and time estimation of COVID-19 developing severe symptoms using chest CT scan
Xiaofeng Zhu 0001, Bin Song 0002, Feng Shi 0001, Yanbo Chen 0003, Rongyao Hu, Jiangzhang Gan, Wenhai Zhang, Liye Wang, Yaozong Gao, Dinggang Shen |
Medical Image Anal. | 12 |
| 2021 | Advanced deep learning methods for biomedical information analysis: An editorial
Yudong Zhang 0001, Francesco Carlo Morabito, Dinggang Shen, Khan Muhammad 0001 |
Neural Networks | 3 |
| 2021 | Synergistic learning of lung lobe segmentation and hierarchical multi-instance classification for automated severity assessment of COVID-19 in CT images
Kelei He, Wei Zhao 0040, Xingzhi Xie, Mingxia Liu 0001, Zhenyu Tang 0002, Yinghuan Shi, Feng Shi 0001, Yang Gao 0001, Jun Liu 0075, Dinggang Shen |
Pattern Recognit. | 12 |
| 2021 | Reducing magnetic resonance image spacing by learning without ground-truth
Kai Xuan, Liping Si, Lichi Zhang, Zhong Xue, Yining Jiao, Weiwu Yao, Dinggang Shen, Dijia Wu, Qian Wang 0001 |
Pattern Recognit. | 7 |
| 2021 | Cascaded MultiTask 3-D Fully Convolutional Networks for Pancreas SegmentationabstractAutomatic pancreas segmentation is crucial to the diagnostic assessment of diabetes or pancreatic cancer. However, the relatively small size of the pancreas in the upper body, as well as large variations of its location and shape in retroperitoneum, make the segmentation task challenging. To alleviate these challenges, in this article, we propose a cascaded multitask 3-D fully convolution network (FCN) to automatically segment the pancreas. Our cascaded network is composed of two parts. The first part focuses on fast locating the region of the pancreas, and the second part uses a multitask FCN with dense connections to refine the segmentation map for fine voxel-wise segmentation. In particular, our multitask FCN with dense connections is implemented to simultaneously complete tasks of the voxel-wise segmentation and skeleton extraction from the pancreas. These two tasks are complementary, that is, the extracted skeleton provides rich information about the shape and size of the pancreas in retroperitoneum, which can boost the segmentation of pancreas. The multitask FCN is also designed to share the low- and mid-level features across the tasks. A feature consistency module is further introduced to enhance the connection and fusion of different levels of feature maps. Evaluations on two pancreas datasets demonstrate the robustness of our proposed method in correctly segmenting the pancreas in various settings. Our experimental results outperform both baseline and state-of-the-art methods. Moreover, the ablation study shows that our proposed parts/modules are critical for effective multitask learning. Jie Xue 0001, Kelei He, Dong Nie, Ehsan Adeli-Mosabbeb, Zhenshan Shi, Seong-Whan Lee, Yuanjie Zheng, Xiyu Liu 0001, Dengwang Li, Dinggang Shen |
IEEE Trans. Cybern. | 10 |
| 2021 | Learning-Based Computer-Aided Prescription Model for Parkinson's Disease: A Data-Driven PerspectiveabstractIn this article, we study a novel problem: "automatic prescription recommendation for PD patients." To realize this goal, we first build a dataset by collecting 1) symptoms of PD patients, and 2) their prescription drug provided by neurologists. Then, we build a novel computer-aided prescription model by learning the relation between observed symptoms and prescription drug. Finally, for the new coming patients, we could recommend (predict) suitable prescription drug on their observed symptoms by our prescription model. From the methodology part, our proposed model, namely Prescription viA Learning lAtent Symptoms (PALAS), could recommend prescription using the multi-modality representation of the data. In PALAS, a latent symptom space is learned to better model the relationship between symptoms and prescription drug, as there is a large semantic gap between them. Moreover, we present an efficient alternating optimization method for PALAS. We evaluated our method using the data collected from 136 PD patients at Nanjing Brain Hospital, which can be regarded as a large dataset in PD research community. The experimental results demonstrate the effectiveness and clinical potential of our method in this recommendation task, if compared with other competing methods. Yinghuan Shi, Wanqi Yang, Kim-Han Thung, Hao Wang 0013, Yang Gao 0001, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 8 |
| 2021 | Fast and Accurate Craniomaxillofacial Landmark Detection via 3D Faster R-CNNabstractAutomatic craniomaxillofacial (CMF) landmark localization from cone-beam computed tomography (CBCT) images is challenging, considering that 1) the number of landmarks in the images may change due to varying deformities and traumatic defects, and 2) the CBCT images used in clinical practice are typically large. In this paper, we propose a two-stage, coarse-to-fine deep learning method to tackle these challenges with both speed and accuracy in mind. Specifically, we first use a 3D faster R-CNN to roughly locate landmarks in down-sampled CBCT images that have varying numbers of landmarks. By converting the landmark point detection problem to a generic object detection problem, our 3D faster R-CNN is formulated to detect virtual, fixed-size objects in small boxes with centers indicating the approximate locations of the landmarks. Based on the rough landmark locations, we then crop 3D patches from the high-resolution images and send them to a multi-scale UNet for the regression of heatmaps, from which the refined landmark locations are finally derived. We evaluated the proposed approach by detecting up to 18 landmarks on a real clinical dataset of CMF CBCT images with various conditions. Experiments show that our approach achieves state-of-the-art accuracy of 0.89 ± 0.64mm in an average time of 26.2 seconds per volume. Xiaoyang Chen 0002, Chunfeng Lian, Hannah H. Deng, Tianshu Kuang, Hung-Ying Lin, Deqiang Xiao, Jaime Gateno, Dinggang Shen, James J. Xia, Pew-Thian Yap |
IEEE Trans. Medical Imaging | 8 |
| 2021 | Anatomy-Regularized Representation Learning for Cross-Modality Medical Image SegmentationabstractAn increasing number of studies are leveraging unsupervised cross-modality synthesis to mitigate the limited label problem in training medical image segmentation models. They typically transfer ground truth annotations from a label-rich imaging modality to a label-lacking imaging modality, under an assumption that different modalities share the same anatomical structure information. However, since these methods commonly use voxel/pixel-wise cycle-consistency to regularize the mappings between modalities, high-level semantic information is not necessarily preserved. In this paper, we propose a novel anatomy-regularized representation learning approach for segmentation-oriented cross-modality image synthesis. It learns a common feature encoding across different modalities to form a shared latent space, where 1) the input and its synthesis present consistent anatomical structure information, and 2) the transformation between two images in one domain is preserved by their syntheses in another domain. We applied our method to the tasks of cross-modality skull segmentation and cardiac substructure segmentation. Experimental results demonstrate the superiority of our method in comparison with state-of-the-art cross-modality medical image segmentation methods. Xu Chen 0020, Chunfeng Lian, Li Wang 0026, Hannah H. Deng, Tianshu Kuang, Steve H. Fung, Jaime Gateno, Pew-Thian Yap, James J. Xia, Dinggang Shen |
IEEE Trans. Medical Imaging | 10 |
| 2021 | Structure-Driven Unsupervised Domain Adaptation for Cross-Modality Cardiac SegmentationabstractPerformance degradation due to domain shift remains a major challenge in medical image analysis. Unsupervised domain adaptation that transfers knowledge learned from the source domain with ground truth labels to the target domain without any annotation is the mainstream solution to resolve this issue. In this paper, we present a novel unsupervised domain adaptation framework for cross-modality cardiac segmentation, by explicitly capturing a common cardiac structure embedded across different modalities to guide cardiac segmentation. In particular, we first extract a set of 3D landmarks, in a self-supervised manner, to represent the cardiac structure of different modalities. The high-level structure information is then combined with another complementary feature, the Canny edges, to produce accurate cardiac segmentation results both in the source and target domains. We extensively evaluate our method on the MICCAI 2017 MM-WHS dataset for cardiac segmentation. The evaluation, comparison and comprehensive ablation studies demonstrate that our approach achieves satisfactory segmentation results and outperforms state-of-the-art unsupervised domain adaptation methods by a significant margin. Zhiming Cui 0001, Changjian Li 0001, Zhixu Du, Nenglun Chen, Guodong Wei, Runnan Chen, Lei Yang 0048, Dinggang Shen, Wenping Wang 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2021 | HF-UNet: Learning Hierarchically Inter-Task Relevance in Multi-Task U-Net for Accurate Prostate Segmentation in CT ImagesabstractAccurate segmentation of the prostate is a key step in external beam radiation therapy treatments. In this paper, we tackle the challenging task of prostate segmentation in CT images by a two-stage network with 1) the first stage to fast localize, and 2) the second stage to accurately segment the prostate. To precisely segment the prostate in the second stage, we formulate prostate segmentation into a multi-task learning framework, which includes a main task to segment the prostate, and an auxiliary task to delineate the prostate boundary. Here, the second task is applied to provide additional guidance of unclear prostate boundary in CT images. Besides, the conventional multi-task deep networks typically share most of the parameters (i.e., feature representations) across all tasks, which may limit their data fitting ability, as the specificity of different tasks are inevitably ignored. By contrast, we solve them by a hierarchically-fused U-Net structure, namely HF-UNet. The HF-UNet has two complementary branches for two tasks, with the novel proposed attention-based task consistency learning block to communicate at each level between the two decoding branches. Therefore, HF-UNet endows the ability to learn hierarchically the shared representations for different tasks, and preserve the specificity of learned representations for different tasks simultaneously. We did extensive evaluations of the proposed method on a large planning CT image dataset and a benchmark prostate zonal dataset. The experimental results show HF-UNet outperforms the conventional multi-task network architectures and the state-of-the-art methods. Kelei He, Chunfeng Lian, Bing Zhang 0012, Xin Zhang 0013, Xiaohuan Cao, Dong Nie, Yang Gao 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 9 |
| 2021 | NHBS-Net: A Feature Fusion Attention Network for Ultrasound Neonatal Hip Bone SegmentationabstractUltrasound is a widely used technology for diagnosing developmental dysplasia of the hip (DDH) because it does not use radiation. Due to its low cost and convenience, 2-D ultrasound is still the most common examination in DDH diagnosis. In clinical usage, the complexity of both ultrasound image standardization and measurement leads to a high error rate for sonographers. The automatic segmentation results of key structures in the hip joint can be used to develop a standard plane detection method that helps sonographers decrease the error rate. However, current automatic segmentation methods still face challenges in robustness and accuracy. Thus, we propose a neonatal hip bone segmentation network (NHBS-Net) for the first time for the segmentation of seven key structures. We design three improvements, an enhanced dual attention module, a two-class feature fusion module, and a coordinate convolution output head, to help segment different structures. Compared with current state-of-the-art networks, NHBS-Net gains outstanding performance accuracy and generalizability, as shown in the experiments. Additionally, image standardization is a common need in ultrasonography. The ability of segmentation-based standard plane detection is tested on a 50-image standard dataset. The experiments show that our method can help healthcare workers decrease their error rate from 6%-10% to 2%. In addition, the segmentation performance in another ultrasound dataset (fetal heart) demonstrates the ability of our network. Ruhan Liu, Mengyao Liu 0004, Bin Sheng 0001, Huating Li, Ping Li 0016, Haitao Song 0001, Ping Zhang 0016, Lixin Jiang, Dinggang Shen |
IEEE Trans. Medical Imaging | 9 |
| 2021 | Multi-Site Infant Brain Segmentation Algorithms: The iSeg-2019 ChallengeabstractTo better understand early brain development in health and disorder, it is critical to accurately segment infant brain magnetic resonance (MR) images into white matter (WM), gray matter (GM), and cerebrospinal fluid (CSF). Deep learning-based methods have achieved state-of-the-art performance; h owever, one of the major limitations is that the learning-based methods may suffer from the multi-site issue, that is, the models trained on a dataset from one site may not be applicable to the datasets acquired from other sites with different imaging protocols/scanners. To promote methodological development in the community, the iSeg-2019 challenge (http://iseg2019.web.unc.edu) provides a set of 6-month infant subjects from multiple sites with different protocols/scanners for the participating methods. T raining/validation subjects are from UNC (MAP) and testing subjects are from UNC/UMN (BCP), Stanford University, and Emory University. By the time of writing, there are 30 automatic segmentation methods participated in the iSeg-2019. In this article, 8 top-ranked methods were reviewed by detailing their pipelines/implementations, presenting experimental results, and evaluating performance across different sites in terms of whole brain, regions of interest, and gyral landmark curves. We further pointed out their limitations and possible directions for addressing the multi-site issue. We find that multi-site consistency is still an open issue. We hope that the multi-site dataset in the iSeg-2019 and this review article will attract more researchers to address the challenging and critical multi-site issue in practice. Yue Sun 0001, Kun Gao 0002, Zhengwang Wu, Xiaopeng Zong, Zhihao Lei, Ying Wei 0007, Jun Ma 0016, Xiaoping Yang 0001, Xue Feng 0001, Li Zhao 0001, Trung Le Phan, Jitae Shin, Tao Zhong 0002, Yu Zhang 0064, Lequan Yu, Caizi Li, Ramesh Basnet, M. Omair Ahmad, M. N. S. Swamy 0001, Wenao Ma, Qi Dou 0001, Toan Duc Bui, Camilo Bermudez, Bennett A. Landman, Ian H. Gotlib, Kathryn L. Humphreys, Sarah Shultz, Longchuan Li, Sijie Niu, Weili Lin, Valerie Jewells, Dinggang Shen, Gang Li 0001, Li Wang 0026 |
IEEE Trans. Medical Imaging | 33 |
| 2021 | Boundary Coding Representation for Organ Segmentation in Prostate Cancer RadiotherapyabstractAccurate segmentation of the prostate and organs at risk (OARs, e.g., bladder and rectum) in male pelvic CT images is a critical step for prostate cancer radiotherapy. Unfortunately, the unclear organ boundary and large shape variation make the segmentation task very challenging. Previous studies usually used representations defined directly on unclear boundaries as context information to guide segmentation. Those boundary representations may not be so discriminative, resulting in limited performance improvement. To this end, we propose a novel boundary coding network (BCnet) to learn a discriminative representation for organ boundary and use it as the context information to guide the segmentation. Specifically, we design a two-stage learning strategy in the proposed BCnet: 1) Boundary coding representation learning. Two sub-networks under the supervision of the dilation and erosion masks transformed from the manually delineated organ mask are first separately trained to learn the spatial-semantic context near the organ boundary. Then we encode the organ boundary based on the predictions of these two sub-networks and design a multi-atlas based refinement strategy by transferring the knowledge from training data to inference. 2) Organ segmentation. The boundary coding representation as context information, in addition to the image patches, are used to train the final segmentation network. Experimental results on a large and diverse male pelvic CT dataset show that our method achieves superior performance compared with several state-of-the-art methods. Shuai Wang 0003, Mingxia Liu 0001, Jun Lian, Dinggang Shen |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Deep Bayesian Hashing With Center Prior for Multi-Modal Neuroimage RetrievalabstractMulti-modal neuroimage retrieval has greatly facilitated the efficiency and accuracy of decision making in clinical practice by providing physicians with previous cases (with visually similar neuroimages) and corresponding treatment records. However, existing methods for image retrieval usually fail when applied directly to multi-modal neuroimage databases, since neuroimages generally have smaller inter-class variation and larger inter-modal discrepancy compared to natural images. To this end, we propose a deep Bayesian hash learning framework, called CenterHash, which can map multi-modal data into a shared Hamming space and learn discriminative hash codes from imbalanced multi-modal neuroimages. The key idea to tackle the small inter-class variation and large inter-modal discrepancy is to learn a common center representation for similar neuroimages from different modalities and encourage hash codes to be explicitly close to their corresponding center representations. Specifically, we measure the similarity between hash codes and their corresponding center representations and treat it as a center prior in the proposed Bayesian learning framework. A weighted contrastive likelihood loss function is also developed to facilitate hash learning from imbalanced neuroimage pairs. Comprehensive empirical evidence shows that our method can generate effective hash codes and yield state-of-the-art performance in cross-modal retrieval on three multi-modal neuroimage datasets. Erkun Yang, Mingxia Liu 0001, Dongren Yao, Bing Cao 0002, Chunfeng Lian, Pew-Thian Yap, Dinggang Shen |
IEEE Trans. Medical Imaging | 7 |
| 2021 | A Mutual Multi-Scale Triplet Graph Convolutional Network for Classification of Brain Disorders Using Functional or Structural ConnectivityabstractBrain connectivity alterations associated with mental disorders have been widely reported in both functional MRI (fMRI) and diffusion MRI (dMRI). However, extracting useful information from the vast amount of information afforded by brain networks remains a great challenge. Capturing network topology, graph convolutional networks (GCNs) have demonstrated to be superior in learning network representations tailored for identifying specific brain disorders. Existing graph construction techniques generally rely on a specific brain parcellation to define regions-of-interest (ROIs) to construct networks, often limiting the analysis into a single spatial scale. In addition, most methods focus on the pairwise relationships between the ROIs and ignore high-order associations between subjects. In this letter, we propose a mutual multi-scale triplet graph convolutional network (MMTGCN) to analyze functional and structural connectivity for brain disorder diagnosis. We first employ several templates with different scales of ROI parcellation to construct coarse-to-fine brain connectivity networks for each subject. Then, a triplet GCN (TGCN) module is developed to learn functional/structural representations of brain connectivity networks at each scale, with the triplet relationship among subjects explicitly incorporated into the learning process. Finally, we propose a template mutual learning strategy to train different scale TGCNs collaboratively for disease classification. Experimental results on 1,160 subjects from three datasets with fMRI or dMRI data demonstrate that our MMTGCN outperforms several state-of-the-art methods in identifying three types of brain disorders. Dongren Yao, Jing Sui, Erkun Yang, Yeerfan Jiaerken, Pew-Thian Yap, Mingxia Liu 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 9 |
| 2021 | Spherical Deformable U-Net: Application to Cortical Surface Parcellation and Development PredictionabstractConvolutional Neural Networks (CNNs) have achieved overwhelming success in learning-related problems for 2D/3D images in the Euclidean space. However, unlike in the Euclidean space, the shapes of many structures in medical imaging have an inherent spherical topology in a manifold space, e.g., the convoluted brain cortical surfaces represented by triangular meshes. There is no consistent neighborhood definition and thus no straightforward convolution/pooling operations for such cortical surface data. In this paper, leveraging the regular and hierarchical geometric structure of the resampled spherical cortical surfaces, we create the 1-ring filter on spherical cortical triangular meshes and accordingly develop convolution/pooling operations for constructing Spherical U-Net for cortical surface data. However, the regular nature of the 1-ring filter makes it inherently limited to model fixed geometric transformations. To further enhance the transformation modeling capability of Spherical U-Net, we introduce the deformable convolution and deformable pooling to cortical surface data and accordingly propose the Spherical Deformable U-Net (SDU-Net). Specifically, spherical offsets are learned to freely deform the 1-ring filter on the sphere to adaptively localize cortical structures with different sizes and shapes. We then apply the SDU-Net to two challenging and scientifically important tasks in neuroimaging: cortical surface parcellation and cortical attribute map prediction. Both applications validate the competitive performance of our approach in accuracy and computational efficiency in comparison with state-of-the-art methods. Fenqiang Zhao, Zhengwang Wu, Li Wang 0026, Weili Lin, John H. Gilmore, Shunren Xia, Dinggang Shen, Gang Li 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2021 | S3Reg: Superfast Spherical Surface Registration Based on Deep LearningabstractCortical surface registration is an essential step and prerequisite for surface-based neuroimaging analysis. It aligns cortical surfaces across individuals and time points to establish cross-sectional and longitudinal cortical correspondences to facilitate neuroimaging studies. Though achieving good performance, available methods are either time consuming or not flexible to extend to multiple or high dimensional features. Considering the explosive availability of large-scale and multimodal brain MRI data, fast surface registration methods that can flexibly handle multimodal features are desired. In this study, we develop a Superfast Spherical Surface Registration (S3Reg) framework for the cerebral cortex. Leveraging an end-to-end unsupervised learning strategy, S3Reg offers great flexibility in the choice of input feature sets and output similarity measures for registration, and meanwhile reduces the registration time significantly. Specifically, we exploit the powerful learning capability of spherical Convolutional Neural Network (CNN) to directly learn the deformation fields in spherical space and implement diffeomorphic design with "scaling and squaring" layers to guarantee topology-preserving deformations. To handle the polar-distortion issue, we construct a novel spherical CNN model using three orthogonal Spherical U-Nets. Experiments are performed on two different datasets to align both adult and infant multimodal cortical features. Results demonstrate that our S3Reg shows superior or comparable performance with state-of-the-art methods, while improving the registration time from 1 min to 10 sec. Fenqiang Zhao, Zhengwang Wu, Fan Wang 0023, Weili Lin, Shunren Xia, Dinggang Shen, Li Wang 0026, Gang Li 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2020 | Auto-GAN: Self-Supervised Collaborative Learning for Medical Image SynthesisabstractIn various clinical scenarios, medical image is crucial in disease diagnosis and treatment. Different modalities of medical images provide complementary information and jointly helps doctors to make accurate clinical decision. However, due to clinical and practical restrictions, certain imaging modalities may be unavailable nor complete. To impute missing data with adequate clinical accuracy, here we propose a framework called self-supervised collaborative learning to synthesize missing modality for medical images. The proposed method comprehensively utilize all available information correlated to the target modality from multi-source-modality images to generate any missing modality in a single model. Different from the existing methods, we introduce an auto-encoder network as a novel, self-supervised constraint, which provides target-modality-specific information to guide generator training. In addition, we design a modality mask vector as the target modality label. With experiments on multiple medical image databases, we demonstrate a great generalization ability as well as specialty of our method compared with other state-of-the-arts. Bing Cao 0002, Han Zhang 0002, Nannan Wang 0001, Xinbo Gao 0001, Dinggang Shen |
AAAI | 5 |
| 2020 | Non-Local U-Nets for Biomedical Image SegmentationabstractDeep learning has shown its great promise in various biomedical image segmentation tasks. Existing models are typically based on U-Net and rely on an encoder-decoder architecture with stacked local operators to aggregate long-range information gradually. However, only using the local operators limits the efficiency and effectiveness. In this work, we propose the non-local U-Nets, which are equipped with flexible global aggregation blocks, for biomedical image segmentation. These blocks can be inserted into U-Net as size-preserving processes, as well as down-sampling and up-sampling layers. We perform thorough experiments on the 3D multimodality isointense infant brain MR image segmentation task to evaluate the non-local U-Nets. Results show that our proposed models achieve top performances with fewer parameters and faster computation. Na Zou 0001, Dinggang Shen, Shuiwang Ji |
AAAI | 3 |
| 2020 | Identifying patch-level MSI from histological images of Colorectal Cancer by a Knowledge Distillation ModelabstractMicrosatellite instability (MSI) is the result of a defective DNA mismatch repair (MMR) system, and its presence occurs in a variety of cancers. The determination of MSI in colorectal cancer (CRC) will have a better prognosis and management of cancer patients. As the routine MSI identification via molecular testing is expensive, time-consuming, and region-restricted, novel methods to detect MSI are of great interest. In this work, we propose a multi-stage convolutional neural network (CNN) based framework to identify MSI status in colorectal cancer patients from histopathological images. A mislabel-aware module is designed to deal with the uncertainty problem in global-local labelling. An auto-grading model is proposed to discriminate patches by the degree of their histopathological correlation with recognizable MSI status, and subsequently aggregate the weights to make slide-level predictions. Our proposed methodology outperforms the existing models in the classification accuracy, and explicitly sorts out patches with representative features. The research outcome has the potential to assist in the interpretation of histopathology as a surrogate for MSI testing and also in the study of recognizable morphology of MSI-H/MSS tumors. Furthermore, this approach can be extended and applied to other cancer types. Jing Ke, Yiqing Shen 0003, Jason D. Wright, Naifeng Jing, Xiaoyao Liang, Dinggang Shen |
BIBM | 6 |
| 2020 | Automatic Data Augmentation Via Deep Reinforcement Learning for Effective Kidney Tumor SegmentationabstractConventional data augmentation realized by performing simple pre-processing operations (e.g., rotation, crop, etc.) has been validated for its advantage in enhancing the performance for medical image segmentation. However, the data generated by these conventional augmentation methods are random and sometimes harmful to the subsequent segmentation. In this paper, we developed a novel automatic learning-based data augmentation method for medical image segmentation which models the augmentation task as a trial-and-error procedure using deep reinforcement learning (DRL). In our method, we innovatively combine the data augmentation module and the subsequent segmentation module in an end-to-end training manner with a consistent loss. Specifically, the best sequential combination of different basic operations is automatically learned by directly maximizing the performance improvement (i.e., Dice ratio) on the available validation set. We extensively evaluated our method on CT kidney tumor segmentation which validated the promising results of our method. Tiexin Qin, Kelei He, Yinghuan Shi, Yang Gao 0001, Dinggang Shen |
ICASSP | 6 |
| 2020 | Fast Correction of Eddy-Current and Susceptibility-Induced Distortions Using Rotation-Invariant Contrasts
Sahar Ahmad, Ye Wu 0001, Khoi Minh Huynh, Kim-Han Thung, Weili Lin, Dinggang Shen, Pew-Thian Yap |
MICCAI (2) | 6 |
| 2020 | Semantic Hierarchy Guided Registration Networks for Intra-subject Pulmonary CT Image Alignment
Liyun Chen, Xiaohuan Cao, Lei Chen 0012, Yaozong Gao, Dinggang Shen, Qian Wang 0001, Zhong Xue |
MICCAI (3) | 5 |
| 2020 | Estimating Tissue Microstructure with Undersampled Diffusion Data via Graph Convolutional Neural Networks
Geng Chen 0001, Yoonmi Hong, Yongqin Zhang, Jaeil Kim, Khoi Minh Huynh, Jiquan Ma, Weili Lin, Dinggang Shen, Pew-Thian Yap |
MICCAI (7) | 8 |
| 2020 | A Deep Spatial Context Guided Framework for Infant Brain Subcortical Segmentation
Liangjun Chen, Zhengwang Wu, Dan Hu 0004, Zhanhao Mo, Li Wang 0026, Weili Lin, Dinggang Shen, Gang Li 0001 |
MICCAI (7) | 8 |
| 2020 | Acceleration of High-Resolution 3D MR Fingerprinting via a Graph Convolutional Network
Yong Chen 0026, Xiaopeng Zong, Weili Lin, Dinggang Shen, Pew-Thian Yap |
MICCAI (2) | 5 |
| 2020 | A New Metric for Characterizing Dynamic Redundancy of Dense Brain Chronnectome and Its Application to Early Detection of Alzheimer's Disease
Maryam Ghanbari, Li-Ming Hsu, Zhen Zhou 0004, Amir Ghanbari, Zhanhao Mo, Pew-Thian Yap, Han Zhang 0002, Dinggang Shen |
MICCAI (7) | 8 |
| 2020 | Pair-Wise and Group-Wise Deformation Consistency in Deep Registration Network
Dongdong Gu, Xiaohuan Cao, Shanshan Ma, Lei Chen 0012, Guocai Liu, Dinggang Shen, Zhong Xue |
MICCAI (3) | 6 |
| 2020 | Disentangled Intensive Triplet Autoencoder for Infant Functional Connectome Fingerprinting
Dan Hu 0004, Fan Wang 0023, Han Zhang 0002, Zhengwang Wu, Li Wang 0026, Weili Lin, Gang Li 0001, Dinggang Shen |
MICCAI (7) | 8 |
| 2020 | Construction of Spatiotemporal Infant Cortical Surface Functional Templates
Ying Huang 0007, Fan Wang 0023, Zhengwang Wu, Zengsi Chen, Han Zhang 0002, Li Wang 0026, Weili Lin, Dinggang Shen, Gang Li 0001 |
MICCAI (7) | 8 |
| 2020 | Characterizing Intra-soma Diffusion with Spherical Mean Spectrum Imaging
Khoi Minh Huynh, Ye Wu 0001, Kim-Han Thung, Sahar Ahmad, Hoyt Patrick Taylor IV, Dinggang Shen, Pew-Thian Yap |
MICCAI (7) | 6 |
| 2020 | Automatic Localization of Landmarks in Craniomaxillofacial CBCT Images Using a Local Attention-Based Graph Convolution Network
Yankun Lang, Chunfeng Lian, Deqiang Xiao, Hannah H. Deng, Peng Yuan 0001, Jaime Gateno, Steve G. Shen, David M. Alfi, Pew-Thian Yap, James J. Xia, Dinggang Shen |
MICCAI (4) | 11 |
| 2020 | Multi-task Dynamic Transformer Network for Concurrent Bone Segmentation and Large-Scale Landmark Localization with Dental CBCT
Chunfeng Lian, Fan Wang 0023, Hannah H. Deng, Li Wang 0026, Deqiang Xiao, Tianshu Kuang, Hung-Ying Lin, Jaime Gateno, Steve G. Shen, Pew-Thian Yap, James J. Xia, Dinggang Shen |
MICCAI (4) | 12 |
| 2020 | Joint Image Quality Assessment and Brain Extraction of Fetal MRI Using Deep Learning
Lufan Liao, Xin Zhang 0013, Fenqiang Zhao, Tao Zhong 0002, Yuchen Pei, Xiangmin Xu 0001, Li Wang 0026, He Zhang 0023, Dinggang Shen, Gang Li 0001 |
MICCAI (6) | 9 |
| 2020 | Generating Dual-Energy Subtraction Soft-Tissue Images from Chest Radiographs via Bone Edge-Guided GAN
Yunbi Liu, Mingxia Liu 0001, Yuhua Xi, Genggeng Qin, Dinggang Shen, Wei Yang 0006 |
MICCAI (2) | 5 |
| 2020 | Joint Neuroimage Synthesis and Representation Learning for Conversion Prediction of Subjective Cognitive Decline
Yunbi Liu, Yongsheng Pan, Wei Yang 0006, Zhenyuan Ning, Ling Yue, Mingxia Liu 0001, Dinggang Shen |
MICCAI (7) | 7 |
| 2020 | A Computational Framework for Dissociating Development-Related from Individually Variable Flexibility in Regional Modularity Assignment in Early Infancy
Mayssa Soussia, Xuyun Wen, Zhen Zhou 0004, Bing Jin, Tae-Eui Kam, Li-Ming Hsu, Zhengwang Wu, Gang Li 0001, Li Wang 0026, Islem Rekik, Weili Lin, Dinggang Shen, Han Zhang 0002 |
MICCAI (7) | 12 |
| 2020 | Globally Optimized Super-Resolution of Diffusion MRI Data via Fiber Continuity
Ye Wu 0001, Yoonmi Hong, Sahar Ahmad, Wei-Tang Chang, Weili Lin, Dinggang Shen, Pew-Thian Yap |
MICCAI (7) | 6 |
| 2020 | Tract Dictionary Learning for Fast and Robust Recognition of Fiber Bundles
Ye Wu 0001, Yoonmi Hong, Sahar Ahmad, Weili Lin, Dinggang Shen, Pew-Thian Yap |
MICCAI (7) | 5 |
| 2020 | Asymmetrical Multi-task Attention U-Net for the Segmentation of Prostate Bed in CT Image
Xuanang Xu, Chunfeng Lian, Shuai Wang 0003, Andrew Z. Wang, Trevor J. Royce, Ronald C. Chen, Jun Lian, Dinggang Shen |
MICCAI (4) | 8 |
| 2020 | Deep Disentangled Hashing with Momentum Triplets for Neuroimage Search
Erkun Yang, Dongren Yao, Bing Cao 0002, Pew-Thian Yap, Dinggang Shen, Mingxia Liu 0001 |
MICCAI (1) | 6 |
| 2020 | Infant Cognitive Scores Prediction with Multi-stream Attention-Based Temporal Path Signature Features
Xin Zhang 0013, Hao Ni 0001, Chenyang Li 0007, Xiangmin Xu 0001, Zhengwang Wu, Li Wang 0026, Weili Lin, Dinggang Shen, Gang Li 0001 |
MICCAI (7) | 9 |
| 2020 | Domain-Invariant Prior Knowledge Guided Attention Networks for Robust Skull Stripping of Developing Macaque Brains
Tao Zhong 0002, Yu Zhang 0064, Fenqiang Zhao, Yuchen Pei, Lufan Liao, Zhenyuan Ning, Li Wang 0026, Dinggang Shen, Gang Li 0001 |
MICCAI (7) | 8 |
| 2020 | Adversarial Confidence Learning for Medical Image Segmentation and Synthesis
Dong Nie, Dinggang Shen |
Int. J. Comput. Vis. | 2 |
| 2020 | Learning longitudinal classification-regression model for infant hippocampus segmentation
Yanrong Guo, Zhengwang Wu, Dinggang Shen |
Neurocomputing | 3 |
| 2020 | A novel approach to multiple anatomical shape analysis: Application to fetal ventriculomegaly
Oualid M. Benkarim, Gemma Piella, Islem Rekik, Nadine Hahner, Elisenda Eixarch, Dinggang Shen, Gang Li 0001, Miguel Ángel González Ballester, Gerard Sanroma |
Medical Image Anal. | 6 |
| 2020 | Designing weighted correlation kernels in convolutional neural networks for functional connectivity based brain disease diagnosis
Biao Jie, Mingxia Liu 0001, Chunfeng Lian, Feng Shi 0001, Dinggang Shen |
Medical Image Anal. | 5 |
| 2020 | Synthesized 7T MRI from 3T MRI via deep learning in spatial and wavelet domains
Liangqiong Qu, Yongqin Zhang, Shuai Wang 0002, Pew-Thian Yap, Dinggang Shen |
Medical Image Anal. | 5 |
| 2020 | Domain-invariant interpretable fundus image quality assessment
Yaxin Shen, Bin Sheng 0001, Ruogu Fang, Huating Li, Skylar E. Stolte, Harry Qin, Weiping Jia, Dinggang Shen |
Medical Image Anal. | 9 |
| 2020 | SLIR: Synthesis, localization, inpainting, and registration for image-guided thermal ablation of liver tumors
Dongming Wei, Sahar Ahmad, Jiayu Huo, Pu Huang 0001, Pew-Thian Yap, Zhong Xue, Jianqi Sun, Dinggang Shen, Qian Wang 0001 |
Medical Image Anal. | 9 |
| 2020 | Mitigating gyral bias in cortical tractography via asymmetric fiber orientation distributions
Ye Wu 0001, Yoonmi Hong, Yuanjing Feng, Dinggang Shen, Pew-Thian Yap |
Medical Image Anal. | 4 |
| 2020 | Context-guided fully convolutional networks for joint craniomaxillofacial bone segmentation and landmark digitization
Jun Zhang 0018, Mingxia Liu 0001, Li Wang 0026, Peng Yuan 0001, Jianfu Li, Steve G. Shen, Ken-Chung Chen, James J. Xia, Dinggang Shen |
Medical Image Anal. | 11 |
| 2020 | Multi-modal latent space inducing ensemble SVM classifier for early dementia diagnosis with neuroimaging data
Tao Zhou 0002, Kim-Han Thung, Mingxia Liu 0001, Feng Shi 0001, Changqing Zhang 0002, Dinggang Shen |
Medical Image Anal. | 6 |
| 2020 | Hierarchical Fully Convolutional Network for Joint Atrophy Localization and Alzheimer's Disease Diagnosis Using Structural MRIabstractStructural magnetic resonance imaging (sMRI) has been widely used for computer-aided diagnosis of neurodegenerative disorders, e.g., Alzheimer's disease (AD), due to its sensitivity to morphological changes caused by brain atrophy. Recently, a few deep learning methods (e.g., convolutional neural networks, CNNs) have been proposed to learn task-oriented features from sMRI for AD diagnosis, and achieved superior performance than the conventional learning-based methods using hand-crafted features. However, these existing CNN-based methods still require the pre-determination of informative locations in sMRI. That is, the stage of discriminative atrophy localization is isolated to the latter stages of feature extraction and classifier construction. In this paper, we propose a hierarchical fully convolutional network (H-FCN) to automatically identify discriminative local patches and regions in the whole brain sMRI, upon which multi-scale feature representations are then jointly learned and fused to construct hierarchical classification models for AD diagnosis. Our proposed H-FCN method was evaluated on a large cohort of subjects from two independent datasets (i.e., ADNI-1 and ADNI-2), demonstrating good performance on joint discriminative atrophy localization and brain disease diagnosis. Chunfeng Lian, Mingxia Liu 0001, Jun Zhang 0018, Dinggang Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Multiple Kernel $k$k-Means with Incomplete KernelsabstractMultiple kernel clustering (MKC) algorithms optimally combine a group of pre-specified base kernel matrices to improve clustering performance. However, existing MKC algorithms cannot efficiently address the situation where some rows and columns of base kernel matrices are absent. This paper proposes two simple yet effective algorithms to address this issue. Different from existing approaches where incomplete kernel matrices are first imputed and a standard MKC algorithm is applied to the imputed kernel matrices, our first algorithm integrates imputation and clustering into a unified learning procedure. Specifically, we perform multiple kernel clustering directly with the presence of incomplete kernel matrices, which are treated as auxiliary variables to be jointly optimized. Our algorithm does not require that there be at least one complete base kernel matrix over all the samples. Also, it adaptively imputes incomplete kernel matrices and combines them to best serve clustering. Moreover, we further improve this algorithm by encouraging these incomplete kernel matrices to mutually complete each other. The three-step iterative algorithm is designed to solve the resultant optimization problems. After that, we theoretically study the generalization bound of the proposed algorithms. Extensive experiments are conducted on 13 benchmark data sets to compare the proposed algorithms with existing imputation-based methods. Our algorithms consistently achieve superior performance and the improvement becomes more significant with increasing missing ratio, verifying the effectiveness and advantages of the proposed joint imputation and clustering. Xinwang Liu 0002, Xinzhong Zhu, Miaomiao Li 0001, Lei Wang 0001, En Zhu, Tongliang Liu, Marius Kloft, Dinggang Shen, Jianping Yin, Wen Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2020 | Deep morphological simplification network (MS-Net) for guided registration of brain magnetic resonance images
Dongming Wei, Lichi Zhang, Zhengwang Wu, Xiaohuan Cao, Gang Li 0001, Dinggang Shen, Qian Wang 0001 |
Pattern Recognit. | 6 |
| 2020 | Population-guided large margin classifier for high-dimension low-sample-size problems
Qingbo Yin, Ehsan Adeli-Mosabbeb, Liran Shen, Dinggang Shen |
Pattern Recognit. | 4 |
| 2020 | Weakly Supervised Deep Learning for Brain Disease Prognosis Using MRI and Incomplete Clinical ScoresabstractAs a hot topic in brain disease prognosis, predicting clinical measures of subjects based on brain magnetic resonance imaging (MRI) data helps to assess the stage of pathology and predict future development of the disease. Due to incomplete clinical labels/scores, previous learning-based studies often simply discard subjects without ground-truth scores. This would result in limited training data for learning reliable and robust models. Also, existing methods focus only on using hand-crafted features (e.g., image intensity or tissue volume) of MRI data, and these features may not be well coordinated with prediction models. In this paper, we propose a weakly supervised densely connected neural network (wiseDNN) for brain disease prognosis using baseline MRI data and incomplete clinical scores. Specifically, we first extract multiscale image patches (located by anatomical landmarks) from MRI to capture local-to-global structural information of images, and then develop a weakly supervised densely connected network for task-oriented extraction of imaging features and joint prediction of multiple clinical measures. A weighted loss function is further employed to make full use of all available subjects (even those without ground-truth scores at certain time-points) for network training. The experimental results on 1469 subjects from both ADNI-1 and ADNI-2 datasets demonstrate that our proposed method can efficiently predict future clinical measures of subjects. Mingxia Liu 0001, Jun Zhang 0018, Chunfeng Lian, Dinggang Shen |
IEEE Trans. Cybern. | 4 |
| 2020 | Real-Time Quality Assessment of Pediatric MRI via Semi-Supervised Deep Nonlocal Residual Neural NetworksabstractIn this paper, we introduce an image quality assessment (IQA) method for pediatric T1- and T2-weighted MR images. IQA is first performed slice-wise using a nonlocal residual neural network (NR-Net) and then volume-wise by agglomerating the slice QA results using random forest. Our method requires only a small amount of quality-annotated images for training and is designed to be robust to annotation noise that might occur due to rater errors and the inevitable mix of good and bad slices in an image volume. Using a small set of quality-assessed images, we pre-train NR-Net to annotate each image slice with an initial quality rating (i.e., pass, questionable, fail), which we then refine by semi-supervised learning and iterative self-training. Experimental results demonstrate that our method, trained using only samples of modest size, exhibit great generalizability, capable of real-time (milliseconds per volume) large-scale IQA with nearperfect accuracy. Siyuan Liu 0004, Kim-Han Thung, Weili Lin, Pew-Thian Yap, Dinggang Shen |
IEEE Trans. Image Process. | 5 |
| 2020 | Task Decomposition and Synchronization for Semantic Biomedical Image SegmentationabstractSemantic segmentation is essentially important to biomedical image analysis. Many recent works mainly focus on integrating the Fully Convolutional Network (FCN) architecture with sophisticated convolution implementation and deep supervision. Such complex networks need large training datasets, a requirement which is challenging for medical image analysis. In this paper, we propose to decompose the single segmentation task into three subsequent sub-tasks, including (1) pixel-wise image semantic segmentation, (2) prediction of the instance class labels of the objects within the image, and (3) classification of the scene the image belonging to. While these three sub-tasks are trained to optimize their individual loss functions at different perceptual levels, we propose to allow their interaction within the task-task context ensemble. Moreover, we propose a novel sync-regularization to penalize the deviation between the outputs of the pixel-wise semantic segmentation and the instance class prediction tasks. These effective regularizations help FCN utilize context information comprehensively and attain accurate segmentation, even though the number of images for training may be limited in many biomedical applications. We have successfully applied our framework to three diverse 2D/3D medical image datasets, including Robotic Scene Segmentation Challenge 18 (ROBOT18), Brain Tumor Segmentation Challenge 18 (BRATS18), and Retinal Fundus Glaucoma Challenge (REFUGE18). We have achieved outperformed or comparable performance in all the three challenges. Our code, typical data and trained models are available athttps://github.com/xuhuaren/TDSNet. Xuhua Ren, Sahar Ahmad, Lichi Zhang, Lei Xiang 0001, Dong Nie, Fan Yang 0054, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Image Process. | 8 |
| 2020 | Multi-Atlas Brain Parcellation Using Squeeze-and-Excitation Fully Convolutional NetworksabstractMulti-atlas parcellation (MAP) is carried out on a brain image by propagating and fusing labelled regions from brain atlases. Typical nonlinear registration-based label propagation is time-consuming and sensitive to inter-subject differences. Recently, deep learning parcellation (DLP) has been proposed to avoid nonlinear registration for better efficiency and robustness than MAP. However, most existing DLP methods neglect using brain atlases, which contain high-level information (e.g., manually labelled brain regions), to provide auxiliary features for improving the parcellation accuracy. In this paper, we propose a novel multi-atlas DLP method for brain parcellation. Our method is based on fully convolutional networks (FCN) and squeeze-and-excitation (SE) modules. It can automatically and adaptively select features from the most relevant brain atlases to guide parcellation. Moreover, our method is trained via a generative adversarial network (GAN), where a convolutional neural network (CNN) with multi-scale l1loss is used as the discriminator. Benefiting from brain atlases, our method outperforms MAP and state-of-the-art DLP methods on two public image datasets (LPBA40 and NIREP-NA0). Zhenyu Tang 0002, Xianli Liu, Yang Li 0010, Pew-Thian Yap, Dinggang Shen |
IEEE Trans. Image Process. | 5 |
| 2020 | High-Resolution Encoder-Decoder Networks for Low-Contrast Medical Image SegmentationabstractAutomatic image segmentation is an essential step for many medical image analysis applications, include computer-aided radiation therapy, disease diagnosis, and treatment effect evaluation. One of the major challenges for this task is the blurry nature of medical images (e.g., CT, MR and, microscopic images), which can often result in low-contrast and vanishing boundaries. With the recent advances in convolutional neural networks, vast improvements have been made for image segmentation, mainly based on the skip-connection-linked encoder-decoder deep architectures. However, in many applications (with adjacent targets in blurry images), these models often fail to accurately locate complex boundaries and properly segment tiny isolated parts. In this paper, we aim to provide a method for blurry medical image segmentation and argue that skip connections are not enough to help accurately locate indistinct boundaries. Accordingly, we propose a novel high-resolution multi-scale encoder-decoder network (HMEDN), in which multi-scale dense connections are introduced for the encoder-decoder structure to finely exploit comprehensive semantic information. Besides skip connections, extra deeply-supervised high-resolution pathways (comprised of densely connected dilated convolutions) are integrated to collect high-resolution semantic information for accurate boundary localization. These pathways are paired with a difficulty-guided cross-entropy loss function and a contour regression task to enhance the quality of boundary detection. Extensive experiments on a pelvic CT image dataset, a multi-modal brain tumor dataset, and a cell segmentation dataset show the effectiveness of our method for 2D/3D semantic segmentation and 2D instance segmentation, respectively. Our experimental results also show that besides increasing the network complexity, raising the resolution of semantic feature maps can largely affect the overall model performance. For different tasks, finding a balance between these two factors can further improve the performance of the corresponding network. Sihang Zhou 0001, Dong Nie, Ehsan Adeli-Mosabbeb, Jianping Yin, Jun Lian, Dinggang Shen |
IEEE Trans. Image Process. | 6 |
| 2020 | Adaptive Feature Selection Guided Deep Forest for COVID-19 Classification With Chest CTabstractChest computed tomography (CT) becomes an effective tool to assist the diagnosis of coronavirus disease-19 (COVID-19). Due to the outbreak of COVID-19 worldwide, using the computed-aided diagnosis technique for COVID-19 classification based on CT images could largely alleviate the burden of clinicians. In this paper, we propose an Adaptive Feature Selection guided Deep Forest (AFS-DF) for COVID-19 classification based on chest CT images. Specifically, we first extract location-specific features from CT images. Then, in order to capture the high-level representation of these features with the relatively small-scale data, we leverage a deep forest model to learn high-level representation of the features. Moreover, we propose a feature selection method based on the trained deep forest model to reduce the redundancy of features, where the feature selection could be adaptively incorporated with the COVID-19 classification model. We evaluated our proposed AFS-DF on COVID-19 dataset with 1495 patients of COVID-19 and 1027 patients of community acquired pneumonia (CAP). The accuracy (ACC), sensitivity (SEN), specificity (SPE), AUC, precision and F1-score achieved by our method are 91.79%, 93.05%, 89.95%, 96.35%, 93.10% and 93.07%, respectively. Experimental results on the COVID-19 dataset suggest that the proposed AFS-DF achieves superior performance in COVID-19 vs. CAP classification, compared with 4 widely used machine learning methods. Liang Sun 0009, Zhanhao Mo, Fuhua Yan, Liming Xia, Zhongxiang Ding, Bin Song 0002, Wanchun Gao, Wei Shao 0005, Feng Shi 0001, Huan Yuan, Huiting Jiang, Dijia Wu, Ying Wei 0009, Yaozong Gao, He Sui, Daoqiang Zhang, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 18 |
| 2020 | Editorial: Predictive Intelligence in Biomedical and Health InformaticsabstractThe papers in this special section examine the use of predictive intelligence for bioinformatics. Big data is fueling diverse research directions in both medical image analysis and computer vision research fields. These can be divided into two main categories: (1) analytical methods, and (2) predictive methods. While analytical methods aim to efficiently analyze, represent, and interpret data, predictive methods leverage the data currently available to predict observations at present (e.g., by completingmissing observations), at previous time-points (e.g., by solving reverse problems), or at later time-points (i.e., forecasting the future). For instance, a method which only focuses on classifying patients with mild cognitive impairment (MCI) and patients with Alzheimer’s disease (AD) is an analytical method, while a method that predicts if a subject diagnosed with MCI will remain stable or convert to AD over time is a predictive method. Similar cases can be established for various neurodegenerative or neuropsychiatric disorders, degenerative arthritis, or cancer studies, in which the disease/disorder develops over time. Ehsan Adeli-Mosabbeb, S. H. Rekik, Sanghyun Park 0004, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | Hierarchical Rough-to-Fine Model for Infant Age Prediction Based on Cortical FeaturesabstractPrediction of the chronological age based on neuroimaging data is important for brain development analysis and brain disease diagnosis. Although many researches have been conducted for age prediction of older children and adults, little work has been dedicated to infants. To this end, this paper focuses on predicting infant age from birth to 2-year old using brain MR images, as well as identifying some related biomarkers. However, brain development during infancy is too rapid and heterogeneous to be accurately modeled by the conventional regression models. To address this issue, a two-stage prediction method is proposed. Specifically, our method first roughly predicts the age range of an infant and then finely predicts the accurate chronological age based on a learned, age-group-specific regression model. Combining this two-stage prediction method with another complementary one-stage prediction method, a hierarchical rough-to-fine (HRtoF) model is built. HRtoF effectively splits the rapid and heterogeneous changes during a long time period into several short time ranges and further mines the discrimination capability of cortical features, thus reaching high accuracy in infant age prediction. Taking 8 types of cortical morphometric features from structural MRI as predictors, the effectiveness of our proposed HRtoF model is validated using an infant dataset including 50 healthy subjects with 251 longitudinal MRI scans from 14 to 797 days. Comparing with five state-of-the-art regression methods, HRtoF model reduces the mean absolute error of the prediction from >48 days to 32.1 days. The correlation coefficient of the predicted age and the chronological age reaches 0.963. Moreover, based on HRtoF, the relative contributions of the eight types of cortical features for age prediction are also studied. Dan Hu 0004, Zhengwang Wu, Weili Lin, Gang Li 0001, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | Toward a Better Estimation of Functional Brain Network for Mild Cognitive Impairment Identification: A Transfer Learning ViewabstractMild cognitive impairment (MCI) is an intermediate stage of brain cognitive decline, associated with increasing risk of developing Alzheimer's disease (AD). It is believed that early treatment of MCI could slow down the progression of AD, and functional brain network (FBN) could provide potential imaging biomarkers for MCI diagnosis and response to treatment. However, there are still some challenges to estimate a “good” FBN, particularly due to the poor quality and limited quantity of functional magnetic resonance imaging (fMRI) data from the target domain (i.e., MCI study). Inspired by the idea of transfer learning, we attempt to transfer information in high-quality data from source domain (e.g., human connectome project in this paper) into the target domain towards a better FBN estimation, and propose a novel method, namely NERTL (Network Estimation via Regularized Transfer Learning). Specifically, we first construct a high-quality network “template” based on the source data, and then use the template to guide or constrain the target of FBN estimation by a weighted l1-norm regularizer. Finally, we conduct experiments to identify subjects with MCI from normal controls (NCs) based on the estimated FBNs. Despite its simplicity, our proposed method is more effective than the baseline methods in modeling discriminative FBNs, as demonstrated by the superior MCI classification accuracy of 82.4% and the area under curve (AUC) of 0.910. Weikai Li 0003, Lishan Qiao, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | An Effective MR-Guided CT Network Training for Segmenting Prostate in CT ImagesabstractSegmentation of prostate in medical imaging data (e.g., CT, MRI, TRUS) is often considered as a critical yet challenging task for radiotherapy treatment. It is relatively easier to segment prostate from MR images than from CT images, due to better soft tissue contrast of the MR images. For segmenting prostate from CT images, most previous methods mainly used CT alone, and thus their performances are often limited by low tissue contrast in the CT images. In this article, we explore the possibility of using indirect guidance from MR images for improving prostate segmentation in the CT images. In particular, we propose a novel deep transfer learning approach, i.e., MR-guided CT network training (namely MICS-NET), which can employ MR images to help better learning of features in CT images for prostate segmentation. In MICS-NET, the guidance from MRI consists of two steps: (1) learning informative and transferable features from MRI and then transferring them to CT images in a cascade manner, and (2) adaptively transferring the prostate likelihood of MRI model (i.e., well-trained convnet by purely using MR images) with a view consistency constraint. To illustrate the effectiveness of our approach, we evaluate MICS-NET on a real CT prostate image set, with the manual delineations available as the ground truth for evaluation. Our methods generate promising segmentation results which achieve (1) six percentages higher Dice Ratio than the CT model purely using CT images and (2) comparable performance with the MRI model purely using MR images. Wanqi Yang, Yinghuan Shi, Sanghyun Park 0004, Ming Yang 0014, Yang Gao 0001, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 6 |
| 2020 | One-Shot Generative Adversarial Learning for MRI Segmentation of Craniomaxillofacial Bony StructuresabstractCompared to computed tomography (CT), magnetic resonance imaging (MRI) delineation of craniomaxillofacial (CMF) bony structures can avoid harmful radiation exposure. However, bony boundaries are blurry in MRI, and structural information needs to be borrowed from CT during the training. This is challenging since paired MRI-CT data are typically scarce. In this paper, we propose to make full use of unpaired data, which are typically abundant, along with a single paired MRI-CT data to construct a one-shot generative adversarial model for automated MRI segmentation of CMF bony structures. Our model consists of a cross-modality image synthesis sub-network, which learns the mapping between CT and MRI, and an MRI segmentation sub-network. These two sub-networks are trained jointly in an end-to-end manner. Moreover, in the training phase, a neighbor-based anchoring method is proposed to reduce the ambiguity problem inherent in cross-modality synthesis, and a feature-matching-based semantic consistency constraint is proposed to encourage segmentation-oriented MRI synthesis. Experimental results demonstrate the superiority of our method both qualitatively and quantitatively in comparison with the state-of-the-art MRI segmentation methods. Xu Chen 0020, James J. Xia, Dinggang Shen, Chunfeng Lian, Li Wang 0026, Hannah H. Deng, Steve H. Fung, Dong Nie, Kim-Han Thung, Pew-Thian Yap, Jaime Gateno |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Erratum to "Deep Learning for Fast and Spatially Constrained Tissue Quantification From Highly Accelerated Data in Magnetic Resonance Fingerprinting"
Zhenghan Fang, Yong Chen 0026, Mingxia Liu 0001, Lei Xiang 0001, Qian Zhang 0066, Qian Wang 0001, Weili Lin, Dinggang Shen |
IEEE Trans. Medical Imaging | 8 |
| 2020 | Disentangled-Multimodal Adversarial Autoencoder: Application to Infant Age Prediction With Incomplete Multimodal NeuroimagesabstractEffective fusion of structural magnetic resonance imaging (sMRI) and functional magnetic resonance imaging (fMRI) data has the potential to boost the accuracy of infant age prediction thanks to the complementary information provided by different imaging modalities. However, functional connectivity measured by fMRI during infancy is largely immature and noisy compared to the morphological features from sMRI, thus making the sMRI and fMRI fusion for infant brain analysis extremely challenging. With the conventional multimodal fusion strategies, adding fMRI data for age prediction has a high risk of introducing more noises than useful features, which would lead to reduced accuracy than that merely using sMRI data. To address this issue, we develop a novel model termed as disentangled-multimodal adversarial autoencoder (DMM-AAE) for infant age prediction based on multimodal brain MRI. Specifically, we disentangle the latent variables of autoencoder into common and specific codes to represent the shared and complementary information among modalities, respectively. Then, cross-reconstruction requirement and common-specific distance ratio loss are designed as regularizations to ensure the effectiveness and thoroughness of the disentanglement. By arranging relatively independent autoencoders to separate the modalities and employing disentanglement under cross-reconstruction requirement to integrate them, our DMM-AAE method effectively restrains the possible interference cross modalities, while realizing effective information fusion. Taking advantage of the latent variable disentanglement, a new strategy is further proposed and embedded into DMM-AAE to address the issue of incompleteness of the multimodal neuroimages, which can also be used as an independent algorithm for missing modality imputation. By taking six types of cortical morphometric features from sMRI and brain functional connectivity from fMRI as predictors, the superiority of the proposed DMM-AAE is validated on infant age (35 to 848 days after birth) prediction using incomplete multimodal neuroimages. The mean absolute error of the prediction based on DMM-AAE reaches 37.6 days, outperforming state-of-the-art methods. Generally, our proposed DMM-AAE can serve as a promising model for prediction with multimodal data. Dan Hu 0004, Han Zhang 0002, Zhengwang Wu, Fan Wang 0023, Li Wang 0026, J. Keith Smith, Weili Lin, Gang Li 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 9 |
| 2020 | Probing Tissue Microarchitecture of the Baby Brain via Spherical Mean Spectrum ImagingabstractDuring the first years of life, the human brain undergoes dynamic spatially-heterogeneous changes, invo- lving differentiation of neuronal types, dendritic arbori- zation, axonal ingrowth, outgrowth and retraction, synaptogenesis, and myelination. To better quantify these changes, this article presents a method for probing tissue microarchitecture by characterizing water diffusion in a spectrum of length scales, factoring out the effects of intra-voxel orientation heterogeneity. Our method is based on the spherical means of the diffusion signal, computed over gradient directions for a set of diffusion weightings (i.e., b -values). We decompose the spherical mean profile at each voxel into a spherical mean spectrum (SMS), which essentially encodes the fractions of spin packets undergoing fine- to coarse-scale diffusion proce- sses, characterizing restricted and hindered diffusion stemming respectively from intra- and extra-cellular water compartments. From the SMS, multiple orientation distribution invariant indices can be computed, allowing for example the quantification of neurite density, microscopic fractional anisotropy ( μ FA), per-axon axial/radial diffusivity, and free/restricted isotropic diffusivity. We show that these indices can be computed for the developing brain for greater sensitivity and specificity to development related changes in tissue microstructure. Also, we demonstrate that our method, called spherical mean spectrum imaging (SMSI), is fast, accurate, and can overcome the biases associated with other state-of-the-art microstructure models. Khoi Minh Huynh, Ye Wu 0001, Xifeng Wang, Geng Chen 0001, Haiyong Wu, Kim-Han Thung, Weili Lin, Dinggang Shen, Pew-Thian Yap |
IEEE Trans. Medical Imaging | 9 |
| 2020 | Deep Learning of Static and Dynamic Brain Functional Networks for Early MCI DetectionabstractWhile convolutional neural network (CNN) has been demonstrating powerful ability to learn hierarchical spatial features from medical images, it is still difficult to apply it directly to resting-state functional MRI (rs-fMRI) and the derived brain functional networks (BFNs). We propose a novel CNN framework to simultaneously learn embedded features from BFNs for brain disease diagnosis. Since BFNs can be built by considering both static and dynamic functional connectivity (FC), we first decompose rs-fMRI into multiple static BFNs with modified independent component analysis. Then, the voxel-wise variability in dynamic FC is used to quantify BFN dynamics. A set of paired 3D images representing static/dynamic BFNs can be fed into 3D CNNs, from which we can hierarchically and simultaneously learn static/dynamic BFN features. As a result, the dynamic BFN features can complement static BFN features and, at the meantime, different BFNs can help each other toward a joint and better classification. We validate our method with a publicly accessible, large cohort of rs-fMRI dataset in early-stage mild cognitive impairment (eMCI) diagnosis, which is one of the most challenging problems to the clinicians. By comparing with a conventional method, our method shows significant diagnostic performance improvement by almost 10%. This result demonstrates the effectiveness of deep learning in preclinical Alzheimer's disease diagnosis, based on the complex and high-dimensional voxel-wise spatiotemporal patterns of the resting-state brain functional connectomics. The framework provides a new but intuitive way to fully exploit deeply embedded diagnostic features from rs-fMRI for a better-individualized diagnosis of various neurological diseases. Tae-Eui Kam, Han Zhang 0002, Zhicheng Jiao, Dinggang Shen |
IEEE Trans. Medical Imaging | 4 |
| 2020 | Diagnosis of Coronavirus Disease 2019 (COVID-19) With Structured Latent Multi-View Representation LearningabstractRecently, the outbreak of Coronavirus Disease 2019 (COVID-19) has spread rapidly across the world. Due to the large number of infected patients and heavy labor for doctors, computer-aided diagnosis with machine learning algorithm is urgently needed, and could largely reduce the efforts of clinicians and accelerate the diagnosis process. Chest computed tomography (CT) has been recognized as an informative tool for diagnosis of the disease. In this study, we propose to conduct the diagnosis of COVID-19 with a series of features extracted from CT images. To fully explore multiple features describing CT images from different views, a unified latent representation is learned which can completely encode information from different aspects of features and is endowed with promising class structure for separability. Specifically, the completeness is guaranteed with a group of backward neural networks (each for one type of features), while by using class labels the representation is enforced to be compact within COVID-19/community-acquired pneumonia (CAP) and also a large margin is guaranteed between different types of pneumonia. In this way, our model can well avoid overfitting compared to the case of directly projecting high-dimensional features into classes. Extensive experimental results show that the proposed method outperforms all comparison methods, and rather stable performances are observed when varying the number of training data. Hengyuan Kang, Liming Xia, Fuhua Yan, Zhibin Wan, Feng Shi 0001, Huan Yuan, Huiting Jiang, Dijia Wu, He Sui, Changqing Zhang 0002, Dinggang Shen |
IEEE Trans. Medical Imaging | 11 |
| 2020 | A Multi-Organ Nucleus Segmentation ChallengeabstractGeneralized nucleus segmentation techniques can contribute greatly to reducing the time to develop and validate visual biomarkers for new digital pathology datasets. We summarize the results of MoNuSeg 2018 Challenge whose objective was to develop generalizable nuclei segmentation techniques in digital pathology. The challenge was an official satellite event of the MICCAI 2018 conference in which 32 teams with more than 80 participants from geographically diverse institutes participated. Contestants were given a training set with 30 images from seven organs with annotations of 21,623 individual nuclei. A test dataset with 14 images taken from seven organs, including two organs that did not appear in the training set was released without annotations. Entries were evaluated based on average aggregated Jaccard index (AJI) on the test set to prioritize accurate instance segmentation as opposed to mere semantic segmentation. More than half the teams that completed the challenge outperformed a previous baseline. Among the trends observed that contributed to increased accuracy were the use of color normalization as well as heavy data augmentation. Additionally, fully convolutional networks inspired by variants of U-Net, FCN, and Mask-RCNN were popularly used, typically based on ResNet or VGG base architectures. Watershed segmentation on predicted semantic segmentation maps was a popular post-processing strategy. Several of the top techniques compared favorably to an individual human annotator and can be used with confidence for nuclear morphometrics. Neeraj Kumar 0002, Ruchika Verma, Deepak Anand, Yanning Zhou 0001, Omer Fahri Onder, Efstratios Tsougenis, Hao Chen 0011, Pheng-Ann Heng, Jiahui Li 0005, Navid Alemi Koohbanani, Mostafa Jahanifar, Neda Zamani Tajeddin, Ali Gooya, Nasir M. Rajpoot, Xuhua Ren, Sihang Zhou 0001, Qian Wang 0001, Dinggang Shen, Cheng-Kun Yang, Chi-Hung Weng, Wei-Hsiang Yu, Chao-Yuan Yeh, Shuoyu Xu, Pak-Hei Yeung, Amirreza Mahbod, Gerald Schaefer, Isabella Ellinger, Rupert Ecker, Örjan Smedby, Chunliang Wang, Benjamin Chidester, Vinh Ton-That, Minh-Triet Tran, Jian Ma 0004, Minh N. Do, Simon Graham, Quoc Dang Vu, Jin Tae Kwak, Akshaykumar Gunda, Raviteja Chunduri, Corey Hu, Dariush Lotfi, Reza Safdari, Antanas Kascenas, Alison O'Neil, Dennis Eschweiler, Johannes Stegmaier, Yanping Cui, Kailin Chen, Xinmei Tian 0001, Philipp Grüning, Erhardt Barth, Elad Arbel, Itay Remer, Amir Ben-Dor, Ekaterina Sirazitdinova, Matthias Kohl, Stefan Braunewell, Yuexiang Li, Xinpeng Xie, LinLin Shen, Jun Ma 0016, Krishanu Das Baksi, Mohammad Azam Khan, Jaegul Choo, Adrián Colomer, Valery Naranjo, Linmin Pei, Khan M. Iftekharuddin, Kaushiki Roy, Debotosh Bhattacharjee, Aníbal Pedraza, Gloria Bueno García, Sabarinathan Devanathan, Saravanan Radhakrishnan, Praveen Koduganty, Zihan Wu 0001, Guanyu Cai, Amit Sethi |
IEEE Trans. Medical Imaging | 20 |
| 2020 | Deep Multi-Scale Mesh Feature Learning for Automated Labeling of Raw Dental Surfaces From 3D Intraoral ScannersabstractPrecisely labeling teeth on digitalized 3D dental surface models is the precondition for tooth position rearrangements in orthodontic treatment planning. However, it is a challenging task primarily due to the abnormal and varying appearance of patients' teeth. The emerging utilization of intraoral scanners (IOSs) in clinics further increases the difficulty in automated tooth labeling, as the raw surfaces acquired by IOS are typically low-quality at gingival and deep intraoral regions. In recent years, some pioneering end-to-end methods (e.g., PointNet) have been proposed in the communities of computer vision and graphics to consume directly raw surface for 3D shape segmentation. Although these methods are potentially applicable to our task, most of them fail to capture fine-grained local geometric context that is critical to the identification of small teeth with varying shapes and appearances. In this paper, we propose an end-to-end deep-learning method, called MeshSegNet, for automated tooth labeling on raw dental surfaces. Using multiple raw surface attributes as inputs, MeshSegNet integrates a series of graph-constrained learning modules along its forward path to hierarchically extract multi-scale local contextual features. Then, a dense fusion strategy is applied to combine local-to-global geometric features for the learning of higher-level features for mesh cell annotation. The predictions produced by our MeshSegNet are further post-processed by a graph-cut refinement step for final segmentation. We evaluated MeshSegNet using a real-patient dataset consisting of raw maxillary surfaces acquired by 3D IOS. Experimental results, performed 5-fold cross-validation, demonstrate that MeshSegNet significantly outperforms state-of-the-art deep learning methods for 3D shape segmentation. Chunfeng Lian, Li Wang 0026, Tai-Hsien Wu, Fan Wang 0023, Pew-Thian Yap, Ching-Chang Ko, Dinggang Shen |
IEEE Trans. Medical Imaging | 7 |
| 2020 | Multi-View Spatial Aggregation Framework for Joint Localization and Segmentation of Organs at Risk in Head and Neck CT ImagesabstractAccurate segmentation of organs at risk (OARs) from head and neck (H&N) CT images is crucial for effective H&N cancer radiotherapy. However, the existing deep learning methods are often not trained in an end-to-end fashion, i.e., they independently predetermine the regions of target organs before organ segmentation, causing limited information sharing between related tasks and thus leading to suboptimal segmentation results. Furthermore, when conventional segmentation network is used to segment all the OARs simultaneously, the results often favor big OARs over small OARs. Thus, the existing methods often train a specific model for each OAR, ignoring the correlation between different segmentation tasks. To address these issues, we propose a new multi-view spatial aggregation framework for joint localization and segmentation of multiple OARs using H&N CT images. The core of our framework is a proposed region-of-interest (ROI)-based fine-grained representation convolutional neural network (CNN), which is used to generate multi-OAR probability maps from each 2D view (i.e., axial, coronal, and sagittal view) of CT images. Specifically, our ROI-based fine-grained representation CNN (1) unifies the OARs localization and segmentation tasks and trains them in an end-to-end fashion, and (2) improves the segmentation results of various-sized OARs via a novel ROI-based fine-grained representation. Our multi-view spatial aggregation framework then spatially aggregates and assembles the generated multi-view multi-OAR probability maps to segment all the OARs simultaneously. We evaluate our framework using two sets of H&N CT images and achieve competitive and highly robust segmentation performance for OARs of various sizes. Shujun Liang, Kim-Han Thung, Dong Nie, Yu Zhang 0064, Dinggang Shen |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Hierarchical Nonlocal Residual Networks for Image Quality Assessment of Pediatric Diffusion MRI With Limited and Noisy AnnotationsabstractFast and automated image quality assessment (IQA) of diffusion MR images is crucial for making timely decisions for rescans. However, learning a model for this task is challenging as the number of annotated data is limited and the annotation labels might not always be correct. As a remedy, we will introduce in this paper an automatic image quality assessment (IQA) method based on hierarchical non-local residual networks for pediatric diffusion MR images. Our IQA is performed in three sequential stages, i.e., 1) slice-wise IQA, where a nonlocal residual network is first pre-trained to annotate each slice with an initial quality rating (i.e., pass/questionable/fail), which is subsequently refined via iterative semi-supervised learning and slice self-training; 2) volume-wise IQA, which agglomerates the features extracted from the slices of a volume, and uses a nonlocal network to annotate the quality rating for each volume via iterative volume self-training; and 3) subject-wise IQA, which ensembles the volumetric IQA results to determine the overall image quality pertaining to a subject. Experimental results demonstrate that our method, trained using only samples of modest size, exhibits great generalizability, and is capable of conducting rapid hierarchical IQA with near-perfect accuracy. Siyuan Liu 0004, Kim-Han Thung, Weili Lin, Dinggang Shen, Pew-Thian Yap |
IEEE Trans. Medical Imaging | 4 |
| 2020 | Dual-Sampling Attention Network for Diagnosis of COVID-19 From Community Acquired PneumoniaabstractThe coronavirus disease (COVID-19) is rapidly spreading all over the world, and has infected more than 1,436,000 people in more than 200 countries and territories as of April 9, 2020. Detecting COVID-19 at early stage is essential to deliver proper healthcare to the patients and also to protect the uninfected population. To this end, we develop a dual-sampling attention network to automatically diagnose COVID-19 from the community acquired pneumonia (CAP) in chest computed tomography (CT). In particular, we propose a novel online attention module with a 3D convolutional network (CNN) to focus on the infection regions in lungs when making decisions of diagnoses. Note that there exists imbalanced distribution of the sizes of the infection regions between COVID-19 and CAP, partially due to fast progress of COVID-19 after symptom onset. Therefore, we develop a dual-sampling strategy to mitigate the imbalanced learning. Our method is evaluated (to our best knowledge) upon the largest multi-center CT data for COVID-19 from 8 hospitals. In the training-validation stage, we collect 2186 CT scans from 1588 patients for a 5-fold cross-validation. In the testing stage, we employ another independent large-scale testing dataset including 2796 CT scans from 2057 patients. Results show that our algorithm can identify the COVID-19 images with the area under the receiver operating characteristic curve (AUC) value of 0.944, accuracy of 87.5%, sensitivity of 86.9%, specificity of 90.1%, and F1-score of 82.0%. With this performance, the proposed algorithm could potentially aid radiologists with COVID-19 diagnosis from CAP, especially in the early stage of the COVID-19 outbreak. Xi Ouyang, Jiayu Huo, Liming Xia, Jun Liu 0075, Zhanhao Mo, Fuhua Yan, Zhongxiang Ding, Bin Song 0002, Feng Shi 0001, Huan Yuan, Ying Wei 0009, Xiaohuan Cao, Yaozong Gao, Dijia Wu, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 18 |
| 2020 | Spatially-Constrained Fisher Representation for Brain Disease Identification With Incomplete Multi-Modal NeuroimagesabstractMulti-modal neuroimages, such as magnetic resonance imaging (MRI) and positron emission tomography (PET), can provide complementary structural and functional information of the brain, thus facilitating automated brain disease identification. Incomplete data problem is unavoidable in multi-modal neuroimage studies due to patient dropouts and/or poor data quality. Conventional methods usually discard data-missing subjects, thus significantly reducing the number of training samples. Even though several deep learning methods have been proposed, they usually rely on pre-defined regions-of-interest in neuroimages, requiring disease-specific expert knowledge. To this end, we propose a spatially-constrained Fisher representation framework for brain disease diagnosis with incomplete multi-modal neuroimages. We first impute missing PET images based on their corresponding MRI scans using a hybrid generative adversarial network. With the complete (after imputation) MRI and PET data, we then develop a spatially-constrained Fisher representation network to extract statistical descriptors of neuroimages for disease diagnosis, assuming that these descriptors follow a Gaussian mixture model with a strong spatial constraint (i.e., images from different subjects have similar anatomical structures). Experimental results on three databases suggest that our method can synthesize reasonable neuroimages and achieve promising results in brain disease identification, compared with several state-of-the-art methods. Yongsheng Pan, Mingxia Liu 0001, Chunfeng Lian, Yong Xia 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Guest Editorial: Special Issue on Imaging-Based Diagnosis of COVID-19abstractThe novel coronavirus 2019 (COVID-19) began infecting humans in late 2019 and then turned into pandemic in the successive months spreading all over the world. At the beginning of July 2020, the global number of confirmed cases reported by the World Health Organization is above 10 million, with more than half million deaths and a rate of new cases of almost 150 000 per day. Dinggang Shen, Yaozong Gao, Arrate Muñoz-Barrutia, Delia Cabrera DeBuc, Gennaro Percannella |
IEEE Trans. Medical Imaging | 1 |
| 2020 | Deep Learning of Imaging Phenotype and Genotype for Predicting Overall Survival Time of Glioblastoma PatientsabstractGlioblastoma (GBM) is the most common and deadly malignant brain tumor. For personalized treatment, an accurate pre-operative prognosis for GBM patients is highly desired. Recently, many machine learning-based methods have been adopted to predict overall survival (OS) time based on the pre-operative mono- or multi-modal imaging phenotype. The genotypic information of GBM has been proven to be strongly indicative of the prognosis; however, this has not been considered in the existing imaging-based OS prediction methods. The main reason is that the tumor genotype is unavailable pre-operatively unless deriving from craniotomy. In this paper, we propose a new deep learning-based OS prediction method for GBM patients, which can derive tumor genotype-related features from pre-operative multimodal magnetic resonance imaging (MRI) brain data and feed them to OS prediction. Specifically, we propose a multi-task convolutional neural network (CNN) to accomplish both tumor genotype and OS prediction tasks jointly. As the network can benefit from learning tumor genotype-related features for genotype prediction, the accuracy of predicting OS time can be prominently improved. In the experiments, multimodal MRI brain dataset of 120 GBM patients, with as many as four different genotypic/molecular biomarkers, are used to evaluate our method. Our method achieves the highest OS prediction accuracy compared to other state-of-the-art methods. Zhenyu Tang 0002, Yuyun Xu, Lei Jin 0006, Abudumijiti Aibaidula, Zhicheng Jiao, Jinsong Wu 0002, Han Zhang 0002, Dinggang Shen |
IEEE Trans. Medical Imaging | 9 |
| 2020 | CT Male Pelvic Organ Segmentation via Hybrid Loss Network With Incomplete AnnotationabstractSufficient data with complete annotation is essential for training deep models to perform automatic and accurate segmentation of CT male pelvic organs, especially when such data is with great challenges such as low contrast and large shape variation. However, manual annotation is expensive in terms of both finance and human effort, which usually results in insufficient completely annotated data in real applications. To this end, we propose a novel deep framework to segment male pelvic organs in CT images with incomplete annotation delineated in a very user-friendly manner. Specifically, we design a hybrid loss network derived from both voxel classification and boundary regression, to jointly improve the organ segmentation performance in an iterative way. Moreover, we introduce a label completion strategy to complete the labels of the rich unannotated voxels and then embed them into the training data to enhance the model capability. To reduce the computation complexity and improve segmentation performance, we locate the pelvic region based on salient bone structures to focus on the candidate segmentation organs. Experimental results on a large planning CT pelvic organ dataset show that our proposed method with incomplete annotation achieves comparable segmentation performance to the state-of-the-art methods with complete annotation. Moreover, our proposed method requires much less effort of manual contouring from medical professionals such that an institutional specific model can be more easily established. Shuai Wang 0003, Dong Nie, Liangqiong Qu, Yeqin Shao, Jun Lian, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 7 |
| 2020 | Identifying Autism Spectrum Disorder With Multi-Site fMRI via Low-Rank Domain AdaptationabstractAutism spectrum disorder (ASD) is a neurodevelopmental disorder that is characterized by a wide range of symptoms. Identifying biomarkers for accurate diagnosis is crucial for early intervention of ASD. While multi-site data increase sample size and statistical power, they suffer from inter-site heterogeneity. To address this issue, we propose a multi-site adaption framework via low-rank representation decomposition (maLRR) for ASD identification based on functional MRI (fMRI). The main idea is to determine a common low-rank representation for data from the multiple sites, aiming to reduce differences in data distributions. Treating one site as a target domain and the remaining sites as source domains, data from these domains are transformed (i.e., adapted) to a common space using low-rank representation. To reduce data heterogeneity between the target and source domains, data from the source domains are linearly represented in the common space by those from the target domain. We evaluated the proposed method on both synthetic and real multi-site fMRI data for ASD identification. The results suggest that our method yields superior performance over several state-of-the-art domain adaptation methods. Daoqiang Zhang, Jiashuang Huang, Pew-Thian Yap, Dinggang Shen, Mingxia Liu 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Multi-Class ASD Classification Based on Functional Connectivity and Functional Correlation Tensor via Multi-Source Domain Adaptation and Multi-View Sparse RepresentationabstractThe resting-state functional magnetic resonance imaging (rs-fMRI) reflects functional activity of brain regions by blood-oxygen-level dependent (BOLD) signals. Up to now, many computer-aided diagnosis methods based on rs-fMRI have been developed for Autism Spectrum Disorder (ASD). These methods are mostly the binary classification approaches to determine whether a subject is an ASD patient or not. However, the disease often consists of several sub-categories, which are complex and thus still confusing to many automatic classification methods. Besides, existing methods usually focus on the functional connectivity (FC) features in grey matter regions, which only account for a small portion of the rs-fMRI data. Recently, the possibility to reveal the connectivity information in the white matter regions of rs-fMRI has drawn high attention. To this end, we propose to use the patch-based functional correlation tensor (PBFCT) features extracted from rs-fMRI in white matter, in addition to the traditional FC features from gray matter, to develop a novel multi-class ASD diagnosis method in this work. Our method has two stages. Specifically, in the first stage of multi-source domain adaptation (MSDA), the source subjects belonging to multiple clinical centers (thus called as source domains) are all transformed into the same target feature space. Thus each subject in the target domain can be linearly reconstructed by the transformed subjects. In the second stage of multi-view sparse representation (MVSR), a multi-view classifier for multi-class ASD diagnosis is developed by jointly using both views of the FC and PBFCT features. The experimental results using the ABIDE dataset verify the effectiveness of our method, which is capable of accurately classifying each subject into a respective ASD sub-category. Jun Wang 0024, Lichi Zhang, Qian Wang 0001, Lei Chen 0011, Jun Shi 0004, Xiaobo Chen 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 8 |
| 2020 | Leveraging Coupled Interaction for Multimodal Alzheimer's Disease DiagnosisabstractAs the population becomes older worldwide, accurate computer-aided diagnosis for Alzheimer's disease (AD) in the early stage has been regarded as a crucial step for neurodegeneration care in recent years. Since it extracts the low-level features from the neuroimaging data, previous methods regarded this computer-aided diagnosis as a classification problem that ignored latent featurewise relation. However, it is known that multiple brain regions in the human brain are anatomically and functionally interlinked according to the current neuroscience perspective. Thus, it is reasonable to assume that the extracted features from different brain regions are related to each other to some extent. Also, the complementary information between different neuroimaging modalities could benefit multimodal fusion. To this end, we consider leveraging the coupled interactions in the feature level and modality level for diagnosis in this paper. First, we propose capturing the feature-level coupled interaction using a coupled feature representation. Then, to model the modality-level coupled interaction, we present two novel methods: 1) the coupled boosting (CB) that models the correlation of pairwise coupled-diversity on both inconsistently and incorrectly classified samples between different modalities and 2) the coupled metric ensemble (CME) that learns an informative feature projection from different modalities by integrating the intrarelation and interrelation of training samples. We systematically evaluated our methods with the AD neuroimaging initiative data set. By comparison with the baseline learning-based methods and the state-of-the-art methods that are specially developed for AD/MCI (mild cognitive impairment) diagnosis, our methods achieved the best performance with accuracy of 95.0% and 80.7% (CB), 94.9% and 79.9% (CME) for AD/NC (normal control), and MCI/NC identification, respectively. Yinghuan Shi, Heung-Il Suk, Yang Gao 0001, Seong-Whan Lee, Dinggang Shen |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2019 | Difficulty-Aware Attention Network with Confidence Learning for Medical Image SegmentationabstractMedical image segmentation is a key step for various applications, such as image-guided radiation therapy and diagnosis. Recently, deep neural networks provided promising solutions for automatic image segmentation; however, they often perform good on regular samples (i.e., easy-to-segment samples), since the datasets are dominated by easy and regular samples. For medical images, due to huge inter-subject variations or disease-specific effects on subjects, there exist several difficult-to-segment cases that are often overlooked by the previous works. To address this challenge, we propose a difficulty-aware deep segmentation network with confidence learning for end-to-end segmentation. The proposed framework has two main contributions: 1) Besides the segmentation network, we also propose a fully convolutional adversarial network for confidence learning to provide voxel-wise and region-wise confidence information for the segmentation network. We relax the adversarial learning to confidence learning by decreasing the priority of adversarial learning, so that we can avoid the training imbalance between generator and discriminator. 2) We propose a difficulty-aware attention mechanism to properly handle hard samples or hard regions considering structural information, which may go beyond the shortcomings of focal loss. We further propose a fusion module to selectively fuse the concatenated feature maps in encoder-decoder architectures. Experimental results on clinical and challenge datasets show that our proposed network can achieve state-of-the-art segmentation accuracy. Further analysis also indicates that each individual component of our proposed network contributes to the overall performance improvement. Dong Nie, Li Wang 0026, Lei Xiang 0001, Sihang Zhou 0001, Ehsan Adeli-Mosabbeb, Dinggang Shen |
AAAI | 6 |
| 2019 | Decoding EEG by Visual-guided Deep Neural NetworksabstractDecoding visual stimuli from brain activities is an interdisciplinary study of neuroscience and computer vision. With the emerging of Human-AI Collaboration, Human-Computer Interaction, and the development of advanced machine learning models, brain decoding based on deep learning attracts more attention. Electroencephalogram (EEG) is a widely used neurophysiology tool. Inspired by the success of deep learning on image representation and neural decoding, we proposed a visual-guided EEG decoding method that contains a decoding stage and a generation stage. In the classification stage, we designed a visual-guided convolutional neural network (CNN) to obtain more discriminative representations from EEG, which are applied to achieve the classification results. In the generation stage, the visual-guided EEG features are input to our improved deep generative model with a visual consistence module to generate corresponding visual stimuli. With the help of our visual-guided strategies, the proposed method outperforms traditional machine learning methods and deep learning models in the EEG decoding task. Zhicheng Jiao, Haoxuan You, Fan Yang 0054, Xin Li 0079, Han Zhang 0002, Dinggang Shen |
IJCAI | 6 |
| 2019 | Inter-modality Dependence Induced Data Recovery for MCI Conversion Prediction
Tao Zhou 0002, Kim-Han Thung, Yu Zhang 0009, Huazhu Fu, Jianbing Shen, Dinggang Shen, Ling Shao 0001 |
MICCAI (4) | 6 |
| 2019 | Interpretable Feature Learning Using Multi-output Takagi-Sugeno-Kang Fuzzy System for Multi-center ASD Diagnosis
Jun Wang 0024, Tao Zhou 0002, Zhaohong Deng, Huifang Huang, Shitong Wang 0001, Jun Shi 0004, Dinggang Shen |
MICCAI (3) | 8 |
| 2019 | Surface-Volume Consistent Construction of Longitudinal Atlases for the Early Developing Brain
Sahar Ahmad, Zhengwang Wu, Gang Li 0001, Li Wang 0026, Weili Lin, Pew-Thian Yap, Dinggang Shen |
MICCAI (2) | 7 |
| 2019 | RCA-U-Net: Residual Channel Attention U-Net for Fast Tissue Quantification in Magnetic Resonance Fingerprinting
Zhenghan Fang, Yong Chen 0026, Dong Nie, Weili Lin, Dinggang Shen |
MICCAI (3) | 5 |
| 2019 | Reconstructing High-Quality Diffusion MRI Data from Orthogonal Slice-Undersampled Data Using Graph Convolutional Neural Networks
Yoonmi Hong, Geng Chen 0001, Pew-Thian Yap, Dinggang Shen |
MICCAI (3) | 4 |
| 2019 | Deep Granular Feature-Label Distribution Learning for Neuroimaging-Based Infant Age Prediction
Dan Hu 0004, Han Zhang 0002, Zhengwang Wu, Weili Lin, Gang Li 0001, Dinggang Shen |
MICCAI (4) | 6 |
| 2019 | CoCa-GAN: Common-Feature-Learning-Based Context-Aware Generative Adversarial Network for Glioma Grading
Pu Huang 0001, Dengwang Li, Zhicheng Jiao, Dongming Wei, Guoshi Li, Qian Wang 0001, Han Zhang 0002, Dinggang Shen |
MICCAI (3) | 8 |
| 2019 | Probing Brain Micro-architecture by Orientation Distribution Invariant Identification of Diffusion Compartments
Khoi Minh Huynh, Ye Wu 0001, Geng Chen 0001, Kim-Han Thung, Haiyong Wu, Weili Lin, Dinggang Shen, Pew-Thian Yap |
MICCAI (3) | 8 |
| 2019 | Characterizing Non-Gaussian Diffusion in Heterogeneously Oriented Tissue Microenvironments
Khoi Minh Huynh, Ye Wu 0001, Kim-Han Thung, Geng Chen 0001, Weili Lin, Dinggang Shen, Pew-Thian Yap |
MICCAI (3) | 7 |
| 2019 | Early Development of Infant Brain Complex Network
Weixiong Jiang, Han Zhang 0002, Li-Ming Hsu, Dan Hu 0004, Guoshi Li, Ye Wu 0001, Dinggang Shen |
MICCAI (2) | 7 |
| 2019 | Dynamic Routing Capsule Networks for Mild Cognitive Impairment Diagnosis
Zhicheng Jiao, Pu Huang 0001, Tae-Eui Kam, Li-Ming Hsu, Ye Wu 0001, Han Zhang 0002, Dinggang Shen |
MICCAI (4) | 7 |
| 2019 | A Deep Learning Framework for Noise Component Detection from Resting-State Functional MRI
Tae-Eui Kam, Xuyun Wen, Bing Jin, Zhicheng Jiao, Li-Ming Hsu, Zhen Zhou 0004, Koji Yamashita, Sheng-Che Hung, Weili Lin, Han Zhang 0002, Dinggang Shen |
MICCAI (3) | 12 |
| 2019 | Identification of Abnormal Circuit Dynamics in Major Depressive Disorder via Multiscale Neural Modeling of Resting-State fMRI
Guoshi Li, Yanting Zheng, Ye Wu 0001, Pew-Thian Yap, Shijun Qiu, Han Zhang 0002, Dinggang Shen |
MICCAI (3) | 8 |
| 2019 | End-to-End Dementia Status Prediction from Brain MRI Using Multi-task Weakly-Supervised Attention Network
Chunfeng Lian, Mingxia Liu 0001, Li Wang 0026, Dinggang Shen |
MICCAI (4) | 4 |
| 2019 | MeshSNet: Deep Multi-scale Mesh Feature Learning for End-to-End Tooth Labeling on 3D Dental Surfaces
Chunfeng Lian, Li Wang 0026, Tai-Hsien Wu, Mingxia Liu 0001, Francisca Durán, Ching-Chang Ko, Dinggang Shen |
MICCAI (6) | 7 |
| 2019 | Multi-stage Image Quality Assessment of Diffusion MRI via Semi-supervised Nonlocal Residual Networks
Siyuan Liu 0004, Kim-Han Thung, Weili Lin, Pew-Thian Yap, Dinggang Shen |
MICCAI (3) | 5 |
| 2019 | Disease-Image Specific Generative Adversarial Network for Brain Disease Diagnosis with Incomplete Multi-modal Neuroimages
Yongsheng Pan, Mingxia Liu 0001, Chunfeng Lian, Yong Xia 0001, Dinggang Shen |
MICCAI (3) | 5 |
| 2019 | Wavelet-based Semi-supervised Adversarial Learning for Synthesizing Realistic 7T from 3T MRI
Liangqiong Qu, Shuai Wang 0002, Pew-Thian Yap, Dinggang Shen |
MICCAI (4) | 4 |
| 2019 | Pre-operative Overall Survival Time Prediction for Glioblastoma Patients Using Deep Learning on Both Imaging Phenotype and Genotype
Zhenyu Tang 0002, Yuyun Xu, Zhicheng Jiao, Lei Jin 0006, Abudumijiti Aibaidula, Jinsong Wu 0002, Qian Wang 0001, Han Zhang 0002, Dinggang Shen |
MICCAI (1) | 10 |
| 2019 | Automated Parcellation of the Cortex Using Structural Connectome Harmonics
Hoyt Patrick Taylor IV, Zhengwang Wu, Ye Wu 0001, Dinggang Shen, Han Zhang 0002, Pew-Thian Yap |
MICCAI (3) | 4 |
| 2019 | Revealing Developmental Regionalization of Infant Cerebral Cortex Based on Multiple Cortical Properties
Fan Wang 0023, Chunfeng Lian, Zhengwang Wu, Li Wang 0026, Weili Lin, John H. Gilmore, Dinggang Shen, Gang Li 0001 |
MICCAI (2) | 7 |
| 2019 | Synthesis and Inpainting-Based MR-CT Registration for Image-Guided Thermal Ablation of Liver Tumors
Dongming Wei, Sahar Ahmad, Jiayu Huo, Wen Peng, Yunhao Ge, Zhong Xue, Pew-Thian Yap, Dinggang Shen, Qian Wang 0001 |
MICCAI (5) | 9 |
| 2019 | Intrinsic Patch-Based Cortical Anatomical Parcellation Using Graph Convolutional Neural Network on Surface Manifold
Zhengwang Wu, Fenqiang Zhao, Li Wang 0026, Weili Lin, John H. Gilmore, Gang Li 0001, Dinggang Shen |
MICCAI (3) | 8 |
| 2019 | Estimating Reference Bony Shape Model for Personalized Surgical Reconstruction of Posttraumatic Facial Defects
Deqiang Xiao, Li Wang 0026, Hannah H. Deng, Kim-Han Thung, Jihua Zhu, Peng Yuan 0001, Yriu L. Rodrigues, Leonel Perez Jr., Christopher E. Crecelius, Jaime Gateno, Tiansku Kuang, Steve G. Shen, Daeseung Kim, David M. Alfi, Pew-Thian Yap, James J. Xia, Dinggang Shen |
MICCAI (5) | 17 |
| 2019 | Harmonization of Infant Cortical Thickness Using Surface-to-Surface Cycle-Consistent Adversarial Networks
Fenqiang Zhao, Zhengwang Wu, Li Wang 0026, Weili Lin, Shunren Xia, Dinggang Shen, Gang Li 0001 |
MICCAI (4) | 6 |
| 2019 | Deep Multi-modal Latent Representation Learning for Automated Dementia Diagnosis
Tao Zhou 0002, Mingxia Liu 0001, Huazhu Fu, Jun Wang 0024, Jianbing Shen, Ling Shao 0001, Dinggang Shen |
MICCAI (4) | 7 |
| 2019 | Multi-layer Temporal Network Analysis Reveals Increasing Temporal Reachability and Spreadability in the First Two Years of Life
Zhen Zhou 0004, Han Zhang 0002, Li-Ming Hsu, Weili Lin, Gang Pan 0001, Dinggang Shen |
MICCAI (3) | 6 |
| 2019 | Robust and Discriminative Brain Genome Association Study
Xiaofeng Zhu 0001, Dinggang Shen |
MICCAI (4) | 2 |
| 2019 | MIDCN: A Multiple Instance Deep Convolutional Network for Image Classification
Kelei He, Jing Huo, Yinghuan Shi, Yang Gao 0001, Dinggang Shen |
PRICAI (1) | 5 |
| 2019 | Surface-constrained volumetric registration for the early developing brain
Sahar Ahmad, Zhengwang Wu, Gang Li 0001, Li Wang 0026, Weili Lin, Pew-Thian Yap, Dinggang Shen |
Medical Image Anal. | 7 |
| 2019 | XQ-SR: Joint x-q space super-resolution with application to infant diffusion MRI
Geng Chen 0001, Bin Dong 0001, Yong Zhang 0004, Weili Lin, Dinggang Shen, Pew-Thian Yap |
Medical Image Anal. | 5 |
| 2019 | Noise reduction in diffusion MRI using non-local self-similar information in joint x-q space
Geng Chen 0001, Yafeng Wu, Dinggang Shen, Pew-Thian Yap |
Medical Image Anal. | 3 |
| 2019 | Adversarial learning for mono- or multi-modal registration
Jingfan Fan, Xiaohuan Cao, Qian Wang 0001, Pew-Thian Yap, Dinggang Shen |
Medical Image Anal. | 5 |
| 2019 | BIRNet: Brain image registration using dual-supervised fully convolutional networks
Jingfan Fan, Xiaohuan Cao, Pew-Thian Yap, Dinggang Shen |
Medical Image Anal. | 4 |
| 2019 | Automatic brain labeling via multi-atlas guided fully convolutional networks
Longwei Fang, Lichi Zhang, Dong Nie, Xiaohuan Cao, Islem Rekik, Seong-Whan Lee, Huiguang He, Dinggang Shen |
Medical Image Anal. | 8 |
| 2019 | Multimodal hyper-connectivity of functional networks using functionally-weighted LASSO for MCI classification
Yang Li 0010, Jingyu Liu 0002, Xinqiang Gao, Biao Jie, Minjeong Kim 0001, Pew-Thian Yap, Chong-Yaw Wee, Dinggang Shen |
Medical Image Anal. | 8 |
| 2019 | Automated detection and classification of thyroid nodules in ultrasound images using clinical-knowledge-guided convolutional neural networks
Qianqian Guo, Chunfeng Lian, Xuhua Ren, Shujun Liang, Jing Yu 0005, Lijuan Niu, Dinggang Shen |
Medical Image Anal. | 9 |
| 2019 | CT male pelvic organ segmentation using fully convolutional networks with boundary sensitive representation
Shuai Wang 0002, Kelei He, Dong Nie, Sihang Zhou 0001, Yaozong Gao, Dinggang Shen |
Medical Image Anal. | 6 |
| 2019 | Multi-task exclusive relationship learning for alzheimer's disease progression prediction with longitudinal data
Daoqiang Zhang, Dinggang Shen, Mingxia Liu 0001 |
Medical Image Anal. | 3 |
| 2019 | Super-resolution reconstruction of neonatal brain magnetic resonance images via residual structured sparse representation
Yongqin Zhang, Pew-Thian Yap, Geng Chen 0001, Weili Lin, Li Wang 0026, Dinggang Shen |
Medical Image Anal. | 6 |
| 2019 | Semi-Supervised Discriminative Classification Robust to Sample-Outliers and Feature-NoisesabstractDiscriminative methods commonly produce models with relatively good generalization abilities. However, this advantage is challenged in real-world applications (e.g., medical image analysis problems), in which there often exist outlier data points (sample-outliers) and noises in the predictor values (feature-noises). Methods robust to both types of these deviations are somewhat overlooked in the literature. We further argue that denoising can be more effective, if we learn the model using all the available labeled and unlabeled samples, as the intrinsic geometry of the sample manifold can be better constructed using more data points. In this paper, we propose a semi-supervised robust discriminative classification method based on the least-squares formulation of linear discriminant analysis to detect sample-outliers and feature-noises simultaneously, using both labeled training and unlabeled testing data. We conduct several experiments on a synthetic, some benchmark semi-supervised learning, and two brain neurodegenerative disease diagnosis datasets (for Parkinson's and Alzheimer's diseases). Specifically for the application of neurodegenerative diseases diagnosis, incorporating robust machine learning methods can be of great benefit, due to the noisy nature of neuroimaging data. Our results show that our method outperforms the baseline and several state-of-the-art methods, in terms of both accuracy and the area under the ROC curve. Ehsan Adeli-Mosabbeb, Kim-Han Thung, Guorong Wu 0001, Feng Shi 0001, Dinggang Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2019 | Late Fusion Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering optimally integrates a group of pre-specified incomplete views to improve clustering performance. Among various excellent solutions, multiple kernel $k$k-means with incomplete kernels forms a benchmark, which redefines the incomplete multi-view clustering as a joint optimization problem where the imputation and clustering are alternatively performed until convergence. However, the comparatively intensive computational and storage complexities preclude it from practical applications. To address these issues, we propose Late Fusion Incomplete Multi-view Clustering (LF-IMVC) which effectively and efficiently integrates the incomplete clustering matrices generated by incomplete views. Specifically, our algorithm jointly learns a consensus clustering matrix, imputes each incomplete base matrix, and optimizes the corresponding permutation matrices. We develop a three-step iterative algorithm to solve the resultant optimization problem with linear computational complexity and theoretically prove its convergence. Further, we conduct comprehensive experiments to study the proposed LF-IMVC in terms of clustering accuracy, running time, advantages of late fusion multi-view clustering, evolution of the learned consensus clustering matrix, parameter sensitivity and convergence. As indicated, our algorithm significantly and consistently outperforms some state-of-the-art algorithms with much less running time and memory. Xinwang Liu 0002, Xinzhong Zhu, Miaomiao Li 0001, Lei Wang 0001, Chang Tang, Jianping Yin, Dinggang Shen, Huaimin Wang 0001, Wen Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2019 | Structured sparsity regularized multiple kernel learning for Alzheimer's disease diagnosis
Jialin Peng, Xiaofeng Zhu 0001, Dinggang Shen |
Pattern Recognit. | 5 |
| 2019 | Weighted graph regularized sparse brain network construction for MCI identification
Renping Yu, Lishan Qiao, Mingming Chen 0005, Seong-Whan Lee, Xuan Fei, Dinggang Shen |
Pattern Recognit. | 6 |
| 2019 | Strength and similarity guided group-level brain functional network construction for MCI diagnosis
Yu Zhang 0009, Han Zhang 0002, Xiaobo Chen 0001, Mingxia Liu 0001, Xiaofeng Zhu 0001, Seong-Whan Lee, Dinggang Shen |
Pattern Recognit. | 7 |
| 2019 | 3-D Fully Convolutional Networks for Multimodal Isointense Infant Brain Image SegmentationabstractAccurate segmentation of infant brain images into different regions of interest is one of the most important fundamental steps in studying early brain development. In the isointense phase (approximately 6-8 months of age), white matter and gray matter exhibit similar levels of intensities in magnetic resonance (MR) images, due to the ongoing myelination and maturation. This results in extremely low tissue contrast and thus makes tissue segmentation very challenging. Existing methods for tissue segmentation in this isointense phase usually employ patch-based sparse labeling on single modality. To address the challenge, we propose a novel 3-D multimodal fully convolutional network (FCN) architecture for segmentation of isointense phase brain MR images. Specifically, we extend the conventional FCN architectures from 2-D to 3-D, and, rather than directly using FCN, we intuitively integrate coarse (naturally high-resolution) and dense (highly semantic) feature maps to better model tiny tissue regions, in addition, we further propose a transformation module to better connect the aggregating layers; we also propose a fusion module to better serve the fusion of feature maps. We compare the performance of our approach with several baseline and state-of-the-art methods on two sets of isointense phase brain images. The comparison results show that our proposed 3-D multimodal FCN model outperforms all previous methods by a large margin in terms of segmentation accuracy. In addition, the proposed framework also achieves faster segmentation results compared to all other methods. Our experiments further demonstrate that: 1) carefully integrating coarse and dense feature maps can considerably improve the segmentation performance; 2) batch normalization can speed up the convergence of the networks, especially when hierarchical feature aggregations occur; and 3) integrating multimodal information can further boost the segmentation performance. Dong Nie, Li Wang 0026, Ehsan Adeli-Mosabbeb, Cuijin Lao, Weili Lin, Dinggang Shen |
IEEE Trans. Cybern. | 6 |
| 2019 | Sparse Multiview Task-Centralized Ensemble Learning for ASD Diagnosis Based on Age- and Sex-Related Functional Connectivity PatternsabstractAutism spectrum disorder (ASD) is an age- and sex-related neurodevelopmental disorder that alters the brain's functional connectivity (FC). The changes caused by ASD are associated with different age- and sex-related patterns in neuroimaging data. However, most contemporary computer-assisted ASD diagnosis methods ignore the aforementioned age-/sex-related patterns. In this paper, we propose a novel sparse multiview task-centralized (Sparse-MVTC) ensemble classification method for image-based ASD diagnosis. Specifically, with the age and sex information of each subject, we formulate the classification as a multitask learning problem, where each task corresponds to learning upon a specific age/sex group. We also extract multiview features per subject to better reveal the FC changes. Then, in Sparse-MVTC learning, we select a certain central task and treat the rest as auxiliary tasks. By considering both task-task and view-view relationships between the central task and each auxiliary task, we can learn better upon the entire dataset. Finally, by selecting the central task, in turn, we are able to derive multiple classifiers for each task/group. An ensemble strategy is further adopted, such that the final diagnosis can be integrated for each subject. Our comprehensive experiments on the ABIDE database demonstrate that our proposed Sparse-MVTC ensemble learning can significantly outperform the state-of-the-art classification methods for ASD diagnosis. Jun Wang 0024, Qian Wang 0001, Han Zhang 0002, Jiawei Chen 0001, Shitong Wang 0001, Dinggang Shen |
IEEE Trans. Cybern. | 6 |
| 2019 | Longitudinally Guided Super-Resolution of Neonatal Brain Magnetic Resonance ImagesabstractNeonatal magnetic resonance (MR) images typically have low spatial resolution and insufficient tissue contrast. Interpolation methods are commonly used to upsample the images for the subsequent analysis. However, the resulting images are often blurry and susceptible to partial volume effects. In this paper, we propose a novel longitudinally guided super-resolution (SR) algorithm for neonatal images. This is motivated by the fact that anatomical structures evolve slowly and smoothly as the brain develops after birth. We propose a strategy involving longitudinal regularization, similar to bilateral filtering, in combination with low-rank and total variation constraints to solve the ill-posed inverse problem associated with image SR. Experimental results on neonatal MR images demonstrate that the proposed algorithm recovers clear structural details and outperforms state-of-the-art methods both qualitatively and quantitatively. Yongqin Zhang, Feng Shi 0001, Jian Cheng 0002, Li Wang 0026, Pew-Thian Yap, Dinggang Shen |
IEEE Trans. Cybern. | 6 |
| 2019 | Foreground Fisher Vector: Encoding Class-Relevant Foreground to Improve Image ClassificationabstractImage classification is an essential and challenging task in computer vision. Despite its prevalence, the combination of the deep convolutional neural network (DCNN) and the Fisher vector (FV) encoding method has limited performance since the class-irrelevant background used in the traditional FV encoding may result in less discriminative image features. In this paper, we propose the foreground FV (fgFV) encoding algorithm and its fast approximation for image classification. We try to separate implicitly the class-relevant foreground from the class-irrelevant background during the encoding process via tuning the weights of the partial gradients corresponding to each Gaussian component under the supervision of image labels and, then, use only those local descriptors extracted from the class-relevant foreground to estimate FVs. We have evaluated our fgFV against the widely used FV and improved FV (iFV) under the combined DCNN-FV framework and also compared them to several state-of-the-art image classification approaches on ten benchmark image datasets for the recognition of fine-grained natural species and artificial manufactures, categorization of course objects, and classification of scenes. Our results indicate that the proposed fgFV encoding algorithm can construct more discriminative image presentations from local descriptors than FV and iFV, and the combined DCNN-fgFV algorithm can improve the performance of image classification. Yongsheng Pan, Yong Xia 0001, Dinggang Shen |
IEEE Trans. Image Process. | 3 |
| 2019 | A New Multi-Atlas Registration Framework for Multimodal Pathological Images Using Conventional Monomodal Normal AtlasesabstractUsing multi-atlas registration (MAR), information carried by atlases can be transferred onto a new input image for the tasks of region of interest (ROI) segmentation, anatomical landmark detection, and so on. Conventional atlases used in MAR methods are monomodal and contain only normal anatomical structures. Therefore, the majority of MAR methods cannot handle input multimodal pathological images, which are often collected in routine image-based diagnosis. This is because registering monomodal atlases with normal appearances to multimodal pathological images involves two major problems: (1) missing imaging modalities in the monomodal atlases, and (2) influence from pathological regions. In this paper, we propose a new MAR framework to tackle these problems. In this framework, a deep learning based image synthesizers are applied for synthesizing multimodal normal atlases from conventional monomodal normal atlases. To reduce the influence from pathological regions, we further propose a multimodal lowrank approach to recover multimodal normal-looking images from multimodal pathological images. Finally, the multimodal normal atlases can be registered to the recovered multimodal images in a multi-channel way. We evaluate our MAR framework via brain ROI segmentation of multimodal tumor brain images. Due to the utilization of multimodal information and the reduced influence from pathological regions, experimental results show that registration based on our method is more accurate and robust, leading to significantly improved brain ROI segmentation compared with state-of-the-art methods. Zhenyu Tang 0002, Pew-Thian Yap, Dinggang Shen |
IEEE Trans. Image Process. | 3 |
| 2019 | Guest Editorial Skin Lesion Image Analysis for Melanoma DetectionabstractThe papers in this special section focus on the use of image analysis to detect Melanoma. Melanoma is deadliest form of skin cancer, with roughly 91,000 new cases reported every year in the US and more than 9,000 deaths. Unlike many other cancer types, the incidence rate of melanoma has been steadily increasing in the past several decades. Early diagnosis is crucial since melanoma can be cured with a simple excision, if detected early. The goals of this special issue are to summarize the state-ofthe- art in the automated analysis of skin lesion images and to provide future directions for this exciting subfield of medical image analysis. The intended audience includes researchers and practicing clinicians, who are increasingly using digital analytic tools. M. Emre Celebi 0001, Noel Codella, Alan Halpern, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 4 |
| 2019 | Functional Brain Network Estimation With Time Series Self-ScrubbingabstractFunctional brain network (FBN) is becoming an increasingly important measurement for exploring cerebral mechanisms and mining informative biomarkers that assist diagnosis of some neurodegenerative disorders. Despite its effectiveness to discover valuable hidden patterns in the human brain, the estimated FBNs are often heavily influenced by the quality of the observed data (e.g., blood oxygen level dependent signal series). In practice, a preprocessing pipeline is usually employed for improving data quality. With this in mind, some data points (volumes or time course in the time series) are still not clean enough, due to artifacts including spurious resting-state processes (head movement, mind-wandering). Therefore, not all volumes in the fMRI time series can contribute to the subsequent FBN estimation. To address this issue, we propose a novel FBN estimation method by introducing a latent variable as an indicator of the data quality, and develop an alternating optimization algorithm for jointly scrubbing the data and estimating FBN simultaneously. To further illustrate the effectiveness of the proposed method, we conduct experiments on two public datasets to identify subjects with mild cognitive impairment from normal controls based on the estimated FBNs, and achieve improved accuracies than the baseline methods. Weikai Li 0003, Lishan Qiao, Zhengxia Wang, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 5 |
| 2019 | Regression Convolutional Neural Network for Automated Pediatric Bone Age Assessment From Hand RadiographabstractSkeletal bone age assessment is a common clinical practice to investigate endocrinology, and genetic and growth disorders of children. However, clinical interpretation and bone age analyses are time-consuming, labor intensive, and often subject to inter-observer variability. This advocates the need of a fully automated method for bone age assessment. We propose a regression convolutional neural network (CNN) to automatically assess the pediatric bone age from hand radiograph. Our network is specifically trained to place more attention to those bone age related regions in the X-ray images. Specifically, we first adopt the attention module to process all images and generate the coarse/fine attention maps as inputs for the regression network. Then, the regression CNN follows the supervision of the dynamic attention loss during training; thus, it can estimate the bone age of the hard (or "outlier") images more accurately. The experimental results show that our method achieves an average discrepancy of 5.2-5.3 months between clinical and automatic bone age evaluations on two large datasets. In conclusion, we propose a fully automated deep learning solution to process X-ray images of the hand for bone age assessment, with the accuracy comparable to human experts but with much better efficiency. Xuhua Ren, Xiujun Yang, Shuai Wang 0003, Sahar Ahmad, Lei Xiang 0001, Shaun Richard Stone, Yiqiang Zhan, Dinggang Shen, Qian Wang 0001 |
IEEE J. Biomed. Health Informatics | 10 |
| 2019 | Machine Learning in Medical ImagingabstractThe papers in this special issue focus on machine learning for use in medical image processing applications. The use of machine learning in this area has become indispensable in diagnosis and treatment of many diseases. With advances in new imaging techniques, the need to take full advantage of abundant images draws more and more attention. Machine learning, including deep learning particularly, provides us a new paradigm to learn and to utilize the overwhelming volume of big imaging data smartly. Nowadays, machine learning in medical imaging has become one of the most promising and growing fields of research. The main aim of this special issue is to help advance the scientific research within the broad field of machine learning in medical imaging. The special issue was planned in conjunction with the International Workshop on Machine Learning in Medical Imaging (MLMI) 2017. Qian Wang 0001, Yinghuan Shi, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 3 |