Ling Zhang 0002

dblp:76/5973-2 · DBLP profile ↗
← Back
38ranked-venue papers
9as first author
26since 2021 · last 2026
0000-0001-8371-5252ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 25 · 9 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 3 first-author · 18 since 2021Artificial intelligence and machine learning · 13 · 12 since 2021
YearPublicationVenuePosition
2026 CT-FineBench: A Diagnostic Fidelity Benchmark for Fine-Grained Evaluation of CT Report Generation
abstract
Ruifeng Yuan, Wanxing Chang, Weiwei Cao, Bowen Shi, Zhongyu Wei, Ling Zhang, Jianpeng Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ruifeng Yuan, Wanxing Chang, Zhongyu Wei, Ling Zhang 0002
ACL (1)6
2026 Non-contrast CT esophageal varices grading through clinical prior-enhanced multi-organ analysis
Xiaoming Zhang 0008, Chunli Li, Jiacheng Hao, Yuan Gao 0017, Danyang Tu, Jianyi Qiao, Xiaoli Yin, Le Lu 0001, Ling Zhang 0002, Ke Yan 0006
Medical Image Anal.9
2026 Preoperative Prediction of Esophageal Cancer Survival in CT via Tumor and Lymph Node Context and Geometry Modeling
abstract
Esophageal cancer is one of the most lethal cancers, with 5-year survival rate of only 20%. Patient outcomes can vary significantly even though they are at the same cancer stage and receive similar treatments. Accurate prognostic prediction for esophageal cancer patients is highly desired to receive personalized precise treatment. Nevertheless, there are very few automated methods yet to fully exploit the preoperative contrast-enhanced computed tomography (CE-CT) imaging for assessing esophageal cancer prognosis. In addition to image patterns, important prognostic factors should encompass tumor size and location, as well as lymph nodes (LNs) involvement, including features such as LN number, size, spatial distribution, and their proximity to tumor. Considering these complexities, we propose a novel Tumor and LN Context-Geometry network for the preoperative prediction of esophageal cancer survival in CE-CT images. Specifically, we 1) focus on learning survival patterns of CT texture via co-attention context modeling at most informative regions, i.e., automatically segmented tumor, LNs and LN-stations; and 2) integrate tumor and LN anatomical and spatial associations into neural geometry modeling for a comprehensive learning of metastatic involvement and tumor invasion to adjacent structures. Empirical studies show our presented framework can improve overall survival prediction performances compared with existing state-of-the-art survival analysis methods, and evidently suggest that incorporating these findings into the existing esophageal cancer staging system would add its clinical values.
Yirui Wang 0002, Haoshen Li, Jiawen Yao, Lianzhen Zhong, Dazhou Guo, Ke Yan 0006, David S. Doermann, Le Lu 0001, Feiran Jiao, Tsung-Ying Ho, Ling Zhang 0002, Abudili Abuduxuku, Xianghua Ye, Dakai Jin
IEEE Trans. Medical Imaging13
2025 Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-Language Pre-Training
Zhongyi Shui, Sinuo Wang, Zeli Chen, Le Lu 0001, Xianghua Ye, Tingbo Liang, Ling Zhang 0002
ICCV11
2025 Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image Understanding
abstract
Artificial intelligence (AI) shows great potential in assisting radiologists to improve the efficiency and accuracy of medical image interpretation and diagnosis. However, a versatile AI model requires large-scale data and comprehensive annotations, which are often impractical in medical settings. Recent studies leverage radiology reports as a naturally high-quality supervision for medical images, using contrastive language-image pre-training (CLIP) to develop language-informed models for radiological image interpretation. Nonetheless, these approaches typically contrast entire images with reports, neglecting the local associations between imaging regions and report sentences, which may undermine model performance and interoperability. In this paper, we propose a fine-grained vision-language model (fVLM) for anatomy-level CT image interpretation. Specifically, we explicitly match anatomical regions of CT images with corresponding descriptions in radiology reports and perform contrastive pre-training for each anatomy individually. Fine-grained alignment, however, faces considerable false-negative challenges, mainly from the abundance of anatomy-level healthy samples and similarly diseased abnormalities, leading to ambiguous patient-level pairings. To tackle this issue, we propose identifying false negatives of both normal and abnormal samples and calibrating contrastive learning from patient-level to disease-aware pairing. We curated the largest CT dataset to date, comprising imaging and report data from 69,086 patients, and conducted a comprehensive evaluation of 54 major and important disease (including several most deadly cancers) diagnosis tasks across 15 main anatomies. Experimental results demonstrate the substantial potential of fVLM in versatile medical image interpretation. In the zero-shot classification task, we achieved an average AUC of 81.3% on 54 diagnosis tasks, surpassing CLIP and supervised methods by 12.9% and 8.0%, respectively. Additionally, on the publicly available CT-RATE and Rad-ChestCT benchmarks, our fVLM outperformed the current state-of-the-art methods with absolute AUC gains of 7.4% and 4.8%, respectively.
Zhongyi Shui, Sinuo Wang, Ruizhe Guo, Le Lu 0001, Lin Yang 0002, Xianghua Ye, Tingbo Liang, Ling Zhang 0002
ICLR11
2025 PLUS: Plug-and-Play Enhanced Liver Lesion Diagnosis Model on Non-contrast CT Scans
Jiacheng Hao, Xiaoming Zhang 0008, Wei Liu 0127, Xiaoli Yin, Yuan Gao 0017, Chunli Li, Ling Zhang 0002, Le Lu 0001, Xu Han 0023, Ke Yan 0006
MICCAI (15)7
2025 Deep Attention Learning for Pre-operative Lymph Node Metastasis Prediction in Pancreatic Cancer via Multi-object Relationship Modeling
Zhilin Zheng, Jiawen Yao, Le Lu 0001, Jianping Lu, Ling Zhang 0002, Chengwei Shao, Yun Bian
Int. J. Comput. Vis.9
2025 Correction: Deep Attention Learning for Pre-operative Lymph Node Metastasis Prediction in Pancreatic Cancer via Multi-object Relationship Modeling
Zhilin Zheng, Jiawen Yao, Le Lu 0001, Jianping Lu, Ling Zhang 0002, Chengwei Shao, Yun Bian
Int. J. Comput. Vis.9
2025 A Colorectal Coordinate-Driven Method for Colorectum and Colorectal Cancer Segmentation in Conventional CT Scans
abstract
Automated colorectal cancer (CRC) segmentation in medical imaging is the key to achieving automation of CRC detection, staging, and treatment response monitoring. Compared with magnetic resonance imaging (MRI) and computed tomography colonography (CTC), conventional computed tomography (CT) has enormous potential because of its broad implementation, superiority for the hollow viscera (colon), and convenience without needing bowel preparation. However, the segmentation of CRC in conventional CT is more challenging due to the difficulties presenting with the unprepared bowel, such as distinguishing the colorectum from other structures with similar appearance and distinguishing the CRC from the contents of the colorectum. To tackle these challenges, we introduce DeepCRC-SL, the first automated segmentation algorithm for CRC and colorectum in conventional contrast-enhanced CT scans. We propose a topology-aware deep learning-based approach, which builds a novel 1-D colorectal coordinate system and encodes each voxel of the colorectum with a relative position along the coordinate system. We then induce an auxiliary regression task to predict the colorectal coordinate value of each voxel, aiming to integrate global topology into the segmentation network and thus improve the colorectum's continuity. Self-attention layers are utilized to capture global contexts for the coordinate regression task and enhance the ability to differentiate CRC and colorectum tissues. Moreover, a coordinate-driven self-learning (SL) strategy is introduced to leverage a large amount of unlabeled data to improve segmentation performance. We validate the proposed approach on a dataset including 227 labeled and 585 unlabeled CRC cases by fivefold cross-validation. Experimental results demonstrate that our method outperforms some recent related segmentation methods and achieves the segmentation accuracy in DSC for CRC of 0.669 and colorectum of 0.892, reaching to the performance (at 0.639 and 0.890, respectively) of a medical resident with two years of specialized CRC imaging fellowship.
Yingda Xia, Suyun Li, Jiawen Yao, Dakai Jin, Yanting Liang, Jiatai Lin, Bingchao Zhao, Chu Han, Le Lu 0001, Ling Zhang 0002, Zaiyi Liu, Xin Chen 0058
IEEE Trans. Neural Networks Learn. Syst.12
2024 Bootstrapping Chest CT Image Understanding by Distilling Knowledge from X-Ray Expert Models
abstract
Radiologists highly desire fully automated versatile AI for medical imaging interpretation. However, the lack of extensively annotated large-scale multi-disease datasets has hindered the achievement of this goal. In this paper, we explore the feasibility of leveraging language as a natu-rally high-quality supervision for chest CT imaging. In light of the limited availability of image-report pairs, we boot-strap the understanding of 3D chest CT images by distilling chest-related diagnostic knowledge from an extensively pre-trained 2D X-ray expert model. Specifically, we propose a language-guided retrieval method to match each 3D CT image with its semantically closest 2D X-ray image, and perform pair-wise and semantic relation knowledge distillation. Subsequently, we use contrastive learning to align images and reports within the same patient while distin-guishing them from the other patients. However, the challenge arises when patients have similar semantic diagnoses, such as healthy patients, potentially confusing if treated as negatives. We introduce a robust contrastive learning that identifies and corrects these false negatives. We train our model with over 12K pairs of chest CT images and radiology reports. Extensive experiments across multiple scenarios, including zero-shot learning, report generation, and fine-tuning processes, demonstrate the model's feasibility in interpreting chest CT images.
Yingda Xia, Tony C. W. Mok, Xianghua Ye, Le Lu 0001, Yuxing Tang, Ling Zhang 0002
CVPR10
2024 CycleINR: Cycle Implicit Neural Representation for Arbitrary-Scale Volumetric Super-Resolution of Medical Data
abstract
In the realm of medical 3D data, such as CT and MRI images, prevalent anisotropic resolution is characterized by high intra-slice but diminished inter-slice resolution. The lowered resolution between adjacent slices poses challenges, hindering optimal viewing experiences and impeding the development of robust downstream analysis algorithms. Various volumetric super-resolution algorithms aim to surmount these challenges, enhancing inter-slice resolution and overall 3D medical imaging quality. However, existing approaches confront inherent challenges: 1) often tailored to specific upsampling factors, lacking flexibility for diverse clinical scenarios; 2) newly generated slices frequently suffer from over-smoothing, degrading fine details, and leading to inter-slice inconsistency. In response, this study presents CycleINR, a novel enhanced Implicit Neural Representation model for 3D medical data volumetric super-resolution. Leveraging the continuity of the learned implicit function, the CycleINR model can achieve results with arbitrary up-sampling rates, eliminating the need for separate training. Additionally, we enhance the grid sampling in CycleINR with a local attention mechanism and mitigate over-smoothing by integrating cycleconsistent loss. We introduce a new metric, Slice-wise Noise Level Inconsistency (SNLI), to quantitatively assess inter-slice noise level inconsistency. The effectiveness of our approach is demonstrated through image quality evaluations on an in-house dataset and a downstream task analysis on the Medical Segmentation Decathlon liver tumor dataset.
Wei Fang 0005, Yuxing Tang, Heng Guo 0008, Mingze Yuan, Tony C. W. Mok, Ke Yan 0006, Jiawen Yao, Xin Chen 0058, Zaiyi Liu, Le Lu 0001, Ling Zhang 0002, Minfeng Xu
CVPR11
2024 Modality-Agnostic Structural Image Representation Learning for Deformable Multi-Modality Medical Image Registration
abstract
Establishing dense anatomical correspondence across distinct imaging modalities is a foundational yet challenging procedure for numerous medical image analysis studies and image-guided radiotherapy. Existing multimodality image registration algorithms rely on statistical-based similarity measures or local structural image representations. However, the former is sensitive to locally varying noise, while the latter is not discriminative enough to cope with complex anatomical structures in multimodal scans, causing ambiguity in determining the anatomical correspon-dence across scans with different modalities. In this paper, we propose a modality-agnostic structural representation learning method, which leverages Deep Neighbour-hood Self-similarity (DNS) and anatomy-aware contrastive learning to learn discriminative and contrast-invariance deep structural image representations (DSIR) without the need for anatomical delineations or pre-aligned training images. We evaluate our method on multiphase CT, abdomen MR-CT, and brain MR T1w-T2w registration. Comprehensive results demonstrate that our method is superior to the conventional local structural representation and statistical-based similarity measures in terms of discriminability and accuracy.
Tony C. W. Mok, Yunhao Bai, Wei Liu 0127, Yan-Jie Zhou, Ke Yan 0006, Dakai Jin, Xiaoli Yin, Le Lu 0001, Ling Zhang 0002
CVPR12
2024 LIDIA: Precise Liver Tumor Diagnosis on Multi-Phase Contrast-Enhanced CT via Iterative Fusion and Asymmetric Contrastive Learning
Wei Liu 0127, Xiaoming Zhang 0008, Xiaoli Yin, Xu Han 0023, Chunli Li, Yuan Gao 0017, Le Lu 0001, Ling Zhang 0002, Lei Zhang 0006, Ke Yan 0006
MICCAI (9)10
2024 Improved Esophageal Varices Assessment from Non-contrast CT Scans
Chunli Li, Xiaoming Zhang 0008, Yuan Gao 0017, Xiaoli Yin, Le Lu 0001, Ling Zhang 0002, Ke Yan 0006
MICCAI (5)6
2024 A Curvature-Guided Coarse-to-Fine Framework for Enhanced Whole Brain Segmentation
Fenqiang Zhao, Yuxing Tang, Le Lu 0001, Ling Zhang 0002
MICCAI (9)4
2023 Devil is in the Queries: Advancing Mask Transformers for Real-world Medical Image Segmentation and Out-of-Distribution Localization
abstract
Real-world medical image segmentation has tremendous long-tailed complexity of objects, among which tail conditions correlate with relatively rare diseases and are clinically significant. A trustworthy medical AI algorithm should demonstrate its effectiveness on tail conditions to avoid clinically dangerous damage in these out-of-distribution (OOD) cases. In this paper, we adopt the concept of object queries in Mask Transformers to formulate semantic segmentation as a soft cluster assignment. The queries fit the feature-level cluster centers of inliers during training. Therefore, when performing inference on a medical image in real-world scenarios, the similarity between pixels and the queries detects and localizes OOD regions. We term this OOD localization as MaxQuery. Furthermore, the foregrounds of real-world medical images, whether OOD objects or inliers, are lesions. The difference between them is less than that between the foreground and background, possibly misleading the object queries to focus redundantly on the background. Thus, we propose a query-distribution (QD) loss to enforce clear boundaries between segmentation targets and other regions at the query level, improving the inlier segmentation and OOD indication. Our proposed framework is tested on two real-world segmentation tasks, i.e., segmentation of pancreatic and liver tumors, outperforming previous state-of-the-art algorithms by an average of 7.39% on AUROC, 14.69% on AUPR, and 13.79% on FPR95 for OOD localization. On the other hand, our framework improves the performance of inlier segmentation by an average of 5.27% DSC when compared with the leading baseline nnUNet.
Mingze Yuan, Yingda Xia, Hexin Dong, Zifan Chen, Jiawen Yao, Mingyan Qiu, Ke Yan 0006, Xiaoli Yin, Xin Chen 0058, Zaiyi Liu, Bin Dong 0001, Jingren Zhou 0001, Le Lu 0001, Ling Zhang 0002, Li Zhang 0047
CVPR15
2023 CancerUniT: Towards a Single Unified Model for Effective Detection, Segmentation, and Diagnosis of Eight Major Cancers Using a Large Collection of CT Scans
abstract
Human readers or radiologists routinely perform full-body multi-organ multi-disease detection and diagnosis in clinical practice, while most medical AI systems are built to focus on single organs with a narrow list of a few diseases. This might severely limit AI’s clinical adoption. A certain number of AI models need to be assembled nontrivially to match the diagnostic process of a human reading a CT scan. In this paper, we construct a Unified Tumor Transformer (CancerUniT) model to jointly detect tumor existence & location and diagnose tumor characteristics for eight major cancers in CT scans. CancerUniT is a query-based Mask Transformer model with the output of multi-tumor prediction. We decouple the object queries into organ queries, tumor detection queries and tumor diagnosis queries, and further establish hierarchical relationships among the three groups. This clinically-inspired architecture effectively assists inter- and intra-organ representation learning of tumors and facilitates the resolution of these complex, anatomically related multi-organ cancer image reading tasks. CancerUniT is trained end-to-end using a curated large-scale CT images of 10,042 patients including eight major types of cancers and occurring non-cancer tumors (all are pathology-confirmed with 3D tumor masks annotated by radiologists). On the test set of 631 patients, CancerUniT has demonstrated strong performance under a set of clinically relevant evaluation metrics, substantially outperforming both multi-disease methods and an assembly of eight single-organ expert models in tumor detection, segmentation, and diagnosis. This moves one step closer towards a universal high performance cancer screening tool.
Jieneng Chen, Yingda Xia, Jiawen Yao, Ke Yan 0006, Le Lu 0001, Fakai Wang, Bo Zhou 0009, Mingyan Qiu, Qihang Yu, Mingze Yuan, Wei Fang 0005, Yuxing Tang, Minfeng Xu, Xianghua Ye, Xiaoli Yin, Xin Chen 0058, Jingren Zhou 0001, Alan L. Yuille, Zaiyi Liu, Ling Zhang 0002
ICCV25
2023 Improved Prognostic Prediction of Pancreatic Cancer Using Multi-phase CT by Integrating Neural Distance and Texture-Aware Transformer
Hexin Dong, Jiawen Yao, Yuxing Tang, Mingze Yuan, Yingda Xia, Jingren Zhou 0001, Bin Dong 0001, Le Lu 0001, Zaiyi Liu, Li Zhang 0047, Ling Zhang 0002
MICCAI (5)14
2023 Liver Tumor Screening and Diagnosis in CT with Pixel-Lesion-Patient Network
Ke Yan 0006, Xiaoli Yin, Yingda Xia, Fakai Wang, Yuan Gao 0017, Jiawen Yao, Chunli Li, Jingren Zhou 0001, Ling Zhang 0002, Le Lu 0001
MICCAI (5)11
2023 Cluster-Induced Mask Transformers for Effective Opportunistic Gastric Cancer Screening on Non-contrast CT Scans
Mingze Yuan, Yingda Xia, Xin Chen 0058, Jiawen Yao, Mingyan Qiu, Hexin Dong, Jingren Zhou 0001, Bin Dong 0001, Le Lu 0001, Li Zhang 0047, Zaiyi Liu, Ling Zhang 0002
MICCAI (5)13
2023 Parse and Recall: Towards Accurate Lung Nodule Malignancy Prediction Like Radiologists
Xianghua Ye, Yuxing Tang, Minfeng Xu, Jianfei Guo, Xin Chen 0058, Zaiyi Liu, Jingren Zhou 0001, Le Lu 0001, Ling Zhang 0002
MICCAI (5)11
2022 DeepCRC: Colorectum and Colorectal Cancer Segmentation in CT Scans via Deep Colorectal Coordinate Transform
Yingda Xia, Jiawen Yao, Dakai Jin, Bingjiang Qiu, Suyun Li, Yanting Liang, Xian-Sheng Hua 0001, Le Lu 0001, Xin Chen 0058, Zaiyi Liu, Ling Zhang 0002
MICCAI (3)14
2022 Effective Opportunistic Esophageal Cancer Screening Using Noncontrast CT Imaging
Jiawen Yao, Xianghua Ye, Yingda Xia, Ke Yan 0006, Lili Lin, Haogang Yu, Xian-Sheng Hua 0001, Le Lu 0001, Dakai Jin, Ling Zhang 0002
MICCAI (3)13
2021 3D Graph Anatomy Geometry-Integrated Network for Pancreatic Mass Segmentation, Diagnosis, and Quantitative Patient Management
abstract
The pancreatic disease taxonomy includes ten types of masses (tumors or cysts) [20], [8]. Previous work focuses on developing segmentation or classification methods only for certain mass types. Differential diagnosis of all mass types is clinically highly desirable [20] but has not been investigated using an automated image understanding approach.We exploit the feasibility to distinguish pancreatic ductal adenocarcinoma (PDAC) from the nine other nonPDAC masses using multi-phase CT imaging. Both image appearance and the 3D organ-mass geometry relationship are critical. We propose a holistic segmentation-mesh-classification network (SMCN) to provide patient-level diagnosis, by fully utilizing the geometry and location information, which is accomplished by combining the anatomical structure and the semantic detection-by-segmentation network. SMCN learns the pancreas and mass segmentation task and builds an anatomical correspondence-aware organ mesh model by progressively deforming a pancreas prototype on the raw segmentation mask (i.e., mask-to-mesh). A new graph-based residual convolutional network (Graph-ResNet), whose nodes fuse the information of the mesh model and feature vectors extracted from the segmentation network, is developed to produce the patient-level differential classification results. Extensive experiments on 661 patients’ CT scans (five phases per patient) show that SMCN can improve the mass segmentation and detection accuracy compared to the strong baseline method nnUNet (e.g., for nonPDAC, Dice: 0.611 vs. 0.478; detection rate: 89% vs. 70%), achieve similar sensitivity and specificity in differentiating PDAC and nonPDAC as expert radiologists (i.e., 94% and 90%), and obtain results comparable to a multimodality test [20] that combines clinical, imaging, and molecular testing for clinical management of patients.
Jiawen Yao, Isabella Nogues, Le Lu 0001, Lingyun Huang, Jing Xiao 0006, Zhaozheng Yin, Ling Zhang 0002
CVPR9
2021 Effective Pancreatic Cancer Screening on Non-contrast CT Scans via Anatomy-Aware Transformers
Yingda Xia, Jiawen Yao, Le Lu 0001, Lingyun Huang, Guo Tong Xie, Jing Xiao 0006, Alan L. Yuille, Ling Zhang 0002
MICCAI (5)9
2021 DeepPrognosis: Preoperative prediction of pancreatic cancer survival and surgical margin via comprehensive understanding of dynamic contrast-enhanced CT imaging and tumor-vascular contact parsing
Jiawen Yao, Le Lu 0001, Jianping Lu, Qike Song, Gang Jin, Jing Xiao 0006, Ling Zhang 0002
Medical Image Anal.10
2020 DeepPrognosis: Preoperative Prediction of Pancreatic Cancer Survival and Surgical Margin via Contrast-Enhanced CT Imaging
Jiawen Yao, Le Lu 0001, Jing Xiao 0006, Ling Zhang 0002
MICCAI (2)5
2020 Robust Pancreatic Ductal Adenocarcinoma Segmentation with Multi-institutional Multi-phase Partially-Annotated CT Scans
Ling Zhang 0002, Jiawen Yao, Yun Bian, Dakai Jin, Jing Xiao 0006, Le Lu 0001
MICCAI (4)1
2020 Spatio-Temporal Convolutional LSTMs for Tumor Growth Prediction by Learning 4D Longitudinal Patient Data
abstract
Prognostic tumor growth modeling via volumetric medical imaging observations can potentially lead to better outcomes of tumor treatment management and surgical planning. Recent advances of convolutional networks (ConvNets) have demonstrated higher accuracy than traditional mathematical models can be achieved in predicting future tumor volumes. This indicates that deep learning based data-driven techniques may have great potentials on addressing such problem. However, current 2D image patch based modeling approaches can not make full use of the spatio-temporal imaging context of the tumor's longitudinal 4D (3D + time) patient data. Moreover, they are incapable to predict clinically-relevant tumor properties, other than the tumor volumes. In this paper, we exploit to formulate the tumor growth process through convolutional Long Short-Term Memory (ConvLSTM) that extract tumor's static imaging appearances and simultaneously capture its temporal dynamic changes within a single network. We extend ConvLSTM into the spatio-temporal domain (ST-ConvLSTM) by jointly learning the inter-slice 3D contexts and the longitudinal or temporal dynamics from multiple patient studies. Our approach can incorporate other non-imaging patient information in an end-to-end trainable manner. Experiments are conducted on the largest 4D longitudinal tumor dataset of 33 patients to date. Results validate that the proposed ST-ConvLSTM model produces a Dice score of 83.2%±5.1% and a RVD of 11.2%±10.8%, both statistically significantly outperforming (p < 0.05) other compared methods of traditional linear model, ConvLSTM, and generative adversarial network (GAN) under the metric of predicting future tumor volumes. Additionally, our new method enables the prediction of both cell density and CT intensity numbers. Last, we demonstrate the generalizability of ST-ConvLSTM by employing it in 4D medical image segmentation task, which achieves an averaged Dice score of 86.3%±1.2% for left-ventricle segmentation in 4D ultrasound with 3 seconds per patient case.
Ling Zhang 0002, Le Lu 0001, Xiaosong Wang 0001, Robert Zhu, Mohammadhadi Bagheri, Ronald M. Summers, Jianhua Yao 0001
IEEE Trans. Medical Imaging1
2020 Generalizing Deep Learning for Medical Image Segmentation to Unseen Domains via Deep Stacked Transformation
abstract
Recent advances in deep learning for medical image segmentation demonstrate expert-level accuracy. However, application of these models in clinically realistic environments can result in poor generalization and decreased accuracy, mainly due to the domain shift across different hospitals, scanner vendors, imaging protocols, and patient populations etc. Common transfer learning and domain adaptation techniques are proposed to address this bottleneck. However, these solutions require data (and annotations) from the target domain to retrain the model, and is therefore restrictive in practice for widespread model deployment. Ideally, we wish to have a trained (locked) model that can work uniformly well across unseen domains without further training. In this paper, we propose a deep stacked transformation approach for domain generalization. Specifically, a series of n stacked transformations are applied to each image during network training. The underlying assumption is that the "expected" domain shift for a specific medical imaging modality could be simulated by applying extensive data augmentation on a single source domain, and consequently, a deep model trained on the augmented "big" data (BigAug) could generalize well on unseen domains. We exploit four surprisingly effective, but previously understudied, image-based characteristics for data augmentation to overcome the domain generalization problem. We train and evaluate the BigAug model (with n=9 transformations) on three different 3D segmentation tasks (prostate gland, left atrial, left ventricle) covering two medical imaging modalities (MRI and ultrasound) involving eight publicly available challenge datasets. The results show that when training on relatively small dataset (n = 10~32 volumes, depending on the size of the available datasets) from a single source domain: (i) BigAug models degrade an average of 11%(Dice score change) from source to unseen domain, substantially better than conventional augmentation (degrading 39%) and CycleGAN-based domain adaptation method (degrading 25%), (ii) BigAug is better than "shallower" stacked transforms (i.e. those with fewer transforms) on unseen domains and demonstrates modest improvement to conventional augmentation on the source domain, (iii) after training with BigAug on one source domain, performance on an unseen domain is similar to training a model from scratch on that domain when using the same number of training samples. When training on large datasets (n = 465 volumes) with BigAug, (iv) application to unseen domains reaches the performance of state-of-the-art fully supervised models that are trained and tested on their source domains. These findings establish a strong benchmark for the study of domain generalization in medical imaging, and can be generalized to the design of highly robust deep segmentation models for clinical deployment.
Ling Zhang 0002, Xiaosong Wang 0001, Dong Yang 0005, Thomas Sanford, Stephanie A. Harmon, Baris Turkbey, Bradford J. Wood, Holger Roth, Andriy Myronenko, Daguang Xu, Ziyue Xu 0001
IEEE Trans. Medical Imaging1
2019 Searching Learning Strategy with Reinforcement Learning for 3D Medical Image Segmentation
Dong Yang 0005, Holger Roth, Ziyue Xu 0001, Fausto Milletari, Ling Zhang 0002, Daguang Xu
MICCAI (2)5
2018 Deep Lesion Graphs in the Wild: Relationship Learning and Organization of Significant Radiology Image Findings in a Diverse Large-Scale Lesion Database
abstract
Radiologists in their daily work routinely find and annotate significant abnormalities on a large number of radiology images. Such abnormalities, or lesions, have collected over years and stored in hospitals' picture archiving and communication systems. However, they are basically unsorted and lack semantic annotations like type and location. In this paper, we aim to organize and explore them by learning a deep feature representation for each lesion. A large-scale and comprehensive dataset, DeepLesion, is introduced for this task. DeepLesion contains bounding boxes and size measurements of over 32K lesions. To model their similarity relationship, we leverage multiple supervision information including types, self-supervised location coordinates, and sizes. They require little manual annotation effort but describe useful attributes of the lesions. Then, a triplet network is utilized to learn lesion embeddings with a sequential sampling strategy to depict their hierarchical similarity structure. Experiments show promising qualitative and quantitative results on lesion retrieval, clustering, and classification. The learned embeddings can be further employed to build a lesion graph for various clinically useful applications. An algorithm for intra-patient lesion matching is proposed and validated with experiments.
Ke Yan 0006, Xiaosong Wang 0001, Le Lu 0001, Ling Zhang 0002, Adam P. Harrison, Mohammadhadi Bagheri, Ronald M. Summers
CVPR4
2018 Convolutional Invasion and Expansion Networks for Tumor Growth Prediction
abstract
Tumor growth is associated with cell invasion and mass-effect, which are traditionally formulated by mathematical models, namely reaction-diffusion equations and biomechanics. Such models can be personalized based on clinical measurements to build the predictive models for tumor growth. In this paper, we investigate the possibility of using deep convolutional neural networks to directly represent and learn the cell invasion and mass-effect, and to predict the subsequent involvement regions of a tumor. The invasion network learns the cell invasion from information related to metabolic rate, cell density, and tumor boundary derived from multimodal imaging data. The expansion network models the mass-effect from the growing motion of tumor mass. We also study different architectures that fuse the invasion and expansion networks, in order to exploit the inherent correlations among them. Our network can easily be trained on population data and personalized to a target patient, unlike most previous mathematical modeling methods that fail to incorporate population data. Quantitative experiments on a pancreatic tumor data set show that the proposed method substantially outperforms a state-of-the-art mathematical model-based approach in both accuracy and efficiency, and that the information captured by each of the two subnetworks is complementary.
Ling Zhang 0002, Le Lu 0001, Ronald M. Summers, Electron Kebebew, Jianhua Yao 0001
IEEE Trans. Medical Imaging1
2018 Predicting Locations of High-Risk Plaques in Coronary Arteries in Patients Receiving Statin Therapy
abstract
Features of high-risk coronary artery plaques prone to major adverse cardiac events (MACE) were identified by intravascular ultrasound (IVUS) virtual histology (VH). These plaque features are: thin-cap fibroatheroma (TCFA), plaque burden PB ≥ 70%, or minimal luminal area MLA ≤ 4 mm2. Identification of arterial locations likely to later develop such high-risk plaques may help prevent MACE. We report a machine learning method for prediction of future high-risk coronary plaque locations and types in patients under statin therapy. Sixty-one patients with stable angina on statin therapy underwent baseline and one-year follow-up VH-IVUS non-culprit vessel examinations followed by quantitative image analysis. For each segmented and registered VH-IVUS frame pair (${n} =6341$ ), location-specific ($\approx 0.5$ mm) vascular features and demographic information at baseline were identified. Seven independent support vector machine classifiers with seven different feature subsets were trained to predict high-risk plaque types one year later. A leave-one-patient-out cross-validation was used to evaluate the prediction power of different feature subsets. The experimental results showed that our machine learning method predicted future TCFA with correctness of 85.9%, 81.7%, and 77.0% (G-mean) for baseline plaque phenotypes of TCFA, thick-cap fibroatheroma, and non-fibroatheroma, respectively. For predicting PB ≥ 70%, correctness was 80.8% for baseline PB ≥ 70% and 85.6% for 50% ≤ PB2was 81.6% for baseline MLA ≤ 4 mm2and 80.2% for 4 mm22. Location-specific prediction of future high-risk coronary artery plaques is feasible through machine learning using focal vascular features and demographic variables. Our approach outperforms previously reported results and shows the importance of local factors on high-risk coronary artery plaque development.
Ling Zhang 0002, Andreas Wahle, Zhi Chen 0025, John J. Lopez, Tomas Kovarnik, Milan Sonka
IEEE Trans. Medical Imaging1
2017 Personalized Pancreatic Tumor Growth Prediction via Group Learning
Ling Zhang 0002, Le Lu 0001, Ronald M. Summers, Electron Kebebew, Jianhua Yao 0001
MICCAI (2)1
2017 DeepPap: Deep Convolutional Networks for Cervical Cell Classification
abstract
Automation-assisted cervical screening via Pap smear or liquid-based cytology (LBC) is a highly effective cell imaging based cancer detection tool, where cells are partitioned into "abnormal" and "normal" categories. However, the success of most traditional classification methods relies on the presence of accurate cell segmentations. Despite sixty years of research in this field, accurate segmentation remains a challenge in the presence of cell clusters and pathologies. Moreover, previous classification methods are only built upon the extraction of hand-crafted features, such as morphology and texture. This paper addresses these limitations by proposing a method to directly classify cervical cells-without prior segmentation-based on deep features, using convolutional neural networks (ConvNets). First, the ConvNet is pretrained on a natural image dataset. It is subsequently fine-tuned on a cervical cell dataset consisting of adaptively resampled image patches coarsely centered on the nuclei. In the testing phase, aggregation is used to average the prediction scores of a similar set of image patches. The proposed method is evaluated on both Pap smear and LBC datasets. Results show that our method outperforms previous algorithms in classification accuracy (98.3%), area under the curve (0.99) values, and especially specificity (98.3%), when applied to the Herlev benchmark Pap smear dataset and evaluated using five-fold cross validation. Similar superior performances are also achieved on the HEMLBC (H&E stained manual LBC) dataset. Our method is promising for the development of automation-assisted reading systems in primary cervical screening.
Ling Zhang 0002, Le Lu 0001, Isabella Nogues, Ronald M. Summers, Shaoxiong Liu, Jianhua Yao 0001
IEEE J. Biomed. Health Informatics1
2015 Prospective Prediction of Thin-Cap Fibroatheromas from Baseline Virtual Histology Intravascular Ultrasound Data
Ling Zhang 0002, Andreas Wahle, Zhi Chen 0025, John J. Lopez, Tomas Kovarnik, Milan Sonka
MICCAI (2)1
2015 Simultaneous Registration of Location and Orientation in Intravascular Ultrasound Pullbacks Pairs Via 3D Graph-Based Optimization
abstract
A novel method is reported for simultaneous registration of location (axial direction) and orientation (circumferential direction) of two intravascular ultrasound (IVUS) pullbacks of the same vessel taken at different times. Monitoring plaque progression or regression (e.g., during lipid treatment) is of high clinical relevance. Our method uses a 3D graph optimization approach, in which the cost function jointly reflects similarity of plaque morphology and plaque/perivascular image appearance. Graph arcs incorporate prior information about temporal correspondence of the two IVUS sequences and limited angular twisting between consecutive IVUS images. Additionally, our approach automatically identifies starting and ending frame pairs in the two IVUS pullbacks. Validation of our method was performed in 29 pairs of IVUS baseline/follow-up pullback sequences consisting of 8 622 IVUS image frames in total. In comparison to manual registration by three experts, the average location and orientation registration errors ranged from 0.72 mm to 0.79 mm and from 7.3(°) to 9.3(°), respectively, all close to the inter-observer variability with no difference being statistically significant (p = NS). Rotation angles determined by our automated approach and expert observers showed high correlation (r(2) of 0.97 to 0.98) and agreed closely (mutual bias between the automated method and expert observers ranged from -1.57(°) to 0.15(°)). Compared with state-of-the-art approaches, the new method offers lower errors in both location and orientation registration. Our method offers highly automated and accurate IVUS pullback registration and can be employed in IVUS-based studies of coronary disease progression, enabling more focal studies of coronary plaque development and transition of vulnerability.
Ling Zhang 0002, Andreas Wahle, Zhi Chen 0025, Li Zhang 0047, Richard W. Downe, Tomas Kovarnik, Milan Sonka
IEEE Trans. Medical Imaging1