EDBT 2026 Demo / reviewers in the wild / expert
Holger Roth
dblp:42/8528 · also Holger R. Roth
· DBLP profile ↗
62ranked-venue papers
10as first author
35since 2021 · last 2026
0000-0002-3662-8743ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 43 · 10 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 6 first-author · 21 since 2021Artificial intelligence and machine learning · 12 · 11 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | End-to-End Spatiotemporal Analysis of Color Doppler Echocardiograms: Application for Rheumatic Heart Disease DetectionabstractRheumatic heart disease (RHD) represents a significant global health challenge, disproportionately affecting over 40 million people in low- and middle-income countries. Early detection through color Doppler echocardiography is crucial for treating RHD, but it requires specialized physicians who are often scarce in resource-limited settings. To address this disparity, artificial intelligence (AI)-driven tools for RHD screening can provide scalable, autonomous solutions to improve access to critical healthcare services in underserved regions. This paper introduces RADAR (Rapid AI-Assisted Echocardiography Detection and Analysis of RHD), a novel and generalizable AI approach for end-to-end spatiotemporal analysis of color Doppler echocardiograms, aimed at detecting early RHD in resource-limited settings. RADAR identifies key imaging views and employs convolutional neural networks to analyze diagnostically relevant phases of the cardiac cycle. It also localizes essential anatomical regions and examines blood flow patterns. It then integrates all findings into a cohesive analytical framework. RADAR was trained and validated on 1,022 echocardiogram videos from 511 Ugandan children, acquired using standard portable ultrasound devices. An independent set of 318 cases, acquired using a handheld ultrasound device with diverse imaging characteristics, was also tested. On the validation set, RADAR outperformed existing methods, achieving an average accuracy of 0.92, sensitivity of 0.94, and specificity of 0.90. In independent testing, it maintained high, clinically acceptable performance, with an average accuracy of 0.79, sensitivity of 0.87, and specificity of 0.70. These results highlight RADAR's potential to improve RHD detection and promote health equity for vulnerable children by enhancing timely, accurate diagnoses in underserved regions. Pooneh Roshanitabrizi, Vishwesh Nath, Kelsey Brown, Taylor Gloria Broudy, Zhifan Jiang, Abhijeet Parida, Joselyn Rwebembera, Emmy Okello, Andrea Z. Beaton, Holger Roth, Craig A. Sable, Marius George Linguraru |
IEEE Trans. Medical Imaging | 10 |
| 2025 | VILA-M3: Enhancing Vision-Language Models with Medical Expert KnowledgeabstractGeneralist vision language models (VLMs) have made significant strides in computer vision, but they fall short in specialized fields like healthcare, where expert knowledge is essential. Current large multimodal models like Gemini and GPT-4o are insufficient for medical tasks due to their reliance on memorized internet knowledge rather than the nuanced expertise required in healthcare. Meanwhile, existing medical VLMs (e.g. Med-Gemini) often lack expert consultation as part of their design, and many rely on outdated, static datasets that were not created with modern, large deep learning models in mind. VLMs are usually trained in three stages: vision pre-training, vision-language pre-training, and instruction fine-tuning (IFT). IFT has been typically applied using a mixture of generic and healthcare data. In contrast, we propose that for medical VLMs, a fourth stage of specialized IFT is necessary, which focuses on medical data and includes information from domain expert models. Domain expert models developed for medical use are crucial because they are specifically trained for certain clinical tasks, e.g. to detect tumors and classify abnormalities through segmentation and classification, which learn fine-grained features of medical data−features that are often too intricate for a VLM to capture effectively. This paper introduces a new framework, VILA-M3, for medical VLMs that utilizes domain knowledge via expert models. We argue that generic VLM architectures alone are not viable for real-world clinical applications and on-demand usage of domain-specialized expert model knowledge is critical for advancing AI in healthcare. Through our experiments, we show an improved state-of-the-art (SOTA) performance with an average improvement of ~9% over the prior SOTA model Med-Gemini and ~6% over models trained on the specific tasks. Our approach emphasizes the importance of domain expertise in creating precise, reliable VLMs for medical applications. Vishwesh Nath, Wenqi Li 0001, Dong Yang 0005, Andriy Myronenko, Mingxin Zheng, Yao Lu 0006, Hongxu Yin, Yee Man Law, Yucheng Tang, Can Zhao 0001, Ziyue Xu 0001, Yufan He, Stephanie A. Harmon, Benjamin Simon, Greg Heinrich, Stephen R. Aylward, Marc Edgar, Michael Zephyr, Pavlo Molchanov 0001, Baris Turkbey, Holger Roth, Daguang Xu |
CVPR | 23 |
| 2025 | Synthesis of Pathological Dual-Channel Color Doppler Echocardiograms for Equitable Diagnosis of Heart Diseases
Pooneh Roshanitabrizi, Artur Arturi Aharonyan, Kelsey Brown, Taylor Gloria Broudy, Abhijeet Parida, Austin Tapp, Zhifan Jiang, Alison Tompsett, Joselyn Rwebembera, Emmy Okello, Andrea Z. Beaton, Holger Roth, Daguang Xu, Syed Muhammad Anwar, Craig A. Sable, Marius George Linguraru |
MICCAI (2) | 13 |
| 2024 | FedBPT: Efficient Federated Black-box Prompt Tuning for Large Language ModelsabstractPre-trained language models (PLM) have revolutionized the NLP landscape, achieving stellar performances across diverse tasks. These models, while benefiting from vast training data, often require fine-tuning on specific data to cater to distinct downstream tasks. However, this data adaptation process has inherent security and privacy concerns, primarily when leveraging user-generated, device-residing data. Federated learning (FL) provides a solution, allowing collaborative model fine-tuning without centralized data collection. However, applying FL to finetune PLMs is hampered by challenges, including restricted model parameter access due to the high encapsulation, high computational requirements, and communication overheads. This paper introduces Federated Black-box Prompt Tuning (FedBPT), a framework designed to address these challenges. FedBPT allows the clients to treat the model as a black-box inference API. By focusing on training optimal prompts and utilizing gradient-free optimization methods, FedBPT reduces the number of exchanged variables, boosts communication efficiency, and minimizes computational and storage costs. Experiments highlight the framework’s ability to drastically cut communication and memory costs while maintaining competitive performance. Ultimately, FedBPT presents a promising solution for efficient, privacy-preserving fine-tuning of PLM in the age of large language models. Jingwei Sun 0002, Ziyue Xu 0001, Hongxu Yin, Dong Yang 0005, Daguang Xu, Zhixu Du, Yiran Chen 0001, Holger Roth |
ICML | 9 |
| 2024 | Super-Field MRI Synthesis for Infant Brains Enhanced by Dual Channel Latent Diffusion
Austin Tapp, Can Zhao 0001, Holger Roth, Jeffrey Tanedo, Syed Muhammad Anwar, Niall J. Bourke, Joseph V. Hajnal, Victoria Nankabirwa, Sean C. L. Deoni, Natasha Leporé, Marius George Linguraru |
MICCAI (3) | 3 |
| 2024 | Federated Black-box Prompt Tuning System for Large Language Models on the EdgeabstractFederated learning (FL) offers a privacy-preserving way to train models across decentralized data. However, fine-tuning pre-trained language models (PLMs) in FL is challenging due to restricted model parameter access, high computational demands, and communication overheads. Our method treats large language models (LLMs) as black-box inference APIs, optimizing prompts with gradient-free methods. This approach, FedBPT, reduces exchanged variables, boosts communication efficiency, and minimizes computational and memory costs. We demonstrate the practical implementation of FedBPT on resource-limited edge devices, showcasing its ability to efficiently achieve collaborative on-device LLM fine-tuning. Jingwei Sun 0002, Ang Li 0005, Beidi Chen, Holger Roth, Daguang Xu, Tingjun Chen, Yiran Chen 0001 |
MobiCom | 7 |
| 2024 | Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?abstractHow can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain. Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 47 |
| 2024 | IR-FRestormer: Iterative Refinement with Fourier-Based Restormer for Accelerated MRI ReconstructionabstractAccelerated magnetic resonance imaging (MRI) aims to reconstruct high-quality MR images from a set of under-sampled measurements. State-of-the-art methods for this task use deep learning, which offers high reconstruction accuracy and fast runtimes. In this work, we propose a new state-of-the-art reconstruction model for accelerated MRI reconstruction. Our model is the first to combine the power of deep neural networks with iterative refinement for this task. For the neural network component of our method, we utilize a transformer-based architecture as transformers are state-of-the-art in various image reconstruction tasks. However, a major drawback of transformers which has limited their emergence among the state-of-the-art MRI models is that they are often memory inefficient for high-resolution inputs. To address this limitation, we propose a transformer-based model which uses parameter-free Fourier-based attention modules, achieving 2× more memory efficiency. We evaluate our model on the largest publicly available MRI dataset, the fastMRI dataset [46], and achieve on-par performance with other state-of-the-art1methods on the dataset’s leaderboard2. Mohammad Zalbagi Darestani, Vishwesh Nath, Wenqi Li 0001, Yufan He, Holger Roth, Ziyue Xu 0001, Daguang Xu, Reinhard Heckel, Can Zhao 0001 |
WACV | 5 |
| 2024 | Learning Quality Labels for Robust Image ClassificationabstractSupervised learning paradigms largely benefit from the tremendous amount of annotated data. However, the quality of the annotations often varies among labelers. Multi-observer studies have been conducted to examine the annotation variances (by labeling the same data multiple times) to see how it affects critical applications like medical image analysis. In this paper, we demonstrate how multiple sets of annotations (either hand-labeled or algorithm-generated) can be utilized together and mutually benefit the learning of classification tasks. A scheme of learning-to-vote is introduced to sample quality label sets for each data entry on-the-fly during the training. Specifically, a label-sampling module is designed to achieve refined labels (weighted sum of attended ones) that benefit the model learning the most through additional back-propagations. We apply the learning-to-vote scheme on the classification task of a synthetic noisy CIFAR-10 to prove the concept and then demonstrate superior results (3-5% increase on average in multiple disease classification AUCs) on the chest x-ray images from a hospital-scale dataset (MIMIC-CXR) and hand-labeled dataset (OpenI) in comparison to regular training paradigms. Xiaosong Wang 0001, Ziyue Xu 0001, Dong Yang 0005, Leo K. Tam, Holger Roth, Daguang Xu |
WACV | 5 |
| 2024 | MONAI Label: A framework for AI-assisted interactive labeling of 3D medical images
Andres Diaz-Pinto, Sachidanand Alle, Vishwesh Nath, Yucheng Tang, Alvin Ihsani, Muhammad Asad 0001, Fernando Pérez-García, Pritesh Mehta, Wenqi Li 0001, Mona Flores, Holger Roth, Tom Vercauteren, Daguang Xu, Prerna Dogra, Sébastien Ourselin, Andrew Feng, Manuel Jorge Cardoso |
Medical Image Anal. | 11 |
| 2024 | Fair evaluation of federated learning algorithms for automated breast density classification: The results of the 2022 ACR-NCI-NVIDIA federated learning challenge
Kendall Schmidt, Ben Bearce, Ken Chang, Laura Coombs, Keyvan Farahani, Marawan Elbatel, Kaouther Mouheb, Robert Martí, Ya Zhang 0002, Yanfeng Wang 0001, Yaojun Hu, Haochao Ying, Yuyang Xu, Conrad Testagrose, Mutlu Demirer, Vikash Gupta, Ünal Akünal, Markus Bujotzek, Klaus H. Maier-Hein, Yi Qin 0006, Xiaomeng Li 0001, Jayashree Kalpathy-Cramer, Holger Roth |
Medical Image Anal. | 24 |
| 2023 | Fair Federated Medical Image Segmentation via Client Contribution EstimationabstractHow to ensure fairness is an important topic in federated learning (FL). Recent studies have investigated how to reward clients based on their contribution (collaboration fairness), and how to achieve uniformity of performance across clients (performance fairness). Despite achieving progress on either one, we argue that it is critical to consider them together, in order to engage and motivate more diverse clients joining FL to derive a high-quality global model. In this work, we propose a novel method to optimize both types of fairness simultaneously. Specifically, we propose to estimate client contribution in gradient and data space. In gradient space, we monitor the gradient direction differences of each client with respect to others. And in data space, we measure the prediction error on client data using an auxiliary model. Based on this contribution estimation, we propose a FL method, federated training via contribution estimation (FedCE), i.e., using estimation as global model aggregation weights. We have theoretically analyzed our method and empirically evaluated it on two real-world medical datasets. The effectiveness of our approach has been validated with significant performance improvements, better collaboration fairness, better performance fairness, and comprehensive analytical studies. Code is available at https://nvidia.github.io/NVFlare/research/fed-ce Meirui Jiang, Holger Roth, Wenqi Li 0001, Dong Yang 0005, Can Zhao 0001, Vishwesh Nath, Daguang Xu, Qi Dou 0001, Ziyue Xu 0001 |
CVPR | 2 |
| 2023 | Communication-Efficient Vertical Federated Learning with Limited Overlapping SamplesabstractFederated learning is a popular collaborative learning approach that enables clients to train a global model without sharing their local data. Vertical federated learning (VFL) deals with scenarios in which the data on clients have different feature spaces but share some overlapping samples. Existing VFL approaches suffer from high communication costs and cannot deal efficiently with limited overlapping samples commonly seen in the real world. We propose a practical VFL framework called one-shot VFL that can solve the communication bottleneck and the problem of limited overlapping samples simultaneously based on semi-supervised learning. We also propose few-shot VFL to improve the accuracy further with just one more communication round between the server and the clients. In our proposed framework, the clients only need to communicate with the server once or only a few times. We evaluate the proposed VFL framework on both image and tabular datasets. Our methods can improve the accuracy by more than 46.5% and reduce the communication cost by more than 330× compared with state-of-the-art VFL methods when evaluated on CIFAR-10. Our code is available at https://nvidia.github.io/NVFlare/research/one-shot-vfl. Jingwei Sun 0002, Ziyue Xu 0001, Dong Yang 0005, Vishwesh Nath, Wenqi Li 0001, Can Zhao 0001, Daguang Xu, Yiran Chen 0001, Holger Roth |
ICCV | 9 |
| 2023 | DAST: Differentiable Architecture Search with Transformer for 3D Medical Image Segmentation
Dong Yang 0005, Ziyue Xu 0001, Yufan He, Vishwesh Nath, Wenqi Li 0001, Andriy Myronenko, Ali Hatamizadeh, Can Zhao 0001, Holger Roth, Daguang Xu |
MICCAI (3) | 9 |
| 2023 | Do Gradient Inversion Attacks Make Federated Learning Unsafe?abstractFederated learning (FL) allows the collaborative training of AI models without needing to share raw data. This capability makes it especially interesting for healthcare applications where patient and data privacy is of utmost concern. However, recent works on the inversion of deep neural networks from model gradients raised concerns about the security of FL in preventing the leakage of training data. In this work, we show that these attacks presented in the literature are impractical in FL use-cases where the clients' training involves updating the Batch Normalization (BN) statistics and provide a new baseline attack that works for such scenarios. Furthermore, we present new ways to measure and visualize potential data leakage in FL. Our work is a step towards establishing reproducible methods of measuring data leakage in FL and could help determine the optimal tradeoffs between privacy-preserving techniques, such as differential privacy, and model accuracy based on quantifiable metrics. Ali Hatamizadeh, Hongxu Yin, Pavlo Molchanov 0001, Andriy Myronenko, Wenqi Li 0001, Prerna Dogra, Andrew Feng, Mona Flores, Jan Kautz, Daguang Xu, Holger Roth |
IEEE Trans. Medical Imaging | 11 |
| 2023 | Guest Editorial Special Issue on Federated Learning for Medical Imaging: Enabling Collaborative Development of Robust AI ModelsabstractFederated Learning (FL) could solve the challenges of training AI models on large datasets for medical imaging due to data privacy and ownership concerns by allowing collaborative training without the need for sharing raw data. This Special Issue on Federated Learning for Medical Imaging features papers covering FL-related topics and discussing their implications for healthcare and medical imaging. The included articles focus on a broad range of federated scenarios and applications, such as semi-supervised and self-supervised learning, histopathology, image reconstruction, graph neural networks, privacy preservation, active learning, data auditing, multi-task learning, personalization, and swarm learning. The importance of training unbiased, privacy-preserving, and generalizable AI models that have the potential to be translated into clinical practice increases the need for collaborative training techniques such as FL. The articles included in this Special Issue have moved the needle markedly forward in this regard. Holger Roth, Nicola Rieke, Shadi Albarqouni, Quanzheng Li |
IEEE Trans. Medical Imaging | 1 |
| 2022 | GradViT: Gradient Inversion of Vision TransformersabstractIn this work we demonstrate the vulnerability of vision transformers (ViTs) to gradient-based inversion attacks. During this attack, the original data batch is reconstructed given model weights and the corresponding gradients. We introduce a method, named GradViT, that optimizes random noise into naturally looking images via an iterative process. The optimization objective consists of (i) a loss on matching the gradients, (ii) image prior in the form of distance to batch-normalization statistics of a pretrained CNN model, and (iii) a total variation regularization on patches to guide correct recovery locations. We propose a unique loss scheduling function to overcome local minima during optimization. We evaluate GadViT on ImageNet1K and MS-Celeb-1M datasets, and observe unprecedentedly high fidelity and closeness to the original (hidden) data. During the analysis we find that vision transformers are significantly more vulnerable than previously studied CNNs due to the presence of the attention mechanism. Our method demonstrates new state-of-the-art results for gradient inversion in both qualitative and quantitative metrics. Project page at https://gradvit.github.io/. Ali Hatamizadeh, Hongxu Yin, Holger Roth, Wenqi Li 0001, Jan Kautz, Daguang Xu, Pavlo Molchanov 0001 |
CVPR | 3 |
| 2022 | Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image AnalysisabstractVision Transformers (ViT)s have shown great performance in self-supervised learning of global and local representations that can be transferred to downstream applications. Inspired by these results, we introduce a novel self-supervised learning framework with tailored proxy tasks for medical image analysis. Specifically, we propose: (i) a new 3D transformer-based model, dubbed Swin UNEt TRansformers (Swin UNETR), with a hierarchical encoder for self-supervised pretraining; (ii) tailored proxy tasks for learning the underlying pattern of human anatomy. We demonstrate successful pre-training of the proposed model on 5,050 publicly available computed tomography (CT) images from various body organs. The effectiveness of our approach is validated by fine-tuning the pre-trained models on the Beyond the Cranial Vault (BTCV) Segmentation Challenge with 13 abdominal organs and segmentation tasks from the Medical Segmentation Decathlon (MSD) dataset. Our model is currently the state-of-the-art on the public test leaderboards of both MSD11https://decathlon-10.grand-challenge.org/evaluation/challenge/leaderboard/ and BTCV22https://www.synapse.org/#!Synapse:syn3193805/wiki/217785/ datasets. Code: https://monai.io/research/swin-unetr. Yucheng Tang, Dong Yang 0005, Wenqi Li 0001, Holger Roth, Bennett A. Landman, Daguang Xu, Vishwesh Nath, Ali Hatamizadeh |
CVPR | 4 |
| 2022 | Closing the Generalization Gap of Cross-silo Federated Medical Image SegmentationabstractCross-silo federated learning (FL) has attracted much attention in medical imaging analysis with deep learning in recent years as it can resolve the critical issues of insufficient data, data privacy, and training efficiency. However, there can be a generalization gap between the model trained from FL and the one from centralized training. This important issue comes from the non-iid data distribution of the local data in the participating clients and is well-known as client drift. In this work, we propose a novel training frame-work FedSM to avoid the client drift issue and successfully close the generalization gap compared with the centralized training for medical image segmentation tasks for the first time. We also propose a novel personalized FL objective formulation and a new method SoftPull to solve it in our proposed framework FedSM. We conduct rigorous theoretical analysis to guarantee its convergence for optimizing the non-convex smooth objective function. Real-world medical image segmentation experiments using deep FL validate the motivations and effectiveness of our proposed method. An Xu, Wenqi Li 0001, Dong Yang 0005, Holger Roth, Ali Hatamizadeh, Can Zhao 0001, Daguang Xu, Heng Huang 0001, Ziyue Xu 0001 |
CVPR | 5 |
| 2022 | Auto-FedRL: Federated Hyperparameter Optimization for Multi-institutional Medical Image Segmentation
Dong Yang 0005, Ali Hatamizadeh, An Xu, Ziyue Xu 0001, Wenqi Li 0001, Can Zhao 0001, Daguang Xu, Stephanie A. Harmon, Evrim Turkbey, Baris Turkbey, Bradford J. Wood, Francesca Patella, Elvira Stellato, Gianpaolo Carrafiello, Vishal M. Patel, Holger Roth |
ECCV (21) | 17 |
| 2022 | Warm Start Active Learning with Proxy Labels and Selection via Semi-supervised Fine-Tuning
Vishwesh Nath, Dong Yang 0005, Holger Roth, Daguang Xu |
MICCAI (8) | 3 |
| 2022 | Ensembled Prediction of Rheumatic Heart Disease from Ungated Doppler Echocardiography Acquired in Low-Resource Settings
Pooneh R. Tabrizi, Holger Roth, Alison Tompsett, Athelia Rosa Paulli, Kelsey Brown, Joselyn Rwebembera, Emmy Okello, Andrea Z. Beaton, Craig A. Sable, Marius George Linguraru |
MICCAI (1) | 2 |
| 2022 | Clinical-Realistic Annotation for Histopathology Images with Probabilistic Semi-supervision: A Worst-Case Study
Ziyue Xu 0001, Andriy Myronenko, Dong Yang 0005, Holger Roth, Can Zhao 0001, Xiaosong Wang 0001, Daguang Xu |
MICCAI (2) | 4 |
| 2022 | UNETR: Transformers for 3D Medical Image SegmentationabstractFully Convolutional Neural Networks (FCNNs) with contracting and expanding paths have shown prominence for the majority of medical image segmentation applications since the past decade. In FCNNs, the encoder plays an integral role by learning both global and local features and contextual representations which can be utilized for semantic output prediction by the decoder. Despite their success, the locality of convolutional layers in FCNNs, limits the capability of learning long-range spatial dependencies. Inspired by the recent success of transformers for Natural Language Processing (NLP) in long-range sequence learning, we reformulate the task of volumetric (3D) medical image segmentation as a sequence-to-sequence prediction problem. We introduce a novel architecture, dubbed as UNEt TRansformers (UNETR), that utilizes a transformer as the encoder to learn sequence representations of the input volume and effectively capture the global multi-scale information, while also following the successful "U-shaped" network design for the encoder and decoder. The transformer encoder is directly connected to a decoder via skip connections at different resolutions to compute the final semantic segmentation output. We have validated the performance of our method on the Multi Atlas Labeling Beyond The Cranial Vault (BTCV) dataset for multi-organ segmentation and the Medical Segmentation Decathlon (MSD) dataset for brain tumor and spleen segmentation tasks. Our benchmarks demonstrate new state-of-the-art performance on the BTCV leaderboard. Ali Hatamizadeh, Yucheng Tang, Vishwesh Nath, Dong Yang 0005, Andriy Myronenko, Bennett A. Landman, Holger Roth, Daguang Xu |
WACV | 7 |
| 2022 | Rapid artificial intelligence solutions in a pandemic - The COVID-19-20 Lung CT Lesion Segmentation Challenge
Holger Roth, Ziyue Xu 0001, Carlos Tor-Díez, Ramon Sánchez-Jacob, Jonathan Zember, Jose Molto, Wenqi Li 0001, Sheng Xu 0001, Baris Turkbey, Evrim Turkbey, Dong Yang 0005, Ahmed Harouni, Nicola Rieke, Shishuai Hu, Fabian Isensee, Claire Tang, Qinji Yu, Jan Sölter, Vitali Liauchuk, Jan Hendrik Moltz, Bruno Oliveira 0002, Yong Xia 0001, Klaus H. Maier-Hein, Qikai Li, Andreas Husch, Vassili Kovalev, Alessa Hering, João L. Vilaça, Mona Flores, Daguang Xu, Bradford J. Wood, Marius George Linguraru |
Medical Image Anal. | 1 |
| 2022 | Cardiac segmentation on late gadolinium enhancement MRI: A benchmark study from multi-sequence cardiac MR segmentation challenge
Xiahai Zhuang, Jiahang Xu, Xinzhe Luo, Chen Chen 0042, Cheng Ouyang, Daniel Rueckert, Víctor M. Campello, Karim Lekadir, Sulaiman Vesal, Nishant Ravikumar, Yashu Liu 0003, Gongning Luo, Jingkun Chen, Hongwei Li 0004, Buntheng Ly, Maxime Sermesant, Holger Roth, Wentao Zhu 0001, Jiexiang Wang, Xinghao Ding, Sen Yang 0006, Lei Li 0020 |
Medical Image Anal. | 17 |
| 2021 | DiNTS: Differentiable Neural Network Topology Search for 3D Medical Image SegmentationabstractRecently, neural architecture search (NAS) has been applied to automatically search high-performance networks for medical image segmentation. The NAS search space usually contains a network topology level (controlling connections among cells with different spatial scales) and a cell level (operations within each cell). Existing methods either require long searching time for large-scale 3D image datasets, or are limited to pre-defined topologies (such as U-shaped or single-path) . In this work, we focus on three important aspects of NAS in 3D medical image segmentation: flexible multi-path network topology, high search efficiency, and budgeted GPU memory usage. A novel differentiable search framework is proposed to support fast gradient-based search within a highly flexible network topology search space. The discretization of the searched optimal continuous model in differentiable scheme may produce a sub-optimal final discrete model (discretization gap). Therefore, we propose a topology loss to alleviate this problem. In addition, the GPU memory usage for the searched 3D model is limited with budget constraints during search. Our Differentiable Network Topology Search scheme (DiNTS) is evaluated on the Medical Segmentation Decathlon (MSD) challenge, which contains ten challenging segmentation tasks. Our method achieves the state-of-the-art performance and the top ranking on the MSD challenge leaderboard. Yufan He, Dong Yang 0005, Holger Roth, Can Zhao 0001, Daguang Xu |
CVPR | 3 |
| 2021 | T-AutoML: Automated Machine Learning for Lesion Segmentation using Transformers in 3D Medical ImagingabstractLesion segmentation in medical imaging has been an important topic in clinical research. Researchers have proposed various detection and segmentation algorithms to address this task. Recently, deep learning-based approaches have significantly improved the performance over conventional methods. However, most state-of-the-art deep learning methods require the manual design of multiple network components and training strategies. In this paper, we propose a new automated machine learning algorithm, T-AutoML, which not only searches for the best neural architecture, but also finds the best combination of hyper-parameters and data augmentation strategies simultaneously. The proposed method utilizes the modern transformer model, which is introduced to adapt to the dynamic length of the search space embedding and can significantly improve the ability of the search. We validate T-AutoML on several large-scale public lesion segmentation data-sets and achieve state-of-the-art performance. Dong Yang 0005, Andriy Myronenko, Xiaosong Wang 0001, Ziyue Xu 0001, Holger Roth, Daguang Xu |
ICCV | 5 |
| 2021 | Accounting for Dependencies in Deep Learning Based Multiple Instance Learning for Whole Slide Imaging
Andriy Myronenko, Ziyue Xu 0001, Dong Yang 0005, Holger Roth, Daguang Xu |
MICCAI (8) | 4 |
| 2021 | The Power of Proxy Data and Proxy Networks for Hyper-parameter Optimization in Medical Image Segmentation
Vishwesh Nath, Dong Yang 0005, Ali Hatamizadeh, Anas A. Abidin, Andriy Myronenko, Holger Roth, Daguang Xu |
MICCAI (3) | 6 |
| 2021 | Federated Whole Prostate Segmentation in MRI with Personalized Neural Architectures
Holger Roth, Dong Yang 0005, Wenqi Li 0001, Andriy Myronenko, Wentao Zhu 0001, Ziyue Xu 0001, Xiaosong Wang 0001, Daguang Xu |
MICCAI (3) | 1 |
| 2021 | Federated learning improves site performance in multicenter deep learning without data sharingabstractOBJECTIVE: To demonstrate enabling multi-institutional training without centralizing or sharing the underlying physical data via federated learning (FL). MATERIALS AND METHODS: Deep learning models were trained at each participating institution using local clinical data, and an additional model was trained using FL across all of the institutions. RESULTS: We found that the FL model exhibited superior performance and generalizability to the models trained at single institutions, with an overall performance level that was significantly better than that of any of the institutional models alone when evaluated on held-out test sets from each institution and an outside challenge dataset. DISCUSSION: The power of FL was successfully demonstrated across 3 academic institutions while avoiding the privacy risk associated with the transfer and pooling of patient data. CONCLUSION: Federated learning is an effective methodology that merits further study to enable accelerated development of models across institutions, enabling greater generalizability in clinical use. Karthik Sarma, Stephanie A. Harmon, Thomas Sanford, Holger Roth, Ziyue Xu 0001, Jesse Tetreault, Daguang Xu, Mona Flores, Alex G. Raman, Rushikesh Kulkarni, Bradford J. Wood, Peter L. Choyke, Alan Priester, Leonard S. Marks, Steven S. Raman, Dieter R. Enzmann, Baris Turkbey, William Speier, Corey W. Arnold |
J. Am. Medical Informatics Assoc. | 4 |
| 2021 | Federated semi-supervised learning for COVID region segmentation in chest CT using multi-national data from China, Italy, Japan
Dong Yang 0005, Ziyue Xu 0001, Wenqi Li 0001, Andriy Myronenko, Holger Roth, Stephanie A. Harmon, Sheng Xu 0001, Baris Turkbey, Evrim Turkbey, Xiaosong Wang 0001, Wentao Zhu 0001, Gianpaolo Carrafiello, Francesca Patella, Maurizio Cariati, Hirofumi Obinata, Hitoshi Mori, Kaku Tamura, Peng An 0002, Bradford J. Wood, Daguang Xu |
Medical Image Anal. | 5 |
| 2021 | Diminishing Uncertainty Within the Training Pool: Active Learning for Medical Image SegmentationabstractActive learning is a unique abstraction of machine learning techniques where the model/algorithm could guide users for annotation of a set of data points that would be beneficial to the model, unlike passive machine learning. The primary advantage being that active learning frameworks select data points that can accelerate the learning process of a model and can reduce the amount of data needed to achieve full accuracy as compared to a model trained on a randomly acquired data set. Multiple frameworks for active learning combined with deep learning have been proposed, and the majority of them are dedicated to classification tasks. Herein, we explore active learning for the task of segmentation of medical imaging data sets. We investigate our proposed framework using two datasets: 1.) MRI scans of the hippocampus, 2.) CT scans of pancreas and tumors. This work presents a query-by-committee approach for active learning where a joint optimizer is used for the committee. At the same time, we propose three new strategies for active learning: 1.) increasing frequency of uncertain data to bias the training data set; 2.) Using mutual information among the input images as a regularizer for acquisition to ensure diversity in the training dataset; 3.) adaptation of Dice log-likelihood for Stein variational gradient descent (SVGD). The results indicate an improvement in terms of data reduction by achieving full accuracy while only using 22.69% and 48.85% of the available data for each dataset, respectively. Vishwesh Nath, Dong Yang 0005, Bennett A. Landman, Daguang Xu, Holger Roth |
IEEE Trans. Medical Imaging | 5 |
| 2021 | Guest Editorial Annotation-Efficient Deep Learning: The Holy Grail of Medical ImagingabstractAnnotation-efficient deep learning refers to methods and practices that yield high-performance deep learning models without the use of massive carefully labeled training datasets. This paradigm has recently attracted attention from the medical imaging research community because (1) it is difficult to collect large, representative medical imaging datasets given the diversity of imaging protocols, imaging devices, and patient populations, (2) it is expensive to acquire accurate annotations from medical experts even for moderately sized medical imaging datasets, and (3) it is infeasible to adapt data-hungry deep learning models to detect and diagnose rare diseases whose low prevalence hinders data collection. Nima Tajbakhsh, Holger Roth, Demetri Terzopoulos, Jianming Liang |
IEEE Trans. Medical Imaging | 2 |
| 2020 | C2FNAS: Coarse-to-Fine Neural Architecture Search for 3D Medical Image Segmentationabstract3D convolution neural networks (CNN) have been proved very successful in parsing organs or tumours in 3D medical images, but it remains sophisticated and time-consuming to choose or design proper 3D networks given different task contexts. Recently, Neural Architecture Search (NAS) is proposed to solve this problem by searching for the best network architecture automatically. However, the inconsistency between search stage and deployment stage often exists in NAS algorithms due to memory constraints and large search space, which could become more serious when applying NAS to some memory and time-consuming tasks, such as 3D medical image segmentation. In this paper, we propose a coarse-to-fine neural architecture search (C2FNAS) to automatically search a 3D segmentation network from scratch without inconsistency on network size or input size. Specifically, we divide the search procedure into two stages: 1) the coarse stage, where we search the macro-level topology of the network, i.e. how each convolution module is connected to other modules; 2) the fine stage, where we search at micro-level for operations in each cell based on previous searched macro-level topology. The coarse-to-fine manner divides the search procedure into two consecutive stages and meanwhile resolves the inconsistency. We evaluate our method on 10 public datasets from Medical Segmentation Decalthon (MSD) challenge, and achieve state-of-the-art performance with the network searched using one dataset, which demonstrates the effectiveness and generalization of our searched models. Qihang Yu, Dong Yang 0005, Holger Roth, Yutong Bai, Yixiao Zhang 0001, Alan L. Yuille, Daguang Xu |
CVPR | 3 |
| 2020 | LAMP: Large Deep Nets with Automated Model Parallelism for Image Segmentation
Wentao Zhu 0001, Can Zhao 0001, Wenqi Li 0001, Holger Roth, Ziyue Xu 0001, Daguang Xu |
MICCAI (4) | 4 |
| 2020 | 3D Semi-Supervised Learning with Uncertainty-Aware Multi-View Co-TrainingabstractWhile making a tremendous impact in various fields, deep neural networks usually require large amounts of labeled data for training which are expensive to collect in many applications, especially in the medical domain. Un-labeled data, on the other hand, is much more abundant. Semi-supervised learning techniques, such as co-training, could provide a powerful tool to leverage unlabeled data. In this paper, we propose a novel framework, uncertainty-aware multi-view co-training (UMCT), to address semi-supervised learning on 3D data, such as volumetric data from medical imaging. In our work, co-training is achieved by exploiting multi-viewpoint consistency of 3D data. We generate different views by rotating or permuting the 3D data and utilize asymmetrical 3D kernels to encourage diversified features in different sub-networks. In addition, we propose an uncertainty-weighted label fusion mechanism to estimate the reliability of each view's prediction with Bayesian deep learning. As one view requires the supervision from other views in co-training, our self-adaptive approach computes a confidence score for the prediction of each unlabeled sample in order to assign a reliable pseudo label. Thus, our approach can take advantage of unlabeled data during training. We show the effectiveness of our proposed semi-supervised method on several public datasets from medical image segmentation tasks (NIH pancreas & LiTS liver tumor dataset). Meanwhile, a fully-supervised method based on our approach achieved state-of-the-art performances on both the LiTS liver tumor segmentation and the Medical Segmentation Decathlon (MSD) challenge, demonstrating the robustness and value of our framework, even when fully supervised training is feasible. Yingda Xia, Fengze Liu, Dong Yang 0005, Jinzheng Cai, Lequan Yu, Zhuotun Zhu, Daguang Xu, Alan L. Yuille, Holger Roth |
WACV | 9 |
| 2020 | NeurReg: Neural Registration and Its Application to Image SegmentationabstractRegistration is a fundamental task in medical image analysis which can be applied to several tasks including image segmentation, intra-operative tracking, multi-modal image alignment, and motion analysis. Popular registration tools such as ANTs and NiftyReg optimize an objective function for each pair of images from scratch which is time-consuming for large images with complicated deformation. Facilitated by the rapid progress of deep learning, learning-based approaches such as VoxelMorph have been emerging for image registration. These approaches can achieve competitive performance in a fraction of a second on advanced GPUs. In this work, we construct a neural registration framework, called NeurReg, with a hybrid loss of displacement fields and data similarity, which substantially improves the current state-of-the-art of registrations. Within the framework, we simulate various transformations by a registration simulator which generates fixed image and displacement field ground truth for training. Furthermore, we design three segmentation frameworks based on the proposed registration framework: 1) atlas-based segmentation, 2) joint learning of both segmentation and registration tasks, and 3) multi-task learning with atlas-based segmentation as an intermediate feature. Extensive experimental results validate the effectiveness of the proposed NeurReg framework based on various metrics: the endpoint error (EPE) of the predicted displacement field, mean square error (MSE), normalized local cross-correlation (NLCC), mutual information (MI), Dice coefficient, uncertainty estimation, and the interpretability of the segmentation. The proposed NeurReg improves registration accuracy with fast inference speed, which can greatly accelerate related medical image analysis tasks. Wentao Zhu 0001, Andriy Myronenko, Ziyue Xu 0001, Wenqi Li 0001, Holger Roth, Yufang Huang, Fausto Milletari, Daguang Xu |
WACV | 5 |
| 2020 | Uncertainty-aware multi-view co-training for semi-supervised medical image segmentation and domain adaptation
Yingda Xia, Dong Yang 0005, Zhiding Yu, Fengze Liu, Jinzheng Cai, Lequan Yu, Zhuotun Zhu, Daguang Xu, Alan L. Yuille, Holger Roth |
Medical Image Anal. | 10 |
| 2020 | Generalizing Deep Learning for Medical Image Segmentation to Unseen Domains via Deep Stacked TransformationabstractRecent advances in deep learning for medical image segmentation demonstrate expert-level accuracy. However, application of these models in clinically realistic environments can result in poor generalization and decreased accuracy, mainly due to the domain shift across different hospitals, scanner vendors, imaging protocols, and patient populations etc. Common transfer learning and domain adaptation techniques are proposed to address this bottleneck. However, these solutions require data (and annotations) from the target domain to retrain the model, and is therefore restrictive in practice for widespread model deployment. Ideally, we wish to have a trained (locked) model that can work uniformly well across unseen domains without further training. In this paper, we propose a deep stacked transformation approach for domain generalization. Specifically, a series of n stacked transformations are applied to each image during network training. The underlying assumption is that the "expected" domain shift for a specific medical imaging modality could be simulated by applying extensive data augmentation on a single source domain, and consequently, a deep model trained on the augmented "big" data (BigAug) could generalize well on unseen domains. We exploit four surprisingly effective, but previously understudied, image-based characteristics for data augmentation to overcome the domain generalization problem. We train and evaluate the BigAug model (with n=9 transformations) on three different 3D segmentation tasks (prostate gland, left atrial, left ventricle) covering two medical imaging modalities (MRI and ultrasound) involving eight publicly available challenge datasets. The results show that when training on relatively small dataset (n = 10~32 volumes, depending on the size of the available datasets) from a single source domain: (i) BigAug models degrade an average of 11%(Dice score change) from source to unseen domain, substantially better than conventional augmentation (degrading 39%) and CycleGAN-based domain adaptation method (degrading 25%), (ii) BigAug is better than "shallower" stacked transforms (i.e. those with fewer transforms) on unseen domains and demonstrates modest improvement to conventional augmentation on the source domain, (iii) after training with BigAug on one source domain, performance on an unseen domain is similar to training a model from scratch on that domain when using the same number of training samples. When training on large datasets (n = 465 volumes) with BigAug, (iv) application to unseen domains reaches the performance of state-of-the-art fully supervised models that are trained and tested on their source domains. These findings establish a strong benchmark for the study of domain generalization in medical imaging, and can be generalized to the design of highly robust deep segmentation models for clinical deployment. Ling Zhang 0002, Xiaosong Wang 0001, Dong Yang 0005, Thomas Sanford, Stephanie A. Harmon, Baris Turkbey, Bradford J. Wood, Holger Roth, Andriy Myronenko, Daguang Xu, Ziyue Xu 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2019 | Unsupervised Segmentation of Micro-CT Images of Lung Cancer Specimen Using Deep Generative Models
Takayasu Moriya, Hirohisa Oda, Midori Mitarai, Shota Nakamura, Holger Roth, Masahiro Oda 0001, Kensaku Mori |
MICCAI (6) | 5 |
| 2019 | Searching Learning Strategy with Reinforcement Learning for 3D Medical Image Segmentation
Dong Yang 0005, Holger Roth, Ziyue Xu 0001, Fausto Milletari, Ling Zhang 0002, Daguang Xu |
MICCAI (2) | 2 |
| 2018 | Towards Automated Colonoscopy Diagnosis: Binary Polyp Size Estimation via Unsupervised Depth Learning
Hayato Itoh, Holger Roth, Le Lu 0001, Masahiro Oda 0001, Masashi Misawa, Yuichi Mori, Shin-ei Kudo, Kensaku Mori |
MICCAI (2) | 2 |
| 2018 | BESNet: Boundary-Enhanced Segmentation of Cells in Histopathological Images
Hirohisa Oda, Holger Roth, Kosuke Chiba, Jure Sokolic, Takayuki Kitasaka, Masahiro Oda 0001, Akinari Hinoki, Hiroo Uchida, Julia A. Schnabel, Kensaku Mori |
MICCAI (2) | 2 |
| 2018 | Colon Shape Estimation Method for Colonoscope Tracking Using Recurrent Neural Networks
Masahiro Oda 0001, Holger Roth, Takayuki Kitasaka, Kazuhiro Furukawa, Ryoji Miyahara, Yoshiki Hirooka, Hidemi Goto, Nassir Navab, Kensaku Mori |
MICCAI (4) | 2 |
| 2018 | A Multi-scale Pyramid of 3D Fully Convolutional Networks for Abdominal Multi-organ Segmentation
Holger Roth, Chen Shen 0002, Hirohisa Oda, Takaaki Sugino, Masahiro Oda 0001, Yuichiro Hayashi, Kazunari Misawa, Kensaku Mori |
MICCAI (4) | 1 |
| 2018 | Spatial aggregation of holistically-nested convolutional neural networks for automated pancreas localization and segmentation
Holger Roth, Le Lu 0001, Nathan Lay, Adam P. Harrison, Amal Farag, Andrew Sohn, Ronald M. Summers |
Medical Image Anal. | 1 |
| 2017 | Tracking and Segmentation of the Airways in Chest CT Using a Fully Convolutional Network
Meng Qier, Holger Roth, Takayuki Kitasaka, Masahiro Oda 0001, Junji Ueno, Kensaku Mori |
MICCAI (2) | 2 |
| 2017 | TBS: Tensor-Based Supervoxels for Unfolding the Heart
Hirohisa Oda, Holger Roth, Kanwal K. Bhatia, Masahiro Oda 0001, Takayuki Kitasaka, Toshiaki Akita, Julia A. Schnabel, Kensaku Mori |
MICCAI (1) | 2 |
| 2017 | A Bottom-Up Approach for Pancreas Segmentation Using Cascaded Superpixels and (Deep) Image Patch LabelingabstractRobust organ segmentation is a prerequisite for computer-aided diagnosis, quantitative imaging analysis, pathology detection, and surgical assistance. For organs with high anatomical variability (e.g., the pancreas), previous segmentation approaches report low accuracies, compared with well-studied organs, such as the liver or heart. We present an automated bottom-up approach for pancreas segmentation in abdominal computed tomography (CT) scans. The method generates a hierarchical cascade of information propagation by classifying image patches at different resolutions and cascading (segments) superpixels. The system contains four steps: 1) decomposition of CT slice images into a set of disjoint boundary-preserving superpixels; 2) computation of pancreas class probability maps via dense patch labeling; 3) superpixel classification by pooling both intensity and probability features to form empirical statistics in cascaded random forest frameworks; and 4) simple connectivity based post-processing. Dense image patch labeling is conducted using two methods: efficient random forest classification on image histogram, location and texture features; and more expensive (but more accurate) deep convolutional neural network classification, on larger image windows (i.e., with more spatial contexts). Over-segmented 2-D CT slices by the simple linear iterative clustering approach are adopted through model/parameter calibration and labeled at the superpixel level for positive (pancreas) or negative (non-pancreas or background) classes. The proposed method is evaluated on a data set of 80 manually segmented CT volumes, using six-fold cross-validation. Its performance equals or surpasses other state-of-the-art methods (evaluated by "leave-one-patient-out"), with a dice coefficient of 70.7% and Jaccard index of 57.9%. In addition, the computational efficiency has improved significantly, requiring a mere 6 ~ 8 min per testing case, versus ≥ 10 h for other methods. The segmentation framework using deep patch labeling confidences is also more numerically stable, as reflected in the smaller performance metric standard deviations. Finally, we implement a multi-atlas label fusion (MALF) approach for pancreas segmentation using the same data set. Under six-fold cross-validation, our bottom-up segmentation method significantly outperforms its MALF counterpart: 70.7±13.0% versus 52.51±20.84% in dice coefficients. Amal Farag, Le Lu 0001, Holger Roth, Evrim Turkbey, Ronald M. Summers |
IEEE Trans. Image Process. | 3 |
| 2016 | Automatic Lymph Node Cluster Segmentation Using Holistically-Nested Neural Networks and Structured Optimization in CT Images
Isabella Nogues, Le Lu 0001, Xiaosong Wang 0001, Holger Roth, Gedas Bertasius, Nathan Lay, Jianbo Shi, Yohannes Tsehay, Ronald M. Summers |
MICCAI (2) | 4 |
| 2016 | Spatial Aggregation of Holistically-Nested Networks for Automated Pancreas Segmentation
Holger Roth, Le Lu 0001, Amal Farag, Andrew Sohn, Ronald M. Summers |
MICCAI (2) | 1 |
| 2016 | Improving Computer-Aided Detection Using Convolutional Neural Networks and Random View AggregationabstractAutomated computer-aided detection (CADe) has been an important tool in clinical practice and research. State-of-the-art methods often show high sensitivities at the cost of high false-positives (FP) per patient rates. We design a two-tiered coarse-to-fine cascade framework that first operates a candidate generation system at sensitivities ∼ 100% of but at high FP levels. By leveraging existing CADe systems, coordinates of regions or volumes of interest (ROI or VOI) are generated and function as input for a second tier, which is our focus in this study. In this second stage, we generate 2D (two-dimensional) or 2.5D views via sampling through scale transformations, random translations and rotations. These random views are used to train deep convolutional neural network (ConvNet) classifiers. In testing, the ConvNets assign class (e.g., lesion, pathology) probabilities for a new set of random views that are then averaged to compute a final per-candidate classification probability. This second tier behaves as a highly selective process to reject difficult false positives while preserving high sensitivities. The methods are evaluated on three data sets: 59 patients for sclerotic metastasis detection, 176 patients for lymph node detection, and 1,186 patients for colonic polyp detection. Experimental results show the ability of ConvNets to generalize well to different medical imaging CADe applications and scale elegantly to various data sets. Our proposed methods improve performance markedly in all cases. Sensitivities improved from 57% to 70%, 43% to 77%, and 58% to 75% at 3 FPs per patient for sclerotic metastases, lymph nodes and colonic polyps, respectively. Holger Roth, Le Lu 0001, Jianhua Yao 0001, Ari Seff, Kevin M. Cherry, Lauren Kim, Ronald M. Summers |
IEEE Trans. Medical Imaging | 1 |
| 2016 | Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer LearningabstractRemarkable progress has been made in image recognition, primarily due to the availability of large-scale annotated datasets and deep convolutional neural networks (CNNs). CNNs enable learning data-driven, highly representative, hierarchical image features from sufficient training data. However, obtaining datasets as comprehensively annotated as ImageNet in the medical imaging domain remains a challenge. There are currently three major techniques that successfully employ CNNs to medical image classification: training the CNN from scratch, using off-the-shelf pre-trained CNN features, and conducting unsupervised CNN pre-training with supervised fine-tuning. Another effective method is transfer learning, i.e., fine-tuning CNN models pre-trained from natural image dataset to medical image tasks. In this paper, we exploit three important, but previously understudied factors of employing deep convolutional neural networks to computer-aided detection problems. We first explore and evaluate different CNN architectures. The studied models contain 5 thousand to 160 million parameters, and vary in numbers of layers. We then evaluate the influence of dataset scale and spatial image context on performance. Finally, we examine when and why transfer learning from pre-trained ImageNet (via fine-tuning) can be useful. We study two specific computer-aided detection (CADe) problems, namely thoraco-abdominal lymph node (LN) detection and interstitial lung disease (ILD) classification. We achieve the state-of-the-art performance on the mediastinal LN detection, and report the first five-fold cross-validation classification results on predicting axial CT slices with ILD categories. Our extensive empirical evaluation, CNN model analysis and valuable insights can be extended to the design of high performance CAD systems for other medical imaging tasks. Hoo-Chang Shin, Holger Roth, Mingchen Gao, Le Lu 0001, Ziyue Xu 0001, Isabella Nogues, Jianhua Yao 0001, Daniel J. Mollura, Ronald M. Summers |
IEEE Trans. Medical Imaging | 2 |
| 2015 | DeepOrgan: Multi-level Deep Convolutional Networks for Automated Pancreas Segmentation
Holger Roth, Le Lu 0001, Amal Farag, Hoo-Chang Shin, Evrim Turkbey, Ronald M. Summers |
MICCAI (1) | 1 |
| 2015 | Leveraging Mid-Level Semantic Boundary Cues for Automated Lymph Node Detection
Ari Seff, Le Lu 0001, Adrian Barbu, Holger Roth, Hoo-Chang Shin, Ronald M. Summers |
MICCAI (2) | 4 |
| 2014 | A New 2.5D Representation for Lymph Node Detection Using Random Sets of Deep Convolutional Neural Network Observations
Holger Roth, Le Lu 0001, Ari Seff, Kevin M. Cherry, Joanne Hoffman, Evrim Turkbey, Ronald M. Summers |
MICCAI (1) | 1 |
| 2014 | 2D View Aggregation for Lymph Node Detection Using a Shallow Hierarchy of Linear Classifiers
Ari Seff, Le Lu 0001, Kevin M. Cherry, Holger Roth, Joanne Hoffman, Evrim Turkbey, Ronald M. Summers |
MICCAI (1) | 4 |
| 2013 | Endoluminal surface registration for CT colonography using haustral fold matchingabstractComputed Tomographic (CT) colonography is a technique used for the detection of bowel cancer or potentially precancerous polyps. The procedure is performed routinely with the patient both prone and supine to differentiate fixed colonic pathology from mobile faecal residue. Matching corresponding locations is difficult and time consuming for radiologists due to colonic deformations that occur during patient repositioning. We propose a novel method to establish correspondence between the two acquisitions automatically. The problem is first simplified by detecting haustral folds using a graph cut method applied to a curvature-based metric applied to a surface mesh generated from segmentation of the colonic lumen. A virtual camera is used to create a set of images that provide a metric for matching pairs of folds between the prone and supine acquisitions. Image patches are generated at the fold positions using depth map renderings of the endoluminal surface and optimised by performing a virtual camera registration over a restricted set of degrees of freedom. The intensity difference between image pairs, along with additional neighbourhood information to enforce geometric constraints over a 2D parameterisation of the 3D space, are used as unary and pair-wise costs respectively, and included in a Markov Random Field (MRF) model to estimate the maximum a posteriori fold labelling assignment. The method achieved fold matching accuracy of 96.0% and 96.1% in patient cases with and without local colonic collapse. Moreover, it improved upon an existing surface-based registration algorithm by providing an initialisation. The set of landmark correspondences is used to non-rigidly transform a 2D source image derived from a conformal mapping process on the 3D endoluminal surface mesh. This achieves full surface correspondence between prone and supine views and can be further refined with an intensity based registration showing a statistically significant improvement (p<0.001), and decreasing mean error from 11.9 mm to 6.0 mm measured at 1743 reference points from 17 CTC datasets. Thomas Hampshire, Holger Roth, Emma Helbren, Andrew Plumb, Darren Boone, Gregory Slabaugh, Steve Halligan, David J. Hawkes |
Medical Image Anal. | 2 |
| 2011 | Automatic Prone to Supine Haustral Fold Matching in CT Colonography Using a Markov Random Field Model
Thomas Hampshire, Holger Roth, Mingxing Hu, Darren Boone, Gregory Slabaugh, Shonit Punwani, Steve Halligan, David J. Hawkes |
MICCAI (1) | 2 |
| 2010 | Establishing Spatial Correspondence between the Inner Colon Surfaces from Prone and Supine CT Colonography
Holger Roth, Jamie McClelland, Marc Modat, Darren Boone, Mingxing Hu, Sébastien Ourselin, Gregory Slabaugh, Steve Halligan, David J. Hawkes |
MICCAI (3) | 1 |