Daguang Xu

dblp:140/3604 · DBLP profile ↗
← Back
67ranked-venue papers
0as first author
49since 2021 · last 2026
0000-0002-4621-881XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 28 since 2021Applied, interdisciplinary, general and emerging computing · 36 · 25 since 2021Artificial intelligence and machine learning · 23 · 19 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MAISI-v2: Accelerated 3D High-Resolution Medical Image Synthesis with Rectified Flow and Region-specific Contrastive Loss
abstract
Medical image synthesis is an important topic for both clinical and research applications. Recently, diffusion models have become a leading approach in this area. Despite their strengths, many existing methods struggle with (1) limited generalizability, only working for specific body regions or voxel spacings, (2) slow inference, which is a common issue for diffusion models, and (3) weak alignment with input conditions, which is a critical issue for medical imaging. MAISI, a previously proposed framework, addresses generalizability issues but still suffers from slow inference and limited condition consistency. In this work, we present MAISI-v2, the first accelerated 3D medical image synthesis framework that integrates rectified flow to enable fast and high-quality generation. To further enhance condition fidelity, we introduce a novel region-specific contrastive loss to improve sensitivity to the region of interest. Our experiments show that MAISI-v2 can achieve state-of-the-art image quality with 33× acceleration for latent diffusion models. We also conducted a downstream segmentation experiment to show that the synthetic images can be used for data augmentation. We release our code, training details, model weights, and a GUI demo to facilitate reproducibility and promote further development within the community.
Can Zhao 0001, Dong Yang 0005, Yufan He, Yucheng Tang, Benjamin Simon, Mason Belue, Stephanie A. Harmon, Baris Turkbey, Daguang Xu
AAAI10
2026 A comprehensive survey of computer vision methods for spatial transcriptomics
abstract
Spatial transcriptomics (ST) enables the simultaneous measurement of gene expression and spatial localization within tissue sections, providing unprecedented opportunities to dissect tissue architecture and functional organization. As a relatively new omics technology, bioinformatics has driven much of the innovation in ST. However, within these frameworks, spatial information is often reduced to locations and relationships between molecular profiles, without fully leveraging the wealth of sub-micron morphological detail and histological knowledge available. Advances in computer vision-based artificial intelligence (AI) are opening exciting new avenues beyond conventional bioinformatics approaches by modeling complex histological patterns and linking morphology to molecular states. More excitingly, they bring fresh perspectives to potentially address key limitations of ST, including its high cost, limited clinical applicability, and reliance on 2D analysis of inherently 3D tissues. For instance, models that predict ST directly from histology images enable virtual sequencing, drastically reducing costs while integrating morphological insights from pathology with molecular biomarkers, thus accelerating clinical translation. Moreover, computer vision techniques can reconstruct pixel-aligned 3D tissue models, overcoming the technical barriers of 2D acquisition and advancing 3D spatial omics analytics. In this paper, we present the first systematic survey of computer vision AI models for ST analytics, categorizing approaches across architectures, learning paradigms, tasks, and datasets, and tracing their technological evolution. We highlight key challenges and future directions, offering a panoramic perspective on vision-driven ST and its potential to transform both basic research and clinical practice. The curated collection of vision-driven ST papers is available at https://github.com/hrlblab/computer_vision_spatial_omics.
Junchao Zhu, Ruining Deng, Junlin Guo, Tianyuan Yao, Siqi Lu, Chongyu Qu, Juming Xiong, Yanfan Zhu, Zhengyi Lu, Yuechen Yang, Marilyn Lionts, Yucheng Tang, Daguang Xu, Shilin Zhao, Haichun Yang, Yuankai Huo
Briefings Bioinform.13
2026 Text-Driven Tumor Synthesis
abstract
Tumor synthesis can generate challenging cases that AI often misses or over-detects. Training on these cases improves AI performance. However, most existing synthesis methods are either unconditional- generating images from random variables-or conditioned only on tumor shape. As a result, they lack control over clinically important tumor characteristics, such as texture, heterogeneity, boundary, and pathology. The generated tumors are therefore overly similar or duplicates of existing training cases, failing to effectively address AI's weaknesses. We propose a new text-driven tumor synthesis approach, termed TextoMorph, that provides textual control over tumor characteristics in conjunction with mask control. This approach is particularly beneficial for examples that confuse the AI the most, such as early tumor detection (improving Sensitivity by + 6.5%), tumor segmentation for precise radiotherapy (improving NSD by + 3.1%), and classification between benign and malignant tumors (improving Sensitivity by + 8.2%). By incorporating text mined from radiology reports into the synthesis process, we increase the variability and controllability of the synthetic tumors to target AI's failure cases more precisely. Moreover, TextoMorph uses contrastive learning across different texts and CT scans, significantly reducing dependence on scarce image-report pairs (only 141 pairs used in this study) by leveraging a large corpus of 34,035 radiology reports. Finally, we have developed rigorous tests to evaluate synthetic tumors, showing that our synthetic tumors is realistic and diverse in texture, heterogeneity, boundary, and pathology. Code and models are available at https://github.com/MrGiovanni/TextoMorph.
Yi Shuai, Qi Chen 0014, Dong Yang 0005, Can Zhao 0001, Pedro R. A. S. Bassi, Daguang Xu, Kang Wang 0016, Yang Yang 0009, Alan L. Yuille, Zongwei Zhou
IEEE Trans. Medical Imaging10
2025 VISTA3D: A Unified Segmentation Foundation Model For 3D Medical Imaging
abstract
Foundation models for interactive segmentation in 2D natural images and videos have sparked significant interest in building 3D foundation models for medical imaging. However, the domain gaps and clinical use cases for 3D medical imaging require a dedicated model that diverges from existing 2D solutions. Specifically, such foundation models should support a full workflow that can actually reduce human effort. Treating 3D medical images as sequences of 2D slices and reusing interactive 2D foundation models seems straightforward, but 2D annotation is too time-consuming for 3D tasks. Moreover, for large cohort analysis, it’s the highly accurate automatic segmentation models that reduce the most human effort. However, these models lack support for interactive corrections and lack zero-shot ability for novel structures, which is a key feature of "foundation". While reusing pre-trained 2D backbones in 3D enhances zero-shot potential, their performance on complex 3D structures still lags behind leading 3D models. To address these issues, we present VISTA3D, Versatile Imaging SegmenTation and Annotation model, that targets to solve all these challenges and requirements with one unified foundation model. VISTA3D is built on top of the well-established 3D segmentation pipeline, and it is the first model to achieve state-of-the-art performance in both 3D automatic (supporting 127 classes) and 3D interactive segmentation, even when compared with top 3D expert models on large and diverse benchmarks. Additionally, VISTA3D’s 3D interactive design allows efficient human correction, and a novel 3D supervoxel method that distills 2D pre-trained backbones grants VISTA3D top 3D zero-shot performance. We believe the model, recipe, and insights represent a promising step towards a clinically useful 3D foundation model. Code and weights are publicly available at https://github.com/Project-MONAI/VISTA.
Yufan He, Yucheng Tang, Andriy Myronenko, Vishwesh Nath, Ziyue Xu 0001, Dong Yang 0005, Can Zhao 0001, Benjamin Simon, Mason Belue, Stephanie A. Harmon, Baris Turkbey, Daguang Xu, Wenqi Li 0001
CVPR13
2025 VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge
abstract
Generalist vision language models (VLMs) have made significant strides in computer vision, but they fall short in specialized fields like healthcare, where expert knowledge is essential. Current large multimodal models like Gemini and GPT-4o are insufficient for medical tasks due to their reliance on memorized internet knowledge rather than the nuanced expertise required in healthcare. Meanwhile, existing medical VLMs (e.g. Med-Gemini) often lack expert consultation as part of their design, and many rely on outdated, static datasets that were not created with modern, large deep learning models in mind. VLMs are usually trained in three stages: vision pre-training, vision-language pre-training, and instruction fine-tuning (IFT). IFT has been typically applied using a mixture of generic and healthcare data. In contrast, we propose that for medical VLMs, a fourth stage of specialized IFT is necessary, which focuses on medical data and includes information from domain expert models. Domain expert models developed for medical use are crucial because they are specifically trained for certain clinical tasks, e.g. to detect tumors and classify abnormalities through segmentation and classification, which learn fine-grained features of medical data−features that are often too intricate for a VLM to capture effectively. This paper introduces a new framework, VILA-M3, for medical VLMs that utilizes domain knowledge via expert models. We argue that generic VLM architectures alone are not viable for real-world clinical applications and on-demand usage of domain-specialized expert model knowledge is critical for advancing AI in healthcare. Through our experiments, we show an improved state-of-the-art (SOTA) performance with an average improvement of ~9% over the prior SOTA model Med-Gemini and ~6% over models trained on the specific tasks. Our approach emphasizes the importance of domain expertise in creating precise, reliable VLMs for medical applications.
Vishwesh Nath, Wenqi Li 0001, Dong Yang 0005, Andriy Myronenko, Mingxin Zheng, Yao Lu 0006, Hongxu Yin, Yee Man Law, Yucheng Tang, Can Zhao 0001, Ziyue Xu 0001, Yufan He, Stephanie A. Harmon, Benjamin Simon, Greg Heinrich, Stephen R. Aylward, Marc Edgar, Michael Zephyr, Pavlo Molchanov 0001, Baris Turkbey, Holger Roth, Daguang Xu
CVPR24
2025 MedSegFactory: Text-Guided Generation of Medical Image-Mask Pairs
abstract
This paper presents MedSegFactory, a versatile medical synthesis framework that generates high-quality paired medical images and segmentation masks across modalities and tasks. It aims to serve as an unlimited data repository, supplying image-mask pairs to enhance existing segmentation tools. The core of MedSegFactory is a dual-stream diffusion model, where one stream synthesizes medical images and the other generates corresponding segmentation masks. To ensure precise alignment between image-mask pairs, we introduce Joint Cross-Attention (JCA), enabling a collaborative denoising paradigm by dynamic cross-conditioning between streams. This bidirectional interaction allows both representations to guide each other's generation, enhancing consistency between generated pairs. MedSegFactory unlocks on-demand generation of paired medical images and segmentation masks through user-defined prompts that specify the target labels, imaging modalities, anatomical regions, and pathological conditions, facilitating scalable and high-quality data generation. This new paradigm of medical image synthesis enables seamless integration into diverse medical imaging workflows, enhancing both efficiency and accuracy. Extensive experiments show that MedSegFactory generates data of superior quality and usability, achieving competitive or state-of-the-art performance in 2D and 3D segmentation tasks while addressing data scarcity and regulatory constraints.
Yuhan Wang 0001, Yucheng Tang, Daguang Xu, Kang Wang 0016, Yang Yang 0009, Zongwei Zhou, Yuyin Zhou
ICCV4
2025 Synthesis of Pathological Dual-Channel Color Doppler Echocardiograms for Equitable Diagnosis of Heart Diseases
Pooneh Roshanitabrizi, Artur Arturi Aharonyan, Kelsey Brown, Taylor Gloria Broudy, Abhijeet Parida, Austin Tapp, Zhifan Jiang, Alison Tompsett, Joselyn Rwebembera, Emmy Okello, Andrea Z. Beaton, Holger Roth, Daguang Xu, Syed Muhammad Anwar, Craig A. Sable, Marius George Linguraru
MICCAI (2)14
2025 Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging
abstract
Recent progress in vision-language modeling for 3D medical imaging has been fueled by large-scale computed tomography (CT) corpora with paired free-text reports, stronger architectures, and powerful pretrained models. This has enabled applications such as automated report generation and text-conditioned 3D image synthesis. Yet, current approaches struggle with high-resolution, long-sequence volumes: contrastive pretraining often yields vision encoders that are misaligned with clinical language, and slice-wise tokenization blurs fine anatomy, reducing diagnostic performance on downstream tasks. We introduce BTB3D (Better Tokens for Better 3D), a causal convolutional encoder-decoder that unifies 2D and 3D training and inference while producing compact, frequency-aware volumetric tokens. A three-stage training curriculum enables (i) local reconstruction, (ii) overlapping-window tiling, and (iii) long-context decoder refinement, during which the model learns from short slice excerpts yet generalizes to scans exceeding $300$ slices without additional memory overhead. BTB3D sets a new state-of-the-art on two key tasks: it improves BLEU scores and increases clinical F1 by 40\% over CT2Rep, CT-CHAT, and Merlin for report generation; and it reduces FID by 75\% and halves FVD compared to GenerateCT and MedSyn for text-to-CT synthesis, producing anatomically consistent $512\times512\times241$ volumes. These results confirm that precise three-dimensional tokenization, rather than larger language backbones alone, is essential for scalable vision-language modeling in 3D medical imaging. The codebase is available at: https://github.com/ibrahimethemhamamci/BTB3D
Ibrahim Ethem Hamamci, Sezgin Er, Suprosanna Shit, Hadrien Reynaud, Dong Yang 0005, Marc Edgar, Daguang Xu, Bernhard Kainz, Bjoern Menze
NeurIPS8
2025 PanTS: The Pancreatic Tumor Segmentation Dataset
abstract
PanTS is a large-scale, multi-institutional dataset curated to advance research in pancreatic CT analysis. It contains 36,390 CT scans from 145 medical centers, with expert-validated, voxel-wise annotations of over 993,000 anatomical structures, covering pancreatic tumors, pancreas head, body, and tail, and 24 surrounding anatomical structures such as vascular/skeletal structures and abdominal/thoracic organs. Each scan includes metadata such as patient age, sex, diagnosis, contrast phase, in-plane spacing, slice thickness, etc. AI models trained on PanTS achieve significantly better performance in pancreatic tumor detection, localization, and segmentation than those trained on existing public datasets. Our analysis indicates that these gains are directly attributable to the 16× larger-scale tumor annotations and indirectly supported by the 24 additional surrounding anatomical structures. As the largest and most comprehensive resource of its kind, PanTS offers a new benchmark for developing and evaluating AI models in pancreatic CT analysis.
Xinze Zhou, Qi Chen 0014, Pedro R. A. S. Bassi, Xiaoxi Chen, Zheren Zhu, Kang Wang 0016, Yang Yang 0009, Yucheng Tang, Daguang Xu, Alan L. Yuille, Zongwei Zhou
NeurIPS14
2025 MAISI: Medical AI for Synthetic Imaging
abstract
Medical imaging analysis faces challenges such as data scarcity, high annotation costs, and privacy concerns. This paper introduces the Medical AI for Synthetic Imaging (MAISI), an innovative approach using the diffusion model to generate synthetic 3D computed tomography (CT) images to address those challenges. MAISI leverages the foundation volume compression network and the latent diffusion model to produce high-resolution CT images (up to a landmark volume dimension of 512 × 512 × 768) with flexible volume dimensions and voxel spacing. By incorporating ControlNet, MAISI can process organ segmentation, including 127 anatomical structures, as additional conditions and enables the generation of accurately annotated synthetic images that can be used for various downstream tasks. Our experiment results show that MAISI's capabilities in generating realistic, anatomically accurate images for diverse regions and conditions reveal its promising potential to mitigate challenges using synthetic data.
Can Zhao 0001, Dong Yang 0005, Ziyue Xu 0001, Vishwesh Nath, Yucheng Tang, Benjamin Simon, Mason Belue, Stephanie A. Harmon, Baris Turkbey, Daguang Xu
WACV11
2025 From image to report: automating lung cancer screening interpretation and reporting with vision-language models
Tien-Yu Chang, Qinglin Gou, Leyi Zhao, Tiancheng Zhou, Dong Yang 0005, Huiwen Ju, Kaleb E. Smith, Chengkun Sun, Jinqian Pan, Yu Huang 0018, Xing He 0003, Xuhong Zhang 0001, Daguang Xu, Jie Xu 0012, Jiang Bian 0001, Aokun Chen
J. Biomed. Informatics14
2025 Multi-Center Fetal Brain Tissue Annotation (FeTA) Challenge 2022 Results
abstract
Segmentation is a critical step in analyzing the developing human fetal brain. There have been vast improvements in automatic segmentation methods in the past several years, and the Fetal Brain Tissue Annotation (FeTA) Challenge 2021 helped to establish an excellent standard of fetal brain segmentation. However, FeTA 2021 was a single center study, limiting real-world clinical applicability and acceptance. The multi-center FeTA Challenge 2022 focused on advancing the generalizability of fetal brain segmentation algorithms for magnetic resonance imaging (MRI). In FeTA 2022, the training dataset contained images and corresponding manually annotated multi-class labels from two imaging centers, and the testing data contained images from these two centers as well as two additional unseen centers. The multi-center data included different MR scanners, imaging parameters, and fetal brain super-resolution algorithms applied. 16 teams participated and 17 algorithms were evaluated. Here, the challenge results are presented, focusing on the generalizability of the submissions. Both in- and out-of-domain, the white matter and ventricles were segmented with the highest accuracy (Top Dice scores: 0.89, 0.87 respectively), while the most challenging structure remains the grey matter (Top Dice score: 0.75) due to anatomical complexity. The top 5 average Dices scores ranged from 0.81-0.82, the top 5 average percentile Hausdorff distance values ranged from 2.3-2.5mm, and the top 5 volumetric similarity scores ranged from 0.90-0.92. The FeTA Challenge 2022 was able to successfully evaluate and advance generalizability of multi-class fetal brain tissue segmentation algorithms for MRI and it continues to benchmark new algorithms.
Kelly Payette, Céline Steger, Roxane Licandro, Priscille de Dumast, Hongwei Li 0004, Matthew J. Barkovich, Liu Li 0001, Maik Dannecker, Chen Chen 0042, Cheng Ouyang, Niccolò McConnell, Alina Dana Miron, Yongmin Li 0001, Alena Uus, Irina Grigorescu, Paula Ramirez Gilliland, Md Mahfuzur Rahman Siddiquee, Daguang Xu, Andriy Myronenko, Haoyu Wang 0010, Ziyan Huang, Jin Ye 0002, Mireia Alenyà, Valentin Comte, Oscar Camara 0001, Jean-Baptiste Masson, Astrid Nilsson, Charlotte Godard, Moona Mazher, Abdul Qayyum 0002, Yibo Gao, Hangqi Zhou, Shangqi Gao, Guiming Dong, Guotai Wang, ZunHyan Rieu, HyeonSik Yang, Szymon Plotka, Michal K. Grzeszczyk, Arkadiusz Sitek, Luisa Vargas Daza, Santiago Usma, Pablo Andrés Arbeláez, Wenying Lu, Romain Valabrègue, Anand A. Joshi, Krishna N. Nayak, Richard M. Leahy, Luca Wilhelmi, Aline Dändliker, Antonio G. Gennari, Anton Jakovcic, Melita Klaic, Ana Adzic, Pavel Markovic, Gracia Grabaric, Gregor Kasprian, Gregor Dovjak, Milan Rados, Lana Vasung, Meritxell Bach Cuadra, András Jakab
IEEE Trans. Medical Imaging18
2024 Perada: Parameter-Efficient Federated Learning Personalization with Generalization Guarantees
abstract
Personalized Federated Learning (pFL) has emerged as a promising solution to tackle data heterogeneity across clients in FL. However, existing pFL methods either (1) introduce high computation and communication costs or (2) overfit to local data, which can be limited in scope and vulnerable to evolved test samples with natural distribution shifts. In this paper, we propose Perada,a parameter-efficient pFL framework that reduces communication and computational costs and exhibits superior generalization performance, es-pecially under test-time distribution shifts. Peradareduces the costs by leveraging the power of pretrained models and only updates and communicates a small number of additional parameters from adapters. Peradaachieves high generalization by regularizing each client's personalized adapter with a global adapter, while the global adapter uses knowledge distillation to aggregate generalized information from all clients. Theoretically, we provide generalization bounds of Perada,and we prove its convergence to station-ary points under non-convex settings. Empirically, Peradademonstrates higher personalized performance (+4.85% on CheXpert) and enables better out-of-distribution generalization (+5.23% on CIFAR-10-C) on different datasets across natural and medical domains compared with baselines, while only updating 12.6% of parameters per model. Our code is available at https://github.com/NVlabsIPerAda.
Chulin Xie, De-An Huang, Wenda Chu, Daguang Xu, Chaowei Xiao, Bo Li 0026, Anima Anandkumar
CVPR4
2024 FedBPT: Efficient Federated Black-box Prompt Tuning for Large Language Models
abstract
Pre-trained language models (PLM) have revolutionized the NLP landscape, achieving stellar performances across diverse tasks. These models, while benefiting from vast training data, often require fine-tuning on specific data to cater to distinct downstream tasks. However, this data adaptation process has inherent security and privacy concerns, primarily when leveraging user-generated, device-residing data. Federated learning (FL) provides a solution, allowing collaborative model fine-tuning without centralized data collection. However, applying FL to finetune PLMs is hampered by challenges, including restricted model parameter access due to the high encapsulation, high computational requirements, and communication overheads. This paper introduces Federated Black-box Prompt Tuning (FedBPT), a framework designed to address these challenges. FedBPT allows the clients to treat the model as a black-box inference API. By focusing on training optimal prompts and utilizing gradient-free optimization methods, FedBPT reduces the number of exchanged variables, boosts communication efficiency, and minimizes computational and storage costs. Experiments highlight the framework’s ability to drastically cut communication and memory costs while maintaining competitive performance. Ultimately, FedBPT presents a promising solution for efficient, privacy-preserving fine-tuning of PLM in the age of large language models.
Jingwei Sun 0002, Ziyue Xu 0001, Hongxu Yin, Dong Yang 0005, Daguang Xu, Zhixu Du, Yiran Chen 0001, Holger Roth
ICML5
2024 Federated Black-box Prompt Tuning System for Large Language Models on the Edge
abstract
Federated learning (FL) offers a privacy-preserving way to train models across decentralized data. However, fine-tuning pre-trained language models (PLMs) in FL is challenging due to restricted model parameter access, high computational demands, and communication overheads. Our method treats large language models (LLMs) as black-box inference APIs, optimizing prompts with gradient-free methods. This approach, FedBPT, reduces exchanged variables, boosts communication efficiency, and minimizes computational and memory costs. We demonstrate the practical implementation of FedBPT on resource-limited edge devices, showcasing its ability to efficiently achieve collaborative on-device LLM fine-tuning.
Jingwei Sun 0002, Ang Li 0005, Beidi Chen, Holger Roth, Daguang Xu, Tingjun Chen, Yiran Chen 0001
MobiCom8
2024 Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?
abstract
How can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain.
Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou
NeurIPS48
2024 IR-FRestormer: Iterative Refinement with Fourier-Based Restormer for Accelerated MRI Reconstruction
abstract
Accelerated magnetic resonance imaging (MRI) aims to reconstruct high-quality MR images from a set of under-sampled measurements. State-of-the-art methods for this task use deep learning, which offers high reconstruction accuracy and fast runtimes. In this work, we propose a new state-of-the-art reconstruction model for accelerated MRI reconstruction. Our model is the first to combine the power of deep neural networks with iterative refinement for this task. For the neural network component of our method, we utilize a transformer-based architecture as transformers are state-of-the-art in various image reconstruction tasks. However, a major drawback of transformers which has limited their emergence among the state-of-the-art MRI models is that they are often memory inefficient for high-resolution inputs. To address this limitation, we propose a transformer-based model which uses parameter-free Fourier-based attention modules, achieving 2× more memory efficiency. We evaluate our model on the largest publicly available MRI dataset, the fastMRI dataset [46], and achieve on-par performance with other state-of-the-art1methods on the dataset’s leaderboard2.
Mohammad Zalbagi Darestani, Vishwesh Nath, Wenqi Li 0001, Yufan He, Holger Roth, Ziyue Xu 0001, Daguang Xu, Reinhard Heckel, Can Zhao 0001
WACV7
2024 Learning Quality Labels for Robust Image Classification
abstract
Supervised learning paradigms largely benefit from the tremendous amount of annotated data. However, the quality of the annotations often varies among labelers. Multi-observer studies have been conducted to examine the annotation variances (by labeling the same data multiple times) to see how it affects critical applications like medical image analysis. In this paper, we demonstrate how multiple sets of annotations (either hand-labeled or algorithm-generated) can be utilized together and mutually benefit the learning of classification tasks. A scheme of learning-to-vote is introduced to sample quality label sets for each data entry on-the-fly during the training. Specifically, a label-sampling module is designed to achieve refined labels (weighted sum of attended ones) that benefit the model learning the most through additional back-propagations. We apply the learning-to-vote scheme on the classification task of a synthetic noisy CIFAR-10 to prove the concept and then demonstrate superior results (3-5% increase on average in multiple disease classification AUCs) on the chest x-ray images from a hospital-scale dataset (MIMIC-CXR) and hand-labeled dataset (OpenI) in comparison to regular training paradigms.
Xiaosong Wang 0001, Ziyue Xu 0001, Dong Yang 0005, Leo K. Tam, Holger Roth, Daguang Xu
WACV6
2024 MONAI Label: A framework for AI-assisted interactive labeling of 3D medical images
Andres Diaz-Pinto, Sachidanand Alle, Vishwesh Nath, Yucheng Tang, Alvin Ihsani, Muhammad Asad 0001, Fernando Pérez-García, Pritesh Mehta, Wenqi Li 0001, Mona Flores, Holger Roth, Tom Vercauteren, Daguang Xu, Prerna Dogra, Sébastien Ourselin, Andrew Feng, Manuel Jorge Cardoso
Medical Image Anal.13
2024 Guest Editorial: Trustworthy Machine Learning for Health Informatics
abstract
Machine learning (ML), the stem of today's artificial intelligence, has shown significant growth in the field of biomedical and health informatics. On the one hand, ML techniques are becoming more complex in order to deal with real-world data. On the other hand, ML is also more and more accessible to broader users. For example, automated machine learning products are enabling users to build their own ML models without writing code [1].
Luyang Luo, Daguang Xu, Harry Qin, Yueming Jin, Hao Chen 0011
IEEE J. Biomed. Health Informatics2
2023 Fair Federated Medical Image Segmentation via Client Contribution Estimation
abstract
How to ensure fairness is an important topic in federated learning (FL). Recent studies have investigated how to reward clients based on their contribution (collaboration fairness), and how to achieve uniformity of performance across clients (performance fairness). Despite achieving progress on either one, we argue that it is critical to consider them together, in order to engage and motivate more diverse clients joining FL to derive a high-quality global model. In this work, we propose a novel method to optimize both types of fairness simultaneously. Specifically, we propose to estimate client contribution in gradient and data space. In gradient space, we monitor the gradient direction differences of each client with respect to others. And in data space, we measure the prediction error on client data using an auxiliary model. Based on this contribution estimation, we propose a FL method, federated training via contribution estimation (FedCE), i.e., using estimation as global model aggregation weights. We have theoretically analyzed our method and empirically evaluated it on two real-world medical datasets. The effectiveness of our approach has been validated with significant performance improvements, better collaboration fairness, better performance fairness, and comprehensive analytical studies. Code is available at https://nvidia.github.io/NVFlare/research/fed-ce
Meirui Jiang, Holger Roth, Wenqi Li 0001, Dong Yang 0005, Can Zhao 0001, Vishwesh Nath, Daguang Xu, Qi Dou 0001, Ziyue Xu 0001
CVPR7
2023 Communication-Efficient Vertical Federated Learning with Limited Overlapping Samples
abstract
Federated learning is a popular collaborative learning approach that enables clients to train a global model without sharing their local data. Vertical federated learning (VFL) deals with scenarios in which the data on clients have different feature spaces but share some overlapping samples. Existing VFL approaches suffer from high communication costs and cannot deal efficiently with limited overlapping samples commonly seen in the real world. We propose a practical VFL framework called one-shot VFL that can solve the communication bottleneck and the problem of limited overlapping samples simultaneously based on semi-supervised learning. We also propose few-shot VFL to improve the accuracy further with just one more communication round between the server and the clients. In our proposed framework, the clients only need to communicate with the server once or only a few times. We evaluate the proposed VFL framework on both image and tabular datasets. Our methods can improve the accuracy by more than 46.5% and reduce the communication cost by more than 330× compared with state-of-the-art VFL methods when evaluated on CIFAR-10. Our code is available at https://nvidia.github.io/NVFlare/research/one-shot-vfl.
Jingwei Sun 0002, Ziyue Xu 0001, Dong Yang 0005, Vishwesh Nath, Wenqi Li 0001, Can Zhao 0001, Daguang Xu, Yiran Chen 0001, Holger Roth
ICCV7
2023 SwinUNETR-V2: Stronger Swin Transformers with Stagewise Convolutions for 3D Medical Image Segmentation
Yufan He, Vishwesh Nath, Dong Yang 0005, Yucheng Tang, Andriy Myronenko, Daguang Xu
MICCAI (4)6
2023 DAST: Differentiable Architecture Search with Transformer for 3D Medical Image Segmentation
Dong Yang 0005, Ziyue Xu 0001, Yufan He, Vishwesh Nath, Wenqi Li 0001, Andriy Myronenko, Ali Hatamizadeh, Can Zhao 0001, Holger Roth, Daguang Xu
MICCAI (3)10
2023 The Liver Tumor Segmentation Benchmark (LiTS)
abstract
In this work, we report the set-up and results of the Liver Tumor Segmentation Benchmark (LiTS), which was organized in conjunction with the IEEE International Symposium on Biomedical Imaging (ISBI) 2017 and the International Conferences on Medical Image Computing and Computer-Assisted Intervention (MICCAI) 2017 and 2018. The image dataset is diverse and contains primary and secondary tumors with varied sizes and appearances with various lesion-to-background levels (hyper-/hypo-dense), created in collaboration with seven hospitals and research institutions. Seventy-five submitted liver and liver tumor segmentation algorithms were trained on a set of 131 computed tomography (CT) volumes and were tested on 70 unseen test images acquired from different patients. We found that not a single algorithm performed best for both liver and liver tumors in the three events. The best liver segmentation algorithm achieved a Dice score of 0.963, whereas, for tumor segmentation, the best algorithms achieved Dices scores of 0.674 (ISBI 2017), 0.702 (MICCAI 2017), and 0.739 (MICCAI 2018). Retrospectively, we performed additional analysis on liver tumor detection and revealed that not all top-performing segmentation algorithms worked well for tumor detection. The best liver tumor detection method achieved a lesion-wise recall of 0.458 (ISBI 2017), 0.515 (MICCAI 2017), and 0.554 (MICCAI 2018), indicating the need for further research. LiTS remains an active benchmark and resource for research, e.g., contributing the liver-related segmentation tasks in http://medicaldecathlon.com/. In addition, both data and online evaluation are accessible via https://competitions.codalab.org/competitions/17094.
Patrick Bilic, Patrick Ferdinand Christ, Hongwei Li 0004, Eugene Vorontsov, Avi Ben-Cohen, Georgios Kaissis, Adi Szeskin, Colin Jacobs, Gabriel Efrain Humpire Mamani, Gabriel Chartrand, Fabian Lohöfer, Julian Walter Holch, Wieland H. Sommer, Felix Hofmann, Alexandre Hostettler, Naama Lev-Cohain, Michal Drozdzal, Michal Amitai, Refael Vivanti, Jacob Sosna, Ivan Ezhov, Anjany Sekuboyina, Fernando Navarro, Florian Kofler, Johannes C. Paetzold, Suprosanna Shit, Xiaobin Hu, Jana Lipková, Markus Rempfler, Marie Piraud, Jan Kirschke, Benedikt Wiestler, Christian Hülsemeyer, Marcel Beetz, Florian Ettlinger, Michela Antonelli, Woong Bae, Miriam Bellver, Lei Bi 0001, Hao Chen 0011, Grzegorz Chlebus, Erik Dam, Qi Dou 0001, Chi-Wing Fu, Bogdan Georgescu, Xavier Giró-i-Nieto, Felix Grün, Xu Han 0009, Pheng-Ann Heng, Jürgen Hesser, Jan Hendrik Moltz, Christian Igel, Fabian Isensee, Paul F. Jaeger, Fucang Jia, Krishna Chaitanya Kaluva, Mahendra Khened, Ildoo Kim, Jae-Hun Kim, Sungwoong Kim, Simon Kohl, Tomasz K. Konopczynski, Avinash Kori, Ganapathy Krishnamurthi, Xiaomeng Li 0001, John S. Lowengrub, Jun Ma 0016, Klaus H. Maier-Hein, Kevis-Kokitsi Maninis, Hans Meine, Dorit Merhof, Akshay Pai, Mathias Perslev, Jens Petersen, Jordi Pont-Tuset, Xiaojuan Qi 0001, Oliver Rippel, Karsten Roth, Ignacio Sarasua, Andrea Schenk, Zengming Shen, Jordi Torres, Christian Wachinger, Chunliang Wang, Leon Weninger, Daguang Xu, Xiaoping Yang 0001, Simon C. H. Yu, Yading Yuan, Miao Yue, Liping Zhang 0009, Manuel Jorge Cardoso, Spyridon Bakas, Rickmer Braren, Volker Heinemann, Christopher Joseph Pal, An Tang, Samuel Kadoury, Luc Soler, Bram van Ginneken, Hayit Greenspan, Leo Joskowicz, Bjoern Menze
Medical Image Anal.92
2023 Fetal brain tissue annotation and segmentation challenge results
abstract
In-utero fetal MRI is emerging as an important tool in the diagnosis and analysis of the developing human brain. Automatic segmentation of the developing fetal brain is a vital step in the quantitative analysis of prenatal neurodevelopment both in the research and clinical context. However, manual segmentation of cerebral structures is time-consuming and prone to error and inter-observer variability. Therefore, we organized the Fetal Tissue Annotation (FeTA) Challenge in 2021 in order to encourage the development of automatic segmentation algorithms on an international level. The challenge utilized FeTA Dataset, an open dataset of fetal brain MRI reconstructions segmented into seven different tissues (external cerebrospinal fluid, gray matter, white matter, ventricles, cerebellum, brainstem, deep gray matter). 20 international teams participated in this challenge, submitting a total of 21 algorithms for evaluation. In this paper, we provide a detailed analysis of the results from both a technical and clinical perspective. All participants relied on deep learning methods, mainly U-Nets, with some variability present in the network architecture, optimization, and image pre- and post-processing. The majority of teams used existing medical imaging deep learning frameworks. The main differences between the submissions were the fine tuning done during training, and the specific pre- and post-processing steps performed. The challenge results showed that almost all submissions performed similarly. Four of the top five teams used ensemble learning methods. However, one team's algorithm performed significantly superior to the other submissions, and consisted of an asymmetrical U-Net network architecture. This paper provides a first of its kind benchmark for future automatic multi-tissue segmentation algorithms for the developing human brain in utero.
Kelly Payette, Hongwei Li 0004, Priscille de Dumast, Roxane Licandro, Md Mahfuzur Rahman Siddiquee, Daguang Xu, Andriy Myronenko, Yuchen Pei, Lisheng Wang, Juanying Xie, Huiquan Zhang, Guiming Dong, Hao Fu 0014, Guotai Wang, ZunHyan Rieu, Hyun Gi Kim, Davood Karimi, Ali Gholipour, Helena R. Torres, Bruno Oliveira 0002, João L. Vilaça, Netanell Avisdris, Ori Ben-Zvi, Dafna Ben-Bashat, Lucas Fidon, Michael Aertsen, Tom Vercauteren, Daniel Sobotka, Georg Langs, Mireia Alenyà, Maria Inmaculada Villanueva, Oscar Camara 0001, Bella Specktor-Fadida, Leo Joskowicz, Liao Weibin, Lv Yi, Xuesong Li 0003, Moona Mazher, Abdul Qayyum 0002, Domenec Puig, Hamza Kebiri, KuanLun Liao, YiXuan Wu, JinTai Chen, Yunzhi Xu, Lana Vasung, Bjoern Menze, Meritxell Bach Cuadra, András Jakab
Medical Image Anal.7
2023 Do Gradient Inversion Attacks Make Federated Learning Unsafe?
abstract
Federated learning (FL) allows the collaborative training of AI models without needing to share raw data. This capability makes it especially interesting for healthcare applications where patient and data privacy is of utmost concern. However, recent works on the inversion of deep neural networks from model gradients raised concerns about the security of FL in preventing the leakage of training data. In this work, we show that these attacks presented in the literature are impractical in FL use-cases where the clients' training involves updating the Batch Normalization (BN) statistics and provide a new baseline attack that works for such scenarios. Furthermore, we present new ways to measure and visualize potential data leakage in FL. Our work is a step towards establishing reproducible methods of measuring data leakage in FL and could help determine the optimal tradeoffs between privacy-preserving techniques, such as differential privacy, and model accuracy based on quantifiable metrics.
Ali Hatamizadeh, Hongxu Yin, Pavlo Molchanov 0001, Andriy Myronenko, Wenqi Li 0001, Prerna Dogra, Andrew Feng, Mona Flores, Jan Kautz, Daguang Xu, Holger Roth
IEEE Trans. Medical Imaging10
2022 GradViT: Gradient Inversion of Vision Transformers
abstract
In this work we demonstrate the vulnerability of vision transformers (ViTs) to gradient-based inversion attacks. During this attack, the original data batch is reconstructed given model weights and the corresponding gradients. We introduce a method, named GradViT, that optimizes random noise into naturally looking images via an iterative process. The optimization objective consists of (i) a loss on matching the gradients, (ii) image prior in the form of distance to batch-normalization statistics of a pretrained CNN model, and (iii) a total variation regularization on patches to guide correct recovery locations. We propose a unique loss scheduling function to overcome local minima during optimization. We evaluate GadViT on ImageNet1K and MS-Celeb-1M datasets, and observe unprecedentedly high fidelity and closeness to the original (hidden) data. During the analysis we find that vision transformers are significantly more vulnerable than previously studied CNNs due to the presence of the attention mechanism. Our method demonstrates new state-of-the-art results for gradient inversion in both qualitative and quantitative metrics. Project page at https://gradvit.github.io/.
Ali Hatamizadeh, Hongxu Yin, Holger Roth, Wenqi Li 0001, Jan Kautz, Daguang Xu, Pavlo Molchanov 0001
CVPR6
2022 HyperSegNAS: Bridging One-Shot Neural Architecture Search with 3D Medical Image Segmentation using HyperNet
abstract
Semantic segmentation of 3D medical images is a challenging task due to the high variability of the shape and pattern of objects (such as organs or tumors). Given the recent success of deep learning in medical image segmentation, Neural Architecture Search (NAS) has been introduced to find high-performance 3D segmentation network architectures. However, because of the massive computational requirements of 3D data and the discrete optimization nature of architecture search, previous NAS methods require a long search time or necessary continuous relaxation, and commonly lead to sub-optimal network architectures. While one-shot NAS can potentially address these disadvantages, its application in the segmentation domain has not been well studied in the expansive multi-scale multi-path search space. To enable one-shot NAS for medical image segmentation, our method, named HyperSegNAS, introduces a HyperNet to assist super-net training by incorporating architecture topology information. Such a HyperNet can be removed once the super-net is trained and introduces no overhead during architecture search. We show that HyperSegNAS yields better performing and more intuitive architectures compared to the previous state-of-the-art (SOTA) segmentation networks; furthermore, it can quickly and accurately find good architecture candidates under different computing constraints. Our method is evaluated on public datasets from the Medical Segmentation Decathlon (MSD) challenge, and achieves SOTA performances.
Cheng Peng 0008, Andriy Myronenko, Ali Hatamizadeh, Vishwesh Nath, Md Mahfuzur Rahman Siddiquee, Yufan He, Daguang Xu, Rama Chellappa, Dong Yang 0005
CVPR7
2022 Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image Analysis
abstract
Vision Transformers (ViT)s have shown great performance in self-supervised learning of global and local representations that can be transferred to downstream applications. Inspired by these results, we introduce a novel self-supervised learning framework with tailored proxy tasks for medical image analysis. Specifically, we propose: (i) a new 3D transformer-based model, dubbed Swin UNEt TRansformers (Swin UNETR), with a hierarchical encoder for self-supervised pretraining; (ii) tailored proxy tasks for learning the underlying pattern of human anatomy. We demonstrate successful pre-training of the proposed model on 5,050 publicly available computed tomography (CT) images from various body organs. The effectiveness of our approach is validated by fine-tuning the pre-trained models on the Beyond the Cranial Vault (BTCV) Segmentation Challenge with 13 abdominal organs and segmentation tasks from the Medical Segmentation Decathlon (MSD) dataset. Our model is currently the state-of-the-art on the public test leaderboards of both MSD11https://decathlon-10.grand-challenge.org/evaluation/challenge/leaderboard/ and BTCV22https://www.synapse.org/#!Synapse:syn3193805/wiki/217785/ datasets. Code: https://monai.io/research/swin-unetr.
Yucheng Tang, Dong Yang 0005, Wenqi Li 0001, Holger Roth, Bennett A. Landman, Daguang Xu, Vishwesh Nath, Ali Hatamizadeh
CVPR6
2022 Closing the Generalization Gap of Cross-silo Federated Medical Image Segmentation
abstract
Cross-silo federated learning (FL) has attracted much attention in medical imaging analysis with deep learning in recent years as it can resolve the critical issues of insufficient data, data privacy, and training efficiency. However, there can be a generalization gap between the model trained from FL and the one from centralized training. This important issue comes from the non-iid data distribution of the local data in the participating clients and is well-known as client drift. In this work, we propose a novel training frame-work FedSM to avoid the client drift issue and successfully close the generalization gap compared with the centralized training for medical image segmentation tasks for the first time. We also propose a novel personalized FL objective formulation and a new method SoftPull to solve it in our proposed framework FedSM. We conduct rigorous theoretical analysis to guarantee its convergence for optimizing the non-convex smooth objective function. Real-world medical image segmentation experiments using deep FL validate the motivations and effectiveness of our proposed method.
An Xu, Wenqi Li 0001, Dong Yang 0005, Holger Roth, Ali Hatamizadeh, Can Zhao 0001, Daguang Xu, Heng Huang 0001, Ziyue Xu 0001
CVPR8
2022 Auto-FedRL: Federated Hyperparameter Optimization for Multi-institutional Medical Image Segmentation
Dong Yang 0005, Ali Hatamizadeh, An Xu, Ziyue Xu 0001, Wenqi Li 0001, Can Zhao 0001, Daguang Xu, Stephanie A. Harmon, Evrim Turkbey, Baris Turkbey, Bradford J. Wood, Francesca Patella, Elvira Stellato, Gianpaolo Carrafiello, Vishal M. Patel, Holger Roth
ECCV (21)8
2022 Efficient Population Based Hyperparameter Scheduling for Medical Image Segmentation
Yufan He, Dong Yang 0005, Andriy Myronenko, Daguang Xu
MICCAI (5)4
2022 Warm Start Active Learning with Proxy Labels and Selection via Semi-supervised Fine-Tuning
Vishwesh Nath, Dong Yang 0005, Holger Roth, Daguang Xu
MICCAI (8)4
2022 Clinical-Realistic Annotation for Histopathology Images with Probabilistic Semi-supervision: A Worst-Case Study
Ziyue Xu 0001, Andriy Myronenko, Dong Yang 0005, Holger Roth, Can Zhao 0001, Xiaosong Wang 0001, Daguang Xu
MICCAI (2)7
2022 UNETR: Transformers for 3D Medical Image Segmentation
abstract
Fully Convolutional Neural Networks (FCNNs) with contracting and expanding paths have shown prominence for the majority of medical image segmentation applications since the past decade. In FCNNs, the encoder plays an integral role by learning both global and local features and contextual representations which can be utilized for semantic output prediction by the decoder. Despite their success, the locality of convolutional layers in FCNNs, limits the capability of learning long-range spatial dependencies. Inspired by the recent success of transformers for Natural Language Processing (NLP) in long-range sequence learning, we reformulate the task of volumetric (3D) medical image segmentation as a sequence-to-sequence prediction problem. We introduce a novel architecture, dubbed as UNEt TRansformers (UNETR), that utilizes a transformer as the encoder to learn sequence representations of the input volume and effectively capture the global multi-scale information, while also following the successful "U-shaped" network design for the encoder and decoder. The transformer encoder is directly connected to a decoder via skip connections at different resolutions to compute the final semantic segmentation output. We have validated the performance of our method on the Multi Atlas Labeling Beyond The Cranial Vault (BTCV) dataset for multi-organ segmentation and the Medical Segmentation Decathlon (MSD) dataset for brain tumor and spleen segmentation tasks. Our benchmarks demonstrate new state-of-the-art performance on the BTCV leaderboard.
Ali Hatamizadeh, Yucheng Tang, Vishwesh Nath, Dong Yang 0005, Andriy Myronenko, Bennett A. Landman, Holger Roth, Daguang Xu
WACV8
2022 Rapid artificial intelligence solutions in a pandemic - The COVID-19-20 Lung CT Lesion Segmentation Challenge
Holger Roth, Ziyue Xu 0001, Carlos Tor-Díez, Ramon Sánchez-Jacob, Jonathan Zember, Jose Molto, Wenqi Li 0001, Sheng Xu 0001, Baris Turkbey, Evrim Turkbey, Dong Yang 0005, Ahmed Harouni, Nicola Rieke, Shishuai Hu, Fabian Isensee, Claire Tang, Qinji Yu, Jan Sölter, Vitali Liauchuk, Jan Hendrik Moltz, Bruno Oliveira 0002, Yong Xia 0001, Klaus H. Maier-Hein, Qikai Li, Andreas Husch, Vassili Kovalev, Alessa Hering, João L. Vilaça, Mona Flores, Daguang Xu, Bradford J. Wood, Marius George Linguraru
Medical Image Anal.34
2021 DiNTS: Differentiable Neural Network Topology Search for 3D Medical Image Segmentation
abstract
Recently, neural architecture search (NAS) has been applied to automatically search high-performance networks for medical image segmentation. The NAS search space usually contains a network topology level (controlling connections among cells with different spatial scales) and a cell level (operations within each cell). Existing methods either require long searching time for large-scale 3D image datasets, or are limited to pre-defined topologies (such as U-shaped or single-path) . In this work, we focus on three important aspects of NAS in 3D medical image segmentation: flexible multi-path network topology, high search efficiency, and budgeted GPU memory usage. A novel differentiable search framework is proposed to support fast gradient-based search within a highly flexible network topology search space. The discretization of the searched optimal continuous model in differentiable scheme may produce a sub-optimal final discrete model (discretization gap). Therefore, we propose a topology loss to alleviate this problem. In addition, the GPU memory usage for the searched 3D model is limited with budget constraints during search. Our Differentiable Network Topology Search scheme (DiNTS) is evaluated on the Medical Segmentation Decathlon (MSD) challenge, which contains ten challenging segmentation tasks. Our method achieves the state-of-the-art performance and the top ranking on the MSD challenge leaderboard.
Yufan He, Dong Yang 0005, Holger Roth, Can Zhao 0001, Daguang Xu
CVPR5
2021 T-AutoML: Automated Machine Learning for Lesion Segmentation using Transformers in 3D Medical Imaging
abstract
Lesion segmentation in medical imaging has been an important topic in clinical research. Researchers have proposed various detection and segmentation algorithms to address this task. Recently, deep learning-based approaches have significantly improved the performance over conventional methods. However, most state-of-the-art deep learning methods require the manual design of multiple network components and training strategies. In this paper, we propose a new automated machine learning algorithm, T-AutoML, which not only searches for the best neural architecture, but also finds the best combination of hyper-parameters and data augmentation strategies simultaneously. The proposed method utilizes the modern transformer model, which is introduced to adapt to the dynamic length of the search space embedding and can significantly improve the ability of the search. We validate T-AutoML on several large-scale public lesion segmentation data-sets and achieve state-of-the-art performance.
Dong Yang 0005, Andriy Myronenko, Xiaosong Wang 0001, Ziyue Xu 0001, Holger Roth, Daguang Xu
ICCV6
2021 Test-Time Training for Deformable Multi-Scale Image Registration
abstract
Registration is a fundamental task in medical robotics and is often a crucial step for many downstream tasks such as motion analysis, intra-operative tracking and image segmentation. Popular registration methods such as ANTs and NiftyReg optimize objective functions for each pair of images from scratch, which are time-consuming for 3D and sequential images with complex deformations. Recently, deep learning-based registration approaches such as VoxelMorph have been emerging and achieve competitive performance. In this work, we construct a test-time training for deep deformable image registration to improve the generalization ability of conventional learning-based registration model. We design multi-scale deep networks to consecutively model the residual deformations, which is effective for high variational deformations. Extensive experiments validate the effectiveness of multi-scale deep registration with test-time training based on Dice coefficient for image segmentation and mean square error (MSE), normalized local cross-correlation (NLCC) for tissue dense tracking tasks.
Wentao Zhu 0001, Yufang Huang, Daguang Xu, Wei Fan 0001, Xiaohui Xie
ICRA3
2021 Improving Pneumonia Localization via Cross-Attention on Medical Images and Reports
Riddhish Bhalodia, Ali Hatamizadeh, Leo K. Tam, Ziyue Xu 0001, Xiaosong Wang 0001, Evrim Turkbey, Daguang Xu
MICCAI (2)7
2021 Accounting for Dependencies in Deep Learning Based Multiple Instance Learning for Whole Slide Imaging
Andriy Myronenko, Ziyue Xu 0001, Dong Yang 0005, Holger Roth, Daguang Xu
MICCAI (8)5
2021 The Power of Proxy Data and Proxy Networks for Hyper-parameter Optimization in Medical Image Segmentation
Vishwesh Nath, Dong Yang 0005, Ali Hatamizadeh, Anas A. Abidin, Andriy Myronenko, Holger Roth, Daguang Xu
MICCAI (3)7
2021 Federated Whole Prostate Segmentation in MRI with Personalized Neural Architectures
Holger Roth, Dong Yang 0005, Wenqi Li 0001, Andriy Myronenko, Wentao Zhu 0001, Ziyue Xu 0001, Xiaosong Wang 0001, Daguang Xu
MICCAI (3)8
2021 Federated learning improves site performance in multicenter deep learning without data sharing
abstract
OBJECTIVE: To demonstrate enabling multi-institutional training without centralizing or sharing the underlying physical data via federated learning (FL). MATERIALS AND METHODS: Deep learning models were trained at each participating institution using local clinical data, and an additional model was trained using FL across all of the institutions. RESULTS: We found that the FL model exhibited superior performance and generalizability to the models trained at single institutions, with an overall performance level that was significantly better than that of any of the institutional models alone when evaluated on held-out test sets from each institution and an outside challenge dataset. DISCUSSION: The power of FL was successfully demonstrated across 3 academic institutions while avoiding the privacy risk associated with the transfer and pooling of patient data. CONCLUSION: Federated learning is an effective methodology that merits further study to enable accelerated development of models across institutions, enabling greater generalizability in clinical use.
Karthik Sarma, Stephanie A. Harmon, Thomas Sanford, Holger Roth, Ziyue Xu 0001, Jesse Tetreault, Daguang Xu, Mona Flores, Alex G. Raman, Rushikesh Kulkarni, Bradford J. Wood, Peter L. Choyke, Alan Priester, Leonard S. Marks, Steven S. Raman, Dieter R. Enzmann, Baris Turkbey, William Speier, Corey W. Arnold
J. Am. Medical Informatics Assoc.7
2021 Federated semi-supervised learning for COVID region segmentation in chest CT using multi-national data from China, Italy, Japan
Dong Yang 0005, Ziyue Xu 0001, Wenqi Li 0001, Andriy Myronenko, Holger Roth, Stephanie A. Harmon, Sheng Xu 0001, Baris Turkbey, Evrim Turkbey, Xiaosong Wang 0001, Wentao Zhu 0001, Gianpaolo Carrafiello, Francesca Patella, Maurizio Cariati, Hirofumi Obinata, Hitoshi Mori, Kaku Tamura, Peng An 0002, Bradford J. Wood, Daguang Xu
Medical Image Anal.20
2021 VerSe: A Vertebrae labelling and segmentation benchmark for multi-detector CT images
Anjany Sekuboyina, Malek El Husseini, Amirhossein Bayat, Maximilian Löffler, Hans Liebl, Hongwei Li 0004, Giles Tetteh, Jan Kukacka, Christian Payer, Darko Stern, Martin Urschler, Maodong Chen, Dalong Cheng, Nikolas Leßmann, Yujin Hu, Tianfu Wang 0001, Dong Yang 0005, Daguang Xu, Felix Ambellan, Tamaz Amiranashvili, Moritz Ehlke, Hans Lamecker, Sebastian Lehnert, Marilia Lirio, Nicolás Pérez de Olaguer, Heiko Ramm, Manish Sahu, Alexander Tack, Stefan Zachow, Xinjun Ma, Christoph Angerman, Xin Wang 0113, Alexandre Kirszenberg, Élodie Puybareau, Yiwei Bai, Brandon H. Rapazzo, Timyoas Yeah, Amber Zhang, Shangliang Xu, Feng Hou, Zhiqiang He 0002, Chan Zeng, Zheng Xiangshang, Xu Liming, Tucker J. Netherton, Raymond P. Mumme, Laurence E. Court, Zixun Huang, Chenhang He, Li-Wen Wang, Sai-Ho Ling, Lê Duy Huynh, Nicolas Boutry, Roman Jakubícek, Jirí Chmelík, Supriti Mulay, Mohanasankar Sivaprakasam, Johannes C. Paetzold, Suprosanna Shit, Ivan Ezhov, Benedikt Wiestler, Ben Glocker, Alexander Valentinitsch, Markus Rempfler, Bjoern Menze, Jan Kirschke
Medical Image Anal.18
2021 Diminishing Uncertainty Within the Training Pool: Active Learning for Medical Image Segmentation
abstract
Active learning is a unique abstraction of machine learning techniques where the model/algorithm could guide users for annotation of a set of data points that would be beneficial to the model, unlike passive machine learning. The primary advantage being that active learning frameworks select data points that can accelerate the learning process of a model and can reduce the amount of data needed to achieve full accuracy as compared to a model trained on a randomly acquired data set. Multiple frameworks for active learning combined with deep learning have been proposed, and the majority of them are dedicated to classification tasks. Herein, we explore active learning for the task of segmentation of medical imaging data sets. We investigate our proposed framework using two datasets: 1.) MRI scans of the hippocampus, 2.) CT scans of pancreas and tumors. This work presents a query-by-committee approach for active learning where a joint optimizer is used for the committee. At the same time, we propose three new strategies for active learning: 1.) increasing frequency of uncertain data to bias the training data set; 2.) Using mutual information among the input images as a regularizer for acquisition to ensure diversity in the training dataset; 3.) adaptation of Dice log-likelihood for Stein variational gradient descent (SVGD). The results indicate an improvement in terms of data reduction by achieving full accuracy while only using 22.69% and 48.85% of the available data for each dataset, respectively.
Vishwesh Nath, Dong Yang 0005, Bennett A. Landman, Daguang Xu, Holger Roth
IEEE Trans. Medical Imaging4
2021 Multi-Domain Image Completion for Random Missing Input Data
abstract
Multi-domain data are widely leveraged in vision applications taking advantage of complementary information from different modalities, e.g., brain tumor segmentation from multi-parametric magnetic resonance imaging (MRI). However, due to possible data corruption and different imaging protocols, the availability of images for each domain could vary amongst multiple data sources in practice, which makes it challenging to build a universal model with a varied set of input data. To tackle this problem, we propose a general approach to complete the random missing domain(s) data in real applications. Specifically, we develop a novel multi-domain image completion method that utilizes a generative adversarial network (GAN) with a representational disentanglement scheme to extract shared content encoding and separate style encoding across multiple domains. We further illustrate that the learned representation in multi-domain image completion could be leveraged for high-level tasks, e.g., segmentation, by introducing a unified framework consisting of image completion and segmentation with a shared content encoder. The experiments demonstrate consistent performance improvement on three datasets for brain tumor segmentation, prostate segmentation, and facial expression image completion respectively.
Liyue Shen, Wentao Zhu 0001, Xiaosong Wang 0001, Lei Xing 0001, John M. Pauly, Baris Turkbey, Stephanie A. Harmon, Thomas Sanford, Sherif Mehralivand, Peter L. Choyke, Bradford J. Wood, Daguang Xu
IEEE Trans. Medical Imaging12
2020 When Radiology Report Generation Meets Knowledge Graph
abstract
Automatic radiology report generation has been an attracting research problem towards computer-aided diagnosis to alleviate the workload of doctors in recent years. Deep learning techniques for natural image captioning are successfully adapted to generating radiology reports. However, radiology image reporting is different from the natural image captioning task in two aspects: 1) the accuracy of positive disease keyword mentions is critical in radiology image reporting in comparison to the equivalent importance of every single word in a natural image caption; 2) the evaluation of reporting quality should focus more on matching the disease keywords and their associated attributes instead of counting the occurrence of N-gram. Based on these concerns, we propose to utilize a pre-constructed graph embedding module (modeled with a graph convolutional neural network) on multiple disease findings to assist the generation of reports in this work. The incorporation of knowledge graph allows for dedicated feature learning for each disease finding and the relationship modeling between them. In addition, we proposed a new evaluation metric for radiology image reporting with the assistance of the same composed graph. Experimental results demonstrate the superior performance of the methods integrated with the proposed graph embedding module on a publicly accessible dataset (IU-RR) of chest radiographs compared with previous approaches using both the conventional evaluation metrics commonly adopted for image captioning and our proposed ones.
Yixiao Zhang 0001, Xiaosong Wang 0001, Ziyue Xu 0001, Qihang Yu, Alan L. Yuille, Daguang Xu
AAAI6
2020 C2FNAS: Coarse-to-Fine Neural Architecture Search for 3D Medical Image Segmentation
abstract
3D convolution neural networks (CNN) have been proved very successful in parsing organs or tumours in 3D medical images, but it remains sophisticated and time-consuming to choose or design proper 3D networks given different task contexts. Recently, Neural Architecture Search (NAS) is proposed to solve this problem by searching for the best network architecture automatically. However, the inconsistency between search stage and deployment stage often exists in NAS algorithms due to memory constraints and large search space, which could become more serious when applying NAS to some memory and time-consuming tasks, such as 3D medical image segmentation. In this paper, we propose a coarse-to-fine neural architecture search (C2FNAS) to automatically search a 3D segmentation network from scratch without inconsistency on network size or input size. Specifically, we divide the search procedure into two stages: 1) the coarse stage, where we search the macro-level topology of the network, i.e. how each convolution module is connected to other modules; 2) the fine stage, where we search at micro-level for operations in each cell based on previous searched macro-level topology. The coarse-to-fine manner divides the search procedure into two consecutive stages and meanwhile resolves the inconsistency. We evaluate our method on 10 public datasets from Medical Segmentation Decalthon (MSD) challenge, and achieve state-of-the-art performance with the network searched using one dataset, which demonstrates the effectiveness and generalization of our searched models.
Qihang Yu, Dong Yang 0005, Holger Roth, Yutong Bai, Yixiao Zhang 0001, Alan L. Yuille, Daguang Xu
CVPR7
2020 Weakly Supervised One-Stage Vision and Language Disease Detection Using Large Scale Pneumonia and Pneumothorax Studies
Leo K. Tam, Xiaosong Wang 0001, Evrim Turkbey, Yuhong Wen, Daguang Xu
MICCAI (4)6
2020 LAMP: Large Deep Nets with Automated Model Parallelism for Image Segmentation
Wentao Zhu 0001, Can Zhao 0001, Wenqi Li 0001, Holger Roth, Ziyue Xu 0001, Daguang Xu
MICCAI (4)6
2020 3D Semi-Supervised Learning with Uncertainty-Aware Multi-View Co-Training
abstract
While making a tremendous impact in various fields, deep neural networks usually require large amounts of labeled data for training which are expensive to collect in many applications, especially in the medical domain. Un-labeled data, on the other hand, is much more abundant. Semi-supervised learning techniques, such as co-training, could provide a powerful tool to leverage unlabeled data. In this paper, we propose a novel framework, uncertainty-aware multi-view co-training (UMCT), to address semi-supervised learning on 3D data, such as volumetric data from medical imaging. In our work, co-training is achieved by exploiting multi-viewpoint consistency of 3D data. We generate different views by rotating or permuting the 3D data and utilize asymmetrical 3D kernels to encourage diversified features in different sub-networks. In addition, we propose an uncertainty-weighted label fusion mechanism to estimate the reliability of each view's prediction with Bayesian deep learning. As one view requires the supervision from other views in co-training, our self-adaptive approach computes a confidence score for the prediction of each unlabeled sample in order to assign a reliable pseudo label. Thus, our approach can take advantage of unlabeled data during training. We show the effectiveness of our proposed semi-supervised method on several public datasets from medical image segmentation tasks (NIH pancreas & LiTS liver tumor dataset). Meanwhile, a fully-supervised method based on our approach achieved state-of-the-art performances on both the LiTS liver tumor segmentation and the Medical Segmentation Decathlon (MSD) challenge, demonstrating the robustness and value of our framework, even when fully supervised training is feasible.
Yingda Xia, Fengze Liu, Dong Yang 0005, Jinzheng Cai, Lequan Yu, Zhuotun Zhu, Daguang Xu, Alan L. Yuille, Holger Roth
WACV7
2020 NeurReg: Neural Registration and Its Application to Image Segmentation
abstract
Registration is a fundamental task in medical image analysis which can be applied to several tasks including image segmentation, intra-operative tracking, multi-modal image alignment, and motion analysis. Popular registration tools such as ANTs and NiftyReg optimize an objective function for each pair of images from scratch which is time-consuming for large images with complicated deformation. Facilitated by the rapid progress of deep learning, learning-based approaches such as VoxelMorph have been emerging for image registration. These approaches can achieve competitive performance in a fraction of a second on advanced GPUs. In this work, we construct a neural registration framework, called NeurReg, with a hybrid loss of displacement fields and data similarity, which substantially improves the current state-of-the-art of registrations. Within the framework, we simulate various transformations by a registration simulator which generates fixed image and displacement field ground truth for training. Furthermore, we design three segmentation frameworks based on the proposed registration framework: 1) atlas-based segmentation, 2) joint learning of both segmentation and registration tasks, and 3) multi-task learning with atlas-based segmentation as an intermediate feature. Extensive experimental results validate the effectiveness of the proposed NeurReg framework based on various metrics: the endpoint error (EPE) of the predicted displacement field, mean square error (MSE), normalized local cross-correlation (NLCC), mutual information (MI), Dice coefficient, uncertainty estimation, and the interpretability of the segmentation. The proposed NeurReg improves registration accuracy with fast inference speed, which can greatly accelerate related medical image analysis tasks.
Wentao Zhu 0001, Andriy Myronenko, Ziyue Xu 0001, Wenqi Li 0001, Holger Roth, Yufang Huang, Fausto Milletari, Daguang Xu
WACV8
2020 Deep hiearchical multi-label classification applied to chest X-ray abnormality taxonomies
Haomin Chen, Shun Miao, Daguang Xu, Gregory D. Hager, Adam P. Harrison
Medical Image Anal.3
2020 Uncertainty-aware multi-view co-training for semi-supervised medical image segmentation and domain adaptation
Yingda Xia, Dong Yang 0005, Zhiding Yu, Fengze Liu, Jinzheng Cai, Lequan Yu, Zhuotun Zhu, Daguang Xu, Alan L. Yuille, Holger Roth
Medical Image Anal.8
2020 Generalizing Deep Learning for Medical Image Segmentation to Unseen Domains via Deep Stacked Transformation
abstract
Recent advances in deep learning for medical image segmentation demonstrate expert-level accuracy. However, application of these models in clinically realistic environments can result in poor generalization and decreased accuracy, mainly due to the domain shift across different hospitals, scanner vendors, imaging protocols, and patient populations etc. Common transfer learning and domain adaptation techniques are proposed to address this bottleneck. However, these solutions require data (and annotations) from the target domain to retrain the model, and is therefore restrictive in practice for widespread model deployment. Ideally, we wish to have a trained (locked) model that can work uniformly well across unseen domains without further training. In this paper, we propose a deep stacked transformation approach for domain generalization. Specifically, a series of n stacked transformations are applied to each image during network training. The underlying assumption is that the "expected" domain shift for a specific medical imaging modality could be simulated by applying extensive data augmentation on a single source domain, and consequently, a deep model trained on the augmented "big" data (BigAug) could generalize well on unseen domains. We exploit four surprisingly effective, but previously understudied, image-based characteristics for data augmentation to overcome the domain generalization problem. We train and evaluate the BigAug model (with n=9 transformations) on three different 3D segmentation tasks (prostate gland, left atrial, left ventricle) covering two medical imaging modalities (MRI and ultrasound) involving eight publicly available challenge datasets. The results show that when training on relatively small dataset (n = 10~32 volumes, depending on the size of the available datasets) from a single source domain: (i) BigAug models degrade an average of 11%(Dice score change) from source to unseen domain, substantially better than conventional augmentation (degrading 39%) and CycleGAN-based domain adaptation method (degrading 25%), (ii) BigAug is better than "shallower" stacked transforms (i.e. those with fewer transforms) on unseen domains and demonstrates modest improvement to conventional augmentation on the source domain, (iii) after training with BigAug on one source domain, performance on an unseen domain is similar to training a model from scratch on that domain when using the same number of training samples. When training on large datasets (n = 465 volumes) with BigAug, (iv) application to unseen domains reaches the performance of state-of-the-art fully supervised models that are trained and tested on their source domains. These findings establish a strong benchmark for the study of domain generalization in medical imaging, and can be generalized to the design of highly robust deep segmentation models for clinical deployment.
Ling Zhang 0002, Xiaosong Wang 0001, Dong Yang 0005, Thomas Sanford, Stephanie A. Harmon, Baris Turkbey, Bradford J. Wood, Holger Roth, Andriy Myronenko, Daguang Xu, Ziyue Xu 0001
IEEE Trans. Medical Imaging10
2019 V-NAS: Neural Architecture Search for Volumetric Medical Image Segmentation
abstract
Deep learning algorithms, in particular 2D and 3D fully convolutional neural networks (FCNs), have rapidly become the mainstream methodology for volumetric medical image segmentation. However, 2D convolutions cannot fully leverage the rich spatial information along the third axis, while 3D convolutions suffer from the demanding computation and high GPU memory consumption. In this paper, we propose to automatically search the network architecture tailoring to volumetric medical image segmentation problem. Concretely, we formulate the structure learning as differentiable neural architecture search, and let the network itself choose between 2D, 3D or Pseudo-3D (P3D) convolutions at each layer. We evaluate our method on 3 public datasets, i.e., the NIH Pancreas dataset, the Lung and Pancreas dataset from the Medical Segmentation Decathlon (MSD) Challenge. Our method, named V-NAS, consistently outperforms other state-of-the-arts on the segmentation tasks of both normal organ (NIH Pancreas) and abnormal organs (MSD Lung tumors and MSD Pancreas tumors), which shows the power of chosen architecture. Moreover, the searched architecture on one dataset can be well generalized to other datasets, which demonstrates the robustness and practical use of our proposed method.
Zhuotun Zhu, Chenxi Liu 0001, Dong Yang 0005, Alan L. Yuille, Daguang Xu
3DV5
2019 An Alarm System for Segmentation Algorithm Based on Shape Model
abstract
It is usually hard for a learning system to predict correctly on rare events that never occur in the training data, and there is no exception for segmentation algorithms. Meanwhile, manual inspection of each case to locate the failures becomes infeasible due to the trend of large data scale and limited human resource. Therefore, we build an alarm system that will set off alerts when the segmentation result is possibly unsatisfactory, assuming no corresponding ground truth mask is provided. One plausible solution is to project the segmentation results into a low dimensional feature space; then learn classifiers/regressors to predict their qualities. Motivated by this, in this paper, we learn a feature space using the shape information which is a strong prior shared among different datasets and robust to the appearance variation of input data. The shape feature is captured using a Variational Auto-Encoder (VAE) network that trained with only the ground truth masks. During testing, the segmentation results with bad shapes shall not fit the shape prior well, resulting in large loss values. Thus, the VAE is able to evaluate the quality of segmentation result on unseen data, without using ground truth. Finally, we learn a regressor in the one-dimensional feature space to predict the qualities of segmentation results. Our alarm system is evaluated on several recent state-of-art segmentation algorithms for 3D medical segmentation tasks. Compared with other standard quality assessment methods, our system consistently provides more reliable prediction on the qualities of segmentation results.
Fengze Liu, Yingda Xia, Dong Yang 0005, Alan L. Yuille, Daguang Xu
ICCV5
2019 Searching Learning Strategy with Reinforcement Learning for 3D Medical Image Segmentation
Dong Yang 0005, Holger Roth, Ziyue Xu 0001, Fausto Milletari, Ling Zhang 0002, Daguang Xu
MICCAI (2)6
2019 Integrating 3D Geometry of Organ for Improving Medical Image Segmentation
Jiawen Yao, Jinzheng Cai, Dong Yang 0005, Daguang Xu, Junzhou Huang
MICCAI (5)4
2018 3D Anisotropic Hybrid Network: Transferring Convolutional Features from 2D Images to 3D Anisotropic Volumes
Siqi Liu 0001, Daguang Xu, Shaohua Kevin Zhou, Olivier Pauly, Sasa Grbic, Thomas Mertelmeier, Julia Wicklein, Anna K. Jerebko, Tom Weidong Cai, Dorin Comaniciu
MICCAI (2)2
2017 Supervised Action Classifier: Approaching Landmark Detection as Image Partitioning
Zhoubing Xu, Qiangui Huang, Jin Hyeong Park, Mingqing Chen, Daguang Xu, Dong Yang 0005, David Liu 0001, Shaohua Kevin Zhou
MICCAI (3)5
2017 Deep Image-to-Image Recurrent Network with Shape Basis Learning for Automatic Vertebra Labeling in Large-Scale 3D CT Volumes
Dong Yang 0005, Daguang Xu, Shaohua Kevin Zhou, Zhoubing Xu, Mingqing Chen, Jin Hyeong Park, Sasa Grbic, Trac D. Tran, Sang (Peter) Chin, Dimitris N. Metaxas, Dorin Comaniciu
MICCAI (3)3
2017 Automatic Liver Segmentation Using an Adversarial Image-to-Image Network
Dong Yang 0005, Daguang Xu, Shaohua Kevin Zhou, Bogdan Georgescu, Mingqing Chen, Sasa Grbic, Dimitris N. Metaxas, Dorin Comaniciu
MICCAI (3)2
2013 Learning to translate with products of novices: a suite of open-ended challenge problems for teaching MT
abstract
Machine translation (MT) draws from several different disciplines, making it a complex subject to teach. There are excellent pedagogical texts, but problems in MT and current algorithms for solving them are best learned by doing. As a centerpiece of our MT course, we devised a series of open-ended challenges for students in which the goal was to improve performance on carefully constrained instances of four key MT tasks: alignment, decoding, evaluation, and reranking. Students brought a diverse set of techniques to the problems, including some novel solutions which performed remarkably well. A surprising and exciting outcome was that student solutions or their combinations fared competitively on some tasks, demonstrating that even newcomers to the field can help improve the state-of-the-art on hard NLP problems while simultaneously learning a great deal. The problems, baseline code, and results are freely available.
Adam Lopez, Matt Post, Chris Callison-Burch, Jonathan Weese, Juri Ganitkevitch, Narges Ahmidi, Olivia Buzek, Leah Hanson, Beaniesh Jamil, Matthias A. Lee, Ya-Ting Lin, Henry Pao, Fatima Rivera, Leili Shahriyari, Debu Sinha, Adam R. Teichert, Stephen Wampler, Michael Weinberger, Daguang Xu, Lin Yang 0002, Shang Zhao 0002
Trans. Assoc. Comput. Linguistics19