VLDB 2026 Research / reviewers in the wild / expert
Ziyue Xu 0001
dblp:59/9160-1
· DBLP profile ↗
43ranked-venue papers
6as first author
25since 2021 · last 2026
0000-0002-5728-6869ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 29 · 5 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Source-Resilient Joint Learning Framework for Preserving Stable Generalization on Diverse Ultrasonic Source ScenariosabstractJoint learning on diverse ultrasonic source scenarios presents a challenge in preserving stable gen-eralization due to the combination of heterogeneity of different sources and the inconsistency of joint learning features. Previous joint learning studies, which are not source-resilient frameworks, may not preserve stable generalization when trained on diverse source scenarios. Furthermore, the limited variations insingle-source data and the interference from ultrasound imaging, which are common in ultrasonic source scenarios, further decrease generalization. To address these problems, we pro posed a source-resilient joint learning framework consisting of three stages: 1) Source transforming, where our 1-to-N transformation unifies diverse source scenarios for source-resiliency. 2) Our feature enhancement modules model the source-resilient joint learning network, including a manifold-constraint normalization module (MCNM) for addressing heterogeneity by minimizing manifold-based loss, a task-consistent attention module (TCAM) shares the multi-scale features with self-attention to address inconsistency, and an adaptive feature-shifting module (AFSM) for feature-level augmentation to overcome single-source data.3) Our ultrasound-hybrid linear mapping (USmapping) cascades speckle randomization and mask-guiding Monge-Kantorovitch linear mapping to achieve ultrasonic style randomization for addressing the interference of ultrasonic data. Our framework was evaluated on eight ultrasound datasets from various scanners at multiple center sand surpassed previous comparable studies in both segmentation (DSCWAvgof 75.7%) and classification (AUROCWAvgof 68.8%) tasks. Our framework has the potential to serve as a general framework for enhancing the performance of joint learning under diverse ultrasonic source scenarios. Bin Huang 0021, Zhong Liu 0004, Ziyue Xu 0001, S. C. Chan 0001, Huiying Wen, Qicai Huang, Meiqin Jiang, Changfeng Dong, Ruhai Zou, Bingsheng Huang, Xin Chen 0025, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | VISTA3D: A Unified Segmentation Foundation Model For 3D Medical ImagingabstractFoundation models for interactive segmentation in 2D natural images and videos have sparked significant interest in building 3D foundation models for medical imaging. However, the domain gaps and clinical use cases for 3D medical imaging require a dedicated model that diverges from existing 2D solutions. Specifically, such foundation models should support a full workflow that can actually reduce human effort. Treating 3D medical images as sequences of 2D slices and reusing interactive 2D foundation models seems straightforward, but 2D annotation is too time-consuming for 3D tasks. Moreover, for large cohort analysis, it’s the highly accurate automatic segmentation models that reduce the most human effort. However, these models lack support for interactive corrections and lack zero-shot ability for novel structures, which is a key feature of "foundation". While reusing pre-trained 2D backbones in 3D enhances zero-shot potential, their performance on complex 3D structures still lags behind leading 3D models. To address these issues, we present VISTA3D, Versatile Imaging SegmenTation and Annotation model, that targets to solve all these challenges and requirements with one unified foundation model. VISTA3D is built on top of the well-established 3D segmentation pipeline, and it is the first model to achieve state-of-the-art performance in both 3D automatic (supporting 127 classes) and 3D interactive segmentation, even when compared with top 3D expert models on large and diverse benchmarks. Additionally, VISTA3D’s 3D interactive design allows efficient human correction, and a novel 3D supervoxel method that distills 2D pre-trained backbones grants VISTA3D top 3D zero-shot performance. We believe the model, recipe, and insights represent a promising step towards a clinically useful 3D foundation model. Code and weights are publicly available at https://github.com/Project-MONAI/VISTA. Yufan He, Yucheng Tang, Andriy Myronenko, Vishwesh Nath, Ziyue Xu 0001, Dong Yang 0005, Can Zhao 0001, Benjamin Simon, Mason Belue, Stephanie A. Harmon, Baris Turkbey, Daguang Xu, Wenqi Li 0001 |
CVPR | 6 |
| 2025 | VILA-M3: Enhancing Vision-Language Models with Medical Expert KnowledgeabstractGeneralist vision language models (VLMs) have made significant strides in computer vision, but they fall short in specialized fields like healthcare, where expert knowledge is essential. Current large multimodal models like Gemini and GPT-4o are insufficient for medical tasks due to their reliance on memorized internet knowledge rather than the nuanced expertise required in healthcare. Meanwhile, existing medical VLMs (e.g. Med-Gemini) often lack expert consultation as part of their design, and many rely on outdated, static datasets that were not created with modern, large deep learning models in mind. VLMs are usually trained in three stages: vision pre-training, vision-language pre-training, and instruction fine-tuning (IFT). IFT has been typically applied using a mixture of generic and healthcare data. In contrast, we propose that for medical VLMs, a fourth stage of specialized IFT is necessary, which focuses on medical data and includes information from domain expert models. Domain expert models developed for medical use are crucial because they are specifically trained for certain clinical tasks, e.g. to detect tumors and classify abnormalities through segmentation and classification, which learn fine-grained features of medical data−features that are often too intricate for a VLM to capture effectively. This paper introduces a new framework, VILA-M3, for medical VLMs that utilizes domain knowledge via expert models. We argue that generic VLM architectures alone are not viable for real-world clinical applications and on-demand usage of domain-specialized expert model knowledge is critical for advancing AI in healthcare. Through our experiments, we show an improved state-of-the-art (SOTA) performance with an average improvement of ~9% over the prior SOTA model Med-Gemini and ~6% over models trained on the specific tasks. Our approach emphasizes the importance of domain expertise in creating precise, reliable VLMs for medical applications. Vishwesh Nath, Wenqi Li 0001, Dong Yang 0005, Andriy Myronenko, Mingxin Zheng, Yao Lu 0006, Hongxu Yin, Yee Man Law, Yucheng Tang, Can Zhao 0001, Ziyue Xu 0001, Yufan He, Stephanie A. Harmon, Benjamin Simon, Greg Heinrich, Stephen R. Aylward, Marc Edgar, Michael Zephyr, Pavlo Molchanov 0001, Baris Turkbey, Holger Roth, Daguang Xu |
CVPR | 13 |
| 2025 | MAISI: Medical AI for Synthetic ImagingabstractMedical imaging analysis faces challenges such as data scarcity, high annotation costs, and privacy concerns. This paper introduces the Medical AI for Synthetic Imaging (MAISI), an innovative approach using the diffusion model to generate synthetic 3D computed tomography (CT) images to address those challenges. MAISI leverages the foundation volume compression network and the latent diffusion model to produce high-resolution CT images (up to a landmark volume dimension of 512 × 512 × 768) with flexible volume dimensions and voxel spacing. By incorporating ControlNet, MAISI can process organ segmentation, including 127 anatomical structures, as additional conditions and enables the generation of accurately annotated synthetic images that can be used for various downstream tasks. Our experiment results show that MAISI's capabilities in generating realistic, anatomically accurate images for diverse regions and conditions reveal its promising potential to mitigate challenges using synthetic data. Can Zhao 0001, Dong Yang 0005, Ziyue Xu 0001, Vishwesh Nath, Yucheng Tang, Benjamin Simon, Mason Belue, Stephanie A. Harmon, Baris Turkbey, Daguang Xu |
WACV | 4 |
| 2024 | FedBPT: Efficient Federated Black-box Prompt Tuning for Large Language ModelsabstractPre-trained language models (PLM) have revolutionized the NLP landscape, achieving stellar performances across diverse tasks. These models, while benefiting from vast training data, often require fine-tuning on specific data to cater to distinct downstream tasks. However, this data adaptation process has inherent security and privacy concerns, primarily when leveraging user-generated, device-residing data. Federated learning (FL) provides a solution, allowing collaborative model fine-tuning without centralized data collection. However, applying FL to finetune PLMs is hampered by challenges, including restricted model parameter access due to the high encapsulation, high computational requirements, and communication overheads. This paper introduces Federated Black-box Prompt Tuning (FedBPT), a framework designed to address these challenges. FedBPT allows the clients to treat the model as a black-box inference API. By focusing on training optimal prompts and utilizing gradient-free optimization methods, FedBPT reduces the number of exchanged variables, boosts communication efficiency, and minimizes computational and storage costs. Experiments highlight the framework’s ability to drastically cut communication and memory costs while maintaining competitive performance. Ultimately, FedBPT presents a promising solution for efficient, privacy-preserving fine-tuning of PLM in the age of large language models. Jingwei Sun 0002, Ziyue Xu 0001, Hongxu Yin, Dong Yang 0005, Daguang Xu, Zhixu Du, Yiran Chen 0001, Holger Roth |
ICML | 2 |
| 2024 | IR-FRestormer: Iterative Refinement with Fourier-Based Restormer for Accelerated MRI ReconstructionabstractAccelerated magnetic resonance imaging (MRI) aims to reconstruct high-quality MR images from a set of under-sampled measurements. State-of-the-art methods for this task use deep learning, which offers high reconstruction accuracy and fast runtimes. In this work, we propose a new state-of-the-art reconstruction model for accelerated MRI reconstruction. Our model is the first to combine the power of deep neural networks with iterative refinement for this task. For the neural network component of our method, we utilize a transformer-based architecture as transformers are state-of-the-art in various image reconstruction tasks. However, a major drawback of transformers which has limited their emergence among the state-of-the-art MRI models is that they are often memory inefficient for high-resolution inputs. To address this limitation, we propose a transformer-based model which uses parameter-free Fourier-based attention modules, achieving 2× more memory efficiency. We evaluate our model on the largest publicly available MRI dataset, the fastMRI dataset [46], and achieve on-par performance with other state-of-the-art1methods on the dataset’s leaderboard2. Mohammad Zalbagi Darestani, Vishwesh Nath, Wenqi Li 0001, Yufan He, Holger Roth, Ziyue Xu 0001, Daguang Xu, Reinhard Heckel, Can Zhao 0001 |
WACV | 6 |
| 2024 | Learning Quality Labels for Robust Image ClassificationabstractSupervised learning paradigms largely benefit from the tremendous amount of annotated data. However, the quality of the annotations often varies among labelers. Multi-observer studies have been conducted to examine the annotation variances (by labeling the same data multiple times) to see how it affects critical applications like medical image analysis. In this paper, we demonstrate how multiple sets of annotations (either hand-labeled or algorithm-generated) can be utilized together and mutually benefit the learning of classification tasks. A scheme of learning-to-vote is introduced to sample quality label sets for each data entry on-the-fly during the training. Specifically, a label-sampling module is designed to achieve refined labels (weighted sum of attended ones) that benefit the model learning the most through additional back-propagations. We apply the learning-to-vote scheme on the classification task of a synthetic noisy CIFAR-10 to prove the concept and then demonstrate superior results (3-5% increase on average in multiple disease classification AUCs) on the chest x-ray images from a hospital-scale dataset (MIMIC-CXR) and hand-labeled dataset (OpenI) in comparison to regular training paradigms. Xiaosong Wang 0001, Ziyue Xu 0001, Dong Yang 0005, Leo K. Tam, Holger Roth, Daguang Xu |
WACV | 2 |
| 2024 | An interpretable two-branch bi-coordinate network based on multi-grained domain knowledge for classification of thyroid nodules in ultrasound images
Ziyue Xu 0001, Weiwei Zhan, Jing Xiao 0006, Yiqing Hou, Bingsheng Huang, Lingyun Huang, Shuo Li 0001 |
Medical Image Anal. | 3 |
| 2024 | Improving Tumor Classification by Reusing Self-Predicted Segmentation of Medical Images as Guiding KnowledgeabstractDifferential diagnosis of tumors is important for computer-aided diagnosis. In computer-aided diagnosis systems, expert knowledge of lesion segmentation masks is limited as it is only used during preprocessing or as supervision to guide feature extraction. To improve the utilization of lesion segmentation masks, this study proposes a simple and effective multitask learning network that improves medical image classification using self-predicted segmentation as guiding knowledge; we call this network RS$^{2}$-net. In RS$^{2}$-net, the predicted segmentation probability map obtained from the initial segmentation inference is added to the original image to form a new input, which is then reinput to the network for the final classification inference. We validated the proposed RS$^{2}$-net using three datasets: the pNENs-Grade dataset, which tested the prediction of pancreatic neuroendocrine neoplasm grading, and the HCC-MVI dataset, which tested the prediction of microvascular invasion of hepatocellular carcinoma, and ISIC 2017 public skin lesion dataset. The experimental results indicate that the proposed strategy of reusing self-predicted segmentation is effective, and RS$^{2}$-net outperforms other popular networks and existing state-of-the-art studies. Interpretive analytics based on feature visualization demonstrates that the improved classification performance of our reuse strategy is due to the semantic information that can be acquired in advance in a shallow network. Xiaoyi Lin, Ziyue Xu 0001, Xin Chen 0025, Chenglang Yuan, Songxiong Wu, Yanji Luo, Jingxian Shen, Shi-Ting Feng, Bingsheng Huang |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Fair Federated Medical Image Segmentation via Client Contribution EstimationabstractHow to ensure fairness is an important topic in federated learning (FL). Recent studies have investigated how to reward clients based on their contribution (collaboration fairness), and how to achieve uniformity of performance across clients (performance fairness). Despite achieving progress on either one, we argue that it is critical to consider them together, in order to engage and motivate more diverse clients joining FL to derive a high-quality global model. In this work, we propose a novel method to optimize both types of fairness simultaneously. Specifically, we propose to estimate client contribution in gradient and data space. In gradient space, we monitor the gradient direction differences of each client with respect to others. And in data space, we measure the prediction error on client data using an auxiliary model. Based on this contribution estimation, we propose a FL method, federated training via contribution estimation (FedCE), i.e., using estimation as global model aggregation weights. We have theoretically analyzed our method and empirically evaluated it on two real-world medical datasets. The effectiveness of our approach has been validated with significant performance improvements, better collaboration fairness, better performance fairness, and comprehensive analytical studies. Code is available at https://nvidia.github.io/NVFlare/research/fed-ce Meirui Jiang, Holger Roth, Wenqi Li 0001, Dong Yang 0005, Can Zhao 0001, Vishwesh Nath, Daguang Xu, Qi Dou 0001, Ziyue Xu 0001 |
CVPR | 9 |
| 2023 | Communication-Efficient Vertical Federated Learning with Limited Overlapping SamplesabstractFederated learning is a popular collaborative learning approach that enables clients to train a global model without sharing their local data. Vertical federated learning (VFL) deals with scenarios in which the data on clients have different feature spaces but share some overlapping samples. Existing VFL approaches suffer from high communication costs and cannot deal efficiently with limited overlapping samples commonly seen in the real world. We propose a practical VFL framework called one-shot VFL that can solve the communication bottleneck and the problem of limited overlapping samples simultaneously based on semi-supervised learning. We also propose few-shot VFL to improve the accuracy further with just one more communication round between the server and the clients. In our proposed framework, the clients only need to communicate with the server once or only a few times. We evaluate the proposed VFL framework on both image and tabular datasets. Our methods can improve the accuracy by more than 46.5% and reduce the communication cost by more than 330× compared with state-of-the-art VFL methods when evaluated on CIFAR-10. Our code is available at https://nvidia.github.io/NVFlare/research/one-shot-vfl. Jingwei Sun 0002, Ziyue Xu 0001, Dong Yang 0005, Vishwesh Nath, Wenqi Li 0001, Can Zhao 0001, Daguang Xu, Yiran Chen 0001, Holger Roth |
ICCV | 2 |
| 2023 | A Style Transfer-Based Augmentation Framework for Improving Segmentation and Classification Performance Across Different Sources in Ultrasound Images
Bin Huang 0021, Ziyue Xu 0001, S. C. Chan 0001, Zhong Liu 0004, Huiying Wen, Qicai Huang, Meiqin Jiang, Changfeng Dong, Ruhai Zou, Bingsheng Huang, Xin Chen 0025, Shuo Li 0001 |
MICCAI (6) | 2 |
| 2023 | DAST: Differentiable Architecture Search with Transformer for 3D Medical Image Segmentation
Dong Yang 0005, Ziyue Xu 0001, Yufan He, Vishwesh Nath, Wenqi Li 0001, Andriy Myronenko, Ali Hatamizadeh, Can Zhao 0001, Holger Roth, Daguang Xu |
MICCAI (3) | 2 |
| 2022 | Closing the Generalization Gap of Cross-silo Federated Medical Image SegmentationabstractCross-silo federated learning (FL) has attracted much attention in medical imaging analysis with deep learning in recent years as it can resolve the critical issues of insufficient data, data privacy, and training efficiency. However, there can be a generalization gap between the model trained from FL and the one from centralized training. This important issue comes from the non-iid data distribution of the local data in the participating clients and is well-known as client drift. In this work, we propose a novel training frame-work FedSM to avoid the client drift issue and successfully close the generalization gap compared with the centralized training for medical image segmentation tasks for the first time. We also propose a novel personalized FL objective formulation and a new method SoftPull to solve it in our proposed framework FedSM. We conduct rigorous theoretical analysis to guarantee its convergence for optimizing the non-convex smooth objective function. Real-world medical image segmentation experiments using deep FL validate the motivations and effectiveness of our proposed method. An Xu, Wenqi Li 0001, Dong Yang 0005, Holger Roth, Ali Hatamizadeh, Can Zhao 0001, Daguang Xu, Heng Huang 0001, Ziyue Xu 0001 |
CVPR | 10 |
| 2022 | Auto-FedRL: Federated Hyperparameter Optimization for Multi-institutional Medical Image Segmentation
Dong Yang 0005, Ali Hatamizadeh, An Xu, Ziyue Xu 0001, Wenqi Li 0001, Can Zhao 0001, Daguang Xu, Stephanie A. Harmon, Evrim Turkbey, Baris Turkbey, Bradford J. Wood, Francesca Patella, Elvira Stellato, Gianpaolo Carrafiello, Vishal M. Patel, Holger Roth |
ECCV (21) | 5 |
| 2022 | Clinical-Realistic Annotation for Histopathology Images with Probabilistic Semi-supervision: A Worst-Case Study
Ziyue Xu 0001, Andriy Myronenko, Dong Yang 0005, Holger Roth, Can Zhao 0001, Xiaosong Wang 0001, Daguang Xu |
MICCAI (2) | 1 |
| 2022 | Rapid artificial intelligence solutions in a pandemic - The COVID-19-20 Lung CT Lesion Segmentation Challenge
Holger Roth, Ziyue Xu 0001, Carlos Tor-Díez, Ramon Sánchez-Jacob, Jonathan Zember, Jose Molto, Wenqi Li 0001, Sheng Xu 0001, Baris Turkbey, Evrim Turkbey, Dong Yang 0005, Ahmed Harouni, Nicola Rieke, Shishuai Hu, Fabian Isensee, Claire Tang, Qinji Yu, Jan Sölter, Vitali Liauchuk, Jan Hendrik Moltz, Bruno Oliveira 0002, Yong Xia 0001, Klaus H. Maier-Hein, Qikai Li, Andreas Husch, Vassili Kovalev, Alessa Hering, João L. Vilaça, Mona Flores, Daguang Xu, Bradford J. Wood, Marius George Linguraru |
Medical Image Anal. | 2 |
| 2022 | 3D Lightweight Network for Simultaneous Registration and Segmentation of Organs-at-Risk in CT Images of Head and Neck CancerabstractImage-guided radiation therapy (IGRT) is the most effective treatment for head and neck cancer. The successful implementation of IGRT requires accurate delineation of organ-at-risk (OAR) in the computed tomography (CT) images. In routine clinical practice, OARs are manually segmented by oncologists, which is time-consuming, laborious, and subjective. To assist oncologists in OAR contouring, we proposed a three-dimensional (3D) lightweight framework for simultaneous OAR registration and segmentation. The registration network was designed to align a selected OAR template to a new image volume for OAR localization. A region of interest (ROI) selection layer then generated ROIs of OARs from the registration results, which were fed into a multiview segmentation network for accurate OAR segmentation. To improve the performance of registration and segmentation networks, a centre distance loss was designed for the registration network, an ROI classification branch was employed for the segmentation network, and further, context information was incorporated to iteratively promote both networks' performance. The segmentation results were further refined with shape information for final delineation. We evaluated registration and segmentation performances of the proposed framework using three datasets. On the internal dataset, the Dice similarity coefficient (DSC) of registration and segmentation was 69.7% and 79.6%, respectively. In addition, our framework was evaluated on two external datasets and gained satisfactory performance. These results showed that the 3D lightweight framework achieved fast, accurate and robust registration and segmentation of OARs in head and neck cancer. The proposed framework has the potential of assisting oncologists in OAR delineation. Bin Huang 0021, Yufeng Ye, Ziyue Xu 0001, Zongyou Cai, Zhangnan Zhong, Lingxiang Liu, Xin Chen 0025, Hanwei Chen, Bingsheng Huang |
IEEE Trans. Medical Imaging | 3 |
| 2021 | T-AutoML: Automated Machine Learning for Lesion Segmentation using Transformers in 3D Medical ImagingabstractLesion segmentation in medical imaging has been an important topic in clinical research. Researchers have proposed various detection and segmentation algorithms to address this task. Recently, deep learning-based approaches have significantly improved the performance over conventional methods. However, most state-of-the-art deep learning methods require the manual design of multiple network components and training strategies. In this paper, we propose a new automated machine learning algorithm, T-AutoML, which not only searches for the best neural architecture, but also finds the best combination of hyper-parameters and data augmentation strategies simultaneously. The proposed method utilizes the modern transformer model, which is introduced to adapt to the dynamic length of the search space embedding and can significantly improve the ability of the search. We validate T-AutoML on several large-scale public lesion segmentation data-sets and achieve state-of-the-art performance. Dong Yang 0005, Andriy Myronenko, Xiaosong Wang 0001, Ziyue Xu 0001, Holger Roth, Daguang Xu |
ICCV | 4 |
| 2021 | Improving Pneumonia Localization via Cross-Attention on Medical Images and Reports
Riddhish Bhalodia, Ali Hatamizadeh, Leo K. Tam, Ziyue Xu 0001, Xiaosong Wang 0001, Evrim Turkbey, Daguang Xu |
MICCAI (2) | 4 |
| 2021 | Accounting for Dependencies in Deep Learning Based Multiple Instance Learning for Whole Slide Imaging
Andriy Myronenko, Ziyue Xu 0001, Dong Yang 0005, Holger Roth, Daguang Xu |
MICCAI (8) | 2 |
| 2021 | Federated Whole Prostate Segmentation in MRI with Personalized Neural Architectures
Holger Roth, Dong Yang 0005, Wenqi Li 0001, Andriy Myronenko, Wentao Zhu 0001, Ziyue Xu 0001, Xiaosong Wang 0001, Daguang Xu |
MICCAI (3) | 6 |
| 2021 | Federated learning improves site performance in multicenter deep learning without data sharingabstractOBJECTIVE: To demonstrate enabling multi-institutional training without centralizing or sharing the underlying physical data via federated learning (FL). MATERIALS AND METHODS: Deep learning models were trained at each participating institution using local clinical data, and an additional model was trained using FL across all of the institutions. RESULTS: We found that the FL model exhibited superior performance and generalizability to the models trained at single institutions, with an overall performance level that was significantly better than that of any of the institutional models alone when evaluated on held-out test sets from each institution and an outside challenge dataset. DISCUSSION: The power of FL was successfully demonstrated across 3 academic institutions while avoiding the privacy risk associated with the transfer and pooling of patient data. CONCLUSION: Federated learning is an effective methodology that merits further study to enable accelerated development of models across institutions, enabling greater generalizability in clinical use. Karthik Sarma, Stephanie A. Harmon, Thomas Sanford, Holger Roth, Ziyue Xu 0001, Jesse Tetreault, Daguang Xu, Mona Flores, Alex G. Raman, Rushikesh Kulkarni, Bradford J. Wood, Peter L. Choyke, Alan Priester, Leonard S. Marks, Steven S. Raman, Dieter R. Enzmann, Baris Turkbey, William Speier, Corey W. Arnold |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | Federated semi-supervised learning for COVID region segmentation in chest CT using multi-national data from China, Italy, Japan
Dong Yang 0005, Ziyue Xu 0001, Wenqi Li 0001, Andriy Myronenko, Holger Roth, Stephanie A. Harmon, Sheng Xu 0001, Baris Turkbey, Evrim Turkbey, Xiaosong Wang 0001, Wentao Zhu 0001, Gianpaolo Carrafiello, Francesca Patella, Maurizio Cariati, Hirofumi Obinata, Hitoshi Mori, Kaku Tamura, Peng An 0002, Bradford J. Wood, Daguang Xu |
Medical Image Anal. | 2 |
| 2021 | Capsules for biomedical image segmentation
Rodney LaLonde, Ziyue Xu 0001, Ismail Irmakci, Sanjay Jain 0002, Ulas Bagci |
Medical Image Anal. | 2 |
| 2020 | When Radiology Report Generation Meets Knowledge GraphabstractAutomatic radiology report generation has been an attracting research problem towards computer-aided diagnosis to alleviate the workload of doctors in recent years. Deep learning techniques for natural image captioning are successfully adapted to generating radiology reports. However, radiology image reporting is different from the natural image captioning task in two aspects: 1) the accuracy of positive disease keyword mentions is critical in radiology image reporting in comparison to the equivalent importance of every single word in a natural image caption; 2) the evaluation of reporting quality should focus more on matching the disease keywords and their associated attributes instead of counting the occurrence of N-gram. Based on these concerns, we propose to utilize a pre-constructed graph embedding module (modeled with a graph convolutional neural network) on multiple disease findings to assist the generation of reports in this work. The incorporation of knowledge graph allows for dedicated feature learning for each disease finding and the relationship modeling between them. In addition, we proposed a new evaluation metric for radiology image reporting with the assistance of the same composed graph. Experimental results demonstrate the superior performance of the methods integrated with the proposed graph embedding module on a publicly accessible dataset (IU-RR) of chest radiographs compared with previous approaches using both the conventional evaluation metrics commonly adopted for image captioning and our proposed ones. Yixiao Zhang 0001, Xiaosong Wang 0001, Ziyue Xu 0001, Qihang Yu, Alan L. Yuille, Daguang Xu |
AAAI | 3 |
| 2020 | LAMP: Large Deep Nets with Automated Model Parallelism for Image Segmentation
Wentao Zhu 0001, Can Zhao 0001, Wenqi Li 0001, Holger Roth, Ziyue Xu 0001, Daguang Xu |
MICCAI (4) | 5 |
| 2020 | NeurReg: Neural Registration and Its Application to Image SegmentationabstractRegistration is a fundamental task in medical image analysis which can be applied to several tasks including image segmentation, intra-operative tracking, multi-modal image alignment, and motion analysis. Popular registration tools such as ANTs and NiftyReg optimize an objective function for each pair of images from scratch which is time-consuming for large images with complicated deformation. Facilitated by the rapid progress of deep learning, learning-based approaches such as VoxelMorph have been emerging for image registration. These approaches can achieve competitive performance in a fraction of a second on advanced GPUs. In this work, we construct a neural registration framework, called NeurReg, with a hybrid loss of displacement fields and data similarity, which substantially improves the current state-of-the-art of registrations. Within the framework, we simulate various transformations by a registration simulator which generates fixed image and displacement field ground truth for training. Furthermore, we design three segmentation frameworks based on the proposed registration framework: 1) atlas-based segmentation, 2) joint learning of both segmentation and registration tasks, and 3) multi-task learning with atlas-based segmentation as an intermediate feature. Extensive experimental results validate the effectiveness of the proposed NeurReg framework based on various metrics: the endpoint error (EPE) of the predicted displacement field, mean square error (MSE), normalized local cross-correlation (NLCC), mutual information (MI), Dice coefficient, uncertainty estimation, and the interpretability of the segmentation. The proposed NeurReg improves registration accuracy with fast inference speed, which can greatly accelerate related medical image analysis tasks. Wentao Zhu 0001, Andriy Myronenko, Ziyue Xu 0001, Wenqi Li 0001, Holger Roth, Yufang Huang, Fausto Milletari, Daguang Xu |
WACV | 3 |
| 2020 | Generalizing Deep Learning for Medical Image Segmentation to Unseen Domains via Deep Stacked TransformationabstractRecent advances in deep learning for medical image segmentation demonstrate expert-level accuracy. However, application of these models in clinically realistic environments can result in poor generalization and decreased accuracy, mainly due to the domain shift across different hospitals, scanner vendors, imaging protocols, and patient populations etc. Common transfer learning and domain adaptation techniques are proposed to address this bottleneck. However, these solutions require data (and annotations) from the target domain to retrain the model, and is therefore restrictive in practice for widespread model deployment. Ideally, we wish to have a trained (locked) model that can work uniformly well across unseen domains without further training. In this paper, we propose a deep stacked transformation approach for domain generalization. Specifically, a series of n stacked transformations are applied to each image during network training. The underlying assumption is that the "expected" domain shift for a specific medical imaging modality could be simulated by applying extensive data augmentation on a single source domain, and consequently, a deep model trained on the augmented "big" data (BigAug) could generalize well on unseen domains. We exploit four surprisingly effective, but previously understudied, image-based characteristics for data augmentation to overcome the domain generalization problem. We train and evaluate the BigAug model (with n=9 transformations) on three different 3D segmentation tasks (prostate gland, left atrial, left ventricle) covering two medical imaging modalities (MRI and ultrasound) involving eight publicly available challenge datasets. The results show that when training on relatively small dataset (n = 10~32 volumes, depending on the size of the available datasets) from a single source domain: (i) BigAug models degrade an average of 11%(Dice score change) from source to unseen domain, substantially better than conventional augmentation (degrading 39%) and CycleGAN-based domain adaptation method (degrading 25%), (ii) BigAug is better than "shallower" stacked transforms (i.e. those with fewer transforms) on unseen domains and demonstrates modest improvement to conventional augmentation on the source domain, (iii) after training with BigAug on one source domain, performance on an unseen domain is similar to training a model from scratch on that domain when using the same number of training samples. When training on large datasets (n = 465 volumes) with BigAug, (iv) application to unseen domains reaches the performance of state-of-the-art fully supervised models that are trained and tested on their source domains. These findings establish a strong benchmark for the study of domain generalization in medical imaging, and can be generalized to the design of highly robust deep segmentation models for clinical deployment. Ling Zhang 0002, Xiaosong Wang 0001, Dong Yang 0005, Thomas Sanford, Stephanie A. Harmon, Baris Turkbey, Bradford J. Wood, Holger Roth, Andriy Myronenko, Daguang Xu, Ziyue Xu 0001 |
IEEE Trans. Medical Imaging | 11 |
| 2019 | Searching Learning Strategy with Reinforcement Learning for 3D Medical Image Segmentation
Dong Yang 0005, Holger Roth, Ziyue Xu 0001, Fausto Milletari, Ling Zhang 0002, Daguang Xu |
MICCAI (2) | 3 |
| 2018 | CT-Realistic Lung Nodule Simulation from 3D Conditional Generative Adversarial Networks for Robust Lung Segmentation
Dakai Jin, Ziyue Xu 0001, Youbao Tang, Adam P. Harrison, Daniel J. Mollura |
MICCAI (2) | 2 |
| 2018 | Joint solution for PET image segmentation, denoising, and partial volume correction
Ziyue Xu 0001, Mingchen Gao, Georgios Z. Papadakis, Brian Luna, Sanjay Jain 0002, Daniel J. Mollura, Ulas Bagci |
Medical Image Anal. | 1 |
| 2017 | Progressive and Multi-path Holistically Nested Neural Networks for Pathological Lung Segmentation from CT Images
Adam P. Harrison, Ziyue Xu 0001, Kevin George, Le Lu 0001, Ronald M. Summers, Daniel J. Mollura |
MICCAI (3) | 2 |
| 2016 | Characterization of Lung Nodule Malignancy Using Hybrid Shape and Appearance Features
Mario Buty, Ziyue Xu 0001, Mingchen Gao, Ulas Bagci, Aaron Wu, Daniel J. Mollura |
MICCAI (1) | 2 |
| 2016 | Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer LearningabstractRemarkable progress has been made in image recognition, primarily due to the availability of large-scale annotated datasets and deep convolutional neural networks (CNNs). CNNs enable learning data-driven, highly representative, hierarchical image features from sufficient training data. However, obtaining datasets as comprehensively annotated as ImageNet in the medical imaging domain remains a challenge. There are currently three major techniques that successfully employ CNNs to medical image classification: training the CNN from scratch, using off-the-shelf pre-trained CNN features, and conducting unsupervised CNN pre-training with supervised fine-tuning. Another effective method is transfer learning, i.e., fine-tuning CNN models pre-trained from natural image dataset to medical image tasks. In this paper, we exploit three important, but previously understudied factors of employing deep convolutional neural networks to computer-aided detection problems. We first explore and evaluate different CNN architectures. The studied models contain 5 thousand to 160 million parameters, and vary in numbers of layers. We then evaluate the influence of dataset scale and spatial image context on performance. Finally, we examine when and why transfer learning from pre-trained ImageNet (via fine-tuning) can be useful. We study two specific computer-aided detection (CADe) problems, namely thoraco-abdominal lymph node (LN) detection and interstitial lung disease (ILD) classification. We achieve the state-of-the-art performance on the mediastinal LN detection, and report the first five-fold cross-validation classification results on predicting axial CT slices with ILD categories. Our extensive empirical evaluation, CNN model analysis and valuable insights can be extended to the design of high performance CAD systems for other medical imaging tasks. Hoo-Chang Shin, Holger Roth, Mingchen Gao, Le Lu 0001, Ziyue Xu 0001, Isabella Nogues, Jianhua Yao 0001, Daniel J. Mollura, Ronald M. Summers |
IEEE Trans. Medical Imaging | 5 |
| 2015 | A hybrid method for airway segmentation and automated measurement of bronchial wall thickness on CT
Ziyue Xu 0001, Ulas Bagci, Brent Foster, Awais Mansoor, Jayaram K. Udupa, Daniel J. Mollura |
Medical Image Anal. | 1 |
| 2015 | Correction to "A Generic Approach to Pathological Lung Segmentation"abstractIn the above paper (ibid., vol. 33, no. 12, pp. 2293-2310, Dec. 2014), Awais Mansoor was incorrectly indicated as the corresponding author. Ulas Bagci should have been indicated as the corresponding author. Awais Mansoor, Ulas Bagci, Ziyue Xu 0001, Brent Foster, Kenneth N. Olivier, Jason M. Elinoff, Anthony F. Suffredini, Jayaram K. Udupa, Daniel J. Mollura |
IEEE Trans. Medical Imaging | 3 |
| 2014 | Segmentation Based Denoising of PET Images: An Iterative Approach via Regional Means and Affinity Propagation
Ziyue Xu 0001, Ulas Bagci, Jurgen Seidel, David Thomasson, Jeffrey M. Solomon, Daniel J. Mollura |
MICCAI (1) | 1 |
| 2014 | A Generic Approach to Pathological Lung SegmentationabstractIn this study, we propose a novel pathological lung segmentation method that takes into account neighbor prior constraints and a novel pathology recognition system. Our proposed framework has two stages; during stage one, we adapted the fuzzy connectedness (FC) image segmentation algorithm to perform initial lung parenchyma extraction. In parallel, we estimate the lung volume using rib-cage information without explicitly delineating lungs. This rudimentary, but intelligent lung volume estimation system allows comparison of volume differences between rib cage and FC based lung volume measurements. Significant volume difference indicates the presence of pathology, which invokes the second stage of the proposed framework for the refinement of segmented lung. In stage two, texture-based features are utilized to detect abnormal imaging patterns (consolidations, ground glass, interstitial thickening, tree-inbud, honeycombing, nodules, and micro-nodules) that might have been missed during the first stage of the algorithm. This refinement stage is further completed by a novel neighboring anatomy-guided segmentation approach to include abnormalities with weak textures, and pleura regions. We evaluated the accuracy and efficiency of the proposed method on more than 400 CT scans with the presence of a wide spectrum of abnormalities. To our best of knowledge, this is the first study to evaluate all abnormal imaging patterns in a single segmentation framework. The quantitative results show that our pathological lung segmentation method improves on current standards because of its high sensitivity and specificity and may have considerable potential to enhance the performance of routine clinical tasks. Awais Mansoor, Ulas Bagci, Ziyue Xu 0001, Brent Foster, Kenneth N. Olivier, Jason M. Elinoff, Anthony F. Suffredini, Jayaram K. Udupa, Daniel J. Mollura |
IEEE Trans. Medical Imaging | 3 |
| 2013 | Spatially Constrained Random Walk Approach for Accurate Estimation of Airway Wall Surfaces
Ziyue Xu 0001, Ulas Bagci, Brent Foster, Awais Mansoor, Daniel J. Mollura |
MICCAI (2) | 1 |
| 2013 | Joint segmentation of anatomical and functional images: Applications in quantification of lesions from PET, PET-CT, MRI-PET, and MRI-PET-CT images
Ulas Bagci, Jayaram K. Udupa, Neil Mendhiratta, Brent Foster, Ziyue Xu 0001, Jianhua Yao 0001, Xinjian Chen 0001, Daniel J. Mollura |
Medical Image Anal. | 5 |
| 2012 | Quantitative Characterization of Trabecular Bone Micro-architecture Using Tensor Scale and Multi-Detector CT Imaging
Yinxiao Liu, Punam K. Saha, Ziyue Xu 0001 |
MICCAI (1) | 3 |
| 2012 | Tensor scale: An analytic approach with efficient computation and applications
Ziyue Xu 0001, Punam K. Saha, Soura Dasgupta |
Comput. Vis. Image Underst. | 1 |