VLDB 2026 Research / reviewers in the wild / expert
Can Zhao 0001
dblp:35/2787-1
· DBLP profile ↗
18ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0001-7286-3452ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MAISI-v2: Accelerated 3D High-Resolution Medical Image Synthesis with Rectified Flow and Region-specific Contrastive LossabstractMedical image synthesis is an important topic for both clinical and research applications. Recently, diffusion models have become a leading approach in this area. Despite their strengths, many existing methods struggle with (1) limited generalizability, only working for specific body regions or voxel spacings, (2) slow inference, which is a common issue for diffusion models, and (3) weak alignment with input conditions, which is a critical issue for medical imaging. MAISI, a previously proposed framework, addresses generalizability issues but still suffers from slow inference and limited condition consistency. In this work, we present MAISI-v2, the first accelerated 3D medical image synthesis framework that integrates rectified flow to enable fast and high-quality generation. To further enhance condition fidelity, we introduce a novel region-specific contrastive loss to improve sensitivity to the region of interest. Our experiments show that MAISI-v2 can achieve state-of-the-art image quality with 33× acceleration for latent diffusion models. We also conducted a downstream segmentation experiment to show that the synthetic images can be used for data augmentation. We release our code, training details, model weights, and a GUI demo to facilitate reproducibility and promote further development within the community. Can Zhao 0001, Dong Yang 0005, Yufan He, Yucheng Tang, Benjamin Simon, Mason Belue, Stephanie A. Harmon, Baris Turkbey, Daguang Xu |
AAAI | 1 |
| 2026 | Text-Driven Tumor SynthesisabstractTumor synthesis can generate challenging cases that AI often misses or over-detects. Training on these cases improves AI performance. However, most existing synthesis methods are either unconditional- generating images from random variables-or conditioned only on tumor shape. As a result, they lack control over clinically important tumor characteristics, such as texture, heterogeneity, boundary, and pathology. The generated tumors are therefore overly similar or duplicates of existing training cases, failing to effectively address AI's weaknesses. We propose a new text-driven tumor synthesis approach, termed TextoMorph, that provides textual control over tumor characteristics in conjunction with mask control. This approach is particularly beneficial for examples that confuse the AI the most, such as early tumor detection (improving Sensitivity by + 6.5%), tumor segmentation for precise radiotherapy (improving NSD by + 3.1%), and classification between benign and malignant tumors (improving Sensitivity by + 8.2%). By incorporating text mined from radiology reports into the synthesis process, we increase the variability and controllability of the synthetic tumors to target AI's failure cases more precisely. Moreover, TextoMorph uses contrastive learning across different texts and CT scans, significantly reducing dependence on scarce image-report pairs (only 141 pairs used in this study) by leveraging a large corpus of 34,035 radiology reports. Finally, we have developed rigorous tests to evaluate synthetic tumors, showing that our synthetic tumors is realistic and diverse in texture, heterogeneity, boundary, and pathology. Code and models are available at https://github.com/MrGiovanni/TextoMorph. Yi Shuai, Qi Chen 0014, Dong Yang 0005, Can Zhao 0001, Pedro R. A. S. Bassi, Daguang Xu, Kang Wang 0016, Yang Yang 0009, Alan L. Yuille, Zongwei Zhou |
IEEE Trans. Medical Imaging | 8 |
| 2025 | VISTA3D: A Unified Segmentation Foundation Model For 3D Medical ImagingabstractFoundation models for interactive segmentation in 2D natural images and videos have sparked significant interest in building 3D foundation models for medical imaging. However, the domain gaps and clinical use cases for 3D medical imaging require a dedicated model that diverges from existing 2D solutions. Specifically, such foundation models should support a full workflow that can actually reduce human effort. Treating 3D medical images as sequences of 2D slices and reusing interactive 2D foundation models seems straightforward, but 2D annotation is too time-consuming for 3D tasks. Moreover, for large cohort analysis, it’s the highly accurate automatic segmentation models that reduce the most human effort. However, these models lack support for interactive corrections and lack zero-shot ability for novel structures, which is a key feature of "foundation". While reusing pre-trained 2D backbones in 3D enhances zero-shot potential, their performance on complex 3D structures still lags behind leading 3D models. To address these issues, we present VISTA3D, Versatile Imaging SegmenTation and Annotation model, that targets to solve all these challenges and requirements with one unified foundation model. VISTA3D is built on top of the well-established 3D segmentation pipeline, and it is the first model to achieve state-of-the-art performance in both 3D automatic (supporting 127 classes) and 3D interactive segmentation, even when compared with top 3D expert models on large and diverse benchmarks. Additionally, VISTA3D’s 3D interactive design allows efficient human correction, and a novel 3D supervoxel method that distills 2D pre-trained backbones grants VISTA3D top 3D zero-shot performance. We believe the model, recipe, and insights represent a promising step towards a clinically useful 3D foundation model. Code and weights are publicly available at https://github.com/Project-MONAI/VISTA. Yufan He, Yucheng Tang, Andriy Myronenko, Vishwesh Nath, Ziyue Xu 0001, Dong Yang 0005, Can Zhao 0001, Benjamin Simon, Mason Belue, Stephanie A. Harmon, Baris Turkbey, Daguang Xu, Wenqi Li 0001 |
CVPR | 8 |
| 2025 | VILA-M3: Enhancing Vision-Language Models with Medical Expert KnowledgeabstractGeneralist vision language models (VLMs) have made significant strides in computer vision, but they fall short in specialized fields like healthcare, where expert knowledge is essential. Current large multimodal models like Gemini and GPT-4o are insufficient for medical tasks due to their reliance on memorized internet knowledge rather than the nuanced expertise required in healthcare. Meanwhile, existing medical VLMs (e.g. Med-Gemini) often lack expert consultation as part of their design, and many rely on outdated, static datasets that were not created with modern, large deep learning models in mind. VLMs are usually trained in three stages: vision pre-training, vision-language pre-training, and instruction fine-tuning (IFT). IFT has been typically applied using a mixture of generic and healthcare data. In contrast, we propose that for medical VLMs, a fourth stage of specialized IFT is necessary, which focuses on medical data and includes information from domain expert models. Domain expert models developed for medical use are crucial because they are specifically trained for certain clinical tasks, e.g. to detect tumors and classify abnormalities through segmentation and classification, which learn fine-grained features of medical data−features that are often too intricate for a VLM to capture effectively. This paper introduces a new framework, VILA-M3, for medical VLMs that utilizes domain knowledge via expert models. We argue that generic VLM architectures alone are not viable for real-world clinical applications and on-demand usage of domain-specialized expert model knowledge is critical for advancing AI in healthcare. Through our experiments, we show an improved state-of-the-art (SOTA) performance with an average improvement of ~9% over the prior SOTA model Med-Gemini and ~6% over models trained on the specific tasks. Our approach emphasizes the importance of domain expertise in creating precise, reliable VLMs for medical applications. Vishwesh Nath, Wenqi Li 0001, Dong Yang 0005, Andriy Myronenko, Mingxin Zheng, Yao Lu 0006, Hongxu Yin, Yee Man Law, Yucheng Tang, Can Zhao 0001, Ziyue Xu 0001, Yufan He, Stephanie A. Harmon, Benjamin Simon, Greg Heinrich, Stephen R. Aylward, Marc Edgar, Michael Zephyr, Pavlo Molchanov 0001, Baris Turkbey, Holger Roth, Daguang Xu |
CVPR | 12 |
| 2025 | MAISI: Medical AI for Synthetic ImagingabstractMedical imaging analysis faces challenges such as data scarcity, high annotation costs, and privacy concerns. This paper introduces the Medical AI for Synthetic Imaging (MAISI), an innovative approach using the diffusion model to generate synthetic 3D computed tomography (CT) images to address those challenges. MAISI leverages the foundation volume compression network and the latent diffusion model to produce high-resolution CT images (up to a landmark volume dimension of 512 × 512 × 768) with flexible volume dimensions and voxel spacing. By incorporating ControlNet, MAISI can process organ segmentation, including 127 anatomical structures, as additional conditions and enables the generation of accurately annotated synthetic images that can be used for various downstream tasks. Our experiment results show that MAISI's capabilities in generating realistic, anatomically accurate images for diverse regions and conditions reveal its promising potential to mitigate challenges using synthetic data. Can Zhao 0001, Dong Yang 0005, Ziyue Xu 0001, Vishwesh Nath, Yucheng Tang, Benjamin Simon, Mason Belue, Stephanie A. Harmon, Baris Turkbey, Daguang Xu |
WACV | 2 |
| 2024 | Super-Field MRI Synthesis for Infant Brains Enhanced by Dual Channel Latent Diffusion
Austin Tapp, Can Zhao 0001, Holger Roth, Jeffrey Tanedo, Syed Muhammad Anwar, Niall J. Bourke, Joseph V. Hajnal, Victoria Nankabirwa, Sean C. L. Deoni, Natasha Leporé, Marius George Linguraru |
MICCAI (3) | 2 |
| 2024 | IR-FRestormer: Iterative Refinement with Fourier-Based Restormer for Accelerated MRI ReconstructionabstractAccelerated magnetic resonance imaging (MRI) aims to reconstruct high-quality MR images from a set of under-sampled measurements. State-of-the-art methods for this task use deep learning, which offers high reconstruction accuracy and fast runtimes. In this work, we propose a new state-of-the-art reconstruction model for accelerated MRI reconstruction. Our model is the first to combine the power of deep neural networks with iterative refinement for this task. For the neural network component of our method, we utilize a transformer-based architecture as transformers are state-of-the-art in various image reconstruction tasks. However, a major drawback of transformers which has limited their emergence among the state-of-the-art MRI models is that they are often memory inefficient for high-resolution inputs. To address this limitation, we propose a transformer-based model which uses parameter-free Fourier-based attention modules, achieving 2× more memory efficiency. We evaluate our model on the largest publicly available MRI dataset, the fastMRI dataset [46], and achieve on-par performance with other state-of-the-art1methods on the dataset’s leaderboard2. Mohammad Zalbagi Darestani, Vishwesh Nath, Wenqi Li 0001, Yufan He, Holger Roth, Ziyue Xu 0001, Daguang Xu, Reinhard Heckel, Can Zhao 0001 |
WACV | 9 |
| 2023 | Fair Federated Medical Image Segmentation via Client Contribution EstimationabstractHow to ensure fairness is an important topic in federated learning (FL). Recent studies have investigated how to reward clients based on their contribution (collaboration fairness), and how to achieve uniformity of performance across clients (performance fairness). Despite achieving progress on either one, we argue that it is critical to consider them together, in order to engage and motivate more diverse clients joining FL to derive a high-quality global model. In this work, we propose a novel method to optimize both types of fairness simultaneously. Specifically, we propose to estimate client contribution in gradient and data space. In gradient space, we monitor the gradient direction differences of each client with respect to others. And in data space, we measure the prediction error on client data using an auxiliary model. Based on this contribution estimation, we propose a FL method, federated training via contribution estimation (FedCE), i.e., using estimation as global model aggregation weights. We have theoretically analyzed our method and empirically evaluated it on two real-world medical datasets. The effectiveness of our approach has been validated with significant performance improvements, better collaboration fairness, better performance fairness, and comprehensive analytical studies. Code is available at https://nvidia.github.io/NVFlare/research/fed-ce Meirui Jiang, Holger Roth, Wenqi Li 0001, Dong Yang 0005, Can Zhao 0001, Vishwesh Nath, Daguang Xu, Qi Dou 0001, Ziyue Xu 0001 |
CVPR | 5 |
| 2023 | Communication-Efficient Vertical Federated Learning with Limited Overlapping SamplesabstractFederated learning is a popular collaborative learning approach that enables clients to train a global model without sharing their local data. Vertical federated learning (VFL) deals with scenarios in which the data on clients have different feature spaces but share some overlapping samples. Existing VFL approaches suffer from high communication costs and cannot deal efficiently with limited overlapping samples commonly seen in the real world. We propose a practical VFL framework called one-shot VFL that can solve the communication bottleneck and the problem of limited overlapping samples simultaneously based on semi-supervised learning. We also propose few-shot VFL to improve the accuracy further with just one more communication round between the server and the clients. In our proposed framework, the clients only need to communicate with the server once or only a few times. We evaluate the proposed VFL framework on both image and tabular datasets. Our methods can improve the accuracy by more than 46.5% and reduce the communication cost by more than 330× compared with state-of-the-art VFL methods when evaluated on CIFAR-10. Our code is available at https://nvidia.github.io/NVFlare/research/one-shot-vfl. Jingwei Sun 0002, Ziyue Xu 0001, Dong Yang 0005, Vishwesh Nath, Wenqi Li 0001, Can Zhao 0001, Daguang Xu, Yiran Chen 0001, Holger Roth |
ICCV | 6 |
| 2023 | DAST: Differentiable Architecture Search with Transformer for 3D Medical Image Segmentation
Dong Yang 0005, Ziyue Xu 0001, Yufan He, Vishwesh Nath, Wenqi Li 0001, Andriy Myronenko, Ali Hatamizadeh, Can Zhao 0001, Holger Roth, Daguang Xu |
MICCAI (3) | 8 |
| 2022 | Closing the Generalization Gap of Cross-silo Federated Medical Image SegmentationabstractCross-silo federated learning (FL) has attracted much attention in medical imaging analysis with deep learning in recent years as it can resolve the critical issues of insufficient data, data privacy, and training efficiency. However, there can be a generalization gap between the model trained from FL and the one from centralized training. This important issue comes from the non-iid data distribution of the local data in the participating clients and is well-known as client drift. In this work, we propose a novel training frame-work FedSM to avoid the client drift issue and successfully close the generalization gap compared with the centralized training for medical image segmentation tasks for the first time. We also propose a novel personalized FL objective formulation and a new method SoftPull to solve it in our proposed framework FedSM. We conduct rigorous theoretical analysis to guarantee its convergence for optimizing the non-convex smooth objective function. Real-world medical image segmentation experiments using deep FL validate the motivations and effectiveness of our proposed method. An Xu, Wenqi Li 0001, Dong Yang 0005, Holger Roth, Ali Hatamizadeh, Can Zhao 0001, Daguang Xu, Heng Huang 0001, Ziyue Xu 0001 |
CVPR | 7 |
| 2022 | Auto-FedRL: Federated Hyperparameter Optimization for Multi-institutional Medical Image Segmentation
Dong Yang 0005, Ali Hatamizadeh, An Xu, Ziyue Xu 0001, Wenqi Li 0001, Can Zhao 0001, Daguang Xu, Stephanie A. Harmon, Evrim Turkbey, Baris Turkbey, Bradford J. Wood, Francesca Patella, Elvira Stellato, Gianpaolo Carrafiello, Vishal M. Patel, Holger Roth |
ECCV (21) | 7 |
| 2022 | Clinical-Realistic Annotation for Histopathology Images with Probabilistic Semi-supervision: A Worst-Case Study
Ziyue Xu 0001, Andriy Myronenko, Dong Yang 0005, Holger Roth, Can Zhao 0001, Xiaosong Wang 0001, Daguang Xu |
MICCAI (2) | 5 |
| 2021 | DiNTS: Differentiable Neural Network Topology Search for 3D Medical Image SegmentationabstractRecently, neural architecture search (NAS) has been applied to automatically search high-performance networks for medical image segmentation. The NAS search space usually contains a network topology level (controlling connections among cells with different spatial scales) and a cell level (operations within each cell). Existing methods either require long searching time for large-scale 3D image datasets, or are limited to pre-defined topologies (such as U-shaped or single-path) . In this work, we focus on three important aspects of NAS in 3D medical image segmentation: flexible multi-path network topology, high search efficiency, and budgeted GPU memory usage. A novel differentiable search framework is proposed to support fast gradient-based search within a highly flexible network topology search space. The discretization of the searched optimal continuous model in differentiable scheme may produce a sub-optimal final discrete model (discretization gap). Therefore, we propose a topology loss to alleviate this problem. In addition, the GPU memory usage for the searched 3D model is limited with budget constraints during search. Our Differentiable Network Topology Search scheme (DiNTS) is evaluated on the Medical Segmentation Decathlon (MSD) challenge, which contains ten challenging segmentation tasks. Our method achieves the state-of-the-art performance and the top ranking on the MSD challenge leaderboard. Yufan He, Dong Yang 0005, Holger Roth, Can Zhao 0001, Daguang Xu |
CVPR | 4 |
| 2021 | SMORE: A Self-Supervised Anti-Aliasing and Super-Resolution Algorithm for MRI Using Deep LearningabstractHigh resolution magnetic resonance (MR) images are desired in many clinical and research applications. Acquiring such images with high signal-to-noise (SNR), however, can require a long scan duration, which is difficult for patient comfort, is more costly, and makes the images susceptible to motion artifacts. A very common practical compromise for both 2D and 3D MR imaging protocols is to acquire volumetric MR images with high in-plane resolution, but lower through-plane resolution. In addition to having poor resolution in one orientation, 2D MRI acquisitions will also have aliasing artifacts, which further degrade the appearance of these images. This paper presents an approach SMORE1 based on convolutional neural networks (CNNs) that restores image quality by improving resolution and reducing aliasing in MR images.2 This approach is self-supervised, which requires no external training data because the high-resolution and low-resolution data that are present in the image itself are used for training. For 3D MRI, the method consists of only one self-supervised super-resolution (SSR) deep CNN that is trained from the volumetric image data. For 2D MRI, there is a self-supervised anti-aliasing (SAA) deep CNN that precedes the SSR CNN, also trained from the volumetric image data. Both methods were evaluated on a broad collection of MR data, including filtered and downsampled images so that quantitative metrics could be computed and compared, and actual acquired low resolution images for which visual and sharpness measures could be computed and compared. The super-resolution method is shown to be visually and quantitatively superior to previously reported methods. Can Zhao 0001, Blake Dewey, Dzung L. Pham, Peter A. Calabresi, Daniel S. Reich, Jerry L. Prince |
IEEE Trans. Medical Imaging | 1 |
| 2020 | LAMP: Large Deep Nets with Automated Model Parallelism for Image Segmentation
Wentao Zhu 0001, Can Zhao 0001, Wenqi Li 0001, Holger Roth, Ziyue Xu 0001, Daguang Xu |
MICCAI (4) | 2 |
| 2020 | Unsupervised MR-to-CT Synthesis Using Structure-Constrained CycleGANabstractSynthesizing a CT image from an available MR image has recently emerged as a key goal in radiotherapy treatment planning for cancer patients. CycleGANs have achieved promising results on unsupervised MR-to-CT image synthesis; however, because they have no direct constraints between input and synthetic images, cycleGANs do not guarantee structural consistency between these two images. This means that anatomical geometry can be shifted in the synthetic CT images, clearly a highly undesirable outcome in the given application. In this paper, we propose a structure-constrained cycleGAN for unsupervised MR-to-CT synthesis by defining an extra structure-consistency loss based on the modality independent neighborhood descriptor. We also utilize a spectral normalization technique to stabilize the training process and a self-attention module to model the long-range spatial dependencies in the synthetic images. Results on unpaired brain and abdomen MR-to-CT image synthesis show that our method produces better synthetic CT images in both accuracy and visual quality as compared to other unsupervised synthesis methods. We also show that an approximate affine pre-registration for unpaired training data can improve synthesis results. Heran Yang, Jian Sun 0009, Aaron Carass, Can Zhao 0001, Jerry L. Prince, Zongben Xu |
IEEE Trans. Medical Imaging | 4 |
| 2018 | A Deep Learning Based Anti-aliasing Self Super-Resolution Algorithm for MRI
Can Zhao 0001, Aaron Carass, Blake Dewey, Jonghye Woo, Jiwon Oh, Peter A. Calabresi, Daniel S. Reich, Pascal Sati, Dzung L. Pham, Jerry L. Prince |
MICCAI (1) | 1 |