EDBT 2026 Demo / reviewers in the wild / expert
Yucheng Tang
dblp:201/0160
· DBLP profile ↗
38ranked-venue papers
10as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 6 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 17 since 2021Artificial intelligence and machine learning · 17 · 4 first-author · 13 since 2021Systems, architecture and hardware · 3 · 3 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MAISI-v2: Accelerated 3D High-Resolution Medical Image Synthesis with Rectified Flow and Region-specific Contrastive LossabstractMedical image synthesis is an important topic for both clinical and research applications. Recently, diffusion models have become a leading approach in this area. Despite their strengths, many existing methods struggle with (1) limited generalizability, only working for specific body regions or voxel spacings, (2) slow inference, which is a common issue for diffusion models, and (3) weak alignment with input conditions, which is a critical issue for medical imaging. MAISI, a previously proposed framework, addresses generalizability issues but still suffers from slow inference and limited condition consistency. In this work, we present MAISI-v2, the first accelerated 3D medical image synthesis framework that integrates rectified flow to enable fast and high-quality generation. To further enhance condition fidelity, we introduce a novel region-specific contrastive loss to improve sensitivity to the region of interest. Our experiments show that MAISI-v2 can achieve state-of-the-art image quality with 33× acceleration for latent diffusion models. We also conducted a downstream segmentation experiment to show that the synthetic images can be used for data augmentation. We release our code, training details, model weights, and a GUI demo to facilitate reproducibility and promote further development within the community. Can Zhao 0001, Dong Yang 0005, Yufan He, Yucheng Tang, Benjamin Simon, Mason Belue, Stephanie A. Harmon, Baris Turkbey, Daguang Xu |
AAAI | 5 |
| 2026 | A comprehensive survey of computer vision methods for spatial transcriptomicsabstractSpatial transcriptomics (ST) enables the simultaneous measurement of gene expression and spatial localization within tissue sections, providing unprecedented opportunities to dissect tissue architecture and functional organization. As a relatively new omics technology, bioinformatics has driven much of the innovation in ST. However, within these frameworks, spatial information is often reduced to locations and relationships between molecular profiles, without fully leveraging the wealth of sub-micron morphological detail and histological knowledge available. Advances in computer vision-based artificial intelligence (AI) are opening exciting new avenues beyond conventional bioinformatics approaches by modeling complex histological patterns and linking morphology to molecular states. More excitingly, they bring fresh perspectives to potentially address key limitations of ST, including its high cost, limited clinical applicability, and reliance on 2D analysis of inherently 3D tissues. For instance, models that predict ST directly from histology images enable virtual sequencing, drastically reducing costs while integrating morphological insights from pathology with molecular biomarkers, thus accelerating clinical translation. Moreover, computer vision techniques can reconstruct pixel-aligned 3D tissue models, overcoming the technical barriers of 2D acquisition and advancing 3D spatial omics analytics. In this paper, we present the first systematic survey of computer vision AI models for ST analytics, categorizing approaches across architectures, learning paradigms, tasks, and datasets, and tracing their technological evolution. We highlight key challenges and future directions, offering a panoramic perspective on vision-driven ST and its potential to transform both basic research and clinical practice. The curated collection of vision-driven ST papers is available at https://github.com/hrlblab/computer_vision_spatial_omics. Junchao Zhu, Ruining Deng, Junlin Guo, Tianyuan Yao, Siqi Lu, Chongyu Qu, Juming Xiong, Yanfan Zhu, Zhengyi Lu, Yuechen Yang, Marilyn Lionts, Yucheng Tang, Daguang Xu, Shilin Zhao, Haichun Yang, Yuankai Huo |
Briefings Bioinform. | 12 |
| 2026 | SEQUAL: Self-refining and effective querying active learning with pseudo label divergence score for carotid intima-media segmentation in ultrasound
Yucheng Tang, Yipeng Hu, Chao Rong, Ciyuan Feng, Xiuzhen Yang, Hongxiang Lin |
Medical Image Anal. | 1 |
| 2025 | VISTA3D: A Unified Segmentation Foundation Model For 3D Medical ImagingabstractFoundation models for interactive segmentation in 2D natural images and videos have sparked significant interest in building 3D foundation models for medical imaging. However, the domain gaps and clinical use cases for 3D medical imaging require a dedicated model that diverges from existing 2D solutions. Specifically, such foundation models should support a full workflow that can actually reduce human effort. Treating 3D medical images as sequences of 2D slices and reusing interactive 2D foundation models seems straightforward, but 2D annotation is too time-consuming for 3D tasks. Moreover, for large cohort analysis, it’s the highly accurate automatic segmentation models that reduce the most human effort. However, these models lack support for interactive corrections and lack zero-shot ability for novel structures, which is a key feature of "foundation". While reusing pre-trained 2D backbones in 3D enhances zero-shot potential, their performance on complex 3D structures still lags behind leading 3D models. To address these issues, we present VISTA3D, Versatile Imaging SegmenTation and Annotation model, that targets to solve all these challenges and requirements with one unified foundation model. VISTA3D is built on top of the well-established 3D segmentation pipeline, and it is the first model to achieve state-of-the-art performance in both 3D automatic (supporting 127 classes) and 3D interactive segmentation, even when compared with top 3D expert models on large and diverse benchmarks. Additionally, VISTA3D’s 3D interactive design allows efficient human correction, and a novel 3D supervoxel method that distills 2D pre-trained backbones grants VISTA3D top 3D zero-shot performance. We believe the model, recipe, and insights represent a promising step towards a clinically useful 3D foundation model. Code and weights are publicly available at https://github.com/Project-MONAI/VISTA. Yufan He, Yucheng Tang, Andriy Myronenko, Vishwesh Nath, Ziyue Xu 0001, Dong Yang 0005, Can Zhao 0001, Benjamin Simon, Mason Belue, Stephanie A. Harmon, Baris Turkbey, Daguang Xu, Wenqi Li 0001 |
CVPR | 3 |
| 2025 | VILA-M3: Enhancing Vision-Language Models with Medical Expert KnowledgeabstractGeneralist vision language models (VLMs) have made significant strides in computer vision, but they fall short in specialized fields like healthcare, where expert knowledge is essential. Current large multimodal models like Gemini and GPT-4o are insufficient for medical tasks due to their reliance on memorized internet knowledge rather than the nuanced expertise required in healthcare. Meanwhile, existing medical VLMs (e.g. Med-Gemini) often lack expert consultation as part of their design, and many rely on outdated, static datasets that were not created with modern, large deep learning models in mind. VLMs are usually trained in three stages: vision pre-training, vision-language pre-training, and instruction fine-tuning (IFT). IFT has been typically applied using a mixture of generic and healthcare data. In contrast, we propose that for medical VLMs, a fourth stage of specialized IFT is necessary, which focuses on medical data and includes information from domain expert models. Domain expert models developed for medical use are crucial because they are specifically trained for certain clinical tasks, e.g. to detect tumors and classify abnormalities through segmentation and classification, which learn fine-grained features of medical data−features that are often too intricate for a VLM to capture effectively. This paper introduces a new framework, VILA-M3, for medical VLMs that utilizes domain knowledge via expert models. We argue that generic VLM architectures alone are not viable for real-world clinical applications and on-demand usage of domain-specialized expert model knowledge is critical for advancing AI in healthcare. Through our experiments, we show an improved state-of-the-art (SOTA) performance with an average improvement of ~9% over the prior SOTA model Med-Gemini and ~6% over models trained on the specific tasks. Our approach emphasizes the importance of domain expertise in creating precise, reliable VLMs for medical applications. Vishwesh Nath, Wenqi Li 0001, Dong Yang 0005, Andriy Myronenko, Mingxin Zheng, Yao Lu 0006, Hongxu Yin, Yee Man Law, Yucheng Tang, Can Zhao 0001, Ziyue Xu 0001, Yufan He, Stephanie A. Harmon, Benjamin Simon, Greg Heinrich, Stephen R. Aylward, Marc Edgar, Michael Zephyr, Pavlo Molchanov 0001, Baris Turkbey, Holger Roth, Daguang Xu |
CVPR | 10 |
| 2025 | MedSegFactory: Text-Guided Generation of Medical Image-Mask PairsabstractThis paper presents MedSegFactory, a versatile medical synthesis framework that generates high-quality paired medical images and segmentation masks across modalities and tasks. It aims to serve as an unlimited data repository, supplying image-mask pairs to enhance existing segmentation tools. The core of MedSegFactory is a dual-stream diffusion model, where one stream synthesizes medical images and the other generates corresponding segmentation masks. To ensure precise alignment between image-mask pairs, we introduce Joint Cross-Attention (JCA), enabling a collaborative denoising paradigm by dynamic cross-conditioning between streams. This bidirectional interaction allows both representations to guide each other's generation, enhancing consistency between generated pairs. MedSegFactory unlocks on-demand generation of paired medical images and segmentation masks through user-defined prompts that specify the target labels, imaging modalities, anatomical regions, and pathological conditions, facilitating scalable and high-quality data generation. This new paradigm of medical image synthesis enables seamless integration into diverse medical imaging workflows, enhancing both efficiency and accuracy. Extensive experiments show that MedSegFactory generates data of superior quality and usability, achieving competitive or state-of-the-art performance in 2D and 3D segmentation tasks while addressing data scarcity and regulatory constraints. Yuhan Wang 0001, Yucheng Tang, Daguang Xu, Kang Wang 0016, Yang Yang 0009, Zongwei Zhou, Yuyin Zhou |
ICCV | 3 |
| 2025 | ETA-IK: Execution-Time-Aware Inverse Kinematics for Dual-Arm SystemsabstractThis paper presents ETA-IK, a novel Execution-Time-Aware Inverse Kinematics method tailored for dual-arm robotic systems. The primary goal is to optimize motion execution time by leveraging the redundancy of the entire system, specifically in tasks where only the relative pose of the robots is constrained, such as dual-arm scanning of unknown objects. Unlike traditional IK methods using surrogate metrics, our approach directly optimizes execution time while implicitly considering collisions. A neural network based execution time approximator is employed to predict time-efficient joint configurations while accounting for potential collisions. Through experimental evaluation on a system composed of a UR5 and a KUKA iiwa robot, we demonstrate significant reductions in execution time. The proposed method outperforms conventional approaches, showing improved motion efficiency without sacrificing positioning accuracy. Yucheng Tang, Xi Huang 0005, Yongzhou Zhang, Ilshat Mamaev, Björn Hein |
IROS | 1 |
| 2025 | Analysis of Image-and-Text Uncertainty Propagation in Multimodal Large Language Models with Cardiac MR-Based Applications
Yucheng Tang, Yunguan Fu, Weixi Yi, Daniel C. Alexander, Rhodri H. Davies, Yipeng Hu |
MICCAI (4) | 1 |
| 2025 | PanTS: The Pancreatic Tumor Segmentation DatasetabstractPanTS is a large-scale, multi-institutional dataset curated to advance research in pancreatic CT analysis. It contains 36,390 CT scans from 145 medical centers, with expert-validated, voxel-wise annotations of over 993,000 anatomical structures, covering pancreatic tumors, pancreas head, body, and tail, and 24 surrounding anatomical structures such as vascular/skeletal structures and abdominal/thoracic organs. Each scan includes metadata such as patient age, sex, diagnosis, contrast phase, in-plane spacing, slice thickness, etc. AI models trained on PanTS achieve significantly better performance in pancreatic tumor detection, localization, and segmentation than those trained on existing public datasets. Our analysis indicates that these gains are directly attributable to the 16× larger-scale tumor annotations and indirectly supported by the 24 additional surrounding anatomical structures. As the largest and most comprehensive resource of its kind, PanTS offers a new benchmark for developing and evaluating AI models in pancreatic CT analysis. Xinze Zhou, Qi Chen 0014, Pedro R. A. S. Bassi, Xiaoxi Chen, Zheren Zhu, Kang Wang 0016, Yang Yang 0009, Yucheng Tang, Daguang Xu, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 13 |
| 2025 | BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation LearningabstractWe present the B-spline Encoded Action Sequence Tokenizer
(BEAST), a novel action tokenizer that encodes action sequences into compact discrete or continuous tokens using B-splines. In contrast to existing action tokenizers based on vector quantization or byte pair encoding, BEAST requires no separate tokenizer training and consistently produces tokens of uniform length, enabling fast action sequence generation via parallel decoding. Leveraging our B-spline formulation, BEAST inherently ensures generating smooth trajectories without discontinuities between adjacent segments. We extensively evaluate BEAST by integrating it with three distinct model architectures: a Variational Autoencoder (VAE) with continuous tokens, a decoder-only Transformer with discrete tokens, and Florence-2, a pretrained Vision-Language Model with an encoder-decoder architecture, demonstrating BEAST's compatibility and scalability with large pretrained models. We evaluate BEAST across three established benchmarks consisting of 166 simulated tasks and on three distinct robot settings with a total of 8 real-world tasks. Experimental results demonstrate that BEAST (i) significantly reduces both training and inference computational costs, and (ii) consistently generates smooth, high-frequency control signals suitable for continuous control tasks while (iii) reliably achieves competitive task success rates compared to state-of-the-art methods. Hongyi Zhou, Weiran Liao, Xi Huang 0005, Yucheng Tang, Fabian Otto, Xiaogang Jia, Xinkai Jiang, Simon Hilber, Ömer Erdinç Yagmurlu, Nils Blank, Moritz Reuss, Rudolf Lioutikov |
NeurIPS | 4 |
| 2025 | MAISI: Medical AI for Synthetic ImagingabstractMedical imaging analysis faces challenges such as data scarcity, high annotation costs, and privacy concerns. This paper introduces the Medical AI for Synthetic Imaging (MAISI), an innovative approach using the diffusion model to generate synthetic 3D computed tomography (CT) images to address those challenges. MAISI leverages the foundation volume compression network and the latent diffusion model to produce high-resolution CT images (up to a landmark volume dimension of 512 × 512 × 768) with flexible volume dimensions and voxel spacing. By incorporating ControlNet, MAISI can process organ segmentation, including 127 anatomical structures, as additional conditions and enables the generation of accurately annotated synthetic images that can be used for various downstream tasks. Our experiment results show that MAISI's capabilities in generating realistic, anatomically accurate images for diverse regions and conditions reveal its promising potential to mitigate challenges using synthetic data. Can Zhao 0001, Dong Yang 0005, Ziyue Xu 0001, Vishwesh Nath, Yucheng Tang, Benjamin Simon, Mason Belue, Stephanie A. Harmon, Baris Turkbey, Daguang Xu |
WACV | 6 |
| 2024 | PrPSeg: Universal Proposition Learning for Panoramic Renal Pathology SegmentationabstractUnderstanding the anatomy of renal pathology is crucial for advancing disease diagnostics, treatment evaluation, and clinical research. The complex kidney system comprises various components across multiple levels, including regions (cortex, medulla), functional units (glomeruli, tubules), and cells (podocytes, mesangial cells in glomerulus). Prior studies have predominantly overlooked the intricate spatial interrelations among objects from clinical knowledge. In this research, we introduce a novel universal proposition learning approach, called panoramic renal pathology segmentation (PrPSeg), designed to segment comprehensively panoramic structures within kidney by integrating extensive knowledge of kidney anatomy. In this paper, we propose (1) the design of a comprehensive universal proposition matrix for renal pathology, facilitating the incorporation of classification and spatial relationships into the segmentation process; (2) a token-based dynamic head single network architecture, with the improvement of the partial label image segmentation and capability for future data enlargement; and (3) an anatomy loss function, quantifying the inter-object relationships across the kidney. Ruining Deng, Quan Liu 0002, Can Cui 0006, Tianyuan Yao, Jialin Yue, Juming Xiong, Lining Yu, Mengmeng Yin, Shilin Zhao, Yucheng Tang, Haichun Yang, Yuankai Huo |
CVPR | 12 |
| 2024 | HATs: Hierarchical Adaptive Taxonomy Segmentation for Panoramic Pathology Image Analysis
Ruining Deng, Quan Liu 0002, Can Cui 0006, Tianyuan Yao, Juming Xiong, Shunxing Bao, Hao Li 0108, Mengmeng Yin, Shilin Zhao, Yucheng Tang, Haichun Yang, Yuankai Huo |
MICCAI (4) | 11 |
| 2024 | Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?abstractHow can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain. Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 3 |
| 2024 | MONAI Label: A framework for AI-assisted interactive labeling of 3D medical images
Andres Diaz-Pinto, Sachidanand Alle, Vishwesh Nath, Yucheng Tang, Alvin Ihsani, Muhammad Asad 0001, Fernando Pérez-García, Pritesh Mehta, Wenqi Li 0001, Mona Flores, Holger Roth, Tom Vercauteren, Daguang Xu, Prerna Dogra, Sébastien Ourselin, Andrew Feng, Manuel Jorge Cardoso |
Medical Image Anal. | 4 |
| 2024 | Universal and extensible language-vision models for organ segmentation and tumor detection from abdominal computed tomography
Jie Liu 0044, Yixiao Zhang 0001, Kang Wang 0016, Mehmet Can Yavuz, Xiaoxi Chen, Yixuan Yuan, Haoliang Li, Yang Yang 0009, Alan L. Yuille, Yucheng Tang, Zongwei Zhou |
Medical Image Anal. | 10 |
| 2023 | CLIP-Driven Universal Model for Organ Segmentation and Tumor DetectionabstractAn increasing number of public datasets have shown a marked impact on automated organ segmentation and tumor detection. However, due to the small size and partially labeled problem of each dataset, as well as a limited investigation of diverse types of tumors, the resulting models are often limited to segmenting specific organs/tumors and ignore the semantics of anatomical structures, nor can they be extended to novel domains. To address these issues, we propose the CLIP-Driven Universal Model, which incorporates text embedding learned from Contrastive Language-Image Pre-training (CLIP) to segmentation models. This CLIP-based label encoding captures anatomical relationships, enabling the model to learn a structured feature embedding and segment 25 organs and 6 types of tumors. The proposed model is developed from an assembly of 14 datasets, using a total of 3,410 CT scans for training and then evaluated on 6,162 external CT scans from 3 additional datasets. We rank first on the Medical Segmentation Decathlon (MSD) public leaderboard and achieve state-of-the-art results on Beyond The Cranial Vault (BTCV). Additionally, the Universal Model is computationally more efficient (6× faster) compared with dataset-specific models, generalized better to CT scans from varying sites, and shows stronger transfer learning performance on novel tasks. Jie Liu 0044, Yixiao Zhang 0001, Jieneng Chen, Junfei Xiao, Yongyi Lu, Bennett A. Landman, Yixuan Yuan, Alan L. Yuille, Yucheng Tang, Zongwei Zhou |
ICCV | 9 |
| 2023 | Reachability-Aware Collision Avoidance for Tractor-Trailer System with Non-Linear MPC and Control Barrier FunctionabstractThis paper proposes a reachability-aware model predictive control with a discrete control barrier function for backward obstacle avoidance for a tractor-trailer system. The framework incorporates the state-variant reachable set obtained through sampling-based reachability analysis and symbolic regression into the objective function of model predictive control. By optimizing the intersection of the reachable set and iterative non-safe region generated by the control barrier function, the system demonstrates better performance in terms of safety with a constant decay rate, while enhancing the feasibility of the optimization problem. The proposed algorithm improves real-time performance due to a shorter horizon and outperforms the state-of-the-art algorithms in the simulation environment and on a real robot. Yucheng Tang, Ilshat Mamaev, Christian Wurll, Björn Hein |
IROS | 1 |
| 2023 | Democratizing Pathological Image Segmentation with Lay Annotators via Molecular-Empowered Learning
Ruining Deng, Peize Li, Jiacheng Wang 0007, Lucas W. Remedios, Saydolimkhon Agzamkhodjaev, Zuhayr Asad, Quan Liu 0002, Can Cui 0006, Yaohong Wang, Yucheng Tang, Haichun Yang, Yuankai Huo |
MICCAI (6) | 12 |
| 2023 | SwinUNETR-V2: Stronger Swin Transformers with Stagewise Convolutions for 3D Medical Image Segmentation
Yufan He, Vishwesh Nath, Dong Yang 0005, Yucheng Tang, Andriy Myronenko, Daguang Xu |
MICCAI (4) | 4 |
| 2023 | PLD-AL: Pseudo-label Divergence-Based Active Learning in Carotid Intima-Media Segmentation for Ultrasound Images
Yucheng Tang, Yipeng Hu, Hongxiang Lin |
MICCAI (2) | 1 |
| 2023 | AbdomenAtlas-8K: Annotating 8, 000 CT Volumes for Multi-Organ Segmentation in Three WeeksabstractAnnotating medical images, particularly for organ segmentation, is laborious and time-consuming. For example, annotating an abdominal organ requires an estimated rate of 30-60 minutes per CT volume based on the expertise of an annotator and the size, visibility, and complexity of the organ. Therefore, publicly available datasets for multi-organ segmentation are often limited in data size and organ diversity. This paper proposes an active learning procedure to expedite the annotation process for organ segmentation and creates the largest multi-organ dataset (by far) with the spleen, liver, kidneys, stomach, gallbladder, pancreas, aorta, and IVC annotated in 8,448 CT volumes, equating to 3.2 million slices. The conventional annotation methods would take an experienced annotator up to 1,600 weeks (or roughly 30.8 years) to complete this task. In contrast, our annotation procedure has accomplished this task in three weeks (based on an 8-hour workday, five days a week) while maintaining a similar or even better annotation quality. This achievement is attributed to three unique properties of our method: (1) label bias reduction using multiple pre-trained segmentation models, (2) effective error detection in the model predictions, and (3) attention guidance for annotators to make corrections on the most salient errors. Furthermore, we summarize the taxonomy of common errors made by AI algorithms and annotators. This allows for continuous improvement of AI and annotations, significantly reducing the annotation costs required to create large-scale datasets for a wider variety of medical imaging tasks. Code and dataset are available at https://github.com/MrGiovanni/AbdomenAtlas Chongyu Qu, Tiezheng Zhang, Hualin Qiao, Jie Liu 0044, Yucheng Tang, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 5 |
| 2023 | Transforming medical imaging with Transformers? A comparative review of key properties, current progresses, and future perspectives
Jun Li 0103, Junyu Chen 0002, Yucheng Tang, Ce Wang 0001, Bennett A. Landman, Shaohua Kevin Zhou |
Medical Image Anal. | 3 |
| 2023 | UNesT: Local spatial representation learning with hierarchical transformer for efficient medical segmentation
Xin Yu 0010, Qi Yang 0004, Yinchi Zhou, Leon Y. Cai, Riqiang Gao, Ho Hin Lee, Thomas Z. Li, Shunxing Bao, Zhoubing Xu, Thomas A. Lasko, Richard G. Abramson, Yuankai Huo, Bennett A. Landman, Yucheng Tang |
Medical Image Anal. | 15 |
| 2023 | Semantic-Aware Contrastive Learning for Multi-Object Medical Image SegmentationabstractMedical image segmentation, or computing voxel-wise semantic masks, is a fundamental yet challenging task in medical imaging domain. To increase the ability of encoder-decoder neural networks to perform this task across large clinical cohorts, contrastive learning provides an opportunity to stabilize model initialization and enhances downstream tasks performance without ground-truth voxel-wise labels. However, multiple target objects with different semantic meanings and contrast level may exist in a single image, which poses a problem for adapting traditional contrastive learning methods from prevalent "image-level classification" to "pixel-level segmentation". In this article, we propose a simple semantic-aware contrastive learning approach leveraging attention masks and image-wise labels to advance multi-object semantic segmentation. Briefly, we embed different semantic objects to different clusters rather than the traditional image-level embeddings. We evaluate our proposed method on a multi-organ medical image segmentation task with both in-house data and MICCAI Challenge 2015 BTCV datasets. Compared with current state-of-the-art training strategies, our proposed pipeline yields a substantial improvement of 5.53% and 6.09% on Dice score for both medical image segmentation cohorts respectively (p-value 0.01). The performance of the proposed method is further assessed on external medical image cohort via MICCAI Challenge FLARE 2021 dataset, and achieves a substantial improvement from Dice 0.922 to 0.933 (p-value 0.01). Ho Hin Lee, Yucheng Tang, Qi Yang 0004, Xin Yu 0010, Leon Y. Cai, Lucas W. Remedios, Shunxing Bao, Bennett A. Landman, Yuankai Huo |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Self-Supervised Pre-Training of Swin Transformers for 3D Medical Image AnalysisabstractVision Transformers (ViT)s have shown great performance in self-supervised learning of global and local representations that can be transferred to downstream applications. Inspired by these results, we introduce a novel self-supervised learning framework with tailored proxy tasks for medical image analysis. Specifically, we propose: (i) a new 3D transformer-based model, dubbed Swin UNEt TRansformers (Swin UNETR), with a hierarchical encoder for self-supervised pretraining; (ii) tailored proxy tasks for learning the underlying pattern of human anatomy. We demonstrate successful pre-training of the proposed model on 5,050 publicly available computed tomography (CT) images from various body organs. The effectiveness of our approach is validated by fine-tuning the pre-trained models on the Beyond the Cranial Vault (BTCV) Segmentation Challenge with 13 abdominal organs and segmentation tasks from the Medical Segmentation Decathlon (MSD) dataset. Our model is currently the state-of-the-art on the public test leaderboards of both MSD11https://decathlon-10.grand-challenge.org/evaluation/challenge/leaderboard/ and BTCV22https://www.synapse.org/#!Synapse:syn3193805/wiki/217785/ datasets. Code: https://monai.io/research/swin-unetr. Yucheng Tang, Dong Yang 0005, Wenqi Li 0001, Holger Roth, Bennett A. Landman, Daguang Xu, Vishwesh Nath, Ali Hatamizadeh |
CVPR | 1 |
| 2022 | Reducing Positional Variance in Cross-sectional Abdominal CT Slices with Deep Conditional Generative Models
Xin Yu 0010, Qi Yang 0004, Yucheng Tang, Riqiang Gao, Shunxing Bao, Leon Y. Cai, Ho Hin Lee, Yuankai Huo, Ann Zenobia Moore, Luigi Ferrucci, Bennett A. Landman |
MICCAI (8) | 3 |
| 2022 | UNETR: Transformers for 3D Medical Image SegmentationabstractFully Convolutional Neural Networks (FCNNs) with contracting and expanding paths have shown prominence for the majority of medical image segmentation applications since the past decade. In FCNNs, the encoder plays an integral role by learning both global and local features and contextual representations which can be utilized for semantic output prediction by the decoder. Despite their success, the locality of convolutional layers in FCNNs, limits the capability of learning long-range spatial dependencies. Inspired by the recent success of transformers for Natural Language Processing (NLP) in long-range sequence learning, we reformulate the task of volumetric (3D) medical image segmentation as a sequence-to-sequence prediction problem. We introduce a novel architecture, dubbed as UNEt TRansformers (UNETR), that utilizes a transformer as the encoder to learn sequence representations of the input volume and effectively capture the global multi-scale information, while also following the successful "U-shaped" network design for the encoder and decoder. The transformer encoder is directly connected to a decoder via skip connections at different resolutions to compute the final semantic segmentation output. We have validated the performance of our method on the Multi Atlas Labeling Beyond The Cranial Vault (BTCV) dataset for multi-organ segmentation and the Medical Segmentation Decathlon (MSD) dataset for brain tumor and spleen segmentation tasks. Our benchmarks demonstrate new state-of-the-art performance on the BTCV leaderboard. Ali Hatamizadeh, Yucheng Tang, Vishwesh Nath, Dong Yang 0005, Andriy Myronenko, Bennett A. Landman, Holger Roth, Daguang Xu |
WACV | 2 |
| 2021 | Lung Cancer Risk Estimation with Incomplete Data: A Joint Missing Imputation Perspective
Riqiang Gao, Yucheng Tang, Kaiwen Xu, Ho Hin Lee, Steve Deppen, Kim L. Sandler, Pierre P. Massion, Thomas A. Lasko, Yuankai Huo, Bennett A. Landman |
MICCAI (5) | 2 |
| 2021 | Pancreas CT Segmentation by Predictive Phenotyping
Yucheng Tang, Riqiang Gao, Ho Hin Lee, Qi Yang 0004, Xin Yu 0010, Yuyin Zhou, Shunxing Bao, Yuankai Huo, Jeffrey M. Spraggins, John Virostko, Zhoubing Xu, Bennett A. Landman |
MICCAI (1) | 1 |
| 2021 | High-resolution 3D abdominal segmentation with random patch network fusion
Yucheng Tang, Riqiang Gao, Ho Hin Lee, Shizhong Han, Yunqiang Chen, Dashan Gao 0001, Vishwesh Nath, Camilo Bermudez, Michael R. Savona, Richard G. Abramson, Shunxing Bao, Ilwoo Lyu, Yuankai Huo, Bennett A. Landman |
Medical Image Anal. | 1 |
| 2021 | Body Part Regression With Self-SupervisionabstractBody part regression is a promising new technique that enables content navigation through self-supervised learning. Using this technique, the global quantitative spatial location for each axial view slice is obtained from computed tomography (CT). However, it is challenging to define a unified global coordinate system for body CT scans due to the large variabilities in image resolution, contrasts, sequences, and patient anatomy. Therefore, the widely used supervised learning approach cannot be easily deployed. To address these concerns, we propose an annotation-free method named blind-unsupervised-supervision network (BUSN). The contributions of the work are in four folds: (1) 1030 multi-center CT scans are used in developing BUSN without any manual annotation. (2) the proposed BUSN corrects the predictions from unsupervised learning and uses the corrected results as the new supervision; (3) to improve the consistency of predictions, we propose a novel neighbor message passing (NMP) scheme that is integrated with BUSN as a statistical learning based correction; and (4) we introduce a new pre-processing pipeline with inclusion of the BUSN, which is validated on 3D multi-organ segmentation. The proposed method is trained on 1,030 whole body CT scans (230,650 slices) from five datasets, as well as an independent external validation cohort with 100 scans. From the body part regression results, the proposed BUSN achieved significantly higher median R-squared score (=0.9089) than the state-of-the-art unsupervised method (=0.7153). When introducing BUSN as a preprocessing stage in volumetric segmentation, the proposed pre-processing pipeline using BUSN approach increases the total mean Dice score of the 3D abdominal multi-organ segmentation from 0.7991 to 0.8145. Yucheng Tang, Riqiang Gao, Shizhong Han, Yunqiang Chen, Dashan Gao 0001, Vishwesh Nath, Camilo Bermudez, Michael R. Savona, Shunxing Bao, Ilwoo Lyu, Yuankai Huo, Bennett A. Landman |
IEEE Trans. Medical Imaging | 1 |
| 2020 | Multi-path x-D recurrent neural networks for collaborative image classification
Riqiang Gao, Yuankai Huo, Shunxing Bao, Yucheng Tang, Sanja Antic, Emily S. Epstein, Steve Deppen, Alexis B. Paulson, Kim L. Sandler, Pierre P. Massion, Bennett A. Landman |
Neurocomputing | 4 |
| 2020 | Time-distanced gates in long short-term memory networks
Riqiang Gao, Yucheng Tang, Kaiwen Xu, Yuankai Huo, Shunxing Bao, Sanja Antic, Emily S. Epstein, Steve Deppen, Alexis B. Paulson, Kim L. Sandler, Pierre P. Massion, Bennett A. Landman |
Medical Image Anal. | 2 |
| 2018 | More Knowledge Is Better: Cross-Modality Volume Completion and 3D+2D Segmentation for Intracardiac Echocardiography Contouring
Haofu Liao, Yucheng Tang, Gareth Funka-Lea, Jiebo Luo 0001, Shaohua Kevin Zhou |
MICCAI (2) | 2 |
| 2018 | Face image retrieval: super-resolution based on sketch-photo transformation
Shu Zhan, Yucheng Tang, Zhenzhu Xie |
Soft Comput. | 3 |
| 2017 | A frog-inspired swimming robot based on dielectric elastomer actuatorsabstractFrogs are capable of multiple locomotion modes including jumping and swimming, which enables them to adapt to various environmental conditions. This paper demonstrates a frog-inspired robot, which can mimic the swimming motion of a natural frog. The robot is developed based on dielectric elastomer actuators, which exhibits muscle-like behavior such as large voltage-induced deformation, high energy density, fast response and low weight. Inspired by the webbed feet of a frog, the foot actuator of the swimming robot is able to increase its projected area by 66% when subject to high voltage. Actuation of the foot actuators can significantly improve the averaged peak thrust by 34.5%. The total mass of the two dielectric elastomer actuators is 14g which only accounts for 13% of its total mass of 108g. The measured average swimming speed for a square wave voltage of 5kV and 0.25Hz is 19mm/s for the swimming robot. Future work of the project includes optimal design and control of this soft robot. Yucheng Tang, Chee-Meng Chew, Jian Zhu 0005 |
IROS | 1 |
| 2017 | Real-time 3D face modeling based on 3D face imaging
Shu Zhan, Lele Chang, Toru Kurihara, Yucheng Tang |
Neurocomputing | 6 |