Yunhe Gao

dblp:237/4741 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-1559-5532ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation
abstract
Accurate disease interpretation from radiology remains challenging due to imaging heterogeneity. Achieving expert-level diagnostic decisions requires integration of subtle image features with clinical knowledge. Yet major vision-language models (VLMs) treat images as holistic entities and overlook fine-grained image details that are vital for disease diagnosis. Clinicians analyze images by utilizing their prior medical knowledge and identify anatomical structures as important region of interests (ROIs). Inspired from this human-centric workflow, we introduce Anatomy-VLM, a fine-grained, vision-language model that incorporates multi-scale information. First, we design a model encoder to localize key anatomical features from entire medical images. Second, these regions are enriched with structured knowledge for contextually-aware interpretation. Finally, the model encoder aligns multi-scale medical information to generate clinically-interpretable disease prediction. Anatomy-VLM achieves outstanding performance on both in- and out-of-distribution datasets. We also validate the performance of Anatomy-VLM on downstream image segmentation tasks, suggesting that its fine-grained alignment captures anatomical and pathology-related knowledge. Furthermore, the Anatomy-VLM’s encoder facilitates zero-shot anatomy-wise interpretation, providing its strong expert-level clinical interpretation capabilities.
Difei Gu, Yunhe Gao, Mu Zhou, Dimitris N. Metaxas
WACV2
2025 Show and Segment: Universal Medical Image Segmentation via In-Context Learning
abstract
Medical image segmentation remains challenging due to the vast diversity of anatomical structures, imaging modalities, and segmentation tasks. While deep learning has made significant advances, current approaches struggle to generalize as they require task-specific training or fine-tuning on unseen classes. We present Iris, a novel In-context Reference Image guided Segmentation framework that enables flexible adaptation to novel tasks through the use of reference examples without fine-tuning. At its core, Iris features a lightweight context task encoding module that distills task-specific information from reference context image-label pairs. This rich context embedding information is used to guide the segmentation of target objects. By decoupling task encoding from inference, Iris supports diverse strategies from one-shot inference and context example ensemble to object-level context example retrieval and in-context tuning. Through comprehensive evaluation across twelve datasets, we demonstrate that Iris performs strongly compared to task-specific models on in-distribution tasks. On seven held-out datasets, Iris shows superior generalization to out-of-distribution data and unseen classes. Further, Iris’s task encoding module can automatically discover anatomical relationships across datasets and modalities, offering insights into medical objects without explicit anatomical supervision.
Yunhe Gao, Di Liu 0003, Zhuowei Li 0002, Yunsheng Li, Mu Zhou, Dimitris N. Metaxas
CVPR1
2025 Implicit In-context Learning
abstract
In-context Learning (ICL) empowers large language models (LLMs) to swiftly adapt to unseen tasks at inference-time by prefixing a few demonstration examples before queries. Despite its versatility, ICL incurs substantial computational and memory overheads compared to zero-shot learning and is sensitive to the selection and order of demonstration examples. In this work, we introduce \textbf{Implicit In-context Learning} (I2CL), an innovative paradigm that reduces the inference cost of ICL to that of zero-shot learning with minimal information loss. I2CL operates by first generating a condensed vector representation, namely a context vector, extracted from the demonstration examples. It then conducts an inference-time intervention through injecting a linear combination of the context vector and query activations back into the model’s residual streams. Empirical evaluation on nine real-world tasks across three model architectures demonstrates that I2CL achieves few-shot level performance at zero-shot inference cost, and it exhibits robustness against variations in demonstration examples. Furthermore, I2CL facilitates a novel representation of ``task-ids'', enhancing task similarity detection and fostering effective transfer learning. We also perform a comprehensive analysis and ablation study on I2CL, offering deeper insights into its internal mechanisms. Code is available at https://github.com/LzVv123456/I2CL.
Zhuowei Li 0002, Zihao Xu 0001, Ligong Han, Yunhe Gao, Song Wen 0001, Di Liu 0003, Hao Wang 0014, Dimitris N. Metaxas
ICLR4
2025 The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models Via Visual Information Steering
abstract
Large Vision-Language Models (LVLMs) can reason effectively over both textual and visual inputs, but they tend to hallucinate syntactically coherent yet visually ungrounded contents. In this paper, we investigate the internal dynamics of hallucination by examining the tokens logits rankings throughout the generation process, revealing three key patterns in how LVLMs process information: (1) gradual visual information loss – visually grounded tokens gradually become less favored throughout generation, and (2) early excitation – semantically meaningful tokens achieve peak activation in the layers earlier than the final layer. (3) hidden genuine information – visually grounded tokens though not being eventually decided still retain relatively high rankings at inference. Based on these insights, we propose VISTA (Visual Information Steering with Token-logit Augmentation), a training-free inference-time intervention framework that reduces hallucination while promoting genuine information. VISTA works by combining two complementary approaches: reinforcing visual information in activation space and leveraging early layer activations to promote semantically meaningful decoding. Compared to existing methods, VISTA requires no external supervision and is applicable to various decoding strategies. Extensive experiments show that VISTA on average reduces hallucination by about 40% on evaluated open-ended generation task, and it consistently outperforms existing methods on four benchmarks across four architectures under three decoding strategies. Code is available at https://github.com/LzVv123456/VISTA.
Zhuowei Li 0002, Haizhou Shi, Yunhe Gao, Di Liu 0003, Zhenting Wang, Yuxiao Chen 0002, Ting Liu 0005, Long Zhao 0003, Hao Wang 0014, Dimitris N. Metaxas
ICML3
2025 RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment
Difei Gu, Yunhe Gao, Yang Zhou 0053, Mu Zhou, Dimitris N. Metaxas
MICCAI (7)2
2024 Training Like a Medical Resident: Context-Prior Learning Toward Universal Medical Image Segmentation
abstract
A major focus of clinical imaging workflow is disease diagnosis and management, leading to medical imaging datasets strongly tied to specific clinical objectives. This scenario has led to the prevailing practice of developing task-specific segmentation models, without gaining insights from widespread imaging cohorts. Inspired by the training program of medical radiology residents, we propose a shift towards universal medical image segmentation, a paradigm aiming to build medical image understanding foundation models by leveraging the diversity and commonality across clinical targets, body regions, and imaging modalities. Towards this goal, we develop Hermes, a novel context-prior learning approach to address the challenges of data heterogeneity and annotation differences in medical image segmentation. In a large collection of eleven diverse datasets (2,438 3D images) across five modalities (CT, PET, T1, T2 and cine MRI) and multiple body regions, we demonstrate the merit of the universal paradigm over the traditional paradigm on addressing multiple tasks within a single model. By exploiting the synergy across tasks, Hermes achieves state-of-the-art performance on all testing datasets and shows superior model scalability. Results on two additional datasets reveals Hermes' strong performance for transfer learning, incremental learning, and generalization to downstream tasks. Hermes's learned priors demonstrate an appealing trait to reflect the intricate relations among tasks and modalities, which aligns with the established anatomical and imaging principles in radiology. The code is available11https:/1github.com/yhygao/universal-medical-image-segmentation.
Yunhe Gao
CVPR1
2024 Aligning Human Knowledge with Visual Concepts Towards Explainable Medical Image Classification
Yunhe Gao, Difei Gu, Mu Zhou, Dimitris N. Metaxas
MICCAI (10)1
2023 LEPARD: Learning Explicit Part Discovery for 3D Articulated Shape Reconstruction
abstract
Reconstructing the 3D articulated shape of an animal from a single in-the-wild image is a challenging task. We propose LEPARD, a learning-based framework that discovers semantically meaningful 3D parts and reconstructs 3D shapes in a part-based manner. This is advantageous as 3D parts are robust to pose variations due to articulations and their shape is typically simpler than the overall shape of the object. In our framework, the parts are explicitly represented as parameterized primitive surfaces with global and local deformations in 3D that deform to match the image evidence. We propose a kinematics-inspired optimization to guide each transformation of the primitive deformation given 2D evidence. Similar to recent approaches, LEPARD is only trained using off-the-shelf deep features from DINO and does not require any form of 2D or 3D annotations. Experiments on 3D animal shape reconstruction, demonstrate significant improvement over existing alternatives in terms of both the overall reconstruction performance as well as the ability to discover semantically meaningful and consistent parts.
Di Liu 0003, Anastasis Stathopoulos, Qilong Zhangli, Yunhe Gao, Dimitris N. Metaxas
NeurIPS4
2022 TransFusion: Multi-view Divergent Fusion for Medical Image Segmentation with Transformers
Di Liu 0003, Yunhe Gao, Qilong Zhangli, Ligong Han, Xiaoxiao He, Zhaoyang Xia, Song Wen 0001, Zhennan Yan, Mu Zhou, Dimitris N. Metaxas
MICCAI (5)2
2022 Region Proposal Rectification Towards Robust Instance Segmentation of Biological Images
Qilong Zhangli, Jingru Yi, Di Liu 0003, Xiaoxiao He, Zhaoyang Xia, Ligong Han, Yunhe Gao, Song Wen 0001, Haiming Tang, He Wang 0016, Mu Zhou, Dimitris N. Metaxas
MICCAI (4)8
2021 CrossNorm and SelfNorm for Generalization under Distribution Shifts
abstract
Traditional normalization techniques (e.g., Batch Normalization and Instance Normalization) generally and simplistically assume that training and test data follow the same distribution. As distribution shifts are inevitable in real-world applications, well-trained models with previous normalization methods can perform badly in new environments. Can we develop new normalization methods to improve generalization robustness under distribution shifts? In this paper, we answer the question by proposing Cross-Norm and SelfNorm. CrossNorm exchanges channel-wise mean and variance between feature maps to enlarge training distribution, while SelfNorm uses attention to recalibrate the statistics to bridge gaps between training and test distributions. CrossNorm and SelfNorm can complement each other, though exploring different directions in statistics usage. Extensive experiments on different fields (vision and language), tasks (classification and segmentation), settings (supervised and semi-supervised), and distribution shift types (synthetic and natural) show the effectiveness. Code is available at https://github.com/amazon-research/crossnorm-selfnorm
Zhiqiang Tang 0001, Yunhe Gao, Yi Zhu 0001, Zhi Zhang 0005, Mu Li 0003, Dimitris N. Metaxas
ICCV2
2021 UTNet: A Hybrid Transformer Architecture for Medical Image Segmentation
Yunhe Gao, Mu Zhou, Dimitris N. Metaxas
MICCAI (3)1
2021 FocusNetv2: Imbalanced large and small organ segmentation with adversarial shape constraint for head and neck CT images
Yunhe Gao, Rui Huang 0001, Yiwei Yang 0001, Kainan Shao, Changjuan Tao, Yuanyuan Chen 0007, Dimitris N. Metaxas, Hongsheng Li 0001, Ming Chen 0030
Medical Image Anal.1
2020 OnlineAugment: Online Data Augmentation with Less Domain Knowledge
Zhiqiang Tang 0001, Yunhe Gao, Leonid Karlinsky, Prasanna Sattigeri, Rogério Feris, Dimitris N. Metaxas
ECCV (7)2
2020 Vertebrae Identification and Localization Utilizing Fully Convolutional Networks and a Hidden Markov Model
abstract
Automated identification and localization of vertebrae in spinal computed tomography (CT) imaging is a complicated hybrid task. This task requires detecting and indexing a long sequence in a 3-D image, and both image feature extraction and sequence modeling are needed to address the problem. In this paper, the powerful fully convolutional neural network (FCN) technique performs both of these tasks simultaneously because FCNs directly encode and decode the spatial interdependence of different components in images. The key module of our proposed framework is a 3-D FCN trained in an end-to-end manner at the spine level to capture the long-range contextual information in CT volumes. The large increase in the calculation due to the full-size image inputs is alleviated by the scale-down of the inputs and the use of an auxiliary FCN to compensate for the loss of details. The composite network pipeline design enables the integration of local image details and global image patterns. Furthermore, explicit spatial and sequential constraints are imposed by the hidden Markov model (HMM) for a higher robustness and a clearer interpretation of network outputs. The proposed framework is quantitatively evaluated on the public dataset from the MICCAI 2014 Computational Challenge on Vertebrae Localization and Identification and demonstrates an identification rate (within 20 mm) of 94.67%, a mean identification rate of 87.97%, and a mean error distance of 2.56 mm on the test set, thus achieving the highest performance reported on this dataset.
Yizhi Chen, Yunhe Gao, Kang Li 0004, Liang Zhao 0018, Jun Zhao 0010
IEEE Trans. Medical Imaging2
2019 FocusNet: Imbalanced Large and Small Organ Segmentation with an End-to-End Deep Neural Network for Head and Neck CT Images
Yunhe Gao, Rui Huang 0001, Ming Chen 0030, Zhe Wang 0006, Jincheng Deng, Yuanyuan Chen 0007, Yiwei Yang 0001, Chanjuan Tao, Hongsheng Li 0001
MICCAI (3)1
2019 Multi-resolution Path CNN with Deep Supervision for Intervertebral Disc Localization and Segmentation
Yunhe Gao, Liang Zhao 0018
MICCAI (2)1