Zihao Zhao 0002

dblp:63/9613-2 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0001-8044-7683ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 BrainSMM: Lifespan Brain Segmentation Model With Metadata-Driven Prompt Learning
abstract
Accurate and automatic segmentation of lifespan brain MRI into regions of interest (ROIs) is crucial for studying brain development, aging, and early diagnosis of neurological diseases. Existing segmentation methods are often tailored to specific age groups, such as infants or adults, resulting in inconsistent performance when processing brain data from different age groups. To overcome this limitation, we introduce BrainSMM, a novel metadata-driven model that incorporates text-based prompts to guide representation learning in a segmentation backbone. These prompts, extracted via a pretrained image-text alignment model, encode valuable prior knowledge (e.g., age, scanner, gender) and are infused into the vision model to condition the features according to domain-specific contexts. We evaluate BrainSMM on a large-scale lifespan brain MRI dataset with 5,565 T1w MR images spanning multiple ages. Our approach achieves an average DSC of 94.59% for tissue segmentation (i.e., gray matter, white matter, and cerebrospinal fluid) and 86.34% for anatomical region segmentation (e.g., hippocampus, putamen, etc.) with corresponding average ASD of 0.20 mm and 0.75 mm, respectively. Notably, BrainSMM shows strong consistency in segmentation accuracy across all age groups and demonstrates improved anatomical detail preservation compared to baseline methods. Additionally, our metadata prompt technique is easily transferable and compatible with multiple backbone architectures, highlighting its adaptability. Overall, BrainSMM offers a robust, generalizable solution for lifespan brain MRI segmentation and lays the groundwork for enhanced clinical and developmental neuroimaging applications.
Zihao Zhao 0002, Feng Shi 0001, Dinggang Shen
IEEE Trans. Medical Imaging2
2025 Med-LEGO: Editing and Adapting Toward Generalist Medical Image Diagnosis
Yitao Zhu, Jiaming Li 0012, Mengjie Xu, Zihao Zhao 0002, Honglin Xiong, Sheng Wang 0014, Qian Wang 0001
MICCAI (6)5
2025 Learning better contrastive view from radiologist's gaze
Sheng Wang 0014, Zihao Zhao 0002, Zixu Zhuang, Xi Ouyang, Lichi Zhang, Zheren Li, Chong Ma 0004, Tianming Liu 0001, Dinggang Shen, Qian Wang 0001
Pattern Recognit.2
2025 Improving Self-Supervised Medical Image Pre-Training by Early Alignment With Human Eye Gaze Information
abstract
Alignment between human knowledge and machine learning models is crucial for achieving efficient and interpretable AI systems. However, conventional self-supervised pre-training methods often suffer from low efficiency, as they do not incorporate human knowledge during the pre-training process and instead rely mainly on post-hoc alignment techniques. We propose Gaze Pre-Training (GzPT), a novel approach that introduces early alignment with human eye gaze information during the pre-training process to enhance both the learning efficiency and performance of self-supervised models. By leveraging contrastive learning to pull together images with similar gaze patterns, GzPT can effectively align the model with human attention during the pre-training. We demonstrate the effectiveness of our approach on three diverse medical image datasets, showing that GzPT can consistently outperform baseline methods and learn more meaningful and interpretable representations. Our findings also highlight the potential of incorporating human eye gaze as a form of passive knowledge to bridge the gap between human and machine learning in the self-supervised pre-training. Our code is available at Github.
Sheng Wang 0014, Zihao Zhao 0002, Zhenrong Shen 0001, Bin Wang 0068, Qian Wang 0001, Dinggang Shen
IEEE Trans. Medical Imaging2
2024 Mining Gaze for Contrastive Learning toward Computer-Assisted Diagnosis
abstract
Obtaining large-scale radiology reports can be difficult for medical images due to ethical concerns, limiting the effectiveness of contrastive pre-training in the medical image domain and underscoring the need for alternative methods. In this paper, we propose eye-tracking as an alternative to text reports, as it allows for the passive collection of gaze signals without ethical issues. By tracking the gaze of radiologists as they read and diagnose medical images, we can understand their visual attention and clinical reasoning. When a radiologist has similar gazes for two medical images, it may indicate semantic similarity for diagnosis, and these images should be treated as positive pairs when pre-training a computer-assisted diagnosis (CAD) network through contrastive learning. Accordingly, we introduce the Medical contrastive Gaze Image Pre-training (McGIP) as a plug-and-play module for contrastive learning frameworks. McGIP uses radiologist gaze to guide contrastive pre-training. We evaluate our method using two representative types of medical images and two common types of gaze data. The experimental results demonstrate the practicality of McGIP, indicating its high potential for various clinical scenarios and applications.
Zihao Zhao 0002, Sheng Wang 0014, Qian Wang 0001, Dinggang Shen
AAAI1
2024 Gaze-DETR: Using Expert Gaze to Reduce False Positives in Vulvovaginal Candidiasis Screening
Yan Kong, Sheng Wang 0014, Jiangdong Cai, Zihao Zhao 0002, Zhenrong Shen 0001, Yonghao Li, Manman Fei, Qian Wang 0001
MICCAI (4)4
2024 Knowledge-Guided Prompt Learning for Lifespan Brain MR Image Segmentation
Zihao Zhao 0002, Zehong Cao, Runqi Meng, Feng Shi 0001, Dinggang Shen
MICCAI (2)2
2024 ChatCAD+: Toward a Universal and Reliable Interactive CAD Using LLMs
abstract
The integration of Computer-Aided Diagnosis (CAD) with Large Language Models (LLMs) presents a promising frontier in clinical applications, notably in automating diagnostic processes akin to those performed by radiologists and providing consultations similar to a virtual family doctor. Despite the promising potential of this integration, current works face at least two limitations: (1) From the perspective of a radiologist, existing studies typically have a restricted scope of applicable imaging domains, failing to meet the diagnostic needs of different patients. Also, the insufficient diagnostic capability of LLMs further undermine the quality and reliability of the generated medical reports. (2) Current LLMs lack the requisite depth in medical expertise, rendering them less effective as virtual family doctors due to the potential unreliability of the advice provided during patient consultations. To address these limitations, we introduce ChatCAD+, to be universal and reliable. Specifically, it is featured by two main modules: (1) Reliable Report Generation and (2) Reliable Interaction. The Reliable Report Generation module is capable of interpreting medical images from diverse domains and generate high-quality medical reports via our proposed hierarchical in-context learning. Concurrently, the interaction module leverages up-to-date information from reputable medical websites to provide reliable medical advice. Together, these designed modules synergize to closely align with the expertise of human medical professionals, offering enhanced consistency and reliability for interpretation and advice. The source code is available at GitHub.
Zihao Zhao 0002, Sheng Wang 0014, Jinchen Gu, Yitao Zhu, Lanzhuju Mei, Zixu Zhuang, Zhiming Cui 0001, Qian Wang 0001, Dinggang Shen
IEEE Trans. Medical Imaging1