Yuhan Zhang 0001

dblp:06/7406-1 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0002-4421-2414ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 From discrete to continuous: A spatiotemporal evolution-aware adversarial diffusion framework for retinal disease progression prediction
Yuhan Zhang 0001, Sijie Niu, Songtao Yuan, Qiang Chen 0004
Neurocomputing2
2026 Test-time generative augmentation for medical image segmentation
Xiao Ma 0011, Yuhui Tao, Zetian Zhang, Yuhan Zhang 0001, Xi Wang 0013, Sheng Zhang 0024, Zexuan Ji, Yizhe Zhang 0001, Qiang Chen 0004, Guang Yang 0006
Medical Image Anal.4
2026 MT-SAM: A Mamba-Transformer Enhanced SAM With Prior-Guided Prompting for Multi-Modal Prostate Cancer Delineation
abstract
Clinically, bi-parametric MRI (bp-MRI), including T2-weighted imaging, diffusion-weighted imaging, and apparent diffusion coefficient map, offers essential prior localization of biopsy and focal therapy for suspicious clinically significant prostate cancer (csPCa), and accurate csPCa delineation from bp-MRI is crucial for better outcomes. However, due to the complexity and high variability in appearance, size, shape, and indistinct boundaries, delineating csPCa remains challenging, time-consuming, and heavily relies on the clinician's experience. To address these issues, we propose MT-SAM, a novel framework that enhances SAM with higher-quality feature extraction and a prior-guided automatic prompting strategy. Specifically, we introduce a mamba-transformer network to extract multi-stage multi-modal features from bp-MRI and fuse them into the SAM encoder via cross-mamba modules. Moreover, we propose a prior-guided pyramid-mamba prompting strategy to strengthen the model's attention on the targets. We extensively evaluate our method on both public and private datasets, and the experimental results show that our method achieves up to 5.6-34.1% higher Dice scores than state-of-the-art methods. Code is available at https://github.com/LuckLT/MT-SAM.
Litao Zhao, Yuhan Zhang 0001, Libiao Ji, Caizi Li, Chi-Fai Ng, Pheng-Ann Heng
IEEE Trans. Medical Imaging2
2024 Memory-Efficient High-Resolution OCT Volume Synthesis with Cascaded Amortized Latent Diffusion Models
Xiao Ma 0011, Yuhan Zhang 0001, Songtao Yuan, Yong Liu 0026, Qiang Chen 0004, Huazhu Fu
MICCAI (7)3
2024 Coarse-to-Fine Latent Diffusion Model for Glaucoma Forecast on Sequential Fundus Images
Yuhan Zhang 0001, Xikai Yang, Xiao Ma 0011, Ningli Wang, Xi Wang 0013, Pheng-Ann Heng
MICCAI (5)1
2024 SAVE: Encoding spatial interactions for vision transformers
Xiao Ma 0011, Zetian Zhang, Zexuan Ji, Mingchao Li 0002, Yuhan Zhang 0001, Qiang Chen 0004
Image Vis. Comput.6
2024 OCTA-500: A retinal dataset for optical coherence tomography angiography study
Mingchao Li 0002, Qiuzhuo Xu, Jiadong Yang, Yuhan Zhang 0001, Zexuan Ji, Keren Xie, Songtao Yuan, Qinghuai Liu, Qiang Chen 0004
Medical Image Anal.5
2024 Diverse Data Generation for Retinal Layer Segmentation With Potential Structure Modeling
abstract
Accurate retinal layer segmentation on optical coherence tomography (OCT) images is hampered by the challenges of collecting OCT images with diverse pathological characterization and balanced distribution. Current generative models can produce high-realistic images and corresponding labels without quantitative limitations by fitting distributions of real collected data. Nevertheless, the diversity of their generated data is still limited due to the inherent imbalance of training data. To address these issues, we propose an image-label pair generation framework that generates diverse and balanced potential data from imbalanced real samples. Specifically, the framework first generates diverse layer masks, and then generates plausible OCT images corresponding to these layer masks using two customized diffusion probabilistic models respectively. To learn from imbalanced data and facilitate balanced generation, we introduce pathological-related conditions to guide the generation processes. To enhance the diversity of the generated image-label pairs, we propose a potential structure modeling technique that transfers the knowledge of diverse sub-structures from lowly- or non-pathological samples to highly pathological samples. We conducted extensive experiments on two public datasets for retinal layer segmentation. Firstly, our method generates OCT images with higher image quality and diversity compared to other generative methods. Furthermore, based on the extensive training with the generated OCT images, downstream retinal layer segmentation tasks demonstrate improved results. The code is publicly available at: https://github.com/nicetomeetu21/GenPSM.
Xiao Ma 0011, Zetian Zhang, Yuhan Zhang 0001, Songtao Yuan, Huazhu Fu, Qiang Chen 0004
IEEE Trans. Medical Imaging4
2024 Semantic-Oriented Visual Prompt Learning for Diabetic Retinopathy Grading on Fundus Images
abstract
Diabetic retinopathy (DR) is a serious ocular condition that requires effective monitoring and treatment by ophthalmologists. However, constructing a reliable DR grading model remains a challenging and costly task, heavily reliant on high-quality training sets and adequate hardware resources. In this paper, we investigate the knowledge transferability of large-scale pre-trained models (LPMs) to fundus images based on prompt learning to construct a DR grading model efficiently. Unlike full-tuning which fine-tunes all parameters of LPMs, prompt learning only involves a minimal number of additional learnable parameters while achieving a competitive effect as full-tuning. Inspired by visual prompt tuning, we propose Semantic-oriented Visual Prompt Learning (SVPL) to enhance the semantic perception ability for better extracting task-specific knowledge from LPMs, without any additional annotations. Specifically, SVPL assigns a group of learnable prompts for each DR level to fit the complex pathological manifestations and then aligns each prompt group to task-specific semantic space via a contrastive group alignment (CGA) module. We also propose a plug-and-play adapter module, Hierarchical Semantic Delivery (HSD), which allows the semantic transition of prompt groups from shallow to deep layers to facilitate efficient knowledge mining and model convergence. Our extensive experiments on three public DR grading datasets demonstrate that SVPL achieves superior results compared to other transfer tuning and DR grading methods. Further analysis suggests that the generalized knowledge from LPMs is advantageous for constructing the DR grading model on fundus images.
Yuhan Zhang 0001, Xiao Ma 0011, Mingchao Li 0002, Pheng-Ann Heng
IEEE Trans. Medical Imaging1
2023 SATTA: Semantic-Aware Test-Time Adaptation for Cross-Domain Medical Image Segmentation
Yuhan Zhang 0001, Cheng Chen 0013, Qiang Chen 0004, Pheng-Ann Heng
MICCAI (2)1
2023 Triplet attention and dual-pool contrastive learning for clinic-driven multi-label medical image classification
Yuhan Zhang 0001, Luyang Luo, Qi Dou 0001, Pheng-Ann Heng
Medical Image Anal.1
2022 Self-Supervised Sequence Recovery for Semi-Supervised Retinal Layer Segmentation
abstract
Automated layer segmentation plays an important role for retinal disease diagnosis in optical coherence tomography (OCT) images. However, the severe retinal diseases result in the performance degeneration of automated layer segmentation approaches. In this paper, we present a robust semi-supervised layer segmentation network to relieve the model failures on abnormal retinas. We obtain the lesion features from the labeled images with disease-balanced distribution, and utilize the unlabeled images to supplement the layer structure information. Specifically, in our method, the cross-consistency training is utilized over the predictions of different decoders, and we enforce a consistency between different decoder predictions to improve the encoder's representation. Then, we propose a sequence prediction branch based on self-supervised manner, which is designed to predict the position of each jigsaw puzzle to obtain sensory perception of the retinal layer structure. To this task, a layer spatial pyramid pooling (LSPP) module is designed to extract multi-scale layer spatial features. Furthermore, we use the optical coherence tomography angiography (OCTA) to supplement the information damaged by diseases. The experimental results illustrate that our method achieves more robust results compared with current supervised segmentation methods. Meanwhile, advanced segmentation performance can be obtained compared with state-of-the-art semi-supervised segmentation methods.
Jiadong Yang, Yuhui Tao, Qiuzhuo Xu, Yuhan Zhang 0001, Xiao Ma 0011, Songtao Yuan, Qiang Chen 0004
IEEE J. Biomed. Health Informatics4
2022 LamNet: A Lesion Attention Maps-Guided Network for the Prediction of Choroidal Neovascularization Volume in SD-OCT Images
abstract
Choroidal neovascularization (CNV) volume prediction has an important clinical significance to predict the therapeutic effect and schedule the follow-up. In this paper, we propose a Lesion Attention Maps-Guided Network (LamNet) to automatically predict the CNV volume of next follow-up visit after therapy based on 3-dimentional spectral-domain optical coherence tomography (SD-OCT) images. In particular, the backbone of LamNet is a 3D convolutional neural network (3D-CNN). In order to guide the network to focus on the local CNV lesion regions, we use CNV attention maps generated by an attention map generator to produce the multi-scale local context features. Then, the multi-scale of both local and global feature maps are fused to achieve the high-precision CNV volume prediction. In addition, we also design a synergistic multi-task predictor, in which a trend-consistent loss ensures that the change trend of the predicted CNV volume is consistent with the real change trend of the CNV volume. The experiments include a total of 541 SD-OCT cubes from 68 patients with two types of CNV captured by two different SD-OCT devices. The results demonstrate that LamNet can provide the reliable and accurate CNV volume prediction, which would further assist the clinical diagnosis and design the treatment options.
Yuhan Zhang 0001, Xiao Ma 0011, Mingchao Li 0002, Zexuan Ji, Songtao Yuan, Qiang Chen 0004
IEEE J. Biomed. Health Informatics1
2021 Twin self-supervision based semi-supervised learning (TS-SSL): Retinal anomaly classification in SD-OCT images
Yuhan Zhang 0001, Mingchao Li 0002, Zexuan Ji, Wen Fan 0003, Songtao Yuan, Qinghuai Liu, Qiang Chen 0004
Neurocomputing1
2021 An integrated time adaptive geographic atrophy prediction model for SD-OCT images
Yuhan Zhang 0001, Zexuan Ji, Sijie Niu, Theodore Leng, Daniel L. Rubin, Songtao Yuan, Qiang Chen 0004
Medical Image Anal.1
2020 Adaptive Dictionary Learning Based Multimodal Branch Retinal Vein Occlusion Fusion
Keren Xie, Yuhan Zhang 0001, Mingchao Li 0002, Qiang Chen 0004
MICCAI (5)3
2020 Robust Layer Segmentation Against Complex Retinal Abnormalities for en face OCTA Generation
Yuhan Zhang 0001, Mingchao Li 0002, Sha Xie, Keren Xie, Zexuan Ji, Songtao Yuan, Qiang Chen 0004
MICCAI (5)1