Yuchong Yao

dblp:311/3645 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0001-8368-4978ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 5 since 2021Theory of computation · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Zero-shot Stroke Lesion Segmentation via CAM-guided Prompting of MedSAM2
abstract
Accurate segmentation of stroke lesions in diffusion-weighted imaging (DWI) is crucial for clinical decision-making. However, automated infarct segmentation remains challenging due to variable infarct sizes and locations, and it is labor-intensive, requiring expert manual annotations for training. We propose a zero-shot framework to eliminate the need for manual segmentation labels by leveraging weak supervision from class activation maps (CAMs) to guide segmentation using MedSAM2, a foundation model for 3D medical image segmentation. By extracting attention maps from a fine-tuned ResNet on DWI scans labeled with stroke etiology (cause) and combining them with intensity information, we identify key regions and generate bounding-box prompts for MedSAM2. Our method achieves a Dice score of 54.2 ± 5.3% without any manual segmentation labels or tuning of the MedSAM2 model, demonstrating its potential as a scalable solution for reliable pseudo-label generation.
Mohammad Javad Shokri, Yuchong Yao, Nandakishor Desai, Aravinda S. Rao, Angelos Sharobeam, Bernard Yan, Marimuthu Palaniswami
CIKM2
2025 RepMedGAN: Self-supervised Representation-guided Medical GAN for Label-free Medical Image Synthesis
abstract
Medical image synthesis addresses healthcare data scarcity by generating realistic samples for clinical support systems, AI training, and research. However, the field faces challenges due to the complexity of imaging data with its diverse modalities, characteristics, and disease variations. To produce high-quality images, medical image synthesis typically relies on conditional generation, where labels and annotations serve as essential conditions that provide critical guidance signals during the generation process to control desired semantics and fidelity. However, in the medical domain, labels are often inaccessible due to the high cost of annotation, requirements for clinical expertise, as well as ethical concerns. To address this critical challenge, we propose RepMedGAN, a novel self-supervised representation-guided image generation framework that enhances label-free medical image synthesis by leveraging self-supervised learning representations, enabling high-quality generation across different modalities without requiring labels or annotations. Our framework incorporates a Self-supervised Guidance Module that provides rich semantic knowledge during training and introduces a Guidance Representation Generator to bridge the train-inference disparity. Through extensive evaluation across four diverse medical datasets including brain MRI, chest X-ray, kidney CT, and eye glaucoma images, we demonstrate that RepMedGAN consistently achieves state-of-the-art results across multiple metrics and produces superior-quality medical images.
Yuchong Yao, Nandakishor Desai, Marimuthu Palaniswami
CIKM1
2025 Rethinking Masked Image Modeling for Ultrasound Image Denoising
abstract
Ultrasound imaging serves as an important clinical diagnostic modality due to its non-invasive, radiation-free, and real-time capabilities. However, ultrasound images suffer from speckle noise that significantly compromises diagnostic accuracy and clinical interpretation. Traditional denoising methods are limited by speckle noise's signal-dependent nature, often removing important diagnostic features. While deep learning performs better, it requires large labelled datasets that are difficult to obtain due to privacy concerns and annotation costs. Self-supervised learning through masked image modeling (MIM) shows potential in addressing data scarcity, but conventional MIM, developed for high-level vision tasks, is unsuitable for low-level tasks like image denoising due to its framework architecture and learning strategy. To this end, we propose Image Denoising Masked Image Modeling (ID-MIM), the first MIM framework for ultrasound image denoising. ID-MIM incorporates a novel high-frequency oriented dual-branch masking and a specialized learning objective for noise reduction. Our encoder-only architecture features a multi-scale hierarchical transformer with dynamic skip connections, where the encoder directly performs denoising rather than relying on separate decoder reconstruction as in conventional MIM approaches. Extensive experiments demonstrate the superior performance of our ID-MIM framework across diverse noise scenarios, establishing new state-of-the-art results.
Yuchong Yao, Nandakishor Desai, Marimuthu Palaniswami
CIKM1
2024 Masked Contrastive Representation Learning for Self-Supervised Visual Pre-Training
abstract
Self-supervised learning has achieved state-of-the-art performance in various tasks and applications. In computer vision, self-supervised learning often employs contrastive learning and masked image modeling, each with its limitations: contrastive learning heavily relies on strong data augmentation and large batch sizes, etc., while masked image modeling struggles to capture high-level semantics and discrimination. In this work, we introduce MAsked Contrastive Representation Learning (MACRL), a novel framework that integrates both paradigms through an asymmetric siamese network design. The online and momentum branches of the network receive asymmetric data augmentation operations and extract features through their encoders. The decoder in the online branch reconstructs the original image, while the projectors in both branches compute the contrastive loss. The online branch and the momentum branch are updated through gradient backpropagation and exponential moving average, respectively. MACRL jointly optimizes the reconstruction and the contrastive objectives to encourage representations with enhanced discrimination and semantics. Experimental results show that MACRL achieves competitive performance in downstream vision tasks, including image classification and semantic segmentation. Moreover, it demonstrates consistent performance across both large-scale and small-scale datasets.
Yuchong Yao, Nandakishor Desai, Marimuthu Palaniswami
DSAA1
2024 MOMA: Contrastive Learning Distills Better Masked Autoencoders
Yuchong Yao, Nandakishor Desai, Marimuthu Palaniswami
ICPR (27)1
2022 Conditional Variational Autoencoder with Balanced Pre-training for Generative Adversarial Networks
abstract
Class imbalance occurs in many real-world applications, including image classification, where the number of images in each class differs significantly. With imbalanced data, the generative adversarial networks (GANs) leans to majority class samples. The two recent methods, Balancing GAN (BAGAN) and improved BAGAN (BAGAN-GP), are proposed as an augmentation tool to handle this problem and restore the balance to the data. The former pre-trains the autoencoder weights in an unsupervised manner. However, it is unstable when the images from different categories have similar features. The latter is improved based on BAGAN by facilitating supervised autoencoder training, but the pre-training is biased towards the majority classes. In this work, we propose a novel Conditional Variational Autoencoder with Balanced Pre-training for Generative Adversarial Networks (CAPGAN) as an augmentation tool to generate realistic synthetic images. In particular, we utilize a conditional convolutional variational autoencoder with supervised and balanced pre-training for the GAN initialization and training with gradient penalty. Our proposed method presents a superior performance of other state-of-the-art methods on the highly imbalanced version of MNIST, Fashion-MNIST, CIFAR-10, and two medical imaging datasets. Our method can synthesize high-quality minority samples in terms of Fréchet inception distance, structural similarity index measure and perceptual quality. The source code is available at https://github.com/alibraytee/CAPGAN.
Yuchong Yao, Yuanbang Ma, Jiaying Wei, Ali Anaissi, Ali Braytee
DSAA1