VLDB 2026 Research / reviewers in the wild / expert
Yijin Huang
dblp:266/6995
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-0988-0134ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hessian-Aware Zeroth-Order Optimization for Quantized Large Language Models
Zhangchi Wang, Yidu Wu, Yijin Huang, Pujin Cheng, Qinghai Guo, Xiaoying Tang 0001 |
ICIC (5) | 3 |
| 2026 | MIRAGE: Medical image-text pre-training for robustness against noisy environments
Pujin Cheng, Yijin Huang, Li Lin 0006, Junyan Lyu, Kenneth K. Y. Wong, Xiaoying Tang 0001 |
Medical Image Anal. | 2 |
| 2025 | Boosting Memory Efficiency in Transfer Learning for High-Resolution Medical Image ClassificationabstractThe success of large-scale pretrained models has established fine-tuning as a standard method for achieving significant improvements in downstream tasks. However, fine-tuning the entire parameter set of a pretrained model is costly. Parameter-efficient transfer learning (PETL) has recently emerged as a cost-effective alternative for adapting pretrained models to downstream tasks. Despite its advantages, the increasing model size and input resolution present challenges for PETL, as the training memory consumption is not reduced as effectively as the parameter usage. In this article, we introduce fine-grained prompt tuning plus (FPT+), a PETL method designed for high-resolution medical image classification, which significantly reduces the training memory consumption compared to other PETL methods. FPT+ performs transfer learning by training a lightweight side network and accessing pretrained knowledge from a large pretrained model (LPM) through fine-grained prompts and fusion modules. Specifically, we freeze the LPM of interest and construct a learnable lightweight side network. The frozen LPM processes high-resolution images to extract fine-grained features, while the side network employs corresponding downsampled low-resolution images to minimize memory usage. To enable the side network to leverage pretrained knowledge, we propose fine-grained prompts and fusion modules, which collaborate to summarize information through the LPM's intermediate activations. We evaluate FPT+ on eight medical image datasets of varying sizes, modalities, and complexities. Experimental results demonstrate that FPT+ outperforms other PETL methods, using only 1.03% of the learnable parameters and 3.18% of the memory required for fine-tuning an entire ViT-B model. Our code is available https://github.com/YijinHuang/FPT. Yijin Huang, Pujin Cheng, Roger C. Tam, Xiaoying Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Fusion Side Tuning: A Parameter and Memory Efficient Fine-tuning Method for High-resolution Medical Image ClassificationabstractParameter-efficient fine-tuning (PEFT) has been proposed as a cost-effective approach for transferring large-scale pre-trained models (LPMs) to downstream tasks, mitigating the high costs associated with updating all parameters of LPMs. However, current PEFT methods encounter the challenge that GPU memory usage during training is not reduced as effectively as parameter usage. In this paper, we propose Fusion Side Tuning (FST), a novel memory-efficient and parameter-efficient fine-tuning method. FST significantly reduces GPU memory consumption during training, particularly at high resolutions. To achieve this, we freeze the backbone LPM and construct a learnable side fusion network that takes intermediate features from the backbone as input. The side fusion network consists of a sequence of fusion modules, which enable it to leverage the knowledge embedded in the intermediate features. Additionally, we employ an important token selection mechanism to further reduce training costs and memory requirements. We evaluate FST on eight medical image datasets of varying modalities and sizes. Experimental results demonstrate that FST outperforms existing PEFT methods, utilizing only 2% of the learnable parameters and 20% of the GPU memory required for full fine-tuning of a ViT-B encoder with an input resolution of 512 × 512. Zhangchi Wang, Yijin Huang, Yidu Wu, Pujin Cheng, Li Lin 0006, Qinghai Guo, Xiaoying Tang 0001 |
BIBM | 2 |
| 2024 | Fine-Grained Prompt Tuning: A Parameter and Memory Efficient Transfer Learning Method for High-Resolution Medical Image Classification
Yijin Huang, Pujin Cheng, Roger C. Tam, Xiaoying Tang 0001 |
MICCAI (12) | 1 |
| 2024 | SSiT: Saliency-Guided Self-Supervised Image Transformer for Diabetic Retinopathy GradingabstractSelf-supervised Learning (SSL) has been widely applied to learn image representations through exploiting unlabeled images. However, it has not been fully explored in the medical image analysis field. In this work, Saliency-guided Self-Supervised image Transformer (SSiT) is proposed for Diabetic Retinopathy (DR) grading from fundus images. We novelly introduce saliency maps into SSL, with a goal of guiding self-supervised pre-training with domain-specific prior knowledge. Specifically, two saliency-guided learning tasks are employed in SSiT: 1) Saliency-guided contrastive learning is conducted based on the momentum contrast, wherein fundus images' saliency maps are utilized to remove trivial patches from the input sequences of the momentum-updated key encoder. Thus, the key encoder is constrained to provide target representations focusing on salient regions, guiding the query encoder to capture salient features. 2) The query encoder is trained to predict the saliency segmentation, encouraging the preservation of fine-grained information in the learned representations. To assess our proposed method, four publicly-accessible fundus image datasets are adopted. One dataset is employed for pre-training, while the three others are used to evaluate the pre-trained models' performance on downstream DR grading. The proposed SSiT significantly outperforms other representative state-of-the-art SSL methods on all downstream datasets and under various evaluation settings. For example, SSiT achieves a Kappa score of 81.88% on the DDR dataset under fine-tuning evaluation, outperforming all other ViT-based SSL methods by at least 9.48%. Yijin Huang, Junyan Lyu, Pujin Cheng, Roger C. Tam, Xiaoying Tang 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2023 | PRIOR: Prototype Representation Joint Learning from Medical Images and ReportsabstractContrastive learning based vision-language joint pre-training has emerged as a successful representation learning strategy. In this paper, we present a prototype representation learning framework incorporating both global and local alignment between medical images and reports. In contrast to standard global multi-modality alignment methods, we employ a local alignment module for fine-grained representation. Furthermore, a cross-modality conditional reconstruction module is designed to interchange information across modalities in the training phase by reconstructing masked images and reports. For reconstructing long reports, a sentence-wise prototype memory bank is constructed, enabling the network to focus on low-level localized visual and high-level clinical linguistic features. Additionally, a non-auto-regressive generation paradigm is proposed for reconstructing non-sequential reports. Experimental results on five downstream tasks, including supervised classification, zero-shot classification, image-to-text retrieval, semantic segmentation, and object detection, show the proposed method outperforms other state-of-the-art methods across multiple datasets and under different dataset size settings. The code is available at https://github.com/QtacierP/PRIOR. Pujin Cheng, Li Lin 0006, Junyan Lyu, Yijin Huang, Wenhan Luo, Xiaoying Tang 0001 |
ICCV | 4 |
| 2022 | AADG: Automatic Augmentation for Domain Generalization on Retinal Image SegmentationabstractConvolutional neural networks have been widely applied to medical image segmentation and have achieved considerable performance. However, the performance may be significantly affected by the domain gap between training data (source domain) and testing data (target domain). To address this issue, we propose a data manipulation based domain generalization method, called Automated Augmentation for Domain Generalization (AADG). Our AADG framework can effectively sample data augmentation policies that generate novel domains and diversify the training set from an appropriate search space. Specifically, we introduce a novel proxy task maximizing the diversity among multiple augmented novel domains as measured by the Sinkhorn distance in a unit sphere space, making automated augmentation tractable. Adversarial training and deep reinforcement learning are employed to efficiently search the objectives. Quantitative and qualitative experiments on 11 publicly-accessible fundus image datasets (four for retinal vessel segmentation, four for optic disc and cup (OD/OC) segmentation and three for retinal lesion segmentation) are comprehensively performed. Two OCTA datasets for retinal vasculature segmentation are further involved to validate cross-modality generalization. Our proposed AADG exhibits state-of-the-art generalization performance and outperforms existing approaches by considerable margins on retinal vessel, OD/OC and lesion segmentation tasks. The learned policies are empirically validated to be model-agnostic and can transfer well to other models. The source code is available at https://github.com/CRazorback/AADG. Junyan Lyu, Yijin Huang, Li Lin 0006, Pujin Cheng, Xiaoying Tang 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2021 | I-SECRET: Importance-Guided Fundus Image Enhancement via Semi-supervised Contrastive Constraining
Pujin Cheng, Li Lin 0006, Yijin Huang, Junyan Lyu, Xiaoying Tang 0001 |
MICCAI (8) | 3 |
| 2021 | Lesion-Based Contrastive Learning for Diabetic Retinopathy Grading from Fundus Images
Yijin Huang, Li Lin 0006, Pujin Cheng, Junyan Lyu, Xiaoying Tang 0001 |
MICCAI (2) | 1 |
| 2021 | BSDA-Net: A Boundary Shape and Distance Aware Joint Learning Framework for Segmenting and Classifying OCTA Images
Li Lin 0006, Jiewei Wu, Yijin Huang, Junyan Lyu, Pujin Cheng, Xiaoying Tang 0001 |
MICCAI (8) | 4 |