Kai Yu 0009

dblp:197/1322-9 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
11since 2021 · last 2025
0000-0001-6752-6216ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Text to Image for Multi-Label Image Recognition With Joint Prompt-Adapter Learning
abstract
Benefited from image-text contrastive learning, pre-trained vision-language models, e.g., CLIP, allow to direct leverage texts as images (TaI) for parameter-efficient fine-tuning (PEFT). While CLIP is capable of making image features to be similar to the corresponding text features, the modality gap remains a nontrivial issue and limits image recognition performance of TaI. Using multi-label image recognition (MLR) as an example, we present a novel method, called T2I-PAL to tackle the modality gap issue when using only text captions for PEFT. The core design of T2I-PAL is to leverage pre-trained text-to-image generation models to generate photo-realistic and diverse images from text captions, thereby reducing the modality gap. To further enhance MLR, T2I-PAL incorporates a class-wise heatmap and learnable prototypes. This aggregates local similarities, making the representation of local visual features more robust and informative for multi-label recognition. For better PEFT, we further combine both prompt tuning and adapter learning to enhance classification performance. T2I-PAL offers significant advantages: it eliminates the need for fully semantically annotated training images, thereby reducing the manual annotation workload, and it preserves the intrinsic mode of the CLIP model, allowing for seamless integration with any existing CLIP framework. Extensive experiments on multiple benchmarks, including MS-COCO, VOC2007, and NUS-WIDE, show that our T2I-PAL can boost recognition performance by 3.47% in average above the top-ranked state-of-the-art methods.
Chun-Mei Feng 0001, Kai Yu 0009, Xinxing Xu, Salman Khan 0001, Rick Siow Mong Goh, Wangmeng Zuo, Yong Liu 0026
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Reliable Federated Disentangling Network for Non-IID Domain Feature
abstract
Federated Learning (FL), as an efficient decentralized distributed learning approach, enables multiple institutions to collaboratively train a model without sharing their local data. Despite its advantages, the performance of FL models is substantially impacted by the domain feature shift arising from different acquisition devices/clients. Moreover, existing FL methods often prioritize accuracy without considering reliability factors such as confidence or uncertainty, leading to unreliable predictions in safety-critical applications. Thus, our goal is to enhance FL performance by addressing non-domain feature issues and ensuring model reliability. In this study, we introduce a novel approach named RFedDis (Reliable Federated Disentangling Network). RFedDis leverages feature disentangling to capture a global domain-invariant cross-client representation while preserving local client-specific feature learning. Additionally, we incorporate an uncertainty-aware decision fusion mechanism to effectively integrate the decoupled features. This ensures dynamic integration at the evidence level, producing reliable predictions accompanied by estimated uncertainties. Therefore, RFedDis is the FL approach to combine evidential uncertainty with feature disentangling, enhancing both performance and reliability in handling non-IID domain features. Extensive experimental results demonstrate that RFedDis outperforms other state-of-the-art FL approaches, providing outstanding performance coupled with a high degree of reliability.
Meng Wang 0038, Kai Yu 0009, Chun-Mei Feng 0001, Yiming Qian, Ke Zou, Lianyu Wang, Rick Siow Mong Goh, Xinxing Xu, Yong Liu 0026, Huazhu Fu
IEEE Trans. Big Data2
2024 UrFound: Towards Universal Retinal Foundation Models via Knowledge-Guided Masked Modeling
Kai Yu 0009, Yang Zhou 0017, Yang Bai 0011, Zhi Da Soh, Xinxing Xu, Rick Siow Mong Goh, Ching Yu Cheng, Yong Liu 0026
MICCAI (12)1
2024 Exploring Separable Attention for Multi-Contrast MR Image Super-Resolution
abstract
Super-resolving the magnetic resonance (MR) image of a target contrast under the guidance of the corresponding auxiliary contrast, which provides additional anatomical information, is a new and effective solution for fast MR imaging. However, current multi-contrast super-resolution (SR) methods tend to concatenate different contrasts directly, ignoring their relationships in different clues, e.g., in the high-and low-intensity regions. In this study, we propose a separable attention network (comprising high-intensity priority (HP) attention and low-intensity separation (LS) attention), named SANet. Our SANet could explore the areas of high-and low-intensity regions in the "forward" and "reverse" directions with the help of the auxiliary contrast while learning clearer anatomical structure and edge information for the SR of a target-contrast MR image. SANet provides three appealing benefits: First, it is the first model to explore a separable attention mechanism that uses the auxiliary contrast to predict the high-and low-intensity regions, diverting more attention to refining any uncertain details between these regions and correcting the fine areas in the reconstructed results. Second, a multistage integration module is proposed to learn the response of multi-contrast fusion at multiple stages, get the dependency between the fused representations, and boost their representation ability. Third, extensive experiments with various state-of-the-art multi-contrast SR methods on fastMRI and clinical in vivo datasets demonstrate the superiority of our model. The code is released at https://github.com/chunmeifeng/SANet.
Chun-Mei Feng 0001, Yunlu Yan, Kai Yu 0009, Yong Xu 0001, Huazhu Fu, Jian Yang 0003, Ling Shao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Diverse Data Augmentation with Diffusions for Effective Test-time Prompt Tuning
abstract
Benefiting from prompt tuning, recent years have witnessed the promising performance of pre-trained vision-language models, e.g., CLIP, on versatile downstream tasks. In this paper, we focus on a particular setting of learning adaptive prompts on the fly for each test sample from an unseen new domain, which is known as test-time prompt tuning (TPT). Existing TPT methods typically rely on data augmentation and confidence selection. However, conventional data augmentation techniques, e.g., random resized crops, suffers from the lack of data diversity, while entropy-based confidence selection alone is not sufficient to guarantee prediction fidelity. To address these issues, we propose a novel TPT method, named DiffTPT, which leverages pre-trained diffusion models to generate diverse and informative new data. Specifically, we incorporate augmented data by both conventional method and pre-trained stable diffusion to exploit their respective merits, improving the model’s ability to adapt to unknown new test data. Moreover, to ensure the prediction fidelity of generated data, we introduce a cosine similarity-based filtration technique to select the generated data with higher similarity to the single test sample. Our experiments on test datasets with distribution shifts and unseen categories demonstrate that DiffTPT improves the zero-shot accuracy by an average of 5.13% compared to the state-of-the-art TPT method.
Chun-Mei Feng 0001, Kai Yu 0009, Yong Liu 0026, Salman Khan 0001, Wangmeng Zuo
ICCV2
2023 Towards Instance-adaptive Inference for Federated Learning
abstract
Federated learning (FL) is a distributed learning paradigm that enables multiple clients to learn a powerful global model by aggregating local training. However, the performance of the global model is often hampered by non-i.i.d. distribution among the clients, requiring extensive efforts to mitigate inter-client data heterogeneity. Going beyond inter-client data heterogeneity, we note that intra-client heterogeneity can also be observed on complex real-world data and seriously deteriorate FL performance. In this paper, we present a novel FL algorithm, i.e., FedIns, to handle intra-client data heterogeneity by enabling instance-adaptive inference in the FL framework. Instead of huge instance-adaptive models, we resort to a parameter-efficient fine-tuning method, i.e., scale and shift deep features (SSF), upon a pre-trained model. Specifically, we first train an SSF pool for each client, and aggregate these SSF pools on the server side, thus still maintaining a low communication cost. To enable instance-adaptive inference, for a given instance, we dynamically find the best-matched SSF subsets from the pool and aggregate them to generate an adaptive SSF specified for the instance, thereby reducing the intra-client as well as the inter-client heterogeneity. Extensive experiments show that our FedIns outperforms state-of-the-art FL algorithms, e.g., a 6.64% improvement against the top-performing method with less than 15% communication cost on Tiny-ImageNet.
Chun-Mei Feng 0001, Kai Yu 0009, Nian Liu 0002, Xinxing Xu, Salman Khan 0001, Wangmeng Zuo
ICCV2
2023 Graph Attention U-Net for Retinal Layer Surface Detection and Choroid Neovascularization Segmentation in OCT Images
abstract
Choroidal neovascularization (CNV) is a typical symptom of age-related macular degeneration (AMD) and is one of the leading causes for blindness. Accurate segmentation of CNV and detection of retinal layers are critical for eye disease diagnosis and monitoring. In this paper, we propose a novel graph attention U-Net (GA-UNet) for retinal layer surface detection and CNV segmentation in optical coherence tomography (OCT) images. Due to retinal layer deformation caused by CNV, it is challenging for existing models to segment CNV and detect retinal layer surfaces with the correct topological order. We propose two novel modules to address the challenge. The first module is a graph attention encoder (GAE) in a U-Net model that automatically integrates topological and pathological knowledge of retinal layers into the U-Net structure to achieve effective feature embedding. The second module is a graph decorrelation module (GDM) that takes reconstructed features by the decoder of the U-Net as inputs, it then decorrelates and removes information unrelated to retinal layer for improved retinal layer surface detection. In addition, we propose a new loss function to maintain the correct topological order of retinal layers and the continuity of their boundaries. The proposed model learns graph attention maps automatically during training and performs retinal layer surface detection and CNV segmentation simultaneously with the attention maps during inference. We evaluated the proposed model on our private AMD dataset and another public dataset. Experiment results show that the proposed model outperformed the competing methods for retinal layer surface detection and CNV segmentation and achieved new state of the arts on the datasets.
Yuhe Shen, Jiang Li 0001, Weifang Zhu, Kai Yu 0009, Meng Wang 0038, Yi Zhou 0024, Liling Guan, Xinjian Chen 0001
IEEE Trans. Medical Imaging4
2022 Speckle Noise Reduction for OCT Images Based on Image Style Transfer and Conditional GAN
abstract
Raw optical coherence tomography (OCT) images typically are of low quality because speckle noise blurs retinal structures, severely compromising visual quality and degrading performances of subsequent image analysis tasks. In our previous study (Ma et al., 2018), we have developed a Conditional Generative Adversarial Network (cGAN) for speckle noise removal in OCT images collected by several commercial OCT scanners, which we collectively refer to as scanner T. In this paper, we improve the cGAN model and apply it to our in-house OCT scanner (scanner B) for speckle noise suppression. The proposed model consists of two steps: 1) We train a Cycle-Consistent GAN (CycleGAN) to learn style transfer between two OCT image datasets collected by different scanners. The purpose of the CycleGAN is to leverage the ground truth dataset created in our previous study. 2) We train a mini-cGAN model based on the PatchGAN mechanism with the ground truth dataset to suppress speckle noise in OCT images. After training, we first apply the CycleGAN model to convert raw images collected by scanner B to match the style of the images from scanner T, and subsequently use the mini-cGAN model to suppress speckle noise in the style transferred images. We evaluate the proposed method on a dataset collected by scanner B. Experimental results show that the improved model outperforms our previous method and other state-of-the-art models in speckle noise removal, retinal structure preservation and contrast enhancement.
Yi Zhou 0024, Kai Yu 0009, Meng Wang 0038, Yuhui Ma, Zhongyue Chen, Weifang Zhu, Xinjian Chen 0001
IEEE J. Biomed. Health Informatics2
2022 MsTGANet: Automatic Drusen Segmentation From Retinal OCT Images
abstract
Drusen is considered as the landmark for diagnosis of AMD and important risk factor for the development of AMD. Therefore, accurate segmentation of drusen in retinal OCT images is crucial for early diagnosis of AMD. However, drusen segmentation in retinal OCT images is still very challenging due to the large variations in size and shape of drusen, blurred boundaries, and speckle noise interference. Moreover, the lack of OCT dataset with pixel-level annotation is also a vital factor hindering the improvement of drusen segmentation accuracy. To solve these problems, a novel multi-scale transformer global attention network (MsTGANet) is proposed for drusen segmentation in retinal OCT images. In MsTGANet, which is based on U-Shape architecture, a novel multi-scale transformer non-local (MsTNL) module is designed and inserted into the top of encoder path, aiming at capturing multi-scale non-local features with long-range dependencies from different layers of encoder. Meanwhile, a novel multi-semantic global channel and spatial joint attention module (MsGCS) between encoder and decoder is proposed to guide the model to fuse different semantic features, thereby improving the model's ability to learn multi-semantic global contextual information. Furthermore, to alleviate the shortage of labeled data, we propose a novel semi-supervised version of MsTGANet (Semi-MsTGANet) based on pseudo-labeled data augmentation strategy, which can leverage a large amount of unlabeled data to further improve the segmentation performance. Finally, comprehensive experiments are conducted to evaluate the performance of the proposed MsTGANet and Semi-MsTGANet. The experimental results show that our proposed methods achieve better segmentation accuracy than other state-of-the-art CNN-based methods.
Meng Wang 0038, Weifang Zhu, Jinzhu Su, Haoyu Chen 0002, Kai Yu 0009, Yi Zhou 0024, Zhongyue Chen, Xinjian Chen 0001
IEEE Trans. Medical Imaging6
2021 Automatic Staging for Retinopathy of Prematurity With Deep Feature Fusion and Ordinal Classification Strategy
abstract
Retinopathy of prematurity (ROP) is a retinal disease which frequently occurs in premature babies with low birth weight and is considered as one of the major preventable causes of childhood blindness. Although automatic and semi-automatic diagnoses of ROP based on fundus image have been researched, most of the previous studies focused on plus disease detection and ROP screening. There are few studies focusing on ROP staging, which is important for the severity evaluation of the disease. To be consistent with clinical 5-level ROP staging, a novel and effective deep neural network based 5-level ROP staging network is proposed, which consists of multi-stream based parallel feature extractor, concatenation based deep feature fuser and clinical practice based ordinal classifier. First, the three-stream parallel framework including ResNet18, DenseNet121 and EfficientNetB2 is proposed as the feature extractor, which can extract rich and diverse high-level features. Second, the features from three streams are deeply fused by concatenation and convolution to generate a more effective and comprehensive feature. Finally, in the classification stage, an ordinal classification strategy is adopted, which can effectively improve the ROP staging performance. The proposed ROP staging network was evaluated with per-image and per-examination strategies. For per-image ROP staging, the proposed method was evaluated on 635 retinal fundus images from 196 examinations, including 303 Normal, 26 Stage 1, 127 Stage 2, 106 Stage 3, 61 Stage 4 and 12 Stage 5, which achieves 0.9055 for weighted recall, 0.9092 for weighted precision, 0.9043 for weighted F1 score, 0.9827 for accuracy with 1 (ACC1) and 0.9786 for Kappa, respectively. While for per-examination ROP staging, 1173 examinations with a 4-fold cross validation strategy were used to evaluate the effectiveness of the proposed method, which prove the validity and advantage of the proposed method.
Weifang Zhu, Zhongyue Chen, Meng Wang 0038, Le Geng, Kai Yu 0009, Yi Zhou 0024, Daoman Xiang, Xinjian Chen 0001
IEEE Trans. Medical Imaging6
2021 Semi-Supervised Capsule cGAN for Speckle Noise Reduction in Retinal OCT Images
abstract
Speckle noise is the main cause of poor optical coherence tomography (OCT) image quality. Convolutional neural networks (CNNs) have shown remarkable performances for speckle noise reduction. However, speckle noise denoising still meets great challenges because the deep learning-based methods need a large amount of labeled data whose acquisition is time-consuming or expensive. Besides, many CNNs-based methods design complex structure based networks with lots of parameters to improve the denoising performance, which consume hardware resources severely and are prone to overfitting. To solve these problems, we propose a novel semi-supervised learning based method for speckle noise denoising in retinal OCT images. First, to improve the model's ability to capture complex and sparse features in OCT images, and avoid the problem of a great increase of parameters, a novel capsule conditional generative adversarial network (Caps-cGAN) with small number of parameters is proposed to construct the semi-supervised learning system. Then, to tackle the problem of retinal structure information loss in OCT images caused by lack of detailed guidance during unsupervised learning, a novel joint semi-supervised loss function composed of unsupervised loss and supervised loss is proposed to train the model. Compared with other state-of-the-art methods, the proposed semi-supervised method is suitable for retinal OCT images collected from different OCT devices and can achieve better performance even only using half of the training data.
Meng Wang 0038, Weifang Zhu, Kai Yu 0009, Zhongyue Chen, Yi Zhou 0024, Yuhui Ma, Dengsen Bao, Shuanglang Feng, Dehui Xiang, Xinjian Chen 0001
IEEE Trans. Medical Imaging3
2020 OCTRexpert: A Feature-Based 3D Registration Method for Retinal OCT Images
abstract
Medical image registration can be used for studying longitudinal and cross-sectional data, quantitatively monitoring disease progression and guiding computer assisted diagnosis and treatments. However, deformable registration which enables more precise and quantitative comparison has not been well developed for retinal optical coherence tomography (OCT) images. This paper proposes a new 3D registration approach for retinal OCT data called OCTRexpert. To the best of our knowledge, the proposed algorithm is the first full 3D registration approach for retinal OCT images which can be applied to longitudinal OCT images for both normal and serious pathological subjects. In this approach, a pre-processing method is first performed to remove eye motion artifact and then a novel design-detection-deformation strategy is applied for the registration. In the design step, a couple of features are designed for each voxel in the image. In the detection step, active voxels are selected and the point-to-point correspondences between the subject and template images are established. In the deformation step, the image is hierarchically deformed according to the detected correspondences in multi-resolution. The proposed method is evaluated on a dataset with longitudinal OCT images from 20 healthy subjects and 4 subjects diagnosed with serious Choroidal Neovascularization (CNV). Experimental results show that the proposed registration algorithm consistently yields statistically significant improvements in both Dice similarity coefficient and the average unsigned surface error compared with the other registration methods.
Lingjiao Pan, Dehui Xiang, Kai Yu 0009, Luwen Duan, Jian Zheng 0001, Xinjian Chen 0001
IEEE Trans. Image Process.4
2019 Surrogate-Assisted Retinal OCT Image Classification Based on Convolutional Neural Networks
abstract
Optical Coherence Tomography (OCT) is beco-ming one of the most important modalities for the noninvasive assessment of retinal eye diseases. As the number of acquired OCT volumes increases, automating the OCT image analysis is becoming increasingly relevant. In this paper, we propose a surrogate-assisted classification method to classify retinal OCT images automatically based on convolutional neural networks (CNNs). Image denoising is first performed to reduce the noise. Thresholding and morphological dilation are applied to extract the masks. The denoised images and the masks are then employed to generate a lot of surrogate images, which are used to train the CNN model. Finally, the prediction for a test image is determined by the average of the outputs from the trained CNN model on the surrogate images. The proposed method has been evaluated on different databases. The results (AUC of 0.9783 in the local database and AUC of 0.9856 in the Duke database) show that the proposed method is a very promising tool for classifying the retinal OCT images automatically.
Yibiao Rong, Dehui Xiang, Weifang Zhu, Kai Yu 0009, Zhun Fan, Xinjian Chen 0001
IEEE J. Biomed. Health Informatics4
2017 Single-Channel Sparse Non-Negative Blind Source Separation Method for Automatic 3-D Delineation of Lung Tumor in PET Images
abstract
In this paper, we propose a novel method for single-channel blind separation of nonoverlapped sources and, to the best of our knowledge, apply it for the first time to automatic segmentation of lung tumors in positron emission tomography (PET) images. Our approach first converts a 3-D PET image into a pseudo-multichannel image. Afterward, regularization free sparseness constrained non-negative matrix factorization is used to separate tumor from other tissues. By using complexity based criterion, we select tumor component as the one with minimal complexity. We have compared the proposed method with threshold based on 40% and 50% maximum standardized uptake value (SUV), graph cuts (GC), random walks (RW), and affinity propagation (AP) algorithms on 18 nonsmall cell lung cancer datasets with respect to ground truth (GT) provided by two radiologists. Dice similarity coefficient averaged with respect to two GTs is: 0.78 ± 0.12 by the proposed algorithm, 0.78 ± 0.1 by GC, 0.77 ± 0.13 by AP, 0.77 ± 0.07 by RW, and 0.75 ± 0.13 by 50% maximum SUV threshold. Since the proposed method achieved performance comparable with interactive methods, considering the unique challenges of lung tumor segmentation from PET images, our findings support possibility of using our fully automated method in routine clinics. The source codes will be available at www.mipav.net/English/research/research.html.
Ivica Kopriva, Wei Ju 0002, Bin Zhang 0049, Dehui Xiang, Kai Yu 0009, Ximing Wang, Ulas Bagci, Xinjian Chen 0001
IEEE J. Biomed. Health Informatics6