Luyang Luo

dblp:242/9268 · DBLP profile ↗
← Back
30ranked-venue papers
7as first author
24since 2021 · last 2026
0000-0002-7485-4151ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 27 · 6 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Knowledge-Enhanced Explainable Prompting for Vision-Language Models
abstract
Large-scale vision-language models (VLMs) embedded with expansive representations and visual concepts have showcased significant potential in image and text understanding. Efficiently adapting VLMs such as CLIP to downstream tasks like few-shot image classification has garnered growing attention, with prompt learning emerging as a representative approach. However, most existing prompt-based adaptation methods, which rely solely on coarse-grained textual prompts, suffer from limited performance and interpretability when handling domain tasks that require specific knowledge. This results in a failure to satisfy the stringent trustworthiness requirements of Explainable Artificial Intelligence (XAI) in high-risk scenarios like healthcare. To address this issue, we propose a Knowledge-Enhanced Explainable Prompting (KEEP) framework that leverages fine-grained domain-specific knowledge to enhance the adaptation process of VLMs across various domains and image modalities. By incorporating retrieval augmented generation and domain foundation models, our framework can provide more reliable image-wise knowledge for prompt learning in various domains, alleviating the lack of fine-grained annotations, while offering both visual and textual explanations. Extensive experiments and explainability analyses conducted on eight datasets of different domains and image modalities demonstrate that our method simultaneously achieves superior performance and interpretability, highlighting the effectiveness of the collaboration between foundation models and XAI.
Yequan Bie, Andong Tan, Zhixuan Chen, Zhiyuan Cai, Luyang Luo, Hao Chen 0011
AAAI5
2026 Learning with less supervision: A survey of label-efficient learning for medical image analysis
Cheng Jin 0003, Zhengrui Guo, Yi Lin 0009, Luyang Luo, Hao Chen 0011
Medical Image Anal.4
2026 SurgPETL: Parameter-Efficient Image-to-Surgical-Video Transfer Learning for Surgical Phase Recognition
abstract
Capitalizing on image-level pre-trained models for various downstream tasks has recently emerged with promising performance. However, the paradigm of "image pre-training followed by video fine-tuning" for high-dimensional video data inevitably introduces significant performance bottlenecks. Furthermore, in the medical domain, many surgical video tasks encounter additional challenges posed by the limited availability of video data and the necessity for comprehensive spatiotemporal modeling. Recently, Parameter-Efficient Image-to-Video Transfer Learning (PEIVTL) has emerged as an efficient and effective paradigm for video action recognition tasks, which employs image-level pre-trained models with promising feature transferability and involves cross-modality temporal modeling with minimal fine-tuning. Nevertheless, the effectiveness and generalizability of this paradigm within intricate surgical domain remain unexplored. In this paper, we delve into a novel problem of efficiently adapting image-level pre-trained models to specialize in fine-grained surgical phase recognition, termed Parameter-Efficient Image-to-Surgical-Video Transfer Learning. First, we develop SurgPETL, a parameter-efficient transfer learning framework for surgical phase recognition, and conduct extensive experiments with three advanced methods based on ViTs of two distinct scales pre-trained on five large-scale natural and medical datasets. Then, we introduce the Adaptive Spatiotemporal Representation Modulation (ASRM) module, integrating a standard spatial adapter with a novel temporal adapter to capture detailed spatial features and establish connections across temporal sequences for robust spatiotemporal modeling. Extensive experiments on three challenging datasets spanning various surgical procedures demonstrate the effectiveness of SurgPETL with ASRM. SurgPETL-ASRM outperforms both parameter-efficient alternatives and state-of-the-art surgical phase recognition methods while maintaining parameter efficiency and minimizing overhead.
Shu Yang 0004, Zhiyuan Cai, Luyang Luo, Shuchang Xu, Hao Chen 0011
IEEE Trans. Medical Imaging3
2025 Dia-LLaMA: Towards Large Language Model-Driven CT Report Generation
Zhixuan Chen, Luyang Luo, Yequan Bie, Hao Chen 0011
MICCAI (7)2
2025 Mitigating medical dataset bias by learning adaptive agreement from a biased council
Luyang Luo, Zhuoyue Wan, Wanteng Ma, Hao Chen 0011
Medical Image Anal.1
2025 Learning robust medical image segmentation from multi-source annotations
Luyang Luo, Mingxiang Wu, Qiong Wang 0001, Hao Chen 0011
Medical Image Anal.2
2025 HMIL: Hierarchical Multi-Instance Learning for Fine-Grained Whole Slide Image Classification
abstract
Fine-grained classification of whole slide images (WSIs) is essential in precision oncology, enabling precise cancer diagnosis and personalized treatment strategies. The core of this task involves distinguishing subtle morphological variations within the same broad category of gigapixel-resolution images, which presents a significant challenge. While the multi-instance learning (MIL) paradigm alleviates the computational burden of WSIs, existing MIL methods often overlook hierarchical label correlations, treating fine-grained classification as a flat multi-class classification task. To overcome these limitations, we introduce a novel hierarchical multi-instance learning (HMIL) framework. By facilitating on the hierarchical alignment of inherent relationships between different hierarchy of labels at instance and bag level, our approach provides a more structured and informative learning process. Specifically, HMIL incorporates a class-wise attention mechanism that aligns hierarchical information at both the instance and bag levels. Furthermore, we introduce supervised contrastive learning to enhance the discriminative capability for fine-grained classification and a curriculum-based dynamic weighting module to adaptively balance the hierarchical feature during training. Extensive experiments on our large-scale cytology cervical cancer (CCC) dataset and two public histology datasets, BRACS and PANDA, demonstrate the state-of-the-art class-wise and overall performance of our HMIL framework. Our source code is available at https://github.com/ChengJin-git/HMIL.
Cheng Jin 0003, Luyang Luo, Huangjing Lin, Hao Chen 0011
IEEE Trans. Medical Imaging2
2025 Advancing Volumetric Medical Image Segmentation via Global-Local Masked Autoencoders
abstract
Masked Autoencoder (MAE) is a self-supervised pre-training technique that holds promise in improving the representation learning of neural networks. However, the current application of MAE directly to volumetric medical images poses two challenges: (i) insufficient global information for clinical context understanding of the holistic data, and (ii) the absence of any assurance of stabilizing the representations learned from randomly masked inputs. To conquer these limitations, we propose the Global-Local Masked AutoEncoders (GL-MAE), a simple yet effective self-supervised pre-training strategy. GL-MAE acquires robust anatomical structure features by incorporating multi-level reconstruction from fine-grained local details to high-level global semantics. Furthermore, a complete global view serves as an anchor to direct anatomical semantic alignment and stabilize the learning process through global-to-global consistency learning and global-to-local consistency learning. Our fine-tuning results on eight mainstream public datasets demonstrate the superiority of our method over other state-of-the-art self-supervised algorithms, highlighting its effectiveness on versatile volumetric medical image segmentation and classification tasks. We will release codes upon acceptance at https://github.com/JiaxinZhuang/GL-MAE.
Jiaxin Zhuang, Luyang Luo, Qiong Wang 0001, Mingxiang Wu, Hao Chen 0011
IEEE Trans. Medical Imaging2
2025 Scale-Aware Super-Resolution Network With Dual Affinity Learning for Lesion Segmentation From Medical Images
abstract
Convolutional neural networks (CNNs) have shown remarkable progress in medical image segmentation. However, the lesion segmentation remains a challenge to state-of-the-art CNN-based algorithms due to the variance in scales and shapes. On the one hand, tiny lesions are hard to delineate precisely from the medical images which are often of low resolutions. On the other hand, segmenting large-size lesions requires large receptive fields, which exacerbates the first challenge. In this article, we present a scale-aware super-resolution (SR) network to adaptively segment lesions of various sizes from low-resolution (LR) medical images. Our proposed network contains dual branches to simultaneously conduct lesion mask SR (LMSR) and lesion image SR (LISR). Meanwhile, we introduce scale-aware dilated convolution (SDC) blocks into the multitask decoders to adaptively adjust the receptive fields of the convolutional kernels according to the lesion sizes. To guide the segmentation branch to learn from richer high-resolution (HR) features, we propose a feature affinity (FA) module and a scale affinity (SA) module to enhance the multitask learning of the dual branches. On multiple challenging lesion segmentation datasets, our proposed network achieved consistent improvements compared with other state-of-the-art methods. Code will be available at: https://github.com/poiuohke/SASR_Net.
Luyang Luo, Yanwen Li, Zhizhong Chai, Huangjing Lin, Pheng-Ann Heng, Hao Chen 0011
IEEE Trans. Neural Networks Learn. Syst.1
2024 MICA: Towards Explainable Skin Lesion Diagnosis via Multi-Level Image-Concept Alignment
abstract
Black-box deep learning approaches have showcased significant potential in the realm of medical image analysis. However, the stringent trustworthiness requirements intrinsic to the medical field have catalyzed research into the utilization of Explainable Artificial Intelligence (XAI), with a particular focus on concept-based methods. Existing concept-based methods predominantly apply concept annotations from a single perspective (e.g., global level), neglecting the nuanced semantic relationships between sub-regions and concepts embedded within medical images. This leads to underutilization of the valuable medical information and may cause models to fall short in harmoniously balancing interpretability and performance when employing inherently interpretable architectures such as Concept Bottlenecks. To mitigate these shortcomings, we propose a multi-modal explainable disease diagnosis framework that meticulously aligns medical images and clinical-related concepts semantically at multiple strata, encompassing the image level, token level, and concept level. Moreover, our method allows for model intervention and offers both textual and visual explanations in terms of human-interpretable concepts. Experimental results on three skin image datasets demonstrate that our method, while preserving model interpretability, attains high performance and label efficiency for concept detection and disease diagnosis. The code is available at https://github.com/Tommy-Bie/MICA.
Yequan Bie, Luyang Luo, Hao Chen 0011
AAAI2
2024 Bootstrapping Radiography Pre-training via Siamese Masked Vision-Language Modeling with Complementary Self-distillation
abstract
Diagnosing thoracic diseases from chest X-rays (CXR) using deep learning faces unique challenges due to the high anatomical similarity across images and the critical nature of minute anomalies. In this paper, we introduce a novel self-supervised learning framework, Siamese Masked Vision-Language Modeling with Complementary Self-distillation (SMVLM), designed to enhance disease diagnosis in CXR by addressing these specific challenges. First, we tackle the problem of high anatomical similarity and subtle variance in CXR by employing a complementary masking strategy in a Siamese network setup, which forces the model to focus on subtle, yet diagnostically relevant features from both global and local perspectives. Secondly, we enrich our model’s learning capabilities by integrating a dual masked image-text contrastive loss that aligns radiographic findings with their corresponding text report, harnessing the synergistic potential of multimodal data. In conjunction with this, cross-modal image-text pre-reconstruction with registers is introduced to deepen the contextual understanding of the CXR images and radiology report, ensuring a comprehensive feature representation. Extensive evaluations on benchmark datasets demonstrate that our method significantly outperforms existing approaches, providing a robust solution for the accurate and reliable diagnosis of thoracic diseases from CXR images.
Luyang Luo, Hao Chen 0011
BIBM2
2024 XCoOp: Explainable Prompt Learning for Computer-Aided Diagnosis via Concept-Guided Context Optimization
Yequan Bie, Luyang Luo, Zhixuan Chen, Hao Chen 0011
MICCAI (12)2
2024 Enable the Right to be Forgotten with Federated Client Unlearning in Medical Imaging
Zhipeng Deng, Luyang Luo, Hao Chen 0011
MICCAI (10)2
2024 Surgformer: Surgical Transformer with Hierarchical Temporal Attention for Surgical Phase Recognition
Shu Yang 0004, Luyang Luo, Qiong Wang 0001, Hao Chen 0011
MICCAI (6)2
2024 Guest Editorial: Trustworthy Machine Learning for Health Informatics
abstract
Machine learning (ML), the stem of today's artificial intelligence, has shown significant growth in the field of biomedical and health informatics. On the one hand, ML techniques are becoming more complex in order to deal with real-world data. On the other hand, ML is also more and more accessible to broader users. For example, automated machine learning products are enabling users to build their own ML models without writing code [1].
Luyang Luo, Daguang Xu, Harry Qin, Yueming Jin, Hao Chen 0011
IEEE J. Biomed. Health Informatics1
2024 Deep Omni-Supervised Learning for Rib Fracture Detection From Chest Radiology Images
abstract
Deep learning (DL)-based rib fracture detection has shown promise of playing an important role in preventing mortality and improving patient outcome. Normally, developing DL-based object detection models requires a huge amount of bounding box annotation. However, annotating medical data is time-consuming and expertise-demanding, making obtaining a large amount of fine-grained annotations extremely infeasible. This poses a pressing need for developing label-efficient detection models to alleviate radiologists' labeling burden. To tackle this challenge, the literature on object detection has witnessed an increase of weakly-supervised and semi-supervised approaches, yet still lacks a unified framework that leverages various forms of fully-labeled, weakly-labeled, and unlabeled data. In this paper, we present a novel omni-supervised object detection network, ORF-Netv2, to leverage as much available supervision as possible. Specifically, a multi-branch omni-supervised detection head is introduced with each branch trained with a specific type of supervision. A co-training-based dynamic label assignment strategy is then proposed to enable flexible and robust learning from the weakly-labeled and unlabeled data. Extensive evaluation was conducted for the proposed framework with three rib fracture datasets on both chest CT and X-ray. By leveraging all forms of supervision, ORF-Netv2 achieves mAPs of 34.7, 44.7, and 19.4 on the three datasets, respectively, surpassing the baseline detector which uses only box annotations by mAP gains of 3.8, 4.8, and 5.0, respectively. Furthermore, ORF-Netv2 consistently outperforms other competitive label-efficient methods over various scenarios, showing a promising framework for label-efficient fracture detection. The code is available at: https://github.com/zhizhongchai/ORF-Net.
Zhizhong Chai, Luyang Luo, Huangjing Lin, Pheng-Ann Heng, Hao Chen 0011
IEEE Trans. Medical Imaging2
2024 Rethinking Multiple Instance Learning for Whole Slide Image Classification: A Bag-Level Classifier is a Good Instance-Level Teacher
abstract
Multiple Instance Learning (MIL) has demonstrated promise in Whole Slide Image (WSI) classification. However, a major challenge persists due to the high computational cost associated with processing these gigapixel images. Existing methods generally adopt a two-stage approach, comprising a non-learnable feature embedding stage and a classifier training stage. Though it can greatly reduce memory consumption by using a fixed feature embedder pre-trained on other domains, such a scheme also results in a disparity between the two stages, leading to suboptimal classification accuracy. To address this issue, we propose that a bag-level classifier can be a good instance-level teacher. Based on this idea, we design Iteratively Coupled Multiple Instance Learning (ICMIL) to couple the embedder and the bag classifier at a low cost. ICMIL initially fixes the patch embedder to train the bag classifier, followed by fixing the bag classifier to fine-tune the patch embedder. The refined embedder can then generate better representations in return, leading to a more accurate classifier for the next iteration. To realize more flexible and more effective embedder fine-tuning, we also introduce a teacher-student framework to efficiently distill the category knowledge in the bag classifier to help the instance-level embedder fine-tuning. Intensive experiments were conducted on four distinct datasets to validate the effectiveness of ICMIL. The experimental results consistently demonstrated that our method significantly improves the performance of existing MIL backbones, achieving state-of-the-art results. The code and the organized datasets can be accessed by: https://github.com/Dootmaan/ICMIL/tree/confidence-based.
Hongyi Wang 0002, Luyang Luo, Fang Wang 0030, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hongjie Hu, Lanfen Lin, Hao Chen 0011
IEEE Trans. Medical Imaging2
2023 Scale Federated Learning for Label Set Mismatch in Medical Image Classification
Zhipeng Deng, Luyang Luo, Hao Chen 0011
MICCAI (3)2
2023 Iteratively Coupled Multiple Instance Learning from Instance to Bag Classifier for Whole Slide Image Classification
Hongyi Wang 0002, Luyang Luo, Fang Wang 0030, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hongjie Hu, Lanfen Lin, Hao Chen 0011
MICCAI (6)2
2023 Triplet attention and dual-pool contrastive learning for clinic-driven multi-label medical image classification
Yuhan Zhang 0001, Luyang Luo, Qi Dou 0001, Pheng-Ann Heng
Medical Image Anal.2
2022 ORF-Net: Deep Omni-Supervised Rib Fracture Detection from Chest CT Scans
Zhizhong Chai, Huangjing Lin, Luyang Luo, Pheng-Ann Heng, Hao Chen 0011
MICCAI (3)3
2022 Pseudo Bias-Balanced Learning for Debiased Chest X-Ray Classification
Luyang Luo, Dunyuan Xu, Hao Chen 0011, Tien-Tsin Wong, Pheng-Ann Heng
MICCAI (8)1
2021 Dual-Consistency Semi-supervised Learning with Uncertainty Quantification for COVID-19 Lesion Segmentation from CT Images
Yanwen Li, Luyang Luo, Huangjing Lin, Hao Chen 0011, Pheng-Ann Heng
MICCAI (2)2
2021 OXnet: Deep Omni-Supervised Thoracic Disease Detection from Chest X-Rays
Luyang Luo, Hao Chen 0011, Yanning Zhou 0001, Huangjing Lin, Pheng-Ann Heng
MICCAI (2)1
2020 Towards multi-center glaucoma OCT image screening with semi-supervised joint structure and function multi-task learning
Xi Wang 0013, Hao Chen 0011, An-ran Ran, Luyang Luo, Poemen P. Chan, Clement C. Tham, Robert T. Chang, Suria S. Mannil, Carol Y. Cheung, Pheng-Ann Heng
Medical Image Anal.4
2020 UD-MIL: Uncertainty-Driven Deep Multiple Instance Learning for OCT Image Classification
abstract
Deep learning has achieved remarkable success in the optical coherence tomography (OCT) image classification task with substantial labelled B-scan images available. However, obtaining such fine-grained expert annotations is usually quite difficult and expensive. How to leverage the volume-level labels to develop a robust classifier is very appealing. In this paper, we propose a weakly supervised deep learning framework with uncertainty estimation to address the macula-related disease classification problem from OCT images with the only volume-level label being available. First, a convolutional neural network (CNN) based instance-level classifier is iteratively refined by using the proposed uncertainty-driven deep multiple instance learning scheme. To our best knowledge, we are the first to incorporate the uncertainty evaluation mechanism into multiple instance learning (MIL) for training a robust instance classifier. The classifier is able to detect suspicious abnormal instances and abstract the corresponding deep embedding with high representation capability simultaneously. Second, a recurrent neural network (RNN) takes instance features from the same bag as input and generates the final bag-level prediction by considering the individually local instance information and globally aggregated bag-level representation. For more comprehensive validation, we built two large diabetic macular edema (DME) OCT datasets from different devices and imaging protocols to evaluate the efficacy of our method, which are composed of 30,151 B-scans in 1,396 volumes from 274 patients (Heidelberg-DME dataset) and 38,976 B-scans in 3,248 volumes from 490 patients (Triton-DME dataset), respectively. We compare the proposed method with the state-of-the-art approaches, and experimentally demonstrate that our method is superior to alternative methods, achieving volume-level accuracy, F1-score and area under the receiver operating characteristic curve (AUC) of 95.1%, 0.939 and 0.990 on Heidelberg-DME and those of 95.1%, 0.935 and 0.986 on Triton-DME, respectively. Furthermore, the proposed method also yields competitive results on another public age-related macular degeneration OCT dataset, indicating the high potential as an effective screening tool in the clinical practice.
Xi Wang 0013, Fangyao Tang, Hao Chen 0011, Luyang Luo, Ziqi Tang, An-ran Ran, Carol Y. Cheung, Pheng-Ann Heng
IEEE J. Biomed. Health Informatics4
2020 Semi-Supervised Medical Image Classification With Relation-Driven Self-Ensembling Model
abstract
Training deep neural networks usually requires a large amount of labeled data to obtain good performance. However, in medical image analysis, obtaining high-quality labels for the data is laborious and expensive, as accurately annotating medical images demands expertise knowledge of the clinicians. In this paper, we present a novel relation-driven semi-supervised framework for medical image classification. It is a consistency-based method which exploits the unlabeled data by encouraging the prediction consistency of given input under perturbations, and leverages a self-ensembling model to produce high-quality consistency targets for the unlabeled data. Considering that human diagnosis often refers to previous analogous cases to make reliable decisions, we introduce a novel sample relation consistency (SRC) paradigm to effectively exploit unlabeled data by modeling the relationship information among different samples. Superior to existing consistency-based methods which simply enforce consistency of individual predictions, our framework explicitly enforces the consistency of semantic relation among different samples under perturbations, encouraging the model to explore extra semantic information from unlabeled data. We have conducted extensive experiments to evaluate our method on two public benchmark medical image classification datasets, i.e., skin lesion diagnosis with ISIC 2018 challenge and thorax disease classification with ChestX-ray14. Our method outperforms many state-of-the-art semi-supervised learning methods on both single-label and multi-label image classification scenarios.
Quande Liu, Lequan Yu, Luyang Luo, Qi Dou 0001, Pheng-Ann Heng
IEEE Trans. Medical Imaging3
2020 Deep Mining External Imperfect Data for Chest X-Ray Disease Screening
abstract
Deep learning approaches have demonstrated remarkable progress in automatic Chest X-ray analysis. The data-driven feature of deep models requires training data to cover a large distribution. Therefore, it is substantial to integrate knowledge from multiple datasets, especially for medical images. However, learning a disease classification model with extra Chest X-ray (CXR) data is yet challenging. Recent researches have demonstrated that performance bottleneck exists in joint training on different CXR datasets, and few made efforts to address the obstacle. In this paper, we argue that incorporating an external CXR dataset leads to imperfect training data, which raises the challenges. Specifically, the imperfect data is in two folds: domain discrepancy, as the image appearances vary across datasets; and label discrepancy, as different datasets are partially labeled. To this end, we formulate the multi-label thoracic disease classification problem as weighted independent binary tasks according to the categories. For common categories shared across domains, we adopt task-specific adversarial training to alleviate the feature differences. For categories existing in a single dataset, we present uncertainty-aware temporal ensembling of model predictions to mine the information from the missing labels further. In this way, our framework simultaneously models and tackles the domain and label discrepancies, enabling superior knowledge mining ability. We conduct extensive experiments on three datasets with more than 360,000 Chest X-ray images. Our method outperforms other competing models and sets state-of-the-art performance on the official NIH test set with 0.8349 AUC, demonstrating its effectiveness of utilizing the external dataset to improve the internal classification.
Luyang Luo, Lequan Yu, Hao Chen 0011, Quande Liu, Xi Wang 0013, Pheng-Ann Heng
IEEE Trans. Medical Imaging1
2019 Deep Angular Embedding and Feature Correlation Attention for Breast MRI Cancer Analysis
Luyang Luo, Hao Chen 0011, Xi Wang 0013, Qi Dou 0001, Huangjing Lin, Gongjie Li, Pheng-Ann Heng
MICCAI (4)1
2019 Unifying Structure Analysis and Surrogate-Driven Function Regression for Glaucoma OCT Image Screening
Xi Wang 0013, Hao Chen 0011, Luyang Luo, An-ran Ran, Poemen P. Chan, Clement C. Tham, Carol Y. Cheung, Pheng-Ann Heng
MICCAI (1)3