VLDB 2026 Research / reviewers in the wild / expert
Linhao Qu
dblp:308/1001
· DBLP profile ↗
25ranked-venue papers
11as first author
25since 2021 · last 2026
0000-0001-8815-7050ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 7 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing zero-shot brain tumor subtype classification via fine-grained patch-text alignment
Lubin Gan, Jing Zhang 0165, Linhao Qu, Siying Wu, Xiaoyan Sun 0001 |
Expert Syst. Appl. | 3 |
| 2026 | Deep Mutual Learning Among Partially Labeled Datasets for Multi-Organ SegmentationabstractLabeling multiple organs for segmentation is a complex and time-consuming process, resulting in a scarcity of comprehensively labeled multi-organ datasets while the emergence of numerous partially labeled datasets. Current methods face three critical limitations: incomplete exploitation of available supervision; complex inference, and insufficient validation of generalization capabilities. This paper proposes a new framework based on mutual learning, aiming to improve multi-organ segmentation performance by complementing information among partially labeled datasets. Specifically, this method consists of three key components: 1) partial-organ segmentation models training with Difference Mutual Learning, 2) pseudo-label generation and filtering, and 3) full-organ segmentation models training enhanced by Similarity Mutual Learning. Difference Mutual Learning enables each partial-organ segmentation model to utilize labels and features from other datasets as complementary signals, improving cross-dataset organ detection for better pseudo labels. Similarity Mutual Learning augments each full-organ segmentation model training with two additional supervision sources: inter-dataset ground truths and dynamic reliable transferred features, significantly boosting segmentation accuracy. The model obtained by this method achieves both high accuracy and efficient inference for multi-organ segmentation. Extensive experiments conducted on nine datasets spanning the head-neck, chest, abdomen, and pelvis demonstrate that the proposed method achieves SOTA performance. Linhao Qu, Ziyue Xie, Yonghong Shi, Zhijian Song |
IEEE Trans. Medical Imaging | 2 |
| 2025 | FANCL: Feature-guided attention network with curriculum learning for brain metastases segmentation
Zijiang Liu, Linhao Qu, Yonghong Shi |
Neurocomputing | 3 |
| 2025 | MSCPT: Few-Shot Whole Slide Image Classification With Multi-Scale and Context-Focused Prompt TuningabstractMultiple instance learning (MIL) has become a standard paradigm for the weakly supervised classification of whole slide images (WSIs). However, this paradigm relies on using a large number of labeled WSIs for training. The lack of training data and the presence of rare diseases pose significant challenges for these methods. Prompt tuning combined with pre-trained Vision-Language models (VLMs) is an effective solution to the Few-shot Weakly Supervised WSI Classification (FSWC) task. Nevertheless, applying prompt tuning methods designed for natural images to WSIs presents three significant challenges: 1) These methods fail to fully leverage the prior knowledge from the VLM's text modality; 2) They overlook the essential multi-scale and contextual information in WSIs, leading to suboptimal results; and 3) They lack exploration of instance aggregation methods. To address these problems, we propose a Multi-Scale and Context-focused Prompt Tuning (MSCPT) method for FSWC task. Specifically, MSCPT employs the frozen large language model to generate pathological visual language prior knowledge at multiple scales, guiding hierarchical prompt tuning. Additionally, we design a graph prompt tuning module to learn essential contextual information within WSI, and finally, a non-parametric cross-guided instance aggregation module has been introduced to derive the WSI-level features. Extensive experiments, visualizations, and interpretability analyses were conducted on five datasets and three downstream tasks using three VLMs, demonstrating the strong performance of our MSCPT. All codes have been made publicly accessible at https://github.com/Hanminghao/MSCPT. Minghao Han, Linhao Qu, Dingkang Yang, Xukun Zhang, Lihua Zhang 0002 |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Local Implicit Wavelet Transformer for Arbitrary-Scale Super-Resolution
Minghong Duan, Linhao Qu, Shaolei Liu, Manning Wang |
BMVC | 2 |
| 2024 | Separate and Conquer: Decoupling Co-occurrence via Decomposition and Representation for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation (WSSS) with image-level labels aims to achieve segmentation tasks with-out dense annotations. However, attributed to the frequent coupling of co-occurring objects and the limited supervision from image-level labels, the challenging co-occurrence problem is widely present and leads to false activation of objects in WSSS. In this work, we devise a ‘Separate and Conquer’ scheme SeCo to tackle this issue from di-mensions of image space and feature space. In the im-age space, we propose to ‘separate’ the co-occurring ob-jects with image decomposition by subdividing images into patches. Importantly, we assign each patch a category tag from Class Activation Maps (CAMs), which spatially helps remove the co-context bias and guide the subsequent rep-resentation. In the feature space, we propose to ‘conquer’ the false activation by enhancing semantic representation with multi-granularity knowledge contrast. To this end, a dual-teacher-single-student architecture is designed and tag-guided contrast is conducted, which guarantee the cor-rectness of knowledge and further facilitate the discrepancy among co-contexts. We streamline the multi-staged WSSS pipeline end-to-end and tackle this issue without external supervision. Extensive experiments are conducted, validating the efficiency of our method and the superiority over previous single-staged and even multi-staged competitors on PASCAL VOC and MS COCO. Code is available here. Kexue Fu 0001, Minghong Duan, Linhao Qu, Shuo Wang 0011, Zhijian Song |
CVPR | 4 |
| 2024 | Pathology-Knowledge Enhanced Multi-instance Prompt Learning for Few-Shot Whole Slide Image Classification
Linhao Qu, Dingkang Yang, Qinhao Guo, Rongkui Luo, Shaoting Zhang 0001, Xiaosong Wang 0001 |
ECCV (11) | 1 |
| 2024 | Multi-modal Data Binding for Survival Analysis Modeling with Incomplete Data and Annotations
Linhao Qu, Shaoting Zhang 0001, Xiaosong Wang 0001 |
MICCAI (5) | 1 |
| 2024 | FAST: A Dual-tier Few-Shot Learning Paradigm for Whole Slide Image ClassificationabstractThe expensive fine-grained annotation and data scarcity
have become the primary obstacles for the widespread adoption of deep learning-based Whole Slide Images (WSI) classification algorithms in clinical practice. Unlike few-shot learning methods in natural images that can leverage the labels of each image, existing few-shot WSI classification methods only utilize a small number of fine-grained labels or weakly supervised slide labels for training in order to avoid expensive fine-grained annotation. They lack sufficient mining of available WSIs, severely limiting WSI classification performance. To address the above issues, we propose a novel and efficient dual-tier few-shot learning paradigm for WSI classification, named FAST. FAST consists of a dual-level annotation strategy and a dual-branch classification framework. Firstly, to avoid expensive fine-grained annotation, we collect a very small number of WSIs at the slide level, and annotate an extremely small number of patches. Then, to fully mining the available WSIs, we use all the patches and available patch labels to build a cache branch, which utilizes the labeled patches to learn the labels of unlabeled patches and through knowledge retrieval for patch classification. In addition to the cache branch, we also construct a prior branch that includes learnable prompt vectors, using the text encoder of visual-language models for patch classification. Finally, we integrate the results from both branches to achieve WSI classification. Extensive experiments on binary and multi-class datasets demonstrate that our proposed method significantly surpasses existing few-shot classification methods and approaches the accuracy of fully supervised methods with only 0.22% annotation costs. All codes and models will be publicly available on https://github.com/fukexue/FAST. Kexue Fu 0001, Xiaoyuan Luo, Linhao Qu, Shuo Wang 0011, Ilias Maglogiannis, Longxiang Gao, Manning Wang |
NeurIPS | 3 |
| 2024 | POS-BERT: Point cloud one-stage BERT pre-training
Kexue Fu 0001, Peng Gao 0007, Shaolei Liu, Linhao Qu, Longxiang Gao, Manning Wang |
Expert Syst. Appl. | 4 |
| 2024 | Trans2Fuse: Empowering image fusion through self-supervised learning and multi-modal transformations via transformer networks
Linhao Qu, Shaolei Liu, Manning Wang, Shiman Li, Siqi Yin, Zhijian Song |
Expert Syst. Appl. | 1 |
| 2024 | Wavelet-based spectrum transfer with collaborative learning for unsupervised bidirectional cross-modality domain adaptation on medical image segmentation
Shaolei Liu, Linhao Qu, Siqi Yin, Manning Wang, Zhijian Song |
Neural Comput. Appl. | 2 |
| 2024 | Rethinking Multiple Instance Learning for Whole Slide Image Classification: A Good Instance Classifier Is All You NeedabstractWeakly supervised whole slide image classification is usually formulated as a multiple instance learning (MIL) problem, where each slide is treated as a bag, and the patches cut out of it are treated as instances. Existing methods either train an instance classifier through pseudo-labeling or aggregate instance features into a bag feature through attention mechanisms and then train a bag classifier, where the attention scores can be used for instance-level classification. However, the pseudo instance labels constructed by the former usually contain a lot of noise, and the attention scores constructed by the latter are not accurate enough, both of which affect their performance. In this paper, we propose an instance-level MIL framework based on contrastive learning and prototype learning to effectively accomplish both instance classification and bag classification tasks. To this end, we propose an instance-level weakly supervised contrastive learning algorithm for the first time under the MIL setting to effectively learn instance feature representation. We also propose an accurate pseudo label generation method through prototype learning. We then develop a joint training strategy for weakly supervised contrastive learning, prototype learning, and instance classifier training. Extensive experiments and visualizations on four datasets demonstrate the powerful performance of our method. Codes will be available. Linhao Qu, Yingfan Ma, Xiaoyuan Luo, Qinhao Guo, Manning Wang, Zhijian Song |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic RepresentationsabstractUnderstanding human intentions (e.g., emotions) from videos has received considerable attention recently. Video streams generally constitute a blend of temporal data stemming from distinct modalities, including natural language, facial expressions, and auditory clues. Despite the impressive advancements of previous works via attention-based paradigms, the inherent temporal asynchrony and modality heterogeneity challenges remain in multimodal sequence fusion, causing adverse performance bottlenecks. To tackle these issues, we propose a Multimodal fusion approach for learning modality-Exclusive and modality-Agnostic representations (MEA) to refine multimodal features and leverage the complementarity across distinct modalities. On the one hand, MEA introduces a predictive self-attention module to capture reliable context dynamics within modalities and reinforce unique features over the modality-exclusive spaces. On the other hand, a hierarchical cross-modal attention module is designed to explore valuable element correlations among modalities over the modality-agnostic space. Meanwhile, a double-discriminator strategy is presented to ensure the production of distinct representations in an adversarial manner. Eventually, we propose a decoupled graph fusion mechanism to enhance knowledge exchange across heterogeneous modalities and learn robust multimodal representations for downstream tasks. Numerous experiments are implemented on three multimodal datasets with asynchronous sequences. Systematic analyses show the necessity of our approach. Dingkang Yang, Mingcheng Li, Linhao Qu, Kun Yang 0010, Peng Zhai, Song Wang 0002, Lihua Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Negative Instance Guided Self-Distillation Framework for Whole Slide Image AnalysisabstractHistopathology image classification is an important clinical task, and current deep learning-based whole-slide image (WSI) classification methods typically cut WSIs into small patches and cast the problem as multi-instance learning. The mainstream approach is to train a bag-level classifier, but their performance on both slide classification and positive patch localization is limited because the instance-level information is not fully explored. In this article, we propose a negative instance-guided, self-distillation framework to directly train an instance-level classifier end-to-end. Instead of depending only on the self-supervised training of the teacher and the student classifiers in a typical self-distillation framework, we input the true negative instances into the student classifier to guide the classifier to better distinguish positive and negative instances. In addition, we propose a prediction bank to constrain the distribution of pseudo instance labels generated by the teacher classifier to prevent the self-distillation from falling into the degeneration of classifying all instances as negative. We conduct extensive experiments and analysis on three publicly available pathological datasets: CAMELYON16, PANDA, and TCGA, as well as an in-house pathological dataset for cervical cancer lymph node metastasis prediction. The results show that our method outperforms existing methods by a large margin. Code will be publicly available. Xiaoyuan Luo, Linhao Qu, Qinhao Guo, Zhijian Song, Manning Wang |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Reducing Domain Gap in Frequency and Spatial Domain for Cross-Modality Domain Adaptation on Medical Image SegmentationabstractUnsupervised domain adaptation (UDA) aims to learn a model trained on source domain and performs well on unlabeled target domain. In medical image segmentation field, most existing UDA methods depend on adversarial learning to address the domain gap between different image modalities, which is ineffective due to its complicated training process. In this paper, we propose a simple yet effective UDA method based on frequency and spatial domain transfer under multi-teacher distillation framework. In the frequency domain, we first introduce non-subsampled contourlet transform for identifying domain-invariant and domain-variant frequency components (DIFs and DVFs), and then keep the DIFs unchanged while replacing the DVFs of the source domain images with that of the target domain images to narrow the domain gap. In the spatial domain, we propose a batch momentum update-based histogram matching strategy to reduce the domain-variant image style bias. Experiments on two commonly used cross-modality medical image segmentation datasets show that our proposed method achieves superior performance compared to state-of-the-art methods. Shaolei Liu, Siqi Yin, Linhao Qu, Manning Wang |
AAAI | 3 |
| 2023 | Boosting Whole Slide Image Classification from the Perspectives of Distribution, Correlation and MagnificationabstractBag-based multiple instance learning (MIL) methods have become the mainstream for Whole Slide Image (WSI) classification. However, there are still three important issues that have not been fully addressed: (1) positive bags with a low positive instance ratio are prone to the influence of a large number of negative instances; (2) the correlation between local and global features of pathology images has not been fully modeled; and (3) there is a lack of effective information interaction between different magnifications. In this paper, we propose MILBooster, a powerful dual-scale multi-stage MIL framework to address these issues from the perspectives of distribution, correlation, and magnification. Specifically, to address issue (1), we propose a plug-and-play bag filter that effectively increases the positive instance ratio of positive bags. For issue (2), we propose a novel window-based Transformer architecture called PiceBlock to model the correlation between local and global features of pathology images. For issue (3), we propose a dual-branch architecture to process different magnifications and design an information interaction module called Scale Mixer for efficient information interaction between them. We conducted extensive experiments on four clinical WSI classification tasks using three datasets. MILBooster achieved new state-of-the-art performance on all these tasks. Codes will be available at https://github.com/miccaiif/MILBooster. Linhao Qu, Minghong Duan, Yingfan Ma, Shuo Wang 0011, Manning Wang, Zhijian Song |
ICCV | 1 |
| 2023 | OpenAL: An Efficient Deep Active Learning Framework for Open-Set Pathology Image Classification
Linhao Qu, Yingfan Ma, Manning Wang, Zhijian Song |
MICCAI (2) | 1 |
| 2023 | The Rise of AI Language Pathologists: Exploring Two-level Prompt Learning for Few-shot Weakly-supervised Whole Slide Image ClassificationabstractThis paper introduces the novel concept of few-shot weakly supervised learning for pathology Whole Slide Image (WSI) classification, denoted as FSWC. A solution is proposed based on prompt learning and the utilization of a large language model, GPT-4. Since a WSI is too large and needs to be divided into patches for processing, WSI classification is commonly approached as a Multiple Instance Learning (MIL) problem. In this context, each WSI is considered a bag, and the obtained patches are treated as instances. The objective of FSWC is to classify both bags and instances with only a limited number of labeled bags. Unlike conventional few-shot learning problems, FSWC poses additional challenges due to its weak bag labels within the MIL framework. Drawing inspiration from the recent achievements of vision-language models (V-L models) in downstream few-shot classification tasks, we propose a two-level prompt learning MIL framework tailored for pathology, incorporating language prior knowledge. Specifically, we leverage CLIP to extract instance features for each patch, and introduce a prompt-guided pooling strategy to aggregate these instance features into a bag feature. Subsequently, we employ a small number of labeled bags to facilitate few-shot prompt learning based on the bag features. Our approach incorporates the utilization of GPT-4 in a question-and-answer mode to obtain language prior knowledge at both the instance and bag levels, which are then integrated into the instance and bag level language prompts. Additionally, a learnable component of the language prompts is trained using the available few-shot labeled data. We conduct extensive experiments on three real WSI datasets encompassing breast cancer, lung cancer, and cervical cancer, demonstrating the notable performance of the proposed method in bag and instance classification. All codes will be made publicly accessible. Linhao Qu, Xiaoyuan Luo, Kexue Fu 0001, Manning Wang, Zhijian Song |
NeurIPS | 1 |
| 2023 | AIM-MEF: Multi-exposure image fusion based on adaptive information mining in both spatial and frequency domains
Linhao Qu, Siqi Yin, Shaolei Liu, Manning Wang, Zhijian Song |
Expert Syst. Appl. | 1 |
| 2023 | A Structure-Aware Framework of Unsupervised Cross-Modality Domain Adaptation via Frequency and Spatial Knowledge DistillationabstractUnsupervised domain adaptation (UDA) aims to train a model on a labeled source domain and adapt it to an unlabeled target domain. In medical image segmentation field, most existing UDA methods rely on adversarial learning to address the domain gap between different image modalities. However, this process is complicated and inefficient. In this paper, we propose a simple yet effective UDA method based on both frequency and spatial domain transfer under a multi-teacher distillation framework. In the frequency domain, we introduce non-subsampled contourlet transform for identifying domain-invariant and domain-variant frequency components (DIFs and DVFs) and replace the DVFs of the source domain images with those of the target domain images while keeping the DIFs unchanged to narrow the domain gap. In the spatial domain, we propose a batch momentum update-based histogram matching strategy to minimize the domain-variant image style bias. Additionally, we further propose a dual contrastive learning module at both image and pixel levels to learn structure-related information. Our proposed method outperforms state-of-the-art methods on two cross-modality medical image segmentation datasets (cardiac and abdominal). Codes are avaliable at https://github.com/slliuEric/FSUDA. Shaolei Liu, Siqi Yin, Linhao Qu, Manning Wang, Zhijian Song |
IEEE Trans. Medical Imaging | 3 |
| 2022 | TransMEF: A Transformer-Based Multi-Exposure Image Fusion Framework Using Self-Supervised Multi-Task LearningabstractIn this paper, we propose TransMEF, a transformer-based multi-exposure image fusion framework that uses self-supervised multi-task learning. The framework is based on an encoder-decoder network, which can be trained on large natural image datasets and does not require ground truth fusion images. We design three self-supervised reconstruction tasks according to the characteristics of multi-exposure images and conduct these tasks simultaneously using multi-task learning; through this process, the network can learn the characteristics of multi-exposure images and extract more generalized features. In addition, to compensate for the defect in establishing long-range dependencies in CNN-based architectures, we design an encoder that combines a CNN module with a transformer module. This combination enables the network to focus on both local and global information. We evaluated our method and compared it to 11 competitive traditional and deep learning-based methods on the latest released multi-exposure image fusion benchmark dataset, and our method achieved the best performance in both subjective and objective evaluations. Code will be available at https://github.com/miccaiif/TransMEF. Linhao Qu, Shaolei Liu, Manning Wang, Zhijian Song |
AAAI | 1 |
| 2022 | DGMIL: Distribution Guided Multiple Instance Learning for Whole Slide Image Classification
Linhao Qu, Xiaoyuan Luo, Shaolei Liu, Manning Wang, Zhijian Song |
MICCAI (2) | 1 |
| 2022 | Bi-directional Weakly Supervised Knowledge Distillation for Whole Slide Image ClassificationabstractComputer-aided pathology diagnosis based on the classification of Whole Slide Image (WSI) plays an important role in clinical practice, and it is often formulated as a weakly-supervised Multiple Instance Learning (MIL) problem. Existing methods solve this problem from either a bag classification or an instance classification perspective. In this paper, we propose an end-to-end weakly supervised knowledge distillation framework (WENO) for WSI classification, which integrates a bag classifier and an instance classifier in a knowledge distillation framework to mutually improve the performance of both classifiers. Specifically, an attention-based bag classifier is used as the teacher network, which is trained with weak bag labels, and an instance classifier is used as the student network, which is trained using the normalized attention scores obtained from the teacher network as soft pseudo labels for the instances in positive bags. An instance feature extractor is shared between the teacher and the student to further enhance the knowledge exchange between them. In addition, we propose a hard positive instance mining strategy based on the output of the student network to force the teacher network to keep mining hard positive instances. WENO is a plug-and-play framework that can be easily applied to any existing attention-based bag classification methods. Extensive experiments on five datasets demonstrate the efficiency of WENO. Code is available at https://github.com/miccaiif/WENO. Linhao Qu, Xiaoyuan Luo, Manning Wang, Zhijian Song |
NeurIPS | 1 |
| 2022 | Wavelet-based self-supervised learning for multi-scene image fusion
Shaolei Liu, Linhao Qu, Qin Qiao, Manning Wang, Zhijian Song |
Neural Comput. Appl. | 2 |