VLDB 2026 Research / reviewers in the wild / expert
Lei Liu 0079
dblp:21/2715-79
· DBLP profile ↗
11ranked-venue papers
7as first author
11since 2021 · last 2026
0009-0006-0962-724XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A knowledge prompt augmented lightweight multimodal language assistant for biomedicine
Lei Liu 0079, Xiangdong Su, Xingxiang Zhou, Guanglai Gao |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | MedVSA: Medical Visual Spoken-Question AnsweringabstractWith the rapid advancement of technology, smart healthcare has made significant progress, particularly in medical visual question answering (MedVQA). However, current MedVQA primarily relies on text, whereas practical applications often involve spoken interactions, such as in medical consultations and mobile-based queries. To bridge this gap, we propose a novel task, medical visual spoken-question answering (MedVSA), extending the conventional medical image and text-based question-answering paradigm to include spoken interactions for enhanced applicability. We expand upon four commonly used MedVQA datasets, namely VQA-RAD, SLAKE, PathVQA, and OVQA, by leveraging Alibaba Cloud speech synthesis technology to convert text questions into spoken questions. Various strategies are incorporated to ensure diverse and realistic speech synthesis. The resulting dataset comprises images, corresponding text, and synthesized speech data. Subsequently, we design a single-stage model and a two-stage model to tackle the MedVSA task. For the single-stage model, we directly input the speech and images into our designed whisper self-distillation model to obtain the results. For the two-stage model, we first use the Whisper model to convert the speech into text, then input the converted text and medical images into the self-distillation model to obtain the results. We provide two solutions for the MedVSA task and establish two baselines. Experimental results show that the two-stage model significantly outperforms the single-stage model, indicating that text conversion is crucial for solving MedVSA. This study advances smart healthcare developments by proposing MedVSA and designing two baselines tailored to its specificities. Source code and MedVSA dataset are available at https://github.com/Alivelei/MedVSA. Lei Liu 0079, Xiangdong Su, Guanglai Gao |
ICMR | 1 |
| 2025 | Fourier Self-Adaptation for Transferring General Pretrained Models to Specific DomainsabstractWhile pre-trained models in the general domain have proliferated, existing methods for transferring these models to specific domains often depend on source domain data for distribution alignment and are typically tailored for single tasks. We propose a source data-free approach, Fourier Self-Adaptation (FSA), which effectively adapts general models to a wide range of specific domains. Our method leverages the distinct properties of Fourier phase and amplitude: phase contains high-level structural and positional information, which is less affected by domain shifts, while amplitude contains details and brightness information, which is more affected by domain shifts. FSA adjusts the image distribution by initializing a trainable adaptive image from a normal distribution. It then interpolates the amplitude of the target domain image with that of the adaptive image, where the interpolation ratio is dynamically controlled by learnable weight and bias. During training, the model captures advanced phase information of the target image and refines the data distribution through amplitude interpolation. Additionally, a dual regularization loss constrains the model representation, encouraging it to focus on the intrinsic relationships of the target domain data while discarding irrelevant knowledge. We evaluate FSA using general pre-trained models on 11 unimodal image classification datasets and 6 multimodal visual question answering datasets, covering specific domains such as radiology, pathology, remote sensing, and art. Our method consistently achieves state-of-the-art performance across multiple datasets, with performance improvements ranging from 1% to 8% compared to basic pre-trained models. Source code are available at https://github.com/Alivelei/FSA. Lei Liu 0079, Xiangdong Su, Guanglai Gao |
ACM Multimedia | 1 |
| 2024 | Leveraging Convolutional Models as Backbone for Medical Visual Question AnsweringabstractConvolutional neural networks (CNNs) have made significant contributions to computer vision and offer the advantages of higher training efficiency and lower model complexity. However, their application as the backbone in medical visual question answering (MedVQA) remains an open question. To address this issue, we employ popular convolutional models, including ResNet, DenseNet, and ShuffleNet, as the foundation for MedVQA, achieving outstanding performance. Different backbones can be tailored to diverse real-world scenarios. The central challenge in utilizing CNNs for visual question answering is effectively managing textual features and integrating multi-modal information. To overcome this challenge, we design a novel global interaction attention (GIA) that facilitates efficient interactions between text and image features. Additionally, we utilize the dot product before the classifier output to enhance visual and textual modal fusions. To further enhance model performance, we propose a novel multi-modal hidden mixup (MHidMix) technique for data augmentation, which involves interpolating hidden states during model training. This data augmentation technique smoothes the decision boundary without the need for complex sample selection, further improving model performance. Experimental results underscore the versatility of our proposed framework across various convolutional models, leading to outstanding performance on four MedVQA datasets. Notably, we achieved an accuracy increase of 9.4% on the PathVQA dataset and 4.5% on the OVQA dataset. Lei Liu 0079, Xiangdong Su, Guanglai Gao |
BIBM | 1 |
| 2024 | Optimizing Transformer and MLP with Hidden States Perturbation for Medical Visual Question AnsweringabstractOptimizing model performance is a crucial objective in medical visual question answering (MedVQA), and a wide range of techniques have been developed to achieve this goal. In this paper, we propose a novel technique called network state perturbation, which distinguishes itself from existing research. Our study designs four innovative methods to modify the hidden states within the network and improve the performance of multimodal models in the MedVQA task. Specifically, we evaluate these methods on both general transformer and multilayer perceptron (MLP) models, which allow for hidden state adjustments at each layer. The four introduced methods are as follows: (1) Randomly Set Zero, which assigns zeros to the hidden states of different modalities in the network; (2) Randomly Replace Content, which performs interpolation between the hidden states of the text sequence and the image sequence; (3) Randomly Add Gaussian Noise, which adds Gaussian noise to the hidden states of different modalities; and (4) Pair Interpolation, which interpolates the hidden states of different modalities of the current sample with those of other samples and performs corresponding label interpolation to facilitate model training. Our experimental results demonstrate that the proposed hidden state perturbation methods significantly enhance the performance of various transformer and MLP models on multiple MedVQA datasets, without requiring additional data or computational resources during training. These findings highlight the potential of hidden state perturbation as a novel model improvement technique for multimodal models in the medical domain. Lei Liu 0079, Xiangdong Su, Guanglai Gao |
BIBM | 1 |
| 2024 | Learning Frequency Adaptation for Cross-domain Medical Image SegmentationabstractIn medical image segmentation, some recent methods improve domain adaptation performance through frequency domain adaptation and frequency mixup. However, these approaches have two limitations: (1) frequency domain adaptation ignores the adverse effects of high-frequency noise on model generalization, and (2) frequency mixup confuses semantic information. To address these issues, we propose a novel frequency adaptation approach for medical image segmentation including low-frequency component alignment (LFCA) and random amplitude cutmix (RAC). Since low-frequency contains major image information and high-frequency noise affects adaptation, we leverage discrete wavelet transform to decompose images into low and high-frequency components. LFCA aligns the domain distribution of low frequencies and high frequencies are passed to the decoder via skip connections. In addition, we design RAC to generate diverse augmented samples through amplitude cutmix while avoiding distortion of the original distributions. Experiments on benchmark datasets validate the efficacy of our proposed approach. Lei Liu 0079, Xiangdong Su, Guanglai Gao |
BIBM | 2 |
| 2024 | FSAM: Fine-tuning SAM encoder and decoder for Medical Image SegmentationabstractRecently, the Segment Anything Model (SAM), a large pre-training model, has achieved excellent results on natural image segmentation tasks and received extensive attention. However, SAM on medical image segmentation is unsatisfactory since there are significant differences between natural and medical images. How to extend the SAM’s powerful segmentation capabilities to the medical domain requires further exploration. To this end, we propose a simple and efficient fine-tuning approach for SAM that does not require large-scale data called FSAM. Specifically, FSAM simultaneously fine-tunes the encoder and decoder while freezing the prompt encoder. This allows the encoder to extract medical image features effectively and guide decoder segmentation predictions. FSAM achieves state-of-the-art results on eight public medical image datasets, outperforming SAM by +15.39%. Moreover, FSAM exhibits weaker sample scale dependence. The proposed framework further improves the segmentation capabilities of SAM in medical images. Lei Liu 0079, Xiangdong Su, Guanglai Gao |
BIBM | 2 |
| 2023 | YOLO-D: Dual-Branch Infrared Distant Target Detection Based on Multi-level Weighted Feature Fusion
Jianqiang Jing, Bing Jia, Baoqi Huang, Lei Liu 0079 |
ICONIP (13) | 4 |
| 2022 | How Well Apply Multimodal Mixup and Simple MLPs Backbone to Medical Visual Question Answering?abstractAlthough current methods have significantly improved the performance of medical visual question answering (Med-VQA), there are still two aspects worth exploring, namely the simplification of model structure and the effective model training on small-scale data. Different from the previous Med-VQA model, this paper only employs multi-layer perceptrons (MLPs) as the backbone network for feature extraction and modal fusion and designs a Med-VQA model on such basis, which achieves superior performance with a simple backbone network. To enhance model generalization, we design multimodal mixup (M-Mixup) to augment images and questions separately, which effectively alleviates the problem of insufficient training samples in the Med-VQA task. To prevent the destruction of the feature relationship when tokenizing the medical image, we design pooling tokens (PTs), a simple downsampling structure to capture fine-grained visual features without affecting the parameters and FLOPs of the entire model. Experimental results demonstrate that our model achieves state-of-the-art on the SLAKE, and obtains a remarkably competitive performance on the VQA-RAD. The source code and models are available at https://github.com/Alivelei/M-Mixup. Lei Liu 0079, Xiangdong Su |
BIBM | 1 |
| 2022 | Medical Visual Question Answering via Targeted Choice Contrast and Multimodal Entity Matching
Lei Liu 0079, Xiangdong Su |
ICONIP (2) | 2 |
| 2022 | A Transformer-based Medical Visual Question Answering ModelabstractWhile the Transformer architecture has been widely used in natural language processing tasks and computer vision tasks, its application in medical visual question answering is still limited. Most current methods rely on an image extractor to obtain visual features and a text extractor to capture semantic features, and then a fusion module to merge the information from the two modalities to predict the final result. In contrast, this paper proposes a novel Transformer-based medical vision question answering model, called MQAT, in which an improved Transformer structure is used for feature extraction and modal fusion to achieve better performance. Experimental results demonstrate that our Transformer structure not only ensures the stability of the model performance, but also accelerates its convergence, and the MQAT model outperforms the existing state-of-the-art methods. Lei Liu 0079, Xiangdong Su, Daobin Zhu |
ICPR | 1 |