Liangrui Pan

dblp:270/5710 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
18since 2021 · last 2025
0000-0003-0565-4217ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 7 first-author · 14 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Accurate Nucleic Acid-Binding Residue Identification Based Domain-Adaptive Protein Language Model and Explainable Geometric Deep Learning
abstract
Protein-nucleic acid interactions play a fundamental and critical role in a wide range of life activities. Accurate identification of nucleic acid-binding residues helps to understand the intrinsic mechanisms of the interactions. However, the accuracy and interpretability of existing computational methods for recognizing nucleic acid-binding residues need to be further improved. Here, we propose a novel method called GeSite based the domain-adaptive protein language model and E(3)-equivariant graph neural network. Prediction results across multiple benchmark test sets demonstrate that GeSite is superior or comparable to state-of-the-art prediction methods. The MCC values of GeSite are 0.522 and 0.326 for the one DNA-binding residue test set and one RNA-binding resi-due test set, which are 0.57 and 38.14% higher than that of the second-best method, respectively. Detailed experi-mental results suggest that the advanced performance of GeSite lies in the well-designed nucleic acid-binding pro-tein adaptive language model. Additionally, interpretabil-ity analysis exposes the perception of the prediction mod-el on various remote and close functional domains, which is the source of its discernment ability.
Wenwu Zeng, Liangrui Pan, Shaoliang Peng
AAAI2
2025 MP-MIL: Multi-View Multiple Instance Learning with Positional Embedding to Predict PIK3CA Mutation
abstract
Phosphatidylinositol-4, 5-Bisphosphate 3-Kinase Catalytic Subunit Alpha (PIK3CA) gene mutations are among the most common somatic mutations in cancer, particularly in hormone receptor-positive breast cancer, with a mutation rate as high as 40 %. They are crucial for guiding targeted therapies in precision medicine. However, traditional detection methods, such as tissue-based next-generation sequencing, are challenged by high costs, time consumption, and insufficient detection of low-frequency mutations. This study proposes a novel multi-view multiple instance learning model, named MP-MIL, for non-invasively predicting PIK3CA mutation status from whole slide images (WSIs). MP-MIL effectively captures the complex spatial relationships and heterogeneity of the tumor microenvironment by fusing multi-scale features with patch sizes of$256 \times 256$and$512 \times 512$. It introducing a regional multi-head self-attention mechanism (RMSA) and a position embedding for attention-based (PEAT). Furthermore, the multi-view feature concatenation (MVC) module integrates microscopic details with macroscopic contextual information, improving the model's adaptability to heterogeneous pathological data. Experimental results on three datasets, TCGA-LUAD, TCGA-LUSC, and TCGA-BRCA, demonstrate that MP-MIL outperforms seven baseline models on most performance metrics. Ablation experiments further validated the key role of multi-view feature integration and the PEAT module in improving prediction performance. MP-MIL provides an efficient and non-invasive method for PIK3CA mutation prediction, providing important support for the advancement of precision medicine.
Guanting Li, Liangrui Pan, Xiaoyu Li 0008, Jiadi Luo, Qingchun Liang, Shaoliang Peng
BIBM2
2025 DLiPath: A Benchmark for the Comprehensive Assessment of Donor Liver Based on Histopathological Image Dataset
abstract
Pathologists' comprehensive evaluation of donor liver biopsies provides crucial information for accepting or discarding potential grafts. However, rapidly and accurately obtaining these assessments intraoperatively poses a significant challenge for pathologists. Features in donor liver biopsies, such as portal tract fibrosis, total steatosis, macrovesicular steatosis, and hepatocellular ballooning are correlated with transplant outcomes, yet quantifying these indicators suffers from substantial inter- and intra-observer variability. To address this, we introduce DLiPath, the first benchmark for comprehensive donor liver assessment based on a histopathology image dataset. We collected and publicly released 636 whole slide images from 304 donor liver patients at the Department of Pathology, the Third Xiangya Hospital, with expert annotations for key pathological features (including cholestasis, portal tract fibrosis, portal inflammation, total steatosis, macrovesicular steatosis, and hepatocellular ballooning). We selected nine state-of-the-art multiple-instance learning (MIL) models based on the DLiPath dataset as baselines for extensive comparative analysis. The experimental results demonstrate that several MIL models achieve high accuracy across donor liver assessment indicators on DLiPath, charting a clear course for future automated and intelligent donor liver assessment research. Data and code are available at https://github.com/panliangrui/liver.
Liangrui Pan, Zhongyi Chen, Chenchen Nie, Ling Chu, Shaoliang Peng
BIBM1
2025 SpaceSeg: Spatially Feature-Aware Segmentation and Classification Model for Cell Nuclei
abstract
Detecting and segmenting cell nuclei in Hematoxylin and Eosin (H&E) stained tissue images is a critical clinical task with broad applications. However, it is a challenging problem due to variations in staining and size, overlapping boundaries, cell clustering, and the high morphological and size variability of lesion regions in medical images. Accurate segmentation in medical imaging requires precise global contour localization and careful handling of local boundaries. Existing CNN-based and Transformer-based models are often limited by high parameter counts and computational complexity, making it difficult to effectively integrate these features. To address this challenge, we propose a spatially feature-aware segmentation and classification model for cell nuclei, named SpaceSeg. The Partial Gated Feed-forward Network module in SpaceSeg enhances feature representations, enabling the model to focus on key features while reducing attention to redundant information. The Spatial Attention Block enhances the spatial information of features passed to the decoder, allowing the network to automatically focus on the most critical regions in the image, especially those related to nuclear boundaries. The Partial Gated CNN module more accurately fuses low-level and high-level features, helping the model learn finer semantic information. Experimental results on the PanNuke and MoNuSeg datasets demonstrate that SpaceSeg outperforms all existing state-of-the-art models.
Yijun Peng, Liangrui Pan, Jiadi Luo, Christopher Wang, Qingchun Liang, Shaoliang Peng
BIBM2
2025 MSI-GNN: A Graph Neural Network for Pathogenicity Prediction of Microsatellite Insertions
abstract
Microsatellite insertions (MSIs) are a common type of genetic variation implicated in various hereditary disorders and cancers. However, due to their repetitive structure, high sequence variability, and limited annotation resources, the pathogenicity of MSIs remains difficult to determine in clinical settings, hindering their utility in genetic diagnosis and disease mechanism studies. To address this challenge, we propose MSIGNN, a graph neural network model that integrates multidimensional annotations and graph attention mechanisms to predict the pathogenicity of MSI events. MSI-GNN constructs a comprehensive feature profile for each MSI by incorporating heterogeneous annotations, including genomic functions, epigenetic signals, deleteriousness scores, functional constraints, evolutionary conservation, splicing effects, and molecular consequences. A multi-layer graph attention network is then employed to model biological correlations among variants and extract informative pathogenicity representations. Experimental results demonstrate that MSI-GNN significantly outperforms existing general-purpose models across multiple evaluation metrics, offering superior predictive performance and interpretability. Our model provides a promising tool for elucidating the pathogenic mechanisms of MSIs and advancing precision medicine.
Yaning Yang, Yadong Fan, Liangrui Pan, Shaoliang Peng
BIBM4
2025 GEMIL: A GELU-Enhanced Multiple-Instance Learning Model for Predicting Gene Mutations in Lung Cancer
abstract
Lung cancer remains a leading cause of cancer mortality globally. Recently, targeted therapies have significantly improved clinical outcomes in lung cancer patients, making accurate identification of driver gene mutations crucial for precision treatment. Artificial intelligence approaches leveraging routinely acquired whole-slide histopathology images (WSIs) offer a promising and cost-effective means for molecular biomarker prediction, potentially enhancing clinical decision-making. However, mutation prediction from WSIs faces substantial technical challenges, including high data dimensionality, weakly supervised labels, and severe class imbalance inherent in clinical datasets. To address these issues, we propose GEMIL, a novel multipleinstance learning (MIL) model. GEMIL features a hierarchical attention-based encoder that employs deep non-linear projections and structured regularization to learn discriminative patch-level representations. These features are then aggregated by a querydriven decoder, which efficiently consolidates instance information into a robust slide-level prediction. We validated GEMIL through extensive experiments on the large-scale PathGene dataset, with further evaluation on an independent TCGA cohort. The results demonstrate that GEMIL consistently outperforms state-of-the-art MIL methods across multiple gene mutation prediction tasks (TP53, EGFR, etc.), improving the average accuracy by 2.8% and the F1-score by 3.1%. Consequently, GEMIL provides a robust and generalizable computational tool for WSI-based biomarker prediction, holding significant potential to advance precision oncology.
Haihua Zhu 0003, Liangrui Pan, Christopher Wang, Jiadi Luo, Qingchun Liang, Shaoliang Peng
BIBM2
2025 SMILE: A Scale-aware Multiple Instance Learning Method for Multicenter STAS Lung Cancer Histopathology Diagnosis
abstract
Spread through air spaces (STAS) represents a newly identified aggressive pattern in lung cancer, which is known to be associated with adverse prognostic factors and complex pathological features. Pathologists currently rely on time-consuming manual assessments, which are highly subjective and prone to variation. This highlights the urgent need for automated and precise diagnostic solutions. 2,970 lung cancer tissue slides are comprised from multiple centers, re-diagnosed them, and constructed and publicly released three lung cancer STAS datasets: STAS-CSU (hospital), STAS-TCGA, and STAS-CPTAC. All STAS datasets provide corresponding pathological feature diagnoses and related clinical data. To address the bias, sparse and heterogeneous nature of STAS, we propose an scale-aware multiple instance learning(SMILE) method for STAS diagnosis of lung cancer. By introducing a scale-adaptive attention mechanism, the SMILE can adaptively adjust high-attention instances, reducing over-reliance on local regions and promoting consistent detection of STAS lesions. Extensive experiments show that SMILE achieved competitive diagnostic results on STAS-CSU, diagnosing 251 and 319 STAS samples in CPTAC and TCGA, respectively, surpassing clinical average AUC. The 11 open baseline results are the first to be established for STAS research, laying the foundation for the future expansion, interpretability, and clinical integration of computational pathology technologies. The datasets and code are available at https://github.com/panliangrui/IJCAI25.
Liangrui Pan, Xiaoyu Li 0008, Yutao Dou, Qiya Song, Jiadi Luo, Qingchun Liang, Shaoliang Peng
IJCAI1
2025 Application of deep learning-based multimodal fusion technology in cancer diagnosis: A survey
Liangrui Pan, Yijun Peng, Xiaoyu Li 0008, Limeng Qu, Qiya Song, Qingchun Liang, Shaoliang Peng
Eng. Appl. Artif. Intell.2
2024 ParaSAT: a scalable parallel framework for sequence alignment
abstract
Sequence alignment is a fundamental step in genomic data analysis. Third-generation sequencing technology facilitates the acquisition of high-quality genomic data but the explosive growth of sequencing data poses huge challenges to current sequence alignment. To reduce sequence alignment time and enhance alignment performance, a parallelization framework based on Minimap2 is proposed, called ParaSAT, aiming to expedite sequence alignment and offer insights to researchers in the field. In order to achieve load balancing among multi-nodes, We design a task pool scheduling strategy which can dynamically distribute tasks according to the status of compute nodes. To evaluate the performance of the framework, We choose 6 dataset to conduct experiments on TH-1 supercomputer. The outcomes confirm that the parallel framework ensures sequence alignment accuracy, and demonstrates a notable speedup on large datasets, approaching linearity, with parallel efficiency consistently above 80%. The framework also exhibits strong and weak scalability, effectively enhancing the efficiency of sequence alignment and offering guidance for high-performance genomic data processing.
Wenjuan Liu, Jianbang Xu, Liangrui Pan, Tie Cai, Shaoliang Peng
BIBM3
2024 FedDP: Privacy-preserving method based on federated learning for histopathology image segmentation
abstract
Hematoxylin and Eosin (H&E) staining of whole slide images (WSIs) is considered the gold standard for pathologists and medical practitioners for tumor diagnosis, surgical planning, and post-operative assessment. With the rapid advancement of deep learning technologies, the development of numerous models based on convolutional neural networks and transformer-based models has been applied to the precise segmentation of WSIs. However, due to privacy regulations and the need to protect patient confidentiality, centralized storage and processing of image data are impractical. Training a centralized model directly is challenging to implement in medical settings due to these privacy concerns.This paper addresses the dispersed nature and privacy sensitivity of medical image data by employing a federated learning framework, allowing medical institutions to collaboratively learn while protecting patient privacy. Additionally, to address the issue of original data reconstruction through gradient inversion during the federated learning training process, differential privacy introduces noise into the model updates, preventing attackers from inferring the contributions of individual samples, thereby protecting the privacy of the training data.Experimental results show that the proposed method, FedDP, minimally impacts model accuracy while effectively safeguarding the privacy of cancer pathology image data, with only a slight decrease in Dice, Jaccard, and Acc indices by 0.55%, 0.63%, and 0.42%, respectively. This approach facilitates cross-institutional collaboration and knowledge sharing while protecting sensitive data privacy, providing a viable solution for further research and application in the medical field.
Liangrui Pan, Mao Huang, Pinle Qin, Shaoliang Peng
BIBM1
2024 TranSVPath: A TabTransformer-Based Model for Predicting the Pathogenicity of Structural Variants
abstract
Genomic structural variants are recognized as critical molecular factors contributing to various major diseases, including cancer and genetic disorders. However, accurately determining whether a variant can cause a disease in clinical settings is extremely challenging due to limited sample sizes, the diversity of variant types, and the complexity of the mechanisms linking variants to diseases. Existing computational tools attempt to predict the pathogenic effects of these variants, but they often consider only single-layer biological data, limiting their ability to comprehensively explain the functional impacts of the variants. To address this, we propose TranSVPath, a pathogenicity scoring tool for structural variants based on the Transformer framework. TranSVPath provides a more comprehensive biological annotation of structural variations by integrating multi-dimensional data, including overlap with specific genomic regions, single nucleotide variant deleteriousness scores, phylogenetic conservation scores, Mendelian clinical application pathogenicity scores, and gene function loss. Additionally, TranSVPath employs a variant of the Transformer architecture tailored to this integrated data. Its attention mechanism captures key biological features, facilitating molecular-level interpretation of the pathogenic mechanisms of variants. Consequently, TranSVPath can accurately predict the pathogenicity of deletions, insertions, tandem duplications, inversions, and microsatellite insertions. Evaluations on large-scale human structural variant datasets demonstrate that TranSVPath outperforms state-of-the-art algorithms across relevant metrics.
Yaning Yang, Liangrui Pan, Shaoliang Peng
BIBM5
2024 Optimization of the parallel semi-Lagrangian scheme to overlap computation with communication based on grouping levels in YHGSM
Dazheng Liu, Wenjuan Liu, Liangrui Pan, Yutao Dou
CCF Trans. High Perform. Comput.3
2023 LDCSF: Local depth convolution-based Swim framework for classifying multi-label histopathology images
abstract
Histopathological images are the gold standard for diagnosing liver cancer. However, the accuracy of fully digital diagnosis in computational pathology needs to be improved. In this paper, in order to solve the problem of multi-label and low classification accuracy of histopathology images, we propose a locally deep convolutional Swim framework (LDCSF) to classify multi-label histopathology images. In order to be able to provide local field of view diagnostic results, we propose the LDCSF model, which consists of a Swin transformer module, a local depth convolution (LDC) module, a feature reconstruction (FR) module, and a ResNet module. The Swin transformer module reduces the amount of computation generated by the attention mechanism by limiting the attention to each window. The LDC then reconstructs the attention map and performs convolution operations in multiple channels, passing the resulting feature map to the next layer. The FR module uses the corresponding weight coefficient vectors obtained from the channels to dot product with the original feature map vector matrix to generate representative feature maps. Finally, the residual network undertakes the final classification task. As a result, the classification accuracy of LDCSF for interstitial area, necrosis, non-tumor and tumor reached 0.9460, 0.9960, 0.9808, 0.9847, respectively.
Liangrui Pan, Guo Chen 0001, Wenjuan Liu, Xuan Liu 0001, Shaoliang Peng
BIBM1
2023 CVFC: Attention-Based Cross-View Feature Consistency for Weakly Supervised Semantic Segmentation of Pathology Images
abstract
Histopathology image segmentation is the gold standard for diagnosing cancer, and can indicate cancer prognosis. However, histopathology image segmentation requires high-quality masks, so many studies now use image-level labels to achieve pixel-level segmentation to reduce the need for fine-grained annotation. To solve this problem, we propose an attention-based cross-view feature consistency end-to-end pseudo-mask generation framework named CVFC based on the attention mechanism. Specifically, CVFC is a three-branch joint framework composed of two Resnet38 and one Resnet50, and the independent branch multi-scale integrated feature map to generate a class activation map (CAM); in each branch, through down-sampling and The expansion method adjusts the size of the CAM; the middle branch projects the feature matrix to the query and key feature spaces, and generates a feature space perception matrix through the connection layer and inner product to adjust and refine the CAM of each branch; finally, through the feature consistency loss and feature cross loss to optimize the parameters of CVFC in co-training mode. After a large number of experiments, An IoU of 0.7122 and a fwIoU of 0.7018 are obtained on the WSSS4LUAD dataset, which outperforms HistoSegNet, SEAM, C-CAM, WSSS-Tissue, and OEEM, respectively.
Liangrui Pan, Keqin Li 0001, Wenjuan Liu, Zhichao Feng, Shaoliang Peng
BIBM1
2023 PACS: Prediction and analysis of cancer subtypes from multi-omics data based on a multi-head attention mechanism model
abstract
Due to the high heterogeneity and clinical characteristics of cancer, there are significant differences in multi-omic data and clinical characteristics among different cancer subtypes. Therefore, accurate classification of cancer subtypes can help doctors choose the most appropriate treatment options, improve treatment outcomes, and provide more accurate patient survival predictions. In this study, we propose a supervised multi-head attention mechanism model (SMA) to classify cancer subtypes successfully. The attention mechanism and feature sharing module of the SMA model can successfully learn the global and local feature information of multi-omics data. Second, it enriches the parameters of the model by deeply fusing multi-head attention encoders from Siamese through the fusion module. Validated by extensive experiments, the SMA model achieves the highest accuracy, F1 macroscopic, F1 weighted, and accurate classification of cancer subtypes in simulated, single-cell, and cancer multi-omics datasets compared to AE, CNN, and GNN-based models. Therefore, we contribute to future research on multi-omics data using our attention-based approach.
Liangrui Pan, Pinle Qin, Pengfei Rong, Xiangxiang Zeng, Dazheng Liu, Shaoliang Peng
BIBM1
2022 MGTUNet: An new UNet for colon nuclei instance segmentation and quantification
abstract
Colorectal cancer (CRC) is among the top three malignant tumor types in terms of morbidity and mortality. Histopathological images are the gold standard for diagnosing colon cancer. Cellular nuclei instance segmentation and classification, and nuclear component regression tasks can aid in the analysis of the tumor microenvironment in colon tissue. Traditional methods are still unable to handle both types of tasks end-to-end at the same time, and have poor prediction accuracy and high application costs. This paper proposes a new UNet model for handling nuclei based on the UNet framework, called MGTUNet, which uses Mish, Group normalization and transposed convolution layer to improve the segmentation model, and a ranger optimizer to adjust the SmoothL1Loss values. Secondly, it uses different channels to segment and classify different types of nucleus, ultimately completing the nuclei instance segmentation and classification task, and the nuclei component regression task simultaneously. Finally, we did extensive comparison experiments using eight segmentation models. By comparing the three evaluation metrics and the parameter sizes of the models, MGTUNet obtained 0.6254 on PQ, 0.6359 on mPQ, and 0.8695 on R2. Thus, the experiments demonstrated that MGTUNet is now a state-of-the-art method for quantifying histopathological images of colon cancer.
Liangrui Pan, Zhichao Feng, Zhujun Xu, Shaoliang Peng
BIBM1
2021 DFL-PiDA: Prediction of Piwi-interacting RNA-Disease Associations based on Deep Feature Learning
abstract
Piwi-interacting RNAs (piRNAs) fulfill the necessary requirements of epigenetic mechanisms, working to regulate gene expression in diseases and homeostasis in a coordinated manner. Hence, predicting new piRNAs that are associated with diseases conduces to understanding the pathogenicity mechanisms. In this study, we presented a deep feature learning model (DFLPiDA) to predict potential piRNA-disease associations based on the multi-model similarity features of piRNAs and diseases and the convolutional denoising auto-encoder. In particular, we firstly calculated four types of similarity features of piRNAs and diseases. Then, the convolutional denoising auto-encoder was utilized to perform deep learning on the fused similarity features. Finally, the extreme learning machine was employed as the training model as well as to predict unknown associations. The empirical results of five-fold cross-validation experiments show that the DFL-PiDA is efficient for predicting potential piRNA-disease associations. Furthermore, we proved the effectiveness of convolutional denoising auto-encoder neural network in piRNA and disease association prediction. Case studies also demonstrate the practical application of DFL-PiDA to discover potential associations.
Jiawei Luo 0001, Liangrui Pan, Shaoliang Peng
BIBM3
2021 FEDI: Few-shot learning based on Earth Mover's Distance algorithm combined with deep residual network to identify diabetic retinopathy
abstract
Diabetic retinopathy(DR) is the main cause of blindness in diabetic patients. However, DR can easily delay the occurrence of blindness through the diagnosis of the fundus. In view of the reality, it is difficult to collect a large amount of diabetic retina data in clinical practice. This paper proposes a few-shot learning model of a deep residual network based on Earth Mover's Distance algorithm to assist in diagnosing DR. We build training and validation classification tasks for few-shot learning based on 39 categories of 1000 sample data, train deep residual networks, and obtain experience maximization pre-training models. Based on the weights of the pre-trained model, the Earth Mover's Distance algorithm calculates the distance between the images, obtains the similarity between the images, and changes the model's parameters to improve the accuracy of the training model. Finally, the experimental construction of the small sample classification task of the test set to optimize the model further, and finally, an accuracy of 93.5667% on the 3wayl0shot task of the diabetic retina test set. For the experimental code and results, please refer to: https://github.com/panliangrui/few-shot-learning-funds.
Liangrui Pan, Peng Zhang 0035, Fei Xia 0003, Wenjuan Liu, Hetian Wang, Mitchai Chongcheawchamnan, Shaoliang Peng
BIBM1
2020 Classification of Hazardous Chemicals with Raman Spectrum by Convolution Neural Network
abstract
Dangerous chemicals have always been the hidden danger of social security, how to accurately identify chemicals is very important. In this experiment, the Raman scattering instrument will provide us with the Raman spectrum signal of about 190 chemical substances, each of which has its own characteristics. However, the traditional methods of identifying and classifying chemicals are not only inefficient, but also lack of security. This study proved the feasibility of using neural network to classify chemical substances. For one-dimensional signal, the experiment mainly uses the semi-supervised learning method to establish the 1D-DCNN model and simulate the real noise environment. One-dimensional signal is used as input and then the model is trained to get the model. The experimental results show that the accuracy of toxic and toxic, flammable, corrosive, environment hazard, health hazard, safe, expansive, harmful classification is 99% ± 1%. This shows that the 1D-DCNN model has strong anti-interference and robustness for signals in noise environments. This rapid classification method will provide reference value for the identification of chemical substances.
Liangrui Pan, Pronthep Pipitsunthonsan, Mitchai Chongcheawchamnan
HSI1