EDBT 2026 Demo / reviewers in the wild / expert
Bo Liao 0001
dblp:11/5830-1
· DBLP profile ↗
15ranked-venue papers
0as first author
15since 2021 · last 2026
0000-0002-3383-5691ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 11 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HSIC-H-FLapSVM: A Kernel Entropy Component Analysis and Multiple Kernel Learning-Based Fuzzy Laplacian SVM Model for Identifying Exosomal ProteinsabstractExosomal proteins participate in many vital biological processes and have great application in clinical diagnosis and prognosis. However, because of the low sequence similarity within exosomal protein datasets. It becomes increasingly urgent to develop computational methods for accurately identifying proteins secreted by exosomes. Therefore, we propose an algorithmic model called HSIC-H-FLapSVM. We extract six feature types from protein sequences. Through experimentation, we determine to use PSSM-DWT, PSSM-AB, and PsePSSM. Subsequently, the features are integrated via multiple kernel learning method based on the Hilbert-Schmidt independence criterion (MKL-HSIC). Following this, fuzzy membership scores for the training samples are derived using kernel entropy component analysis (KECA). Ultimately, HSIC-H-FLapSVM achieves superior performance on testing set, outperforming all competing methods with an ACC of 0.8623 and an MCC of 0.6347. These results demonstrate that HSIC-H-FLapSVM is an effective tool for predicting exosomal proteins. Jiajia He, Shaoyou Yu, Bo Liao 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2026 | Low-Count PET Image Reconstruction With Generalized Sparsity Priors via Unrolled Deep NetworksabstractDeep learning has demonstrated remarkable efficacy in reconstructing low-count PET (Positron EmissionTomography) images, attracting considerable attention in the medical imaging community. However, most existing deep learning approaches have not fully exploited the unique physical characteristics of PET imaging in the design of fidelity and prior regularization terms, resulting in constrained model performance and interpretability. In light of these considerations, we introduce an unrolled deep network based on maximum likelihood estimation for the Poisson distribution and a Generalized domain transformation for Sparsity learning, dubbed GS-Net. To address this complex optimization challenge, we employ the Alternating Direction Method of Multipliers (ADMM) framework, integrating a modified Expectation Maximization (EM) approach to address the primary objective and utilize the shrinkage thresholding approach to optimize the L1 norm term. Additionally, within this unrolled deep network, all hyperparameters are adaptively adjusted through end-to-end learning to eliminate the need for manual parameter tuning. Through extensive experiments on simulated patient brain datasets and real patient whole-body clinical datasets with multiple count levels, our method has demonstrated advanced performance compared to traditional non-iterative and iterative reconstruction, deep learning-based direct reconstruction, and hybrid unrolled methods, as demonstrated by qualitative and quantitative evaluations. Minghan Fu, Bo Liao 0001, Dong Liang 0001, Zhanli Hu, Fang-Xiang Wu |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | SG-Fusion: A swin-transformer and graph convolution-based multi-modal deep neural network for glioma prognosisabstractThe integration of morphological attributes extracted from histopathological images and genomic data holds significant importance in advancing tumor diagnosis, prognosis, and grading. Histopathological images are acquired through microscopic examination of tissue slices, providing valuable insights into cellular structures and pathological features. On the other hand, genomic data provides information about tumor gene expression and functionality. The fusion of these two distinct data types is crucial for gaining a more comprehensive understanding of tumor characteristics and progression. In the past, many studies relied on single-modal approaches for tumor diagnosis. However, these approaches had limitations as they were unable to fully harness the information from multiple data sources. To address these limitations, researchers have turned to multi-modal methods that concurrently leverage both histopathological images and genomic data. These methods better capture the multifaceted nature of tumors and enhance diagnostic accuracy. Nonetheless, existing multi-modal methods have, to some extent, oversimplified the extraction processes for both modalities and the fusion process. In this study, we presented a dual-branch neural network, namely SG-Fusion. Specifically, for the histopathological modality, we utilize the Swin-Transformer structure to capture both local and global features and incorporate contrastive learning to encourage the model to discern commonalities and differences in the representation space. For the genomic modality, we developed a graph convolutional network based on gene functional and expression level similarities. Additionally, our model integrates a cross-attention module to enhance information interaction and employs divergence-based regularization to enhance the model's generalization performance. Validation conducted on glioma datasets from the Cancer Genome Atlas unequivocally demonstrates that our SG-Fusion model outperforms both single-modal methods and existing multi-modal approaches in both survival analysis and tumor grading. Minghan Fu, Rayyan Azam Khan, Bo Liao 0001, Zhanli Hu, Fang-Xiang Wu |
Artif. Intell. Medicine | 4 |
| 2024 | A systematic review on deep learning based methods for cervical cell image analysisabstractCervical cytology image analysis is indispensable for the detection of abnormal cervical cells. Traditionally, manual screening is time-consuming and labor-intensive. Therefore, a lot of deep learning (DL)-based automatic detection methods have been employed in this field to provide timely, accurate and objective results. In this study, we systematically review the current developments in cervical cell image analysis with DL methods. Specifically, we first present the most popular DL models that are widely applied in cervical cell analysis. Second, we describe the methodology for conducting this review. Third, we provide all publicly available datasets related to cervical cell images to the best of our knowledge. Then, we introduce relevant evaluation metrics and loss functions. Next, we summarize and assort the applications for cervical cell classification and segmentation. Afterwards, we discuss about current challenges and future research directions in this field. Finally, we draw the conclusion of this review. According to the analysis, we conclude that the studies based on DL models have maintained an increasing trend in recent years, which indicates the potential of DL in cervical cell image analysis. In cervical cell image classification, CNN is the most commonly used DL model. Among CNN models, we can find that VGGNet and ResNet are the most popular network architectures for the classification of cervical cells. Transformer is the second commonly used DL model. Moreover, Herlev and SIPaKMeD are the most popular public datasets used for cervical cell classification. In cervical cell segmentation, U-Net and FCN are the two most popular DL architectures. In addition, ISBI2014 and Herlev datasets are the most frequently used among the existing publicly available segmentation datasets. However, there are some issues in this field, such as poor cervical cell classification performance as a result of similar pathological properties between different cell categories. Therefore, it is necessary to develop more effective methods with DL models to improve these issues in the future research. Bo Liao 0001, Xiujuan Lei, Fang-Xiang Wu |
Neurocomputing | 2 |
| 2023 | Imputing single-cell RNA-seq data by graph autoencoder with multi-kernelabstractSingle-cell RNA-sequencing (scRNA-seq) technology has revolutionized the field by enabling the profiling of transcriptomes in cell resolution. However, it is flawed by the sparsity caused by low mRNA capture efficiency during sequencing. This results in "dropout" events where genes are expressed but not detected. Dropout can hinder downstream analyses like differential expression and clustering. To tackle this issue, we present a novel imputation approach called MKGAE, which utilizes graph convolution and autoencoder techniques to construct a generative model for imputing missing values within scRNA-seq data. Meanwhile, considering the intricate relationships between genes, merging them into a single graph might lead to the loss of important insights. To address this, we utilize two gene-to-gene graph kernels for graph convolution. Experiments across both simulated and real scRNA-seq datasets illustrate MKGAE’s superiority over other state-of-the-art methods in terms of clustering analysis and differentially expressed gene identification. Bo Liao 0001, Petros Papagerakis, Fang-Xiang Wu |
BIBM | 2 |
| 2023 | Biomarker Identification via a Factorization Machine-Based Neural Network With Binary Pairwise EncodingabstractBiomolecules, microRNAs (miRNAs) and long non-coding RNAs (lncRNAs), play critical roles in diverse fundamental and vital biological processes. They can serve as disease biomarkers as their dysregulations could cause complex human diseases. Identifying those biomarkers is helpful with the diagnosis, treatment, prognosis, and prevention of diseases. In this study, we propose a factorization machine-based deep neural network with binary pairwise encoding, DFMbpe, to identify the disease-related biomarkers. First, to comprehensively consider the interdependence of features, a binary pairwise encoding method is designed to obtain the raw feature representations for each biomarker-disease pair. Second, the raw features are mapped into their corresponding embedding vectors. Then, the factorization machine is conducted to get the wide low-order feature interdependence, while the deep neural network is applied to obtain the deep high-order feature interdependence. Finally, two kinds of features are combined to get the final prediction results. Unlike other biomarker identification models, the binary pairwise encoding considers the interdependence of features even though they never appear in the same sample, and the DFMbpe architecture emphasizes both low-order and high-order feature interactions simultaneously. The experimental results show that DFMbpe greatly outperforms the state-of-the-art identification models on both cross-validation and independent dataset evaluation. Besides, three types of case studies further demonstrate the effectiveness of this model. Yulian Ding, Xiujuan Lei, Bo Liao 0001, Fang-Xiang Wu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | A Robust Oversampling Approach for Class Imbalance Problem With Small DisjunctsabstractClass imbalance is one of the important challenges for machine learning because of it's learning to bias toward the majority classes. The oversampling method is a fundamental imbalance-learning technique with many real-world applications. However, when the small disjuncts problem occurs, how to effectively avoiding the negative oversampling results rather than using clusters previously, remains a challenging task. Thus, this study introduces a disjuncts-robust oversampling (DROS) method. The novel method shows that the data filling of new synthetic samples to the minority class areas in data space can be thought of as the searchlight illuminating with light cones to the restricted areas in real life. In the first step, DROS computes a series of light-cone structures that is first started from the inner minority class area, then passes through the boundary minority class area, last is stopped by the majority class area. In the second step, DROS generates new synthetic samples in those light-cone structures. Experiments considering both real-world and 2D emulational datasets demonstrate that our method outperforms the current state-of-the-art oversampling methods and suggest that our method is able to deal with the small disjuncts. Bo Liao 0001, Wen Zhu, Junlin Xu |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | MLRDFM: a multi-view Laplacian regularized DeepFM model for predicting miRNA-disease associationsabstractMOTIVATION: MicroRNAs (miRNAs), as critical regulators, are involved in various fundamental and vital biological processes, and their abnormalities are closely related to human diseases. Predicting disease-related miRNAs is beneficial to uncovering new biomarkers for the prevention, detection, prognosis, diagnosis and treatment of complex diseases. RESULTS: In this study, we propose a multi-view Laplacian regularized deep factorization machine (DeepFM) model, MLRDFM, to predict novel miRNA-disease associations while improving the standard DeepFM. Specifically, MLRDFM improves DeepFM from two aspects: first, MLRDFM takes the relationships among items into consideration by regularizing their embedding features via their similarity-based Laplacians. In this study, miRNA Laplacian regularization integrates four types of miRNA similarity, while disease Laplacian regularization integrates two types of disease similarity. Second, to judiciously train our model, Laplacian eigenmaps are utilized to initialize the weights in the dense embedding layer. The experimental results on the latest HMDD v3.2 dataset show that MLRDFM improves the performance and reduces the overfitting phenomenon of DeepFM. Besides, MLRDFM is greatly superior to the state-of-the-art models in miRNA-disease association prediction in terms of different evaluation metrics with the 5-fold cross-validation. Furthermore, case studies further demonstrate the effectiveness of MLRDFM. Yulian Ding, Xiujuan Lei, Bo Liao 0001, Fang-Xiang Wu |
Briefings Bioinform. | 3 |
| 2022 | Predicting drug-drug interactions by graph convolutional network with multi-kernelabstractDrug repositioning is proposed to find novel usages for existing drugs. Among many types of drug repositioning approaches, predicting drug-drug interactions (DDIs) helps explore the pharmacological functions of drugs and achieves potential drugs for novel treatments. A number of models have been applied to predict DDIs. The DDI network, which is constructed from the known DDIs, is a common part in many of the existing methods. However, the functions of DDIs are different, and thus integrating them in a single DDI graph may overlook some useful information. We propose a graph convolutional network with multi-kernel (GCNMK) to predict potential DDIs. GCNMK adopts two DDI graph kernels for the graph convolutional layers, namely, increased DDI graph consisting of 'increase'-related DDIs and decreased DDI graph consisting of 'decrease'-related DDIs. The learned drug features are fed into a block with three fully connected layers for the DDI prediction. We compare various types of drug features, whereas the target feature of drugs outperforms all other types of features and their concatenated features. In comparison with three different DDI prediction methods, our proposed GCNMK achieves the best performance in terms of area under receiver operating characteristic curve and area under precision-recall curve. In case studies, we identify the top 20 potential DDIs from all unknown DDIs, and the top 10 potential DDIs from the unknown DDIs among breast, colorectal and lung neoplasms-related drugs. Most of them have evidence to support the existence of their interactions. [email protected]. Fei Wang 0095, Xiujuan Lei, Bo Liao 0001, Fang-Xiang Wu |
Briefings Bioinform. | 3 |
| 2022 | ACP_MS: prediction of anticancer peptides based on feature extractionabstractAnticancer peptides (ACPs) are bioactive peptides with antitumor activity and have become the most promising drugs in the treatment of cancer. Therefore, the accurate prediction of ACPs is of great significance to the research of cancer diseases. In the paper, we developed a more efficient prediction model called ACP_MS. Firstly, the monoMonoKGap method is used to extract the characteristic of anticancer peptide sequences and form the digital features. Then, the AdaBoost model is used to select the most discriminating features from the digital features. Finally, a stochastic gradient descent algorithm is introduced to identify anticancer peptide sequences. We adopt 7-fold cross-validation and independent test set validation, and the final accuracy of the main dataset reached 92.653% and 91.597%, respectively. The accuracy of the alternate dataset reached 98.678% and 98.317%, respectively. Compared with other advanced prediction models, the ACP_MS model improves the identification ability of anticancer peptide sequences. The data of this model can be downloaded from the public website for free https://github.com/Zhoucaimao1998/Zc. Caimao Zhou, Dejun Peng, Bo Liao 0001, Ranran Jia, Fang-Xiang Wu |
Briefings Bioinform. | 3 |
| 2022 | Identifying Gene Signatures for Cancer Drug Repositioning Based on Sample ClusteringabstractDrug repositioning is an important approach for drug discovery. Computational drug repositioning approaches typically use a gene signature to represent a particular disease and connect the gene signature with drug perturbation profiles. Although disease samples, especially from cancer, may be heterogeneous, most existing methods consider them as a homogeneous set to identify differentially expressed genes (DEGs)for further determining a gene signature. As a result, some genes that should be in a gene signature may be averaged off. In this study, we propose a new framework to identify gene signatures for cancer drug repositioning based on sample clustering (GS4CDRSC). GS4CDRSC first groups samples into several clusters based on their gene expression profiles. Second, an existing method is applied to the samples in each cluster for generating a list of DEGs. Then a weighting approach is used to identify an intergrated gene signature from all the lists of DEGs. The integrated gene signature is used to connect with drug perturbation profiles in the Connectivity Map (CMap)database to generate a list of drug candidates. GS4CDRSC has been tested with several cancer datasets and existing methods. The computational results show that GS4CDRSC outperforms those methods without the sample clustering and weighting approaches in terms of both number and rate of predicted known drugs for specific cancers. Fei Wang 0095, Yulian Ding, Xiujuan Lei, Bo Liao 0001, Fang-Xiang Wu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | Predicting miRNA-Disease Associations Based On Multi-View Variational Graph Auto-Encoder With Matrix FactorizationabstractMicroRNAs (miRNAs) have been proved to play critical roles in diverse biological processes, including the human disease development process. Exploring the potential associations between miRNAs and diseases can help us better understand complex disease mechanisms. Given that traditional biological experiments are expensive and time-consuming, computational models can serve as efficient means to uncover potential miRNA-disease associations. This study presents a new computational model based on variational graph auto-encoder with matrix factorization (VGAMF) for miRNA-disease association prediction. More specifically, VGAMF first integrates four different types of information about miRNAs into an miRNA comprehensive similarity network and two types of information about diseases into a disease comprehensive similarity network, respectively. Then, VGAMF gets the non-linear representations of miRNAs and diseases, respectively, from those two comprehensive similarity networks with variational graph auto-encoders. Simultaneously, a non-negative matrix factorization is conducted on the miRNA-disease association matrix to get the linear representations of miRNAs and diseases. Finally, a fully connected neural network combines linear and non-linear representations of miRNAs and diseases to get the final predicted association score for all miRNA-disease pairs. In the 10-fold cross-validation experiments, VGAMF achieves an average AUC of 0.9280 on HMDD v2.0 and 0.9470 on HMDD v3.2, which outperforms other competing methods. Besides, the case studies on colon cancer and esophageal cancer further demonstrate the effectiveness of VGAMF in predicting novel miRNA-disease associations. Yulian Ding, Xiujuan Lei, Bo Liao 0001, Fang-Xiang Wu |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | Minority Sub-Region Estimation-Based Oversampling for Imbalance LearningabstractClass imbalance problem that characterized with the skew distribution towards the majority arises as one challenge in recent years. Many oversampling techniques have been proposed to cope with this problem and some of them combine the oversampling procedure with the clustering algorithm which guaranteeing new synthetic samples being generated in clusters. However far-away samples but with the same minority sub-region are generally clustered into different groups owing to the characteristic of clustering algorithm itself. Therefore, the following oversampling procedure is mostly carried in incomplete minority sub-regions that synthetic samples not well cover the integral minority region. And to our best knowledge, none of existing algorithm is designed to directly estimate minority sub-regions for class imbalance problem. Thus, one new grouping algorithm, named Direction Distribution-based Minority Sub-region Estimation (DDMSE), is first proposed. The new algorithm exploits the intuitive observation, that the minority with the same sub-region almost distribute within the same direction when compared to other majority, to estimate minority sub-regions that tactfully ignoring negative impacts brought by the distance factor like in clustering algorithms. Finally, new synthetic samples are generated in those minority sub-regions. And experimental results on real-world datasets show the comparable performance with other state-of-the-art oversampling methods. Bo Liao 0001, Wen Zhu |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2021 | A Novel Approach for Potential Human LncRNA-Disease Association Prediction Based on Local Random WalkabstractIn recent years, lncRNAs (long non-coding RNAs) have been proved to be closely related to many diseases that are seriously harmful to human health. Although researches on clarifying the relationships between lncRNAs and diseases are developing rapidly, associations between the lncRNAs and diseases are still remaining largely unknown. In this manuscript, a novel Local Random Walk based prediction model called LRWHLDA is proposed for inferring potential associations between human lncRNAs and diseases. In LRWHLDA, a new heterogeneous network is established first, which allows that LRWHLDA can be implemented in the case of lacking known lncRNA-disease associations. And then, an improved local random walk method is designed for prediction of novel lncRNA-disease associations, which can help LRWHLDA achieve high prediction accuracy but with low time complexity. Finally, in order to evaluate the prediction performance of LRWHLDA, different frameworks such as LOOCV, 2-folds CV, and 5-folds CV have been implemented, simulation results indicate that LRWHLDA can achieve reliable AUCs of 0.8037, 0.8354, and 0.8556 under the frameworks of 2-fold CV, 5-fold CV, and LOOCV, respectively. Hence, it is easy to know that LRWHLDA contains the potential to be a representative of emerging methods in the field of research on potential lncRNA-disease associations prediction. Jiechen Li, Zhanwei Xuan, Jingwen Yu, Bo Liao 0001, Lei Wang 0069 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2021 | Human Protein Complex-Based Drug Signatures for Personalized Cancer MedicineabstractDisease signature-based drug repositioning approaches typically first identify a disease signature from gene expression profiles of disease samples to represent a particular disease. Then such a disease signature is connected with the drug-induced gene expression profiles to find potential drugs for the particular disease. In order to obtain reliable disease signatures, the size of disease samples should be large enough, which is not always a single case in practice, especially for personalized medicine. On the other hand, the sample sizes of drug-induced gene expression profiles are generally large. In this study, we propose a new drug repositioning approach (HDgS), in which the drug signature is first identified from drug-induced gene expression profiles, and then connected to the gene expression profiles of disease samples to find the potential drugs for patients. In order to take the dependencies among genes into account, the human protein complexes (HPC) are used to define the drug signature. The proposed HDgS is applied to the drug-induced gene expression profiles in LINCS and several types of cancer samples. The results indicate that the HPC-based drug signature can effectively find drug candidates for patients and that the proposed HDgS can be applied for personalized medicine with even one patient sample. Fei Wang 0095, Yulian Ding, Xiujuan Lei, Bo Liao 0001, Fang-Xiang Wu |
IEEE J. Biomed. Health Informatics | 4 |