VLDB 2026 Research / reviewers in the wild / expert
Rui Gao 0006
dblp:43/2694-6
· DBLP profile ↗
13ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-3599-7678ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Grained Line Graph Neural Network With Hierarchical Contrastive Learning for Predicting Drug-Disease AssociationsabstractPredicting drug-disease associations is a crucial step in drug repositioning, especially with computational methods that quickly locate potential drug-disease pairs. Heterogenous network is a common tool for introducing multiple type relation information about drugs and diseases. However, the diversity of relations is ignored in most of existing methods, which makes them difficult to explore type semantic information with structure properties. Therefore, we propose a relation-centric GNN framework to encode critical association patterns. Firstly, we utilize a relation-centric graph, line graph, to represent the context of a drug-disease pair identified as the center node. The prediction problem is modeled to learn the embedding vector of the center node. Secondly, a multi-grained line graph neural network (MGLGNN) is designed to excavate fine-grained features that encapsulate local graph structures. We theoretically define a handful of typical nodes that can be regarded as high-order abstractions of relations in each type. Then, MGLGNN distills the local information and passes it to typical nodes from a global perspective. With learned multi-grained features, the center node automatically captures heterogenous relation semantics and structure patterns. Thirdly, a hierarchical contrastive learning (HCL) mechanism is proposed to ensure the quality of multi-grained features in an unsupervised way. Extensive experiments show the great potential of our model in mining drug-disease associations. Bao-Min Liu, Ling-Yun Dai, Junliang Shang, Chun-Hou Zheng 0001, Ying-Lian Gao, Rui Gao 0006, Jin-Xing Liu 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | soFusion: facilitating tissue structure identification via spatial multi-omics data fusionabstractThe rapid advancement of spatial multi-omics technologies has opened new avenues for dissecting tissue architecture with unprecedented resolution. However, inherent disparities across omics modalities, such as differences in biological hierarchy and resolution, pose significant challenges for integrative analysis. To address this, we present soFusion, a method for representation learning on spatial multi-omics data that enables automated identification of tissue compartmentalization. soFusion employs a graph convolutional network (GCN) to extract latent embeddings from spatial omics profiles. To simultaneously capture both cross-modality relationships and modality-specific features, we introduce a novel strategy for intra- and inter-omics feature learning. Moreover, modality-specific decoders are designed to preserve the unique information embedded in each omics type. We evaluated soFusion on multiple datasets including gene expression, protein expression, and epigenetic features. Across all benchmarks, soFusion consistently outperformed existing methods in delineating anatomical structures and identifying spatial domains with improved continuity and reduced noise. Collectively, soFusion offers an effective solution for spatial multi-omics integration, substantially enhancing the robustness of spatial domain identification. Na Yu 0004, Qi Zou 0003, Daoliang Zhang, Wei Zhang 0241, Rui Gao 0006 |
Briefings Bioinform. | 9 |
| 2025 | Multi-scale cancer driver gene prediction by flexible data selection and network topology guidanceabstractOBJECTIVE: Efficient and comprehensive prioritization of cancer driver genes across individual patients, cancer cohorts, and pan-cancer is crucial for advancing cancer diagnosis and treatment. The existing methods are effective, but they seem to have reached a plateau in accuracy enhancement and lack broad-scale joint analysis, flexibility in adapting to cancer and interpretability. METHODS: Here, we introduce GenMorw, a heterogeneous network framework that discovers a novel association score between patients and their mutated genes, enabling the estimation of the likelihood of the mutated genes acting as drivers in patients. GenMorw flexibly integrates or fully utilize collected mutation, gene/miRNA expression, methylation data and PPI networks to classify patient groups based on data-specific characteristics and identify potential drivers at the individual, cancer and pan-cancer levels. RESULTS: GenMorw outperforms existing algorithms with an average cohort AUC improvement of 17.66% and higher overall accuracy by a cumulative ranking strategy in patient-gene heterogeneous networks. Except for AUC evaluation, other various comparative strategies consistently demonstrate the superior performance of GenMorw across multiple cancers, outperforming other algorithms. Some uniquely predicted genes, such as ANK3, CENPF, and COL7A1, which are absent from standard databases and not identified by other methods, were validated as highly cancer-related through literature review and survival analysis. Based on GenMorw-derived heterogeneous networks, the strongly connected components and cliques, which are extracted from them, capture most of the predicted or known driver genes to help predict driver genes. CONCLUSION: We conclude that GenMorw, with its novel gene-patient score mechanism, offers a significant advance in cancer driver gene discovery by capturing both population-wide and patient-specific network signals, thereby improving predictive power and enabling deeper insights into cancer heterogeneity. Jian Liu 0039, Yingzan Ren, Guodong Xiao, Ponian Li, Chuanqi Sun, Fubin Ma, Rui Gao 0006, Haiyan Cong, Yusen Zhang 0002 |
J. Biomed. Informatics | 8 |
| 2025 | BIGFormer: A Graph Transformer With Local Structure Awareness for Diagnosis and Pathogenesis Identification of Alzheimer's Disease Using Imaging Genetic DataabstractAlzheimer's disease (AD) is a highly inheritable neurological disorder, and brain imaging genetics (BIG) has become a rapidly advancing field for comprehensive understanding its pathogenesis. However, most of the existing approaches underestimate the complexity of the interactions among factors that cause AD. To take full appreciate of these complexity interactions, we propose BIGFormer, a graph Transformer with local structural awareness, for AD diagnosis and identification of pathogenic mechanisms. Specifically, the factors interaction graph is constructed with lesion brain regions and risk genes as nodes, where the connection between nodes intuitively represents the interaction between nodes. After that, a perception with local structure awareness is built to extract local structure around nodes, which is then injected into node representation. Then, the global reliance inference component assembles the local structure into higher-order structure, and multi-level interaction structures are jointly aggregated into a classification projection head for disease state prediction. Experimental results show that BIGFormer demonstrated superiority in four classification tasks on the AD neuroimaging initiative dataset and proved to identify biomarkers closely intimately related to AD. Qi Zou 0003, Junliang Shang, Jin-Xing Liu 0001, Rui Gao 0006 |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | SCSMD: Single Cell Consistent Clustering based on Spectral Matrix DecompositionabstractCluster analysis, a pivotal step in single-cell sequencing data analysis, presents substantial opportunities to effectively unveil the molecular mechanisms underlying cellular heterogeneity and intercellular phenotypic variations. However, the inherent imperfections arise as different clustering algorithms yield diverse estimates of cluster numbers and cluster assignments. This study introduces Single Cell Consistent Clustering based on Spectral Matrix Decomposition (SCSMD), a comprehensive clustering approach that integrates the strengths of multiple methods to determine the optimal clustering scheme. Testing the performance of SCSMD across different distances and employing the bespoke evaluation metric, the methodological selection undergoes validation to ensure the optimal efficacy of the SCSMD. A consistent clustering test is conducted on 15 authentic scRNA-seq datasets. The application of SCSMD to human embryonic stem cell scRNA-seq data successfully identifies known cell types and delineates their developmental trajectories. Similarly, when applied to glioblastoma cells, SCSMD accurately detects pre-existing cell types and provides finer sub-division within one of the original clusters. The results affirm the robust performance of our SCSMD method in terms of both the number of clusters and cluster assignments. Moreover, we have broadened the application scope of SCSMD to encompass larger datasets, thereby furnishing additional evidence of its superiority. The findings suggest that SCSMD is poised for application to additional scRNA-seq datasets and for further downstream analyses. Ran Jia, Ying-Zan Ren, Ponian Li, Rui Gao 0006, Yusen Zhang 0002 |
Briefings Bioinform. | 4 |
| 2024 | LogicGep: Boolean networks inference using symbolic regression from time-series transcriptomic profiling dataabstractReconstructing the topology of gene regulatory network from gene expression data has been extensively studied. With the abundance functional transcriptomic data available, it is now feasible to systematically decipher regulatory interaction dynamics in a logic form such as a Boolean network (BN) framework, which qualitatively indicates how multiple regulators aggregated to affect a common target gene. However, inferring both the network topology and gene interaction dynamics simultaneously is still a challenging problem since gene expression data are typically noisy and data discretization is prone to information loss. We propose a new method for BN inference from time-series transcriptional profiles, called LogicGep. LogicGep formulates the identification of Boolean functions as a symbolic regression problem that learns the Boolean function expression and solve it efficiently through multi-objective optimization using an improved gene expression programming algorithm. To avoid overly emphasizing dynamic characteristics at the expense of topology structure ones, as traditional methods often do, a set of promising Boolean formulas for each target gene is evolved firstly, and a feed-forward neural network trained with continuous expression data is subsequently employed to pick out the final solution. We validated the efficacy of LogicGep using multiple datasets including both synthetic and real-world experimental data. The results elucidate that LogicGep adeptly infers accurate BN models, outperforming other representative BN inference algorithms in both network topology reconstruction and the identification of Boolean functions. Moreover, the execution of LogicGep is hundreds of times faster than other methods, especially in the case of large network inference. Dezhen Zhang, Shuhua Gao, Zhi-Ping Liu, Rui Gao 0006 |
Briefings Bioinform. | 4 |
| 2023 | MaxCLK: discovery of cancer driver genes via maximal clique and information entropy of modulesabstractMOTIVATION: Cancer is caused by the accumulation of somatic mutations in multiple pathways, in which driver mutations are typically of the properties of high coverage and high exclusivity in patients. Identifying cancer driver genes has a pivotal role in understanding the mechanisms of oncogenesis and treatment. RESULTS: Here, we introduced MaxCLK, an algorithm for identifying cancer driver genes, which was developed by an integrated analysis of somatic mutation data and protein-protein interaction (PPI) networks and further improved by an information entropy index. Tested on pancancer and single cancers, MaxCLK outperformed other existing methods with higher accuracy. About pancancer, we predicted 154 driver genes and 787 driver modules. The analysis of co-occurrence and exclusivity between modules and pathways reveals the correlation of their combinations. Overall, our study has deepened the understanding of driver mechanism in PPI topology and found novel driver genes. AVAILABILITY AND IMPLEMENTATION: The source codes for MaxCLK are freely available at https://github.com/ShandongUniversityMasterMa/MaxCLK-main. Jian Liu 0039, Fubin Ma, Yongdi Zhu, Naiqian Zhang, Lingming Kong, Haiyan Cong, Rui Gao 0006, Yusen Zhang 0002 |
Bioinform. | 8 |
| 2023 | A Tensor Method Based on Enhanced Tensor Nuclear Norm and Hypergraph Laplacian Regularization for Pan-Cancer Omics Data AnalysisabstractAs a powerful data representation technique, tensor robust principal component analysis (TRPCA) has been widely used for clustering and feature selection tasks. However, it ignores the significant difference in singular values of tensor data and the manifold information contained in different views, thereby causing serious degradation of conventional TRPCA performance. In this paper, a novel tensor method based on enhanced tensor nuclear norm and hypergraph Laplacian regularization (ETHLR) is developed to address the above problem. ETHLR can jointly learn the prior knowledge of singular values and high-order manifold structures in the unified tensor space and the view-specific feature spaces, respectively. Specifically, the enhanced tensor nuclear norm, namely, the weighted tensor Schatten p-norm, is used to shrink the singular values by fully considering the salient difference information of singular values and the complementary information embedded in the tensor space; the hypergraph Laplacian constraint helps encode high-order geometric structures among multiple samples in the nonlinear view-specific feature space. Furthermore, we employ inexact augmented Lagrange multipliers (ALM) to optimize the ETHLR method. Numerous experiments on pan-cancer omics data show that the superiority of ETHLR over several state-of-the-art competitors. Na Yu 0004, Yusen Zhang 0002, Rui Gao 0006 |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | A Semi-Supervised Learning Algorithm for Predicting MiRNA-Disease AssociationabstractThe dysregulation of miRNAs can lead to disease. Modeling miRNA-disease association is important in understanding the pathogenesis of diseases. The generation of multiple types of miRNA-disease association data provides us an opportunity to in-depth study of the interaction mechanism between miRNA and disease. However, the existing methods usually only focus on miRNA-disease interaction, which ignores the relationship types of their association. In this paper, we studied the feasibility of tensor robust principal component analysis (TRPCA) to model miRNA-disease-type associations. Unlike the binary association of, we use the triple association ofto better describe the pathogenesis of the disease. Multi-view auxiliary information (semantics of diseases, functions of miRNAs and Gaussian interaction profile kernel) are applied to fully consider the complexity of biological processes. The proposed semi-supervised learning model combines TRPCA and label propagation. The 10-fold cross validation (CV) and case study experiments on HMDD v2.0 dataset demonstrate that the proposed model is superior to the other benchmark methods. Na Yu 0004, Zhi-Ping Liu, Rui Gao 0006 |
BIBM | 3 |
| 2021 | FS-GBDT: identification multicancer-risk module via a feature selection algorithm by integrating Fisher score and GBDTabstractCancer is a highly heterogeneous disease caused by dysregulation in different cell types and tissues. However, different cancers may share common mechanisms. It is critical to identify decisive genes involved in the development and progression of cancer, and joint analysis of multiple cancers may help to discover overlapping mechanisms among different cancers. In this study, we proposed a fusion feature selection framework attributed to ensemble method named Fisher score and Gradient Boosting Decision Tree (FS-GBDT) to select robust and decisive feature genes in high-dimensional gene expression datasets. Joint analysis of 11 human cancers types was conducted to explore the key feature genes subset of cancer. To verify the efficacy of FS-GBDT, we compared it with four other common feature selection algorithms by Support Vector Machine (SVM) classifier. The algorithm achieved highest indicators, outperforms other four methods. In addition, we performed gene ontology analysis and literature validation of the key gene subset, and this subset were classified into several functional modules. Functional modules can be used as markers of disease to replace single gene which is difficult to be found repeatedly in applications of gene chip, and to study the core mechanisms of cancer. Da Xu 0005, Kaijing Hao, Yusen Zhang 0002, Wei Chen 0039, Jiaguo Liu, Rui Gao 0006, Chuanyan Wu, Yang De Marinis |
Briefings Bioinform. | 7 |
| 2019 | PTPD: predicting therapeutic peptides by deep learning and word2vecabstract*: Background In the search for therapeutic peptides for disease treatments, many efforts have been made to identify various functional peptides from large numbers of peptide sequence databases. In this paper, we propose an effective computational model that uses deep learning and word2vec to predict therapeutic peptides (PTPD). *: Results Representation vectors of all k-mers were obtained through word2vec based on k-mer co-existence information. The original peptide sequences were then divided into k-mers using the windowing method. The peptide sequences were mapped to the input layer by the embedding vector obtained by word2vec. Three types of filters in the convolutional layers, as well as dropout and max-pooling operations, were applied to construct feature maps. These feature maps were concatenated into a fully connected dense layer, and rectified linear units (ReLU) and dropout operations were included to avoid over-fitting of PTPD. The classification probabilities were generated by a sigmoid function. PTPD was then validated using two datasets: an independent anticancer peptide dataset and a virulent protein dataset, on which it achieved accuracies of 96% and 94%, respectively. *: Conclusions PTPD identified novel therapeutic peptides efficiently, and it is suitable for application as a useful tool in therapeutic peptide design. Chuanyan Wu, Rui Gao 0006, Yusen Zhang 0002, Yang De Marinis |
BMC Bioinform. | 2 |
| 2019 | Predicting FAD Interacting Residues with Feature Selection and Comprehensive Sequence DescriptorsabstractThe function of a flavoprotein is determined to a great extent by the binding sites on its surface that interacts with flavin adenine dinucleotide (FAD). Malfunction or dysregulation of FAD binding leads to a series of diseases. Therefore, accurately identifying FAD interacting residues (FIRs) provides insights into the molecular mechanisms of flavoprotein-related biological processes and disease progression. In this paper, a new computational method is proposed for identifying FIRs from protein sequences. Various sequence-derived discriminative features are explored. We analyze the distinctions of these features between FIRs and non-FIRs. We also investigate the predictive capabilities of both individual features and combinations of features. A relief algorithm followed by incremental feature selection (relief-IFS) is then adopted to search the optimal features. Finally, a random forest (RF) module is used to predict FIRs based on the optimal features. Using a 5-fold cross-validation test, the proposed method performs well, with a sensitivity of 0.847, a specificity of 0.933, an accuracy of 0.890, and a Matthews correlation coefficient (MCC) of 0.782, thereby outperforming previous methods. These results indicate that our method is relatively successful at predicting FIRs. Runtao Yang, Chengjin Zhang, Rui Gao 0006, Lina Zhang 0001, Qing Song 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2016 | Prediction of aptamer-protein interacting pairs using an ensemble classifier in combination with various protein sequence attributesabstractBACKGROUND: Aptamer-protein interacting pairs play a variety of physiological functions and therapeutic potentials in organisms. Rapidly and effectively predicting aptamer-protein interacting pairs is significant to design aptamers binding to certain interested proteins, which will give insight into understanding mechanisms of aptamer-protein interacting pairs and developing aptamer-based therapies. RESULTS: In this study, an ensemble method is presented to predict aptamer-protein interacting pairs with hybrid features. The features for aptamers are extracted from Pseudo K-tuple Nucleotide Composition (PseKNC) while the features for proteins incorporate Discrete Cosine Transformation (DCT), disorder information, and bi-gram Position Specific Scoring Matrix (PSSM). We investigate predictive capabilities of various feature spaces. The proposed ensemble method obtains the best performance with Youden's Index of 0.380, using the hybrid feature space of PseKNC, DCT, bi-gram PSSM, and disorder information by 10-fold cross validation. The Relief-Incremental Feature Selection (IFS) method is adopted to obtain the optimal feature set. Based on the optimal feature set, the proposed method achieves a balanced performance with a sensitivity of 0.753 and a specificity of 0.725 on the training dataset, which indicates that this method can solve the imbalanced data problem effectively. To evaluate the prediction performance objectively, an independent testing dataset is used to evaluate the proposed method. Encouragingly, our proposed method performs better than previous study with a sensitivity of 0.738 and a Youden's Index of 0.451. CONCLUSIONS: These results suggest that the proposed method can be a potential candidate for aptamer-protein interacting pair prediction, which may contribute to finding novel aptamer-protein interacting pairs and understanding the relationship between aptamers and proteins. Lina Zhang 0001, Chengjin Zhang, Rui Gao 0006, Runtao Yang, Qing Song 0002 |
BMC Bioinform. | 3 |