VLDB 2026 Research / reviewers in the wild / expert
Gaoshi Li
dblp:189/1390
· DBLP profile ↗
37ranked-venue papers
2as first author
34since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 2 first-author · 16 since 2021Artificial intelligence and machine learning · 9 · 8 since 2021Theory of computation · 5 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Two-Round Probabilistic Diagnosis Algorithm for fault identification
Wenfei Liu, Jiafei Liu 0001, Chia-Wei Lee, Sun-Yuan Hsieh, Jingli Wu, Gaoshi Li |
Discret. Appl. Math. | 6 |
| 2026 | Learning-based network diagnostics: Handling high fault densities with PMC/MM* model
Wenfei Liu, Jiafei Liu 0001, Jingli Wu, Chia-Wei Lee, Dajin Wang, Gaoshi Li |
Expert Syst. Appl. | 6 |
| 2026 | CDMI-NTDI: Cancer driver module identification via network topology and deep interaction features
Jingli Wu, Yanhua Huang, Gaoshi Li, Jiafei Liu 0001, Haize Hu |
Neurocomputing | 3 |
| 2026 | A novel influence rank algorithm in complex networks
Xinbang Cheng, Jiafei Liu 0001, Chia-Wei Lee, Sun-Yuan Hsieh, Jingli Wu, Gaoshi Li |
Inf. Sci. | 6 |
| 2026 | GDCC: scRNA-seq data imputation via Graph-cGAN based dual conditional guidance with constraint training
Gaoshi Li, Jingli Wu, Jiafei Liu 0001 |
Knowl. Based Syst. | 2 |
| 2026 | A Multi-Attribute Adaptive Fault Diagnosis Framework for Star NetworksabstractWith the proliferation of interconnection networks in mission-critical systems ranging from cloud computing infrastructures to large-scale data centers, the escalating structural complexity has intensified network vulnerability to malicious attacks and cyber warfare incidents. This article establishes a theoretical framework for evaluating network self-diagnostic capability through a novelh-extrar-component diagnosability metric, denoted as$\widehat{ec}_{r}^{h}(G)$, which quantifies a network’s resilience under compound fault patterns. The proposed metric requires that after removing specific nodes, the remaining subgraph is required to preserve at leastrconnected components where every component maintains a node count exceedingh. Through rigorous combinatorial analysis, we derive closed-form expressions for star networks$S_{n}$,$\widehat{ec}_{2}^{1}(S_{n}) = 4n - 9$and$\widehat{ec}_{3}^{1}(S_{n}) = 6n - 15$when$n \ge 6$, establishing the tight diagnosability bounds for this fundamental network topology. To enable practical implementation, we design a Trial System-based Fault Diagnosis Algorithm (TSFD) that features adaptive syndrome verification and parallel fault localization mechanisms. Extensive simulations demonstrate the accuracy of 98.99% fault detection with linear-time complexity$O(Nd)$inn-dimensional star networks. This work advances network reliability theory by introducing a multi-feature diagnosability measure for system-level diagnosis and developing an efficient diagnosis algorithm validated through large-scale network emulation. Wenfei Liu, Jiafei Liu 0001, Eddie Cheng 0001, Sun-Yuan Hsieh, Jingli Wu, Gaoshi Li |
IEEE Trans. Computers | 6 |
| 2026 | Essential Proteins Prediction Using Features Synergy Model and GO Pure CentralityabstractEssential proteins are a crucial component of living organisms, and their absence will lead to cell death or reproductive arrest. Discovering these proteins can propel advancements in synthetic biology and facilitate the development of novel antibiotics and therapies for various diseases. However, current computational methods suffer from two major drawbacks that hinder their discovery rate: one is the significant noise in protein-protein interaction (PPI) data, and the other is the inadequate consideration of feature relationships. To enhance identification capabilities, this study proposes a novel essential protein prediction method, Feature Synergy Method (FSM), which leverages a features synergy model and GO pure centrality. The FSM is described as follows:Firstly, based on the principle of co-expression, gene expression data are integrated with the original PPI network to construct a pure PPI network (PPIN). Subsequently, GO annotation data are employed to calculate GO_sim weights for the interactions within the original PPI network, forming a GS_PIN. The PPIN and GS_PIN are then fused to establish the GS_PPIN, which helps mitigate the impact of noise in PPI data. Secondly, a new centrality measure, GO pure centrality (GPC), is designed based on this GO similarity-weighted pure PPI network. Thirdly, an evolutionary conservation score (ECS) is extracted from subcellular localization and orthologous proteins data. Fourthly, after analyzing the relationship between GPC and ECS, a novel fusion model, the features synergy model, is developed to integrate GPC and ECS, ultimately leading to the proposal of the new essential protein prediction method, FSM. To validate the performance of FSM, six computational methods (PeC, WDC, ION, NCCO, E_POC, and JDC) and six centrality measures (NC, IC, EC, SC, CC, and DC) were evaluated on three distinct yeast datasets. The results demonstrate that FSM achieves a higher essential protein identification rate. Similarly, GPC identifies more essential proteins compared to the six centrality-based approaches (NC, IC, EC, SC, CC, and DC). Xinlong Luo 0002, Gaoshi Li, Zhipeng Hu, Jingli Wu, Wei Peng 0004, Jiafei Liu 0001, Xiaoshu Zhu |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2026 | A Deep Learning Framework for Identifying Essential Proteins Based on Vision TransformerabstractEssential proteins are fundamental to the reproduction and survival of cells, and if they are killed, the cells will stop reproducing or die. Many computational methods of identifying essential proteins are proposed which fuse a large number of features from multi-omics data. Some of them extract features from subcellular localization data by subjectively selecting certain subcellular locations. Meanwhile, there is still room to improve the identification rate of essential proteins. In this paper, a new deep learning framework for identifying essential proteins based on Vision Transformer is proposed, named EPViT. Firstly, topological features are extracted from the protein-protein interaction network. Secondly, a feature matrix is designed from the subcellular localization information without subjective factor. Then, the two classes of features are fused into a new feature matrix by outer product operation. Finally, the new feature matrix is input into the Vision Transformer model to discover essential proteins. The results show that EPViT has the highest recognition rate among the comparison experiments on yeast data. Gaoshi Li, Jingli Wu, Wei Peng 0004, Jiafei Liu 0001, Xiaoshu Zhu |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2026 | The h-extra r-component connectivity for a class of interconnection networks
Jiafei Liu 0001, Dajin Wang, Jingli Wu, Gaoshi Li |
Theor. Comput. Sci. | 5 |
| 2026 | An efficient two-stage diagnostic algorithm for assessing system reliability
Chunjian Liang, Jiafei Liu 0001, Chia-Wei Lee, Jingli Wu, Gaoshi Li |
Theor. Comput. Sci. | 5 |
| 2026 | A Novel Conditional Diagnostic Scheme for Hypercube-Based Multiprocessor SystemsabstractWith the scale of multiprocessor systems constantly increasing, the large number of interconnected processors (or nodes) makes faulty nodes inevitable. The fault diagnosis of multiprocessor systems therefore is a key technique for the system’s robustness. In this paper, we first propose a novel diagnostic metric, the$h$-extra$r$-component diagnosability, denoted$ECD^{h}_{r}(G)$, which characterizes one special pattern of faults. We derive some theoretical results for the ECD of hypercube, denoted$ECD^{h}_{r}(Q_{n})$, under the PMC model. Diagnostic algorithms is proposed and implemented to detect faulty nodes that will disconnect hypercube$Q_{n}$into$r$components each containing at least$h+1$nodes. We also test the ECD-PMC algorithm to the hypercube network with different number of faulty processors satisfying the$h$-extra$r$-component condition. Extensive simulation results show that our proposed method achieves very good performance in terms of ACCR, TPR, FPR, and TNR. Jiafei Liu 0001, Dajin Wang, Wenfei Liu, Jingli Wu, Gaoshi Li |
IEEE Trans. Netw. | 6 |
| 2026 | The $g$-Good-Neighbor $r$-Component Diagnosability of Hypercube - Theoretical and Algorithmic ApproachesabstractThe proliferation of interconnection networks has intensified the demand for robust fault diagnosis methodologies. Although existing research focuses predominantly on single-condition diagnosability metrics, these approaches often fail to capture hybrid failure scenarios in large-scale networks. To provide a more comprehensive and realistic resilience assessment, this article introduces a new diagnosability metric termed the$g$-good-neighbor$r$-component diagnosability, denoted by$D_{g,r}(G)$. This metric imposes two stringent constraints on the network after removing a faulty node set$F$: i) the residual network must contain at least$r$connected components, and ii) every fault-free node must retain at least$g$fault-free neighbors. We focus on the hypercube ($Q_{n}$), a prevalent interconnection architecture renowned for its high symmetry, scalability, and fault tolerance. Under the PMC and MM* diagnostic models, we establish the exact value$D_{2,2}(Q_{n}) = 8n - 21$for$n \geq 24$. Leveraging the distinct characteristics of the PMC and MM* models, we propose two scalable fault localization algorithms tailored for hypercube architectures. Simulation experiments on$Q_{n}$networks demonstrate that the proposed framework achieves approximately 100% true positive rate (TPR) when faulty nodes constitute$\leq 20\%$of the network, maintaining TPR$> 98.7\%$even as fault densities approach 50% . Yuankang Mao, Jiafei Liu 0001, Sun-Yuan Hsieh, Jingli Wu, Gaoshi Li |
IEEE Trans. Reliab. | 5 |
| 2025 | Identification of Potential Cancer Driver Genes Based on Modular Dysregulated GenesabstractThe occurrence and progression of cancer are generally considered to result from the accumulation of driver gene mutations. Accurately identifying these genes is crucial for elucidating tumor mechanisms and developing targeted therapies. We propose IMDG, a method for identifying cancer driver genes based on modular dysregulated gene networks that integrates multi-layer scoring. IMDG first constructs a weighted proteinprotein interaction (PPI) network based on the gene-miRNA network. It then extracts features using node embedding to build a relationship network for dysregulated genes and scores genes based on the association between potential driver genes and modules of the dysregulated gene network. Subsequently, it reassesses gene mutations through random walks on the weighted PPI network and integrates network connectivity to derive the final gene scores. We compared IMDG with six existing driver gene prioritization methods on two real-world cancer datasets. Results demonstrate that IMDG exhibits optimal identification performance in most cases. Genes prioritized by IMDG not only show higher concordance with benchmark databases compared to those identified by other methods, but some also exhibit strong relevance to cancer. Zheng Deng, Jingli Wu, Xiaorong Chen, Gaoshi Li |
BIBM | 4 |
| 2025 | MNMO: discover driver genes from a multi-omics data based-multi-layer networkabstractMOTIVATION: Cancer as a public health problem is driven by genomic variations in "cancer driver" genes. The identification of driver genes is critical for the discovery of key biomarkers and the development of personalized therapy. RESULTS: We propose a prediction method MNMO: a multi-layer network model based on multi-omics data. MNMO firstly constructs a dynamically adjusted four-layer network composed of miRNAs and three kinds of genes with different features. Then three kinds of scores, i.e. control capacity, mutation score, and network score, are devised and calculated by harmonic mean to produce the integrated gene score. Experiments were performed on three kinds of real cancer data to compare the identification performance of method MNMO with that of six state-of-the-art ones. The results indicate that method MNMO presents the best identification performance under most circumstances. The genes prioritized by method MNMO not only have a better match to the benchmark ones than those identified by the other methods, but also are all associated with the development and progression of cancers. In addition, some extended versions of method MNMO can further achieve better performance on most evaluation metrics for some specific datasets. They may be more conducive to identifying tissue-specific genes, which has been verified through a number of experiments. AVAILABILITY AND IMPLEMENTATION: The source code and the R package "MNMO" are available at https://github.com/Zheng-D/MNMO. The dataset and code are archived at https://doi.org/10.5281/zenodo.14969986. Zheng Deng, Jingli Wu, Xiaorong Chen, Gaoshi Li, Jiafei Liu 0001, Zhipeng Hu, Rongyuan Li, Wansu Deng |
Bioinform. | 4 |
| 2025 | A Deep Learning Framework for Identifying Cancer Driver Genes Based on Transformer and Graph Convolutional NetworkabstractCorrect identification of cancer driver genes plays a significant role in cancer research. The advancement of graph neural network (GNN) research has led to the emergence of many high-performance cancer driver gene prediction methods. However, GNN-based methods frequently overlook the importance of capturing global information. Additionally, as GNN layers increase, the feature representation of genes begins to become overly smooth. These problems hinder the effectiveness of GNN-based identification methods. In this study, we introduce TGCN, a method integrating Transformer and graph convolutional network (GCN), aiming to address these issues and improve cancer driver gene identification. First, we composed multivariate feature matrices of genes from multi-omics data and multi-dimensional gene association networks. Second, we constructed a Transformer module to enrich gene feature representations. Finally, we utilized Chebyshev GCN to yield the identification results. The experimental results demonstrate that TGCN outperforms representative methods in identifying driver genes for both pan-cancer and single-type cancers. Gaoshi Li, Jingli Wu, Jiafei Liu 0001, Haize Hu, Qiyong Zhu |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | Using Multi-Feature Weak Consensus Model to Discover Essential ProteinsabstractEssential proteins play an essential role in cell survival and replication. Currently, more and more computational methods are developed to identify essential proteins, which overcome the time-consuming, costly and inefficient shortcomings with biological experimental methods. In order to improve the recognition rate, some new methods by fusing multiple features are developed, but they seldom consider the connection among features. After analyzing a large number of methods based on multi-feature fusion, a phenomenon among features is found, called weak consensus, then a weak consensus model to fuse these features is proposed in this paper. After analyzing the relationship between a protein and its neighbors in protein-protein interaction networks, a new centrality, namely neighborhood aggregation centrality(NAC) is developed in this paper. Then, a Max-Min strategy is used to integrate NAC with Pearson correlation coefficient and Jaccard similarity coefficient based on gene expression data to obtain local importance score. In addition, orthologous feature score is used to measure proteins conservation. Finally, by using the weak consensus model to fuse orthologous feature score with local importance score, a new method WOL is proposed in this paper. Then experiments are performed on S.cerevisiae data. The results show that compared with WDC, PeC, ION, JDC, NCCO and E_POC, WOL has a higher recognition rate. Zhipeng Hu, Gaoshi Li, Xinlong Luo 0002, Jiafei Liu 0001, Jingli Wu, Wei Peng 0004, Xiaoshu Zhu |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | scPEGEnhanced Graph Convolutional Sparse Subspace Clustering Method for scRNA-Seq DataabstractThe identification of cell types by clustering single-cell RNA sequencing (scRNA-seq) data is a fundamental step in the downstream analysis of single-cell data. However, great challenges remain owing to the inherent characteristics of scRNA-seq data, including high dimensionality, high noise, and high sparsity. In this study, we propose a proximity enhanced graph convolutional sparse subspace clustering method scPEGSSC for scRNA-seq data. Method scPEGSSC generates the similarity matrix with the self-expression matrix (SEM) learned from a graph autoencoder, and enhances it further through its square. Experiments were performed on thirteen real biological datasets. The experimental results indicate compared with eleven state-of-the-art single-cell clustering methods, method scPEGSSC have attained superior performance across most datasets. Jingli Wu, Xiaopeng Wei, Gaoshi Li, Jiafei Liu 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | Identifying Cancer Driver Genes Using a Neural Network Framework With Cross-Attention MechanismabstractIdentifying cancer driver genes can accelerate the discovery of drug targets and the development of cancer therapies. Recent research methods improve the accuracy of identifying cancer driver genes by using deep learning framework. However, due to ignore the connection among learned features, they usually have weak feature representations that limits further improvement in the accuracy of identifying cancer driver genes. In this work, we propose a graph neural network framework combining graph convolutional network, Transformer with cross-attention, and multi-layer perceptron classifier, called GTCM, to improve the accuracy of identifying cancer driver genes. Specifically, GTCM first uses graph convolutional network to learn gene feature representations from three different gene association networks. Second, to enhance the feature representations of cancer driver genes, GTCM adopts Transformer with cross-attention to dynamically learn the connections between different feature sets. Finally, GTCM predicts cancer driver genes using multi-layer perceptron classifier. Ablation experiments prove that Transformer with cross-attention effectively improves the feature representations learned from graph convolutional network and further improves the identification rate. Compared with existing representative methods, GTCM exhibits excellent performance in terms of area under the receiver operating characteristic curves and area under precision-recall curves. Gaoshi Li, Jingli Wu, Jiafei Liu 0001, Haize Hu, Qiyong Zhu |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | The reliability of (n,k)-star network in terms of non-inclusive fault pattern
Qigong Chen, Jiafei Liu 0001, Chia-Wei Lee, Jingli Wu, Gaoshi Li |
Theor. Comput. Sci. | 5 |
| 2025 | A novel fault diagnostic algorithm with multiple characteristics for multiprocessor systems
Gaotao Ge, Jiafei Liu 0001, Dajin Wang, Jingli Wu, Gaoshi Li |
Theor. Comput. Sci. | 5 |
| 2025 | An analysis on component reliability of (n, k)-star networks
Zhihang Wang, Jiafei Liu 0001, Chia-Wei Lee, Jingli Wu, Gaoshi Li |
J. Supercomput. | 5 |
| 2025 | A novel ranking scheme for identifying influential nodes in complex networks
Jiafei Liu 0001, Dajin Wang, Jingli Wu, Gaoshi Li |
J. Supercomput. | 5 |
| 2024 | IntroGRN: Gene Regulatory Network Inference from Single-Cell RNA Data Based on Introspective VAE
Rongyuan Li, Jingli Wu, Gaoshi Li, Jiafei Liu 0001, Jinlu Liu, Junbo Xuan, Zheng Deng |
ISBRA (1) | 3 |
| 2024 | AGImpute: imputation of scRNA-seq data based on a hybrid GAN with dropouts identificationabstractMOTIVATION: Dropout events bring challenges in analyzing single-cell RNA sequencing data as they introduce noise and distort the true distributions of gene expression profiles. Recent studies focus on estimating dropout probability and imputing dropout events by leveraging information from similar cells or genes. However, the number of dropout events differs in different cells, due to the complex factors, such as different sequencing protocols, cell types, and batch effects. The dropout event differences are not fully considered in assessing the similarities between cells and genes, which compromises the reliability of downstream analysis. RESULTS: This work proposes a hybrid Generative Adversarial Network with dropouts identification to impute single-cell RNA sequencing data, named AGImpute. First, the numbers of dropout events in different cells in scRNA-seq data are differentially estimated by using a dynamic threshold estimation strategy. Next, the identified dropout events are imputed by a hybrid deep learning model, combining Autoencoder with a Generative Adversarial Network. To validate the efficiency of the AGImpute, it is compared with seven state-of-the-art dropout imputation methods on two simulated datasets and seven real single-cell RNA sequencing datasets. The results show that AGImpute imputes the least number of dropout events than other methods. Moreover, AGImpute enhances the performance of downstream analysis, including clustering performance, identifying cell-specific marker genes, and inferring trajectory in the time-course dataset. AVAILABILITY AND IMPLEMENTATION: The source code can be obtained from https://github.com/xszhu-lab/AGImpute. Xiaoshu Zhu, Shuang Meng, Gaoshi Li, Jianxin Wang 0001, Xiaoqing Peng |
Bioinform. | 3 |
| 2024 | A model and multi-core parallel co-evolution algorithm for identifying cancer driver pathways
Xiaorong Chen, Jingli Wu, Zheng Deng, Gaoshi Li |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Identification of Cancer Driver Genes based on Dynamic Incentive ModelabstractCancer is a complex genomic mutation disease, and identifying cancer driver genes promotes the development of targeted drugs and personalized therapies. The current computational method takes less consideration of the relationship among features and the effect of noise in protein-protein interaction(PPI) data, resulting in a low recognition rate. In this paper, we propose a cancer driver genes identification method based on dynamic incentive model, DIM. This method firstly constructs a hypergraph to reduce the impact of false positive data in PPI. Then, the importance of genes in each hyperedge in hypergraph is considered from three perspectives, network and functional score(NFS) is proposed. By analyzing the relation among features, the dynamic incentive model is proposed to fuse NFS, the differential expression score of mRNA and the differential expression score of miRNA. DIM is compared with some classical methods on breast cancer, lung cancer, prostate cancer, and pan-cancer datasets. The results show that DIM has the best performance on statistical evaluation indicators, functional consistency and the partial area under the ROC curve, and has good cross-cancer capability. Zhipeng Hu, Gaoshi Li, Xinlong Luo 0002, Wei Peng 0004, Jiafei Liu 0001, Xiaoshu Zhu, Jingli Wu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | Essential proteins identification based on weak consensus model and neighborhood aggregation centralityabstractEssential proteins play an essential role in cell survival and replication. Currently, more and more computational methods are developed to identify essential proteins, which overcome the time-consuming, costly and inefficient shortcomings with biological experimental methods. In order to improve the recognition rate, some new methods by fusing multiple features are developed, but they seldom consider the connection among features. After analyzing a large number of methods based on multi-feature fusion, a weak consensus model to fuse features is proposed in this paper. Then, this paper uses the weak consensus model to fuse protein-protein interaction network, gene expression data, and orthologous data, thus proposing a new method, WOL. Then experiments are performed on one S.cerevisiae dataset. The results show that compared with WDC, PeC, ION, JDC, NCCO and E_POC, WOL has a higher recognition rate. Zhipeng Hu, Gaoshi Li, Jingli Wu, Xinlong Luo 0002, Jiafei Liu 0001, Wei Peng 0004, Xiaoshu Zhu |
BIBM | 2 |
| 2023 | Mdwgan-gp: data augmentation for gene expression data based on multiple discriminator WGAN-GPabstractBACKGROUND: Although gene expression data play significant roles in biological and medical studies, their applications are hampered due to the difficulty and high expenses of gathering them through biological experiments. It is an urgent problem to generate high quality gene expression data with computational methods. WGAN-GP, a generative adversarial network-based method, has been successfully applied in augmenting gene expression data. However, mode collapse or over-fitting may take place for small training samples due to just one discriminator is adopted in the method. RESULTS: In this study, an improved data augmentation approach MDWGAN-GP, a generative adversarial network model with multiple discriminators, is proposed. In addition, a novel method is devised for enriching training samples based on linear graph convolutional network. Extensive experiments were implemented on real biological data. CONCLUSIONS: The experimental results have demonstrated that compared with other state-of-the-art methods, the MDWGAN-GP method can produce higher quality generated gene expression data in most cases. Rongyuan Li, Jingli Wu, Gaoshi Li, Jiafei Liu 0001, Junbo Xuan |
BMC Bioinform. | 3 |
| 2023 | Identifying driver pathways based on a parameter-free model and a partheno-genetic algorithmabstractBACKGROUND: Tremendous amounts of omics data accumulated have made it possible to identify cancer driver pathways through computational methods, which is believed to be able to offer critical information in such downstream research as ascertaining cancer pathogenesis, developing anti-cancer drugs, and so on. It is a challenging problem to identify cancer driver pathways by integrating multiple omics data. RESULTS: In this study, a parameter-free identification model SMCMN, incorporating both pathway features and gene associations in Protein-Protein Interaction (PPI) network, is proposed. A novel measurement of mutual exclusivity is devised to exclude some gene sets with "inclusion" relationship. By introducing gene clustering based operators, a partheno-genetic algorithm CPGA is put forward for solving the SMCMN model. Experiments were implemented on three real cancer datasets to compare the identification performance of models and methods. The comparisons of models demonstrate that the SMCMN model does eliminate the "inclusion" relationship, and produces gene sets with better enrichment performance compared with the classical model MWSM in most cases. CONCLUSIONS: The gene sets recognized by the proposed CPGA-SMCMN method possess more genes engaging in known cancer related pathways, as well as stronger connectivity in PPI network. All of which have been demonstrated through extensive contrast experiments among the CPGA-SMCMN method and six state-of-the-art ones. Jingli Wu, Qinghua Nie, Gaoshi Li, Kai Zhu 0009 |
BMC Bioinform. | 3 |
| 2023 | A model and cooperative co-evolution algorithm for identifying driver pathways based on the integrated data and PPI network
Kai Zhu 0009, Jingli Wu, Gaoshi Li, Xiaorong Chen, Michael Y. Luo |
Expert Syst. Appl. | 3 |
| 2022 | Identifying driver genes in cancer based on Pareto optimality consensusabstractAn important issue in cancer genomics is the identification of driver genes. It is significant for the discovery of key biomarkers and the development of effective personalized therapies. In this paper, a computated method PGScore is proposed. It scores genes at multilayer and integrates the scores to identify cancer driver genes based on Pareto Optimality Consensus(POC) strategy. PGScore uses random walks to reevaluate gene mutations, and integrates differential expression of mRNA and miRNA in normal and cancer samples. It measures the centrality of the gene in the network according to the weight of its direct and indirect neighbors, and finally integrates the above layers to get the final priority of the genes. We compare PGScore with state-of-the-art cancer driver genes prioritization methods on two real cancer datasets. The results show that PGScore can obtain better performance in identification accuracy and the partial area under the ROC(pAUC) curve on multiple reference databases. Zheng Deng, Jingli Wu, Xiaorong Chen, Gaoshi Li |
BIBM | 4 |
| 2022 | A model and algorithm for identifying driver pathways based on weighted non-binary mutation matrixabstractAbstract It is generally acknowledged that driver pathway plays a decisive role in the occurrence and progress of tumors, and the identification of driver pathways has become imperative for precision medicine or personalized medicine. Due to the inevitable sequencing error, the noise contained in single omics cancer data usually plays a negative effect on identification. It is a feasible approach to take advantage of multi-omics cancer data rather than a single one now that large amounts of multi-omics cancer data have become available. The identification of driver pathways by integrating multi-omics cancer data has attracted attention of researchers in bioinformatics recently. In this paper, a weighted non-binary mutation matrix is constructed by integrating copy number variations, somatic mutations and gene expressions. Based on the weighted non-binary mutation matrix, a new identification model is proposed through defining new measurements of coverage and exclusivity. Then, a cooperative coevolutionary algorithm CGA-MWS is put forward for solving the presented model. Both real cancer data and simulated one were used to conduct comparisons among methods Dendrix, GA, iMCMC, MOGA, PGA-MWS and CGA-MWS. Compared with the pathways identified by the other five methods, more genes, belonging to the pathway identified by the CGA-MWS method, are enriched in a known signaling pathway in most cases. Simultaneously, the high efficiency of method CGA-MWS makes it practical in realistic applications. All of which have been verified through a number of experiments. Jingli Wu, Kai Zhu 0009, Gaoshi Li, Qirong Cai |
Appl. Intell. | 3 |
| 2022 | Identifying common driver modules by equilibrating coverage and mutual exclusivity across pan-cancer data
Jingli Wu, Gaoshi Li |
Neurocomputing | 3 |
| 2021 | IDM-SPS: Identifying driver module with somatic mutation, PPI network and subcellular localizationabstractMutation profiles together with prior knowledge such as interactions between genes/proteins provide abundant critical information for the identification of driver modules, which is very important for analyzing mutational heterogeneity in human cancers. Due to the negative effects of inevitable false positive interactions in the PPI network, subcellular localization data are exerted to filter out them firstly, and somatic mutation profiles are used to weight the retained interactions. Five novel recombination operators are introduced basing on the vertex degrees and the edge weights in the PPI network, and a parthenogenetic algorithm is devised for solving the presented identification model which takes into account network connectivity, mutual exclusivity, coverage, and hops between genes within a module. Extensive experimental results indicate that compared with two state-of-the-art computational methods Hotnet2 and MEXCOwalk, the proposed method exhibits competitive performance in most cases in terms of recovering known cancer genes, providing modules that have satisfied coverage and mutual exclusivity, and are enriched for mutations in specific cancer types. Many identified gene sets are involved in known signaling pathways, most of the implicated genes are oncogenes or tumor suppressors previously reported in the literature. In addition, the proposed method does identify many cancer related genes missed by methods Hotnet2 and MEXCOwalk, including some recognized genes covering many types of cancers but having low mutation frequency. Jingli Wu, Jifan Yang, Gaoshi Li |
Eng. Appl. Artif. Intell. | 3 |
| 2020 | Two novel models and a parthenogenetic algorithm for detecting common driver pathways from pan-cancer data
Jingli Wu, Gaoshi Li, Kai Zhu 0009, Qirong Cai |
Eng. Appl. Artif. Intell. | 3 |
| 2020 | United Neighborhood Closeness Centrality and Orthology for Predicting Essential ProteinsabstractIdentifying essential proteins plays an important role in disease study, drug design, and understanding the minimal requirement for cellular life. Computational methods for essential proteins discovery overcome the disadvantages of biological experimental methods that are often time-consuming, expensive, and inefficient. The topological features of protein-protein interaction (PPI) networks are often used to design computational prediction methods, such as Degree Centrality (DC), Betweenness Centrality (BC), Closeness Centrality (CC), Subgraph Centrality (SC), Eigenvector Centrality (EC), Information Centrality (IC), and Neighborhood Centrality (NC). However, the prediction accuracies of these individual methods still have space to be improved. Studies show that additional information, such as orthologous relations, helps discover essential proteins. Many researchers have proposed different methods by combining multiple information sources to gain improvement of prediction accuracy. In this study, we find that essential proteins appear in triangular structure in PPI network significantly more often than nonessential ones. Based on this phenomenon, we propose a novel pure centrality measure, so-called Neighborhood Closeness Centrality (NCC). Accordingly, we develop a new combination model, Extended Pareto Optimality Consensus model, named EPOC, to fuse NCC and Orthology information and a novel essential proteins identification method, NCCO, is fully proposed. Compared with seven existing classic centrality methods (DC, BC, IC, CC, SC, EC, and NC) and three consensus methods (PeC, ION, and CSC), our results on S.cerevisiae and E.coli datasets show that NCCO has clear advantages. As a consensus method, EPOC also yields better performance than the random walk model. Gaoshi Li, Min Li 0007, Jianxin Wang 0001, Yaohang Li, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2016 | Predicting essential proteins based on subcellular localization, orthology and PPI networksabstractBACKGROUND: Essential proteins play an indispensable role in the cellular survival and development. There have been a series of biological experimental methods for finding essential proteins; however they are time-consuming, expensive and inefficient. In order to overcome the shortcomings of biological experimental methods, many computational methods have been proposed to predict essential proteins. The computational methods can be roughly divided into two categories, the topology-based methods and the sequence-based ones. The former use the topological features of protein-protein interaction (PPI) networks while the latter use the sequence features of proteins to predict essential proteins. Nevertheless, it is still challenging to improve the prediction accuracy of the computational methods. RESULTS: Comparing with nonessential proteins, essential proteins appear more frequently in certain subcellular locations and their evolution more conservative. By integrating the information of subcellular localization, orthologous proteins and PPI networks, we propose a novel essential protein prediction method, named SON, in this study. The experimental results on S.cerevisiae data show that the prediction accuracy of SON clearly exceeds that of nine competing methods: DC, BC, IC, CC, SC, EC, NC, PeC and ION. CONCLUSIONS: We demonstrate that, by integrating the information of subcellular localization, orthologous proteins with PPI networks, the accuracy of predicting essential proteins can be improved. Our proposed method SON is effective for predicting essential proteins. Gaoshi Li, Min Li 0007, Jianxin Wang 0001, Jingli Wu, Fang-Xiang Wu, Yi Pan 0001 |
BMC Bioinform. | 1 |