EDBT 2026 Demo / reviewers in the wild / expert
Zhi-Ping Liu
dblp:14/6575
· DBLP profile ↗
46ranked-venue papers
9as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 41 · 8 first-author · 25 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prognostic biomarker discovery via a connected network-constrained Cox proportional hazards modelabstractAbstract Biomarker discovery in biomedical sciences can be framed as feature selection in machine learning [1]. However, existing methods often overlook gene co-localization within regulatory interaction networks, leading to the identification of isolated biomarkers with limited biological interpretability [2]. Here, we present the Connected Network-regularized Cox proportional hazards model (CNet-Cox), which incorporates network connectivity constraints into sparse regularization to identify prognostic biomarkers for breast cancer (BRCA) on the discovery dataset from TCGA (1,092 patients), while explicitly accounting for patient survival time. CNet-Cox reveals the network structures of prognostic genes, evaluated in the internal validation dataset with a concordance index of 0.913, surpassing traditional regularized Cox methods. CNet-Cox shifts biomarker recognition from isolated to connected features within biomolecular networks and offers new biological insights. Furthermore, we established a six-gene BRCA prognostic risk scoring (PRS) metric and validated its robustness across six independent external validation datasets comprising 1,829 patients, and one spatial transcriptomic dataset containing 4,992 spots. The PRS score consistently demonstrated superior performance in patient/sample stratification across extensive and diverse validation datasets. Overall, our comprehensive downstream analyses underscore that CNet-Cox offers a novel approach for embedding network topology into feature selection, enabling the systematic discovery of key connected prognostic biomarkers. This significantly advances early detection and prognosis prediction, facilitating precision medicine for BRCA. References 1. Li L, Liu Z P. “Biomarker discovery from high-throughput data by connected network-constrained support vector machine.” Expert Systems with Applications 2023; 226: 120179. 2. Hartman E, Scott A M, Karlsson C, et al. “Interpreting biologically informed neural networks for enhanced proteomic biomarker discovery and pathway analysis.” Nature Communications 2023; 14(1): 5359. Wai-Ki Ching, Zhi-Ping Liu |
Briefings Bioinform. | 3 |
| 2026 | MTPrior: A Multi-Task Hierarchical Graph Embedding Framework for Prioritizing Hepatocellular Carcinoma-Associated Genes and Long Noncoding RNAsabstractHepatocellular carcinoma (HCC) represents a highly prevalent liver cancer, posing a substantial global health challenge. The prioritization of both coding genes and noncoding RNAs, such as long noncoding RNAs (lncRNAs), is paramount in unraveling the mechanisms of HCC and advancing diagnostics, prognostics and therapeutic strategies. The development of computational models for prioritizing cancer-associated RNAs plays a pivotal role in reducing reliance on costly and time-consuming experimental methodologies. However, most existing approaches focus on a single factor, such as genes, lncRNAs, or microRNAs (miRNAs), neglecting the interactions between coding genes and noncoding RNAs as well as their combined influence. Models capable of prioritizing multiple RNA types while accounting for these interactions remain scarce. In this study, we introduce MTPrior, a multi-task graph embedding prioritization model. Our approach is designed to achieve multi-task prioritization by constructing an adaptable framework that accommodates diverse tasks and refines the network structure tailored to specific tasks. It meticulously considers interactions between coding and noncoding RNAs, navigating efficient biological pathways to discover the most pertinent results. By analyzing extensive datasets from HCC patients, alongside a comprehensive inventory of genes and lncRNAs, we have developed a model that proficiently prioritizes and identifies the most relevant genes and lncRNAs associated with HCC, thereby streamlining research efforts towards key candidates for further investigation. Furthermore, an ablation study underscores the effectiveness of each component within our proposed method. The convincing results demonstrate that MTPrior outperforms other state-of-the-art methods in predicting disease-related genes and lncRNAs, highlighting its efficiency and advantages. Fatemeh Keikha, Zhi-Ping Liu |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | getDNB: identifying dynamic network biomarkers from time-varying gene regulations utilizing graph embedding techniquesabstractAbstract Aim Complex diseases remain difficult to detect early because conventional diagnostic strategies rely on static biomarkers that typically emerge at advanced stages. We aimed to develop a computational framework to systematically identify dynamic network biomarkers (DNBs) from temporally evolving gene regulatory networks. Methods We present getDNB, a graph embedding technique framework with three main steps: (1) constructing stage-specific regulatory networks to capture dynamic alterations in molecular interactions during disease progression; (2) employing graph convolutional networks (GCN) to project these high-dimensional networks into low-dimensional embeddings while preserving topological structure; (3) quantifying gene-level abnormality via K-means clustering and outlier scores, followed by network refinement using minimum dominating set and shortest path algorithms to ensure network connectivity and reduce redundancy. Additionally, a dynamic network index (DNI) was introduced to quantify temporal fluctuations in network disorder, providing a quantitative signal for critical transition states. Results Applied to a hepatocellular carcinoma (HCC) dataset, getDNB identified 33 robust DNBs and their interconnected network, achieving high predictive accuracy (AUROC = 0.929). The DNI curve exhibited a pronounced increase at the pre-disease stage, consistent with complex system transition theory predictions. Functional enrichment analysis revealed significant associations of these DNBs with key oncogenic pathways, including hepatocellular carcinoma, hepatitis B infection, and cell cycle regulation. Conclusion getDNB provides a powerful and generalizable approach for dynamic biomarker discovery. By integrating graph neural networks, anomaly detection, and network optimization, it offers mechanistic insights into complex disease progression and enables identification of early-warning signals with potential clinical translational value. References 1. Chen L, Liu R, Liu Z, Li M, Aihara K. ‘Detecting early-warning signals for sudden deterioration of complex diseases by dynamical network biomarkers.’ Scientific Reports 2012; 2: 342. 2. Wang T, Liu Z. ‘getDNB: identifying dynamic network biomarkers of hepatocellular carcinoma from time-varying gene regulations utilizing graph embedding techniques for anomaly detection.’ Bioinformatics 2025; 41(9): btaf518. Zhi-Ping Liu |
Briefings Bioinform. | 3 |
| 2025 | LogicSR: prior-guided symbolic regression for gene regulatory network inference from single-cell transcriptomics dataabstractDeciphering gene regulatory mechanisms from high-dimensional biology data remains a central challenge in modern systems biology, despite the growing availability of single-cell datasets. The difficulty stems partly from the sparsity and noise inherent in single-cell data and partly from the complexity of dynamic combinatorial regulation mediated by transcription factors. In this work, we introduce LogicSR, a computational framework that reconstructs gene regulatory networks from single-cell gene expression data with high accuracy by integrating the mechanistic interpretability of Boolean logical models with the equation-discovery capabilities of symbolic regression. It incorporates prior knowledge into a multi-objective Monte Carlo tree search (MCTS) framework, leveraging it to ensure biological plausibility and accelerate the search for optimal governing equations. LogicSR outperforms existing methods on both synthetic and real-world benchmark datasets. When applied to a human embryonic stem cell dataset, it demonstrates superior performance in elucidating complex combinatorial TF-target gene regulations and identifying key regulators. Dezhen Zhang, Zhi-Ping Liu |
Briefings Bioinform. | 2 |
| 2025 | Inferring cell-type-specific gene regulatory network from cellular transcriptomics data with GeneLink+abstractDeciphering cell-type-specific gene regulatory networks (ctGRNs) is crucial for elucidating fundamental biological processes, such as tissue development and cancer progression. However, accurately inferring ctGRNs from high-dimensional transcriptomic data poses a significant challenge, primarily due to issues like data sparsity, cell heterogeneity, and over-smoothing (i.e. the tendency of node features to become indistinguishable after many graph convolution layers) in deep learning models. To tackle these obstacles, we present GeneLink+, an innovative framework for ctGRN inference leveraging directed graph link prediction (i.e. inferring causal regulator-target edges) tasks. Building upon the robust predictive capabilities of its primary version, GENELink, GeneLink+ incorporates residual-GATv2 blocks, which synergize dynamic attention mechanisms with residual connections. This architecture effectively mitigates information loss during the aggregation process and preserves cell-type-specific gene features, thereby enhancing the identification of regulatory mechanisms as well as the model's interpretability. Furthermore, GeneLink+ uses a modified dot product scheme with learnable weight parameters to adaptively prioritize informative gene pairs when scoring regulatory relationships, thus enabling more precise causal edge attribution. Comprehensive benchmarking across seven datasets demonstrated that GeneLink+ either outperforms or matches the performance of existing state-of-the-art methods in terms of predictive accuracy and biological relevance. Additionally, applications to a wide array of transcriptomic data, encompassing single-cell ribonucleic acid sequencing, small nuclear ribonucleic acid sequencing, and spatially resolved transcriptomics, have unveiled pivotal causal regulatory relationships in blood immune cells, Alzheimer's disease, and breast cancer. Wei Zhang 0241, Bowen Shao, Wenbo Guo 0010, Jiaxin Lyu, Chuanyuan Wang, Zhi-Ping Liu |
Briefings Bioinform. | 8 |
| 2025 | getDNB: identifying dynamic network biomarkers of hepatocellular carcinoma from time-varying gene regulations utilizing graph embedding techniques for anomaly detectionabstractMOTIVATION: Early detection and timely intervention of hepatocellular carcinoma (HCC) are pivotal for improving patient prognosis. Current diagnostic approaches often detect HCC at later stages, thereby diminishing treatment efficacy. Recent advancements in high-throughput sequencing technology have vastly improved the identification of molecular markers via biological networks. However, existing methodologies frequently overlook the intricate gene interaction information in temporal gene regulatory networks. Therefore, our study proposes an algorithm model, getDNB, leveraging graph embedding technique (get) for anomaly detection in time-varying dynamic networks. The model aims to facilitate early HCC detection and propel precision medicine by recognizing dynamic network biomarker (DNB). RESULTS: We proposed the getDNB model, which utilizes graph convolutional networks for graph embedding, mapping high-dimensional gene regulatory networks to low-dimensional feature vector spaces. By calculating gene anomaly degrees through an outlier score, and using the minimum dominant set algorithm alongside with the shortest path algorithm, we discovered DNBs and their associated networks in HCC. The getDNB model successfully pinpointed 33 HCC DNBs, effectively differentiating various temporal stages of HCC progression, and demonstrated robustness across numerous real HCC datasets. Functional enrichment analysis unveiled that these DNBs play critical roles in HCC occurrence and development, outperforming widely used feature selection algorithms. AVAILABILITY AND IMPLEMENTATION: The source code and data can be found at https://github.com/zpliulab/getDNB. Zhi-Ping Liu |
Bioinform. | 2 |
| 2025 | Quantum-enhanced blockchain federated learning via quantum Byzantine agreement
Hao-Wen Liu, Zhi-Ping Liu, Hua-Lei Yin, Zeng-Bing Chen |
Sci. China Inf. Sci. | 2 |
| 2025 | TransMarker: Unveiling dynamic network biomarkers in cancer progression through cross-state graph alignment and optimal transportabstractThe identification of state specific biomarkers that reflect dynamic changes in gene regulatory networks is critical for understanding cancer progression and enhancing diagnostic precision. While multilayer network models have been proposed for analyzing disease evolution, most existing methods rely solely on topological features, neglecting structural rewiring and expression variability across disease states. In this study, we introduce TransMarker, a framework designed to detect genes with regulatory role transitions, those with meaningful shifts in regulatory roles during disease progression, as dynamic biomarkers via cross-state alignment of multi-state single-cell data. TransMarker encodes each disease state as a distinct layer in a multilayer graph, integrating prior interaction data with state-specific expression to construct attributed gene networks. Contextualized embeddings for each stage are generated for each state using Graph Attention Networks (GATs), and structural shifts are quantified via Gromov-Wasserstein optimal transport. Genes with significant changes are ranked using a Dynamic Network Index (DNI), which captures their regulatory variability. These prioritized biomarkers are then applied in a deep neural network for disease state classification. We validate our approach on synthetic simulated and real world dataset of gastric adenocarcinoma (GAC), to evaluate performance across diverse scenarios and assess generalizability. TransMarker outperforms existing multilayer network ranking techniques in classification accuracy, robustness, and biomarker relevance. Ablation studies confirm the contribution of each step to overall performance. Our findings suggest that combining regulatory rewiring, temporal expression dynamics, and cross-state alignment provides a powerful strategy for identifying biologically meaningful biomarkers and modeling disease progression at single cell resolution. Fatemeh Keikha, Chuanyuan Wang, Zhi-Ping Liu |
PLoS Comput. Biol. | 4 |
| 2025 | NetWalkRank: Cancer Driver Gene Prioritization in Multiplex Gene Regulatory Networks by a Random Walk ApproachabstractFinding and prioritizing cancer driver genes (CDGs) that disrupt normal cell functionality and contribute to cancer occurrence and development is a significant challenge in oncology. Integrating multiple information pertaining to the characteristics of each gene at different stages of the disease and incorporating multiple steps as individual layers in the model provides a more comprehensive understanding of each node or gene. Thus, it is reasonable to organize them into multiplex gene regulatory networks (GRNs). In this work, we present a network-based framework called NetWalkRank, for prioritizing CDGs in the multiplex GRNs with gene expression profiling data. The framework applies the concept of network propagation to calculate the relative impact of each gene in spreading abnormality throughout the multiplex GRNs. It was employed to give priority to the driver genes of hepatocellular carcinoma (HCC) in humans. The performance of NetWalkRank was demonstrated through the ranks and classifications assigned to the known CDGs, which validated its effectiveness. To showcase the predictive capabilities of our proposed framework, we trained a random forest model that utilizes the obtained scores to accurately predict CDGs. We compared the advantage and efficiency of our method with other well-known driver gene ranking methods through numerical experiments. The findings show that the usage of GRNs across various steps of multiplex networks in prioritizing and predicting CDGs is significant, as demonstrated by the efficiency and effectiveness of NetWalkRank. Fateme Keikha, Wai-Ki Ching, Zhi-Ping Liu |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | PGDTA: Predicting Drug-Target Affinity Using Three-Dimensional Structure of Protein Pocket and Graph Neural NetworkabstractDrug-Target Affinity (DTA) prediction plays a crucial role in drug discovery, and accurate DTA prediction can significantly reduce the cost of drug development. While most studies focus on the entire protein structure, they often overlook the local structure of protein pockets which play a vital role in DTA due to their direct interaction with drugs. At the methodological level, numerous deep learning approaches have been developed to predict DTA using protein and drug sequences or structures, yet the effective utilization of protein and drug features remains a pressing challenge. Our study proposes leveraging pre-trained models to represent sequence features of protein and drug separately. Subsequently, we construct a geometric graph neural network module capable of parallelizing diverse spatial structural information. We conducted experiments on three public datasets and compared our approach with current state-of-the-art (SOTA) methods, validating the effectiveness of our method. Furthermore, we compared the impact of entire proteins versus protein pockets on DTA, further affirming the reliability of our approach. Consequently, our method (called PGDTA) enhances the accuracy of DTA prediction, thereby aiding in improving the efficiency of the drug discovery process. Yunhai Li, Pengpai Li, Duanchen Sun, Zhi-Ping Liu |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | Predicting Drug-Target Affinity Using Protein Pocket and Graph Convolution Network
Yunhai Li, Pengpai Li, Duanchen Sun, Zhi-Ping Liu |
ISBRA (1) | 4 |
| 2024 | LogicGep: Boolean networks inference using symbolic regression from time-series transcriptomic profiling dataabstractReconstructing the topology of gene regulatory network from gene expression data has been extensively studied. With the abundance functional transcriptomic data available, it is now feasible to systematically decipher regulatory interaction dynamics in a logic form such as a Boolean network (BN) framework, which qualitatively indicates how multiple regulators aggregated to affect a common target gene. However, inferring both the network topology and gene interaction dynamics simultaneously is still a challenging problem since gene expression data are typically noisy and data discretization is prone to information loss. We propose a new method for BN inference from time-series transcriptional profiles, called LogicGep. LogicGep formulates the identification of Boolean functions as a symbolic regression problem that learns the Boolean function expression and solve it efficiently through multi-objective optimization using an improved gene expression programming algorithm. To avoid overly emphasizing dynamic characteristics at the expense of topology structure ones, as traditional methods often do, a set of promising Boolean formulas for each target gene is evolved firstly, and a feed-forward neural network trained with continuous expression data is subsequently employed to pick out the final solution. We validated the efficacy of LogicGep using multiple datasets including both synthetic and real-world experimental data. The results elucidate that LogicGep adeptly infers accurate BN models, outperforming other representative BN inference algorithms in both network topology reconstruction and the identification of Boolean functions. Moreover, the execution of LogicGep is hundreds of times faster than other methods, especially in the case of large network inference. Dezhen Zhang, Shuhua Gao, Zhi-Ping Liu, Rui Gao 0006 |
Briefings Bioinform. | 3 |
| 2024 | Network Activity Evaluation Reveals Significant Gene Regulatory Architectures During SARS-CoV-2 Viral Infection From Dynamic scRNA-seq DataabstractThe key to understand COVID-19 caused by SARS-CoV-2, which has caused massive deaths worldwide, is to reveal the gene activities at molecular level. Single-cell RNA-sequencing (scRNA-seq) technology allows us to capture gene expression at high resolution, thereby delineating cell-specific gene regulatory network (GRN). Network activity refers to the degree of consistency between GRN architectures and gene expression profiles in a specific condition or cellular microenvironment. Currently, numerous experimentally determined molecular interactions, including regulatory relationships closely related to SARS-CoV-2 infection, are documented in knowledge-bases. However, GRN activity is closely related to the cell dynamic environment and the heterogeneity of cell clusters. Therefore, to evaluate the consistency of GRN with gene expression profiles, we propose a single-cell Network Activity Evaluation framework, called scNAE. First, scNAE performs ODE modeling of time-course gene expression data. Then, the loss function with regularization penalty terms is constructed for formulating GRN inference rules from transcriptomic data. Furthermore, we have devised a rapid-convergence alternating direction method of multipliers to solve the regularized and constrained programs. Finally, an empirical P-value is derived based on a permutation statistical testing procedure to quantify the likelihood significance of the network matching with the data. The efficiency and advantage of scNAE have also been demonstrated by extensive numerical experiments, which can clearly depict the dynamic responses underlying GRN architectures triggered by the infection of SARS-CoV-2 in cells. The code and data of scNAE are available at https://github.com/zpliulab/scNAE. Chuan-Yuan Wang, Zhi-Ping Liu |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Multi-objective Optimization-Based Approach for Detection of Breast Cancer Biomarkers
Chuan-Yuan Wang, Duanchen Sun, Zhi-Ping Liu |
ICIC (3) | 4 |
| 2023 | MOFNet: A Deep Learning Framework of Integrating Multi-omics Data for Breast Cancer Diagnosis
Chunxiao Zhang, Pengpai Li, Duanchen Sun, Zhi-Ping Liu |
ICIC (3) | 4 |
| 2023 | LogBTF: gene regulatory network inference using Boolean threshold network model from single-cell gene expression dataabstractMOTIVATION: From a systematic perspective, it is crucial to infer and analyze gene regulatory network (GRN) from high-throughput single-cell RNA sequencing data. However, most existing GRN inference methods mainly focus on the network topology, only few of them consider how to explicitly describe the updated logic rules of regulation in GRNs to obtain their dynamics. Moreover, some inference methods also fail to deal with the over-fitting problem caused by the noise in time series data. RESULTS: In this article, we propose a novel embedded Boolean threshold network method called LogBTF, which effectively infers GRN by integrating regularized logistic regression and Boolean threshold function. First, the continuous gene expression values are converted into Boolean values and the elastic net regression model is adopted to fit the binarized time series data. Then, the estimated regression coefficients are applied to represent the unknown Boolean threshold function of the candidate Boolean threshold network as the dynamical equations. To overcome the multi-collinearity and over-fitting problems, a new and effective approach is designed to optimize the network topology by adding a perturbation design matrix to the input data and thereafter setting sufficiently small elements of the output coefficient vector to zeros. In addition, the cross-validation procedure is implemented into the Boolean threshold network model framework to strengthen the inference capability. Finally, extensive experiments on one simulated Boolean value dataset, dozens of simulation datasets, and three real single-cell RNA sequencing datasets demonstrate that the LogBTF method can infer GRNs from time series data more accurately than some other alternative methods for GRN inference. AVAILABILITY AND IMPLEMENTATION: The source data and code are available at https://github.com/zpliulab/LogBTF. Liangjie Sun, Chi-Wing Wong, Wai-Ki Ching, Zhi-Ping Liu |
Bioinform. | 6 |
| 2023 | ActivePPI: quantifying protein-protein interaction network activity with Markov random fieldsabstractMOTIVATION: Protein-protein interactions (PPI) are crucial components of the biomolecular networks that enable cells to function. Biological experiments have identified a large number of PPI, and these interactions are stored in knowledge bases. However, these interactions are often restricted to specific cellular environments and conditions. Network activity can be characterized as the extent of agreement between a PPI network (PPIN) and a distinct cellular environment measured by protein mass spectrometry, and it can also be quantified as a statistical significance score. Without knowing the activity of these PPI in the cellular environments or specific phenotypes, it is impossible to reveal how these PPI perform and affect cellular functioning. RESULTS: To calculate the activity of PPIN in different cellular conditions, we proposed a PPIN activity evaluation framework named ActivePPI to measure the consistency between network architecture and protein measurement data. ActivePPI estimates the probability density of protein mass spectrometry abundance and models PPIN using a Markov-random-field-based method. Furthermore, empirical P-value is derived based on a nonparametric permutation test to quantify the likelihood significance of the match between PPIN structure and protein abundance data. Extensive numerical experiments demonstrate the superior performance of ActivePPI and result in network activity evaluation, pathway activity assessment, and optimal network architecture tuning tasks. To summarize it succinctly, ActivePPI is a versatile tool for evaluating PPI network that can uncover the functional significance of protein interactions in crucial cellular biological processes and offer further insights into physiological phenomena. AVAILABILITY AND IMPLEMENTATION: All source code and data are freely available at https://github.com/zpliulab/ActivePPI. Chuan-Yuan Wang, Duanchen Sun, Zhi-Ping Liu |
Bioinform. | 4 |
| 2023 | Biomarker discovery from high-throughput data by connected network-constrained support vector machine
Zhi-Ping Liu |
Expert Syst. Appl. | 2 |
| 2022 | A Comparison Study of Predicting lncRNA-Protein Interactions via Representative Network Embedding Methods
Pengpai Li, Zhi-Ping Liu |
ICIC (2) | 3 |
| 2022 | A connected network-regularized logistic regression model for feature selection
Zhi-Ping Liu |
Appl. Intell. | 2 |
| 2022 | Graph attention network for link prediction of gene regulations from single-cell RNA-sequencing dataabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) data provides unprecedented opportunities to reconstruct gene regulatory networks (GRNs) at fine-grained resolution. Numerous unsupervised or self-supervised models have been proposed to infer GRN from bulk RNA-seq data, but few of them are appropriate for scRNA-seq data under the circumstance of low signal-to-noise ratio and dropout. Fortunately, the surging of TF-DNA binding data (e.g. ChIP-seq) makes supervised GRN inference possible. We regard supervised GRN inference as a graph-based link prediction problem that expects to learn gene low-dimensional vectorized representations to predict potential regulatory interactions. RESULTS: In this paper, we present GENELink to infer latent interactions between transcription factors (TFs) and target genes in GRN using graph attention network. GENELink projects the single-cell gene expression with observed TF-gene pairs to a low-dimensional space. Then, the specific gene representations are learned to serve for downstream similarity measurement or causal inference of pairwise genes by optimizing the embedding space. Compared to eight existing GRN reconstruction methods, GENELink achieves comparable or better performance on seven scRNA-seq datasets with four types of ground-truth networks. We further apply GENELink on scRNA-seq of human breast cancer metastasis and reveal regulatory heterogeneity of Notch and Wnt signalling pathways between primary tumour and lung metastasis. Moreover, the ontology enrichment results of unique lung metastasis GRN indicate that mitochondrial oxidative phosphorylation (OXPHOS) is functionally important during the seeding step of the cancer metastatic cascade, which is validated by pharmacological assays. AVAILABILITY AND IMPLEMENTATION: The code and data are available at https://github.com/zpliulab/GENELink. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhi-Ping Liu |
Bioinform. | 2 |
| 2022 | PST-PRNA: prediction of RNA-binding sites using protein surface topography and deep learningabstractMOTIVATION: Protein-RNA interactions play essential roles in many biological processes, including pre-mRNA processing, post-transcriptional gene regulation and RNA degradation. Accurate identification of binding sites on RNA-binding proteins (RBPs) is important for functional annotation and site-directed mutagenesis. Experimental assays to sparse RBPs are precise and convincing but also costly and time consuming. Therefore, flexible and reliable computational methods are required to recognize RNA-binding residues. RESULTS: In this work, we propose PST-PRNA, a novel model for predicting RNA-binding sites (PRNA) based on protein surface topography (PST). Taking full advantage of the 3D structural information of protein, PST-PRNA creates representative topography images of the entire protein surface by mapping it onto a unit spherical surface. Four kinds of descriptors are encoded to represent residues on the surface. Then, the potential features are integrated and optimized by using deep learning models. We compile a comprehensive non-redundant RBP dataset to train and test PST-PRNA using 10-fold cross-validation. Numerous experiments demonstrate PST-PRNA learns successfully the latent structural information of protein surface. On the non-redundant dataset with sequence identity of 0.3, PST-PRNA achieves area under the receiver operating characteristic curves (AUC) value of 0.860 and Matthew's correlation coefficient value of 0.420. Furthermore, we construct a completely independent test dataset for justification and comparison. PST-PRNA achieves AUC value of 0.913 on the independent dataset, which is superior to the other state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: The code and data are available at https://www.github.com/zpliulab/PST-PRNA. A web server is freely available at http://www.zpliulab.cn/PSTPRNA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Pengpai Li, Zhi-Ping Liu |
Bioinform. | 2 |
| 2022 | Identifying biomarkers for breast cancer by gene regulatory network rewiringabstractBACKGROUND: Mining gene regulatory network (GRN) is an important avenue for addressing cancer mechanism. Mutations in cancer genome perturb GRN and cause a rewiring in an orchestrated network. Hence, the exploration of gene regulatory network rewiring is significant to discover potential biomarkers and indicators for discriminating cancer phenotypes. RESULTS: Here, we propose a new bioinformatics method of identifying biomarkers based on network rewiring in different states. It firstly reconstructs GRN in different phenotypic conditions from gene expression data with a priori background network. We employ the algorithm based on path consistency algorithm and conditional mutual information to delete false-positive regulatory interactions between independent nodes/genes or not closely related gene pairs. And then a differential gene regulatory network (D-GRN) is constructed from the rewiring parts in the two phenotype-specific GRNs. Community detection technique is then applied for D-GRN to detect functional modules. Finally, we apply logistic regression classifier with recursive feature elimination to select biomarker genes in each module individually. The extracted feature genes result in a gene set of biomarkers with impressing ability to distinguish normal samples from controls. We verify the identified biomarkers in external independent validation datasets. For a proof-of-concept study, we apply the framework to identify diagnostic biomarkers of breast cancer. The identified biomarkers obtain a maximum AUC of 0.985 in the internal sample classification experiments. And these biomarkers achieve a maximum AUC of 0.989 in the external validations. CONCLUSION: In conclusion, network rewiring reveals significant differences between different phenotypes, which indicating cancer dysfunctional mechanisms. With the development of sequencing technology, the amount and quality of gene expression data become available. Condition-specific gene regulatory networks that are close to the real regulations in different states will be established. Revealing the network rewiring will greatly benefit the discovery of biomarkers or signatures for phenotypes. D-GRN is a general method to meet this demand of deciphering the high-throughput data for biomarker discovery. It is also easy to be extended for identifying biomarkers of other complex diseases beyond breast cancer. Yijuan Wang, Zhi-Ping Liu |
BMC Bioinform. | 2 |
| 2022 | Evaluating Gene Regulatory Network Activity From Dynamic Expression Data by Regularized Constraint ProgrammingabstractBy extracting molecular interactions identified by experiments, gene regulatory networks or gene circuits have documented in a large number of knowledge-based repositories. They provide systematic information and guidance of the functional connections between regulators, e.g., transcription factor proteins and miRNAs, and target genes. Network activity is defined as the degree of consistency between a regulatory network architecture and a specific cellular context of gene expression and can also be measured as a score of statistical significance. The gene network activities are closely related to the dynamics of cell states. To evaluate the activity of regulatory events in the form of network, we propose a network activity evaluation (NAE) framework by measuring the consistency between network architecture and gene expression data across specific states based on mathematical programming. NAE firstly employs the dynamic Bayesian network model to formulate the network structure with time series profiling data. For the constraints of prior knowledge about gene regulatory network, NAE introduces an interpretable general loss function with regularization penalties to calculate the degree of consistency between gene network and gene expression data. Moreover, we design a fast and convergent alternating direction method of multipliers algorithm to optimize the regularized constraint programming. The efficiency and advantage of the NAE framework is deduced through numerous experiments and comparison studies. It reflects the possibility and potential of the match between network and data, thereby helping to reveal the network activity and to explain the dynamic responds underlying the network structure caused by changes in molecular environment of living cells. The code of NAE is freely available for academic use (https://github.com/zpliulab/NAE). Chuan-Yuan Wang, Zhi-Ping Liu |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | A Semi-Supervised Learning Algorithm for Predicting MiRNA-Disease AssociationabstractThe dysregulation of miRNAs can lead to disease. Modeling miRNA-disease association is important in understanding the pathogenesis of diseases. The generation of multiple types of miRNA-disease association data provides us an opportunity to in-depth study of the interaction mechanism between miRNA and disease. However, the existing methods usually only focus on miRNA-disease interaction, which ignores the relationship types of their association. In this paper, we studied the feasibility of tensor robust principal component analysis (TRPCA) to model miRNA-disease-type associations. Unlike the binary association of, we use the triple association ofto better describe the pathogenesis of the disease. Multi-view auxiliary information (semantics of diseases, functions of miRNAs and Gaussian interaction profile kernel) are applied to fully consider the complexity of biological processes. The proposed semi-supervised learning model combines TRPCA and label propagation. The 10-fold cross validation (CV) and case study experiments on HMDD v2.0 dataset demonstrate that the proposed model is superior to the other benchmark methods. Na Yu 0004, Zhi-Ping Liu, Rui Gao 0006 |
BIBM | 2 |
| 2021 | Inference of Gene Regulatory Network from Time Series Expression Data by Combining Local Geometric Similarity and Multivariate Regression
Zhi-Ping Liu |
ICIC (3) | 2 |
| 2021 | Prioritizing Type 2 Diabetes Genes by Weighted PageRank on Bilayer Heterogeneous NetworksabstractThe prevalence of diabetes mellitus has been increasing rapidly in recent years. Type 2 diabetes makes up about 90 percent cases of diabetes. The interacting mixed effects of genetics and environments build possible interpretable pathogenesis. Thus, finding the causal disease genes is crucial in its clinical diagnosis and medical treatment. Currently, network-based computational method becomes a powerful tool of systematically analyzing complex diseases, such as the identification of candidate disease genes from networks. In this paper, we propose a bioinformatics framework of prioritizing type 2 diabetes genes by leveraging the modified PageRank algorithm on bilayer biomolecular networks consisting an ensemble gene-gene regulatory network and an integrative protein-protein interaction network. We specifically weigh the networks by differential mutual information for measuring the context specificities between genes and between proteins by transcriptomic and proteomic datasets, respectively. After formulating the network into two components of known disease genes and the other normal healthy genes, we rank the diabetes genes and others by bringing the orders in the bilayer network via an improved PageRank algorithm. We conclude that these known disease genes achieve significantly higher ranks compared to these randomly-selected normal genes, and the ranks are robust and consistent in multiple validation scenarios. In functional analysis, these high-ranked genes are identified to perform relevant risks and dysfunctions of type 2 diabetes. Haixia Shang, Zhi-Ping Liu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2019 | Identifying Cancer Biomarkers from High-Throughput RNA Sequencing Data by Machine Learning
Zishuang Zhang, Zhi-Ping Liu |
ICIC (2) | 2 |
| 2018 | Detecting pathway biomarkers of diabetic progression with differential entropy
Zhi-Ping Liu |
J. Biomed. Informatics | 1 |
| 2017 | Structure alignment-based classification of RNA-binding pockets reveals regional RNA recognition motifs on protein surfacesabstractBACKGROUND: Many critical biological processes are strongly related to protein-RNA interactions. Revealing the protein structure motifs for RNA-binding will provide valuable information for deciphering protein-RNA recognition mechanisms and benefit complementary structural design in bioengineering. RNA-binding events often take place at pockets on protein surfaces. The structural classification of local binding pockets determines the major patterns of RNA recognition. RESULTS: In this work, we provide a novel framework for systematically identifying the structure motifs of protein-RNA binding sites in the form of pockets on regional protein surfaces via a structure alignment-based method. We first construct a similarity network of RNA-binding pockets based on a non-sequential-order structure alignment method for local structure alignment. By using network community decomposition, the RNA-binding pockets on protein surfaces are clustered into groups with structural similarity. With a multiple structure alignment strategy, the consensus RNA-binding pockets in each group are identified. The crucial recognition patterns, as well as the protein-RNA binding motifs, are then identified and analyzed. CONCLUSIONS: Large-scale RNA-binding pockets on protein surfaces are grouped by measuring their structural similarities. This similarity network-based framework provides a convenient method for modeling the structural relationships of functional pockets. The local structural patterns identified serve as structure motifs for the recognition with RNA on protein surfaces. Zhi-Ping Liu, Shutang Liu, Ruitang Chen, Xiaopeng Huang, Ling-Yun Wu |
BMC Bioinform. | 1 |
| 2017 | Multi-dimensional data representation using linear tensor codingabstractLinear coding is widely used to concisely represent data sets by discovering basis functions of capturing high‐level features. However, the efficient identification of linear codes for representing multi‐dimensional data remains very challenging. In this study, the authors address the problem by proposing a linear tensor coding algorithm to represent multi‐dimensional data succinctly via a linear combination of tensor‐formed bases without data expansion. Motivated by the amalgamation of linear image coding and multi‐linear algebra, each basis function in the authors’ algorithm captures some specific variabilities. The basis‐associated coefficients can be used for data representation, compression and classification. When the authors apply the algorithm on both simulated phantom data and real facial data, the experimental results demonstrate their algorithm not only preserves the original information of input data, but also produces localised bases with concrete physical meanings. Xu Qiao, Yen-Wei Chen 0001, Zhi-Ping Liu |
IET Image Process. | 4 |
| 2016 | CMIP: a software package capable of reconstructing genome-wide regulatory networks using gene expression dataabstractBACKGROUND: A gene regulatory network (GRN) represents interactions of genes inside a cell or tissue, in which vertexes and edges stand for genes and their regulatory interactions respectively. Reconstruction of gene regulatory networks, in particular, genome-scale networks, is essential for comparative exploration of different species and mechanistic investigation of biological processes. Currently, most of network inference methods are computationally intensive, which are usually effective for small-scale tasks (e.g., networks with a few hundred genes), but are difficult to construct GRNs at genome-scale. RESULTS: Here, we present a software package for gene regulatory network reconstruction at a genomic level, in which gene interaction is measured by the conditional mutual information measurement using a parallel computing framework (so the package is named CMIP). The package is a greatly improved implementation of our previous PCA-CMI algorithm. In CMIP, we provide not only an automatic threshold determination method but also an effective parallel computing framework for network inference. Performance tests on benchmark datasets show that the accuracy of CMIP is comparable to most current network inference methods. Moreover, running tests on synthetic datasets demonstrate that CMIP can handle large datasets especially genome-wide datasets within an acceptable time period. In addition, successful application on a real genomic dataset confirms its practical applicability of the package. CONCLUSIONS: This new software package provides a powerful tool for genomic network reconstruction to biological community. The software can be accessed at http://www.picb.ac.cn/CMIP/ . Guangyong Zheng, Yaochen Xu, Zhi-Ping Liu, Luonan Chen, Xin-Guang Zhu |
BMC Bioinform. | 4 |
| 2016 | Prediction of protein-RNA interactions using sequence and structure descriptors
Zhi-Ping Liu, Hongyu Miao 0001 |
Neurocomputing | 1 |
| 2015 | Identifying module biomarker in type 2 diabetes mellitus by discriminative area of functional activityabstractBACKGROUND: Identifying diagnosis and prognosis biomarkers from expression profiling data is of great significance for achieving personalized medicine and designing therapeutic strategy in complex diseases. However, the reproducibility of identified biomarkers across tissues and experiments is still a challenge for this issue. RESULTS: We propose a strategy based on discriminative area of module activities to identify gene biomarkers which interconnect as a subnetwork or module by integrating gene expression data and protein-protein interactions. Then, we implement the procedure in T2DM as a case study and identify a module biomarker with 32 genes from mRNA expression data in skeletal muscle for T2DM. This module biomarker is enriched with known causal genes and related functions of T2DM. Further analysis shows that the module biomarker is of superior performance in classification, and has consistently high accuracies across tissues and experiments. CONCLUSION: The proposed approach can efficiently identify robust and functionally meaningful module biomarkers in T2DM, and could be employed in biomarker discovery of other complex diseases characterized by expression profiles. Lin Gao 0006, Zhi-Ping Liu, Luonan Chen |
BMC Bioinform. | 3 |
| 2014 | Systematic identification of transcriptional and post-transcriptional regulations in human respiratory epithelial cells during influenza A virus infectionabstractBACKGROUND: Respiratory epithelial cells are the primary target of influenza virus infection in human. However, the molecular mechanisms of airway epithelial cell responses to viral infection are not fully understood. Revealing genome-wide transcriptional and post-transcriptional regulatory relationships can further advance our understanding of this problem, which motivates the development of novel and more efficient computational methods to simultaneously infer the transcriptional and post-transcriptional regulatory networks. RESULTS: Here we propose a novel framework named SITPR to investigate the interactions among transcription factors (TFs), microRNAs (miRNAs) and target genes. Briefly, a background regulatory network on a genome-wide scale (~23,000 nodes and ~370,000 potential interactions) is constructed from curated knowledge and algorithm predictions, to which the identification of transcriptional and post-transcriptional regulatory relationships is anchored. To reduce the dimension of the associated computing problem down to an affordable size, several topological and data-based approaches are used. Furthermore, we propose the constrained LASSO formulation and combine it with the dynamic Bayesian network (DBN) model to identify the activated regulatory relationships from time-course expression data. Our simulation studies on networks of different sizes suggest that the proposed framework can effectively determine the genuine regulations among TFs, miRNAs and target genes; also, we compare SITPR with several selected state-of-the-art algorithms to further evaluate its performance. By applying the SITPR framework to mRNA and miRNA expression data generated from human lung epithelial A549 cells in response to A/Mexico/InDRE4487/2009 (H1N1) virus infection, we are able to detect the activated transcriptional and post-transcriptional regulatory relationships as well as the significant regulatory motifs. CONCLUSION: Compared with other representative state-of-the-art algorithms, the proposed SITPR framework can more effectively identify the activated transcriptional and post-transcriptional regulations simultaneously from a given background network. The idea of SITPR is generally applicable to the analysis of gene regulatory networks in human cells. The results obtained for human respiratory epithelial cells suggest the importance of the transcriptional, post-transcriptional regulations as well as their synergies in the innate immune responses against IAV infection. Zhi-Ping Liu, Hulin Wu, Hongyu Miao 0001 |
BMC Bioinform. | 1 |
| 2013 | NARROMI: a noise and redundancy reduction technique improves accuracy of gene regulatory network inferenceabstractMOTIVATION: Reconstruction of gene regulatory networks (GRNs) is of utmost interest to biologists and is vital for understanding the complex regulatory mechanisms within the cell. Despite various methods developed for reconstruction of GRNs from gene expression profiles, they are notorious for high false positive rate owing to the noise inherited in the data, especially for the dataset with a large number of genes but a small number of samples. RESULTS: In this work, we present a novel method, namely NARROMI, to improve the accuracy of GRN inference by combining ordinary differential equation-based recursive optimization (RO) and information theory-based mutual information (MI). In the proposed algorithm, the noisy regulations with low pairwise correlations are first removed by using MI, and the redundant regulations from indirect regulators are further excluded by RO to improve the accuracy of inferred GRNs. In particular, the RO step can help to determine regulatory directions without prior knowledge of regulators. The results on benchmark datasets from Dialogue for Reverse Engineering Assessments and Methods challenge and experimentally determined GRN of Escherichia coli show that NARROMI significantly outperforms other popular methods in terms of false positive rates and accuracy. AVAILABILITY: All the source data and code are available at: http://csb.shu.edu.cn/narromi.htm. Keqin Liu, Zhi-Ping Liu, Béatrice Duval, Jean-Michel Richer, Xing-Ming Zhao, Jin-Kao Hao, Luonan Chen |
Bioinform. | 3 |
| 2013 | Research and applications: An integrated approach to identify causal network modules of complex diseases with application to colorectal cancerabstractBACKGROUND: Many methods have been developed to identify disease genes and further module biomarkers of complex diseases based on gene expression data. It is generally difficult to distinguish whether the variations in gene expression are causative or merely the effect of a disease. The limitation of relying on gene expression data alone highlights the need to develop new approaches that can explore various data to reflect the casual relationship between network modules and disease traits. METHODS: In this work, we developed a novel network-based approach to identify putative causal module biomarkers of complex diseases by integrating heterogeneous information, for example, epigenomic data, gene expression data, and protein-protein interaction network. We first formulated the identification of modules as a mathematical programming problem, which can be solved efficiently and effectively in an accurate manner. Then, we applied our approach to colorectal cancer (CRC) and identified several network modules that can serve as potential module biomarkers for characterizing CRC. Further validations using three additional gene expression datasets verified their candidate biomarker properties and the effectiveness of the method. Functional enrichment analysis also revealed that the identified modules are strongly related to hallmarks of cancer, and the enriched functions, such as inflammatory response, receptor and signaling pathways, are specific to CRC. RESULTS: Through constructing a transcription factor (TF)-module network, we found that aberrant DNA methylation of genes encoding TF considerably contributes to the activity change of some genes, which may function as causal genes of CRC, and that can also be exploited to develop efficient therapies or effective drugs. CONCLUSION: Our method can potentially be extended to the study of other complex diseases and the multiclassification problem. Zhenshu Wen, Zhi-Ping Liu, Zhengrong Liu, Luonan Chen |
J. Am. Medical Informatics Assoc. | 2 |
| 2012 | Inferring gene regulatory networks from gene expression data by path consistency algorithm based on conditional mutual informationabstractMOTIVATION: Reconstruction of gene regulatory networks (GRNs), which explicitly represent the causality of developmental or regulatory process, is of utmost interest and has become a challenging computational problem for understanding the complex regulatory mechanisms in cellular systems. However, all existing methods of inferring GRNs from gene expression profiles have their strengths and weaknesses. In particular, many properties of GRNs, such as topology sparseness and non-linear dependence, are generally in regulation mechanism but seldom are taken into account simultaneously in one computational method. RESULTS: In this work, we present a novel method for inferring GRNs from gene expression data considering the non-linear dependence and topological structure of GRNs by employing path consistency algorithm (PCA) based on conditional mutual information (CMI). In this algorithm, the conditional dependence between a pair of genes is represented by the CMI between them. With the general hypothesis of Gaussian distribution underlying gene expression data, CMI between a pair of genes is computed by a concise formula involving the covariance matrices of the related gene expression profiles. The method is validated on the benchmark GRNs from the DREAM challenge and the widely used SOS DNA repair network in Escherichia coli. The cross-validation results confirmed the effectiveness of our method (PCA-CMI), which outperforms significantly other previous methods. Besides its high accuracy, our method is able to distinguish direct (or causal) interactions from indirect associations. AVAILABILITY: All the source data and code are available at: http://csb.shu.edu.cn/subweb/grn.htm. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xing-Ming Zhao, Kun He 0007, Le Lu 0001, Yongwei Cao, Jingdong Liu, Jin-Kao Hao, Zhi-Ping Liu, Luonan Chen |
Bioinform. | 8 |
| 2012 | Identifying dysregulated pathways in cancers from pathway interaction networksabstractBACKGROUND: Cancers, a group of multifactorial complex diseases, are generally caused by mutation of multiple genes or dysregulation of pathways. Identifying biomarkers that can characterize cancers would help to understand and diagnose cancers. Traditional computational methods that detect genes differentially expressed between cancer and normal samples fail to work due to small sample size and independent assumption among genes. On the other hand, genes work in concert to perform their functions. Therefore, it is expected that dysregulated pathways will serve as better biomarkers compared with single genes. RESULTS: In this paper, we propose a novel approach to identify dysregulated pathways in cancer based on a pathway interaction network. Our contribution is three-fold. Firstly, we present a new method to construct pathway interaction network based on gene expression, protein-protein interactions and cellular pathways. Secondly, the identification of dysregulated pathways in cancer is treated as a feature selection problem, which is biologically reasonable and easy to interpret. Thirdly, the dysregulated pathways are identified as subnetworks from the pathway interaction networks, where the subnetworks characterize very well the functional dependency or crosstalk between pathways. The benchmarking results on several distinct cancer datasets demonstrate that our method can obtain more reliable and accurate results compared with existing state of the art methods. Further functional analysis and independent literature evidence also confirm that our identified potential pathogenic pathways are biologically reasonable, indicating the effectiveness of our method. CONCLUSIONS: Dysregulated pathways can serve as better biomarkers compared with single genes. In this work, by utilizing pathway interaction networks and gene expression data, we propose a novel approach that effectively identifies dysregulated pathways, which can not only be used as biomarkers to diagnose cancers but also serve as potential drug targets in the future. Keqin Liu, Zhi-Ping Liu, Jin-Kao Hao, Luonan Chen, Xing-Ming Zhao |
BMC Bioinform. | 2 |
| 2012 | Inferring a protein interaction map of Mycobacterium tuberculosis based on sequences and interologsabstractBACKGROUND: Mycobacterium tuberculosis is an infectious bacterium posing serious threats to human health. Due to the difficulty in performing molecular biology experiments to detect protein interactions, reconstruction of a protein interaction map of M. tuberculosis by computational methods will provide crucial information to understand the biological processes in the pathogenic microorganism, as well as provide the framework upon which new therapeutic approaches can be developed. RESULTS: In this paper, we constructed an integrated M. tuberculosis protein interaction network by machine learning and ortholog-based methods. Firstly, we built a support vector machine (SVM) method to infer the protein interactions of M. tuberculosis H37Rv by gene sequence information. We tested our predictors in Escherichia coli and mapped the genetic codon features underlying its protein interactions to M. tuberculosis. Moreover, the documented interactions of 14 other species were mapped to the interactome of M. tuberculosis by the interolog method. The ensemble protein interactions were validated by various functional relationships, i.e., gene coexpression, evolutionary relationship and functional similarity, extracted from heterogeneous data sources. The accuracy and validation demonstrate the effectiveness and efficiency of our framework. CONCLUSIONS: A protein interaction map of M. tuberculosis is inferred from genetic codons and interologs. The prediction accuracy and numerically experimental validation demonstrate the effectiveness and efficiency of our method. Furthermore, our methods can be straightforwardly extended to infer the protein interactions of other bacterial species. Zhi-Ping Liu, Yu-Qing Qiu, Ross K. K. Leung, Xiang-Sun Zhang, Stephen Kwok-Wing Tsui, Luonan Chen |
BMC Bioinform. | 1 |
| 2012 | Multiple-resource and multiple-depot emergency response problem considering secondary disasters
Jianghua Zhang, Zhi-Ping Liu |
Expert Syst. Appl. | 3 |
| 2012 | Identifying disease genes and module biomarkers by differential interactionsabstractOBJECTIVE: A complex disease is generally caused by the mutation of multiple genes or by the dysfunction of multiple biological processes. Systematic identification of causal disease genes and module biomarkers can provide insights into the mechanisms underlying complex diseases, and help develop efficient therapies or effective drugs. MATERIALS AND METHODS: In this paper, we present a novel approach to predict disease genes and identify dysfunctional networks or modules, based on the analysis of differential interactions between disease and control samples, in contrast to the analysis of differential gene or protein expressions widely adopted in existing methods. RESULTS AND DISCUSSION: As an example, we applied our method to the study of three-stage microarray data for gastric cancer. We identified network modules or module biomarkers that include a set of genes related to gastric cancer, implying the predictive power of our method. The results on holdout validation data sets show that our identified module can serve as an effective module biomarker for accurately detecting or diagnosing gastric cancer, thereby validating the efficiency of our method. CONCLUSION: We proposed a new approach to detect module biomarkers for diseases, and the results on gastric cancer demonstrated that the differential interactions are useful to detect dysfunctional modules in the molecular interaction network, which in turn can be used as robust module biomarkers. Xiaoping Liu 0002, Zhi-Ping Liu, Xing-Ming Zhao, Luonan Chen |
J. Am. Medical Informatics Assoc. | 2 |
| 2011 | Inferring Protein-Protein Interactions Based on Sequences and Interologs in Mycobacterium Tuberculosis
Zhi-Ping Liu, Yu-Qing Qiu, Ross K. K. Leung, Xiang-Sun Zhang, Stephen Kwok-Wing Tsui, Luonan Chen |
ICIC (3) | 1 |
| 2010 | Prediction of protein-RNA binding sites by a random forest method with combined featuresabstractMOTIVATION: Protein-RNA interactions play a key role in a number of biological processes, such as protein synthesis, mRNA processing, mRNA assembly, ribosome function and eukaryotic spliceosomes. As a result, a reliable identification of RNA binding site of a protein is important for functional annotation and site-directed mutagenesis. Accumulated data of experimental protein-RNA interactions reveal that a RNA binding residue with different neighbor amino acids often exhibits different preferences for its RNA partners, which in turn can be assessed by the interacting interdependence of the amino acid fragment and RNA nucleotide. RESULTS: In this work, we propose a novel classification method to identify the RNA binding sites in proteins by combining a new interacting feature (interaction propensity) with other sequence- and structure-based features. Specifically, the interaction propensity represents a binding specificity of a protein residue to the interacting RNA nucleotide by considering its two-side neighborhood in a protein residue triplet. The sequence as well as the structure-based features of the residues are combined together to discriminate the interaction propensity of amino acids with RNA. We predict RNA interacting residues in proteins by implementing a well-built random forest classifier. The experiments show that our method is able to detect the annotated protein-RNA interaction sites in a high accuracy. Our method achieves an accuracy of 84.5%, F-measure of 0.85 and AUC of 0.92 prediction of the RNA binding residues for a dataset containing 205 non-homologous RNA binding proteins, and also outperforms several existing RNA binding residue predictors, such as RNABindR, BindN, RNAProB and PPRint, and some alternative machine learning methods, such as support vector machine, naive Bayes and neural network in the comparison study. Furthermore, we provide some biological insights into the roles of sequences and structures in protein-RNA interactions by both evaluating the importance of features for their contributions in predictive accuracy and analyzing the binding patterns of interacting residues. AVAILABILITY: All the source data and code are available at http://www.aporc.org/doc/wiki/PRNA or http://www.sysbio.ac.cn/datatools.asp CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhi-Ping Liu, Ling-Yun Wu, Yong Wang 0001, Xiang-Sun Zhang, Luonan Chen |
Bioinform. | 1 |
| 2009 | Dynamically dysfunctional protein interactions in the development of Alzheimer's diseaseabstractAlzheimer's disease usually causes dementia in the old people and the symptom progression of the disease phenotype displays certain patterns. One possible reason is that the nerve cells in the brains of the patients degenerate at different stages. Here, we analyze the dynamics of disease progression based on its biomolecular network. We develop a novel computational method to integrate an ensemble protein network and the hippocampal gene expression data. Specifically, we construct the induced dynamical pathways which present particular characteristics at different disease stages from the control to disease samples. Based on the network-based method, we reveal that the active pathways tend to be more complicated during the development of disease. Also we find that the disease proteins performing important functions are always located in the cooperations of the identified pathways. These results also demonstrate that the network-based analysis can provide knowledge and evidences on the dynamics and pathological pathways of the complex Alzheimer's disease. Zhi-Ping Liu, Yong Wang 0001, Tieqiao Wen, Xiang-Sun Zhang, Weiming Xia, Luonan Chen |
SMC | 1 |
| 2007 | Predicting gene ontology functions from protein's regional surface structuresabstractBACKGROUND: Annotation of protein functions is an important task in the post-genomic era. Most early approaches for this task exploit only the sequence or global structure information. However, protein surfaces are believed to be crucial to protein functions because they are the main interfaces to facilitate biological interactions. Recently, several databases related to structural surfaces, such as pockets and cavities, have been constructed with a comprehensive library of identified surface structures. For example, CASTp provides identification and measurements of surface accessible pockets as well as interior inaccessible cavities. RESULTS: A novel method was proposed to predict the Gene Ontology (GO) functions of proteins from the pocket similarity network, which is constructed according to the structure similarities of pockets. The statistics of the networks were presented to explore the relationship between the similar pockets and GO functions of proteins. Cross-validation experiments were conducted to evaluate the performance of the proposed method. Results and codes are available at: http://zhangroup.aporc.org/bioinfo/PSN/. CONCLUSION: The computational results demonstrate that the proposed method based on the pocket similarity network is effective and efficient for predicting GO functions of proteins in terms of both computational complexity and prediction accuracy. The proposed method revealed strong relationship between small surface patterns (or pockets) and GO functions, which can be further used to identify active sites or functional motifs. The high quality performance of the prediction method together with the statistics also indicates that pockets play essential roles in biological interactions or the GO functions. Moreover, in addition to pockets, the proposed network framework can also be used for adopting other protein spatial surface patterns to predict the protein functions. Zhi-Ping Liu, Ling-Yun Wu, Yong Wang 0001, Luonan Chen, Xiang-Sun Zhang |
BMC Bioinform. | 1 |