Jie Zheng 0002

dblp:94/190-2 · DBLP profile ↗
← Back
46ranked-venue papers
6as first author
17since 2021 · last 2025
0000-0001-6774-9786ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 43 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Prompt-Driven Context-Aware Prediction of Cell Line-Specific Synthetic Lethality via Biomedical Large Language Models
abstract
Synthetic lethality (SL) refers to a genetic interaction in which the simultaneous dysfunction of two genes leads to cell death, while inactivation of either gene alone has no significant impact, thereby offering a promising avenue for targeted cancer therapy. However, the heterogeneity of SL relationships across cancer cell lines and the scarcity of contextspecific SL labels present challenges for accurately predicting SL relationships using computational methods. In this work, we propose PromptCASL, a prompt-driven, context-aware framework that leverages biomedical large language models (LLMs) for cell line-specific synthetic lethality prediction. To our knowledge, this is the first approach to leverage natural language prompts and LLMs for this task, beyond traditional structured data. Extensive experiments across various scenarios demonstrate that PromptCASL outperforms both pan-cancer and cell line-specific baseline models in prediction accuracy and generalizability. Furthermore, interpretability analysis reveals the model's ability to capture context-specific SL patterns, underscoring its potential to distinguish heterogeneous gene interactions across diverse biological contexts.
Yimiao Feng, Ruoyin Zhang, Yingfan Rui, Jie Zheng 0002
BIBM5
2025 Interpretable high-order knowledge graph neural network for predicting synthetic lethality in human cancers
abstract
Synthetic lethality (SL) is a promising gene interaction for cancer therapy. Recent SL prediction methods integrate knowledge graphs (KGs) into graph neural networks (GNNs) and employ attention mechanisms to extract local subgraphs as explanations for target gene pairs. However, attention mechanisms often lack fidelity, typically generate a single explanation per gene pair, and fail to ensure trustworthy high-order structures in their explanations. To overcome these limitations, we propose Diverse Graph Information Bottleneck for Synthetic Lethality (DGIB4SL), a KG-based GNN that generates multiple faithful explanations for the same gene pair and effectively encodes high-order structures. Specifically, we introduce a novel DGIB objective, integrating a determinant point process constraint into the standard information bottleneck objective, and employ 13 motif-based adjacency matrices to capture high-order structures in gene representations. Experimental results show that DGIB4SL outperforms state-of-the-art baselines and provides multiple explanations for SL prediction, revealing diverse biological mechanisms underlying SL inference.
Xuexin Chen, Ruichu Cai, Zhengting Huang, Zijian Li 0001, Jie Zheng 0002, Min Wu 0008
Briefings Bioinform.5
2025 Learning Universal Knowledge Graph Embedding for Predicting Biomedical Pairwise Interactions
abstract
Predicting biomedical interactions is crucial for understanding various biological processes and drug discovery. Graph neural networks (GNNs) are promising in identifying novel interactions when extensive labeled data are available. However, labeling biomedical interactions is often time-consuming and labor-intensive, resulting in low-data scenarios. Furthermore, distribution shifts between training and test data in real-world applications pose a challenge to the generalizability of GNN models. Recent studies suggest that pre-training GNN models with self-supervised learning on unlabeled data can enhance their performance in predicting biomedical interactions. Here, we propose LukePi, a novel self-supervised pre-training framework that pre-trains GNN models on biomedical knowledge graphs (BKGs). LukePi is trained with two self-supervised tasks: topology-based node degree classification and semantics-based edge recovery. The former is to predict the degree of a node from its topological context and the latter is to infer both type and existence of a candidate edge by learning semantic information in the BKG. By integrating the two complementary tasks, LukePi effectively captures the rich information from the BKG, thereby enhancing the quality of node representations. We evaluate the performance of LukePi on two critical link prediction tasks: predicting synthetic lethality and drug-target interactions, using four benchmark datasets. In both distribution-shift and low-data scenarios, LukePi significantly outperforms 22 baseline models, demonstrating the power of the graph pre-training strategy when labeled data are sparse.
Yang Yang 0161, Xin Liu 0027, Yimiao Feng, Jie Zheng 0002
IEEE Trans. Comput. Biol. Bioinform.5
2025 SLInterpreter: An Exploratory and Iterative Human-AI Collaborative System for GNN-Based Synthetic Lethal Prediction
abstract
Synthetic Lethal (SL) relationships, though rare among the vast array of gene combinations, hold substantial promise for targeted cancer therapy. Despite advancements in AI model accuracy, there is still a significant need among domain experts for interpretive paths and mechanism explorations that align better with domain-specific knowledge, particularly due to the high costs of experimentation. To address this gap, we propose an iterative Human-AI collaborative framework with two key components: 1) Human-Engaged Knowledge Graph Refinement based on Metapath Strategies, which leverages insights from interpretive paths and domain expertise to refine the knowledge graph through metapath strategies with appropriate granularity. 2) Cross-Granularity SL Interpretation Enhancement and Mechanism Analysis, which aids experts in organizing and comparing predictions and interpretive paths across different granularities, uncovering new SL relationships, enhancing result interpretation, and elucidating potential mechanisms inferred by Graph Neural Network (GNN) models. These components cyclically optimize model predictions and mechanism explorations, enhancing expert involvement and intervention to build trust. Facilitated by SLInterpreter, this framework ensures that newly generated interpretive paths increasingly align with domain knowledge and adhere more closely to real-world biological principles through iterative Human-AI collaboration. We evaluate the framework's efficacy through a case study and expert interviews.
Shaohan Shi, Jie Zheng 0002, Quan Li 0002
IEEE Trans. Vis. Comput. Graph.4
2024 Prompt-based Generation of Natural Language Explanations of Synthetic Lethality for Cancer Drug Discovery
abstract
Synthetic lethality (SL) offers a promising approach for targeted anti-cancer therapy. Deeply understanding SL gene pair mechanisms is vital for anti-cancer drug discovery. However, current wet-lab and machine learning-based SL prediction methods lack user-friendly and quantitatively evaluable explanations. To address these problems, we propose a prompt-based pipeline for generating natural language explanations. We first construct a natural language dataset named NexLeth. This dataset is derived from New Bing through prompt-based queries and expert annotations and contains 707 instances. NexLeth enhances the understanding of SL mechanisms and it is a benchmark for evaluating SL explanation methods. For the task of natural language generation for SL explanations, we combine subgraph explanations from an SL knowledge graph (KG) with instructions to construct novel personalized prompts, so as to inject the domain knowledge into the generation process. We then leverage the prompts to fine-tune pre-trained biomedical language models on our dataset. Experimental results show that the fine-tuned model equipped with designed prompts performs better than existing biomedical language models in terms of text quality and explainability, suggesting the potential of our dataset and the fine-tuned model for generating understandable and reliable explanations of SL mechanisms.
Yimiao Feng, Jie Zheng 0002
LREC/COLING3
2024 TEFDTA: a transformer encoder and fingerprint representation combined prediction method for bonded and non-bonded drug-target affinities
abstract
MOTIVATION: The prediction of binding affinity between drug and target is crucial in drug discovery. However, the accuracy of current methods still needs to be improved. On the other hand, most deep learning methods focus only on the prediction of non-covalent (non-bonded) binding molecular systems, but neglect the cases of covalent binding, which has gained increasing attention in the field of drug development. RESULTS: In this work, a new attention-based model, A Transformer Encoder and Fingerprint combined Prediction method for Drug-Target Affinity (TEFDTA) is proposed to predict the binding affinity for bonded and non-bonded drug-target interactions. To deal with such complicated problems, we used different representations for protein and drug molecules, respectively. In detail, an initial framework was built by training our model using the datasets of non-bonded protein-ligand interactions. For the widely used dataset Davis, an additional contribution of this study is that we provide a manually corrected Davis database. The model was subsequently fine-tuned on a smaller dataset of covalent interactions from the CovalentInDB database to optimize performance. The results demonstrate a significant improvement over existing approaches, with an average improvement of 7.6% in predicting non-covalent binding affinity and a remarkable average improvement of 62.9% in predicting covalent binding affinity compared to using BindingDB data alone. At the end, the potential ability of our model to identify activity cliffs was investigated through a case study. The prediction results indicate that our model is sensitive to discriminate the difference of binding affinities arising from small variances in the structures of compounds. AVAILABILITY AND IMPLEMENTATION: The codes and datasets of TEFDTA are available at https://github.com/lizongquan01/TEFDTA.
Zongquan Li, Pengxuan Ren, Hao Yang 0060, Jie Zheng 0002, Fang Bai
Bioinform.4
2024 SL-Miner: a web server for mining evidence and prioritization of cancer-specific synthetic lethality
abstract
SUMMARY: Synthetic lethality (SL) refers to a type of genetic interaction in which the simultaneous inactivation of two genes leads to cell death, while the inactivation of a single gene does not affect cell viability. It significantly expands the range of potential therapeutic targets for anti-cancer treatments. SL interactions are primarily identified through experimental screening and computational prediction. Although various computational methods have been proposed, they tend to ignore providing evidence to support their predictions of SL. Besides, they are rarely user-friendly for biologists who likely have limited programming skills. Moreover, the genetic context specificity of SL interactions is often not taken into consideration. Here, we introduce a web server called SL-Miner, which is designed to mine the evidence of SL relationships between a primary gene and a few candidate SL partner genes in a specific type of cancer, and to prioritize these candidate genes by integrating various types of evidence. For intuitive data visualization, SL-Miner provides a range of charts (e.g. volcano plot and box plot) to help users get insights from the data. AVAILABILITY AND IMPLEMENTATION: SL-Miner is available at https://slminer.sist.shanghaitech.edu.cn.
Xin Liu 0027, Jieni Hu, Jie Zheng 0002
Bioinform.3
2024 TMELand: An End-to-End Pipeline for Quantification and Visualization of Waddington's Epigenetic Landscape Based on Gene Regulatory Network
abstract
Waddington's epigenetic landscape is a framework depicting the processes of cell differentiation and reprogramming under the control of a gene regulatory network (GRN). Traditional model-driven methods for landscape quantification focus on the Boolean network or differential equation-based models of GRN, which need sophisticated prior knowledge and hence hamper their practical applications. To resolve this problem, we combine data-driven methods for inferring GRNs from gene expression data with model-driven approach to the landscape mapping. Specifically, we build an end-to-end pipeline to link data-driven and model-driven methods and develop a software tool named TMELand for GRN inference, visualizing Waddington's epigenetic landscape, and calculating state transition paths between attractors to uncover the intrinsic mechanism of cellular transition dynamics. By integrating GRN inference from real transcriptomic data with landscape modeling, TMELand can facilitate studies of computational systems biology, such as predicting cellular states and visualizing the dynamical trends of cell fate determination and transition dynamics from single-cell transcriptomic data.
Chunhe Li 0001, Jie Zheng 0002
IEEE ACM Trans. Comput. Biol. Bioinform.4
2024 GENNDTI: Drug-Target Interaction Prediction Using Graph Neural Network Enhanced by Router Nodes
abstract
Identifying drug-target interactions (DTI) is crucial in drug discovery and repurposing, and in silico techniques for DTI predictions are becoming increasingly important for reducing time and cost. Most interaction-based DTI models rely on the guilt-by-association principle that "similar drugs can interact with similar targets". However, such methods utilize precomputed similarity matrices and cannot dynamically discover intricate correlations. Meanwhile, some methods enrich DTI networks by incorporating additional networks like DDI and PPI networks, enriching biological signals to enhance DTI prediction. While these approaches have achieved promising performance in DTI prediction, such coarse-grained association data do not explain the specific biological mechanisms underlying DTIs. In this work, we propose GENNDTI, which constructs biologically meaningful routers to represent and integrate the salient properties of drugs and targets. Similar drugs or targets connect to more same router nodes, capturing property sharing. In addition, heterogeneous encoders are designed to distinguish different types of interactions, modeling both real and constructed interactions. This strategy enriches graph topology and enhances prediction efficiency as well. We evaluate the proposed method on benchmark datasets, demonstrating comparative performance over existing methods. We specifically analyze router nodes to validate their efficacy in improving predictions and providing biological explanations.
Beiyuan Yang, Yule Liu, Fang Bai, Mingyue Zheng, Jie Zheng 0002
IEEE J. Biomed. Health Informatics6
2023 KR4SL: knowledge graph reasoning for explainable prediction of synthetic lethality
abstract
MOTIVATION: Synthetic lethality (SL) is a promising strategy for anticancer therapy, as inhibiting SL partners of genes with cancer-specific mutations can selectively kill the cancer cells without harming the normal cells. Wet-lab techniques for SL screening have issues like high cost and off-target effects. Computational methods can help address these issues. Previous machine learning methods leverage known SL pairs, and the use of knowledge graphs (KGs) can significantly enhance the prediction performance. However, the subgraph structures of KG have not been fully explored. Besides, most machine learning methods lack interpretability, which is an obstacle for wide applications of machine learning to SL identification. RESULTS: We present a model named KR4SL to predict SL partners for a given primary gene. It captures the structural semantics of a KG by efficiently constructing and learning from relational digraphs in the KG. To encode the semantic information of the relational digraphs, we fuse textual semantics of entities into propagated messages and enhance the sequential semantics of paths using a recurrent neural network. Moreover, we design an attentive aggregator to identify critical subgraph structures that contribute the most to the SL prediction as explanations. Extensive experiments under different settings show that KR4SL significantly outperforms all the baselines. The explanatory subgraphs for the predicted gene pairs can unveil prediction process and mechanisms underlying synthetic lethality. The improved predictive power and interpretability indicate that deep learning is practically useful for SL-based cancer drug target discovery. AVAILABILITY AND IMPLEMENTATION: The source code is freely available at https://github.com/JieZheng-ShanghaiTech/KR4SL.
Min Wu 0008, Yong Liu 0020, Yimiao Feng, Jie Zheng 0002
Bioinform.5
2022 PiLSL: pairwise interaction learning-based graph neural network for synthetic lethality prediction in human cancers
abstract
MOTIVATION: Synthetic lethality (SL) is a type of genetic interaction in which the simultaneous inactivation of two genes leads to cell death, while the inactivation of a single gene does not affect the cell viability. It can effectively expand the range of anti-cancer therapeutic targets. SL interactions are identified mainly by experimental screening and computational prediction. Recent machine-learning methods mostly learn the representation of each gene individually, ignoring the representation of the pairwise interaction between two genes. In addition, the mechanisms of SL, the key to translating SL into cancer therapeutics, are often unclear. RESULTS: To fill the gaps, we propose a pairwise interaction learning-based graph neural network (GNN) named PiLSL to learn the representation of pairwise interaction between two genes for SL prediction. First, we construct an enclosing graph for each pair of genes from a knowledge graph. Secondly, we design an attentive embedding propagation layer in a GNN to discriminate the importance among the edges in the enclosing graph and to learn the latent features of the pairwise interaction from the weighted enclosing graph. Finally, we further fuse the latent features with explicit features extracted from multi-omics data to obtain powerful gene representations for SL prediction. Extensive experimental results demonstrate that PiLSL outperforms the best baseline by a large margin and generalizes well under three realistic scenarios. Besides, PiLSL provides an explanation of SL mechanisms via the weighted paths in the enclosing graphs by attention mechanism. AVAILABILITY AND IMPLEMENTATION: Our source code is available at https://github.com/JieZheng-ShanghaiTech/PiLSL.
Xin Liu 0027, Jiale Yu, Beiyuan Yang, Shike Wang, Fang Bai, Jie Zheng 0002
Bioinform.8
2022 NSF4SL: negative-sample-free contrastive learning for ranking synthetic lethal partner genes in human cancers
abstract
MOTIVATION: Detecting synthetic lethality (SL) is a promising strategy for identifying anti-cancer drug targets. Targeting SL partners of a primary gene mutated in cancer is selectively lethal to cancer cells. Due to high cost of wet-lab experiments and availability of gold standard SL data, supervised machine learning for SL prediction has been popular. However, most of the methods are based on binary classification and thus limited by the lack of reliable negative data. Contrastive learning can train models without any negative sample and is thus promising for finding novel SLs. RESULTS: We propose NSF4SL, a negative-sample-free SL prediction model based on a contrastive learning framework. It captures the characteristics of positive SL samples by using two branches of neural networks that interact with each other to learn SL-related gene representations. Moreover, a feature-wise data augmentation strategy is used to mitigate the sparsity of SL data. NSF4SL significantly outperforms all baselines which require negative samples, even in challenging experimental settings. To the best of our knowledge, this is the first time that SL prediction is formulated as a gene ranking problem, which is more practical than the current formulation as binary classification. NSF4SL is the first contrastive learning method for SL prediction and its success points to a new direction of machine-learning methods for identifying novel SLs. AVAILABILITY AND IMPLEMENTATION: Our source code is available at https://github.com/JieZheng-ShanghaiTech/NSF4SL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Shike Wang, Yimiao Feng, Xin Liu 0027, Yong Liu 0020, Min Wu 0008, Jie Zheng 0002
Bioinform.6
2022 Prediction of gene co-expression from chromatin contacts with graph attention network
abstract
MOTIVATION: The technology of high-throughput chromatin conformation capture (Hi-C) allows genome-wide measurement of chromatin interactions. Several studies have shown statistically significant relationships between gene-gene spatial contacts and their co-expression. It is desirable to uncover epigenetic mechanisms of transcriptional regulation behind such relationships using computational modeling. Existing methods for predicting gene co-expression from Hi-C data use manual feature engineering or unsupervised learning, which either limits the prediction accuracy or lacks interpretability. RESULTS: To address these issues, we propose HiCoEx (Hi-C predicts gene co-expression), a novel end-to-end framework for explainable prediction of gene co-expression from Hi-C data based on graph neural network. We apply graph attention mechanism to a gene contact network inferred from Hi-C data to distinguish the importance among different neighboring genes of each gene, and learn the gene representation to predict co-expression in a supervised and task-specific manner. Then, from the trained model, we extract the learned gene embeddings as a model interpretation to distill biological insights. Experimental results show that HiCoEx can learn gene representation from 3D genomics signals automatically to improve prediction accuracy, and make the black box model explainable by capturing some biologically meaningful patterns, e.g., in a gene contact network, the common neighbors of two central genes might contribute to the co-expression of the two central genes through sharing enhancers. AVAILABILITY AND IMPLEMENTATION: The source code is freely available at https://github.com/JieZheng-ShanghaiTech/HiCoEx. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jie Zheng 0002
Bioinform.4
2021 Graph contextualized attention network for predicting synthetic lethality in human cancers
abstract
MOTIVATION: Synthetic Lethality (SL) plays an increasingly critical role in the targeted anticancer therapeutics. In addition, identifying SL interactions can create opportunities to selectively kill cancer cells without harming normal cells. Given the high cost of wet-lab experiments, in silico prediction of SL interactions as an alternative can be a rapid and cost-effective way to guide the experimental screening of candidate SL pairs. Several matrix factorization-based methods have recently been proposed for human SL prediction. However, they are limited in capturing the dependencies of neighbors. In addition, it is also highly challenging to make accurate predictions for new genes without any known SL partners. RESULTS: In this work, we propose a novel graph contextualized attention network named GCATSL to learn gene representations for SL prediction. First, we leverage different data sources to construct multiple feature graphs for genes, which serve as the feature inputs for our GCATSL method. Second, for each feature graph, we design node-level attention mechanism to effectively capture the importance of local and global neighbors and learn local and global representations for the nodes, respectively. We further exploit multi-layer perceptron (MLP) to aggregate the original features with the local and global representations and then derive the feature-specific representations. Third, to derive the final representations, we design feature-level attention to integrate feature-specific representations by taking the importance of different feature graphs into account. Extensive experimental results on three datasets under different settings demonstrated that our GCATSL model outperforms 14 state-of-the-art methods consistently. In addition, case studies further validated the effectiveness of our proposed model in identifying novel SL pairs. AVAILABILITYAND IMPLEMENTATION: Python codes and dataset are freely available on GitHub (https://github.com/longyahui/GCATSL) and Zenodo (https://zenodo.org/record/4522679) under the MIT license. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yahui Long, Min Wu 0008, Yong Liu 0020, Jie Zheng 0002, Chee Keong Kwoh 0001, Jiawei Luo 0001, Xiaoli Li 0001
Bioinform.4
2021 KG4SL: knowledge graph neural network for synthetic lethality prediction in human cancers
abstract
MOTIVATION: Synthetic lethality (SL) is a promising gold mine for the discovery of anti-cancer drug targets. Wet-lab screening of SL pairs is afflicted with high cost, batch-effect, and off-target problems. Current computational methods for SL prediction include gene knock-out simulation, knowledge-based data mining and machine learning methods. Most of the existing methods tend to assume that SL pairs are independent of each other, without taking into account the shared biological mechanisms underlying the SL pairs. Although several methods have incorporated genomic and proteomic data to aid SL prediction, these methods involve manual feature engineering that heavily relies on domain knowledge. RESULTS: Here, we propose a novel graph neural network (GNN)-based model, named KG4SL, by incorporating knowledge graph (KG) message-passing into SL prediction. The KG was constructed using 11 kinds of entities including genes, compounds, diseases, biological processes and 24 kinds of relationships that could be pertinent to SL. The integration of KG can help harness the independence issue and circumvent manual feature engineering by conducting message-passing on the KG. Our model outperformed all the state-of-the-art baselines in area under the curve, area under precision-recall curve and F1. Extensive experiments, including the comparison of our model with an unsupervised TransE model, a vanilla graph convolutional network model, and their combination, demonstrated the significant impact of incorporating KG into GNN for SL prediction. AVAILABILITY AND IMPLEMENTATION: : KG4SL is freely available at https://github.com/JieZheng-ShanghaiTech/KG4SL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Shike Wang, Yunyang Li, Yong Liu 0020, Min Wu 0008, Jie Zheng 0002
Bioinform.8
2021 PIKE-R2P: Protein-protein interaction network-based knowledge embedding with graph neural network for single-cell RNA to protein prediction
abstract
BACKGROUND: Recent advances in simultaneous measurement of RNA and protein abundances at single-cell level provide a unique opportunity to predict protein abundance from scRNA-seq data using machine learning models. However, existing machine learning methods have not considered relationship among the proteins sufficiently. RESULTS: We formulate this task in a multi-label prediction framework where multiple proteins are linked to each other at the single-cell level. Then, we propose a novel method for single-cell RNA to protein prediction named PIKE-R2P, which incorporates protein-protein interactions (PPI) and prior knowledge embedding into a graph neural network. Compared with existing methods, PIKE-R2P could significantly improve prediction performance in terms of smaller errors and higher correlations with the gold standard measurements. CONCLUSION: The superior performance of PIKE-R2P indicates that adding the prior knowledge of PPI to graph neural networks can be a powerful strategy for cross-modality prediction of protein abundances at the single-cell level.
Xinnan Dai, Shike Wang, Piyushkumar A. Mundra, Jie Zheng 0002
BMC Bioinform.5
2021 Velo-Predictor: an ensemble learning pipeline for RNA velocity prediction
abstract
BACKGROUND: RNA velocity is a novel and powerful concept which enables the inference of dynamical cell state changes from seemingly static single-cell RNA sequencing (scRNA-seq) data. However, accurate estimation of RNA velocity is still a challenging problem, and the underlying kinetic mechanisms of transcriptional and splicing regulations are not fully clear. Moreover, scRNA-seq data tend to be sparse compared with possible cell states, and a given dataset of estimated RNA velocities needs imputation for some cell states not yet covered. RESULTS: We formulate RNA velocity prediction as a supervised learning problem of classification for the first time, where a cell state space is divided into equal-sized segments by directions as classes, and the estimated RNA velocity vectors are considered as ground truth. We propose Velo-Predictor, an ensemble learning pipeline for predicting RNA velocities from scRNA-seq data. We test different models on two real datasets, Velo-Predictor exhibits good performance, especially when XGBoost was used as the base predictor. Parameter analysis and visualization also show that the method is robust and able to make biologically meaningful predictions. CONCLUSION: The accurate result shows that Velo-Predictor can effectively simplify the procedure by learning a predictive model from gene expression data, which could help to construct a continous landscape and give biologists an intuitive picture about the trend of cellular dynamics.
Jie Zheng 0002
BMC Bioinform.2
2020 Bayesian Data Fusion of Gene Expression and Histone Modification Profiles for Inference of Gene Regulatory Network
abstract
Accurately reconstructing gene regulatory networks (GRNs) from high-throughput gene expression data has been a major challenge in systems biology for decades. Many approaches have been proposed to solve this problem. However, there is still much room for the improvement of GRN inference. Integrating data from different sources is a promising strategy. Epigenetic modifications have a close relationship with gene regulation. Hence, epigenetic data such as histone modification profiles can provide useful information for uncovering regulatory interactions between genes. In this paper, we propose a method to integrate epigenetic data into the inference of GRNs. In particular, a dynamic Bayesian network (DBN) is employed to infer gene regulations from time-series gene expression data. Epigenetic data (histone modification profiles here) are integrated into the prior probability distribution of the Bayesian model. Our method has been validated on both synthetic and real datasets. Experimental results show that the integration of epigenetic data can significantly improve the performance of GRN inference. As more epigenetic datasets become available, our method would be useful for elucidating the gene regulatory mechanisms driving various cellular activities. The source code and testing datasets are available at https://github.com/Zheng-Lab/MetaGRN/tree/master/histonePrior.
Haifen Chen, D. A. K. Maduranga, Piyushkumar A. Mundra, Jie Zheng 0002
IEEE ACM Trans. Comput. Biol. Bioinform.4
2020 SL2MF: Predicting Synthetic Lethality in Human Cancers via Logistic Matrix Factorization
abstract
Synthetic lethality (SL) is a promising concept for novel discovery of anti-cancer drug targets. However, wet-lab experiments for detecting SLs are faced with various challenges, such as high cost, low consistency across platforms, or cell lines. Therefore, computational prediction methods are needed to address these issues. This paper proposes a novel SL prediction method, named SL2MF, which employs logistic matrix factorization to learn latent representations of genes from the observed SL data. The probability that two genes are likely to form SL is modeled by the linear combination of gene latent vectors. As known SL pairs are more trustworthy than unknown pairs, we design importance weighting schemes to assign higher importance weights for known SL pairs and lower importance weights for unknown pairs in SL2MF. Moreover, we also incorporate biological knowledge about genes from protein-protein interaction (PPI) data and Gene Ontology (GO). In particular, we calculate the similarity between genes based on their GO annotations and topological properties in the PPI network. Extensive experiments on the SL interaction data from SynLethDB database have been conducted to demonstrate the effectiveness of SL2MF.
Yong Liu 0020, Min Wu 0008, Xiaoli Li 0001, Jie Zheng 0002
IEEE ACM Trans. Comput. Biol. Bioinform.5
2020 Guest Editorial for the 29th International Conference on Genome Informatics (GIW 2018)
abstract
The six papers in this special section were presented at the 29th International Conference on Genome Informatics (GIW 2018) that was held at Kunming University of Science and Technology, Kunming, China on December 3-5, 2018.
Jie Zheng 0002, Jinyan Li 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1
2018 TROVE: a user-friendly tool for visualizing and analyzing cancer hallmarks in signaling networks
abstract
SUMMARY: Cancer hallmarks, a concept that seeks to explain the complexity of cancer initiation and development, provide a new perspective of studying cancer signaling which could lead to a greater understanding of this complex disease. However, to the best of our knowledge, there is currently a lack of tools that support such hallmark-based study of the cancer signaling network, thereby impeding the gain of knowledge in this area. We present TROVE, an user-friendly software that facilitates hallmark annotation, visualization and analysis in cancer signaling networks. In particular, TROVE facilitates hallmark analysis specific to particular cancer types. AVAILABILITY AND IMPLEMENTATION: Available under the Eclipse Public License from: https://sites.google.com/site/cosbyntu/softwares/trove and https://github.com/trove2017/Trove.
Huey-Eng Chua, Sourav S. Bhowmick, Jie Zheng 0002
Bioinform.3
2018 Single-cell gene expression analysis reveals β-cell dysfunction and deficit mechanisms in type 2 diabetes
abstract
BACKGROUND: Type 2 diabetes (T2D) is one of the most common chronic diseases. Studies on T2D are mainly built upon bulk-cell data analysis, which measures the average gene expression levels for a population of cells and cannot capture the inter-cell heterogeneity. The single-cell RNA-sequencing technology can provide additional information about the molecular mechanisms of T2D at single-cell level. RESULTS: In this work, we analyze three datasets of single-cell transcriptomes to reveal β-cell dysfunction and deficit mechanisms in T2D. Focused on the expression levels of key genes, we conduct discrimination of healthy and T2D β-cells using five machine learning classifiers, and extracted major influential factors by calculating correlation coefficients and mutual information. Our analysis shows that T2D β-cells are normal in insulin gene expression in the scenario of low cellular stress (especially oxidative stress), but appear dysfunctional under the circumstances of high cellular stress. Remarkably, oxidative stress plays an important role in affecting the expression of insulin gene. In addition, by analyzing the genes related to apoptosis, we found that the TNFR1-, BAX-, CAPN1- and CAPN2-dependent pathways may be crucial for β-cell apoptosis in T2D. Finally, personalized analysis indicates cell heterogeneity and individual-specific insulin gene expression. CONCLUSIONS: Oxidative stress is an important influential factor on insulin gene expression in T2D. Based on the uncovered mechanism of β-cell dysfunction and deficit, targeting key genes in the apoptosis pathway along with alleviating oxidative stress could be a potential treatment strategy for T2D.
Lichun Ma, Jie Zheng 0002
BMC Bioinform.2
2017 NetLand: quantitative modeling and visualization of Waddington's epigenetic landscape using probabilistic potential
abstract
SUMMARY: Waddington's epigenetic landscape is a powerful metaphor for cellular dynamics driven by gene regulatory networks (GRNs). Its quantitative modeling and visualization, however, remains a challenge, especially when there are more than two genes in the network. A software tool for Waddington's landscape has not been available in the literature. We present NetLand, an open-source software tool for modeling and simulating the kinetic dynamics of GRNs, and visualizing the corresponding Waddington's epigenetic landscape in three dimensions without restriction on the number of genes in a GRN. With an interactive and graphical user interface, NetLand can facilitate the knowledge discovery and experimental design in the study of cell fate regulation (e.g. stem cell differentiation and reprogramming). AVAILABILITY AND IMPLEMENTATION: NetLand can run under operating systems including Windows, Linux and OS X. The executive files and source code of NetLand as well as a user manual, example models etc. can be downloaded from http://netland-ntu.github.io/NetLand/ . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vivek Tanavde, Jie Zheng 0002
Bioinform.5
2017 HopLand: single-cell pseudotime recovery using continuous Hopfield network-based modeling of Waddington's epigenetic landscape
abstract
MOTIVATION: The interpretation of transcriptional dynamics in single-cell data, especially pseudotime estimation, could help understand the transition of gene expression profiles. The recovery of pseudotime increases the temporal resolution of single-cell transcriptional data, but is challenging due to the high variability in gene expression between individual cells. Here, we introduce HopLand, a pseudotime recovery method using continuous Hopfield network to map cells to a Waddington's epigenetic landscape. It reveals from the single-cell data the combinatorial regulatory interactions among genes that control the dynamic progression through successive cell states. RESULTS: We applied HopLand to different types of single-cell transcriptomic data. It achieved high accuracies of pseudotime prediction compared with existing methods. Moreover, a kinetic model can be extracted from each dataset. Through the analysis of such a model, we identified key genes and regulatory interactions driving the transition of cell states. Therefore, our method has the potential to generate fundamental insights into cell fate regulation. AVAILABILITY AND IMPLEMENTATION: The MATLAB implementation of HopLand is available at https://github.com/NetLand-NTU/HopLand . CONTACT: [email protected].
Jie Zheng 0002
Bioinform.2
2016 The Max-Min High-Order Dynamic Bayesian Network for Learning Gene Regulatory Networks with Time-Delayed Regulations
abstract
Accurately reconstructing gene regulatory network (GRN) from gene expression data is a challenging task in systems biology. Although some progresses have been made, the performance of GRN reconstruction still has much room for improvement. Because many regulatory events are asynchronous, learning gene interactions with multiple time delays is an effective way to improve the accuracy of GRN reconstruction. Here, we propose a new approach, called Max-Min high-order dynamic Bayesian network (MMHO-DBN) by extending the Max-Min hill-climbing Bayesian network technique originally devised for learning a Bayesian network's structure from static data. Our MMHO-DBN can explicitly model the time lags between regulators and targets in an efficient manner. It first uses constraint-based ideas to limit the space of potential structures, and then applies search-and-score ideas to search for an optimal HO-DBN structure. The performance of MMHO-DBN to GRN reconstruction was evaluated using both synthetic and real gene expression time-series data. Results show that MMHO-DBN is more accurate than current time-delayed GRN learning methods, and has an intermediate computing performance. Furthermore, it is able to learn long time-delayed relationships between genes. We applied sensitivity analysis on our model to study the performance variation along different parameter settings. The result provides hints on the setting of parameters of MMHO-DBN.
Yifeng Li 0001, Haifen Chen, Jie Zheng 0002, Alioune Ngom
IEEE ACM Trans. Comput. Biol. Bioinform.3
2015 Single-cell transcriptional analysis to uncover regulatory circuits driving cell fate decisions in early mouse development
abstract
MOTIVATION: Transcriptional regulatory networks controlling cell fate decisions in mammalian embryonic development remain elusive despite a long time of research. The recent emergence of single-cell RNA profiling technology raises hope for new discovery. Although experimental works have obtained intriguing insights into the mouse early development, a holistic and systematic view is still missing. Mathematical models of cell fates tend to be concept-based, not designed to learn from real data. To elucidate the regulatory mechanisms behind cell fate decisions, it is highly desirable to synthesize the data-driven and knowledge-driven modeling approaches. RESULTS: We propose a novel method that integrates the structure of a cell lineage tree with transcriptional patterns from single-cell data. This method adopts probabilistic Boolean network (PBN) for network modeling, and genetic algorithm as search strategy. Guided by the 'directionality' of cell development along branches of the cell lineage tree, our method is able to accurately infer the regulatory circuits from single-cell gene expression data, in a holistic way. Applied on the single-cell transcriptional data of mouse preimplantation development, our algorithm outperforms conventional methods of network inference. Given the network topology, our method can also identify the operational interactions in the gene regulatory network (GRN), corresponding to specific cell fate determination. This is one of the first attempts to infer GRNs from single-cell transcriptional data, incorporating dynamics of cell development along a cell lineage tree. AVAILABILITY AND IMPLEMENTATION: Implementation of our algorithm is available from the authors upon request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Haifen Chen, Shital K. Mishra, Paul Robson 0001, Mahesan Niranjan, Jie Zheng 0002
Bioinform.6
2015 Improving compound-protein interaction prediction by building up highly credible negative samples
abstract
MOTIVATION: Computational prediction of compound-protein interactions (CPIs) is of great importance for drug design and development, as genome-scale experimental validation of CPIs is not only time-consuming but also prohibitively expensive. With the availability of an increasing number of validated interactions, the performance of computational prediction approaches is severely impended by the lack of reliable negative CPI samples. A systematic method of screening reliable negative sample becomes critical to improving the performance of in silico prediction methods. RESULTS: This article aims at building up a set of highly credible negative samples of CPIs via an in silico screening method. As most existing computational models assume that similar compounds are likely to interact with similar target proteins and achieve remarkable performance, it is rational to identify potential negative samples based on the converse negative proposition that the proteins dissimilar to every known/predicted target of a compound are not much likely to be targeted by the compound and vice versa. We integrated various resources, including chemical structures, chemical expression profiles and side effects of compounds, amino acid sequences, protein-protein interaction network and functional annotations of proteins, into a systematic screening framework. We first tested the screened negative samples on six classical classifiers, and all these classifiers achieved remarkably higher performance on our negative samples than on randomly generated negative samples for both human and Caenorhabditis elegans. We then verified the negative samples on three existing prediction models, including bipartite local model, Gaussian kernel profile and Bayesian matrix factorization, and found that the performances of these models are also significantly improved on the screened negative samples. Moreover, we validated the screened negative samples on a drug bioactivity dataset. Finally, we derived two sets of new interactions by training an support vector machine classifier on the positive interactions annotated in DrugBank and our screened negative interactions. The screened negative samples and the predicted interactions provide the research community with a useful resource for identifying new drug targets and a helpful supplement to the current curated compound-protein databases. AVAILABILITY: Supplementary files are available at: http://admis.fudan.edu.cn/negative-cpi/.
Hui Liu 0026, Jianjiang Sun, Jihong Guan, Jie Zheng 0002, Shuigeng Zhou
Bioinform.4
2014 Extracting rate changes in transcriptional regulation from MEDLINE abstracts
abstract
BACKGROUND: Time delays are important factors that are often neglected in gene regulatory network (GRN) inference models. Validating time delays from knowledge bases is a challenge since the vast majority of biological databases do not record temporal information of gene regulations. Biological knowledge and facts on gene regulations are typically extracted from bio-literature with specialized methods that depend on the regulation task. In this paper, we mine evidences for time delays related to the transcriptional regulation of yeast from the PubMed abstracts. RESULTS: Since the vast majority of abstracts lack quantitative time information, we can only collect qualitative evidences of time delays. Specifically, the speed-up or delay in transcriptional regulation rate can provide evidences for time delays (shorter or longer) in GRN. Thus, we focus on deriving events related to rate changes in transcriptional regulation. A corpus of yeast regulation related abstracts was manually labeled with such events. In order to capture these events automatically, we create an ontology of sub-processes that are likely to result in transcription rate changes by combining textual patterns and biological knowledge. We also propose effective feature extraction methods based on the created ontology to identify the direct evidences with specific details of these events. Our ontologies outperform existing state-of-the-art gene regulation ontologies in the automatic rule learning method applied to our corpus. The proposed deterministic ontology rule-based method can achieve comparable performance to the automatic rule learning method based on decision trees. This demonstrates the effectiveness of our ontology in identifying rate-changing events. We also tested the effectiveness of the proposed feature mining methods on detecting direct evidence of events. Experimental results show that the machine learning method on these features achieves an F1-score of 71.43%. CONCLUSIONS: The manually labeled corpus of events relating to rate changes in transcriptional regulation for yeast is available in https://sites.google.com/site/wentingntu/data. The created ontologies summarized both biological causes of rate changes in transcriptional regulation and corresponding positive and negative textual patterns from the corpus. They are demonstrated to be effective in identifying rate-changing events, which shows the benefits of combining textual patterns and biological knowledge on extracting complex biological events.
Kui Miao, Guangxia Li, Kuiyu Chang, Jie Zheng 0002, Jagath C. Rajapakse
BMC Bioinform.5
2014 IFACEwat: the interfacial water-implemented re-ranking algorithm to improve the discrimination of near native structures for protein rigid docking
abstract
Protein-protein docking is an in silico method to predict the formation of protein complexes. Due to limited computational resources, the protein-protein docking approach has been developed under the assumption of rigid docking, in which one of the two protein partners remains rigid during the protein associations and water contribution is ignored or implicitly presented. Despite obtaining a number of acceptable complex predictions, it seems to-date that most initial rigid docking algorithms still find it difficult or even fail to discriminate successfully the correct predictions from the other incorrect or false positive ones. To improve the rigid docking results, re-ranking is one of the effective methods that help re-locate the correct predictions in top high ranks, discriminating them from the other incorrect ones. In this paper, we propose a new re-ranking technique using a new energy-based scoring function, namely IFACEwat - a combined Interface Atomic Contact Energy (IFACE) and water effect. The IFACEwat aims to further improve the discrimination of the near-native structures of the initial rigid docking algorithm ZDOCK3.0.2. Unlike other re-ranking techniques, the IFACEwat explicitly implements interfacial water into the protein interfaces to account for the water-mediated contacts during the protein interactions. Our results showed that the IFACEwat increased both the numbers of the near-native structures and improved their ranks as compared to the initial rigid docking ZDOCK3.0.2. In fact, the IFACEwat achieved a success rate of 83.8% for Antigen/Antibody complexes, which is 10% better than ZDOCK3.0.2. As compared to another re-ranking technique ZRANK, the IFACEwat obtains success rates of 92.3% (8% better) and 90% (5% better) respectively for medium and difficult cases. When comparing with the latest published re-ranking method F 2 Dock, the IFACEwat performed equivalently well or even better for several Antigen/Antibody complexes. With the inclusion of interfacial water, the IFACEwat improves mostly results of the initial rigid docking, especially for Antigen/Antibody complexes. The improvement is achieved by explicitly taking into account the contribution of water during the protein interactions, which was ignored or not fully presented by the initial rigid docking and other re-ranking techniques. In addition, the IFACEwat maintains sufficient computational efficiency of the initial docking algorithm, yet improves the ranks as well as the number of the near native structures found. As our implementation so far targeted to improve the results of ZDOCK3.0.2, and particularly for the Antigen/Antibody complexes, it is expected in the near future that more implementations will be conducted to be applicable for other initial rigid docking algorithms.
Chinh Tran To Su, Thuy-Diem Nguyen, Jie Zheng 0002, Chee Keong Kwoh 0001
BMC Bioinform.3
2014 LDsplit: screening for cis-regulatory motifs stimulating meiotic recombination hotspots by analysis of DNA sequence polymorphisms
abstract
BACKGROUND: As a fundamental genomic element, meiotic recombination hotspot plays important roles in life sciences. Thus uncovering its regulatory mechanisms has broad impact on biomedical research. Despite the recent identification of the zinc finger protein PRDM9 and its 13-mer binding motif as major regulators for meiotic recombination hotspots, other regulators remain to be discovered. Existing methods for finding DNA sequence motifs of recombination hotspots often rely on the enrichment of co-localizations between hotspots and short DNA patterns, which ignore the cross-individual variation of recombination rates and sequence polymorphisms in the population. Our objective in this paper is to capture signals encoded in genetic variations for the discovery of recombination-associated DNA motifs. RESULTS: Recently, an algorithm called "LDsplit" has been designed to detect the association between single nucleotide polymorphisms (SNPs) and proximal meiotic recombination hotspots. The association is measured by the difference of population recombination rates at a hotspot between two alleles of a candidate SNP. Here we present an open source software tool of LDsplit, with integrative data visualization for recombination hotspots and their proximal SNPs. Applying LDsplit on SNPs inside an established 7-mer motif bound by PRDM9 we observed that SNP alleles preserving the original motif tend to have higher recombination rates than the opposite alleles that disrupt the motif. Running on SNP windows around hotspots each containing an occurrence of the 7-mer motif, LDsplit is able to guide the established motif finding algorithm of MEME to recover the 7-mer motif. In contrast, without LDsplit the 7-mer motif could not be identified. CONCLUSIONS: LDsplit is a software tool for the discovery of cis-regulatory DNA sequence motifs stimulating meiotic recombination hotspots by screening and narrowing down to hotspot associated SNPs. It is the first computational method that utilizes the genetic variation of recombination hotspots among individuals, opening a new avenue for motif finding. Tested on an established motif and simulated datasets, LDsplit shows promise to discover novel DNA motifs for meiotic recombination hotspots.
Peng Yang 0010, Min Wu 0008, Chee Keong Kwoh 0001, Teresa M. Przytycka, Jie Zheng 0002
BMC Bioinform.6
2014 Reliable and Fast Estimation of Recombination Rates by Convergence Diagnosis and Parallel Markov Chain Monte Carlo
abstract
Genetic recombination is an essential event during the process of meiosis resulting in an exchange of segments between paired chromosomes. Estimating recombination rate is crucial for understanding the process of recombination. Experimental methods are normally difficult and limited to small scale estimations. Thus statistical methods using population genetics data are important for large-scale analysis. LDhat is an extensively used statistical method using rjMCMC algorithm to predict recombination rates. Due to the complexity of rjMCMC scheme, LDhat may take a long time for large SNP data sets. In addition, rjMCMC parameters should be manually defined in the original program which directly impact results. To address these issues, we designed an improved algorithm based on LDhat implementing MCMC convergence diagnostic algorithms to automatically predict values of parameters and monitor the mixing process. Then parallel computation methods were employed to further accelerate the new program. The new algorithms have been tested on ten samples from HapMap phase 2 data set. The results were compared with previous code and showed nearly identical output. However, our new methods achieved significant acceleration proving that they are more efficient and reliable for the estimation of recombination rates. The stand-alone package is freely available for download http://www.ntu.edu.sg/home/zhengjie/software/CPLDhat.
Ritika Jain, Peng Yang 0010, Chee Keong Kwoh 0001, Jie Zheng 0002
IEEE ACM Trans. Comput. Biol. Bioinform.6
2013 Syn-Lethality: An integrative knowledge base of synthetic lethality towards discovery of selective anticancer therapies
abstract
Synthetic lethality (SL) is a novel strategy for anticancer therapies, whereby mutations of two genes will kill a cell but mutation of a single gene will not. Therefore, a cancer-specific mutation combined with a drug-induced mutation, if they have SL interactions, will selectively kill cancer cells. While numerous SL interactions have been identified in yeast, only a few have been known in human. There is a pressing need to systematically discover and understand SL interactions specific to human cancer. In this paper, we present Syn-Lethality, the first integrative knowledge base of SL that is dedicated to human cancer. It integrates experimentally discovered and verified human SL gene pairs into a network, associated with annotations of gene functions, pathways and molecular mechanisms. It also includes yeast SL genes from high-throughput screenings which are mapped to orthologous human genes. Such an integrative knowledge base, organized as a relational database with user interface for searching and network visualization, will greatly expedite the discovery of novel anticancer drug targets based on synthetic lethality interactions. The database can be downloaded as a stand-alone Java application from: http://www.ntu.edu.sg/home/zhengjie/software/Syn-Lethality/.
Xuejuan Li, Shital K. Mishra, Min Wu 0008, Fan Zhang 0089, Jie Zheng 0002
BIBM5
2013 Integrating epigenetic prior in dynamic Bayesian network for gene regulatory network inference
abstract
Gene regulatory network (GRN) inference from high throughput biological data has drawn a lot of research interest in the last decade. However, due to the complexity of gene regulation and lack of sufficient data, GRN inference still has much space to improve. One way to improve the inference of GRN is by developing methods to accurately combine various types of data. Here we apply dynamic Bayesian network (DBN) to infer GRN from time-series gene expression data where the Bayesian prior is derived from epigenetic data of histone modifications. We propose several kinds of prior from histone modification data, and use both real and synthetic data to compare their performance. Parameters of prior integration are also studied to achieve better results. Experiments on gene expression data of yeast cell cycle show that our methods increase the accuracy of GRN inference significantly.
Haifen Chen, D. A. K. Maduranga, Piyushkumar A. Mundra, Jie Zheng 0002
CIBCB4
2013 Gene Regulatory Networks from Gene Ontology
Kuiyu Chang, Jie Zheng 0002, Jain Divya, Jung-Jae Kim 0001, Jagath C. Rajapakse
ISBRA3
2013 Inferring Time-Delayed Gene Regulatory Networks Using Cross-Correlation and Sparse Regression
Piyushkumar A. Mundra, Jie Zheng 0002, Mahesan Niranjan, Roy E. Welsch, Jagath C. Rajapakse
ISBRA2
2013 Drug-target interaction prediction by learning from local information and neighbors
abstract
MOTIVATION: In silico methods provide efficient ways to predict possible interactions between drugs and targets. Supervised learning approach, bipartite local model (BLM), has recently been shown to be effective in prediction of drug-target interactions. However, for drug-candidate compounds or target-candidate proteins that currently have no known interactions available, its pure 'local' model is not able to be learned and hence BLM may fail to make correct prediction when involving such kind of new candidates. RESULTS: We present a simple procedure called neighbor-based interaction-profile inferring (NII) and integrate it into the existing BLM method to handle the new candidate problem. Specifically, the inferred interaction profile is treated as label information and is used for model learning of new candidates. This functionality is particularly important in practice to find targets for new drug-candidate compounds and identify targeting drugs for new target-candidate proteins. Consistent good performance of the new BLM-NII approach has been observed in the experiment for the prediction of interactions between drugs and four categories of target proteins. Especially for nuclear receptors, BLM-NII achieves the most significant improvement as this dataset contains many drugs/targets with no interactions in the cross-validation. This demonstrates the effectiveness of the NII strategy and also shows the great potential of BLM-NII for prediction of compound-protein interactions. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jian-Ping Mei, Chee Keong Kwoh 0001, Peng Yang 0010, Xiaoli Li 0001, Jie Zheng 0002
Bioinform.5
2013 Structural analysis of the novel influenza A (H7N9) viral Neuraminidase interactions with current approved neuraminidase inhibitors Oseltamivir, Zanamivir, and Peramivir in the presence of mutation R289K
abstract
BACKGROUND: Since late March 2013, there has been another global health concern with a sudden wave of flu infections by a novel strain of avian influenza A (H7N9) virus in China. To-date, there have been more than 100 infections with 23 deaths. It is more worrying as this viral strain has never been detected in humans and only been found to be of low-pathogenicity. Currently, there are 3 effective neuraminidase inhibitors for this H7N9 virus strain, i.e. oseltamivir, zanamivir, and peramivir. These drugs have been used for treatment of the H7N9 influenza in China. However, how these inhibitors work and affect the binding cavity of the novel H7N9 neuraminidase in the presence of potential mutations has not been disclosed. In our study, we investigate steric effects and subsequently show the conformational restraints of the inhibitor-binding site of the non-mutated and mutated H7N9 neuraminidase structures to different drug compounds. RESULTS: Combination of molecular docking and Molecular Dynamics simulation reveal that zanamivir forms more favorable and stable complex than oseltamivir and peramivir when binding to the active site of the H7N9 neuraminidase. And it is likely that the novel influenza A (H7N9) virus adopts a higher probability to acquire resistance to peramivir than the other two inhibitors. Conformational changes induced by the mutation R289K causes loss of number of hydrogen bonds between the inhibitors and the H7N9 viral neuraminidase in 2 out of 3 complexes. In addition, our results of binding-affinity relationships of the 3 inhibitors with the viral neuraminidase proteins of previous pandemics (H1N1, H5N1) and the current novel H7N9 reflected the extent of binding effectiveness of the 3 inhibitors to the novel H7N9 neuraminidase. CONCLUSIONS: The results are novel and specific for the A/Hangzhou/1/2013(H7N9) influenza strain. Furthermore, the protocol could be useful for further drug-binding analysis and prediction of future viral mutations to which the virus evolves through adaptation and acquires resistance to the current available drugs.
Chinh Tran To Su, Xuchang Ouyang, Jie Zheng 0002, Chee Keong Kwoh 0001
BMC Bioinform.3
2012 Teasing Apart Translational and Transcriptional Components of Stochastic Variations in Eukaryotic Gene Expression
abstract
The intrinsic stochasticity of gene expression leads to cell-to-cell variations, noise, in protein abundance. Several processes, including transcription, translation, and degradation of mRNA and proteins, can contribute to these variations. Recent single cell analyses of gene expression in yeast have uncovered a general trend where expression noise scales with protein abundance. This trend is consistent with a stochastic model of gene expression where mRNA copy number follows the random birth and death process. However, some deviations from this basic trend have also been observed, prompting questions about the contribution of gene-specific features to such deviations. For example, recent studies have pointed to the TATA box as a sequence feature that can influence expression noise by facilitating expression bursts. Transcription-originated noise can be potentially further amplified in translation. Therefore, we asked the question of to what extent sequence features known or postulated to accompany translation efficiency can also be associated with increase in noise strength and, on average, how such increase compares to the amplification associated with the TATA box. Untangling different components of expression noise is highly nontrivial, as they may be gene or gene-module specific. In particular, focusing on codon usage as one of the sequence features associated with efficient translation, we found that ribosomal genes display a different relationship between expression noise and codon usage as compared to other genes. Within nonribosomal genes we found that sequence high codon usage is correlated with increased noise relative to the average noise of proteins with the same abundance. Interestingly, by projecting the data on a theoretical model of gene expression, we found that the amplification of noise strength associated with codon usage is comparable to that of the TATA box, suggesting that the effect of translation on noise in eukaryotic gene expression might be more prominent than previously appreciated.
Raheleh Salari, Damian Wójtowicz, Jie Zheng 0002, David Levens, Yitzhak Pilpel, Teresa M. Przytycka
PLoS Comput. Biol.3
2011 Prediction of Trans-regulators of Recombination Hotspots in Mouse Genome
abstract
The regulatory mechanism of recombination is a fundamental problem in genomics, with wide applications in genome wide association studies, birth-defect diseases, molecular evolution, cancer research, etc. In mammalian genomes, recombination events cluster into short genomic regions called ¡§recombination hotspots¡¨. Recently, a 13-mer motif enriched in hotspots is identified as a candidate cis-regulatory element of human recombination hotspots, moreover, a zinc finger protein, PRDM9, binds to this motif and is associated with variation of recombination phenotype in human and mouse genomes, thus is a trans-acting regulator of recombination hotspots. However, this pair of cis and trans-regulators covers only a fraction of hotspots, thus other regulators of recombination hotspots remain to be discovered. In this paper, we propose an approach to predicting additional trans-regulators from DNA-binding proteins by comparing their enrichment of binding sites in hotspots. Applying this approach on newly mapped mouse hotspots genome-wide, we confirmed that PRDM9 is a major trans-regulator of hotspots. In addition, a list of top candidate trans-regulators of mouse hotspots is reported. Using GO analysis we observed that the top genes are enriched with function of his tone modification, highlighting the epigenetic regulatory mechanisms of recombination hotspots.
Min Wu 0008, Chee Keong Kwoh 0001, Teresa M. Przytycka, Jing Li 0002, Jie Zheng 0002
BIBM5
2010 SimBoolNet - a Cytoscape plugin for dynamic simulation of signaling networks
abstract
SUMMARY: SimBoolNet is an open source Cytoscape plugin that simulates the dynamics of signaling transduction using Boolean networks. Given a user-specified level of stimulation to signal receptors, SimBoolNet simulates the response of downstream molecules and visualizes with animation and records the dynamic changes of the network. It can be used to generate hypotheses and facilitate experimental studies about causal relations and crosstalk among cellular signaling pathways. AVAILABILITY: SimBoolNet package (with manual) is freely available at http://www.ncbi.nlm.nih.gov/CBBresearch/Przytycka/SimBoolNet
Jie Zheng 0002, Pawel F. Przytycki, Rafal Zielinski, Jacek Capala, Teresa M. Przytycka
Bioinform.1
2006 OligoSpawn: a software tool for the design of overgo probes from large unigene datasets
abstract
BACKGROUND: Expressed sequence tag (EST) datasets represent perhaps the largest collection of genetic information. ESTs can be exploited in a variety of biological experiments and analysis. Here we are interested in the design of overlapping oligonucleotide (overgo) probes from large unigene (EST-contigs) datasets. RESULTS: OLIGOSPAWN is a suite of software tools that offers two complementary services, namely (1) the selection of "unique" oligos each of which appears in one unigene but does not occur (exactly or approximately) in any other and (2) the selection of "popular" oligos each of which occurs (exactly or approximately) in as many unigenes as possible. In this paper, we describe the functionalities of OLIGOSPAWN and the computational methods it employs, and we report on experimental results for the overgo probes designed with it. CONCLUSION: The algorithms we designed are highly efficient and capable of processing unigene datasets of sizes on the order of several tens of Mb in a few hours on a regular PC. The software has been used to design overgo probes employed to screen a barley BAC library (Hordeum vulgare). OLIGOSPAWN is freely available at http://oligospawn.ucr.edu/.
Jie Zheng 0002, Jan T. Svensson, Kavitha Madishetty, Timothy J. Close, Tao Jiang 0001, Stefano Lonardi
BMC Bioinform.1
2005 Computing the Assignment of Orthologous Genes via Genome Rearrangement
Xin Chen 0037, Jie Zheng 0002, Zheng Fu, Peng Nan, Stefano Lonardi, Tao Jiang 0001
APBC2
2005 Discovery of Repetitive Patterns in DNA with Accurate Boundaries
abstract
The accurate identification of repeats remains a challenging open problem in bioinformatics. Most existing methods of repeat identification either depend on annotated repeat databases or restrict repeats to pairs of similar sequences that are maximal in length. The fundamental flaw in most of the available methods is the lack of a definition that correctly balances the importance of the length and the frequency. In this paper, we propose a new definition of repeats that satisfies both criteria. We give a novel characterization of the building blocks of repeats, called elementary repeats, which leads to a natural definition of repeat boundaries. We design efficient algorithms and test them on synthetic and real biological data. Experimental results show that our method is highly accurate.
Jie Zheng 0002, Stefano Lonardi
BIBE1
2005 Assignment of Orthologous Genes via Genome Rearrangement
abstract
The assignment of orthologous genes between a pair of genomes is a fundamental and challenging problem in comparative genomics. Existing methods that assign orthologs based on the similarity between DNA or protein sequences may make erroneous assignments when sequence similarity does not clearly delineate the evolutionary relationship among genes of the same families. In this paper, we present a new approach to ortholog assignment that takes into account both sequence similarity and evolutionary events at a genome level, where orthologous genes are assumed to correspond to each other in the most parsimonious evolving scenario under genome rearrangement. First, the problem is formulated as that of computing the signed reversal distance with duplicates between the two genomes of interest. Then, the problem is decomposed into two new optimization problems, called minimum common partition and maximum cycle decomposition, for which efficient heuristic algorithms are given. Following this approach, we have implemented a high-throughput system for assigning orthologs on a genome scale, called SOAR, and tested it on both simulated data and real genome sequence data. Compared to a recent ortholog assignment method based entirely on homology search (called INPARANOID), SOAR shows a marginally better performance in terms of sensitivity on the real data set because it is able to identify several correct orthologous pairs that are missed by INPARANOID. The simulation results demonstrate that SOAR, in general, performs better than the iterated exemplar algorithm in terms of computing the reversal distance and assigning correct orthologs.
Xin Chen 0037, Jie Zheng 0002, Zheng Fu, Peng Nan, Stefano Lonardi, Tao Jiang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2004 Efficient selection of unique and popular oligos for large EST databases
abstract
MOTIVATION: Expressed sequence tag (EST) databases have grown exponentially in recent years and now represent the largest collection of genetic sequences. An important application of these databases is that they contain information useful for the design of gene-specific oligonucleotides (or simply, oligos) that can be used in PCR primer design, microarray experiments and genomic library screening. RESULTS: In this paper, we study two complementary problems concerning the selection of short oligos, e.g. 20-50 bases, from a large database of tens of thousands of ESTs: (i) selection of oligos each of which appears (exactly) in one unigene but does not appear (exactly or approximately) in any other unigene and (ii) selection of oligos that appear (exactly or approximately) in many unigenes. The first problem is called the unique oligo problem and has applications in PCR primer and microarray probe designs, and library screening for gene-rich clones. The second is called the popular oligo problem and is also useful in screening genomic libraries. We present an efficient algorithm to identify all unique oligos in the unigenes and an efficient heuristic algorithm to enumerate the most popular oligos. By taking into account the distribution of the frequencies of the words in the unigene database, the algorithms have been engineered carefully to achieve remarkable running times on regular PCs. Each of the algorithms takes only a couple of hours (on a 1.2 GHz CPU, 1 GB RAM machine) to run on a dataset 28 Mb of barley unigenes from the HarvEST database. We present simulation results on the synthetic data and a preliminary analysis of the barley unigene database. AVAILABILITY: Available on request from the authors.
Jie Zheng 0002, Timothy J. Close, Tao Jiang 0001, Stefano Lonardi
Bioinform.1
2003 Efficient Selection of Unique and Popular Oligos for Large EST Databases
Jie Zheng 0002, Timothy J. Close, Tao Jiang 0001, Stefano Lonardi
CPM1