Jianyang Zeng 0001

dblp:59/6722-1 · DBLP profile ↗
← Back
36ranked-venue papers
5as first author
11since 2021 · last 2024
0000-0003-0950-7716ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 31 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2024 PRS-Net: Interpretable Polygenic Risk Scores via Geometric Learning
Han Li 0018, Jianyang Zeng 0001, Michael Snyder 0001
RECOMB2
2024 DrugRepPT: a deep pretraining and fine-tuning framework for drug repositioning based on drug's expression perturbation and treatment effectiveness
abstract
MOTIVATION: Drug repositioning (DR), identifying novel indications for approved drugs, is a cost-effective strategy in drug discovery. Despite numerous proposed DR models, integrating network-based features, differential gene expression, and chemical structures for high-performance DR remains challenging. RESULTS: We propose a comprehensive deep pretraining and fine-tuning framework for DR, termed DrugRepPT. Initially, we design a graph pretraining module employing model-augmented contrastive learning on a vast drug-disease heterogeneous graph to capture nuanced interactions and expression perturbations after intervention. Subsequently, we introduce a fine-tuning module leveraging a graph residual-like convolution network to elucidate intricate interactions between diseases and drugs. Moreover, a Bayesian multiloss approach is introduced to balance the existence and effectiveness of drug treatment effectively. Extensive experiments showcase the efficacy of our framework, with DrugRepPT exhibiting remarkable performance improvements compared to SOTA (state of the arts) baseline methods (improvement 106.13% on Hit@1 and 54.45% on mean reciprocal rank). The reliability of predicted results is further validated through two case studies, i.e. gastritis and fatty liver, via literature validation, network medicine analysis, and docking screening. AVAILABILITY AND IMPLEMENTATION: The code and results are available at https://github.com/2020MEAI/DrugRepPT.
Shuyue Fan, Kuo Yang 0001, Kezhi Lu, Xin Dong 0017, Xianan Li, Shao Li, Jianyang Zeng 0001, Xuezhong Zhou
Bioinform.8
2024 A probabilistic knowledge graph for target identification
abstract
Early identification of safe and efficacious disease targets is crucial to alleviating the tremendous cost of drug discovery projects. However, existing experimental methods for identifying new targets are generally labor-intensive and failure-prone. On the other hand, computational approaches, especially machine learning-based frameworks, have shown remarkable application potential in drug discovery. In this work, we propose Progeni, a novel machine learning-based framework for target identification. In addition to fully exploiting the known heterogeneous biological networks from various sources, Progeni integrates literature evidence about the relations between biological entities to construct a probabilistic knowledge graph. Graph neural networks are then employed in Progeni to learn the feature embeddings of biological entities to facilitate the identification of biologically relevant target candidates. A comprehensive evaluation of Progeni demonstrated its superior predictive power over the baseline methods on the target identification task. In addition, our extensive tests showed that Progeni exhibited high robustness to the negative effect of exposure bias, a common phenomenon in recommendation systems, and effectively identified new targets that can be strongly supported by the literature. Moreover, our wet lab experiments successfully validated the biological significance of the top target candidates predicted by Progeni for melanoma and colorectal cancer. All these results suggested that Progeni can identify biologically effective targets and thus provide a powerful and useful tool for advancing the drug discovery process.
Chang Liu 0082, Kaimin Xiao, Cuinan Yu, Yipin Lei, Kangbo Lyu, Tingzhong Tian, Dan Zhao 0004, Fengfeng Zhou, Haidong Tang, Jianyang Zeng 0001
PLoS Comput. Biol.10
2024 Analyzing Large-Scale Single-Cell RNA-Seq Data Using Coreset
abstract
The recent boom in single-cell sequencing technologies provides valuable insights into the transcriptomes of individual cells. Through single-cell data analyses, a number of biological discoveries, such as novel cell types, developmental cell lineage trajectories, and gene regulatory networks, have been uncovered. However, the massive and increasingly accumulated single-cell datasets have also posed a seriously computational and analytical challenge for researchers. To address this issue, one typically applies dimensionality reduction approaches to reduce the large-scale datasets. However, these approaches are generally computationally infeasible for tall matrices. In addition, the downstream data analysis tasks such as clustering still take a large time complexity even on the dimension-reduced datasets. We present single-cell Coreset (scCoreset), a data summarization framework that extracts a small weighted subset of cells from a huge sparse single-cell RNA-seq data to facilitate the downstream data analysis tasks. Single-cell data analyses run on the extracted subset yield similar results to those derived from the original uncompressed data. Tests on various single-cell datasets show that scCoreset outperforms the existing data summarization approaches for common downstream tasks such as visualization and clustering. We believe that scCoreset can serve as a useful plug-in tool to improve the efficiency of current single-cell RNA-seq data analyses.
Khalid Usman, Fangping Wan, Dan Zhao 0004, Jian Peng 0001, Jianyang Zeng 0001
IEEE ACM Trans. Comput. Biol. Bioinform.5
2023 Improving comparative analyses of Hi-C data via contrastive self-supervised learning
abstract
Hi-C is a widely applied chromosome conformation capture (3C)-based technique, which has produced a large number of genomic contact maps with high sequencing depths for a wide range of cell types, enabling comprehensive analyses of the relationships between biological functionalities (e.g. gene regulation and expression) and the three-dimensional genome structure. Comparative analyses play significant roles in Hi-C data studies, which are designed to make comparisons between Hi-C contact maps, thus evaluating the consistency of replicate Hi-C experiments (i.e. reproducibility measurement) and detecting statistically differential interacting regions with biological significance (i.e. differential chromatin interaction detection). However, due to the complex and hierarchical nature of Hi-C contact maps, it remains challenging to conduct systematic and reliable comparative analyses of Hi-C data. Here, we proposed sslHiC, a contrastive self-supervised representation learning framework, for precisely modeling the multi-level features of chromosome conformation and automatically producing informative feature embeddings for genomic loci and their interactions to facilitate comparative analyses of Hi-C contact maps. Comprehensive computational experiments on both simulated and real datasets demonstrated that our method consistently outperformed the state-of-the-art baseline methods in providing reliable measurements of reproducibility and detecting differential interactions with biological meanings.
Han Li 0018, Lawrence Kurowski, Ruotian Zhang, Dan Zhao 0004, Jianyang Zeng 0001
Briefings Bioinform.6
2023 DrugAI: a multi-view deep learning model for predicting drug-target activating/inhibiting mechanisms
abstract
Understanding the mechanisms of candidate drugs play an important role in drug discovery. The activating/inhibiting mechanisms between drugs and targets are major types of mechanisms of drugs. Owing to the complexity of drug-target (DT) mechanisms and data scarcity, modelling this problem based on deep learning methods to accurately predict DT activating/inhibiting mechanisms remains a considerable challenge. Here, by considering network pharmacology, we propose a multi-view deep learning model, DrugAI, which combines four modules, i.e. a graph neural network for drugs, a convolutional neural network for targets, a network embedding module for drugs and targets and a deep neural network for predicting activating/inhibiting mechanisms between drugs and targets. Computational experiments show that DrugAI performs better than state-of-the-art methods and has good robustness and generalization. To demonstrate the reliability of the predictive results of DrugAI, bioassay experiments are conducted to validate two drugs (notopterol and alpha-asarone) predicted to activate TRPV1. Moreover, external validation bears out 61 pairs of mechanism relationships between natural products and their targets predicted by DrugAI based on independent literatures and PubChem bioassays. DrugAI, for the first time, provides a powerful multi-view deep learning framework for robust prediction of DT activating/inhibiting mechanisms.
Siqin Zhang, Kuo Yang 0001, Xinxing Lai, Jianyang Zeng 0001, Shao Li
Briefings Bioinform.6
2022 KPGT: Knowledge-Guided Pre-training of Graph Transformer for Molecular Property Prediction
abstract
Designing accurate deep learning models for molecular property prediction plays an increasingly essential role in drug and material discovery. Recently, due to the scarcity of labeled molecules, self-supervised learning methods for learning generalizable and transferable representations of molecular graphs have attracted lots of attention. In this paper, we argue that there exist two major issues hindering current self-supervised learning methods from obtaining desired performance on molecular property prediction, that is, the ill-defined pre-training tasks and the limited model capacity. To this end, we introduce Knowledge-guided Pre-training of Graph Transformer (KPGT), a novel self-supervised learning framework for molecular graph representation learning, to alleviate the aforementioned issues and improve the performance on the downstream molecular property prediction tasks. More specifically, we first introduce a high-capacity model, named Line Graph Transformer (LiGhT), which emphasizes the importance of chemical bonds and is mainly designed to model the structural information of molecular graphs. Then, a knowledge-guided pre-training strategy is proposed to exploit the additional knowledge of molecules to guide the model to capture the abundant structural and semantic information from large-scale unlabeled molecular graphs. Extensive computational tests demonstrated that KPGT can offer superior performance over current state-of-the-art methods on several molecular property prediction tasks.
Han Li 0018, Dan Zhao 0004, Jianyang Zeng 0001
KDD3
2022 Boosting single-cell gene regulatory network reconstruction via bulk-cell transcriptomic data
abstract
Computational recovery of gene regulatory network (GRN) has recently undergone a great shift from bulk-cell towards designing algorithms targeting single-cell data. In this work, we investigate whether the widely available bulk-cell data could be leveraged to assist the GRN predictions for single cells. We infer cell-type-specific GRNs from both the single-cell RNA sequencing data and the generic GRN derived from the bulk cells by constructing a weakly supervised learning framework based on the axial transformer. We verify our assumption that the bulk-cell transcriptomic data are a valuable resource, which could improve the prediction of single-cell GRN by conducting extensive experiments. Our GRN-transformer achieves the state-of-the-art prediction accuracy in comparison to existing supervised and unsupervised approaches. In addition, we show that our method can identify important transcription factors and potential regulations for Alzheimer's disease risk genes by using the predicted GRN. Availability: The implementation of GRN-transformer is available at https://github.com/HantaoShu/GRN-Transformer.
Hantao Shu, Jingtian Zhou, Yexiang Xue, Dan Zhao 0004, Jianyang Zeng 0001, Jianzhu Ma
Briefings Bioinform.6
2021 Riboexp: an interpretable reinforcement learning framework for ribosome density modeling
abstract
Translation elongation is a crucial phase during protein biosynthesis. In this study, we develop a novel deep reinforcement learning-based framework, named Riboexp, to model the determinants of the uneven distribution of ribosomes on mRNA transcripts during translation elongation. In particular, our model employs a policy network to perform a context-dependent feature selection in the setting of ribosome density prediction. Our extensive tests demonstrated that Riboexp can significantly outperform the state-of-the-art methods in predicting ribosome density by up to 5.9% in terms of per-gene Pearson correlation coefficient on the datasets from three species. In addition, Riboexp can indicate more informative sequence features for the prediction task than other commonly used attribution methods in deep learning. In-depth analyses also revealed the meaningful biological insights generated by the Riboexp framework. Moreover, the application of Riboexp in codon optimization resulted in an increase of protein production by around 31% over the previous state-of-the-art method that models ribosome density. These results have established Riboexp as a powerful and useful computational tool in the studies of translation dynamics and protein synthesis. Availability: The data and code of this study are available on GitHub: https://github.com/Liuxg16/Riboexp. Contact:[email protected]; [email protected].
Hailin Hu 0002, Xianggen Liu, An Xiao, Chengdong Zhang, Tao Jiang 0001, Dan Zhao 0004, Sen Song, Jianyang Zeng 0001
Briefings Bioinform.9
2021 Predicting MHC-peptide binding affinity by differential boundary tree
abstract
MOTIVATION: The prediction of the binding between peptides and major histocompatibility complex (MHC) molecules plays an important role in neoantigen identification. Although a large number of computational methods have been developed to address this problem, they produce high false-positive rates in practical applications, since in most cases, a single residue mutation may largely alter the binding affinity of a peptide binding to MHC which cannot be identified by conventional deep learning methods. RESULTS: We developed a differential boundary tree-based model, named DBTpred, to address this problem. We demonstrated that DBTpred can accurately predict MHC class I binding affinity compared to the state-of-art deep learning methods. We also presented a parallel training algorithm to accelerate the training and inference process which enables DBTpred to be applied to large datasets. By investigating the statistical properties of differential boundary trees and the prediction paths to test samples, we revealed that DBTpred can provide an intuitive interpretation and possible hints in detecting important residue mutations that can largely influence binding affinity. AVAILABILITY AND IMPLEMENTATION: The DBTpred package is implemented in Python and freely available at: https://github.com/fpy94/DBT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Peiyuan Feng, Jianyang Zeng 0001, Jianzhu Ma
Bioinform.2
2021 Full-length ribosome density prediction by a multi-input and multi-output model
abstract
Translation elongation is regulated by a series of complicated mechanisms in both prokaryotes and eukaryotes. Although recent advance in ribosome profiling techniques has enabled one to capture the genome-wide ribosome footprints along transcripts at codon resolution, the regulatory codes of elongation dynamics are still not fully understood. Most of the existing computational approaches for modeling translation elongation from ribosome profiling data mainly focus on local contextual patterns, while ignoring the continuity of the elongation process and relations between ribosome densities of remote codons. Modeling the translation elongation process in full-length coding sequence (CDS) level has not been studied to the best of our knowledge. In this paper, we developed a deep learning based approach with a multi-input and multi-output framework, named RiboMIMO, for modeling the ribosome density distributions of full-length mRNA CDS regions. Through considering the underlying correlations in translation efficiency among neighboring and remote codons and extracting hidden features from the input full-length coding sequence, RiboMIMO can greatly outperform the state-of-the-art baseline approaches and accurately predict the ribosome density distributions along the whole mRNA CDS regions. In addition, RiboMIMO explores the contributions of individual input codons to the predictions of output ribosome densities, which thus can help reveal important biological factors influencing the translation elongation process. The analyses, based on our interpretable metric named codon impact score, not only identified several patterns consistent with the previously-published literatures, but also for the first time (to the best of our knowledge) revealed that the codons located at a long distance from the ribosomal A site may also have an association on the translation elongation rate. This finding of long-range impact on translation elongation velocity may shed new light on the regulatory mechanisms of protein synthesis. Overall, these results indicated that RiboMIMO can provide a useful tool for studying the regulation of translation elongation in the range of full-length CDS.
Tingzhong Tian, Shuya Li, Peng Lang, Dan Zhao 0004, Jianyang Zeng 0001
PLoS Comput. Biol.5
2020 MONN: A Multi-objective Neural Network for Predicting Pairwise Non-covalent Interactions and Binding Affinities Between Compounds and Proteins
Shuya Li, Fangping Wan, Hantao Shu, Tao Jiang 0001, Dan Zhao 0004, Jianyang Zeng 0001
RECOMB6
2020 Secure multiparty computation for privacy-preserving drug discovery
abstract
MOTIVATION: Quantitative structure-activity relationship (QSAR) and drug-target interaction (DTI) prediction are both commonly used in drug discovery. Collaboration among pharmaceutical institutions can lead to better performance in both QSAR and DTI prediction. However, the drug-related data privacy and intellectual property issues have become a noticeable hindrance for inter-institutional collaboration in drug discovery. RESULTS: We have developed two novel algorithms under secure multiparty computation (MPC), including QSARMPC and DTIMPC, which enable pharmaceutical institutions to achieve high-quality collaboration to advance drug discovery without divulging private drug-related information. QSARMPC, a neural network model under MPC, displays good scalability and performance and is feasible for privacy-preserving collaboration on large-scale QSAR prediction. DTIMPC integrates drug-related heterogeneous network data and accurately predicts novel DTIs, while keeping the drug information confidential. Under several experimental settings that reflect the situations in real drug discovery scenarios, we have demonstrated that DTIMPC possesses significant performance improvement over the baseline methods, generates novel DTI predictions with supporting evidence from the literature and shows the feasible scalability to handle growing DTI data. All these results indicate that QSARMPC and DTIMPC can provide practically useful tools for advancing privacy-preserving drug discovery. AVAILABILITY AND IMPLEMENTATION: The source codes of QSARMPC and DTIMPC are available on the GitHub: https://github.com/rongma6/QSARMPC_DTIMPC.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yi Li 0005, Chenxing Li, Fangping Wan, Hailin Hu 0002, Wei Xu 0005, Jianyang Zeng 0001
Bioinform.7
2019 DIFFUSE: predicting isoform functions from sequences and expression profiles via deep learning
abstract
MOTIVATION: Alternative splicing generates multiple isoforms from a single gene, greatly increasing the functional diversity of a genome. Although gene functions have been well studied, little is known about the specific functions of isoforms, making accurate prediction of isoform functions highly desirable. However, the existing approaches to predicting isoform functions are far from satisfactory due to at least two reasons: (i) unlike genes, isoform-level functional annotations are scarce. (ii) The information of isoform functions is concealed in various types of data including isoform sequences, co-expression relationship among isoforms, etc. RESULTS: In this study, we present a novel approach, DIFFUSE (Deep learning-based prediction of IsoForm FUnctions from Sequences and Expression), to predict isoform functions. To integrate various types of data, our approach adopts a hybrid framework by first using a deep neural network (DNN) to predict the functions of isoforms from their genomic sequences and then refining the prediction using a conditional random field (CRF) based on co-expression relationship. To overcome the lack of isoform-level ground truth labels, we further propose an iterative semi-supervised learning algorithm to train both the DNN and CRF together. Our extensive computational experiments demonstrate that DIFFUSE could effectively predict the functions of isoforms and genes. It achieves an average area under the receiver operating characteristics curve of 0.840 and area under the precision-recall curve of 0.581 over 4184 GO functional categories, which are significantly higher than the state-of-the-art methods. We further validate the prediction results by analyzing the correlation between functional similarity, sequence similarity, expression similarity and structural similarity, as well as the consistency between the predicted functions and some well-studied functional features of isoform sequences. AVAILABILITY AND IMPLEMENTATION: https://github.com/haochenucr/DIFFUSE. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hao Chen 0097, Dipan Shaw, Jianyang Zeng 0001, Dongbo Bu, Tao Jiang 0001
Bioinform.3
2019 ACME: pan-specific peptide-MHC class I binding prediction through attention-based deep neural networks
abstract
MOTIVATION: Prediction of peptide binding to the major histocompatibility complex (MHC) plays a vital role in the development of therapeutic vaccines for the treatment of cancer. Algorithms with improved correlations between predicted and actual binding affinities are needed to increase precision and reduce the number of false positive predictions. RESULTS: We present ACME (Attention-based Convolutional neural networks for MHC Epitope binding prediction), a new pan-specific algorithm to accurately predict the binding affinities between peptides and MHC class I molecules, even for those new alleles that are not seen in the training data. Extensive tests have demonstrated that ACME can significantly outperform other state-of-the-art prediction methods with an increase of the Pearson correlation coefficient between predicted and measured binding affinities by up to 23 percentage points. In addition, its ability to identify strong-binding peptides has been experimentally validated. Moreover, by integrating the convolutional neural network with attention mechanism, ACME is able to extract interpretable patterns that can provide useful and detailed insights into the binding preferences between peptides and their MHC partners. All these results have demonstrated that ACME can provide a powerful and practically useful tool for the studies of peptide-MHC class I interactions. AVAILABILITY AND IMPLEMENTATION: ACME is available as an open source software at https://github.com/HYsxe/ACME. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hailin Hu 0002, Fangping Wan, Yuanpeng Xiong, Dan Zhao 0004, Weiren Huang, Jianyang Zeng 0001
Bioinform.10
2019 DeepHINT: understanding HIV-1 integration via deep learning with attention
abstract
MOTIVATION: Human immunodeficiency virus type 1 (HIV-1) genome integration is closely related to clinical latency and viral rebound. In addition to human DNA sequences that directly interact with the integration machinery, the selection of HIV integration sites has also been shown to depend on the heterogeneous genomic context around a large region, which greatly hinders the prediction and mechanistic studies of HIV integration. RESULTS: We have developed an attention-based deep learning framework, named DeepHINT, to simultaneously provide accurate prediction of HIV integration sites and mechanistic explanations of the detected sites. Extensive tests on a high-density HIV integration site dataset showed that DeepHINT can outperform conventional modeling strategies by automatically learning the genomic context of HIV integration from primary DNA sequence alone or together with epigenetic information. Systematic analyses on diverse known factors of HIV integration further validated the biological relevance of the prediction results. More importantly, in-depth analyses of the attention values output by DeepHINT revealed intriguing mechanistic implications in the selection of HIV integration sites, including potential roles of several DNA-binding proteins. These results established DeepHINT as an effective and explainable deep learning framework for the prediction and mechanistic study of HIV integration. AVAILABILITY AND IMPLEMENTATION: DeepHINT is available as an open-source software and can be downloaded from https://github.com/nonnerdling/DeepHINT. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hailin Hu 0002, An Xiao, Xuanling Shi, Tao Jiang 0001, Linqi Zhang, Lei Zhang 0095, Jianyang Zeng 0001
Bioinform.9
2019 Metagenomic binning through low-density hashing
abstract
Motivation: Vastly greater quantities of microbial genome data are being generated where environmental samples mix together the DNA from many different species. Here, we present Opal for metagenomic binning, the task of identifying the origin species of DNA sequencing reads. We introduce 'low-density' locality sensitive hashing to bioinformatics, with the addition of Gallager codes for even coverage, enabling quick and accurate metagenomic binning. Results: On public benchmarks, Opal halves the error on precision/recall (F1-score) as compared with both alignment-based and alignment-free methods for species classification. We demonstrate even more marked improvement at higher taxonomic levels, allowing for the discovery of novel lineages. Furthermore, the innovation of low-density, even-coverage hashing should itself prove an essential methodological advance as it enables the application of machine learning to other bioinformatic challenges. Availability and implementation: Full source code and datasets are available at http://opal.csail.mit.edu and https://github.com/yunwilliamyu/opal. Supplementary information: Supplementary data are available at Bioinformatics online.
Yunan Luo, Yun William Yu, Jianyang Zeng 0001, Bonnie Berger, Jian Peng 0001
Bioinform.3
2019 NeoDTI: neural integration of neighbor information from a heterogeneous network for discovering new drug-target interactions
abstract
Motivation: Accurately predicting drug-target interactions (DTIs) in silico can guide the drug discovery process and thus facilitate drug development. Computational approaches for DTI prediction that adopt the systems biology perspective generally exploit the rationale that the properties of drugs and targets can be characterized by their functional roles in biological networks. Results: Inspired by recent advance of information passing and aggregation techniques that generalize the convolution neural networks to mine large-scale graph data and greatly improve the performance of many network-related prediction tasks, we develop a new nonlinear end-to-end learning model, called NeoDTI, that integrates diverse information from heterogeneous network data and automatically learns topology-preserving representations of drugs and targets to facilitate DTI prediction. The substantial prediction performance improvement over other state-of-the-art DTI prediction methods as well as several novel predicted DTIs with evidence supports from previous studies have demonstrated the superior predictive power of NeoDTI. In addition, NeoDTI is robust against a wide range of choices of hyperparameters and is ready to integrate more drug and target related information (e.g. compound-protein binding affinity data). All these results suggest that NeoDTI can offer a powerful and robust tool for drug development and drug repositioning. Availability and implementation: The source code and data used in NeoDTI are available at: https://github.com/FangpingWan/NeoDTI. Supplementary information: Supplementary data are available at Bioinformatics online.
Fangping Wan, Lixiang Hong, An Xiao, Tao Jiang 0001, Jianyang Zeng 0001
Bioinform.5
2019 DeepShape: estimating isoform-level ribosome abundance and distribution with Ribo-seq data
abstract
BACKGROUND: Ribosome profiling brings insight to the process of translation. A basic step in profile construction at transcript level is to map Ribo-seq data to transcripts, and then assign a huge number of multiple-mapped reads to similar isoforms. Existing methods either discard the multiple mapped-reads, or allocate them randomly, or assign them proportionally according to transcript abundance estimated from RNA-seq data. RESULTS: Here we present DeepShape, an RNA-seq free computational method to estimate ribosome abundance of isoforms, and simultaneously compute their ribosome profiles using a deep learning model. Our simulation results demonstrate that DeepShape can provide more accurate estimations on both ribosome abundance and profiles when compared to state-of-the-art methods. We applied DeepShape to a set of Ribo-seq data from PC3 human prostate cancer cells with and without PP242 treatment. In the four cell invasion/metastasis genes that are translationally regulated by PP242 treatment, different isoforms show very different characteristics of translational efficiency and regulation patterns. Transcript level ribosome distributions were analyzed by "Codon Residence Index (CRI)" proposed in this study to investigate the relative speed that a ribosome moves on a codon compared to its synonymous codons. We observe consistent CRI patterns in PC3 cells. We found that the translation of several codons could be regulated by PP242 treatment. CONCLUSION: In summary, we demonstrate that DeepShape can serve as a powerful tool for Ribo-seq data analysis.
Hongfei Cui, Hailin Hu 0002, Jianyang Zeng 0001, Ting Chen 0006
BMC Bioinform.3
2017 A Network Integration Approach for Drug-Target Interaction Prediction and Computational Drug Repositioning from Heterogeneous Information
Yunan Luo, Xinbin Zhao, Jingtian Zhou, Jinling Yang, Yanqing Zhang 0010, Wenhua Kuang, Jian Peng 0001, Ligong Chen, Jianyang Zeng 0001
RECOMB9
2017 ROSE: A Deep Learning Based Framework for Predicting Ribosome Stalling
Hailin Hu 0002, Jingtian Zhou, Tao Jiang 0001, Jianyang Zeng 0001
RECOMB6
2017 A deep learning framework for improving long-range residue-residue contact prediction using a hierarchical strategy
abstract
MOTIVATION: Residue-residue contacts are of great value for protein structure prediction, since contact information, especially from those long-range residue pairs, can significantly reduce the complexity of conformational sampling for protein structure prediction in practice. Despite progresses in the past decade on protein targets with abundant homologous sequences, accurate contact prediction for proteins with limited sequence information is still far from satisfaction. Methodologies for these hard targets still need further improvement. RESULTS: We presented a computational program DeepConPred, which includes a pipeline of two novel deep-learning-based methods (DeepCCon and DeepRCon) as well as a contact refinement step, to improve the prediction of long-range residue contacts from primary sequences. When compared with previous prediction approaches, our framework employed an effective scheme to identify optimal and important features for contact prediction, and was only trained with coevolutionary information derived from a limited number of homologous sequences to ensure robustness and usefulness for hard targets. Independent tests showed that 59.33%/49.97%, 64.39%/54.01% and 70.00%/59.81% of the top L/5, top L/10 and top 5 predictions were correct for CASP10/CASP11 proteins, respectively. In general, our algorithm ranked as one of the best methods for CASP targets. AVAILABILITY AND IMPLEMENTATION: All source data and codes are available at http://166.111.152.91/Downloads.html . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Dapeng Xiong, Jianyang Zeng 0001, Haipeng Gong
Bioinform.2
2017 TITER: predicting translation initiation sites by deep learning
abstract
MOTIVATION: Translation initiation is a key step in the regulation of gene expression. In addition to the annotated translation initiation sites (TISs), the translation process may also start at multiple alternative TISs (including both AUG and non-AUG codons), which makes it challenging to predict TISs and study the underlying regulatory mechanisms. Meanwhile, the advent of several high-throughput sequencing techniques for profiling initiating ribosomes at single-nucleotide resolution, e.g. GTI-seq and QTI-seq, provides abundant data for systematically studying the general principles of translation initiation and the development of computational method for TIS identification. METHODS: We have developed a deep learning-based framework, named TITER, for accurately predicting TISs on a genome-wide scale based on QTI-seq data. TITER extracts the sequence features of translation initiation from the surrounding sequence contexts of TISs using a hybrid neural network and further integrates the prior preference of TIS codon composition into a unified prediction framework. RESULTS: Extensive tests demonstrated that TITER can greatly outperform the state-of-the-art prediction methods in identifying TISs. In addition, TITER was able to identify important sequence signatures for individual types of TIS codons, including a Kozak-sequence-like motif for AUG start codon. Furthermore, the TITER prediction score can be related to the strength of translation initiation in various biological scenarios, including the repressive effect of the upstream open reading frames on gene expression and the mutational effects influencing translation initiation efficiency. AVAILABILITY AND IMPLEMENTATION: TITER is available as an open-source software and can be downloaded from https://github.com/zhangsaithu/titer . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hailin Hu 0002, Tao Jiang 0001, Lei Zhang 0095, Jianyang Zeng 0001
Bioinform.5
2016 Low-Density Locality-Sensitive Hashing Boosts Metagenomic Binning
Yunan Luo, Jianyang Zeng 0001, Bonnie Berger, Jian Peng 0001
RECOMB2
2015 Massively Parallel A* Search on a GPU
abstract
A* search is a fundamental topic in artificial intelligence. Recently, the general purpose computation on graphics processing units (GPGPU) has been widely used to accelerate numerous computational tasks. In this paper, we propose the first parallel variant of the A* search algorithm such that the search process of an agent can be accelerated by a single GPU processor in a massively parallel fashion. Our experiments have demonstrated that the GPU-accelerated A* search is efficient in solving multiple real-world search tasks, including combinatorial optimization problems, pathfinding and game solving. Compared to the traditional sequential CPU-based A* implementation, our GPU-based A* algorithm can achieve a significant speedup by up to 45x on large-scale search problems.
Jianyang Zeng 0001
AAAI2
2015 Constructing Structure Ensembles of Intrinsically Disordered Proteins from Chemical Shift Data
Huichao Gong, Jiangdian Wang, Haipeng Gong, Jianyang Zeng 0001
RECOMB5
2015 Computational Protein Design Using AND/OR Branch-and-Bound Search
Yuexin Wu, Jianyang Zeng 0001
RECOMB3
2015 Integrative Data Analysis of Multi-Platform Cancer Data with a Multimodal Deep Learning Approach
abstract
Identification of cancer subtypes plays an important role in revealing useful insights into disease pathogenesis and advancing personalized therapy. The recent development of high-throughput sequencing technologies has enabled the rapid collection of multi-platform genomic data (e.g., gene expression, miRNA expression, and DNA methylation) for the same set of tumor samples. Although numerous integrative clustering approaches have been developed to analyze cancer data, few of them are particularly designed to exploit both deep intrinsic statistical properties of each input modality and complex cross-modality correlations among multi-platform input data. In this paper, we propose a new machine learning model, called multimodal deep belief network (DBN), to cluster cancer patients from multi-platform observation data. In our integrative clustering framework, relationships among inherent features of each single modality are first encoded into multiple layers of hidden variables, and then a joint latent model is employed to fuse common features derived from multiple input modalities. A practical learning algorithm, called contrastive divergence (CD), is applied to infer the parameters of our multimodal DBN model in an unsupervised manner. Tests on two available cancer datasets show that our integrative data analysis approach can effectively extract a unified representation of latent features to capture both intra- and cross-modality correlations, and identify meaningful disease subtypes from multi-platform cancer data. In addition, our approach can identify key genes and miRNAs that may play distinct roles in the pathogenesis of different cancer subtypes. Among those key miRNAs, we found that the expression level of miR-29a is highly correlated with survival time in ovarian cancer patients. These results indicate that our multimodal DBN based data analysis approach may have practical applications in cancer pathogenesis studies and provide useful guidelines for personalized cancer therapy.
Muxuan Liang, Ting Chen 0006, Jianyang Zeng 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2014 An efficient parallel algorithm for accelerating computational protein design
abstract
MOTIVATION: Structure-based computational protein design (SCPR) is an important topic in protein engineering. Under the assumption of a rigid backbone and a finite set of discrete conformations of side-chains, various methods have been proposed to address this problem. A popular method is to combine the dead-end elimination (DEE) and A* tree search algorithms, which provably finds the global minimum energy conformation (GMEC) solution. RESULTS: In this article, we improve the efficiency of computing A* heuristic functions for protein design and propose a variant of A* algorithm in which the search process can be performed on a single GPU in a massively parallel fashion. In addition, we make some efforts to address the memory exceeding problem in A* search. As a result, our enhancements can achieve a significant speedup of the A*-based protein design algorithm by four orders of magnitude on large-scale test data through pre-computation and parallelization, while still maintaining an acceptable memory overhead. We also show that our parallel A* search algorithm could be successfully combined with iMinDEE, a state-of-the-art DEE criterion, for rotamer pruning to further improve SCPR with the consideration of continuous side-chain flexibility. AVAILABILITY: Our software is available and distributed open-source under the GNU Lesser General License Version 2.1 (GNU, February 1999). The source code can be downloaded from http://www.cs.duke.edu/donaldlab/osprey.php or http://iiis.tsinghua.edu.cn/∼compbio/software.html.
Wei Xu 0005, Bruce Randall Donald, Jianyang Zeng 0001
Bioinform.4
2013 Predicting drug-target interactions using restricted Boltzmann machines
abstract
MOTIVATION: In silico prediction of drug-target interactions plays an important role toward identifying and developing new uses of existing or abandoned drugs. Network-based approaches have recently become a popular tool for discovering new drug-target interactions (DTIs). Unfortunately, most of these network-based approaches can only predict binary interactions between drugs and targets, and information about different types of interactions has not been well exploited for DTI prediction in previous studies. On the other hand, incorporating additional information about drug-target relationships or drug modes of action can improve prediction of DTIs. Furthermore, the predicted types of DTIs can broaden our understanding about the molecular basis of drug action. RESULTS: We propose a first machine learning approach to integrate multiple types of DTIs and predict unknown drug-target relationships or drug modes of action. We cast the new DTI prediction problem into a two-layer graphical model, called restricted Boltzmann machine, and apply a practical learning algorithm to train our model and make predictions. Tests on two public databases show that our restricted Boltzmann machine model can effectively capture the latent features of a DTI network and achieve excellent performance on predicting different types of DTIs, with the area under precision-recall curve up to 89.6. In addition, we demonstrate that integrating multiple types of DTIs can significantly outperform other predictions either by simply mixing multiple types of interactions without distinction or using only a single interaction type. Further tests show that our approach can infer a high fraction of novel DTIs that has been validated by known experiments in the literature or other databases. These results indicate that our approach can have highly practical relevance to DTI prediction and drug repositioning, and hence advance the drug discovery process. AVAILABILITY: Software and datasets are available on request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jianyang Zeng 0001
Bioinform.2
2011 Protein Loop Closure Using Orientational Restraints from NMR Data
Chittaranjan Tripathy, Jianyang Zeng 0001, Pei Zhou 0001, Bruce Randall Donald
RECOMB2
2011 A Bayesian Approach for Determining Protein Side-Chain Rotamer Conformations Using Unassigned NOE Data
Jianyang Zeng 0001, Kyle E. Roberts, Pei Zhou 0001, Bruce Randall Donald
RECOMB1
2010 A Markov Random Field Framework for Protein Side-Chain Resonance Assignment
Jianyang Zeng 0001, Pei Zhou 0001, Bruce Randall Donald
RECOMB1
2005 ReCord: A Distributed Hash Table with Recursive Structure
abstract
We propose a simple distributed hash table called ReCord, which is a generalized version of Randomized- Chord and offers improved tradeoffs in performance and topology maintenance over existing P2P systems. ReCord is scalable and can be easily implemented as an overlay network, and offers a good tradeoff between the node degree and query latency. For instance, an n-node ReCord with O(log n) node degree has an expected latency of \theta (\log n) hops. Alternatively, it can also offer \theta (\frac{{\log n}}{{\log \log n}}) hops latency at a higher cost of o(\frac{{\log ^2 n}}{{\log \log n}}) node degree. Meanwhile, simulations of the dynamic behaviors of ReCord are studied.
Jianyang Zeng 0001, Wen-Jing Hsu
PDCAT1
2005 Optimal Routing in a Small-World Network
abstract
Recently a bulk of research [14, 5, 15, 9] has been done on the modelling of the small-world phenomenon, which has been shown to be pervasive in social and nature networks, and engineering systems [16, 1, 2, 11]. In order to examine the navigating aspects of small-world graphs, Kleinberg [9] proposes a network model based on a d-dimensional torus lattice with long-range links chosen at random according to the d-harmonic distribution. Kleinberg shows that the greedy routing algorithm, by using only local information, performs in O(lg^2 n) expected number of hops.We extend Kleinberg’s small-world model in that each node x has two more random links to nodes chosen uniformly and randomly within (lg n)\frac{2}{a} Manhattan distance from x, where d denotes the dimension of the model. Based on this extended model, we then propose an oblivious algorithm that can route messages between any two nodes in O(lg n) expected number of hops, which is an optimal expected bound for routing. Our routing algorithm keeps only O((\lg n)^{\beta + 1} ) bits of information on each node, where 1 \le \beta \le 2, thus being scalable with the network size. To our knowledge, our result is the first to achieve the optimal routing complexity while stillkeeping a poly-logarithmic number of bits of information stored on each node in the small-world networks. Our results may be applied to the design of the logical overlay structure of large-scale distributed systems, such as peer-to-peer networks, in the same spirit as Symphony [11].
Jianyang Zeng 0001, Wen-Jing Hsu
PDCAT1
2003 Conflict-free routing of AGVs on the mesh topology based on a discrete-time model
abstract
Automated guided vehicles (or AGVs) have become an important option in material handling. In many applications, such as container terminals, the service area is often arranged into rectangular blocks, which leads to a mesh-like path topology. Therefore, developing efficient algorithms for AGV routing on the mesh topology has become an important research topic. In this paper, we present a discrete time model, based on which a simple routing algorithm on the mesh topology is presented. The algorithm works by carefully choosing suitable parameters such that the vehicles using a same junction will arrive at different points in time, and hence no conflicts will occur during the routing; meanwhile, high routing performance can be achieved. Analyses of the task completion time and the requirements on timing control during the AGV routing are also presented.
Jianyang Zeng 0001, Wen-Jing Hsu
ICRA1