Tamir Tuller

dblp:79/6712 · DBLP profile ↗
← Back
49ranked-venue papers
5as first author
4since 2021 · last 2025
0000-0003-4194-7068ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 49 · 5 first-author · 4 since 2021
YearPublicationVenuePosition
2025 The bioinformatics of the finding that the hepatitis delta virus RNA editing mechanism by a conformational switch exists in genotype 7 in addition to genotype 3
abstract
Hepatitis delta virus (HDV) is geographically classified according to eight known genotypes. The combined hepatitis B-hepatitis D (HEPB-HEPD) disease is the severest form of chronic viral hepatitis in humans and is characterized by mortality rates of ~20%. Hepatitis delta virus has no FDA approved therapy and its only available vaccine is the one for HEPB. Because it is the smallest RNA virus known to infect humans, RNA folding predictions by energy minimization of the whole genome can reveal important information on functional RNA secondary structure elements within the genome. A public HDV database (HDVdb) contains 512 HDV strains on which various bioinformatics methods can be applied, aiming to detect strains that could perform RNA editing via conformational switching. Up to date, only one such strain from HDVdb was known to perform that, in HDV genotype 3. Our goal was to locate more such strains, both in genotype 3 and in other possible HDV genotypes. In past work, by an eigenvalue mathematical analysis, we made an initial prediction that this peculiar RNA editing mechanism also exists in HDV genotype 7. We hereby extend our earlier findings and present newly discovered HDV strains from multiple genotypes for further analysis of RNA editing sites within the virus. The relevant strains taken from HDVdb are from both genotype 3 of Peru and genotype 7 of Cameroon. Additionally, the new strains have a variety of optional RNA editing sites that we report, many of which are unknown to date.
Rami Zakh, Alexander Churkin, Marina Parr, Tamir Tuller, Ohad Etzion, Harel Dahari, Danny Barash
Briefings Bioinform.4
2024 Strong association between genomic 3D structure and CRISPR cleavage efficiency
abstract
CRISPR is a gene editing technology which enables precise in-vivo genome editing; but its potential is hampered by its relatively low specificity and sensitivity. Improving CRISPR's on-target and off-target effects requires a better understanding of its mechanism and determinants. Here we demonstrate, for the first time, the chromosomal 3D spatial structure's association with CRISPR's cleavage efficiency, and its predictive capabilities. We used high-resolution Hi-C data to estimate the 3D distance between different regions in the human genome and utilized these spatial properties to generate 3D-based features, characterizing each region's density. We evaluated these features based on empirical, in-vivo CRISPR efficiency data and compared them to 425 features used in state-of-the-art models. The 3D features ranked in the top 13% of the features, and significantly improved the predictive power of LASSO and xgboost models trained with these features. The features indicated that sites with lower spatial density demonstrated higher efficiency. Understanding how CRISPR is affected by the 3D DNA structure provides insight into CRISPR's mechanism in general and improves our ability to correctly predict CRISPR's cleavage as well as design sgRNAs for therapeutic and scientific use.
Shaked Bergman, Tamir Tuller
PLoS Comput. Biol.2
2024 Computational Analysis of MDR1 Variants Predicts Effect on Cancer Cells via their Effect on mRNA Folding
abstract
The P-glycoprotein efflux pump, encoded by the MDR1 gene, is an ATP-driven transporter capable of expelling a diverse array of compounds from cells. Overexpression of this protein is implicated in the multi-drug resistant phenotype observed in various cancers. Numerous studies have attempted to decipher the impact of genetic variants within MDR1 on P-glycoprotein expression, functional activity, and clinical outcomes in cancer patients. Among these, three specific single nucleotide polymorphisms-T1236C, T2677G, and T3435C - have been the focus of extensive research efforts, primarily through in vitro cell line models and clinical cohort analyses. However, the findings from these studies have been remarkably contradictory. In this study, we employ a computational, data-driven approach to systematically evaluate the effects of these three variants on principal stages of the gene expression process. Leveraging current knowledge of gene regulatory mechanisms, we elucidate potential mechanisms by which these variants could modulate P-glycoprotein levels and function. Our findings suggest that all three variants significantly change the mRNA folding in their vicinity. This change in mRNA structure is predicted to increase local translation elongation rates, but not to change the protein expression. Nonetheless, the increased translation rate near T3435C is predicted to affect the protein's co-translational folding trajectory in the region of the second ATP binding domain. This potentially impacts P-glycoprotein conformation and function. Our study demonstrates the value of computational approaches in elucidating the functional consequences of genetic variants. This framework provides new insights into the molecular mechanisms of MDR1 variants and their potential impact on cancer prognosis and treatment resistance. Furthermore, we introduce an approach which can be systematically applied to identify mutations potentially affecting mRNA folding in pathology. We demonstrate the utility of this approach on both ClinVar and TCGA and identify hundreds of disease related variants that modify mRNA folding at essential positions.
Tal Gutman, Tamir Tuller
PLoS Comput. Biol.2
2021 New computational model for miRNA-mediated repression reveals novel regulatory roles of miRNA bindings inside the coding region
abstract
MOTIVATION: MicroRNAs (miRNAs) are short (∼24nt), non-coding RNAs, which downregulate gene expression in many species and physiological processes. Many details regarding the mechanism which governs miRNA-mediated repression continue to elude researchers. RESULTS: We elucidate the interplay between the coding sequence and the 3'UTR, by using elastic net regularization and incorporating translation-related features to predict miRNA-mediated repression. We find that miRNA binding sites at the end of the coding sequence contribute to repression, and that weak binding sites are linked to effective de-repression, possibly as a result of competing with stronger binding sites. Furthermore, we propose a recycling model for miRNAs dissociated from the open reading frame (ORF) by traversing ribosomes, explaining the observed link between increased ribosome density/traversal speed and increased repression. We uncover a novel layer of interaction between the coding sequence and the 3'UTR (untranslated region) and suggest the ORF has a larger role than previously thought in the mechanism of miRNA-mediated repression. AVAILABILITY AND IMPLEMENTATION: The code is freely available at https://github.com/aescrdni/miRNA_model. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Shaked Bergman, Alon Diament, Tamir Tuller
Bioinform.3
2020 CSN: unsupervised approach for inferring biological networks based on the genome alone
abstract
BACKGROUND: Most organisms cannot be cultivated, as they live in unique ecological conditions that cannot be mimicked in the lab. Understanding the functionality of those organisms' genes and their interactions by performing large-scale measurements of transcription levels, protein-protein interactions or metabolism, is extremely difficult and, in some cases, impossible. Thus, efficient algorithms for deciphering genome functionality based only on the genomic sequences with no other experimental measurements are needed. RESULTS: In this study, we describe a novel algorithm that infers gene networks that we name Common Substring Network (CSN). The algorithm enables inferring novel regulatory relations among genes based only on the genomic sequence of a given organism and partial homolog/ortholog-based functional annotation. It can specifically infer the functional annotation of genes with unknown homology. This approach is based on the assumption that related genes, not necessarily homologs, tend to share sub-sequences, which may be related to common regulatory mechanisms, similar functionality of encoded proteins, common evolutionary history, and more. We demonstrate that CSNs, which are based on S. cerevisiae and E. coli genomes, have properties similar to 'traditional' biological networks inferred from experiments. Highly expressed genes tend to have higher degree nodes in the CSN, genes with similar protein functionality tend to be closer, and the CSN graph exhibits a power-law degree distribution. Also, we show how the CSN can be used for predicting gene interactions and functions. CONCLUSIONS: The reported results suggest that 'silent' code inside the transcript can help to predict central features of biological networks and gene function. This approach can help researchers to understand the genome of novel microorganisms, analyze metagenomic data, and can help to decipher new gene functions. AVAILABILITY: Our MATLAB implementation of CSN is available at https://www.cs.tau.ac.il/~tamirtul/CSN-Autogen.
Maya Galili, Tamir Tuller
BMC Bioinform.2
2020 Whole cell biophysical modeling of codon-tRNA competition reveals novel insights related to translation dynamics
abstract
The importance of mRNA translation models has been demonstrated across many fields of science and biotechnology. However, a whole cell model with codon resolution and biophysical dynamics is still lacking. We describe a whole cell model of translation for E. coli. The model simulates all major translation components in the cell: ribosomes, mRNAs and tRNAs. It also includes, for the first time, fundamental aspects of translation, such as competition for ribosomes and tRNAs at a codon resolution while considering tRNAs wobble interactions and tRNA recycling. The model uses parameters that are tightly inferred from large scale measurements of translation. Furthermore, we demonstrate a robust modelling approach which relies on state-of-the-art practices of translation modelling and also provides a framework for easy generalizations. This novel approach allows simulation of thousands of mRNAs that undergo translation in the same cell with common resources such as ribosomes and tRNAs in feasible time. Based on this model, we demonstrate, for the first time, the direct importance of competition for resources on translation and its accurate modelling. An effective supply-demand ratio (ESDR) measure, which is related to translation factors such as tRNAs, has been devised and utilized to show superior predictive power in complex scenarios of heterologous gene expression. The devised model is not only more accurate than the existing models, but, more importantly, provides a framework for analyzing complex whole cell translation problems and variables that haven't been explored before, making it important in various biomedical fields.
Doron Levin, Tamir Tuller
PLoS Comput. Biol.2
2019 ChimeraUGEM: unsupervised gene expression modeling in any given organism
abstract
MOTIVATION: Regulation of the amount of protein that is synthesized from genes has proved to be a serious challenge in terms of analysis and prediction, and in terms of engineering and optimization, due to the large diversity in expression machinery across species. RESULTS: To address this challenge, we developed a methodology and a software tool (ChimeraUGEM) for predicting gene expression as well as adapting the coding sequence of a target gene to any host organism. We demonstrate these methods by predicting protein levels in seven organisms, in seven human tissues, and by increasing in vivo the expression of a synthetic gene up to 26-fold in the single-cell green alga Chlamydomonas reinhardtii. The underlying model is designed to capture sequence patterns and regulatory signals with minimal prior knowledge on the host organism and can be applied to a multitude of species and applications. AVAILABILITY AND IMPLEMENTATION: Source code (MATLAB, C) and binaries are freely available for download for non-commercial use at http://www.cs.tau.ac.il/~tamirtul/ChimeraUGEM/, and supported on macOS, Linux and Windows. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alon Diament, Iddo Weiner, Noam Shahar, Shira Landman, Yael Feldman, Shimshi Atar, Meital Avitan, Shira Schweitzer, Iftach Yacoby, Tamir Tuller
Bioinform.10
2019 The COP9 signalosome influences the epigenetic landscape of Arabidopsis thaliana
abstract
MOTIVATION: The COP9 signalosome is a highly conserved multi-protein complex consisting of eight subunits, which influences key developmental pathways through its regulation of protein stability and transcription. In Arabidopsis thaliana, mutations in the COP9 signalosome exhibit a number of diverse pleiotropic phenotypes. Total or partial loss of COP9 signalosome function in Arabidopsis leads to misregulation of a number of genes involved in DNA methylation, suggesting that part of the pleiotropic phenotype is due to global effects on DNA methylation. RESULTS: We determined and analyzed the methylomes and transcriptomes of both partial- and total-loss-of-function Arabidopsis mutants of the COP9 signalosome. Our results support the hypothesis that the COP9 signalosome has a global genome-wide effect on methylation and that this effect is at least partially encoded in the DNA. Our analyses suggest that COP9 signalosome-dependent methylation is related to gene expression regulation in various ways. Differentially methylated regions tend to be closer in the 3D conformation of the genome to differentially expressed genes. These results suggest that the COP9 signalosome has a more comprehensive effect on gene expression than thought before, and this is partially related to regulation of methylation. The high level of COP9 signalosome conservation among eukaryotes may also suggest that COP9 signalosome regulates methylation not only in plants but also in other eukaryotes, including humans. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tamir Tuller, Alon Diament, Avital Yahalom, Assaf Zemach, Shimshi Atar, Daniel A. Chamovitz
Bioinform.1
2019 Correction: Computational analysis of the oscillatory behavior at the translation level induced by mRNA levels oscillations due to finite intracellular resources
abstract
[This corrects the article DOI: 10.1371/journal.pcbi.1006055.].
Yoram Zarai, Tamir Tuller
PLoS Comput. Biol.2
2018 Universal evolutionary selection for high dimensional silent patterns of information hidden in the redundancy of viral genetic code
abstract
Motivation: Understanding how viruses co-evolve with their hosts and adapt various genomic level strategies in order to ensure their fitness may have essential implications in unveiling the secrets of viral evolution, and in developing new vaccines and therapeutic approaches. Here, based on a novel genomic analysis of 2625 different viruses and 439 corresponding host organisms, we provide evidence of universal evolutionary selection for high dimensional 'silent' patterns of information hidden in the redundancy of viral genetic code. Results: Our model suggests that long substrings of nucleotides in the coding regions of viruses from all classes, often also repeat in the corresponding viral hosts from all domains of life. Selection for these substrings cannot be explained only by such phenomena as codon usage bias, horizontal gene transfer and the encoded proteins. Genes encoding structural proteins responsible for building the core of the viral particles were found to include more host-repeating substrings, and these substrings tend to appear in the middle parts of the viral coding regions. In addition, in human viruses these substrings tend to be enriched with motives related to transcription factors and RNA binding proteins. The host-repeating substrings are possibly related to the evolutionary pressure on the viruses to effectively interact with host's intracellular factors and to efficiently escape from the host's immune system. Supplementary information: Supplementary data are available at Bioinformatics online.
Eli Goz, Zohar Zafrir, Tamir Tuller
Bioinform.3
2018 Generation and comparative genomics of synthetic dengue viruses
abstract
BACKGROUND: Synthetic virology is an important multidisciplinary scientific field, with emerging applications in biotechnology and medicine, aiming at developing methods to generate and engineer synthetic viruses. In particular, many of the RNA viruses, including among others the Dengue and Zika, are widespread pathogens of significant importance to human health. The ability to design and synthesize such viruses may contribute to exploring novel approaches for developing vaccines and virus based therapies. RESULTS: Here we develop a full multidisciplinary pipeline for generation and analysis of synthetic RNA viruses and specifically apply it to Dengue virus serotype 2 (DENV-2). The major steps of the pipeline include comparative genomics of endogenous and synthetic viral strains. Specifically, we show that although the synthetic DENV-2 viruses were found to have lower nucleotide variability, their phenotype, as reflected in the study of the AG129 mouse model morbidity, RNA levels, and neutralization antibodies, is similar or even more pathogenic in comparison to the wildtype master strain. Additionally, the highly variable positions, identified in the analyzed DENV-2 population, were found to overlap with less conserved homologous positions in Zika virus and other Dengue serotypes. These results may suggest that synthetic DENV-2 could enhance virulence if the correct sequence is selected. CONCLUSIONS: The approach reported in this study can be used to generate and analyze synthetic RNA viruses both on genotypic and on phenotypic level. It could be applied for understanding the functionality and the fitness effects of any set of mutations in viral RNA and for editing RNA viruses for various target applications.
Eli Goz, Yael Tsalenchuck, Rony Oren Benaroya, Zohar Zafrir, Shimshi Atar, Tahel Altman, Justin Julander, Tamir Tuller
BMC Bioinform.8
2018 The extent of ribosome queuing in budding yeast
abstract
Ribosome queuing is a fundamental phenomenon suggested to be related to topics such as genome evolution, synthetic biology, gene expression regulation, intracellular biophysics, and more. However, this phenomenon hasn't been quantified yet at a genomic level. Nevertheless, methodologies for studying translation (e.g. ribosome footprints) are usually calibrated to capture only single ribosome protected footprints (mRPFs) and thus limited in their ability to detect ribosome queuing. On the other hand, most of the models in the field assume and analyze a certain level of queuing. Here we present an experimental-computational approach for studying ribosome queuing based on sequencing of RNA footprints extracted from pairs of ribosomes (dRPFs) using a modified ribosome profiling protocol. We combine our approach with traditional ribosome profiling to generate a detailed profile of ribosome traffic. The data are analyzed using computational models of translation dynamics. The approach was implemented on the Saccharomyces cerevisiae transcriptome. Our data shows that ribosome queuing is more frequent than previously thought: the measured ratio of ribosomes within dRPFs to mRPFs is 0.2-0.35, suggesting that at least one to five translating ribosomes is in a traffic jam; these queued ribosomes cannot be captured by traditional methods. We found that specific regions are enriched with queued ribosomes, such as the 5'-end of ORFs, and regions upstream to mRPF peaks, among others. While queuing is related to higher density of ribosomes on the transcript (characteristic of highly translated genes), we report cases where traffic jams are relatively more severe in lowly expressed genes and possibly even selected for. In addition, our analysis demonstrates that higher adaptation of the coding region to the intracellular tRNA levels is associated with lower queuing levels. Our analysis also suggests that the Saccharomyces cerevisiae transcriptome undergoes selection for eliminating traffic jams. Thus, our proposed approach is an essential tool for high resolution analysis of ribosome traffic during mRNA translation and understanding its evolution.
Alon Diament, Anna Feldman, Elisheva Schochet, Martin Kupiec, Yoav Arava, Tamir Tuller
PLoS Comput. Biol.6
2018 Computational analysis of the oscillatory behavior at the translation level induced by mRNA levels oscillations due to finite intracellular resources
abstract
Recent studies have demonstrated how the competition for the finite pool of available gene expression factors has important effect on fundamental gene expression aspects. In this study, based on a whole-cell model simulation of translation in S. cerevisiae, we evaluate for the first time the expected effect of mRNA levels fluctuations on translation due to the finite pool of ribosomes. We show that fluctuations of a single gene or a group of genes mRNA levels induce periodic behavior in all S. cerevisiae translation factors and aspects: the ribosomal densities and the translation rates of all S. cerevisiae mRNAs oscillate. We numerically measure the oscillation amplitudes demonstrating that fluctuations of endogenous and heterologous genes can cause a significant fluctuation of up to 50% in the steady-state translation rates of the rest of the genes. Furthermore, we demonstrate by synonymous mutations that oscillating the levels of mRNAs that experience high ribosomal occupancy (e.g. ribosomal "traffic jam") induces the largest impact on the translation of the S. cerevisiae genome. The results reported here should provide novel insights and principles related to the design of synthetic gene expression circuits and related to the evolutionary constraints shaping gene expression of endogenous genes.
Yoram Zarai, Tamir Tuller
PLoS Comput. Biol.2
2018 Controllability Analysis and Control Synthesis for the Ribosome Flow Model
abstract
The ribosomal density along different parts of the coding regions of the mRNA molecule affect various fundamental intracellular phenomena including: protein production rates, global ribosome allocation and organismal fitness, ribosomal drop off, co-translational protein folding, mRNA degradation, and more. Thus, regulating translation in order to obtain a desired ribosomal profile along the mRNA molecule is an important biological problem. We study this problem using a dynamical model for mRNA translation, called the ribosome flow model (RFM). In the RFM, the mRNA molecule is modeled as chain of n sites. The n state-variables describe the ribosomal density profile along the mRNA molecule, whereas the transition rates from each site to the next are controlled by n + 1 positive constants. To study the problem of controlling the density profile, we consider some or all of the transition rates as time-varying controls. We consider the following problem: given an initial and a desired ribosomal density profile in the RFM, determine the time-varying values of the transition rates that steer the system to this density profile, if they exist. More specifically, we consider two control problems. In the first, all transition rates can be regulated and the goal is to steer the ribosomal density profile and the protein production rate from a given initial value to a desired value. In the second, a single transition rate is controlled and the goal is to steer the production rate to a desired value. In the first case, we show that the system is controllable, i.e. the control is powerful enough to steer the system to any desired value, and we provide simple closed-form expressions for constant control functions (or transition rates) that asymptotically steer the system to the desired value. In the second case, we show that we can steer the production rate to any desired value in a feasible region determined by the other constant transition rates. We discuss some of the biological implications of these results.
Yoram Zarai, Michael Margaliot, Eduardo D. Sontag, Tamir Tuller
IEEE ACM Trans. Comput. Biol. Bioinform.4
2017 stAIcalc: tRNA adaptation index calculator based on species-specific weights
abstract
Summary: The tRNA Adaptation Index (tAI) is a tRNA-centric measure of translation efficiency which includes weights that take into account the efficiencies of the different wobble interactions. To enable the calculation of the index based on a species-specific inference of these weights, we created the stAI calc . The calculator includes optimized tAI weights for 100 species from the three domains of life along with a standalone software package that optimizes the weights for new organisms. The tAI with the optimized weights should enable performing large scale studies in disciplines such as molecular evolution, genomics, systems biology and synthetic biology. Availability and Implementation: The calculator is publicly available at http://www.cs.tau.ac.il/∼tamirtul/stAIcalc/stAIcalc.html. Contact: [email protected].
Renana Sabi, Renana Volvovitch Daniel, Tamir Tuller
Bioinform.3
2017 Unsupervised detection of regulatory gene expression information in different genomic regions enables gene expression ranking
abstract
BACKGROUND: The regulation of all gene expression steps (e.g., Transcription, RNA processing, Translation, and mRNA Degradation) is known to be primarily encoded in different parts of genes and in genomic regions in proximity to genes (e.g., promoters, untranslated regions, coding regions, introns, etc.). However, the entire gene expression codes and the genomic regions where they are encoded are still unknown. RESULTS: Here, we employ an unsupervised approach to estimate the concentration of gene expression codes in different non-coding parts of genes and transcripts, such as introns and untranslated regions, focusing on three model organisms (Escherichia coli, Saccharomyces cerevisiae, and Schizosaccharomyces pombe). Our analyses support the conjecture that regions adjacent to the beginning and end of ORFs and the beginning and end of introns tend to include higher concentration of gene expression information relatively to regions further away. In addition, we report the exact regions with elevated concentration of gene expression codes. Furthermore, we demonstrate that the concentration of these codes in different genetic regions is correlated with the expression levels of the corresponding genes, and with splicing efficiency measurements and meiotic stage gene expression measurements in S. cerevisiae. CONCLUSION: We suggest that these discoveries improve our understanding of gene expression regulation and evolution; they can also be used for developing improved models of genome/gene evolution and for engineering gene expression in various biotechnological and synthetic biology applications.
Zohar Zafrir, Tamir Tuller
BMC Bioinform.2
2015 Exploiting hidden information interleaved in the redundancy of the genetic code without prior knowledge
abstract
MOTIVATION: Dozens of studies in recent years have demonstrated that codon usage encodes various aspects related to all stages of gene expression regulation. When relevant high-quality large-scale gene expression data are available, it is possible to statistically infer and model these signals, enabling analysing and engineering gene expression. However, when these data are not available, it is impossible to infer and validate such models. RESULTS: In this current study, we suggest Chimera-an unsupervised computationally efficient approach for exploiting hidden high-dimensional information related to the way gene expression is encoded in the open reading frame (ORF), based solely on the genome of the analysed organism. One version of the approach, named Chimera Average Repetitive Substring (ChimeraARS), estimates the adaptability of an ORF to the intracellular gene expression machinery of a genome (host), by computing its tendency to include long substrings that appear in its coding sequences; the second version, named ChimeraMap, engineers the codons of a protein such that it will include long substrings of codons that appear in the host coding sequences, improving its adaptation to a new host's gene expression machinery. We demonstrate the applicability of the new approach for analysing and engineering heterologous genes and for analysing endogenous genes. Specifically, focusing on Escherichia coli, we show that it can exploit information that cannot be detected by conventional approaches (e.g. the CAI-Codon Adaptation Index), which only consider single codon distributions; for example, we report correlations of up to 0.67 for the ChimeraARS measure with heterologous gene expression, when the CAI yielded no correlation. AVAILABILITY AND IMPLEMENTATION: For non-commercial purposes, the code of the Chimera approach can be downloaded from http://www.cs.tau.ac.il/∼tamirtul/Chimera/download.htm. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hadas Zur, Tamir Tuller
Bioinform.2
2015 Improving 3D Genome Reconstructions Using Orthologous and Functional Constraints
abstract
The study of the 3D architecture of chromosomes has been advancing rapidly in recent years. While a number of methods for 3D reconstruction of genomic models based on Hi-C data were proposed, most of the analyses in the field have been performed on different 3D representation forms (such as graphs). Here, we reproduce most of the previous results on the 3D genomic organization of the eukaryote Saccharomyces cerevisiae using analysis of 3D reconstructions. We show that many of these results can be reproduced in sparse reconstructions, generated from a small fraction of the experimental data (5% of the data), and study the properties of such models. Finally, we propose for the first time a novel approach for improving the accuracy of 3D reconstructions by introducing additional predicted physical interactions to the model, based on orthologous interactions in an evolutionary-related organism and based on predicted functional interactions between genes. We demonstrate that this approach indeed leads to the reconstruction of improved models.
Alon Diament, Tamir Tuller
PLoS Comput. Biol.2
2015 Ribosome Flow Model on a Ring
abstract
The asymmetric simple exclusion process (ASEP) is an important model from statistical physics describing particles that hop randomly from one site to the next along an ordered lattice of sites, but only if the next site is empty. ASEP has been used to model and analyze numerous multiagent systems with local interactions including the flow of ribosomes along the mRNA strand. In ASEP with periodic boundary conditions a particle that hops from the last site returns to the first one. The mean field approximation of this model is referred to as the ribosome flow model on a ring (RFMR). The RFMR may be used to model both synthetic and endogenous gene expression regimes. We analyze the RFMR using the theory of monotone dynamical systems. We show that it admits a continuum of equilibrium points and that every trajectory converges to an equilibrium point. Furthermore, we show that it entrains to periodic transition rates between the sites. We describe the implications of the analysis results to understanding and engineering cyclic mRNA translation in-vitro and in-vivo.
Alon Raveh, Yoram Zarai, Michael Margaliot, Tamir Tuller
IEEE ACM Trans. Comput. Biol. Bioinform.4
2014 Maximizing Protein Translation Rate in the Ribosome Flow Model: The Homogeneous Case
abstract
Gene translation is the process in which intracellular macro-molecules, called ribosomes, decode genetic information in the mRNA chain into the corresponding proteins. Gene translation includes several steps. During the elongation step, ribosomes move along the mRNA in a sequential manner and link amino-acids together in the corresponding order to produce the proteins. The homogeneous ribosome flow model (HRFM) is a deterministic computational model for translation-elongation under the assumption of constant elongation rates along the mRNA chain. The HRFM is described by a set of n first-order nonlinear ordinary differential equations, where n represents the number of sites along the mRNA chain. The HRFM also includes two positive parameters: ribosomal initiation rate and the (constant) elongation rate. In this paper, we show that the steady-state translation rate in the HRFM is a concave function of its parameters. This means that the problem of determining the parameter values that maximize the translation rate is relatively simple. Our results may contribute to a better understanding of the mechanisms and evolution of translation-elongation. We demonstrate this by using the theoretical results to estimate the initiation rate in M. musculus embryonic stem cell. The underlying assumption is that evolution optimized the translation mechanism. For the infinite-dimensional HRFM, we derive a closed-form solution to the problem of determining the initiation and transition rates that maximize the protein translation rate. We show that these expressions provide good approximations for the optimal values in the n-dimensional HRFM already for relatively small values of n. These results may have applications for synthetic biology where an important problem is to re-engineer genomic systems in order to maximize the protein production rate.
Yoram Zarai, Michael Margaliot, Tamir Tuller
IEEE ACM Trans. Comput. Biol. Bioinform.3
2013 Transcript features alone enable accurate prediction and understanding of gene expression in S. cerevisiae
abstract
BACKGROUND: Gene expression is a central process in all living organisms. Central questions in the field are related to the way the expression levels of genes are encoded in the transcripts and affect their evolution, and the potential to predict expression levels solely by transcript features. In this study we analyze S. cerevisiae, a model organism with the most abundant relevant cellular and genomic measurements, to evaluate the accuracy in which expression levels can be predicted by different parts of the transcript. To this end, we perform various types of regression analyses based on a total of 5323 features of the transcript. The main advantage of the proposed predictors over previous ones is related to the accurate and comprehensive definitions of the relevant transcript features, which are based on biophysical knowledge of the gene transcription and translation processes, their modeling and evolution. RESULTS: Cross validation analyses of our predictors demonstrate that they achieve a correlation of 0.68/0.68/0.70/0.61/0.81 with mRNA levels, ribosomal density, protein levels, proteins per mRNA molecule (PPR), and ribosomal load (RL) respectively (all p-values <10(-140)). When we consider predictors that are based exclusively on the features related to different parts of the transcript (5'UTR, ORF, 3'UTR), the correlations with protein levels were 0.27/0.71/0.25 (all p-values <10(-5)), suggesting that the information in the UTRs is redundant, and features of the ORF alone yield similar predictions to the ones obtained based on the entire transcript. CONCLUSIONS: The reported results demonstrate that in the analyzed model organism the expression levels of a gene are encoded in the transcript. Specifically, the prediction of a large fraction of the variance of the different gene expression steps based on transcript features alone is feasible in S. cerevisiae. We report dozens of novel transcript features related to expression levels predictions, demonstrating how such analyses can aid in understanding the gene expression process and its evolution, and how such predictors can be designed for other organisms in the future.
Hadas Zur, Tamir Tuller
BMC Bioinform.2
2013 New Universal Rules of Eukaryotic Translation Initiation Fidelity
abstract
The accepted model of eukaryotic translation initiation begins with the scanning of the transcript by the pre-initiation complex from the 5'end until an ATG codon with a specific nucleotide (nt) context surrounding it is recognized (Kozak rule). According to this model, ATG codons upstream to the beginning of the ORF should affect translation. We perform for the first time, a genome-wide statistical analysis, uncovering a new, more comprehensive and quantitative, set of initiation rules for improving the cost of translation and its efficiency. Analyzing dozens of eukaryotic genomes, we find that in all frames there is a universal trend of selection for low numbers of ATG codons; specifically, 16-27 codons upstream, but also 5-11 codons downstream of the START ATG, include less ATG codons than expected. We further suggest that there is selection for anti optimal ATG contexts in the vicinity of the START ATG. Thus, the efficiency and fidelity of translation initiation is encoded in the 5'UTR as required by the scanning model, but also at the beginning of the ORF. The observed nt patterns suggest that in all the analyzed organisms the pre-initiation complex often misses the START ATG of the ORF, and may start translation from an alternative initiation start-site. Thus, to prevent the translation of undesired proteins, there is selection for nucleotide sequences with low affinity to the pre-initiation complex near the beginning of the ORF. With the new suggested rules we were able to obtain a twice higher correlation with ribosomal density and protein levels in comparison to the Kozak rule alone (e.g. for protein levels r=0.7 vs. r=0.31; p<10(-12)).
Hadas Zur, Tamir Tuller
PLoS Comput. Biol.2
2013 Explicit Expression for the Steady-State Translation Rate in the Infinite-Dimensional Homogeneous Ribosome Flow Model
abstract
Gene translation is a central stage in the intracellular process of protein synthesis. Gene translation proceeds in three major stages: initiation, elongation, and termination. During the elongation step, ribosomes (intracellular macromolecules) link amino acids together in the order specified by messenger RNA (mRNA) molecules. The homogeneous ribosome flow model (HRFM) is a mathematical model of translation-elongation under the assumption of constant elongation rate along the mRNA sequence. The HRFM includes $(n)$ first-order nonlinear ordinary differential equations, where $(n)$ represents the length of the mRNA sequence, and two positive parameters: ribosomal initiation rate and the (constant) elongation rate. Here, we analyze the HRFM when $(n)$ goes to infinity and derive a simple expression for the steady-state protein synthesis rate. We also derive bounds that show that the behavior of the HRFM for finite, and relatively small, values of $(n)$ is already in good agreement with the closed-form result in the infinite-dimensional case. For example, for $(n=15)$, the relative error is already less than 4 percent. Our results can, thus, be used in practice for analyzing the behavior of finite-dimensional HRFMs that model translation. To demonstrate this, we apply our approach to estimate the mean initiation rate in M. musculus, finding it to be around 0.17 codons per second.
Yoram Zarai, Michael Margaliot, Tamir Tuller
IEEE ACM Trans. Comput. Biol. Bioinform.3
2012 RFMapp: ribosome flow model application
abstract
UNLABELLED: The RFMapp is a graphical user interface application based on the RFM (ribosome flow model), enabling the estimation of the translation elongation rates of messenger ribonucleic acids (mRNAs) and the profile of ribosomal densities along the mRNAs, in a computationally efficient way. The RFMapp is based on the approach previously described by Reuveni et al., and unlike other traditional approaches in the field, which are mainly related to the genes' mean codon translation efficiency, the RFM additionally considers the codon order, the ribosomes' size and their order. Thus, it has been shown that RFM outperforms traditional predictors when analyzing both heterologous and endogenous genes. AVAILABILITY AND IMPLEMENTATION: Distributable cross-platform application and guideline are available for download at: http://www.cs.tau.ac.il/~tamirtul/RFM_Installers/install.htm.
Hadas Zur, Tamir Tuller
Bioinform.2
2012 Determinants of Translation Elongation Speed and Ribosomal Profiling Biases in Mouse Embryonic Stem Cells
abstract
Ribosomal profiling is a promising approach with increasing popularity for studying translation. This approach enables monitoring the ribosomal density along genes at a resolution of single nucleotides.In this study, we focused on ribosomal density profiles of mouse embryonic stem cells. Our analysis suggests, for the first time, that even in mammals such as M. musculus the elongation speed is significantly and directly affected by determinants of the coding sequence such as: 1) the adaptation of codons to the tRNA pool; 2) the local mRNA folding of the coding sequence; 3) the local charge of amino acids encoded in the codon sequence. In addition, our analyses suggest that in general, the translation velocity of ribosomes is slower at the beginning of the coding sequence and tends to increase downstream.Finally, a comparison of these data to the expected biophysical behavior of translation suggests that it suffers from some unknown biases. Specifically, the ribosomal flux measured on the experimental data increases along the coding sequence; however, according to any biophysical model of ribosomal movement lacking internal initiation sites, the flux is expected to remain constant or decrease. Thus, developing experimental and/or statistical methods for understanding, detecting and dealing with such biases is of high importance.
Alexandra Dana, Tamir Tuller
PLoS Comput. Biol.2
2012 Stability Analysis of the Ribosome Flow Model
abstract
Gene translation is a central process in all living organisms. Developing a better understanding of this complex process may have ramifications to almost every biomedical discipline. Recently, Reuveni et al. proposed a new computational model of this process called the ribosome flow model (RFM). In this study, we show that the dynamical behavior of the RFM is relatively simple. There exists a unique equilibrium point e and every trajectory converges to e. Furthermore, convergence is monotone in the sense that the distance to e can never increase. This qualitative behavior is maintained for any feasible set of parameter values, suggesting that the RFM is highly robust. Our analysis is based on a contraction principle and the theory of monotone dynamical systems. These analysis tools may prove useful in studying other properties of the RFM as well as additional intracellular biological processes.
Michael Margaliot, Tamir Tuller
IEEE ACM Trans. Comput. Biol. Bioinform.2
2012 On the Steady-State Distribution in the Homogeneous Ribosome Flow Model
abstract
A central biological process in all living organisms is gene translation. Developing a deeper understanding of this complex process may have ramifications to almost every biomedical discipline. Reuveni et al. recently proposed a new computational model of gene translation called the Ribosome Flow Model (RFM). In this paper, we consider a particular case of this model, called the Homogeneous Ribosome Flow Model (HRFM). From a biological viewpoint, this corresponds to the case where the transition rates of all the coding sequence codons are identical. This regime has been suggested recently based on experiments in mouse embryonic cells. We consider the steady-state distribution of the HRFM. We provide formulas that relate the different parameters of the model in steady state. We prove the following properties: 1) the ribosomal density profile is monotonically decreasing along the coding sequence; 2) the ribosomal density at each codon monotonically increases with the initiation rate; and 3) for a constant initiation rate, the translation rate monotonically decreases with the length of the coding sequence. In addition, we analyze the translation rate of the HRFM at the limit of very high and very low initiation rate, and provide explicit formulas for the translation rate in these two cases. We discuss the relationship between these theoretical results and biological findings on the translation process.
Michael Margaliot, Tamir Tuller
IEEE ACM Trans. Comput. Biol. Bioinform.2
2011 A Ribosome Flow Model for Analyzing Translation Elongation - (Extended Abstract)
Shlomi Reuveni, Isaac Meilijson, Martin Kupiec, Eytan Ruppin, Tamir Tuller
RECOMB5
2011 Efficient algorithms for reconstructing gene content by co-evolution
abstract
BACKGROUND: In a previous study we demonstrated that co-evolutionary information can be utilized for improving the accuracy of ancestral gene content reconstruction. To this end, we defined a new computational problem, the Ancestral Co-Evolutionary (ACE) problem, and developed algorithms for solving it. RESULTS: In the current paper we generalize our previous study in various ways. First, we describe new efficient computational approaches for solving the ACE problem. The new approaches are based on reductions to classical methods such as linear programming relaxation, quadratic programming, and min-cut. Second, we report new computational hardness results related to the ACE, including practical cases where it can be solved in polynomial time.Third, we generalize the ACE problem and demonstrate how our approach can be used for inferring parts of the genomes of non-ancestral organisms. To this end, we describe a heuristic for finding the portion of the genome ('dominant set') that can be used to reconstruct the rest of the genome with the lowest error rate. This heuristic utilizes both evolutionary information and co-evolutionary information.We implemented these algorithms on a large input of the ACE problem (95 unicellular organisms, 4,873 protein families, and 10, 576 of co-evolutionary relations), demonstrating that some of these algorithms can outperform the algorithm used in our previous study. In addition, we show that based on our approach a 'dominant set' cab be used reconstruct a major fraction of a genome (up to 79%) with relatively low error-rate (e.g. 0.11). We find that the 'dominant set' tends to include metabolic and regulatory genes, with high evolutionary rate, and low protein abundance and number of protein-protein interactions. CONCLUSIONS: The ACE problem can be efficiently extended for inferring the genomes of organisms that exist today. In addition, it may be solved in polynomial time in many practical cases. Metabolic and regulatory genes were found to be the most important groups of genes necessary for reconstructing gene content of an organism based on other related genomes.
Hadas Birin, Tamir Tuller
BMC Bioinform.2
2011 Genome-Scale Analysis of Translation Elongation with a Ribosome Flow Model
abstract
We describe the first large scale analysis of gene translation that is based on a model that takes into account the physical and dynamical nature of this process. The Ribosomal Flow Model (RFM) predicts fundamental features of the translation process, including translation rates, protein abundance levels, ribosomal densities and the relation between all these variables, better than alternative ('non-physical') approaches. In addition, we show that the RFM can be used for accurate inference of various other quantities including genes' initiation rates and translation costs. These quantities could not be inferred by previous predictors. We find that increasing the number of available ribosomes (or equivalently the initiation rate) increases the genomic translation rate and the mean ribosome density only up to a certain point, beyond which both saturate. Strikingly, assuming that the translation system is tuned to work at the pre-saturation point maximizes the predictive power of the model with respect to experimental data. This result suggests that in all organisms that were analyzed (from bacteria to Human), the global initiation rate is optimized to attain the pre-saturation point. The fact that similar results were not observed for heterologous genes indicates that this feature is under selection. Remarkably, the gap between the performance of the RFM and alternative predictors is strikingly large in the case of heterologous genes, testifying to the model's promising biotechnological value in predicting the abundance of heterologous proteins before expressing them in the desired host.
Shlomi Reuveni, Isaac Meilijson, Martin Kupiec, Eytan Ruppin, Tamir Tuller
PLoS Comput. Biol.5
2011 Co-evolution Is Incompatible with the Markov Assumption in Phylogenetics
abstract
Markov models are extensively used in the analysis of molecular evolution. A recent line of research suggests that pairs of proteins with functional and physical interactions co-evolve with each other. Here, by analyzing hundreds of orthologous sets of three fungi and their co-evolutionary relations, we demonstrate that co-evolutionary assumption may violate the Markov assumption. Our results encourage developing alternative probabilistic models for the cases of extreme co-evolution.
Tamir Tuller, Elchanan Mossel
IEEE ACM Trans. Comput. Biol. Bioinform.1
2010 Comparative classification of species and the study of pathway evolution based on the alignment of metabolic pathways
abstract
BACKGROUND: Pathways provide topical descriptions of cellular circuitry. Comparing analogous pathways reveals intricate insights into individual functional differences among species. While previous works in the field performed genomic comparisons and evolutionary studies that were based on specific genes or proteins, whole genomic sequence, or even single pathways, none of them described a genomic system level comparative analysis of metabolic pathways. In order to properly implement such an analysis one should overcome two specific challenges: how to combine the effect of many pathways under a unified framework and how to appropriately analyze co-evolution of pathways. Here we present a computational approach for solving these two challenges. First, we describe a comprehensive, scalable, information theory based computational pipeline that calculates pathway alignment information and then compiles it in a novel manner that allows further analysis. This approach can be used for building phylogenies and for pointing out specific differences that can then be analyzed in depth. Second, we describe a new approach for comparing the evolution of metabolic pathways. This approach can be used for detecting co-evolutionary relationships between metabolic pathways. RESULTS: We demonstrate the advantages of our approach by applying our pipeline to data from the MetaCyc repository (which includes a total of 205 organisms and 660 metabolic pathways). Our analysis revealed several surprising biological observations. For example, we show that the different habitats in which Archaea organisms reside are reflected by a pathway based phylogeny. In addition, we discover two striking clusters of metabolic pathways, each cluster includes pathways that have very similar evolution. CONCLUSION: We demonstrate that distance measures that are based on the topology and the content of metabolic networks are useful for studying evolution and co-evolution.
Adi Mano, Tamir Tuller, Oded Béjà, Ron Y. Pinter
BMC Bioinform.2
2010 Discovering local patterns of co - evolution: computational aspects and biological examples
abstract
BACKGROUND: Co-evolution is the process in which two (or more) sets of orthologs exhibit a similar or correlative pattern of evolution. Co-evolution is a powerful way to learn about the functional interdependencies between sets of genes and cellular functions and to predict physical interactions. More generally, it can be used for answering fundamental questions about the evolution of biological systems.Orthologs that exhibit a strong signal of co-evolution in a certain part of the evolutionary tree may show a mild signal of co-evolution in other branches of the tree. The major reasons for this phenomenon are noise in the biological input, genes that gain or lose functions, and the fact that some measures of co-evolution relate to rare events such as positive selection. Previous publications in the field dealt with the problem of finding sets of genes that co-evolved along an entire underlying phylogenetic tree, without considering the fact that often co-evolution is local. RESULTS: In this work, we describe a new set of biological problems that are related to finding patterns of local co-evolution. We discuss their computational complexity and design algorithms for solving them. These algorithms outperform other bi-clustering methods as they are designed specifically for solving the set of problems mentioned above.We use our approach to trace the co-evolution of fungal, eukaryotic, and mammalian genes at high resolution across the different parts of the corresponding phylogenetic trees. Specifically, we discover regions in the fungi tree that are enriched with positive evolution. We show that metabolic genes exhibit a remarkable level of co-evolution and different patterns of co-evolution in various biological datasets.In addition, we find that protein complexes that are related to gene expression exhibit non-homogenous levels of co-evolution across different parts of the fungi evolutionary line. In the case of mammalian evolution, signaling pathways that are related to neurotransmission exhibit a relatively higher level of co-evolution along the primate subtree. CONCLUSIONS: We show that finding local patterns of co-evolution is a computationally challenging task and we offer novel algorithms that allow us to solve this problem, thus opening a new approach for analyzing the evolution of biological systems.
Tamir Tuller, Yifat Felder, Martin Kupiec
BMC Bioinform.1
2009 Parsimony Score of Phylogenetic Networks: Hardness Results and a Linear-Time Heuristic
abstract
Phylogenies-the evolutionary histories of groups of organisms-play a major role in representing the interrelationships among biological entities. Many methods for reconstructing and studying such phylogenies have been proposed, almost all of which assume that the underlying history of a given set of species can be represented by a binary tree. Although many biological processes can be effectively modeled and summarized in this fashion, others cannot: recombination, hybrid speciation, and horizontal gene transfer result in networks of relationships rather than trees of relationships. In previous works, we formulated a maximum parsimony (MP) criterion for reconstructing and evaluating phylogenetic networks, and demonstrated its quality on biological as well as synthetic data sets. In this paper, we provide further theoretical results as well as a very fast heuristic algorithm for the MP criterion of phylogenetic networks. In particular, we provide a novel combinatorial definition of phylogenetic networks in terms of "forbidden cycles," and provide detailed hardness and hardness of approximation proofs for the "small" MP problem. We demonstrate the performance of our heuristic in terms of time and accuracy on both biological and synthetic data sets. Finally, we explain the difference between our model and a similar one formulated by Nguyen et al., and describe the implications of this difference on the hardness and approximation results.
Guohua Jin, Luay Nakhleh, Sagi Snir, Tamir Tuller
IEEE ACM Trans. Comput. Biol. Bioinform.4
2008 Editing Bayesian Networks: A New Approach for Combining Prior Knowledge and Gene Expression Measurements for Researching Diseases
abstract
We describe a new approach for combining and comparing: (1) The information that gene expression measurements represent. (2) Prior biological knowledge (that is modeled as a weighted graph). Our approach includes translating the prior biological knowledge to a Bayesian Network (BN), and searching solutions with a small edit distance from that BN but with a significant better fitness to the gene expression measurements. This method can be useful for analyzing gene expression of patients whose regulatory processes are slightly different than those that are known from the literature (e.g. for healthy subjects).We demonstrate the viability of our method by analyzing synthetic and biological examples.
Udi Rubinstein, Yifat Felder, Nana Ginzbourg, Michael Gurevich, Tamir Tuller
BIBM5
2008 Novel Phylogenetic Network Inference by Combining Maximum Likelihood and Hidden Markov Models
Sagi Snir, Tamir Tuller
WABI2
2008 Inferring horizontal transfers in the presence of rearrangements by the minimum evolution criterion
abstract
MOTIVATION: The evolution of viruses is very rapid and in addition to local point mutations (insertion, deletion, substitution) it also includes frequent recombinations, genome rearrangements and horizontal transfer of genetic materials (HGTS). Evolutionary analysis of viral sequences is therefore a complicated matter for two main reasons: First, due to HGTs and recombinations, the right model of evolution is a network and not a tree. Second, due to genome rearrangements, an alignment of the input sequences is not guaranteed. These facts encourage developing methods for inferring phylogenetic networks that do not require aligned sequences as input. RESULTS: In this work, we present the first computational approach which deals with both genome rearrangements and horizontal gene transfers and does not require a multiple alignment as input. We formalize a new set of computational problems which involve analyzing such complex models of evolution. We investigate their computational complexity, and devise algorithms for solving them. Moreover, we demonstrate the viability of our methods on several synthetic datasets as well as four biological datasets. AVAILABILITY: The code is available from the authors upon request.
Hadas Birin, Zohar Gal-Or, Isaac Elias, Tamir Tuller
Bioinform.4
2007 A New Linear-Time Heuristic Algorithm for Computing the Parsimony Score of Phylogenetic Networks: Theoretical Bounds and Empirical Performance
Guohua Jin, Luay Nakhleh, Sagi Snir, Tamir Tuller
ISBRA4
2007 Inferring Models of Rearrangements, Recombinations, and Horizontal Transfers by the Minimum Evolution Criterion
Hadas Birin, Zohar Gal-Or, Isaac Elias, Tamir Tuller
WABI4
2007 Efficient parsimony-based methods for phylogenetic network reconstruction
abstract
MOTIVATION: Phylogenies--the evolutionary histories of groups of organisms-play a major role in representing relationships among biological entities. Although many biological processes can be effectively modeled as tree-like relationships, others, such as hybrid speciation and horizontal gene transfer (HGT), result in networks, rather than trees, of relationships. Hybrid speciation is a significant evolutionary mechanism in plants, fish and other groups of species. HGT plays a major role in bacterial genome diversification and is a significant mechanism by which bacteria develop resistance to antibiotics. Maximum parsimony is one of the most commonly used criteria for phylogenetic tree inference. Roughly speaking, inference based on this criterion seeks the tree that minimizes the amount of evolution. In 1990, Jotun Hein proposed using this criterion for inferring the evolution of sequences subject to recombination. Preliminary results on small synthetic datasets. Nakhleh et al. (2005) demonstrated the criterion's application to phylogenetic network reconstruction in general and HGT detection in particular. However, the naive algorithms used by the authors are inapplicable to large datasets due to their demanding computational requirements. Further, no rigorous theoretical analysis of computing the criterion was given, nor was it tested on biological data. RESULTS: In the present work we prove that the problem of scoring the parsimony of a phylogenetic network is NP-hard and provide an improved fixed parameter tractable algorithm for it. Further, we devise efficient heuristics for parsimony-based reconstruction of phylogenetic networks. We test our methods on both synthetic and biological data (rbcL gene in bacteria) and obtain very promising results.
Guohua Jin, Luay Nakhleh, Sagi Snir, Tamir Tuller
Bioinform.4
2007 Maximum likelihood of phylogenetic networks
abstract
Bioinformatics (2006) 22(21), 2604–2611 The authors would like to apologize for errors of graph misplacement in Figures 4–6, and an error in the caption of Figure 6. The correct figures and their captions are shown below. The species tree of the 14 organisms, as reported by Tailliez et al. (2002), and the four HGT edges inferred by our heuristic for the big ML problem. Likelihood criterion = ancestral, and tree criterion = all. The improvement in the likelihood score as a function of the number of HGT edges added to the species tree (shown in Fig. 4). The improvement achieved by adding the fourth HGT edge is smaller compared to that achieved by adding the first three events, which indicates that three HGT events suffice to model the evolution of the ribosomal protein rpl12e gene for these 14 archaea. Likelihood criterion = ancestral, and tree criterion = all. The improvement in the likelihood score as a function of the number of HGT edges added to the species tree (shown in Fig. 3). The improvement achieved by adding the sixth HGT edge is smaller compared to that achieved by adding the first five events, which indicates that five HGT events suffice to model the evolution of the rbcL gene for these 15 organisms. Likelihood criterion = ancestral, and tree criterion = all.
Guohua Jin, Luay Nakhleh, Sagi Snir, Tamir Tuller
Bioinform.4
2007 Determinants of Protein Abundance and Translation Efficiency in S. cerevisiae
abstract
The translation efficiency of most Saccharomyces cerevisiae genes remains fairly constant across poor and rich growth media. This observation has led us to revisit the available data and to examine the potential utility of a protein abundance predictor in reinterpreting existing mRNA expression data. Our predictor is based on large-scale data of mRNA levels, the tRNA adaptation index, and the evolutionary rate. It attains a correlation of 0.76 with experimentally determined protein abundance levels on unseen data and successfully cross-predicts protein abundance levels in another yeast species (Schizosaccharomyces pombe). The predicted abundance levels of proteins in known S. cerevisiae complexes, and of interacting proteins, are significantly more coherent than their corresponding mRNA expression levels. Analysis of gene expression measurement experiments using the predicted protein abundance levels yields new insights that are not readily discernable when clustering the corresponding mRNA expression levels. Comparing protein abundance levels across poor and rich media, we find a general trend for homeostatic regulation where transcription and translation change in a reciprocal manner. This phenomenon is more prominent near origins of replications. Our analysis shows that in parallel to the adaptation occurring at the tRNA level via the codon bias, proteins do undergo a complementary adaptation at the amino acid level to further increase their abundance.
Tamir Tuller, Martin Kupiec, Eytan Ruppin
PLoS Comput. Biol.1
2006 Biological Networks: Comparison, Conservation, and Evolutionary Trees
Benny Chor, Tamir Tuller
RECOMB2
2006 Maximum likelihood of phylogenetic networks
abstract
MOTIVATION: Horizontal gene transfer (HGT) is believed to be ubiquitous among bacteria, and plays a major role in their genome diversification as well as their ability to develop resistance to antibiotics. In light of its evolutionary significance and implications for human health, developing accurate and efficient methods for detecting and reconstructing HGT is imperative. RESULTS: In this article we provide a new HGT-oriented likelihood framework for many problems that involve phylogeny-based HGT detection and reconstruction. Beside the formulation of various likelihood criteria, we show that most of these problems are NP-hard, and offer heuristics for efficient and accurate reconstruction of HGT under these criteria. We implemented our heuristics and used them to analyze biological as well as synthetic data. In both cases, our criteria and heuristics exhibited very good performance with respect to identifying the correct number of HGT events as well as inferring their correct location on the species tree. AVAILABILITY: Implementation of the criteria as well as heuristics and hardness proofs are available from the authors upon request. Hardness proofs can also be downloaded at http://www.cs.tau.ac.il/~tamirtul/MLNET/Supp-ML.pdf
Guohua Jin, Luay Nakhleh, Sagi Snir, Tamir Tuller
Bioinform.4
2006 Finding a maximum likelihood tree is hard
abstract
Maximum likelihood (ML) is an increasingly popular optimality criterion for selecting evolutionary trees [Felsenstein 1981]. Finding optimal ML trees appears to be a very hard computational task, but for tractable cases, ML is the method of choice. In particular, algorithms and heuristics for ML take longer to run than algorithms and heuristics for the second major character based criterion, maximum parsimony (MP). However, while MP has been known to be NP-complete for over 20 years [Foulds and Graham, 1982; Day et al. 1986], such a hardness result for ML has so far eluded researchers in the field.An important work by Tuffley and Steel [1997] proves quantitative relations between the parsimony values of given sequences and the corresponding log likelihood values. However, a direct application of their work would only give an exponential time reduction from MP to ML. Another step in this direction has recently been made by Addario-Berry et al. [2004], who proved that ancestral maximum likelihood (AML) is NP-complete. AML “lies in between” the two problems, having some properties of MP and some properties of ML. Still, the AML proof is not directly applicable to the ML problem.We resolve the question, showing that “regular” ML on phylogenetic trees is indeed intractable. Our reduction follows the vertex cover reductions for MP [Day et al. 1986] and AML [Addario-Berry et al. 2004], but its starting point is an approximation version of vertex cover, known as gap vc. The crux of our work is not the reduction, but its correctness proof. The proof goes through a series of tree modifications, while controlling the likelihood losses at each step, using the bounds of Tuffley and Steel [1997]. The proof can be viewed as correlating the value of any ML solution to an arbitrarily close approximation to vertex cover.
Benny Chor, Tamir Tuller
J. ACM2
2005 Information Theoretic Approaches to Whole Genome Phylogenies
David Burstein, Igor Ulitsky, Tamir Tuller, Benny Chor
RECOMB3
2005 Maximum Likelihood of Evolutionary Trees Is Hard
Benny Chor, Tamir Tuller
RECOMB2
2005 Time-Window Analysis of Developmental Gene Expression Data with Multiple Genetic Backgrounds
Tamir Tuller, Efrat Oron, Erez Makavy, Daniel A. Chamovitz, Benny Chor
WABI1
2004 Adding Hidden Nodes to Gene Networks (Extended Abstract)
Benny Chor, Tamir Tuller
WABI2