EDBT 2026 Demo / reviewers in the wild / expert
Tal Pupko
dblp:70/3084
· DBLP profile ↗
25ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0001-9463-2575ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient algorithms for simulating sequences along a phylogenetic treeabstractMOTIVATION: Sequence simulations along phylogenetic trees play an important role in numerous molecular evolution studies such as benchmarking algorithms for ancestral sequence reconstruction, multiple sequence alignment, and phylogeny inference. They are also used in phylogenetic model-selection tasks, including the inference of selective forces. Recently, Approximate Bayesian Computation (ABC)-based approaches have been developed for inferring parameters of complex evolutionary models, which rely on massive generation of simulated data. For all these applications, computationally efficient sequence simulators are essential. RESULTS: In this study, we investigate fast algorithms for simulating sequences along a phylogenetic tree, focusing on accelerating the speed-limiting component of the simulation process: handling insertion and deletion (indel) events. We demonstrate that data structures which efficiently store indel events along a tree can substantially accelerate the simulation process compared to a naive approach. To illustrate the utility of this efficient simulator, we integrated it into an ABC-based algorithm for inferring indel model parameters and applied it to study indel dynamics within Chiroptera. AVAILABILITY AND IMPLEMENTATION: The source code for the different simulation algorithms, alongside the data used, is available at: https://github.com/nimrodSerokTAU/evo-sim. The simulator has also been integrated into SpartaABC, a website for the inference of indel parameters, accessible at: https://spartaabc.tau.ac.il/. Elya Wygoda, Asher Moshe, Nimrod Serok, Edo Dotan, Noa Ecker, Naiel Jabareen, Omer Israeli, Itsik Pe'er, Tal Pupko |
Bioinform. | 9 |
| 2025 | BetaAlign: a deep learning approach for multiple sequence alignmentabstractMOTIVATION: Multiple sequence alignments (MSAs) are extensively used in biology, from phylogenetic reconstruction to structure and function prediction. Here, we suggest an out-of-the-box approach for the inference of MSAs, which relies on algorithms developed for processing natural languages. We show that our artificial intelligence (AI)-based methodology can be trained to align sequences by processing alignments that are generated via simulations, and thus different aligners can be easily generated for datasets with specific evolutionary dynamics attributes. We expect that natural language processing (NLP) solutions will replace or augment classic solutions for computing alignments, and more generally, challenging inference tasks in phylogenomics. RESULTS: The MSA problem is a fundamental pillar in bioinformatics, comparative genomics, and phylogenetics. Here, we characterize and improve BetaAlign, the first deep learning aligner, which substantially deviates from conventional algorithms of alignment computation. BetaAlign draws on NLP techniques and trains transformers to map a set of unaligned biological sequences to an MSA. We show that our approach is highly accurate, comparable and sometimes better than state-of-the-art alignment tools. We characterize the performance of BetaAlign and the effect of various aspects on accuracy; for example, the size of the training data, the effect of different transformer architectures, and the effect of learning on a subspace of indel-model parameters (subspace learning). We also introduce a new technique that leads to improved performance compared to our previous approach. Our findings further uncover the potential of NLP-based methods for sequence alignment, highlighting that AI-based algorithms can substantially challenge classic approaches in phylogenomics and bioinformatics. AVAILABILITY AND IMPLEMENTATION: Datasets used in this work are available on HuggingFace (Wolf et al. Transformers: state-of-the-art natural language processing. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. p.38-45. 2020) at: https://huggingface.co/dotan1111. Source code is available at: https://github.com/idotan286/SimulateAlignments. Edo Dotan, Elya Wygoda, Noa Ecker, Michael Alburquerque, Oren Avram, Yonatan Belinkov, Tal Pupko |
Bioinform. | 7 |
| 2025 | Effectidor II: a pan-genomic AI-based algorithm for the prediction of type III secretion system effectorsabstractMOTIVATION: Type III secretion systems are used by many Gram-negative bacteria to inject type 3 effectors (T3Es) directly into eukaryotic cells, promoting disease or provoking immune response. Because of these opposing evolutionary forces, T3E repertoires often vary within taxonomic groups. Identifying the full effector gene repertoire in genomes of related individuals is crucial for determining core and specialized effectors, understanding the disease dynamics, and developing appropriate management strategies against pathogens. It can also help uncover novel T3Es that have recently emerged in a population. Our previously published Effectidor web server successfully addressed the challenge of identifying T3Es in a single bacterial genome. Here, we enriched the web server with various novel capabilities, including the identification of T3Es from multiple genome sequences simultaneously. RESULTS: We present Effectidor II, a web server that relies on machine learning to predict T3E-encoding genes within bacterial pan-genomes. We demonstrate the benefit of learning based on features extracted from the entire sequences comprising the pan-genome and report a novel T3E discovered by it in Xanthomonas euroxanthea. AVAILABILITY AND IMPLEMENTATION: Effectidor II is available at: https://effectidor.tau.ac.il and the source code is available at: https://github.com/naamawagner/Effectidor. A stand-alone version of Effectidor II is available at: https://github.com/naamawagner/Effectidor/tree/StandAlone. The source code for the standalone version and the data used in this work are also provided in https://doi.org/10.5281/zenodo.15081636. Naama Wagner, Ella Baumer, Iris Lyubman, Yair Shimony, Noam Bracha, Leonor Martins, Neha Potnis, Jeff H. Chang, Doron Teper, Ralf Koebnik, Tal Pupko |
Bioinform. | 11 |
| 2024 | Effect of tokenization on transformers for biological sequencesabstractMOTIVATION: Deep-learning models are transforming biological research, including many bioinformatics and comparative genomics algorithms, such as sequence alignments, phylogenetic tree inference, and automatic classification of protein functions. Among these deep-learning algorithms, models for processing natural languages, developed in the natural language processing (NLP) community, were recently applied to biological sequences. However, biological sequences are different from natural languages, such as English, and French, in which segmentation of the text to separate words is relatively straightforward. Moreover, biological sequences are characterized by extremely long sentences, which hamper their processing by current machine-learning models, notably the transformer architecture. In NLP, one of the first processing steps is to transform the raw text to a list of tokens. Deep-learning applications to biological sequence data mostly segment proteins and DNA to single characters. In this work, we study the effect of alternative tokenization algorithms on eight different tasks in biology, from predicting the function of proteins and their stability, through nucleotide sequence alignment, to classifying proteins to specific families. RESULTS: We demonstrate that applying alternative tokenization algorithms can increase accuracy and at the same time, substantially reduce the input length compared to the trivial tokenizer in which each character is a token. Furthermore, applying these tokenization algorithms allows interpreting trained models, taking into account dependencies among positions. Finally, we trained these tokenizers on a large dataset of protein sequences containing more than 400 billion amino acids, which resulted in over a 3-fold decrease in the number of tokens. We then tested these tokenizers trained on large-scale data on the above specific tasks and showed that for some tasks it is highly beneficial to train database-specific tokenizers. Our study suggests that tokenizers are likely to be a critical component in future deep-network analysis of biological sequence data. AVAILABILITY AND IMPLEMENTATION: Code, data, and trained tokenizers are available on https://github.com/technion-cs-nlp/BiologicalTokenizers. Edo Dotan, Gal Jaschek, Tal Pupko, Yonatan Belinkov |
Bioinform. | 3 |
| 2024 | A machine-learning-based alternative to phylogenetic bootstrapabstractMOTIVATION: Currently used methods for estimating branch support in phylogenetic analyses often rely on the classic Felsenstein's bootstrap, parametric tests, or their approximations. As these branch support scores are widely used in phylogenetic analyses, having accurate, fast, and interpretable scores is of high importance. RESULTS: Here, we employed a data-driven approach to estimate branch support values with a probabilistic interpretation. To this end, we simulated thousands of realistic phylogenetic trees and the corresponding multiple sequence alignments. Each of the obtained alignments was used to infer the phylogeny using state-of-the-art phylogenetic inference software, which was then compared to the true tree. Using these extensive data, we trained machine-learning algorithms to estimate branch support values for each bipartition within the maximum-likelihood trees obtained by each software. Our results demonstrate that our model provides fast and more accurate probability-based branch support values than commonly used procedures. We demonstrate the applicability of our approach on empirical datasets. AVAILABILITY AND IMPLEMENTATION: The data supporting this work are available in the Figshare repository at https://doi.org/10.6084/m9.figshare.25050554.v1, and the underlying code is accessible via GitHub at https://github.com/noaeker/bootstrap_repo. Noa Ecker, Dorothée Huchon, Yishay Mansour, Itay Mayrose, Tal Pupko |
Bioinform. | 5 |
| 2024 | Statistical framework to determine indel-length distributionabstractMOTIVATION: Insertions and deletions (indels) of short DNA segments, along with substitutions, are the most frequent molecular evolutionary events. Indels were shown to affect numerous macro-evolutionary processes. Because indels may span multiple positions, their impact is a product of both their rate and their length distribution. An accurate inference of indel-length distribution is important for multiple evolutionary and bioinformatics applications, most notably for alignment software. Previous studies counted the number of continuous gap characters in alignments to determine the best-fitting length distribution. However, gap-counting methods are not statistically rigorous, as gap blocks are not synonymous with indels. Furthermore, such methods rely on alignments that regularly contain errors and are biased due to the assumption of alignment methods that indels lengths follow a geometric distribution. RESULTS: We aimed to determine which indel-length distribution best characterizes alignments using statistical rigorous methodologies. To this end, we reduced the alignment bias using a machine-learning algorithm and applied an Approximate Bayesian Computation methodology for model selection. Moreover, we developed a novel method to test if current indel models provide an adequate representation of the evolutionary process. We found that the best-fitting model varies among alignments, with a Zipf length distribution fitting the vast majority of them. AVAILABILITY AND IMPLEMENTATION: The data underlying this article are available in Github, at https://github.com/elyawy/SpartaSim and https://github.com/elyawy/SpartaPipeline. Elya Wygoda, Gil Loewenthal, Asher Moshe, Michael Alburquerque, Itay Mayrose, Tal Pupko |
Bioinform. | 6 |
| 2023 | Multiple sequence alignment as a sequence-to-sequence learning problem
Edo Dotan, Yonatan Belinkov, Oren Avram, Elya Wygoda, Noa Ecker, Michael Alburquerque, Omri Keren, Gil Loewenthal, Tal Pupko |
ICLR | 9 |
| 2022 | A LASSO-based approach to sample sites for phylogenetic tree searchabstractMOTIVATION: In recent years, full-genome sequences have become increasingly available and as a result many modern phylogenetic analyses are based on very long sequences, often with over 100 000 sites. Phylogenetic reconstructions of large-scale alignments are challenging for likelihood-based phylogenetic inference programs and usually require using a powerful computer cluster. Current tools for alignment trimming prior to phylogenetic analysis do not promise a significant reduction in the alignment size and are claimed to have a negative effect on the accuracy of the obtained tree. RESULTS: Here, we propose an artificial-intelligence-based approach, which provides means to select the optimal subset of sites and a formula by which one can compute the log-likelihood of the entire data based on this subset. Our approach is based on training a regularized Lasso-regression model that optimizes the log-likelihood prediction accuracy while putting a constraint on the number of sites used for the approximation. We show that computing the likelihood based on 5% of the sites already provides accurate approximation of the tree likelihood based on the entire data. Furthermore, we show that using this Lasso-based approximation during a tree search decreased running-time substantially while retaining the same tree-search performance. AVAILABILITY AND IMPLEMENTATION: The code was implemented in Python version 3.8 and is available through GitHub (https://github.com/noaeker/lasso_positions_sampling). The datasets used in this paper were retrieved from Zhou et al. (2018) as described in section 3. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Noa Ecker, Dana Azouri, Ben Bettisworth, Alexandros Stamatakis, Yishay Mansour, Itay Mayrose, Tal Pupko |
Bioinform. | 7 |
| 2022 | Effectidor: an automated machine-learning-based web server for the prediction of type-III secretion system effectorsabstractMOTIVATION: Type-III secretion systems are utilized by many Gram-negative bacteria to inject type-3 effectors (T3Es) to eukaryotic cells. These effectors manipulate host processes for the benefit of the bacteria and thus promote disease. They can also function as host-specificity determinants through their recognition as avirulence proteins that elicit immune response. Identifying the full effector repertoire within a set of bacterial genomes is of great importance to develop appropriate treatments against the associated pathogens. RESULTS: We present Effectidor, a user-friendly web server that harnesses several machine-learning techniques to predict T3Es within bacterial genomes. We compared the performance of Effectidor to other available tools for the same task on three pathogenic bacteria. Effectidor outperformed these tools in terms of classification accuracy (area under the precision-recall curve above 0.98 in all cases). AVAILABILITY AND IMPLEMENTATION: Effectidor is available at: https://effectidor.tau.ac.il, and the source code is available at: https://github.com/naamawagner/Effectidor. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Naama Wagner, Oren Avram, Dafna Gold-Binshtok, Ben Zerah, Doron Teper, Tal Pupko |
Bioinform. | 6 |
| 2021 | PASA: Proteomic analysis of serum antibodies web serverabstractMOTIVATION: A comprehensive characterization of the humoral response towards a specific antigen requires quantification of the B-cell receptor repertoire by next-generation sequencing (BCR-Seq), as well as the analysis of serum antibodies against this antigen, using proteomics. The proteomic analysis is challenging since it necessitates the mapping of antigen-specific peptides to individual B-cell clones. RESULTS: The PASA web server provides a robust computational platform for the analysis and integration of data obtained from proteomics of serum antibodies. PASA maps peptides derived from antibodies raised against a specific antigen to corresponding antibody sequences. It then analyzes and integrates proteomics and BCR-Seq data, thus providing a comprehensive characterization of the humoral response. The PASA web server is freely available at https://pasa.tau.ac.il and open to all users without a login requirement. Oren Avram, Aya Kigel, Anna Vaisman-Mentesh, Sharon Kligsberg, Shai Rosenstein, Yael Dror, Tal Pupko, Yariv Wine |
PLoS Comput. Biol. | 7 |
| 2019 | Ancestral sequence reconstruction: accounting for structural information by averaging over replacement matricesabstractMOTIVATION: Ancestral sequence reconstruction (ASR) is widely used to understand protein evolution, structure and function. Current ASR methodologies do not fully consider differences in evolutionary constraints among positions imposed by the three-dimensional (3D) structure of the protein. Here, we developed an ASR algorithm that allows different protein sites to evolve according to different mixtures of replacement matrices. We show that assigning replacement matrices to protein positions based on their solvent accessibility leads to ASR with higher log-likelihoods compared to naïve models that assume a single replacement matrix for all sites. Improved ASR log-likelihoods are also demonstrated when solvent accessibility is predicted from protein sequences rather than inferred from a known 3D structure. Finally, we show that using such structure-aware mixture models results in substantial differences in the inferred ancestral sequences. AVAILABILITY AND IMPLEMENTATION: http://fastml.tau.ac.il. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Asher Moshe, Tal Pupko |
Bioinform. | 2 |
| 2012 | Uncovering the co-evolutionary network among prokaryotic genesabstractMOTIVATION: Correlated events of gains and losses enable inference of co-evolution relations. The reconstruction of the co-evolutionary interactions network in prokaryotic species may elucidate functional associations among genes. RESULTS: We developed a novel probabilistic methodology for the detection of co-evolutionary interactions between pairs of genes. Using this method we inferred the co-evolutionary network among 4593 Clusters of Orthologous Genes (COGs). The number of co-evolutionary interactions substantially differed among COGs. Over 40% were found to co-evolve with at least one partner. We partitioned the network of co-evolutionary relations into clusters and uncovered multiple modular assemblies of genes with clearly defined functions. Finally, we measured the extent to which co-evolutionary relations coincide with other cellular relations such as genomic proximity, gene fusion propensity, co-expression, protein-protein interactions and metabolic connections. Our results show that co-evolutionary relations only partially overlap with these other types of networks. Our results suggest that the inferred co-evolutionary network in prokaryotes is highly informative towards revealing functional relations among genes, often showing signals that cannot be extracted from other network types. AVAILABILITY AND IMPLEMENTATION: Available under GPL license as open source. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ofir Cohen, Haim Ashkenazy, David Burstein, Tal Pupko |
Bioinform. | 4 |
| 2010 | GLOOME: gain loss mapping engineabstractUNLABELLED: The evolutionary analysis of presence and absence profiles (phyletic patterns) is widely used in biology. It is assumed that the observed phyletic pattern is the result of gain and loss dynamics along a phylogenetic tree. Examples of characters that are represented by phyletic patterns include restriction sites, gene families, introns and indels, to name a few. Here, we present a user-friendly web server that accurately infers branch-specific and site-specific gain and loss events. The novel inference methodology is based on a stochastic mapping approach utilizing models that reliably capture the underlying evolutionary processes. A variety of features are available including the ability to analyze the data with various evolutionary models, to infer gain and loss events using either stochastic mapping or maximum parsimony, and to estimate gain and loss rates for each character analyzed. AVAILABILITY: Freely available for use at http://gloome.tau.ac.il/. Ofir Cohen, Haim Ashkenazy, Frida Belinky, Dorothée Huchon, Tal Pupko |
Bioinform. | 5 |
| 2009 | Epitopia: a web-server for predicting B-cell epitopesabstractBACKGROUND: Detecting candidate B-cell epitopes in a protein is a basic and fundamental step in many immunological applications. Due to the impracticality of experimental approaches to systematically scan the entire protein, a computational tool that predicts the most probable epitope regions is desirable. RESULTS: The Epitopia server is a web-based tool that aims to predict immunogenic regions in either a protein three-dimensional structure or a linear sequence. Epitopia implements a machine-learning algorithm that was trained to discern antigenic features within a given protein. The Epitopia algorithm has been compared to other available epitope prediction tools and was found to have higher predictive power. A special emphasis was put on the development of a user-friendly graphical interface for displaying the results. CONCLUSION: Epitopia is a user-friendly web-server that predicts immunogenic regions for both a protein structure and a protein sequence. Its accuracy and functionality make it a highly useful tool. Epitopia is available at http://epitopia.tau.ac.il and includes extensive explanations and example predictions. Nimrod D. Rubinstein, Itay Mayrose, Eric Martz, Tal Pupko |
BMC Bioinform. | 4 |
| 2008 | Evolutionary Modeling of Rate Shifts Reveals Specificity Determinants in HIV-1 SubtypesabstractA hallmark of the human immunodeficiency virus 1 (HIV-1) is its rapid rate of evolution within and among its various subtypes. Two complementary hypotheses are suggested to explain the sequence variability among HIV-1 subtypes. The first suggests that the functional constraints at each site remain the same across all subtypes, and the differences among subtypes are a direct reflection of random substitutions, which have occurred during the time elapsed since their divergence. The alternative hypothesis suggests that the functional constraints themselves have evolved, and thus sequence differences among subtypes in some sites reflect shifts in function. To determine the contribution of each of these two alternatives to HIV-1 subtype evolution, we have developed a novel Bayesian method for testing and detecting site-specific rate shifts. The RAte Shift EstimatoR (RASER) method determines whether or not site-specific functional shifts characterize the evolution of a protein and, if so, points to the specific sites and lineages in which these shifts have most likely occurred. Applying RASER to a dataset composed of large samples of HIV-1 sequences from different group M subtypes, we reveal rampant evolutionary shifts throughout the HIV-1 proteome. Most of these rate shifts have occurred during the divergence of the major subtypes, establishing that subtype divergence occurred together with functional diversification. We report further evidence for the emergence of a new sub-subtype, characterized by abundant rate-shifting sites. When focusing on the rate-shifting sites detected, we find that many are associated with known function relating to viral life cycle and drug resistance. Finally, we discuss mechanisms of covariation of rate-shifting sites. Osnat Penn, Adi Stern, Nimrod D. Rubinstein, Julien Dutheil, Eran Bacharach, Nicolas Galtier, Tal Pupko |
PLoS Comput. Biol. | 7 |
| 2007 | Pepitope: epitope mapping from affinity-selected peptidesabstractUNLABELLED: Identifying the epitope to which an antibody binds is central for many immunological applications such as drug design and vaccine development. The Pepitope server is a web-based tool that aims at predicting discontinuous epitopes based on a set of peptides that were affinity-selected against a monoclonal antibody of interest. The server implements three different algorithms for epitope mapping: PepSurf, Mapitope, and a combination of the two. The rationale behind these algorithms is that the set of peptides mimics the genuine epitope in terms of physicochemical properties and spatial organization. When the three-dimensional (3D) structure of the antigen is known, the information in these peptides can be used to computationally infer the corresponding epitope. A user-friendly web interface and a graphical tool that allows viewing the predicted epitopes were developed. Pepitope can also be applied for inferring other types of protein-protein interactions beyond the immunological context, and as a general tool for aligning linear sequences to a 3D structure. AVAILABILITY: http://pepitope.tau.ac.il/ Itay Mayrose, Osnat Penn, Elana Erez, Nimrod D. Rubinstein, Tomer Shlomi, Natalia Tarnovitski Freund, Erez M. Bublil, Eytan Ruppin, Roded Sharan, Jonathan M. Gershoni, Eric Martz, Tal Pupko |
Bioinform. | 12 |
| 2007 | Phylogeny reconstruction: increasing the accuracy of pairwise distance estimation using Bayesian inference of evolutionary ratesabstractDistance-based methods for phylogeny reconstruction are the fastest and easiest to use, and their popularity is accordingly high. They are also the only known methods that can cope with huge datasets of thousands of sequences. These methods rely on evolutionary distance estimation and are sensitive to errors in such estimations. In this study, a novel Bayesian method for estimation of evolutionary distances is developed. The proposed method enables the use of a sophisticated evolutionary model that better accounts for among-site rate variation (ASRV), thereby improving the accuracy of distance estimation. Rate variations are estimated within a Bayesian framework by extracting information from the entire dataset of sequences, unlike standard methods that can only use one pair of sequences at a time. We compare the accuracy of a cascade of distance estimation methods, starting from commonly used methods and moving towards the more sophisticated novel method. Simulation studies show significant improvements in the accuracy of distance estimation by the novel method over the commonly used ones. We demonstrate the effect of the improved accuracy on tree reconstruction using both real and simulated protein sequence alignments. An implementation of this method is available as part of the SEMPHY package. Matan Ninio, Eyal Privman, Tal Pupko, Nir Friedman |
Bioinform. | 3 |
| 2005 | Selecton: a server for detecting evolutionary forces at a single amino-acid siteabstractUNLABELLED: We present an algorithmic tool for the identification of biologically significant amino acids in proteins of known three dimensional structure. We estimate the degree of purifying selection and positive Darwinian selection at each site and project these estimates onto the molecular surface of the protein. Thus, patches of functional residues (undergoing either positive or purifying selection), which may be discontinuous in the linear sequence, are revealed. We test for the statistical significance of the site-specific scores in order to obtain reliable and valid estimates. AVAILABILITY: The Selecton web server is available at: http://selecton.bioinfo.tau.ac.il SUPPLEMENTARY INFORMATION: More information is available at http://selecton.bioinfo.tau.ac.il/overview.html. A set of examples is available at http://selecton.bioinfo.tau.ac.il/gallery.html. Adi Doron-Faigenboim, Adi Stern, Itay Mayrose, Eran Bacharach, Tal Pupko |
Bioinform. | 5 |
| 2004 | ConSeq: the identification of functionally and structurally important residues in protein sequencesabstractMOTIVATION: ConSeq is a web server for the identification of biologically important residues in protein sequences. Functionally important residues that take part, e.g. in ligand binding and protein-protein interactions, are often evolutionarily conserved and are most likely to be solvent-accessible, whereas conserved residues within the protein core most probably have an important structural role in maintaining the protein's fold. Thus, estimated evolutionary rates, as well as relative solvent accessibility predictions, are assigned to each amino acid in the sequence; both are subsequently used to indicate residues that have potential structural or functional importance. AVAILABILITY: The ConSeq web server is available at http://conseq.bioinfo.tau.ac.il/ SUPPLEMENTARY INFORMATION: The ConSeq methodology, a description of its performance in a set of five well-documented proteins, a comparison to other methods, and the outcome of its application to a set of 111 proteins of unknown function, are presented at http://conseq.bioinfo.tau.ac.il/ under 'OVERVIEW', 'VALIDATION', 'COMPARISON' and 'PREDICTIONS', respectively. Carine Berezin, Fabian Glaser, Josef Rosenberg, Inbal Paz, Tal Pupko, Piero Fariselli, Rita Casadio, Nir Ben-Tal |
Bioinform. | 5 |
| 2004 | Incomplete Directed Perfect PhylogenyabstractPerfect phylogeny is one of the fundamental models for studying evolution. We investigate the following variant of the model: The input is a species-characters matrix. The characters are binary and directed; i.e., a species can only gain characters. The difference from standard perfect phylogeny is that for some species the states of some characters are unknown.The question is whether one can complete the missing states in a way that admits a perfect phylogeny. The problem arises in classical phylogenetic studies, when some states are missing or undetermined. Quite recently, studies that infer phylogenies using inserted repeat elements in DNA gave rise to the same problem. Extant solutions for it take time O(n 2m ) for n species and m characters. We provide a graph theoretic formulation of the problem as a graph sandwich problem, and give near-optimal $\tilde{O}(nm)$-time algorithms for the problem. We also study the problem of finding a single, general solution tree, from which any other solution can be obtained by node splitting. We provide an algorithm to construct such a tree, or determine that none exists. Itsik Pe'er, Tal Pupko, Ron Shamir, Roded Sharan |
SIAM J. Comput. | 2 |
| 2003 | ConSurf: Identification of Functional Regions in Proteins by Surface-Mapping of Phylogenetic InformationabstractAbstract Summary: We recently developed algorithmic tools for the identification of functionally important regions in proteins of known three dimensional structure by estimating the degree of conservation of the amino-acid sites among their close sequence homologues. Projecting the conservation grades onto the molecular surface of these proteins reveals patches of highly conserved (or occasionally highly variable) residues that are often of important biological function. We present a new web server, ConSurf, which automates these algorithmic tools. ConSurf may be used for high-throughput characterization of functional regions in proteins. Availability: The ConSurf web server is available at:http://consurf.tau.ac.il. Contact: [email protected]@[email protected]@[email protected]@[email protected] Supplementary Information: A set of examples is available at http://consurf.tau.ac.il under ‘GALLERY’. * To whom correspondence should be addressed. Fabian Glaser, Tal Pupko, Inbal Paz, Rachel E. Bell, Dalit Bechor-Shental, Eric Martz, Nir Ben-Tal |
Bioinform. | 2 |
| 2002 | Rate4Site: an algorithmic tool for the identification of functional regions in proteins by surface mapping of evolutionary determinants within their homologuesabstractAbstract Motivation: A number of proteins of known three-dimensional (3D) structure exist, with yet unknown function. In light of the recent progress in structure determination methodology, this number is likely to increase rapidly. A novel method is presented here: ‘Rate4Site’, which maps the rate of evolution among homologous proteins onto the molecular surface of one of the homologues whose 3D-structure is known. Functionally important regions often correspond to surface patches of slowly evolving residues. Results: Rate4Site estimates the rate of evolution of amino acid sites using the maximum likelihood (ML) principle. The ML estimate of the rates considers the topology and branch lengths of the phylogenetic tree, as well as the underlying stochastic process. To demonstrate its potency, we study the Src SH2 domain. Like previously established methods, Rate4Site detected the SH2 peptide-binding groove. Interestingly, it also detected inter-domain interactions between the SH2 domain and the rest of the Src protein that other methods failed to detect. Availability: Rate4Site can be downloaded at: http://ashtoret.tau.ac.il/ It is implemented as a web server at: bioinfo.tau.ac.il/ConSurf Contact: [email protected]@[email protected]@ashtoret.tau.ac.il Supplementary Information: Multiple sequence alignment of homologous SH2 domains, the corresponding phylogenetic tree and additional examples are available at http://ashtoret.tau.ac.il/~rebell Keywords: rate variation among sites; evolutionary conservation; protein evolution; maximum likelihood; SH2 domains. *To whom correspondence should be addressed. †These authors contributed equally. Tal Pupko, Rachel E. Bell, Itay Mayrose, Fabian Glaser, Nir Ben-Tal |
ISMB | 1 |
| 2002 | A branch-and-bound algorithm for the inference of ancestral amino-acid sequences when the replacement rate varies among sites: Application to the evolution of five gene familiesabstractMOTIVATION: We developed an algorithm to reconstruct ancestral sequences, taking into account the rate variation among sites of the protein sequences. Our algorithm maximizes the joint probability of the ancestral sequences, assuming that the rate is gamma distributed among sites. Our algorithm probably finds the global maximum. The use of 'joint' reconstruction is motivated by studies that use the sequences at all the internal nodes in a phylogenetic tree, such as, for instance, the inference of patterns of amino-acid replacement, or tracing the biochemical changes that occurred during the evolution of a given protein family. RESULTS: We give an algorithm that guarantees finding the global maximum. The efficient search method makes our method applicable to datasets with large number sequences. We analyze ancestral sequences of five gene families, exploring the effect of the amount of among-site-rate-variation, and the degree of sequence divergence on the resulting ancestral states. AVAILABILITY AND SUPPLEMENTARY INFORMATION: http://evolu3.ism.ac.jp/~tal/ CONTACT: [email protected] Tal Pupko, Itsik Pe'er, Masami Hasegawa, Dan Graur, Nir Friedman |
Bioinform. | 1 |
| 2001 | A structural EM algorithm for phylogenetic inferenceabstractA central task in the study of evolution is the reconstruction of a phylogenetic tree from sequences of current-day taxa. A well supported approach to tree reconstruction performs maximum likelihood (ML) analysis. Unfortunately, searching for the maximum likelihood phylogenetic tree is computationally expensive. In this paper, we describe a new algorithm that uses Structural-EM for learning maximum likelihood trees. This algorithm is similar to the standard EM method for estimating branch lengths, except that during iterations of this algorithms the topology is improved as well as the branch length. The algorithm performs iterations of two steps. In the E-Step, we use the current tree topology and branch lengths to compute expected sufficient statistics, which summarize the data. In the M-Step, we search for a topology that maximizes the likelihood with respect to these expected sufficient statistics. As we show, searching for better topologies inside the M-step can be done efficiently, as opposed to standard search over topologies. We prove that each iteration of this procedure increases the likelihood of the topology, and thus the procedure must converge. We evaluate our new algorithm on both synthetic and real sequence data, and show that it is both dramatically faster and finds more plausible trees than standard search for maximum likelihood phylogenies. Nir Friedman, Matan Ninio, Itsik Pe'er, Tal Pupko |
RECOMB | 4 |
| 2001 | A Chemical-Distance-Based Test for Positive Darwinian Selection
Tal Pupko, Roded Sharan, Masami Hasegawa, Ron Shamir, Dan Graur |
WABI | 1 |