EDBT 2026 Demo / reviewers in the wild / expert
Christophe Dessimoz
dblp:22/229
· DBLP profile ↗
35ranked-venue papers
7as first author
7since 2021 · last 2025
0000-0002-2170-853XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 31 · 7 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Annotation matters: the effect of structural gene annotation on orthology inferenceabstractMOTIVATION: In silico gene annotation, the process of identifying the genes present in a genome, remains a challenging task. As genome assemblies rapidly increase, the corresponding gene models and repertoires often fall short in quality. Despite advances in annotation methods, a lack of community standards means that most published gene annotations result from ad hoc pipelines. As a result, only a few species have nearly complete and accurate gene models. This annotation quality is thought to affect downstream analyses, including orthology inference, often the first step of comparative genomics studies. RESULTS: We show that different annotation methods yield markedly distinct orthology inferences. We compared orthology assignments of gene models obtained by four prominent protein-coding gene model sources: the NCBI Eukaryotic Genome Annotation Pipeline, the Ensembl Gene Annotation System, the UniProt Reference Proteomes, and Augustus 3.4 (an ab initio pipeline). We observe significant discrepancies between sources, namely in the proportion of orthologous genes per genome, the completeness of Hierarchical Orthologous Groups, and the accuracy and recall of the predicted orthologs on a standard orthology benchmark. Silvia Prieto-Baños, Yannis Nevers, Adrian M. Altenhoff, Alex Warwick Vesztrocy, Christophe Dessimoz, Natasha M. Glover |
Bioinform. | 5 |
| 2022 | ISMB 2022 proceedingsabstractThis special issue of Bioinformatics serves as the proceedings of the 30th annual conference on Intelligent Systems for Molecular Biology (ISMB), which took place on July 10–14, 2022, in Madison, WI, USA. ISMB is the leading international forum for presenting new research results, disseminating methods and techniques and facilitating discussions among leading researchers, practitioners and students in the field. In addition, ISMB is the flagship conference of the International Society for Computational Biology. Due to the worldwide COVID-19 pandemic, the ISMB 2022 meeting was run as a hybrid conference, with online participants from all around the world complementing the on-site participants. The papers published in this volume were selected from 243 submitted full-length papers featuring original research. The submitted papers were thoroughly reviewed with each paper receiving 4.08 reviews on average. For the review purpose, the submitted manuscripts were assigned to one of the 11 scientific areas according to the authors’ preference and research topic, allowing for minor adjustments to avoid conflicts of interest. In addition to selecting one of the 11 areas, the authors could also designate a particular Community of Special Interest (COSI; Table 1) that would provide the best forum for the presentation of their paper. The 11 research areas covered a broad spectrum of topics (Table 2) and also included a special General Computational Biology area intended for submissions on emerging topics or for those manuscripts that did not fit well in other reviewing areas. This year we also sought papers in a new area, equity-focused research. COSI distribution of accepted ISMB 2022 proceedings papers COSI distribution of accepted ISMB 2022 proceedings papers Thematic areas of ISMB 2022 Note: The table lists the Area Chairs for each theme, the number of reviewed papers, the number of accepted papers and the acceptance rate for each area. Thematic areas of ISMB 2022 Note: The table lists the Area Chairs for each theme, the number of reviewed papers, the number of accepted papers and the acceptance rate for each area. The reviewing of the submissions was overseen by the Senior Program Committee (SPC), which included the Proceedings Chairs (this editorial’s authors) and Area Chairs (AC, listed in Table 2). Several of the ACs were nominated by COSIs or by the ISMB Steering Committee. Members of the SPC were responsible for recruiting the Program Committee. Reviewers (Program Committee Members and sub-reviewers recruited by them) judged the papers based on the novelty of computational approaches, the relevance of biological questions, the importance of biological insights, clarity of presentation, correctness and completeness of the study and expected impact. After submitting their reviews, the reviewers had the opportunity to discuss the papers and refine the scores. The ACs facilitated these discussions, oversaw the review process and made sure reviews were detailed and consistent with the overall decision. Final acceptance decisions were made by the entire SPC. Throughout the reviewing process, we followed a stringent policy of guarding conflicts of interest. Submissions that had any association with an Area Chair were reassigned to a different area. Care was taken not to assign papers to reviewers with a perceived conflict either (e.g. co-authorship or same institution). The definition of conflict was defined as broadly as possible—including collaborators (present and past few years), same institution (present, past few years and planned future moves), family relations, advisees/advisors, as well as any personal conflicts that could cause the appearance of a conflict of interest, or that could genuinely interfere with objective reviewing. Finally, as Proceedings Chairs, we refrained from submitting papers to the conference. Among the 243 submissions, 48 were accepted for presentation at ISMB 2022 and publication in the proceedings, conditioned on revisions properly addressing the comments of the reviewers. In a few cases, the authors had to be reminded to release their source code alongside their manuscript. This year, all 48 conditionally accepted papers were revised and subsequently judged to have properly addressed the concerns of the reviewers and were accepted for the conference proceedings, resulting in a 19.8% acceptance rate overall. The acceptance rates for individual areas are shown in Table 2. Accepted papers were assigned to COSIs based on the preferences of both authors and COSI organizers (Table 1). We are deeply grateful to the Area Chairs, the 304 members of the Program Committee and the 270 sub-reviewers for their outstanding efforts in conducting thorough and timely reviews. Their contribution is at the core of the scientific quality of the conference. We also thank Steven Leard, Diane Kovats and Seth Munholland for their support, guidance and handling logistical questions and all the other members of the ISMB Steering Committee for their expert advice and supervision. We also thank the team at Oxford University Press for producing this special proceedings volume. We also thank all the authors for submitting their work. These proceedings would not be possible without the scientific ingenuity of the contributors of all the papers. We recognize that, despite our best efforts, the selection process is necessarily imperfect, and some outstanding work will have been missed. Nonetheless, we hope that all authors received helpful feedback on their work. Finally, we want to thank all the keynote speakers, presenters, and all conference participants. Thank you all for making this meeting possible and the entire ISMB community to continue to thrive. No new data were generated or analysed in support of this research. This work was supported by the James McDonnell Foundation to S.R.; Swiss National Science Foundation grants (205085 and 186397 to C.D.). Conflict of Interest: none declared. Christophe Dessimoz, Sushmita Roy |
Bioinform. | 1 |
| 2022 | OMAMO: orthology-based alternative model organism selectionabstractSUMMARY: The conservation of pathways and genes across species has allowed scientists to use non-human model organisms to gain a deeper understanding of human biology. However, the use of traditional model systems such as mice, rats and zebrafish is costly, time-consuming and increasingly raises ethical concerns, which highlights the need to search for less complex model organisms. Existing tools only focus on the few well-studied model systems, most of which are complex animals. To address these issues, we have developed Orthologous Matrix and Alternative Model Organism (OMAMO), a software and a web service that provides the user with the best non-complex organism for research into a biological process of interest based on orthologous relationships between human and the species. The outputs provided by OMAMO were supported by a systematic literature review. AVAILABILITY AND IMPLEMENTATION: https://omabrowser.org/omamo/, https://github.com/DessimozLab/omamo. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alina Nicheperovich, Adrian M. Altenhoff, Christophe Dessimoz, Sina Majidian |
Bioinform. | 3 |
| 2022 | Bio-SODA UX: enabling natural language question answering over knowledge graphs with user disambiguationabstractAbstract The problem of natural language processing over structured data has become a growing research field, both within the relational database and the Semantic Web community, with significant efforts involved in question answering over knowledge graphs (KGQA). However, many of these approaches are either specifically targeted at open-domain question answering using DBpedia, or require large training datasets to translate a natural language question to SPARQL in order to query the knowledge graph. Hence, these approaches often cannot be applied directly to complex scientific datasets where no prior training data is available. In this paper, we focus on the challenges of natural language processing over knowledge graphs of scientific datasets. In particular, we introduce Bio-SODA, a natural language processing engine that does not require training data in the form of question-answer pairs for generating SPARQL queries. Bio-SODA uses a generic graph-based approach for translating user questions to a ranked list of SPARQL candidate queries. Furthermore, Bio-SODA uses a novel ranking algorithm that includes node centrality as a measure of relevance for selecting the best SPARQL candidate query. Our experiments with real-world datasets across several scientific domains, including the official bioinformatics Question Answering over Linked Data (QALD) challenge, as well as the CORDIS dataset of European projects, show that Bio-SODA outperforms publicly available KGQA systems by an F1-score of least 20% and by an even higher factor on more complex bioinformatics datasets. Finally, we introduce Bio-SODA UX, a graphical user interface designed to assist users in the exploration of large knowledge graphs and in dynamically disambiguating natural language questions that target the data available in these graphs. Ana Claudia Sima, Tarcisio M. Farias, Maria Anisimova, Christophe Dessimoz, Marc Robinson-Rechavi, Erich Zbinden, Kurt Stockinger |
Distributed Parallel Databases | 4 |
| 2021 | Bio-SODA: Enabling Natural Language Question Answering over Knowledge Graphs without Training DataabstractThe problem of natural language processing over structured data has become a growing research field, both within the relational database and the Semantic Web community, with significant efforts involved in question answering over knowledge graphs (KGQA). However, many of these approaches are either specifically targeted at open-domain question answering using DBpedia, or require large training datasets to translate a natural language question to SPARQL in order to query the knowledge graph. Hence, these approaches often cannot be applied directly to complex scientific datasets where no prior training data is available. Ana Claudia Sima, Tarcisio M. Farias, Maria Anisimova, Christophe Dessimoz, Marc Robinson-Rechavi, Erich Zbinden, Kurt Stockinger |
SSDBM | 4 |
| 2021 | ISMB/ECCB 2021 proceedingsabstractThis special issue of Bioinformatics serves as the proceedings of the biennial joint meeting of ISMB (29th annual conference on Intelligent Systems for Molecular Biology) and ECCB (20th European Conference on Computational Biology, which took place July 25–30, 2021). ISMB/ECCB is the leading international forum for presenting new research results, disseminating methods and techniques and facilitating discussions among leading researchers, practitioners and students in the field. In addition, ISMB is the flagship conference of the International Society for Computational Biology (ISCB). Due to the worldwide COVID-19 pandemic, the ISMB/ECCB 2021 meeting, initially intended to be held in Lyon, France, was for the second time run as a fully virtual conference. The organizers of the virtual meeting made the best effort to bridge time zones—enabling the global bioinformatics community to gather and fully participate in the meeting. The papers published in this volume were selected from 289 submitted full length papers featuring original research. The submitted papers where thoroughly reviewed with each paper receiving at least 3 reviews (3.9 reviews on average). For the review purpose, the submitted manuscripts were assigned to 1 of 10 scientific areas according to authors’ preference and research topic, allowing for minor adjustments to avoid conflicts of interest. In addition to selecting 1 of the 10 areas, the authors could also designate a particular Community of Special Interest (COSI; Table 1) that would provide the best forum for presentation of their paper. The 10 research areas covered a broad spectrum of topics (Table 2) and also included a special General Computational Biology area intended for submissions on emerging topics or for those manuscripts that did not fit well in other reviewing areas. Finally, the area of Bioinformatics Education made a return to this year’s ISMB/ECCB. COSI distribution of accepted ISMB/ECCB 2021 proceedings papers COSI distribution of accepted ISMB/ECCB 2021 proceedings papers Thematic areas of ISMB/ECCB 2021 Note. The table lists the Area Chairs for each theme, the number of reviewed papers, the number of accepted papers and the acceptance rate for each area. Thematic areas of ISMB/ECCB 2021 Note. The table lists the Area Chairs for each theme, the number of reviewed papers, the number of accepted papers and the acceptance rate for each area. The reviewing of the submissions was overseen by the Senior Program Committee (SPC), which included the Proceedings Chairs (this editorial’s authors) and Area Chairs (AC, listed in Table 2). Several of the ACs were nominated by COSIs or by the ISMB/ECCB Steering Committee. Members of SPC were responsible for recruiting Program Committee members. Reviewers (Program Committee members and subreviewers recruited by them) judged the papers based on the novelty of computational approaches, relevance of biological questions, importance of biological insights, clarity of presentation, correctness and completeness of the study and expected impact. After submitting their reviews, the reviewers had the opportunity to discuss the papers and refine the scores. Final acceptance decisions were made by the entire SPC. Throughout the reviewing process, we adopted a stringent policy against conflicts of interest. Submissions that had any link to an Area Chair were reassigned to a different Area. Care was taken not to assign papers to reviewers with a link either. The definition of link was explicitly defined as broad—including collaborators (present and past few years), same institution (present, past few years, planned future moves), family relations, advisees/advisors, as well as any personal conflicts that could cause the appearance of conflict of interest, or that could genuinely interfere with objective reviewing. Finally, as Proceedings Chairs, we refrained from submitting papers to the conference. Among the 289 submissions, 55 were accepted for presentation at ISMB/ECCB and publication in the Proceedings, conditioned on revisions properly addressing the comments of the reviewers. This year all 55 conditionally accepted papers were revised and subsequently judged to have properly addressed the concerns of the reviewers and were accepted for the conference proceedings, resulting in a 19% acceptance rate overall. The acceptance rates for individual areas are shown in Table 2. Accepted papers were assigned to COSIs based on the preferences of both authors and COSI organizers (Table 1). We are deeply grateful to the Area Chairs, the 514 members of the Program Committee and the 291 subreviewers for their outstanding efforts in conducting a thorough and timely review process. Their contribution is at the core of the scientific quality of the conference. We also thank Steven Leard, and Seth Munholland for their support, guidance and handling logistical questions and all the other members of the ISMB/ECCB Steering Committee for their expert advice and supervision. We also thank the team at Oxford University Press for producing these special proceedings volume. We also thank all the authors for submitting their work. These proceedings would not be possible without the scientific ingenuity of the contributors of all the papers. We recognize that, despite our best efforts, the selection process is necessarily imperfect, and some outstanding work will have been missed. Nonetheless, we hope that all authors received helpful feedback on their work. Finally, we want to thank all the keynote speakers, presenters and all conference participants. Thank you all for allowing this meeting and the whole ISMB/ECCB community to continue to thrive. T.M.P. is supported by the Intramural Research Program of the National Library of Medicine, NIH. C.D. is supported by the Swiss National Science Foundation grant 183723. Conflict of Interest: none declared. Christophe Dessimoz, Teresa M. Przytycka |
Bioinform. | 1 |
| 2021 | OMAmer: tree-driven and alignment-free protein assignment to subfamilies outperforms closest sequence approachesabstractMOTIVATION: Assigning new sequences to known protein families and subfamilies is a prerequisite for many functional, comparative and evolutionary genomics analyses. Such assignment is commonly achieved by looking for the closest sequence in a reference database, using a method such as BLAST. However, ignoring the gene phylogeny can be misleading because a query sequence does not necessarily belong to the same subfamily as its closest sequence. For example, a hemoglobin which branched out prior to the hemoglobin alpha/beta duplication could be closest to a hemoglobin alpha or beta sequence, whereas it is neither. To overcome this problem, phylogeny-driven tools have emerged but rely on gene trees, whose inference is computationally expensive. RESULTS: Here, we first show that in multiple animal and plant datasets, 18-62% of assignments by closest sequence are misassigned, typically to an over-specific subfamily. Then, we introduce OMAmer, a novel alignment-free protein subfamily assignment method, which limits over-specific subfamily assignments and is suited to phylogenomic databases with thousands of genomes. OMAmer is based on an innovative method using evolutionarily informed k-mers for alignment-free mapping to ancestral protein subfamilies. Whilst able to reject non-homologous family-level assignments, we show that OMAmer provides better and quicker subfamily-level assignments than approaches relying on the closest sequence, whether inferred exactly by Smith-Waterman or by the fast heuristic DIAMOND. AVAILABILITYAND IMPLEMENTATION: OMAmer is available from the Python Package Index (as omamer), with the source code and a precomputed database available at https://github.com/DessimozLab/omamer. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Victor Rossier, Alex Warwick Vesztrocy, Marc Robinson-Rechavi, Christophe Dessimoz |
Bioinform. | 4 |
| 2020 | Parallel and Scalable Precise ClusteringabstractThis paper describes a new technique for parallelizing protein clustering, an important bioinformatics computation for the analysis of protein sequences. Protein clustering identifies groups of proteins that are similar because they share long sequences of similar amino acids. Given a collection of protein sequences, clustering can significantly reduce the computational effort required to identify all similar sequences by avoiding many negative comparisons. The challenge, however, is to build a clustering that misses as few similar sequences (or elements, more generally) as possible. Stuart Byma, Akash Balasaheb Dhasade, Adrian M. Altenhoff, Christophe Dessimoz, James R. Larus |
PACT | 4 |
| 2020 | Benchmarking gene ontology function predictions using negative annotationsabstractMOTIVATION: With the ever-increasing number and diversity of sequenced species, the challenge to characterize genes with functional information is even more important. In most species, this characterization almost entirely relies on automated electronic methods. As such, it is critical to benchmark the various methods. The Critical Assessment of protein Function Annotation algorithms (CAFA) series of community experiments provide the most comprehensive benchmark, with a time-delayed analysis leveraging newly curated experimentally supported annotations. However, the definition of a false positive in CAFA has not fully accounted for the open world assumption (OWA), leading to a systematic underestimation of precision. The main reason for this limitation is the relative paucity of negative experimental annotations. RESULTS: This article introduces a new, OWA-compliant, benchmark based on a balanced test set of positive and negative annotations. The negative annotations are derived from expert-curated annotations of protein families on phylogenetic trees. This approach results in a large increase in the average information content of negative annotations. The benchmark has been tested using the naïve and BLAST baseline methods, as well as two orthology-based methods. This new benchmark could complement existing ones in future CAFA experiments. AVAILABILITY AND IMPLEMENTATION: All data, as well as code used for analysis, is available from https://lab.dessimoz.org/20_not. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alex Warwick Vesztrocy, Christophe Dessimoz |
Bioinform. | 2 |
| 2020 | Scalable phylogenetic profiling using MinHash uncovers likely eukaryotic sexual reproduction genesabstractPhylogenetic profiling is a computational method to predict genes involved in the same biological process by identifying protein families which tend to be jointly lost or retained across the tree of life. Phylogenetic profiling has customarily been more widely used with prokaryotes than eukaryotes, because the method is thought to require many diverse genomes. There are now many eukaryotic genomes available, but these are considerably larger, and typical phylogenetic profiling methods require at least quadratic time as a function of the number of genes. We introduce a fast, scalable phylogenetic profiling approach entitled HogProf, which leverages hierarchical orthologous groups for the construction of large profiles and locality-sensitive hashing for efficient retrieval of similar profiles. We show that the approach outperforms Enhanced Phylogenetic Tree, a phylogeny-based method, and use the tool to reconstruct networks and query for interactors of the kinetochore complex as well as conserved proteins involved in sexual reproduction: Hap2, Spo11 and Gex1. HogProf enables large-scale phylogenetic profiling across the three domains of life, and will be useful to predict biological pathways among the hundreds of thousands of eukaryotic species that will become available in the coming few years. HogProf is available at https://github.com/DessimozLab/HogProf. David Moi, Laurent Kilchoer, Pablo S. Aguilar, Christophe Dessimoz |
PLoS Comput. Biol. | 4 |
| 2019 | Phylogenetic approaches to identifying fragments of the same gene, with application to the wheat genomeabstractMOTIVATION: As the time and cost of sequencing decrease, the number of available genomes and transcriptomes rapidly increases. Yet the quality of the assemblies and the gene annotations varies considerably and often remains poor, affecting downstream analyses. This is particularly true when fragments of the same gene are annotated as distinct genes, which may cause them to be mistaken as paralogs. RESULTS: In this study, we introduce two novel phylogenetic tests to infer non-overlapping or partially overlapping genes that are in fact parts of the same gene. One approach collapses branches with low bootstrap support and the other computes a likelihood ratio test. We extensively validated these methods by (i) introducing and recovering fragmentation on the bread wheat, Triticum aestivum cv. Chinese Spring, chromosome 3B; (ii) by applying the methods to the low-quality 3B assembly and validating predictions against the high-quality 3B assembly; and (iii) by comparing the performance of the proposed methods to the performance of existing methods, namely Ensembl Compara and ESPRIT. Application of this combination to a draft shotgun assembly of the entire bread wheat genome revealed 1221 pairs of genes that are highly likely to be fragments of the same gene. Our approach demonstrates the power of fine-grained evolutionary inferences across multiple species to improving genome assemblies and annotations. AVAILABILITY AND IMPLEMENTATION: An open source software tool is available at https://github.com/DessimozLab/esprit2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ivana Pilizota, Clément-Marie Train, Adrian M. Altenhoff, Henning Redestig, Christophe Dessimoz |
Bioinform. | 5 |
| 2019 | iHam and pyHam: visualizing and processing hierarchical orthologous groupsabstractSUMMARY: The evolutionary history of gene families can be complex due to duplications and losses. This complexity is compounded by the large number of genomes simultaneously considered in contemporary comparative genomic analyses. As provided by several orthology databases, hierarchical orthologous groups (HOGs) are sets of genes that are inferred to have descended from a common ancestral gene within a species clade. This implies that the set of HOGs defined for a particular clade correspond to the ancestral genes found in its last common ancestor. Furthermore, by keeping track of HOG composition along the species tree, it is possible to infer the emergence, duplications and losses of genes within a gene family of interest. However, the lack of tools to manipulate and analyse HOGs has made it difficult to extract, display and interpret this type of information. To address this, we introduce interactive HOG analysis method, an interactive JavaScript widget to visualize and explore gene family history encoded in HOGs and python HOG analysis method, a python library for programmatic processing of genes families. These complementary open source tools greatly ease adoption of HOGs as a scalable and interpretable concept to relate genes across multiple species. AVAILABILITY AND IMPLEMENTATION: iHam's code is available at https://github.com/DessimozLab/iHam or can be loaded dynamically. pyHam's code is available at https://github.com/DessimozLab/pyHam and or via the pip package 'pyham'. Clément-Marie Train, Miguel Pignatelli, Adrian M. Altenhoff, Christophe Dessimoz |
Bioinform. | 4 |
| 2018 | RecPhyloXML: a format for reconciled gene treesabstractMotivation: A reconciliation is an annotation of the nodes of a gene tree with evolutionary events-for example, speciation, gene duplication, transfer, loss, etc.-along with a mapping onto a species tree. Many algorithms and software produce or use reconciliations but often using different reconciliation formats, regarding the type of events considered or whether the species tree is dated or not. This complicates the comparison and communication between different programs. Results: Here, we gather a consortium of software developers in gene tree species tree reconciliation to propose and endorse a format that aims to promote an integrative-albeit flexible-specification of phylogenetic reconciliations. This format, named recPhyloXML, is accompanied by several tools such as a reconciled tree visualizer and conversion utilities. Availability and implementation: http://phylariane.univ-lyon1.fr/recphyloxml/. Wandrille Duchemin, Guillaume Gence, Anne-Muriel Arigon Chifolleau, Lars Arvestad, Mukul S. Bansal, Vincent Berry, Bastien Boussau, François Chevenet, Nicolas Comte, Adrián A. Davín, Christophe Dessimoz, David Dylus, Damir Hasic, Diego Mallo, Rémi Planel, David Posada, Céline Scornavacca, Gergely J. Szöllosi, Louxin Zhang, Eric Tannier, Vincent Daubin |
Bioinform. | 11 |
| 2018 | Gearing up to handle the mosaic nature of life in the quest for orthologsabstractThe Quest for Orthologs (QfO) is an open collaboration framework for experts in comparative phylogenomics and related research areas who have an interest in highly accurate orthology predictions and their applications. We here report highlights and discussion points from the QfO meeting 2015 held in Barcelona. Achievements in recent years have established a basis to support developments for improved orthology prediction and to explore new approaches. Central to the QfO effort is proper benchmarking of methods and services, as well as design of standardized datasets and standardized formats to allow sharing and comparison of results. Simultaneously, analysis pipelines have been improved, evaluated and adapted to handle large datasets. All this would not have occurred without the long-term collaboration of Consortium members. Meeting regularly to review and coordinate complementary activities from a broad spectrum of innovative researchers clearly benefits the community. Highlights of the meeting include addressing sources of and legitimacy of disagreements between orthology calls, the context dependency of orthology definitions, special challenges encountered when analyzing very anciently rooted orthologies, orthology in the light of whole-genome duplications, and the concept of orthologous versus paralogous relationships at different levels, including domain-level orthology. Furthermore, particular needs for different applications (e.g. plant genomics, ancient gene families and others) and the infrastructure for making orthology inferences available (e.g. interfaces with model organism databases) were discussed, with several ongoing efforts that are expected to be reported on during the upcoming 2017 QfO meeting. Sofia K. Forslund, Cécile Pereira, Salvador Capella-Gutiérrez, Alan W. Sousa da Silva, Adrian M. Altenhoff, Jaime Huerta-Cepas, Matthieu Muffato, Mateus Patricio, Klaas Vandepoele, Ingo Ebersberger, Judith A. Blake, Jesualdo Tomás Fernández-Breis, Brigitte Boeckmann, Toni Gabaldón, Erik L. L. Sonnhammer, Christophe Dessimoz, Suzanna Lewis |
Bioinform. | 17 |
| 2018 | Prioritising candidate genes causing QTL using hierarchical orthologous groupsabstractMotivation: A key goal in plant biotechnology applications is the identification of genes associated to particular phenotypic traits (for example: yield, fruit size, root length). Quantitative Trait Loci (QTL) studies identify genomic regions associated with a trait of interest. However, to infer potential causal genes in these regions, each of which can contain hundreds of genes, these data are usually intersected with prior functional knowledge of the genes. This process is however laborious, particularly if the experiment is performed in a non-model species, and the statistical significance of the inferred candidates is typically unknown. Results: This paper introduces QTLSearch, a method and software tool to search for candidate causal genes in QTL studies by combining Gene Ontology annotations across many species, leveraging hierarchical orthologous groups. The usefulness of this approach is demonstrated by re-analysing two metabolic QTL studies: one in Arabidopsis thaliana, the other in Oryza sativa subsp. indica. Even after controlling for statistical significance, QTLSearch inferred potential causal genes for more QTL than BLAST-based functional propagation against UniProtKB/Swiss-Prot, and for more QTL than in the original studies. Availability and implementation: QTLSearch is distributed under the LGPLv3 license. It is available to install from the Python Package Index (as qtlsearch), with the source available from https://bitbucket.org/alex-warwickvesztrocy/qtlsearch. Supplementary information: Supplementary data are available at Bioinformatics online. Alex Warwick Vesztrocy, Christophe Dessimoz, Henning Redestig |
Bioinform. | 2 |
| 2018 | Submit a Topic Page to PLOS Computational Biology and WikipediaabstractPLOS Computational Biology launched its 'Topic Pages' project as a way to help fill important gaps in Wikipedia's coverage of computational biology content and to credit authors for their contributions.Topic Pages are written in the style of a Wikipedia article and are openly and publicly peer reviewed on the PLOS Wiki before being published in our PLOS journals, with a second, editable version posted to Wikipedia.Six years on, PLOS Computational Biology has published 11 Topic Pages covering a good range of subjects, from the Hypercycle to Approximate Bayesian Computation.The published articles have been widely viewed on Wikipedia as well as in the journal and well received by the community.We are welcoming submissions for further PLOS Computational Biology Topic Pages.We are looking for topics in computational biology that are of interest to our readership, the broader scientific community, and the public at large and that are not yet covered or insufficiently covered (i.e., exist as a 'stub') in Wikipedia.Last year, PLOS Genetics joined the Topic Pages initiative, as detailed in this blog post.We are also exploring how the Topic Pages approach could be extended to include Wikidata, the community-curated database connecting concepts covered in any Wikipedia article with the Semantic Web [1].For instance, data from more and more research-related databases are being integrated with Wikidata or its semantic core, Wikibase.This creates the need to formalize data models: How should concepts like a disease outbreak, a cell-cycle checkpoint, a sequencer, biomineralization, or a functional magnetic resonance imaging (fMRI) data set be modelled in Wikidata or Wikibase?Conversely, what workflows allow us to collect information about such concepts in Wikidata, to interlink it with related information, to validate it, and to keep it up to date?Or, how can the data from Wikidata be explored or put to use in other contexts relevant to computational biology?We are working on establishing the editorial workflows to handle such Wikidata-focused Topic Pages and would welcome submissions to test these waters.For some inspiration, we suggest taking a look at Wikidata-based tools for browsing microbial genomes [2], scholarly publications [3], or software and file formats [4].The Author Guidelines for Wikipedia-focused Topic Pages are available here.If you've noticed a gap in Wikipedia's coverage of particular computational biology topics, we want to hear from you! Please send ideas for Topic Pages to [email protected]. Daniel Mietchen, Shoshana J. Wodak, Szymon Wasik, Natalia Szostak, Christophe Dessimoz |
PLoS Comput. Biol. | 5 |
| 2017 | Orthologous Matrix (OMA) algorithm 2.0: more robust to asymmetric evolutionary rates and more scalable hierarchical orthologous group inferenceabstractMOTIVATION: Accurate orthology inference is a fundamental step in many phylogenetics and comparative analysis. Many methods have been proposed, including OMA (Orthologous MAtrix). Yet substantial challenges remain, in particular in coping with fragmented genes or genes evolving at different rates after duplication, and in scaling to large datasets. With more and more genomes available, it is necessary to improve the scalability and robustness of orthology inference methods. RESULTS: We present improvements in the OMA algorithm: (i) refining the pairwise orthology inference step to account for same-species paralogs evolving at different rates, and (ii) minimizing errors in the pairwise orthology verification step by testing the consistency of pairwise distance estimates, which can be problematic in the presence of fragmentary sequences. In addition we introduce a more scalable procedure for hierarchical orthologous group (HOG) clustering, which are several orders of magnitude faster on large datasets. Using the Quest for Orthologs consortium orthology benchmark service, we show that these changes translate into substantial improvement on multiple empirical datasets. AVAILABILITY AND IMPLEMENTATION: This new OMA 2.0 algorithm is used in the OMA database ( http://omabrowser.org ) from the March 2017 release onwards, and can be run on custom genomes using OMA standalone version 2.0 and above ( http://omabrowser.org/standalone ). CONTACT: [email protected] or [email protected]. Clément-Marie Train, Natasha M. Glover, Gaston H. Gonnet, Adrian M. Altenhoff, Christophe Dessimoz |
Bioinform. | 5 |
| 2015 | Inferring Horizontal Gene TransferabstractHorizontal or Lateral Gene Transfer (HGT or LGT) is the transmission of portions of genomic DNA between organisms through a process decoupled from vertical inheritance. In the presence of HGT events, different fragments of the genome are the result of different evolutionary histories. This can therefore complicate the investigations of evolutionary relatedness of lineages and species. Also, as HGT can bring into genomes radically different genotypes from distant lineages, or even new genes bearing new functions, it is a major source of phenotypic innovation and a mechanism of niche adaptation. For example, of particular relevance to human health is the lateral transfer of antibiotic resistance and pathogenicity determinants, leading to the emergence of pathogenic lineages. Computational identification of HGT events relies upon the investigation of sequence composition or evolutionary history of genes. Sequence composition-based ("parametric") methods search for deviations from the genomic average, whereas evolutionary history-based ("phylogenetic") approaches identify genes whose evolutionary history significantly differs from that of the host species. The evaluation and benchmarking of HGT inference methods typically rely upon simulated genomes, for which the true history is known. On real data, different methods tend to infer different HGT events, and as a result it can be difficult to ascertain all but simple and clear-cut HGT events. Matt Ravenhall, Nives Skunca, Florent Lassalle, Christophe Dessimoz |
PLoS Comput. Biol. | 4 |
| 2014 | Big data and other challenges in the quest for orthologsabstractUNLABELLED: Given the rapid increase of species with a sequenced genome, the need to identify orthologous genes between them has emerged as a central bioinformatics task. Many different methods exist for orthology detection, which makes it difficult to decide which one to choose for a particular application. Here, we review the latest developments and issues in the orthology field, and summarize the most recent results reported at the third 'Quest for Orthologs' meeting. We focus on community efforts such as the adoption of reference proteomes, standard file formats and benchmarking. Progress in these areas is good, and they are already beneficial to both orthology consumers and providers. However, a major current issue is that the massive increase in complete proteomes poses computational challenges to many of the ortholog database providers, as most orthology inference algorithms scale at least quadratically with the number of proteomes. The Quest for Orthologs consortium is an open community with a number of working groups that join efforts to enhance various aspects of orthology analysis, such as defining standard formats and datasets, documenting community resources and benchmarking. AVAILABILITY AND IMPLEMENTATION: All such materials are available at http://questfororthologs.org. Erik L. L. Sonnhammer, Toni Gabaldón, Alan W. Sousa da Silva, Maria Jesus Martin, Marc Robinson-Rechavi, Brigitte Boeckmann, Paul D. Thomas, Christophe Dessimoz |
Bioinform. | 8 |
| 2013 | Approximate Bayesian ComputationabstractApproximate Bayesian computation (ABC) constitutes a class of computational methods rooted in Bayesian statistics. In all model-based statistical inference, the likelihood function is of central importance, since it expresses the probability of the observed data under a particular statistical model, and thus quantifies the support data lend to particular values of parameters and to choices among different models. For simple models, an analytical formula for the likelihood function can typically be derived. However, for more complex models, an analytical formula might be elusive or the likelihood function might be computationally very costly to evaluate. ABC methods bypass the evaluation of the likelihood function. In this way, ABC methods widen the realm of models for which statistical inference can be considered. ABC methods are mathematically well-founded, but they inevitably make assumptions and approximations whose impact needs to be carefully assessed. Furthermore, the wider application domain of ABC exacerbates the challenges of parameter estimation and model selection. ABC has rapidly gained popularity over the last years and in particular for the analysis of complex problems arising in biological sciences (e.g., in population genetics, ecology, epidemiology, and systems biology). Mikael Sunnåker, Alberto Giovanni Busetto, Elina Numminen, Jukka Corander, Matthieu Foll, Christophe Dessimoz |
PLoS Comput. Biol. | 6 |
| 2012 | Toward community standards in the quest for orthologsabstractThe identification of orthologs-genes pairs descended from a common ancestor through speciation, rather than duplication-has emerged as an essential component of many bioinformatics applications, ranging from the annotation of new genomes to experimental target prioritization. Yet, the development and application of orthology inference methods is hampered by the lack of consensus on source proteomes, file formats and benchmarks. The second 'Quest for Orthologs' meeting brought together stakeholders from various communities to address these challenges. We report on achievements and outcomes of this meeting, focusing on topics of particular relevance to the research community at large. The Quest for Orthologs consortium is an open community that welcomes contributions from all researchers interested in orthology research and applications. Christophe Dessimoz, Toni Gabaldón, David S. Roos, Erik L. L. Sonnhammer, Javier Herrero |
Bioinform. | 1 |
| 2012 | Resolving the Ortholog Conjecture: Orthologs Tend to Be Weakly, but Significantly, More Similar in Function than ParalogsabstractThe function of most proteins is not determined experimentally, but is extrapolated from homologs. According to the "ortholog conjecture", or standard model of phylogenomics, protein function changes rapidly after duplication, leading to paralogs with different functions, while orthologs retain the ancestral function. We report here that a comparison of experimentally supported functional annotations among homologs from 13 genomes mostly supports this model. We show that to analyze GO annotation effectively, several confounding factors need to be controlled: authorship bias, variation of GO term frequency among species, variation of background similarity among species pairs, and propagated annotation bias. After controlling for these biases, we observe that orthologs have generally more similar functional annotations than paralogs. This is especially strong for sub-cellular localization. We observe only a weak decrease in functional similarity with increasing sequence divergence. These findings hold over a large diversity of species; notably orthologs from model organisms such as E. coli, yeast or mouse have conserved function with human proteins. Adrian M. Altenhoff, Romain A. Studer, Marc Robinson-Rechavi, Christophe Dessimoz |
PLoS Comput. Biol. | 4 |
| 2012 | Quality of Computationally Inferred Gene Ontology AnnotationsabstractGene Ontology (GO) has established itself as the undisputed standard for protein function annotation. Most annotations are inferred electronically, i.e. without individual curator supervision, but they are widely considered unreliable. At the same time, we crucially depend on those automated annotations, as most newly sequenced genomes are non-model organisms. Here, we introduce a methodology to systematically and quantitatively evaluate electronic annotations. By exploiting changes in successive releases of the UniProt Gene Ontology Annotation database, we assessed the quality of electronic annotations in terms of specificity, reliability, and coverage. Overall, we not only found that electronic annotations have significantly improved in recent years, but also that their reliability now rivals that of annotations inferred by curators when they use evidence other than experiments from primary literature. This work provides the means to identify the subset of electronic annotations that can be relied upon-an important outcome given that >98% of all annotations are inferred without direct curation. Nives Skunca, Adrian M. Altenhoff, Christophe Dessimoz |
PLoS Comput. Biol. | 3 |
| 2011 | Conceptual framework and pilot study to benchmark phylogenomic databases based on reference gene treesabstractPhylogenomic databases provide orthology predictions for species with fully sequenced genomes. Although the goal seems well-defined, the content of these databases differs greatly. Seven ortholog databases (Ensembl Compara, eggNOG, HOGENOM, InParanoid, OMA, OrthoDB, Panther) were compared on the basis of reference trees. For three well-conserved protein families, we observed a generally high specificity of orthology assignments for these databases. We show that differences in the completeness of predicted gene relationships and in the phylogenetic information are, for the great majority, not due to the methods used, but to differences in the underlying database concepts. According to our metrics, none of the databases provides a fully correct and comprehensive protein classification. Our results provide a framework for meaningful and systematic comparisons of phylogenomic databases. In the future, a sustainable set of 'Gold standard' phylogenetic trees could provide a robust method for phylogenomic databases to assess their current quality status, measure changes following new database releases and diagnose improvements subsequent to an upgrade of the analysis procedure. Brigitte Boeckmann, Marc Robinson-Rechavi, Ioannis Xenarios, Christophe Dessimoz |
Briefings Bioinform. | 4 |
| 2011 | Editorial: Orthology and applicationsabstractThe accurate inference of orthologous genes underpins almost all biological studies that consider more than a single genome. Indeed, orthology formalizes the intuitive notion of corresponding genes in different species. As such, orthology finds applications in a broad range of research areas, such as functional genomics, comparative genomics, phylogenetics or pharmacology. Accordingly, well over 30 orthology databases have been developed (http://q4o.org/orthology_databases) and many thousands of scientific papers containing the keyword ‘ortholog’ are published each year. But success also comes with new challenges. In particular, each area of orthology applications entails its own constraints and trade-offs. This has given rise to multiple and at times conflicting definitions of orthology and associated relations—a common source of confusion even among long-time practitioners. Hence, to effectively call, interpret or apply ortholog predictions, knowledge of the problem in question is indispensable. The aim of this special issue is to provide a survey of the current state of orthology through multiple lenses, in form of reviews and original research papers. We start with a tribute to Walter M. Fitch, who passed away earlier this year. In his note of remembrance, Eugene Koonin provides a retrospective on Fitch's founding role in orthology. The rest of the first part focuses on definitions and methods for orthology inference. Kristensen et al. review the numerous computational methods that have been developed in recent years. They discuss the relative merits of the various approaches, both in theory and in practice. Doyon et al. consider orthology inference in the context of the more general problem of gene and species tree reconciliation. Indeed, orthology can be viewed as a byproduct of tree reconciliation. The authors review latest developments in parsimony and likelihood approaches. In particular, they report on models accounting not only for speciations and gene duplications, but also for lateral gene transfers. The contribution by Colin Dewey addresses the notion of positional orthology, which he formally defines in terms of past evolutionary events—not in terms of conserved gene neighborhood in present genomes (as in e.g. [1]). Sjölander et al. discuss the challenges of orthology inference when the underlying genes have heterogeneous domain architectures. They discuss a protocol for phylogenetic orthology inference based on domains instead of full-length protein sequences. They argue that the denser taxon sampling afforded by domain-level analyses counterbalances the phylogenetic uncertainty caused by shorter domain sequences. Boeckmann et al. compare seven well-established phylogenomic databases from the perspective of the user. We describe conceptual differences, and what they mean in terms of the orthology, paralogy and tree structure conveyed by each database. The paper shows how measuring these three aspects can allow for effective benchmarking based on reference gene trees. In the second part of this special issue, our attention shifts to applications of orthology. The perhaps greatest impact of orthology studies lies with gene function characterization. Though it might be tempting to systematically ascribe the same function to orthologous genes, Gharib and Robinson-Rechavi remind us that even for the relatively short human–mouse evolutionary distance, there are numerous instances of orthologs that have diverged functionally. Overall, however, Huerta-Cepas et al. report that human–mouse orthologs are significantly more conserved in expression pattern than their paralogous counterparts. These observations underscore the need for differentiated and prudent approaches to propagating function annotations. In their manuscript, Gaudet et al. describe the method used by curators of the Gene Ontology consortium to integrate and transfer function annotations based on the evolutionary history of gene families. In medical research, the focus is not so much on gene function as on gene dysfunction. Schreiber et al. present a major update of OrthoDisease, a database of human disease-associated genes and their orthologs in nearly 100 other species. Based on their data, they observe that disease-associated genes tend to have fewer close paralogs than other human genes, thereby supporting the notion that close paralogs can compensate each other functionally [2, 3]. In another orthology application, Dessimoz et al. examine the problem of split genes, endemic in low-coverage genome assemblies. We present a comparative genomics approach aimed at detecting these gene fragments present on multiple, unassembled contigs. This is of particular relevance here, because such pseudo-paralogs can confound orthology prediction and other phylogenetic analyses. The special issue closes with a letter by Schmitt et al., in which they define and motivate new XML formats for protein sequences and orthology predictions. This initiative epitomizes recent community efforts toward better interoperability and joint standards [4], and its outcome should facilitate the interpretation of results provided by the various orthology databases. CD is supported by an SNSF advanced researcher fellowship (#136461). Christophe Dessimoz |
Briefings Bioinform. | 1 |
| 2011 | Comparative genomics approach to detecting split-coding regions in a low-coverage genome: lessons from the chimaera Callorhinchus milii (Holocephali, Chondrichthyes)abstractRecent development of deep sequencing technologies has facilitated de novo genome sequencing projects, now conducted even by individual laboratories. However, this will yield more and more genome sequences that are not well assembled, and will hinder thorough annotation when no closely related reference genome is available. One of the challenging issues is the identification of protein-coding sequences split into multiple unassembled genomic segments, which can confound orthology assignment and various laboratory experiments requiring the identification of individual genes. In this study, using the genome of a cartilaginous fish, Callorhinchus milii, as test case, we performed gene prediction using a model specifically trained for this genome. We implemented an algorithm, designated ESPRIT, to identify possible linkages between multiple protein-coding portions derived from a single genomic locus split into multiple unassembled genomic segments. We developed a validation framework based on an artificially fragmented human genome, improvements between early and recent mouse genome assemblies, comparison with experimentally validated sequences from GenBank, and phylogenetic analyses. Our strategy provided insights into practical solutions for efficient annotation of only partially sequenced (low-coverage) genomes. To our knowledge, our study is the first formulation of a method to link unassembled genomic segments based on proteomes of relatively distantly related species as references. Christophe Dessimoz, Stefan Zoller, Tereza Manousaki, Huan Qiu, Axel Meyer, Shigehiro Kuraku |
Briefings Bioinform. | 1 |
| 2011 | Base-calling for next-generation sequencing platformsabstractNext-generation sequencing platforms are dramatically reducing the cost of DNA sequencing. With these technologies, bases are inferred from light intensity signals, a process commonly referred to as base-calling. Thus, understanding and improving the quality of sequence data generated using these approaches are of high interest. Recently, a number of papers have characterized the biases associated with base-calling and proposed methodological improvements. In this review, we summarize recent development of base-calling approaches for the Illumina and Roche 454 sequencing platforms. Christian Ledergerber, Christophe Dessimoz |
Briefings Bioinform. | 2 |
| 2011 | The what, where, how and why of gene ontology - a primer for bioinformaticiansabstractWith high-throughput technologies providing vast amounts of data, it has become more important to provide systematic, quality annotations. The Gene Ontology (GO) project is the largest resource for cataloguing gene function. Nonetheless, its use is not yet ubiquitous and is still fraught with pitfalls. In this review, we provide a short primer to the GO for bioinformaticians. We summarize important aspects of the structure of the ontology, describe sources and types of functional annotations, survey measures of GO annotation similarity, review typical uses of GO and discuss other important considerations pertaining to the use of GO in bioinformatics applications. Louis du Plessis, Nives Skunca, Christophe Dessimoz |
Briefings Bioinform. | 3 |
| 2009 | Algorithm of OMA for large-scale orthology inferenceabstractSince the publication of our article (Roth, Gonnet, and Dessimoz: BMC Bioinformatics 2008 9: 518), we have noticed several errors, which we correct in the following. Alexander C. J. Roth, Gaston H. Gonnet, Christophe Dessimoz |
BMC Bioinform. | 3 |
| 2009 | Phylogenetic and Functional Assessment of Orthologs Inference Projects and MethodsabstractAccurate genome-wide identification of orthologs is a central problem in comparative genomics, a fact reflected by the numerous orthology identification projects developed in recent years. However, only a few reports have compared their accuracy, and indeed, several recent efforts have not yet been systematically evaluated. Furthermore, orthology is typically only assessed in terms of function conservation, despite the phylogeny-based original definition of Fitch. We collected and mapped the results of nine leading orthology projects and methods (COG, KOG, Inparanoid, OrthoMCL, Ensembl Compara, Homologene, RoundUp, EggNOG, and OMA) and two standard methods (bidirectional best-hit and reciprocal smallest distance). We systematically compared their predictions with respect to both phylogeny and function, using six different tests. This required the mapping of millions of sequences, the handling of hundreds of millions of predicted pairs of orthologs, and the computation of tens of thousands of trees. In phylogenetic analysis or in functional analysis where high specificity is required, we find that OMA and Homologene perform best. At lower functional specificity but higher coverage level, OrthoMCL outperforms Ensembl Compara, and to a lesser extent Inparanoid. Lastly, the large coverage of the recent EggNOG can be of interest to build broad functional grouping, but the method is not specific enough for phylogenetic or detailed function analyses. In terms of general methodology, we observe that the more sophisticated tree reconstruction/reconciliation approach of Ensembl Compara was at times outperformed by pairwise comparison approaches, even in phylogenetic tests. Furthermore, we show that standard bidirectional best-hit often outperforms projects with more complex algorithms. First, the present study provides guidance for the broad community of orthology data users as to which database best suits their needs. Second, it introduces new methodology to verify orthology. And third, it sets performance standards for current and future approaches. Adrian M. Altenhoff, Christophe Dessimoz |
PLoS Comput. Biol. | 2 |
| 2008 | DLIGHT - Lateral Gene Transfer Detection Using Pairwise Evolutionary Distances in a Statistical Framework
Christophe Dessimoz, Daniel Margadant, Gaston H. Gonnet |
RECOMB | 1 |
| 2008 | Algorithm of OMA for large-scale orthology inferenceabstractBACKGROUND: OMA is a project that aims to identify orthologs within publicly available, complete genomes. With 657 genomes analyzed to date, OMA is one of the largest projects of its kind. RESULTS: The algorithm of OMA improves upon standard bidirectional best-hit approach in several respects: it uses evolutionary distances instead of scores, considers distance inference uncertainty, includes many-to-many orthologous relations, and accounts for differential gene losses. Herein, we describe in detail the algorithm for inference of orthology and provide the rationale for parameter selection through multiple tests. CONCLUSION: OMA contains several novel improvement ideas for orthology inference and provides a unique dataset of large-scale orthology assignments. Alexander C. J. Roth, Gaston H. Gonnet, Christophe Dessimoz |
BMC Bioinform. | 3 |
| 2007 | Alignments with Non-overlapping Moves, Inversions and Tandem Duplications in O ( n 4) Time
Christian Ledergerber, Christophe Dessimoz |
COCOON | 2 |
| 2007 | OMA Browser - Exploring orthologous relations across 352 complete genomesabstractMOTIVATION: Inference of the evolutionary relation between proteins, in particular the identification of orthologs, is a central problem in comparative genomics. Several large-scale efforts with various methodologies and scope tackle this problem, including OMA (the Orthologous MAtrix project). RESULTS: Based on the results of the OMA project, we introduce here the OMA Browser, a web-based tool allowing the exploration of orthologous relations over 352 complete genomes. Orthologs can be viewed as groups across species, but also at the level of sequence pairs, allowing the distinction among one-to-one, one-to-many and many-to-many orthologs. AVAILABILITY: http://omabrowser.org. Adrian Schneider, Christophe Dessimoz, Gaston H. Gonnet |
Bioinform. | 2 |
| 2006 | Fast estimation of the difference between two PAM/JTT evolutionary distances in triplets of homologous sequencesabstractBACKGROUND: The estimation of the difference between two evolutionary distances within a triplet of homologs is a common operation that is used for example to determine which of two sequences is closer to a third one. The most accurate method is currently maximum likelihood over the entire triplet. However, this approach is relatively time consuming. RESULTS: We show that an alternative estimator, based on pairwise estimates and therefore much faster to compute, has almost the same statistical power as the maximum likelihood estimator. We also provide a numerical approximation for its variance, which could otherwise only be estimated through an expensive re-sampling approach such as bootstrapping. An extensive simulation demonstrates that the approximation delivers precise confidence intervals. To illustrate the possible applications of these results, we show how they improve the detection of asymmetric evolution, and the identification of the closest relative to a given sequence in a group of homologs. CONCLUSION: The results presented in this paper constitute a basis for large-scale protein cross-comparisons of pairwise evolutionary distances. Christophe Dessimoz, Manuel Gil, Adrian Schneider, Gaston H. Gonnet |
BMC Bioinform. | 1 |