Erik L. L. Sonnhammer

dblp:46/4992 · DBLP profile ↗
← Back
49ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-9015-5588ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 49 · 5 first-author · 8 since 2021
YearPublicationVenuePosition
2025 The FunCoup Cytoscape App: multi-species network analysis and visualization
abstract
MOTIVATION: Functional association networks, such as FunCoup, are crucial for analyzing complex gene interactions. To facilitate the analysis and visualization of such genome-wide networks, there is a need for seamless integration with powerful network analysis tools like Cytoscape. RESULTS: The FunCoup Cytoscape App integrates the FunCoup web service API with Cytoscape, allowing users to visualize and analyze gene interaction networks for 640 species. Users can input gene identifiers and customize search parameters, using various network expansion algorithms like group or independent gene search, MaxLink, and TOPAS. The app maintains consistent visualizations with the FunCoup website, providing detailed node and link information, including tissue and pathway gene annotations. The integration with Cytoscape plugins, such as ClusterMaker2, enhances the analytical capabilities of FunCoup, as exemplified by the identification of the Myasthenia gravis disease module along with potential new therapeutic targets. AVAILABILITY AND IMPLEMENTATION: The FunCoup Cytoscape App is developed using the Java OSGi framework, with UI components implemented in Java Swing and build support from Maven. The App is available as a JAR file at https://bitbucket.org/sonnhammergroup/funcoup_cytoscape/ repository, and can be downloaded from the Cytoscape App store https://apps.cytoscape.org/apps/funcoup.
Davide Buzzao, Lukas Steininger, Dimitri Guala, Erik L. L. Sonnhammer
Bioinform.4
2025 Topology-based metrics for finding the optimal sparsity in gene regulatory network inference
abstract
MOTIVATION: Gene regulatory network (GRN) inference is a complex task aiming to unravel regulatory interactions between genes in a cell. A major shortcoming of most GRN inference methods is that they do not attempt to find the optimal sparsity, i.e. the single best GRN, which is important when applying GRN inference in a real situation. Instead, the sparsity tends to be controlled by an arbitrarily set hyperparameter. RESULTS: In this paper, two new methods for predicting the optimal sparsity of GRNs are formulated and benchmarked on simulated perturbation-based gene expression data using four GRN inference methods: LASSO, Zscore, LSCON, and GENIE3. Both sparsity prediction methods are defined using the hypothesis that the topology of real GRNs is scale-free, and are evaluated based on their ability to predict the sparsity of the true GRN. The results show that the new topology-based approaches reliably predict a sparsity close to the true one. This ability is valuable for real-world applications where a single GRN is inferred from real data. In such situations, it is vital to be able to infer a GRN with the correct sparsity. AVAILABILITY AND IMPLEMENTATION: https://bitbucket.org/sonnhammergrni/powerlaw_sparsity/ and https://codeocean.com/capsule/4393635/.
Nils Lundqvist, Mateusz Garbulowski, Thomas Hillerton, Erik L. L. Sonnhammer
Bioinform.4
2025 BiGSM: Bayesian inference of gene regulatory network via sparse modelling
abstract
MOTIVATION: Inference of gene regulatory network (GRN) is challenging due to the inherent sparsity of the GRN matrix and noisy expression data, often leading to a high possibility of false positive or negative predictions. To address this, it is essential to leverage the sparsity of the GRN matrix and develop a robust method capable of handling varying levels of noise in the data. Moreover, most existing GRN inference methods produce only fixed point estimates, which lack the flexibility and informativeness for comprehensive network analysis. In contrast, a Bayesian approach that yields closed-form posterior distributions allows probabilistic link selection, offering insights into the statistical confidence of each possible link. Consequently, it is important to engineer a Bayesian GRN inference method and rigorously execute a benchmark evaluation compared to state-of-the-art methods. RESULTS: We propose a method-Bayesian inference of GRN via Sparse Modelling (BiGSM). BiGSM effectively exploits the sparsity of the GRN matrix and infers the posterior distributions of GRN links from noisy expression data by using the maximum likelihood based learning. We thoroughly benchmarked BiGSM using biological and simulated datasets including GeneNetWeaver, GeneSPIDER, and GRNbenchmark. The benchmark test evaluates its accuracy and robustness across varying noise levels and data models. Using point-estimate based performance measures, BiGSM provides an overall best performance in comparison with several state-of-the-art methods including GENIE3, LASSO, LSCON, and Zscore. Additionally, BiGSM is the only method in the set of competing methods that provides posteriors for the GRN weights, helping to decipher confidence across predictions. AVAILABILITY AND IMPLEMENTATION: Code implemented via MATLAB and Python are available at Github: https://github.com/SachLab/BiGSM and archived at zenodo.
Hang Qin, Mateusz Garbulowski, Erik L. L. Sonnhammer, Saikat Chatterjee
Bioinform.3
2024 Benchmarking enrichment analysis methods with the disease pathway network
abstract
Enrichment analysis (EA) is a common approach to gain functional insights from genome-scale experiments. As a consequence, a large number of EA methods have been developed, yet it is unclear from previous studies which method is the best for a given dataset. The main issues with previous benchmarks include the complexity of correctly assigning true pathways to a test dataset, and lack of generality of the evaluation metrics, for which the rank of a single target pathway is commonly used. We here provide a generalized EA benchmark and apply it to the most widely used EA methods, representing all four categories of current approaches. The benchmark employs a new set of 82 curated gene expression datasets from DNA microarray and RNA-Seq experiments for 26 diseases, of which only 13 are cancers. In order to address the shortcomings of the single target pathway approach and to enhance the sensitivity evaluation, we present the Disease Pathway Network, in which related Kyoto Encyclopedia of Genes and Genomes pathways are linked. We introduce a novel approach to evaluate pathway EA by combining sensitivity and specificity to provide a balanced evaluation of EA methods. This approach identifies Network Enrichment Analysis methods as the overall top performers compared with overlap-based methods. By using randomized gene expression datasets, we explore the null hypothesis bias of each method, revealing that most of them produce skewed P-values.
Davide Buzzao, Miguel Castresana Aguirre, Dimitri Guala, Erik L. L. Sonnhammer
Briefings Bioinform.4
2022 Fast and accurate gene regulatory network inference by normalized least squares regression
abstract
MOTIVATION: Inferring an accurate gene regulatory network (GRN) has long been a key goal in the field of systems biology. To do this, it is important to find a suitable balance between the maximum number of true positive and the minimum number of false-positive interactions. Another key feature is that the inference method can handle the large size of modern experimental data, meaning the method needs to be both fast and accurate. The Least Squares Cut-Off (LSCO) method can fulfill both these criteria, however as it is based on least squares it is vulnerable to known issues of amplifying extreme values, small or large. In GRN this manifests itself with genes that are erroneously hyper-connected to a large fraction of all genes due to extremely low value fold changes. RESULTS: We developed a GRN inference method called Least Squares Cut-Off with Normalization (LSCON) that tackles this problem. LSCON extends the LSCO algorithm by regularization to avoid hyper-connected genes and thereby reduce false positives. The regularization used is based on normalization, which removes effects of extreme values on the fit. We benchmarked LSCON and compared it to Genie3, LASSO, LSCO and Ridge regression, in terms of accuracy, speed and tendency to predict hyper-connected genes. The results show that LSCON achieves better or equal accuracy compared to LASSO, the best existing method, especially for data with extreme values. Thanks to the speed of least squares regression, LSCON does this an order of magnitude faster than LASSO. AVAILABILITY AND IMPLEMENTATION: Data: https://bitbucket.org/sonnhammergrni/lscon; Code: https://bitbucket.org/sonnhammergrni/genespider. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Thomas Hillerton, Deniz Seçilmis, Sven Nelander, Erik L. L. Sonnhammer
Bioinform.4
2022 PathwAX II: network-based pathway analysis with interactive visualization of network crosstalk
abstract
MOTIVATION: Pathway annotation tools are indispensable for the interpretation of a wide range of experiments in life sciences. Network-based algorithms have recently been developed which are more sensitive than traditional overlap-based algorithms, but there is still a lack of good online tools for network-based pathway analysis. RESULTS: We present PathwAX II-a pathway analysis web tool based on network crosstalk analysis using the BinoX algorithm. It offers several new features compared with the first version, including interactive graphical network visualization of the crosstalk between a query gene set and an enriched pathway, and the addition of Reactome pathways. AVAILABILITY AND IMPLEMENTATION: PathwAX II is available at http://pathwax.sbc.su.se. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Christoph Ogris, Miguel Castresana Aguirre, Erik L. L. Sonnhammer
Bioinform.3
2022 InParanoid-DIAMOND: faster orthology analysis with the InParanoid algorithm
abstract
SUMMARY: Predicting orthologs, genes in different species having shared ancestry, is an important task in bioinformatics. Orthology prediction tools are required to make accurate and fast predictions, in order to analyze large amounts of data within a feasible time frame. InParanoid is a well-known algorithm for orthology analysis, shown to perform well in benchmarks, but having the major limitation of long runtimes on large datasets. Here, we present an update to the InParanoid algorithm that can use the faster tool DIAMOND instead of BLAST for the homolog search step. We show that it reduces the runtime by 94%, while still obtaining similar performance in the Quest for Orthologs benchmark. AVAILABILITY AND IMPLEMENTATION: The source code is available at (https://bitbucket.org/sonnhammergroup/inparanoid). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Emma Persson, Erik L. L. Sonnhammer
Bioinform.2
2021 Inferring the experimental design for accurate gene regulatory network inference
abstract
MOTIVATION: Accurate inference of gene regulatory interactions is of importance for understanding the mechanisms of underlying biological processes. For gene expression data gathered from targeted perturbations, gene regulatory network (GRN) inference methods that use the perturbation design are the top performing methods. However, the connection between the perturbation design and gene expression can be obfuscated due to problems, such as experimental noise or off-target effects, limiting the methods' ability to reconstruct the true GRN. RESULTS: In this study, we propose an algorithm, IDEMAX, to infer the effective perturbation design from gene expression data in order to eliminate the potential risk of fitting a disconnected perturbation design to gene expression. We applied IDEMAX to synthetic data from two different data generation tools, GeneNetWeaver and GeneSPIDER, and assessed its effect on the experiment design matrix as well as the accuracy of the GRN inference, followed by application to a real dataset. The results show that our approach consistently improves the accuracy of GRN inference compared to using the intended perturbation design when much of the signal is hidden by noise, which is often the case for real data. AVAILABILITY AND IMPLEMENTATION: https://bitbucket.org/sonnhammergrni/idemax. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Deniz Seçilmis, Thomas Hillerton, Sven Nelander, Erik L. L. Sonnhammer
Bioinform.4
2020 Genome-wide functional association networks: background, data & state-of-the-art resources
abstract
The vast amount of experimental data from recent advances in the field of high-throughput biology begs for integration into more complex data structures such as genome-wide functional association networks. Such networks have been used for elucidation of the interplay of intra-cellular molecules to make advances ranging from the basic science understanding of evolutionary processes to the more translational field of precision medicine. The allure of the field has resulted in rapid growth of the number of available network resources, each with unique attributes exploitable to answer different biological questions. Unfortunately, the high volume of network resources makes it impossible for the intended user to select an appropriate tool for their particular research question. The aim of this paper is to provide an overview of the underlying data and representative network resources as well as to mention methods of integration, allowing a customized approach to resource selection. Additionally, this report will provide a primer for researchers venturing into the field of network integration.
Dimitri Guala, Christoph Ogris, Nikola S. Müller, Erik L. L. Sonnhammer
Briefings Bioinform.4
2019 A generalized framework for controlling FDR in gene regulatory network inference
abstract
MOTIVATION: Inference of gene regulatory networks (GRNs) from perturbation data can give detailed mechanistic insights of a biological system. Many inference methods exist, but the resulting GRN is generally sensitive to the choice of method-specific parameters. Even though the inferred GRN is optimal given the parameters, many links may be wrong or missing if the data is not informative. To make GRN inference reliable, a method is needed to estimate the support of each predicted link as the method parameters are varied. RESULTS: To achieve this we have developed a method called nested bootstrapping, which applies a bootstrapping protocol to GRN inference, and by repeated bootstrap runs assesses the stability of the estimated support values. To translate bootstrap support values to false discovery rates we run the same pipeline with shuffled data as input. This provides a general method to control the false discovery rate of GRN inference that can be applied to any setting of inference parameters, noise level, or data properties. We evaluated nested bootstrapping on a simulated dataset spanning a range of such properties, using the LASSO, Least Squares, RNI, GENIE3 and CLR inference methods. An improved inference accuracy was observed in almost all situations. Nested bootstrapping was incorporated into the GeneSPIDER package, which was also used for generating the simulated networks and data, as well as running and analyzing the inferences. AVAILABILITY AND IMPLEMENTATION: https://bitbucket.org/sonnhammergrni/genespider/src/NB/%2B Methods/NestBoot.m.
Daniel Morgan, Andreas Tjärnberg, Torbjörn E. M. Nordling, Erik L. L. Sonnhammer
Bioinform.4
2019 Domainoid: domain-oriented orthology inference
abstract
BACKGROUND: Orthology inference is normally based on full-length protein sequences. However, most proteins contain independently folding and recurring regions, domains. The domain architecture of a protein is vital for its function, and recombination events mean individual domains can have different evolutionary histories. It has previously been shown that orthologous proteins may differ in domain architecture, creating challenges for orthology inference methods operating on full-length sequences. We have developed Domainoid, a new tool aiming to overcome these challenges faced by full-length orthology methods by inferring orthology on the domain level. It employs the InParanoid algorithm on single domains separately, to infer groups of orthologous domains. RESULTS: This domain-oriented approach allows detection of discordant domain orthologs, cases where different domains on the same protein have different evolutionary histories. In addition to domain level analysis, protein level orthology based on the fraction of domains that are orthologous can be inferred. Domainoid orthology assignments were compared to those yielded by the conventional full-length approach InParanoid, and were validated in a standard benchmark. CONCLUSIONS: Our results show that domain-based orthology inference can reveal many orthologous relationships that are not found by full-length sequence approaches. AVAILABILITY: https://bitbucket.org/sonnhammergroup/domainoid/.
Emma Persson, Mateusz Kaduk, Sofia K. Forslund, Erik L. L. Sonnhammer
BMC Bioinform.4
2018 Gearing up to handle the mosaic nature of life in the quest for orthologs
abstract
The Quest for Orthologs (QfO) is an open collaboration framework for experts in comparative phylogenomics and related research areas who have an interest in highly accurate orthology predictions and their applications. We here report highlights and discussion points from the QfO meeting 2015 held in Barcelona. Achievements in recent years have established a basis to support developments for improved orthology prediction and to explore new approaches. Central to the QfO effort is proper benchmarking of methods and services, as well as design of standardized datasets and standardized formats to allow sharing and comparison of results. Simultaneously, analysis pipelines have been improved, evaluated and adapted to handle large datasets. All this would not have occurred without the long-term collaboration of Consortium members. Meeting regularly to review and coordinate complementary activities from a broad spectrum of innovative researchers clearly benefits the community. Highlights of the meeting include addressing sources of and legitimacy of disagreements between orthology calls, the context dependency of orthology definitions, special challenges encountered when analyzing very anciently rooted orthologies, orthology in the light of whole-genome duplications, and the concept of orthologous versus paralogous relationships at different levels, including domain-level orthology. Furthermore, particular needs for different applications (e.g. plant genomics, ancient gene families and others) and the infrastructure for making orthology inferences available (e.g. interfaces with model organism databases) were discussed, with several ongoing efforts that are expected to be reported on during the upcoming 2017 QfO meeting.
Sofia K. Forslund, Cécile Pereira, Salvador Capella-Gutiérrez, Alan W. Sousa da Silva, Adrian M. Altenhoff, Jaime Huerta-Cepas, Matthieu Muffato, Mateus Patricio, Klaas Vandepoele, Ingo Ebersberger, Judith A. Blake, Jesualdo Tomás Fernández-Breis, Brigitte Boeckmann, Toni Gabaldón, Erik L. L. Sonnhammer, Christophe Dessimoz, Suzanna Lewis
Bioinform.16
2017 Improved orthology inference with Hieranoid 2
abstract
Motivation: The initial step in many orthology inference methods is the computationally demanding establishment of all pairwise protein similarities across all analysed proteomes. The quadratic scaling with proteomes has become a major bottleneck. A remedy is offered by the Hieranoid algorithm which reduces the complexity to linear by hierarchically aggregating ortholog groups from InParanoid along a species tree. Results: We have further developed the Hieranoid algorithm in many ways. Major improvements have been made to the construction of multiple sequence alignments and consensus sequences. Hieranoid version 2 was evaluated with standard benchmarks that reveal a dramatic increase in the coverage/accuracy tradeoff over version 1, such that it now compares favourably with the best methods. The new parallelized cluster mode allows Hieranoid to be run on large data sets in a much shorter timespan than InParanoid, yet at similar accuracy. Contact: [email protected]. Availability and Implementation: Perl code freely available at http://hieranoid.sbc.su.se/ . Supplementary information: Supplementary data are available at Bioinformatics online.
Mateusz Kaduk, Erik L. L. Sonnhammer
Bioinform.2
2016 TreeDom: a graphical web tool for analysing domain architecture evolution
abstract
UNLABELLED: We present TreeDom, a web tool for graphically analysing the evolutionary history of domains in multi-domain proteins. Individual domains on the same protein chain may have distinct evolutionary histories, which is important to grasp in order to understand protein function. For instance, it may be important to know whether a domain was duplicated recently or long ago, to know the origin of inserted domains, or to know the pattern of domain loss within a protein family. TreeDom uses the Pfam database as the source of domain annotations, and displays these on a sequence tree. An advantage of TreeDom is that the user can limit the analysis to N sequences that are most similar to a query, or provide a list of sequence IDs to include. Using the Pfam alignment of the selected sequences, a tree is built and displayed together with the domain architecture of each sequence.Availablility and implementation: http://TreeDom.sbc.su.se CONTACT: [email protected].
Christian Haider, Marina Kavic, Erik L. L. Sonnhammer
Bioinform.3
2016 Benchmarking the next generation of homology inference tools
abstract
MOTIVATION: Over the last decades, vast numbers of sequences were deposited in public databases. Bioinformatics tools allow homology and consequently functional inference for these sequences. New profile-based homology search tools have been introduced, allowing reliable detection of remote homologs, but have not been systematically benchmarked. To provide such a comparison, which can guide bioinformatics workflows, we extend and apply our previously developed benchmark approach to evaluate the 'next generation' of profile-based approaches, including CS-BLAST, HHSEARCH and PHMMER, in comparison with the non-profile based search tools NCBI-BLAST, USEARCH, UBLAST and FASTA. METHOD: We generated challenging benchmark datasets based on protein domain architectures within either the PFAM + Clan, SCOP/Superfamily or CATH/Gene3D domain definition schemes. From each dataset, homologous and non-homologous protein pairs were aligned using each tool, and standard performance metrics calculated. We further measured congruence of domain architecture assignments in the three domain databases. RESULTS: CSBLAST and PHMMER had overall highest accuracy. FASTA, UBLAST and USEARCH showed large trade-offs of accuracy for speed optimization. CONCLUSION: Profile methods are superior at inferring remote homologs but the difference in accuracy between methods is relatively small. PHMMER and CSBLAST stand out with the highest accuracy, yet still at a reasonable computational cost. Additionally, we show that less than 0.1% of Swiss-Prot protein pairs considered homologous by one database are considered non-homologous by another, implying that these classifications represent equivalent underlying biological phenomena, differing mostly in coverage and granularity. AVAILABILITY AND IMPLEMENTATION: Benchmark datasets and all scripts are placed at (http://sonnhammer.org/download/Homology_benchmark). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ganapathi Varma Saripella, Erik L. L. Sonnhammer, Sofia K. Forslund
Bioinform.2
2014 MaxLink: network-based prioritization of genes tightly linked to a disease seed set
abstract
UNLABELLED: MaxLink, a guilt-by-association network search algorithm, has been made available as a web resource and a stand-alone version. Based on a user-supplied list of query genes, MaxLink identifies and ranks genes that are tightly linked to the query list. This functionality can be used to predict potential disease genes from an initial set of genes with known association to a disease. The original algorithm, used to identify and rank novel genes potentially involved in cancer, has been updated to use a more statistically sound method for selection of candidate genes and made applicable to other areas than cancer. The algorithm has also been made faster by re-implementation in C++, and the Web site uses FunCoup 3.0 as the underlying network. AVAILABILITY AND IMPLEMENTATION: MaxLink is freely available at http://maxlink.sbc.su.se both as a web service and a stand-alone application for download.
Dimitri Guala, Erik Sjölund, Erik L. L. Sonnhammer
Bioinform.3
2014 Big data and other challenges in the quest for orthologs
abstract
UNLABELLED: Given the rapid increase of species with a sequenced genome, the need to identify orthologous genes between them has emerged as a central bioinformatics task. Many different methods exist for orthology detection, which makes it difficult to decide which one to choose for a particular application. Here, we review the latest developments and issues in the orthology field, and summarize the most recent results reported at the third 'Quest for Orthologs' meeting. We focus on community efforts such as the adoption of reference proteomes, standard file formats and benchmarking. Progress in these areas is good, and they are already beneficial to both orthology consumers and providers. However, a major current issue is that the massive increase in complete proteomes poses computational challenges to many of the ortholog database providers, as most orthology inference algorithms scale at least quadratically with the number of proteomes. The Quest for Orthologs consortium is an open community with a number of working groups that join efforts to enhance various aspects of orthology analysis, such as defining standard formats and datasets, documenting community resources and benchmarking. AVAILABILITY AND IMPLEMENTATION: All such materials are available at http://questfororthologs.org.
Erik L. L. Sonnhammer, Toni Gabaldón, Alan W. Sousa da Silva, Maria Jesus Martin, Marc Robinson-Rechavi, Brigitte Boeckmann, Paul D. Thomas, Christophe Dessimoz
Bioinform.1
2014 Functional association networks as priors for gene regulatory network inference
abstract
MOTIVATION: Gene regulatory network (GRN) inference reveals the influences genes have on one another in cellular regulatory systems. If the experimental data are inadequate for reliable inference of the network, informative priors have been shown to improve the accuracy of inferences. RESULTS: This study explores the potential of undirected, confidence-weighted networks, such as those in functional association databases, as a prior source for GRN inference. Such networks often erroneously indicate symmetric interaction between genes and may contain mostly correlation-based interaction information. Despite these drawbacks, our testing on synthetic datasets indicates that even noisy priors reflect some causal information that can improve GRN inference accuracy. Our analysis on yeast data indicates that using the functional association databases FunCoup and STRING as priors can give a small improvement in GRN inference accuracy with biological data.
Matthew Studham, Andreas Tjärnberg, Torbjörn E. M. Nordling, Sven Nelander, Erik L. L. Sonnhammer
Bioinform.5
2012 Toward community standards in the quest for orthologs
abstract
The identification of orthologs-genes pairs descended from a common ancestor through speciation, rather than duplication-has emerged as an essential component of many bioinformatics applications, ranging from the annotation of new genomes to experimental target prioritization. Yet, the development and application of orthology inference methods is hampered by the lack of consensus on source proteomes, file formats and benchmarks. The second 'Quest for Orthologs' meeting brought together stakeholders from various communities to address these challenges. We report on achievements and outcomes of this meeting, focusing on topics of particular relevance to the research community at large. The Quest for Orthologs consortium is an open community that welcomes contributions from all researchers interested in orthology research and applications.
Christophe Dessimoz, Toni Gabaldón, David S. Roos, Erik L. L. Sonnhammer, Javier Herrero
Bioinform.4
2011 OrthoDisease: tracking disease gene orthologs across 100 species
abstract
Orthology is one of the most important tools available to modern biology, as it allows making inferences from easily studied model systems to much less tractable systems of interest, such as ourselves. This becomes important not least in the study of genetic diseases. We here review work on the orthology of disease-associated genes and also present an updated version of the InParanoid-based disease orthology database and web site OrthoDisease, with 14-fold increased species coverage since the previous version. Using this resource, we survey the taxonomic distribution of orthologs of human genes involved in different disease categories. The hypothesis that paralogs can mask the effect of deleterious mutations predicts that known heritable disease genes should have fewer close paralogs. We found large-scale support for this hypothesis as significantly fewer duplications were observed for disease genes in the OrthoDisease ortholog groups.
Sofia K. Forslund, Fabian Schreiber, Nattaphon Thanintorn, Erik L. L. Sonnhammer
Briefings Bioinform.4
2011 Letter to the Editor: SeqXML and OrthoXML: standards for sequence and orthology information
abstract
There is a great need for standards in the orthology field. Users must contend with different ortholog data representations from each provider, and the providers themselves must independently gather and parse the input sequence data. These burdensome and redundant procedures make data comparison and integration difficult. We have designed two XML-based formats, SeqXML and OrthoXML, to solve these problems. SeqXML is a lightweight format for sequence records-the input for orthology prediction. It stores the same sequence and metadata as typical FASTA format records, but overcomes common problems such as unstructured metadata in the header and erroneous sequence content. XML provides validation to prevent data integrity problems that are frequent in FASTA files. The range of applications for SeqXML is broad and not limited to ortholog prediction. We provide read/write functions for BioJava, BioPerl, and Biopython. OrthoXML was designed to represent ortholog assignments from any source in a consistent and structured way, yet cater to specific needs such as scoring schemes or meta-information. A unified format is particularly valuable for ortholog consumers that want to integrate data from numerous resources, e.g. for gene annotation projects. Reference proteomes for 61 organisms are already available in SeqXML, and 10 orthology databases have signed on to OrthoXML. Adoption by the entire field would substantially facilitate exchange and quality control of sequence and orthology information.
Thomas Schmitt 0003, David N. Messina, Fabian Schreiber, Erik L. L. Sonnhammer
Briefings Bioinform.4
2011 Domain architecture conservation in orthologs
abstract
BACKGROUND: As orthologous proteins are expected to retain function more often than other homologs, they are often used for functional annotation transfer between species. However, ortholog identification methods do not take into account changes in domain architecture, which are likely to modify a protein's function. By domain architecture we refer to the sequential arrangement of domains along a protein sequence.To assess the level of domain architecture conservation among orthologs, we carried out a large-scale study of such events between human and 40 other species spanning the entire evolutionary range. We designed a score to measure domain architecture similarity and used it to analyze differences in domain architecture conservation between orthologs and paralogs relative to the conservation of primary sequence. We also statistically characterized the extents of different types of domain swapping events across pairs of orthologs and paralogs. RESULTS: The analysis shows that orthologs exhibit greater domain architecture conservation than paralogous homologs, even when differences in average sequence divergence are compensated for, for homologs that have diverged beyond a certain threshold. We interpret this as an indication of a stronger selective pressure on orthologs than paralogs to retain the domain architecture required for the proteins to perform a specific function. In general, orthologs as well as the closest paralogous homologs have very similar domain architectures, even at large evolutionary separation.The most common domain architecture changes observed in both ortholog and paralog pairs involved insertion/deletion of new domains, while domain shuffling and segment duplication/deletion were very infrequent. CONCLUSIONS: On the whole, our results support the hypothesis that function conservation between orthologs demands higher domain architecture conservation than other types of homologs, relative to primary sequence conservation. This supports the notion that orthologs are functionally more similar than other types of homologs at the same evolutionary distance.
Sofia K. Forslund, Isabella Pekkari, Erik L. L. Sonnhammer
BMC Bioinform.3
2009 Comparative analysis and unification of domain-domain interaction networks
abstract
MOTIVATION: Certain protein domains are known to preferentially interact with other domains. Several approaches have been proposed to predict domain-domain interactions, and over nine datasets are available. Our aim is to analyse the coverage and quality of the existing resources, as well as the extent of their overlap. With this knowledge, we have the opportunity to merge individual domain interaction networks to construct a comprehensive and reliable database. RESULTS: In this article we introduce a new approach towards comparing domain-domain interaction networks. This approach is used to compare nine predicted domain and protein interaction networks. The networks were used to generate a database of unified domain interactions, UniDomInt. Each interaction in the dataset is scored according to the benchmarked reliability of the sources. The performance of UniDomInt is an improvement compared to the underlying source networks and to another composite resource, Domine. AVAILABILITY: http://sonnhammer.sbc.su.se/download/UniDomInt/
Patrik Björkholm, Erik L. L. Sonnhammer
Bioinform.2
2009 Predicting protein function from domain content
abstract
Bioinformatics (2008) 24(15), 1681–1687. The authors regret that there were the following errors in the above paper. At the bottom of page 2, in Section 2.2: ‘where the product is taken over the i = 0…K subsets of D. There are K = 2N − 1 such subsets for N unique domains in D.’ should be ‘where the product is taken over the i = 1…K subsets of D. There are K = 2N − 1 such subsets for N unique domains in D.’
Sofia K. Forslund, Erik L. L. Sonnhammer
Bioinform.2
2009 Benchmarking homology detection procedures with low complexity filters
abstract
BACKGROUND: Low-complexity sequence regions present a common problem in finding true homologs to a protein query sequence. Several solutions to this have been suggested, but a detailed comparison between these on challenging data has so far been lacking. A common benchmark for homology detection procedures is to use SCOP/ASTRAL domain sequences belonging to the same or different superfamilies, but these contain almost no low complexity sequences. RESULTS: We here introduce an alternative benchmarking strategy based around Pfam domains and clans on whole-proteome data sets. This gives a realistic level of low complexity sequences. We used it to evaluate all six built-in BLAST low complexity filter settings as well as a range of settings in the MSPcrunch post-processing filter. The effect on alignment length was also assessed. CONCLUSION: Score matrix adjustment methods provide a low false positive rate at a relatively small loss in sensitivity relative to no filtering, across the range of test conditions we apply. MSPcrunch achieved even less loss in sensitivity, but at a higher false positive rate. A drawback of the score matrix adjustment methods is however that the alignments often become truncated. AVAILABILITY: Perl scripts for MSPcrunch BLAST filtering and for generating the benchmark dataset are available at http://sonnhammer.sbc.su.se/download/software/MSPcrunch+Blixem/benchmark.tar.gz
Sofia K. Forslund, Erik L. L. Sonnhammer
Bioinform.2
2009 DASher: a stand-alone protein sequence client for DAS, the Distributed Annotation System
abstract
SUMMARY: The rise in biological sequence data has led to a proliferation of separate, specialized databases. While there is great value in having many independent annotations, it is critical that there be a way to integrate them in one combined view. The Distributed Annotation System (DAS) was developed for that very purpose. There are currently no DAS clients that are open source, specialized for aggregating and comparing protein sequence annotation, and that can run as a self-contained application outside of a web browser. The speed, flexibility and extensibility that come with a stand-alone application motivated us to create DASher, an open-source Java DAS client. Given a UniProt sequence identifier, DASher automatically queries DAS-supporting servers worldwide for any information on that sequence and then displays the annotations in an interactive viewer for easy comparison. DASher is a fast, Java-based DAS client optimized for viewing protein sequence annotation and compliant with the latest DAS protocol specification 1.53E. AVAILABILITY: DASher is available for direct use and download at http://dasher.sbc.su.se including examples and source code under the GPLv3 licence. Java version 6 or higher is required.
David N. Messina, Erik L. L. Sonnhammer
Bioinform.2
2009 MetaTM - a consensus method for transmembrane protein topology prediction
abstract
BACKGROUND: Transmembrane (TM) proteins are proteins that span a biological membrane one or more times. As their 3-D structures are hard to determine, experiments focus on identifying their topology (i. e. which parts of the amino acid sequence are buried in the membrane and which are located on either side of the membrane), but only a few topologies are known. Consequently, various computational TM topology predictors have been developed, but their accuracies are far from perfect. The prediction quality can be improved by applying a consensus approach, which combines results of several predictors to yield a more reliable result. RESULTS: A novel TM consensus method, named MetaTM, is proposed in this work. MetaTM is based on support vector machine models and combines the results of six TM topology predictors and two signal peptide predictors. On a large data set comprising 1460 sequences of TM proteins with known topologies and 2362 globular protein sequences it correctly predicts 86.7% of all topologies. CONCLUSION: Combining several TM predictors in a consensus prediction framework improves overall accuracy compared to any of the individual methods. Our proposed SVM-based system also has higher accuracy than a previous consensus predictor. MetaTM is made available both as downloadable source code and as DAS server at http://MetaTM.sbc.su.se.
Martin Klammer, David N. Messina, Thomas Schmitt 0003, Erik L. L. Sonnhammer
BMC Bioinform.4
2008 siRNA specificity searching incorporating mismatch tolerance data
abstract
UNLABELLED: Artificially synthesized short interfering RNAs (siRNAs) are widely used in functional genomics to knock down specific target genes. One ongoing challenge is to guarantee that the siRNA does not elicit off-target effects. Initial reports suggested that siRNAs were highly sequence-specific; however, subsequent data indicates that this is not necessarily the case. It is still uncertain what level of similarity and other rules are required for an off-target effect to be observed, and scoring schemes have not been developed to look beyond simple measures such as the number of mismatches or the number of consecutive matching bases present. We created design rules for predicting the likelihood of a non-specific effect and present a web server that allows the user to check the specificity of a given siRNA in a flexible manner using a combination of methods. The server finds potential off-target matches in the corresponding RefSeq database and ranks them according to a scoring system based on experimental studies of specificity. AVAILABILITY: The server is available at http://informatics-eskitis.griffith.edu.au/SpecificityServer.
Alistair M. Chalk, Erik L. L. Sonnhammer
Bioinform.2
2008 Predicting protein function from domain content
abstract
MOTIVATION: Computational assignment of protein function may be the single most vital application of bioinformatics in the post-genome era. These assignments are made based on various protein features, where one is the presence of identifiable domains. The relationship between protein domain content and function is important to investigate, to understand how domain combinations encode complex functions. RESULTS: Two different models are presented on how protein domain combinations yield specific functions: one rule-based and one probabilistic. We demonstrate how these are useful for Gene Ontology annotation transfer. The first is an intuitive generalization of the Pfam2GO mapping, and detects cases of strict functional implications of sets of domains. The second uses a probabilistic model to represent the relationship between domain content and annotation terms, and was found to be better suited for incomplete training sets. We implemented these models as predictors of Gene Ontology functional annotation terms. Both predictors were more accurate than conventional best BLAST-hit annotation transfer and more sensitive than a single-domain model on a large-scale dataset. We present a number of cases where combinations of Pfam-A protein domains predict functional terms that do not follow from the individual domains. AVAILABILITY: Scripts and documentation are available for download at http://sonnhammer.sbc.su.se/multipfam2go_source_docs.tar
Sofia K. Forslund, Erik L. L. Sonnhammer
Bioinform.2
2008 Predicting protein function from domain content
abstract
Bioinformatics 2008; Vol. 24 no. 15: 1681–1687. The authors regret that there was an error in the above paper. At the bottom of page 2, in section 2.2: ‘where the product is taken over the i=0.K subsets of D. There are K=2N – 1 such subsets for N unique domains in D.’ should be ‘where the product is taken over the i=1.K subsets of D. There are K=2N – 1 such subsets for N unique domains in D.’
Sofia K. Forslund, Erik L. L. Sonnhammer
Bioinform.2
2008 jSquid: a Java applet for graphical on-line network exploration
abstract
UNLABELLED: jSquid is a graph visualization tool for exploring graphs from protein-protein interaction or functional coupling networks. The tool was designed for the FunCoup web site, but can be used for any similar network exploring purpose. The program offers various visualization and graph manipulation techniques to increase the utility for the user. AVAILABILITY: jSquid is available for direct usage and download at http://jSquid.sbc.su.se including source code under the GPLv3 license, and input examples. It requires Java version 5 or higher to run properly. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Martin Klammer, Sanjit Roopra, Erik L. L. Sonnhammer
Bioinform.3
2007 PfamAlyzer: domain-centric homology search
abstract
UNLABELLED: PfamAlyzer is a Java applet that enables exploration of Pfam domain architectures using a user-friendly graphical interface. It can search the UniProt protein database for a domain pattern. Domain patterns similar to the query are presented graphically by PfamAlyzer either in a ranked list or pinned to the tree of life. Such domain-centric homology search can assist identification of distant homologs with shared domain architecture. AVAILABILITY: PfamAlyzer has been integrated with the Pfam database and can be accessed at http://pfam.cgb.ki.se/pfamalyzer.
Volker Hollich, Erik L. L. Sonnhammer
Bioinform.2
2007 Automatic extraction of reliable regions from multiple sequence alignments
abstract
BACKGROUND: High quality multiple alignments are crucial in the transfer of annotation from one genome to another. Multiple alignment methods strive to achieve ever increasing levels of average accuracy on benchmark sets while the accuracy of individual alignments is often overlooked. RESULTS: We have previously developed a method to automatically assess the accuracy and overall difficulty of multiple alignments. This was achieved by a per-residue comparison between alternate alignments of the same sequences. Here we present a key extension to this method, an algorithm to extract similarly aligned regions from several alignments and merge them into a new consensus alignment. CONCLUSION: We demonstrate that the fraction of correctly aligned residues within the resulting alignments is increased by 25-100 percent compared to the original input alignments, as only the most reliably aligned parts are considered.
Timo Lassmann, Erik L. L. Sonnhammer
BMC Bioinform.2
2005 Kalign - an accurate and fast multiple sequence alignment algorithm
abstract
BACKGROUND: The alignment of multiple protein sequences is a fundamental step in the analysis of biological data. It has traditionally been applied to analyzing protein families for conserved motifs, phylogeny, structural properties, and to improve sensitivity in homology searching. The availability of complete genome sequences has increased the demands on multiple sequence alignment (MSA) programs. Current MSA methods suffer from being either too inaccurate or too computationally expensive to be applied effectively in large-scale comparative genomics. RESULTS: We developed Kalign, a method employing the Wu-Manber string-matching algorithm, to improve both the accuracy and speed of multiple sequence alignment. We compared the speed and accuracy of Kalign to other popular methods using Balibase, Prefab, and a new large test set. Kalign was as accurate as the best other methods on small alignments, but significantly more accurate when aligning large and distantly related sets of sequences. In our comparisons, Kalign was about 10 times faster than ClustalW and, depending on the alignment size, up to 50 times faster than popular iterative methods. CONCLUSION: Kalign is a fast and robust alignment method. It is especially well suited for the increasingly important task of aligning large numbers of sequences.
Timo Lassmann, Erik L. L. Sonnhammer
BMC Bioinform.2
2005 Scoredist: A simple and robust protein sequence distance estimator
abstract
BACKGROUND: Distance-based methods are popular for reconstructing evolutionary trees thanks to their speed and generality. A number of methods exist for estimating distances from sequence alignments, which often involves some sort of correction for multiple substitutions. The problem is to accurately estimate the number of true substitutions given an observed alignment. So far, the most accurate protein distance estimators have looked for the optimal matrix in a series of transition probability matrices, e.g. the Dayhoff series. The evolutionary distance between two aligned sequences is here estimated as the evolutionary distance of the optimal matrix. The optimal matrix can be found either by an iterative search for the Maximum Likelihood matrix, or by integration to find the Expected Distance. As a consequence, these methods are more complex to implement and computationally heavier than correction-based methods. Another problem is that the result may vary substantially depending on the evolutionary model used for the matrices. An ideal distance estimator should produce consistent and accurate distances independent of the evolutionary model used. RESULTS: We propose a correction-based protein sequence estimator called Scoredist. It uses a logarithmic correction of observed divergence based on the alignment score according to the BLOSUM62 score matrix. We evaluated Scoredist and a number of optimal matrix methods using three evolutionary models for both training and testing Dayhoff, Jones-Taylor-Thornton, and Muller-Vingron, as well as Whelan and Goldman solely for testing. Test alignments with known distances between 0.01 and 2 substitutions per position (1-200 PAM) were simulated using ROSE. Scoredist proved as accurate as the optimal matrix methods, yet substantially more robust. When trained on one model but tested on another one, Scoredist was nearly always more accurate. The Jukes-Cantor and Kimura correction methods were also tested, but were substantially less accurate. CONCLUSION: The Scoredist distance estimator is fast to implement and run, and combines robustness with accuracy. Scoredist has been incorporated into the Belvu alignment viewer, which is available at ftp://ftp.cgb.ki.se/pub/prog/belvu/.
Erik L. L. Sonnhammer, Volker Hollich
BMC Bioinform.1
2005 Improved profile HMM performance by assessment of critical algorithmic features in SAM and HMMER
abstract
BACKGROUND: Profile hidden Markov model (HMM) techniques are among the most powerful methods for protein homology detection. Yet, the critical features for successful modelling are not fully known. In the present work we approached this by using two of the most popular HMM packages: SAM and HMMER. The programs' abilities to build models and score sequences were compared on a SCOP/Pfam based test set. The comparison was done separately for local and global HMM scoring. RESULTS: Using default settings, SAM was overall more sensitive. SAM's model estimation was superior, while HMMER's model scoring was more accurate. Critical features for model building were then analysed by comparing the two packages' algorithmic choices and parameters. The weighting between prior probabilities and multiple alignment counts held the primary explanation why SAM's model building was superior. Our analysis suggests that HMMER gives too much weight to the sequence counts. SAM's emission prior probabilities were also shown to be more sensitive. The relative sequence weighting schemes are different in the two packages but performed equivalently. CONCLUSION: SAM model estimation was more sensitive, while HMMER model scoring was more accurate. By combining the best algorithmic features from both packages the accuracy was substantially improved compared to their default performance.
Markus Wistrand, Erik L. L. Sonnhammer
BMC Bioinform.2
2004 Sfixem - graphical sequence feature display in Java
abstract
UNLABELLED: Sfixem is an sequence feature series (SFS) visualization tool implemented in Java. It is designed to visualize data from sequence analysis programs, allowing the user to view multiple sets of computationally generated analysis to assist the analysis process. SFS is used as the data exchange format. AVAILABILITY: Sfixem is available for direct usage or download for local usage at http://sfixem.cgb.ki.se. A protein sequence analysis workbench using Sfixem is available at http://sfinx.cgb.ki.se.
Alistair M. Chalk, Martin Wennerberg, Erik L. L. Sonnhammer
Bioinform.3
2004 ChromoWheel: a new spin on eukaryotic chromosome visualization
abstract
ChromoWheel is an Internet browser application for generating whole-genome illustrations. It can be used to depict chromosomes, genes and relations between chromosomal loci. The circular layout of chromosomes is advantageous for showing relationships between different chromosomes, as the connecting line never crosses over a chromosome. All graphical image components are in the vector-based format Scalable Vector Graphics, which are highly scaleable and admit user interaction. ChromoWheel can either be run with user-provided data in the generic SFS format, or as a browser front-end for precompiled genomic data.
Sven Ekdahl, Erik L. L. Sonnhammer
Bioinform.2
2004 Profiled support vector machines for antisense oligonucleotide efficacy prediction
abstract
BACKGROUND: This paper presents the use of Support Vector Machines (SVMs) for prediction and analysis of antisense oligonucleotide (AO) efficacy. The collected database comprises 315 AO molecules including 68 features each, inducing a problem well-suited to SVMs. The task of feature selection is crucial given the presence of noisy or redundant features, and the well-known problem of the curse of dimensionality. We propose a two-stage strategy to develop an optimal model: (1) feature selection using correlation analysis, mutual information, and SVM-based recursive feature elimination (SVM-RFE), and (2) AO prediction using standard and profiled SVM formulations. A profiled SVM gives different weights to different parts of the training data to focus the training on the most important regions. RESULTS: In the first stage, the SVM-RFE technique was most efficient and robust in the presence of low number of samples and high input space dimension. This method yielded an optimal subset of 14 representative features, which were all related to energy and sequence motifs. The second stage evaluated the performance of the predictors (overall correlation coefficient between observed and predicted efficacy, r; mean error, ME; and root-mean-square-error, RMSE) using 8-fold and minus-one-RNA cross-validation methods. The profiled SVM produced the best results (r = 0.44, ME = 0.022, and RMSE= 0.278) and predicted high (>75% inhibition of gene expression) and low efficacy (<25%) AOs with a success rate of 83.3% and 82.9%, respectively, which is better than by previous approaches. A web server for AO prediction is available online at http://aosvm.cgb.ki.se/. CONCLUSIONS: The SVM approach is well suited to the AO prediction problem, and yields a prediction accuracy superior to previous methods. The profiled SVM was found to perform better than the standard SVM, suggesting that it could lead to improvements in other prediction problems as well.
Gustau Camps-Valls, Alistair M. Chalk, Antonio J. Serrano, José D. Martín-Guerrero, Erik L. L. Sonnhammer
BMC Bioinform.5
2002 Computational antisense oligo prediction with a neural network model
abstract
MOTIVATION: The expression of a gene can be selectively inhibited by antisense oligonucleotides (AOs) targeting the mRNA. However, if the target site in the mRNA is picked randomly, typically 20% or less of the AOs are effective inhibitors in vivo. The sequence properties that make an AO effective are not well understood, thus many AOs need to be tested to find good inhibitors, which is time consuming and costly. So far computational models have been based exclusively on RNA structure prediction or motif searches while ignoring information from other aspects of AO design into the model. RESULTS: We present a computational model for AO prediction based on a neural network approach using a broad range of input parameters. Collecting sequence and efficacy data from AO scanning experiments in the literature generated a database of 490 AO molecules. Using a set of derived parameters based on AO sequence properties we trained a neural network model. The best model, an ensemble of 10 networks, gave an overall correlation coefficient of 0.30 (p=10(-8)). This model can predict effective AOs (>50% inhibition of gene expression) with a success rate of 92%. Using these thresholds the model predicts on average 12 effective AOs per 1000 base pairs, making it a stringent yet practical method for AO prediction.
Alistair M. Chalk, Erik L. L. Sonnhammer
Bioinform.2
2002 OrthoGUI: graphical presentation of Orthostrapper results
abstract
SUMMARY: Orthostrapper is a program that calculates orthology support values for pairs of sequences in a multiple alignment (Storm and Sonnhammer, Bioinformatics, 18, 92-99, 2002). Here we present OrthoGUI, a web interface and display tool for Orthostrapper analysis. OrthoGUI visualizes the Orthostrapper output in both tabular and tree representations, and can also apply a clustering algorithm to identify groups of multiple orthologs, which are indicated by colour coding. AVAILABILITY: http://www.cgb.ki.se/OrthoGUI CONTACT: [email protected]
Volker Hollich, Christian E. V. Storm, Erik L. L. Sonnhammer
Bioinform.3
2002 Automated ortholog inference from phylogenetic trees and calculation of orthology reliability
abstract
MOTIVATION: Orthologous proteins in different species are likely to have similar biochemical function and biological role. When annotating a newly sequenced genome by sequence homology, the most precise and reliable functional information can thus be derived from orthologs in other species. A standard method of finding orthologs is to compare the sequence tree with the species tree. However, since the topology of phylogenetic tree is not always reliable one might get incorrect assignments. RESULTS: Here we present a novel method that resolves this problem by analyzing a set of bootstrap trees instead of the optimal tree. The frequency of orthology assignments in the bootstrap trees can be interpreted as a support value for the possible orthology of the sequences. Our method is efficient enough to analyze data in the scale of whole genomes. It is implemented in Java and calculates orthology support levels for all pairwise combinations of homologous sequences of two species. The method was tested on simulated datasets and on real data of homologous proteins.
Christian E. V. Storm, Erik L. L. Sonnhammer
Bioinform.2
2001 MEDUSA: large scale automatic selection and visual assessment of PCR primer pairs
abstract
UNLABELLED: MEDUSA is a tool for automatic selection and visual assessment of PCR primer pairs, developed to assist large scale gene expression analysis projects. The system allows specification of constraints of the location and distances between the primers in a pair. For instance, primers in coding, non-coding, exon/intron-spanning regions might be selected. Medusa applies these constraints as a filter to primers predicted by three external programs, and displays the resulting primer pairs graphically in the Blixem (Sonnhammer and Durbin, COMPUT: Appl. Biosci. 10, 301-307, 1994; http://www.cgr.ki.se/cgr/groups/sonnhammer/Blixem.html) viewer. AVAILABILITY: The MEDUSA web server is available at http://www.cgr.ki.se/cgr/MEDUSA. The source code and user information are available at ftp://ftp.cgr.ki.se/pub/prog/medusa.
Raf M. Podowski, Erik L. L. Sonnhammer
Bioinform.2
2001 NIFAS: visual analysis of domain evolution in proteins
abstract
MOTIVATION: Multi-domain proteins have evolved by insertions or deletions of distinct protein domains. Tracing the history of a certain domain combination can be important for functional annotation of multi-domain proteins, and for understanding the function of individual domains. In order to analyze the evolutionary history of the domains in modular proteins it is desirable to inspect a phylogenetic tree based on sequence divergence with the modular architecture of the sequences superimposed on the tree. RESULT: A Java applet, NIFAS, that integrates graphical domain schematics for each sequence in an evolutionary tree was developed. NIFAS retrieves domain information from the Pfam database and uses CLUSTAL W to calculate a tree for a given Pfam domain. The tree can be displayed with symbolic bootstrap values, and to allow the user to focus on a part of the tree, the layout can be altered by swapping nodes, changing the outgroup, and showing/collapsing subtrees. NIFAS is integrated with the Pfam database and is accessible over the internet (http://www.cgr.ki.se/Pfam). As an example, we use NIFAS to analyze the evolution of domains in Protein Kinases C.
Christian E. V. Storm, Erik L. L. Sonnhammer
Bioinform.2
1999 A comparison of sequence and structure protein domain families as a basis for structural genomics
abstract
MOTIVATION: Protein families can be defined based on structure or sequence similarity. We wanted to compare two protein family databases, one based on structural and one on sequence similarity, to investigate to what extent they overlap, the similarity in definition of corresponding families, and to create a list of large protein families with unknown structure as a resource for structural genomics. We also wanted to increase the sensitivity of fold assignment by exploiting protein family HMMs. RESULTS: We compared Pfam, a protein family database based on sequence similarity, to Scop, which is based on structural similarity. We found that 70% of the Scop families exist in Pfam while 57% of the Pfam families exist in Scop. Most families that occur in both databases correspond well to each other, but in some cases they are different. Such cases highlight situations in which structure and sequence approaches differ significantly. The comparison enabled us to compile a list of the largest families that do not occur in Scop; these are suitable targets for structure prediction and determination, and may be useful to guide projects in structural genomics. It can be noted that 13 out of the 20 largest protein families without a known structure are likely transmembrane proteins. We also exploited Pfam to increase the sensitivity of detecting homologs of proteins with known structure, by comparing query sequences to Pfam HMMs that correspond to Scop families. For SWISSPROT+TREMBL, this yielded an increase in fold assignment from 31% to 42% compared to using FASTA only. This method assigned a structure to 22% of the proteins in Saccharomyces cerevisiae, 24% in Escherichia coli, and 16% in Methanococcus jannaschii.
Arne Elofsson, Erik L. L. Sonnhammer
Bioinform.2
1998 A Hidden Markov Model for Predicting Transmembrane Helices in Protein Sequences
Erik L. L. Sonnhammer, Gunnar von Heijne, Anders Krogh
ISMB1
1994 An Expert System for Processing Sequence Homology Data
Erik L. L. Sonnhammer, Richard Durbin
ISMB1
1994 A workbench for large-scale sequence homology analysis
abstract
When routinely analysing very long stretches of DNA sequences produced by genome sequencing projects, detailed analysis of database search results becomes exceedingly time consuming. To reduce the tedious browsing of large quantities of protein similarities, two programs, MSPcrunch and Blixem, were developed, which assist in processing the results from the database search programs in the BLAST suite. MSPcrunch removes biased composition and redundant matches while keeping weak matches that are consistent with a larger gapped alignment. This makes BLAST searching in practice more sensitive and reduces the risk of overlooking distant similarities. Blixem is a multiple sequence alignment viewer for X-windows which makes it significantly easier to scan and evaluate the matches ratified by MSPcrunch. In Blixem, matches to the translated DNA query sequence are simultaneously aligned in three frames. Also, the distribution of matches over the whole DNA query is displayed. Examples of usage are drawn from 36 C. elegans cosmid clones totalling 1.2 megabases, to which these tools were applied.
Erik L. L. Sonnhammer, Richard Durbin
Comput. Appl. Biosci.1
1992 METASIM: object-oriented modelling of cell regulation
abstract
Enzymatic processes and substances are modelled as distinct objects, belonging to a limited number of classes. A set of class definitions in C++ is presented that constitutes an object-oriented programming platform. The latter supports 'biological' data types and functions and facilitates simulation of metabolic and regulatory pathways in living cells. To compute the time-evolution, Euler or Runge-Kutta methods are used, though the latter method compromises a strict object-oriented philosophy. As an example, histone gene expression during embryogenesis of Xenopus laevis is modelled. This object-oriented programming system forms a modelling 'language' which is readily understood by both biochemists and programmers. It allows biological problems to be programmed more easily and correctly and brings the program closer to the biological reality, hence making it more meaningful to bioscientists. Moreover, it can readily be extended to new models by class derivation.
H. J. Stoffers, Erik L. L. Sonnhammer, G. J. Blommestijn, N. J. Raat, Hans V. Westerhoff
Comput. Appl. Biosci.2