EDBT 2026 Demo / reviewers in the wild / expert
Burkhard Rost
dblp:66/159
· DBLP profile ↗
63ranked-venue papers
9as first author
10since 2021 · last 2026
0000-0003-0179-8424ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 61 · 8 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Whole-genome prediction of bacterial pathogenic capacity on novel bacteria using protein language models with PathogenFinder2abstractMOTIVATION: Infectious diseases continue to be a leading cause of mortality and pose a significant global health threat. Thus, the development of tools for surveillance and early detection of emerging pathogens is needed. RESULTS: We introduce PathogenFinder2, a novel, alignment-free, taxonomy-agnostic model for predicting bacterial pathogenic capacity in humans using protein language models. It outperforms previous methods, particularly for novel taxa, and provides interpretable outputs by highlighting proteins most relevant to pathogenic potential. These insights aid the identification of virulence factors, vaccine targets, and infection-related metabolic pathways. Furthermore, we introduce the Bacterial Pathogenic Capacity Landscape, which reveals patterns linked to host condition, infection site, microbial antagonism, and environmental origin. AVAILABILITY: The model is freely available online at https://genepi.dk/pathogenfinder2, or as a standalone program (https://github.com/genomicepidemiology/PathogenFinder2). Alfred Ferrer Florensa, José Juan Almagro Armenteros, Rolf S. Kaas, Philip Thomas Lanken Conradsen Clausen, Henrik Nielsen, Burkhard Rost, Frank Møller Aarestrup |
Bioinform. | 6 |
| 2025 | FlatProt: 2D visualization eases protein structure comparisonabstractBACKGROUND: Understanding and comparing three-dimensional (3D) structures of proteins can advance bioinformatics, molecular biology, and drug discovery. While 3D models offer detailed insights, comparing multiple structures simultaneously remains challenging, especially on two-dimensional (2D) displays. Existing 2D visualization tools lack standardized approaches for pipelined inspection of large protein sets, limiting their utility in large-scale pre-filtering. RESULTS: We introduce FlatProt, a tool designed to complement 3D viewers by enabling standardized 2D visualization of individual protein structures or large sets thereof. By including Foldseek-based family rotation alignment or an inertia-based fallback, FlatProt creates consistent and scalable visual representations for user-defined protein structures. It supports domain-aware decomposition, family-level overlays, and lightweight visual abstraction of secondary structures. FlatProt processes proteins efficiently, as showcased on a subset of the human-proteome. CONCLUSION: FlatProt provides clear, consistent, user-friendly visualizations that support rapid, comparative inspection of protein structures at scale. By bridging the gap between interactive 3D tools and static visual summaries, it enables users to explore conserved features, detect outliers, and prioritize structures for further analysis. AVAILABILITY: GitHub ( https://github.com/t03i/FlatProt ); Zenodo ( https://doi.org/10.5281/zenodo.15697296 ). Tobias Olenyi, Constantin Carl, Tobias Senoner, Ivan Koludarov, Burkhard Rost |
BMC Bioinform. | 5 |
| 2024 | Expert-guided protein language models enable accurate and blazingly fast fitness predictionabstractMOTIVATION: Exhaustive experimental annotation of the effect of all known protein variants remains daunting and expensive, stressing the need for scalable effect predictions. We introduce VespaG, a blazingly fast missense amino acid variant effect predictor, leveraging protein language model (pLM) embeddings as input to a minimal deep learning model. RESULTS: To overcome the sparsity of experimental training data, we created a dataset of 39 million single amino acid variants from the human proteome applying the multiple sequence alignment-based effect predictor GEMME as a pseudo standard-of-truth. This setup increases interpretability compared to the baseline pLM and is easily retrainable with novel or updated pLMs. Assessed against the ProteinGym benchmark (217 multiplex assays of variant effect-MAVE-with 2.5 million variants), VespaG achieved a mean Spearman correlation of 0.48 ± 0.02, matching top-performing methods evaluated on the same data. VespaG has the advantage of being orders of magnitude faster, predicting all mutational landscapes of all proteins in proteomes such as Homo sapiens or Drosophila melanogaster in under 30 min on a consumer laptop (12-core CPU, 16 GB RAM). AVAILABILITY AND IMPLEMENTATION: VespaG is available freely at https://github.com/jschlensok/vespag. The associated training data and predictions are available at https://doi.org/10.5281/zenodo.11085958. Céline Marquet, Julius Schlensok, Marina Abakarova, Burkhard Rost, Elodie Laine |
Bioinform. | 4 |
| 2023 | CATHe: detection of remote homologues for CATH superfamilies using embeddings from protein language modelsabstractMOTIVATION: CATH is a protein domain classification resource that exploits an automated workflow of structure and sequence comparison alongside expert manual curation to construct a hierarchical classification of evolutionary and structural relationships. The aim of this study was to develop algorithms for detecting remote homologues missed by state-of-the-art hidden Markov model (HMM)-based approaches. The method developed (CATHe) combines a neural network with sequence representations obtained from protein language models. It was assessed using a dataset of remote homologues having less than 20% sequence identity to any domain in the training set. RESULTS: The CATHe models trained on 1773 largest and 50 largest CATH superfamilies had an accuracy of 85.6 ± 0.4% and 98.2 ± 0.3%, respectively. As a further test of the power of CATHe to detect more remote homologues missed by HMMs derived from CATH domains, we used a dataset consisting of protein domains that had annotations in Pfam, but not in CATH. By using highly reliable CATHe predictions (expected error rate <0.5%), we were able to provide CATH annotations for 4.62 million Pfam domains. For a subset of these domains from Homo sapiens, we structurally validated 90.86% of the predictions by comparing their corresponding AlphaFold2 structures with structures from the CATH superfamilies to which they were assigned. AVAILABILITY AND IMPLEMENTATION: The code for the developed models is available on https://github.com/vam-sin/CATHe, and the datasets developed in this study can be accessed on https://zenodo.org/record/6327572. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Vamsi Nallapareddy, Nicola Bordin, Ian Sillitoe, Michael Heinzinger, Maria Littmann, Vaishali P. Waman, Neeladri Sen, Burkhard Rost, Christine A. Orengo |
Bioinform. | 8 |
| 2023 | Rendering protein mutation movies with MutAmoreabstractBACKGROUND: The success of AlphaFold2 in reliable protein three-dimensional (3D) structure prediction, assists the move of structural biology toward studies of protein dynamics and mutational impact on structure and function. This transition needs tools that qualitatively assess alternative 3D conformations. RESULTS: We introduce MutAmore, a bioinformatics tool that renders individual images of protein 3D structures for, e.g., sequence mutations into a visually intuitive movie format. MutAmore streamlines a pipeline casting single amino-acid variations (SAVs) into a dynamic 3D mutation movie providing a qualitative perspective on the mutational landscape of a protein. By default, the tool first generates all possible variants of the sequence reachable through SAVs (L*19 for proteins with L residues). Next, it predicts the structural conformation for all L*19 variants using state-of-the-art models. Finally, it visualizes the mutation matrix and produces a color-coded 3D animation. Alternatively, users can input other types of variants, e.g., from experimental structures. CONCLUSION: MutAmore samples alternative protein configurations to study the dynamical space accessible from SAVs in the post-AlphaFold2 era of structural biology. As the field shifts towards the exploration of alternative conformations of proteins, MutAmore aids in the understanding of the structural impact of mutations by providing a flexible pipeline for the generation of protein mutation movies using current and future structure prediction models. Konstantin Weissenow, Burkhard Rost |
BMC Bioinform. | 2 |
| 2022 | TMbed: transmembrane proteins predicted through language model embeddingsabstractBACKGROUND: Despite the immense importance of transmembrane proteins (TMP) for molecular biology and medicine, experimental 3D structures for TMPs remain about 4-5 times underrepresented compared to non-TMPs. Today's top methods such as AlphaFold2 accurately predict 3D structures for many TMPs, but annotating transmembrane regions remains a limiting step for proteome-wide predictions. RESULTS: Here, we present TMbed, a novel method inputting embeddings from protein Language Models (pLMs, here ProtT5), to predict for each residue one of four classes: transmembrane helix (TMH), transmembrane strand (TMB), signal peptide, or other. TMbed completes predictions for entire proteomes within hours on a single consumer-grade desktop machine at performance levels similar or better than methods, which are using evolutionary information from multiple sequence alignments (MSAs) of protein families. On the per-protein level, TMbed correctly identified 94 ± 8% of the beta barrel TMPs (53 of 57) and 98 ± 1% of the alpha helical TMPs (557 of 571) in a non-redundant data set, at false positive rates well below 1% (erred on 30 of 5654 non-membrane proteins). On the per-segment level, TMbed correctly placed, on average, 9 of 10 transmembrane segments within five residues of the experimental observation. Our method can handle sequences of up to 4200 residues on standard graphics cards used in desktop PCs (e.g., NVIDIA GeForce RTX 3060). CONCLUSIONS: Based on embeddings from pLMs and two novel filters (Gaussian and Viterbi), TMbed predicts alpha helical and beta barrel TMPs at least as accurately as any other method but at lower false positive rates. Given the few false positives and its outstanding speed, TMbed might be ideal to sieve through millions of 3D structures soon to be predicted, e.g., by AlphaFold2. Michael Bernhofer, Burkhard Rost |
BMC Bioinform. | 2 |
| 2022 | ProtTrans: Toward Understanding the Language of Life Through Self-Supervised LearningabstractComputational biology and bioinformatics provide vast data gold-mines from protein sequences, ideal for Language Models (LMs) taken from Natural Language Processing (NLP). These LMs reach for new prediction frontiers at low inference costs. Here, we trained two auto-regressive models (Transformer-XL, XLNet) and four auto-encoder models (BERT, Albert, Electra, T5) on data from UniRef and BFD containing up to 393 billion amino acids. The protein LMs (pLMs) were trained on the Summit supercomputer using 5616 GPUs and TPU Pod up-to 1024 cores. Dimensionality reduction revealed that the raw pLM-embeddings from unlabeled data captured some biophysical features of protein sequences. We validated the advantage of using the embeddings as exclusive input for several subsequent tasks: (1) a per-residue (per-token) prediction of protein secondary structure (3-state accuracy Q3=81%-87%); (2) per-protein (pooling) predictions of protein sub-cellular location (ten-state accuracy: Q10=81%) and membrane versus water-soluble (2-state accuracy Q2=91%). For secondary structure, the most informative embeddings (ProtT5) for the first time outperformed the state-of-the-art without multiple sequence alignments (MSAs) or evolutionary information thereby bypassing expensive database searches. Taken together, the results implied that pLMs learned some of the grammar of the language of life. All our models are available through https://github.com/agemagician/ProtTrans. Ahmed Elnaggar, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Yu Wang 0008, Llion Jones, Tom Gibbs, Tamas Feher, Christoph Angerer, Martin Steinegger, Debsindhu Bhowmik, Burkhard Rost |
IEEE Trans. Pattern Anal. Mach. Intell. | 12 |
| 2022 | Engineering indel and substitution variants of diverse and ancient enzymes using Graphical Representation of Ancestral Sequence Predictions (GRASP)abstractAncestral sequence reconstruction is a technique that is gaining widespread use in molecular evolution studies and protein engineering. Accurate reconstruction requires the ability to handle appropriately large numbers of sequences, as well as insertion and deletion (indel) events, but available approaches exhibit limitations. To address these limitations, we developed Graphical Representation of Ancestral Sequence Predictions (GRASP), which efficiently implements maximum likelihood methods to enable the inference of ancestors of families with more than 10,000 members. GRASP implements partial order graphs (POGs) to represent and infer insertion and deletion events across ancestors, enabling the identification of building blocks for protein engineering. To validate the capacity to engineer novel proteins from realistic data, we predicted ancestor sequences across three distinct enzyme families: glucose-methanol-choline (GMC) oxidoreductases, cytochromes P450, and dihydroxy/sugar acid dehydratases (DHAD). All tested ancestors demonstrated enzymatic activity. Our study demonstrates the ability of GRASP (1) to support large data sets over 10,000 sequences and (2) to employ insertions and deletions to identify building blocks for engineering biologically active ancestors, by exploring variation over evolutionary time. Gabriel Foley, Ariane Mora, Connie M. Ross, Scott Bottoms, Leander Sützl, Marnie L. Lamprecht, Julian Zaugg, Alexandra Essebier, Brad Balderson, Rhys Newell, Raine E. S. Thomson, Bostjan Kobe, Ross T. Barnard, Luke Guddat, Gerhard Schenk, Jörg Carsten, Yosephine Gumulya, Burkhard Rost, Dietmar Haltrich, Volker Sieber, Elizabeth M. J. Gillam, Mikael Bodén |
PLoS Comput. Biol. | 18 |
| 2021 | Mutations in transmembrane proteins: diseases, evolutionary insights, prediction and comparison with globular proteinsabstractMembrane proteins are unique in that they interact with lipid bilayers, making them indispensable for transporting molecules and relaying signals between and across cells. Due to the significance of the protein's functions, mutations often have profound effects on the fitness of the host. This is apparent both from experimental studies, which implicated numerous missense variants in diseases, as well as from evolutionary signals that allow elucidating the physicochemical constraints that intermembrane and aqueous environments bring. In this review, we report on the current state of knowledge acquired on missense variants (referred to as to single amino acid variants) affecting membrane proteins as well as the insights that can be extrapolated from data already available. This includes an overview of the annotations for membrane protein variants that have been collated within databases dedicated to the topic, bioinformatics approaches that leverage evolutionary information in order to shed light on previously uncharacterized membrane protein structures or interaction interfaces, tools for predicting the effects of mutations tailored specifically towards the characteristics of membrane proteins as well as two clinically relevant case studies explaining the implications of mutated membrane proteins in cancer and cardiomyopathy. Jan Zaucha, Michael Heinzinger, A. Kulandaisamy, Evans Kataka, Óscar Llorian Salvádor, Petr Popov, Burkhard Rost, M. Michael Gromiha, Boris S. Zhorov, Dmitrij Frishman |
Briefings Bioinform. | 7 |
| 2021 | Clustering FunFams using sequence embeddings improves EC purityabstractMOTIVATION: Classifying proteins into functional families can improve our understanding of protein function and can allow transferring annotations within one family. For this, functional families need to be 'pure', i.e., contain only proteins with identical function. Functional Families (FunFams) cluster proteins within CATH superfamilies into such groups of proteins sharing function. 11% of all FunFams (22 830 of 203 639) contain EC annotations and of those, 7% (1526 of 22 830) have inconsistent functional annotations. RESULTS: We propose an approach to further cluster FunFams into functionally more consistent sub-families by encoding their sequences through embeddings. These embeddings originate from language models transferring knowledge gained from predicting missing amino acids in a sequence (ProtBERT) and have been further optimized to distinguish between proteins belonging to the same or a different CATH superfamily (PB-Tucker). Using distances between embeddings and DBSCAN to cluster FunFams and identify outliers, doubled the number of pure clusters per FunFam compared to random clustering. Our approach was not limited to FunFams but also succeeded on families created using sequence similarity alone. Complementing EC annotations, we observed similar results for binding annotations. Thus, we expect an increased purity also for other aspects of function. Our results can help generating FunFams; the resulting clusters with improved functional consistency allow more reliable inference of annotations. We expect this approach to succeed equally for any other grouping of proteins by their phenotypes. AVAILABILITY AND IMPLEMENTATION: Code and embeddings are available via GitHub: https://github.com/Rostlab/FunFamsClustering. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Maria Littmann, Nicola Bordin, Michael Heinzinger, Konstantin Schütze, Christian Dallago, Christine A. Orengo, Burkhard Rost |
Bioinform. | 7 |
| 2020 | Protein-protein and protein-nucleic acid binding residues important for common and rare sequence variants in humanabstractAbstract Background Any two unrelated people differ by about 20,000 missense mutations (also referred to as SAVs: Single Amino acid Variants or missense SNV). Many SAVs have been predicted to strongly affect molecular protein function. Common SAVs (> 5% of population) were predicted to have, on average, more effect on molecular protein function than rare SAVs (< 1% of population). We hypothesized that the prevalence of effect in common over rare SAVs might partially be caused by common SAVs more often occurring at interfaces of proteins with other proteins, DNA, or RNA, thereby creating subgroup-specific phenotypes. We analyzed SAVs from 60,706 people through the lens of two prediction methods, one (SNAP2) predicting the effects of SAVs on molecular protein function, the other (ProNA2020) predicting residues in DNA-, RNA- and protein-binding interfaces. Results Three results stood out. Firstly, SAVs predicted to occur at binding interfaces were predicted to more likely affect molecular function than those predicted as not binding (p value < 2.2 × 10–16). Secondly, for SAVs predicted to occur at binding interfaces, common SAVs were predicted more strongly with effect on protein function than rare SAVs (p value < 2.2 × 10–16). Restriction to SAVs with experimental annotations confirmed all results, although the resulting subsets were too small to establish statistical significance for any result. Thirdly, the fraction of SAVs predicted at binding interfaces differed significantly between tissues, e.g. urinary bladder tissue was found abundant in SAVs predicted at protein-binding interfaces, and reproductive tissues (ovary, testis, vagina, seminal vesicle and endometrium) in SAVs predicted at DNA-binding interfaces. Conclusions Overall, the results suggested that residues at protein-, DNA-, and RNA-binding interfaces contributed toward predicting that common SAVs more likely affect molecular function than rare SAVs. Jiajun Qiu, Dmitrii Nechaev, Burkhard Rost |
BMC Bioinform. | 3 |
| 2020 | Variant effect predictions capture some aspects of deep mutational scanning experimentsabstractBACKGROUND: Deep mutational scanning (DMS) studies exploit the mutational landscape of sequence variation by systematically and comprehensively assaying the effect of single amino acid variants (SAVs; also referred to as missense mutations, or non-synonymous Single Nucleotide Variants - missense SNVs or nsSNVs) for particular proteins. We assembled SAV annotations from 22 different DMS experiments and normalized the effect scores to evaluate variant effect prediction methods. Three trained on traditional variant effect data (PolyPhen-2, SIFT, SNAP2), a regression method optimized on DMS data (Envision), and a naïve prediction using conservation information from homologs. RESULTS: On a set of 32,981 SAVs, all methods captured some aspects of the experimental effect scores, albeit not the same. Traditional methods such as SNAP2 correlated slightly more with measurements and better classified binary states (effect or neutral). Envision appeared to better estimate the precise degree of effect. Most surprising was that the simple naïve conservation approach using PSI-BLAST in many cases outperformed other methods. All methods captured beneficial effects (gain-of-function) significantly worse than deleterious (loss-of-function). For the few proteins with multiple independent experimental measurements, experiments differed substantially, but agreed more with each other than with predictions. CONCLUSIONS: DMS provides a new powerful experimental means of understanding the dynamics of the protein sequence space. As always, promising new beginnings have to overcome challenges. While our results demonstrated that DMS will be crucial to improve variant effect prediction methods, data diversity hindered simplification and generalization. Jonas Reeb, Theresa Wirth, Burkhard Rost |
BMC Bioinform. | 3 |
| 2019 | Modeling aspects of the language of life through transfer-learning protein sequencesabstractBACKGROUND: Predicting protein function and structure from sequence is one important challenge for computational biology. For 26 years, most state-of-the-art approaches combined machine learning and evolutionary information. However, for some applications retrieving related proteins is becoming too time-consuming. Additionally, evolutionary information is less powerful for small families, e.g. for proteins from the Dark Proteome. Both these problems are addressed by the new methodology introduced here. RESULTS: We introduced a novel way to represent protein sequences as continuous vectors (embeddings) by using the language model ELMo taken from natural language processing. By modeling protein sequences, ELMo effectively captured the biophysical properties of the language of life from unlabeled big data (UniRef50). We refer to these new embeddings as SeqVec (Sequence-to-Vector) and demonstrate their effectiveness by training simple neural networks for two different tasks. At the per-residue level, secondary structure (Q3 = 79% ± 1, Q8 = 68% ± 1) and regions with intrinsic disorder (MCC = 0.59 ± 0.03) were predicted significantly better than through one-hot encoding or through Word2vec-like approaches. At the per-protein level, subcellular localization was predicted in ten classes (Q10 = 68% ± 1) and membrane-bound were distinguished from water-soluble proteins (Q2 = 87% ± 1). Although SeqVec embeddings generated the best predictions from single sequences, no solution improved over the best existing method using evolutionary information. Nevertheless, our approach improved over some popular methods using evolutionary information and for some proteins even did beat the best. Thus, they prove to condense the underlying principles of protein sequences. Overall, the important novelty is speed: where the lightning-fast HHblits needed on average about two minutes to generate the evolutionary information for a target protein, SeqVec created embeddings on average in 0.03 s. As this speed-up is independent of the size of growing sequence databases, SeqVec provides a highly scalable approach for the analysis of big data in proteomics, i.e. microbiome or metaproteome analysis. CONCLUSION: Transfer-learning succeeded to extract information from unlabeled sequence databases relevant for various protein prediction tasks. SeqVec modeled the language of life, namely the principles underlying protein sequences better than any features suggested by textbooks and prediction methods. The exception is evolutionary information, however, that information is not available on the level of a single sequence. Michael Heinzinger, Ahmed Elnaggar, Yu Wang 0008, Christian Dallago, Dmitrii Nechaev, Florian Matthes, Burkhard Rost |
BMC Bioinform. | 7 |
| 2019 | Detailed prediction of protein sub-nuclear localizationabstractBACKGROUND: Sub-nuclear structures or locations are associated with various nuclear processes. Proteins localized in these substructures are important to understand the interior nuclear mechanisms. Despite advances in high-throughput methods, experimental protein annotations remain limited. Predictions of cellular compartments have become very accurate, largely at the expense of leaving out substructures inside the nucleus making a fine-grained analysis impossible. RESULTS: Here, we present a new method (LocNuclei) that predicts nuclear substructures from sequence alone. LocNuclei used a string-based Profile Kernel with Support Vector Machines (SVMs). It distinguishes sub-nuclear localization in 13 distinct substructures and distinguishes between nuclear proteins confined to the nucleus and those that are also native to other compartments (traveler proteins). High performance was achieved by implicitly leveraging a large biological knowledge-base in creating predictions by homology-based inference through BLAST. Using this approach, the performance reached AUC = 0.70-0.74 and Q13 = 59-65%. Travelling proteins (nucleus and other) were identified at Q2 = 70-74%. A Gene Ontology (GO) analysis of the enrichment of biological processes revealed that the predicted sub-nuclear compartments matched the expected functionality. Analysis of protein-protein interactions (PPI) show that formation of compartments and functionality of proteins in these compartments highly rely on interactions between proteins. This suggested that the LocNuclei predictions carry important information about function. The source code and data sets are available through GitHub: https://github.com/Rostlab/LocNuclei . CONCLUSIONS: LocNuclei predicts subnuclear compartments and traveler proteins accurately. These predictions carry important information about functionality and PPIs. Maria Littmann, Tatyana Goldberg, Sebastian Seitz, Mikael Bodén, Burkhard Rost |
BMC Bioinform. | 5 |
| 2019 | Correction to: Detailed prediction of protein sub-nuclear localizationabstractFollowing publication of the original article [1], the author reported that an incorrect figure has been published as Figure 2. The correct Figure 2 is shown below. Maria Littmann, Tatyana Goldberg, Sebastian Seitz, Mikael Bodén, Burkhard Rost |
BMC Bioinform. | 5 |
| 2019 | FunFam protein families improve residue level molecular function predictionabstractBACKGROUND: The CATH database provides a hierarchical classification of protein domain structures including a sub-classification of superfamilies into functional families (FunFams). We analyzed the similarity of binding site annotations in these FunFams and incorporated FunFams into the prediction of protein binding residues. RESULTS: FunFam members agreed, on average, in 36.9 ± 0.6% of their binding residue annotations. This constituted a 6.7-fold increase over randomly grouped proteins and a 1.2-fold increase (1.1-fold on the same dataset) over proteins with the same enzymatic function (identical Enzyme Commission, EC, number). Mapping de novo binding residue prediction methods (BindPredict-CCS, BindPredict-CC) onto FunFam resulted in consensus predictions for those residues that were aligned and predicted alike (binding/non-binding) within a FunFam. This simple consensus increased the F1-score (for binding) 1.5-fold over the original prediction method. Variation of the threshold for how many proteins in the consensus prediction had to agree provided a convenient control of accuracy/precision and coverage/recall, e.g. reaching a precision as high as 60.8 ± 0.4% for a stringent threshold. CONCLUSIONS: The FunFams outperformed even the carefully curated EC numbers in terms of agreement of binding site residues. Additionally, we assume that our proof-of-principle through the prediction of protein binding residues will be relevant for many other solutions profiting from FunFams to infer functional information at the residue level. Linus Scheibenreif, Maria Littmann, Christine A. Orengo, Burkhard Rost |
BMC Bioinform. | 4 |
| 2018 | HFSP: high speed homology-driven function annotation of proteinsabstractMotivation: The rapid drop in sequencing costs has produced many more (predicted) protein sequences than can feasibly be functionally annotated with wet-lab experiments. Thus, many computational methods have been developed for this purpose. Most of these methods employ homology-based inference, approximated via sequence alignments, to transfer functional annotations between proteins. The increase in the number of available sequences, however, has drastically increased the search space, thus significantly slowing down alignment methods. Results: Here we describe homology-derived functional similarity of proteins (HFSP), a novel computational method that uses results of a high-speed alignment algorithm, MMseqs2, to infer functional similarity of proteins on the basis of their alignment length and sequence identity. We show that our method is accurate (85% precision) and fast (more than 40-fold speed increase over state-of-the-art). HFSP can help correct at least a 16% error in legacy curations, even for a resource of as high quality as Swiss-Prot. These findings suggest HFSP as an ideal resource for large-scale functional annotation efforts. Supplementary information: Supplementary data are available at Bioinformatics online. Yannick Mahlich, Martin Steinegger, Burkhard Rost, Yana Bromberg |
Bioinform. | 3 |
| 2018 | Correcting mistakes in predicting distributionsabstractMotivation: Many applications monitor predictions of a whole range of features for biological datasets, e.g. the fraction of secreted human proteins in the human proteome. Results and error estimates are typically derived from publications. Results: Here, we present a simple, alternative approximation that uses performance estimates of methods to error-correct the predicted distributions. This approximation uses the confusion matrix (TP true positives, TN true negatives, FP false positives and FN false negatives) describing the performance of the prediction tool for correction. As proof-of-principle, the correction was applied to a two-class (membrane/not) and to a seven-class (localization) prediction. Availability and implementation: Datasets and a simple JavaScript tool available freely for all users at http://www.rostlab.org/services/distributions. Supplementary information: Supplementary data are available at Bioinformatics online. Valérie Marot-Lassauzaie, Michael Bernhofer, Burkhard Rost |
Bioinform. | 3 |
| 2018 | LocText: relation extraction of protein localizations to assist database curationabstractBACKGROUND: The subcellular localization of a protein is an important aspect of its function. However, the experimental annotation of locations is not even complete for well-studied model organisms. Text mining might aid database curators to add experimental annotations from the scientific literature. Existing extraction methods have difficulties to distinguish relationships between proteins and cellular locations co-mentioned in the same sentence. RESULTS: LocText was created as a new method to extract protein locations from abstracts and full texts. LocText learned patterns from syntax parse trees and was trained and evaluated on a newly improved LocTextCorpus. Combined with an automatic named-entity recognizer, LocText achieved high precision (P = 86%±4). After completing development, we mined the latest research publications for three organisms: human (Homo sapiens), budding yeast (Saccharomyces cerevisiae), and thale cress (Arabidopsis thaliana). Examining 60 novel, text-mined annotations, we found that 65% (human), 85% (yeast), and 80% (cress) were correct. Of all validated annotations, 40% were completely novel, i.e. did neither appear in the annotations nor the text descriptions of Swiss-Prot. CONCLUSIONS: LocText provides a cost-effective, semi-automated workflow to assist database curators in identifying novel protein localization annotations. The annotations suggested through text-mining would be verified by experts to guarantee high-quality standards of manually-curated databases such as Swiss-Prot. Juan Miguel Cejuela, Shrikant Vinchurkar, Tatyana Goldberg, Madhukar Sollepura Prabhu Shankar, Ashish Baghudana, Aleksandar Bojchevski, Carsten Uhlig, André Ofner, Pandu Raharja-Liu, Lars Juhl Jensen, Burkhard Rost |
BMC Bioinform. | 11 |
| 2017 | nala: text mining natural language mutation mentionsabstractMOTIVATION: The extraction of sequence variants from the literature remains an important task. Existing methods primarily target standard (ST) mutation mentions (e.g. 'E6V'), leaving relevant mentions natural language (NL) largely untapped (e.g. 'glutamic acid was substituted by valine at residue 6'). RESULTS: We introduced three new corpora suggesting named-entity recognition (NER) to be more challenging than anticipated: 28-77% of all articles contained mentions only available in NL. Our new method nala captured NL and ST by combining conditional random fields with word embedding features learned unsupervised from the entire PubMed. In our hands, nala substantially outperformed the state-of-the-art. For instance, we compared all unique mentions in new discoveries correctly detected by any of three methods (SETH, tmVar, or nala ). Neither SETH nor tmVar discovered anything missed by nala , while nala uniquely tagged 33% mentions. For NL mentions the corresponding value shot up to 100% nala -only. AVAILABILITY AND IMPLEMENTATION: Source code, API and corpora freely available at: http://tagtog.net/-corpora/IDP4+ . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Juan Miguel Cejuela, Aleksandar Bojchevski, Carsten Uhlig, Rustem Bekmukhametov, Sanjeev Kumar Karn, Shpend Mahmuti, Ashish Baghudana, Ankit Dubey, Venkata P. Satagopam, Burkhard Rost |
Bioinform. | 10 |
| 2016 | MSAViewer: interactive JavaScript visualization of multiple sequence alignmentsabstractThe MSAViewer is a quick and easy visualization and analysis JavaScript component for Multiple Sequence Alignment data of any size. Core features include interactive navigation through the alignment, application of popular color schemes, sorting, selecting and filtering. The MSAViewer is 'web ready': written entirely in JavaScript, compatible with modern web browsers and does not require any specialized software. The MSAViewer is part of the BioJS collection of components. AVAILABILITY AND IMPLEMENTATION: The MSAViewer is released as open source software under the Boost Software License 1.0. Documentation, source code and the viewer are available at http://msa.biojs.net/Supplementary information: Supplementary data are available at Bioinformatics online. CONTACT: [email protected]. Guy Yachdav, Sebastian Wilzbach, Benedikt Rauscher, Robert Sheridan, Ian Sillitoe, James B. Procter, Suzanna Lewis, Burkhard Rost, Tatyana Goldberg |
Bioinform. | 8 |
| 2016 | Predicted Molecular Effects of Sequence Variants Link to System Level of DiseaseabstractDevelopments in experimental and computational biology are advancing our understanding of how protein sequence variation impacts molecular protein function. However, the leap from the micro level of molecular function to the macro level of the whole organism, e.g. disease, remains barred. Here, we present new results emphasizing earlier work that suggested some links from molecular function to disease. We focused on non-synonymous single nucleotide variants, also referred to as single amino acid variants (SAVs). Building upon OMIA (Online Mendelian Inheritance in Animals), we introduced a curated set of 117 disease-causing SAVs in animals. Methods optimized to capture effects upon molecular function often correctly predict human (OMIM) and animal (OMIA) Mendelian disease-causing variants. We also predicted effects of human disease-causing variants in the mouse model, i.e. we put OMIM SAVs into mouse orthologs. Overall, fewer variants were predicted with effect in the model organism than in the original organism. Our results, along with other recent studies, demonstrate that predictions of molecular effects capture some important aspects of disease. Thus, in silico methods focusing on the micro level of molecular function can help to understand the macro system level of disease. Jonas Reeb, Maximilian Hecht, Yannick Mahlich, Yana Bromberg, Burkhard Rost |
PLoS Comput. Biol. | 5 |
| 2015 | More challenges for machine-learning protein interactionsabstractMOTIVATION: Machine learning may be the most popular computational tool in molecular biology. Providing sustained performance estimates is challenging. The standard cross-validation protocols usually fail in biology. Park and Marcotte found that even refined protocols fail for protein-protein interactions (PPIs). RESULTS: Here, we sketch additional problems for the prediction of PPIs from sequence alone. First, it not only matters whether proteins A or B of a target interaction A-B are similar to proteins of training interactions (positives), but also whether A or B are similar to proteins of non-interactions (negatives). Second, training on multiple interaction partners per protein did not improve performance for new proteins (not used to train). In contrary, a strictly non-redundant training that ignored good data slightly improved the prediction of difficult cases. Third, which prediction method appears to be best crucially depends on the sequence similarity between the test and the training set, how many true interactions should be found and the expected ratio of negatives to positives. The correct assessment of performance is the most complicated task in the development of prediction methods. Our analyses suggest that PPIs square the challenge for this task. Tobias Hamp, Burkhard Rost |
Bioinform. | 2 |
| 2015 | Evolutionary profiles improve protein-protein interaction prediction from sequenceabstractMOTIVATION: Many methods predict the physical interaction between two proteins (protein-protein interactions; PPIs) from sequence alone. Their performance drops substantially for proteins not used for training. RESULTS: Here, we introduce a new approach to predict PPIs from sequence alone which is based on evolutionary profiles and profile-kernel support vector machines. It improved over the state-of-the-art, in particular for proteins that are sequence-dissimilar to proteins with known interaction partners. Filtering by gene expression data increased accuracy further for the few, most reliably predicted interactions (low recall). The overall improvement was so substantial that we compiled a list of the most reliably predicted PPIs in human. Our method makes a significant difference for biology because it improves most for the majority of proteins without experimental annotations. AVAILABILITY AND IMPLEMENTATION: Implementation and most reliably predicted human PPIs available at https://rostlab.org/owiki/index.php/Profppikernel. Tobias Hamp, Burkhard Rost |
Bioinform. | 2 |
| 2015 | Message from the ISCB: ISCB Ebola award for important future research on the computational biology of Ebola virusabstractUNLABELLED: Speed is of the essence in combating Ebola; thus, computational approaches should form a significant component of Ebola research. As for the development of any modern drug, computational biology is uniquely positioned to contribute through comparative analysis of the genome sequences of Ebola strains and three-dimensional protein modeling. Other computational approaches to Ebola may include large-scale docking studies of Ebola proteins with human proteins and with small-molecule libraries, computational modeling of the spread of the virus, computational mining of the Ebola literature and creation of a curated Ebola database. Taken together, such computational efforts could significantly accelerate traditional scientific approaches. In recognition of the need for important and immediate solutions from the field of computational biology against Ebola, the International Society for Computational Biology (ISCB) announces a prize for an important computational advance in fighting the Ebola virus. ISCB will confer the ISCB Fight against Ebola Award, along with a prize of US$2000, at its July 2016 annual meeting (ISCB Intelligent Systems for Molecular Biology 2016, Orlando, FL). CONTACT: [email protected] or [email protected]. Peter D. Karp, Bonnie Berger, Diane E. Kovats, Thomas Lengauer, Michal Linial, Pardis Sabeti, Winston Hide, Burkhard Rost |
Bioinform. | 8 |
| 2015 | ISCB Ebola Award for Important Future Research on the Computational Biology of Ebola VirusabstractSpeed is of the essence in combating Ebola; thus, computational approaches should form a significant component of Ebola research.As for the development of any modern drug, computational biology is uniquely positioned to contribute through comparative analysis of the genome sequences of Ebola strains as well as 3-D protein modeling.Other computational approaches to Ebola may include large-scale docking studies of Ebola proteins with human proteins and with small-molecule libraries, computational modeling of the spread of the virus, computational mining of the Ebola literature, and creation of a curated Ebola database.Taken together, such computational efforts could significantly accelerate traditional scientific approaches.In recognition of the need for important and immediate solutions from the field of computational biology against Ebola, the International Society for Computational Biology (ISCB) announces a prize for an important computational advance in fighting the Ebola virus.ISCB will confer the ISCB Fight against Ebola Award, along with a prize of US$2,000, at its July 2016 annual meeting (ISCB Intelligent Systems for Molecular Biology [ISMB] 2016, Orlando, Florida). Peter D. Karp, Bonnie Berger, Diane E. Kovats, Thomas Lengauer, Michal Linial, Pardis Sabeti, Winston Hide, Burkhard Rost |
PLoS Comput. Biol. | 8 |
| 2014 | ISCB: past-present perspective for the International Society for Computational BiologyabstractAbstract Since its establishment in 1997, International Society for Computational Biology (ISCB) has contributed importantly toward advancing the understanding of living systems through computation. The ISCB represents nearly 3000 members working in >70 countries. It has doubled the number of members since 2007. At the same time, the number of meetings organized by the ISCB has increased from two in 2007 to eight in 2013, and the society has cemented many lasting alliances with regional societies and specialist groups. ISCB is ready to grow into a challenging and promising future. The progress over the past 7 years has resulted from the vision, and possibly more importantly, the passion and hard working dedication of many individuals. Burkhard Rost |
Bioinform. | 1 |
| 2014 | FreeContact: fast and free software for protein contact prediction from residue co-evolutionabstractBACKGROUND: 20 years of improved technology and growing sequences now renders residue-residue contact constraints in large protein families through correlated mutations accurate enough to drive de novo predictions of protein three-dimensional structure. The method EVfold broke new ground using mean-field Direct Coupling Analysis (EVfold-mfDCA); the method PSICOV applied a related concept by estimating a sparse inverse covariance matrix. Both methods (EVfold-mfDCA and PSICOV) are publicly available, but both require too much CPU time for interactive applications. On top, EVfold-mfDCA depends on proprietary software. RESULTS: Here, we present FreeContact, a fast, open source implementation of EVfold-mfDCA and PSICOV. On a test set of 140 proteins, FreeContact was almost eight times faster than PSICOV without decreasing prediction performance. The EVfold-mfDCA implementation of FreeContact was over 220 times faster than PSICOV with negligible performance decrease. EVfold-mfDCA was unavailable for testing due to its dependency on proprietary software. FreeContact is implemented as the free C++ library "libfreecontact", complete with command line tool "freecontact", as well as Perl and Python modules. All components are available as Debian packages. FreeContact supports the BioXSD format for interoperability. CONCLUSIONS: FreeContact provides the opportunity to compute reliable contact predictions in any environment (desktop or cloud). László Kaján, Thomas A. Hopf, Matús Kalas, Debora S. Marks, Burkhard Rost |
BMC Bioinform. | 5 |
| 2013 | ISCB: past-present perspective for the International Society for Computational BiologyabstractSince its establishment in 1997, International Society for Computational Biology (ISCB) has contributed importantly toward advancing the understanding of living systems through computation. The ISCB represents nearly 3000 members working in >70 countries. It has doubled the number of members since 2007. At the same time, the number of meetings organized by the ISCB has increased from two in 2007 to eight in 2013, and the society has cemented many lasting alliances with regional societies and specialist groups. ISCB is ready to grow into a challenging and promising future. The progress over the past 7 years has resulted from the vision, and possibly more importantly, the passion and hard working dedication of many individuals. Burkhard Rost |
Bioinform. | 1 |
| 2013 | Homology-based inference sets the bar high for protein function predictionabstractBACKGROUND: Any method that de novo predicts protein function should do better than random. More challenging, it also ought to outperform simple homology-based inference. METHODS: Here, we describe a few methods that predict protein function exclusively through homology. Together, they set the bar or lower limit for future improvements. RESULTS AND CONCLUSIONS: During the development of these methods, we faced two surprises. Firstly, our most successful implementation for the baseline ranked very high at CAFA1. In fact, our best combination of homology-based methods fared only slightly worse than the top-of-the-line prediction method from the Jones group. Secondly, although the concept of homology-based inference is simple, this work revealed that the precise details of the implementation are crucial: not only did the methods span from top to bottom performers at CAFA, but also the reasons for these differences were unexpected. In this work, we also propose a new rigorous measure to compare predicted and experimental annotations. It puts more emphasis on the details of protein function than the other measures employed by CAFA and may best reflect the expectations of users. Clearly, the definition of proper goals remains one major objective for CAFA. Tobias Hamp, Rebecca Kassner, Stefan Seemayer, Esmeralda Vicedo, Christian Schaefer, Dominik Achten, Florian Auer, Ariane Boehm, Tatjana Braun, Maximilian Hecht, Mark Heron, Peter Hönigschmid, Thomas A. Hopf, Stefanie Kaufmann, Michael Kiening, Denis Krompass, Cedric Landerer, Yannick Mahlich, Manfred Roos, Burkhard Rost |
BMC Bioinform. | 20 |
| 2013 | ISCB Computational Biology Wikipedia CompetitionabstractThe International Society for Computational Biology is pleased to announce the 2013 ISCB Computational Biology Wikipedia competition. The competition, in which entrants create or improve the content of any Wikipedia article in the field of computational biology, is open to all students and trainees. Further information about the competition can be found here: http://en.wikipedia.org/wiki/Wikipedia:WikiProject_Computational_Biology/ISCB_competition_announcement_2013
The mission of the ISCB is to promote the use of computational biology and to help educate the next generation of computational biologists. The society has numerous activities that help to address these aims, including conferences, training and mentoring initiatives, and an active student council.
As the world's largest online encyclopedia, Wikipedia has become an indispensable resource for those seeking information on all scientific and technical topics. The English language version of Wikipedia contains over 4.2 million articles, and Wikipedia is now available in 286 languages. The global rise in smartphone use, which allows access to Wikipedia, means that a large fraction of the world's population can now gain access to the world's knowledge. Wikipedia is the most successful example of crowd-sourcing with about 80,000 active editors updating its content each month.
But is Wikipedia a good source of information for computational biology? Certainly, many people are reading the articles. For example, the Bioinformatics article has been visited 1,600 times per day over the last 3 months. Wikipedia contains articles on algorithms, biological databases, software packages, and biographies of eminent computational biologists. The computational biology content ranges from incomplete, a mere “stub” of an article in Wikipedia parlance, to highly detailed Featured Articles. A group of Wikipedia editors have formed the Computational Biology Wikiproject (http://en.wikipedia.org/wiki/Wikipedia:WikiProject_Computational_Biology). This group oversees the computational biology articles and rates them for their importance and their quality. Figure 1 shows the current state of the articles (see also Figure S1). In total, there are over 1,140 articles that have been considered as falling under Computational Biology. There are a small number of articles that have been brought up to the highest levels of quality (Featured Article and Good Article) such as Multiple Sequence Alignment, Genome Wide Association Study, and Folding@home.
Figure 1
The computational biology articles rated by quality and importance by the Wikipedia Computational Biology Wikiproject.
The 2012 competition began 9th September 2012 (coinciding with the start of the European Conference on Computational Biology) and finished four months later on the 10th January 2013. Each article entered in the competition was reviewed for a difference in article quality between these two dates. In 2012, there were 13 substantive entries into the competition. Six of these articles were shortlisted by members of the ISCB Student Council and then considered by the judging panel. The judging panel considered articles based on the criteria of clarity of the writing, depth of knowledge of the subject, and quality of figures and images used. In one case, it was clear that the article was largely derived from a published review, and was not considered further. For the other entries, the quantity and quality of the contributions were very good, and it was a challenge to rank the articles. After much deliberation, the judging panel selected the following as the winners of the 2012 ISCB Wikipedia competition:
1st prize: James Estevez for improvements to the Genomics Article.
2nd prize: Benjamin Moore for improvements to the European Nucleotide Archive article.
3rd prize: Luis Pedro Coelho for improvements to the Bioimage Analysis article.
We are keen to grow the depth and quality of computational biology articles and wish to encourage the widest possible range of students and trainees to take part. We envisage that teachers, tutors, and lecturers could use the competition as an opportunity to train students in literature research on topics of computational biology. This approach to literature review provides the students with a thorough grounding in the subject area of the article. In addition, the collaborative writing environment of Wikipedia encourages critical thinking and improves literature research skills. Furthermore, compared to traditional literature reviews carried out by students, which typically end up unread in a filing cabinet, contributing to Wikipedia means that the students' scholarly contributions will be publicly visible.
We hope that the ISCB Wikipedia competition will continue to grow and help improve the quality of Computational Biology information freely available on the Internet. We are interested in improving not just the articles in Wikipedia, but also the associated media, such as images and figures on Wikimedia Commons, and data through Wikidata. We encourage you to get involved by either entering the competition if you are a student or trainee, or getting your own students to participate. Alex Bateman, Janet Kelso, Daniel Mietchen, Geoff MacIntyre, Tomás Di Domenico, Thomas Abeel, Darren W. Logan, Predrag Radivojac, Burkhard Rost |
PLoS Comput. Biol. | 9 |
| 2012 | LocTree2 predicts localization for all domains of lifeabstractMOTIVATION: Subcellular localization is one aspect of protein function. Despite advances in high-throughput imaging, localization maps remain incomplete. Several methods accurately predict localization, but many challenges remain to be tackled. RESULTS: In this study, we introduced a framework to predict localization in life's three domains, including globular and membrane proteins (3 classes for archaea; 6 for bacteria and 18 for eukaryota). The resulting method, LocTree2, works well even for protein fragments. It uses a hierarchical system of support vector machines that imitates the cascading mechanism of cellular sorting. The method reaches high levels of sustained performance (eukaryota: Q18=65%, bacteria: Q6=84%). LocTree2 also accurately distinguishes membrane and non-membrane proteins. In our hands, it compared favorably with top methods when tested on new data. AVAILABILITY: Online through PredictProtein (predictprotein.org); as standalone version at http://www.rostlab.org/services/loctree2. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tatyana Goldberg, Tobias Hamp, Burkhard Rost |
Bioinform. | 3 |
| 2012 | Paving the future: finding suitable ISMB venuesabstractThe International Society for Computational Biology, ISCB, organizes the largest event in the field of computational biology and bioinformatics, namely the annual international conference on Intelligent Systems for Molecular Biology, the ISMB. This year at ISMB 2012 in Long Beach, ISCB celebrated the 20th anniversary of its flagship meeting. ISCB is a young, lean and efficient society that aspires to make a significant impact with only limited resources. Many constraints make the choice of venues for ISMB a tough challenge. Here, we describe those challenges and invite the contribution of ideas for solutions. Burkhard Rost, Terry Gaasterland, Thomas Lengauer, Michal Linial, Scott Markel, B. J. Morrison McKay, Reinhard Schneider 0002, Paul Horton, Janet Kelso |
Bioinform. | 1 |
| 2012 | SNPdbe: constructing an nsSNP functional impacts databaseabstractUNLABELLED: Many existing databases annotate experimentally characterized single nucleotide polymorphisms (SNPs). Each non-synonymous SNP (nsSNP) changes one amino acid in the gene product (single amino acid substitution;SAAS). This change can either affect protein function or be neutral in that respect. Most polymorphisms lack experimental annotation of their functional impact. Here, we introduce SNPdbe-SNP database of effects, with predictions of computationally annotated functional impacts of SNPs. Database entries represent nsSNPs in dbSNP and 1000 Genomes collection, as well as variants from UniProt and PMD. SAASs come from >2600 organisms; 'human' being the most prevalent. The impact of each SAAS on protein function is predicted using the SNAP and SIFT algorithms and augmented with experimentally derived function/structure information and disease associations from PMD, OMIM and UniProt. SNPdbe is consistently updated and easily augmented with new sources of information. The database is available as an MySQL dump and via a web front end that allows searches with any combination of organism names, sequences and mutation IDs. AVAILABILITY: http://www.rostlab.org/services/snpdbe. Christian Schaefer, Alice Meier, Burkhard Rost, Yana Bromberg |
Bioinform. | 3 |
| 2012 | Alternative Protein-Protein Interfaces Are Frequent ExceptionsabstractThe intricate molecular details of protein-protein interactions (PPIs) are crucial for function. Therefore, measuring the same interacting protein pair again, we expect the same result. This work measured the similarity in the molecular details of interaction for the same and for homologous protein pairs between different experiments. All scores analyzed suggested that different experiments often find exceptions in the interfaces of similar PPIs: up to 22% of all comparisons revealed some differences even for sequence-identical pairs of proteins. The corresponding number for pairs of close homologs reached 68%. Conversely, the interfaces differed entirely for 12-29% of all comparisons. All these estimates were calculated after redundancy reduction. The magnitude of interface differences ranged from subtle to the extreme, as illustrated by a few examples. An extreme case was a change of the interacting domains between two observations of the same biological interaction. One reason for different interfaces was the number of copies of an interaction in the same complex: the probability of observing alternative binding modes increases with the number of copies. Even after removing the special cases with alternative hetero-interfaces to the same homomer, a substantial variability remained. Our results strongly support the surprising notion that there are many alternative solutions to make the intricate molecular details of PPIs crucial for function. Tobias Hamp, Burkhard Rost |
PLoS Comput. Biol. | 2 |
| 2011 | Towards big data science in the decade ahead from ten years of InCoB and the 1st ISCB-Asia Joint ConferenceabstractThe 2011 International Conference on Bioinformatics (InCoB) conference, which is the annual scientific conference of the Asia-Pacific Bioinformatics Network (APBioNet), is hosted by Kuala Lumpur, Malaysia, is co-organized with the first ISCB-Asia conference of the International Society for Computational Biology (ISCB). InCoB and the sequencing of the human genome are both celebrating their tenth anniversaries and InCoB's goalposts for the next decade, implementing standards in bioinformatics and globally distributed computational networks, will be discussed and adopted at this conference. Of the 49 manuscripts (selected from 104 submissions) accepted to BMC Genomics and BMC Bioinformatics conference supplements, 24 are featured in this issue, covering software tools, genome/proteome analysis, systems biology (networks, pathways, bioimaging) and drug discovery and design. Shoba Ranganathan, Christian Schönbach, Janet Kelso, Burkhard Rost, Sheila Nathan, Tin Wee Tan |
BMC Bioinform. | 4 |
| 2011 | ISCB Public Policy Statement on Open Access to Scientific and Technical Research LiteratureabstractThe International Society for Computational Biology (ISCB) is dedicated to advancing human knowledge at the intersection of computation and life sciences.On behalf of the ISCB members, this public policy statement expresses strong support for open access, reuse, integration, and distillation of the publicly funded archival scientific and technical research literature, and for the infrastructure to achieve that goal.Knowledge is the fruit of the research endeavor, and the archival scientific and technical research literature is its practical expression and means of communication.Shared knowledge multiplies in utility because every new scientific discovery is built upon previous scientific knowledge.Access to knowledge is access to the power to solve new problems and make informed decisions.Free, open, public, online access to the archival scientific and technical research literature will empower citizens and scientists to solve more problems and make better, more informed decisions.Attribution to the original authors will maintain consistency and accountability within the knowledge base.Computational reuse, integration, and distillation of that literature will produce new and as yet unforeseen knowledge.We strongly encourage open software, data, and databases, issues that are not addressed here.A prior ISCB public policy statement on sharing software provides very clear support for open source/open access (http://www.iscb.org/iscb-policy- Richard H. Lathrop, Burkhard Rost |
PLoS Comput. Biol. | 2 |
| 2010 | Protein secondary structure appears to be robust under in silico evolution while protein disorder appears not to beabstractMOTIVATION: The mutation of amino acids often impacts protein function and structure. Mutations without negative effect sustain evolutionary pressure. We study a particular aspect of structural robustness with respect to mutations: regular protein secondary structure and natively unstructured (intrinsically disordered) regions. Is the formation of regular secondary structure an intrinsic feature of amino acid sequences, or is it a feature that is lost upon mutation and is maintained by evolution against the odds? Similarly, is disorder an intrinsic sequence feature or is it difficult to maintain? To tackle these questions, we in silico mutated native protein sequences into random sequence-like ensembles and monitored the change in predicted secondary structure and disorder. RESULTS: We established that by our coarse-grained measures for change, predictions and observations were similar, suggesting that our results were not biased by prediction mistakes. Changes in secondary structure and disorder predictions were linearly proportional to the change in sequence. Surprisingly, neither the content nor the length distribution for the predicted secondary structure changed substantially. Regions with long disorder behaved differently in that significantly fewer such regions were predicted after a few mutation steps. Our findings suggest that the formation of regular secondary structure is an intrinsic feature of random amino acid sequences, while the formation of long-disordered regions is not an intrinsic feature of proteins with disordered regions. Put differently, helices and strands appear to be maintained easily by evolution, whereas maintaining disordered regions appears difficult. Neutral mutations with respect to disorder are therefore very unlikely. Christian Schaefer, Avner Schlessinger, Burkhard Rost |
Bioinform. | 3 |
| 2009 | Correlating protein function and stability through the analysis of single amino acid substitutionsabstractBACKGROUND: Mutations resulting in the disruption of protein function are the underlying causes of many genetic diseases. Some mutations affect the number of expressed proteins while others alter the activity on a per-molecule basis. Single amino acid substitutions as caused by non-synonymous Single Nucleotide Polymorphisms (nsSNPs) often disrupt function by altering protein structure and/or stability, but can also wreak havoc by directly impacting functional binding sites. Given the experimental three-dimensional (3D) structure of a protein, we can try to differentiate between the "effect on structure/stability" and the "effect on binding". However, experimental 3D structures are available for only 1% of all known proteins; the magnitude of stability change caused by a given mutation is more widely available. RESULTS: Here, we analyze to which extent the functional effect of a mutation can be predicted from the effect on protein stability. We find that simple sequence-based methods succeed in predicting functional effects of nsSNPs. In fact, such methods consistently outperform approaches that predict functional change through the application of binary thresholds to stability change. We also observed that if stability is affected, functional change is easier to predict than when stability is not affected. CONCLUSION: Our results confirmed that stability change is somehow related to function change. However, we also show that the knowledge of stability changes in no way suffices to predict functional changes and that many function changing mutations have no effect on stability. Yana Bromberg, Burkhard Rost |
BMC Bioinform. | 2 |
| 2008 | SNAP predicts effect of mutations on protein functionabstractAbstract Summary: Many non-synonymous single nucleotide polymor-phisms (nsSNPs) in humans are suspected to impact protein function. Here, we present a publicly available server implementation of the method SNAP (screening for non-acceptable polymorphisms) that predicts the functional effects of single amino acid substitutions. SNAP identifies over 80% of the non-neutral mutations at 77% accuracy and over 76% of the neutral mutations at 80% accuracy at its default threshold. Each prediction is associated with a reliability index that correlates with accuracy and thereby enables experimentalists to zoom into the most promising predictions. Availability: Web-server: http://www.rostlab.org/services/SNAP; downloadable program available upon request. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Yana Bromberg, Guy Yachdav, Burkhard Rost |
Bioinform. | 3 |
| 2008 | MetalDetector: a web server for predicting metal-binding sites and disulfide bridges in proteins from sequenceabstractUNLABELLED: The web server MetalDetector classifies histidine residues in proteins into one of two states (free or metal bound) and cysteines into one of three states (free, metal bound or disulfide bridged). A decision tree integrates predictions from two previously developed methods (DISULFIND and Metal Ligand Predictor). Cross-validated performance assessment indicates that our server predicts disulfide bonding state at 88.6% precision and 85.1% recall, while it identifies cysteines and histidines in transition metal-binding sites at 79.9% precision and 76.8% recall, and at 60.8% precision and 40.7% recall, respectively. AVAILABILITY: Freely available at http://metaldetector.dsi.unifi.it. SUPPLEMENTARY INFORMATION: Details and data can be found at http://metaldetector.dsi.unifi.it/help.php. Marco Lippi 0001, Andrea Passerini, Marco Punta, Burkhard Rost, Paolo Frasconi |
Bioinform. | 4 |
| 2008 | Powerful fusion: PSI-BLAST and consensus sequencesabstractMOTIVATION: A typical PSI-BLAST search consists of iterative scanning and alignment of a large sequence database during which a scoring profile is progressively built and refined. Such a profile can also be stored and used to search against a different database of sequences. Using it to search against a database of consensus rather than native sequences is a simple add-on that boosts performance surprisingly well. The improvement comes at a price: we hypothesized that random alignment score statistics would differ between native and consensus sequences. Thus PSI-BLAST-based profile searches against consensus sequences might incorrectly estimate statistical significance of alignment scores. In addition, iterative searches against consensus databases may fail. Here, we addressed these challenges in an attempt to harness the full power of the combination of PSI-BLAST and consensus sequences. RESULTS: We studied alignment score statistics for various types of consensus sequences. In general, the score distribution parameters of profile-based consensus sequence alignments differed significantly from those derived for the native sequences. PSI-BLAST partially compensated for the parameter variation. We have identified a protocol for building specialized consensus sequences that significantly improved search sensitivity and preserved score distribution parameters. As a result, PSI-BLAST profiles can be used to search specialized consensus sequences without sacrificing estimates of statistical significance. We also provided results indicating that iterative PSI-BLAST searches against consensus sequences could work very well. Overall, we showed how a very popular and effective method could be used to identify significantly more relevant similarities among protein sequences. AVAILABILITY: http://www.rostlab.org/services/consensus/. Dariusz Przybylski, Burkhard Rost |
Bioinform. | 2 |
| 2008 | Physical protein-protein interactions predicted from microarraysabstractMOTIVATION: Microarray expression data reveal functionally associated proteins. However, most proteins that are associated are not actually in direct physical contact. Predicting physical interactions directly from microarrays is both a challenging and important task that we addressed by developing a novel machine learning method optimized for this task. RESULTS: We validated our support vector machine-based method on several independent datasets. At the same levels of accuracy, our method recovered more experimentally observed physical interactions than a conventional correlation-based approach. Pairs predicted by our method to very likely interact were close in the overall network of interaction, suggesting our method as an aid for functional annotation. We applied the method to predict interactions in yeast (Saccharomyces cerevisiae). A Gene Ontology function annotation analysis and literature search revealed several probable and novel predictions worthy of future experimental validation. We therefore hope our new method will improve the annotation of interactions as one component of multi-source integrated systems. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ta-tsen Soong, Kazimierz O. Wrzeszczynski, Burkhard Rost |
Bioinform. | 3 |
| 2008 | ISMB 2008 TorontoabstractISCB) presents the Sixteenth International Conference on Intelligent Systems for Molecular Biology (ISMB 2008), to be held in Toronto, Canada, July 19-23, 2008.Now in the final phases of scheduling selected presentations, demonstrations, and posters, the organizers are preparing what will likely be recognized as the premier conference on computational biology in 2008.ISMB 2008 (http://www.iscb.org/ismb2008/)will follow the road paved by the ISMB/ ECCB 2007 (http://www.iscb.org/ismbeccb2007/) in Vienna in the attempt to specifically encourage increased participation from previously under-represented disciplines of computational biology.This conference will feature the best of the computer and life sciences through a variety of core sessions running in multiple parallel tracks, along with single-tracked Keynote Presentations, posters on display throughout the duration of the conference, and an extensive commercial exposition.The first day (July 18) of the meeting is reserved for two-day Special Interest Group (SIG) and Satellite meetings, the second day (July 19) runs SIGs for the first time in parallel with Tutorials and the Student Council Symposium, and for the first time two SIGs are running in parallel with the main ISMB meeting (July 20-23). Michal Linial, Jill P. Mesirov, B. J. Morrison McKay, Burkhard Rost |
PLoS Comput. Biol. | 4 |
| 2007 | ISIS: interaction sites identified from sequenceabstractMOTIVATION: Large-scale experiments reveal pairs of interacting proteins but leave the residues involved in the interactions unknown. These interface residues are essential for understanding the mechanism of interaction and are often desired drug targets. Reliable identification of residues that reside in protein-protein interface typically requires analysis of protein structure. Therefore, for the vast majority of proteins, for which there is no high-resolution structure, there is no effective way of identifying interface residues. RESULTS: Here we present a machine learning-based method that identifies interacting residues from sequence alone. Although the method is developed using transient protein-protein interfaces from complexes of experimentally known 3D structures, it never explicitly uses 3D information. Instead, we combine predicted structural features with evolutionary information. The strongest predictions of the method reached over 90% accuracy in a cross-validation experiment. Our results suggest that despite the significant diversity in the nature of protein-protein interactions, they all share common basic principles and that these principles are identifiable from sequence alone. Yanay Ofran, Burkhard Rost |
Bioinform. | 2 |
| 2007 | Natively unstructured regions in proteins identified from contact predictionsabstractMOTIVATION: Natively unstructured (also dubbed intrinsically disordered) regions in proteins lack a defined 3D structure under physiological conditions and often adopt regular structures under particular conditions. Proteins with such regions are overly abundant in eukaryotes, they may increase functional complexity of organisms and they usually evade structure determination in the unbound form. Low propensity for the formation of internal residue contacts has been previously used to predict natively unstructured regions. RESULTS: We combined PROFcon predictions for protein-specific contacts with a generic pairwise potential to predict unstructured regions. This novel method, Ucon, outperformed the best available methods in predicting proteins with long unstructured regions. Furthermore, Ucon correctly identified cases missed by other methods. By computing the difference between predictions based on specific contacts (approach introduced here) and those based on generic potentials (realized in other methods), we might identify unstructured regions that are involved in protein-protein binding. We discussed one example to illustrate this ambitious aim. Overall, Ucon added quality and an orthogonal aspect that may help in the experimental study of unstructured regions in network hubs. AVAILABILITY: http://www.predictprotein.org/submit_ucon.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Avner Schlessinger, Marco Punta, Burkhard Rost |
Bioinform. | 3 |
| 2007 | ISMB/ECCB 2007: The Premier Conference on Computational BiologyabstractThe International Society for Computational Biology (ISCB) presents ISMB/ECCB 2007, the Fifteenth International Conference on Intelligent Systems for Molecular Biology (ISMB 2007), held jointly with the Sixth European Conference on Computational Biology (ECCB 2007) in Vienna, Austria, July 21–25, 2007 (http://www.iscb.org/ismbeccb2007). Now in the final phases of selecting papers, presentations, demonstrations, and posters, the organizers are preparing what will likely be recognized as the premier conference on computational biology in 2007. ISMB/ECCB 2007 has expanded in ways to specifically encourage increased participation from previously underrepresented disciplines of computational biology. This conference will feature the best of the computer and life sciences through a variety of new and core sessions running in multiple parallel tracks, along with an increase in keynote presentations, posters on display throughout the duration of the conference, and an extensive industry exhibition. Special interest group meetings, a satellite meeting, and tutorials all will precede the main conference dates. Thomas Lengauer, B. J. Morrison McKay, Burkhard Rost |
PLoS Comput. Biol. | 3 |
| 2007 | Protein-Protein Interaction Hotspots Carved into SequencesabstractProtein-protein interactions, a key to almost any biological process, are mediated by molecular mechanisms that are not entirely clear. The study of these mechanisms often focuses on all residues at protein-protein interfaces. However, only a small subset of all interface residues is actually essential for recognition or binding. Commonly referred to as "hotspots," these essential residues are defined as residues that impede protein-protein interactions if mutated. While no in silico tool identifies hotspots in unbound chains, numerous prediction methods were designed to identify all the residues in a protein that are likely to be a part of protein-protein interfaces. These methods typically identify successfully only a small fraction of all interface residues. Here, we analyzed the hypothesis that the two subsets correspond (i.e., that in silico methods may predict few residues because they preferentially predict hotspots). We demonstrate that this is indeed the case and that we can therefore predict directly from the sequence of a single protein which residues are interaction hotspots (without knowledge of the interaction partner). Our results suggested that most protein complexes are stabilized by similar basic principles. The ability to accurately and efficiently identify hotspots from sequence enables the annotation and analysis of protein-protein interaction hotspots in entire organisms and thus may benefit function prediction and drug development. The server for prediction is available at http://www.rostlab.org/services/isis. Yanay Ofran, Burkhard Rost |
PLoS Comput. Biol. | 2 |
| 2007 | Natively Unstructured Loops Differ from Other LoopsabstractNatively unstructured or disordered protein regions may increase the functional complexity of an organism; they are particularly abundant in eukaryotes and often evade structure determination. Many computational methods predict unstructured regions by training on outliers in otherwise well-ordered structures. Here, we introduce an approach that uses a neural network in a very different and novel way. We hypothesize that very long contiguous segments with nonregular secondary structure (NORS regions) differ significantly from regular, well-structured loops, and that a method detecting such features could predict natively unstructured regions. Training our new method, NORSnet, on predicted information rather than on experimental data yielded three major advantages: it removed the overlap between testing and training, it systematically covered entire proteomes, and it explicitly focused on one particular aspect of unstructured regions with a simple structural interpretation, namely that they are loops. Our hypothesis was correct: well-structured and unstructured loops differ so substantially that NORSnet succeeded in their distinction. Benchmarks on previously used and new experimental data of unstructured regions revealed that NORSnet performed very well. Although it was not the best single prediction method, NORSnet was sufficiently accurate to flag unstructured regions in proteins that were previously not annotated. In one application, NORSnet revealed previously undetected unstructured regions in putative targets for structural genomics and may thereby contribute to increasing structural coverage of large eukaryotic families. NORSnet found unstructured regions more often in domain boundaries than expected at random. In another application, we estimated that 50%-70% of all worm proteins observed to have more than seven protein-protein interaction partners have unstructured regions. The comparative analysis between NORSnet and DISOPRED2 suggested that long unstructured loops are a major part of unstructured regions in molecular networks. Avner Schlessinger, Jinfeng Liu 0003, Burkhard Rost |
PLoS Comput. Biol. | 3 |
| 2006 | PROFbval: predict flexible and rigid residues in proteinsabstractUNLABELLED: The mobility of a residue on the protein surface is closely linked to its function. The identification of extremely rigid or flexible surface residues can therefore contribute information crucial for solving the complex problem of identifying functionally important residues in proteins. Mobility is commonly measured by B-value data from high-resolution three-dimensional X-ray structures. Few methods predict B-values from sequence. Here, we present PROFbval, the first web server to predict normalized B-values from amino acid sequence. The server handles amino acid sequences (or alignments) as input and outputs normalized B-value and two-state (flexible/rigid) predictions. The server also assigns a reliability index for each prediction. For example, PROFbval correctly identifies residues in active sites on the surface of enzymes as particularly rigid. AVAILABILITY: http://www.rostlab.org/services/profbval CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Avner Schlessinger, Guy Yachdav, Burkhard Rost |
Bioinform. | 3 |
| 2006 | Protein-Protein Interactions More Conserved within Species than across SpeciesabstractExperimental high-throughput studies of protein-protein interactions are beginning to provide enough data for comprehensive computational studies. Today, about ten large data sets, each with thousands of interacting pairs, coarsely sample the interactions in fly, human, worm, and yeast. Another about 55,000 pairs of interacting proteins have been identified by more careful, detailed biochemical experiments. Most interactions are experimentally observed in prokaryotes and simple eukaryotes; very few interactions are observed in higher eukaryotes such as mammals. It is commonly assumed that pathways in mammals can be inferred through homology to model organisms, e.g. the experimental observation that two yeast proteins interact is transferred to infer that the two corresponding proteins in human also interact. Two pairs for which the interaction is conserved are often described as interologs. The goal of this investigation was a large-scale comprehensive analysis of such inferences, i.e. of the evolutionary conservation of interologs. Here, we introduced a novel score for measuring the overlap between protein-protein interaction data sets. This measure appeared to reflect the overall quality of the data and was the basis for our two surprising results from our large-scale analysis. Firstly, homology-based inferences of physical protein-protein interactions appeared far less successful than expected. In fact, such inferences were accurate only for extremely high levels of sequence similarity. Secondly, and most surprisingly, the identification of interacting partners through sequence similarity was significantly more reliable for protein pairs within the same organism than for pairs between species. Our analysis underlined that the discrepancies between different datasets are large, even when using the same type of experiment on the same organism. This reality considerably constrains the power of homology-based transfer of interactions. In particular, the experimental probing of interactions in distant model organisms has to be undertaken with some caution. More comprehensive images of protein-protein networks will require the combination of many high-throughput methods, including in silico inferences and predictions. http://www.rostlab.org/results/2006/ppi_homology/ Sven Mika, Burkhard Rost |
PLoS Comput. Biol. | 2 |
| 2005 | PROFcon: novel prediction of long-range contactsabstractMOTIVATION: Despite the continuing advance in the experimental determination of protein structures, the gap between the number of known protein sequences and structures continues to increase. Prediction methods can bridge this sequence-structure gap only partially. Better predictions of non-local contacts between residues could improve comparative modeling, fold recognition and could assist in the experimental structure determination. RESULTS: Here, we introduced PROFcon, a novel contact prediction method that combines information from alignments, from predictions of secondary structure and solvent accessibility, from the region between two residues and from the average properties of the entire protein. In contrast to some other methods, PROFcon predicted short and long proteins at similar levels of accuracy. As expected, PROFcon was clearly less accurate when tested on sparse evolutionary profiles, that is, on families with few homologs. Prediction accuracy was highest for proteins belonging to the SCOP alpha/beta class. PROFcon compared favorably with state-of-the-art prediction methods at the CASP6 meeting. While the performance may still be perceived as low, our method clearly pushed the mark higher. Furthermore, predictions are already accurate enough to seed predictions of global features of protein structure. Marco Punta, Burkhard Rost |
Bioinform. | 2 |
| 2002 | ISMB 2002
Janice I. Glasgow, Burkhard Rost |
ISMB | 2 |
| 2002 | Inferring sub-cellular localization through automated lexical analysisabstractAbstract Motivation: The SWISS-PROT sequence database contains keywords of functional annotations for many proteins. In contrast, information about the sub-cellular localization is available for only a few proteins. Experts can often infer localization from keywords describing protein function. We developed LOCkey, a fully automated method for lexical analysis of SWISS-PROT keywords that assigns sub-cellular localization. With the rapid growth in sequence data, the biochemical characterisation of sequences has been falling behind. Our method may be a useful tool for supplementing functional information already automatically available. Results: The method reached a level of more than 82% accuracy in a full cross-validation test. Due to a lack of functional annotations, we could infer localization for fewer than half of all proteins in SWISS-PROT. We applied LOCkey to annotate five entirely sequenced proteomes, namely Saccharomyces cerevisiae (yeast), Caenorhabditis elegans (worm), Drosophila melanogaster (fly), Arabidopsis thaliana (plant) and a subset of all human proteins. LOCkey found about 8000 new annotations of sub-cellular localization for these eukaryotes. Availability: Annotations of localization for eukaryotes at: http://cubic.bioc.columbia.edu/services/LOCkey Contact: [email protected]@columbia.edu Keywords: genome sequence analysis; predicting sub-cellular localization; protein function; lexical analysis. *To whom correspondence should be addressed. Rajesh Nair, Burkhard Rost |
ISMB | 2 |
| 2002 | Target space for structural genomics revisitedabstractMOTIVATION: Structural genomics eventually aims at determining structures for all proteins. However, in the beginning experimentalists are likely to focus on globular proteins to achieve a rapid basic coverage of protein sequence space. How many proteins will structural genomics have to target? How many proteins will be excluded since we already have structural information for these or since they are not globular? We have to answer these questions in the context of our target selection for the North-East Structural Genomics Consortium (NESG). RESULTS: We estimated that structural information is available for about 6-38% of all proteins; 6% if we require high accuracy in comparative modelling, 38% if we are satisfied with having a rough idea about the fold. Excluding all regions that are not globular, we found that structural genomics may have to target about 48% of all proteins. This corresponded to a similar percentage of residues of the entire proteomes (52%). We explored a number of different strategies to cluster protein space in order to find the number of families representing these 48% of structurally unknown proteins. For the subset of all entirely sequenced eukaryotes, we found over 18 000 fragment clusters each of which may be a suitable target for structural genomics. AVAILABILITY: All data are available from the authors, most results are summarized at: http://cubic.bioc.columbia.edu/genomes/RES/2002_bioinformatics/ Jinfeng Liu 0003, Burkhard Rost |
Bioinform. | 2 |
| 2002 | Bioinformatics in structural genomics - EditorialabstractBurkhard Rost, Barry Honig, Alfonso Valencia; Bioinformatics in structural genomics, Bioinformatics, Volume 18, Issue 7, 1 July 2002, Pages 897, https://doi.org Burkhard Rost, Barry Honig, Alfonso Valencia |
Bioinform. | 1 |
| 2001 | EVA: continuous automatic evaluation of protein structure prediction serversabstractUNLABELLED: Evaluation of protein structure prediction methods is difficult and time-consuming. Here, we describe EVA, a web server for assessing protein structure prediction methods, in an automated, continuous and large-scale fashion. Currently, EVA evaluates the performance of a variety of prediction methods available through the internet. Every week, the sequences of the latest experimentally determined protein structures are sent to prediction servers, results are collected, performance is evaluated, and a summary is published on the web. EVA has so far collected data for more than 3000 protein chains. These results may provide valuable insight to both developers and users of prediction methods. AVAILABILITY: http://cubic.bioc.columbia.edu/eva. CONTACT: [email protected] Volker A. Eyrich, Marc A. Martí-Renom, Dariusz Przybylski, M. S. Madhusudhan 0001, András Fiser, Florencio Pazos, Alfonso Valencia, Andrej Sali, Burkhard Rost |
Bioinform. | 9 |
| 1999 | A platform for integrating threading results with protein family analysesabstractAbstract Summary: We have developed a package for the interactive visualization of results from different threading programs. Additionally, we have integrated relevant information about protein sequence, function, evolution, and structure into the interface. Availability: A detailed documentation of THREADLIZE, and the binaries for IRIX, SunOS and Linux are available at http://www.cnb.uam.es/~pazos/threadlize. The package is free for academic users. Contact: [email protected] Supplementary information: http://www.cnb.uam.es/~pazos/threadlize Florencio Pazos, Burkhard Rost, Alfonso Valencia |
Bioinform. | 2 |
| 1997 | Sisyphus and prediction of protein structureabstractThe problem of predicting protein structure from the sequence remains fundamentally unsolved despite more than three decades of intensive research effort. However, new and promising methods in three-dimensional (3D), 2D and 1D prediction have reopened the field. Mean-force-potentials derived from the protein databases can distinguish between correct and incorrect models (3D). Inter-residue contacts (2D) can be detected by analysis of correlated mutations, albeit with low accuracy. Secondary structure, solvent accessibility and transmembrane helices (1D) can be predicted with significantly improved accuracy using multiple sequence alignments. Some of these new prediction methods have proven accurate and reliable enough to be useful in genome analysis, and in experimental structure determination. Moreover, the new generation of theoretical methods is increasingly influencing experiments in molecular biology. Burkhard Rost, Seán I. O'Donoghue |
Comput. Appl. Biosci. | 1 |
| 1996 | Refining Neural Network Predictions for Helical Transmembrane Proteins by Dynamic Programming
Burkhard Rost, Rita Casadio, Piero Fariselli |
ISMB | 1 |
| 1995 | TOPITS: Threading One-Dimensional Predictions Into Three-Dimensional Structures
Burkhard Rost |
ISMB | 1 |
| 1994 | PHD - an automatic mail server for protein secondary structure predictionabstractBy the middle of 1993, > 30,000 protein sequences has been listed. For 1000 of these, the three-dimensional (tertiary) structure has been experimentally solved. Another 7000 can be modelled by homology. For the remaining 21,000 sequences, secondary structure prediction provides a rough estimate of structural features. Predictions in three states range between 35% (random) and 88% (homology modelling) overall accuracy. Using information about evolutionary conservation as contained in multiple sequence alignments, the secondary structure of 4700 protein sequences was predicted by the automatic e-mail server PHD. For proteins with at least one known homologue, the method has an expected overall three-state accuracy of 71.4% for proteins with at least one known homologue (evaluated on 126 unique protein chains). Burkhard Rost, Chris Sander, Reinhard Schneider 0002 |
Comput. Appl. Biosci. | 1 |
| 1992 | Exercising Multi-Layered Networks on Protein Secondary StructureabstractThe quality of a multi-layered network predicting the secondary structure of proteins is improved substantially by: (i) using information about evolutionarily conserved amino acids (increase of overall accuracy by six percentage points), (ii) balancing the training dynamics (increase of accuracy for strand), and (iii) combining uncorrelated networks in a jury (increase two percentage points). In addition, appending a second level structure-to-structure network results in better reproduction of the length of secondary structure segments. Burkhard Rost, Chris Sander |
Int. J. Neural Syst. | 1 |