Sanne Abeln

dblp:62/893 · DBLP profile ↗
← Back
25ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-2779-7174ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 24 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Bridging the gap between performance and interpretability: An explainable disentangled multimodal framework for cancer survival prediction
abstract
While multimodal survival prediction models are increasingly accurate, their complexity often reduces interpretability, limiting insight into how different data sources influence predictions. To address this, we introduce DIMAFx, an explainable multimodal framework for cancer survival prediction that produces disentangled, interpretable modality-specific and modality-shared representations from histopathology whole-slide images and transcriptomics data. Across four TCGA cancer cohorts, DIMAFx achieves survival prediction performance competitive with the state of the art and consistently stronger representation disentanglement. Leveraging its interpretable design, SHapley Additive exPlanations, and pathologist-in-the-loop annotations, DIMAFx facilitates systematic investigation of key multimodal interactions and the biological information encoded in the multimodal, disentangled representations. In breast cancer survival prediction, the most predictive features contain modality-shared information, including one capturing solid tumor morphology contextualized primarily by late estrogen response, where higher-grade morphology aligned with pathway downregulation was associated with increased risk, consistent with known breast cancer biology. Key modality-specific features capture microenvironmental signals from interacting adipose and stromal morphologies. These results show that DIMAFx substantially narrows the gap between performance and interpretability, supporting the application of such models in precision oncology.
Aniek Eijpe, Soufyan Lakbir, Melis Erdal Cesur, Sara Pires de Oliveira, Angelos Chatzimparmpas, Sanne Abeln, Wilson Silva
Artif. Intell. Medicine6
2026 PLM-eXplain: divide and conquer the protein embedding space
abstract
MOTIVATION: Protein language models (PLMs) have revolutionized computational biology through their ability to generate powerful sequence representations for diverse prediction tasks. However, their black-box nature limits biological interpretation and translation to actionable insights. Bridging this gap requires approaches that maintain predictive performance while providing interpretable explanations of model behaviour. RESULTS: We present PLM-eXplain (PLM-X), an explainable adapter layer that bridges this gap by factoring PLM embeddings into two complementary components: an interpretable subspace based on established biochemical features, and a residual subspace that retains predictive, non-interpretable information. Using embeddings from ESM2 and ProtBert, PLM-X incorporates well-established properties, including secondary structure and hydropathy, while maintaining high predictive performance. We demonstrate the effectiveness of our approach across three biologically relevant classification tasks: extracellular vesicle association, transmembrane helix prediction, and aggregation propensity prediction. PLM-X enables biological interpretation of model decisions without sacrificing accuracy, offering a generalizable solution for enhancing PLM interpretability across various downstream applications. AVAILABILITY AND IMPLEMENTATION: Source code and models are available at https://github.com/AIT4LIFE-UU/PLM-eXplain/.
Jan van Eck, Dea Gogishvili, Wilson Silva, Sanne Abeln
Bioinform.4
2025 Disentangled and Interpretable Multimodal Attention Fusion for Cancer Survival Prediction
Aniek Eijpe, Soufyan Lakbir, Melis Erdal Cesur, Sara Pires de Oliveira, Sanne Abeln, Wilson Silva
MICCAI (14)5
2025 Comparison of sequence- and structure-based antibody clustering approaches on simulated repertoire sequencing data
abstract
Repertoire sequencing allows us to investigate the antibody-mediated immune response. The clustering of sequences is a crucial step in the data analysis pipeline, aiding in the identification of functionally related antibodies. The conventional clustering approach of clonotyping relies on sequence information, particularly CDRH3 sequence identity and V/J gene usage, to group sequences into clonotypes. It has been suggested that the limitations of sequence-based approaches to identify sequence-dissimilar but functionally converged antibodies can be overcome by using structure information to group antibodies. Recent advances have made structure-based methods feasible on a repertoire level. However, so far, their performance has only been evaluated on single-antigen sets of antibodies. A comprehensive comparison of the benefits and limitations of structure-based tools on realistic and diverse repertoire data is missing. Here, we aim to explore the promise of structure-based clustering algorithms to replace or augment the standard sequence-based approach, specifically by identifying low-sequence identity groups. Two methods, SAAB+ and SPACE2, are evaluated against clonotyping. We curated a dataset of well-annotated pairs of antibodies that show high overlap in epitope residues and thus bind the same region within their respective antigen. This set of antibodies was introduced into a simulated repertoire to compare the performance of clustering approaches on a diverse antibody set. Our analysis reveals that structure-based methods do group more antibodies together compared to clonotyping. However, it also highlights the limitations associated with the need for same-length CDR regions by SPACE2. This work thoroughly compares the utility of different clustering methods and provides insights into what further steps are required to effectively use antibody structural information to group immune repertoire data.
Katharina Waury, Stefan Lelieveld, Sanne Abeln, Henk-Jan van den Ham
PLoS Comput. Biol.3
2024 CIBRA identifies genomic alterations with a system-wide impact on tumor biology
abstract
MOTIVATION: Genomic instability is a hallmark of cancer, leading to many somatic alterations. Identifying which alterations have a system-wide impact is a challenging task. Nevertheless, this is an essential first step for prioritizing potential biomarkers. We developed CIBRA (Computational Identification of Biologically Relevant Alterations), a method that determines the system-wide impact of genomic alterations on tumor biology by integrating two distinct omics data types: one indicating genomic alterations (e.g. genomics), and another defining a system-wide expression response (e.g. transcriptomics). CIBRA was evaluated with genome-wide screens in 33 cancer types using primary and metastatic cancer data from the Cancer Genome Atlas and Hartwig Medical Foundation. RESULTS: We demonstrate the capability of CIBRA by successfully confirming the impact of point mutations in experimentally validated oncogenes and tumor suppressor genes (0.79 AUC). Surprisingly, many genes affected by structural variants were identified to have a strong system-wide impact (30.3%), suggesting that their role in cancer development has thus far been largely under-reported. Additionally, CIBRA can identify impact with only 10 cases and controls, providing a novel way to prioritize genomic alterations with a prominent role in cancer biology. Our findings demonstrate that CIBRA can identify cancer drivers by combining genomics and transcriptomics data. Moreover, our work shows an unexpected substantial system-wide impact of structural variants in cancer. Hence, CIBRA has the potential to preselect and refine current definitions of genomic alterations to derive more nuanced biomarkers for diagnostics, disease progression, and treatment response. AVAILABILITY AND IMPLEMENTATION: The R package CIBRA is available at https://github.com/AIT4LIFE-UU/CIBRA.
Soufyan Lakbir, Caterina Buranelli, Gerrit Meijer, Jaap Heringa, Remond J. A. Fijneman, Sanne Abeln
Bioinform.6
2022 PIPENN: protein interface prediction from sequence with an ensemble of neural nets
abstract
MOTIVATION: The interactions between proteins and other molecules are essential to many biological and cellular processes. Experimental identification of interface residues is a time-consuming, costly and challenging task, while protein sequence data are ubiquitous. Consequently, many computational and machine learning approaches have been developed over the years to predict such interface residues from sequence. However, the effectiveness of different Deep Learning (DL) architectures and learning strategies for protein-protein, protein-nucleotide and protein-small molecule interface prediction has not yet been investigated in great detail. Therefore, we here explore the prediction of protein interface residues using six DL architectures and various learning strategies with sequence-derived input features. RESULTS: We constructed a large dataset dubbed BioDL, comprising protein-protein interactions from the PDB, and DNA/RNA and small molecule interactions from the BioLip database. We also constructed six DL architectures, and evaluated them on the BioDL benchmarks. This shows that no single architecture performs best on all instances. An ensemble architecture, which combines all six architectures, does consistently achieve peak prediction accuracy. We confirmed these results on the published benchmark set by Zhang and Kurgan (ZK448), and on our own existing curated homo- and heteromeric protein interaction dataset. Our PIPENN sequence-based ensemble predictor outperforms current state-of-the-art sequence-based protein interface predictors on ZK448 on all interaction types, achieving an AUC-ROC of 0.718 for protein-protein, 0.823 for protein-nucleotide and 0.842 for protein-small molecule. AVAILABILITY AND IMPLEMENTATION: Source code and datasets are available at https://github.com/ibivu/pipenn/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bas Stringer, Hans de Ferrante, Sanne Abeln, Jaap Heringa, K. Anton Feenstra, Reza Haydarlou
Bioinform.3
2022 INFLECT: an R-package for cytometry cluster evaluation using marker modality
abstract
BACKGROUND: Current methods of high-dimensional unsupervised clustering of mass cytometry data lack means to monitor and evaluate clustering results. Whether unsupervised clustering is correct is typically evaluated by agreement with dimensionality reduction techniques or based on benchmarking with manually classified cells. The ambiguity and lack of reproducibility of sequential gating has been replaced with ambiguity in interpretation of clustering results. On the other hand, spurious overclustering of data leads to loss of statistical power. We have developed INFLECT, an R-package designed to give insight in clustering results and provide an optimal number of clusters. In our approach, a mass cytometry dataset is overclustered intentionally to ensure the smallest phenotypically different subsets are captured using FlowSOM. A range of metacluster number endpoints are generated and evaluated using marker interquartile range and distribution unimodality checks. The fraction of marker distributions that pass these checks is taken as a measure of clustering success. The fraction of unimodal distributions within metaclusters is plotted against the number of generated metaclusters and reaches a plateau of diminishing returns. The inflection point at which this occurs gives an optimal point of capturing cellular heterogeneity versus statistical power. RESULTS: We applied INFLECT to four publically available mass cytometry datasets of different size and number of markers. The unimodality score consistently reached a plateau, with an inflection point dependent on dataset size and number of dimensions. We tested both ConsenusClusterPlus metaclustering and hierarchical clustering. While hierarchical clustering is less computationally expensive and thus faster, it achieved similar results to ConsensusClusterPlus. The four datasets consisted of labeled data and we compared INFLECT metaclustering to published results. INFLECT identified a higher optimal number of metaclusters for all datasets. We illustrated the underlying heterogeneity within labels, showing that these labels encompass distinct types of cells. CONCLUSION: INFLECT addresses a knowledge gap in high-dimensional cytometry analysis, namely assessing clustering results. This is done through monitoring marker distributions for interquartile range and unimodality across a range of metacluster numbers. The inflection point is the optimal trade-off between cellular heterogeneity and statistical power, applied in this work for FlowSOM clustering on mass cytometry datasets.
Jan Verhoeff, Sanne Abeln, Juan J. Garcia-Vallejo
BMC Bioinform.2
2021 SeRenDIP-CE: sequence-based interface prediction for conformational epitopes
abstract
MOTIVATION: Antibodies play an important role in clinical research and biotechnology, with their specificity determined by the interaction with the antigen's epitope region, as a special type of protein-protein interaction (PPI) interface. The ubiquitous availability of sequence data, allows us to predict epitopes from sequence in order to focus time-consuming wet-lab experiments toward the most promising epitope regions. Here, we extend our previously developed sequence-based predictors for homodimer and heterodimer PPI interfaces to predict epitope residues that have the potential to bind an antibody. RESULTS: We collected and curated a high quality epitope dataset from the SAbDab database. Our generic PPI heterodimer predictor obtained an AUC-ROC of 0.666 when evaluated on the epitope test set. We then trained a random forest model specifically on the epitope dataset, reaching AUC 0.694. Further training on the combined heterodimer and epitope datasets, improves our final predictor to AUC 0.703 on the epitope test set. This is better than the best state-of-the-art sequence-based epitope predictor BepiPred-2.0. On one solved antibody-antigen structure of the COVID19 virus spike receptor binding domain, our predictor reaches AUC 0.778. We added the SeRenDIP-CE Conformational Epitope predictors to our webserver, which is simple to use and only requires a single antigen sequence as input, which will help make the method immediately applicable in a wide range of biomedical and biomolecular research. AVAILABILITY AND IMPLEMENTATION: Webserver, source code and datasets at www.ibi.vu.nl/programs/serendipwww/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Qingzhen Hou, Bas Stringer, Katharina Waury, Henriette Capel, Reza Haydarlou, Fuzhong Xue, Sanne Abeln, Jaap Heringa, K. Anton Feenstra
Bioinform.7
2020 The hydrophobic effect characterises the thermodynamic signature of amyloid fibril growth
abstract
Many proteins have the potential to aggregate into amyloid fibrils, protein polymers associated with a wide range of human disorders such as Alzheimer's and Parkinson's disease. The thermodynamic stability of amyloid fibrils, in contrast to that of folded proteins, is not well understood: the balance between entropic and enthalpic terms, including the chain entropy and the hydrophobic effect, are poorly characterised. Using a combination of theory, in vitro experiments, simulations of a coarse-grained protein model and meta-data analysis, we delineate the enthalpic and entropic contributions that dominate amyloid fibril elongation. Our prediction of a characteristic temperature-dependent enthalpic signature is confirmed by the performed calorimetric experiments and a meta-analysis over published data. From these results we are able to define the necessary conditions to observe cold denaturation of amyloid fibrils. Overall, we show that amyloid fibril elongation is associated with a negative heat capacity, the magnitude of which correlates closely with the hydrophobic surface area that is buried upon fibril formation, highlighting the importance of hydrophobicity for fibril stability.
Juami Hermine Mariama van Gils, Erik van Dijk, Alessia Peduzzo, Alexander Hofmann, Nicola Vettore, Marie P. Schützmann, Georg Groth, Halima Mouhib, Daniel E. Otzen, Alexander K. Buell, Sanne Abeln
PLoS Comput. Biol.11
2019 Tailor-made multiple sequence alignments using the PRALINE 2 alignment toolkit
abstract
SUMMARY: PRALINE 2 is a toolkit for custom multiple sequence alignment workflows. It can be used to incorporate sequence annotations, such as secondary structure or (DNA) motifs, into the alignment scoring, as well as to customize many other aspects of a progressive multiple alignment workflow. AVAILABILITY AND IMPLEMENTATION: PRALINE 2 is implemented in Python and available as open source software on GitHub: https://github.com/ibivu/PRALINE/.
Maurits J. J. Dijkstra, Atze van der Ploeg, K. Anton Feenstra, Wan J. Fokkink, Sanne Abeln, Jaap Heringa
Bioinform.5
2019 SeRenDIP: SEquential REmasteriNg to DerIve profiles for fast and accurate predictions of PPI interface positions
abstract
MOTIVATION: Interpretation of ubiquitous protein sequence data has become a bottleneck in biomolecular research, due to a lack of structural and other experimental annotation data for these proteins. Prediction of protein interaction sites from sequence may be a viable substitute. We therefore recently developed a sequence-based random forest method for protein-protein interface prediction, which yielded a significantly increased performance than other methods on both homomeric and heteromeric protein-protein interactions. Here, we present a webserver that implements this method efficiently. RESULTS: With the aim of accelerating our previous approach, we obtained sequence conservation profiles by re-mastering the alignment of homologous sequences found by PSI-BLAST. This yielded a more than 10-fold speedup and at least the same accuracy, as reported previously for our method; these results allowed us to offer the method as a webserver. The web-server interface is targeted to the non-expert user. The input is simply a sequence of the protein of interest, and the output a table with scores indicating the likelihood of having an interaction interface at a certain position. As the method is sequence-based and not sensitive to the type of protein interaction, we expect this webserver to be of interest to many biological researchers in academia and in industry. AVAILABILITY AND IMPLEMENTATION: Webserver, source code and datasets are available at www.ibi.vu.nl/programs/serendipwww/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Qingzhen Hou, Paul F. G. De Geest, Christian J. Griffioen, Sanne Abeln, Jaap Heringa, K. Anton Feenstra
Bioinform.4
2018 Training for translation between disciplines: a philosophy for life and data sciences curricula
abstract
Motivation: Our society has become data-rich to the extent that research in many areas has become impossible without computational approaches. Educational programmes seem to be lagging behind this development. At the same time, there is a growing need not only for strong data science skills, but foremost for the ability to both translate between tools and methods on the one hand, and application and problems on the other. Results: Here we present our experiences with shaping and running a masters' programme in bioinformatics and systems biology in Amsterdam. From this, we have developed a comprehensive philosophy on how translation in training may be achieved in a dynamic and multidisciplinary research area, which is described here. We furthermore describe two requirements that enable translation, which we have found to be crucial: sufficient depth and focus on multidisciplinary topic areas, coupled with a balanced breadth from adjacent disciplines. Finally, we present concrete suggestions on how this may be implemented in practice, which may be relevant for the effectiveness of life science and data science curricula in general, and of particular interest to those who are in the process of setting up such curricula. Supplementary information: Supplementary data are available at Bioinformatics online.
K. Anton Feenstra, Sanne Abeln, Johan A. Westerhuis, Filipe Brancos dos Santos, Douwe Molenaar, Bas Teusink, Huub C. J. Hoefsloot, Jaap Heringa
Bioinform.2
2018 Motif-Aware PRALINE: Improving the alignment of motif regions
abstract
Protein or DNA motifs are sequence regions which possess biological importance. These regions are often highly conserved among homologous sequences. The generation of multiple sequence alignments (MSAs) with a correct alignment of the conserved sequence motifs is still difficult to achieve, due to the fact that the contribution of these typically short fragments is overshadowed by the rest of the sequence. Here we extended the PRALINE multiple sequence alignment program with a novel motif-aware MSA algorithm in order to address this shortcoming. This method can incorporate explicit information about the presence of externally provided sequence motifs, which is then used in the dynamic programming step by boosting the amino acid substitution matrix towards the motif. The strength of the boost is controlled by a parameter, α. Using a benchmark set of alignments we confirm that a good compromise can be found that improves the matching of motif regions while not significantly reducing the overall alignment quality. By estimating α on an unrelated set of reference alignments we find there is indeed a strong conservation signal for motifs. A number of typical but difficult MSA use cases are explored to exemplify the problems in correctly aligning functional sequence motifs and how the motif-aware alignment method can be employed to alleviate these problems.
Maurits J. J. Dijkstra, Punto Bawono, Sanne Abeln, K. Anton Feenstra, Wan J. Fokkink, Jaap Heringa
PLoS Comput. Biol.3
2016 BioASF: a framework for automatically generating executable pathway models specified in BioPAX
abstract
MOTIVATION: Biological pathways play a key role in most cellular functions. To better understand these functions, diverse computational and cell biology researchers use biological pathway data for various analysis and modeling purposes. For specifying these biological pathways, a community of researchers has defined BioPAX and provided various tools for creating, validating and visualizing BioPAX models. However, a generic software framework for simulating BioPAX models is missing. Here, we attempt to fill this gap by introducing a generic simulation framework for BioPAX. The framework explicitly separates the execution model from the model structure as provided by BioPAX, with the advantage that the modelling process becomes more reproducible and intrinsically more modular; this ensures natural biological constraints are satisfied upon execution. The framework is based on the principles of discrete event systems and multi-agent systems, and is capable of automatically generating a hierarchical multi-agent system for a given BioPAX model. RESULTS: To demonstrate the applicability of the framework, we simulated two types of biological network models: a gene regulatory network modeling the haematopoietic stem cell regulators and a signal transduction network modeling the Wnt/β-catenin signaling pathway. We observed that the results of the simulations performed using our framework were entirely consistent with the simulation results reported by the researchers who developed the original models in a proprietary language. AVAILABILITY AND IMPLEMENTATION: The framework, implemented in Java, is open source and its source code, documentation and tutorial are available at http://www.ibi.vu.nl/programs/BioASF CONTACT: [email protected].
Reza Haydarlou, Annika Jacobsen, Nicola Bonzanni, K. Anton Feenstra, Sanne Abeln, Jaap Heringa
Bioinform.5
2016 ECCB 2016: The 15th European Conference on Computational Biology
abstract
This special issue includes the proceeding papers accepted for presentation at the 15th European Conference on Computational Biology (ECCB 2016), to be held from September 3 to 7, 2016 at the World Forum Convention Center in The Hague, The Netherlands. Details of the conference are available on the conference web site (www.eccb2016.org) and will later be archived at eccb.iscb.org/2016/. ECCB is the premier European conference in computational biology and bioinformatics, and together with ISMB (Intelligent Systems in Molecular Biology) and RECOMB (Research in Computational Molecular Biology), it is one of the major international conference series in this domain. The ECCB conferences are gathering about a thousand scientists and industry staff working at the intersection of a broad range of disciplines including computer science, mathematics, biology and medicine. New challenges are now emerging in these fields with the recent advances in low-cost ultra-fast sequencing, bio-imaging and big data. Computational analysis platforms are challenged by the enormous complexity of biological systems and the sheer amount of data resulting from high-throughput measuring techniques, such as single-cell or single-molecule measurements. As a consequence, databases and software are evolving rapidly, and new algorithms are required to improve computational analyses of massive biological or biomedical datasets. Recent advances are presented at the conference in the field of data interoperability and machine learning, in particular ‘deep learning’ and network-based analysis techniques. The impact on the field of public data repositories such as TCGA (Weinstein et al., 2013), ENCODE (ENCODE Project Consortium, 2012) and GDSC (Yang et al., 2013) is palpable with many submissions revolving around these resources. ECCB is held annually in a different country, while it is held jointly with the ISMB conference biennially. Going back in time, the fourteen previous editions of ECCB have been held in: Dublin, Ireland, together with ISMB (Moreau and Beerenwinkel, 2015); Strasbourg, France (Devignes, 2014); Berlin, Germany, together with ISMB (Ben-Tal, 2013); Basel, Switzerland (Schwede and Iber, 2012); Vienna, Austria, together with ISMB (Gaasterland and Vingron, 2011); Ghent, Belgium (Moreau and Heringa, 2010); Stockholm, Sweden, together with ISMB (Gusfield and Tramontano, 2009); Cagliari, Italy (Tramontano, 2008); Vienna, Austria, together with ISMB (Lengauer et al., 2007); Eilat, Israel (Wolfson and Safer, 2007); Madrid, Spain (Guigo et al., 2005); Glasgow, United Kingdom, together with ISMB (Thornton et al., 2004); Paris, France (Lenhof and Sagot, 2003); and Saarbrücken, Germany (Lengauer, 2002). The ECCB 2016 edition features keynote lectures by distinguished speakers. The opening keynote lecture will be delivered by the 2013 Breakthrough Prize-winner Hans Clevers (Hubrecht Institute and Princess Maxima Centre for Pediatric Oncology, Utrecht, The Netherlands). Further keynote presentations will be given by Amos Tanay (Weizmann Institute, Tel Aviv, Israel), John Marioni (EMBL-EBI, Hinxton, UK), Nuria Lopez-Bigas (Universitat Pompeu Fabra, Barcelona, Spain), Benedict Paten (UCSC, Santa Cruz, USA), Pauline Hogeweg (Utrecht University, Utrecht, The Netherlands) and Christina Leslie (Memorial Sloan Kettering Cancer Center, New York, USA). The conference topics span all areas of methodological developments for computational biology and innovative applications of computational methods to molecular biology and biomedicine. To present a more unified view of where the science has gone over recent years, new to ECCB this year is that the conference presentations are divided over five broad themes: (i) Data (organization, management, categorization, integration, analysis of data, knowledge discovery); (ii) Genome (sequence analysis, alignment, evolution, phylogeny, genetics, epigenetics, 3D conformation); (iii) Genes (expression, function, regulation, transcription, translation, geno/phenotype); (iv) Proteins (structure, function, alterations, assemblies, interactions, design, proteomics) and (v) Systems (systems biology, pathways, molecular networks, dynamics, signalling, multi-scale modelling). This year four different tracks were created for ECCB: two scientific and two application-oriented ones. On the scientific side, the Proceedings Track presents novel scientific contributions, while the Highlights Track showcases already published high-impact science in computational biology. These two tracks were coordinated and managed across the five themes by a board of 5 × 4 = 20 co-chairs, both overseen by a dedicated track chair. A novelty of ECCB this year is that we have given a prominent platform to applications in the new Application and ELIXIR tracks. The Application Track is an initiative to promote application of computational biology in industry and other fields beyond academia. Submissions to this track should cross the boundaries of traditional academic science, or show developments that are directly relevant beyond academia or have potential for it. Consequently, submissions relating to (pure) research were deemed out of scope. The ELIXIR Track, running for the first time at ECCB 2016, is managed by ELIXIR and focuses on developments relating to services and infrastructure within the ELIXIR nodes. ELIXIR is the pan-European life science infrastructural network (‘Data for the Life Sciences’). It coordinates, integrates and sustains bioinformatics resources across its member states, enabling users in academia and industry to access vital data, tools, standards, compute and training services for research. ELIXIR has chosen ECCB as their major dissemination platform, acting as co-organising sponsor. Following the call for Proceedings papers, we received 150 submissions. Submission authors were asked to rank the five themes for fit of their paper. These rankings were subsequently used to assign papers to themes and to arrive at an optimally balanced distribution of submissions over the themes. Within each theme, the theme co-chairs assigned papers to expert referees, taking care to avoid any conflict of interest. Together, the Programme Committee (PC) was composed of 217 reviewers and 35 co-reviewers. The reviewing and selection process was carried out using the EasyChair multi-track conference reviewing system (www.easychair.org). The review form explicitly differentiated between impact and suitability for ECCB on the one hand, and scientific quality and reproducibility on the other. This distinction was intended to create more clarity for both the reviewers and authors; the reviewing criteria were therefore explicitly stated in the submission guidelines. The added focus on reproducibility resulted in many authors opting to make their methods open source. After the reviewers reached a consensus, a final ranking and selection for each theme was carried out by the theme (co-)chairs. A total of 48 papers were conditionally accepted (acceptance ratio of 32%) based on a predefined number of acceptances per theme based on the distribution of initial assignments over the themes. The authors had two weeks to modify their papers according to the suggestions made by the reviewers and to respond to the reviewers’ comments, which was checked by the theme (co-)chairs. We thank the authors for incorporating these suggestions as they were given little time to carry out (minor) revisions. Essentially, these efforts contribute to the success and reputation of the ECCB conference! It is worth noting that many authors expressed their gratitude to the reviewers for their comments and suggestions. We also gratefully acknowledge the hard and diligent work performed by the reviewers over a short period of time. We believe that for all of the rejected submissions, the reviewers provided high-quality reports. We hope that authors of rejected manuscripts will benefit from these remarks in their future research. We thank all theme co-chairs for their availability throughout the reviewing process, for their very positive attitude in the final selection and for their valuable help in re-examining the modified submissions. The 48 accepted papers are included in this special issue. The Proceedings Track papers with their supplementary files are available free-for-view in electronic format from Oxford’s press journal Bioinformatics from September 1st, 2016. Highlight presentations were introduced at ISMB/ECCB 2007 in Vienna (Lengauer et al., 2007), and immediately became one of the most popular features of the conference. All original research papers that had been published in peer-review journals between 1 March, 2015, and the submission deadline of 29 March 2016, were eligible to be presented as a Highlight talk. After thorough consideration and discussion, the ECCB 2016 theme (co-)chairs selected 24 proposals out of 71 submissions, mainly on the criteria of compatibility with the ECCB objectives, wide impact in the life sciences and the potential for attracting a large audience to the conference. The Applications Track features 11 presentations, which were selected out of 31 submissions. At the time of writing, three Sponsored Talks will be delivered as part of the Applications Track by respectively The Hyve, Keygene and Data Computing. Sponsored Talks enable sponsors to showcase their innovations in computational biology and to highlight their scientific value. Finally, the ELIXIR track proved to be a very popular addition, with 50 high-quality submissions for only 12 presentation slots. Submissions to the Poster Track were evaluated based on a 250-word abstract and will be shown at the conference along the five main conference themes. In line with the new Application and ELIXIR tracks, there will also be special Application and ELIXIR poster tracks. A dedicated Education Poster Track was created to devote in-depth attention to education that is so crucial for the next generation of bioinformaticians. All poster abstracts are available on the conference web site. We also arranged with F1000Research (http://f1000research.com/) to publish the posters via a new ECCB2016 channel; submission is on a voluntary basis. At least sixteen exhibitor booths will be open throughout the conference in the central conference hall, which is well connected to the other activities at ECCB: The Hyve, EMBL-EBI, TimeLogic, Springer, ISCB Student Council, Goblet, ELIXIR, ISCB, Oxford University Press, ENPICOM, Data Computing, CRC Press, SIB, ELIXIR Denmark and Cambridge University Press. They will be presenting the latest scientific literature in the field of computational biology, bioinformatics, data stewardship, modeling and simulation, as well as new hardware, software and technology developments. The Hague Tourist Office, and the four organizing institutions (DTL, Netherlands Bioinformatics and System Biology Research School (BioSB), VU University Amsterdam and Delft University of Technology) will also be represented. During the weekend before the conference, a satellite meeting, 14 workshops and nine tutorials will take place. The Student Council of the International Society for Computational Biology (ISCB) organizes its 4th European Student Council Symposium (ESCS), chaired by Annika Jacobsen (Vrije Universiteit Amsterdam) and Kevin Schwahn (Universität Potsdam, Max Planck Institute of Molecular Plant Physiology). ESCS highlights will be published in F1000Research via the ISCB Student Council channel. The ECCB 2016 organizing committee congratulates all these dynamic young scientists for their enthusiasm, which is essential to the future of research in computational biology. The 14 workshops preceding the ECCB 2016 main meeting were selected out of a total of 29 applications, showing how popular ECCB has become as a venue for dissemination of computational biology research. The workshops, running all but one for a single day, provide participants with an informal setting to discuss technical issues, exchange research ideas, and to share practical experiences on a range of focused or emerging topics in computational biology. Taken together, the workshops demonstrate how extensively technologies have found their way into large-scale practical applications: • (W1) The 10th International Workshop on Machine Learning in Systems Biology, organised by (Juho Rousu, Aalto University, Finland), Dick de Ridder Wageningen University, The Netherlands), Harri Lähdesmäki (Aalto University, Finland) and Aalt—Jan van Dijk (Wageningen University, The Netherlands), is a two-day workshop with its own proceedings track, where accepted papers will be published in BMC Bioinformatics, • (W2) Network Inference: New Methods and New Data, organised by Anagha Joshi (Roslin Institute, University of Edinburgh, UK), Tom Michoel (Roslin Institute, University of Edinburgh, UK) and Eric Bonnet (Centre National de Génotypage, CEA, Paris, France) • (W3) Getting the Most out of Your Methods and Algorithms: A Workshop on How to Use Existing Datasets to Gain Novel Insight, chaired by Morris Swertz and Lude Franke (both at University Medical Centre Groningen, The Netherlands), • (W4) Digital Pathology Meets Bioinformatics, organised by Yves Sucaet (Vrije Universiteit Brussel, Belgium), Jeroen Van der Laak (UMC Radboud, Nijmegen, The Netherlands), Marius Nap (HistoGeneX, Belgium and Rigshospitalet Copenhagen, Denmark), Zev Leifer (New York College of Podiatric Medicine, USA), Yukako Yagi (Harvard Medical School, Cambridge and Massachusetts General Hospital, Boston, USA) and Raphaël Marée (Université de Liège, Belgium), • (W5) RepSeq 2016: Immune Repertoire Sequencing—Bioinformatics and Applications in Hematology and Immunology, organised by Jack Bartram (University College London, UK), Eva Froňková (Charles University Prague, Czech Republic), Mathieu Giraud, (CNRS, Lille, France—program co-chair), Peter N. Robinson (Charité Berlin, Germany), Mikaël Salson (Université de Lille, France—proceedings chair), Mikhail Shugay (Shemyakin and Ovchinnikov Institute of Bioorganic Chemistry, Moscow, Russia—program co-chair) and Andrew P. Stubbs (Erasmus MC, Rotterdam, The Netherlands), • (W6) FAIR Data and Data Stewardship, organised by Erik Schultes, Mark Thompson Marco Roos (all three at Leiden University Medical Centre, The Netherlands), Mark Wilkinson (Universidad Politecnica de Madrid, Spain), Luiz Olavo Bonino da Silva Santos (Dutch Techcentre for Life Sciences and Vrije Universiteit Amsterdam, The Netherlands), • (W7) Challenges and Approaches in Comprehensive and Informative Complex Network Analysis for Precision Medicine, organised by Igor Jurisica (University of Toronto, Canada), Natasa Przulj (University College London, UK) and Tijana Milenkovic (University of Notre Dame, USA), • (W8) Computing a Tissue: Modeling Multicellular Systems, organised by Walter de Back (TU Dresden, Germany), Sara Montagna (University of Bologna, Italy) and Roeland Merks (Center for Mathematics and Computer Science (CWI), Amsterdam and Leiden University, The Netherlands), • (W9) Computational Pan-Genomics, organised by Zamin Iqbal (University of Oxford, UK), Tobias Marschall (Saarland University and Max Planck Institute for Informatics, Saarbrücken, Germany) and Benedict Paten (3UC Santa Cruz Genomics Institute, Santa Cruz, CA, USA), • (W10) BioNetVisA: from Biological Network Reconstruction to data visualisation and analysis in Molecular Biology and Medicine, organised by Inna Kuperstein, Emmanuel Barillot, Andrei Zinovyev (all three at Institut Curie, Paris, France), Hiroaki Kitano (Okinawa Institute of Science and Technology Graduate University, RIKEN Center for Integrative Medical Sciences, Japan), Minoru Kanehisa (Kyoto University, Japan), Samik Ghosh (Systems Biology Institute, Tokyo, Japan), Nicolas Le Novère (Babraham Institute, Cambridge, UK), Robin Haw (Ontario Institute for Cancer Research, Canada), Alfonso Valencia (Spanish National Bioinformatics Institute, Madrid, Spain) and Lodewyk Wessels (Netherlands Cancer Institute, Amsterdam, and Technical University Delft, The Netherlands), • (W11) Recent Computational Advances in Metagenomics, organised by Sophie Schbath, Valentin Loux and Mahendra Mariadassou (all at INRA, Jouy-en-Josas, France), • (W12) BioExcel: Advanced Simulations for Biomolecular Research, a SIG workshop organised by Rossen Apostolov (KTH, Stockholm, Sweden)), Alexandre Bonvin (Utrecht University, The Netherlands), Cath Brooksbank (EMBL-EBI, Hinxton, UK) and Ian Harrow (Ian Harrow Consulting, Whitstable, UK), • (W13) Clinical Bioinformatics as a Service, organised by Niko Beerenwinkel (ETH Zurich, and SIB Swiss Institute of Bioinformatics, Switzerland), Wolfgang Huber (EMBL, Heidelberg), Simon Tavaré (Cancer Research UK Cambridge Institute, UK) and Daniel Stekhoven and SIB Swiss Institute of Bioinformatics, Switzerland), • Computational Challenges of Data organised by University Medical Center, The Netherlands), (UMC Utrecht, The Netherlands) and Hans The Netherlands). • Data Analysis with and organised by Jeroen and The Netherlands), • Genome organised by (University of Spain) and (University of Spain), • and organised by (University of Spain), • Analysis and with organised by United • and for Modeling Biological Systems, organised by (University of Germany) and (University of Germany), • Analysis and for Data, organised by Medical Center, Amsterdam and University of Amsterdam, The Netherlands), • A to Research on organised by and (both at University Pompeu Fabra, Barcelona, Spain), • and their in Genome Data, organised by N. (University of Cambridge, UK), University of Science and and (University of UK), • into and its Applications in Bioinformatics and in Network organised by and (both at Hans Institute, and University Hospital, It is to thank all the and that are ECCB 2016 a and high-quality conference. of we are to the committee composed of the theme (co-)chairs and reviewers for their crucial and dedicated We are to the ECCB committee for their and to the of the conference. In the and by Committee and ECCB ECCB Yves ECCB ECCB and ECCB 2012) were Yves and for and The and of ISCB Society for Computational Biology) in the about ECCB 2016 at the international and for were We thank all provided to the conference. In to the co-organising ELIXIR, we gratefully acknowledge sponsors The Netherlands and USA) and the National Centre for Research (CNRS, We also as where also the opening will be the of The Hague, the and The Bioinformatics Centre the Poster We are to the Oxford University for the ECCB 2016 special issue. The F1000Research is for and the ECCB 2016 posters on their web The ECCB 2016 science has been in the of the efforts have been van (Netherlands Bioinformatics and Systems Biology Research Techcentre for Life ECCB 2016 for the selection (TU Delft, The Netherlands), ECCB 2016 for out the workshop selection (EMBL, Germany), ECCB 2016 Applications Track for and the new Applications Andrew and (all at ELIXIR Hinxton, UK) for taking care in of the first ELIXIR Applications and (University of Finland), ECCB 2016 Poster Track for and the ECCB 2016 to the of the conference, beyond the call of and we a in particular and van both at the Techcentre for Life Sciences for their and in this are also to and van of by Netherlands), ECCB 2016 (Dutch Techcentre for Life had a in the while van is more for also in Finally, all these efforts be the many participants from all over the of will to the conference in of scientific contributions, applications, or poster presentations and all for there and for to science at ECCB 2016 in The
Jaap Heringa, Marcel J. T. Reinders, Sanne Abeln, Jeroen de Ridder
Bioinform.3
2016 metaModules identifies key functional subnetworks in microbiome-related disease
abstract
MOTIVATION: The human microbiome plays a key role in health and disease. Thanks to comparative metatranscriptomics, the cellular functions that are deregulated by the microbiome in disease can now be computationally explored. Unlike gene-centric approaches, pathway-based methods provide a systemic view of such functions; however, they typically consider each pathway in isolation and in its entirety. They can therefore overlook the key differences that (i) span multiple pathways, (ii) contain bidirectionally deregulated components, (iii) are confined to a pathway region. To capture these properties, computational methods that reach beyond the scope of predefined pathways are needed. RESULTS: By integrating an existing module discovery algorithm into comparative metatranscriptomic analysis, we developed metaModules, a novel computational framework for automated identification of the key functional differences between health- and disease-associated communities. Using this framework, we recovered significantly deregulated subnetworks that were indeed recognized to be involved in two well-studied, microbiome-mediated oral diseases, such as butanoate production in periodontal disease and metabolism of sugar alcohols in dental caries. More importantly, our results indicate that our method can be used for hypothesis generation based on automated discovery of novel, disease-related functional subnetworks, which would otherwise require extensive and laborious manual assessment. AVAILABILITY AND IMPLEMENTATION: metaModules is available at https://bitbucket.org/alimay/metamodules/ CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ali May, Bernd W. Brandt, Mohammed El-Kebir, Gunnar W. Klau, Egija Zaura, Wim Crielaard, Jaap Heringa, Sanne Abeln
Bioinform.8
2015 The Hydrophobic Temperature Dependence of Amino Acids Directly Calculated from Protein Structures
abstract
The hydrophobic effect is the main driving force in protein folding. One can estimate the relative strength of this hydrophobic effect for each amino acid by mining a large set of experimentally determined protein structures. However, the hydrophobic force is known to be strongly temperature dependent. This temperature dependence is thought to explain the denaturation of proteins at low temperatures. Here we investigate if it is possible to extract this temperature dependence directly from a large set of protein structures determined at different temperatures. Using NMR structures filtered for sequence identity, we were able to extract hydrophobicity propensities for all amino acids at five different temperature ranges (spanning 265-340 K). These propensities show that the hydrophobicity becomes weaker at lower temperatures, in line with current theory. Alternatively, one can conclude that the temperature dependence of the hydrophobic effect has a measurable influence on protein structures. Moreover, this work provides a method for probing the individual temperature dependence of the different amino acid types, which is difficult to obtain by direct experiment.
Erik van Dijk, Arlo Hoogeveen, Sanne Abeln
PLoS Comput. Biol.3
2015 Mapping the Protein Fold Universe Using the CamTube Force Field in Molecular Dynamics Simulations
abstract
It has been recently shown that the coarse-graining of the structures of polypeptide chains as self-avoiding tubes can provide an effective representation of the conformational space of proteins. In order to fully exploit the opportunities offered by such a 'tube model' approach, we present here a strategy to combine it with molecular dynamics simulations. This strategy is based on the incorporation of the 'CamTube' force field into the Gromacs molecular dynamics package. By considering the case of a 60-residue polyvaline chain, we show that CamTube molecular dynamics simulations can comprehensively explore the conformational space of proteins. We obtain this result by a 20 μs metadynamics simulation of the polyvaline chain that recapitulates the currently known protein fold universe. We further show that, if residue-specific interaction potentials are added to the CamTube force field, it is possible to fold a protein into a topology close to that of its native state. These results illustrate how the CamTube force field can be used to explore efficiently the universe of protein folds with good accuracy and very limited computational cost.
Predrag Kukic, Arvind Kannan, Maurits J. J. Dijkstra, Sanne Abeln, Carlo Camilloni, Michele Vendruscolo
PLoS Comput. Biol.4
2014 Unraveling the outcome of 16S rDNA-based taxonomy analysis through mock data and simulations
abstract
MOTIVATION: 16S rDNA pyrosequencing is a powerful approach that requires extensive usage of computational methods for delineating microbial compositions. Previously, it was shown that outcomes of studies relying on this approach vastly depend on the choice of pre-processing and clustering algorithms used. However, obtaining insights into the effects and accuracy of these algorithms is challenging due to difficulties in generating samples of known composition with high enough diversity. Here, we use in silico microbial datasets to better understand how the experimental data are transformed into taxonomic clusters by computational methods. RESULTS: We were able to qualitatively replicate the raw experimental pyrosequencing data after rigorously adjusting existing simulation software. This allowed us to simulate datasets of real-life complexity, which we used to assess the influence and performance of two widely used pre-processing methods along with 11 clustering algorithms. We show that the choice, order and mode of the pre-processing methods have a larger impact on the accuracy of the clustering pipeline than the clustering methods themselves. Without pre-processing, the difference between the performances of clustering methods is large. Depending on the clustering algorithm, the most optimal analysis pipeline resulted in significant underestimations of the expected number of clusters (minimum: 3.4%; maximum: 13.6%), allowing us to make quantitative estimations of the bacterial complexity of real microbiome samples.
Ali May, Sanne Abeln, Wim Crielaard, Jaap Heringa, Bernd W. Brandt
Bioinform.2
2014 Coarse-grained versus atomistic simulations: realistic interaction free energies for real proteins
abstract
MOTIVATION: To assess whether two proteins will interact under physiological conditions, information on the interaction free energy is needed. Statistical learning techniques and docking methods for predicting protein-protein interactions cannot quantitatively estimate binding free energies. Full atomistic molecular simulation methods do have this potential, but are completely unfeasible for large-scale applications in terms of computational cost required. Here we investigate whether applying coarse-grained (CG) molecular dynamics simulations is a viable alternative for complexes of known structure. RESULTS: We calculate the free energy barrier with respect to the bound state based on molecular dynamics simulations using both a full atomistic and a CG force field for the TCR-pMHC complex and the MP1-p14 scaffolding complex. We find that the free energy barriers from the CG simulations are of similar accuracy as those from the full atomistic ones, while achieving a speedup of >500-fold. We also observe that extensive sampling is extremely important to obtain accurate free energy barriers, which is only within reach for the CG models. Finally, we show that the CG model preserves biological relevance of the interactions: (i) we observe a strong correlation between evolutionary likelihood of mutations and the impact on the free energy barrier with respect to the bound state; and (ii) we confirm the dominant role of the interface core in these interactions. Therefore, our results suggest that CG molecular simulations can realistically be used for the accurate prediction of protein-protein interaction strength. AVAILABILITY AND IMPLEMENTATION: The python analysis framework and data files are available for download at http://www.ibi.vu.nl/downloads/bioinformatics-2013-btt675.tgz.
Ali May, René Pool, Erik van Dijk, Jochem Bijlard, Sanne Abeln, Jaap Heringa, K. Anton Feenstra
Bioinform.5
2013 Bioinformatics and Systems Biology: bridging the gap between heterogeneous student backgrounds
abstract
Teaching students with very diverse backgrounds can be extremely challenging. This article uses the Bioinformatics and Systems Biology MSc in Amsterdam as a case study to describe how the knowledge gap for students with heterogeneous backgrounds can be bridged. We show that a mix in backgrounds can be turned into an advantage by creating a stimulating learning environment for the students. In the MSc Programme, conversion classes help to bridge differences between students, by mending initial knowledge and skill gaps. Mixing students from different backgrounds in a group to solve a complex task creates an opportunity for the students to reflect on their own abilities. We explain how a truly interdisciplinary approach to teaching helps students of all backgrounds to achieve the MSc end terms. Moreover, transferable skills obtained by the students in such a mixed study environment are invaluable for their later careers.
Sanne Abeln, Douwe Molenaar, K. Anton Feenstra, Huub C. J. Hoefsloot, Bas Teusink, Jaap Heringa
Briefings Bioinform.1
2013 Exploring Fold Space Preferences of New-born and Ancient Protein Superfamilies
abstract
The evolution of proteins is one of the fundamental processes that has delivered the diversity and complexity of life we see around ourselves today. While we tend to define protein evolution in terms of sequence level mutations, insertions and deletions, it is hard to translate these processes to a more complete picture incorporating a polypeptide's structure and function. By considering how protein structures change over time we can gain an entirely new appreciation of their long-term evolutionary dynamics. In this work we seek to identify how populations of proteins at different stages of evolution explore their possible structure space. We use an annotation of superfamily age to this space and explore the relationship between these ages and a diverse set of properties pertaining to a superfamily's sequence, structure and function. We note several marked differences between the populations of newly evolved and ancient structures, such as in their length distributions, secondary structure content and tertiary packing arrangements. In particular, many of these differences suggest a less elaborate structure for newly evolved superfamilies when compared with their ancient counterparts. We show that the structural preferences we report are not a residual effect of a more fundamental relationship with function. Furthermore, we demonstrate the robustness of our results, using significant variation in the algorithm used to estimate the ages. We present these age estimates as a useful tool to analyse protein populations. In particularly, we apply this in a comparison of domains containing greek key or jelly roll motifs.
Hannah Edwards, Sanne Abeln, Charlotte M. Deane
PLoS Comput. Biol.2
2012 Comparing clustering and pre-processing in taxonomy analysis
abstract
MOTIVATION: Massively parallel sequencing allows for rapid sequencing of large numbers of sequences in just a single run. Thus, 16S ribosomal RNA (rRNA) amplicon sequencing of complex microbial communities has become possible. The sequenced 16S rRNA fragments (reads) are clustered into operational taxonomic units and taxonomic categories are assigned. Recent reports suggest that data pre-processing should be performed before clustering. We assessed combinations of data pre-processing steps and clustering algorithms on cluster accuracy for oral microbial sequence data. RESULTS: The number of clusters varied up to two orders of magnitude depending on pre-processing. Pre-processing using both denoising and chimera checking resulted in a number of clusters that was closest to the number of species in the mock dataset (25 versus 15). Based on run time, purity and normalized mutual information, we could not identify a single best clustering algorithm. The differences in clustering accuracy among the algorithms after the same pre-processing were minor compared with the differences in accuracy among different pre-processing steps. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected] or [email protected]
Marc J. Bonder, Sanne Abeln, Egija Zaura, Bernd W. Brandt
Bioinform.2
2008 Disordered Flanks Prevent Peptide Aggregation
abstract
Natively unstructured or disordered regions appear to be abundant in eukaryotic proteins. Many such regions have been found alongside small linear binding motifs. We report a Monte Carlo study that aims to elucidate the role of disordered regions adjacent to such binding motifs. The coarse-grained simulations show that small hydrophobic peptides without disordered flanks tend to aggregate under conditions where peptides embedded in unstructured peptide sequences are stable as monomers or as part of small micelle-like clusters. Surprisingly, the binding free energy of the motif is barely decreased by the presence of disordered flanking regions, although it is sensitive to the loss of entropy of the motif itself upon binding. This latter effect allows for reversible binding of the signalling motif to the substrate. The work provides insights into a mechanism that prevents the aggregation of signalling peptides, distinct from the general mechanism of protein folding, and provides a testable hypothesis to explain the abundance of disordered regions in proteins.
Sanne Abeln, Daan Frenkel
PLoS Comput. Biol.1
2007 Using Phylogeny to Improve Genome-Wide Distant Homology Recognition
abstract
The gap between the number of known protein sequences and structures continues to widen, particularly as a result of sequencing projects for entire genomes. Recently there have been many attempts to generate structural assignments to all genes on sets of completed genomes using fold-recognition methods. We developed a method that detects false positives made by these genome-wide structural assignment experiments by identifying isolated occurrences. The method was tested using two sets of assignments, generated by SUPERFAMILY and PSI-BLAST, on 150 completed genomes. A phylogeny of these genomes was built and a parsimony algorithm was used to identify isolated occurrences by detecting occurrences that cause a gain at leaf level. Isolated occurrences tend to have high e-values, and in both sets of assignments, a sudden increase in isolated occurrences is observed for e-values >10(-8) for SUPERFAMILY and >10(-4) for PSI-BLAST. Conditions to predict false positives are based on these results. Independent tests confirm that the predicted false positives are indeed more likely to be incorrectly assigned. Evaluation of the predicted false positives also showed that the accuracy of profile-based fold-recognition methods might depend on secondary structure content and sequence length. We show that false positives generated by fold-recognition methods can be identified by considering structural occurrence patterns on completed genomes; occurrences that are isolated within the phylogeny tend to be less reliable. The method provides a new independent way to examine the quality of fold assignments and may be used to improve the output of any genome-wide fold assignment method.
Sanne Abeln, Carlo Teubner, Charlotte M. Deane
PLoS Comput. Biol.1