VLDB 2026 Research / reviewers in the wild / expert
Jaap Heringa
dblp:85/5966
· DBLP profile ↗
40ranked-venue papers
3as first author
4since 2021 · last 2024
0000-0001-8641-4930ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 38 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | CIBRA identifies genomic alterations with a system-wide impact on tumor biologyabstractMOTIVATION: Genomic instability is a hallmark of cancer, leading to many somatic alterations. Identifying which alterations have a system-wide impact is a challenging task. Nevertheless, this is an essential first step for prioritizing potential biomarkers. We developed CIBRA (Computational Identification of Biologically Relevant Alterations), a method that determines the system-wide impact of genomic alterations on tumor biology by integrating two distinct omics data types: one indicating genomic alterations (e.g. genomics), and another defining a system-wide expression response (e.g. transcriptomics). CIBRA was evaluated with genome-wide screens in 33 cancer types using primary and metastatic cancer data from the Cancer Genome Atlas and Hartwig Medical Foundation. RESULTS: We demonstrate the capability of CIBRA by successfully confirming the impact of point mutations in experimentally validated oncogenes and tumor suppressor genes (0.79 AUC). Surprisingly, many genes affected by structural variants were identified to have a strong system-wide impact (30.3%), suggesting that their role in cancer development has thus far been largely under-reported. Additionally, CIBRA can identify impact with only 10 cases and controls, providing a novel way to prioritize genomic alterations with a prominent role in cancer biology. Our findings demonstrate that CIBRA can identify cancer drivers by combining genomics and transcriptomics data. Moreover, our work shows an unexpected substantial system-wide impact of structural variants in cancer. Hence, CIBRA has the potential to preselect and refine current definitions of genomic alterations to derive more nuanced biomarkers for diagnostics, disease progression, and treatment response. AVAILABILITY AND IMPLEMENTATION: The R package CIBRA is available at https://github.com/AIT4LIFE-UU/CIBRA. Soufyan Lakbir, Caterina Buranelli, Gerrit Meijer, Jaap Heringa, Remond J. A. Fijneman, Sanne Abeln |
Bioinform. | 4 |
| 2024 | Mining literature and pathway data to explore the relations of ketamine with neurotransmitters and gut microbiota using a knowledge-graphabstractMOTIVATION: Up-to-date pathway knowledge is usually presented in scientific publications for human reading, making it difficult to utilize these resources for semantic integration and computational analysis of biological pathways. We here present an approach to mining knowledge graphs by combining manual curation with automated named entity recognition and automated relation extraction. This approach allows us to study pathway-related questions in detail, which we here show using the ketamine pathway, aiming to help improve understanding of the role of gut microbiota in the antidepressant effects of ketamine. RESULTS: The thus devised ketamine pathway 'KetPath' knowledge graph comprises five parts: (i) manually curated pathway facts from images; (ii) recognized named entities in biomedical texts; (iii) identified relations between named entities; (iv) our previously constructed microbiota and pre-/probiotics knowledge bases; and (v) multiple community-accepted public databases. We first assessed the performance of automated extraction of relations between named entities using the specially designed state-of-the-art tool BioKetBERT. The query results show that we can retrieve drug actions, pathway relations, co-occurring entities, and their relations. These results uncover several biological findings, such as various gut microbes leading to increased expression of BDNF, which may contribute to the sustained antidepressant effects of ketamine. We envision that the methods and findings from this research will aid researchers who wish to integrate and query data and knowledge from multiple biomedical databases and literature simultaneously. AVAILABILITY AND IMPLEMENTATION: Data and query protocols are available in the KetPath repository at https://dx.doi.org/10.5281/zenodo.8398941 and https://github.com/tingcosmos/KetPath. K. Anton Feenstra, Zhisheng Huang, Jaap Heringa |
Bioinform. | 4 |
| 2022 | PIPENN: protein interface prediction from sequence with an ensemble of neural netsabstractMOTIVATION: The interactions between proteins and other molecules are essential to many biological and cellular processes. Experimental identification of interface residues is a time-consuming, costly and challenging task, while protein sequence data are ubiquitous. Consequently, many computational and machine learning approaches have been developed over the years to predict such interface residues from sequence. However, the effectiveness of different Deep Learning (DL) architectures and learning strategies for protein-protein, protein-nucleotide and protein-small molecule interface prediction has not yet been investigated in great detail. Therefore, we here explore the prediction of protein interface residues using six DL architectures and various learning strategies with sequence-derived input features. RESULTS: We constructed a large dataset dubbed BioDL, comprising protein-protein interactions from the PDB, and DNA/RNA and small molecule interactions from the BioLip database. We also constructed six DL architectures, and evaluated them on the BioDL benchmarks. This shows that no single architecture performs best on all instances. An ensemble architecture, which combines all six architectures, does consistently achieve peak prediction accuracy. We confirmed these results on the published benchmark set by Zhang and Kurgan (ZK448), and on our own existing curated homo- and heteromeric protein interaction dataset. Our PIPENN sequence-based ensemble predictor outperforms current state-of-the-art sequence-based protein interface predictors on ZK448 on all interaction types, achieving an AUC-ROC of 0.718 for protein-protein, 0.823 for protein-nucleotide and 0.842 for protein-small molecule. AVAILABILITY AND IMPLEMENTATION: Source code and datasets are available at https://github.com/ibivu/pipenn/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bas Stringer, Hans de Ferrante, Sanne Abeln, Jaap Heringa, K. Anton Feenstra, Reza Haydarlou |
Bioinform. | 4 |
| 2021 | SeRenDIP-CE: sequence-based interface prediction for conformational epitopesabstractMOTIVATION: Antibodies play an important role in clinical research and biotechnology, with their specificity determined by the interaction with the antigen's epitope region, as a special type of protein-protein interaction (PPI) interface. The ubiquitous availability of sequence data, allows us to predict epitopes from sequence in order to focus time-consuming wet-lab experiments toward the most promising epitope regions. Here, we extend our previously developed sequence-based predictors for homodimer and heterodimer PPI interfaces to predict epitope residues that have the potential to bind an antibody. RESULTS: We collected and curated a high quality epitope dataset from the SAbDab database. Our generic PPI heterodimer predictor obtained an AUC-ROC of 0.666 when evaluated on the epitope test set. We then trained a random forest model specifically on the epitope dataset, reaching AUC 0.694. Further training on the combined heterodimer and epitope datasets, improves our final predictor to AUC 0.703 on the epitope test set. This is better than the best state-of-the-art sequence-based epitope predictor BepiPred-2.0. On one solved antibody-antigen structure of the COVID19 virus spike receptor binding domain, our predictor reaches AUC 0.778. We added the SeRenDIP-CE Conformational Epitope predictors to our webserver, which is simple to use and only requires a single antigen sequence as input, which will help make the method immediately applicable in a wide range of biomedical and biomolecular research. AVAILABILITY AND IMPLEMENTATION: Webserver, source code and datasets at www.ibi.vu.nl/programs/serendipwww/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qingzhen Hou, Bas Stringer, Katharina Waury, Henriette Capel, Reza Haydarlou, Fuzhong Xue, Sanne Abeln, Jaap Heringa, K. Anton Feenstra |
Bioinform. | 8 |
| 2020 | A framework for exhaustive modelling of genetic interaction patterns using Petri netsabstractMOTIVATION: Genetic interaction (GI) patterns are characterized by the phenotypes of interacting single and double mutated gene pairs. Uncovering the regulatory mechanisms of GIs would provide a better understanding of their role in biological processes, diseases and drug response. Computational analyses can provide insights into the underpinning mechanisms of GIs. RESULTS: In this study, we present a framework for exhaustive modelling of GI patterns using Petri nets (PN). Four-node models were defined and generated on three levels with restrictions, to enable an exhaustive approach. Simulations suggest ∼5 million models of GIs. Generalizing these we propose putative mechanisms for the GI patterns, inversion and suppression. We demonstrate that exhaustive PN modelling enables reasoning about mechanisms of GIs when only the phenotypes of gene pairs are known. The framework can be applied to other GI or genetic regulatory datasets. AVAILABILITY AND IMPLEMENTATION: The framework is available at http://www.ibi.vu.nl/programs/ExhMod. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Annika Jacobsen, Olga Ivanova, Saman Amini, Jaap Heringa, Patrick Kemmeren, K. Anton Feenstra |
Bioinform. | 4 |
| 2019 | Bioinformatics in the Netherlands: the value of a nationwide communityabstractThis review provides a historical overview of the inception and development of bioinformatics research in the Netherlands. Rooted in theoretical biology by foundational figures such as Paulien Hogeweg (at Utrecht University since the 1970s), the developments leading to organizational structures supporting a relatively large Dutch bioinformatics community will be reviewed. We will show that the most valuable resource that we have built over these years is the close-knit national expert community that is well engaged in basic and translational life science research programmes. The Dutch bioinformatics community is accustomed to facing the ever-changing landscape of data challenges and working towards solutions together. In addition, this community is the stable factor on the road towards sustainability, especially in times where existing funding models are challenged and change rapidly. Celia W. G. van Gelder, Rob W. W. Hooft, Merlijn N. van Rijswijk, Linda van den Berg, Ruben G. Kok, Marcel J. T. Reinders, Barend Mons, Jaap Heringa |
Briefings Bioinform. | 8 |
| 2019 | Tailor-made multiple sequence alignments using the PRALINE 2 alignment toolkitabstractSUMMARY: PRALINE 2 is a toolkit for custom multiple sequence alignment workflows. It can be used to incorporate sequence annotations, such as secondary structure or (DNA) motifs, into the alignment scoring, as well as to customize many other aspects of a progressive multiple alignment workflow. AVAILABILITY AND IMPLEMENTATION: PRALINE 2 is implemented in Python and available as open source software on GitHub: https://github.com/ibivu/PRALINE/. Maurits J. J. Dijkstra, Atze van der Ploeg, K. Anton Feenstra, Wan J. Fokkink, Sanne Abeln, Jaap Heringa |
Bioinform. | 6 |
| 2019 | SeRenDIP: SEquential REmasteriNg to DerIve profiles for fast and accurate predictions of PPI interface positionsabstractMOTIVATION: Interpretation of ubiquitous protein sequence data has become a bottleneck in biomolecular research, due to a lack of structural and other experimental annotation data for these proteins. Prediction of protein interaction sites from sequence may be a viable substitute. We therefore recently developed a sequence-based random forest method for protein-protein interface prediction, which yielded a significantly increased performance than other methods on both homomeric and heteromeric protein-protein interactions. Here, we present a webserver that implements this method efficiently. RESULTS: With the aim of accelerating our previous approach, we obtained sequence conservation profiles by re-mastering the alignment of homologous sequences found by PSI-BLAST. This yielded a more than 10-fold speedup and at least the same accuracy, as reported previously for our method; these results allowed us to offer the method as a webserver. The web-server interface is targeted to the non-expert user. The input is simply a sequence of the protein of interest, and the output a table with scores indicating the likelihood of having an interaction interface at a certain position. As the method is sequence-based and not sensitive to the type of protein interaction, we expect this webserver to be of interest to many biological researchers in academia and in industry. AVAILABILITY AND IMPLEMENTATION: Webserver, source code and datasets are available at www.ibi.vu.nl/programs/serendipwww/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qingzhen Hou, Paul F. G. De Geest, Christian J. Griffioen, Sanne Abeln, Jaap Heringa, K. Anton Feenstra |
Bioinform. | 5 |
| 2019 | The ability of transcription factors to differentially regulate gene expression is a crucial component of the mechanism underlying inversion, a frequently observed genetic interaction patternabstractGenetic interactions, a phenomenon whereby combinations of mutations lead to unexpected effects, reflect how cellular processes are wired and play an important role in complex genetic diseases. Understanding the molecular basis of genetic interactions is crucial for deciphering pathway organization as well as understanding the relationship between genetic variation and disease. Several hypothetical molecular mechanisms have been linked to different genetic interaction types. However, differences in genetic interaction patterns and their underlying mechanisms have not yet been compared systematically between different functional gene classes. Here, differences in the occurrence and types of genetic interactions are compared for two classes, gene-specific transcription factors (GSTFs) and signaling genes (kinases and phosphatases). Genome-wide gene expression data for 63 single and double deletion mutants in baker's yeast reveals that the two most common genetic interaction patterns are buffering and inversion. Buffering is typically associated with redundancy and is well understood. In inversion, genes show opposite behavior in the double mutant compared to the corresponding single mutants. The underlying mechanism is poorly understood. Although both classes show buffering and inversion patterns, the prevalence of inversion is much stronger in GSTFs. To decipher potential mechanisms, a Petri Net modeling approach was employed, where genes are represented as nodes and relationships between genes as edges. This allowed over 9 million possible three and four node models to be exhaustively enumerated. The models show that a quantitative difference in interaction strength is a strict requirement for obtaining inversion. In addition, this difference is frequently accompanied with a second gene that shows buffering. Taken together, these results provide a mechanistic explanation for inversion. Furthermore, the ability of transcription factors to differentially regulate expression of their targets provides a likely explanation why inversion is more prevalent for GSTFs compared to kinases and phosphatases. Saman Amini, Annika Jacobsen, Olga Ivanova, Philip Lijnzaad, Jaap Heringa, Frank C. P. Holstege, K. Anton Feenstra, Patrick Kemmeren |
PLoS Comput. Biol. | 5 |
| 2018 | Message from the eScience 2018 Program Committee Chairs for the Focused Session on Data Handling and Analytics for HealthabstractPresents the introductory welcome message from the conference proceedings. May include the conference officers' congratulations to all involved with the conference event and publication of the proceedings record. Jaap Heringa, Vincent T. van Hees |
eScience | 1 |
| 2018 | Bioinformatics in the Netherlands: the value of a nationwide communityabstractBriefings in Bioinformatics, 2017. https://doi.org/10.1093/bib/bbx087 The authors have corrected the acknowledgements section to include Antoine van Kampen. The section has been corrected online and now reads as follows: Huge thanks are due to Gert Vriend, Jacob de Vlieg and Bob Hertzberger for securing funding for NBIC, initiating its activities, and for leading NBIC during its initial years. We are indebted in the same vein to Antoine van Kampen, who chaired and directed NBIC from 2006 to 2010. The authors would also like to acknowledge the Dutch bioinformatics community at large for its commitment and energy to drive the field further. Celia W. G. van Gelder, Rob W. W. Hooft, Merlijn N. van Rijswijk, Linda van den Berg, Ruben G. Kok, Marcel J. T. Reinders, Barend Mons, Jaap Heringa |
Briefings Bioinform. | 8 |
| 2018 | Training for translation between disciplines: a philosophy for life and data sciences curriculaabstractMotivation: Our society has become data-rich to the extent that research in many areas has become impossible without computational approaches. Educational programmes seem to be lagging behind this development. At the same time, there is a growing need not only for strong data science skills, but foremost for the ability to both translate between tools and methods on the one hand, and application and problems on the other. Results: Here we present our experiences with shaping and running a masters' programme in bioinformatics and systems biology in Amsterdam. From this, we have developed a comprehensive philosophy on how translation in training may be achieved in a dynamic and multidisciplinary research area, which is described here. We furthermore describe two requirements that enable translation, which we have found to be crucial: sufficient depth and focus on multidisciplinary topic areas, coupled with a balanced breadth from adjacent disciplines. Finally, we present concrete suggestions on how this may be implemented in practice, which may be relevant for the effectiveness of life science and data science curricula in general, and of particular interest to those who are in the process of setting up such curricula. Supplementary information: Supplementary data are available at Bioinformatics online. K. Anton Feenstra, Sanne Abeln, Johan A. Westerhuis, Filipe Brancos dos Santos, Douwe Molenaar, Bas Teusink, Huub C. J. Hoefsloot, Jaap Heringa |
Bioinform. | 8 |
| 2018 | Motif-Aware PRALINE: Improving the alignment of motif regionsabstractProtein or DNA motifs are sequence regions which possess biological importance. These regions are often highly conserved among homologous sequences. The generation of multiple sequence alignments (MSAs) with a correct alignment of the conserved sequence motifs is still difficult to achieve, due to the fact that the contribution of these typically short fragments is overshadowed by the rest of the sequence. Here we extended the PRALINE multiple sequence alignment program with a novel motif-aware MSA algorithm in order to address this shortcoming. This method can incorporate explicit information about the presence of externally provided sequence motifs, which is then used in the dynamic programming step by boosting the amino acid substitution matrix towards the motif. The strength of the boost is controlled by a parameter, α. Using a benchmark set of alignments we confirm that a good compromise can be found that improves the matching of motif regions while not significantly reducing the overall alignment quality. By estimating α on an unrelated set of reference alignments we find there is indeed a strong conservation signal for motifs. A number of typical but difficult MSA use cases are explored to exemplify the problems in correctly aligning functional sequence motifs and how the motif-aware alignment method can be employed to alleviate these problems. Maurits J. J. Dijkstra, Punto Bawono, Sanne Abeln, K. Anton Feenstra, Wan J. Fokkink, Jaap Heringa |
PLoS Comput. Biol. | 6 |
| 2017 | Seeing the trees through the forest: sequence-based homo- and heteromeric protein-protein interaction sites prediction using random forestabstractMOTIVATION: Genome sequencing is producing an ever-increasing amount of associated protein sequences. Few of these sequences have experimentally validated annotations, however, and computational predictions are becoming increasingly successful in producing such annotations. One key challenge remains the prediction of the amino acids in a given protein sequence that are involved in protein-protein interactions. Such predictions are typically based on machine learning methods that take advantage of the properties and sequence positions of amino acids that are known to be involved in interaction. In this paper, we evaluate the importance of various features using Random Forest (RF), and include as a novel feature backbone flexibility predicted from sequences to further optimise protein interface prediction. RESULTS: We observe that there is no single sequence feature that enables pinpointing interacting sites in our Random Forest models. However, combining different properties does increase the performance of interface prediction. Our homomeric-trained RF interface predictor is able to distinguish interface from non-interface residues with an area under the ROC curve of 0.72 in a homomeric test-set. The heteromeric-trained RF interface predictor performs better than existing predictors on a independent heteromeric test-set. We trained a more general predictor on the combined homomeric and heteromeric dataset, and show that in addition to predicting homomeric interfaces, it is also able to pinpoint interface residues in heterodimers. This suggests that our random forest model and the features included capture common properties of both homodimer and heterodimer interfaces. AVAILABILITY AND IMPLEMENTATION: The predictors and test datasets used in our analyses are freely available ( http://www.ibi.vu.nl/downloads/RF_PPI/ ). CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qingzhen Hou, Paul F. G. De Geest, Wim F. Vranken, Jaap Heringa, K. Anton Feenstra |
Bioinform. | 4 |
| 2016 | BioASF: a framework for automatically generating executable pathway models specified in BioPAXabstractMOTIVATION: Biological pathways play a key role in most cellular functions. To better understand these functions, diverse computational and cell biology researchers use biological pathway data for various analysis and modeling purposes. For specifying these biological pathways, a community of researchers has defined BioPAX and provided various tools for creating, validating and visualizing BioPAX models. However, a generic software framework for simulating BioPAX models is missing. Here, we attempt to fill this gap by introducing a generic simulation framework for BioPAX. The framework explicitly separates the execution model from the model structure as provided by BioPAX, with the advantage that the modelling process becomes more reproducible and intrinsically more modular; this ensures natural biological constraints are satisfied upon execution. The framework is based on the principles of discrete event systems and multi-agent systems, and is capable of automatically generating a hierarchical multi-agent system for a given BioPAX model. RESULTS: To demonstrate the applicability of the framework, we simulated two types of biological network models: a gene regulatory network modeling the haematopoietic stem cell regulators and a signal transduction network modeling the Wnt/β-catenin signaling pathway. We observed that the results of the simulations performed using our framework were entirely consistent with the simulation results reported by the researchers who developed the original models in a proprietary language. AVAILABILITY AND IMPLEMENTATION: The framework, implemented in Java, is open source and its source code, documentation and tutorial are available at http://www.ibi.vu.nl/programs/BioASF CONTACT: [email protected]. Reza Haydarlou, Annika Jacobsen, Nicola Bonzanni, K. Anton Feenstra, Sanne Abeln, Jaap Heringa |
Bioinform. | 6 |
| 2016 | ECCB 2016: The 15th European Conference on Computational BiologyabstractThis special issue includes the proceeding papers accepted for presentation at the 15th European Conference on Computational Biology (ECCB 2016), to be held from September 3 to 7, 2016 at the World Forum Convention Center in The Hague, The Netherlands. Details of the conference are available on the conference web site (www.eccb2016.org) and will later be archived at eccb.iscb.org/2016/. ECCB is the premier European conference in computational biology and bioinformatics, and together with ISMB (Intelligent Systems in Molecular Biology) and RECOMB (Research in Computational Molecular Biology), it is one of the major international conference series in this domain. The ECCB conferences are gathering about a thousand scientists and industry staff working at the intersection of a broad range of disciplines including computer science, mathematics, biology and medicine. New challenges are now emerging in these fields with the recent advances in low-cost ultra-fast sequencing, bio-imaging and big data. Computational analysis platforms are challenged by the enormous complexity of biological systems and the sheer amount of data resulting from high-throughput measuring techniques, such as single-cell or single-molecule measurements. As a consequence, databases and software are evolving rapidly, and new algorithms are required to improve computational analyses of massive biological or biomedical datasets. Recent advances are presented at the conference in the field of data interoperability and machine learning, in particular ‘deep learning’ and network-based analysis techniques. The impact on the field of public data repositories such as TCGA (Weinstein et al., 2013), ENCODE (ENCODE Project Consortium, 2012) and GDSC (Yang et al., 2013) is palpable with many submissions revolving around these resources. ECCB is held annually in a different country, while it is held jointly with the ISMB conference biennially. Going back in time, the fourteen previous editions of ECCB have been held in: Dublin, Ireland, together with ISMB (Moreau and Beerenwinkel, 2015); Strasbourg, France (Devignes, 2014); Berlin, Germany, together with ISMB (Ben-Tal, 2013); Basel, Switzerland (Schwede and Iber, 2012); Vienna, Austria, together with ISMB (Gaasterland and Vingron, 2011); Ghent, Belgium (Moreau and Heringa, 2010); Stockholm, Sweden, together with ISMB (Gusfield and Tramontano, 2009); Cagliari, Italy (Tramontano, 2008); Vienna, Austria, together with ISMB (Lengauer et al., 2007); Eilat, Israel (Wolfson and Safer, 2007); Madrid, Spain (Guigo et al., 2005); Glasgow, United Kingdom, together with ISMB (Thornton et al., 2004); Paris, France (Lenhof and Sagot, 2003); and Saarbrücken, Germany (Lengauer, 2002). The ECCB 2016 edition features keynote lectures by distinguished speakers. The opening keynote lecture will be delivered by the 2013 Breakthrough Prize-winner Hans Clevers (Hubrecht Institute and Princess Maxima Centre for Pediatric Oncology, Utrecht, The Netherlands). Further keynote presentations will be given by Amos Tanay (Weizmann Institute, Tel Aviv, Israel), John Marioni (EMBL-EBI, Hinxton, UK), Nuria Lopez-Bigas (Universitat Pompeu Fabra, Barcelona, Spain), Benedict Paten (UCSC, Santa Cruz, USA), Pauline Hogeweg (Utrecht University, Utrecht, The Netherlands) and Christina Leslie (Memorial Sloan Kettering Cancer Center, New York, USA). The conference topics span all areas of methodological developments for computational biology and innovative applications of computational methods to molecular biology and biomedicine. To present a more unified view of where the science has gone over recent years, new to ECCB this year is that the conference presentations are divided over five broad themes: (i) Data (organization, management, categorization, integration, analysis of data, knowledge discovery); (ii) Genome (sequence analysis, alignment, evolution, phylogeny, genetics, epigenetics, 3D conformation); (iii) Genes (expression, function, regulation, transcription, translation, geno/phenotype); (iv) Proteins (structure, function, alterations, assemblies, interactions, design, proteomics) and (v) Systems (systems biology, pathways, molecular networks, dynamics, signalling, multi-scale modelling). This year four different tracks were created for ECCB: two scientific and two application-oriented ones. On the scientific side, the Proceedings Track presents novel scientific contributions, while the Highlights Track showcases already published high-impact science in computational biology. These two tracks were coordinated and managed across the five themes by a board of 5 × 4 = 20 co-chairs, both overseen by a dedicated track chair. A novelty of ECCB this year is that we have given a prominent platform to applications in the new Application and ELIXIR tracks. The Application Track is an initiative to promote application of computational biology in industry and other fields beyond academia. Submissions to this track should cross the boundaries of traditional academic science, or show developments that are directly relevant beyond academia or have potential for it. Consequently, submissions relating to (pure) research were deemed out of scope. The ELIXIR Track, running for the first time at ECCB 2016, is managed by ELIXIR and focuses on developments relating to services and infrastructure within the ELIXIR nodes. ELIXIR is the pan-European life science infrastructural network (‘Data for the Life Sciences’). It coordinates, integrates and sustains bioinformatics resources across its member states, enabling users in academia and industry to access vital data, tools, standards, compute and training services for research. ELIXIR has chosen ECCB as their major dissemination platform, acting as co-organising sponsor. Following the call for Proceedings papers, we received 150 submissions. Submission authors were asked to rank the five themes for fit of their paper. These rankings were subsequently used to assign papers to themes and to arrive at an optimally balanced distribution of submissions over the themes. Within each theme, the theme co-chairs assigned papers to expert referees, taking care to avoid any conflict of interest. Together, the Programme Committee (PC) was composed of 217 reviewers and 35 co-reviewers. The reviewing and selection process was carried out using the EasyChair multi-track conference reviewing system (www.easychair.org). The review form explicitly differentiated between impact and suitability for ECCB on the one hand, and scientific quality and reproducibility on the other. This distinction was intended to create more clarity for both the reviewers and authors; the reviewing criteria were therefore explicitly stated in the submission guidelines. The added focus on reproducibility resulted in many authors opting to make their methods open source. After the reviewers reached a consensus, a final ranking and selection for each theme was carried out by the theme (co-)chairs. A total of 48 papers were conditionally accepted (acceptance ratio of 32%) based on a predefined number of acceptances per theme based on the distribution of initial assignments over the themes. The authors had two weeks to modify their papers according to the suggestions made by the reviewers and to respond to the reviewers’ comments, which was checked by the theme (co-)chairs. We thank the authors for incorporating these suggestions as they were given little time to carry out (minor) revisions. Essentially, these efforts contribute to the success and reputation of the ECCB conference! It is worth noting that many authors expressed their gratitude to the reviewers for their comments and suggestions. We also gratefully acknowledge the hard and diligent work performed by the reviewers over a short period of time. We believe that for all of the rejected submissions, the reviewers provided high-quality reports. We hope that authors of rejected manuscripts will benefit from these remarks in their future research. We thank all theme co-chairs for their availability throughout the reviewing process, for their very positive attitude in the final selection and for their valuable help in re-examining the modified submissions. The 48 accepted papers are included in this special issue. The Proceedings Track papers with their supplementary files are available free-for-view in electronic format from Oxford’s press journal Bioinformatics from September 1st, 2016. Highlight presentations were introduced at ISMB/ECCB 2007 in Vienna (Lengauer et al., 2007), and immediately became one of the most popular features of the conference. All original research papers that had been published in peer-review journals between 1 March, 2015, and the submission deadline of 29 March 2016, were eligible to be presented as a Highlight talk. After thorough consideration and discussion, the ECCB 2016 theme (co-)chairs selected 24 proposals out of 71 submissions, mainly on the criteria of compatibility with the ECCB objectives, wide impact in the life sciences and the potential for attracting a large audience to the conference. The Applications Track features 11 presentations, which were selected out of 31 submissions. At the time of writing, three Sponsored Talks will be delivered as part of the Applications Track by respectively The Hyve, Keygene and Data Computing. Sponsored Talks enable sponsors to showcase their innovations in computational biology and to highlight their scientific value. Finally, the ELIXIR track proved to be a very popular addition, with 50 high-quality submissions for only 12 presentation slots. Submissions to the Poster Track were evaluated based on a 250-word abstract and will be shown at the conference along the five main conference themes. In line with the new Application and ELIXIR tracks, there will also be special Application and ELIXIR poster tracks. A dedicated Education Poster Track was created to devote in-depth attention to education that is so crucial for the next generation of bioinformaticians. All poster abstracts are available on the conference web site. We also arranged with F1000Research (http://f1000research.com/) to publish the posters via a new ECCB2016 channel; submission is on a voluntary basis. At least sixteen exhibitor booths will be open throughout the conference in the central conference hall, which is well connected to the other activities at ECCB: The Hyve, EMBL-EBI, TimeLogic, Springer, ISCB Student Council, Goblet, ELIXIR, ISCB, Oxford University Press, ENPICOM, Data Computing, CRC Press, SIB, ELIXIR Denmark and Cambridge University Press. They will be presenting the latest scientific literature in the field of computational biology, bioinformatics, data stewardship, modeling and simulation, as well as new hardware, software and technology developments. The Hague Tourist Office, and the four organizing institutions (DTL, Netherlands Bioinformatics and System Biology Research School (BioSB), VU University Amsterdam and Delft University of Technology) will also be represented. During the weekend before the conference, a satellite meeting, 14 workshops and nine tutorials will take place. The Student Council of the International Society for Computational Biology (ISCB) organizes its 4th European Student Council Symposium (ESCS), chaired by Annika Jacobsen (Vrije Universiteit Amsterdam) and Kevin Schwahn (Universität Potsdam, Max Planck Institute of Molecular Plant Physiology). ESCS highlights will be published in F1000Research via the ISCB Student Council channel. The ECCB 2016 organizing committee congratulates all these dynamic young scientists for their enthusiasm, which is essential to the future of research in computational biology. The 14 workshops preceding the ECCB 2016 main meeting were selected out of a total of 29 applications, showing how popular ECCB has become as a venue for dissemination of computational biology research. The workshops, running all but one for a single day, provide participants with an informal setting to discuss technical issues, exchange research ideas, and to share practical experiences on a range of focused or emerging topics in computational biology. Taken together, the workshops demonstrate how extensively technologies have found their way into large-scale practical applications: • (W1) The 10th International Workshop on Machine Learning in Systems Biology, organised by (Juho Rousu, Aalto University, Finland), Dick de Ridder Wageningen University, The Netherlands), Harri Lähdesmäki (Aalto University, Finland) and Aalt—Jan van Dijk (Wageningen University, The Netherlands), is a two-day workshop with its own proceedings track, where accepted papers will be published in BMC Bioinformatics, • (W2) Network Inference: New Methods and New Data, organised by Anagha Joshi (Roslin Institute, University of Edinburgh, UK), Tom Michoel (Roslin Institute, University of Edinburgh, UK) and Eric Bonnet (Centre National de Génotypage, CEA, Paris, France) • (W3) Getting the Most out of Your Methods and Algorithms: A Workshop on How to Use Existing Datasets to Gain Novel Insight, chaired by Morris Swertz and Lude Franke (both at University Medical Centre Groningen, The Netherlands), • (W4) Digital Pathology Meets Bioinformatics, organised by Yves Sucaet (Vrije Universiteit Brussel, Belgium), Jeroen Van der Laak (UMC Radboud, Nijmegen, The Netherlands), Marius Nap (HistoGeneX, Belgium and Rigshospitalet Copenhagen, Denmark), Zev Leifer (New York College of Podiatric Medicine, USA), Yukako Yagi (Harvard Medical School, Cambridge and Massachusetts General Hospital, Boston, USA) and Raphaël Marée (Université de Liège, Belgium), • (W5) RepSeq 2016: Immune Repertoire Sequencing—Bioinformatics and Applications in Hematology and Immunology, organised by Jack Bartram (University College London, UK), Eva Froňková (Charles University Prague, Czech Republic), Mathieu Giraud, (CNRS, Lille, France—program co-chair), Peter N. Robinson (Charité Berlin, Germany), Mikaël Salson (Université de Lille, France—proceedings chair), Mikhail Shugay (Shemyakin and Ovchinnikov Institute of Bioorganic Chemistry, Moscow, Russia—program co-chair) and Andrew P. Stubbs (Erasmus MC, Rotterdam, The Netherlands), • (W6) FAIR Data and Data Stewardship, organised by Erik Schultes, Mark Thompson Marco Roos (all three at Leiden University Medical Centre, The Netherlands), Mark Wilkinson (Universidad Politecnica de Madrid, Spain), Luiz Olavo Bonino da Silva Santos (Dutch Techcentre for Life Sciences and Vrije Universiteit Amsterdam, The Netherlands), • (W7) Challenges and Approaches in Comprehensive and Informative Complex Network Analysis for Precision Medicine, organised by Igor Jurisica (University of Toronto, Canada), Natasa Przulj (University College London, UK) and Tijana Milenkovic (University of Notre Dame, USA), • (W8) Computing a Tissue: Modeling Multicellular Systems, organised by Walter de Back (TU Dresden, Germany), Sara Montagna (University of Bologna, Italy) and Roeland Merks (Center for Mathematics and Computer Science (CWI), Amsterdam and Leiden University, The Netherlands), • (W9) Computational Pan-Genomics, organised by Zamin Iqbal (University of Oxford, UK), Tobias Marschall (Saarland University and Max Planck Institute for Informatics, Saarbrücken, Germany) and Benedict Paten (3UC Santa Cruz Genomics Institute, Santa Cruz, CA, USA), • (W10) BioNetVisA: from Biological Network Reconstruction to data visualisation and analysis in Molecular Biology and Medicine, organised by Inna Kuperstein, Emmanuel Barillot, Andrei Zinovyev (all three at Institut Curie, Paris, France), Hiroaki Kitano (Okinawa Institute of Science and Technology Graduate University, RIKEN Center for Integrative Medical Sciences, Japan), Minoru Kanehisa (Kyoto University, Japan), Samik Ghosh (Systems Biology Institute, Tokyo, Japan), Nicolas Le Novère (Babraham Institute, Cambridge, UK), Robin Haw (Ontario Institute for Cancer Research, Canada), Alfonso Valencia (Spanish National Bioinformatics Institute, Madrid, Spain) and Lodewyk Wessels (Netherlands Cancer Institute, Amsterdam, and Technical University Delft, The Netherlands), • (W11) Recent Computational Advances in Metagenomics, organised by Sophie Schbath, Valentin Loux and Mahendra Mariadassou (all at INRA, Jouy-en-Josas, France), • (W12) BioExcel: Advanced Simulations for Biomolecular Research, a SIG workshop organised by Rossen Apostolov (KTH, Stockholm, Sweden)), Alexandre Bonvin (Utrecht University, The Netherlands), Cath Brooksbank (EMBL-EBI, Hinxton, UK) and Ian Harrow (Ian Harrow Consulting, Whitstable, UK), • (W13) Clinical Bioinformatics as a Service, organised by Niko Beerenwinkel (ETH Zurich, and SIB Swiss Institute of Bioinformatics, Switzerland), Wolfgang Huber (EMBL, Heidelberg), Simon Tavaré (Cancer Research UK Cambridge Institute, UK) and Daniel Stekhoven and SIB Swiss Institute of Bioinformatics, Switzerland), • Computational Challenges of Data organised by University Medical Center, The Netherlands), (UMC Utrecht, The Netherlands) and Hans The Netherlands). • Data Analysis with and organised by Jeroen and The Netherlands), • Genome organised by (University of Spain) and (University of Spain), • and organised by (University of Spain), • Analysis and with organised by United • and for Modeling Biological Systems, organised by (University of Germany) and (University of Germany), • Analysis and for Data, organised by Medical Center, Amsterdam and University of Amsterdam, The Netherlands), • A to Research on organised by and (both at University Pompeu Fabra, Barcelona, Spain), • and their in Genome Data, organised by N. (University of Cambridge, UK), University of Science and and (University of UK), • into and its Applications in Bioinformatics and in Network organised by and (both at Hans Institute, and University Hospital, It is to thank all the and that are ECCB 2016 a and high-quality conference. of we are to the committee composed of the theme (co-)chairs and reviewers for their crucial and dedicated We are to the ECCB committee for their and to the of the conference. In the and by Committee and ECCB ECCB Yves ECCB ECCB and ECCB 2012) were Yves and for and The and of ISCB Society for Computational Biology) in the about ECCB 2016 at the international and for were We thank all provided to the conference. In to the co-organising ELIXIR, we gratefully acknowledge sponsors The Netherlands and USA) and the National Centre for Research (CNRS, We also as where also the opening will be the of The Hague, the and The Bioinformatics Centre the Poster We are to the Oxford University for the ECCB 2016 special issue. The F1000Research is for and the ECCB 2016 posters on their web The ECCB 2016 science has been in the of the efforts have been van (Netherlands Bioinformatics and Systems Biology Research Techcentre for Life ECCB 2016 for the selection (TU Delft, The Netherlands), ECCB 2016 for out the workshop selection (EMBL, Germany), ECCB 2016 Applications Track for and the new Applications Andrew and (all at ELIXIR Hinxton, UK) for taking care in of the first ELIXIR Applications and (University of Finland), ECCB 2016 Poster Track for and the ECCB 2016 to the of the conference, beyond the call of and we a in particular and van both at the Techcentre for Life Sciences for their and in this are also to and van of by Netherlands), ECCB 2016 (Dutch Techcentre for Life had a in the while van is more for also in Finally, all these efforts be the many participants from all over the of will to the conference in of scientific contributions, applications, or poster presentations and all for there and for to science at ECCB 2016 in The Jaap Heringa, Marcel J. T. Reinders, Sanne Abeln, Jeroen de Ridder |
Bioinform. | 1 |
| 2016 | metaModules identifies key functional subnetworks in microbiome-related diseaseabstractMOTIVATION: The human microbiome plays a key role in health and disease. Thanks to comparative metatranscriptomics, the cellular functions that are deregulated by the microbiome in disease can now be computationally explored. Unlike gene-centric approaches, pathway-based methods provide a systemic view of such functions; however, they typically consider each pathway in isolation and in its entirety. They can therefore overlook the key differences that (i) span multiple pathways, (ii) contain bidirectionally deregulated components, (iii) are confined to a pathway region. To capture these properties, computational methods that reach beyond the scope of predefined pathways are needed. RESULTS: By integrating an existing module discovery algorithm into comparative metatranscriptomic analysis, we developed metaModules, a novel computational framework for automated identification of the key functional differences between health- and disease-associated communities. Using this framework, we recovered significantly deregulated subnetworks that were indeed recognized to be involved in two well-studied, microbiome-mediated oral diseases, such as butanoate production in periodontal disease and metabolism of sugar alcohols in dental caries. More importantly, our results indicate that our method can be used for hypothesis generation based on automated discovery of novel, disease-related functional subnetworks, which would otherwise require extensive and laborious manual assessment. AVAILABILITY AND IMPLEMENTATION: metaModules is available at https://bitbucket.org/alimay/metamodules/ CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ali May, Bernd W. Brandt, Mohammed El-Kebir, Gunnar W. Klau, Egija Zaura, Wim Crielaard, Jaap Heringa, Sanne Abeln |
Bioinform. | 7 |
| 2015 | xHeinz: an algorithm for mining cross-species network modules under a flexible conservation modelabstractMOTIVATION: Integrative network analysis methods provide robust interpretations of differential high-throughput molecular profile measurements. They are often used in a biomedical context-to generate novel hypotheses about the underlying cellular processes or to derive biomarkers for classification and subtyping. The underlying molecular profiles are frequently measured and validated on animal or cellular models. Therefore the results are not immediately transferable to human. In particular, this is also the case in a study of the recently discovered interleukin-17 producing helper T cells (Th17), which are fundamental for anti-microbial immunity but also known to contribute to autoimmune diseases. RESULTS: We propose a mathematical model for finding active subnetwork modules that are conserved between two species. These are sets of genes, one for each species, which (i) induce a connected subnetwork in a species-specific interaction network, (ii) show overall differential behavior and (iii) contain a large number of orthologous genes. We propose a flexible notion of conservation, which turns out to be crucial for the quality of the resulting modules in terms of biological interpretability. We propose an algorithm that finds provably optimal or near-optimal conserved active modules in our model. We apply our algorithm to understand the mechanisms underlying Th17 T cell differentiation in both mouse and human. As a main biological result, we find that the key regulation of Th17 differentiation is conserved between human and mouse. AVAILABILITY AND IMPLEMENTATION: xHeinz, an implementation of our algorithm, as well as all input data and results, are available at http://software.cwi.nl/xheinz and as a Galaxy service at http://services.cbib.u-bordeaux2.fr/galaxy in CBiB Tools. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mohammed El-Kebir, Hayssam Soueidan, Thomas Hume, Daniela Beisser, Marcus T. Dittrich, Tobias Müller 0001, Guillaume Blin, Jaap Heringa, Macha Nikolski, Lodewyk F. A. Wessels, Gunnar W. Klau |
Bioinform. | 8 |
| 2015 | Sequence specificity between interacting and non-interacting homologs identifies interface residues - a homodimer and monomer use caseabstractBACKGROUND: Protein families participating in protein-protein interactions may contain sub-families that have different binding characteristics, ranging from right binding to showing no interaction at all. Composition differences at the sequence level in these sub-families are often decisive to their differential functional interaction. Methods to predict interface sites from protein sequences typically exploit conservation as a signal. Here, instead, we provide proof of concept that the sequence specificity between interacting versus non-interacting groups can be exploited to recognise interaction sites. RESULTS: We collected homodimeric and monomeric proteins and formed homologous groups, each having an interacting (homodimer) subgroup and a non-interacting (monomer) subgroup. We then compiled multiple sequence alignments of the proteins in the homologous groups and identified compositional differences between the homodimeric and monomeric subgroups for each of the alignment positions. Our results show that this specificity signal distinguishes interface and other surface residues with 40.9% recall and up to 25.1% precision. CONCLUSIONS: To our best knowledge, this is the first large scale study that exploits sequence specificity between interacting and non-interacting homologs to predict interaction sites from sequence information only. The performance obtained indicates that this signal contains valuable information to identify protein-protein interaction sites. Qingzhen Hou, Bas E. Dutilh, Martijn A. Huynen, Jaap Heringa, K. Anton Feenstra |
BMC Bioinform. | 4 |
| 2014 | Unraveling the outcome of 16S rDNA-based taxonomy analysis through mock data and simulationsabstractMOTIVATION: 16S rDNA pyrosequencing is a powerful approach that requires extensive usage of computational methods for delineating microbial compositions. Previously, it was shown that outcomes of studies relying on this approach vastly depend on the choice of pre-processing and clustering algorithms used. However, obtaining insights into the effects and accuracy of these algorithms is challenging due to difficulties in generating samples of known composition with high enough diversity. Here, we use in silico microbial datasets to better understand how the experimental data are transformed into taxonomic clusters by computational methods. RESULTS: We were able to qualitatively replicate the raw experimental pyrosequencing data after rigorously adjusting existing simulation software. This allowed us to simulate datasets of real-life complexity, which we used to assess the influence and performance of two widely used pre-processing methods along with 11 clustering algorithms. We show that the choice, order and mode of the pre-processing methods have a larger impact on the accuracy of the clustering pipeline than the clustering methods themselves. Without pre-processing, the difference between the performances of clustering methods is large. Depending on the clustering algorithm, the most optimal analysis pipeline resulted in significant underestimations of the expected number of clusters (minimum: 3.4%; maximum: 13.6%), allowing us to make quantitative estimations of the bacterial complexity of real microbiome samples. Ali May, Sanne Abeln, Wim Crielaard, Jaap Heringa, Bernd W. Brandt |
Bioinform. | 4 |
| 2014 | Coarse-grained versus atomistic simulations: realistic interaction free energies for real proteinsabstractMOTIVATION: To assess whether two proteins will interact under physiological conditions, information on the interaction free energy is needed. Statistical learning techniques and docking methods for predicting protein-protein interactions cannot quantitatively estimate binding free energies. Full atomistic molecular simulation methods do have this potential, but are completely unfeasible for large-scale applications in terms of computational cost required. Here we investigate whether applying coarse-grained (CG) molecular dynamics simulations is a viable alternative for complexes of known structure. RESULTS: We calculate the free energy barrier with respect to the bound state based on molecular dynamics simulations using both a full atomistic and a CG force field for the TCR-pMHC complex and the MP1-p14 scaffolding complex. We find that the free energy barriers from the CG simulations are of similar accuracy as those from the full atomistic ones, while achieving a speedup of >500-fold. We also observe that extensive sampling is extremely important to obtain accurate free energy barriers, which is only within reach for the CG models. Finally, we show that the CG model preserves biological relevance of the interactions: (i) we observe a strong correlation between evolutionary likelihood of mutations and the impact on the free energy barrier with respect to the bound state; and (ii) we confirm the dominant role of the interface core in these interactions. Therefore, our results suggest that CG molecular simulations can realistically be used for the accurate prediction of protein-protein interaction strength. AVAILABILITY AND IMPLEMENTATION: The python analysis framework and data files are available for download at http://www.ibi.vu.nl/downloads/bioinformatics-2013-btt675.tgz. Ali May, René Pool, Erik van Dijk, Jochem Bijlard, Sanne Abeln, Jaap Heringa, K. Anton Feenstra |
Bioinform. | 6 |
| 2013 | Bioinformatics and Systems Biology: bridging the gap between heterogeneous student backgroundsabstractTeaching students with very diverse backgrounds can be extremely challenging. This article uses the Bioinformatics and Systems Biology MSc in Amsterdam as a case study to describe how the knowledge gap for students with heterogeneous backgrounds can be bridged. We show that a mix in backgrounds can be turned into an advantage by creating a stimulating learning environment for the students. In the MSc Programme, conversion classes help to bridge differences between students, by mending initial knowledge and skill gaps. Mixing students from different backgrounds in a group to solve a complex task creates an opportunity for the students to reflect on their own abilities. We explain how a truly interdisciplinary approach to teaching helps students of all backgrounds to achieve the MSc end terms. Moreover, transferable skills obtained by the students in such a mixed study environment are invaluable for their later careers. Sanne Abeln, Douwe Molenaar, K. Anton Feenstra, Huub C. J. Hoefsloot, Bas Teusink, Jaap Heringa |
Briefings Bioinform. | 6 |
| 2013 | Hard-wired heterogeneity in blood stem cells revealed using a dynamic regulatory network modelabstractMOTIVATION: Combinatorial interactions of transcription factors with cis-regulatory elements control the dynamic progression through successive cellular states and thus underpin all metazoan development. The construction of network models of cis-regulatory elements, therefore, has the potential to generate fundamental insights into cellular fate and differentiation. Haematopoiesis has long served as a model system to study mammalian differentiation, yet modelling based on experimentally informed cis-regulatory interactions has so far been restricted to pairs of interacting factors. Here, we have generated a Boolean network model based on detailed cis-regulatory functional data connecting 11 haematopoietic stem/progenitor cell (HSPC) regulator genes. RESULTS: Despite its apparent simplicity, the model exhibits surprisingly complex behaviour that we charted using strongly connected components and shortest-path analysis in its Boolean state space. This analysis of our model predicts that HSPCs display heterogeneous expression patterns and possess many intermediate states that can act as 'stepping stones' for the HSPC to achieve a final differentiated state. Importantly, an external perturbation or 'trigger' is required to exit the stem cell state, with distinct triggers characterizing maturation into the various different lineages. By focusing on intermediate states occurring during erythrocyte differentiation, from our model we predicted a novel negative regulation of Fli1 by Gata1, which we confirmed experimentally thus validating our model. In conclusion, we demonstrate that an advanced mammalian regulatory network model based on experimentally validated cis-regulatory interactions has allowed us to make novel, experimentally testable hypotheses about transcriptional mechanisms that control differentiation of mammalian stem cells. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Nicola Bonzanni, Abhishek Garg, K. Anton Feenstra, Judith Schütte, Sarah Kinston, Diego Miranda-Saavedra, Jaap Heringa, Ioannis Xenarios, Berthold Göttgens |
Bioinform. | 7 |
| 2013 | Mapping proteins in the presence of paralogs using units of coevolutionabstractBACKGROUND: We study the problem of mapping proteins between two protein families in the presence of paralogs. This problem occurs as a difficult subproblem in coevolution-based computational approaches for protein-protein interaction prediction. RESULTS: Similar to prior approaches, our method is based on the idea that coevolution implies equal rates of sequence evolution among the interacting proteins, and we provide a first attempt to quantify this notion in a formal statistical manner. We call the units that are central to this quantification scheme the units of coevolution. A unit consists of two mapped protein pairs and its score quantifies the coevolution of the pairs. This quantification allows us to provide a maximum likelihood formulation of the paralog mapping problem and to cast it into a binary quadratic programming formulation. CONCLUSION: CUPID, our software tool based on a Lagrangian relaxation of this formulation, makes it, for the first time, possible to compute state-of-the-art quality pairings in a few minutes of runtime. In summary, we suggest a novel alternative to the earlier available approaches, which is statistically sound and computationally feasible. Mohammed El-Kebir, Tobias Marschall, Inken Wohlers, Murray Patterson, Jaap Heringa, Alexander Schönhuth, Gunnar W. Klau |
BMC Bioinform. | 5 |
| 2010 | Computational quantification of metabolic fluxes from a single isotope snapshot: application to an animal biopsyabstractMOTIVATION: Quantitative determination of metabolic fluxes in single tissue biopsies is difficult. We report a novel analysis approach and software package for in vivo flux quantification using stable isotope labeling. RESULTS: We developed a protocol based on brief, timed infusion of (13)C isotope-enriched substrates for the tricarboxylic acid (TCA) cycle followed by quick freezing of tissue biopsies. NMR measurements of tissue extracts were used for flux estimation based on a computational model of carbon transitions between TCA cycle metabolites and related amino acids. To this end, we developed a computational framework in which metabolic systems can be flexibly assembled, simulated and analyzed. Flux parameters were quantified from NMR multiplets by a partial grid search followed by repeated Nelder-Mead optimizations implemented on a computer grid. We implemented a model of the TCA cycle and showed by extensive simulations that the timed infusion protocol reliably quantitates multiple fluxes. Experimental validation of the method was done in vivo on hearts of anesthetized pigs under two different conditions: basal state (n = 7) and cardiac stress caused by infusion of dobutamine (n = 7). About nine tissue samples (40-200 mg dry-weight) were taken per heart. TCA cycle flux was 6.11 +/- 0.28 (SEM) micromol/min x gdw at baseline versus 9.29 +/- 1.03 micromol/min x gdw for dobutamine stress. Oxygen consumption calculated from the TCA cycle flux and from 'gold standard' blood gas-based measurements were close, correlating with r=0.88 (P < 10(-4)). Spatial heterogeneity in metabolic fluxes is detectable amongst the small samples. We propose that our novel isotope snapshot methodology is suitable for flux measurements in biopsies in vivo. AVAILABILITY: Non-profit organizations will, upon request, be granted a non-exclusive license to use the software for internal research and teaching purposes at no charge. A web interface for using the software on our computer grid is available under http://www.ibi.vu.nl/programs/ Thomas W. Binsl, David J. C. Alders, Jaap Heringa, A. B. Johan Groeneveld, Johannes H. G. M. van Beek |
Bioinform. | 3 |
| 2010 | CGHnormaliter: a Bioconductor package for normalization of array CGH data with many CNAsabstractSUMMARY: CGHnormaliter is a package for normalization of array comparative genomic hybridization (aCGH) data. It uses an iterative procedure that effectively eliminates the influence of imbalanced copy numbers. This leads to a more reliable assessment of copy number alterations (CNAs). CGHnormaliter is integrated in the Bioconductor environment allowing a smooth link to visualization tools and further data analysis. AVAILABILITY AND IMPLEMENTATION: The CGHnormaliter package is implemented in R and under GPL 3.0 license available at Bioconductor: http://www.bioconductor.org CONTACT: [email protected] Bart P. P. van Houte, Thomas W. Binsl, Hannes Hettling, Jaap Heringa |
Bioinform. | 4 |
| 2010 | Accurate confidence aware clustering of array CGH tumor profilesabstractMOTIVATION: Chromosomal aberrations tend to be characteristic for given (sub)types of cancer. Such aberrations can be detected with array comparative genomic hybridization (aCGH). Clustering aCGH tumor profiles aids in identifying chromosomal regions of interest and provides useful diagnostic information on the cancer type. An important issue here is to what extent individual aCGH tumor profiles can be reliably assigned to clusters associated with a given cancer type. RESULTS: We introduce a novel evolutionary fuzzy clustering (EFC) algorithm, which is able to deal with overlapping clusterings. Our method assesses these overlaps by using cluster membership degrees, which we use here as a confidence measure for individual samples to be assigned to a given tumor type. We first demonstrate the usefulness of our method using a synthetic aCGH dataset and subsequently show that EFC outperforms existing methods on four real datasets of aCGH tumor profiles involving four different cancer types. We also show that in general best performance is obtained using 1- Pearson correlation coefficient as a distance measure and that extra preprocessing steps, such as segmentation and calling, lead to decreased clustering performance. AVAILABILITY: The source code of the program is available from http://ibi.vu.nl/programs/efcwww Bart P. P. van Houte, Jaap Heringa |
Bioinform. | 2 |
| 2010 | Analysis of the functional properties of the creatine kinase system using a multiscale 'sloppy' modeling approachabstractBackground Distinct functions have been hypothesized for the creatine kinase (CK) enzyme catalyzing the reversible transfer of the high-energy phosphate group of ATP to creatine. In muscle cells, two CK isoforms might mediate (i) temporal energy buffering to maintain ATP homeostasis and (ii) energy transport from mitochondria to myofibrils via the “phosphocreatine shuttle” mechanism. Here we investigate the relative importance of the two roles using a computational model [1]. Model simulations predict a contribution of CK to the net transcytosolic energy transport of less than 1/3. Hannes Hettling, Jaap Heringa, Johannes H. G. M. van Beek |
BMC Bioinform. | 2 |
| 2009 | Executing multicellular differentiation: quantitative predictive modelling of C.elegans vulval developmentabstractMOTIVATION: Understanding the processes involved in multi-cellular pattern formation is a central problem of developmental biology, hopefully leading to many new insights, e.g. in the treatment of various diseases. Defining suitable computational techniques for development modelling, able to perform in silico simulation experiments, is an open and challenging problem. RESULTS: Previously, we proposed a coarse-grained, quantitative approach based on the basic Petri net formalism, to mimic the behaviour of the biological processes during multicellular differentiation. Here, we apply our modelling approach to the well-studied process of Caenorhabditis elegans vulval development. We show that our model correctly reproduces a large set of in vivo experiments with statistical accuracy. It also generates gene expression time series in accordance with recent biological evidence. Finally, we modelled the role of microRNA mir-61 during vulval development and predict its contribution in stabilizing cell pattern formation. Nicola Bonzanni, Elzbieta Krepska, K. Anton Feenstra, Wan J. Fokkink, Thilo Kielmann, Henri E. Bal, Jaap Heringa |
Bioinform. | 7 |
| 2009 | Executing multicellular differentiation: quantitative predictive modelling of C.elegans vulval developmentabstractBioinformatics 25(16), 2049–2056 We regret that Figure 5 on page 4 of this paper was incorrect and should appear as below. Nicola Bonzanni, Elzbieta Krepska, K. Anton Feenstra, Wan J. Fokkink, Thilo Kielmann, Henri E. Bal, Jaap Heringa |
Bioinform. | 7 |
| 2009 | Structure and function analysis of flexible alignment regions in proteins
Walter Pirovano, Anneke van der Reijden, K. Anton Feenstra, Jaap Heringa |
BMC Bioinform. | 4 |
| 2008 | PRALINETM: a strategy for improved multiple alignment of transmembrane proteinsabstractMOTIVATION: Membrane-bound proteins are a special class of proteins. The regions that insert into the cell-membrane have a profoundly different hydrophobicity pattern compared with soluble proteins. Multiple alignment techniques use scoring schemes tailored for sequences of soluble proteins and are therefore in principle not optimal to align membrane-bound proteins. RESULTS: Transmembrane (TM) regions in protein sequences can be reliably recognized using state-of-the-art sequence prediction techniques. Furthermore, membrane-specific scoring matrices are available. We have developed a new alignment method, called PRALINETM, which integrates these two features to enhance multiple sequence alignment. We tested our algorithm on the TM alignment benchmark set by Bahr et al. (2001), and showed that the quality of TM alignments can be significantly improved compared with the quality produced by a standard multiple alignment technique. The results clearly indicate that the incorporation of these new elements into current state-of-the-art alignment methods is crucial for optimizing the alignment of TM proteins. AVAILABILITY: A webserver is available at http://www.ibi.vu.nl/programs/pralinewww. Walter Pirovano, K. Anton Feenstra, Jaap Heringa |
Bioinform. | 3 |
| 2008 | Multi-RELIEF: a method to recognize specificity determining residues from multiple sequence alignments using a Machine-Learning approach for feature weightingabstractMOTIVATION: Identification of residues that account for protein function specificity is crucial, not only for understanding the nature of functional specificity, but also for protein engineering experiments aimed at switching the specificity of an enzyme, regulator or transporter. Available algorithms generally use multiple sequence alignments to identify residue positions conserved within subfamilies but divergent in between. However, many biological examples show a much subtler picture than simple intra-group conservation versus inter-group divergence. RESULTS: We present multi-RELIEF, a novel approach for identifying specificity residues that is based on RELIEF, a state-of-the-art Machine-Learning technique for feature weighting. It estimates the expected 'local' functional specificity of residues from an alignment divided in multiple classes. Optionally, 3D structure information is exploited by increasing the weight of residues that have high-weight neighbors. Using ROC curves over a large body of experimental reference data, we show that (a) multi-RELIEF identifies specificity residues for the seven test sets used, (b) incorporating structural information improves prediction for specificity of interaction with small molecules and (c) comparison of multi-RELIEF with four other state-of-the-art algorithms indicates its robustness and best overall performance. AVAILABILITY: A web-server implementation of multi-RELIEF is available at www.ibi.vu.nl/programs/multirelief. Matlab source code of the algorithm and data sets are available on request for academic use. Kai Ye 0001, K. Anton Feenstra, Jaap Heringa, Adriaan P. IJzerman, Elena Marchiori |
Bioinform. | 3 |
| 2008 | The meaning of alignment: lessons from structural diversityabstractBACKGROUND: Protein structural alignment provides a fundamental basis for deriving principles of functional and evolutionary relationships. It is routinely used for structural classification and functional characterization of proteins and for the construction of sequence alignment benchmarks. However, the available techniques do not fully consider the implications of protein structural diversity and typically generate a single alignment between sequences. RESULTS: We have taken alternative protein crystal structures and generated simulation snapshots to explicitly investigate the impact of structural changes on the alignments. We show that structural diversity has a significant effect on structural alignment. Moreover, we observe alignment inconsistencies even for modest spatial divergence, implying that the biological interpretation of alignments is less straightforward than commonly assumed. A salient example is the GroES 'mobile loop' where sub-Angstrom variations give rise to contradictory sequence alignments. CONCLUSION: A comprehensive treatment of ambiguous alignment regions is crucial for further development of structural alignment applications and for the representation of alignments in general. For this purpose we have developed an on-line database containing our data and new ways of visualizing alignment inconsistencies, which can be found at http://www.ibi.vu.nl/databases/stralivari. Walter Pirovano, K. Anton Feenstra, Jaap Heringa |
BMC Bioinform. | 3 |
| 2006 | A Feature Selection Algorithm for Detecting Subtype Specific Functional Sites from Protein Sequences for Smad Receptor BindingabstractMultiple sequence alignments are often used to reveal functionally important residues within a protein family. In particular, they can be very useful for identification of key residues that determine functional differences between protein subclasses (subtype specific sites). This paper proposes a new algorithm for selecting subtype specific sites from a set of aligned protein sequences. The algorithm combines a feature selection technique with neighbor position information for selecting and ranking a set of putative relevant sites. The algorithm is applied to a dataset of protein sequences from the MH2 domain of the SMAD family of transcriptor factors. Validation of the results on the basis of the known interaction and function of the sites shows that the algorithm successfully identifies the known (from literature) subtype specific sites and new putative ones Elena Marchiori, Walter Pirovano, Jaap Heringa, K. Anton Feenstra |
ICMLA | 3 |
| 2006 | AuberGene - a sensitive genome alignment toolabstractMOTIVATION: The accumulation of genome sequences will only accelerate in the coming years. We aim to use this abundance of data to improve the quality of genomic alignments and devise a method which is capable of detecting regions evolving under weak or no evolutionary constraints. RESULTS: We describe a genome alignment program AuberGene, which explores the idea of transitivity of local alignments. Assessment of the program was done based on a 2 Mbp genomic region containing the CFTR gene of 13 species. In this region, we can identify 53% of human sequence sharing common ancestry with mouse, as compared with 44% found using the usual pairwise alignment. Between human and tetraodon 93 orthologous exons are found, as compared with 77 detected by the pairwise human-tetraodon comparison. AuberGene allows the user to (1) identify distant, previously undetected, conserved orthogonal regions such as ORFs or regulatory regions; (2) identify neutrally evolving regions in related species which are often overlooked by other alignment programs; (3) recognize false orthologous genomic regions. The increased sensitivity of the method is not obtained at the cost of reduced specificity. Our results suggest that, over the CFTR region, human shares 10% more sequence with mouse than previously thought ( approximately 50%, instead of 40% found with the pairwise alignment). Radek Szklarczyk, Jaap Heringa |
Bioinform. | 2 |
| 2005 | A simple and fast secondary structure prediction method using hidden neural networksabstractMOTIVATION: In this paper, we present a secondary structure prediction method YASPIN that unlike the current state-of-the-art methods utilizes a single neural network for predicting the secondary structure elements in a 7-state local structure scheme and then optimizes the output using a hidden Markov model, which results in providing more information for the prediction. RESULTS: YASPIN was compared with the current top-performing secondary structure prediction methods, such as PHDpsi, PROFsec, SSPro2, JNET and PSIPRED. The overall prediction accuracy on the independent EVA5 sequence set is comparable with that of the top performers, according to the Q3, SOV and Matthew's correlations accuracy measures. YASPIN shows the highest accuracy in terms of Q3 and SOV scores for strand prediction. AVAILABILITY: YASPIN is available on-line at the Centre for Integrative Bioinformatics website (http://ibivu.cs.vu.nl/programs/yaspinwww/) at the Vrije University in Amsterdam and will soon be mirrored on the Mathematical Biology website (http://www.mathbio.nimr.mrc.ac.uk) at the NIMR in London. CONTACT: [email protected] Kuang Lin, Victor A. Simossis, William R. Taylor, Jaap Heringa |
Bioinform. | 4 |
| 2003 | A Million-Fold Speed Improvement in Genomic Repeats DetectionabstractThis paper presents a novel, parallel algorithm for generating top alignments. Top alignments are used for finding internal repeats in biological sequences like proteins and genes. Our algorithm replaces an older, sequential algorithm (Repro), which was prohibitively slow for sequence lengths higher than 2000. The new algorithm is an order of magnitude faster (O n 3 rather than O n 4). The paper presents a three-level parallel implementation of the algorithm: using SIMD multimedia extensions found on present-day processors (a novel technique that can be used to parallelize any application that performs many sequence alignments), using shared-memory parallelism, and using distributed-memory parallelism. It allows processing the longest known proteins (nearly 35000 amino acids). We show exceptionally high speed improvements: between 500 and 831 on a cluster of 64 dualprocessor machines, compared to the new sequential algorithm. Especially for long sequences, extreme speed improvements over the old algorithm are obtained. 1 John W. Romein, Jaap Heringa, Henri E. Bal |
SC | 2 |
| 2002 | Parallelized multiple alignmentabstractUNLABELLED: Multiple sequence alignment is a frequently used technique for analyzing sequence relationships. Compilation of large alignments is computationally expensive, but processing time can be considerably reduced when the computational load is distributed over many processors. Parallel processing functionality in the form of single-instruction multiple-data (SIMD) technology was implemented into the multiple alignment program Praline by using 'message passing interface' (MPI) routines. Over the alignments tested here, the parallelized program performed up to ten times faster on 25 processors compared to the single processor version. AVAILABILITY: Example program code for parallelizing pairwise alignment loops is available from http://mathbio.nimr.mrc.ac.uk/~jkleinj/tools/mpicode. The 'message passing interface' package (MPICH) is available from http:/www.unix.mcs.anl.gov/mpi/mpich. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Praline is accessible at http://mathbio.nimr.mrc.ac.uk/praline. Jens Kleinjung, Nigel Douglas, Jaap Heringa |
Bioinform. | 3 |
| 1992 | OBSTRUCT: a program to obtain largest cliques from a protein sequence set according to structural resolution and sequence similarityabstractA program OBSTRUCT has been developed to obtain the largest possible subset according to specific constraints from a set of protein sequences whose tertiary structures have been determined crystallographically. The user can request a range in sequence similarity level and/or structural resolution. The program optionally includes sequences with known three-dimensional folds elicited from NMR data. Jaap Heringa, Hubert Sommerfeldt, Desmond G. Higgins, Patrick Argos |
Comput. Appl. Biosci. | 1 |