Martin Kuiper

dblp:26/3716 · also Martin T. R. Kuiper · DBLP profile ↗
← Back
27ranked-venue papers
0as first author
4since 2021 · last 2021
0000-0002-1171-9876ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 26 · 4 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2021 Setting the basis of best practices and standards for curation and annotation of logical models in biology - highlights of the [BC]2 2019 CoLoMoTo/SysMod Workshop
abstract
The fast accumulation of biological data calls for their integration, analysis and exploitation through more systematic approaches. The generation of novel, relevant hypotheses from this enormous quantity of data remains challenging. Logical models have long been used to answer a variety of questions regarding the dynamical behaviours of regulatory networks. As the number of published logical models increases, there is a pressing need for systematic model annotation, referencing and curation in community-supported and standardised formats. This article summarises the key topics and future directions of a meeting entitled 'Annotation and curation of computational models in biology', organised as part of the 2019 [BC]2 conference. The purpose of the meeting was to develop and drive forward a plan towards the standardised annotation of logical models, review and connect various ongoing projects of experts from different communities involved in the modelling and annotation of molecular biological entities, interactions, pathways and models. This article defines a roadmap towards the annotation and curation of logical models, including milestones for best practices and minimum standard requirements.
Anna Niarakis, Martin Kuiper, Marek Ostaszewski, Rahuman S. Malik-Sheriff, Cristina Casals-Casas, Denis Thieffry, Tom C. Freeman, Paul D. Thomas, Vasundra Touré, Vincent Noel, Gautier Stoll, Julio Saez-Rodriguez, Aurélien Naldi, Eugenia Oshurko, Ioannis Xenarios, Sylvain Soliman, Claudine Chaouiya, Tomás Helikar, Laurence Calzone
Briefings Bioinform.2
2021 The status of causality in biological databases: data resources and data retrieval possibilities to support logical modeling
abstract
Causal molecular interactions represent key building blocks used in computational modeling, where they facilitate the assembly of regulatory networks. Logical regulatory networks can be used to predict biological and cellular behaviors by system perturbations and in silico simulations. Today, broad sets of causal interactions are available in a variety of biological knowledge resources. However, different visions, based on distinct biological interests, have led to the development of multiple ways to describe and annotate causal molecular interactions. It can therefore be challenging to efficiently explore various resources of causal interaction and maintain an overview of recorded contextual information that ensures valid use of the data. This review lists the different types of public resources with causal interactions, the different views on biological processes that they represent, the various data formats they use for data representation and storage, and the data exchange and conversion procedures that are available to extract and download these interactions. This may further raise awareness among the targeted audience, i.e. logical modelers and other scientists interested in molecular causal interactions, but also database managers and curators, about the abundance and variety of causal molecular interaction data, and the variety of tools and approaches to convert them into one interoperable resource.
Vasundra Touré, Åsmund Flobak, Anna Niarakis, Steven Vercruysse, Martin Kuiper
Briefings Bioinform.5
2021 The Minimum Information about a Molecular Interaction CAusal STatement (MI2CAST)
abstract
MOTIVATION: A large variety of molecular interactions occurs between biomolecular components in cells. When a molecular interaction results in a regulatory effect, exerted by one component onto a downstream component, a so-called 'causal interaction' takes place. Causal interactions constitute the building blocks in our understanding of larger regulatory networks in cells. These causal interactions and the biological processes they enable (e.g. gene regulation) need to be described with a careful appreciation of the underlying molecular reactions. A proper description of this information enables archiving, sharing and reuse by humans and for automated computational processing. Various representations of causal relationships between biological components are currently used in a variety of resources. RESULTS: Here, we propose a checklist that accommodates current representations, called the Minimum Information about a Molecular Interaction CAusal STatement (MI2CAST). This checklist defines both the required core information, as well as a comprehensive set of other contextual details valuable to the end user and relevant for reusing and reproducing causal molecular interaction information. The MI2CAST checklist can be used as reporting guidelines when annotating and curating causal statements, while fostering uniformity and interoperability of the data across resources. AVAILABILITY AND IMPLEMENTATION: The checklist together with examples is accessible at https://github.com/MI2CAST/MI2CAST. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vasundra Touré, Steven Vercruysse, Marcio Luis Acencio, Ruth C. Lovering, Sandra E. Orchard, Glyn Bradley, Cristina Casals-Casas, Claudine Chaouiya, Noemi del-Toro, Åsmund Flobak, Pascale Gaudet, Henning Hermjakob, Charles Tapley Hoyt, Luana Licata, Astrid Lægreid, Chris Mungall, Anne Niknejad, Simona Panni, Livia Perfetto, Pablo Porras, Dexter Pratt, Julio Saez-Rodriguez, Denis Thieffry, Paul D. Thomas, Dénes Türei, Martin Kuiper
Bioinform.26
2021 UniBioDicts: Unified access to Biological Dictionaries
abstract
SUMMARY: We present a set of software packages that provide uniform access to diverse biological vocabulary resources that are instrumental for current biocuration efforts and tools. The Unified Biological Dictionaries (UniBioDicts or UBDs) provide a single query-interface for accessing the online API services of leading biological data providers. Given a search string, UBDs return a list of matching term, identifier and metadata units from databases (e.g. UniProt), controlled vocabularies (e.g. PSI-MI) and ontologies (e.g. GO, via BioPortal). This functionality can be connected to input fields (user-interface components) that offer autocomplete lookup for these dictionaries. UBDs create a unified gateway for accessing life science concepts, helping curators find annotation terms across resources (based on descriptive metadata and unambiguous identifiers), and helping data users search and retrieve the right query terms. AVAILABILITY AND IMPLEMENTATION: The UBDs are available through npm and the code is available in the GitHub organisation UniBioDicts (https://github.com/UniBioDicts) under the Affero GPL license. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
John Zobolas, Vasundra Touré, Martin Kuiper, Steven Vercruysse
Bioinform.3
2020 The Cytoscape BioGateway App: explorative network building from an RDF store
abstract
SUMMARY: The BioGateway App is a Cytoscape (version 3) plugin designed to provide easy query access to the BioGateway RDF triple store, which contains functional and interaction information for proteins from several curated resources. For explorative network building, we have added a comprehensive dataset with regulatory relationships of mammalian DNA binding transcription factors and their target genes, compiled both from curated resources and from a text mining effort. Query results are visualised using the inherent flexibility of the Cytoscape framework, and network links can be checked against curated database records or against the original publication. AVAILABILITY: Install through the Cytoscape application manager or visit www.biogateway.eu for download and tutorial documents. SUPPLEMENTARY INFORMATION: Supplementary information is available at Bioinformatics online.
Stian Holmås, Rafael Riudavets Puig, Marcio Luis Acencio, Vladimir Mironov, Martin Kuiper
Bioinform.5
2020 A foundation for reference models for drug combinations with an application to Loewe's reference model
abstract
BACKGROUND: Treating patients with combinations of drugs that have synergistic effects has become widespread practice in the clinic. Drugs work synergistically when the observed effect of a drug combination is larger than the effect predicted by the reference model. The reference model is a theoretical null model that returns the combined effect of given doses of drugs under the assumption that these drugs do not interact. There is ongoing debate on what it means for drugs to not interact. The controversy transcends mathematical punctuality, as different non-interaction principles result in different reference models. A famous reference model that has been in existence for already a long time is Loewe's reference model. Loewe's vision on non-interaction was purely intuitive: two drugs do not interact if all combinations of doses that result in a certain given effect lie on a straight line. RESULTS: We show that Loewe's reference model can be obtained from much more fundamental principles. First, we introduce the new notion of complementary dose. Secondly, we reformulate the existing concept of equivalent dose, whereby our formulation is more general than existing ones. Finally, a very general non-interaction principle is put forward. The proposed non-interaction principle represents a certain interplay between complementary and equivalent doses: drugs are non-interacting if complementarity is preserved under equivalence. It is then shown that Loewe's reference model naturally follows from these principles by an appropriate choice of complementarity. CONCLUSIONS: The presented work increases insight into Loewe's reference model for drug combinations, which is realized by the introduction of a very general non-interaction principle that does not refer to any specific dose-response curve, nor to any property of applicable dose-response curves.
Wim De Mulder, Martin Kuiper
BMC Bioinform.2
2019 CausalTAB: the PSI-MITAB 2.8 updated format for signalling data representation and dissemination
abstract
MOTIVATION: Combining multiple layers of information underlying biological complexity into a structured framework represent a challenge in systems biology. A key task is the formalization of such information in models describing how biological entities interact to mediate the response to external and internal signals. Several databases with signalling information, focus on capturing, organizing and displaying signalling interactions by representing them as binary, causal relationships between biological entities. The curation efforts that build these individual databases demand a concerted effort to ensure interoperability among resources. RESULTS: Aware of the enormous benefits of standardization efforts in the molecular interaction research field, representatives of the signalling network community agreed to extend the PSI-MI controlled vocabulary to include additional terms representing aspects of causal interactions. Here, we present a common standard for the representation and dissemination of signalling information: the PSI Causal Interaction tabular format (CausalTAB) which is an extension of the existing PSI-MI tab-delimited format, now designated PSI-MITAB 2.8. We define the new term 'causal interaction', and related child terms, which are children of the PSI-MI 'molecular interaction' term. The new vocabulary terms in this extended PSI-MI format will enable systems biologists to model large-scale signalling networks more precisely and with higher coverage than before. AVAILABILITY AND IMPLEMENTATION: PSI-MITAB 2.8 format and the new reference implementation of PSICQUIC are available online (https://psicquic.github.io/ and https://psicquic.github.io/MITAB28Format.html). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Livia Perfetto, Marcio Luis Acencio, Glyn Bradley, Gianni Cesareni, Noemi del-Toro, Dávid Fazekas, Henning Hermjakob, Tamás Korcsmáros, Martin Kuiper, Astrid Lægreid, Prisca Lo Surdo, Ruth C. Lovering, Sandra E. Orchard, Pablo Porras, Paul D. Thomas, Vasundra Touré, John Zobolas, Luana Licata
Bioinform.9
2015 Discovery of Drug Synergies in Gastric Cancer Cells Predicted by Logical Modeling
abstract
Discovery of efficient anti-cancer drug combinations is a major challenge, since experimental testing of all possible combinations is clearly impossible. Recent efforts to computationally predict drug combination responses retain this experimental search space, as model definitions typically rely on extensive drug perturbation data. We developed a dynamical model representing a cell fate decision network in the AGS gastric cancer cell line, relying on background knowledge extracted from literature and databases. We defined a set of logical equations recapitulating AGS data observed in cells in their baseline proliferative state. Using the modeling software GINsim, model reduction and simulation compression techniques were applied to cope with the vast state space of large logical models and enable simulations of pairwise applications of specific signaling inhibitory chemical substances. Our simulations predicted synergistic growth inhibitory action of five combinations from a total of 21 possible pairs. Four of the predicted synergies were confirmed in AGS cell growth real-time assays, including known effects of combined MEK-AKT or MEK-PI3K inhibitions, along with novel synergistic effects of combined TAK1-AKT or TAK1-PI3K inhibitions. Our strategy reduces the dependence on a priori drug perturbation experimentation for well-characterized signaling networks, by demonstrating that a model predictive of combinatorial drug effects can be inferred from background knowledge on unperturbed and proliferating cancer cells. Our modeling approach can thus contribute to preclinical discovery of efficient anticancer drug combinations, and thereby to development of strategies to tailor treatment to individual cancer patients.
Åsmund Flobak, Anaïs Baudot, Elisabeth Remy, Liv Thommesen, Denis Thieffry, Martin Kuiper, Astrid Lægreid
PLoS Comput. Biol.6
2014 orthAgogue: an agile tool for the rapid prediction of orthology relations
abstract
MOTIVATION: The comparison of genes and gene products across species depends on high-quality tools to determine the relationships between gene or protein sequences from various species. Although some excellent applications are available and widely used, their performance leaves room for improvement. RESULTS: We developed orthAgogue: a multithreaded C application for high-speed estimation of homology relations in massive datasets, operated via a flexible and easy command-line interface. AVAILABILITY: The orthAgogue software is distributed under the GNU license. The source code and binaries compiled for Linux are available at https://code.google.com/p/orthagogue/.
Ole Kristian Ekseth, Martin Kuiper, Vladimir Mironov
Bioinform.2
2014 Finding gene regulatory network candidates using the gene expression knowledge base
abstract
BACKGROUND: Network-based approaches for the analysis of large-scale genomics data have become well established. Biological networks provide a knowledge scaffold against which the patterns and dynamics of 'omics' data can be interpreted. The background information required for the construction of such networks is often dispersed across a multitude of knowledge bases in a variety of formats. The seamless integration of this information is one of the main challenges in bioinformatics. The Semantic Web offers powerful technologies for the assembly of integrated knowledge bases that are computationally comprehensible, thereby providing a potentially powerful resource for constructing biological networks and network-based analysis. RESULTS: We have developed the Gene eXpression Knowledge Base (GeXKB), a semantic web technology based resource that contains integrated knowledge about gene expression regulation. To affirm the utility of GeXKB we demonstrate how this resource can be exploited for the identification of candidate regulatory network proteins. We present four use cases that were designed from a biological perspective in order to find candidate members relevant for the gastrin hormone signaling network model. We show how a combination of specific query definitions and additional selection criteria derived from gene expression data and prior knowledge concerning candidate proteins can be used to retrieve a set of proteins that constitute valid candidates for regulatory network extensions. CONCLUSIONS: Semantic web technologies provide the means for processing and integrating various heterogeneous information sources. The GeXKB offers biologists such an integrated knowledge resource, allowing them to address complex biological questions pertaining to gene expression. This work illustrates how GeXKB can be used in combination with gene expression results and literature information to identify new potential candidates that may be considered for extending a gene regulatory network.
Aravind Venkatesan, Sushil Tripathi, Alejandro Sanz de Galdeano, Ward Blondé, Astrid Lægreid, Vladimir Mironov, Martin Kuiper
BMC Bioinform.7
2013 TFcheckpoint: a curated compendium of specific DNA-binding RNA polymerase II transcription factors
abstract
SUMMARY: Gene regulatory network assembly and analysis requires high-quality knowledge sources that cover functional aspects of the various components of the gene regulatory machinery. A multiplicity of resources exists with information about mammalian transcription factors (TFs); yet, only few of these provide sufficiently accurate classifications of the functional roles of individual TFs, or standardized evidence that would justify the information on which these functional classifications are based. We compiled the list of all putative TFs from nine different resources, ignored factors such as general TFs, mediator complexes and chromatin modifiers, and for the remaining factors checked the available literature for references that support their function as a true sequence-specific DNA-binding RNA polymerase II TF (DbTF). The results are available in the TFcheckpoint database, an exhaustive collection of TFs annotated according to experimental and other evidence on their function as true DbTFs. TFcheckpoint.org provides a high-quality and comprehensive knowledge source for genome-scale regulatory network studies. AVAILABILITY: The TFcheckpoint database is freely available at www.tfcheckpoint.org
Konika Chawla, Sushil Tripathi, Liv Thommesen, Astrid Lægreid, Martin Kuiper
Bioinform.5
2012 Gauging triple stores with actual biological data
abstract
Abstract Background Semantic Web technologies have been developed to overcome the limitations of the current Web and conventional data integration solutions. The Semantic Web is expected to link all the data present on the Internet instead of linking just documents. One of the foundations of the Semantic Web technologies is the knowledge representation language Resource Description Framework (RDF). Knowledge expressed in RDF is typically stored in so-called triple stores (also known as RDF stores), from which it can be retrieved with SPARQL, a language designed for querying RDF-based models. The Semantic Web technologies should allow federated queries over multiple triple stores. In this paper we compare the efficiency of a set of biologically relevant queries as applied to a number of different triple store implementations. Results Previously we developed a library of queries to guide the use of our knowledge base Cell Cycle Ontology implemented as a triple store. We have now compared the performance of these queries on five non-commercial triple stores: OpenLink Virtuoso (Open-Source Edition), Jena SDB, Jena TDB, SwiftOWLIM and 4Store. We examined three performance aspects: the data uploading time, the query execution time and the scalability. The queries we had chosen addressed diverse ontological or biological questions, and we found that individual store performance was quite query-specific. We identified three groups of queries displaying similar behaviour across the different stores: 1) relatively short response time queries, 2) moderate response time queries and 3) relatively long response time queries. SwiftOWLIM proved to be a winner in the first group, 4Store in the second one and Virtuoso in the third one. Conclusions Our analysis showed that some queries behaved idiosyncratically, in a triple store specific manner, mainly with SwiftOWLIM and 4Store. Virtuoso, as expected, displayed a very balanced performance - its load time and its response time for all the tested queries were better than average among the selected stores; it showed a very good scalability and a reasonable run-to-run reproducibility. Jena SDB and Jena TDB were consistently slower than the other three implementations. Our analysis demonstrated that most queries developed for Virtuoso could be successfully used for other implementations.
Vladimir Mironov, Nirmala Seethappan, Ward Blondé, Erick Antezana, Andrea Splendiani, Martin Kuiper
BMC Bioinform.6
2012 OLSVis: an animated, interactive visual browser for bio-ontologies
abstract
BACKGROUND: More than one million terms from biomedical ontologies and controlled vocabularies are available through the Ontology Lookup Service (OLS). Although OLS provides ample possibility for querying and browsing terms, the visualization of parts of the ontology graphs is rather limited and inflexible. RESULTS: We created the OLSVis web application, a visualiser for browsing all ontologies available in the OLS database. OLSVis shows customisable subgraphs of the OLS ontologies. Subgraphs are animated via a real-time force-based layout algorithm which is fully interactive: each time the user makes a change, e.g. browsing to a new term, hiding, adding, or dragging terms, the algorithm performs smooth and only essential reorganisations of the graph. This assures an optimal viewing experience, because subsequent screen layouts are not grossly altered, and users can easily navigate through the graph. URL: http://ols.wordvis.com CONCLUSIONS: The OLSVis web application provides a user-friendly tool to visualise ontologies from the OLS repository. It broadens the possibilities to investigate and select ontology subgraphs through a smooth visualisation method.
Steven Vercruysse, Aravind Venkatesan, Martin Kuiper
BMC Bioinform.3
2011 Reasoning with bio-ontologies: using relational closure rules to enable practical querying
abstract
MOTIVATION: Ontologies have become indispensable in the Life Sciences for managing large amounts of knowledge. The use of logics in ontologies ranges from sound modelling to practical querying of that knowledge, thus adding a considerable value. We conceive reasoning on bio-ontologies as a semi-automated process in three steps: (i) defining a logic-based representation language; (ii) building a consistent ontology using that language; and (iii) exploiting the ontology through querying. RESULTS: Here, we report on how we have implemented this approach to reasoning on the OBO Foundry ontologies within BioGateway, a biological Resource Description Framework knowledge base. By separating the three steps in a manual curation effort on Metarel, a vocabulary that specifies relation semantics, we were able to apply reasoning on a large scale. Starting from an initial 401 million triples, we inferred about 158 million knowledge statements that allow for a myriad of prospective queries, potentially leading to new hypotheses about for instance gene products, processes, interactions or diseases. AVAILABILITY: SPARUL code, a query end point and curated relation types in OBO Format, RDF and OWL 2 DL are freely available at http://www.semantic-systems-biology.org/metarel.
Ward Blondé, Vladimir Mironov, Aravind Venkatesan, Erick Antezana, Bernard De Baets, Martin Kuiper
Bioinform.6
2010 DASS-GUI: a user interface for identification and analysis of significant patterns in non-sequential data
abstract
SUMMARY: Many large 'omics' datasets have been published and many more are expected in the near future. New analysis methods are needed for best exploitation. We have developed a graphical user interface (GUI) for easy data analysis. Our discovery of all significant substructures (DASS) approach elucidates the underlying modularity, a typical feature of complex biological data. It is related to biclustering and other data mining approaches. Importantly, DASS-GUI also allows handling of multi-sets and calculation of statistical significances. DASS-GUI contains tools for further analysis of the identified patterns: analysis of the pattern hierarchy, enrichment analysis, module validation, analysis of additional numerical data, easy handling of synonymous names, clustering, filtering and merging. Different export options allow easy usage of additional tools such as Cytoscape. AVAILABILITY: Source code, pre-compiled binaries for different systems, a comprehensive tutorial, case studies and many additional datasets are freely available at http://www.ifr.ac.uk/dass/gui/. DASS-GUI is implemented in Qt.
Jens Hollunder, Maik Friedel, Martin Kuiper, Thomas Wilhelm 0003
Bioinform.3
2010 ONTO-ToolKit: enabling bio-ontology engineering via Galaxy
abstract
BACKGROUND: The biosciences increasingly face the challenge of integrating a wide variety of available data, information and knowledge in order to gain an understanding of biological systems. Data integration is supported by a diverse series of tools, but the lack of a consistent terminology to label these data still presents significant hurdles. As a consequence, much of the available biological data remains disconnected or worse: becomes misconnected. The need to address this terminology problem has spawned the building of a large number of bio-ontologies. OBOF, RDF and OWL are among the most used ontology formats to capture terms and relationships in the Life Sciences, opening the potential to use the Semantic Web to support data integration and further exploitation of integrated resources via automated retrieval and reasoning procedures. METHODS: We extended the Perl suite ONTO-PERL and functionally integrated it into the Galaxy platform. The resulting ONTO-ToolKit supports the analysis and handling of OBO-formatted ontologies via the Galaxy interface, and we demonstrated its functionality in different use cases that illustrate the flexibility to obtain sets of ontology terms that match specific search criteria. RESULTS: ONTO-ToolKit is available as a tool suite for Galaxy. Galaxy not only provides a user friendly interface allowing the interested biologist to manipulate OBO ontologies, it also opens up the possibility to perform further biological (and ontological) analyses by using other tools available within the Galaxy environment. Moreover, it provides tools to translate OBO-formatted ontologies into Semantic Web formats such as RDF and OWL. CONCLUSIONS: ONTO-ToolKit reaches out to researchers in the biosciences, by providing a user-friendly way to analyse and manipulate ontologies. This type of functionality will become increasingly important given the wealth of information that is becoming available based on ontologies.
Erick Antezana, Aravind Venkatesan, Chris Mungall, Vladimir Mironov, Martin Kuiper
BMC Bioinform.5
2010 Flexible network reconstruction from relational databases with Cytoscape and CytoSQL
abstract
BACKGROUND: Molecular interaction networks can be efficiently studied using network visualization software such as Cytoscape. The relevant nodes, edges and their attributes can be imported in Cytoscape in various file formats, or directly from external databases through specialized third party plugins. However, molecular data are often stored in relational databases with their own specific structure, for which dedicated plugins do not exist. Therefore, a more generic solution is presented. RESULTS: A new Cytoscape plugin 'CytoSQL' is developed to connect Cytoscape to any relational database. It allows to launch SQL ('Structured Query Language') queries from within Cytoscape, with the option to inject node or edge features of an existing network as SQL arguments, and to convert the retrieved data to Cytoscape network components. Supported by a set of case studies we demonstrate the flexibility and the power of the CytoSQL plugin in converting specific data subsets into meaningful network representations. CONCLUSIONS: CytoSQL offers a unified approach to let Cytoscape interact with relational databases. Thanks to the power of the SQL syntax, this tool can rapidly generate and enrich networks according to very complex criteria. The plugin is available at http://www.ptools.ua.ac.be/CytoSQL.
Kris Laukens, Jens Hollunder, Geert De Jaeger, Martin Kuiper, Erwin Witters, Alain Verschoren, Koenraad Van Leemput
BMC Bioinform.5
2009 Biological knowledge management: the emerging role of the Semantic Web technologies
abstract
New knowledge is produced at a continuously increasing speed, and the list of papers, databases and other knowledge sources that a researcher in the life sciences needs to cope with is actually turning into a problem rather than an asset. The adequate management of knowledge is therefore becoming fundamentally important for life scientists, especially if they work with approaches that thoroughly depend on knowledge integration, such as systems biology. Several initiatives to organize biological knowledge sources into a readily exploitable resourceome are presently being carried out. Ontologies and Semantic Web technologies revolutionize these efforts. Here, we review the benefits, trends, current possibilities, and the potential this holds for the biosciences.
Erick Antezana, Martin Kuiper, Vladimir Mironov
Briefings Bioinform.2
2009 Gene expression trends and protein features effectively complement each other in gene function prediction
abstract
MOTIVATION: Genome-scale 'omics' data constitute a potentially rich source of information about biological systems and their function. There is a plethora of tools and methods available to mine omics data. However, the diversity and complexity of different omics data types is a stumbling block for multi-data integration, hence there is a dire need for additional methods to exploit potential synergy from integrated orthogonal data. Rough Sets provide an efficient means to use complex information in classification approaches. Here, we set out to explore the possibilities of Rough Sets to incorporate diverse information sources in a functional classification of unknown genes. RESULTS: We explored the use of Rough Sets for a novel data integration strategy where gene expression data, protein features and Gene Ontology (GO) annotations were combined to describe general and biologically relevant patterns represented by If-Then rules. The descriptive rules were used to predict the function of unknown genes in Arabidopsis thaliana and Schizosaccharomyces pombe. The If-Then rule models showed success rates of up to 0.89 (discriminative and predictive power for both modeled organisms); whereas, models built solely of one data type (protein features or gene expression data) yielded success rates varying from 0.68 to 0.78. Our models were applied to generate classifications for many unknown genes, of which a sizeable number were confirmed either by PubMed literature reports or electronically interfered annotations. Finally, we studied cell cycle protein-protein interactions derived from both tandem affinity purification experiments and in silico experiments in the BioGRID interactome database and found strong experimental evidence for the predictions generated by our models. The results show that our approach can be used to build very robust models that create synergy from integrating gene expression data and protein features. AVAILABILITY: The Rough Set-based method is implemented in the Rosetta toolkit kernel version 1.0.1 available at: http://rosetta.lcb.uu.se/
Krzysztof Wabnik, Torgeir R. Hvidsten, Anna M. Kedzierska, Jelle Van Leene, Geert De Jaeger, Gerrit T. S. Beemster, Jan Komorowski, Martin Kuiper
Bioinform.8
2009 BioGateway: a semantic systems biology tool for the life sciences
abstract
BACKGROUND: Life scientists need help in coping with the plethora of fast growing and scattered knowledge resources. Ideally, this knowledge should be integrated in a form that allows them to pose complex questions that address the properties of biological systems, independently from the origin of the knowledge. Semantic Web technologies prove to be well suited for knowledge integration, knowledge production (hypothesis formulation), knowledge querying and knowledge maintenance. RESULTS: We implemented a semantically integrated resource named BioGateway, comprising the entire set of the OBO foundry candidate ontologies, the GO annotation files, the SWISS-PROT protein set, the NCBI taxonomy and several in-house ontologies. BioGateway provides a single entry point to query these resources through SPARQL. It constitutes a key component for a Semantic Systems Biology approach to generate new hypotheses concerning systems properties. In the course of developing BioGateway, we faced challenges that are common to other projects that involve large datasets in diverse representations. We present a detailed analysis of the obstacles that had to be overcome in creating BioGateway. We demonstrate the potential of a comprehensive application of Semantic Web technologies to global biomedical data. CONCLUSION: The time is ripe for launching a community effort aimed at a wider acceptance and application of Semantic Web technologies in the life sciences. We call for the creation of a forum that strives to implement a truly semantic life science foundation for Semantic Systems Biology. Access to the system and supplementary information (such as a listing of the data sources in RDF, and sample queries) can be found at http://www.semantic-systems-biology.org/biogateway.
Erick Antezana, Ward Blondé, Mikel Egaña Aranguren, Alistair Rutherford, Robert Stevens 0001, Bernard De Baets, Vladimir Mironov, Martin Kuiper
BMC Bioinform.8
2008 Initialization Dependence of Clustering Algorithms
Wim De Mulder, Stefan Schliebs, René K. Boel, Martin Kuiper
ICONIP (2)4
2008 ONTO-PERL: An API for supporting the development and analysis of bio-ontologies
abstract
MOTIVATION: Many biomedical ontologies use OBO or OWL as knowledge representation language. The rapid increase of such ontologies calls for adequate tools to facilitate their use. In particular, there is a pressing need to programmatically deal with such ontologies in many applications, including data integration, text mining, as well as semantic applications supporting translational research. RESULTS: We present an Application Programming Interface (API) called ONTO-PERL. This API significantly extends the repertoire of available tools supporting the development and analysis of bio-ontologies. AVAILABILITY: The source code code as well as sample usage scripts can be found at: http://search.cpan.org/dist/ONTO-PERL/
Erick Antezana, Mikel Egaña Aranguren, Bernard De Baets, Martin Kuiper, Vladimir Mironov
Bioinform.4
2008 Ontology Design Patterns for bio-ontologies: a case study on the Cell Cycle Ontology
abstract
BACKGROUND: Bio-ontologies are key elements of knowledge management in bioinformatics. Rich and rigorous bio-ontologies should represent biological knowledge with high fidelity and robustness. The richness in bio-ontologies is a prior condition for diverse and efficient reasoning, and hence querying and hypothesis validation. Rigour allows a more consistent maintenance. Modelling such bio-ontologies is, however, a difficult task for bio-ontologists, because the necessary richness and rigour is difficult to achieve without extensive training. RESULTS: Analogous to design patterns in software engineering, Ontology Design Patterns are solutions to typical modelling problems that bio-ontologists can use when building bio-ontologies. They offer a means of creating rich and rigorous bio-ontologies with reduced effort. The concept of Ontology Design Patterns is described and documentation and application methodologies for Ontology Design Patterns are presented. Some real-world use cases of Ontology Design Patterns are provided and tested in the Cell Cycle Ontology. Ontology Design Patterns, including those tested in the Cell Cycle Ontology, can be explored in the Ontology Design Patterns public catalogue that has been created based on the documentation system presented (http://odps.sourceforge.net/). CONCLUSIONS: Ontology Design Patterns provide a method for rich and rigorous modelling in bio-ontologies. They also offer advantages at different development levels (such as design, implementation and communication) enabling, if used, a more modular, well-founded and richer representation of the biological knowledge. This representation will produce a more efficient knowledge management in the long term.
Mikel Egaña Aranguren, Erick Antezana, Martin Kuiper, Robert Stevens 0001
BMC Bioinform.3
2007 Validating module network learning algorithms using simulated data
abstract
BACKGROUND: In recent years, several authors have used probabilistic graphical models to learn expression modules and their regulatory programs from gene expression data. Despite the demonstrated success of such algorithms in uncovering biologically relevant regulatory relations, further developments in the area are hampered by a lack of tools to compare the performance of alternative module network learning strategies. Here, we demonstrate the use of the synthetic data generator SynTReN for the purpose of testing and comparing module network learning algorithms. We introduce a software package for learning module networks, called LeMoNe, which incorporates a novel strategy for learning regulatory programs. Novelties include the use of a bottom-up Bayesian hierarchical clustering to construct the regulatory programs, and the use of a conditional entropy measure to assign regulators to the regulation program nodes. Using SynTReN data, we test the performance of LeMoNe in a completely controlled situation and assess the effect of the methodological changes we made with respect to an existing software package, namely Genomica. Additionally, we assess the effect of various parameters, such as the size of the data set and the amount of noise, on the inference performance. RESULTS: Overall, application of Genomica and LeMoNe to simulated data sets gave comparable results. However, LeMoNe offers some advantages, one of them being that the learning process is considerably faster for larger data sets. Additionally, we show that the location of the regulators in the LeMoNe regulation programs and their conditional entropy may be used to prioritize regulators for functional validation, and that the combination of the bottom-up clustering strategy with the conditional entropy-based assignment of regulators improves the handling of missing or hidden regulators. CONCLUSION: We show that data simulators such as SynTReN are very well suited for the purpose of developing, testing and improving module network algorithms. We used SynTReN data to develop and test an alternative module network learning strategy, which is incorporated in the software package LeMoNe, and we provide evidence that this alternative strategy has several advantages with respect to existing methods.
Tom Michoel, Steven Maere, Eric Bonnet, Anagha Joshi, Yvan Saeys, Tim Van den Bulcke, Koenraad Van Leemput, Piet van Remortel, Martin Kuiper, Kathleen Marchal, Yves Van de Peer
BMC Bioinform.9
2007 CATMA, a comprehensive genome-scale resource for silencing and transcript profiling of Arabidopsis genes
abstract
BACKGROUND: The Complete Arabidopsis Transcript MicroArray (CATMA) initiative combines the efforts of laboratories in eight European countries 1 to deliver gene-specific sequence tags (GSTs) for the Arabidopsis research community. The CATMA initiative offers the power and flexibility to regularly update the GST collection according to evolving knowledge about the gene repertoire. These GST amplicons can easily be reamplified and shared, subsets can be picked at will to print dedicated arrays, and the GSTs can be cloned and used for other functional studies. This ongoing initiative has already produced approximately 24,000 GSTs that have been made publicly available for spotted microarray printing and RNA interference. RESULTS: GSTs from the CATMA version 2 repertoire (CATMAv2, created in 2002) were mapped onto the gene models from two independent Arabidopsis nuclear genome annotation efforts, TIGR5 and PSB-EuGène, to consolidate a list of genes that were targeted by previously designed CATMA tags. A total of 9,027 gene models were not tagged by any amplified CATMAv2 GST, and 2,533 amplified GSTs were no longer predicted to tag an updated gene model. To validate the efficacy of GST mapping criteria and design rules, the predicted and experimentally observed hybridization characteristics associated to GST features were correlated in transcript profiling datasets obtained with the CATMAv2 microarray, confirming the reliability of this platform. To complete the CATMA repertoire, all 9,027 gene models for which no GST had yet been designed were processed with an adjusted version of the Specific Primer and Amplicon Design Software (SPADS). A total of 5,756 novel GSTs were designed and amplified by PCR from genomic DNA. Together with the pre-existing GST collection, this new addition constitutes the CATMAv3 repertoire. It comprises 30,343 unique amplified sequences that tag 24,202 and 23,009 protein-encoding nuclear gene models in the TAIR6 and EuGène genome annotations, respectively. To cover the remaining untagged genes, we identified 543 additional GSTs using less stringent design criteria and designed 990 sequence tags matching multiple members of gene families (Gene Family Tags or GFTs) to cover any remaining untagged genes. These latter 1,533 features constitute the CATMAv4 addition. CONCLUSION: To update the CATMA GST repertoire, we designed 7,289 additional sequence tags, bringing the total number of tagged TAIR6-annotated Arabidopsis nuclear protein-coding genes to 26,173. This resource is used both for the production of spotted microarrays and the large-scale cloning of hairpin RNA silencing vectors. All information about the resulting updated CATMA repertoire is available through the CATMA database http://www.catma.org.
Gert Sclep, Joke Allemeersch, Robin Liechti, Björn De Meyer, Jim Beynon, Rishikesh Bhalerao, Yves Moreau, Wilfried Nietfeld, Jean-Pierre Renou, Philippe Reymond, Martin Kuiper, Pierre Hilson
BMC Bioinform.11
2005 BiNGO: a Cytoscape plugin to assess overrepresentation of Gene Ontology categories in Biological Networks
abstract
The Biological Networks Gene Ontology tool (BiNGO) is an open-source Java tool to determine which Gene Ontology (GO) terms are significantly overrepresented in a set of genes. BiNGO can be used either on a list of genes, pasted as text, or interactively on subgraphs of biological networks visualized in Cytoscape. BiNGO maps the predominant functional themes of the tested gene set on the GO hierarchy, and takes advantage of Cytoscape's versatile visualization environment to produce an intuitive and customizable visual representation of the results.
Steven Maere, Karel Heymans, Martin Kuiper
Bioinform.3
2005 Simulating genetic networks made easy: network construction with simple building blocks
abstract
UNLABELLED: We present SIM-plex, a genetic network simulator with a very intuitive interface in which a user can easily specify interactions as simple 'if-then' statements. The simulator is based on the mathematical model of Piecewise Linear Differential Equations (PLDEs). With PLDEs, genetic interactions are approximated as acting in a switch-like manner. AVAILABILITY: The Java program, examples and a tutorial are available at http://www.psb.ugent.be/cbd/ CONTACT: {stcru,makui}@psb.ugent.be
Steven Vercruysse, Martin Kuiper
Bioinform.2