Reinhard Schneider 0002

dblp:45/2531-2 · DBLP profile ↗
← Back
39ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0002-8278-1618ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 38 · 5 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 COBREXA 2: tidy and scalable construction of complex metabolic models
abstract
SUMMARY: Constraint-based metabolic models offer a scalable framework to investigate biological systems using optimality principles. Construction and simulation of detailed models that utilize multiple kinds of constraint systems pose a significant coding overhead, complicating implementation of new types of analyses. We present an improved version of the constraint-based metabolic modeling package COBREXA, which utilizes a hierarchical model construction framework that decouples the implemented analysis algorithms into independent, yet re-combinable, building blocks. By removing the need to re-implement modeling components, assembly of complex metabolic models is simplified, which we demonstrate on use-cases of resource-balanced models, and enzyme-constrained flux balance models of interacting bacterial communities. Notably, these models show improved predictive capabilities in both monoculture and community settings. In perspective, the re-usable model-building components in COBREXA 2 provide a sustainable way to handle increasingly complex models in constraint-based modeling. AVAILABILITY AND IMPLEMENTATION: COBREXA 2 is available from https://github.com/COBREXA/COBREXA.jl, and from Julia package repositories. COBREXA 2 works on all major operating systems and computer architectures. Documentation is available at https://cobrexa.github.io/COBREXA.jl/.
Miroslav Kratochvíl, St. Elmo Wilken, Oliver Ebenhöh, Reinhard Schneider 0002, Venkata P. Satagopam
Bioinform.4
2024 Graph databases in systems biology: a systematic review
abstract
Graph databases are becoming increasingly popular across scientific disciplines, being highly suitable for storing and connecting complex heterogeneous data. In systems biology, they are used as a backend solution for biological data repositories, ontologies, networks, pathways, and knowledge graph databases. In this review, we analyse all publications using or mentioning graph databases retrieved from PubMed and PubMed Central full-text search, focusing on the top 16 available graph databases, Publications are categorized according to their domain and application, focusing on pathway and network biology and relevant ontologies and tools. We detail different approaches and highlight the advantages of outstanding resources, such as UniProtKB, Disease Ontology, and Reactome, which provide graph-based solutions. We discuss ongoing efforts of the systems biology community to standardize and harmonize knowledge graph creation and the maintenance of integrated resources. Outlining prospects, including the use of graph databases as a way of communication between biological data repositories, we conclude that efficient design, querying, and maintenance of graph databases will be key for knowledge generation in systems biology and other research fields with heterogeneous data.
Ilya Mazein, Adrien Rougny, Alexander Mazein, Ron Henkel, Lea Gütebier, Lea Michaelis, Marek Ostaszewski, Reinhard Schneider 0002, Venkata P. Satagopam, Lars Juhl Jensen, Dagmar Waltemath, Judith A. H. Wodke, Irina Balaur
Briefings Bioinform.8
2022 COBREXA.jl: constraint-based reconstruction and exascale analysis
abstract
SUMMARY: COBREXA.jl is a Julia package for scalable, high-performance constraint-based reconstruction and analysis of very large-scale biological models. Its primary purpose is to facilitate the integration of modern high performance computing environments with the processing and analysis of large-scale metabolic models of challenging complexity. We report the architecture of the package, and demonstrate how the design promotes analysis scalability on several use-cases with multi-organism community models. AVAILABILITY AND IMPLEMENTATION: https://doi.org/10.17881/ZKCR-BT30. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Miroslav Kratochvíl, Laurent Heirendt, St. Elmo Wilken, Taneli Pusa, Sylvain Arreckx, Alberto Noronha, Marvin van Aalst, Venkata P. Satagopam, Oliver Ebenhöh, Reinhard Schneider 0002, Christophe Trefois
Bioinform.10
2021 Closing the gap between formats for storing layout information in systems biology
abstract
The first version of this article listed one of its authors as Jan Hausenauer rather than Jan Hasenauer. This has now been corrected. The authors regret the error.
David Hoksza, Piotr Gawron, Marek Ostaszewski, Jan Hasenauer, Reinhard Schneider 0002
Briefings Bioinform.5
2021 Reusability and composability in process description maps: RAS-RAF-MEK-ERK signalling
abstract
Detailed maps of the molecular basis of the disease are powerful tools for interpreting data and building predictive models. Modularity and composability are considered necessary network features for large-scale collaborative efforts to build comprehensive molecular descriptions of disease mechanisms. An effective way to create and manage large systems is to compose multiple subsystems. Composable network components could effectively harness the contributions of many individuals and enable teams to seamlessly assemble many individual components into comprehensive maps. We examine manually built versions of the RAS-RAF-MEK-ERK cascade from the Atlas of Cancer Signalling Network, PANTHER and Reactome databases and review them in terms of their reusability and composability for assembling new disease models. We identify design principles for managing complex systems that could make it easier for investigators to share and reuse network components. We demonstrate the main challenges including incompatible levels of detail and ambiguous representation of complexes and highlight the need to address these challenges.
Alexander Mazein, Adrien Rougny, Jonathan R. Karr, Julio Saez-Rodriguez, Marek Ostaszewski, Reinhard Schneider 0002
Briefings Bioinform.6
2020 Closing the gap between formats for storing layout information in systems biology
abstract
The understanding of complex biological networks often relies on both a dedicated layout and a topology. Currently, there are three major competing layout-aware systems biology formats, but there are no software tools or software libraries supporting all of them. This complicates the management of molecular network layouts and hinders their reuse and extension. In this paper, we present a high-level overview of the layout formats in systems biology, focusing on their commonalities and differences, review their support in existing software tools, libraries and repositories and finally introduce a new conversion module within the MINERVA platform. The module is available via a REST API and offers, besides the ability to convert between layout-aware systems biology formats, the possibility to export layouts into several graphical formats. The module enables conversion of very large networks with thousands of elements, such as disease maps or metabolic reconstructions, rendering it widely applicable in systems biology.
David Hoksza, Piotr Gawron, Marek Ostaszewski, Jan Hasenauer, Reinhard Schneider 0002
Briefings Bioinform.5
2020 LAITOR4HPC: A text mining pipeline based on HPC for building interaction networks
abstract
BACKGROUND: The amount of published full-text articles has increased dramatically. Text mining tools configure an essential approach to building biological networks, updating databases and providing annotation for new pathways. PESCADOR is an online web server based on LAITOR and NLProt text mining tools, which retrieves protein-protein co-occurrences in a tabular-based format, adding a network schema. Here we present an HPC-oriented version of PESCADOR's native text mining tool, renamed to LAITOR4HPC, aiming to access an unlimited abstract amount in a short time to enrich available networks, build new ones and possibly highlight whether fields of research have been exhaustively studied. RESULTS: By taking advantage of parallel computing HPC infrastructure, the full collection of MEDLINE abstracts available until June 2017 was analyzed in a shorter period (6 days) when compared to the original online implementation (with an estimated 2 years to run the same data). Additionally, three case studies were presented to illustrate LAITOR4HPC usage possibilities. The first case study targeted soybean and was used to retrieve an overview of published co-occurrences in a single organism, retrieving 15,788 proteins in 7894 co-occurrences. In the second case study, a target gene family was searched in many organisms, by analyzing 15 species under biotic stress. Most co-occurrences regarded Arabidopsis thaliana and Zea mays. The third case study concerned the construction and enrichment of an available pathway. Choosing A. thaliana for further analysis, the defensin pathway was enriched, showing additional signaling and regulation molecules, and how they respond to each other in the modulation of this complex plant defense response. CONCLUSIONS: LAITOR4HPC can be used for an efficient text mining based construction of biological networks derived from big data sources, such as MEDLINE abstracts. Time consumption and data input limitations will depend on the available resources at the HPC facility. LAITOR4HPC enables enough flexibility for different approaches and data amounts targeted to an organism, a subject, or a specific pathway. Additionally, it can deliver comprehensive results where interactions are classified into four types, according to their reliability.
Bruna Piereck, Marx Oliveira-Lima, Ana Maria Benko-Iseppon, Sarah Diehl, Reinhard Schneider 0002, Ana Christina Brasileiro-Vidal, Adriano Barbosa-Silva
BMC Bioinform.5
2019 Community-driven roadmap for integrated disease maps
abstract
The Disease Maps Project builds on a network of scientific and clinical groups that exchange best practices, share information and develop systems biomedicine tools. The project aims for an integrated, highly curated and user-friendly platform for disease-related knowledge. The primary focus of disease maps is on interconnected signaling, metabolic and gene regulatory network pathways represented in standard formats. The involvement of domain experts ensures that the key disease hallmarks are covered and relevant, up-to-date knowledge is adequately represented. Expert-curated and computer readable, disease maps may serve as a compendium of knowledge, allow for data-supported hypothesis generation or serve as a scaffold for the generation of predictive mathematical models. This article summarizes the 2nd Disease Maps Community meeting, highlighting its important topics and outcomes. We outline milestones on the roadmap for the future development of disease maps, including creating and maintaining standardized disease maps; sharing parts of maps that encode common human disease mechanisms; providing technical solutions for complexity management of maps; and Web tools for in-depth exploration of such maps. A dedicated discussion was focused on mathematical modeling approaches, as one of the main goals of disease map development is the generation of mathematically interpretable representations to predict disease comorbidity or drug response and to suggest drug repositioning, altogether supporting clinical decisions.
Marek Ostaszewski, Stephan Gebel, Inna Kuperstein, Alexander Mazein, Andrei Yu. Zinovyev, Ugur Dogrusoz, Jan Hasenauer, Ronan M. T. Fleming, Nicolas Le Novère, Piotr Gawron, Thomas S. Ligon, Anna Niarakis, David P. Nickerson, Daniel Weindl, Rudi Balling, Emmanuel Barillot, Charles Auffray, Reinhard Schneider 0002
Briefings Bioinform.18
2019 MINERVA API and plugins: opening molecular network analysis and visualization to the community
abstract
SUMMARY: The complexity of molecular networks makes them difficult to navigate and interpret, creating a need for specialized software. MINERVA is a web platform for visualization, exploration and management of molecular networks. Here, we introduce an extension to MINERVA architecture that greatly facilitates the access and use of the stored molecular network data. It allows to incorporate such data in analytical pipelines via a programmatic access interface, and to extend the platform's visual exploration and analytics functionality via plugin architecture. This is possible for any molecular network hosted by the MINERVA platform encoded in well-recognized systems biology formats. To showcase the possibilities of the plugin architecture, we have developed several plugins extending the MINERVA core functionalities. In the article, we demonstrate the plugins for interactive tree traversal of molecular networks, for enrichment analysis and for mapping and visualization of known disease variants or known adverse drug reactions to molecules in the network. AVAILABILITY AND IMPLEMENTATION: Plugins developed and maintained by the MINERVA team are available under the AGPL v3 license at https://git-r3lab.uni.lu/minerva/plugins/. The MINERVA API and plugin documentation is available at https://minerva-web.lcsb.uni.lu.
David Hoksza, Piotr Gawron, Marek Ostaszewski, Ewa Smula, Reinhard Schneider 0002
Bioinform.5
2019 Data and knowledge management in translational research: implementation of the eTRIKS platform for the IMI OncoTrack consortium
abstract
BACKGROUND: For large international research consortia, such as those funded by the European Union's Horizon 2020 programme or the Innovative Medicines Initiative, good data coordination practices and tools are essential for the successful collection, organization and analysis of the resulting data. Research consortia are attempting ever more ambitious science to better understand disease, by leveraging technologies such as whole genome sequencing, proteomics, patient-derived biological models and computer-based systems biology simulations. RESULTS: The IMI eTRIKS consortium is charged with the task of developing an integrated knowledge management platform capable of supporting the complexity of the data generated by such research programmes. In this paper, using the example of the OncoTrack consortium, we describe a typical use case in translational medicine. The tranSMART knowledge management platform was implemented to support data from observational clinical cohorts, drug response data from cell culture models and drug response data from mouse xenograft tumour models. The high dimensional (omics) data from the molecular analyses of the corresponding biological materials were linked to these collections, so that users could browse and analyse these to derive candidate biomarkers. CONCLUSIONS: In all these steps, data mapping, linking and preparation are handled automatically by the tranSMART integration platform. Therefore, researchers without specialist data handling skills can focus directly on the scientific questions, without spending undue effort on processing the data and data integration, which are otherwise a burden and the most time-consuming part of translational research data analysis.
Reha Yildirimman, Emmanuel Van der Stuyft, Denny Verbeeck, Sascha Herzinger, Venkata P. Satagopam, Adriano Barbosa-Silva, Reinhard Schneider 0002, Bodo M. H. Lange, Hans Lehrach, Yike Guo, David Henderson, Anthony Rowe 0002
BMC Bioinform.8
2018 MolArt: a molecular structure annotation and visualization tool
abstract
Summary: MolArt fills the gap between sequence and structure visualization by providing a light-weight, interactive environment enabling exploration of sequence annotations in the context of available experimental or predicted protein structures. Provided a UniProt ID, MolArt downloads and displays sequence annotations, sequence-structure mapping and relevant structures. The sequence and structure views are interlinked, enabling sequence annotations being color overlaid over the mapped structures, thus providing an enhanced understanding and interpretation of the available molecular data. Availability and implementation: MolArt is released under the Apache 2 license and is available at https://github.com/davidhoksza/MolArt. The project web page https://davidhoksza.github.io/MolArt/ features examples and applications of the tool.
David Hoksza, Piotr Gawron, Marek Ostaszewski, Reinhard Schneider 0002
Bioinform.4
2018 Clustering approaches for visual knowledge exploration in molecular interaction networks
abstract
BACKGROUND: Biomedical knowledge grows in complexity, and becomes encoded in network-based repositories, which include focused, expert-drawn diagrams, networks of evidence-based associations and established ontologies. Combining these structured information sources is an important computational challenge, as large graphs are difficult to analyze visually. RESULTS: We investigate knowledge discovery in manually curated and annotated molecular interaction diagrams. To evaluate similarity of content we use: i) Euclidean distance in expert-drawn diagrams, ii) shortest path distance using the underlying network and iii) ontology-based distance. We employ clustering with these metrics used separately and in pairwise combinations. We propose a novel bi-level optimization approach together with an evolutionary algorithm for informative combination of distance metrics. We compare the enrichment of the obtained clusters between the solutions and with expert knowledge. We calculate the number of Gene and Disease Ontology terms discovered by different solutions as a measure of cluster quality. Our results show that combining distance metrics can improve clustering accuracy, based on the comparison with expert-provided clusters. Also, the performance of specific combinations of distance functions depends on the clustering depth (number of clusters). By employing bi-level optimization approach we evaluated relative importance of distance functions and we found that indeed the order by which they are combined affects clustering performance. Next, with the enrichment analysis of clustering results we found that both hierarchical and bi-level clustering schemes discovered more Gene and Disease Ontology terms than expert-provided clusters for the same knowledge repository. Moreover, bi-level clustering found more enriched terms than the best hierarchical clustering solution for three distinct distance metric combinations in three different instances of disease maps. CONCLUSIONS: In this work we examined the impact of different distance functions on clustering of a visual biomedical knowledge repository. We found that combining distance functions may be beneficial for clustering, and improve exploration of such repositories. We proposed bi-level optimization to evaluate the importance of order by which the distance functions are combined. Both combination and order of these functions affected clustering quality and knowledge recognition in the considered benchmarks. We propose that multiple dimensions can be utilized simultaneously for visual knowledge exploration.
Marek Ostaszewski, Emmanuel Kieffer, Grégoire Danoy, Reinhard Schneider 0002, Pascal Bouvry
BMC Bioinform.4
2017 SmartR: an open-source platform for interactive visual analytics for translational research data
abstract
SUMMARY: In translational research, efficient knowledge exchange between the different fields of expertise is crucial. An open platform that is capable of storing a multitude of data types such as clinical, pre-clinical or OMICS data combined with strong visual analytical capabilities will significantly accelerate the scientific progress by making data more accessible and hypothesis generation easier. The open data warehouse tranSMART is capable of storing a variety of data types and has a growing user community including both academic institutions and pharmaceutical companies. tranSMART, however, currently lacks interactive and dynamic visual analytics and does not permit any post-processing interaction or exploration. For this reason, we developed SmartR , a plugin for tranSMART, that equips the platform not only with several dynamic visual analytical workflows, but also provides its own framework for the addition of new custom workflows. Modern web technologies such as D3.js or AngularJS were used to build a set of standard visualizations that were heavily improved with dynamic elements. AVAILABILITY AND IMPLEMENTATION: The source code is licensed under the Apache 2.0 License and is freely available on GitHub: https://github.com/transmart/SmartR . CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sascha Herzinger, Venkata P. Satagopam, Serge Eifes, Kavita Rege, Adriano Barbosa-Silva, Reinhard Schneider 0002
Bioinform.7
2017 ReconMap: an interactive visualization of human metabolism
abstract
Motivation: A genome-scale reconstruction of human metabolism, Recon 2, is available but no interface exists to interactively visualize its content integrated with omics data and simulation results. Results: We manually drew a comprehensive map, ReconMap 2.0, that is consistent with the content of Recon 2. We present it within a web interface that allows content query, visualization of custom datasets and submission of feedback to manual curators. Availability and Implementation: ReconMap can be accessed via http://vmh.uni.lu , with network export in a Systems Biology Graphical Notation compliant format released under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. A Constraint-Based Reconstruction and Analysis (COBRA) Toolbox extension to interact with ReconMap is available via https://github.com/opencobra/cobratoolbox . Contact: [email protected].
Alberto Noronha, Anna Dröfn Daníelsdóttir, Piotr Gawron, Freyr Jóhannsson, Soffía Jónsdóttir, Sindri Jarlsson, Jón Pétur Gunnarsson, Sigurður Brynjólfsson, Reinhard Schneider 0002, Ines Thiele, Ronan M. T. Fleming
Bioinform.9
2015 A Novel Multi-objectivisation Approach for Optimising the Protein Inverse Folding Problem
Sune S. Nielsen, Grégoire Danoy, Wiktor Jurkowski, Juan Luis Jiménez Laredo, Reinhard Schneider 0002, El-Ghazali Talbi, Pascal Bouvry
EvoApplications5
2015 RepExplore: addressing technical replicate variance in proteomics and metabolomics data analysis
abstract
UNLABELLED: High-throughput omics datasets often contain technical replicates included to account for technical sources of noise in the measurement process. Although summarizing these replicate measurements by using robust averages may help to reduce the influence of noise on downstream data analysis, the information on the variance across the replicate measurements is lost in the averaging process and therefore typically disregarded in subsequent statistical analyses.We introduce RepExplore, a web-service dedicated to exploit the information captured in the technical replicate variance to provide more reliable and informative differential expression and abundance statistics for omics datasets. The software builds on previously published statistical methods, which have been applied successfully to biomedical omics data but are difficult to use without prior experience in programming or scripting. RepExplore facilitates the analysis by providing a fully automated data processing and interactive ranking tables, whisker plot, heat map and principal component analysis visualizations to interpret omics data and derived statistics. AVAILABILITY AND IMPLEMENTATION: Freely available at http://www.repexplore.tk CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Enrico Glaab, Reinhard Schneider 0002
Bioinform.2
2015 BioTextQuest+: a knowledge integration platform for literature mining and concept discovery
abstract
Bioinformatics (2014); 30(22), 3249–3256 doi: 10.1093/bioinformatics/btu524 The above article contained an incorrect email address for the corresponding author, Ioannis Iliopoulos, the correct email address is [email protected]
Nikolas Papanikolaou, Georgios A. Pavlopoulos, Evangelos Pafilis, Theodosios Theodosiou, Reinhard Schneider 0002, Venkata P. Satagopam, Christos A. Ouzounis, Aristides G. Eliopoulos, Vasilis J. Promponas
Bioinform.5
2014 Visualizing time-related data in biology, a review
abstract
Time is of the essence in biology as in so much else. For example, monitoring disease progression or the timing of developmental defects is important for the processes of drug discovery and therapy trials. Furthermore, an understanding of the basic dynamics of biological phenomena that are often strictly time regulated (e.g. circadian rhythms) is needed to make accurate inferences about the evolution of biological processes. Recent advances in technologies have enabled us to measure timing effects more accurately and in more detail. This has driven related advances in visualization and analysis tools that try to effectively exploit this data. Beyond timeline plots, notable attempts at more involved temporal interpretation have been made in recent years, but awareness of the available resources is still limited within the scientific community. Here, we review some advances in biological visualization of time-driven processes and consider how they aid data analysis and interpretation.
Maria Secrier, Reinhard Schneider 0002
Briefings Bioinform.2
2014 BioTextQuest+: a knowledge integration platform for literature mining and concept discovery
abstract
SUMMARY: The iterative process of finding relevant information in biomedical literature and performing bioinformatics analyses might result in an endless loop for an inexperienced user, considering the exponential growth of scientific corpora and the plethora of tools designed to mine PubMed(®) and related biological databases. Herein, we describe BioTextQuest(+), a web-based interactive knowledge exploration platform with significant advances to its predecessor (BioTextQuest), aiming to bridge processes such as bioentity recognition, functional annotation, document clustering and data integration towards literature mining and concept discovery. BioTextQuest(+) enables PubMed and OMIM querying, retrieval of abstracts related to a targeted request and optimal detection of genes, proteins, molecular functions, pathways and biological processes within the retrieved documents. The front-end interface facilitates the browsing of document clustering per subject, the analysis of term co-occurrence, the generation of tag clouds containing highly represented terms per cluster and at-a-glance popup windows with information about relevant genes and proteins. Moreover, to support experimental research, BioTextQuest(+) addresses integration of its primary functionality with biological repositories and software tools able to deliver further bioinformatics services. The Google-like interface extends beyond simple use by offering a range of advanced parameterization for expert users. We demonstrate the functionality of BioTextQuest(+) through several exemplary research scenarios including author disambiguation, functional term enrichment, knowledge acquisition and concept discovery linking major human diseases, such as obesity and ageing. AVAILABILITY: The service is accessible at http://bioinformatics.med.uoc.gr/biotextquest. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Nikolas Papanikolaou, Georgios A. Pavlopoulos, Evangelos Pafilis, Theodosios Theodosiou, Reinhard Schneider 0002, Venkata P. Satagopam, Christos A. Ouzounis, Aristides G. Eliopoulos, Vasilis J. Promponas
Bioinform.5
2014 The HIV Mutation Browser: A Resource for Human Immunodeficiency Virus Mutagenesis and Polymorphism Data
abstract
Huge research effort has been invested over many years to determine the phenotypes of natural or artificial mutations in HIV proteins--interpretation of mutation phenotypes is an invaluable source of new knowledge. The results of this research effort are recorded in the scientific literature, but it is difficult for virologists to rapidly find it. Manually locating data on phenotypic variation within the approximately 270,000 available HIV-related research articles, or the further 1,500 articles that are published each month is a daunting task. Accordingly, the HIV research community would benefit from a resource cataloguing the available HIV mutation literature. We have applied computational text-mining techniques to parse and map mutagenesis and polymorphism information from the HIV literature, have enriched the data with ancillary information and have developed a public, web-based interface through which it can be intuitively explored: the HIV mutation browser. The current release of the HIV mutation browser describes the phenotypes of 7,608 unique mutations at 2,520 sites in the HIV proteome, resulting from the analysis of 120,899 papers. The mutation information for each protein is organised in a residue-centric manner and each residue is linked to the relevant experimental literature. The importance of HIV as a global health burden advocates extensive effort to maximise the efficiency of HIV research. The HIV mutation browser provides a valuable new resource for the research community. The HIV mutation browser is available at: http://hivmut.org.
Norman E. Davey, Venkata P. Satagopam, Salvador Santiago-Mozos, Carlos Villacorta-Martin, Tanmay A. M. Bharat, Reinhard Schneider 0002, John A. G. Briggs
PLoS Comput. Biol.6
2013 OnTheFly 2.0: A tool for automatic annotation of files and biological information extraction
abstract
Retrieving all of the necessary information from databases about bioentities mentioned in an article is not a trivial or an easy task. Following the daily literature about a specific biological topic and collecting all the necessary information about the bioentities mentioned in the literature manually is tedious and time consuming. OnTheFly 2.0 is a web application mainly designed for non-computer experts which aims to automate data collection and knowledge extraction from biological literature in a user friendly and efficient way. OnTheFly 2.0 is able to extract bioentities from individual articles such as text, Microsoft Word, Excel and PDF files. With a simple drag-and-drop motion, the text of a document is extensively parsed for bioentities such as protein/gene names and chemical compound names. Utilizing high quality data integration platforms, OnTheFly allows the generation of informative summaries, interaction networks and at-a-glance popup windows containing knowledge related to the bioentities found in documents. OnTheFly 2.0 provides a concise application to automate the extraction of bioentities hidden in various documents and is offered as a web based application. It can be found at: http://onthefly.embl.de, http://onthefly.med.uoc.gr or http://onthefly.hcmr.gr.
Evangelos Pafilis, Georgios A. Pavlopoulos, Venkata P. Satagopam, Nikolas Papanikolaou, Heiko Horn, Christos Arvanitidis, Lars Juhl Jensen, Reinhard Schneider 0002
BIBE8
2013 iAnn: an event sharing platform for the life sciences
abstract
SUMMARY: We present iAnn, an open source community-driven platform for dissemination of life science events, such as courses, conferences and workshops. iAnn allows automatic visualisation and integration of customised event reports. A central repository lies at the core of the platform: curators add submitted events, and these are subsequently accessed via web services. Thus, once an iAnn widget is incorporated into a website, it permanently shows timely relevant information as if it were native to the remote site. At the same time, announcements submitted to the repository are automatically disseminated to all portals that query the system. To facilitate the visualization of announcements, iAnn provides powerful filtering options and views, integrated in Google Maps and Google Calendar. All iAnn widgets are freely available. AVAILABILITY: http://iann.pro/iannviewer CONTACT: [email protected].
Rafael C. Jiménez, Juan P. Albar, Jong Bhak, Marie-Claude Blatter, Thomas Blicher, Michelle D. Brazas, Catherine Brooksbank, Aidan Budd, Javier De Las Rivas, Jacqueline Dreyer, Marc A. van Driel, Michael J. Dunn, Pedro L. Fernandes, Celia W. G. van Gelder, Henning Hermjakob, Vassilios Ioannidis, David Phillip Judge, Pascal Kahlem, Eija Korpelainen, Hans-Joachim Kraus, Jane E. Loveland, Christine Mayer, Jennifer McDowall, Federico Morán, Nicola J. Mulder, Tommi H. Nyrönen, Kristian Rother, Gustavo A. Salazar, Reinhard Schneider 0002, Allegra Via, Jose M. Villaveces, Maria Victoria Schneider, Terri K. Attwood, Manuel Corpas
Bioinform.29
2012 EnrichNet: network-based gene set enrichment analysis
abstract
MOTIVATION: Assessing functional associations between an experimentally derived gene or protein set of interest and a database of known gene/protein sets is a common task in the analysis of large-scale functional genomics data. For this purpose, a frequently used approach is to apply an over-representation-based enrichment analysis. However, this approach has four drawbacks: (i) it can only score functional associations of overlapping gene/proteins sets; (ii) it disregards genes with missing annotations; (iii) it does not take into account the network structure of physical interactions between the gene/protein sets of interest and (iv) tissue-specific gene/protein set associations cannot be recognized. RESULTS: To address these limitations, we introduce an integrative analysis approach and web-application called EnrichNet. It combines a novel graph-based statistic with an interactive sub-network visualization to accomplish two complementary goals: improving the prioritization of putative functional gene/protein set associations by exploiting information from molecular interaction networks and tissue-specific gene expression data and enabling a direct biological interpretation of the results. By using the approach to analyse sets of genes with known involvement in human diseases, new pathway associations are identified, reflecting a dense sub-network of interactions between their corresponding proteins. AVAILABILITY: EnrichNet is freely available at http://www.enrichnet.org. CONTACT: [email protected], [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics Online.
Enrico Glaab, Anaïs Baudot, Natalio Krasnogor, Reinhard Schneider 0002, Alfonso Valencia
Bioinform.4
2012 PathVar: analysis of gene and protein expression variance in cellular pathways using microarray data
abstract
SUMMARY: Finding significant differences between the expression levels of genes or proteins across diverse biological conditions is one of the primary goals in the analysis of functional genomics data. However, existing methods for identifying differentially expressed genes or sets of genes by comparing measures of the average expression across predefined sample groups do not detect differential variance in the expression levels across genes in cellular pathways. Since corresponding pathway deregulations occur frequently in microarray gene or protein expression data, we present a new dedicated web application, PathVar, to analyze these data sources. The software ranks pathway-representing gene/protein sets in terms of the differences of the variance in the within-pathway expression levels across different biological conditions. Apart from identifying new pathway deregulation patterns, the tool exploits these patterns by combining different machine learning methods to find clusters of similar samples and build sample classification models. AVAILABILITY: freely available at http://pathvar.embl.de CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Enrico Glaab, Reinhard Schneider 0002
Bioinform.2
2012 ReLiance: a machine learning and literature-based prioritization of receptor - ligand pairings
abstract
MOTIVATION: The prediction of receptor-ligand pairings is an important area of research as intercellular communications are mediated by the successful interaction of these key proteins. As the exhaustive assaying of receptor-ligand pairs is impractical, a computational approach to predict pairings is necessary. We propose a workflow to carry out this interaction prediction task, using a text mining approach in conjunction with a state of the art prediction method, as well as a widely accessible and comprehensive dataset. Among several modern classifiers, random forests have been found to be the best at this prediction task. The training of this classifier was carried out using an experimentally validated dataset of Database of Ligand-Receptor Partners (DLRP) receptor-ligand pairs. New examples, co-cited with the training receptors and ligands, are then classified using the trained classifier. After applying our method, we find that we are able to successfully predict receptor-ligand pairs within the GPCR family with a balanced accuracy of 0.96. Upon further inspection, we find several supported interactions that were not present in the Database of Interacting Proteins (DIPdatabase). We have measured the balanced accuracy of our method resulting in high quality predictions stored in the available database ReLiance. AVAILABILITY: http://homes.esat.kuleuven.be/~bioiuser/ReLianceDB/index.php CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ernesto Iacucci, Léon-Charles Tranchevent, Dusan Popovic, Georgios A. Pavlopoulos, Bart De Moor, Reinhard Schneider 0002, Yves Moreau
Bioinform.6
2012 Paving the future: finding suitable ISMB venues
abstract
The International Society for Computational Biology, ISCB, organizes the largest event in the field of computational biology and bioinformatics, namely the annual international conference on Intelligent Systems for Molecular Biology, the ISMB. This year at ISMB 2012 in Long Beach, ISCB celebrated the 20th anniversary of its flagship meeting. ISCB is a young, lean and efficient society that aspires to make a significant impact with only limited resources. Many constraints make the choice of venues for ISMB a tough challenge. Here, we describe those challenges and invite the contribution of ideas for solutions.
Burkhard Rost, Terry Gaasterland, Thomas Lengauer, Michal Linial, Scott Markel, B. J. Morrison McKay, Reinhard Schneider 0002, Paul Horton, Janet Kelso
Bioinform.7
2012 A bioinformatics e-dating story: computational prediction and prioritization of receptor-ligand pairs
abstract
Regulation of cellular events is initiated, often, via extracellular signaling when a circulating protein ligand interacts with one or more membrane-bound protein receptors. Identification of receptor-ligand pairs is thus an important and difficult task to address as this form of interaction is transient and not well studied. In order to address this problem, we collect the most readily available data from repositories (expression, domain, pathway, sequence, and text-based), and apply a high through-put analysis to this problem. We have worked on the receptor-ligand pairing problem in three main studies. In our first study, using a LS-SVM classifier, we show that we are able to more aptly match members of the chemokine and tgfβ families than a previously published method [ 1 ]. Notably, we are able to achieve an increase in recall of 0.76 over the 0.44 for the matching of receptor-ligands in the tgfβ family. In our subsequent study, we benchmarked several machine learning techniques, and essayed several parameters, on the receptior-ligand interaction prediction task. We found that we could reach a balanced accuracy of 0.84. In our final work, we produce a publicly available database of our results with respect to a text-based in silico prediction workflow. The resulting database, contains several key findings, particularly predictions in the GPCR family with a balanced accuracy of 0.96. The receptor-ligand prediction task is an essential one, as the challenge of predicting such pairs is an important issue in wet-labs, biotech, and pharmaceutical companies. Through several studies, we have determined the most appropriate methodology to predict the receptor-ligand pairs and have made available high-quality predictions at our ReLianceDB website ( http://homes.esat.kuleuven.be/~bioiuser/ReLianceDB ), a tool to aid in performing effective and targeted research.
Ernesto Iacucci, Léon-Charles Tranchevent, Dusan Popovic, Georgios A. Pavlopoulos, Bart De Moor, Reinhard Schneider 0002, Yves Moreau
BMC Bioinform.6
2012 Arena3D: visualizing time-driven phenotypic differences in biological systems
abstract
BACKGROUND: Elucidating the genotype-phenotype connection is one of the big challenges of modern molecular biology. To fully understand this connection, it is necessary to consider the underlying networks and the time factor. In this context of data deluge and heterogeneous information, visualization plays an essential role in interpreting complex and dynamic topologies. Thus, software that is able to bring the network, phenotypic and temporal information together is needed. Arena3D has been previously introduced as a tool that facilitates link discovery between processes. It uses a layered display to separate different levels of information while emphasizing the connections between them. We present novel developments of the tool for the visualization and analysis of dynamic genotype-phenotype landscapes. RESULTS: Version 2.0 introduces novel features that allow handling time course data in a phenotypic context. Gene expression levels or other measures can be loaded and visualized at different time points and phenotypic comparison is facilitated through clustering and correlation display or highlighting of impacting changes through time. Similarity scoring allows the identification of global patterns in dynamic heterogeneous data. In this paper we demonstrate the utility of the tool on two distinct biological problems of different scales. First, we analyze a medium scale dataset that looks at perturbation effects of the pluripotency regulator Nanog in murine embryonic stem cells. Dynamic cluster analysis suggests alternative indirect links between Nanog and other proteins in the core stem cell network. Moreover, recurrent correlations from the epigenetic to the translational level are identified. Second, we investigate a large scale dataset consisting of genome-wide knockdown screens for human genes essential in the mitotic process. Here, a potential new role for the gene lsm14a in cytokinesis is suggested. We also show how phenotypic patterning allows for extensive comparison and identification of high impact knockdown targets. CONCLUSIONS: We present a new visualization approach for perturbation screens with multiple phenotypic outcomes. The novel functionality implemented in Arena3D enables effective understanding and comparison of temporal patterns within morphological layers, to help with the system-wide analysis of dynamic processes. Arena3D is available free of charge for academics as a downloadable standalone application from: http://arena3d.org/.
Maria Secrier, Georgios A. Pavlopoulos, Jan Aerts, Reinhard Schneider 0002
BMC Bioinform.4
2010 LAITOR - Literature Assistant for Identification of Terms co-Occurrences and Relationships
abstract
BACKGROUND: Biological knowledge is represented in scientific literature that often describes the function of genes/proteins (bioentities) in terms of their interactions (biointeractions). Such bioentities are often related to biological concepts of interest that are specific of a determined research field. Therefore, the study of the current literature about a selected topic deposited in public databases, facilitates the generation of novel hypotheses associating a set of bioentities to a common context. RESULTS: We created a text mining system (LAITOR: Literature Assistant for Identification of Terms co-Occurrences and Relationships) that analyses co-occurrences of bioentities, biointeractions, and other biological terms in MEDLINE abstracts. The method accounts for the position of the co-occurring terms within sentences or abstracts. The system detected abstracts mentioning protein-protein interactions in a standard test (BioCreative II IAS test data) with a precision of 0.82-0.89 and a recall of 0.48-0.70. We illustrate the application of LAITOR to the detection of plant response genes in a dataset of 1000 abstracts relevant to the topic. CONCLUSIONS: Text mining tools combining the extraction of interacting bioentities and biological concepts with network displays can be helpful in developing reasonable hypotheses in different scientific backgrounds.
Adriano Barbosa-Silva, Theodoros G. Soldatos, Ivan L. F. Magalhães, Georgios A. Pavlopoulos, Jean-Fred Fontaine, Miguel A. Andrade-Navarro, Reinhard Schneider 0002, José Miguel Ortega
BMC Bioinform.7
2010 Live Coverage of Intelligent Systems for Molecular Biology/European Conference on Computational Biology (ISMB/ECCB) 2009
abstract
peer reviewed
Allyson L. Lister, Ruchira S. Datta, Oliver Hofmann 0001, Roland Krause, Michael Kuhn 0004, Bettina Roth, Reinhard Schneider 0002
PLoS Comput. Biol.7
2010 Live Coverage of Scientific Conferences Using Web Technologies
abstract
peer reviewed
Allyson L. Lister, Ruchira S. Datta, Oliver Hofmann 0001, Roland Krause, Michael Kuhn 0004, Bettina Roth, Reinhard Schneider 0002
PLoS Comput. Biol.7
2010 Reflect: A practical approach to web semantics
Seán I. O'Donoghue, Heiko Horn, Evangelos Pafilis, Sven Haag, Michael Kuhn 0004, Venkata P. Satagopam, Reinhard Schneider 0002, Lars Juhl Jensen
J. Web Semant.7
2009 jClust: a clustering and visualization toolbox
abstract
UNLABELLED: jClust is a user-friendly application which provides access to a set of widely used clustering and clique finding algorithms. The toolbox allows a range of filtering procedures to be applied and is combined with an advanced implementation of the Medusa interactive visualization module. These implemented algorithms are k-Means, Affinity propagation, Bron-Kerbosch, MULIC, Restricted neighborhood search cluster algorithm, Markov clustering and Spectral clustering, while the supported filtering procedures are haircut, outside-inside, best neighbors and density control operations. The combination of a simple input file format, a set of clustering and filtering algorithms linked together with the visualization tool provides a powerful tool for data analysis and information extraction. AVAILABILITY: http://jclust.embl.de/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Georgios A. Pavlopoulos, Charalampos N. Moschopoulos, Sean D. Hooper, Reinhard Schneider 0002, Sophia Kossida
Bioinform.4
2009 OnTheFly: a tool for automated document-based text annotation, data linking and network generation
abstract
UNLABELLED: OnTheFly is a web-based application that applies biological named entity recognition to enrich Microsoft Office, PDF and plain text documents. The input files are converted into the HTML format and then sent to the Reflect tagging server, which highlights biological entity names like genes, proteins and chemicals, and attaches to them JavaScript code to invoke a summary pop-up window. The window provides an overview of relevant information about the entity, such as a protein description, the domain composition, a link to the 3D structure and links to other relevant online resources. OnTheFly is also able to extract the bioentities mentioned in a set of files and to produce a graphical representation of the networks of the known and predicted associations of these entities by retrieving the information from the STITCH database. AVAILABILITY: http://onthefly.embl.de, http://onthefly.embl.de/FAQ.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Georgios A. Pavlopoulos, Evangelos Pafilis, Michael Kuhn 0004, Sean D. Hooper, Reinhard Schneider 0002
Bioinform.5
2009 GIBA: a clustering tool for detecting protein complexes
abstract
BACKGROUND: During the last years, high throughput experimental methods have been developed which generate large datasets of protein - protein interactions (PPIs). However, due to the experimental methodologies these datasets contain errors mainly in terms of false positive data sets and reducing therefore the quality of any derived information. Typically these datasets can be modeled as graphs, where vertices represent proteins and edges the pairwise PPIs, making it easy to apply automated clustering methods to detect protein complexes or other biological significant functional groupings. METHODS: In this paper, a clustering tool, called GIBA (named by the first characters of its developers' nicknames), is presented. GIBA implements a two step procedure to a given dataset of protein-protein interaction data. First, a clustering algorithm is applied to the interaction data, which is then followed by a filtering step to generate the final candidate list of predicted complexes. RESULTS: The efficiency of GIBA is demonstrated through the analysis of 6 different yeast protein interaction datasets in comparison to four other available algorithms. We compared the results of the different methods by applying five different performance measurement metrices. Moreover, the parameters of the methods that constitute the filter have been checked on how they affect the final results. CONCLUSION: GIBA is an effective and easy to use tool for the detection of protein complexes out of experimentally measured protein - protein interaction networks. The results show that GIBA has superior prediction accuracy than previously published methods.
Charalampos N. Moschopoulos, Georgios A. Pavlopoulos, Reinhard Schneider 0002, Spiridon D. Likothanassis, Sophia Kossida
BMC Bioinform.3
2008 Clustering of cognate proteins among distinct proteomes derived from multiple links to a single seed sequence
abstract
BACKGROUND: Modern proteomes evolved by modification of pre-existing ones. It is extremely important to comparative biology that related proteins be identified as members of the same cognate group, since a characterized putative homolog could be used to find clues about the function of uncharacterized proteins from the same group. Typically, databases of related proteins focus on those from completely-sequenced genomes. Unfortunately, relatively few organisms have had their genomes fully sequenced; accordingly, many proteins are ignored by the currently available databases of cognate proteins, despite the high amount of important genes that are functionally described only for these incomplete proteomes. RESULTS: We have developed a method to cluster cognate proteins from multiple organisms beginning with only one sequence, through connectivity saturation with that Seed sequence. We show that the generated clusters are in agreement with some other approaches based on full genome comparison. CONCLUSION: The method produced results that are as reliable as those produced by conventional clustering approaches. Generating clusters based only on individual proteins of interest is less time consuming than generating clusters for whole proteomes.
Adriano Barbosa-Silva, Venkata P. Satagopam, Reinhard Schneider 0002, José Miguel Ortega
BMC Bioinform.3
1997 Sequence analysis of the Methanococcus jannaschii genome and the prediction of protein function
abstract
Miguel Andrade, Georg Casari, Antoine de Daruvar, Chris Sander, Reinhard Schneider, Javier Tamames, Alfonso Valencia, Christos Ouzounis; Sequence analysis
Miguel A. Andrade-Navarro, Georg Casari, Antoine de Daruvar, Chris Sander, Reinhard Schneider 0002, Javier Tamames, Alfonso Valencia, Christos A. Ouzounis
Comput. Appl. Biosci.5
1994 GeneQuiz: A Workbench for Sequence Analysis
Michael Scharf, Reinhard Schneider 0002, Georg Casari, Peer Bork, Alfonso Valencia, Christos A. Ouzounis, Chris Sander
ISMB2
1994 PHD - an automatic mail server for protein secondary structure prediction
abstract
By the middle of 1993, > 30,000 protein sequences has been listed. For 1000 of these, the three-dimensional (tertiary) structure has been experimentally solved. Another 7000 can be modelled by homology. For the remaining 21,000 sequences, secondary structure prediction provides a rough estimate of structural features. Predictions in three states range between 35% (random) and 88% (homology modelling) overall accuracy. Using information about evolutionary conservation as contained in multiple sequence alignments, the secondary structure of 4700 protein sequences was predicted by the automatic e-mail server PHD. For proteins with at least one known homologue, the method has an expected overall three-state accuracy of 71.4% for proteins with at least one known homologue (evaluated on 126 unique protein chains).
Burkhard Rost, Chris Sander, Reinhard Schneider 0002
Comput. Appl. Biosci.3