EDBT 2026 Demo / reviewers in the wild / expert
Philippe Rocca-Serra
dblp:99/588
· DBLP profile ↗
22ranked-venue papers
1as first author
1since 2021 · last 2022
0000-0001-9853-5668ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
10 papers |
Bioinformatics and computational biology · 97% Computational science and engineering · 3% | |
| Computer graphics and multimedia
2 papers |
Visualization and visual analytics · 81% Image and video coding · 19% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 100% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
metabolomics |
0.7 | 2 | 2019 | Interoperable and scalable data analysis with microservices: applications in metabolomics · Bioinform. 2019 mzML2ISA & nmrML2ISA: generating enriched ISA-Tab metadata files from metabolomics XML data · Bioinform. 2017 |
Cloud and datacenter computing
container orchestration |
0.4 | 1 | 2019 | Interoperable and scalable data analysis with microservices: applications in metabolomics · Bioinform. 2019 |
Visualization and visual analytics › visual encoding
glyph design |
0.3 | 2 | 2013 | Visual Compression of Workflow Visualizations with Automated Detection of Macro Motifs · IEEE Trans. Vis. Comput. Graph. 2013 Taxonomy-Based Glyph Design - with a Case Study on Visualizing Workflows of Biological Experiments · IEEE Trans. Vis. Comput. Graph. 2012 |
Bioinformatics and computational biology
data integration |
0.2 | 1 | 2022 | ELIXIR biovalidator for semantic validation of life science metadata · Bioinform. 2022 |
Bioinformatics and computational biology › genome annotation
ontology-based annotation |
0.2 | 1 | 2013 | OntoMaton: a Bioportal powered ontology widget for Google Spreadsheets · Bioinform. 2013 |
Visualization and visual analytics › software visualization
workflow visualization |
0.2 | 1 | 2013 | Visual Compression of Workflow Visualizations with Automated Detection of Macro Motifs · IEEE Trans. Vis. Comput. Graph. 2013 |
Visualization and visual analytics › visual encoding
glyph-based visualization |
0.1 | 1 | 2012 | Taxonomy-Based Glyph Design - with a Case Study on Visualizing Workflows of Biological Experiments · IEEE Trans. Vis. Comput. Graph. 2012 |
Image and video coding
visual communication coding |
0.1 | 1 | 2012 | Taxonomy-Based Glyph Design - with a Case Study on Visualizing Workflows of Biological Experiments · IEEE Trans. Vis. Comput. Graph. 2012 |
Bioinformatics and computational biology › ontology
ontology development |
0.1 | 2 | 2013 | The MGED Ontology: a resource for semantics-based description of microarray experiments · Bioinform. 2006 OntoMaton: a Bioportal powered ontology widget for Google Spreadsheets · Bioinform. 2013 |
Computational science and engineering › scientific data management
data standards |
0.1 | 1 | 2006 | The MGED Ontology: a resource for semantics-based description of microarray experiments · Bioinform. 2006 |
Bioinformatics and computational biology › biological database
gene expression database |
0.1 | 1 | 2005 | The ArrayExpress gene expression database: a software engineering and implementation perspective · Bioinform. 2005 |
Bioinformatics and computational biology › biological database
microarray data management |
0.1 | 1 | 2005 | The ArrayExpress gene expression database: a software engineering and implementation perspective · Bioinform. 2005 |
Methods — techniques the papers use, named apart from their topics
microservice architecture · 0.8kubernetes · 0.8docker · 0.8ontology-based validation · 0.6JSON schema validation · 0.6visual encoding · 0.3graph motif extraction · 0.3XML parsing · 0.3ISA-Tab format conversion · 0.3ontology lookup web services · 0.2visual channel ordering · 0.1taxonomy construction · 0.1size-optimized encoding · 0.1graph transformation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | ELIXIR biovalidator for semantic validation of life science metadataabstractSUMMARY: To advance biomedical research, increasingly large amounts of complex data need to be discovered and integrated. This requires syntactic and semantic validation to ensure shared understanding of relevant entities. This article describes the ELIXIR biovalidator, which extends the syntactic validation of the widely used AJV library with ontology-based validation of JSON documents. AVAILABILITY AND IMPLEMENTATION: Source code: https://github.com/elixir-europe/biovalidator, Release: v1.9.1, License: Apache License 2.0, Deployed at: https://www.ebi.ac.uk/biosamples/schema/validator/validate. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Isuru Udara Liyanage, Tony Burdett, Bert Droesbeke, Karoly Erdos, Rolando Fernandez, Alasdair J. G. Gray, Simon Jupp, Flavia Penim, Cyril Pommier, Philippe Rocca-Serra, Mélanie Courtot, Frederik Coppens |
Bioinform. | 11 |
| 2019 | Interoperable and scalable data analysis with microservices: applications in metabolomicsabstractMOTIVATION: Developing a robust and performant data analysis workflow that integrates all necessary components whilst still being able to scale over multiple compute nodes is a challenging task. We introduce a generic method based on the microservice architecture, where software tools are encapsulated as Docker containers that can be connected into scientific workflows and executed using the Kubernetes container orchestrator. RESULTS: We developed a Virtual Research Environment (VRE) which facilitates rapid integration of new tools and developing scalable and interoperable workflows for performing metabolomics data analysis. The environment can be launched on-demand on cloud resources and desktop computers. IT-expertise requirements on the user side are kept to a minimum, and workflows can be re-used effortlessly by any novice user. We validate our method in the field of metabolomics on two mass spectrometry, one nuclear magnetic resonance spectroscopy and one fluxomics study. We showed that the method scales dynamically with increasing availability of computational resources. We demonstrated that the method facilitates interoperability using integration of the major software suites resulting in a turn-key workflow encompassing all steps for mass-spectrometry-based metabolomics including preprocessing, statistics and identification. Microservices is a generic methodology that can serve any scientific discipline and opens up for new types of large-scale integrative science. AVAILABILITY AND IMPLEMENTATION: The PhenoMeNal consortium maintains a web portal (https://portal.phenomenal-h2020.eu) providing a GUI for launching the Virtual Research Environment. The GitHub repository https://github.com/phnmnl/ hosts the source code of all projects. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Payam Emami Khoonsari, Pablo A. Moreno, Sven Bergmann, Joachim Burman, Marco Capuccini, Matteo Carone, Marta Cascante, Pedro de Atauri, Carles Foguet, Alejandra N. González-Beltrán, Thomas Hankemeier, Kenneth Haug, Sijin He, Stephanie Herman, David Johnson 0006, Namrata Kale, Anders Larsson, Steffen Neumann, Kristian Peters, Luca Pireddu, Philippe Rocca-Serra, Pierrick Roger, Rico Rueedi, Christoph Ruttkies, Noureddin Sadawi, Reza M. Salek, Susanna-Assunta Sansone, Daniel Schober, Vitaly A. Selivanov, Etienne A. Thévenot, Michael van Vliet, Gianluigi Zanetti, Christoph Steinbeck, Kim Kultima, Ola Spjuth |
Bioinform. | 21 |
| 2018 | DataMed - an open source discovery index for finding biomedical datasetsabstractOBJECTIVE: Finding relevant datasets is important for promoting data reuse in the biomedical domain, but it is challenging given the volume and complexity of biomedical data. Here we describe the development of an open source biomedical data discovery system called DataMed, with the goal of promoting the building of additional data indexes in the biomedical domain. MATERIALS AND METHODS: DataMed, which can efficiently index and search diverse types of biomedical datasets across repositories, is developed through the National Institutes of Health-funded biomedical and healthCAre Data Discovery Index Ecosystem (bioCADDIE) consortium. It consists of 2 main components: (1) a data ingestion pipeline that collects and transforms original metadata information to a unified metadata model, called DatA Tag Suite (DATS), and (2) a search engine that finds relevant datasets based on user-entered queries. In addition to describing its architecture and techniques, we evaluated individual components within DataMed, including the accuracy of the ingestion pipeline, the prevalence of the DATS model across repositories, and the overall performance of the dataset retrieval engine. RESULTS AND CONCLUSION: Our manual review shows that the ingestion pipeline could achieve an accuracy of 90% and core elements of DATS had varied frequency across repositories. On a manually curated benchmark dataset, the DataMed search engine achieved an inferred average precision of 0.2033 and a precision at 10 (P@10, the number of relevant results in the top 10 search results) of 0.6022, by implementing advanced natural language processing and terminology services. Currently, we have made the DataMed system publically available as an open source package for the biomedical community. Anupama E. Gururaj, Ibrahim Burak Özyurt, Ruiling Liu, Ergin Soysal, Trevor Cohen, Firat Tiryaki, Yueling Li, Nansu Zong, Min Jiang 0007, Deevakar Rogith, Mandana Salimi, Hyeon-Eui Kim, Philippe Rocca-Serra, Alejandra N. González-Beltrán, Claudiu Farcas, Todd Johnson, Ronald Margolis, George Alter, Susanna-Assunta Sansone, Ian Fore, Lucila Ohno-Machado, Jeffrey S. Grethe, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 14 |
| 2018 | Data discovery with DATS: exemplar adoptions and lessons learnedabstractThe DAta Tag Suite (DATS) is a model supporting dataset description, indexing, and discovery. It is available as an annotated serialization with schema.org, a vocabulary used by major search engines, thus making the datasets discoverable on the web. DATS underlies DataMed, the National Institutes of Health Big Data to Knowledge Data Discovery Index prototype, which aims to provide a "PubMed for datasets." The experience gained while indexing a heterogeneous range of >60 repositories in DataMed helped in evaluating DATS's entities, attributes, and scope. In this work, 3 additional exemplary and diverse data sources were mapped to DATS by their representatives or experts, offering a deep scan of DATS fitness against a new set of existing data. The procedure, including feedback from users and implementers, resulted in DATS implementation guidelines and best practices, and identification of a path for evolving and optimizing the model. Finally, the work exposed additional needs when defining datasets for indexing, especially in the context of clinical and observational information. Alejandra N. González-Beltrán, John Campbell 0004, Patrick J. Dunn, Diana Guijarro, Sanda Ionescu, Hyeon-Eui Kim, Jared Lyle, Jeffrey A. Wiser, Susanna-Assunta Sansone, Philippe Rocca-Serra |
J. Am. Medical Informatics Assoc. | 10 |
| 2017 | mzML2ISA & nmrML2ISA: generating enriched ISA-Tab metadata files from metabolomics XML dataabstractSUMMARY: Submission to the MetaboLights repository for metabolomics data currently places the burden of reporting instrument and acquisition parameters in ISA-Tab format on users, who have to do it manually, a process that is time consuming and prone to user input error. Since the large majority of these parameters are embedded in instrument raw data files, an opportunity exists to capture this metadata more accurately. Here we report a set of Python packages that can automatically generate ISA-Tab metadata file stubs from raw XML metabolomics data files. The parsing packages are separated into mzML2ISA (encompassing mzML and imzML formats) and nmrML2ISA (nmrML format only). Overall, the use of mzML2ISA & nmrML2ISA reduces the time needed to capture metadata substantially (capturing 90% of metadata on assay and sample levels), is much less prone to user input errors, improves compliance with minimum information reporting guidelines and facilitates more finely grained data exploration and querying of datasets. AVAILABILITY AND IMPLEMENTATION: mzML2ISA & nmrML2ISA are available under version 3 of the GNU General Public Licence at https://github.com/ISA-tools. Documentation is available from http://2isa.readthedocs.io/en/latest/. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Martin Larralde, Thomas N. Lawson, Ralf J. M. Weber, Pablo A. Moreno, Kenneth Haug, Philippe Rocca-Serra, Mark R. Viant, Christoph Steinbeck, Reza M. Salek |
Bioinform. | 6 |
| 2016 | A Scalable Dataset Indexing Infrastructure for the bioCADDIE Data Discovery System
Jeffrey S. Grethe, Ibrahim Burak Özyurt, Hua Xu 0001, Ruiling Liu, Ergin Soysal, Anupama E. Gururaj, Hyeon-Eui Kim, Trevor Cohen, Todd R. Johnson, Mandana Salimi, Saeid Pournejati, Min Jiang 0007, Claudiu Farcas, Alejandra N. González-Beltrán, Philippe Rocca-Serra, Muhamamd F. Amith, Cui Tao, Ian Fore, Ronald Margolis, George Alter, Susanna-Assunta Sansone, Lucila Ohno-Machado |
AMIA | 16 |
| 2016 | Development of DataMed, a Data Discovery Index Prototype by bioCADDIE: Laying the Groundwork for Biomedical Data Discovery
Hua Xu 0001, Jeffrey S. Grethe, Ruiling Liu, Ergin Soysal, Anupama E. Gururaj, Yueling Li, Ibrahim Burak Özyurt, Hyeon-Eui Kim, Trevor Cohen, Todd R. Johnson, Mandana Salimi, Saeid Pournejati, Min Jiang 0007, Claudiu Farcas, Alejandra N. González-Beltrán, Philippe Rocca-Serra, Muhamamd F. Amith, Cui Tao, Ian Fore, Ronald Margolis, George Alter, Susanna-Assunta Sansone, Lucila Ohno-Machado |
AMIA | 17 |
| 2015 | The center for expanded data annotation and retrievalabstractThe Center for Expanded Data Annotation and Retrieval is studying the creation of comprehensive and expressive metadata for biomedical datasets to facilitate data discovery, data interpretation, and data reuse. We take advantage of emerging community-based standard templates for describing different kinds of biomedical datasets, and we investigate the use of computational techniques to help investigators to assemble templates and to fill in their values. We are creating a repository of metadata from which we plan to identify metadata patterns that will drive predictive data entry when filling in metadata templates. The metadata repository not only will capture annotations specified when experimental datasets are initially created, but also will incorporate links to the published literature, including secondary analyses and possible refinements or retractions of experimental interpretations. By working initially with the Human Immunology Project Consortium and the developers of the ImmPort data repository, we are developing and evaluating an end-to-end solution to the problems of metadata authoring and management that will generalize to other data-management environments. Mark A. Musen, Carol A. Bean, Kei-Hoi Cheung, Michel Dumontier, Kim A. Durante, Olivier Gevaert, Alejandra N. González-Beltrán, Purvesh Khatri, Steven H. Kleinstein, Martin J. O'Connor, Yannick Pouliot, Philippe Rocca-Serra, Susanna-Assunta Sansone, Jeffrey A. Wiser |
J. Am. Medical Informatics Assoc. | 12 |
| 2014 | linkedISA: semantic representation of ISA-Tab experimental metadataabstractBACKGROUND: Reporting and sharing experimental metadata- such as the experimental design, characteristics of the samples, and procedures applied, along with the analysis results, in a standardised manner ensures that datasets are comprehensible and, in principle, reproducible, comparable and reusable. Furthermore, sharing datasets in formats designed for consumption by humans and machines will also maximize their use. The Investigation/Study/Assay (ISA) open source metadata tracking framework facilitates standards-compliant collection, curation, visualization, storage and sharing of datasets, leveraging on other platforms to enable analysis and publication. The ISA software suite includes several components used in increasingly diverse set of life science and biomedical domains; it is underpinned by a general-purpose format, ISA-Tab, and conversions exist into formats required by public repositories. While ISA-Tab works well mainly as a human readable format, we have also implemented a linked data approach to semantically define the ISA-Tab syntax. RESULTS: We present a semantic web representation of the ISA-Tab syntax that complements ISA-Tab's syntactic interoperability with semantic interoperability. We introduce the linkedISA conversion tool from ISA-Tab to the Resource Description Framework (RDF), supporting mappings from the ISA syntax to multiple community-defined, open ontologies and capitalising on user-provided ontology annotations in the experimental metadata. We describe insights of the implementation and how annotations can be expanded driven by the metadata. We applied the conversion tool as part of Bio-GraphIIn, a web-based application supporting integration of the semantically-rich experimental descriptions. Designed in a user-friendly manner, the Bio-GraphIIn interface hides most of the complexities to the users, exposing a familiar tabular view of the experimental description to allow seamless interaction with the RDF representation, and visualising descriptors to drive the query over the semantic representation of the experimental design. In addition, we defined queries over the linkedISA RDF representation and demonstrated its use over the linkedISA conversion of datasets from Nature' Scientific Data online publication. CONCLUSIONS: Our linked data approach has allowed us to: 1) make the ISA-Tab semantics explicit and machine-processable, 2) exploit the existing ontology-based annotations in the ISA-Tab experimental descriptions, 3) augment the ISA-Tab syntax with new descriptive elements, 4) visualise and query elements related to the experimental design. Reasoning over ISA-Tab metadata and associated data will facilitate data integration and knowledge discovery. Alejandra N. González-Beltrán, Eamonn Maguire, Susanna-Assunta Sansone, Philippe Rocca-Serra |
BMC Bioinform. | 4 |
| 2014 | The Risa R/Bioconductor package: integrative data analysis from experimental metadata and back againabstractBACKGROUND: The ISA-Tab format and software suite have been developed to break the silo effect induced by technology-specific formats for a variety of data types and to better support experimental metadata tracking. Experimentalists seldom use a single technique to monitor biological signals. Providing a multi-purpose, pragmatic and accessible format that abstracts away common constructs for describing Investigations, Studies and Assays, ISA is increasingly popular. To attract further interest towards the format and extend support to ensure reproducible research and reusable data, we present the Risa package, which delivers a central component to support the ISA format by enabling effortless integration with R, the popular, open source data crunching environment. RESULTS: The Risa package bridges the gap between the metadata collection and curation in an ISA-compliant way and the data analysis using the widely used statistical computing environment R. The package offers functionality for: i) parsing ISA-Tab datasets into R objects, ii) augmenting annotation with extra metadata not explicitly stated in the ISA syntax; iii) interfacing with domain specific R packages iv) suggesting potentially useful R packages available in Bioconductor for subsequent processing of the experimental data described in the ISA format; and finally v) saving back to ISA-Tab files augmented with analysis specific metadata from R. We demonstrate these features by presenting use cases for mass spectrometry data and DNA microarray data. CONCLUSIONS: The Risa package is open source (with LGPL license) and freely available through Bioconductor. By making Risa available, we aim to facilitate the task of processing experimental data, encouraging a uniform representation of experimental information and results while delivering tools for ensuring traceability and provenance tracking. SOFTWARE AVAILABILITY: The Risa package is available since Bioconductor 2.11 (version 1.0.0) and version 1.2.1 appeared in Bioconductor 2.12, both along with documentation and examples. The latest version of the code is at the development branch in Bioconductor and can also be accessed from GitHub https://github.com/ISA-tools/Risa, where the issue tracker allows users to report bugs or feature requests. Alejandra N. González-Beltrán, Steffen Neumann, Eamonn Maguire, Susanna-Assunta Sansone, Philippe Rocca-Serra |
BMC Bioinform. | 5 |
| 2013 | OntoMaton: a Bioportal powered ontology widget for Google SpreadsheetsabstractMOTIVATION: Data collection in spreadsheets is ubiquitous, but current solutions lack support for collaborative semantic annotation that would promote shared and interdisciplinary annotation practices, supporting geographically distributed players. RESULTS: OntoMaton is an open source solution that brings ontology lookup and tagging capabilities into a cloud-based collaborative editing environment, harnessing Google Spreadsheets and the NCBO Web services. It is a general purpose, format-agnostic tool that may serve as a component of the ISA software suite. OntoMaton can also be used to assist the ontology development process. AVAILABILITY: OntoMaton is freely available from Google widgets under the CPAL open source license; documentation and examples at: https://github.com/ISA-tools/OntoMaton. Eamonn Maguire, Alejandra N. González-Beltrán, Patricia L. Whetzel, Susanna-Assunta Sansone, Philippe Rocca-Serra |
Bioinform. | 5 |
| 2013 | Visual Compression of Workflow Visualizations with Automated Detection of Macro MotifsabstractThis paper is concerned with the creation of 'macros' in workflow visualization as a support tool to increase the efficiency of data curation tasks. We propose computation of candidate macros based on their usage in large collections of workflows in data repositories. We describe an efficient algorithm for extracting macro motifs from workflow graphs. We discovered that the state transition information, used to identify macro candidates, characterizes the structural pattern of the macro and can be harnessed as part of the visual design of the corresponding macro glyph. This facilitates partial automation and consistency in glyph design applicable to a large set of macro glyphs. We tested this approach against a repository of biological data holding some 9,670 workflows and found that the algorithmically generated candidate macros are in keeping with domain expert expectations. Eamonn Maguire, Philippe Rocca-Serra, Susanna-Assunta Sansone, Jim Davies, Min Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2012 | graph2tab, a library to convert experimental workflow graphs into tabular formatsabstractAbstract Motivations: Spreadsheet-like tabular formats are ever more popular in the biomedical field as a mean for experimental reporting. The problem of converting the graph of an experimental workflow into a table-based representation occurs in many such formats and is not easy to solve. Results: We describe graph2tab, a library that implements methods to realise such a conversion in a size-optimised way. Our solution is generic and can be adapted to specific cases of data exporters or data converters that need to be implemented. Availability and Implementation: The library source code and documentation are available at http://github.com/ISA-tools/graph2tab. Contact: [email protected]. Supplementary Information: A supplementary document describes the theoretical and technical details about the library implementation. Marco Brandizi, Natalja Kurbatova, Ugis Sarkans, Philippe Rocca-Serra |
Bioinform. | 4 |
| 2012 | Translating standards into practice - One Semantic Web API for Gene Expression
Helena F. Deus, Eric Prud'hommeaux, Michael Miller 0001, Jun Zhao 0003, James Malone, Tomasz Adamusiak, Jamie P. McCusker, Sudeshna Das 0001, Philippe Rocca-Serra, Ronan Fox, M. Scott Marshall |
J. Biomed. Informatics | 9 |
| 2012 | Taxonomy-Based Glyph Design - with a Case Study on Visualizing Workflows of Biological ExperimentsabstractGlyph-based visualization can offer elegant and concise presentation of multivariate information while enhancing speed and ease in visual search experienced by users. As with icon designs, glyphs are usually created based on the designers' experience and intuition, often in a spontaneous manner. Such a process does not scale well with the requirements of applications where a large number of concepts are to be encoded using glyphs. To alleviate such limitations, we propose a new systematic process for glyph design by exploring the parallel between the hierarchy of concept categorization and the ordering of discriminative capacity of visual channels. We examine the feasibility of this approach in an application where there is a pressing need for an efficient and effective means to visualize workflows of biological experiments. By processing thousands of workflow records in a public archive of biological experiments, we demonstrate that a cost-effective glyph design can be obtained by following a process of formulating a taxonomy with the aid of computation, identifying visual channels hierarchically, and defining application-specific abstraction and metaphors. Eamonn Maguire, Philippe Rocca-Serra, Susanna-Assunta Sansone, Jim Davies, Min Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2010 | ISA software suite: supporting standards-compliant experimental annotation and enabling curation at the community levelabstractUNLABELLED: The first open source software suite for experimentalists and curators that (i) assists in the annotation and local management of experimental metadata from high-throughput studies employing one or a combination of omics and other technologies; (ii) empowers users to uptake community-defined checklists and ontologies; and (iii) facilitates submission to international public repositories. AVAILABILITY AND IMPLEMENTATION: Software, documentation, case studies and implementations at http://www.isa-tools.org. Philippe Rocca-Serra, Marco Brandizi, Eamonn Maguire, Nataliya Sklyar, Chris F. Taylor, Kimberly Begley, Dawn Field, Stephen C. Harris, Winston Hide, Oliver Hofmann 0001, Steffen Neumann, Peter Sterk, Weida Tong, Susanna-Assunta Sansone |
Bioinform. | 1 |
| 2009 | Survey-based naming conventions for use in OBO Foundry ontology developmentabstractBACKGROUND: A wide variety of ontologies relevant to the biological and medical domains are available through the OBO Foundry portal, and their number is growing rapidly. Integration of these ontologies, while requiring considerable effort, is extremely desirable. However, heterogeneities in format and style pose serious obstacles to such integration. In particular, inconsistencies in naming conventions can impair the readability and navigability of ontology class hierarchies, and hinder their alignment and integration. While other sources of diversity are tremendously complex and challenging, agreeing a set of common naming conventions is an achievable goal, particularly if those conventions are based on lessons drawn from pooled practical experience and surveys of community opinion. RESULTS: We summarize a review of existing naming conventions and highlight certain disadvantages with respect to general applicability in the biological domain. We also present the results of a survey carried out to establish which naming conventions are currently employed by OBO Foundry ontologies and to determine what their special requirements regarding the naming of entities might be. Lastly, we propose an initial set of typographic, syntactic and semantic conventions for labelling classes in OBO Foundry ontologies. CONCLUSION: Adherence to common naming conventions is more than just a matter of aesthetics. Such conventions provide guidance to ontology creators, help developers avoid flaws and inaccuracies when editing, and especially when interlinking, ontologies. Common naming conventions will also assist consumers of ontologies to more readily understand what meanings were intended by the authors of ontologies used in annotating bodies of data. Daniel Schober, Barry Smith 0001, Suzanna Lewis, Waclaw Kusnierczyk, Jane Lomax, Chris Mungall, Chris F. Taylor, Philippe Rocca-Serra, Susanna-Assunta Sansone |
BMC Bioinform. | 8 |
| 2006 | The MGED Ontology: a resource for semantics-based description of microarray experimentsabstractMOTIVATION: The generation of large amounts of microarray data and the need to share these data bring challenges for both data management and annotation and highlights the need for standards. MIAME specifies the minimum information needed to describe a microarray experiment and the Microarray Gene Expression Object Model (MAGE-OM) and resulting MAGE-ML provide a mechanism to standardize data representation for data exchange, however a common terminology for data annotation is needed to support these standards. RESULTS: Here we describe the MGED Ontology (MO) developed by the Ontology Working Group of the Microarray Gene Expression Data (MGED) Society. The MO provides terms for annotating all aspects of a microarray experiment from the design of the experiment and array layout, through to the preparation of the biological sample and the protocols used to hybridize the RNA and analyze the data. The MO was developed to provide terms for annotating experiments in line with the MIAME guidelines, i.e. to provide the semantics to describe a microarray experiment according to the concepts specified in MIAME. The MO does not attempt to incorporate terms from existing ontologies, e.g. those that deal with anatomical parts or developmental stages terms, but provides a framework to reference terms in other ontologies and therefore facilitates the use of ontologies in microarray data annotation. AVAILABILITY: The MGED Ontology version.1.2.0 is available as a file in both DAML and OWL formats at http://mged.sourceforge.net/ontologies/index.php. Release notes and annotation examples are provided. The MO is also provided via the NCICB's Enterprise Vocabulary System (http://nciterms.nci.nih.gov/NCIBrowser/Dictionary.do). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Patricia L. Whetzel, Helen E. Parkinson, Helen C. Causton, Liju Fan, Jennifer Fostel, Gilberto Fragoso, Laurence Game, Mervi Heiskanen, Norman Morrison, Philippe Rocca-Serra, Susanna-Assunta Sansone, Chris F. Taylor, Joseph White, Christian J. Stoeckert Jr. |
Bioinform. | 10 |
| 2006 | The use of concept maps during knowledge elicitation in ontology development processes - the nutrigenomics use caseabstractBACKGROUND: Incorporation of ontologies into annotations has enabled 'semantic integration' of complex data, making explicit the knowledge within a certain field. One of the major bottlenecks in developing bio-ontologies is the lack of a unified methodology. Different methodologies have been proposed for different scenarios, but there is no agreed-upon standard methodology for building ontologies. The involvement of geographically distributed domain experts, the need for domain experts to lead the design process, the application of the ontologies and the life cycles of bio-ontologies are amongst the features not considered by previously proposed methodologies. RESULTS: Here, we present a methodology for developing ontologies within the biological domain. We describe our scenario, competency questions, results and milestones for each methodological stage. We introduce the use of concept maps during knowledge acquisition phases as a feasible transition between domain expert and knowledge engineer. CONCLUSION: The contributions of this paper are the thorough description of the steps we suggest when building an ontology, example use of concept maps, consideration of applicability to the development of lower-level ontologies and application to decentralised environments. We have found that within our scenario conceptual maps played an important role in the development process. Alexander García Castro, Philippe Rocca-Serra, Robert Stevens 0001, Chris F. Taylor, Karim Nashar, Mark A. Ragan, Susanna-Assunta Sansone |
BMC Bioinform. | 2 |
| 2006 | A simple spreadsheet-based, MIAME-supportive format for microarray data: MAGE-TABabstractBACKGROUND: Sharing of microarray data within the research community has been greatly facilitated by the development of the disclosure and communication standards MIAME and MAGE-ML by the MGED Society. However, the complexity of the MAGE-ML format has made its use impractical for laboratories lacking dedicated bioinformatics support. RESULTS: We propose a simple tab-delimited, spreadsheet-based format, MAGE-TAB, which will become a part of the MAGE microarray data standard and can be used for annotating and communicating microarray data in a MIAME compliant fashion. CONCLUSION: MAGE-TAB will enable laboratories without bioinformatics experience or support to manage, exchange and submit well-annotated microarray data in a standard format using a spreadsheet. The MAGE-TAB format is self-contained, and does not require an understanding of MAGE-ML or XML. Tim F. Rayner, Philippe Rocca-Serra, Paul T. Spellman, Helen C. Causton, Anna Farne, Ele Holloway, Rafael A. Irizarry, Junmin Liu, Donald Maier, Michael Miller 0001, Kjell Petersen, John Quackenbush, Gavin Sherlock, Christian J. Stoeckert Jr., Joseph White, Patricia L. Whetzel, Farrell Wymore, Helen E. Parkinson, Ugis Sarkans, Catherine A. Ball, Alvis Brazma |
BMC Bioinform. | 2 |
| 2005 | The ArrayExpress gene expression database: a software engineering and implementation perspectiveabstractMOTIVATION: The lack of microarray data management systems and databases is still one of the major problems faced by many life sciences laboratories. While developing the public repository for microarray data ArrayExpress we had to find novel solutions to many non-trivial software engineering problems. Our experience will be both relevant and useful for most bioinformaticians involved in developing information systems for a wide range of high-throughput technologies. RESULTS: ArrayExpress has been online since February 2002, growing exponentially to well over 10,000 hybridizations (as of September 2004). It has been demonstrated that our chosen design and implementation works for databases aimed at storage, access and sharing of high-throughput data. AVAILABILITY: The ArrayExpress database is available at http://www.ebi.ac.uk/arrayexpress/. The software is open source. CONTACT: [email protected]. Ugis Sarkans, Helen E. Parkinson, Gonzalo Garcia Lara, Ahmet Oezcimen, Anjan Sharma, Niran Abeygunawardena, Sergio Contrino, Ele Holloway, Philippe Rocca-Serra, Gaurab Mukherjee, Mohammadreza Shojatalab, Misha Kapushesky, Susanna-Assunta Sansone, Anna Farne, Tim F. Rayner, Alvis Brazma |
Bioinform. | 9 |
| 2002 | An open letter to the scientific journalsabstractCatherine A. Ball, Gavin Sherlock, Helen Parkinson, Philippe Rocca-Sera, Catherine Brooksbank, Helen C. Causton, Duccio Cavalieri, Terry Gaasterland, Pascal Hin Catherine A. Ball, Gavin Sherlock, Helen E. Parkinson, Philippe Rocca-Serra, Catherine Brooksbank, Helen C. Causton, Duccio Cavalieri, Terry Gaasterland, Pascal Hingamp, Frank C. P. Holstege, Martin Ringwald, Paul T. Spellman, Christian J. Stoeckert Jr., Jason E. Stewart, Ronald C. Taylor, Alvis Brazma, John Quackenbush |
Bioinform. | 4 |