Philippe Rocca-Serra

dblp:99/588 · DBLP profile ↗
← Back
22ranked-venue papers
1as first author
1since 2021 · last 2022
0000-0001-9853-5668ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 20 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
10 papers
Bioinformatics and computational biology · 97% Computational science and engineering · 3%
Computer graphics and multimedia
2 papers
Visualization and visual analytics · 81% Image and video coding · 19%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
metabolomics
0.722019
Interoperable and scalable data analysis with microservices: applications in metabolomics · Bioinform. 2019
mzML2ISA & nmrML2ISA: generating enriched ISA-Tab metadata files from metabolomics XML data · Bioinform. 2017
Cloud and datacenter computing
container orchestration
0.412019
Interoperable and scalable data analysis with microservices: applications in metabolomics · Bioinform. 2019
Visualization and visual analytics › visual encoding
glyph design
0.322013
Visual Compression of Workflow Visualizations with Automated Detection of Macro Motifs · IEEE Trans. Vis. Comput. Graph. 2013
Taxonomy-Based Glyph Design - with a Case Study on Visualizing Workflows of Biological Experiments · IEEE Trans. Vis. Comput. Graph. 2012
Bioinformatics and computational biology
data integration
0.212022
ELIXIR biovalidator for semantic validation of life science metadata · Bioinform. 2022
Bioinformatics and computational biology › genome annotation
ontology-based annotation
0.212013
OntoMaton: a Bioportal powered ontology widget for Google Spreadsheets · Bioinform. 2013
Visualization and visual analytics › software visualization
workflow visualization
0.212013
Visual Compression of Workflow Visualizations with Automated Detection of Macro Motifs · IEEE Trans. Vis. Comput. Graph. 2013
Visualization and visual analytics › visual encoding
glyph-based visualization
0.112012
Taxonomy-Based Glyph Design - with a Case Study on Visualizing Workflows of Biological Experiments · IEEE Trans. Vis. Comput. Graph. 2012
Image and video coding
visual communication coding
0.112012
Taxonomy-Based Glyph Design - with a Case Study on Visualizing Workflows of Biological Experiments · IEEE Trans. Vis. Comput. Graph. 2012
Bioinformatics and computational biology › ontology
ontology development
0.122013
The MGED Ontology: a resource for semantics-based description of microarray experiments · Bioinform. 2006
OntoMaton: a Bioportal powered ontology widget for Google Spreadsheets · Bioinform. 2013
Computational science and engineering › scientific data management
data standards
0.112006
The MGED Ontology: a resource for semantics-based description of microarray experiments · Bioinform. 2006
Bioinformatics and computational biology › biological database
gene expression database
0.112005
The ArrayExpress gene expression database: a software engineering and implementation perspective · Bioinform. 2005
Bioinformatics and computational biology › biological database
microarray data management
0.112005
The ArrayExpress gene expression database: a software engineering and implementation perspective · Bioinform. 2005

Methods — techniques the papers use, named apart from their topics

microservice architecture · 0.8kubernetes · 0.8docker · 0.8ontology-based validation · 0.6JSON schema validation · 0.6visual encoding · 0.3graph motif extraction · 0.3XML parsing · 0.3ISA-Tab format conversion · 0.3ontology lookup web services · 0.2visual channel ordering · 0.1taxonomy construction · 0.1size-optimized encoding · 0.1graph transformation · 0.1
YearPublicationVenuePosition
2022 ELIXIR biovalidator for semantic validation of life science metadata
abstract
SUMMARY: To advance biomedical research, increasingly large amounts of complex data need to be discovered and integrated. This requires syntactic and semantic validation to ensure shared understanding of relevant entities. This article describes the ELIXIR biovalidator, which extends the syntactic validation of the widely used AJV library with ontology-based validation of JSON documents. AVAILABILITY AND IMPLEMENTATION: Source code: https://github.com/elixir-europe/biovalidator, Release: v1.9.1, License: Apache License 2.0, Deployed at: https://www.ebi.ac.uk/biosamples/schema/validator/validate. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Isuru Udara Liyanage, Tony Burdett, Bert Droesbeke, Karoly Erdos, Rolando Fernandez, Alasdair J. G. Gray, Simon Jupp, Flavia Penim, Cyril Pommier, Philippe Rocca-Serra, Mélanie Courtot, Frederik Coppens
Bioinform.11
2019 Interoperable and scalable data analysis with microservices: applications in metabolomics
abstract
MOTIVATION: Developing a robust and performant data analysis workflow that integrates all necessary components whilst still being able to scale over multiple compute nodes is a challenging task. We introduce a generic method based on the microservice architecture, where software tools are encapsulated as Docker containers that can be connected into scientific workflows and executed using the Kubernetes container orchestrator. RESULTS: We developed a Virtual Research Environment (VRE) which facilitates rapid integration of new tools and developing scalable and interoperable workflows for performing metabolomics data analysis. The environment can be launched on-demand on cloud resources and desktop computers. IT-expertise requirements on the user side are kept to a minimum, and workflows can be re-used effortlessly by any novice user. We validate our method in the field of metabolomics on two mass spectrometry, one nuclear magnetic resonance spectroscopy and one fluxomics study. We showed that the method scales dynamically with increasing availability of computational resources. We demonstrated that the method facilitates interoperability using integration of the major software suites resulting in a turn-key workflow encompassing all steps for mass-spectrometry-based metabolomics including preprocessing, statistics and identification. Microservices is a generic methodology that can serve any scientific discipline and opens up for new types of large-scale integrative science. AVAILABILITY AND IMPLEMENTATION: The PhenoMeNal consortium maintains a web portal (https://portal.phenomenal-h2020.eu) providing a GUI for launching the Virtual Research Environment. The GitHub repository https://github.com/phnmnl/ hosts the source code of all projects. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Payam Emami Khoonsari, Pablo A. Moreno, Sven Bergmann, Joachim Burman, Marco Capuccini, Matteo Carone, Marta Cascante, Pedro de Atauri, Carles Foguet, Alejandra N. González-Beltrán, Thomas Hankemeier, Kenneth Haug, Sijin He, Stephanie Herman, David Johnson 0006, Namrata Kale, Anders Larsson, Steffen Neumann, Kristian Peters, Luca Pireddu, Philippe Rocca-Serra, Pierrick Roger, Rico Rueedi, Christoph Ruttkies, Noureddin Sadawi, Reza M. Salek, Susanna-Assunta Sansone, Daniel Schober, Vitaly A. Selivanov, Etienne A. Thévenot, Michael van Vliet, Gianluigi Zanetti, Christoph Steinbeck, Kim Kultima, Ola Spjuth
Bioinform.21
2018 DataMed - an open source discovery index for finding biomedical datasets
abstract
OBJECTIVE: Finding relevant datasets is important for promoting data reuse in the biomedical domain, but it is challenging given the volume and complexity of biomedical data. Here we describe the development of an open source biomedical data discovery system called DataMed, with the goal of promoting the building of additional data indexes in the biomedical domain. MATERIALS AND METHODS: DataMed, which can efficiently index and search diverse types of biomedical datasets across repositories, is developed through the National Institutes of Health-funded biomedical and healthCAre Data Discovery Index Ecosystem (bioCADDIE) consortium. It consists of 2 main components: (1) a data ingestion pipeline that collects and transforms original metadata information to a unified metadata model, called DatA Tag Suite (DATS), and (2) a search engine that finds relevant datasets based on user-entered queries. In addition to describing its architecture and techniques, we evaluated individual components within DataMed, including the accuracy of the ingestion pipeline, the prevalence of the DATS model across repositories, and the overall performance of the dataset retrieval engine. RESULTS AND CONCLUSION: Our manual review shows that the ingestion pipeline could achieve an accuracy of 90% and core elements of DATS had varied frequency across repositories. On a manually curated benchmark dataset, the DataMed search engine achieved an inferred average precision of 0.2033 and a precision at 10 (P@10, the number of relevant results in the top 10 search results) of 0.6022, by implementing advanced natural language processing and terminology services. Currently, we have made the DataMed system publically available as an open source package for the biomedical community.
Anupama E. Gururaj, Ibrahim Burak Özyurt, Ruiling Liu, Ergin Soysal, Trevor Cohen, Firat Tiryaki, Yueling Li, Nansu Zong, Min Jiang 0007, Deevakar Rogith, Mandana Salimi, Hyeon-Eui Kim, Philippe Rocca-Serra, Alejandra N. González-Beltrán, Claudiu Farcas, Todd Johnson, Ronald Margolis, George Alter, Susanna-Assunta Sansone, Ian Fore, Lucila Ohno-Machado, Jeffrey S. Grethe, Hua Xu 0001
J. Am. Medical Informatics Assoc.14
2018 Data discovery with DATS: exemplar adoptions and lessons learned
abstract
The DAta Tag Suite (DATS) is a model supporting dataset description, indexing, and discovery. It is available as an annotated serialization with schema.org, a vocabulary used by major search engines, thus making the datasets discoverable on the web. DATS underlies DataMed, the National Institutes of Health Big Data to Knowledge Data Discovery Index prototype, which aims to provide a "PubMed for datasets." The experience gained while indexing a heterogeneous range of >60 repositories in DataMed helped in evaluating DATS's entities, attributes, and scope. In this work, 3 additional exemplary and diverse data sources were mapped to DATS by their representatives or experts, offering a deep scan of DATS fitness against a new set of existing data. The procedure, including feedback from users and implementers, resulted in DATS implementation guidelines and best practices, and identification of a path for evolving and optimizing the model. Finally, the work exposed additional needs when defining datasets for indexing, especially in the context of clinical and observational information.
Alejandra N. González-Beltrán, John Campbell 0004, Patrick J. Dunn, Diana Guijarro, Sanda Ionescu, Hyeon-Eui Kim, Jared Lyle, Jeffrey A. Wiser, Susanna-Assunta Sansone, Philippe Rocca-Serra
J. Am. Medical Informatics Assoc.10
2017 mzML2ISA & nmrML2ISA: generating enriched ISA-Tab metadata files from metabolomics XML data
abstract
SUMMARY: Submission to the MetaboLights repository for metabolomics data currently places the burden of reporting instrument and acquisition parameters in ISA-Tab format on users, who have to do it manually, a process that is time consuming and prone to user input error. Since the large majority of these parameters are embedded in instrument raw data files, an opportunity exists to capture this metadata more accurately. Here we report a set of Python packages that can automatically generate ISA-Tab metadata file stubs from raw XML metabolomics data files. The parsing packages are separated into mzML2ISA (encompassing mzML and imzML formats) and nmrML2ISA (nmrML format only). Overall, the use of mzML2ISA & nmrML2ISA reduces the time needed to capture metadata substantially (capturing 90% of metadata on assay and sample levels), is much less prone to user input errors, improves compliance with minimum information reporting guidelines and facilitates more finely grained data exploration and querying of datasets. AVAILABILITY AND IMPLEMENTATION: mzML2ISA & nmrML2ISA are available under version 3 of the GNU General Public Licence at https://github.com/ISA-tools. Documentation is available from http://2isa.readthedocs.io/en/latest/. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Martin Larralde, Thomas N. Lawson, Ralf J. M. Weber, Pablo A. Moreno, Kenneth Haug, Philippe Rocca-Serra, Mark R. Viant, Christoph Steinbeck, Reza M. Salek
Bioinform.6
2016 A Scalable Dataset Indexing Infrastructure for the bioCADDIE Data Discovery System
Jeffrey S. Grethe, Ibrahim Burak Özyurt, Hua Xu 0001, Ruiling Liu, Ergin Soysal, Anupama E. Gururaj, Hyeon-Eui Kim, Trevor Cohen, Todd R. Johnson, Mandana Salimi, Saeid Pournejati, Min Jiang 0007, Claudiu Farcas, Alejandra N. González-Beltrán, Philippe Rocca-Serra, Muhamamd F. Amith, Cui Tao, Ian Fore, Ronald Margolis, George Alter, Susanna-Assunta Sansone, Lucila Ohno-Machado
AMIA16
2016 Development of DataMed, a Data Discovery Index Prototype by bioCADDIE: Laying the Groundwork for Biomedical Data Discovery
Hua Xu 0001, Jeffrey S. Grethe, Ruiling Liu, Ergin Soysal, Anupama E. Gururaj, Yueling Li, Ibrahim Burak Özyurt, Hyeon-Eui Kim, Trevor Cohen, Todd R. Johnson, Mandana Salimi, Saeid Pournejati, Min Jiang 0007, Claudiu Farcas, Alejandra N. González-Beltrán, Philippe Rocca-Serra, Muhamamd F. Amith, Cui Tao, Ian Fore, Ronald Margolis, George Alter, Susanna-Assunta Sansone, Lucila Ohno-Machado
AMIA17
2015 The center for expanded data annotation and retrieval
abstract
The Center for Expanded Data Annotation and Retrieval is studying the creation of comprehensive and expressive metadata for biomedical datasets to facilitate data discovery, data interpretation, and data reuse. We take advantage of emerging community-based standard templates for describing different kinds of biomedical datasets, and we investigate the use of computational techniques to help investigators to assemble templates and to fill in their values. We are creating a repository of metadata from which we plan to identify metadata patterns that will drive predictive data entry when filling in metadata templates. The metadata repository not only will capture annotations specified when experimental datasets are initially created, but also will incorporate links to the published literature, including secondary analyses and possible refinements or retractions of experimental interpretations. By working initially with the Human Immunology Project Consortium and the developers of the ImmPort data repository, we are developing and evaluating an end-to-end solution to the problems of metadata authoring and management that will generalize to other data-management environments.
Mark A. Musen, Carol A. Bean, Kei-Hoi Cheung, Michel Dumontier, Kim A. Durante, Olivier Gevaert, Alejandra N. González-Beltrán, Purvesh Khatri, Steven H. Kleinstein, Martin J. O'Connor, Yannick Pouliot, Philippe Rocca-Serra, Susanna-Assunta Sansone, Jeffrey A. Wiser
J. Am. Medical Informatics Assoc.12
2014 linkedISA: semantic representation of ISA-Tab experimental metadata
abstract
BACKGROUND: Reporting and sharing experimental metadata- such as the experimental design, characteristics of the samples, and procedures applied, along with the analysis results, in a standardised manner ensures that datasets are comprehensible and, in principle, reproducible, comparable and reusable. Furthermore, sharing datasets in formats designed for consumption by humans and machines will also maximize their use. The Investigation/Study/Assay (ISA) open source metadata tracking framework facilitates standards-compliant collection, curation, visualization, storage and sharing of datasets, leveraging on other platforms to enable analysis and publication. The ISA software suite includes several components used in increasingly diverse set of life science and biomedical domains; it is underpinned by a general-purpose format, ISA-Tab, and conversions exist into formats required by public repositories. While ISA-Tab works well mainly as a human readable format, we have also implemented a linked data approach to semantically define the ISA-Tab syntax. RESULTS: We present a semantic web representation of the ISA-Tab syntax that complements ISA-Tab's syntactic interoperability with semantic interoperability. We introduce the linkedISA conversion tool from ISA-Tab to the Resource Description Framework (RDF), supporting mappings from the ISA syntax to multiple community-defined, open ontologies and capitalising on user-provided ontology annotations in the experimental metadata. We describe insights of the implementation and how annotations can be expanded driven by the metadata. We applied the conversion tool as part of Bio-GraphIIn, a web-based application supporting integration of the semantically-rich experimental descriptions. Designed in a user-friendly manner, the Bio-GraphIIn interface hides most of the complexities to the users, exposing a familiar tabular view of the experimental description to allow seamless interaction with the RDF representation, and visualising descriptors to drive the query over the semantic representation of the experimental design. In addition, we defined queries over the linkedISA RDF representation and demonstrated its use over the linkedISA conversion of datasets from Nature' Scientific Data online publication. CONCLUSIONS: Our linked data approach has allowed us to: 1) make the ISA-Tab semantics explicit and machine-processable, 2) exploit the existing ontology-based annotations in the ISA-Tab experimental descriptions, 3) augment the ISA-Tab syntax with new descriptive elements, 4) visualise and query elements related to the experimental design. Reasoning over ISA-Tab metadata and associated data will facilitate data integration and knowledge discovery.
Alejandra N. González-Beltrán, Eamonn Maguire, Susanna-Assunta Sansone, Philippe Rocca-Serra
BMC Bioinform.4
2014 The Risa R/Bioconductor package: integrative data analysis from experimental metadata and back again
abstract
BACKGROUND: The ISA-Tab format and software suite have been developed to break the silo effect induced by technology-specific formats for a variety of data types and to better support experimental metadata tracking. Experimentalists seldom use a single technique to monitor biological signals. Providing a multi-purpose, pragmatic and accessible format that abstracts away common constructs for describing Investigations, Studies and Assays, ISA is increasingly popular. To attract further interest towards the format and extend support to ensure reproducible research and reusable data, we present the Risa package, which delivers a central component to support the ISA format by enabling effortless integration with R, the popular, open source data crunching environment. RESULTS: The Risa package bridges the gap between the metadata collection and curation in an ISA-compliant way and the data analysis using the widely used statistical computing environment R. The package offers functionality for: i) parsing ISA-Tab datasets into R objects, ii) augmenting annotation with extra metadata not explicitly stated in the ISA syntax; iii) interfacing with domain specific R packages iv) suggesting potentially useful R packages available in Bioconductor for subsequent processing of the experimental data described in the ISA format; and finally v) saving back to ISA-Tab files augmented with analysis specific metadata from R. We demonstrate these features by presenting use cases for mass spectrometry data and DNA microarray data. CONCLUSIONS: The Risa package is open source (with LGPL license) and freely available through Bioconductor. By making Risa available, we aim to facilitate the task of processing experimental data, encouraging a uniform representation of experimental information and results while delivering tools for ensuring traceability and provenance tracking. SOFTWARE AVAILABILITY: The Risa package is available since Bioconductor 2.11 (version 1.0.0) and version 1.2.1 appeared in Bioconductor 2.12, both along with documentation and examples. The latest version of the code is at the development branch in Bioconductor and can also be accessed from GitHub https://github.com/ISA-tools/Risa, where the issue tracker allows users to report bugs or feature requests.
Alejandra N. González-Beltrán, Steffen Neumann, Eamonn Maguire, Susanna-Assunta Sansone, Philippe Rocca-Serra
BMC Bioinform.5
2013 OntoMaton: a Bioportal powered ontology widget for Google Spreadsheets
abstract
MOTIVATION: Data collection in spreadsheets is ubiquitous, but current solutions lack support for collaborative semantic annotation that would promote shared and interdisciplinary annotation practices, supporting geographically distributed players. RESULTS: OntoMaton is an open source solution that brings ontology lookup and tagging capabilities into a cloud-based collaborative editing environment, harnessing Google Spreadsheets and the NCBO Web services. It is a general purpose, format-agnostic tool that may serve as a component of the ISA software suite. OntoMaton can also be used to assist the ontology development process. AVAILABILITY: OntoMaton is freely available from Google widgets under the CPAL open source license; documentation and examples at: https://github.com/ISA-tools/OntoMaton.
Eamonn Maguire, Alejandra N. González-Beltrán, Patricia L. Whetzel, Susanna-Assunta Sansone, Philippe Rocca-Serra
Bioinform.5
2013 Visual Compression of Workflow Visualizations with Automated Detection of Macro Motifs
abstract
This paper is concerned with the creation of 'macros' in workflow visualization as a support tool to increase the efficiency of data curation tasks. We propose computation of candidate macros based on their usage in large collections of workflows in data repositories. We describe an efficient algorithm for extracting macro motifs from workflow graphs. We discovered that the state transition information, used to identify macro candidates, characterizes the structural pattern of the macro and can be harnessed as part of the visual design of the corresponding macro glyph. This facilitates partial automation and consistency in glyph design applicable to a large set of macro glyphs. We tested this approach against a repository of biological data holding some 9,670 workflows and found that the algorithmically generated candidate macros are in keeping with domain expert expectations.
Eamonn Maguire, Philippe Rocca-Serra, Susanna-Assunta Sansone, Jim Davies, Min Chen 0001
IEEE Trans. Vis. Comput. Graph.2
2012 graph2tab, a library to convert experimental workflow graphs into tabular formats
abstract
Abstract Motivations: Spreadsheet-like tabular formats are ever more popular in the biomedical field as a mean for experimental reporting. The problem of converting the graph of an experimental workflow into a table-based representation occurs in many such formats and is not easy to solve. Results: We describe graph2tab, a library that implements methods to realise such a conversion in a size-optimised way. Our solution is generic and can be adapted to specific cases of data exporters or data converters that need to be implemented. Availability and Implementation: The library source code and documentation are available at http://github.com/ISA-tools/graph2tab. Contact: [email protected]. Supplementary Information: A supplementary document describes the theoretical and technical details about the library implementation.
Marco Brandizi, Natalja Kurbatova, Ugis Sarkans, Philippe Rocca-Serra
Bioinform.4
2012 Translating standards into practice - One Semantic Web API for Gene Expression
Helena F. Deus, Eric Prud'hommeaux, Michael Miller 0001, Jun Zhao 0003, James Malone, Tomasz Adamusiak, Jamie P. McCusker, Sudeshna Das 0001, Philippe Rocca-Serra, Ronan Fox, M. Scott Marshall
J. Biomed. Informatics9
2012 Taxonomy-Based Glyph Design - with a Case Study on Visualizing Workflows of Biological Experiments
abstract
Glyph-based visualization can offer elegant and concise presentation of multivariate information while enhancing speed and ease in visual search experienced by users. As with icon designs, glyphs are usually created based on the designers' experience and intuition, often in a spontaneous manner. Such a process does not scale well with the requirements of applications where a large number of concepts are to be encoded using glyphs. To alleviate such limitations, we propose a new systematic process for glyph design by exploring the parallel between the hierarchy of concept categorization and the ordering of discriminative capacity of visual channels. We examine the feasibility of this approach in an application where there is a pressing need for an efficient and effective means to visualize workflows of biological experiments. By processing thousands of workflow records in a public archive of biological experiments, we demonstrate that a cost-effective glyph design can be obtained by following a process of formulating a taxonomy with the aid of computation, identifying visual channels hierarchically, and defining application-specific abstraction and metaphors.
Eamonn Maguire, Philippe Rocca-Serra, Susanna-Assunta Sansone, Jim Davies, Min Chen 0001
IEEE Trans. Vis. Comput. Graph.2
2010 ISA software suite: supporting standards-compliant experimental annotation and enabling curation at the community level
abstract
UNLABELLED: The first open source software suite for experimentalists and curators that (i) assists in the annotation and local management of experimental metadata from high-throughput studies employing one or a combination of omics and other technologies; (ii) empowers users to uptake community-defined checklists and ontologies; and (iii) facilitates submission to international public repositories. AVAILABILITY AND IMPLEMENTATION: Software, documentation, case studies and implementations at http://www.isa-tools.org.
Philippe Rocca-Serra, Marco Brandizi, Eamonn Maguire, Nataliya Sklyar, Chris F. Taylor, Kimberly Begley, Dawn Field, Stephen C. Harris, Winston Hide, Oliver Hofmann 0001, Steffen Neumann, Peter Sterk, Weida Tong, Susanna-Assunta Sansone
Bioinform.1
2009 Survey-based naming conventions for use in OBO Foundry ontology development
abstract
BACKGROUND: A wide variety of ontologies relevant to the biological and medical domains are available through the OBO Foundry portal, and their number is growing rapidly. Integration of these ontologies, while requiring considerable effort, is extremely desirable. However, heterogeneities in format and style pose serious obstacles to such integration. In particular, inconsistencies in naming conventions can impair the readability and navigability of ontology class hierarchies, and hinder their alignment and integration. While other sources of diversity are tremendously complex and challenging, agreeing a set of common naming conventions is an achievable goal, particularly if those conventions are based on lessons drawn from pooled practical experience and surveys of community opinion. RESULTS: We summarize a review of existing naming conventions and highlight certain disadvantages with respect to general applicability in the biological domain. We also present the results of a survey carried out to establish which naming conventions are currently employed by OBO Foundry ontologies and to determine what their special requirements regarding the naming of entities might be. Lastly, we propose an initial set of typographic, syntactic and semantic conventions for labelling classes in OBO Foundry ontologies. CONCLUSION: Adherence to common naming conventions is more than just a matter of aesthetics. Such conventions provide guidance to ontology creators, help developers avoid flaws and inaccuracies when editing, and especially when interlinking, ontologies. Common naming conventions will also assist consumers of ontologies to more readily understand what meanings were intended by the authors of ontologies used in annotating bodies of data.
Daniel Schober, Barry Smith 0001, Suzanna Lewis, Waclaw Kusnierczyk, Jane Lomax, Chris Mungall, Chris F. Taylor, Philippe Rocca-Serra, Susanna-Assunta Sansone
BMC Bioinform.8
2006 The MGED Ontology: a resource for semantics-based description of microarray experiments
abstract
MOTIVATION: The generation of large amounts of microarray data and the need to share these data bring challenges for both data management and annotation and highlights the need for standards. MIAME specifies the minimum information needed to describe a microarray experiment and the Microarray Gene Expression Object Model (MAGE-OM) and resulting MAGE-ML provide a mechanism to standardize data representation for data exchange, however a common terminology for data annotation is needed to support these standards. RESULTS: Here we describe the MGED Ontology (MO) developed by the Ontology Working Group of the Microarray Gene Expression Data (MGED) Society. The MO provides terms for annotating all aspects of a microarray experiment from the design of the experiment and array layout, through to the preparation of the biological sample and the protocols used to hybridize the RNA and analyze the data. The MO was developed to provide terms for annotating experiments in line with the MIAME guidelines, i.e. to provide the semantics to describe a microarray experiment according to the concepts specified in MIAME. The MO does not attempt to incorporate terms from existing ontologies, e.g. those that deal with anatomical parts or developmental stages terms, but provides a framework to reference terms in other ontologies and therefore facilitates the use of ontologies in microarray data annotation. AVAILABILITY: The MGED Ontology version.1.2.0 is available as a file in both DAML and OWL formats at http://mged.sourceforge.net/ontologies/index.php. Release notes and annotation examples are provided. The MO is also provided via the NCICB's Enterprise Vocabulary System (http://nciterms.nci.nih.gov/NCIBrowser/Dictionary.do). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Patricia L. Whetzel, Helen E. Parkinson, Helen C. Causton, Liju Fan, Jennifer Fostel, Gilberto Fragoso, Laurence Game, Mervi Heiskanen, Norman Morrison, Philippe Rocca-Serra, Susanna-Assunta Sansone, Chris F. Taylor, Joseph White, Christian J. Stoeckert Jr.
Bioinform.10
2006 The use of concept maps during knowledge elicitation in ontology development processes - the nutrigenomics use case
abstract
BACKGROUND: Incorporation of ontologies into annotations has enabled 'semantic integration' of complex data, making explicit the knowledge within a certain field. One of the major bottlenecks in developing bio-ontologies is the lack of a unified methodology. Different methodologies have been proposed for different scenarios, but there is no agreed-upon standard methodology for building ontologies. The involvement of geographically distributed domain experts, the need for domain experts to lead the design process, the application of the ontologies and the life cycles of bio-ontologies are amongst the features not considered by previously proposed methodologies. RESULTS: Here, we present a methodology for developing ontologies within the biological domain. We describe our scenario, competency questions, results and milestones for each methodological stage. We introduce the use of concept maps during knowledge acquisition phases as a feasible transition between domain expert and knowledge engineer. CONCLUSION: The contributions of this paper are the thorough description of the steps we suggest when building an ontology, example use of concept maps, consideration of applicability to the development of lower-level ontologies and application to decentralised environments. We have found that within our scenario conceptual maps played an important role in the development process.
Alexander García Castro, Philippe Rocca-Serra, Robert Stevens 0001, Chris F. Taylor, Karim Nashar, Mark A. Ragan, Susanna-Assunta Sansone
BMC Bioinform.2
2006 A simple spreadsheet-based, MIAME-supportive format for microarray data: MAGE-TAB
abstract
BACKGROUND: Sharing of microarray data within the research community has been greatly facilitated by the development of the disclosure and communication standards MIAME and MAGE-ML by the MGED Society. However, the complexity of the MAGE-ML format has made its use impractical for laboratories lacking dedicated bioinformatics support. RESULTS: We propose a simple tab-delimited, spreadsheet-based format, MAGE-TAB, which will become a part of the MAGE microarray data standard and can be used for annotating and communicating microarray data in a MIAME compliant fashion. CONCLUSION: MAGE-TAB will enable laboratories without bioinformatics experience or support to manage, exchange and submit well-annotated microarray data in a standard format using a spreadsheet. The MAGE-TAB format is self-contained, and does not require an understanding of MAGE-ML or XML.
Tim F. Rayner, Philippe Rocca-Serra, Paul T. Spellman, Helen C. Causton, Anna Farne, Ele Holloway, Rafael A. Irizarry, Junmin Liu, Donald Maier, Michael Miller 0001, Kjell Petersen, John Quackenbush, Gavin Sherlock, Christian J. Stoeckert Jr., Joseph White, Patricia L. Whetzel, Farrell Wymore, Helen E. Parkinson, Ugis Sarkans, Catherine A. Ball, Alvis Brazma
BMC Bioinform.2
2005 The ArrayExpress gene expression database: a software engineering and implementation perspective
abstract
MOTIVATION: The lack of microarray data management systems and databases is still one of the major problems faced by many life sciences laboratories. While developing the public repository for microarray data ArrayExpress we had to find novel solutions to many non-trivial software engineering problems. Our experience will be both relevant and useful for most bioinformaticians involved in developing information systems for a wide range of high-throughput technologies. RESULTS: ArrayExpress has been online since February 2002, growing exponentially to well over 10,000 hybridizations (as of September 2004). It has been demonstrated that our chosen design and implementation works for databases aimed at storage, access and sharing of high-throughput data. AVAILABILITY: The ArrayExpress database is available at http://www.ebi.ac.uk/arrayexpress/. The software is open source. CONTACT: [email protected].
Ugis Sarkans, Helen E. Parkinson, Gonzalo Garcia Lara, Ahmet Oezcimen, Anjan Sharma, Niran Abeygunawardena, Sergio Contrino, Ele Holloway, Philippe Rocca-Serra, Gaurab Mukherjee, Mohammadreza Shojatalab, Misha Kapushesky, Susanna-Assunta Sansone, Anna Farne, Tim F. Rayner, Alvis Brazma
Bioinform.9
2002 An open letter to the scientific journals
abstract
Catherine A. Ball, Gavin Sherlock, Helen Parkinson, Philippe Rocca-Sera, Catherine Brooksbank, Helen C. Causton, Duccio Cavalieri, Terry Gaasterland, Pascal Hin
Catherine A. Ball, Gavin Sherlock, Helen E. Parkinson, Philippe Rocca-Serra, Catherine Brooksbank, Helen C. Causton, Duccio Cavalieri, Terry Gaasterland, Pascal Hingamp, Frank C. P. Holstege, Martin Ringwald, Paul T. Spellman, Christian J. Stoeckert Jr., Jason E. Stewart, Ronald C. Taylor, Alvis Brazma, John Quackenbush
Bioinform.4