EDBT 2026 Demo / reviewers in the wild / expert
Susanna-Assunta Sansone
dblp:71/5798
· DBLP profile ↗
19ranked-venue papers
0as first author
0since 2021 · last 2020
0000-0001-5306-5690ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17Graphics, computer vision, multimedia, augmented reality and games · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
8 papers |
Bioinformatics and computational biology · 96% Computational science and engineering · 4% | |
| Computer graphics and multimedia
2 papers |
Visualization and visual analytics · 81% Image and video coding · 19% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 100% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
metabolomics |
0.4 | 1 | 2019 | Interoperable and scalable data analysis with microservices: applications in metabolomics · Bioinform. 2019 |
Cloud and datacenter computing
container orchestration |
0.4 | 1 | 2019 | Interoperable and scalable data analysis with microservices: applications in metabolomics · Bioinform. 2019 |
Visualization and visual analytics › visual encoding
glyph design |
0.3 | 2 | 2013 | Visual Compression of Workflow Visualizations with Automated Detection of Macro Motifs · IEEE Trans. Vis. Comput. Graph. 2013 Taxonomy-Based Glyph Design - with a Case Study on Visualizing Workflows of Biological Experiments · IEEE Trans. Vis. Comput. Graph. 2012 |
Bioinformatics and computational biology › genome annotation
ontology-based annotation |
0.2 | 1 | 2013 | OntoMaton: a Bioportal powered ontology widget for Google Spreadsheets · Bioinform. 2013 |
Visualization and visual analytics › software visualization
workflow visualization |
0.2 | 1 | 2013 | Visual Compression of Workflow Visualizations with Automated Detection of Macro Motifs · IEEE Trans. Vis. Comput. Graph. 2013 |
Visualization and visual analytics › visual encoding
glyph-based visualization |
0.1 | 1 | 2012 | Taxonomy-Based Glyph Design - with a Case Study on Visualizing Workflows of Biological Experiments · IEEE Trans. Vis. Comput. Graph. 2012 |
Image and video coding
visual communication coding |
0.1 | 1 | 2012 | Taxonomy-Based Glyph Design - with a Case Study on Visualizing Workflows of Biological Experiments · IEEE Trans. Vis. Comput. Graph. 2012 |
Bioinformatics and computational biology › ontology
ontology development |
0.1 | 2 | 2013 | The MGED Ontology: a resource for semantics-based description of microarray experiments · Bioinform. 2006 OntoMaton: a Bioportal powered ontology widget for Google Spreadsheets · Bioinform. 2013 |
Computational science and engineering › scientific data management
data standards |
0.1 | 1 | 2006 | The MGED Ontology: a resource for semantics-based description of microarray experiments · Bioinform. 2006 |
Bioinformatics and computational biology › biological database
gene expression database |
0.1 | 1 | 2005 | The ArrayExpress gene expression database: a software engineering and implementation perspective · Bioinform. 2005 |
Bioinformatics and computational biology › biological database
microarray data management |
0.1 | 1 | 2005 | The ArrayExpress gene expression database: a software engineering and implementation perspective · Bioinform. 2005 |
Methods — techniques the papers use, named apart from their topics
microservice architecture · 0.8kubernetes · 0.8docker · 0.8web scraping · 0.4search interface · 0.4data API · 0.4visual encoding · 0.3graph motif extraction · 0.3taxonomy construction · 0.3ontology lookup web services · 0.2visual channel ordering · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | TeSS: a platform for discovering life-science training opportunitiesabstractSUMMARY: Dispersed across the Internet is an abundance of disparate, disconnected training information, making it hard for researchers to find training opportunities that are relevant to them. To address this issue, we have developed a new platform-TeSS-which aggregates geographically distributed information and presents it in a central, feature-rich portal. Data are gathered automatically from content providers via bespoke scripts. These resources are cross-linked with related data and tools registries, and made available via a search interface, a data API and through widgets. AVAILABILITY AND IMPLEMENTATION: https://tess.elixir-europe.org. Niall Beard, Finn Bacall, Aleksandra Nenadic, Milo Thurston, Carole A. Goble, Susanna-Assunta Sansone, Terri K. Attwood |
Bioinform. | 6 |
| 2019 | Interoperable and scalable data analysis with microservices: applications in metabolomicsabstractMOTIVATION: Developing a robust and performant data analysis workflow that integrates all necessary components whilst still being able to scale over multiple compute nodes is a challenging task. We introduce a generic method based on the microservice architecture, where software tools are encapsulated as Docker containers that can be connected into scientific workflows and executed using the Kubernetes container orchestrator. RESULTS: We developed a Virtual Research Environment (VRE) which facilitates rapid integration of new tools and developing scalable and interoperable workflows for performing metabolomics data analysis. The environment can be launched on-demand on cloud resources and desktop computers. IT-expertise requirements on the user side are kept to a minimum, and workflows can be re-used effortlessly by any novice user. We validate our method in the field of metabolomics on two mass spectrometry, one nuclear magnetic resonance spectroscopy and one fluxomics study. We showed that the method scales dynamically with increasing availability of computational resources. We demonstrated that the method facilitates interoperability using integration of the major software suites resulting in a turn-key workflow encompassing all steps for mass-spectrometry-based metabolomics including preprocessing, statistics and identification. Microservices is a generic methodology that can serve any scientific discipline and opens up for new types of large-scale integrative science. AVAILABILITY AND IMPLEMENTATION: The PhenoMeNal consortium maintains a web portal (https://portal.phenomenal-h2020.eu) providing a GUI for launching the Virtual Research Environment. The GitHub repository https://github.com/phnmnl/ hosts the source code of all projects. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Payam Emami Khoonsari, Pablo A. Moreno, Sven Bergmann, Joachim Burman, Marco Capuccini, Matteo Carone, Marta Cascante, Pedro de Atauri, Carles Foguet, Alejandra N. González-Beltrán, Thomas Hankemeier, Kenneth Haug, Sijin He, Stephanie Herman, David Johnson 0006, Namrata Kale, Anders Larsson, Steffen Neumann, Kristian Peters, Luca Pireddu, Philippe Rocca-Serra, Pierrick Roger, Rico Rueedi, Christoph Ruttkies, Noureddin Sadawi, Reza M. Salek, Susanna-Assunta Sansone, Daniel Schober, Vitaly A. Selivanov, Etienne A. Thévenot, Michael van Vliet, Gianluigi Zanetti, Christoph Steinbeck, Kim Kultima, Ola Spjuth |
Bioinform. | 27 |
| 2018 | DataMed - an open source discovery index for finding biomedical datasetsabstractOBJECTIVE: Finding relevant datasets is important for promoting data reuse in the biomedical domain, but it is challenging given the volume and complexity of biomedical data. Here we describe the development of an open source biomedical data discovery system called DataMed, with the goal of promoting the building of additional data indexes in the biomedical domain. MATERIALS AND METHODS: DataMed, which can efficiently index and search diverse types of biomedical datasets across repositories, is developed through the National Institutes of Health-funded biomedical and healthCAre Data Discovery Index Ecosystem (bioCADDIE) consortium. It consists of 2 main components: (1) a data ingestion pipeline that collects and transforms original metadata information to a unified metadata model, called DatA Tag Suite (DATS), and (2) a search engine that finds relevant datasets based on user-entered queries. In addition to describing its architecture and techniques, we evaluated individual components within DataMed, including the accuracy of the ingestion pipeline, the prevalence of the DATS model across repositories, and the overall performance of the dataset retrieval engine. RESULTS AND CONCLUSION: Our manual review shows that the ingestion pipeline could achieve an accuracy of 90% and core elements of DATS had varied frequency across repositories. On a manually curated benchmark dataset, the DataMed search engine achieved an inferred average precision of 0.2033 and a precision at 10 (P@10, the number of relevant results in the top 10 search results) of 0.6022, by implementing advanced natural language processing and terminology services. Currently, we have made the DataMed system publically available as an open source package for the biomedical community. Anupama E. Gururaj, Ibrahim Burak Özyurt, Ruiling Liu, Ergin Soysal, Trevor Cohen, Firat Tiryaki, Yueling Li, Nansu Zong, Min Jiang 0007, Deevakar Rogith, Mandana Salimi, Hyeon-Eui Kim, Philippe Rocca-Serra, Alejandra N. González-Beltrán, Claudiu Farcas, Todd Johnson, Ronald Margolis, George Alter, Susanna-Assunta Sansone, Ian Fore, Lucila Ohno-Machado, Jeffrey S. Grethe, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 20 |
| 2018 | Data discovery with DATS: exemplar adoptions and lessons learnedabstractThe DAta Tag Suite (DATS) is a model supporting dataset description, indexing, and discovery. It is available as an annotated serialization with schema.org, a vocabulary used by major search engines, thus making the datasets discoverable on the web. DATS underlies DataMed, the National Institutes of Health Big Data to Knowledge Data Discovery Index prototype, which aims to provide a "PubMed for datasets." The experience gained while indexing a heterogeneous range of >60 repositories in DataMed helped in evaluating DATS's entities, attributes, and scope. In this work, 3 additional exemplary and diverse data sources were mapped to DATS by their representatives or experts, offering a deep scan of DATS fitness against a new set of existing data. The procedure, including feedback from users and implementers, resulted in DATS implementation guidelines and best practices, and identification of a path for evolving and optimizing the model. Finally, the work exposed additional needs when defining datasets for indexing, especially in the context of clinical and observational information. Alejandra N. González-Beltrán, John Campbell 0004, Patrick J. Dunn, Diana Guijarro, Sanda Ionescu, Hyeon-Eui Kim, Jared Lyle, Jeffrey A. Wiser, Susanna-Assunta Sansone, Philippe Rocca-Serra |
J. Am. Medical Informatics Assoc. | 9 |
| 2016 | A Scalable Dataset Indexing Infrastructure for the bioCADDIE Data Discovery System
Jeffrey S. Grethe, Ibrahim Burak Özyurt, Hua Xu 0001, Ruiling Liu, Ergin Soysal, Anupama E. Gururaj, Hyeon-Eui Kim, Trevor Cohen, Todd R. Johnson, Mandana Salimi, Saeid Pournejati, Min Jiang 0007, Claudiu Farcas, Alejandra N. González-Beltrán, Philippe Rocca-Serra, Muhamamd F. Amith, Cui Tao, Ian Fore, Ronald Margolis, George Alter, Susanna-Assunta Sansone, Lucila Ohno-Machado |
AMIA | 22 |
| 2016 | Development of DataMed, a Data Discovery Index Prototype by bioCADDIE: Laying the Groundwork for Biomedical Data Discovery
Hua Xu 0001, Jeffrey S. Grethe, Ruiling Liu, Ergin Soysal, Anupama E. Gururaj, Yueling Li, Ibrahim Burak Özyurt, Hyeon-Eui Kim, Trevor Cohen, Todd R. Johnson, Mandana Salimi, Saeid Pournejati, Min Jiang 0007, Claudiu Farcas, Alejandra N. González-Beltrán, Philippe Rocca-Serra, Muhamamd F. Amith, Cui Tao, Ian Fore, Ronald Margolis, George Alter, Susanna-Assunta Sansone, Lucila Ohno-Machado |
AMIA | 23 |
| 2015 | The center for expanded data annotation and retrievalabstractThe Center for Expanded Data Annotation and Retrieval is studying the creation of comprehensive and expressive metadata for biomedical datasets to facilitate data discovery, data interpretation, and data reuse. We take advantage of emerging community-based standard templates for describing different kinds of biomedical datasets, and we investigate the use of computational techniques to help investigators to assemble templates and to fill in their values. We are creating a repository of metadata from which we plan to identify metadata patterns that will drive predictive data entry when filling in metadata templates. The metadata repository not only will capture annotations specified when experimental datasets are initially created, but also will incorporate links to the published literature, including secondary analyses and possible refinements or retractions of experimental interpretations. By working initially with the Human Immunology Project Consortium and the developers of the ImmPort data repository, we are developing and evaluating an end-to-end solution to the problems of metadata authoring and management that will generalize to other data-management environments. Mark A. Musen, Carol A. Bean, Kei-Hoi Cheung, Michel Dumontier, Kim A. Durante, Olivier Gevaert, Alejandra N. González-Beltrán, Purvesh Khatri, Steven H. Kleinstein, Martin J. O'Connor, Yannick Pouliot, Philippe Rocca-Serra, Susanna-Assunta Sansone, Jeffrey A. Wiser |
J. Am. Medical Informatics Assoc. | 13 |
| 2014 | linkedISA: semantic representation of ISA-Tab experimental metadataabstractBACKGROUND: Reporting and sharing experimental metadata- such as the experimental design, characteristics of the samples, and procedures applied, along with the analysis results, in a standardised manner ensures that datasets are comprehensible and, in principle, reproducible, comparable and reusable. Furthermore, sharing datasets in formats designed for consumption by humans and machines will also maximize their use. The Investigation/Study/Assay (ISA) open source metadata tracking framework facilitates standards-compliant collection, curation, visualization, storage and sharing of datasets, leveraging on other platforms to enable analysis and publication. The ISA software suite includes several components used in increasingly diverse set of life science and biomedical domains; it is underpinned by a general-purpose format, ISA-Tab, and conversions exist into formats required by public repositories. While ISA-Tab works well mainly as a human readable format, we have also implemented a linked data approach to semantically define the ISA-Tab syntax. RESULTS: We present a semantic web representation of the ISA-Tab syntax that complements ISA-Tab's syntactic interoperability with semantic interoperability. We introduce the linkedISA conversion tool from ISA-Tab to the Resource Description Framework (RDF), supporting mappings from the ISA syntax to multiple community-defined, open ontologies and capitalising on user-provided ontology annotations in the experimental metadata. We describe insights of the implementation and how annotations can be expanded driven by the metadata. We applied the conversion tool as part of Bio-GraphIIn, a web-based application supporting integration of the semantically-rich experimental descriptions. Designed in a user-friendly manner, the Bio-GraphIIn interface hides most of the complexities to the users, exposing a familiar tabular view of the experimental description to allow seamless interaction with the RDF representation, and visualising descriptors to drive the query over the semantic representation of the experimental design. In addition, we defined queries over the linkedISA RDF representation and demonstrated its use over the linkedISA conversion of datasets from Nature' Scientific Data online publication. CONCLUSIONS: Our linked data approach has allowed us to: 1) make the ISA-Tab semantics explicit and machine-processable, 2) exploit the existing ontology-based annotations in the ISA-Tab experimental descriptions, 3) augment the ISA-Tab syntax with new descriptive elements, 4) visualise and query elements related to the experimental design. Reasoning over ISA-Tab metadata and associated data will facilitate data integration and knowledge discovery. Alejandra N. González-Beltrán, Eamonn Maguire, Susanna-Assunta Sansone, Philippe Rocca-Serra |
BMC Bioinform. | 3 |
| 2014 | The Risa R/Bioconductor package: integrative data analysis from experimental metadata and back againabstractBACKGROUND: The ISA-Tab format and software suite have been developed to break the silo effect induced by technology-specific formats for a variety of data types and to better support experimental metadata tracking. Experimentalists seldom use a single technique to monitor biological signals. Providing a multi-purpose, pragmatic and accessible format that abstracts away common constructs for describing Investigations, Studies and Assays, ISA is increasingly popular. To attract further interest towards the format and extend support to ensure reproducible research and reusable data, we present the Risa package, which delivers a central component to support the ISA format by enabling effortless integration with R, the popular, open source data crunching environment. RESULTS: The Risa package bridges the gap between the metadata collection and curation in an ISA-compliant way and the data analysis using the widely used statistical computing environment R. The package offers functionality for: i) parsing ISA-Tab datasets into R objects, ii) augmenting annotation with extra metadata not explicitly stated in the ISA syntax; iii) interfacing with domain specific R packages iv) suggesting potentially useful R packages available in Bioconductor for subsequent processing of the experimental data described in the ISA format; and finally v) saving back to ISA-Tab files augmented with analysis specific metadata from R. We demonstrate these features by presenting use cases for mass spectrometry data and DNA microarray data. CONCLUSIONS: The Risa package is open source (with LGPL license) and freely available through Bioconductor. By making Risa available, we aim to facilitate the task of processing experimental data, encouraging a uniform representation of experimental information and results while delivering tools for ensuring traceability and provenance tracking. SOFTWARE AVAILABILITY: The Risa package is available since Bioconductor 2.11 (version 1.0.0) and version 1.2.1 appeared in Bioconductor 2.12, both along with documentation and examples. The latest version of the code is at the development branch in Bioconductor and can also be accessed from GitHub https://github.com/ISA-tools/Risa, where the issue tracker allows users to report bugs or feature requests. Alejandra N. González-Beltrán, Steffen Neumann, Eamonn Maguire, Susanna-Assunta Sansone, Philippe Rocca-Serra |
BMC Bioinform. | 4 |
| 2014 | A sea of standards for omics data: sink or swim?abstractIn the era of Big Data, omic-scale technologies, and increasing calls for data sharing, it is generally agreed that the use of community-developed, open data standards is critical. Far less agreed upon is exactly which data standards should be used, the criteria by which one should choose a standard, or even what constitutes a data standard. It is impossible simply to choose a domain and have it naturally follow which data standards should be used in all cases. The 'right' standards to use is often dependent on the use case scenarios for a given project. Potential downstream applications for the data, however, may not always be apparent at the time the data are generated. Similarly, technology evolves, adding further complexity. Would-be standards adopters must strike a balance between planning for the future and minimizing the burden of compliance. Better tools and resources are required to help guide this balancing act. Jessica D. Tenenbaum, Susanna-Assunta Sansone, Melissa A. Haendel |
J. Am. Medical Informatics Assoc. | 2 |
| 2013 | OntoMaton: a Bioportal powered ontology widget for Google SpreadsheetsabstractMOTIVATION: Data collection in spreadsheets is ubiquitous, but current solutions lack support for collaborative semantic annotation that would promote shared and interdisciplinary annotation practices, supporting geographically distributed players. RESULTS: OntoMaton is an open source solution that brings ontology lookup and tagging capabilities into a cloud-based collaborative editing environment, harnessing Google Spreadsheets and the NCBO Web services. It is a general purpose, format-agnostic tool that may serve as a component of the ISA software suite. OntoMaton can also be used to assist the ontology development process. AVAILABILITY: OntoMaton is freely available from Google widgets under the CPAL open source license; documentation and examples at: https://github.com/ISA-tools/OntoMaton. Eamonn Maguire, Alejandra N. González-Beltrán, Patricia L. Whetzel, Susanna-Assunta Sansone, Philippe Rocca-Serra |
Bioinform. | 4 |
| 2013 | Visual Compression of Workflow Visualizations with Automated Detection of Macro MotifsabstractThis paper is concerned with the creation of 'macros' in workflow visualization as a support tool to increase the efficiency of data curation tasks. We propose computation of candidate macros based on their usage in large collections of workflows in data repositories. We describe an efficient algorithm for extracting macro motifs from workflow graphs. We discovered that the state transition information, used to identify macro candidates, characterizes the structural pattern of the macro and can be harnessed as part of the visual design of the corresponding macro glyph. This facilitates partial automation and consistency in glyph design applicable to a large set of macro glyphs. We tested this approach against a repository of biological data holding some 9,670 workflows and found that the algorithmically generated candidate macros are in keeping with domain expert expectations. Eamonn Maguire, Philippe Rocca-Serra, Susanna-Assunta Sansone, Jim Davies, Min Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2012 | Taxonomy-Based Glyph Design - with a Case Study on Visualizing Workflows of Biological ExperimentsabstractGlyph-based visualization can offer elegant and concise presentation of multivariate information while enhancing speed and ease in visual search experienced by users. As with icon designs, glyphs are usually created based on the designers' experience and intuition, often in a spontaneous manner. Such a process does not scale well with the requirements of applications where a large number of concepts are to be encoded using glyphs. To alleviate such limitations, we propose a new systematic process for glyph design by exploring the parallel between the hierarchy of concept categorization and the ordering of discriminative capacity of visual channels. We examine the feasibility of this approach in an application where there is a pressing need for an efficient and effective means to visualize workflows of biological experiments. By processing thousands of workflow records in a public archive of biological experiments, we demonstrate that a cost-effective glyph design can be obtained by following a process of formulating a taxonomy with the aid of computation, identifying visual channels hierarchically, and defining application-specific abstraction and metaphors. Eamonn Maguire, Philippe Rocca-Serra, Susanna-Assunta Sansone, Jim Davies, Min Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2010 | ISA software suite: supporting standards-compliant experimental annotation and enabling curation at the community levelabstractUNLABELLED: The first open source software suite for experimentalists and curators that (i) assists in the annotation and local management of experimental metadata from high-throughput studies employing one or a combination of omics and other technologies; (ii) empowers users to uptake community-defined checklists and ontologies; and (iii) facilitates submission to international public repositories. AVAILABILITY AND IMPLEMENTATION: Software, documentation, case studies and implementations at http://www.isa-tools.org. Philippe Rocca-Serra, Marco Brandizi, Eamonn Maguire, Nataliya Sklyar, Chris F. Taylor, Kimberly Begley, Dawn Field, Stephen C. Harris, Winston Hide, Oliver Hofmann 0001, Steffen Neumann, Peter Sterk, Weida Tong, Susanna-Assunta Sansone |
Bioinform. | 14 |
| 2009 | Survey-based naming conventions for use in OBO Foundry ontology developmentabstractBACKGROUND: A wide variety of ontologies relevant to the biological and medical domains are available through the OBO Foundry portal, and their number is growing rapidly. Integration of these ontologies, while requiring considerable effort, is extremely desirable. However, heterogeneities in format and style pose serious obstacles to such integration. In particular, inconsistencies in naming conventions can impair the readability and navigability of ontology class hierarchies, and hinder their alignment and integration. While other sources of diversity are tremendously complex and challenging, agreeing a set of common naming conventions is an achievable goal, particularly if those conventions are based on lessons drawn from pooled practical experience and surveys of community opinion. RESULTS: We summarize a review of existing naming conventions and highlight certain disadvantages with respect to general applicability in the biological domain. We also present the results of a survey carried out to establish which naming conventions are currently employed by OBO Foundry ontologies and to determine what their special requirements regarding the naming of entities might be. Lastly, we propose an initial set of typographic, syntactic and semantic conventions for labelling classes in OBO Foundry ontologies. CONCLUSION: Adherence to common naming conventions is more than just a matter of aesthetics. Such conventions provide guidance to ontology creators, help developers avoid flaws and inaccuracies when editing, and especially when interlinking, ontologies. Common naming conventions will also assist consumers of ontologies to more readily understand what meanings were intended by the authors of ontologies used in annotating bodies of data. Daniel Schober, Barry Smith 0001, Suzanna Lewis, Waclaw Kusnierczyk, Jane Lomax, Chris Mungall, Chris F. Taylor, Philippe Rocca-Serra, Susanna-Assunta Sansone |
BMC Bioinform. | 9 |
| 2008 | Facilitating the development of controlled vocabularies for metabolomics technologies with text miningabstractBACKGROUND: Many bioinformatics applications rely on controlled vocabularies or ontologies to consistently interpret and seamlessly integrate information scattered across public resources. Experimental data sets from metabolomics studies need to be integrated with one another, but also with data produced by other types of omics studies in the spirit of systems biology, hence the pressing need for vocabularies and ontologies in metabolomics. However, it is time-consuming and non trivial to construct these resources manually. RESULTS: We describe a methodology for rapid development of controlled vocabularies, a study originally motivated by the needs for vocabularies describing metabolomics technologies. We present case studies involving two controlled vocabularies (for nuclear magnetic resonance spectroscopy and gas chromatography) whose development is currently underway as part of the Metabolomics Standards Initiative. The initial vocabularies were compiled manually, providing a total of 243 and 152 terms. A total of 5,699 and 2,612 new terms were acquired automatically from the literature. The analysis of the results showed that full-text articles (especially the Materials and Methods sections) are the major source of technology-specific terms as opposed to paper abstracts. CONCLUSIONS: We suggest a text mining method for efficient corpus-based term acquisition as a way of rapidly expanding a set of controlled vocabularies with the terms used in the scientific literature. We adopted an integrative approach, combining relatively generic software and data resources for time- and cost-effective development of a text mining tool for expansion of controlled vocabularies across various domains, as a practical alternative to both manual term collection and tailor-made named entity recognition methods. Irena Spasic, Daniel Schober, Susanna-Assunta Sansone, Dietrich Rebholz-Schuhmann, Douglas B. Kell, Norman W. Paton |
BMC Bioinform. | 3 |
| 2006 | The MGED Ontology: a resource for semantics-based description of microarray experimentsabstractMOTIVATION: The generation of large amounts of microarray data and the need to share these data bring challenges for both data management and annotation and highlights the need for standards. MIAME specifies the minimum information needed to describe a microarray experiment and the Microarray Gene Expression Object Model (MAGE-OM) and resulting MAGE-ML provide a mechanism to standardize data representation for data exchange, however a common terminology for data annotation is needed to support these standards. RESULTS: Here we describe the MGED Ontology (MO) developed by the Ontology Working Group of the Microarray Gene Expression Data (MGED) Society. The MO provides terms for annotating all aspects of a microarray experiment from the design of the experiment and array layout, through to the preparation of the biological sample and the protocols used to hybridize the RNA and analyze the data. The MO was developed to provide terms for annotating experiments in line with the MIAME guidelines, i.e. to provide the semantics to describe a microarray experiment according to the concepts specified in MIAME. The MO does not attempt to incorporate terms from existing ontologies, e.g. those that deal with anatomical parts or developmental stages terms, but provides a framework to reference terms in other ontologies and therefore facilitates the use of ontologies in microarray data annotation. AVAILABILITY: The MGED Ontology version.1.2.0 is available as a file in both DAML and OWL formats at http://mged.sourceforge.net/ontologies/index.php. Release notes and annotation examples are provided. The MO is also provided via the NCICB's Enterprise Vocabulary System (http://nciterms.nci.nih.gov/NCIBrowser/Dictionary.do). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Patricia L. Whetzel, Helen E. Parkinson, Helen C. Causton, Liju Fan, Jennifer Fostel, Gilberto Fragoso, Laurence Game, Mervi Heiskanen, Norman Morrison, Philippe Rocca-Serra, Susanna-Assunta Sansone, Chris F. Taylor, Joseph White, Christian J. Stoeckert Jr. |
Bioinform. | 11 |
| 2006 | The use of concept maps during knowledge elicitation in ontology development processes - the nutrigenomics use caseabstractBACKGROUND: Incorporation of ontologies into annotations has enabled 'semantic integration' of complex data, making explicit the knowledge within a certain field. One of the major bottlenecks in developing bio-ontologies is the lack of a unified methodology. Different methodologies have been proposed for different scenarios, but there is no agreed-upon standard methodology for building ontologies. The involvement of geographically distributed domain experts, the need for domain experts to lead the design process, the application of the ontologies and the life cycles of bio-ontologies are amongst the features not considered by previously proposed methodologies. RESULTS: Here, we present a methodology for developing ontologies within the biological domain. We describe our scenario, competency questions, results and milestones for each methodological stage. We introduce the use of concept maps during knowledge acquisition phases as a feasible transition between domain expert and knowledge engineer. CONCLUSION: The contributions of this paper are the thorough description of the steps we suggest when building an ontology, example use of concept maps, consideration of applicability to the development of lower-level ontologies and application to decentralised environments. We have found that within our scenario conceptual maps played an important role in the development process. Alexander García Castro, Philippe Rocca-Serra, Robert Stevens 0001, Chris F. Taylor, Karim Nashar, Mark A. Ragan, Susanna-Assunta Sansone |
BMC Bioinform. | 7 |
| 2005 | The ArrayExpress gene expression database: a software engineering and implementation perspectiveabstractMOTIVATION: The lack of microarray data management systems and databases is still one of the major problems faced by many life sciences laboratories. While developing the public repository for microarray data ArrayExpress we had to find novel solutions to many non-trivial software engineering problems. Our experience will be both relevant and useful for most bioinformaticians involved in developing information systems for a wide range of high-throughput technologies. RESULTS: ArrayExpress has been online since February 2002, growing exponentially to well over 10,000 hybridizations (as of September 2004). It has been demonstrated that our chosen design and implementation works for databases aimed at storage, access and sharing of high-throughput data. AVAILABILITY: The ArrayExpress database is available at http://www.ebi.ac.uk/arrayexpress/. The software is open source. CONTACT: [email protected]. Ugis Sarkans, Helen E. Parkinson, Gonzalo Garcia Lara, Ahmet Oezcimen, Anjan Sharma, Niran Abeygunawardena, Sergio Contrino, Ele Holloway, Philippe Rocca-Serra, Gaurab Mukherjee, Mohammadreza Shojatalab, Misha Kapushesky, Susanna-Assunta Sansone, Anna Farne, Tim F. Rayner, Alvis Brazma |
Bioinform. | 13 |