Chris T. A. Evelo

dblp:74/5894 · DBLP profile ↗
← Back
19ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0002-5301-3142ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 5 since 2021Databases, data management, data science and information retrieval · 2Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 Response to Letter to Editor by A. Derbalah et al.: the role of automation in enhancing reproducibility and interoperability of PBPK models
abstract
Dear colleagues, We thank you for your comments on our manuscript [1] and for raising the important question of automation [2]. The publications by Sepp et al. (2019) [3] and Liu et al. (2024) [4] focus on one type of PBPK models, denominated ‘for biologics’. In these models, each compartment (tissue/organ) is divided into vascular, endothelial endosomal and interstitial (sub-)compartments, and lymphatic flows are also represented [3, 4]. In these examples, the pharmacokinetics of proteins are represented. Notably, the protein distribution within tissues includes processes of passive transport across two types of pores via diffusion or fluid convection, processes of pinocytosis, binding, recycling and degradation [4]. This type of PBPK models is complex and their authors successfully used specific tools for automated code generation. In particular, Liu et al. (2024) used ‘mathematical sets’ available in the Ubiquity package, which facilitates assembling model components in R-language [4]. Sepp et al. (2019) used the MATLAB script PBPKassembler.m for automated code generation, where files in Simbiology (MATLAB) and Excel formats were combined [3]. Briefly, automation can help build models, which is highly valuable notably for complex models. Having a robust, thoroughly validated platform for automated code generation will indeed contribute to accessibility and reproducibility. Also, automation of model building seems to be a natural way to address the case of large models, in which manual scripting will be prone to introducing errors. Smaller models or those with fewer equations do not necessarily require automation. From our experience, writing the script manually provides a deep understanding of the underlying model mechanics. Manually written scripts also allow to give a clear view of how equations and parameters are applied, which we consider especially valuable for teaching and knowledge transfer. To conclude, as observed in other niches of systems biology, automation brings significant benefits to the modelling field, e.g. there is a plethora of tools that reconstruct and benchmark genome-scale metabolic models. As the correspondence authors argue, there is more scope for automation in the current practices of PBPK modelling. To facilitate this, we have opted to begin the journey of standardization via open collaboration through ELIXIR. We are looking forward to continuing to engage the research community towards automatic construction, validation and deposition of PBPK models. Thank you. Best regards, Elena Domínguez-Romero, Stanislav Mazurenko, Martin Scheringer, Vítor Martins dos Santos, Chris Evelo, Mihail Anton, John M. Hancock, Anže Županič, and Maria Suarez-Diez. None declared.
Elena Domínguez Romero, Stanislav Mazurenko, Martin Scheringer, Vítor A. P. Martins dos Santos, Chris T. A. Evelo, Mihail Anton, John M. Hancock, Anze Zupanic, María Suárez-Diez
Briefings Bioinform.5
2024 Making PBPK models more reproducible in practice
abstract
Systems biology aims to understand living organisms through mathematically modeling their behaviors at different organizational levels, ranging from molecules to populations. Modeling involves several steps, from determining the model purpose to developing the mathematical model, implementing it computationally, simulating the model's behavior, evaluating, and refining the model. Importantly, model simulation results must be reproducible, ensuring that other researchers can obtain the same results after writing the code de novo and/or using different software tools. Guidelines to increase model reproducibility have been published. However, reproducibility remains a major challenge in this field. In this paper, we tackle this challenge for physiologically-based pharmacokinetic (PBPK) models, which represent the pharmacokinetics of chemicals following exposure in humans or animals. We summarize recommendations for PBPK model reporting that should apply during model development and implementation, in order to ensure model reproducibility and comprehensibility. We make a proposal aiming to harmonize abbreviations used in PBPK models. To illustrate these recommendations, we present an original and reproducible PBPK model code in MATLAB, alongside an example of MATLAB code converted to Systems Biology Markup Language format using MOCCASIN. As directions for future improvement, more tools to convert computational PBPK models from different software platforms into standard formats would increase the interoperability of these models. The application of other systems biology standards to PBPK models is encouraged. This work is the result of an interdisciplinary collaboration involving the ELIXIR systems biology community. More interdisciplinary collaborations like this would facilitate further harmonization and application of good modeling practices in different systems biology fields.
Elena Domínguez Romero, Stanislav Mazurenko, Martin Scheringer, Vítor A. P. Martins dos Santos, Chris T. A. Evelo, Mihail Anton, John M. Hancock, Anze Zupanic, María Suárez-Diez
Briefings Bioinform.5
2021 Ten simple rules to make your publication look better
Friederike Ehrhart, Chris T. A. Evelo
PLoS Comput. Biol.2
2021 Ten simple rules for creating reusable pathway models for computational analysis and visualization
abstract
Pathway models are an effective way to capture and share our current understanding of biological processes.A pathway model is defined here as a set of interactions among biological entities (e.g., proteins and metabolites) relevant to a particular context, curated and organized to illustrate a particular process.Properly modeled pathways can be used in the analysis and visualization of diverse types of omics and other biomedical data [1,2].The modeling process involves taking our knowledge about biological pathways-however messy and incompleteand encoding it in standardized data formats that can be shared, reused, and synthesized with other knowledge in accordance with the Findable, Accessible, Interoperable, and Reusable (FAIR) principles [3].The rules presented here serve as an introduction and guide to the pathway modeling process, leveraging freely available tools and resources.Biological pathway information is often conveyed as published figures, and rules for better figures, in general, are also relevant when producing pathway models [4].However, pathway models are more than just figures.In addition to providing an intuitive depiction of a biological process that is easy to understand for humans, they also provide relevant annotations and metadata to be processed by computers.Similarly, rules for network visualizations can be applied to pathway models [5], but the distinct context, layout, and usage of pathway models necessitate specific guidelines and rules.Pathway models are described with specific languages, such as the Systems Biology Graphical Notation (SBGN) [6], the Systems Biology Markup Language (SBML) [7], the Biological Pathway Exchange (BioPAX) format [8], the Graphical Pathway Markup Language (GPML) [9], and many more.Here, we describe a set of rules for constructing pathway models to optimize their use both as graphical representations for human consumption and as FAIR resources for computational analysis.The rules range from reusability and dissemination (Rules 1, 9, and 10), intuitive visual concepts (Rules 2, 7, and 8), to enabling computational analysis (Rules 3 to 6).We hope that these rules provide a simple framework for pathway model curators who want to create (re)usable resources for the scientific community.
Kristina Hanspers, Martina Kutmon, Susan L. Coort, Daniela Digles, Lauren J. Dupuis, Friederike Ehrhart, Finterly Hu, Elisson N. Lopes, Marvin Martens, Nhung Pham, Woosub Shin, Denise N. Slenter, Andra Waagmeester, Egon L. Willighagen, Laurent A. Winckers, Chris T. A. Evelo, Alexander R. Pico
PLoS Comput. Biol.16
2021 Comparison of metabolic states using genome-scale metabolic models
abstract
Genome-scale metabolic models (GEMs) are comprehensive knowledge bases of cellular metabolism and serve as mathematical tools for studying biological phenotypes and metabolic states or conditions in various organisms and cell types. Given the sheer size and complexity of human metabolism, selecting parameters for existing analysis methods such as metabolic objective functions and model constraints is not straightforward in human GEMs. In particular, comparing several conditions in large GEMs to identify condition- or disease-specific metabolic features is challenging. In this study, we showcase a scalable, model-driven approach for an in-depth investigation and comparison of metabolic states in large GEMs which enables identifying the underlying functional differences. Using a combination of flux space sampling and network analysis, our approach enables extraction and visualisation of metabolically distinct network modules. Importantly, it does not rely on known or assumed objective functions. We apply this novel approach to extract the biochemical differences in adipocytes arising due to unlimited vs blocked uptake of branched-chain amino acids (BCAAs, considered as biomarkers in obesity) using a human adipocyte GEM (iAdipocytes1809). The biological significance of our approach is corroborated by literature reports confirming our identified metabolic processes (TCA cycle and Fatty acid metabolism) to be functionally related to BCAA metabolism. Additionally, our analysis predicts a specific altered uptake and secretion profile indicating a compensation for the unavailability of BCAAs. Taken together, our approach facilitates determining functional differences between any metabolic conditions of interest by offering a versatile platform for analysing and comparing flux spaces of large metabolic networks.
Chaitra Sarathy, Marian Breuer, Martina Kutmon, Michiel E. Adriaens, Chris T. A. Evelo, Ilja C. W. Arts
PLoS Comput. Biol.5
2018 Nanopublications: A Growing Resource of Provenance-Centric Scientific Linked Data
abstract
Nanopublications are a Linked Data format for scholarly data publishing that has received considerable uptake in the last few years. In contrast to the common Linked Data publishing practice, nanopublications work at the granular level of atomic information snippets and provide a consistent container format to attach provenance and metadata at this atomic level. While the nanopublications format is domain-independent, the datasets that have become available in this format are mostly from Life Science domains, including data about diseases, genes, proteins, drugs, biological pathways, and biotic interactions. More than 10 million such nanopublications have been published, which now form a valuable resource for studies on the domain level of the given Life Science domains as well as on the more technical levels of provenance modeling and heterogeneous Linked Data. We provide here an overview of this combined nanopublication dataset, show the results of some overarching analyses, and describe how it can be accessed and queried.
Tobias Kuhn, Albert Meroño-Peñuela, Alexander Malic, Jorrit H. Poelen, Allen H. Hurlbert, Emilio Centeno, Laura Inés Furlong, Núria Queralt-Rosinach, Christine Chichester, Juan M. Banda, Egon L. Willighagen, Friederike Ehrhart, Chris T. A. Evelo, Tareq B. Malas, Michel Dumontier
eScience13
2017 Reliable Granular References to Changing Linked Data
Tobias Kuhn, Egon L. Willighagen, Chris T. A. Evelo, Núria Queralt-Rosinach, Emilio Centeno, Laura Inés Furlong
ISWC (1)3
2016 The systems biology format converter
abstract
BACKGROUND: Interoperability between formats is a recurring problem in systems biology research. Many tools have been developed to convert computational models from one format to another. However, they have been developed independently, resulting in redundancy of efforts and lack of synergy. RESULTS: Here we present the System Biology Format Converter (SBFC), which provide a generic framework to potentially convert any format into another. The framework currently includes several converters translating between the following formats: SBML, BioPAX, SBGN-ML, Matlab, Octave, XPP, GPML, Dot, MDL and APM. This software is written in Java and can be used as a standalone executable or web service. CONCLUSIONS: The SBFC framework is an evolving software project. Existing converters can be used and improved, and new converters can be easily added, making SBFC useful to both modellers and developers. The source code and documentation of the framework are freely available from the project web site.
Nicolas Rodriguez 0001, Jean-Baptiste Pettit, Piero Dalle Pezze, Lu Li 0002, Arnaud Henry, Martijn P. van Iersel, Gaël Jalowicki, Martina Kutmon, Kedar Nath Natarajan, David Tolnay, Melanie I. Stefan, Chris T. A. Evelo, Nicolas Le Novère
BMC Bioinform.12
2016 Reactome from a WikiPathways Perspective
abstract
Reactome and WikiPathways are two of the most popular freely available databases for biological pathways. Reactome pathways are centrally curated with periodic input from selected domain experts. WikiPathways is a community-based platform where pathways are created and continually curated by any interested party. The nascent collaboration between WikiPathways and Reactome illustrates the mutual benefits of combining these two approaches. We created a format converter that converts Reactome pathways to the GPML format used in WikiPathways. In addition, we developed the ComplexViz plugin for PathVisio which simplifies looking up complex components. The plugin can also score the complexes on a pathway based on a user defined criterion. This score can then be visualized on the complex nodes using the visualization options provided by the plugin. Using the merged collection of curated and converted Reactome pathways, we demonstrate improved pathway coverage of relevant biological processes for the analysis of a previously described polycystic ovary syndrome gene expression dataset. Additionally, this conversion allows researchers to visualize their data on Reactome pathways using PathVisio's advanced data visualization functionalities. WikiPathways benefits from the dedicated focus and attention provided to the content converted from Reactome and the wealth of semantic information about interactions. Reactome in turn benefits from the continuous community curation available on WikiPathways. The research community at large benefits from the availability of a larger set of pathways for analysis in PathVisio and Cytoscape. The pathway statistics results obtained from PathVisio are significantly better when using a larger set of candidate pathways for analysis. The conversion serves as a general model for integration of multiple pathway resources developed using different approaches.
Anwesha Bohler, Guanming Wu, Martina Kutmon, Leontius Adhika Pradhana, Susan L. Coort, Kristina Hanspers, Robin Haw, Alexander R. Pico, Chris T. A. Evelo
PLoS Comput. Biol.9
2016 Using the Semantic Web for Rapid Integration of WikiPathways with Other Biological Online Data Resources
abstract
The diversity of online resources storing biological data in different formats provides a challenge for bioinformaticians to integrate and analyse their biological data. The semantic web provides a standard to facilitate knowledge integration using statements built as triples describing a relation between two objects. WikiPathways, an online collaborative pathway resource, is now available in the semantic web through a SPARQL endpoint at http://sparql.wikipathways.org. Having biological pathways in the semantic web allows rapid integration with data from other resources that contain information about elements present in pathways using SPARQL queries. In order to convert WikiPathways content into meaningful triples we developed two new vocabularies that capture the graphical representation and the pathway logic, respectively. Each gene, protein, and metabolite in a given pathway is defined with a standard set of identifiers to support linking to several other biological resources in the semantic web. WikiPathways triples were loaded into the Open PHACTS discovery platform and are available through its Web API (https://dev.openphacts.org/docs) to be used in various tools for drug development. We combined various semantic web resources with the newly converted WikiPathways content using a variety of SPARQL query types and third-party resources, such as the Open PHACTS API. The ability to use pathway information to form new links across diverse biological data highlights the utility of integrating WikiPathways in the semantic web.
Andra Waagmeester, Martina Kutmon, Anders Riutta, Ryan A. Miller, Egon L. Willighagen, Chris T. A. Evelo, Alexander R. Pico
PLoS Comput. Biol.6
2015 diXa: a data infrastructure for chemical safety assessment
abstract
MOTIVATION: The field of toxicogenomics (the application of '-omics' technologies to risk assessment of compound toxicities) has expanded in the last decade, partly driven by new legislation, aimed at reducing animal testing in chemical risk assessment but mainly as a result of a paradigm change in toxicology towards the use and integration of genome wide data. Many research groups worldwide have generated large amounts of such toxicogenomics data. However, there is no centralized repository for archiving and making these data and associated tools for their analysis easily available. RESULTS: The Data Infrastructure for Chemical Safety Assessment (diXa) is a robust and sustainable infrastructure storing toxicogenomics data. A central data warehouse is connected to a portal with links to chemical information and molecular and phenotype data. diXa is publicly available through a user-friendly web interface. New data can be readily deposited into diXa using guidelines and templates available online. Analysis descriptions and tools for interrogating the data are available via the diXa portal. AVAILABILITY AND IMPLEMENTATION: http://www.dixa-fp7.eu CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Diana M. Hendrickx, Hugo J. W. L. Aerts, Florian Caiment, Dominic Clark, Timothy M. D. Ebbels, Chris T. A. Evelo, Hans Gmuender, Dennie G. A. J. Hebels, Ralf Herwig, Jürgen Hescheler, Danyel Jennen, Marlon J. A. Jetten, Stathis Kanterakis, Hector C. Keun, Vera Matser, John P. Overington, Ekaterina Pilicheva, Ugis Sarkans, Marcelo P. Segura-Lepe, Isaia Sotiriadou, Timo Wittenberger, Clemens Wittwehr, Antonella Zanzi, Jos Kleinjans
Bioinform.6
2015 Automatically visualise and analyse data on pathways using PathVisioRPC from any programming environment
abstract
BACKGROUND: Biological pathways are descriptive diagrams of biological processes widely used for functional analysis of differentially expressed genes or proteins. Primary data analysis, such as quality control, normalisation, and statistical analysis, is often performed in scripting languages like R, Perl, and Python. Subsequent pathway analysis is usually performed using dedicated external applications. Workflows involving manual use of multiple environments are time consuming and error prone. Therefore, tools are needed that enable pathway analysis directly within the same scripting languages used for primary data analyses. Existing tools have limited capability in terms of available pathway content, pathway editing and visualisation options, and export file formats. Consequently, making the full-fledged pathway analysis tool PathVisio available from various scripting languages will benefit researchers. RESULTS: We developed PathVisioRPC, an XMLRPC interface for the pathway analysis software PathVisio. PathVisioRPC enables creating and editing biological pathways, visualising data on pathways, performing pathway statistics, and exporting results in several image formats in multiple programming environments. We demonstrate PathVisioRPC functionalities using examples in Python. Subsequently, we analyse a publicly available NCBI GEO gene expression dataset studying tumour bearing mice treated with cyclophosphamide in R. The R scripts demonstrate how calls to existing R packages for data processing and calls to PathVisioRPC can directly work together. To further support R users, we have created RPathVisio simplifying the use of PathVisioRPC in this environment. We have also created a pathway module for the microarray data analysis portal ArrayAnalysis.org that calls the PathVisioRPC interface to perform pathway analysis. This module allows users to use PathVisio functionality online without having to download and install the software and exemplifies how the PathVisioRPC interface can be used by data analysis pipelines for functional analysis of processed genomics data. CONCLUSIONS: PathVisioRPC enables data visualisation and pathway analysis directly from within various analytical environments used for preliminary analyses. It supports the use of existing pathways from WikiPathways or pathways created using the RPC itself. It also enables automation of tasks performed using PathVisio, making it useful to PathVisio users performing repeated visualisation and analysis tasks. PathVisioRPC is freely available for academic and commercial use at http://projects.bigcat.unimaas.nl/pathvisiorpc.
Anwesha Bohler, Lars M. T. Eijssen, Martijn P. van Iersel, Christ Leemans, Egon L. Willighagen, Martina Kutmon, Magali Jaillard, Chris T. A. Evelo
BMC Bioinform.8
2015 PathVisio 3: An Extendable Pathway Analysis Toolbox
abstract
PathVisio is a commonly used pathway editor, visualization and analysis software. Biological pathways have been used by biologists for many years to describe the detailed steps in biological processes. Those powerful, visual representations help researchers to better understand, share and discuss knowledge. Since the first publication of PathVisio in 2008, the original paper was cited more than 170 times and PathVisio was used in many different biological studies. As an online editor PathVisio is also integrated in the community curated pathway database WikiPathways. Here we present the third version of PathVisio with the newest additions and improvements of the application. The core features of PathVisio are pathway drawing, advanced data visualization and pathway statistics. Additionally, PathVisio 3 introduces a new powerful extension systems that allows other developers to contribute additional functionality in form of plugins without changing the core application. PathVisio can be downloaded from http://www.pathvisio.org and in 2014 PathVisio 3 has been downloaded over 5,500 times. There are already more than 15 plugins available in the central plugin repository. PathVisio is a freely available, open-source tool published under the Apache 2.0 license (http://www.apache.org/licenses/LICENSE-2.0). It is implemented in Java and thus runs on all major operating systems. The code repository is available at http://svn.bigcat.unimaas.nl/pathvisio. The support mailing list for users is available on https://groups.google.com/forum/#!forum/wikipathways-discuss and for developers on https://groups.google.com/forum/#!forum/wikipathways-devel.
Martina Kutmon, Martijn P. van Iersel, Anwesha Bohler, Thomas Kelder, Nuno Nunes 0002, Alexander R. Pico, Chris T. A. Evelo
PLoS Comput. Biol.7
2014 Scientific Lenses to Support Multiple Views over Linked Chemistry Data
abstract
When are two entries about a small molecule in different datasets the same? If they have the same drug name, chemical structure, or some other criteria? The choice depends upon the application to which the data will be put. However, existing Linked Data approaches provide a single global view over the data with no way of varying the notion of equivalence to be applied.In this paper, we present an approach to enable applications to choose the equivalence criteria to apply between datasets. Thus, supporting multiple dynamic views over the Linked Data. For chemical data, we show that multiple sets of links can be automatically generated according to different equivalence criteria and published with semantic descriptions capturing their context and interpretation. This approach has been applied within a large scale public-private data integration platform for drug discovery. To cater for different use cases, the platform allows the application of different lenses which vary the equivalence rules to be applied based on the context and interpretation of the links.
Colin R. Batchelor, Christian Y. A. Brenninkmeijer, Christine Chichester, Mark Davies, Daniela Digles, Ian Dunlop, Chris T. A. Evelo, Anna Gaulton, Carole A. Goble, Alasdair J. G. Gray, Paul Groth, Lee Harland, Karen Karapetyan, Antonis Loizou, John P. Overington, Steve Pettifer, Jon Steele, Robert Stevens 0001, Valery Tkachenko, Andra Waagmeester, Antony J. Williams, Egon L. Willighagen
ISWC (1)7
2012 GO-Elite: a flexible solution for pathway and ontology over-representation
abstract
UNLABELLED: We introduce GO-Elite, a flexible and powerful pathway analysis tool for a wide array of species, identifiers (IDs), pathways, ontologies and gene sets. In addition to the Gene Ontology (GO), GO-Elite allows the user to perform over-representation analysis on any structured ontology annotations, pathway database or biological IDs (e.g. gene, protein or metabolite). GO-Elite exploits the structured nature of biological ontologies to report a minimal set of non-overlapping terms. The results can be visualized on WikiPathways or as networks. Built-in support is provided for over 60 species and 50 ID systems, covering gene, disease and phenotype ontologies, multiple pathway databases, biomarkers, and transcription factor and microRNA targets. GO-Elite is available as a web interface, GenMAPP-CS plugin and as a cross-platform application. AVAILABILITY: http://www.genmapp.org/go_elite
Alexander C. Zambon, Stan Gaj, Isaac Ho, Kristina Hanspers, Karen Vranizan, Chris T. A. Evelo, Bruce R. Conklin, Alexander R. Pico, Nathan Salomonis
Bioinform.6
2010 The BridgeDb framework: standardized access to gene, protein and metabolite identifier mapping services
abstract
BACKGROUND: Many complementary solutions are available for the identifier mapping problem. This creates an opportunity for bioinformatics tool developers. Tools can be made to flexibly support multiple mapping services or mapping services could be combined to get broader coverage. This approach requires an interface layer between tools and mapping services. RESULTS: Here we present BridgeDb, a software framework for gene, protein and metabolite identifier mapping. This framework provides a standardized interface layer through which bioinformatics tools can be connected to different identifier mapping services. This approach makes it easier for tool developers to support identifier mapping. Mapping services can be combined or merged to support multi-omics experiments or to integrate custom microarray annotations. BridgeDb provides its own ready-to-go mapping services, both in webservice and local database forms. However, the framework is intended for customization and adaptation to any identifier mapping service. BridgeDb has already been integrated into several bioinformatics applications. CONCLUSION: By uncoupling bioinformatics tools from mapping services, BridgeDb improves capability and flexibility of those tools. All described software is open source and available at http://www.bridgedb.org.
Martijn P. van Iersel, Alexander R. Pico, Thomas Kelder, Jianjiong Gao, Isaac Ho, Kristina Hanspers, Bruce R. Conklin, Chris T. A. Evelo
BMC Bioinform.8
2008 Presenting and exploring biological pathways with PathVisio
abstract
BACKGROUND: Biological pathways are a useful abstraction of biological concepts, and software tools to deal with pathway diagrams can help biological research. PathVisio is a new visualization tool for biological pathways that mimics the popular GenMAPP tool with a completely new Java implementation that allows better integration with other open source projects. The GenMAPP MAPP file format is replaced by GPML, a new XML file format that provides seamless exchange of graphical pathway information among multiple programs. RESULTS: PathVisio can be combined with other bioinformatics tools to open up three possible uses: visual compilation of biological knowledge, interpretation of high-throughput expression datasets, and computational augmentation of pathways with interaction information. PathVisio is open source software and available at http://www.pathvisio.org. CONCLUSION: PathVisio is a graphical editor for biological pathways, with flexibility and ease of use as primary goals.
Martijn P. van Iersel, Thomas Kelder, Alexander R. Pico, Kristina Hanspers, Susan L. Coort, Bruce R. Conklin, Chris T. A. Evelo
BMC Bioinform.7
2007 Linking microarray reporters with protein functions
abstract
BACKGROUND: The analysis of microarray experiments requires accurate and up-to-date functional annotation of the microarray reporters to optimize the interpretation of the biological processes involved. Pathway visualization tools are used to connect gene expression data with existing biological pathways by using specific database identifiers that link reporters with elements in the pathways. RESULTS: This paper proposes a novel method that aims to improve microarray reporter annotation by BLASTing the original reporter sequences against a species-specific EMBL subset, that was derived from and crosslinked back to the highly curated UniProt database. The resulting alignments were filtered using high quality alignment criteria and further compared with the outcome of a more traditional approach, where reporter sequences were BLASTed against EnsEMBL followed by locating the corresponding protein (UniProt) entry for the high quality hits. Combining the results of both methods resulted in successful annotation of > 58% of all reporter sequences with UniProt IDs on two commercial array platforms, increasing the amount of Incyte reporters that could be coupled to Gene Ontology terms from 32.7% to 58.3% and to a local GenMAPP pathway from 9.6% to 16.7%. For Agilent, 35.3% of the total reporters are now linked towards GO nodes and 7.1% on local pathways. CONCLUSION: Our methods increased the annotation quality of microarray reporter sequences and allowed us to visualize more reporters using pathway visualization tools. Even in cases where the original reporter annotation showed the correct description the new identifiers often allowed improved pathway and Gene Ontology linking. These methods are freely available at http://www.bigcat.unimaas.nl/public/publications/Gaj_Annotation/.
Stan Gaj, Arie van Erk, Rachel I. M. van Haaften, Chris T. A. Evelo
BMC Bioinform.4
2006 Biologically relevant effects of mRNA amplification on gene expression profiles
abstract
BACKGROUND: Gene expression microarray technology permits the analysis of global gene expression profiles. The amount of sample needed limits the use of small excision biopsies and/or needle biopsies from human or animal tissues. Linear amplification techniques have been developed to increase the amount of sample derived cDNA. These amplified samples can be hybridised on microarrays. However, little information is available whether microarrays based on amplified and unamplified material yield comparable results. In the present study we compared microarray data obtained from amplified mRNA derived from biopsies of rat cardiac left ventricle and non-amplified mRNA derived from the same organ. Biopsies were linearly amplified to acquire enough material for a microarray experiment. Both amplified and unamplified samples were hybridized to the Rat Expression Set 230 Array of Affymetrix. RESULTS: Analysis of the microarray data showed that unamplified material of two different left ventricles had 99.6% identical gene expression. Gene expression patterns of two biopsies obtained from the same parental organ were 96.3% identical. Similarly, gene expression pattern of two biopsies from dissimilar organs were 92.8% identical to each other.Twenty-one percent of reporters called present in parental left ventricular tissue disappeared after amplification in the biopsies. Those reporters were predominantly seen in the low intensity range. Sequence analysis showed that reporters that disappeared after amplification had a GC-content of 53.7+/-4.0%, while reporters called present in biopsy- and whole LV-samples had an average GC content of 47.8+/-5.5% (P <0.001). Those reporters were also predicted to form significantly more (0.76+/-0.07 versus 0.38+/-0.1) and longer (9.4+/-0.3 versus 8.4+/-0.4) hairpins as compared to representative control reporters present before and after amplification. CONCLUSION: This study establishes that the gene expression profile obtained after amplification of mRNA of left ventricular biopsies is representative for the whole left ventricle of the rat heart. However, specific gene transcripts present in parental tissues were undetectable in the minute left ventricular biopsies. Transcripts that were lost due to the amplification process were not randomly distributed, but had higher GC-content and hairpins in the sequence and were mainly found in the lower intensity range which includes many transcription factors from specific signalling pathways.
Rachel I. M. van Haaften, Blanche Schroen, Ben J. A. Janssen, Arie van Erk, Jacques J. M. Debets, Hubert J. M. Smeets, Jos F. M. Smits, Arthur van den Wijngaard, Yigal M. Pinto, Chris T. A. Evelo
BMC Bioinform.10