VLDB 2026 Research / reviewers in the wild / expert
Rainer Breitling
dblp:b/RainerBreitling
· DBLP profile ↗
34ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0001-7173-0922ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 33 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Proscope: Modeling Biosynthetic Diversity Via Probabilistic Structure-Aware Co-EmbeddingabstractA key challenge of natural product discovery is to retrieve the corresponding natural product of a given biosynthetic gene cluster (BGC). As a single BGC may produce multiple structural analogs, modeling such one-to-many relationships between gene clusters and compounds remains challenging. Existing retrieval approaches often learn a deterministic embedding space, which struggle to capture the inherent biological complexity and fail to generalize to novel biosynthetic classes. To address this issue, we introduce ProSCOPE, a Probabilistic Structureaware Co-Embedding framework that models both BGCs and compounds as diagonal Gaussian distributions bridges and bridge these two modalities in a shared embedding space. To measure similarity between distributions in an efficient way, we employ the closed-form 2-Wasserstein distance as the retrieval distance metric. Furthermore, retrieving structurally similar analogs is often valuable in real-word drug discovery applications. To improve the retrieval performance of chemically relevant compounds, we introduce a structure-aware training strategy that modulates the contribution of negative samples according to their chemical similarity to the ground truth. We evaluate our framework across two realistic scenarios: standard cross-validation and class holdout settings. Experiments on the MIBiG dataset demonstrate that ProSCOPE substantially outperforms the previous pointbased co-embedding model. ProSCOPE achieves a 7.5 % relative improvement in in-distribution Recall@10 and a significant$\mathbf{1 6. 3 \%}$relative improvement in the challenging out-of-distribution Recall@10, highlighting the benefit of modeling distributional uncertainty and domain structure in linking gene clusters to natural products. Eriko Takano, Ben Draper, Wayne Wei Zhong Yeo, Rainer Breitling |
BIBM | 6 |
| 2025 | Cross-Granularity Representations for Biological Sequences: Insights From ESM and BiGCARPabstractRecent advances in general-purpose foundation models have stimulated the development of large biological sequence models. While natural language shows symbolic granularity (characters, words, sentences), biological sequences exhibit hierarchical granularity whose levels (nucleotides, amino acids, protein domains, genes) further encode biologically functional information. In this paper, we investigate the integration of crossgranularity knowledge from models through a case study of BiGCARP, a Pfam domain-level model for biosynthetic gene clusters, and ESM, an amino acid-level protein language model. Using representation analysis tools and a set of probe tasks, we first explain why a straightforward cross-model embedding initialization fails to improve downstream performance in BiGCARP, and show that deeper-layer embeddings capture a more contextual and faithful representation of the model's learned knowledge. Furthermore, we demonstrate that representations at different granularities encode complementary biological knowledge, and that combining them yields measurable performance gains in intermediate-level prediction tasks. Our findings highlight crossgranularity integration as a promising strategy for improving both the performance and interpretability of biological foundation models. Our code is available at https://github.com/Nugkta/cgrep. Hanlin Xiao, Rainer Breitling, Eriko Takano, Mauricio A. Álvarez |
BIBM | 2 |
| 2025 | Synteny plot quality control with SyntenyQCabstractSUMMARY: SyntenyQC is a data pre-processing tool for the construction of synteny plots. It supports genomic data collection, annotation and dereplication to facilitate (and in some cases fundamentally enable) the construction of informative synteny plots. AVAILABILITY AND IMPLEMENTATION: SyntenyQC is a command line app developed using Python version 3.10 and tested using pytest. SyntenyQC is available on PyPI (https://pypi.org/project/SyntenyQC) under the MIT License, along with a detailed user tutorial. Package tests can be viewed at https://github.com/Tim-Kirkwood/SyntenyQC. Timothy D. J. Kirkwood, Jack A. Connolly, Ee Lui Ang, Huimin Zhao 0007, Eriko Takano, Rainer Breitling |
Bioinform. | 6 |
| 2023 | ipaPy2: Integrated Probabilistic Annotation (IPA) 2.0 - an improved Bayesian-based method for the annotation of LC-MS/MS untargeted metabolomics dataabstractSUMMARY: The Integrated Probabilistic Annotation (IPA) is an automated annotation method for LC-MS-based untargeted metabolomics experiments that provides statistically rigorous estimates of the probabilities associated with each annotation. Here, we introduce ipaPy2, a substantially improved and completely refactored Python implementation of the IPA method. The revised method is now able to integrate tandem MS fragmentation data, which increases the accuracy of the identifications. Moreover, ipaPy2 provides a much more user-friendly interface, and isotope peaks are no longer treated as individual features but integrated into isotope fingerprints, greatly speeding up the calculations. The method has also been fully integrated with the mzMatch pipeline, so that the results of the annotation can be explored through the newly developed PeakMLViewerPy tool available at https://github.com/UoMMIB/PeakMLViewerPy. AVAILABILITY AND IMPLEMENTATION: The source code, extensive documentation, and tutorials are freely available on GitHub at https://github.com/francescodc87/ipaPy2. Francesco Del Carratore, William Eagles, Juraj Borka, Rainer Breitling |
Bioinform. | 4 |
| 2020 | Unravelling the γ-butyrolactone network in Streptomyces coelicolor by computational ensemble modellingabstractAntibiotic production is coordinated in the Streptomyces coelicolor population through the use of diffusible signaling molecules of the γ-butyrolactone (GBL) family. The GBL regulatory system involves a small, and not completely defined two-gene network which governs a potentially bi-stable switch between the "on" and "off" states of antibiotic production. The use of this circuit as a tool for synthetic biology has been hampered by a lack of mechanistic understanding of its functionality. We here present the creation and analysis of a versatile and adaptable ensemble model of the Streptomyces GBL system (detailed information on all model mechanisms and parameters is documented in http://www.systemsbiology.ls.manchester.ac.uk/wiki/index.php/Main_Page). We use the model to explore a range of previously proposed mechanistic hypotheses, including transcriptional interference, antisense RNA interactions between the mRNAs of the two genes, and various alternative regulatory activities. Our results suggest that transcriptional interference alone is not sufficient to explain the system's behavior. Instead, antisense RNA interactions seem to be the system's driving force, combined with an aggressive scbR promoter. The computational model can be used to further challenge and refine our understanding of the system's activity and guide future experimentation. Areti Tsigkinopoulou, Eriko Takano, Rainer Breitling |
PLoS Comput. Biol. | 3 |
| 2018 | Selenzyme: enzyme selection tool for pathway designabstractSummary: Synthetic biology applies the principles of engineering to biology in order to create biological functionalities not seen before in nature. One of the most exciting applications of synthetic biology is the design of new organisms with the ability to produce valuable chemicals including pharmaceuticals and biomaterials in a greener; sustainable fashion. Selecting the right enzymes to catalyze each reaction step in order to produce a desired target compound is, however, not trivial. Here, we present Selenzyme, a free online enzyme selection tool for metabolic pathway design. The user is guided through several decision steps in order to shortlist the best candidates for a given pathway step. The tool graphically presents key information about enzymes based on existing databases and tools such as: similarity of sequences and of catalyzed reactions; phylogenetic distance between source organism and intended host species; multiple alignment highlighting conserved regions, predicted catalytic site, and active regions and relevant properties such as predicted solubility and transmembrane regions. Selenzyme provides bespoke sequence selection for automated workflows in biofoundries. Availability and implementation: The tool is integrated as part of the pathway design stage into the design-build-test-learn SYNBIOCHEM pipeline. The Selenzyme web server is available at http://selenzyme.synbiochem.co.uk. Supplementary information: Supplementary data are available at Bioinformatics online. Pablo Carbonell, Jerry Wong, Neil Swainston, Eriko Takano, Nicholas J. Turner, Nigel S. Scrutton, Douglas B. Kell, Rainer Breitling, Jean-Loup Faulon |
Bioinform. | 8 |
| 2017 | RankProd 2.0: a refactored bioconductor package for detecting differentially expressed features in molecular profiling datasetsabstractMOTIVATION: The Rank Product (RP) is a statistical technique widely used to detect differentially expressed features in molecular profiling experiments such as transcriptomics, metabolomics and proteomics studies. An implementation of the RP and the closely related Rank Sum (RS) statistics has been available in the RankProd Bioconductor package for several years. However, several recent advances in the understanding of the statistical foundations of the method have made a complete refactoring of the existing package desirable. RESULTS: We implemented a completely refactored version of the RankProd package, which provides a more principled implementation of the statistics for unpaired datasets. Moreover, the permutation-based P -value estimation methods have been replaced by exact methods, providing faster and more accurate results. AVAILABILITY AND IMPLEMENTATION: RankProd 2.0 is available at Bioconductor ( https://www.bioconductor.org/packages/devel/bioc/html/RankProd.html ) and as part of the mzMatch pipeline ( http://www.mzmatch.sourceforge.net ). CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Francesco Del Carratore, Andris Jankevics, Rob Eisinga, Tom Heskes, Fangxin Hong, Rainer Breitling |
Bioinform. | 6 |
| 2015 | Incorporating peak grouping information for alignment of multiple liquid chromatography-mass spectrometry datasetsabstractMOTIVATION: The combination of liquid chromatography and mass spectrometry (LC/MS) has been widely used for large-scale comparative studies in systems biology, including proteomics, glycomics and metabolomics. In almost all experimental design, it is necessary to compare chromatograms across biological or technical replicates and across sample groups. Central to this is the peak alignment step, which is one of the most important but challenging preprocessing steps. Existing alignment tools do not take into account the structural dependencies between related peaks that coelute and are derived from the same metabolite or peptide. We propose a direct matching peak alignment method for LC/MS data that incorporates related peaks information (within each LC/MS run) and investigate its effect on alignment performance (across runs). The groupings of related peaks necessary for our method can be obtained from any peak clustering method and are built into a pair-wise peak similarity score function. The similarity score matrix produced is used by an approximation algorithm for the weighted matching problem to produce the actual alignment result. RESULTS: We demonstrate that related peak information can improve alignment performance. The performance is evaluated on a set of benchmark datasets, where our method performs competitively compared to other popular alignment tools. AVAILABILITY: The proposed alignment method has been implemented as a stand-alone application in Python, available for download at http://github.com/joewandy/peak-grouping-alignment. Joe Wandy, Rónán Daly, Rainer Breitling, Simon Rogers |
Bioinform. | 3 |
| 2014 | MetAssign: probabilistic annotation of metabolites from LC-MS data using a Bayesian clustering approachabstractMOTIVATION: The use of liquid chromatography coupled to mass spectrometry has enabled the high-throughput profiling of the metabolite composition of biological samples. However, the large amount of data obtained can be difficult to analyse and often requires computational processing to understand which metabolites are present in a sample. This article looks at the dual problem of annotating peaks in a sample with a metabolite, together with putatively annotating whether a metabolite is present in the sample. The starting point of the approach is a Bayesian clustering of peaks into groups, each corresponding to putative adducts and isotopes of a single metabolite. RESULTS: The Bayesian modelling introduced here combines information from the mass-to-charge ratio, retention time and intensity of each peak, together with a model of the inter-peak dependency structure, to increase the accuracy of peak annotation. The results inherently contain a quantitative estimate of confidence in the peak annotations and allow an accurate trade-off between precision and recall. Extensive validation experiments using authentic chemical standards show that this system is able to produce more accurate putative identifications than other state-of-the-art systems, while at the same time giving a probabilistic measure of confidence in the annotations. AVAILABILITY AND IMPLEMENTATION: The software has been implemented as part of the mzMatch metabolomics analysis pipeline, which is available for download at http://mzmatch.sourceforge.net/. Rónán Daly, Simon Rogers, Joe Wandy, Andris Jankevics, Karl E. V. Burgess, Rainer Breitling |
Bioinform. | 6 |
| 2014 | A fast algorithm for determining bounds and accurate approximate p-values of the rank product statistic for replicate experimentsabstractBACKGROUND: The rank product method is a powerful statistical technique for identifying differentially expressed molecules in replicated experiments. A critical issue in molecule selection is accurate calculation of the p-value of the rank product statistic to adequately address multiple testing. Both exact calculation and permutation and gamma approximations have been proposed to determine molecule-level significance. These current approaches have serious drawbacks as they are either computationally burdensome or provide inaccurate estimates in the tail of the p-value distribution. RESULTS: We derive strict lower and upper bounds to the exact p-value along with an accurate approximation that can be used to assess the significance of the rank product statistic in a computationally fast manner. The bounds and the proposed approximation are shown to provide far better accuracy over existing approximate methods in determining tail probabilities, with the slightly conservative upper bound protecting against false positives. We illustrate the proposed method in the context of a recently published analysis on transcriptomic profiling performed in blood. CONCLUSIONS: We provide a method to determine upper bounds and accurate approximate p-values of the rank product statistic. The proposed algorithm provides an order of magnitude increase in throughput as compared with current approaches and offers the opportunity to explore new application domains with even larger multiple testing issue. The R code is published in one of the Additional files and is available at http://www.ru.nl/publish/pages/726696/rankprodbounds.zip . Tom Heskes, Rob Eisinga, Rainer Breitling |
BMC Bioinform. | 3 |
| 2014 | Pep2Path: Automated Mass Spectrometry-Guided Genome Mining of Peptidic Natural ProductsabstractNonribosomally and ribosomally synthesized bioactive peptides constitute a source of molecules of great biomedical importance, including antibiotics such as penicillin, immunosuppressants such as cyclosporine, and cytostatics such as bleomycin. Recently, an innovative mass-spectrometry-based strategy, peptidogenomics, has been pioneered to effectively mine microbial strains for novel peptidic metabolites. Even though mass-spectrometric peptide detection can be performed quite fast, true high-throughput natural product discovery approaches have still been limited by the inability to rapidly match the identified tandem mass spectra to the gene clusters responsible for the biosynthesis of the corresponding compounds. With Pep2Path, we introduce a software package to fully automate the peptidogenomics approach through the rapid Bayesian probabilistic matching of mass spectra to their corresponding biosynthetic gene clusters. Detailed benchmarking of the method shows that the approach is powerful enough to correctly identify gene clusters even in data sets that consist of hundreds of genomes, which also makes it possible to match compounds from unsequenced organisms to closely related biosynthetic gene clusters in other genomes. Applying Pep2Path to a data set of compounds without known biosynthesis routes, we were able to identify candidate gene clusters for the biosynthesis of five important compounds. Notably, one of these clusters was detected in a genome from a different subphylum of Proteobacteria than that in which the molecule had first been identified. All in all, our approach paves the way towards high-throughput discovery of novel peptidic natural products. Pep2Path is freely available from http://pep2path.sourceforge.net/, implemented in Python, licensed under the GNU General Public License v3 and supported on MS Windows, Linux and Mac OS X. Marnix H. Medema, Yared Paalvast, Don D. Nguyen, Alexey Melnik, Pieter C. Dorrestein, Eriko Takano, Rainer Breitling |
PLoS Comput. Biol. | 7 |
| 2013 | mzMatch-ISO: an R tool for the annotation and relative quantification of isotope-labelled mass spectrometry dataabstractMOTIVATION: Stable isotope-labelling experiments have recently gained increasing popularity in metabolomics studies, providing unique insights into the dynamics of metabolic fluxes, beyond the steady-state information gathered by routine mass spectrometry. However, most liquid chromatography-mass spectrometry data analysis software lacks features that enable automated annotation and relative quantification of labelled metabolite peaks. Here, we describe mzMatch-ISO, a new extension to the metabolomics analysis pipeline mzMatch.R. RESULTS: Targeted and untargeted isotope profiling using mzMatch-ISO provides a convenient visual summary of the quality and quantity of labelling for every metabolite through four types of diagnostic plots that show (i) the chromatograms of the isotope peaks of each compound in each sample group; (ii) the ratio of mono-isotopic and labelled peaks indicating the fraction of labelling; (iii) the average peak area of mono-isotopic and labelled peaks in each sample group; and (iv) the trend in the relative amount of labelling in a predetermined isotopomer. To aid further statistical analyses, the values used for generating these plots are also provided as a tab-delimited file. We demonstrate the power and versatility of mzMatch-ISO by analysing a (13)C-labelled metabolome dataset from trypanosomal parasites. AVAILABILITY: mzMatch.R and mzMatch-ISO are available free of charge from http://mzmatch.sourceforge.net and can be used on Linux and Windows platforms running the latest version of R. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Achuthanunni Chokkathukalam, Andris Jankevics, Darren J. Creek, Fiona Achcar, Michael P. Barrett, Rainer Breitling |
Bioinform. | 6 |
| 2013 | Handling Uncertainty in Dynamic Models: The Pentose Phosphate Pathway in Trypanosoma bruceiabstractDynamic models of metabolism can be useful in identifying potential drug targets, especially in unicellular organisms. A model of glycolysis in the causative agent of human African trypanosomiasis, Trypanosoma brucei, has already shown the utility of this approach. Here we add the pentose phosphate pathway (PPP) of T. brucei to the glycolytic model. The PPP is localized to both the cytosol and the glycosome and adding it to the glycolytic model without further adjustments leads to a draining of the essential bound-phosphate moiety within the glycosome. This phosphate "leak" must be resolved for the model to be a reasonable representation of parasite physiology. Two main types of theoretical solution to the problem could be identified: (i) including additional enzymatic reactions in the glycosome, or (ii) adding a mechanism to transfer bound phosphates between cytosol and glycosome. One example of the first type of solution would be the presence of a glycosomal ribokinase to regenerate ATP from ribose 5-phosphate and ADP. Experimental characterization of ribokinase in T. brucei showed that very low enzyme levels are sufficient for parasite survival, indicating that other mechanisms are required in controlling the phosphate leak. Examples of the second type would involve the presence of an ATP:ADP exchanger or recently described permeability pores in the glycosomal membrane, although the current absence of identified genes encoding such molecules impedes experimental testing by genetic manipulation. Confronted with this uncertainty, we present a modeling strategy that identifies robust predictions in the context of incomplete system characterization. We illustrate this strategy by exploring the mechanism underlying the essential function of one of the PPP enzymes, and validate it by confirming the model predictions experimentally. Eduard J. Kerkhoven, Fiona Achcar, Vincent P. Alibu, Richard J. Burchmore, Ian H. Gilbert, Maciej Trybilo, Nicole N. Driessen, David R. Gilbert, Rainer Breitling, Barbara M. Bakker, Michael P. Barrett |
PLoS Comput. Biol. | 9 |
| 2012 | IDEOM: an Excel interface for analysis of LC-MS-based metabolomics dataabstractSUMMARY: The application of emerging metabolomics technologies to the comprehensive investigation of cellular biochemistry has been limited by bottlenecks in data processing, particularly noise filtering and metabolite identification. IDEOM provides a user-friendly data processing application that automates filtering and identification of metabolite peaks, paying particular attention to common sources of noise and false identifications generated by liquid chromatography-mass spectrometry (LC-MS) platforms. Building on advanced processing tools such as mzMatch and XCMS, it allows users to run a comprehensive pipeline for data analysis and visualization from a graphical user interface within Microsoft Excel, a familiar program for most biological scientists. AVAILABILITY AND IMPLEMENTATION: IDEOM is provided free of charge at http://mzmatch.sourceforge.net/ideom.html, as a macro-enabled spreadsheet (.xlsb). Implementation requires Microsoft Excel (2007 or later). R is also required for full functionality. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Darren J. Creek, Andris Jankevics, Karl E. V. Burgess, Rainer Breitling, Michael P. Barrett |
Bioinform. | 4 |
| 2012 | Dynamic Modelling under Uncertainty: The Case of Trypanosoma brucei Energy MetabolismabstractKinetic models of metabolism require detailed knowledge of kinetic parameters. However, due to measurement errors or lack of data this knowledge is often uncertain. The model of glycolysis in the parasitic protozoan Trypanosoma brucei is a particularly well analysed example of a quantitative metabolic model, but so far it has been studied with a fixed set of parameters only. Here we evaluate the effect of parameter uncertainty. In order to define probability distributions for each parameter, information about the experimental sources and confidence intervals for all parameters were collected. We created a wiki-based website dedicated to the detailed documentation of this information: the SilicoTryp wiki (http://silicotryp.ibls.gla.ac.uk/wiki/Glycolysis). Using information collected in the wiki, we then assigned probability distributions to all parameters of the model. This allowed us to sample sets of alternative models, accurately representing our degree of uncertainty. Some properties of the model, such as the repartition of the glycolytic flux between the glycerol and pyruvate producing branches, are robust to these uncertainties. However, our analysis also allowed us to identify fragilities of the model leading to the accumulation of 3-phosphoglycerate and/or pyruvate. The analysis of the control coefficients revealed the importance of taking into account the uncertainties about the parameters, as the ranking of the reactions can be greatly affected. This work will now form the basis for a comprehensive Bayesian analysis and extension of the model considering alternative topologies. Fiona Achcar, Eduard J. Kerkhoven, Barbara M. Bakker, Michael P. Barrett, Rainer Breitling |
PLoS Comput. Biol. | 5 |
| 2010 | DiffCoEx: a simple and sensitive method to find differentially coexpressed gene modulesabstractBACKGROUND: Large microarray datasets have enabled gene regulation to be studied through coexpression analysis. While numerous methods have been developed for identifying differentially expressed genes between two conditions, the field of differential coexpression analysis is still relatively new. More specifically, there is so far no sensitive and untargeted method to identify gene modules (also known as gene sets or clusters) that are differentially coexpressed between two conditions. Here, sensitive and untargeted means that the method should be able to construct de novo modules by grouping genes based on shared, but subtle, differential correlation patterns. RESULTS: We present DiffCoEx, a novel method for identifying correlation pattern changes, which builds on the commonly used Weighted Gene Coexpression Network Analysis (WGCNA) framework for coexpression analysis. We demonstrate its usefulness by identifying biologically relevant, differentially coexpressed modules in a rat cancer dataset. CONCLUSIONS: DiffCoEx is a simple and sensitive method to identify gene coexpression differences between multiple conditions. Bruno M. Tesson, Rainer Breitling, Ritsert C. Jansen |
BMC Bioinform. | 2 |
| 2009 | Probabilistic assignment of formulas to mass peaks in metabolomics experimentsabstractMOTIVATION: High-accuracy mass spectrometry is a popular technology for high-throughput measurements of cellular metabolites (metabolomics). One of the major challenges is the correct identification of the observed mass peaks, including the assignment of their empirical formula, based on the measured mass. RESULTS: We propose a novel probabilistic method for the assignment of empirical formulas to mass peaks in high-throughput metabolomics mass spectrometry measurements. The method incorporates information about possible biochemical transformations between the empirical formulas to assign higher probability to formulas that could be created from other metabolites in the sample. In a series of experiments, we show that the method performs well and provides greater insight than assignments based on mass alone. In addition, we extend the model to incorporate isotope information to achieve even more reliable formula identification. AVAILABILITY: A supplementary document, Matlab code, data and further information are available from http://www.dcs.gla.ac.uk/inference/metsamp. Simon Rogers, Richard A. Scheltema, Mark A. Girolami, Rainer Breitling |
Bioinform. | 4 |
| 2009 | designGG: an R-package and web tool for the optimal design of genetical genomics experimentsabstractBACKGROUND: High-dimensional biomolecular profiling of genetically different individuals in one or more environmental conditions is an increasingly popular strategy for exploring the functioning of complex biological systems. The optimal design of such genetical genomics experiments in a cost-efficient and effective way is not trivial. RESULTS: This paper presents designGG, an R package for designing optimal genetical genomics experiments. A web implementation for designGG is available at http://gbic.biol.rug.nl/designGG. All software, including source code and documentation, is freely available. CONCLUSION: DesignGG allows users to intelligently select and allocate individuals to experimental units and conditions such as drug treatment. The user can maximize the power and resolution of detecting genetic, environmental and interaction effects in a genome-wide or local mode by giving more weight to genome regions of special interest, such as previously detected phenotypic quantitative trait loci. This will help to achieve high power and more accurate estimates of the effects of interesting factors, and thus yield a more reliable biological interpretation of data. DesignGG is applicable to linkage analysis of experimental crosses, e.g. recombinant inbred lines, as well as to association analysis of natural populations. Yang Li 0036, Morris A. Swertz, Gonzalo Vera, Jingyuan Fu, Rainer Breitling, Ritsert C. Jansen |
BMC Bioinform. | 5 |
| 2008 | A structured approach for the engineering of biochemical network models, illustrated for signalling pathwaysabstractQuantitative models of biochemical networks (signal transduction cascades, metabolic pathways, gene regulatory circuits) are a central component of modern systems biology. Building and managing these complex models is a major challenge that can benefit from the application of formal methods adopted from theoretical computing science. Here we provide a general introduction to the field of formal modelling, which emphasizes the intuitive biochemical basis of the modelling process, but is also accessible for an audience with a background in computing science and/or model engineering. We show how signal transduction cascades can be modelled in a modular fashion, using both a qualitative approach--qualitative Petri nets, and quantitative approaches--continuous Petri nets and ordinary differential equations (ODEs). We review the major elementary building blocks of a cellular signalling model, discuss which critical design decisions have to be made during model building, and present a number of novel computational tools that can help to explore alternative modular models in an easy and intuitive manner. These tools, which are based on Petri net theory, offer convenient ways of composing hierarchical ODE models, and permit a qualitative analysis of their behaviour. We illustrate the central concepts using signal transduction as our main example. The ultimate aim is to introduce a general approach that provides the foundations for a structured formal engineering of large-scale models of biochemical networks. Rainer Breitling, David R. Gilbert, Monika Heiner, Richard J. Orton |
Briefings Bioinform. | 1 |
| 2008 | A comparison of meta-analysis methods for detecting differentially expressed genes in microarray experimentsabstractMOTIVATION: The proliferation of public data repositories creates a need for meta-analysis methods to efficiently evaluate, integrate and validate related datasets produced by independent groups. A t-based approach has been proposed to integrate effect size from multiple studies by modeling both intra- and between-study variation. Recently, a non-parametric 'rank product' method, which is derived based on biological reasoning of fold-change criteria, has been applied to directly combine multiple datasets into one meta study. Fisher's Inverse chi(2) method, which only depends on P-values from individual analyses of each dataset, has been used in a couple of medical studies. While these methods address the question from different angles, it is not clear how they compare with each other. RESULTS: We comparatively evaluate the three methods; t-based hierarchical modeling, rank products and Fisher's Inverse chi(2) test with P-values from either the t-based or the rank product method. A simulation study shows that the rank product method, in general, has higher sensitivity and selectivity than the t-based method in both individual and meta-analysis, especially in the setting of small sample size and/or large between-study variation. Not surprisingly, Fisher's chi(2) method highly depends on the method used in the individual analysis. Application to real datasets demonstrates that meta-analysis achieves more reliable identification than an individual analysis, and rank products are more robust in gene ranking, which leads to a much higher reproducibility among independent studies. Though t-based meta-analysis greatly improves over the individual analysis, it suffers from a potentially large amount of false positives when P-values serve as threshold. We conclude that careful meta-analysis is a powerful tool for integrating multiple array studies. Fangxin Hong, Rainer Breitling |
Bioinform. | 2 |
| 2008 | MetaNetter: inference and visualization of high-resolution metabolomic networksabstractAbstract Summary: We present a Cytoscape plugin for the inference and visualization of networks from high-resolution mass spectrometry metabolomic data. The software also provides access to basic topological analysis. This open source, multi-platform software has been successfully used to interpret metabolomic experiments and will enable others using filtered, high mass accuracy mass spectrometric data sets to build and analyse networks. Availability: http://compbio.dcs.gla.ac.uk/fabien/abinitio/abinitio.html Contact: [email protected] Supplementary information: http://compbio.dcs.gla.ac.uk/fabien/abinitio/doc/Supplementary.pdf Fabien Jourdan, Rainer Breitling, Michael P. Barrett, David R. Gilbert |
Bioinform. | 2 |
| 2007 | Discriminating Microbial Species Using Protein Sequence Properties and Machine Learning
Ali Al-Shahib, David R. Gilbert, Rainer Breitling |
IDEAL | 3 |
| 2007 | Analysis of Tiling Microarray Data by Learning Vector Quantization and Relevance Learning
Michael Biehl, Rainer Breitling, Yang Li 0036 |
IDEAL | 2 |
| 2007 | FIVA: Functional Information Viewer and Analyzer extracting biological knowledge from transcriptome data of prokaryotesabstractAbstract Summary: FIVA (Function Information Viewer and Analyzer) aids researchers in the prokaryotic community to quickly identify relevant biological processes following transcriptome analysis. Our software assists in functional profiling of large sets of genes and generates a comprehensive overview of affected biological processes. Availability: http://bioinformatics.biol.rug.nl/standalone/fiva/ Contact: [email protected] Supplementary information: http://bioinformatics.biol.rug.nl/standalone/fiva/suppMaterials.php Evert-Jan Blom, Dinne W. J. Bosman, Sacha A. F. T. van Hijum, Rainer Breitling, Lars Tijsma, Remko Silvis, Jos B. T. M. Roerdink, Oscar P. Kuipers |
Bioinform. | 4 |
| 2007 | A verification protocol for the probe sequences of Affymetrix genome arrays reveals high probe accuracy for studies in mouse, human and ratabstractBACKGROUND: The Affymetrix GeneChip technology uses multiple probes per gene to measure its expression level. Individual probe signals can vary widely, which hampers proper interpretation. This variation can be caused by probes that do not properly match their target gene or that match multiple genes. To determine the accuracy of Affymetrix arrays, we developed an extensive verification protocol, for mouse arrays incorporating the NCBI RefSeq, NCBI UniGene Unique, NIA Mouse Gene Index, and UCSC mouse genome databases. RESULTS: Applying this protocol to Affymetrix Mouse Genome arrays (the earlier U74Av2 and the newer 430 2.0 array), the number of sequence-verified probes with perfect matches was no less than 85% and 95%, respectively; and for 74% and 85% of the probe sets all probes were sequence verified. The latter percentages increased to 80% and 94% after discarding one or two unverifiable probes per probe set, and even further to 84% and 97% when, in addition, allowing for one or two mismatches between probe and target gene. Similar results were obtained for other mouse arrays, as well as for human and rat arrays. Based on these data, refined chip definition files for all arrays are provided online. Researchers can choose the version appropriate for their study to (re)analyze expression data. CONCLUSION: The accuracy of Affymetrix probe sequences is higher than previously reported, particularly on newer arrays. Yet, refined probe set definitions have clear effects on the detection of differentially expressed genes. We demonstrate that the interpretation of the results of Affymetrix arrays is improved when the new chip definition files are used. Rudi Alberts, Peter Terpstra, Menno Hardonk, Leonid V. Bystrykh, Gerald de Haan, Rainer Breitling, Jan Peter Nap, Ritsert C. Jansen |
BMC Bioinform. | 6 |
| 2006 | RankProd: a bioconductor package for detecting differentially expressed genes in meta-analysisabstractUNLABELLED: While meta-analysis provides a powerful tool for analyzing microarray experiments by combining data from multiple studies, it presents unique computational challenges. The Bioconductor package RankProd provides a new and intuitive tool for this purpose in detecting differentially expressed genes under two experimental conditions. The package modifies and extends the rank product method proposed by Breitling et al., [(2004) FEBS Lett., 573, 83-92] to integrate multiple microarray studies from different laboratories and/or platforms. It offers several advantages over t-test based methods and accepts pre-processed expression datasets produced from a wide variety of platforms. The significance of the detection is assessed by a non-parametric permutation test, and the associated P-value and false discovery rate (FDR) are included in the output alongside the genes that are detected by user-defined criteria. A visualization plot is provided to view actual expression levels for each gene with estimated significance measurements. AVAILABILITY: RankProd is available at Bioconductor http://www.bioconductor.org. A web-based interface will soon be available at http://cactus.salk.edu/RankProd Fangxin Hong, Rainer Breitling, Connor W. McEntee, Ben S. Wittner, Jennifer L. Nemhauser, Joanne Chory |
Bioinform. | 2 |
| 2006 | A lock-and-key model for protein-protein interactionsabstractMOTIVATION: Protein-protein interaction networks are one of the major post-genomic data sources available to molecular biologists. They provide a comprehensive view of the global interaction structure of an organism's proteome, as well as detailed information on specific interactions. Here we suggest a physical model of protein interactions that can be used to extract additional information at an intermediate level: It enables us to identify proteins which share biological interaction motifs, and also to identify potentially missing or spurious interactions. RESULTS: Our new graph model explains observed interactions between proteins by an underlying interaction of complementary binding domains (lock-and-key model). This leads to a novel graph-theoretical algorithm to identify bipartite subgraphs within protein-protein interaction networks where the underlying data are taken from yeast two-hybrid experimental results. By testing on synthetic data, we demonstrate that under certain modelling assumptions, the algorithm will return correct domain information about each protein in the network. Tests on data from various model organisms show that the local and global patterns predicted by the model are indeed found in experimental data. Using functional and protein structure annotations, we show that bipartite subnetworks can be identified that correspond to biologically relevant interaction motifs. Some of these are novel and we discuss an example involving SH3 domains from the Saccharomyces cerevisiae interactome. AVAILABILITY: The algorithm (in Matlab format) is available (see http://www.maths.strath.ac.uk/~aas96106/lock_key.html). Julie L. Morrison, Rainer Breitling, Desmond J. Higham, David R. Gilbert |
Bioinform. | 2 |
| 2005 | Vector analysis as a fast and easy method to compare gene expression responses between different experimental backgroundsabstractBACKGROUND: Gene expression studies increasingly compare expression responses between different experimental backgrounds (genetic, physiological, or phylogenetic). By focusing on dynamic responses rather than a direct comparison of static expression levels, this type of study allows a finer dissection of primary and secondary regulatory effects in the various backgrounds. Usually, results of such experiments are presented in the form of Venn diagrams, which are intuitive and visually appealing, but lack a statistical foundation. RESULTS: Here we introduce Vector Analysis (VA) as a simple, yet principled, approach to comparing expression responses in different experimental backgrounds. VA enables the automatic assignment of genes to response prototypes and provides statistical significance estimates to eliminate spurious response patterns. The application of VA to a real dataset, comparing nutrient starvation responses in wild type and mutant Arabidopsis plants, reveals that consistent patterns of expression behavior are present in the data and are reliably detected by the algorithm. CONCLUSION: Vector analysis is a flexible, easy-to-use technique to compare gene expression patterns in different experimental backgrounds. It compares favorably with the classical Venn diagram approach and can be implemented manually using spreadsheets, such as Excel, or automatically by using the supplied software. Rainer Breitling, Patrick Armengaud, Anna Amtmann |
BMC Bioinform. | 1 |
| 2005 | GeneRank: Using search engine technology for the analysis of microarray experimentsabstractBACKGROUND: Interpretation of simple microarray experiments is usually based on the fold-change of gene expression between a reference and a "treated" sample where the treatment can be of many types from drug exposure to genetic variation. Interpretation of the results usually combines lists of differentially expressed genes with previous knowledge about their biological function. Here we evaluate a method--based on the PageRank algorithm employed by the popular search engine Google--that tries to automate some of this procedure to generate prioritized gene lists by exploiting biological background information. RESULTS: GeneRank is an intuitive modification of PageRank that maintains many of its mathematical properties. It combines gene expression information with a network structure derived from gene annotations (gene ontologies) or expression profile correlations. Using both simulated and real data we find that the algorithm offers an improved ranking of genes compared to pure expression change rankings. CONCLUSION: Our modification of the PageRank algorithm provides an alternative method of evaluating microarray experimental results which combines prior knowledge about the underlying network. GeneRank offers an improvement compared to assessing the importance of a gene based on its experimentally observed fold-change alone and may be used as a basis for further analytical developments. Julie L. Morrison, Rainer Breitling, Desmond J. Higham, David R. Gilbert |
BMC Bioinform. | 2 |
| 2005 | Franksum: new feature selection method for protein function predictionabstractIn the study of in silico functional genomics, improving the performance of protein function prediction is the ultimate goal for identifying proteins associated with defined cellular functions. The classical prediction approach is to employ pairwise sequence alignments. However this method often faces difficulties when no statistically significant homologous sequences are identified. An alternative way is to predict protein function from sequence-derived features using machine learning. In this case the choice of possible features which can be derived from the sequence is of vital importance to ensure adequate discrimination to predict function. In this paper we have successfully selected biologically significant features for protein function prediction. This was performed using a new feature selection method (FrankSum) that avoids data distribution assumptions, uses a data independent measurement (p-value) within the feature, identifies redundancy between features and uses an appropriate ranking criterion for feature selection. We have shown that classifiers generated from features selected by FrankSum outperforms classifiers generated from full feature sets, randomly selected features and features selected from the Wrapper method. We have also shown the features are concordant across all species and top ranking features are biologically informative. We conclude that feature selection is vital for successful protein function prediction and FrankSum is one of the feature selection methods that can be applied successfully to such a domain. Ali Al-Shahib, Rainer Breitling, David R. Gilbert |
Int. J. Neural Syst. | 2 |
| 2005 | The Latent Process Decomposition of cDNA Microarray Data SetsabstractWe present a new computational technique (a software implementation, data sets, and supplementary information are available at http://www.enm.bris.ac.uk/lpd/) which enables the probabilistic analysis of cDNA microarray data and we demonstrate its effectiveness in identifying features of biomedical importance. A hierarchical Bayesian model, called Latent Process Decomposition (LPD), is introduced in which each sample in the data set is represented as a combinatorial mixture over a finite set of latent processes, which are expected to correspond to biological processes. Parameters in the model are estimated using efficient variational methods. This type of probabilistic model is most appropriate for the interpretation of measurement data generated by cDNA microarray technology. For determining informative substructure in such data sets, the proposed model has several important advantages over the standard use of dendrograms. First, the ability to objectively assess the optimal number of sample clusters. Second, the ability to represent samples and gene expression levels using a common set of latent variables (dendrograms cluster samples and gene expression values separately which amounts to two distinct reduced space representations). Third, in constrast to standard cluster models, observations are not assigned to a single cluster and, thus, for example, gene expression levels are modeled via combinations of the latent processes identified by the algorithm. We show this new method compares favorably with alternative cluster analysis methods. To illustrate its potential, we apply the proposed technique to several microarray data sets for cancer. For these data sets it successfully decomposes the data into known subtypes and indicates possible further taxonomic subdivision in addition to highlighting, in a wholly unsupervised manner, the importance of certain genes which are known to be medically significant. To illustrate its wider applicability, we also illustrate its performance on a microarray data set for yeast. Simon Rogers, Mark A. Girolami, Colin Campbell, Rainer Breitling |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2004 | Biologically valid linear factor models of gene expressionabstractAbstract Motivation: The identification of physiological processes underlying and generating the expression pattern observed in microarray experiments is a major challenge. Principal component analysis (PCA) is a linear multivariate statistical method that is regularly employed for that purpose as it provides a reduced-dimensional representation for subsequent study of possible biological processes responding to the particular experimental conditions. Making explicit the data assumptions underlying PCA highlights their lack of biological validity thus making biological interpretation of the principal components problematic. A microarray data representation which enables clear biological interpretation is a desirable analysis tool. Results: We address this issue by employing the probabilistic interpretation of PCA and proposing alternative linear factor models which are based on refined biological assumptions. A practical study on two well-understood microarray datasets highlights the weakness of PCA and the greater biological interpretability of the linear models we have developed. Availability: The model estimation routines are currently implemented as Matlab routines and these, as well as data and results reported, are available from the following URL: http://www.dcs.gla.ac.uk/~girolami/lfm/index.html Mark A. Girolami, Rainer Breitling |
Bioinform. | 2 |
| 2004 | Iterative Group Analysis (iGA): A simple tool to enhance sensitivity and facilitate interpretation of microarray experimentsabstractBACKGROUND: The biological interpretation of even a simple microarray experiment can be a challenging and highly complex task. Here we present a new method (Iterative Group Analysis) to facilitate, improve, and accelerate this process. RESULTS: Our Iterative Group Analysis approach (iGA) uses elementary statistics to identify those functional classes of genes that are significantly changed in an experiment and at the same time determines which of the class members are most likely to be differentially expressed. iGA does not require that all members of a class change and is therefore robust against imperfect class assignments, which can be derived from public sources (e.g. GeneOntologies) or automated processes (e.g. key word extraction from gene names). In contrast to previous non-iterative approaches, iGA does not depend on the availability of fixed lists of differentially expressed genes, and thus can be used to increase the sensitivity of gene detection especially in very noisy or small data sets. In the extreme, iGA can even produce statistically meaningful results without any experimental replication. The automated functional annotation provided by iGA greatly reduces the complexity of microarray results and facilitates the interpretation process. In addition, iGA can be used as a fast and efficient tool for the platform-independent comparison of a microarray experiment to the vast number of published results, automatically highlighting shared genes of potential interest. CONCLUSIONS: By applying iGA to a wide variety of data from diverse organisms and platforms we show that this approach enhances and accelerates the interpretation of microarray experiments. Rainer Breitling, Anna Amtmann, Pawel Herzyk |
BMC Bioinform. | 1 |
| 2004 | Graph-based iterative Group Analysis enhances microarray interpretationabstractBACKGROUND: One of the most time-consuming tasks after performing a gene expression experiment is the biological interpretation of the results by identifying physiologically important associations between the differentially expressed genes. A large part of the relevant functional evidence can be represented in the form of graphs, e.g. metabolic and signaling pathways, protein interaction maps, shared GeneOntology annotations, or literature co-citation relations. Such graphs are easily constructed from available genome annotation data. The problem of biological interpretation can then be described as identifying the subgraphs showing the most significant patterns of gene expression. We applied a graph-based extension of our iterative Group Analysis (iGA) approach to obtain a statistically rigorous identification of the subgraphs of interest in any evidence graph. RESULTS: We validated the Graph-based iterative Group Analysis (GiGA) by applying it to the classic yeast diauxic shift experiment of DeRisi et al., using GeneOntology and metabolic network information. GiGA reliably identified and summarized all the biological processes discussed in the original publication. Visualization of the detected subgraphs allowed the convenient exploration of the results. The method also identified several processes that were not presented in the original paper but are of obvious relevance to the yeast starvation response. CONCLUSIONS: GiGA provides a fast and flexible delimitation of the most interesting areas in a microarray experiment, and leads to a considerable speed-up and improvement of the interpretation process. Rainer Breitling, Anna Amtmann, Pawel Herzyk |
BMC Bioinform. | 1 |