EDBT 2026 Demo / reviewers in the wild / expert
Wolfgang Huber
dblp:75/2422
· DBLP profile ↗
34ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0002-0474-2218ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 33 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Feature selection by replicate reproducibility and non-redundancyabstractMOTIVATION: A fundamental step in many analyses of high-dimensional data is dimension reduction. Two basic approaches are introduction of new synthetic coordinates and selection of extant features. Advantages of the latter include interpretability, simplicity, transferability, and modularity. A common criterion for unsupervized feature selection is variance or dynamic range. However, in practice, it can occur that high-variance features are noisy, that important features have low variance, or that variances are simply not comparable across features because they are measured in unrelated numeric scales or physical units. Moreover, users may want to include measures of signal-to-noise ratio and non-redundancy into feature selection. RESULTS: Here, we introduce the RNR algorithm, which selects features based on (i) the reproducibility of their signal across replicates and (ii) their non-redundancy, measured by linear dependence. It takes as input a typically large set of features measured on a collection of objects with two or more replicates per object. It returns an ordered list of features, i1,i2,…,ik, where feature i1 is the one with the highest reproducibility across replicates, i2 that with the highest reproducibility across replicates after projecting out the dimension spanned by i1, and so on. Applications to microscopy-based imaging of cells and proteomics highlight benefits of the approach. AVAILABILITY AND IMPLEMENTATION: The RNR method is available via Bioconductor (Huber W, Carey VJ, Gentleman R et al. (Orchestrating high-throughput genomic analysis with bioconductor. Nat Methods 2015;12:115-21.) in the R package FeatSeekR. Its source code is also available at https://github.com/tcapraz/FeatSeekR under the GPL-3 open source license. Tümay Capraz, Wolfgang Huber |
Bioinform. | 2 |
| 2023 | MsQuality: an interoperable open-source package for the calculation of standardized quality metrics of mass spectrometry dataabstractMOTIVATION: Multiple factors can impact accuracy and reproducibility of mass spectrometry data. There is a need to integrate quality assessment and control into data analytic workflows. RESULTS: The MsQuality package calculates 43 low-level quality metrics based on the controlled mzQC vocabulary defined by the HUPO-PSI on a single mass spectrometry-based measurement of a sample. It helps to identify low-quality measurements and track data quality. Its use of community-standard quality metrics facilitates comparability of quality assessment and control (QA/QC) criteria across datasets. AVAILABILITY AND IMPLEMENTATION: The R package MsQuality is available through Bioconductor at https://bioconductor.org/packages/MsQuality. Thomas Naake, Johannes Rainer, Wolfgang Huber |
Bioinform. | 3 |
| 2023 | htseq-clip: a toolset for the preprocessing of eCLIP/iCLIP datasetsabstractSUMMARY: Transcriptome-wide detection of binding sites of RNA-binding proteins is achieved using Individual-nucleotide crosslinking and immunoprecipitation (iCLIP) and its derivative enhanced CLIP (eCLIP) sequencing methods. Here, we introduce htseq-clip, a python package developed for preprocessing, extracting and summarizing crosslink site counts from i/eCLIP experimental data. The package delivers crosslink site count matrices along with other metrics, which can be directly used for filtering and downstream analyses such as the identification of differential binding sites. AVAILABILITY AND IMPLEMENTATION: The Python package htseq-clip is available via pypi (python package index), bioconda and the Galaxy Tool Shed under the open source MIT License. The code is hosted at https://github.com/EMBL-Hentze-group/htseq-clip and documentation is available under https://htseq-clip.readthedocs.io/en/latest. Sudeep Sahadevan, Thileepan Sekaran, Nadia Ashaf, Marko Fritz, Matthias W. Hentze, Wolfgang Huber, Thomas Schwarzl |
Bioinform. | 6 |
| 2022 | MatrixQCvis: shiny-based interactive data quality exploration for omics dataabstractMOTIVATION: First-line data quality assessment and exploratory data analysis are integral parts of any data analysis workflow. In high-throughput quantitative omics experiments (e.g. transcriptomics, proteomics and metabolomics), after initial processing, the data are typically presented as a matrix of numbers (feature IDs × samples). Efficient and standardized data quality metrics calculation and visualization are key to track the within-experiment quality of these rectangular data types and to guarantee for high-quality datasets and subsequent biological question-driven inference. RESULTS: We present MatrixQCvis, which provides interactive visualization of data quality metrics at the per-sample and per-feature level using R's shiny framework. It provides efficient and standardized ways to analyze data quality of quantitative omics data types that come in a matrix-like format (features IDs × samples). MatrixQCvis builds upon the Bioconductor SummarizedExperiment S4 class and thus facilitates the integration into existing workflows. AVAILABILITY AND IMPLEMENTATION: MatrixQCVis is implemented in R. It is available via Bioconductor and released under the GPL v3.0 license. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Thomas Naake, Wolfgang Huber |
Bioinform. | 2 |
| 2022 | Inferring tumor-specific cancer dependencies through integrating ex vivo drug response assays and drug-protein profilingabstractThe development of cancer therapies may be improved by the discovery of tumor-specific molecular dependencies. The requisite tools include genetic and chemical perturbations, each with its strengths and limitations. Chemical perturbations can be readily applied to primary cancer samples at large scale, but mechanistic understanding of hits and further pharmaceutical development is often complicated by the fact that a chemical compound has affinities to multiple proteins. To computationally infer specific molecular dependencies of individual cancers from their ex vivo drug sensitivity profiles, we developed a mathematical model that deconvolutes these data using measurements of protein-drug affinity profiles. Through integrating a drug-kinase profiling dataset and several drug response datasets, our method, DepInfeR, correctly identified known protein kinase dependencies, including the EGFR dependence of HER2+ breast cancer cell lines, the FLT3 dependence of acute myeloid leukemia (AML) with FLT3-ITD mutations and the differential dependencies on the B-cell receptor pathway in the two major subtypes of chronic lymphocytic leukemia (CLL). Furthermore, our method uncovered new subgroup-specific dependencies, including a previously unreported dependence of high-risk CLL on Checkpoint kinase 1 (CHEK1). The method also produced a detailed map of the kinase dependencies in a heterogeneous set of 117 CLL samples. The ability to deconvolute polypharmacological phenotypes into underlying causal molecular dependencies should increase the utility of high-throughput drug response assays for functional precision oncology. Alina Batzilla, Junyan Lu, Jarno Kivioja, Kerstin Putzker, Joe Lewis, Thorsten Zenz, Wolfgang Huber |
PLoS Comput. Biol. | 7 |
| 2021 | glmGamPoi: fitting Gamma-Poisson generalized linear models on single cell count dataabstractMOTIVATION: The Gamma-Poisson distribution is a theoretically and empirically motivated model for the sampling variability of single cell RNA-sequencing counts and an essential building block for analysis approaches including differential expression analysis, principal component analysis and factor analysis. Existing implementations for inferring its parameters from data often struggle with the size of single cell datasets, which can comprise millions of cells; at the same time, they do not take full advantage of the fact that zero and other small numbers are frequent in the data. These limitations have hampered uptake of the model, leaving room for statistically inferior approaches such as logarithm(-like) transformation. RESULTS: We present a new R package for fitting the Gamma-Poisson distribution to data with the characteristics of modern single cell datasets more quickly and more accurately than existing methods. The software can work with data on disk without having to load them into RAM simultaneously. AVAILABILITYAND IMPLEMENTATION: The package glmGamPoi is available from Bioconductor for Windows, macOS and Linux, and source code is available on github.com/const-ae/glmGamPoi under a GPL-3 license. The scripts to reproduce the results of this paper are available on github.com/const-ae/glmGamPoi-Paper. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Constantin Ahlmann-Eltze, Wolfgang Huber |
Bioinform. | 2 |
| 2015 | HTSeq - a Python framework to work with high-throughput sequencing dataabstractMOTIVATION: A large choice of tools exists for many standard tasks in the analysis of high-throughput sequencing (HTS) data. However, once a project deviates from standard workflows, custom scripts are needed. RESULTS: We present HTSeq, a Python library to facilitate the rapid development of such scripts. HTSeq offers parsers for many common data formats in HTS projects, as well as classes to represent data, such as genomic coordinates, sequences, sequencing reads, alignments, gene model information and variant calls, and provides data structures that allow for querying via genomic coordinates. We also present htseq-count, a tool developed with HTSeq that preprocesses RNA-Seq data for differential expression analysis by counting the overlap of reads with genes. AVAILABILITY AND IMPLEMENTATION: HTSeq is released as an open-source software under the GNU General Public Licence and available from http://www-huber.embl.de/HTSeq or from the Python Package Index at https://pypi.python.org/pypi/HTSeq. Simon Anders, Paul Theodor Pyl, Wolfgang Huber |
Bioinform. | 3 |
| 2015 | SomaticSignatures: inferring mutational signatures from single-nucleotide variantsabstractUNLABELLED: Mutational signatures are patterns in the occurrence of somatic single-nucleotide variants that can reflect underlying mutational processes. The SomaticSignatures package provides flexible, interoperable and easy-to-use tools that identify such signatures in cancer sequencing data. It facilitates large-scale, cross-dataset estimation of mutational signatures, implements existing methods for pattern decomposition, supports extension through user-defined approaches and integrates with existing Bioconductor workflows. AVAILABILITY AND IMPLEMENTATION: The R package SomaticSignatures is available as part of the Bioconductor project. Its documentation provides additional details on the methods and demonstrates applications to biological datasets. CONTACT: [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Julian Gehring, Bernd Fischer 0003, Michael F. Lawrence, Wolfgang Huber |
Bioinform. | 4 |
| 2015 | FourCSeq: analysis of 4C sequencing dataabstractMOTIVATION: Circularized Chromosome Conformation Capture (4C) is a powerful technique for studying the spatial interactions of a specific genomic region called the 'viewpoint' with the rest of the genome, both in a single condition or comparing different experimental conditions or cell types. Observed ligation frequencies typically show a strong, regular dependence on genomic distance from the viewpoint, on top of which specific interaction peaks are superimposed. Here, we address the computational task to find these specific peaks and to detect changes between different biological conditions. RESULTS: We model the overall trend of decreasing interaction frequency with genomic distance by fitting a smooth monotonically decreasing function to suitably transformed count data. Based on the fit, z-scores are calculated from the residuals, and high z-scores are interpreted as peaks providing evidence for specific interactions. To compare different conditions, we normalize fragment counts between samples, and call for differential contact frequencies using the statistical method DESEQ2: adapted from RNA-Seq analysis. AVAILABILITY AND IMPLEMENTATION: A full end-to-end analysis pipeline is implemented in the R package FourCSeq available at www.bioconductor.org. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Felix A. Klein, Tibor Pakozdi, Simon Anders, Yad Ghavi-Helm, Eileen E. M. Furlong, Wolfgang Huber |
Bioinform. | 6 |
| 2014 | h5vc: scalable nucleotide tallies with HDF5abstractSUMMARY: As applications of genome sequencing, including exomes and whole genomes, are expanding, there is a need for analysis tools that are scalable to large sets of samples and/or ultra-deep coverage. Many current tool chains are based on the widely used file formats BAM and VCF or VCF-derivatives. However, for some desirable analyses, data management with these formats creates substantial implementation overhead, and much time is spent parsing files and collating data. We observe that a tally data structure, i.e. the table of counts of nucleotides × samples × strands × genomic positions, provides a reasonable intermediate level of abstraction for many genomics analyses, including single nucleotide variant (SNV) and InDel calling, copy-number estimation and mutation spectrum analysis. Here we present h5vc, a data structure and associated software for managing tallies. The software contains functionality for creating tallies from BAM files, flexible and scalable data visualization, data quality assessment, computing statistics relevant to variant calling and other applications. Through the simplicity of its API, we envision making low-level analysis of large sets of genome sequencing data accessible to a wider range of researchers. AVAILABILITY AND IMPLEMENTATION: The package H5VC for the statistical environment R is available through the Bioconductor project. The HDF5 system is used as the core of our implementation. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Paul Theodor Pyl, Julian Gehring, Bernd Fischer 0003, Wolfgang Huber |
Bioinform. | 4 |
| 2013 | CellH5: a format for data exchange in high-content screeningabstractUNLABELLED: High-throughput microscopy data require a diversity of analytical approaches. However, the construction of workflows that use algorithms from different software packages is difficult owing to a lack of interoperability. To overcome this limitation, we present CellH5, an HDF5 data format for cell-based assays in high-throughput microscopy, which stores high-dimensional image data along with inter-object relations in graphs. CellH5Browser, an interactive gallery image browser, demonstrates the versatility and performance of the file format on live imaging data of dividing human cells. CellH5 provides new opportunities for integrated data analysis by multiple software platforms. AVAILABILITY: Source code is freely available at www.github.com/cellh5 under the GPL license and at www.bioconductor.org/packages/release/bioc/html/rhdf5.html under the Artistic-2.0 license. Demo datasets and the CellH5Browser are available at www.cellh5.org. A Fiji importer for cellh5 will be released soon. Christoph Sommer 0002, Michael Held, Bernd Fischer 0003, Wolfgang Huber, Daniel Gerlich |
Bioinform. | 4 |
| 2013 | Shrinkage estimation of dispersion in Negative Binomial models for RNA-seq experiments with small sample sizeabstractMOTIVATION: RNA-seq experiments produce digital counts of reads that are affected by both biological and technical variation. To distinguish the systematic changes in expression between conditions from noise, the counts are frequently modeled by the Negative Binomial distribution. However, in experiments with small sample size, the per-gene estimates of the dispersion parameter are unreliable. METHOD: We propose a simple and effective approach for estimating the dispersions. First, we obtain the initial estimates for each gene using the method of moments. Second, the estimates are regularized, i.e. shrunk towards a common value that minimizes the average squared difference between the initial estimates and the shrinkage estimates. The approach does not require extra modeling assumptions, is easy to compute and is compatible with the exact test of differential expression. RESULTS: We evaluated the proposed approach using 10 simulated and experimental datasets and compared its performance with that of currently popular packages edgeR, DESeq, baySeq, BBSeq and SAMseq. For these datasets, sSeq performed favorably for experiments with small sample size in sensitivity, specificity and computational time. AVAILABILITY: http://www.stat.purdue.edu/∼ovitek/Software.html and Bioconductor. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Danni Yu, Wolfgang Huber, Olga Vitek |
Bioinform. | 2 |
| 2013 | Dynamical modelling of phenotypes in a genome-wide RNAi live-cell imaging assayabstractBACKGROUND: The combination of time-lapse imaging of live cells with high-throughput perturbation assays is a powerful tool for genetics and cell biology. The Mitocheck project employed this technique to associate thousands of genes with transient biological phenotypes in cell division, cell death and migration. The original analysis of these data proceeded by assigning nuclear morphologies to cells at each time-point using automated image classification, followed by description of population frequencies and temporal distribution of cellular states through event-order maps. One of the choices made by that analysis was not to rely on temporal tracking of the individual cells, due to the relatively low image sampling frequency, and to focus on effects that could be discerned from population-level behaviour. RESULTS: Here, we present a variation of this approach that employs explicit modelling by dynamic differential equations of the cellular state populations. Model fitting to the time course data allowed reliable estimation of the penetrance and time of appearance of four types of disruption of the cell cycle: quiescence, mitotic arrest, polynucleation and cell death. Model parameters yielded estimates of the duration of the interphase and mitosis phases. We identified 2190 siRNAs that induced a disruption of the cell cycle at reproducible times, or increased the durations of the interphase or mitosis phases. CONCLUSIONS: We quantified the dynamic effects of the siRNAs and compiled them as a resource that can be used to characterize the role of their target genes in cell death, mitosis and cell cycle regulation. The described population-based modelling method might be applicable to other large-scale cell-based assays with temporal readout when only population-level measures are available. Grégoire Pau, Thomas Walter 0003, Beate Neumann, Jean-Karim Hériché, Jan Ellenberg, Wolfgang Huber |
BMC Bioinform. | 6 |
| 2013 | Software for Computing and Annotating Genomic RangesabstractWe describe Bioconductor infrastructure for representing and computing on annotated genomic ranges and integrating genomic data with the statistical computing features of R and its extensions. At the core of the infrastructure are three packages: IRanges, GenomicRanges, and GenomicFeatures. These packages provide scalable data structures for representing annotated ranges on the genome, with special support for transcript structures, read alignments and coverage vectors. Computational facilities include efficient algorithms for overlap and nearest neighbor detection, coverage calculation and other range operations. This infrastructure directly supports more than 80 other Bioconductor packages, including those for sequence analysis, differential expression analysis and visualization. Michael F. Lawrence, Wolfgang Huber, Hervé Pagès, Patrick Aboyoun, Marc Carlson, Robert Gentleman, Martin Morgan, Vincent Carey |
PLoS Comput. Biol. | 2 |
| 2011 | Extracting quantitative genetic interaction phenotypes from matrix combinatorial RNAiabstractBACKGROUND: Systematic measurement of genetic interactions by combinatorial RNAi (co-RNAi) is a powerful tool for mapping functional modules and discovering components. It also provides insights into the role of epistasis on the way from genotype to phenotype. The interpretation of co-RNAi data requires computational and statistical analysis in order to detect interactions reliably and sensitively. RESULTS: We present a comprehensive approach to the analysis of univariate phenotype measurements, such as cell growth. The method is based on a quantitative model and is demonstrated on two example Drosophila cell culture data sets. We discuss adjustments for technical variability, data quality assessment, model parameter fitting and fit diagnostics, choice of scale, and assessment of statistical significance. CONCLUSIONS: As a result, we obtain quantitative genetic interactions and interaction networks reflecting known biological relationships between target genes. The reliable extraction of presence, absence, and strength of interactions provides insights into molecular mechanisms. Elin Axelsson, Thomas Sandmann, Thomas Horn, Michael Boutros, Wolfgang Huber, Bernd Fischer 0003 |
BMC Bioinform. | 5 |
| 2011 | Assessments of Affymetrix GeneChip Microarray Quality for Laboratories and Single SamplesabstractBACKGROUND: Microarray technology has become a widely used tool in the biological sciences. Over the past decade, the number of users has grown exponentially, and with the number of applications and secondary data analyses rapidly increasing, we expect this rate to continue. Various initiatives such as the External RNA Control Consortium (ERCC) and the MicroArray Quality Control (MAQC) project have explored ways to provide standards for the technology. For microarrays to become generally accepted as a reliable technology, statistical methods for assessing quality will be an indispensable component; however, there remains a lack of consensus in both defining and measuring microarray quality. RESULTS: We begin by providing a precise definition of microarray quality and reviewing existing Affymetrix GeneChip quality metrics in light of this definition. We show that the best-performing metrics require multiple arrays to be assessed simultaneously. While such multi-array quality metrics are adequate for bench science, as microarrays begin to be used in clinical settings, single-array quality metrics will be indispensable. To this end, we define a single-array version of one of the best multi-array quality metrics and show that this metric performs as well as the best multi-array metrics. We then use this new quality metric to assess the quality of microarry data available via the Gene Expression Omnibus (GEO) using more than 22,000 Affymetrix HGU133a and HGU133plus2 arrays from 809 studies. CONCLUSIONS: We find that approximately 10 percent of these publicly available arrays are of poor quality. Moreover, the quality of microarray measurements varies greatly from hybridization to hybridization, study to study, and lab to lab, with some experiments producing unusable data. Many of the concepts described here are applicable to other high-throughput technologies. Matthew N. McCall, Peter N. Murakami, Margus Lukk, Wolfgang Huber, Rafael A. Irizarry |
BMC Bioinform. | 4 |
| 2010 | EBImage - an R package for image processing with applications to cellular phenotypesabstractSUMMARY: EBImage provides general purpose functionality for reading, writing, processing and analysis of images. Furthermore, in the context of microscopy-based cellular assays, EBImage offers tools to segment cells and extract quantitative cellular descriptors. This allows the automation of such tasks using the R programming language and use of existing tools in the R environment for signal processing, statistical modeling, machine learning and data visualization. AVAILABILITY: EBImage is free and open source, released under the LGPL license and available from the Bioconductor project (http://www.bioconductor.org/packages/release/bioc/html/EBImage.html). Grégoire Pau, Florian Fuchs 0003, Oleg Sklyar, Michael Boutros, Wolfgang Huber |
Bioinform. | 5 |
| 2009 | Array-based genotyping in S.cerevisiae using semi-supervised clusteringabstractAbstract Motivation: Microarrays provide an accurate and cost-effective method for genotyping large numbers of individuals at high resolution. The resulting data permit the identification of loci at which genetic variation is associated with quantitative traits, or fine mapping of meiotic recombination, which is a key determinant of genetic diversity among individuals. Several issues inherent to short oligonucleotide arrays—cross-hybridization, or variability in probe response to target—have the potential to produce genotyping errors. There is a need for improved statistical methods for array-based genotyping. Results: We developed ssGenotyping (ssG), a multivariate, semi-supervised approach for using microarrays to genotype haploid individuals at thousands of polymorphic sites. Using a meiotic recombination dataset, we show that ssG is more accurate than existing supervised classification methods, and that it produces denser marker coverage. The ssG algorithm is able to fit probe-specific affinity differences and to detect and filter spurious signal, permitting high-confidence genotyping at nucleotide resolution. We also demonstrate that oligonucleotide probe response depends significantly on genomic background, even when the probe's specific target sequence is unchanged. As a result, supervised classifiers trained on reference strains may not generalize well to diverged strains; ssG's semi-supervised approach, on the other hand, adapts automatically. Availability: The ssGenotyping software is implemented in R. It is currently available for download (www.ebi.ac.uk/∼bourgon/yeast_genotyping/ssG) and is being submitted to Bioconductor. Contact: [email protected] Supplementary information: Supplementary data and a version including color figures are available at Bioinformatics online. Richard Bourgon, Eugenio Mancera, Alessandro Brozzi, Lars M. Steinmetz, Wolfgang Huber |
Bioinform. | 5 |
| 2009 | arrayQualityMetrics - a bioconductor package for quality assessment of microarray dataabstractSUMMARY: The assessment of data quality is a major concern in microarray analysis. arrayQualityMetrics is a Bioconductor package that provides a report with diagnostic plots for one or two colour microarray data. The quality metrics assess reproducibility, identify apparent outlier arrays and compute measures of signal-to-noise ratio. The tool handles most current microarray technologies and is amenable to use in automated analysis pipelines or for automatic report generation, as well as for use by individuals. The diagnosis of quality remains, in principle, a context-dependent judgement, but our tool provides powerful, automated, objective and comprehensive instruments on which to base a decision. AVAILABILITY: arrayQualityMetrics is a free and open source package, under LGPL license, available from the Bioconductor project at www.bioconductor.org. A users guide and examples are provided with the package. Some examples of HTML reports generated by arrayQualityMetrics can be found at http://www.microarray-quality.org Audrey Kauffmann, Robert Gentleman, Wolfgang Huber |
Bioinform. | 3 |
| 2009 | Importing ArrayExpress datasets into R/BioconductorabstractAbstract Summary:ArrayExpress is one of the largest public repositories of microarray datasets. R/Bioconductor provides a comprehensive suite of microarray analysis and integrative bioinformatics software. However, easy ways for importing datasets from ArrayExpress into R/Bioconductor have been lacking. Here, we present such a tool that is suitable for both interactive and automated use. Availability: The ArrayExpress package is available from the Bioconductor project at http://www.bioconductor.org. A users guide and examples are provided with the package. Contact: [email protected] Supplementary information: Supplementary data are available Bioinformatics online. Audrey Kauffmann, Tim F. Rayner, Helen E. Parkinson, Misha Kapushesky, Margus Lukk, Alvis Brazma, Wolfgang Huber |
Bioinform. | 7 |
| 2008 | Rintact: enabling computational analysis of molecular interaction data from the IntAct repositoryabstractMOTIVATION: The IntAct repository is one of the largest and most widely used databases for the curation and storage of molecular interaction data. These datasets need to be analyzed by computational methods. Software packages in the statistical environment R provide powerful tools for conducting such analyses. RESULTS: We introduce Rintact, a Bioconductor package that allows users to transform PSI-MI XML2.5 interaction data files from IntAct into R graph objects. On these, they can use methods from R and Bioconductor for a variety of tasks: determining cohesive subgraphs, computing summary statistics, fitting mathematical models to the data or rendering graphical layouts. Rintact provides a programmatic interface to the IntAct repository and allows the use of the analytic methods provided by R and Bioconductor. AVAILABILITY: Rintact is freely available at http://bioconductor.org Tony Chiang, Nianhua Li, Sandra E. Orchard, Samuel Kerrien, Henning Hermjakob, Robert Gentleman, Wolfgang Huber |
Bioinform. | 7 |
| 2008 | Estimating node degree in bait-prey graphsabstractMOTIVATION: Proteins work together to drive biological processes in cellular machines. Summarizing global and local properties of the set of protein interactions, the interactome, is necessary for describing cellular systems. We consider a relatively simple per-protein feature of the interactome: the number of interaction partners for a protein, which in graph terminology is the degree of the protein. RESULTS: Using data subject to both stochastic and systematic sources of false positive and false negative observations, we develop an explicit probability model and resultant likelihood method to estimate node degree on portions of the interactome assayed by bait-prey technologies. This approach yields substantial improvement in degree estimation over the current practice that naively sums observed edges. Accurate modeling of observed data in relation to true but unknown parameters of interest gives a formal point of reference from which to draw conclusions about the system under study. AVAILABILITY: All analyses discussed in this text can be performed using the ppiStats and ppiData packages available through the Bioconductor project (http://www.bioconductor.org). Denise M. Scholtens, Tony Chiang, Wolfgang Huber, Robert Gentleman |
Bioinform. | 3 |
| 2008 | Analyzing ChIP-chip Data Using BioconductorabstractChIP-chip, chromatin immunoprecipitation combined with DNA microarrays, is a widely used assay for DNA–protein binding and chromatin plasticity, which are of fundamental interest for the understanding of gene regulation.
The interpretation of ChIP-chip data poses two computational challenges: first, what can be termed primary statistical analysis, which includes quality assessment, data normalization and transformation, and the calling of regions of interest; second, integrative bioinformatic analysis, which interprets the data in the context of existing genome annotation and of related experimental results obtained, for example, from other ChIP-chip or (m)RNA abundance microarray experiments.
Both tasks rely heavily on visualization, which helps to explore the data as well as to present the analysis results. For the primary statistical analysis, some standardization is possible and desirable: commonly used experimental designs and microarray platforms allow the development of relatively standard workflows and statistical procedures. Most software available for ChIP-chip data analysis can be employed in such standardized approaches [1]–[6]. Yet even for primary analysis steps, it may be beneficial to adapt them to specific experiments, and hence it is desirable that software offers flexibility in the choice of algorithms for normalization, visualization, and identification of enriched regions.
For the second task, integrative bioinformatic analysis, the datasets, questions, and applicable methods are diverse, and a degree of flexibility is needed that often can only be achieved in a programmable environment. In such an environment, users are not limited to predefined functions, such as the ones made available as “buttons” in a GUI, but can supply custom functions that are designed toward the analysis at hand.
Bioconductor [7] is an open source and open development software project for the analysis and comprehension of genomic data, and it offers tools that cover a broad range of computational methods, visualizations, and experimental data types, and is designed to allow the construction of scalable, reproducible, and interoperable workflows. A consequence of the wide range of functionality of Bioconductor and its concurrency with research progress in biology and computational statistics is that using its tools can be daunting for a new user. Various books provide a good general introduction to R and Bioconductor (e.g., [8]–[10]), and most Bioconductor packages are accompanied by extensive documentation. This tutorial covers basic ChIP-chip data analysis with Bioconductor. Among the packages used are Ringo [5], biomaRt [11], and topGO [12].
We wrote this document in the Sweave [13] format, which combines explanatory text and the actual R source code used in this analysis [14]. Thus, the analysis can be reproduced by the reader. An R package ccTutorial that contains the data, the text, and code presented here, and supplementary text and code, is available from the Bioconductor Web site.
> library(“Ringo”)
> library(“biomaRt”)
> library(“topGO”)
> library(“ccTutorial”)
Terminology. Reporters are the DNA sequences fixed to the microarray; they are designed to specifically hybridize with corresponding genomic fragments from the immunoprecipitate. A reporter has a unique identifier and a unique sequence, and it can appear in one or multiple features on the array surface [15]. The sample is the aliquot of immunoprecipitated or input DNA that is hybridized to the microarray. We shall call a genomic region apparently enriched by ChIP a ChIP-enriched region.
The data. We consider a ChIP-chip dataset on a post-translational modification of histone protein H3, namely tri-methylation of its Lysine residue 4, in short H3K4me3. H3K4me3 has been associated with active transcription (e.g., [16],[17]). Here, enrichment for H3K4me3 was investigated in Mus musculus brain and heart cells. The microarray platform is a set of four arrays manufactured by NimbleGen containing 390 k reporters each. The reporters were designed to tile 32,482 selected regions of the Mus musculus genome (assembly mm5) with one base every 100 bp, with a different set of promoters represented on each of the four arrays ([18], Methods: Condensed array ChIP-chip). We obtained the data from the GEO repository [19] (accession {type:entrez-geo,attrs:{text:GSE7688,term_id:7688}}GSE7688). Joern Toedling, Wolfgang Huber |
PLoS Comput. Biol. | 2 |
| 2007 | CoCo: a web application to display, store and curate ChIP-on-chip data integrated with diverse types of gene expression dataabstractMOTIVATION: CoCo, ChIP-on-Chip online, is an open-source web application that supports the annotation and curation of regulatory regions and associated target genes discovered in ChIP-on-chip experiments. CoCo integrates ChIP-on-chip results with diverse types of gene expression data (expression profiling, in situ hybridization) and displays them within a genomic context. Regulatory relationships between the transcription factor-bound regions and putative target genes can be stored and expanded throughout different sessions. AVAILABILITY: http://furlonglab.embl.de/methods/tools/coco. Charles Girardot, Oleg Sklyar, Sophie Grosz, Wolfgang Huber, Eileen E. M. Furlong |
Bioinform. | 4 |
| 2007 | In situ analysis of cross-hybridisation on microarrays and the inference of expression correlationabstractBACKGROUND: Microarray co-expression signatures are an important tool for studying gene function and relations between genes. In addition to genuine biological co-expression, correlated signals can result from technical deficiencies like hybridization of reporters with off-target transcripts. An approach that is able to distinguish these factors permits the detection of more biologically relevant co-expression signatures. RESULTS: We demonstrate a positive relation between off-target reporter alignment strength and expression correlation in data from oligonucleotide genechips. Furthermore, we describe a method that allows the identification, from their expression data, of individual probe sets affected by off-target hybridization. CONCLUSION: The effects of off-target hybridization on expression correlation coefficients can be substantial, and can be alleviated by more accurate mapping between microarray reporters and the target transcriptome. We recommend attention to the mapping for any microarray analysis of gene expression patterns. Tineke Casneuf, Yves Van de Peer, Wolfgang Huber |
BMC Bioinform. | 3 |
| 2007 | Graphs in molecular biologyabstractGraph theoretical concepts are useful for the description and analysis of interactions and relationships in biological systems. We give a brief introduction into some of the concepts and their areas of application in molecular biology. We discuss software that is available through the Bioconductor project and present a simple example application to the integration of a protein-protein interaction and a co-expression network. Wolfgang Huber, Vincent Carey, Li Long, Seth Falcon, Robert Gentleman |
BMC Bioinform. | 1 |
| 2007 | Ringo - an R/Bioconductor package for analyzing ChIP-chip readoutsabstractBACKGROUND: Chromatin immunoprecipitation combined with DNA microarrays (ChIP-chip) is a high-throughput assay for DNA-protein-binding or post-translational chromatin/histone modifications. However, the raw microarray intensity readings themselves are not immediately useful to researchers, but require a number of bioinformatic analysis steps. Identified enriched regions need to be bioinformatically annotated and compared to related datasets by statistical methods. RESULTS: We present a free, open-source R package Ringo that facilitates the analysis of ChIP-chip experiments by providing functionality for data import, quality assessment, normalization and visualization of the data, and the detection of ChIP-enriched genomic regions. CONCLUSION: Ringo integrates with other packages of the Bioconductor project, uses common data structures and is accompanied by ample documentation. It facilitates the construction of programmed analysis workflows, offers benefits in scalability, reproducibility and methodical scope of the analyses and opens up a broad selection of follow-up statistical and bioinformatic methods. Joern Toedling, Oleg Sklyar, Wolfgang Huber |
BMC Bioinform. | 3 |
| 2007 | Ringo - an R/Bioconductor package for analyzing ChIP-chip readoutsabstractDOAJ is a unique and extensive index of diverse open access journals from around the world, driven by a growing community, committed to ensuring quality content is freely available online for everyone. Joern Toedling, Oleg Sklyar, Tammo Krueger, Jenny J. Fischer, Silke Sperling, Wolfgang Huber |
BMC Bioinform. | 6 |
| 2006 | Transcript mapping with high-density oligonucleotide tiling arraysabstractMOTIVATION: High-density DNA tiling microarrays are a powerful tool for the characterization of complete transcriptomes. The two major analytical challenges are the segmentation of the hybridization signal along genomic coordinates to accurately determine transcript boundaries and the adjustment of the sequence-dependent response of the oligonucleotide probes to achieve quantitative comparability of the signal between different probes. RESULTS: We describe a dynamic programming algorithm for finding a globally optimal fit of a piecewise constant expression profile along genomic coordinates. We developed a probe-specific background correction and scaling method that employs empirical probe response parameters determined from reference hybridizations with no need for paired mismatch probes. This combined analysis approach allows the accurate determination of dynamical changes in transcription architectures from hybridization data and will help to study the biological significance of complex transcriptional phenomena in eukaryotic genomes. AVAILABILITY: R package tilingArray at http://www.bioconductor.org. Wolfgang Huber, Joern Toedling, Lars M. Steinmetz |
Bioinform. | 1 |
| 2005 | arrayMagic: two-colour cDNA microarray quality control, preprocessingabstractUNLABELLED: arrayMagic is a software package for quality control and preprocessing of two-colour cDNA microarray data. The automated analysis pipeline comprises data import, normalization, replica merging, quality diagnostics and data export. The script-based processing combines reproducibility and flexibility at high-throughput and provides quality-assured and preprocessed microarray data to high-level follow-up analysis. AVAILABILITY: The R package arrayMagic is available with BSD license at http://www.bioconductor.org CONTACT: [email protected] SUPPLEMENTARY INFORMATION: The package contains documentation in the form of manual pages and a vignette with a guided tour of a typical workflow. Andreas Buness, Wolfgang Huber, Klaus Steiner, Holger Sültmann, Annemarie Poustka |
Bioinform. | 2 |
| 2005 | BioMart and Bioconductor: a powerful link between biological databases and microarray data analysisabstractbiomaRt is a new Bioconductor package that integrates BioMart data resources with data analysis software in Bioconductor. It can annotate a wide range of gene or gene product identifiers (e.g. Entrez-Gene and Affymetrix probe identifiers) with information such as gene symbol, chromosomal coordinates, Gene Ontology and OMIM annotation. Furthermore biomaRt enables retrieval of genomic sequences and single nucleotide polymorphism information, which can be used in data analysis. Fast and up-to-date data retrieval is possible as the package executes direct SQL queries to the BioMart databases (e.g. Ensembl). The biomaRt package provides a tight integration of large, public or locally installed BioMart databases with data analysis in Bioconductor creating a powerful environment for biological data mining. Steffen Durinck, Yves Moreau, Arek Kasprzyk, Sean R. Davis, Bart De Moor, Alvis Brazma, Wolfgang Huber |
Bioinform. | 7 |
| 2004 | matchprobes: a Bioconductor package for the sequence-matching of microarray probe elementsabstractUNLABELLED: The nucleotide sequences of the probes on a microarray can be used for a variety of purposes in the analysis of microarray experiments. We describe software and a paradigm for the creation of data packages for curating, distributing and working with probe sequence data in a uniform, across-types-of-microarrays manner. While the implementation is specific to the Bioconductor project, the ideas and general strategies are more general and could be easily adopted by other projects. AVAILABILITY: The R package matchprobes is available under LGPL at http://www.bioconductor.org SUPPLEMENTARY INFORMATION: The package contains documentation in the form of a vignette and manual pages. Wolfgang Huber, Robert Gentleman |
Bioinform. | 1 |
| 2003 | Processor Support for Temporal Predictability - The SPEAR Design ExampleabstractThe demand for predictable timing behavior is characteristic for real-time applications. Experience has shown that this property cannot be achieved by software alone but rather requires support from the processor. This situation is analyzed and mapped to a design rationale for SPEAR (Scalable Processor for Embedded Applications in Real-time Environments), a processor that has been designed to meet the specific temporal demands of real-time systems. At the hardware level, SPEAR guarantees interrupt response with minimum temporal jitter and minimum delay. Furthermore, the processor provides an instruction set that only has constant-time instructions. At the software level, SPEAR supports the implementation of temporally predictable code according to the single-path programming paradigm. Altogether, these features support writing of code with minimal jitter and provide the basis for exact temporal predictability. Experimental results show that SPEAR indeed exhibits the anticipated highly predictable timing behavior. Martin Delvai, Wolfgang Huber, Peter P. Puschner, Andreas Steininger |
ECRTS | 2 |
| 2002 | Variance stabilization applied to microarray data calibration and to the quantification of differential expressionabstractWe introduce a statistical model for microarray gene expression data that comprises data calibration, the quantification of differential expression, and the quantification of measurement error. In particular, we derive a transformation h for intensity measurements, and a difference statistic Deltah whose variance is approximately constant along the whole intensity range. This forms a basis for statistical inference from microarray data, and provides a rational data pre-processing strategy for multivariate analyses. For the transformation h, the parametric form h(x)=arsinh(a+bx) is derived from a model of the variance-versus-mean dependence for microarray intensity data, using the method of variance stabilizing transformations. For large intensities, h coincides with the logarithmic transformation, and Deltah with the log-ratio. The parameters of h together with those of the calibration between experiments are estimated with a robust variant of maximum-likelihood estimation. We demonstrate our approach on data sets from different experimental platforms, including two-colour cDNA arrays and a series of Affymetrix oligonucleotide arrays. Wolfgang Huber, Anja von Heydebreck, Holger Sültmann, Annemarie Poustka, Martin Vingron |
ISMB | 1 |