VLDB 2026 Research / reviewers in the wild / expert
Rafael A. Irizarry
dblp:64/2364
· DBLP profile ↗
28ranked-venue papers
1as first author
0since 2021 · last 2014
0000-0002-3944-4309ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 25 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
16 papers |
Bioinformatics and computational biology · 97% Computational science and engineering · 3% |
Topics — the 23 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › gene expression analysis
microarray data analysis |
0.5 | 5 | 2012 | fRMA ST: frozen robust multiarray analysis for Affymetrix Exon and Gene ST arrays · Bioinform. 2012 A framework for oligonucleotide microarray preprocessing · Bioinform. 2010 R/Bioconductor software for Illumina's Infinium whole-genome genotyping BeadChips · Bioinform. 2009 |
Bioinformatics and computational biology
gene expression analysis |
0.4 | 7 | 2013 | ChIP-PED enhances the analysis of ChIP-seq and ChIP-chip data · Bioinform. 2013 Comparison of Affymetrix GeneChip expression measures · Bioinform. 2006 affy - analysis of Affymetrix GeneChip data at the probe level · Bioinform. 2004 |
Bioinformatics and computational biology › statistical genetics
genotype calling |
0.2 | 2 | 2010 | Quantifying uncertainty in genotype calls · Bioinform. 2010 R/Bioconductor software for Illumina's Infinium whole-genome genotyping BeadChips · Bioinform. 2009 |
Bioinformatics and computational biology › epigenomics › DNA methylation
DNA methylation analysis |
0.2 | 1 | 2014 | Minfi: a flexible and comprehensive Bioconductor package for the analysis of Infinium DNA methylation microarrays · Bioinform. 2014 |
Bioinformatics and computational biology
epigenomics |
0.2 | 1 | 2014 | Minfi: a flexible and comprehensive Bioconductor package for the analysis of Infinium DNA methylation microarrays · Bioinform. 2014 |
Bioinformatics and computational biology › sequence analysis
genomic sequence analysis |
0.2 | 1 | 2014 | Visualization and probability-based scoring of structural variants within repetitive sequences · Bioinform. 2014 |
Bioinformatics and computational biology › genomics › structural variation
structural variant detection |
0.2 | 1 | 2014 | Visualization and probability-based scoring of structural variants within repetitive sequences · Bioinform. 2014 |
Bioinformatics and computational biology
genomics |
0.2 | 2 | 2011 | Performance assessment of copy number microarray platforms using a spike-in experiment · Bioinform. 2011 Stochastic models inspired by hybridization theory for short oligonucleotide arrays · RECOMB 2004 |
Bioinformatics and computational biology › epigenomics
ChIP-seq analysis |
0.2 | 1 | 2013 | ChIP-PED enhances the analysis of ChIP-seq and ChIP-chip data · Bioinform. 2013 |
Bioinformatics and computational biology › epigenomics › ChIP-seq analysis
ChIP-seq data integration |
0.2 | 1 | 2013 | ChIP-PED enhances the analysis of ChIP-seq and ChIP-chip data · Bioinform. 2013 |
Bioinformatics and computational biology › genomics
genotyping |
0.1 | 2 | 2010 | R/Bioconductor software for Illumina's Infinium whole-genome genotyping BeadChips · Bioinform. 2009 Quantifying uncertainty in genotype calls · Bioinform. 2010 |
Bioinformatics and computational biology › genomics › structural variation
copy number variant analysis |
0.1 | 1 | 2011 | Performance assessment of copy number microarray platforms using a spike-in experiment · Bioinform. 2011 |
Computational science and engineering
uncertainty quantification |
0.1 | 1 | 2010 | Quantifying uncertainty in genotype calls · Bioinform. 2010 |
Bioinformatics and computational biology › genomics
single nucleotide polymorphism |
0.1 | 1 | 2009 | R/Bioconductor software for Illumina's Infinium whole-genome genotyping BeadChips · Bioinform. 2009 |
Bioinformatics and computational biology › cancer genomics
copy number analysis |
0.1 | 1 | 2008 | Estimation and assessment of raw copy numbers at the single locus level · Bioinform. 2008 |
Bioinformatics and computational biology › genomics › structural variation
copy number variation |
0.1 | 1 | 2007 | Estimating Genome-Wide Copy Number Using Allele Specific Mixture Models · RECOMB 2007 |
Bioinformatics and computational biology › gene expression analysis
microarray data preprocessing |
0.1 | 1 | 2006 | Comparison of Affymetrix GeneChip expression measures · Bioinform. 2006 |
Bioinformatics and computational biology › epigenomics › differential methylation analysis
differentially methylated region detection |
0.1 | 1 | 2014 | Minfi: a flexible and comprehensive Bioconductor package for the analysis of Infinium DNA methylation microarrays · Bioinform. 2014 |
Bioinformatics and computational biology › gene expression analysis › microarray data preprocessing
probe-level analysis |
0.0 | 1 | 2004 | affy - analysis of Affymetrix GeneChip data at the probe level · Bioinform. 2004 |
Bioinformatics and computational biology
transcriptomics |
0.0 | 1 | 2004 | Stochastic models inspired by hybridization theory for short oligonucleotide arrays · RECOMB 2004 |
Bioinformatics and computational biology › gene expression analysis › microarray data analysis
microarray gene expression analysis |
0.0 | 1 | 2003 | A comparison of normalization methods for high density oligonucleotide array data based on variance and bias · Bioinform. 2003 |
Bioinformatics and computational biology › gene expression analysis › microarray data preprocessing › probe-level analysis
probe signal normalization |
0.0 | 1 | 2003 | A comparison of normalization methods for high density oligonucleotide array data based on variance and bias · Bioinform. 2003 |
Bioinformatics and computational biology › genomics
genome-wide association study |
0.0 | 1 | 2010 | Quantifying uncertainty in genotype calls · Bioinform. 2010 |
Methods — techniques the papers use, named apart from their topics
probe-level model · 0.2visualization · 0.2statistical preprocessing · 0.2quality assessment · 0.2probability-based scoring · 0.2permutation testing · 0.2kernel smoothing · 0.2batch effect correction · 0.1spike-in experiment · 0.1sensitivity and specificity analysis · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Context Aware Group Nearest Shrunken Centroids in Large-Scale Genomic StudiesabstractRecent genomic studies have identified genes related to specific phenotypes. In addition to marginal association analysis for individual genes, analyzing gene pathways (functionally related sets of genes) may yield additional valuable insights. We have devised an approach to phenotype classification from gene expression profiling. Our method named “group Nearest Shrunken Centroids (gNSC)” is an enhancement of the Nearest Shrunken Centroids (NSC) which is a popular and scalable method to analyze big data. While fully utilizing the variable structure of gene pathways, gNSC shares comparable computational speed as NSC if the group size is small. Comparing with NSC, gNSC improves the power of classification by utilizing the gene pathway information. In practice, we investigate the performance of gNSC on one of the largest microarray datasets aggregated from the internet. We show the effectiveness of our method by comparing the misclassification rate of gNSC with that of NSC. Additionally, we present a novel application of NSC/gNSC on context analysis of association between pathways and certain medical words. Some newest biological findings are rediscovered. Juemin Yang, Rafael A. Irizarry, Han Liu 0001 |
AISTATS | 3 |
| 2014 | Minfi: a flexible and comprehensive Bioconductor package for the analysis of Infinium DNA methylation microarraysabstractMOTIVATION: The recently released Infinium HumanMethylation450 array (the '450k' array) provides a high-throughput assay to quantify DNA methylation (DNAm) at ∼450 000 loci across a range of genomic features. Although less comprehensive than high-throughput sequencing-based techniques, this product is more cost-effective and promises to be the most widely used DNAm high-throughput measurement technology over the next several years. RESULTS: Here we describe a suite of computational tools that incorporate state-of-the-art statistical techniques for the analysis of DNAm data. The software is structured to easily adapt to future versions of the technology. We include methods for preprocessing, quality assessment and detection of differentially methylated regions from the kilobase to the megabase scale. We show how our software provides a powerful and flexible development platform for future methods. We also illustrate how our methods empower the technology to make discoveries previously thought to be possible only with sequencing-based methods. AVAILABILITY AND IMPLEMENTATION: http://bioconductor.org/packages/release/bioc/html/minfi.html. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Martin J. Aryee, Andrew E. Jaffe, Héctor Corrada Bravo, Christine Ladd-Acosta, Andrew P. Feinberg, Kasper D. Hansen, Rafael A. Irizarry |
Bioinform. | 7 |
| 2014 | Visualization and probability-based scoring of structural variants within repetitive sequencesabstractMOTIVATION: Repetitive sequences account for approximately half of the human genome. Accurately ascertaining sequences in these regions with next generation sequencers is challenging, and requires a different set of analytical techniques than for reads originating from unique sequences. Complicating the matter are repetitive regions subject to programmed rearrangements, as is the case with the antigen-binding domains in the Immunoglobulin (Ig) and T-cell receptor (TCR) loci. RESULTS: We developed a probability-based score and visualization method to aid in distinguishing true structural variants from alignment artifacts. We demonstrate the usefulness of this method in its ability to separate real structural variants from false positives generated with existing upstream analysis tools. We validated our approach using both target-capture and whole-genome experiments. Capture sequencing reads were generated from primary lymphoid tumors, cancer cell lines and an EBV-transformed lymphoblast cell line over the Ig and TCR loci. Whole-genome sequencing reads were from a lymphoblastoid cell-line. AVAILABILITY: We implement our method as an R package available at https://github.com/Eitan177/targetSeqView. Code to reproduce the figures and results are also available. Eitan Halper-Stromberg, Jared Steranka, Kathleen H. Burns, Sarven Sabunciyan, Rafael A. Irizarry |
Bioinform. | 5 |
| 2014 | KRLMM: an adaptive genotype calling method for common and low frequency variantsabstractBACKGROUND: SNP genotyping microarrays have revolutionized the study of complex disease. The current range of commercially available genotyping products contain extensive catalogues of low frequency and rare variants. Existing SNP calling algorithms have difficulty dealing with these low frequency variants, as the underlying models rely on each genotype having a reasonable number of observations to ensure accurate clustering. RESULTS: Here we develop KRLMM, a new method for converting raw intensities into genotype calls that aims to overcome this issue. Our method is unique in that it applies careful between sample normalization and allows a variable number of clusters k (1, 2 or 3) for each SNP, where k is predicted using the available data. We compare our method to four genotyping algorithms (GenCall, GenoSNP, Illuminus and OptiCall) on several Illumina data sets that include samples from the HapMap project where the true genotypes are known in advance. All methods were found to have high overall accuracy (> 98%), with KRLMM consistently amongst the best. At low minor allele frequency, the KRLMM, OptiCall and GenoSNP algorithms were observed to be consistently more accurate than GenCall and Illuminus on our test data. CONCLUSIONS: Methods that tailor their approach to calling low frequency variants by either varying the number of clusters (KRLMM) or using information from other SNPs (OptiCall and GenoSNP) offer improved accuracy over methods that do not (GenCall and Illuminus). The KRLMM algorithm is implemented in the open-source crlmm package distributed via the Bioconductor project (http://www.bioconductor.org). Zhiyin Dai, Meredith Yeager, Rafael A. Irizarry, Matthew E. Ritchie |
BMC Bioinform. | 4 |
| 2013 | ChIP-PED enhances the analysis of ChIP-seq and ChIP-chip dataabstractMOTIVATION: Although chromatin immunoprecipitation coupled with high-throughput sequencing (ChIP-seq) or tiling array hybridization (ChIP-chip) is increasingly used to map genome-wide-binding sites of transcription factors (TFs), it still remains difficult to generate a quality ChIPx (i.e. ChIP-seq or ChIP-chip) dataset because of the tremendous amount of effort required to develop effective antibodies and efficient protocols. Moreover, most laboratories are unable to easily obtain ChIPx data for one or more TF(s) in more than a handful of biological contexts. Thus, standard ChIPx analyses primarily focus on analyzing data from one experiment, and the discoveries are restricted to a specific biological context. RESULTS: We propose to enrich this existing data analysis paradigm by developing a novel approach, ChIP-PED, which superimposes ChIPx data on large amounts of publicly available human and mouse gene expression data containing a diverse collection of cell types, tissues and disease conditions to discover new biological contexts with potential TF regulatory activities. We demonstrate ChIP-PED using a number of examples, including a novel discovery that MYC, a human TF, plays an important functional role in pediatric Ewing sarcoma cell lines. These examples show that ChIP-PED increases the value of ChIPx data by allowing one to expand the scope of possible discoveries made from a ChIPx experiment. AVAILABILITY: http://www.biostat.jhsph.edu/~gewu/ChIPPED/ George Wu, Jason T. Yustein, Matthew N. McCall, Michael J. Zilliox, Rafael A. Irizarry, Karen Zeller, Chi V. Dang, Hongkai Ji |
Bioinform. | 5 |
| 2012 | fRMA ST: frozen robust multiarray analysis for Affymetrix Exon and Gene ST arraysabstractSUMMARY: Frozen robust multiarray analysis (fRMA) is a single-array preprocessing algorithm that retains the advantages of multiarray algorithms and removes certain batch effects by downweighting probes that have high between-batch residual variance. Here, we extend the fRMA algorithm to two new microarray platforms--Affymetrix Human Exon and Gene 1.0 ST--by modifying the fRMA probe-level model and extending the frma package to work with oligo ExonFeatureSet and GeneFeatureSet objects. AVAILABILITY AND IMPLEMENTATION: All packages are implemented in R. Source code and binaries are freely available through the Bioconductor project. Convenient links to all software and data packages can be found at http://mnmccall.com/software CONTACT: [email protected]. Matthew N. McCall, Harris A. Jaffee, Rafael A. Irizarry |
Bioinform. | 3 |
| 2012 | Improved base-calling and quality scores for 454 sequencing based on a Hurdle Poisson modelabstractBACKGROUND: 454 pyrosequencing is a commonly used massively parallel DNA sequencing technology with a wide variety of application fields such as epigenetics, metagenomics and transcriptomics. A well-known problem of this platform is its sensitivity to base-calling insertion and deletion errors, particularly in the presence of long homopolymers. In addition, the base-call quality scores are not informative with respect to whether an insertion or a deletion error is more likely. Surprisingly, not much effort has been devoted to the development of improved base-calling methods and more intuitive quality scores for this platform. RESULTS: We present HPCall, a 454 base-calling method based on a weighted Hurdle Poisson model. HPCall uses a probabilistic framework to call the homopolymer lengths in the sequence by modeling well-known 454 noise predictors. Base-calling quality is assessed based on estimated probabilities for each homopolymer length, which are easily transformed to useful quality scores. CONCLUSIONS: Using a reference data set of the Escherichia coli K-12 strain, we show that HPCall produces superior quality scores that are very informative towards possible insertion and deletion errors, while maintaining a base-calling accuracy that is better than the current one. Given the generality of the framework, HPCall has the potential to also adapt to other homopolymer-sensitive sequencing technologies. Kristof De Beuf, Joachim M. De Schrijver, Olivier Thas, Wim Van Criekinge, Rafael A. Irizarry, Lieven Clement |
BMC Bioinform. | 5 |
| 2012 | Gene expression anti-profiles as a basis for accurate universal cancer signaturesabstractBACKGROUND: Early screening for cancer is arguably one of the greatest public health advances over the last fifty years. However, many cancer screening tests are invasive (digital rectal exams), expensive (mammograms, imaging) or both (colonoscopies). This has spurred growing interest in developing genomic signatures that can be used for cancer diagnosis and prognosis. However, progress has been slowed by heterogeneity in cancer profiles and the lack of effective computational prediction tools for this type of data. RESULTS: We developed anti-profiles as a first step towards translating experimental findings suggesting that stochastic across-sample hyper-variability in the expression of specific genes is a stable and general property of cancer into predictive and diagnostic signatures. Using single-chip microarray normalization and quality assessment methods, we developed an anti-profile for colon cancer in tissue biopsy samples. To demonstrate the translational potential of our findings, we applied the signature developed in the tissue samples, without any further retraining or normalization, to screen patients for colon cancer based on genomic measurements from peripheral blood in an independent study (AUC of 0.89). This method achieved higher accuracy than the signature underlying commercially available peripheral blood screening tests for colon cancer (AUC of 0.81). We also confirmed the existence of hyper-variable genes across a range of cancer types and found that a significant proportion of tissue-specific genes are hyper-variable in cancer. Based on these observations, we developed a universal cancer anti-profile that accurately distinguishes cancer from normal regardless of tissue type (ten-fold cross-validation AUC > 0.92). CONCLUSIONS: We have introduced anti-profiles as a new approach for developing cancer genomic signatures that specifically takes advantage of gene expression heterogeneity. We have demonstrated that anti-profiles can be successfully applied to develop peripheral-blood based diagnostics for cancer and used anti-profiles to develop a highly accurate universal cancer signature. By using single-chip normalization and quality assessment methods, no further retraining of signatures developed by the anti-profile approach would be required before their application in clinical settings. Our results suggest that anti-profiles may be used to develop inexpensive and non-invasive universal cancer screening tests. Héctor Corrada Bravo, Vasyl Pihur, Matthew N. McCall, Rafael A. Irizarry, Jeffrey T. Leek |
BMC Bioinform. | 4 |
| 2012 | The partitioned LASSO-patternsearch algorithm with application to gene expression dataabstractBACKGROUND: In systems biology, the task of reverse engineering gene pathways from data has been limited not just by the curse of dimensionality (the interaction space is huge) but also by systematic error in the data. The gene expression barcode reduces spurious association driven by batch effects and probe effects. The binary nature of the resulting expression calls lends itself perfectly to modern regularization approaches that thrive in high-dimensional settings. RESULTS: The Partitioned LASSO-Patternsearch algorithm is proposed to identify patterns of multiple dichotomous risk factors for outcomes of interest in genomic studies. A partitioning scheme is used to identify promising patterns by solving many LASSO-Patternsearch subproblems in parallel. All variables that survive this stage proceed to an aggregation stage where the most significant patterns are identified by solving a reduced LASSO-Patternsearch problem in just these variables. This approach was applied to genetic data sets with expression levels dichotomized by gene expression bar code. Most of the genes and second-order interactions thus selected and are known to be related to the outcomes. CONCLUSIONS: We demonstrate with simulations and data analyses that the proposed method not only selects variables and patterns more accurately, but also provides smaller models with better prediction accuracy, in comparison to several alternative methodologies. Weiliang Shi, Grace Wahba, Rafael A. Irizarry, Héctor Corrada Bravo, Stephen J. Wright 0001 |
BMC Bioinform. | 3 |
| 2011 | Performance assessment of copy number microarray platforms using a spike-in experimentabstractMOTIVATION: Changes in the copy number of chromosomal DNA segments [copy number variants (CNVs)] have been implicated in human variation, heritable diseases and cancers. Microarray-based platforms are the current established technology of choice for studies reporting these discoveries and constitute the benchmark against which emergent sequence-based approaches will be evaluated. Research that depends on CNV analysis is rapidly increasing, and systematic platform assessments that distinguish strengths and weaknesses are needed to guide informed choice. RESULTS: We evaluated the sensitivity and specificity of six platforms, provided by four leading vendors, using a spike-in experiment. NimbleGen and Agilent platforms outperformed Illumina and Affymetrix in accuracy and precision of copy number dosage estimates. However, Illumina and Affymetrix algorithms that leverage single nucleotide polymorphism (SNP) information make up for this disadvantage and perform well at variant detection. Overall, the NimbleGen 2.1M platform outperformed others, but only with the use of an alternative data analysis pipeline to the one offered by the manufacturer. AVAILABILITY: The data is available from http://rafalab.jhsph.edu/cnvcomp/. CONTACT: [email protected]; [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Eitan Halper-Stromberg, Laurence Frelin, Ingo Ruczinski, Robert B. Scharpf, Chunfa Jie, Benilton S. Carvalho, Haiping Hao, Kurt N. Hetrick, Anne Jedlicka, Amanda Dziedzic, Kim Doheny, Alan F. Scott, Steve Baylin, Jonathan Pevsner, Forrest Spencer, Rafael A. Irizarry |
Bioinform. | 16 |
| 2011 | Thawing Frozen Robust Multi-array Analysis (fRMA)abstractBACKGROUND: A novel method of microarray preprocessing--Frozen Robust Multi-array Analysis (fRMA)--has recently been developed. This algorithm allows the user to preprocess arrays individually while retaining the advantages of multi-array preprocessing methods. The frozen parameter estimates required by this algorithm are generated using a large database of publicly available arrays. Curation of such a database and creation of the frozen parameter estimates is time-consuming; therefore, fRMA has only been implemented on the most widely used Affymetrix platforms. RESULTS: We present an R package, frmaTools, that allows the user to quickly create his or her own frozen parameter vectors. We describe how this package fits into a preprocessing workflow and explore the size of the training dataset needed to generate reliable frozen parameter estimates. This is followed by a discussion of specific situations in which one might wish to create one's own fRMA implementation. For a few specific scenarios, we demonstrate that fRMA performs well even when a large database of arrays in unavailable. CONCLUSIONS: By allowing the user to easily create his or her own fRMA implementation, the frmaTools package greatly increases the applicability of the fRMA algorithm. The frmaTools package is freely available as part of the Bioconductor project. Matthew N. McCall, Rafael A. Irizarry |
BMC Bioinform. | 2 |
| 2011 | Assessments of Affymetrix GeneChip Microarray Quality for Laboratories and Single SamplesabstractBACKGROUND: Microarray technology has become a widely used tool in the biological sciences. Over the past decade, the number of users has grown exponentially, and with the number of applications and secondary data analyses rapidly increasing, we expect this rate to continue. Various initiatives such as the External RNA Control Consortium (ERCC) and the MicroArray Quality Control (MAQC) project have explored ways to provide standards for the technology. For microarrays to become generally accepted as a reliable technology, statistical methods for assessing quality will be an indispensable component; however, there remains a lack of consensus in both defining and measuring microarray quality. RESULTS: We begin by providing a precise definition of microarray quality and reviewing existing Affymetrix GeneChip quality metrics in light of this definition. We show that the best-performing metrics require multiple arrays to be assessed simultaneously. While such multi-array quality metrics are adequate for bench science, as microarrays begin to be used in clinical settings, single-array quality metrics will be indispensable. To this end, we define a single-array version of one of the best multi-array quality metrics and show that this metric performs as well as the best multi-array metrics. We then use this new quality metric to assess the quality of microarry data available via the Gene Expression Omnibus (GEO) using more than 22,000 Affymetrix HGU133a and HGU133plus2 arrays from 809 studies. CONCLUSIONS: We find that approximately 10 percent of these publicly available arrays are of poor quality. Moreover, the quality of microarray measurements varies greatly from hybridization to hybridization, study to study, and lab to lab, with some experiments producing unusable data. Many of the concepts described here are applicable to other high-throughput technologies. Matthew N. McCall, Peter N. Murakami, Margus Lukk, Wolfgang Huber, Rafael A. Irizarry |
BMC Bioinform. | 5 |
| 2011 | Comparing genotyping algorithms for Illumina's Infinium whole-genome SNP BeadChipsabstractBACKGROUND: Illumina's Infinium SNP BeadChips are extensively used in both small and large-scale genetic studies. A fundamental step in any analysis is the processing of raw allele A and allele B intensities from each SNP into genotype calls (AA, AB, BB). Various algorithms which make use of different statistical models are available for this task. We compare four methods (GenCall, Illuminus, GenoSNP and CRLMM) on data where the true genotypes are known in advance and data from a recently published genome-wide association study. RESULTS: In general, differences in accuracy are relatively small between the methods evaluated, although CRLMM and GenoSNP were found to consistently outperform GenCall. The performance of Illuminus is heavily dependent on sample size, with lower no call rates and improved accuracy as the number of samples available increases. For X chromosome SNPs, methods with sex-dependent models (Illuminus, CRLMM) perform better than methods which ignore gender information (GenCall, GenoSNP). We observe that CRLMM and GenoSNP are more accurate at calling SNPs with low minor allele frequency than GenCall or Illuminus. The sample quality metrics from each of the four methods were found to have a high level of agreement at flagging samples with unusual signal characteristics. CONCLUSIONS: CRLMM, GenoSNP and GenCall can be applied with confidence in studies of any size, as their performance was shown to be invariant to the number of samples available. Illuminus on the other hand requires a larger number of samples to achieve comparable levels of accuracy and its use in smaller studies (50 or fewer individuals) is not recommended. Matthew E. Ritchie, Benilton S. Carvalho, Rafael A. Irizarry |
BMC Bioinform. | 4 |
| 2010 | A framework for oligonucleotide microarray preprocessingabstractMOTIVATION: The availability of flexible open source software for the analysis of gene expression raw level data has greatly facilitated the development of widely used preprocessing methods for these technologies. However, the expansion of microarray applications has exposed the limitation of existing tools. RESULTS: We developed the oligo package to provide a more general solution that supports a wide range of applications. The package is based on the BioConductor principles of transparency, reproducibility and efficiency of development. It extends the existing tools and leverages existing code for visualization, accessing data and widely used preprocessing routines. The oligo package implements a unified paradigm for preprocessing data and interfaces with other BioConductor tools for downstream analysis. Our infrastructure is general and can be used by other BioConductor packages. AVAILABILITY: The oligo package is freely available through BioConductor, http://www.bioconductor.org. Benilton S. Carvalho, Rafael A. Irizarry |
Bioinform. | 2 |
| 2010 | Quantifying uncertainty in genotype callsabstractMOTIVATION: Genome-wide association studies (GWAS) are used to discover genes underlying complex, heritable disorders for which less powerful study designs have failed in the past. The number of GWAS has skyrocketed recently with findings reported in top journals and the mainstream media. Microarrays are the genotype calling technology of choice in GWAS as they permit exploration of more than a million single nucleotide polymorphisms (SNPs) simultaneously. The starting point for the statistical analyses used by GWAS to determine association between loci and disease is making genotype calls (AA, AB or BB). However, the raw data, microarray probe intensities, are heavily processed before arriving at these calls. Various sophisticated statistical procedures have been proposed for transforming raw data into genotype calls. We find that variability in microarray output quality across different SNPs, different arrays and different sample batches have substantial influence on the accuracy of genotype calls made by existing algorithms. Failure to account for these sources of variability can adversely affect the quality of findings reported by the GWAS. RESULTS: We developed a method based on an enhanced version of the multi-level model used by CRLMM version 1. Two key differences are that we now account for variability across batches and improve the call-specific assessment of each call. The new model permits the development of quality metrics for SNPs, samples and batches of samples. Using three independent datasets, we demonstrate that the CRLMM version 2 outperforms CRLMM version 1 and the algorithm provided by Affymetrix, Birdseed. The main advantage of the new approach is that it enables the identification of low-quality SNPs, samples and batches. AVAILABILITY: Software implementing of the method described in this article is available as free and open source code in the crlmm R/BioConductor package. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Benilton S. Carvalho, Thomas A. Louis, Rafael A. Irizarry |
Bioinform. | 3 |
| 2009 | R/Bioconductor software for Illumina's Infinium whole-genome genotyping BeadChipsabstractUNLABELLED: Illumina produces a number of microarray-based technologies for human genotyping. An Infinium BeadChip is a two-color platform that types between 10(5) and 10(6) single nucleotide polymorphisms (SNPs) per sample. Despite being widely used, there is a shortage of open source software to process the raw intensities from this platform into genotype calls. To this end, we have developed the R/Bioconductor package crlmm for analyzing BeadChip data. After careful preprocessing, our software applies the CRLMM algorithm to produce genotype calls, confidence scores and other quality metrics at both the SNP and sample levels. We provide access to the raw summary-level intensity data, allowing users to develop their own methods for genotype calling or copy number analysis if they wish. AVAILABILITY AND IMPLEMENTATION: The crlmm Bioconductor package is available from http://www.bioconductor.org. Data packages and documentation are available from http://rafalab.jhsph.edu/software.html. Matthew E. Ritchie, Benilton S. Carvalho, Kurt N. Hetrick, Simon Tavaré, Rafael A. Irizarry |
Bioinform. | 5 |
| 2008 | Estimation and assessment of raw copy numbers at the single locus levelabstractMOTIVATION: Although copy-number aberrations are known to contribute to the diversity of the human DNA and cause various diseases, many aberrations and their phenotypes are still to be explored. The recent development of single-nucleotide polymorphism (SNP) arrays provides researchers with tools for calling genotypes and identifying chromosomal aberrations at an order-of-magnitude greater resolution than possible a few years ago. The fundamental problem in array-based copy-number (CN) analysis is to obtain CN estimates at a single-locus resolution with high accuracy and precision such that downstream segmentation methods are more likely to succeed. RESULTS: We propose a preprocessing method for estimating raw CNs from Affymetrix SNP arrays. Its core utilizes a multichip probe-level model analogous to that for high-density oligonucleotide expression arrays. We extend this model by adding an adjustment for sequence-specific allelic imbalances such as cross-hybridization between allele A and allele B probes. We focus on total CN estimates, which allows us to further constrain the probe-level model to increase the signal-to-noise ratio of CN estimates. Further improvement is obtained by controlling for PCR effects. Each part of the model is fitted robustly. The performance is assessed by quantifying how well raw CNs alone differentiate between one and two copies on Chromosome X (ChrX) at a single-locus resolution (27kb) up to a 200kb resolution. The evaluation is done with publicly available HapMap data. AVAILABILITY: The proposed method is available as part of an open-source R package named aroma.affymetrix. Because it is a bounded-memory algorithm, any number of arrays can be analyzed. Henrik Bengtsson, Rafael A. Irizarry, Benilton S. Carvalho, Terence P. Speed |
Bioinform. | 2 |
| 2007 | Estimating Genome-Wide Copy Number Using Allele Specific Mixture Models
Wenyi Wang 0001, Benilton S. Carvalho, Nathaniel D. Miller, Jonathan Pevsner, Aravinda Chakravarti, Rafael A. Irizarry |
RECOMB | 6 |
| 2006 | Comparison of Affymetrix GeneChip expression measuresabstractMOTIVATION: In the Affymetrix GeneChip system, preprocessing occurs before one obtains expression level measurements. Because the number of competing preprocessing methods was large and growing we developed a benchmark to help users identify the best method for their application. A webtool was made available for developers to benchmark their procedures. At the time of writing over 50 methods had been submitted. RESULTS: We benchmarked 31 probe set algorithms using a U95A dataset of spike in controls. Using this dataset, we found that background correction, one of the main steps in preprocessing, has the largest effect on performance. In particular, background correction appears to improve accuracy but, in general, worsen precision. The benchmark results put this balance in perspective. Furthermore, we have improved some of the original benchmark metrics to provide more detailed information regarding precision and accuracy. A handful of methods stand out as providing the best balance using spike-in data with the older U95A array, although different experiments on more current arrays may benchmark differently. AVAILABILITY: The affycomp package, now version 1.5.2, continues to be available as part of the Bioconductor project (http://www.bioconductor.org). The webtool continues to be available at http://affycomp.biostat.jhsph.edu CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rafael A. Irizarry, Zhijin Wu, Harris A. Jaffee |
Bioinform. | 1 |
| 2006 | High-resolution spatial normalization for microarrays containing embedded technical replicatesabstractMOTIVATION: Microarray data are susceptible to a wide-range of artifacts, many of which occur on physical scales comparable to the spatial dimensions of the array. These artifacts introduce biases that are spatially correlated. The ability of current methodologies to detect and correct such biases is limited. RESULTS: We introduce a new approach for analyzing spatial artifacts, termed 'conditional residual analysis for microarrays' (CRAM). CRAM requires a microarray design that contains technical replicates of representative features and a limited number of negative controls, but is free of the assumptions that constrain existing analytical procedures. The key idea is to extract residuals from sets of matched replicates to generate residual images. The residual images reveal spatial artifacts with single-feature resolution. Surprisingly, spatial artifacts were found to coexist independently as additive and multiplicative errors. Efficient procedures for bias estimation were devised to correct the spatial artifacts on both intensity scales. In a survey of 484 published single-channel datasets, variance fell 4- to 12-fold in 5% of the datasets after bias correction. Thus, inclusion of technical replicates in a microarray design affords benefits far beyond what one might expect with a conventional 'n = 5' averaging, and should be considered when designing any microarray for which randomization is feasible. AVAILABILITY: CRAM is implemented as version 2 of the hoptag software package for R, which is included in the Supplementary information. Daniel S. Yuan, Rafael A. Irizarry |
Bioinform. | 2 |
| 2006 | A summarization approach for Affymetrix GeneChip data using a reference training set from a large, biologically diverse databaseabstractBACKGROUND: Many of the most popular pre-processing methods for Affymetrix expression arrays, such as RMA, gcRMA, and PLIER, simultaneously analyze data across a set of predetermined arrays to improve precision of the final measures of expression. One problem associated with these algorithms is that expression measurements for a particular sample are highly dependent on the set of samples used for normalization and results obtained by normalization with a different set may not be comparable. A related problem is that an organization producing and/or storing large amounts of data in a sequential fashion will need to either re-run the pre-processing algorithm every time an array is added or store them in batches that are pre-processed together. Furthermore, pre-processing of large numbers of arrays requires loading all the feature-level data into memory which is a difficult task even with modern computers. We utilize a scheme that produces all the information necessary for pre-processing using a very large training set that can be used for summarization of samples outside of the training set. All subsequent pre-processing tasks can be done on an individual array basis. We demonstrate the utility of this approach by defining a new version of the Robust Multi-chip Averaging (RMA) algorithm which we refer to as refRMA. RESULTS: We assess performance based on multiple sets of samples processed over HG U133A Affymetrix GeneChip arrays. We show that the refRMA workflow, when used in conjunction with a large, biologically diverse training set, results in the same general characteristics as that of RMA in its classic form when comparing overall data structure, sample-to-sample correlation, and variation. Further, we demonstrate that the refRMA workflow and reference set can be robustly applied to naïve organ types and to benchmark data where its performance indicates respectable results. CONCLUSION: Our results indicate that a biologically diverse reference database can be used to train a model for estimating probe set intensities of exclusive test sets, while retaining the overall characteristics of the base algorithm. Although the results we present are specific for RMA, similar versions of other multi-array normalization and summarization schemes can be developed. Simon Katz, Rafael A. Irizarry, Mark Tripputi, Mark W. Porter |
BMC Bioinform. | 2 |
| 2006 | A simple spreadsheet-based, MIAME-supportive format for microarray data: MAGE-TABabstractBACKGROUND: Sharing of microarray data within the research community has been greatly facilitated by the development of the disclosure and communication standards MIAME and MAGE-ML by the MGED Society. However, the complexity of the MAGE-ML format has made its use impractical for laboratories lacking dedicated bioinformatics support. RESULTS: We propose a simple tab-delimited, spreadsheet-based format, MAGE-TAB, which will become a part of the MAGE microarray data standard and can be used for annotating and communicating microarray data in a MIAME compliant fashion. CONCLUSION: MAGE-TAB will enable laboratories without bioinformatics experience or support to manage, exchange and submit well-annotated microarray data in a standard format using a spreadsheet. The MAGE-TAB format is self-contained, and does not require an understanding of MAGE-ML or XML. Tim F. Rayner, Philippe Rocca-Serra, Paul T. Spellman, Helen C. Causton, Anna Farne, Ele Holloway, Rafael A. Irizarry, Junmin Liu, Donald Maier, Michael Miller 0001, Kjell Petersen, John Quackenbush, Gavin Sherlock, Christian J. Stoeckert Jr., Joseph White, Patricia L. Whetzel, Farrell Wymore, Helen E. Parkinson, Ugis Sarkans, Catherine A. Ball, Alvis Brazma |
BMC Bioinform. | 7 |
| 2004 | Stochastic models inspired by hybridization theory for short oligonucleotide arraysabstractHigh density oligonucleotide expression arrays are a widely used tool for the measurement of gene expression on a large scale. Affymetrix GeneChip arrays appear to dominate this market. These arrays use short oligonucleotides to probe for genes in an RNA sample. Due to optical noise, non-specific hybridization, probe-specific effects, and measurement error, ad-hoc measures of expression, that summarize probe intensities, can lead to imprecise and inaccurate results. Various researchers have demonstrated that expression measures based on simple statistical models can provide great improvements over the ad-hoc procedure offered by Affymetrix. Recently, physical models based on molecular hybridization theory, have been proposed as useful tools for prediction of, for example, non-specific hybridization. These physical models show great potential in terms of improving existing expression measures. In this paper we suggest that the system producing the measured intensities is too complex to be fully described with these relatively simple physical models and we propose empirically motivated stochastic models that compliment the above mentioned molecular hybridization theory to provide a comprehensive description of the data. We discuss how the proposed model can be used to obtain improved measures of expression useful for the data analysts. Zhijin Wu, Rafael A. Irizarry |
RECOMB | 2 |
| 2004 | A benchmark for Affymetrix GeneChip expression measuresabstractAbstract Motivation: The defining feature of oligonucleotide expression arrays is the use of several probes to assay each targeted transcript. This is a bonanza for the statistical geneticist, who can create probeset summaries with specific characteristics. There are now several methods available for summarizing probe level data from the popular Affymetrix GeneChips, but it is difficult to identify the best method for a given inquiry. Results: We have developed a graphical tool to evaluate summaries of Affymetrix probe level data. Plots and summary statistics offer a picture of how an expression measure performs in several important areas. This picture facilitates the comparison of competing expression measures and the selection of methods suitable for a specific investigation. The key is a benchmark data set consisting of a dilution study and a spike-in study. Because the truth is known for these data, we can identify statistical features of the data for which the expected outcome is known in advance. Those features highlighted in our suite of graphs are justified by questions of biological interest and motivated by the presence of appropriate data. Availability: In conjunction with the release of a graphics toolbox as part of the Bioconductor project (http://www.bioconductor.org), a webtool is available at http://affycomp.biostat.jhsph.edu. Supplemental material is available at http://www.biostat.jhsph.edu/~ririzarr/papers/suppaffycomp.pdf Leslie Cope, Rafael A. Irizarry, Harris A. Jaffee, Zhijin Wu, Terence P. Speed |
Bioinform. | 2 |
| 2004 | affy - analysis of Affymetrix GeneChip data at the probe levelabstractMOTIVATION: The processing of the Affymetrix GeneChip data has been a recent focus for data analysts. Alternatives to the original procedure have been proposed and some of these new methods are widely used. RESULTS: The affy package is an R package of functions and classes for the analysis of oligonucleotide arrays manufactured by Affymetrix. The package is currently in its second release, affy provides the user with extreme flexibility when carrying out an analysis and make it possible to access and manipulate probe intensity data. In this paper, we present the main classes and functions in the package and demonstrate how they can be used to process probe-level data. We also demonstrate the importance of probe-level analysis when using the Affymetrix GeneChip platform. Laurent Gautier, Leslie Cope, Benjamin M. Bolstad, Rafael A. Irizarry |
Bioinform. | 4 |
| 2003 | A comparison of normalization methods for high density oligonucleotide array data based on variance and biasabstractMOTIVATION: When running experiments that involve multiple high density oligonucleotide arrays, it is important to remove sources of variation between arrays of non-biological origin. Normalization is a process for reducing this variation. It is common to see non-linear relations between arrays and the standard normalization provided by Affymetrix does not perform well in these situations. RESULTS: We present three methods of performing normalization at the probe intensity level. These methods are called complete data methods because they make use of data from all arrays in an experiment to form the normalizing relation. These algorithms are compared to two methods that make use of a baseline array: a one number scaling based algorithm and a method that uses a non-linear normalizing relation by comparing the variability and bias of an expression measure. Two publicly available datasets are used to carry out the comparisons. The simplest and quickest complete data method is found to perform favorably. AVAILABILITY: Software implementing all three of the complete data normalization methods is available as part of the R package Affy, which is a part of the Bioconductor project http://www.bioconductor.org. SUPPLEMENTARY INFORMATION: Additional figures may be found at http://www.stat.berkeley.edu/~bolstad/normalize/index.html Benjamin M. Bolstad, Rafael A. Irizarry, Magnus Åstrand, Terence P. Speed |
Bioinform. | 2 |
| 2000 | Some wavelet-based analyses of Markov chain data
David R. Brillinger, Pedro Alberto Morettin, Rafael A. Irizarry, Chang Chiann |
Signal Process. | 3 |
| 1998 | An investigation of the second- and higher-order spectra of music
David R. Brillinger, Rafael A. Irizarry |
Signal Process. | 2 |