Matthew N. McCall

dblp:55/9438 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
5since 2021 · last 2024
0000-0002-2473-0943ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 7 first-author · 5 since 2021
YearPublicationVenuePosition
2024 A model-based hierarchical Bayesian approach to Sholl analysis
abstract
MOTIVATION: Due to the link between microglial morphology and function, morphological changes in microglia are frequently used to identify pathological immune responses in the central nervous system. In the absence of pathology, microglia are responsible for maintaining homeostasis, and their morphology can be indicative of how the healthy brain behaves in the presence of external stimuli and genetic differences. Despite recent interest in high throughput methods for morphological analysis, Sholl analysis is still widely used for quantifying microglia morphology via imaging data. Often, the raw data are naturally hierarchical, minimally including many cells per image and many images per animal. However, existing methods for performing downstream inference on Sholl data rely on truncating this hierarchy so rudimentary statistical testing procedures can be used. RESULTS: To fill this longstanding gap, we introduce a parametric hierarchical Bayesian model-based approach for analyzing Sholl data, so that inference can be performed without aggressive reduction of otherwise very rich data. We apply our model to real data and perform simulation studies comparing the proposed method with a popular alternative. AVAILABILITY AND IMPLEMENTATION: Software to reproduce the results presented in this article is available at: https://github.com/vonkaenelerik/hierarchical_sholl. An R package implementing the proposed models is available at: https://github.com/vonkaenelerik/ShollBayes.
Erik Vonkaenel, Alexis Feidler, Rebecca Lowery, Katherine Andersh, Tanzy Love, Ania Majewska, Matthew N. McCall
Bioinform.7
2024 Correction: Multiple imputation and direct estimation for qPCR data with non-detects
Valeriia Sherina, Helene R. McMurray, Winslow Powers, Hartmut Land, Tanzy M. T. Love, Matthew N. McCall
BMC Bioinform.6
2021 The effect of tissue composition on gene co-expression
abstract
Variable cellular composition of tissue samples represents a significant challenge for the interpretation of genomic profiling studies. Substantial effort has been devoted to modeling and adjusting for compositional differences when estimating differential expression between sample types. However, relatively little attention has been given to the effect of tissue composition on co-expression estimates. In this study, we illustrate the effect of variable cell-type composition on correlation-based network estimation and provide a mathematical decomposition of the tissue-level correlation. We show that a class of deconvolution methods developed to separate tumor and stromal signatures can be applied to two component cell-type mixtures. In simulated and real data, we identify conditions in which a deconvolution approach would be beneficial. Our results suggest that uncorrelated cell-type-specific markers are ideally suited to deconvolute both the expression and co-expression patterns of an individual cell type. We provide a Shiny application for users to interactively explore the effect of cell-type composition on correlation-based co-expression estimation for any cell types of interest.
Yun Zhang 0021, Jonavelle Cuerdo, Marc K. Halushka, Matthew N. McCall
Briefings Bioinform.4
2021 Autoregressive modeling and diagnostics for qPCR amplification
abstract
MOTIVATION: Current methods used to analyze real-time quantitative polymerase chain reaction (qPCR) data exhibit systematic deviations from the assumed model over the progression of the reaction. Slight variations in the amount of the initial target molecule or in early amplifications are likely responsible for these deviations. Commonly used 4- and 5-parameter sigmoidal models appear to be particularly susceptible to this issue, often displaying patterns of autocorrelation in the residuals. The presence of this phenomenon, even for technical replicates, suggests that these parametric models may be misspecified. Specifically, they do not account for the sequential dependent nature of the amplification process that underlies qPCR fluorescence measurements. RESULTS: We demonstrate that a Smooth Transition Autoregressive (STAR) model addresses this limitation by explicitly modeling the dependence between cycles and the gradual transition between amplification regimes. In summary, application of a STAR model to qPCR amplification data improves model fit and reduces autocorrelation in the residuals. AVAILABILITY AND IMPLEMENTATION: R scripts to reproduce all the analyses and results described in this manuscript can be found at: https://github.com/bhsu4/GAPDH.SO. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Benjamin Hsu, Valeriia Sherina, Matthew N. McCall
Bioinform.3
2021 A systems genomics approach uncovers molecular associates of RSV severity
abstract
Respiratory syncytial virus (RSV) infection results in millions of hospitalizations and thousands of deaths each year. Variations in the adaptive and innate immune response appear to be associated with RSV severity. To investigate the host response to RSV infection in infants, we performed a systems-level study of RSV pathophysiology, incorporating high-throughput measurements of the peripheral innate and adaptive immune systems and the airway epithelium and microbiota. We implemented a novel multi-omic data integration method based on multilayered principal component analysis, penalized regression, and feature weight back-propagation, which enabled us to identify cellular pathways associated with RSV severity. In both airway and immune cells, we found an association between RSV severity and activation of pathways controlling Th17 and acute phase response signaling, as well as inhibition of B cell receptor signaling. Dysregulation of both the humoral and mucosal response to RSV may play a critical role in determining illness severity.
Matthew N. McCall, Chin-Yi Chu, Lauren Benoodt, Juilee Thakar, Anthony Corbett, Jeanne Holden-Wiltse, Christopher Slaunwhite, Alex Grier, Steven R. Gill, Ann R. Falsey, David J. Topham, Mary T. Caserta, Edward E. Walsh, Xing Qiu, Thomas J. Mariani
PLoS Comput. Biol.1
2020 Multiple imputation and direct estimation for qPCR data with non-detects
abstract
BACKGROUND: Quantitative real-time PCR (qPCR) is one of the most widely used methods to measure gene expression. An important aspect of qPCR data that has been largely ignored is the presence of non-detects: reactions failing to exceed the quantification threshold and therefore lacking a measurement of expression. While most current software replaces these non-detects with a value representing the limit of detection, this introduces substantial bias in the estimation of both absolute and differential expression. Single imputation procedures, while an improvement on previously used methods, underestimate residual variance, which can lead to anti-conservative inference. RESULTS: We propose to treat non-detects as non-random missing data, model the missing data mechanism, and use this model to impute missing values or obtain direct estimates of model parameters. To account for the uncertainty inherent in the imputation, we propose a multiple imputation procedure, which provides a set of plausible values for each non-detect. We assess the proposed methods via simulation studies and demonstrate the applicability of these methods to three experimental data sets. We compare our methods to mean imputation, single imputation, and a penalized EM algorithm incorporating non-random missingness (PEMM). The developed methods are implemented in the R/Bioconductor package nondetects. CONCLUSIONS: The statistical methods introduced here reduce discrepancies in gene expression values derived from qPCR experiments in the presence of non-detects, providing increased confidence in downstream analyses.
Valeriia Sherina, Helene R. McMurray, Winslow Powers, Hartmut Land, Tanzy M. T. Love, Matthew N. McCall
BMC Bioinform.6
2016 A benchmark for microRNA quantification algorithms using the OpenArray platform
abstract
BACKGROUND: Several techniques have been tailored to the quantification of microRNA expression, including hybridization arrays, quantitative PCR (qPCR), and high-throughput sequencing. Each of these has certain strengths and limitations depending both on the technology itself and the algorithm used to convert raw data into expression estimates. Reliable quantification of microRNA expression is challenging in part due to the relatively low abundance and short length of the miRNAs. While substantial research has been devoted to the development of methods to quantify mRNA expression, relatively little effort has been spent on microRNA expression. RESULTS: In this work, we focus on the Life Technologies TaqMan OpenArray(Ⓡ) system, a qPCR-based platform to measure microRNA expression. Several algorithms currently exist to estimate expression from the raw amplification data produced by qPCR-based technologies. To assess and compare the performance of these methods, we performed a set of dilution/mixture experiments to create a benchmark data set. We also developed a suite of statistical assessments that evaluate many different aspects of performance: accuracy, precision, titration response, number of complete features, limit of detection, and data quality. The benchmark data and software are freely available via two R/Bioconductor packages, miRcomp and miRcompData. Finally, we demonstrate use of our software by comparing two widely used algorithms and providing assessments for four other algorithms. CONCLUSIONS: Benchmark data sets and software are crucial tools for the assessment and comparison of competing algorithms. We believe that the miRcomp and miRcompData packages will facilitate the development of new methodology for microRNA expression estimation.
Matthew N. McCall, Alexander S. Baras, Alexander Crits-Christoph, Roxann Ingersoll, Melissa A. McAlexander, Kenneth W. Witwer, Marc K. Halushka
BMC Bioinform.1
2014 On non-detects in qPCR data
abstract
MOTIVATION: Quantitative real-time PCR (qPCR) is one of the most widely used methods to measure gene expression. Despite extensive research in qPCR laboratory protocols, normalization and statistical analysis, little attention has been given to qPCR non-detects-those reactions failing to produce a minimum amount of signal. RESULTS: We show that the common methods of handling qPCR non-detects lead to biased inference. Furthermore, we show that non-detects do not represent data missing completely at random and likely represent missing data occurring not at random. We propose a model of the missing data mechanism and develop a method to directly model non-detects as missing data. Finally, we show that our approach results in a sizeable reduction in bias when estimating both absolute and differential gene expression. AVAILABILITY AND IMPLEMENTATION: The proposed algorithm is implemented in the R package, nondetects. This package also contains the raw data for the three example datasets used in this manuscript. The package is freely available at http://mnmccall.com/software and as part of the Bioconductor project.
Matthew N. McCall, Helene R. McMurray, Hartmut Land, Anthony Almudevar
Bioinform.1
2013 ChIP-PED enhances the analysis of ChIP-seq and ChIP-chip data
abstract
MOTIVATION: Although chromatin immunoprecipitation coupled with high-throughput sequencing (ChIP-seq) or tiling array hybridization (ChIP-chip) is increasingly used to map genome-wide-binding sites of transcription factors (TFs), it still remains difficult to generate a quality ChIPx (i.e. ChIP-seq or ChIP-chip) dataset because of the tremendous amount of effort required to develop effective antibodies and efficient protocols. Moreover, most laboratories are unable to easily obtain ChIPx data for one or more TF(s) in more than a handful of biological contexts. Thus, standard ChIPx analyses primarily focus on analyzing data from one experiment, and the discoveries are restricted to a specific biological context. RESULTS: We propose to enrich this existing data analysis paradigm by developing a novel approach, ChIP-PED, which superimposes ChIPx data on large amounts of publicly available human and mouse gene expression data containing a diverse collection of cell types, tissues and disease conditions to discover new biological contexts with potential TF regulatory activities. We demonstrate ChIP-PED using a number of examples, including a novel discovery that MYC, a human TF, plays an important functional role in pediatric Ewing sarcoma cell lines. These examples show that ChIP-PED increases the value of ChIPx data by allowing one to expand the scope of possible discoveries made from a ChIPx experiment. AVAILABILITY: http://www.biostat.jhsph.edu/~gewu/ChIPPED/
George Wu, Jason T. Yustein, Matthew N. McCall, Michael J. Zilliox, Rafael A. Irizarry, Karen Zeller, Chi V. Dang, Hongkai Ji
Bioinform.3
2012 Affymetrix GeneChip microarray preprocessing for multivariate analyses
abstract
Affymetrix GeneChip microarrays are the most widely used high-throughput technology to measure gene expression, and a wide variety of preprocessing methods have been developed to transform probe intensities reported by a microarray scanner into gene expression estimates. There have been numerous comparisons of these preprocessing methods, focusing on the most common analyses-detection of differential expression and gene or sample clustering. Recently, more complex multivariate analyses, such as gene co-expression, differential co-expression, gene set analysis and network modeling, are becoming more common; however, the same preprocessing methods are typically applied. In this article, we examine the effect of preprocessing methods on some of these multivariate analyses and provide guidance to the user as to which methods are most appropriate.
Matthew N. McCall, Anthony Almudevar
Briefings Bioinform.1
2012 fRMA ST: frozen robust multiarray analysis for Affymetrix Exon and Gene ST arrays
abstract
SUMMARY: Frozen robust multiarray analysis (fRMA) is a single-array preprocessing algorithm that retains the advantages of multiarray algorithms and removes certain batch effects by downweighting probes that have high between-batch residual variance. Here, we extend the fRMA algorithm to two new microarray platforms--Affymetrix Human Exon and Gene 1.0 ST--by modifying the fRMA probe-level model and extending the frma package to work with oligo ExonFeatureSet and GeneFeatureSet objects. AVAILABILITY AND IMPLEMENTATION: All packages are implemented in R. Source code and binaries are freely available through the Bioconductor project. Convenient links to all software and data packages can be found at http://mnmccall.com/software CONTACT: [email protected].
Matthew N. McCall, Harris A. Jaffee, Rafael A. Irizarry
Bioinform.1
2012 Gene expression anti-profiles as a basis for accurate universal cancer signatures
abstract
BACKGROUND: Early screening for cancer is arguably one of the greatest public health advances over the last fifty years. However, many cancer screening tests are invasive (digital rectal exams), expensive (mammograms, imaging) or both (colonoscopies). This has spurred growing interest in developing genomic signatures that can be used for cancer diagnosis and prognosis. However, progress has been slowed by heterogeneity in cancer profiles and the lack of effective computational prediction tools for this type of data. RESULTS: We developed anti-profiles as a first step towards translating experimental findings suggesting that stochastic across-sample hyper-variability in the expression of specific genes is a stable and general property of cancer into predictive and diagnostic signatures. Using single-chip microarray normalization and quality assessment methods, we developed an anti-profile for colon cancer in tissue biopsy samples. To demonstrate the translational potential of our findings, we applied the signature developed in the tissue samples, without any further retraining or normalization, to screen patients for colon cancer based on genomic measurements from peripheral blood in an independent study (AUC of 0.89). This method achieved higher accuracy than the signature underlying commercially available peripheral blood screening tests for colon cancer (AUC of 0.81). We also confirmed the existence of hyper-variable genes across a range of cancer types and found that a significant proportion of tissue-specific genes are hyper-variable in cancer. Based on these observations, we developed a universal cancer anti-profile that accurately distinguishes cancer from normal regardless of tissue type (ten-fold cross-validation AUC > 0.92). CONCLUSIONS: We have introduced anti-profiles as a new approach for developing cancer genomic signatures that specifically takes advantage of gene expression heterogeneity. We have demonstrated that anti-profiles can be successfully applied to develop peripheral-blood based diagnostics for cancer and used anti-profiles to develop a highly accurate universal cancer signature. By using single-chip normalization and quality assessment methods, no further retraining of signatures developed by the anti-profile approach would be required before their application in clinical settings. Our results suggest that anti-profiles may be used to develop inexpensive and non-invasive universal cancer screening tests.
Héctor Corrada Bravo, Vasyl Pihur, Matthew N. McCall, Rafael A. Irizarry, Jeffrey T. Leek
BMC Bioinform.3
2011 Thawing Frozen Robust Multi-array Analysis (fRMA)
abstract
BACKGROUND: A novel method of microarray preprocessing--Frozen Robust Multi-array Analysis (fRMA)--has recently been developed. This algorithm allows the user to preprocess arrays individually while retaining the advantages of multi-array preprocessing methods. The frozen parameter estimates required by this algorithm are generated using a large database of publicly available arrays. Curation of such a database and creation of the frozen parameter estimates is time-consuming; therefore, fRMA has only been implemented on the most widely used Affymetrix platforms. RESULTS: We present an R package, frmaTools, that allows the user to quickly create his or her own frozen parameter vectors. We describe how this package fits into a preprocessing workflow and explore the size of the training dataset needed to generate reliable frozen parameter estimates. This is followed by a discussion of specific situations in which one might wish to create one's own fRMA implementation. For a few specific scenarios, we demonstrate that fRMA performs well even when a large database of arrays in unavailable. CONCLUSIONS: By allowing the user to easily create his or her own fRMA implementation, the frmaTools package greatly increases the applicability of the fRMA algorithm. The frmaTools package is freely available as part of the Bioconductor project.
Matthew N. McCall, Rafael A. Irizarry
BMC Bioinform.1
2011 Assessments of Affymetrix GeneChip Microarray Quality for Laboratories and Single Samples
abstract
BACKGROUND: Microarray technology has become a widely used tool in the biological sciences. Over the past decade, the number of users has grown exponentially, and with the number of applications and secondary data analyses rapidly increasing, we expect this rate to continue. Various initiatives such as the External RNA Control Consortium (ERCC) and the MicroArray Quality Control (MAQC) project have explored ways to provide standards for the technology. For microarrays to become generally accepted as a reliable technology, statistical methods for assessing quality will be an indispensable component; however, there remains a lack of consensus in both defining and measuring microarray quality. RESULTS: We begin by providing a precise definition of microarray quality and reviewing existing Affymetrix GeneChip quality metrics in light of this definition. We show that the best-performing metrics require multiple arrays to be assessed simultaneously. While such multi-array quality metrics are adequate for bench science, as microarrays begin to be used in clinical settings, single-array quality metrics will be indispensable. To this end, we define a single-array version of one of the best multi-array quality metrics and show that this metric performs as well as the best multi-array metrics. We then use this new quality metric to assess the quality of microarry data available via the Gene Expression Omnibus (GEO) using more than 22,000 Affymetrix HGU133a and HGU133plus2 arrays from 809 studies. CONCLUSIONS: We find that approximately 10 percent of these publicly available arrays are of poor quality. Moreover, the quality of microarray measurements varies greatly from hybridization to hybridization, study to study, and lab to lab, with some experiments producing unusable data. Many of the concepts described here are applicable to other high-throughput technologies.
Matthew N. McCall, Peter N. Murakami, Margus Lukk, Wolfgang Huber, Rafael A. Irizarry
BMC Bioinform.1