Vincent Carey

dblp:98/246 · also Vincent J. Carey · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0003-4046-0063ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 20 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Eleven quick tips for writing a Bioconductor package
abstract
international
Charlotte Soneson, Lori A. Shepherd, Marcel Ramos, Kevin C. Rue-Albrecht, Johannes Rainer, Hervé Pagès, Vincent Carey
PLoS Comput. Biol.7
2024 Partial correlation network analysis identifies coordinated gene expression within a regional cluster of COPD genome-wide association signals
abstract
Chronic obstructive pulmonary disease (COPD) is a complex disease influenced by well-established environmental exposures (most notably, cigarette smoking) and incompletely defined genetic factors. The chromosome 4q region harbors multiple genetic risk loci for COPD, including signals near HHIP, FAM13A, GSTCD, TET2, and BTC. Leveraging RNA-Seq data from lung tissue in COPD cases and controls, we estimated the co-expression network for genes in the 4q region bounded by HHIP and BTC (~70MB), through partial correlations informed by protein-protein interactions. We identified several co-expressed gene pairs based on partial correlations, including NPNT-HHIP, BTC-NPNT and FAM13A-TET2, which were replicated in independent lung tissue cohorts. Upon clustering the co-expression network, we observed that four genes previously associated to COPD: BTC, HHIP, NPNT and PPM1K appeared in the same network community. Finally, we discovered a sub-network of genes differentially co-expressed between COPD vs controls (including FAM13A, PPA2, PPM1K and TET2). Many of these genes were previously implicated in cell-based knock-out experiments, including the knocking out of SPP1 which belongs to the same genomic region and could be a potential local key regulatory gene. These analyses identify chromosome 4q as a region enriched for COPD genetic susceptibility and differential co-expression.
Michele Gentili, Kimberly Glass, Enrico Maiorino, Brian D. Hobbs, Zhonghui Xu, Peter J. Castaldi, Michael H. Cho, Craig P. Hersh, Dandi Qiao, Jarrett D. Morrow, Vincent Carey, John Platig, Edwin K. Silverman
PLoS Comput. Biol.11
2023 RaggedExperiment: the missing link between genomic ranges and matrices in Bioconductor
abstract
SUMMARY: The RaggedExperiment R / Bioconductor package provides lossless representation of disparate genomic ranges across multiple specimens or cells, in conjunction with efficient and flexible calculations of rectangular-shaped summaries for downstream analysis. Applications include statistical analysis of somatic mutations, copy number, methylation, and open chromatin data. RaggedExperiment is compatible with multimodal data analysis as a component of MultiAssayExperiment data objects, and simplifies data representation and transformation for software developers and analysts. MOTIVATION AND RESULTS: Measurement of copy number, mutation, single nucleotide polymorphism, and other genomic attributes that may be stored as VCF files produce "ragged" genomic ranges data: i.e. across different genomic coordinates in each sample. Ragged data are not rectangular or matrix-like, presenting informatics challenges for downstream statistical analyses. We present the RaggedExperiment R/Bioconductor data structure for lossless representation of ragged genomic data, with associated reshaping tools for flexible and efficient calculation of tabular representations to support a wide range of downstream statistical analyses. We demonstrate its applicability to copy number and somatic mutation data across 33 TCGA cancer datasets.
Marcel Ramos, Martin Morgan, Ludwig Geistlinger, Vincent Carey, Levi Waldron
Bioinform.4
2023 Curated single cell multimodal landmark datasets for R/Bioconductor
abstract
BACKGROUND: The majority of high-throughput single-cell molecular profiling methods quantify RNA expression; however, recent multimodal profiling methods add simultaneous measurement of genomic, proteomic, epigenetic, and/or spatial information on the same cells. The development of new statistical and computational methods in Bioconductor for such data will be facilitated by easy availability of landmark datasets using standard data classes. RESULTS: We collected, processed, and packaged publicly available landmark datasets from important single-cell multimodal protocols, including CITE-Seq, ECCITE-Seq, SCoPE2, scNMT, 10X Multiome, seqFISH, and G&T. We integrate data modalities via the MultiAssayExperiment Bioconductor class, document and re-distribute datasets as the SingleCellMultiModal package in Bioconductor's Cloud-based ExperimentHub. The result is single-command actualization of landmark datasets from seven single-cell multimodal data generation technologies, without need for further data processing or wrangling in order to analyze and develop methods within Bioconductor's ecosystem of hundreds of packages for single-cell and multimodal data. CONCLUSIONS: We provide two examples of integrative analyses that are greatly simplified by SingleCellMultiModal. The package will facilitate development of bioinformatic and statistical methods in Bioconductor to meet the challenges of integrating molecular layers and analyzing phenotypic outputs including cell differentiation, activity, and disease.
Kelly B. Eckenrode, Dario Righelli, Marcel Ramos, Ricard Argelaguet, Christophe Vanderaa, Ludwig Geistlinger, Aedín C. Culhane, Laurent Gatto, Vincent Carey, Martin Morgan, Davide Risso, Levi Waldron
PLoS Comput. Biol.9
2021 Toward a gold standard for benchmarking gene set enrichment analysis
abstract
MOTIVATION: Although gene set enrichment analysis has become an integral part of high-throughput gene expression data analysis, the assessment of enrichment methods remains rudimentary and ad hoc. In the absence of suitable gold standards, evaluations are commonly restricted to selected datasets and biological reasoning on the relevance of resulting enriched gene sets. RESULTS: We develop an extensible framework for reproducible benchmarking of enrichment methods based on defined criteria for applicability, gene set prioritization and detection of relevant processes. This framework incorporates a curated compendium of 75 expression datasets investigating 42 human diseases. The compendium features microarray and RNA-seq measurements, and each dataset is associated with a precompiled GO/KEGG relevance ranking for the corresponding disease under investigation. We perform a comprehensive assessment of 10 major enrichment methods, identifying significant differences in runtime and applicability to RNA-seq data, fraction of enriched gene sets depending on the null hypothesis tested and recovery of the predefined relevance rankings. We make practical recommendations on how methods originally developed for microarray data can efficiently be applied to RNA-seq data, how to interpret results depending on the type of gene set test conducted and which methods are best suited to effectively prioritize gene sets with high phenotype relevance. AVAILABILITY: http://bioconductor.org/packages/GSEABenchmarkeR. CONTACT: [email protected].
Ludwig Geistlinger, Gergely Csaba, Mara Santarelli, Marcel Ramos, Lucas Schiffer, Nitesh Turaga, Charity W. Law, Sean R. Davis, Vincent Carey, Martin Morgan, Ralf Zimmer, Levi Waldron
Briefings Bioinform.9
2019 deconvSeq: deconvolution of cell mixture distribution in sequencing data
abstract
MOTIVATION: Although single-cell sequencing is becoming more widely available, many tissue samples such as intracranial aneurysms are both fibrous and minute, and therefore not easily dissociated into single cells. To account for the cell type heterogeneity in such tissues therefore requires a computational method. We present a computational deconvolution method, deconvSeq, for sequencing data (RNA and bisulfite) obtained from bulk tissue. This method can also be applied to single-cell RNA sequencing data. RESULTS: DeconvSeq utilizes a generalized linear model to model effects of tissue type on feature quantification, which is specific to the data structure of the sequencing type used. Estimated model coefficients can then be used to predict the cell type mixture within a tissue. Predicted cell type mixtures were validated against actual cell counts in whole blood samples. Using this method, we obtained a mean correlation of 0.998 (95% CI 0.995-0.999) from the RNA sequencing data of 35 whole blood samples and 0.95 (95% CI 0.91-0.98) from the reduced representation bisulfite sequencing data from 35 whole blood samples. Using symmetric balances to obtain the correlation between compositional parts, we found that the lowest correlation occurred for monocytes for both RNA and bisulfite sequencing. Comparison with other methods of decomposition such as deconRNAseq, CIBERSORT, MuSiC and EpiDISH showed that deconvSeq is able to achieve good prediction using mean correlation with far fewer genes or CpG sites in the signature set. AVAILABILITY AND IMPLEMENTATION: Software implementing deconvSeq is available at https://github.com/rosedu1/deconvSeq. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Rose Du, Vincent Carey, Scott T. Weiss
Bioinform.2
2016 Public data and open source tools for multi-assay genomic investigation of disease
abstract
Molecular interrogation of a biological sample through DNA sequencing, RNA and microRNA profiling, proteomics and other assays, has the potential to provide a systems level approach to predicting treatment response and disease progression, and to developing precision therapies. Large publicly funded projects have generated extensive and freely available multi-assay data resources; however, bioinformatic and statistical methods for the analysis of such experiments are still nascent. We review multi-assay genomic data resources in the areas of clinical oncology, pharmacogenomics and other perturbation experiments, population genomics and regulatory genomics and other areas, and tools for data acquisition. Finally, we review bioinformatic tools that are explicitly geared toward integrative genomic data visualization and analysis. This review provides starting points for accessing publicly available data and tools to support development of needed integrative methods.
Lavanya Kannan, Marcel Ramos, Angela Re, Nehme Hachem, Zhaleh Safikhani, Deena M. A. Gendoo, Sean R. Davis, David Gomez-Cabrero, Robert Castelo, Kasper D. Hansen, Vincent Carey, Martin Morgan, Aedín C. Culhane, Benjamin Haibe-Kains, Levi Waldron
Briefings Bioinform.11
2014 VariantAnnotation: a Bioconductor package for exploration and annotation of genetic variants
abstract
UNLABELLED: VariantAnnotation is an R / Bioconductor package for the exploration and annotation of genetic variants. Capabilities exist for reading, writing and filtering variant call format (VCF) files. VariantAnnotation allows ready access to additional R / Bioconductor facilities for advanced statistical analysis, data transformation, visualization and integration with diverse genomic resources. AVAILABILITY AND IMPLEMENTATION: This package is implemented in R and available for download at the Bioconductor Web site (http://bioconductor.org/packages/2.13/bioc/html/VariantAnnotation.html). The package contains extensive help pages for individual functions and a 'vignette' outlining typical work flows; it is made available under the open source 'Artistic-2.0' license. Version 1.9.38 was used in this article.
Valerie Obenchain, Michael F. Lawrence, Vincent Carey, Stephanie M. Gogarten, Paul Shannon, Martin Morgan
Bioinform.3
2013 Software for Computing and Annotating Genomic Ranges
abstract
We describe Bioconductor infrastructure for representing and computing on annotated genomic ranges and integrating genomic data with the statistical computing features of R and its extensions. At the core of the infrastructure are three packages: IRanges, GenomicRanges, and GenomicFeatures. These packages provide scalable data structures for representing annotated ranges on the genome, with special support for transcript structures, read alignments and coverage vectors. Computational facilities include efficient algorithms for overlap and nearest neighbor detection, coverage calculation and other range operations. This infrastructure directly supports more than 80 other Bioconductor packages, including those for sequence analysis, differential expression analysis and visualization.
Michael F. Lawrence, Wolfgang Huber, Hervé Pagès, Patrick Aboyoun, Marc Carlson, Robert Gentleman, Martin Morgan, Vincent Carey
PLoS Comput. Biol.8
2009 Power enhancement via multivariate outlier testing with gene expression arrays
abstract
MOTIVATION: As the use of microarrays in human studies continues to increase, stringent quality assurance is necessary to ensure accurate experimental interpretation. We present a formal approach for microarray quality assessment that is based on dimension reduction of established measures of signal and noise components of expression followed by parametric multivariate outlier testing. RESULTS: We applied our approach to several data resources. First, as a negative control, we found that the Affymetrix and Illumina contributions to MAQC data were free from outliers at a nominal outlier flagging rate of alpha=0.01. Second, we created a tunable framework for artificially corrupting intensity data from the Affymetrix Latin Square spike-in experiment to allow investigation of sensitivity and specificity of quality assurance (QA) criteria. Third, we applied the procedure to 507 Affymetrix microarray GeneChips processed with RNA from human peripheral blood samples. We show that exclusion of arrays by this approach substantially increases inferential power, or the ability to detect differential expression, in large clinical studies. AVAILABILITY: http://bioconductor.org/packages/2.3/bioc/html/arrayMvout.html and http://bioconductor.org/packages/2.3/bioc/html/affyContam.html affyContam (credentials: readonly/readonly)
Adam L. Asare, Zhong Gao, Vincent Carey, Vicki Seyfert-Margolis
Bioinform.3
2009 Data structures and algorithms for analysis of genetics of gene expression with Bioconductor: GGtools 3.x
abstract
SUMMARY: Associations between DNA polymorphisms and mRNA abundance are a natural target of genetic investigations, and microarrays facilitate genome-wide and transcriptome-wide surveys of these associations. This work is motivated by emerging requirements for data architectures and algorithm interfaces to allow flexible exploration of public and private archives of genotyping and expression arrays. Using R/Bioconductor facilities, Phase II HapMap genotypes and Illumina 47K expression assay results archived on multiple populations may be interactively explored and analyzed using commodity hardware. AVAILABILITY AND IMPLEMENTATION: Open Source. Bioconductor 2.3 packages GGtools, GGBase, GGdata, hmyriB36. Freely available on the web at (http://www.bioconductor.org).
Vincent Carey, Adam R. Davis, Michael F. Lawrence, Robert Gentleman, Benjamin A. Raby
Bioinform.1
2009 rtracklayer: an R package for interfacing with genome browsers
abstract
SUMMARY: The rtracklayer package supports the integration of existing genome browsers with experimental data analyses performed in R. The user may (i) transfer annotation tracks to and from a genome browser and (ii) create and manipulate browser views to focus on a particular set of annotations in a specific genomic region. Currently, the UCSC genome browser is supported. AVAILABILITY: The package is freely available from http://www.bioconductor.org/. A quick-start vignette is included with the package.
Michael F. Lawrence, Robert Gentleman, Vincent Carey
Bioinform.3
2009 Platform dependence of inference on gene-wise and gene-set involvement in human lung development
abstract
BACKGROUND: With the recent development of microarray technologies, the comparability of gene expression data obtained from different platforms poses an important problem. We evaluated two widely used platforms, Affymetrix U133 Plus 2.0 and the Illumina HumanRef-8 v2 Expression Bead Chips, for comparability in a biological system in which changes may be subtle, namely fetal lung tissue as a function of gestational age. RESULTS: We performed the comparison via sequence-based probe matching between the two platforms. "Significance grouping" was defined as a measure of comparability. Using both expression correlation and significance grouping as measures of comparability, we demonstrated that despite overall cross-platform differences at the single gene level, increased correlation between the two platforms was found in genes with higher expression level, higher probe overlap, and lower p-value. We also demonstrated that biological function as determined via KEGG pathways or GO categories is more consistent across platforms than single gene analysis. CONCLUSION: We conclude that while the comparability of the platforms at the single gene level may be increased by increasing sample size, they are highly comparable ontologically even for subtle differences in a relatively small sample size. Biologically relevant inference should therefore be reproducible across laboratories using different platforms.
Rose Du, Kelan Tantisira, Vincent Carey, Soumyaroop Bhattacharya, Stephanie Metje, Alvin T. Kho, Barbara J. Klanderman, Roger Gaedigk, Ross Lazarus, Thomas J. Mariani, J. Steven Leeder, Scott T. Weiss
BMC Bioinform.3
2007 GGtools: analysis of genetics of gene expression in bioconductor
abstract
UNLABELLED: This paper reviews the central concepts and implementation of data structures and methods for studying genetics of gene expression with the GGtools package of Bioconductor. Illustration with a HapMap+expression dataset is provided. AVAILABILITY: Package GGtools is part of Bioconductor 1.9 (http://bioconductor.org). Open source with Artistic License.
Vincent Carey, Martin Morgan, Seth Falcon, Ross Lazarus, Robert Gentleman
Bioinform.1
2007 Graphs in molecular biology
abstract
Graph theoretical concepts are useful for the description and analysis of interactions and relationships in biological systems. We give a brief introduction into some of the concepts and their areas of application in molecular biology. We discuss software that is available through the Bioconductor project and present a simple example application to the integration of a protein-protein interaction and a co-expression network.
Wolfgang Huber, Vincent Carey, Li Long, Seth Falcon, Robert Gentleman
BMC Bioinform.2
2007 Machine Learning and Its Applications to Biology
abstract
The term machine learning refers to a set of topics dealing with the creation and evaluation of algorithms that facilitate pattern recognition, classification, and prediction, based on models derived from existing data. Two facets of mechanization should be acknowledged when considering machine learning in broad terms. Firstly, it is intended that the classification and prediction tasks can be accomplished by a suitably programmed computing machine. That is, the product of machine learning is a classifier that can be feasibly used on available hardware. Secondly, it is intended that the creation of the classifier should itself be highly mechanized, and should not involve too much human input. This second facet is inevitably vague, but the basic objective is that the use of automatic algorithm construction methods can minimize the possibility that human biases could affect the selection and performance of the algorithm. Both the creation of the algorithm and its operation to classify objects or predict events are to be based on concrete, observable data. The history of relations between biology and the field of machine learning is long and complex. An early technique [1] for machine learning called the perceptron constituted an attempt to model actual neuronal behavior, and the field of artificial neural network (ANN) design emerged from this attempt. Early work on the analysis of translation initiation sequences [2] employed the perceptron to define criteria for start sites in Escherichia coli. Further artificial neural network architectures such as the adaptive resonance theory (ART) [3] and neocognitron [4] were inspired from the organization of the visual nervous system. In the intervening years, the flexibility of machine learning techniques has grown along with mathematical frameworks for measuring their reliability, and it is natural to hope that machine learning methods will improve the efficiency of discovery and understanding in the mounting volume and complexity of biological data. This tutorial is structured in four main components. Firstly, a brief section reviews definitions and mathematical prerequisites. Secondly, the field of supervised learning is described. Thirdly, methods of unsupervised learning are reviewed. Finally, a section reviews methods and examples as implemented in the open source data analysis and visualization language R (http://www.r-project.org).
Adi L. Tarca, Vincent Carey, Xue-wen Chen 0001, Roberto Romero, Sorin Draghici
PLoS Comput. Biol.2
2005 Network structures and algorithms in Bioconductor
abstract
UNLABELLED: In this paper, we review the central concepts and implementations of tools for working with network structures in Bioconductor. Interfaces to open source resources for visualization (AT&T Graphviz) and network algorithms (Boost) have been developed to support analysis of graphical structures in genomics and computational biology. AVAILABILITY: Packages graph, Rgraphviz, RBGL of Bioconductor (www.bioconductor.org).
Vincent Carey, Jeff Gentry, Elizabeth Whalen, Robert Gentleman
Bioinform.1
2004 Importing MAGE-ML format microarray data into BioConductor
abstract
UNLABELLED: The microarray gene expression markup language (MAGE-ML) is a widely used XML (eXtensible Markup Language) standard for describing and exchanging information about microarray experiments. It can describe microarray designs, microarray experiment designs, gene expression data and data analysis results. We describe RMAGEML, a new Bioconductor package that provides a link between cDNA microarray data stored in MAGE-ML format and the Bioconductor framework for preprocessing, visualization and analysis of microarray experiments. AVAILABILITY: http://www.bioconductor.org. Open Source.
Steffen Durinck, Joke Allemeersch, Vincent Carey, Yves Moreau, Bart De Moor
Bioinform.3
2003 An extensible application for assembling annotation for genomic data
abstract
SUMMARY: AnnBuilder is an R package for assembling genomic annotation data. The system currently provides parsers to process annotation data from LocusLink, Gene Ontology Consortium, and Human Gene Project and can be extended to new data sources via user defined parsers. AnnBuilder differs from other existing systems in that it provides users with unlimited ability to assemble data from user selected sources. The products of AnnBuilder are files in XML format that can be easily used by different systems. AVAILABILITY: (http://www.bioconductor.org). Open source.
Vincent Carey, Robert Gentleman
Bioinform.2
1999 Evaluation of Guidelines to Help Patients Recognize Myocardial Infarction
Hamish S. F. Fraser, Lucila Ohno-Machado, Vincent Carey, Gregory Sharp, Shapur Naimi
AMIA3