EDBT 2026 Demo / reviewers in the wild / expert
Juan R. González
dblp:49/2636
· DBLP profile ↗
22ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0003-3267-2146ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 5 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | dsOMOP: bridging OMOP CDM and DataSHIELD for secure federated analysis of standardized clinical dataabstractMOTIVATION: Collaborative clinical research projects face several challenges related to data sharing. The disparity between data standards and strict privacy regulations become more relevant as the number of involved institutions increases. To address these challenges, the scientific community has progressively adopted common data models like the Observational Medical Outcomes Partnership Common Data Model (OMOP CDM) for multicenter data standardization and implemented federated data analysis platforms like DataSHIELD to perform remote analyses without transferring individual-level data between centers, thus mitigating disclosure risks. However, there is no native implementation that automatically combines both solutions, revealing the need for a tool that enables interoperability between these systems. RESULTS: We present dsOMOP, a collection of DataSHIELD packages that facilitates automated extraction and transformation of OMOP CDM data into DataSHIELD-compatible datasets, enabling disclosure-controlled federated analyses of standardized clinical data. dsOMOP allows research institutions to provide access to their data for collaborative projects in a format that is interoperable with the project's available data, thus facilitating the analysis of large-scale, multicenter clinical data. It incorporates OMOP data directly into the DataSHIELD workflow, where all analyses occur entirely in a federated environment subject to rigorous disclosure controls, ensuring that only aggregated, non-disclosive results are ever returned to analysts. AVAILABILITY AND IMPLEMENTATION: The general information page for the dsOMOP environment is available at https://isglobal-brge.github.io/dsOMOP, where the most recent installation instructions and usage guides for all dsOMOP packages and their extensions can be found in the "Packages" section.The dsOMOP package and its complementary tools are fully available under the MIT license on GitHub: dsOMOP (https://github.com/isglobal-brge/dsOMOP), dsOMOPClient (https://github.com/isglobal-brge/dsOMOPClient), dsOMOPHelper (https://github.com/isglobal-brge/dsOMOPHelper), and dsOMOP.oracle (https://github.com/isglobal-brge/dsOMOP.oracle).Usage vignettes for the client-side packages are available at the websites of dsOMOPClient (https://isglobal-brge.github.io/dsOMOPClient) and dsOMOPHelper (https://isglobal-brge.github.io/dsOMOPHelper). A permanent archival snapshot of the exact code used in this manuscript is deposited at Figshare: https://doi.org/10.6084/m9.figshare.28607186. David Sarrat-González, Xavier Escriba-Montagut, Jared Houghtaling, Juan R. González |
Bioinform. | 4 |
| 2024 | Federated privacy-protected meta- and mega-omics data analysis in multi-center studies with a fully open-source analytic platformabstractThe importance of maintaining data privacy and complying with regulatory requirements is highlighted especially when sharing omic data between different research centers. This challenge is even more pronounced in the scenario where a multi-center effort for collaborative omics studies is necessary. OmicSHIELD is introduced as an open-source tool aimed at overcoming these challenges by enabling privacy-protected federated analysis of sensitive omic data. In order to ensure this, multiple security mechanisms have been included in the software. This innovative tool is capable of managing a wide range of omic data analyses specifically tailored to biomedical research. These include genome and epigenome wide association studies and differential gene expression analyses. OmicSHIELD is designed to support both meta- and mega-analysis, so that it offers a wide range of capabilities for different analysis designs. We present a series of use cases illustrating some examples of how the software addresses real-world analyses of omic data. Xavier Escriba-Montagut, Yannick Marcon, Augusto Anguita-Ruiz, Demetris Avraam, Jose Urquiza, Andrei S. Morgan, Rebecca C. Wilson, Paul R. Burton, Juan R. González |
PLoS Comput. Biol. | 9 |
| 2022 | Fully exploiting SNP arrays: a systematic review on the tools to extract underlying genomic structureabstractSingle nucleotide polymorphisms (SNPs) are the most abundant type of genomic variation and the most accessible to genotype in large cohorts. However, they individually explain a small proportion of phenotypic differences between individuals. Ancestry, collective SNP effects, structural variants, somatic mutations or even differences in historic recombination can potentially explain a high percentage of genomic divergence. These genetic differences can be infrequent or laborious to characterize; however, many of them leave distinctive marks on the SNPs across the genome allowing their study in large population samples. Consequently, several methods have been developed over the last decade to detect and analyze different genomic structures using SNP arrays, to complement genome-wide association studies and determine the contribution of these structures to explain the phenotypic differences between individuals. We present an up-to-date collection of available bioinformatics tools that can be used to extract relevant genomic information from SNP array data including population structure and ancestry; polygenic risk scores; identity-by-descent fragments; linkage disequilibrium; heritability and structural variants such as inversions, copy number variants, genetic mosaicisms and recombination histories. From a systematic review of recently published applications of the methods, we describe the main characteristics of R packages, command-line tools and desktop applications, both free and commercial, to help make the most of a large amount of publicly available SNP data. Laura Balagué-Dobón, Alejandro Cáceres, Juan R. González |
Briefings Bioinform. | 3 |
| 2022 | teff: estimation of Treatment EFFects on transcriptomic data using causal random forestabstractMOTIVATION: Causal inference on high-dimensional feature data can be used to find a profile of patients who will benefit the most from treatment rather than no treatment. However, there is a need for usable implementations for transcriptomic data. We developed teff that applies random causal forest on gene expression data to target individuals with high expected treatment effects. RESULTS: We extracted a profile of high benefit of treating psoriasis with brodalumab and observed that it was associated with higher T cell abundance in non-lesional skin at baseline and a lower response for etanercept in an independent study. Individual patient targeting with causal inference profiling can inform patients on choosing between treatments before the intervention begins. AVAILABILITY AND IMPLEMENTATION: teff is an R package available at https://teff-package.github.io. The data underlying this article are available in GEO, at https://www.ncbi.nlm.nih.gov/geo/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alejandro Cáceres, Juan R. González |
Bioinform. | 2 |
| 2021 | methylclock: a Bioconductor package to estimate DNA methylation ageabstractMOTIVATION: Ageing is a biological and psychosocial process related to diseases and mortality. It correlates with changes in DNA methylation (DNAm) in all human tissues. Therefore, epigenetic markers can be used to estimate biological age using DNAm profiling across tissues. RESULTS: We developed a Bioconductor package that allows computation of several existing DNAm adult/childhood and gestational age clocks. Functions to visualize the DNAm age prediction versus chronological age and the correlation between DNAm clocks are also available as well as other features, such as missing data imputation of cell types' estimates, that are required for DNAm age clocks. AVAILABILITY AND IMPLEMENTATION: https://github.com/isglobal-brge/methylclock. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dolors Pelegí-Sisó, Paula de Prado, Justiina Ronkainen, Mariona Bustamante, Juan R. González |
Bioinform. | 5 |
| 2021 | Orchestrating privacy-protected big data analyses of data from different resources with R and DataSHIELDabstractCombined analysis of multiple, large datasets is a common objective in the health- and biosciences. Existing methods tend to require researchers to physically bring data together in one place or follow an analysis plan and share results. Developed over the last 10 years, the DataSHIELD platform is a collection of R packages that reduce the challenges of these methods. These include ethico-legal constraints which limit researchers' ability to physically bring data together and the analytical inflexibility associated with conventional approaches to sharing results. The key feature of DataSHIELD is that data from research studies stay on a server at each of the institutions that are responsible for the data. Each institution has control over who can access their data. The platform allows an analyst to pass commands to each server and the analyst receives results that do not disclose the individual-level data of any study participants. DataSHIELD uses Opal which is a data integration system used by epidemiological studies and developed by the OBiBa open source project in the domain of bioinformatics. However, until now the analysis of big data with DataSHIELD has been limited by the storage formats available in Opal and the analysis capabilities available in the DataSHIELD R packages. We present a new architecture ("resources") for DataSHIELD and Opal to allow large, complex datasets to be used at their original location, in their original format and with external computing facilities. We provide some real big data analysis examples in genomics and geospatial projects. For genomic data analyses, we also illustrate how to extend the resources concept to address specific big data infrastructures such as GA4GH or EGA, and make use of shell commands. Our new infrastructure will help researchers to perform data analyses in a privacy-protected way from existing data sharing initiatives or projects. To help researchers use this framework, we describe selected packages and present an online book (https://isglobal-brge.github.io/resource_bookdown). Yannick Marcon, Tom Bishop, Demetris Avraam, Xavier Escriba-Montagut, Patricia Ryser-Welch, Stuart Wheater, Paul R. Burton, Juan R. González |
PLoS Comput. Biol. | 8 |
| 2020 | MADloy: robust detection of mosaic loss of chromosome Y from genotype-array-intensity dataabstractBACKGROUND: Accurate protocols and methods to robustly detect the mosaic loss of chromosome Y (mLOY) are needed given its reported role in cancer, several age-related disorders and overall male mortality. Intensity SNP-array data have been used to infer mLOY status and to determine its prominent role in male disease. However, discrepancies of reported findings can be due to the uncertainty and variability of the methods used for mLOY detection and to the differences in the tissue-matrix used. RESULTS: We created a publicly available software tool called MADloy (Mosaic Alteration Detection for LOY) that incorporates existing methods and includes a new robust approach, allowing efficient calling in large studies and comparisons between methods. MADloy optimizes mLOY calling by correctly modeling the underlying reference population with no-mLOY status and incorporating B-deviation information. We observed improvements in the calling accuracy to previous methods, using experimentally validated samples, and an increment in the statistical power to detect associations with disease and mortality, using simulation studies and real dataset analyses. To understand discrepancies in mLOY detection across different tissues, we applied MADloy to detect the increment of mLOY cellularity in blood on 18 individuals after 3 years and to confirm that its detection in saliva was sub-optimal (41%). We additionally applied MADloy to detect the down-regulation genes in the chromosome Y in kidney and bladder tumors with mLOY, and to perform pathway analyses for the detection of mLOY in blood. CONCLUSIONS: MADloy is a new software tool implemented in R for the easy and robust calling of mLOY status across different tissues aimed to facilitate its study in large epidemiological studies. Juan R. González, Marcos López-Sánchez, Alejandro Cáceres, Pere Puig, Tõnu Esko, Luis A. Pérez-Jurado |
BMC Bioinform. | 1 |
| 2019 | Comprehensive study of the exposome and omic data using rexposome Bioconductor PackagesabstractSUMMARY: Genomics has dramatically improved our understanding of the molecular origins of certain human diseases. Nonetheless, our health is also influenced by the cumulative impact of exposures experienced across the life course (termed 'exposome'). The study of the high-dimensional exposome offers a new paradigm for investigating environmental contributions to disease etiology. However, there is a lack of bioinformatics tools for managing, visualizing and analyzing the exposome. The analysis data should include both association with health outcomes and integration with omic layers. We provide a generic framework called rexposome project, developed in the R/Bioconductor architecture that includes object-oriented classes and methods to leverage high-dimensional exposome data in disease association studies including its integration with a variety of high-throughput data types. The usefulness of the package is illustrated by analyzing a real dataset including exposome data, three health outcomes related to respiratory diseases and its integration with the transcriptome and methylome. AVAILABILITY AND IMPLEMENTATION: rexposome project is available at https://isglobal-brge.github.io/rexposome/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Carles Hernandez-Ferrer, Gregory Wellenius, Ibon Tamayo, Xavier Basagaña, Jordi Sunyer, Martine Vrijheid, Juan R. González |
Bioinform. | 7 |
| 2018 | CTDquerier: a bioconductor R package for Comparative Toxicogenomics DatabaseTM data extraction, visualization and enrichment of environmental and toxicological studiesabstractSummary: Biomedical studies currently include a large volume of genomic and environmental factors for studying the etiology of human diseases. R/Bioconductor projects provide several tools for performing enrichment analysis at gene-pathway level, allowing researchers to develop novel hypotheses. However, there is a need to perform similar analyses at the chemicals-genes or chemicals-diseases levels to provide complementary knowledge of the causal path between chemicals and diseases. While the Comparative Toxicogenomics DatabaseTM (CTD) provides information about these relationships, there is no software for integrating it into R/Bioconductor analysis pipelines. CTDquerier helps users to easily download CTD data and integrate it in the R/Bioconductor framework. The package also contains functions for visualizing CTD data and performing enrichment analyses. We illustrate how to use the package with a real data analysis of asthma-related genes. CTDquerier is a flexible and easy-to-use Bioconductor package that provides novel hypothesis about the relationships between chemicals and diseases. Availability and implementation: CTDquerier R package is available through Bioconductor and its development version at https://github.com/isglobal-brge/CTDquerier. Supplementary information: Supplementary data are available at Bioinformatics online. Carles Hernandez-Ferrer, Juan R. González |
Bioinform. | 2 |
| 2017 | psygenet2r: a R/Bioconductor package for the analysis of psychiatric disease genesabstractMOTIVATION: Psychiatric disorders have a great impact on morbidity and mortality. Genotype-phenotype resources for psychiatric diseases are key to enable the translation of research findings to a better care of patients. PsyGeNET is a knowledge resource on psychiatric diseases and their genes, developed by text mining and curated by domain experts. RESULTS: We present psygenet2r, an R package that contains a variety of functions for leveraging PsyGeNET database and facilitating its analysis and interpretation. The package offers different types of queries to the database along with variety of analysis and visualization tools, including the study of the anatomical structures in which the genes are expressed and gaining insight of gene's molecular function. Psygenet2r is especially suited for network medicine analysis of psychiatric disorders. AVAILABILITY AND IMPLEMENTATION: The package is implemented in R and is available under MIT license from Bioconductor (http://bioconductor.org/packages/release/bioc/html/psygenet2r.html). CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alba Gutiérrez-Sacristán, Carles Hernandez-Ferrer, Juan R. González, Laura Inés Furlong |
Bioinform. | 3 |
| 2017 | MultiDataSet: an R package for encapsulating multiple data sets with application to omic data integrationabstractBACKGROUND: Reduction in the cost of genomic assays has generated large amounts of biomedical-related data. As a result, current studies perform multiple experiments in the same subjects. While Bioconductor's methods and classes implemented in different packages manage individual experiments, there is not a standard class to properly manage different omic datasets from the same subjects. In addition, most R/Bioconductor packages that have been designed to integrate and visualize biological data often use basic data structures with no clear general methods, such as subsetting or selecting samples. RESULTS: To cover this need, we have developed MultiDataSet, a new R class based on Bioconductor standards, designed to encapsulate multiple data sets. MultiDataSet deals with the usual difficulties of managing multiple and non-complete data sets while offering a simple and general way of subsetting features and selecting samples. We illustrate the use of MultiDataSet in three common situations: 1) performing integration analysis with third party packages; 2) creating new methods and functions for omic data integration; 3) encapsulating new unimplemented data from any biological experiment. CONCLUSIONS: MultiDataSet is a suitable class for data integration under R and Bioconductor framework. Carles Hernandez-Ferrer, Carlos Ruiz-Arenas, Alba Beltran-Gomila, Juan R. González |
BMC Bioinform. | 4 |
| 2017 | Redundancy analysis allows improved detection of methylation changes in large genomic regionsabstractBACKGROUND: DNA methylation is an epigenetic process that regulates gene expression. Methylation can be modified by environmental exposures and changes in the methylation patterns have been associated with diseases. Methylation microarrays measure methylation levels at more than 450,000 CpGs in a single experiment, and the most common analysis strategy is to perform a single probe analysis to find methylation probes associated with the outcome of interest. However, methylation changes usually occur at the regional level: for example, genomic structural variants can affect methylation patterns in regions up to several megabases in length. Existing DMR methods provide lists of Differentially Methylated Regions (DMRs) of up to only few kilobases in length, and cannot check if a target region is differentially methylated. Therefore, these methods are not suitable to evaluate methylation changes in large regions. To address these limitations, we developed a new DMR approach based on redundancy analysis (RDA) that assesses whether a target region is differentially methylated. RESULTS: Using simulated and real datasets, we compared our approach to three common DMR detection methods (Bumphunter, blockFinder, and DMRcate). We found that Bumphunter underestimated methylation changes and blockFinder showed poor performance. DMRcate showed poor power in the simulated datasets and low specificity in the real data analysis. Our method showed very high performance in all simulation settings, even with small sample sizes and subtle methylation changes, while controlling type I error. Other advantages of our method are: 1) it estimates the degree of association between the DMR and the outcome; 2) it can analyze a targeted or region of interest; and 3) it can evaluate the simultaneous effects of different variables. The proposed methodology is implemented in MEAL, a Bioconductor package designed to facilitate the analysis of methylation data. CONCLUSIONS: We propose a multivariate approach to decipher whether an outcome of interest alters the methylation pattern of a region of interest. The method is designed to analyze large target genomic regions and outperforms the three most popular methods for detecting DMRs. Our method can evaluate factors with more than two levels or the simultaneous effect of more than one continuous variable, which is not possible with the state-of-the-art methods. Carlos Ruiz-Arenas, Juan R. González |
BMC Bioinform. | 2 |
| 2016 | Efficient and Powerful Method for Combining P-Values in Genome-Wide Association StudiesabstractThe goal of Genome-wide Association Studies (GWAS) is the identification of genetic variants, usually single nucleotide polymorphisms (SNPs), that are associated with disease risk. However, SNPs detected so far with GWAS for most common diseases only explain a small proportion of their total heritability. Gene set analysis (GSA) has been proposed as an alternative to single-SNP analysis with the aim of improving the power of genetic association studies. Nevertheless, most GSA methods rely on expensive computational procedures that make unfeasible their implementation in GWAS. We propose a new GSA method, referred as globalEVT, which uses the extreme value theory to derive gene-level p-values. GlobalEVT reduces dramatically the computational requirements compared to other GSA approaches. In addition, this new approach improves the power by allowing different inheritance models for each genetic variant as illustrated in the simulation study performed and allows the existence of correlation between the SNPs. Real data analysis of an Attention-deficit/hyperactivity disorder (ADHD) study illustrates the importance of using GSA approaches for exploring new susceptibility genes. Specifically, the globalEVT method is able to detect genes related to Cyclophilin A like domain proteins which is known to play an important role in the mechanisms of ADHD development. Natalia Vilor-Tejedor, Juan R. González, Malu Luz Calle |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2015 | affy2svaffy2sv: an R package to pre-process Affymetrix CytoScan HD and 750K arrays for SNP, CNV, inversion and mosaicism callingabstractBACKGROUND: The well-known Genome-Wide Association Studies (GWAS) had led to many scientific discoveries using SNP data. Even so, they were not able to explain the full heritability of complex diseases. Now, other structural variants like copy number variants or DNA inversions, either germ-line or in mosaicism events, are being studies. We present the R package affy2sv to pre-process Affymetrix CytoScan HD/750k array (also for Genome-Wide SNP 5.0/6.0 and Axiom) in structural variant studies. RESULTS: We illustrate the capabilities of affy2sv using two different complete pipelines on real data. The first one performing a GWAS and a mosaic alterations detection study, and the other detecting CNVs and performing an inversion calling. CONCLUSION: Both examples presented in the article show up how affy2sv can be used as part of more complex pipelines aimed to analyze Affymetrix SNP arrays data in genetic association studies, where different types of structural variants are considered. Carles Hernandez-Ferrer, Ines Quintela Garcia, Katharina Danielski, Ángel Carracedo, Luis A. Pérez-Jurado, Juan R. González |
BMC Bioinform. | 6 |
| 2013 | A flexible count data model to fit the wide diversity of expression profiles arising from extensively replicated RNA-seq experimentsabstractBACKGROUND: High-throughput RNA sequencing (RNA-seq) offers unprecedented power to capture the real dynamics of gene expression. Experimental designs with extensive biological replication present a unique opportunity to exploit this feature and distinguish expression profiles with higher resolution. RNA-seq data analysis methods so far have been mostly applied to data sets with few replicates and their default settings try to provide the best performance under this constraint. These methods are based on two well-known count data distributions: the Poisson and the negative binomial. The way to properly calibrate them with large RNA-seq data sets is not trivial for the non-expert bioinformatics user. RESULTS: Here we show that expression profiles produced by extensively-replicated RNA-seq experiments lead to a rich diversity of count data distributions beyond the Poisson and the negative binomial, such as Poisson-Inverse Gaussian or Pólya-Aeppli, which can be captured by a more general family of count data distributions called the Poisson-Tweedie. The flexibility of the Poisson-Tweedie family enables a direct fitting of emerging features of large expression profiles, such as heavy-tails or zero-inflation, without the need to alter a single configuration parameter. We provide a software package for R called tweeDEseq implementing a new test for differential expression based on the Poisson-Tweedie family. Using simulations on synthetic and real RNA-seq data we show that tweeDEseq yields P-values that are equally or more accurate than competing methods under different configuration parameters. By surveying the tiny fraction of sex-specific gene expression changes in human lymphoblastoid cell lines, we also show that tweeDEseq accurately detects differentially expressed genes in a real large RNA-seq data set with improved performance and reproducibility over the previously compared methodologies. Finally, we compared the results with those obtained from microarrays in order to check for reproducibility. CONCLUSIONS: RNA-seq data with many replicates leads to a handful of count data distributions which can be accurately estimated with the statistical model illustrated in this paper. This method provides a better fit to the underlying biological variability; this may be critical when comparing groups of RNA-seq samples with markedly different count data distributions. The tweeDEseq package forms part of the Bioconductor project and it is available for download at http://www.bioconductor.org. Mikel Esnaola, Pedro Puig, Robert Castelo, Juan R. González |
BMC Bioinform. | 5 |
| 2012 | Identification of polymorphic inversions from genotypesabstractBACKGROUND: Polymorphic inversions are a source of genetic variability with a direct impact on recombination frequencies. Given the difficulty of their experimental study, computational methods have been developed to infer their existence in a large number of individuals using genome-wide data of nucleotide variation. Methods based on haplotype tagging of known inversions attempt to classify individuals as having a normal or inverted allele. Other methods that measure differences between linkage disequilibrium attempt to identify regions with inversions but unable to classify subjects accurately, an essential requirement for association studies. RESULTS: We present a novel method to both identify polymorphic inversions from genome-wide genotype data and classify individuals as containing a normal or inverted allele. Our method, a generalization of a published method for haplotype data 1, utilizes linkage between groups of SNPs to partition a set of individuals into normal and inverted subpopulations. We employ a sliding window scan to identify regions likely to have an inversion, and accumulation of evidence from neighboring SNPs is used to accurately determine the inversion status of each subject. Further, our approach detects inversions directly from genotype data, thus increasing its usability to current genome-wide association studies (GWAS). CONCLUSIONS: We demonstrate the accuracy of our method to detect inversions and classify individuals on principled-simulated genotypes, produced by the evolution of an inversion event within a coalescent model 2. We applied our method to real genotype data from HapMap Phase III to characterize the inversion status of two known inversions within the regions 17q21 and 8p23 across 1184 individuals. Finally, we scan the full genomes of the European Origin (CEU) and Yoruba (YRI) HapMap samples. We find population-based evidence for 9 out of 15 well-established autosomic inversions, and for 52 regions previously predicted by independent experimental methods in ten (9+1) individuals 34. We provide efficient implementations of both genotype and haplotype methods as a unified R package inveRsion. Alejandro Cáceres, Suzanne Sindi, Benjamin J. Raphael, Mario Cáceres, Juan R. González |
BMC Bioinform. | 5 |
| 2011 | MLPAstats: An R GUI package for the integrated analysis of copy number alterations using MLPA dataabstractBACKGROUND: Multiplex-Dependent Probe Amplification (MLPA) is a cost-effective experimental method for candidate gene studies, aimed at the identification of copy number alterations. The analysis of such genetic variants, from electropherogram peak intensities, involves two main stages. First, peak normalization for each probe is required to remove the contribution of probe size to peak intensity. Second, the statistical significance of peak alteration between case and control samples is estimated. A number of methods have been proposed in each step with varying levels of complexity and precision. However, there is no single framework from which the results of each method and possible combinations at each step can be assessed. RESULTS: We present MLPAstats, an R package designed to integrate the methods for exploring different analysis scenarios in a reliable way. A GUI has been developed to allow researchers to find their optimal analysis strategy. CONCLUSIONS: MLPAstats is an analysis tool that promotes the use of cost-effective MLPA suitable for candidate gene studies. Its R implementation allows future methods to be easily incorporated, while its GUI will facilitate its use by non-expert analysts. A vignette describing a set-by-step tutorial is also available with the package. Alejandro Cáceres, Lluís Armengol, Sergi Villatoro, Juan R. González |
BMC Bioinform. | 4 |
| 2011 | A fast and accurate method to detect allelic genomic imbalances underlying mosaic rearrangements using SNP array dataabstractBACKGROUND: Mosaicism for copy number and copy neutral chromosomal rearrangements has been recently identified as a relatively common source of genetic variation in the normal population. However its prevalence is poorly defined since it has been only studied systematically in one large-scale study and by using non optimal ad-hoc SNP array data analysis tools, uncovering rather large alterations (> 1 Mb) and affecting a high proportion of cells. Here we propose a novel methodology, Mosaic Alteration Detection-MAD, by providing a software tool that is effective for capturing previously described alterations as wells as new variants that are smaller in size and/or affecting a low percentage of cells. RESULTS: The developed method identified all previously known mosaic abnormalities reported in SNP array data obtained from controls, bladder cancer and HapMap individuals. In addition MAD tool was able to detect new mosaic variants not reported before that were smaller in size and with lower percentage of cells affected. The performance of the tool was analysed by studying simulated data for different scenarios. Our method showed high sensitivity and specificity for all assessed scenarios. CONCLUSIONS: The tool presented here has the ability to identify mosaic abnormalities with high sensitivity and specificity. Our results confirm the lack of sensitivity of former methods by identifying new mosaic variants not reported in previously utilised datasets. Our work suggests that the prevalence of mosaic alterations could be higher than initially thought. The use of appropriate SNP array data analysis methods would help in defining the human genome mosaic map. Juan R. González, Benjamin Rodriguez-Santiago, Alejandro Cáceres, Roger Pique-Regi, Nathaniel Rothman, Stephen J. Chanock, Lluís Armengol, Luis A. Pérez-Jurado |
BMC Bioinform. | 1 |
| 2010 | R-Gada: a fast and flexible pipeline for copy number analysis in association studiesabstractBACKGROUND: Genome-wide association studies (GWAS) using Copy Number Variation (CNV) are becoming a central focus of genetic research. CNVs have successfully provided target genome regions for some disease conditions where simple genetic variation (i.e., SNPs) has previously failed to provide a clear association. RESULTS: Here we present a new R package, that integrates: (i) data import from most common formats of Affymetrix, Illumina and aCGH arrays; (ii) a fast and accurate segmentation algorithm to call CNVs based on Genome Alteration Detection Analysis (GADA); and (iii) functions for displaying and exporting the Copy Number calls, identification of recurrent CNVs, multivariate analysis of population structure, and tools for performing association studies. Using a large dataset containing 270 HapMap individuals (Affymetrix Human SNP Array 6.0 Sample Dataset) we demonstrate a flexible pipeline implemented with the package. It requires less than one minute per sample (3 million probe arrays) on a single core computer, and provides a flexible parallelization for very large datasets. Case-control data were generated from the HapMap dataset to demonstrate a GWAS analysis. CONCLUSIONS: The package provides the tools for creating a complete integrated pipeline from data normalization to statistical association. It can efficiently handle a massive volume of data consisting of millions of genetic markers and hundreds or thousands of samples with very accurate results. Roger Pique-Regi, Alejandro Cáceres, Juan R. González |
BMC Bioinform. | 3 |
| 2009 | Accounting for uncertainty when assessing association between copy number and disease: a latent class modelabstractBACKGROUND: Copy number variations (CNVs) may play an important role in disease risk by altering dosage of genes and other regulatory elements, which may have functional and, ultimately, phenotypic consequences. Therefore, determining whether a CNV is associated or not with a given disease might be relevant in understanding the genesis and progression of human diseases. Current stage technology give CNV probe signal from which copy number status is inferred. Incorporating uncertainty of CNV calling in the statistical analysis is therefore a highly important aspect. In this paper, we present a framework for assessing association between CNVs and disease in case-control studies where uncertainty is taken into account. We also indicate how to use the model to analyze continuous traits and adjust for confounding covariates. RESULTS: Through simulation studies, we show that our method outperforms other simple methods based on inferring the underlying CNV and assessing association using regular tests that do not propagate call uncertainty. We apply the method to a real data set in a controlled MLPA experiment showing good results. The methodology is also extended to illustrate how to analyze aCGH data. CONCLUSION: We demonstrate that our method is robust and achieves maximal theoretical power since it accommodates uncertainty when copy number status are inferred. We have made R functions freely available. Juan R. González, Isaac Subirana, Geòrgia Escaramís, Solymar Peraza, Alejandro Cáceres, Xavier Estivill, Lluís Armengol |
BMC Bioinform. | 1 |
| 2008 | Probe-specific mixed-model approach to detect copy number differences using multiplex ligation-dependent probe amplification (MLPA)abstractBACKGROUND: MLPA method is a potentially useful semi-quantitative method to detect copy number alterations in targeted regions. In this paper, we propose a method for the normalization procedure based on a non-linear mixed-model, as well as a new approach for determining the statistical significance of altered probes based on linear mixed-model. This method establishes a threshold by using different tolerance intervals that accommodates the specific random error variability observed in each test sample. RESULTS: Through simulation studies we have shown that our proposed method outperforms two existing methods that are based on simple threshold rules or iterative regression. We have illustrated the method using a controlled MLPA assay in which targeted regions are variable in copy number in individuals suffering from different disorders such as Prader-Willi, DiGeorge or Autism showing the best performace. CONCLUSION: Using the proposed mixed-model, we are able to determine thresholds to decide whether a region is altered. These threholds are specific for each individual, incorporating experimental variability, resulting in improved sensitivity and specificity as the examples with real data have revealed. Juan R. González, Josep L. Carrasco, Lluís Armengol, Sergi Villatoro, Lluís Jover, Yutaka Yasui, Xavier Estivill |
BMC Bioinform. | 1 |
| 2007 | SNPassoc: an R package to perform whole genome association studiesabstractUNLABELLED: The popularization of large-scale genotyping projects has led to the widespread adoption of genetic association studies as the tool of choice in the search for single nucleotide polymorphisms (SNPs) underlying susceptibility to complex diseases. Although the analysis of individual SNPs is a relatively trivial task, when the number is large and multiple genetic models need to be explored it becomes necessary a tool to automate the analyses. In order to address this issue, we developed SNPassoc, an R package to carry out most common analyses in whole genome association studies. These analyses include descriptive statistics and exploratory analysis of missing values, calculation of Hardy-Weinberg equilibrium, analysis of association based on generalized linear models (either for quantitative or binary traits), and analysis of multiple SNPs (haplotype and epistasis analysis). AVAILABILITY: Package SNPassoc is available at CRAN from http://cran.r-project.org. SUPPLEMENTARY INFORMATION: A tutorial is available on Bioinformatics online and in http://davinci.crg.es/estivill_lab/snpassoc. Juan R. González, Lluís Armengol, Xavier Solé, Elisabet Guinó, Josep M. Mercader, Xavier Estivill, Víctor Moreno |
Bioinform. | 1 |