Miguel Pérez-Enciso

dblp:57/7041 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
0since 2021 · last 2019
0000-0003-3524-995XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
gene expression analysis
0.112010
Regulatory impact factors: unraveling the transcriptional regulation of complex traits from expression data · Bioinform. 2010
Bioinformatics and computational biology › gene regulation › transcription factor analysis
transcription factor identification
0.112010
Regulatory impact factors: unraveling the transcriptional regulation of complex traits from expression data · Bioinform. 2010
Bioinformatics and computational biology › statistical genetics › gene-gene interaction
epistasis analysis
0.112006
Multiple association analysis via simulated annealing (MASSA) · Bioinform. 2006
Bioinformatics and computational biology › genomics
genome-wide association study
0.112006
Multiple association analysis via simulated annealing (MASSA) · Bioinform. 2006
Bioinformatics and computational biology › statistical genetics
association analysis
0.012004
Qxpak: a versatile mixed model application for genetical genomics and QTL analyses · Bioinform. 2004
Bioinformatics and computational biology › statistical genetics › quantitative trait locus mapping
genetical genomics
0.012004
Qxpak: a versatile mixed model application for genetical genomics and QTL analyses · Bioinform. 2004
Bioinformatics and computational biology › statistical genetics
quantitative trait locus analysis
0.012004
Qxpak: a versatile mixed model application for genetical genomics and QTL analyses · Bioinform. 2004

Methods — techniques the papers use, named apart from their topics

mixed model · 0.1differential expression analysis · 0.1differential co-expression analysis · 0.1simulated annealing · 0.1bayesian information criterion · 0.1multitrait analysis · 0.0
YearPublicationVenuePosition
2019 HIV drug resistance prediction with weighted categorical kernel functions
abstract
BACKGROUND: Antiretroviral drugs are a very effective therapy against HIV infection. However, the high mutation rate of HIV permits the emergence of variants that can be resistant to the drug treatment. Predicting drug resistance to previously unobserved variants is therefore very important for an optimum medical treatment. In this paper, we propose the use of weighted categorical kernel functions to predict drug resistance from virus sequence data. These kernel functions are very simple to implement and are able to take into account HIV data particularities, such as allele mixtures, and to weigh the different importance of each protein residue, as it is known that not all positions contribute equally to the resistance. RESULTS: We analyzed 21 drugs of four classes: protease inhibitors (PI), integrase inhibitors (INI), nucleoside reverse transcriptase inhibitors (NRTI) and non-nucleoside reverse transcriptase inhibitors (NNRTI). We compared two categorical kernel functions, Overlap and Jaccard, against two well-known noncategorical kernel functions (Linear and RBF) and Random Forest (RF). Weighted versions of these kernels were also considered, where the weights were obtained from the RF decrease in node impurity. The Jaccard kernel was the best method, either in its weighted or unweighted form, for 20 out of the 21 drugs. CONCLUSIONS: Results show that kernels that take into account both the categorical nature of the data and the presence of mixtures consistently result in the best prediction model. The advantage of including weights depended on the protein targeted by the drug. In the case of reverse transcriptase, weights based in the relative importance of each position clearly increased the prediction performance, while the improvement in the protease was much smaller. This seems to be related to the distribution of weights, as measured by the Gini index. All methods described, together with documentation and examples, are freely available at https://bitbucket.org/elies_ramon/catkern.
Elies Ramon, Lluís A. Belanche Muñoz, Miguel Pérez-Enciso
BMC Bioinform.3
2012 SNP calling by sequencing pooled samples
abstract
BACKGROUND: Performing high throughput sequencing on samples pooled from different individuals is a strategy to characterize genetic variability at a small fraction of the cost required for individual sequencing. In certain circumstances some variability estimators have even lower variance than those obtained with individual sequencing. SNP calling and estimating the frequency of the minor allele from pooled samples, though, is a subtle exercise for at least three reasons. First, sequencing errors may have a much larger relevance than in individual SNP calling: while their impact in individual sequencing can be reduced by setting a restriction on a minimum number of reads per allele, this would have a strong and undesired effect in pools because it is unlikely that alleles at low frequency in the pool will be read many times. Second, the prior allele frequency for heterozygous sites in individuals is usually 0.5 (assuming one is not analyzing sequences coming from, e.g. cancer tissues), but this is not true in pools: in fact, under the standard neutral model, singletons (i.e. alleles of minimum frequency) are the most common class of variants because P(f) ∝ 1/f and they occur more often as the sample size increases. Third, an allele appearing only once in the reads from a pool does not necessarily correspond to a singleton in the set of individuals making up the pool, and vice versa, there can be more than one read - or, more likely, none - from a true singleton. RESULTS: To improve upon existing theory and software packages, we have developed a Bayesian approach for minor allele frequency (MAF) computation and SNP calling in pools (and implemented it in a program called snape): the approach takes into account sequencing errors and allows users to choose different priors. We also set up a pipeline which can simulate the coalescence process giving rise to the SNPs, the pooling procedure and the sequencing. We used it to compare the performance of snape to that of other packages. CONCLUSIONS: We present a software which helps in calling SNPs in pooled samples: it has good power while retaining a low false discovery rate (FDR). The method also provides the posterior probability that a SNP is segregating and the full posterior distribution of f for every SNP. In order to test the behaviour of our software, we generated (through simulated coalescence) artificial genomes and computed the effect of a pooled sequencing protocol, followed by SNP calling. In this setting, snape has better power and False Discovery Rate (FDR) than the comparable packages samtools, PoPoolation, Varscan : for N = 50 chromosomes, snape has power ≈ 35%and FDR ≈ 2.5%. snape is available at http://code.google.com/p/snape-pooled/ (source code and precompiled binaries).
Emanuele Raineri, Luca Ferretti, Anna Esteve-Codina, Bruno Nevado, Simon Heath, Miguel Pérez-Enciso
BMC Bioinform.6
2012 Disease Liability Prediction from Large Scale Genotyping Data Using Classifiers with a Reject Option
abstract
Genome-wide association studies (GWA) try to identify the genetic polymorphisms associated with variation in phenotypes. However, the most significant genetic variants may have a small predictive power to forecast the future development of common diseases. We study the prediction of the risk of developing a disease given genome-wide genotypic data using classifiers with a reject option, which only make a prediction when they are sufficiently certain, but in doubtful situations may reject making a classification. To test the reliability of our proposal, we used the Wellcome Trust Case Control Consortium (WTCCC) data set, comprising 14,000 cases of seven common human diseases and 3,000 shared controls.
José Ramón Quevedo, Antonio Bahamonde, Miguel Pérez-Enciso, Oscar Luaces
IEEE ACM Trans. Comput. Biol. Bioinform.3
2011 Qxpak.5: Old mixed model solutions for new genomics problems
abstract
BACKGROUND: Mixed models have a long and fruitful history in statistics. They are pertinent to genomics problems because they are highly versatile, accommodating a wide variety of situations within the same theoretical and algorithmic framework. RESULTS: Qxpak is a package for versatile statistical genomics, specifically designed for sophisticated quantitative trait loci and association analyses. Multiple loci, multiple trait, infinitesimal genetic effects, imprinting, epistasis or sex linked loci can be fitted. The new version (v. 5) allows us, among other new features, to include either relationship matrices obtained with molecular information or user defined matrices that can be read from an input file. This feature can be used for genome selection or - more importantly - to correct for population structure in association studies. In crosses, two parental lines, not necessarily inbred, can be accommodated. CONCLUSIONS: This software aims at simplifying statistical genetic analyses implementing a coherent and unified approach by mixed models. It provides a tool that can be used in a wide variety of situations with ample genetic and statistical modeling flexibility. The software, a complete manual and examples are available at http://www.icrea.cat/Web/OtherSectionViewer.aspx?key=485&titol=Software:Qxpak.
Miguel Pérez-Enciso, Ignacy Misztal
BMC Bioinform.1
2010 Regulatory impact factors: unraveling the transcriptional regulation of complex traits from expression data
abstract
MOTIVATION: Although transcription factors (TF) play a central regulatory role, their detection from expression data is limited due to their low, and often sparse, expression. In order to fill this gap, we propose a regulatory impact factor (RIF) metric to identify critical TF from gene expression data. RESULTS: To substantiate the generality of RIF, we explore a set of experiments spanning a wide range of scenarios including breast cancer survival, fat, gonads and sex differentiation. We show that the strength of RIF lies in its ability to simultaneously integrate three sources of information into a single measure: (i) the change in correlation existing between the TF and the differentially expressed (DE) genes; (ii) the amount of differential expression of DE genes; and (iii) the abundance of DE genes. As a result, RIF analysis assigns an extreme score to those TF that are consistently most differentially co-expressed with the highly abundant and highly DE genes (RIF1), and to those TF with the most altered ability to predict the abundance of DE genes (RIF2). We show that RIF analysis alone recovers well-known experimentally validated TF for the processes studied. The TF identified confirm the importance of PPAR signaling in adipose development and the importance of transduction of estrogen signals in breast cancer survival and sexual differentiation. We argue that RIF has universal applicability, and advocate its use as a promising hypotheses generating tool for the systematic identification of novel TF not yet documented as critical.
Antonio Reverter, Nicholas J. Hudson 0001, Shivashankar H. Nagaraj, Miguel Pérez-Enciso, Brian P. Dalrymple
Bioinform.4
2006 Multiple association analysis via simulated annealing (MASSA)
abstract
SUMMARY: Genome-wide association studies are now technically feasible and likely to become a fundamental tool in unraveling the ultimate genetic basis of complex traits. However, new statistical and computational methods need to be developed to extract the maximum information in a realistic computing time. Here we propose a new method for multiple association analysis via simulated annealing that allows for epistasis and any number of markers. It consists of finding the model with lowest Bayesian information criterion using simulated annealing. The data are described by means of a mixed model and new alternative models are proposed using a set of rules, e.g. new sites can be added (or deleted), or new epistatic interactions can be included between existing genetic factors. The method is illustrated with simulated and real data. AVAILABILITY: An executable version of the program (MASSA) running under the Linux OS is freely available, together with documentation, at http://www.icrea.es/pag.asp?id=Miguel.Perez.
Miguel Pérez-Enciso
Bioinform.1
2004 Qxpak: a versatile mixed model application for genetical genomics and QTL analyses
abstract
MOTIVATION: Current methodology and software for quantitative trait loci (QTL) analyses do not use all available information and are inadequate to deal with the huge amount of QTL analyses to be needed in forecoming genetical genomics' studies. RESULTS: We show that a mixed model statistical framework provides a very flexible tool for QTL modeling in a variety of populations, be it a cross between inbred lines, a within population study, or experiments involving a mixture of populations or crosses. The software allows multitrait and multiQTL analyses, inclusion of infinitesimal genetic value and a batch multitrait option suitable for genetical genomics studies. It also allows massive association studies between single nucleotide polymorphisms and the trait(s) of interest. AVAILABILITY: A software (Qxpak), together with a manual and example files, is freely available for research purposes. So far, the compiled program is available for linux systems, the windows version will follow soon. See http://www.icrea.es/pag.asp?id=Miguel.Perez
Miguel Pérez-Enciso, Ignacy Misztal
Bioinform.1