Abel González-Pérez

dblp:43/2641 · also Abel González Pérez · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
3since 2021 · last 2024
0000-0002-8582-4660ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
8 papers
Bioinformatics and computational biology · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
cancer genomics
0.942019
OncodriveCLUSTL: a sequence-based clustering method to identify cancer drivers · Bioinform. 2019
OncodriveROLE classifies cancer driver genes in loss of function and activating mode of action · Bioinform. 2014
OncodriveCLUST: exploiting the positional clustering of somatic mutations to identify cancer genes · Bioinform. 2013
Bioinformatics and computational biology › genomics
genomic variant analysis
0.812024
OpenVariant: a toolkit to parse and operate multiple input file formats · Bioinform. 2024
Bioinformatics and computational biology › epigenomics › DNA methylation
DNA methylation detection
0.612022
DeepMP: a deep learning tool to detect DNA base modifications on Nanopore sequencing data · Bioinform. 2022
Bioinformatics and computational biology › sequence analysis › sequencing data analysis
nanopore signal analysis
0.612022
DeepMP: a deep learning tool to detect DNA base modifications on Nanopore sequencing data · Bioinform. 2022
Bioinformatics and computational biology › cancer genomics
cancer driver gene identification
0.522019
OncodriveCLUSTL: a sequence-based clustering method to identify cancer drivers · Bioinform. 2019
OncodriveCLUST: exploiting the positional clustering of somatic mutations to identify cancer genes · Bioinform. 2013
Bioinformatics and computational biology › cancer genomics
mutation clustering analysis
0.522019
OncodriveCLUSTL: a sequence-based clustering method to identify cancer drivers · Bioinform. 2019
OncodriveCLUST: exploiting the positional clustering of somatic mutations to identify cancer genes · Bioinform. 2013
Bioinformatics and computational biology › genomics
computational genomics
0.412020
BnpC: Bayesian non-parametric clustering of single-cell mutation profiles · Bioinform. 2020
Bioinformatics and computational biology › cancer genomics › tumor heterogeneity
intra-tumor heterogeneity
0.412020
BnpC: Bayesian non-parametric clustering of single-cell mutation profiles · Bioinform. 2020
Bioinformatics and computational biology
single-cell analysis
0.412020
Bayesian Non-parametric Clustering of Single-Cell Mutation Profiles · RECOMB 2020
Bioinformatics and computational biology › single-cell analysis › single-cell genomics
single-cell DNA sequencing
0.412020
BnpC: Bayesian non-parametric clustering of single-cell mutation profiles · Bioinform. 2020
Bioinformatics and computational biology › statistical genetics
variant effect prediction
0.112012
PARADIGM-SHIFT predicts the function of mutations in multiple cancers using pathway impact analysis · Bioinform. 2012
Bioinformatics and computational biology › cancer genomics
somatic mutation analysis
0.112019
OncodriveCLUSTL: a sequence-based clustering method to identify cancer drivers · Bioinform. 2019

Methods — techniques the papers use, named apart from their topics

bayesian non-parametric clustering · 0.9python package · 0.8convolutional neural network · 0.6genotype inference · 0.4mutation simulation · 0.4local background model · 0.4machine learning · 0.2statistical modeling · 0.2belief propagation · 0.1
YearPublicationVenuePosition
2024 OpenVariant: a toolkit to parse and operate multiple input file formats
abstract
SUMMARY: Advances in high-throughput DNA sequencing technologies and decreasing costs have fueled the identification of small genetic variants (such as single nucleotide variants and indels) across tumors. Despite efforts to standardize variant formats and vocabularies, many sources of variability persist across databases and computational tools that annotate variants, hindering their integration within cancer genomic analyses. In this context, we present OpenVariant, an easily extendable Python package that facilitates seamless reading, parsing and refinement of diverse input file formats in a customizable structure, all within a single process. AVAILABILITY AND IMPLEMENTATION: OpenVariant is an open-source package available at https://github.com/bbglab/openvariant. Documentation may be found at https://openvariant.readthedocs.io.
David Martínez-Millán, Federica Brando, Miguel L. Grau, Mònica Sánchez-Guixé, Carlos López-Elorduy, Iker Reyes-Salazar, Jordi Deu-Pons, Núria López-Bigas, Abel González-Pérez
Bioinform.9
2022 DeepMP: a deep learning tool to detect DNA base modifications on Nanopore sequencing data
abstract
MOTIVATION: DNA methylation plays a key role in a variety of biological processes. Recently, Nanopore long-read sequencing has enabled direct detection of these modifications. As a consequence, a range of computational methods have been developed to exploit Nanopore data for methylation detection. However, current approaches rely on a human-defined threshold to detect the methylation status of a genomic position and are not optimized to detect sites methylated at low frequency. Furthermore, most methods use either the Nanopore signals or the basecalling errors as the model input and do not take advantage of their combination. RESULTS: Here, we present DeepMP, a convolutional neural network-based model that takes information from Nanopore signals and basecalling errors to detect whether a given motif in a read is methylated or not. Besides, DeepMP introduces a threshold-free position modification calling model sensitive to sites methylated at low frequency across cells. We comprehensively benchmarked DeepMP against state-of-the-art methods on Escherichia coli, human and pUC19 datasets. DeepMP outperforms current approaches at read-based and position-based methylation detection across sites methylated at different frequencies in the three datasets. AVAILABILITY AND IMPLEMENTATION: DeepMP is implemented and freely available under MIT license at https://github.com/pepebonet/DeepMP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
José Bonet 0001, Mandi Chen, Marc Dabad, Simon Heath, Abel González-Pérez, Núria López-Bigas, Jens Lagergren
Bioinform.5
2022 Ten simple rules for a successful international consortium in big data omics
abstract
An African proverb says that "If you want to go fast, go alone, if you want to go far, go together."There are many scientific challenges that exceed the possibilities of an individual laboratory, a single country or even a continent.Proving the existence of the Higgs boson particle would not have been possible without the huge international effort of the Large Hadron Collider (LHC) project [1].This project brought together over 10,000 scientists from more than 100 countries around the world.Biology, however, in contrast to physics, has for a long time largely been a discipline of individual achievements.The Human Genome Project (HGP) [2] marked a key departure from an individualistic to a more collaborative approach in the field of biology.The HGP involved 20 institutes from 6 countries.Since then further large consortium projects in biology have been initiated.Each consortium united under a common goal ranging from providing fundamental information about genomes, e.g., International HapMap project [3], to tackling genomics of disease, e.g., the International Cancer Genome Consortium (ICGC) [4].The cost of sequencing DNA has decreased dramatically over the last decade and consequently the ambition of consortia to generate even larger datasets has increased.In addition, existing datasets are being combined in new studies to tackle more complex questions.In these large, international consortia funders often support their scientists locally and for many of these projects, participation depends on the level of funding committed.If a consortium does not have central funding, this poses an additional challenge of ensuring researchers live up to their promises without a "carrot and stick" at hand.The project also needs to rely on the resources consortium members provide, such as computing power, storage of the data, etc.The most critical resource, however, remains time.Time participants dedicate to the project is not enforceable when there is no central funding.The consortium may also not be the main project of the participants and the time they dedicate to it may vary.Personal motivation and engagement for the topic of the consortium may be the only things that keep the consortium going.In hindsight, vision is always 20/20, therefore we looked back at consortia related to analysing omics data that we have been part of to see what we can learn from them.One consortium in particular in which we gained a lot of experience with a large international effort in the field of big omics, without central funding, was the ICGC/TCGA PanCancer Analysis of Whole Genomes (PCAWG) project [5].This has been one of the largest biological projects aimed at getting the most out of combining existing datasets, jointly analysing nearly 2,700 cancer genomes.More than 1,300 scientists were involved from 37 different countries
Miranda D. Stobbe, Abel González-Pérez, Núria López-Bigas, Ivo Glynne Gut
PLoS Comput. Biol.2
2020 Bayesian Non-parametric Clustering of Single-Cell Mutation Profiles
Nico Borgsmüller, José Bonet 0001, Francesco Marass, Abel González-Pérez, Núria López-Bigas, Niko Beerenwinkel
RECOMB4
2020 BnpC: Bayesian non-parametric clustering of single-cell mutation profiles
abstract
MOTIVATION: The high resolution of single-cell DNA sequencing (scDNA-seq) offers great potential to resolve intratumor heterogeneity (ITH) by distinguishing clonal populations based on their mutation profiles. However, the increasing size of scDNA-seq datasets and technical limitations, such as high error rates and a large proportion of missing values, complicate this task and limit the applicability of existing methods. RESULTS: Here, we introduce BnpC, a novel non-parametric method to cluster individual cells into clones and infer their genotypes based on their noisy mutation profiles. We benchmarked our method comprehensively against state-of-the-art methods on simulated data using various data sizes, and applied it to three cancer scDNA-seq datasets. On simulated data, BnpC compared favorably against current methods in terms of accuracy, runtime and scalability. Its inferred genotypes were the most accurate, especially on highly heterogeneous data, and it was the only method able to run and produce results on datasets with 5000 cells. On tumor scDNA-seq data, BnpC was able to identify clonal populations missed by the original cluster analysis but supported by Supplementary Experimental Data. With ever growing scDNA-seq datasets, scalable and accurate methods such as BnpC will become increasingly relevant, not only to resolve ITH but also as a preprocessing step to reduce data size. AVAILABILITY AND IMPLEMENTATION: BnpC is freely available under MIT license at https://github.com/cbg-ethz/BnpC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Nico Borgsmüller, José Bonet 0001, Francesco Marass, Abel González-Pérez, Núria López-Bigas, Niko Beerenwinkel
Bioinform.4
2019 OncodriveCLUSTL: a sequence-based clustering method to identify cancer drivers
abstract
MOTIVATION: Identification of the genomic alterations driving tumorigenesis is one of the main goals in oncogenomics research. Given the evolutionary principles of cancer development, computational methods that detect signals of positive selection in the pattern of tumor mutations have been effectively applied in the search for cancer genes. One of these signals is the abnormal clustering of mutations, which has been shown to be complementary to other signals in the detection of driver genes. RESULTS: We have developed OncodriveCLUSTL, a new sequence-based clustering algorithm to detect significant clustering signals across genomic regions. OncodriveCLUSTL is based on a local background model derived from the simulation of mutations accounting for the composition of tri- or penta-nucleotide context substitutions observed in the cohort under study. Our method can identify known clusters and bona-fide cancer drivers across cohorts of tumor whole-exomes, outperforming the existing OncodriveCLUST algorithm and complementing other methods based on different signals of positive selection. Our results indicate that OncodriveCLUSTL can be applied to the analysis of non-coding genomic elements and non-human mutations data. AVAILABILITY AND IMPLEMENTATION: OncodriveCLUSTL is available as an installable Python 3.5 package. The source code and running examples are freely available at https://bitbucket.org/bbglab/oncodriveclustl under GNU Affero General Public License. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Claudia Arnedo-Pac, Loris Mularoni, Ferran Muiños, Abel González-Pérez, Núria López-Bigas
Bioinform.4
2019 OncodriveCLUSTL: a sequence-based clustering method to identify cancer drivers
abstract
Bioinformatics (2019) doi: 10.1093/bioinformatics/btz501 In the above article the final acknowledgement in the funding statement on page 3 has been changed to read ‘C.A.-P. is supported by “la Caixa” Foundation (ID 100010434) with code [LCF/BQ/ES18/11670011]’ The author apologises for this error.
Claudia Arnedo-Pac, Loris Mularoni, Ferran Muiños, Abel González-Pérez, Núria López-Bigas
Bioinform.4
2014 OncodriveROLE classifies cancer driver genes in loss of function and activating mode of action
abstract
MOTIVATION: Several computational methods have been developed to identify cancer drivers genes-genes responsible for cancer development upon specific alterations. These alterations can cause the loss of function (LoF) of the gene product, for instance, in tumor suppressors, or increase or change its activity or function, if it is an oncogene. Distinguishing between these two classes is important to understand tumorigenesis in patients and has implications for therapy decision making. Here, we assess the capacity of multiple gene features related to the pattern of genomic alterations across tumors to distinguish between activating and LoF cancer genes, and we present an automated approach to aid the classification of novel cancer drivers according to their role. RESULT: OncodriveROLE is a machine learning-based approach that classifies driver genes according to their role, using several properties related to the pattern of alterations across tumors. The method shows an accuracy of 0.93 and Matthew's correlation coefficient of 0.84 classifying genes in the Cancer Gene Census. The OncodriveROLE classifier, its results when applied to two lists of predicted cancer drivers and TCGA-derived mutation and copy number features used by the classifier are available at http://bg.upf.edu/oncodrive-role. AVAILABILITY AND IMPLEMENTATION: The R implementation of the OncodriveROLE classifier is available at http://bg.upf.edu/oncodrive-role. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Michael P. Schroeder, Carlota Rubio-Perez, David Tamborero, Abel González-Pérez, Núria López-Bigas
Bioinform.4
2013 OncodriveCLUST: exploiting the positional clustering of somatic mutations to identify cancer genes
abstract
MOTIVATION: Gain-of-function mutations often cluster in specific protein regions, a signal that those mutations provide an adaptive advantage to cancer cells and consequently are positively selected during clonal evolution of tumours. We sought to determine the overall extent of this feature in cancer and the possibility to use this feature to identify drivers. RESULTS: We have developed OncodriveCLUST, a method to identify genes with a significant bias towards mutation clustering within the protein sequence. This method constructs the background model by assessing coding-silent mutations, which are assumed not to be under positive selection and thus may reflect the baseline tendency of somatic mutations to be clustered. OncodriveCLUST analysis of the Catalogue of Somatic Mutations in Cancer retrieved a list of genes enriched by the Cancer Gene Census, prioritizing those with dominant phenotypes but also highlighting some recessive cancer genes, which showed wider but still delimited mutation clusters. Assessment of datasets from The Cancer Genome Atlas demonstrated that OncodriveCLUST selected cancer genes that were nevertheless missed by methods based on frequency and functional impact criteria. This stressed the benefit of combining approaches based on complementary principles to identify driver mutations. We propose OncodriveCLUST as an effective tool for that purpose. AVAILABILITY: OncodriveCLUST has been implemented as a Python script and is freely available from http://bg.upf.edu/oncodriveclust CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
David Tamborero, Abel González-Pérez, Núria López-Bigas
Bioinform.2
2012 PARADIGM-SHIFT predicts the function of mutations in multiple cancers using pathway impact analysis
abstract
MOTIVATION: A current challenge in understanding cancer processes is to pinpoint which mutations influence the onset and progression of disease. Toward this goal, we describe a method called PARADIGM-SHIFT that can predict whether a mutational event is neutral, gain-or loss-of-function in a tumor sample. The method uses a belief-propagation algorithm to infer gene activity from gene expression and copy number data in the context of a set of pathway interactions. RESULTS: The method was found to be both sensitive and specific on a set of positive and negative controls for multiple cancers for which pathway information was available. Application to the Cancer Genome Atlas glioblastoma, ovarian and lung squamous cancer datasets revealed several novel mutations with predicted high impact including several genes mutated at low frequency suggesting the approach will be complementary to current approaches that rely on the prevalence of events to reach statistical significance. AVAILABILITY: All source code is available at the github repository http:github.org/paradigmshift. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sam Ng, Eric A. Collisson, Artem Sokolov 0003, Theodore Goldstein, Abel González-Pérez, Núria López-Bigas, Christopher Benz, David Haussler, Joshua M. Stuart
Bioinform.5
2008 Prediction of TF target sites based on atomistic models of protein-DNA complexes
abstract
BACKGROUND: The specific recognition of genomic cis-regulatory elements by transcription factors (TFs) plays an essential role in the regulation of coordinated gene expression. Studying the mechanisms determining binding specificity in protein-DNA interactions is thus an important goal. Most current approaches for modeling TF specific recognition rely on the knowledge of large sets of cognate target sites and consider only the information contained in their primary sequence. RESULTS: Here we describe a structure-based methodology for predicting sequence motifs starting from the coordinates of a TF-DNA complex. Our algorithm combines information regarding the direct and indirect readout of DNA into an atomistic statistical model, which is used to estimate the interaction potential. We first measure the ability of our method to correctly estimate the binding specificities of eight prokaryotic and eukaryotic TFs that belong to different structural superfamilies. Secondly, the method is applied to two homology models, finding that sampling of interface side-chain rotamers remarkably improves the results. Thirdly, the algorithm is compared with a reference structural method based on contact counts, obtaining comparable predictions for the experimental complexes and more accurate sequence motifs for the homology models. CONCLUSION: Our results demonstrate that atomic-detail structural information can be feasibly used to predict TF binding sites. The computational method presented here is universal and might be applied to other systems involving protein-DNA recognition.
Vladimir Espinosa Angarica, Abel González-Pérez, Ana Tereza Ribeiro de Vasconcelos, Julio Collado-Vides, Bruno Contreras-Moreira
BMC Bioinform.2