Núria López-Bigas

dblp:62/106 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
3since 2021 · last 2024
0000-0003-4925-8988ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
13 papers
Bioinformatics and computational biology · 96% Computational science and engineering · 2% Computational social science and digital humanities · 2%

Topics — the 23 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
cancer genomics
1.152019
OncodriveCLUSTL: a sequence-based clustering method to identify cancer drivers · Bioinform. 2019
OncodriveROLE classifies cancer driver genes in loss of function and activating mode of action · Bioinform. 2014
Fast randomization of large genomic datasets while preserving alteration counts · Bioinform. 2014
Bioinformatics and computational biology › genomics
genomic variant analysis
0.812024
OpenVariant: a toolkit to parse and operate multiple input file formats · Bioinform. 2024
Bioinformatics and computational biology › epigenomics › DNA methylation
DNA methylation detection
0.612022
DeepMP: a deep learning tool to detect DNA base modifications on Nanopore sequencing data · Bioinform. 2022
Bioinformatics and computational biology › sequence analysis › sequencing data analysis
nanopore signal analysis
0.612022
DeepMP: a deep learning tool to detect DNA base modifications on Nanopore sequencing data · Bioinform. 2022
Bioinformatics and computational biology › cancer genomics
cancer driver gene identification
0.522019
OncodriveCLUSTL: a sequence-based clustering method to identify cancer drivers · Bioinform. 2019
OncodriveCLUST: exploiting the positional clustering of somatic mutations to identify cancer genes · Bioinform. 2013
Bioinformatics and computational biology › cancer genomics
mutation clustering analysis
0.522019
OncodriveCLUSTL: a sequence-based clustering method to identify cancer drivers · Bioinform. 2019
OncodriveCLUST: exploiting the positional clustering of somatic mutations to identify cancer genes · Bioinform. 2013
Bioinformatics and computational biology › genomics
computational genomics
0.412020
BnpC: Bayesian non-parametric clustering of single-cell mutation profiles · Bioinform. 2020
Bioinformatics and computational biology › cancer genomics › tumor heterogeneity
intra-tumor heterogeneity
0.412020
BnpC: Bayesian non-parametric clustering of single-cell mutation profiles · Bioinform. 2020
Bioinformatics and computational biology
single-cell analysis
0.412020
Bayesian Non-parametric Clustering of Single-Cell Mutation Profiles · RECOMB 2020
Bioinformatics and computational biology › single-cell analysis › single-cell genomics
single-cell DNA sequencing
0.412020
BnpC: Bayesian non-parametric clustering of single-cell mutation profiles · Bioinform. 2020
Computational science and engineering › scientific visualization
interactive heatmap
0.212014
jHeatmap: an interactive heatmap viewer for the web · Bioinform. 2014
Bioinformatics and computational biology › cancer genomics
mutual exclusivity analysis
0.212014
Fast randomization of large genomic datasets while preserving alteration counts · Bioinform. 2014
Bioinformatics and computational biology › omics data analysis
omics data visualization
0.212014
jHeatmap: an interactive heatmap viewer for the web · Bioinform. 2014
Bioinformatics and computational biology
biological data visualization
0.112012
SVGMap: configurable image browser for experimental data · Bioinform. 2012
Computational social science and digital humanities
spatial data visualization
0.112012
SVGMap: configurable image browser for experimental data · Bioinform. 2012
Bioinformatics and computational biology › statistical genetics
variant effect prediction
0.112012
PARADIGM-SHIFT predicts the function of mutations in multiple cancers using pathway impact analysis · Bioinform. 2012
Bioinformatics and computational biology › cancer genomics
somatic mutation analysis
0.112019
OncodriveCLUSTL: a sequence-based clustering method to identify cancer drivers · Bioinform. 2019
Bioinformatics and computational biology › genomics
disease genomics
0.112006
Highly consistent patterns for inherited human diseases at the molecular level · Bioinform. 2006
Bioinformatics and computational biology › statistical genetics › genotype-phenotype association
genotype-phenotype correlation
0.112006
Highly consistent patterns for inherited human diseases at the molecular level · Bioinform. 2006
Bioinformatics and computational biology
statistical genetics
0.112014
Fast randomization of large genomic datasets while preserving alteration counts · Bioinform. 2014
Bioinformatics and computational biology
comparative genomics
0.112005
CoGenT++: an extensive and extensible data environment for computational genomics · Bioinform. 2005
Bioinformatics and computational biology
functional genomics
0.112005
CoGenT++: an extensive and extensible data environment for computational genomics · Bioinform. 2005
Bioinformatics and computational biology
gene expression analysis
0.012012
SVGMap: configurable image browser for experimental data · Bioinform. 2012

Methods — techniques the papers use, named apart from their topics

bayesian non-parametric clustering · 0.9python package · 0.8convolutional neural network · 0.6genotype inference · 0.4mutation simulation · 0.4local background model · 0.4web-based visualization · 0.2monte carlo method · 0.2javascript library · 0.2bipartite network · 0.2
YearPublicationVenuePosition
2024 OpenVariant: a toolkit to parse and operate multiple input file formats
abstract
SUMMARY: Advances in high-throughput DNA sequencing technologies and decreasing costs have fueled the identification of small genetic variants (such as single nucleotide variants and indels) across tumors. Despite efforts to standardize variant formats and vocabularies, many sources of variability persist across databases and computational tools that annotate variants, hindering their integration within cancer genomic analyses. In this context, we present OpenVariant, an easily extendable Python package that facilitates seamless reading, parsing and refinement of diverse input file formats in a customizable structure, all within a single process. AVAILABILITY AND IMPLEMENTATION: OpenVariant is an open-source package available at https://github.com/bbglab/openvariant. Documentation may be found at https://openvariant.readthedocs.io.
David Martínez-Millán, Federica Brando, Miguel L. Grau, Mònica Sánchez-Guixé, Carlos López-Elorduy, Iker Reyes-Salazar, Jordi Deu-Pons, Núria López-Bigas, Abel González-Pérez
Bioinform.8
2022 DeepMP: a deep learning tool to detect DNA base modifications on Nanopore sequencing data
abstract
MOTIVATION: DNA methylation plays a key role in a variety of biological processes. Recently, Nanopore long-read sequencing has enabled direct detection of these modifications. As a consequence, a range of computational methods have been developed to exploit Nanopore data for methylation detection. However, current approaches rely on a human-defined threshold to detect the methylation status of a genomic position and are not optimized to detect sites methylated at low frequency. Furthermore, most methods use either the Nanopore signals or the basecalling errors as the model input and do not take advantage of their combination. RESULTS: Here, we present DeepMP, a convolutional neural network-based model that takes information from Nanopore signals and basecalling errors to detect whether a given motif in a read is methylated or not. Besides, DeepMP introduces a threshold-free position modification calling model sensitive to sites methylated at low frequency across cells. We comprehensively benchmarked DeepMP against state-of-the-art methods on Escherichia coli, human and pUC19 datasets. DeepMP outperforms current approaches at read-based and position-based methylation detection across sites methylated at different frequencies in the three datasets. AVAILABILITY AND IMPLEMENTATION: DeepMP is implemented and freely available under MIT license at https://github.com/pepebonet/DeepMP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
José Bonet 0001, Mandi Chen, Marc Dabad, Simon Heath, Abel González-Pérez, Núria López-Bigas, Jens Lagergren
Bioinform.6
2022 Ten simple rules for a successful international consortium in big data omics
abstract
An African proverb says that "If you want to go fast, go alone, if you want to go far, go together."There are many scientific challenges that exceed the possibilities of an individual laboratory, a single country or even a continent.Proving the existence of the Higgs boson particle would not have been possible without the huge international effort of the Large Hadron Collider (LHC) project [1].This project brought together over 10,000 scientists from more than 100 countries around the world.Biology, however, in contrast to physics, has for a long time largely been a discipline of individual achievements.The Human Genome Project (HGP) [2] marked a key departure from an individualistic to a more collaborative approach in the field of biology.The HGP involved 20 institutes from 6 countries.Since then further large consortium projects in biology have been initiated.Each consortium united under a common goal ranging from providing fundamental information about genomes, e.g., International HapMap project [3], to tackling genomics of disease, e.g., the International Cancer Genome Consortium (ICGC) [4].The cost of sequencing DNA has decreased dramatically over the last decade and consequently the ambition of consortia to generate even larger datasets has increased.In addition, existing datasets are being combined in new studies to tackle more complex questions.In these large, international consortia funders often support their scientists locally and for many of these projects, participation depends on the level of funding committed.If a consortium does not have central funding, this poses an additional challenge of ensuring researchers live up to their promises without a "carrot and stick" at hand.The project also needs to rely on the resources consortium members provide, such as computing power, storage of the data, etc.The most critical resource, however, remains time.Time participants dedicate to the project is not enforceable when there is no central funding.The consortium may also not be the main project of the participants and the time they dedicate to it may vary.Personal motivation and engagement for the topic of the consortium may be the only things that keep the consortium going.In hindsight, vision is always 20/20, therefore we looked back at consortia related to analysing omics data that we have been part of to see what we can learn from them.One consortium in particular in which we gained a lot of experience with a large international effort in the field of big omics, without central funding, was the ICGC/TCGA PanCancer Analysis of Whole Genomes (PCAWG) project [5].This has been one of the largest biological projects aimed at getting the most out of combining existing datasets, jointly analysing nearly 2,700 cancer genomes.More than 1,300 scientists were involved from 37 different countries
Miranda D. Stobbe, Abel González-Pérez, Núria López-Bigas, Ivo Glynne Gut
PLoS Comput. Biol.3
2020 Bayesian Non-parametric Clustering of Single-Cell Mutation Profiles
Nico Borgsmüller, José Bonet 0001, Francesco Marass, Abel González-Pérez, Núria López-Bigas, Niko Beerenwinkel
RECOMB5
2020 BnpC: Bayesian non-parametric clustering of single-cell mutation profiles
abstract
MOTIVATION: The high resolution of single-cell DNA sequencing (scDNA-seq) offers great potential to resolve intratumor heterogeneity (ITH) by distinguishing clonal populations based on their mutation profiles. However, the increasing size of scDNA-seq datasets and technical limitations, such as high error rates and a large proportion of missing values, complicate this task and limit the applicability of existing methods. RESULTS: Here, we introduce BnpC, a novel non-parametric method to cluster individual cells into clones and infer their genotypes based on their noisy mutation profiles. We benchmarked our method comprehensively against state-of-the-art methods on simulated data using various data sizes, and applied it to three cancer scDNA-seq datasets. On simulated data, BnpC compared favorably against current methods in terms of accuracy, runtime and scalability. Its inferred genotypes were the most accurate, especially on highly heterogeneous data, and it was the only method able to run and produce results on datasets with 5000 cells. On tumor scDNA-seq data, BnpC was able to identify clonal populations missed by the original cluster analysis but supported by Supplementary Experimental Data. With ever growing scDNA-seq datasets, scalable and accurate methods such as BnpC will become increasingly relevant, not only to resolve ITH but also as a preprocessing step to reduce data size. AVAILABILITY AND IMPLEMENTATION: BnpC is freely available under MIT license at https://github.com/cbg-ethz/BnpC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Nico Borgsmüller, José Bonet 0001, Francesco Marass, Abel González-Pérez, Núria López-Bigas, Niko Beerenwinkel
Bioinform.5
2019 OncodriveCLUSTL: a sequence-based clustering method to identify cancer drivers
abstract
MOTIVATION: Identification of the genomic alterations driving tumorigenesis is one of the main goals in oncogenomics research. Given the evolutionary principles of cancer development, computational methods that detect signals of positive selection in the pattern of tumor mutations have been effectively applied in the search for cancer genes. One of these signals is the abnormal clustering of mutations, which has been shown to be complementary to other signals in the detection of driver genes. RESULTS: We have developed OncodriveCLUSTL, a new sequence-based clustering algorithm to detect significant clustering signals across genomic regions. OncodriveCLUSTL is based on a local background model derived from the simulation of mutations accounting for the composition of tri- or penta-nucleotide context substitutions observed in the cohort under study. Our method can identify known clusters and bona-fide cancer drivers across cohorts of tumor whole-exomes, outperforming the existing OncodriveCLUST algorithm and complementing other methods based on different signals of positive selection. Our results indicate that OncodriveCLUSTL can be applied to the analysis of non-coding genomic elements and non-human mutations data. AVAILABILITY AND IMPLEMENTATION: OncodriveCLUSTL is available as an installable Python 3.5 package. The source code and running examples are freely available at https://bitbucket.org/bbglab/oncodriveclustl under GNU Affero General Public License. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Claudia Arnedo-Pac, Loris Mularoni, Ferran Muiños, Abel González-Pérez, Núria López-Bigas
Bioinform.5
2019 OncodriveCLUSTL: a sequence-based clustering method to identify cancer drivers
abstract
Bioinformatics (2019) doi: 10.1093/bioinformatics/btz501 In the above article the final acknowledgement in the funding statement on page 3 has been changed to read ‘C.A.-P. is supported by “la Caixa” Foundation (ID 100010434) with code [LCF/BQ/ES18/11670011]’ The author apologises for this error.
Claudia Arnedo-Pac, Loris Mularoni, Ferran Muiños, Abel González-Pérez, Núria López-Bigas
Bioinform.5
2014 jHeatmap: an interactive heatmap viewer for the web
abstract
SUMMARY: The generation of large volumes of omics data to conduct exploratory studies has become feasible and is now extensively used to gain new insights in life sciences. The effective exploration of the generated data by experts is a crucial step for the successful extraction of knowledge from these datasets. This requires availability of intuitive and interactive visualization tools that can display complex data. Matrix heatmaps are graphical representations frequently used for the description of complex omics data. Here, we present jHeatmap, a web-based tool that allows interactive matrix heatmap visualization and exploration. It is an adaptable javascript library designed to be embedded by means of basic coding skills into web portals to visualize data matrices as interactive and customizable heatmaps. AVAILABILITY: jHeatmap is freely available at the GitHub code repository at https://github.com/jheatmap/jheatmap. Working examples and the documentation may be found at http://jheatmap.github.io/jheatmap.
Jordi Deu-Pons, Michael P. Schroeder, Núria López-Bigas
Bioinform.3
2014 Fast randomization of large genomic datasets while preserving alteration counts
abstract
MOTIVATION: Studying combinatorial patterns in cancer genomic datasets has recently emerged as a tool for identifying novel cancer driver networks. Approaches have been devised to quantify, for example, the tendency of a set of genes to be mutated in a 'mutually exclusive' manner. The significance of the proposed metrics is usually evaluated by computing P-values under appropriate null models. To this end, a Monte Carlo method (the switching-algorithm) is used to sample simulated datasets under a null model that preserves patient- and gene-wise mutation rates. In this method, a genomic dataset is represented as a bipartite network, to which Markov chain updates (switching-steps) are applied. These steps modify the network topology, and a minimal number of them must be executed to draw simulated datasets independently under the null model. This number has previously been deducted empirically to be a linear function of the total number of variants, making this process computationally expensive. RESULTS: We present a novel approximate lower bound for the number of switching-steps, derived analytically. Additionally, we have developed the R package BiRewire, including new efficient implementations of the switching-algorithm. We illustrate the performances of BiRewire by applying it to large real cancer genomics datasets. We report vast reductions in time requirement, with respect to existing implementations/bounds and equivalent P-value computations. Thus, we propose BiRewire to study statistical properties in genomic datasets, and other data that can be modeled as bipartite networks. AVAILABILITY AND IMPLEMENTATION: BiRewire is available on BioConductor at http://www.bioconductor.org/packages/2.13/bioc/html/BiRewire.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Andrea Gobbi, Francesco Iorio, Kevin J. Dawson, David C. Wedge, David Tamborero, Ludmil B. Alexandrov, Núria López-Bigas, Mathew Garnett, Giuseppe Jurman, Julio Saez-Rodriguez
Bioinform.7
2014 OncodriveROLE classifies cancer driver genes in loss of function and activating mode of action
abstract
MOTIVATION: Several computational methods have been developed to identify cancer drivers genes-genes responsible for cancer development upon specific alterations. These alterations can cause the loss of function (LoF) of the gene product, for instance, in tumor suppressors, or increase or change its activity or function, if it is an oncogene. Distinguishing between these two classes is important to understand tumorigenesis in patients and has implications for therapy decision making. Here, we assess the capacity of multiple gene features related to the pattern of genomic alterations across tumors to distinguish between activating and LoF cancer genes, and we present an automated approach to aid the classification of novel cancer drivers according to their role. RESULT: OncodriveROLE is a machine learning-based approach that classifies driver genes according to their role, using several properties related to the pattern of alterations across tumors. The method shows an accuracy of 0.93 and Matthew's correlation coefficient of 0.84 classifying genes in the Cancer Gene Census. The OncodriveROLE classifier, its results when applied to two lists of predicted cancer drivers and TCGA-derived mutation and copy number features used by the classifier are available at http://bg.upf.edu/oncodrive-role. AVAILABILITY AND IMPLEMENTATION: The R implementation of the OncodriveROLE classifier is available at http://bg.upf.edu/oncodrive-role. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Michael P. Schroeder, Carlota Rubio-Perez, David Tamborero, Abel González-Pérez, Núria López-Bigas
Bioinform.5
2013 OncodriveCLUST: exploiting the positional clustering of somatic mutations to identify cancer genes
abstract
MOTIVATION: Gain-of-function mutations often cluster in specific protein regions, a signal that those mutations provide an adaptive advantage to cancer cells and consequently are positively selected during clonal evolution of tumours. We sought to determine the overall extent of this feature in cancer and the possibility to use this feature to identify drivers. RESULTS: We have developed OncodriveCLUST, a method to identify genes with a significant bias towards mutation clustering within the protein sequence. This method constructs the background model by assessing coding-silent mutations, which are assumed not to be under positive selection and thus may reflect the baseline tendency of somatic mutations to be clustered. OncodriveCLUST analysis of the Catalogue of Somatic Mutations in Cancer retrieved a list of genes enriched by the Cancer Gene Census, prioritizing those with dominant phenotypes but also highlighting some recessive cancer genes, which showed wider but still delimited mutation clusters. Assessment of datasets from The Cancer Genome Atlas demonstrated that OncodriveCLUST selected cancer genes that were nevertheless missed by methods based on frequency and functional impact criteria. This stressed the benefit of combining approaches based on complementary principles to identify driver mutations. We propose OncodriveCLUST as an effective tool for that purpose. AVAILABILITY: OncodriveCLUST has been implemented as a Python script and is freely available from http://bg.upf.edu/oncodriveclust CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
David Tamborero, Abel González-Pérez, Núria López-Bigas
Bioinform.3
2012 PARADIGM-SHIFT predicts the function of mutations in multiple cancers using pathway impact analysis
abstract
MOTIVATION: A current challenge in understanding cancer processes is to pinpoint which mutations influence the onset and progression of disease. Toward this goal, we describe a method called PARADIGM-SHIFT that can predict whether a mutational event is neutral, gain-or loss-of-function in a tumor sample. The method uses a belief-propagation algorithm to infer gene activity from gene expression and copy number data in the context of a set of pathway interactions. RESULTS: The method was found to be both sensitive and specific on a set of positive and negative controls for multiple cancers for which pathway information was available. Application to the Cancer Genome Atlas glioblastoma, ovarian and lung squamous cancer datasets revealed several novel mutations with predicted high impact including several genes mutated at low frequency suggesting the approach will be complementary to current approaches that rely on the prevalence of events to reach statistical significance. AVAILABILITY: All source code is available at the github repository http:github.org/paradigmshift. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sam Ng, Eric A. Collisson, Artem Sokolov 0003, Theodore Goldstein, Abel González-Pérez, Núria López-Bigas, Christopher Benz, David Haussler, Joshua M. Stuart
Bioinform.6
2012 SVGMap: configurable image browser for experimental data
abstract
SUMMARY: Spatial data visualization is very useful to represent biological data and quickly interpret the results. For instance, to show the expression pattern of a gene in different tissues of a fly, an intuitive approach is to draw the fly with the corresponding tissues and color the expression of the gene in each of them. However, the creation of these visual representations may be a burdensome task. Here we present SVGMap, a java application that automatizes the generation of high-quality graphics for singular data items (e.g. genes) and biological conditions. SVGMap contains a browser that allows the user to navigate the different images created and can be used as a web-based results publishing tool. AVAILABILITY: SVGMap is freely available as precompiled java package as well as source code at http://bg.upf.edu/svgmap. It requires Java 6 and any recent web browser with JavaScript enabled. The software can be run on Linux, Mac OS X and Windows systems. CONTACT: [email protected]
Xavier Rafael Palou, Michael P. Schroeder, Núria López-Bigas
Bioinform.3
2006 Highly consistent patterns for inherited human diseases at the molecular level
abstract
Over 1600 mammalian genes are known to cause an inherited disorder, when subjected to one or more mutations. These disease genes represent a unique resource for the identification and quantification of relationships between phenotypic attributes of a disease and the molecular features of the associated disease genes, including their ascribed annotated functional classes and expression patterns. Such analyses can provide a more global perspective and a deeper understanding of the probable causes underlying human hereditary diseases. In this perspective and critical view of disease genomics, we present a comparative analysis of genes reported to cause inherited diseases in humans in terms of their causative effects on physiology, their genetics and inheritance modes, the functional processes they are involved in and their expression profiles across a wide spectrum of tissues. Our analysis reveals that there are more extensive correlations between these attributes of genetic disease genes than previously appreciated. For instance, the functional pattern of genes causing dominant and recessive diseases is markedly different. Also, the function of the genes and their expression correlate with the type of disease they cause when mutated. The results further indicate that a comparative genomics approach for the analysis of genes linked to human genetic diseases will facilitate the elucidation of the underlying molecular and cellular mechanisms.
Núria López-Bigas, Benjamin J. Blencowe, Christos A. Ouzounis
Bioinform.1
2005 CoGenT++: an extensive and extensible data environment for computational genomics
abstract
MOTIVATION: CoGenT++ is a data environment for computational research in comparative and functional genomics, designed to address issues of consistency, reproducibility, scalability and accessibility. DESCRIPTION: CoGenT++ facilitates the re-distribution of all fully sequenced and published genomes, storing information about species, gene names and protein sequences. We describe our scalable implementation of ProXSim, a continually updated all-against-all similarity database, which stores pairwise relationships between all genome sequences. Based on these similarities, derived databases are generated for gene fusions--AllFuse, putative orthologs--OFAM, protein families--TRIBES, phylogenetic profiles--ProfUse and phylogenetic trees. Extensions based on the CoGenT++ environment include disease gene prediction, pattern discovery, automated domain detection, genome annotation and ancestral reconstruction. CONCLUSION: CoGenT++ provides a comprehensive environment for computational genomics, accessible primarily for large-scale analyses as well as manual browsing.
Leon Goldovsky, Paul J. Janssen, Dag G. Ahrén, Benjamin Audit, Ildefonso Cases, Nikos Darzentas, Anton J. Enright, Núria López-Bigas, José M. Peregrín-Alvarez, Mike Smith 0001, Sophia Tsoka, Victor Kunin, Christos A. Ouzounis
Bioinform.8