VLDB 2026 Research / reviewers in the wild / expert
Alfonso Buil
dblp:54/6778
· DBLP profile ↗
6ranked-venue papers
0as first author
1since 2021 · last 2021
0000-0002-3097-1014ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 100% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › biological network
network biology |
0.8 | 2 | 2021 | The effect of statistical normalization on network propagation scores · Bioinform. 2021 diffuStats: an R package to compute diffusion-based scores on biological networks · Bioinform. 2018 |
Bioinformatics and computational biology › network bioinformatics › biological network analysis
network propagation |
0.8 | 2 | 2021 | The effect of statistical normalization on network propagation scores · Bioinform. 2021 diffuStats: an R package to compute diffusion-based scores on biological networks · Bioinform. 2018 |
Bioinformatics and computational biology › genomics
genome-wide association study |
0.2 | 1 | 2016 | Fast and efficient QTL mapper for thousands of molecular phenotypes · Bioinform. 2016 |
Bioinformatics and computational biology › statistical genetics
quantitative genetics |
0.2 | 1 | 2016 | solarius: an R interface to SOLAR for variance component analysis in pedigrees · Bioinform. 2016 |
Bioinformatics and computational biology › statistical genetics
quantitative trait locus mapping |
0.2 | 1 | 2016 | Fast and efficient QTL mapper for thousands of molecular phenotypes · Bioinform. 2016 |
Bioinformatics and computational biology
variance component analysis |
0.2 | 1 | 2016 | solarius: an R interface to SOLAR for variance component analysis in pedigrees · Bioinform. 2016 |
Bioinformatics and computational biology › statistical genetics
genetic association study |
0.1 | 1 | 2010 | MISS: a non-linear methodology based on mutual information for genetic association studies in both population and sib-pairs analysis · Bioinform. 2010 |
Bioinformatics and computational biology › statistical genetics › genetic association study
multi-locus association mapping |
0.1 | 1 | 2010 | MISS: a non-linear methodology based on mutual information for genetic association studies in both population and sib-pairs analysis · Bioinform. 2010 |
Bioinformatics and computational biology › statistical genetics › gene-gene interaction
SNP interaction analysis |
0.1 | 1 | 2010 | MISS: a non-linear methodology based on mutual information for genetic association studies in both population and sib-pairs analysis · Bioinform. 2010 |
Parallel and multicore computing
parallel computing |
0.1 | 1 | 2016 | solarius: an R interface to SOLAR for variance component analysis in pedigrees · Bioinform. 2016 |
Methods — techniques the papers use, named apart from their topics
permutation analysis · 0.8variance component analysis · 0.5graph spectral analysis · 0.5statistical normalization · 0.3graph kernel · 0.3permutation procedure · 0.2beta distribution modeling · 0.2multiple linear regression · 0.1maximum entropy conditional probability modelling · 0.1information theory · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | The effect of statistical normalization on network propagation scoresabstractMOTIVATION: Network diffusion and label propagation are fundamental tools in computational biology, with applications like gene-disease association, protein function prediction and module discovery. More recently, several publications have introduced a permutation analysis after the propagation process, due to concerns that network topology can bias diffusion scores. This opens the question of the statistical properties and the presence of bias of such diffusion processes in each of its applications. In this work, we characterized some common null models behind the permutation analysis and the statistical properties of the diffusion scores. We benchmarked seven diffusion scores on three case studies: synthetic signals on a yeast interactome, simulated differential gene expression on a protein-protein interaction network and prospective gene set prediction on another interaction network. For clarity, all the datasets were based on binary labels, but we also present theoretical results for quantitative labels. RESULTS: Diffusion scores starting from binary labels were affected by the label codification and exhibited a problem-dependent topological bias that could be removed by the statistical normalization. Parametric and non-parametric normalization addressed both points by being codification-independent and by equalizing the bias. We identified and quantified two sources of bias-mean value and variance-that yielded performance differences when normalizing the scores. We provided closed formulae for both and showed how the null covariance is related to the spectral properties of the graph. Despite none of the proposed scores systematically outperformed the others, normalization was preferred when the sought positive labels were not aligned with the bias. We conclude that the decision on bias removal should be problem and data-driven, i.e. based on a quantitative analysis of the bias and its relation to the positive entities. AVAILABILITY: The code is publicly available at https://github.com/b2slab/diffuBench and the data underlying this article are available at https://github.com/b2slab/retroData. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sergio Picart-Armada, Wesley K. Thompson, Alfonso Buil, Alexandre Perera-Lluna |
Bioinform. | 3 |
| 2018 | diffuStats: an R package to compute diffusion-based scores on biological networksabstractSummary: Label propagation and diffusion over biological networks are a common mathematical formalism in computational biology for giving context to molecular entities and prioritizing novel candidates in the area of study. There are several choices in conceiving the diffusion process-involving the graph kernel, the score definitions and the presence of a posterior statistical normalization-which have an impact on the results. This manuscript describes diffuStats, an R package that provides a collection of graph kernels and diffusion scores, as well as a parallel permutation analysis for the normalized scores, that eases the computation of the scores and their benchmarking for an optimal choice. Availability and implementation: The R package diffuStats is publicly available in Bioconductor, https://bioconductor.org, under the GPL-3 license. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Sergio Picart-Armada, Wesley K. Thompson, Alfonso Buil, Alexandre Perera-Lluna |
Bioinform. | 3 |
| 2016 | Fast and efficient QTL mapper for thousands of molecular phenotypesabstractMOTIVATION: In order to discover quantitative trait loci, multi-dimensional genomic datasets combining DNA-seq and ChiP-/RNA-seq require methods that rapidly correlate tens of thousands of molecular phenotypes with millions of genetic variants while appropriately controlling for multiple testing. RESULTS: We have developed FastQTL, a method that implements a popular cis-QTL mapping strategy in a user- and cluster-friendly tool. FastQTL also proposes an efficient permutation procedure to control for multiple testing. The outcome of permutations is modeled using beta distributions trained from a few permutations and from which adjusted P-values can be estimated at any level of significance with little computational cost. The Geuvadis & GTEx pilot datasets can be now easily analyzed an order of magnitude faster than previous approaches. AVAILABILITY AND IMPLEMENTATION: Source code, binaries and comprehensive documentation of FastQTL are freely available to download at http://fastqtl.sourceforge.net/ CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Halit Ongen, Alfonso Buil, Andrew Anand Brown, Emmanouil T. Dermitzakis, Olivier Delaneau |
Bioinform. | 2 |
| 2016 | solarius: an R interface to SOLAR for variance component analysis in pedigreesabstractUNLABELLED: : The open source environment R is one of the most widely used software for statistical computing. It provides a variety of applications including statistical genetics. Most of the powerful tools for quantitative genetic analyses are stand-alone free programs developed by researchers in academia. SOLAR is one of the standard software programs to perform linkage and association mappings of the quantitative trait loci (QTLs) in pedigrees of arbitrary size and complexity. solarius allows the user to exploit the variance component methods implemented in SOLAR. It automates such routine operations as formatting pedigree and phenotype data. It parses also the model output and contains summary and plotting functions for exploration of the results. In addition, solarius enables parallel computing of the linkage and association analyses that makes the calculation of genome-wide scans more efficient. AVAILABILITY AND IMPLEMENTATION: solarius is available on CRAN and on GitHub https://github.com/ugcd/solarius CONTACT: : [email protected]. Andrey Ziyatdinov, Helena Brunel, Angel Martinez-Perez, Alfonso Buil, Alexandre Perera-Lluna, José Manuel Soria |
Bioinform. | 4 |
| 2010 | MISS: a non-linear methodology based on mutual information for genetic association studies in both population and sib-pairs analysisabstractMOTIVATION: Finding association between genetic variants and phenotypes related to disease has become an important vehicle for the study of complex disorders. In this context, multi-loci genetic association might unravel additional information when compared with single loci search. The main goal of this work is to propose a non-linear methodology based on information theory for finding combinatorial association between multi-SNPs and a given phenotype. RESULTS: The proposed methodology, called MISS (mutual information statistical significance), has been integrated jointly with a feature selection algorithm and has been tested on a synthetic dataset with a controlled phenotype and in the particular case of the F7 gene. The MISS methodology has been contrasted with a multiple linear regression (MLR) method used for genetic association in both, a population-based study and a sib-pairs analysis and with the maximum entropy conditional probability modelling (MECPM) method, which searches for predictive multi-locus interactions. Several sets of SNPs within the F7 gene region have been found to show a significant correlation with the FVII levels in blood. The proposed multi-site approach unveils combinations of SNPs that explain more significant information of the phenotype than their individual polymorphisms. MISS is able to find more correlations between SNPs and the phenotype than MLR and MECPM. Most of the marked SNPs appear in the literature as functional variants with real effect on the protein FVII levels in blood. AVAILABILITY: The code is available at http://sisbio.recerca.upc.edu/R/MISS_0.2.tar.gz Helena Brunel, Joan-Josep Gallardo-Chacón, Alfonso Buil, Montserrat Vallverdú, José Manuel Soria, Pere Caminal, Alexandre Perera-Lluna |
Bioinform. | 3 |
| 2008 | Floating Feature Selection for multiloci association of quantitative traits in sib-pairs analysisabstractFinding association between genotypic differences and disease traits has become one of the main objectives in current genetic research. It has been published that some of the underlying factors in the dynamics of the coagulation process have a genetic compound, showing significant hereditability. This is the case of the Factor VII. In this work, we propose a method for selecting sets of Single Nucleotide Polymorphisms (SNPs) of the F7 gene that are significantly related with the phenotype (Factor VII levels). The methodology is applied to the sib pairs from the GAIT project sample. The method consists of an adapted Sequential Floating Feature Selection (SFFS) algorithm. This algorithm is applied with two relevance criteria, one linear and one non linear. The SNPs sets found with linear models are included in the sets found with non linear techniques. The results fit in with previous results in clinical area. Helena Brunel, Alexandre Perera-Lluna, Alfonso Buil, Maria Sabater Lleal, Juan Carlos Souto, J. Fontcuberta, Montserrat Vallverdú, José Manuel Soria, Pere Caminal |
BIBE | 3 |