EDBT 2026 Demo / reviewers in the wild / expert
Veronica Vinciotti
dblp:86/2258
· DBLP profile ↗
15ranked-venue papers
4as first author
1since 2021 · last 2024
0000-0002-2625-7977ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 3 first-authorArtificial intelligence and machine learning · 6 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Probabilistic and Bayesian machine learning · 75% Graph learning · 25% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › structure learning
bayesian network structure learning |
0.8 | 1 | 2024 | Bayesian Structural Learning with Parametric Marginals for Count Data: An Application to Microbiota Systems · J. Mach. Learn. Res. 2024 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.8 | 1 | 2024 | Bayesian Structural Learning with Parametric Marginals for Count Data: An Application to Microbiota Systems · J. Mach. Learn. Res. 2024 |
Machine learning › Graph learning
graph inference |
0.8 | 1 | 2024 | Bayesian Structural Learning with Parametric Marginals for Count Data: An Application to Microbiota Systems · J. Mach. Learn. Res. 2024 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
structure learning |
0.8 | 1 | 2024 | Bayesian Structural Learning with Parametric Marginals for Count Data: An Application to Microbiota Systems · J. Mach. Learn. Res. 2024 |
Bioinformatics and computational biology › computational microbiology › microbiome analysis
microbial interaction network |
0.2 | 1 | 2024 | Bayesian Structural Learning with Parametric Marginals for Count Data: An Application to Microbiota Systems · J. Mach. Learn. Res. 2024 |
Bioinformatics and computational biology › computational microbiology
microbiome analysis |
0.2 | 1 | 2024 | Bayesian Structural Learning with Parametric Marginals for Count Data: An Application to Microbiota Systems · J. Mach. Learn. Res. 2024 |
Bioinformatics and computational biology › gene expression analysis
microarray experimental design |
0.1 | 1 | 2005 | An experimental evaluation of a loop versus a reference design for two-channel microarrays · Bioinform. 2005 |
Bioinformatics and computational biology
gene expression analysis |
0.0 | 1 | 2005 | An experimental evaluation of a loop versus a reference design for two-channel microarrays · Bioinform. 2005 |
Methods — techniques the papers use, named apart from their topics
parametric marginals · 1.5gaussian copula · 1.5bayesian inference · 1.5simulation study · 0.1reference design · 0.1loop design · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Bayesian Structural Learning with Parametric Marginals for Count Data: An Application to Microbiota SystemsabstractHigh dimensional and heterogeneous count data are collected in various applied fields. In this paper, we look closely at high-resolution sequencing data on the microbiome, which have enabled researchers to study the genomes of entire microbial communities. Revealing the underlying interactions between these communities is of vital importance to learn how microbes influence human health. To perform structural learning from multivariate count data such as these, we develop a novel Gaussian copula graphical model with two key elements. Firstly, we employ parametric regression to characterize the marginal distributions. This step is crucial for accommodating the impact of external covariates. Neglecting this adjustment could potentially introduce distortions in the inference of the underlying network of dependences. Secondly, we advance a Bayesian structure learning framework, based on a computationally efficient search algorithm that is suited to high dimensionality. The approach returns simultaneous inference of the marginal effects and of the dependence structure, including graph uncertainty estimates. A simulation study and a real data analysis of microbiome data highlight the applicability of the proposed approach at inferring networks from multivariate count data in general, and its relevance to microbiome analyses in particular. The proposed method is implemented in the R package BDgraph. Veronica Vinciotti, Pariya Behrouzi, Reza Mohammadi 0002 |
J. Mach. Learn. Res. | 1 |
| 2018 | Predicting gene expression from genome wide protein binding profilesabstractHigh-throughput technologies such as chromatin immunoprecipitation (IP) followed by next generation sequencing (ChIP-seq) in combination with gene expression studies have enabled researchers to investigate relationships between the distribution of chromosome-associated proteins and the regulation of gene transcription on a genome-wide scale. Several attempts at integrative analyses have identified direct relationships between the two processes. However, a comprehensive understanding of the regulatory events remains elusive. This is in part due to the scarcity of robust analytical methods for the detection of binding regions from ChIP-seq data. In this paper, we have applied a recently proposed Markov random field model for the detection of enriched binding regions under different biological conditions and time points. The method accounts for spatial dependencies and IP efficiencies, which can vary significantly between different experiments. We further defined the enriched chromosomal binding regions as distinct genomic features, such as promoter, exon, intron, and distal intergenic, and then investigated how predictive each of these features are of gene expression activity using machine learning techniques, including neural networks, decision trees and random forest. The analysis of a ChIP-seq time-series dataset comprising six protein markers and associated microarray data, obtained from the same biological samples, shows promising results and identified biologically plausible relationships between the protein profiles and gene regulation. Mohsina Mahmuda Ferdous, Yanchun Bao, Veronica Vinciotti, Xiaohui Liu 0001 |
Neurocomputing | 3 |
| 2016 | NEAT: an efficient network enrichment analysis testabstractBACKGROUND: Network enrichment analysis is a powerful method, which allows to integrate gene enrichment analysis with the information on relationships between genes that is provided by gene networks. Existing tests for network enrichment analysis deal only with undirected networks, they can be computationally slow and are based on normality assumptions. RESULTS: We propose NEAT, a test for network enrichment analysis. The test is based on the hypergeometric distribution, which naturally arises as the null distribution in this context. NEAT can be applied not only to undirected, but to directed and partially directed networks as well. Our simulations indicate that NEAT is considerably faster than alternative resampling-based methods, and that its capacity to detect enrichments is at least as good as the one of alternative tests. We discuss applications of NEAT to network analyses in yeast by testing for enrichment of the Environmental Stress Response target gene set with GO Slim and KEGG functional gene sets, and also by inspecting associations between functional sets themselves. CONCLUSIONS: NEAT is a flexible and efficient test for network enrichment analysis that aims to overcome some limitations of existing resampling-based tests. The method is implemented in the R package neat, which can be freely downloaded from CRAN ( https://cran.r-project.org/package=neat ). Mirko Signorelli, Veronica Vinciotti, Ernst Wit |
BMC Bioinform. | 2 |
| 2016 | Consistency of biological networks inferred from microarray and sequencing dataabstractBACKGROUND: Sparse Gaussian graphical models are popular for inferring biological networks, such as gene regulatory networks. In this paper, we investigate the consistency of these models across different data platforms, such as microarray and next generation sequencing, on the basis of a rich dataset containing samples that are profiled under both techniques as well as a large set of independent samples. RESULTS: Our analysis shows that individual node variances can have a remarkable effect on the connectivity of the resulting network. Their inconsistency across platforms and the fact that the variability level of a node may not be linked to its regulatory role mean that, failing to scale the data prior to the network analysis, leads to networks that are not reproducible across different platforms and that may be misleading. Moreover, we show how the reproducibility of networks across different platforms is significantly higher if networks are summarised in terms of enrichment amongst functional groups of interest, such as pathways, rather than at the level of individual edges. CONCLUSIONS: Careful pre-processing of transcriptional data and summaries of networks beyond individual edges can improve the consistency of network inference across platforms. However, caution is needed at this stage in the (over)interpretation of gene regulatory networks inferred from biological data. Veronica Vinciotti, Ernst Wit, Rick Jansen, Eco J. C. de Geus, Brenda W. J. H. Penninx, Dorret I. Boomsma, Peter A. C. 't Hoen |
BMC Bioinform. | 1 |
| 2016 | Multi-objective optimisation for regression testing
Wei Zheng 0006, Robert M. Hierons, Miqing Li, Xiaohui Liu 0001, Veronica Vinciotti |
Inf. Sci. | 5 |
| 2013 | Accounting for immunoprecipitation efficiencies in the statistical analysis of ChIP-seq dataabstractBACKGROUND: ImmunoPrecipitation (IP) efficiencies may vary largely between different antibodies and between repeated experiments with the same antibody. These differences have a large impact on the quality of ChIP-seq data: a more efficient experiment will necessarily lead to a higher signal to background ratio, and therefore to an apparent larger number of enriched regions, compared to a less efficient experiment. In this paper, we show how IP efficiencies can be explicitly accounted for in the joint statistical modelling of ChIP-seq data. RESULTS: We fit a latent mixture model to eight experiments on two proteins, from two laboratories where different antibodies are used for the two proteins. We use the model parameters to estimate the efficiencies of individual experiments, and find that these are clearly different for the different laboratories, and amongst technical replicates from the same lab. When we account for ChIP efficiency, we find more regions bound in the more efficient experiments than in the less efficient ones, at the same false discovery rate. A priori knowledge of the same number of binding sites across experiments can also be included in the model for a more robust detection of differentially bound regions among two different proteins. CONCLUSIONS: We propose a statistical model for the detection of enriched and differentially bound regions from multiple ChIP-seq data sets. The framework that we present accounts explicitly for IP efficiencies in ChIP-seq data, and allows to model jointly, rather than individually, replicates and experiments from different proteins, leading to more robust biological conclusions. Yanchun Bao, Veronica Vinciotti, Ernst Wit, Peter A. C. 't Hoen |
BMC Bioinform. | 2 |
| 2011 | Interspecies Translation of Disease Networks Increases Robustness and Predictive AccuracyabstractGene regulatory networks give important insights into the mechanisms underlying physiology and pathophysiology. The derivation of gene regulatory networks from high-throughput expression data via machine learning strategies is problematic as the reliability of these models is often compromised by limited and highly variable samples, heterogeneity in transcript isoforms, noise, and other artifacts. Here, we develop a novel algorithm, dubbed Dandelion, in which we construct and train intraspecies Bayesian networks that are translated and assessed on independent test sets from other species in a reiterative procedure. The interspecies disease networks are subjected to multi-layers of analysis and evaluation, leading to the identification of the most consistent relationships within the network structure. In this study, we demonstrate the performance of our algorithms on datasets from animal models of oculopharyngeal muscular dystrophy (OPMD) and patient materials. We show that the interspecies network of genes coding for the proteasome provide highly accurate predictions on gene expression levels and disease phenotype. Moreover, the cross-species translation increases the stability and robustness of these networks. Unlike existing modeling approaches, our algorithms do not require assumptions on notoriously difficult one-to-one mapping of protein orthologues or alternative transcripts and can deal with missing data. We show that the identified key components of the OPMD disease network can be confirmed in an unseen and independent disease model. This study presents a state-of-the-art strategy in constructing interspecies disease networks that provide crucial information on regulatory relationships among genes, leading to better understanding of the disease molecular mechanisms. Seyed Yahya Anvar, Allan Tucker, Veronica Vinciotti, Andrea Venema, Gert-Jan B. van Ommen, Silvere M. van der Maarel, Vered Raz, Peter A. C. 't Hoen |
PLoS Comput. Biol. | 3 |
| 2009 | An Extended Kalman Filtering Approach to Modeling Nonlinear Dynamic Gene Regulatory Networks via Short Gene Expression Time SeriesabstractIn this paper, the extended Kalman filter (EKF) algorithm is applied to model the gene regulatory network from gene time series data. The gene regulatory network is considered as a nonlinear dynamic stochastic model that consists of the gene measurement equation and the gene regulation equation. After specifying the model structure, we apply the EKF algorithm for identifying both the model parameters and the actual value of gene expression levels. It is shown that the EKF algorithm is an online estimation algorithm that can identify a large number of parameters (including parameters of nonlinear functions) through iterative procedure by using a small number of observations. Four real-world gene expression data sets are employed to demonstrate the effectiveness of the EKF algorithm, and the obtained models are evaluated from the viewpoint of bioinformatics. Zidong Wang 0001, Xiaohui Liu 0001, Yurong Liu, Jinling Liang, Veronica Vinciotti |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2006 | Exploiting the full power of temporal gene expression profiling through a new statistical test: Application to the analysis of muscular dystrophy dataabstractBACKGROUND: The identification of biologically interesting genes in a temporal expression profiling dataset is challenging and complicated by high levels of experimental noise. Most statistical methods used in the literature do not fully exploit the temporal ordering in the dataset and are not suited to the case where temporal profiles are measured for a number of different biological conditions. We present a statistical test that makes explicit use of the temporal order in the data by fitting polynomial functions to the temporal profile of each gene and for each biological condition. A Hotelling T2-statistic is derived to detect the genes for which the parameters of these polynomials are significantly different from each other. RESULTS: We validate the temporal Hotelling T2-test on muscular gene expression data from four mouse strains which were profiled at different ages: dystrophin-, beta-sarcoglycan and gamma-sarcoglycan deficient mice, and wild-type mice. The first three are animal models for different muscular dystrophies. Extensive biological validation shows that the method is capable of finding genes with temporal profiles significantly different across the four strains, as well as identifying potential biomarkers for each form of the disease. The added value of the temporal test compared to an identical test which does not make use of temporal ordering is demonstrated via a simulation study, and through confirmation of the expression profiles from selected genes by quantitative PCR experiments. The proposed method maximises the detection of the biologically interesting genes, whilst minimising false detections. CONCLUSION: The temporal Hotelling T2-test is capable of finding relatively small and robust sets of genes that display different temporal profiles between the conditions of interest. The test is simple, it can be used on gene expression data generated from any experimental design and for any number of conditions, and it allows fast interpretation of the temporal behaviour of genes. The R code is available from V.V. The microarray data have been submitted to GEO under series GSE1574 and GSE3523. Veronica Vinciotti, Xiaohui Liu 0001, Rolf Turk, Emile J. de Meijer, Peter A. C. 't Hoen |
BMC Bioinform. | 1 |
| 2006 | Temporal Bayesian classifiers for modelling muscular dystrophy expression data
Allan Tucker, Peter A. C. 't Hoen, Veronica Vinciotti, Xiaohui Liu 0001 |
Intell. Data Anal. | 3 |
| 2005 | Bayesian Network Classifiers for Time-Series Microarray Data
Allan Tucker, Veronica Vinciotti, Peter A. C. 't Hoen, Xiaohui Liu 0001 |
IDA | 2 |
| 2005 | A spatio-temporal Bayesian network classifier for understanding visual field deterioration
Allan Tucker, Veronica Vinciotti, Xiaohui Liu 0001, David F. Garway-Heath |
Artif. Intell. Medicine | 2 |
| 2005 | An experimental evaluation of a loop versus a reference design for two-channel microarraysabstractMOTIVATION: Despite theoretical arguments that so-called 'loop designs' for two-channel DNA microarray experiments are more efficient, biologists continue to use 'reference designs'. We describe two sets of microarray experiments with RNA from two different biological systems (TPA-stimulated mammalian cells and Streptomyces coelicolor). In each case, both a loop and a reference design were used with the same RNA preparations with the aim of studying their relative efficiency. RESULTS: The results of these experiments show that (1) the loop design attains a much higher precision than the reference design, (2) multiplicative spot effects are a large source of variability, and if they are not accounted for in the mathematical model, for example, by taking log-ratios or including spot effects, then the model will perform poorly. The first result is reinforced by a simulation study. Practical recommendations are given on how simple loop designs can be extended to more realistic experimental designs and how standard statistical methods allow the experimentalist to use and interpret the results from loop designs in practice. AVAILABILITY: The data and R code are available at http://exgen.ma.umist.ac.uk CONTACT: [email protected]. Veronica Vinciotti, Raya Khanin, Davide D'Alimonte, Xiaohui Liu 0001, N. Cattini, G. Hotchkiss, Giselda Bucca, O. de Jesus, J. Rasaiyaah, Colin P. Smith, Paul Kellam, Ernst Wit |
Bioinform. | 1 |
| 2005 | Computational inference of regulator activity in a single input motif from gene expression data
Raya Khanin, Veronica Vinciotti, Vassilis Mersinias, Colin P. Smith, Ernst Wit |
BMC Bioinform. | 2 |
| 2003 | Choosing k for two-class nearest neighbour classifiers with unbalanced classes
David J. Hand, Veronica Vinciotti |
Pattern Recognit. Lett. | 2 |