VLDB 2026 Research / reviewers in the wild / expert
Peter Horvatovich
dblp:63/2159
· DBLP profile ↗
6ranked-venue papers
0as first author
3since 2021 · last 2023
0000-0003-2218-1140ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 98% Computational science and engineering · 2% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › biological network › network biology › network inference › gene regulatory network inference
gaussian graphical model |
1.0 | 2 | 2022 | GeneNetTools: tests for Gaussian graphical models with shrinkage · Bioinform. 2022 Exact hypothesis testing for shrinkage-based Gaussian graphical models · Bioinform. 2019 |
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference |
1.0 | 2 | 2022 | GeneNetTools: tests for Gaussian graphical models with shrinkage · Bioinform. 2022 Exact hypothesis testing for shrinkage-based Gaussian graphical models · Bioinform. 2019 |
Bioinformatics and computational biology
statistical genetics |
1.0 | 2 | 2022 | GeneNetTools: tests for Gaussian graphical models with shrinkage · Bioinform. 2022 Exact hypothesis testing for shrinkage-based Gaussian graphical models · Bioinform. 2019 |
Bioinformatics and computational biology › network bioinformatics › biological network analysis
differential network analysis |
0.6 | 1 | 2022 | GeneNetTools: tests for Gaussian graphical models with shrinkage · Bioinform. 2022 |
Bioinformatics and computational biology
proteomics |
0.2 | 2 | 2011 | A high-throughput processing service for retention time alignment of complex proteomics and metabolomics LC-MS data · Bioinform. 2011 A noise model for mass spectrometry based proteomics · Bioinform. 2008 |
Bioinformatics and computational biology
metabolomics |
0.1 | 1 | 2011 | A high-throughput processing service for retention time alignment of complex proteomics and metabolomics LC-MS data · Bioinform. 2011 |
Bioinformatics and computational biology › metabolomics
retention time alignment |
0.1 | 1 | 2011 | A high-throughput processing service for retention time alignment of complex proteomics and metabolomics LC-MS data · Bioinform. 2011 |
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis |
0.1 | 1 | 2008 | A noise model for mass spectrometry based proteomics · Bioinform. 2008 |
Computational science and engineering
noise modeling |
0.1 | 1 | 2008 | A noise model for mass spectrometry based proteomics · Bioinform. 2008 |
Bioinformatics and computational biology › proteomics › peptide analysis
peptide detection |
0.0 | 1 | 2008 | A noise model for mass spectrometry based proteomics · Bioinform. 2008 |
Methods — techniques the papers use, named apart from their topics
parametric hypothesis testing · 0.6ledoit-wolf shrinkage · 0.6shrinkage covariance estimation · 0.4monte carlo estimation · 0.4gaussian graphical model · 0.4poisson model · 0.1multinomial model · 0.1detector dead-time correction · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Ten quick tips for building FAIR workflowsabstractResearch data is accumulating rapidly and with it the challenge of fully reproducible science. As a consequence, implementation of high-quality management of scientific data has become a global priority. The FAIR (Findable, Accesible, Interoperable and Reusable) principles provide practical guidelines for maximizing the value of research data; however, processing data using workflows-systematic executions of a series of computational tools-is equally important for good data management. The FAIR principles have recently been adapted to Research Software (FAIR4RS Principles) to promote the reproducibility and reusability of any type of research software. Here, we propose a set of 10 quick tips, drafted by experienced workflow developers that will help researchers to apply FAIR4RS principles to workflows. The tips have been arranged according to the FAIR acronym, clarifying the purpose of each tip with respect to the FAIR4RS principles. Altogether, these tips can be seen as practical guidelines for workflow developers who aim to contribute to more reproducible and sustainable computational science, aiming to positively impact the open science and FAIR community. Casper de Visser, Lennart F. Johansson, Purva Kulkarni, Hailiang Mei, Pieter B. T. Neerincx, K. Joeri van der Velde, Peter Horvatovich, Alain J. van Gool, Morris A. Swertz, Peter A. C. 't Hoen, Anna Niehues |
PLoS Comput. Biol. | 7 |
| 2022 | GeneNetTools: tests for Gaussian graphical models with shrinkageabstractMOTIVATION: Gaussian graphical models (GGMs) are network representations of random variables (as nodes) and their partial correlations (as edges). GGMs overcome the challenges of high-dimensional data analysis by using shrinkage methodologies. Therefore, they have become useful to reconstruct gene regulatory networks from gene-expression profiles. However, it is often ignored that the partial correlations are 'shrunk' and that they cannot be compared/assessed directly. Therefore, accurate (differential) network analyses need to account for the number of variables, the sample size, and also the shrinkage value, otherwise, the analysis and its biological interpretation would turn biased. To date, there are no appropriate methods to account for these factors and address these issues. RESULTS: We derive the statistical properties of the partial correlation obtained with the Ledoit-Wolf shrinkage. Our result provides a toolbox for (differential) network analyses as (i) confidence intervals, (ii) a test for zero partial correlation (null-effects) and (iii) a test to compare partial correlations. Our novel (parametric) methods account for the number of variables, the sample size and the shrinkage values. Additionally, they are computationally fast, simple to implement and require only basic statistical knowledge. Our simulations show that the novel tests perform better than DiffNetFDR-a recently published alternative-in terms of the trade-off between true and false positives. The methods are demonstrated on synthetic data and two gene-expression datasets from Escherichia coli and Mus musculus. AVAILABILITY AND IMPLEMENTATION: The R package with the methods and the R script with the analysis are available in https://github.com/V-Bernal/GeneNetTools. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Victor Bernal, Venustiano Soancatl, Jonas Bulthuis, Victor Guryev, Peter Horvatovich, Marco Grzegorczyk |
Bioinform. | 5 |
| 2021 | The 'un-shrunk' partial correlation in Gaussian graphical modelsabstractBACKGROUND: In systems biology, it is important to reconstruct regulatory networks from quantitative molecular profiles. Gaussian graphical models (GGMs) are one of the most popular methods to this end. A GGM consists of nodes (representing the transcripts, metabolites or proteins) inter-connected by edges (reflecting their partial correlations). Learning the edges from quantitative molecular profiles is statistically challenging, as there are usually fewer samples than nodes ('high dimensional problem'). Shrinkage methods address this issue by learning a regularized GGM. However, it remains open to study how the shrinkage affects the final result and its interpretation. RESULTS: We show that the shrinkage biases the partial correlation in a non-linear way. This bias does not only change the magnitudes of the partial correlations but also affects their order. Furthermore, it makes networks obtained from different experiments incomparable and hinders their biological interpretation. We propose a method, referred to as 'un-shrinking' the partial correlation, which corrects for this non-linear bias. Unlike traditional methods, which use a fixed shrinkage value, the new approach provides partial correlations that are closer to the actual (population) values and that are easier to interpret. This is demonstrated on two gene expression datasets from Escherichia coli and Mus musculus. CONCLUSIONS: GGMs are popular undirected graphical models based on partial correlations. The application of GGMs to reconstruct regulatory networks is commonly performed using shrinkage to overcome the 'high-dimensional problem'. Besides it advantages, we have identified that the shrinkage introduces a non-linear bias in the partial correlations. Ignoring this type of effects caused by the shrinkage can obscure the interpretation of the network, and impede the validation of earlier reported results. Victor Bernal, Rainer Bischoff 0003, Peter Horvatovich, Victor Guryev, Marco Grzegorczyk |
BMC Bioinform. | 3 |
| 2019 | Exact hypothesis testing for shrinkage-based Gaussian graphical modelsabstractMOTIVATION: One of the main goals in systems biology is to learn molecular regulatory networks from quantitative profile data. In particular, Gaussian graphical models (GGMs) are widely used network models in bioinformatics where variables (e.g. transcripts, metabolites or proteins) are represented by nodes, and pairs of nodes are connected with an edge according to their partial correlation. Reconstructing a GGM from data is a challenging task when the sample size is smaller than the number of variables. The main problem consists in finding the inverse of the covariance estimator which is ill-conditioned in this case. Shrinkage-based covariance estimators are a popular approach, producing an invertible 'shrunk' covariance. However, a proper significance test for the 'shrunk' partial correlation (i.e. the GGM edges) is an open challenge as a probability density including the shrinkage is unknown. In this article, we present (i) a geometric reformulation of the shrinkage-based GGM, and (ii) a probability density that naturally includes the shrinkage parameter. RESULTS: Our results show that the inference using this new 'shrunk' probability density is as accurate as Monte Carlo estimation (an unbiased non-parametric method) for any shrinkage value, while being computationally more efficient. We show on synthetic data how the novel test for significance allows an accurate control of the Type I error and outperforms the network reconstruction obtained by the widely used R package GeneNet. This is further highlighted in two gene expression datasets from stress response in Eschericha coli, and the effect of influenza infection in Mus musculus. AVAILABILITY AND IMPLEMENTATION: https://github.com/V-Bernal/GGM-Shrinkage. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Victor Bernal, Rainer Bischoff 0003, Victor Guryev, Marco Grzegorczyk, Peter Horvatovich |
Bioinform. | 5 |
| 2011 | A high-throughput processing service for retention time alignment of complex proteomics and metabolomics LC-MS dataabstractUNLABELLED: Warp2D is a novel time alignment approach, which uses the overlapping peak volume of the reference and sample peak lists to correct misleading peak shifts. Here, we present an easy-to-use web interface for high-throughput Warp2D batch processing time alignment service using the Dutch Life Science Grid, reducing processing time from days to hours. This service provides the warping function, the sample chromatogram peak list with adjusted retention times and normalized quality scores based on the sum of overlapping peak volume of all peaks. Heat maps before and after time alignment are created from the arithmetic mean of the sum of overlapping peak area rearranged with hierarchical clustering, allowing the quality control of the time alignment procedure. Taverna workflow and command line tool are provided for remote processing of local user data. AVAILABILITY: online data processing service is available at http://www.nbpp.nl/warp2d.html. Taverna workflow is available at myExperiment with title '2D Time Alignment-Webservice and Workflow' at http://www.myexperiment.org/workflows/1283.html. Command line tool is available at http://www.nbpp.nl/Warp2D_commandline.zip. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Isthiaq Ahmad, Frank Suits, Berend Hoekman, Morris A. Swertz, Heorhiy Byelas, Martijn Dijkstra, Rob W. W. Hooft, Dmitry Katsubo, Bas van Breukelen, Rainer Bischoff 0003, Peter Horvatovich |
Bioinform. | 11 |
| 2008 | A noise model for mass spectrometry based proteomicsabstractMOTIVATION: Mass spectrometry data are subjected to considerable noise. Good noise models are required for proper detection and quantification of peptides. We have characterized noise in both quadrupole time-of-flight (Q-TOF) and ion trap data, and have constructed models for the noise. RESULTS: We find that the noise in Q-TOF data from Applied Biosystems QSTAR fits well to a combination of multinomial and Poisson model with detector dead-time correction. In comparison, ion trap noise from Agilent MSD-Trap-SL is larger than the Q-TOF noise and is proportional to Poisson noise. We then demonstrate that the noise model can be used to improve deisotoping for peptide detection, by estimating appropriate cutoffs of the goodness of fit parameter at prescribed error rates. The noise models also have implications in noise reduction, retention time alignment and significance testing for biomarker discovery. Peicheng Du, Gustavo Stolovitzky, Peter Horvatovich, Rainer Bischoff 0003, Jihyeon Lim, Frank Suits |
Bioinform. | 3 |