VLDB 2026 Research / reviewers in the wild / expert
Florian Erhard
dblp:88/3908
· DBLP profile ↗
8ranked-venue papers
4as first author
2since 2021 · last 2021
0000-0002-3574-6983ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › gene regulation › promoter analysis
transcription start site identification |
0.5 | 1 | 2021 | Integrative transcription start site identification with iTiSS · Bioinform. 2021 |
Bioinformatics and computational biology
gene expression analysis |
0.5 | 2 | 2018 | Estimating pseudocounts and fold changes for digital expression measurements · Bioinform. 2018 RIP-chip enrichment analysis · Bioinform. 2013 |
Bioinformatics and computational biology › gene expression analysis
differential expression analysis |
0.3 | 1 | 2018 | Estimating pseudocounts and fold changes for digital expression measurements · Bioinform. 2018 |
Bioinformatics and computational biology
RNA sequencing |
0.3 | 1 | 2018 | Dissecting newly transcribed and old RNA using GRAND-SLAM · Bioinform. 2018 |
Bioinformatics and computational biology › gene expression analysis
microarray data analysis |
0.2 | 1 | 2013 | RIP-chip enrichment analysis · Bioinform. 2013 |
Bioinformatics and computational biology › transcriptomics
non-coding RNA analysis |
0.1 | 1 | 2010 | Classification of ncRNAs using position and size information in deep sequencing data · Bioinform. 2010 |
Bioinformatics and computational biology › sequence analysis › high-throughput sequencing data analysis
deep sequencing data analysis |
0.0 | 1 | 2010 | Classification of ncRNAs using position and size information in deep sequencing data · Bioinform. 2010 |
Methods — techniques the papers use, named apart from their topics
bayesian inference · 0.7metabolic labeling · 0.3empirical bayes · 0.3principal component analysis · 0.2gaussian mixture model · 0.2false discovery rate · 0.2pattern matrix scoring · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Machine Learning Model Update Strategies for Hard Disk Drive Failure PredictionabstractThe growing size of today’s data centers and the expectation of 24/7 availability continuously increase the complexity of hardware administration. To this end, the Self-Monitoring, Analysis, and Reporting Technology has been developed to provide insights into the health state of hard disk drives. Many approaches to predicting hard disk drive failures based on such monitoring data have been proposed in recent years. Nevertheless, most approaches consider this problem only as a static task, i.e., they train a static machine learning model on a given training set and evaluate its performance on a test set. However, due to model aging and changes in failure patterns, previously learned prediction models must be updated during runtime, which requires a time-dependent evaluation. Therefore, we present four machine learning model updating strategies, build multiple models for hard disk drive failure prediction using four machine learning algorithms, and compare the prediction quality of the different model update strategies and machine learning algorithms. Experimental results using a real-world data set of hard disk drives demonstrate the need for model update strategies, with XGBoost using the Hoeffding bound update trigger achieving the overall best prediction performance concerning prediction quality and number of updates required. Marwin Züfle, Florian Erhard, Samuel Kounev |
ICMLA | 2 |
| 2021 | Integrative transcription start site identification with iTiSSabstractSUMMARY: Many experimental approaches have been developed to identify transcription start sites (TSS) from genomic scale data. However, experiment specific biases lead to large numbers of false-positive calls. Here, we present our integrative approach iTiSS, which is an accurate and generic TSS caller for any TSS profiling experiment in eukaryotes, and substantially reduces the number of false positives by a joint analysis of several complementary datasets. AVAILABILITY AND IMPLEMENTATION: iTiSS is platform independent and implemented in Java (v1.8) and is freely available at https://www.erhard-lab.de/software and https://github.com/erhard-lab/iTiSS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Christopher Jürges, Lars Dölken, Florian Erhard |
Bioinform. | 3 |
| 2018 | Estimating pseudocounts and fold changes for digital expression measurementsabstractMotivation: Fold changes from count based high-throughput experiments such as RNA-seq suffer from a zero-frequency problem. To circumvent division by zero, so-called pseudocounts are added to make all observed counts strictly positive. The magnitude of pseudocounts for digital expression measurements and on which stage of the analysis they are introduced remained an arbitrary choice. Moreover, in the strict sense, fold changes are not quantities that can be computed. Instead, due to the stochasticity involved in the experiments, they must be estimated by statistical inference. Results: Here, we build on a statistical framework for fold changes, where pseudocounts correspond to the parameters of the prior distribution used for Bayesian inference of the fold change. We show that arbitrary and widely used choices for applying pseudocounts can lead to biased results. As a statistical rigorous alternative, we propose and test an empirical Bayes procedure to choose appropriate pseudocounts. Moreover, we introduce the novel estimator Ψ LFC for fold changes showing favorable properties with small counts and smaller deviations from the truth in simulations and real data compared to existing methods. Our results have direct implications for entities with few reads in sequencing experiments, and indirectly also affect results for entities with many reads. Availability and implementation: Ψ LFC is available as an R package under https://github.com/erhard-lab/lfc (Apache 2.0 license); R scripts to generate all figures are available at zenodo (doi: 10.5281/zenodo.1163029). Supplementary information: Supplementary data are available at Bioinformatics online. Florian Erhard |
Bioinform. | 1 |
| 2018 | Dissecting newly transcribed and old RNA using GRAND-SLAMabstractSummary: Global quantification of total RNA is used to investigate steady state levels of gene expression. However, being able to differentiate pre-existing RNA (that has been synthesized prior to a defined point in time) and newly transcribed RNA can provide invaluable information e.g. to estimate RNA half-lives or identify fast and complex regulatory processes. Recently, new techniques based on metabolic labeling and RNA-seq have emerged that allow to quantify new and old RNA: Nucleoside analogs are incorporated into newly transcribed RNA and are made detectable as point mutations in mapped reads. However, relatively infrequent incorporation events and significant sequencing error rates make the differentiation between old and new RNA a highly challenging task. We developed a statistical approach termed GRAND-SLAM that, for the first time, allows to estimate the proportion of old and new RNA in such an experiment. Uncertainty in the estimates is quantified in a Bayesian framework. Simulation experiments show our approach to be unbiased and highly accurate. Furthermore, we analyze how uncertainty in the proportion translates into uncertainty in estimating RNA half-lives and give guidelines for planning experiments. Finally, we demonstrate that our estimates of RNA half-lives compare favorably to other experimental approaches and that biological processes affecting RNA half-lives can be investigated with greater power than offered by any other method. GRAND-SLAM is freely available for non-commercial use at http://software.erhard-lab.de; R scripts to generate all figures are available at zenodo (doi: 10.5281/zenodo.1162340). Christopher Jürges, Lars Dölken, Florian Erhard |
Bioinform. | 3 |
| 2013 | RIP-chip enrichment analysisabstractMOTIVATION: RIP-chip is a high-throughput method to identify mRNAs that are targeted by RNA-binding proteins. The protein of interest is immunoprecipitated, and the identity and relative amount of mRNA associated with it is measured on microarrays. Even if a variety of methods is available to analyse microarray data, e.g. to detect differentially regulated genes, the additional experimental steps in RIP-chip require specialized methods. Here, we focus on two aspects of RIP-chip data: First, the efficiency of the immunoprecipitation step performed in the RIP-chip protocol varies in between different experiments introducing bias not existing in standard microarray experiments. This requires an additional normalization step to compare different samples and even technical replicates. Second, in contrast to standard differential gene expression experiments, the distribution of measurements is not normal. We exploit this fact to define a set of biologically relevant genes in a statistically meaningful way. RESULTS: Here, we propose two methods to analyse RIP-chip data: We model the measurement distribution as a gaussian mixture distribution, which allows us to compute false discovery rates (FDRs) for any cut-off. Thus, cut-offs can be chosen for any desired FDR. Furthermore, we use principal component analysis to determine the normalization factors necessary to remove immunoprecipitation bias. Both methods are evaluated on a large RIP-chip dataset measuring targets of Ago2, the major component of the microRNA guided RNA-induced silencing complex (RISC). Using published HITS-CLIP experiments performed with the same cell line as used for RIP-chip, we show that the mixture modelling approach is a necessary step to remove background, which computed FDRs are valid, and that the additional normalization is a necessary step to make experiments comparable. AVAILABILITY: An R implementation of REA is available on the project website (http://www.bio.ifi.lmu.de/REA) and as supplementary data file. Florian Erhard, Lars Dölken, Ralf Zimmer |
Bioinform. | 1 |
| 2010 | Classification of ncRNAs using position and size information in deep sequencing dataabstractMOTIVATION: Small non-coding RNAs (ncRNAs) play important roles in various cellular functions in all clades of life. With next-generation sequencing techniques, it has become possible to study ncRNAs in a high-throughput manner and by using specialized algorithms ncRNA classes such as miRNAs can be detected in deep sequencing data. Typically, such methods are targeted to a certain class of ncRNA. Many methods rely on RNA secondary structure prediction, which is not always accurate and not all ncRNA classes are characterized by a common secondary structure. Unbiased classification methods for ncRNAs could be important to improve accuracy and to detect new ncRNA classes in sequencing data. RESULTS: Here, we present a scoring system called ALPS (alignment of pattern matrices score) that only uses primary information from a deep sequencing experiment, i.e. the relative positions and lengths of reads, to classify ncRNAs. ALPS makes no further assumptions, e.g. about common structural properties in the ncRNA class and is nevertheless able to identify ncRNA classes with high accuracy. Since ALPS is not designed to recognize a certain class of ncRNA, it can be used to detect novel ncRNA classes, as long as these unknown ncRNAs have a characteristic pattern of deep sequencing read lengths and positions. We evaluate our scoring system on publicly available deep sequencing data and show that it is able to classify known ncRNAs with high sensitivity and specificity. AVAILABILITY: Calculated pattern matrices of the datasets hESC and EB are available at the project web site http://www.bio.ifi.lmu.de/ALPS. An implementation of the described method is available upon request from the authors. Florian Erhard, Ralf Zimmer |
Bioinform. | 1 |
| 2008 | FERN - a Java framework for stochastic simulation and evaluation of reaction networksabstractBACKGROUND: Stochastic simulation can be used to illustrate the development of biological systems over time and the stochastic nature of these processes. Currently available programs for stochastic simulation, however, are limited in that they either a) do not provide the most efficient simulation algorithms and are difficult to extend, b) cannot be easily integrated into other applications or c) do not allow to monitor and intervene during the simulation process in an easy and intuitive way. Thus, in order to use stochastic simulation in innovative high-level modeling and analysis approaches more flexible tools are necessary. RESULTS: In this article, we present FERN (Framework for Evaluation of Reaction Networks), a Java framework for the efficient simulation of chemical reaction networks. FERN is subdivided into three layers for network representation, simulation and visualization of the simulation results each of which can be easily extended. It provides efficient and accurate state-of-the-art stochastic simulation algorithms for well-mixed chemical systems and a powerful observer system, which makes it possible to track and control the simulation progress on every level. To illustrate how FERN can be easily integrated into other systems biology applications, plugins to Cytoscape and CellDesigner are included. These plugins make it possible to run simulations and to observe the simulation progress in a reaction network in real-time from within the Cytoscape or CellDesigner environment. CONCLUSION: FERN addresses shortcomings of currently available stochastic simulation programs in several ways. First, it provides a broad range of efficient and accurate algorithms both for exact and approximate stochastic simulation and a simple interface for extending to new algorithms. FERN's implementations are considerably faster than the C implementations of gillespie2 or the Java implementations of ISBJava. Second, it can be used in a straightforward way both as a stand-alone program and within new systems biology applications. Finally, complex scenarios requiring intervention during the simulation progress can be modelled easily with FERN. Florian Erhard, Caroline C. Friedel, Ralf Zimmer |
BMC Bioinform. | 1 |
| 2000 | GSFS - A New Group-Aware Cryptographic File System
Claudia Eckert 0001, Florian Erhard, Johannes Geiger |
SEC | 2 |