Sonja Zehetmayer

dblp:30/592 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
1since 2021 · last 2022
0000-0001-6863-7997ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 4 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
statistical genetics
0.222010
Post hoc power estimation in large-scale multiple testing problems · Bioinform. 2010
Two-stage designs for experiments with a large number of hypotheses · Bioinform. 2005
Bioinformatics and computational biology
gene expression analysis
0.112010
Post hoc power estimation in large-scale multiple testing problems · Bioinform. 2010
Bioinformatics and computational biology › gene expression analysis
microarray data analysis
0.112010
Post hoc power estimation in large-scale multiple testing problems · Bioinform. 2010
Bioinformatics and computational biology
power analysis
0.112010
Post hoc power estimation in large-scale multiple testing problems · Bioinform. 2010

Methods — techniques the papers use, named apart from their topics

simulation · 0.2empirical distribution estimation · 0.1sequential testing · 0.1
YearPublicationVenuePosition
2022 Impact of adaptive filtering on power and false discovery rate in RNA-seq experiments
abstract
BACKGROUND: In RNA-sequencing studies a large number of hypothesis tests are performed to compare the differential expression of genes between several conditions. Filtering has been proposed to remove candidate genes with a low expression level which may not be relevant and have little or no chance of showing a difference between conditions. This step may reduce the multiple testing burden and increase power. RESULTS: We show in a simulation study that filtering can lead to some increase in power for RNA-sequencing data, too aggressive filtering, however, can lead to a decline. No uniformly optimal filter in terms of power exists. Depending on the scenario different filters may be optimal. We propose an adaptive filtering strategy which selects one of several filters to maximise the number of rejections. No additional adjustment for multiplicity has to be included, but a rule has to be considered if the number of rejections is too small. CONCLUSIONS: For a large range of simulation scenarios, the adaptive filter maximises the power while the simulated False Discovery Rate is bounded by the pre-defined significance level. Using the adaptive filter, it is not necessary to pre-specify a single individual filtering method optimised for a specific scenario.
Sonja Zehetmayer, Martin Posch, Alexandra Graf
BMC Bioinform.1
2012 False discovery rate control in two-stage designs
abstract
BACKGROUND: For gene expression or gene association studies with a large number of hypotheses the number of measurements per marker in a conventional single-stage design is often low due to limited resources. Two-stage designs have been proposed where in a first stage promising hypotheses are identified and further investigated in the second stage with larger sample sizes. For two types of two-stage designs proposed in the literature we derive multiple testing procedures controlling the False Discovery Rate (FDR) demonstrating FDR control by simulations: designs where a fixed number of top-ranked hypotheses are selected and designs where the selection in the interim analysis is based on an FDR threshold. In contrast to earlier approaches which use only the second-stage data in the hypothesis tests (pilot approach), the proposed testing procedures are based on the pooled data from both stages (integrated approach). RESULTS: For both selection rules the multiple testing procedures control the FDR in the considered simulation scenarios. This holds for the case of independent observations across hypotheses as well as for certain correlation structures. Additionally, we show that in scenarios with small effect sizes the testing procedures based on the pooled data from both stages can give a considerable improvement in power compared to tests based on the second-stage data only. CONCLUSION: The proposed hypothesis tests provide a tool for FDR control for the considered two-stage designs. Comparing the integrated approaches for both selection rules with the corresponding pilot approaches showed an advantage of the integrated approach in many simulation scenarios.
Sonja Zehetmayer, Martin Posch
BMC Bioinform.1
2010 Post hoc power estimation in large-scale multiple testing problems
abstract
BACKGROUND: The statistical power or multiple Type II error rate in large-scale multiple testing problems as, for example, in gene expression microarray experiments, depends on typically unknown parameters and is therefore difficult to assess a priori. However, it has been suggested to estimate the multiple Type II error rate post hoc, based on the observed data. METHODS: We consider a class of post hoc estimators that are functions of the estimated proportion of true null hypotheses among all hypotheses. Numerous estimators for this proportion have been proposed and we investigate the statistical properties of the derived multiple Type II error rate estimators in an extensive simulation study. RESULTS: The performance of the estimators in terms of the mean squared error depends sensitively on the distributional scenario. Estimators based on empirical distributions of the null hypotheses are superior in the presence of strongly correlated test statistics. AVAILABILITY: R-code to compute all considered estimators based on P-values and supplementary material is available on the authors web page http://statistics.msi.meduniwien.ac.at/index.php?page=pageszfnr.
Sonja Zehetmayer, Martin Posch
Bioinform.1
2005 Two-stage designs for experiments with a large number of hypotheses
abstract
Motivation: When a large number of hypotheses are investigated the false discovery rate (FDR) is commonly applied in gene expression analysis or gene association studies. Conventional single-stage designs may lack power due to low sample sizes for the individual hypotheses. We propose two-stage designs where the first stage is used to screen the ‘promising’ hypotheses which are further investigated at the second stage with an increased sample size. A multiple test procedure based on sequential individual P-values is proposed to control the FDR for the case of independent normal distributions with known variance. Results: The power of optimal two-stage designs is impressively larger than the power of the corresponding single–stage design with equal costs. Extensions to the case of unknown variances and correlated test statistics are investigated by simulations. Moreover, it is shown that the simple multiple test procedure using first stage data for screening purposes and deriving the test decisions only from second stage data is a very powerful option. Availability: An R-program is available at http://www.meduniwien.ac.at/medstat/research/fdr/application.R Contact: [email protected] Supplementary information: Supplementary data for this paper is available at Bioinformatics online.
Sonja Zehetmayer, Peter Bauer, Martin Posch
Bioinform.1