Anne-Laure Boulesteix

dblp:58/635 · DBLP profile ↗
← Back
33ranked-venue papers
14as first author
5since 2021 · last 2024
0000-0002-2729-0947ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 30 · 14 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021
YearPublicationVenuePosition
2024 Position: Why We Must Rethink Empirical Research in Machine Learning
abstract
We warn against a common but incomplete understanding of empirical research in machine learning that leads to non-replicable results, makes findings unreliable, and threatens to undermine progress in the field. To overcome this alarming situation, we call for more awareness of the plurality of ways of gaining knowledge experimentally but also of some epistemic limitations. In particular, we argue most current empirical machine learning research is fashioned as confirmatory research while it should rather be considered exploratory.
Moritz Herrmann, F. Julian D. Lange, Katharina Eggensperger, Giuseppe Casalicchio, Marcel Wever, Matthias Feurer 0001, David Rügamer, Eyke Hüllermeier, Anne-Laure Boulesteix, Bernd Bischl
ICML9
2024 Raising awareness of uncertain choices in empirical data analysis: A teaching concept toward replicable research practices
abstract
Throughout their education and when reading the scientific literature, students may get the impression that there is a unique and correct analysis strategy for every data analysis task and that this analysis strategy will always yield a significant and noteworthy result. This expectation conflicts with a growing realization that there is a multiplicity of possible analysis strategies in empirical research, which will lead to overoptimism and nonreplicable research findings if it is combined with result-dependent selective reporting. Here, we argue that students are often ill-equipped for real-world data analysis tasks and unprepared for the dangers of selectively reporting the most promising results. We present a seminar course intended for advanced undergraduates and beginning graduate students of data analysis fields such as statistics, data science, or bioinformatics that aims to increase the awareness of uncertain choices in the analysis of empirical data and present ways to deal with these choices through theoretical modules and practical hands-on sessions.
Maximilian M. Mandl, Sabine Hoffmann, Sebastian Bieringer, Anna E. Jacob, Marie Kraft, Simon Lemster, Anne-Laure Boulesteix
PLoS Comput. Biol.7
2023 Over-optimism in unsupervised microbiome analysis: Insights from network learning and clustering
abstract
In recent years, unsupervised analysis of microbiome data, such as microbial network analysis and clustering, has increased in popularity. Many new statistical and computational methods have been proposed for these tasks. This multiplicity of analysis strategies poses a challenge for researchers, who are often unsure which method(s) to use and might be tempted to try different methods on their dataset to look for the "best" ones. However, if only the best results are selectively reported, this may cause over-optimism: the "best" method is overly fitted to the specific dataset, and the results might be non-replicable on validation data. Such effects will ultimately hinder research progress. Yet so far, these topics have been given little attention in the context of unsupervised microbiome analysis. In our illustrative study, we aim to quantify over-optimism effects in this context. We model the approach of a hypothetical microbiome researcher who undertakes four unsupervised research tasks: clustering of bacterial genera, hub detection in microbial networks, differential microbial network analysis, and clustering of samples. While these tasks are unsupervised, the researcher might still have certain expectations as to what constitutes interesting results. We translate these expectations into concrete evaluation criteria that the hypothetical researcher might want to optimize. We then randomly split an exemplary dataset from the American Gut Project into discovery and validation sets multiple times. For each research task, multiple method combinations (e.g., methods for data normalization, network generation, and/or clustering) are tried on the discovery data, and the combination that yields the best result according to the evaluation criterion is chosen. While the hypothetical researcher might only report this result, we also apply the "best" method combination to the validation dataset. The results are then compared between discovery and validation data. In all four research tasks, there are notable over-optimism effects; the results on the validation data set are worse compared to the discovery data, averaged over multiple random splits into discovery/validation data. Our study thus highlights the importance of validation and replication in microbiome analysis to obtain reliable results and demonstrates that the issue of over-optimism goes beyond the context of statistical testing and fishing for significance.
Theresa Ullmann, Stefanie Peschel, Philipp F. M. Baumann, Christian L. Müller, Anne-Laure Boulesteix
PLoS Comput. Biol.5
2021 Large-scale benchmark study of survival prediction methods using multi-omics data
abstract
Multi-omics data, that is, datasets containing different types of high-dimensional molecular variables, are increasingly often generated for the investigation of various diseases. Nevertheless, questions remain regarding the usefulness of multi-omics data for the prediction of disease outcomes such as survival time. It is also unclear which methods are most appropriate to derive such prediction models. We aim to give some answers to these questions through a large-scale benchmark study using real data. Different prediction methods from machine learning and statistics were applied on 18 multi-omics cancer datasets (35 to 1000 observations, up to 100 000 variables) from the database 'The Cancer Genome Atlas' (TCGA). The considered outcome was the (censored) survival time. Eleven methods based on boosting, penalized regression and random forest were compared, comprising both methods that do and that do not take the group structure of the omics variables into account. The Kaplan-Meier estimate and a Cox model using only clinical variables were used as reference methods. The methods were compared using several repetitions of 5-fold cross-validation. Uno's C-index and the integrated Brier score served as performance metrics. The results indicate that methods taking into account the multi-omics structure have a slightly better prediction performance. Taking this structure into account can protect the predictive information in low-dimensional groups-especially clinical variables-from not being exploited during prediction. Moreover, only the block forest method outperformed the Cox model on average, and only slightly. This indicates, as a by-product of our study, that in the considered TCGA studies the utility of multi-omics data for prediction purposes was limited. Contact:[email protected], +49 89 2180 3198 Supplementary information: Supplementary data are available at Briefings in Bioinformatics online. All analyses are reproducible using R code freely available on Github.
Moritz Herrmann, Philipp Probst, Roman Hornung, Vindi Jurinovic, Anne-Laure Boulesteix
Briefings Bioinform.5
2021 NetCoMi: network construction and comparison for microbiome data in R
abstract
MOTIVATION: Estimating microbial association networks from high-throughput sequencing data is a common exploratory data analysis approach aiming at understanding the complex interplay of microbial communities in their natural habitat. Statistical network estimation workflows comprise several analysis steps, including methods for zero handling, data normalization and computing microbial associations. Since microbial interactions are likely to change between conditions, e.g. between healthy individuals and patients, identifying network differences between groups is often an integral secondary analysis step. Thus far, however, no unifying computational tool is available that facilitates the whole analysis workflow of constructing, analysing and comparing microbial association networks from high-throughput sequencing data. RESULTS: Here, we introduce NetCoMi (Network Construction and comparison for Microbiome data), an R package that integrates existing methods for each analysis step in a single reproducible computational workflow. The package offers functionality for constructing and analysing single microbial association networks as well as quantifying network differences. This enables insights into whether single taxa, groups of taxa or the overall network structure change between groups. NetCoMi also contains functionality for constructing differential networks, thus allowing to assess whether single pairs of taxa are differentially associated between two groups. Furthermore, NetCoMi facilitates the construction and analysis of dissimilarity networks of microbiome samples, enabling a high-level graphical summary of the heterogeneity of an entire microbiome sample collection. We illustrate NetCoMi's wide applicability using data sets from the GABRIELA study to compare microbial associations in settled dust from children's rooms between samples from two study centers (Ulm and Munich). AVAILABILITY: R scripts used for producing the examples shown in this manuscript are provided as supplementary data. The NetCoMi package, together with a tutorial, is available at https://github.com/stefpeschel/NetCoMi. CONTACT: Tel:+49 89 3187 43258; [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Briefings in Bioinformatics online.
Stefanie Peschel, Christian L. Müller, Erika von Mutius, Anne-Laure Boulesteix, Martin Depner
Briefings Bioinform.4
2020 Combining clinical and molecular data in regression prediction models: insights from a simulation study
abstract
Data integration, i.e. the use of different sources of information for data analysis, is becoming one of the most important topics in modern statistics. Especially in, but not limited to, biomedical applications, a relevant issue is the combination of low-dimensional (e.g. clinical data) and high-dimensional (e.g. molecular data such as gene expressions) data sources in a prediction model. Not only the different characteristics of the data, but also the complex correlation structure within and between the two data sources, pose challenging issues. In this paper, we investigate these issues via simulations, providing some useful insight into strategies to combine low- and high-dimensional data in a regression prediction model. In particular, we focus on the effect of the correlation structure on the results, while accounting for the influence of our specific choices in the design of the simulation study.
Riccardo De Bin, Anne-Laure Boulesteix, Axel Benner, Natalia Becker, Willi Sauerbrei
Briefings Bioinform.2
2019 Tunability: Importance of Hyperparameters of Machine Learning Algorithms
abstract
Modern supervised machine learning algorithms involve hyperparameters that have to be set before running them. Options for setting hyperparameters are default values from the software package, manual configuration by the user or configuring them for optimal predictive performance by a tuning procedure. The goal of this paper is two-fold. Firstly, we formalize the problem of tuning from a statistical point of view, define data-based defaults and suggest general measures quantifying the tunability of hyperparameters of algorithms. Secondly, we conduct a large-scale benchmarking study based on 38 datasets from the OpenML platform and six common machine learning algorithms. We apply our measures to assess the tunability of their parameters. Our results yield default values for hyperparameters and enable users to decide whether it is worth conducting a possibly time consuming tuning strategy, to focus on the most important hyperparameters and to choose adequate hyperparameter spaces for tuning.
Philipp Probst, Anne-Laure Boulesteix, Bernd Bischl
J. Mach. Learn. Res.2
2018 Random forest versus logistic regression: a large-scale benchmark experiment
abstract
BACKGROUND AND GOAL: The Random Forest (RF) algorithm for regression and classification has considerably gained popularity since its introduction in 2001. Meanwhile, it has grown to a standard classification approach competing with logistic regression in many innovation-friendly scientific fields. RESULTS: In this context, we present a large scale benchmarking experiment based on 243 real datasets comparing the prediction performance of the original version of RF with default parameters and LR as binary classification tools. Most importantly, the design of our benchmark experiment is inspired from clinical trial methodology, thus avoiding common pitfalls and major sources of biases. CONCLUSION: RF performed better than LR according to the considered accuracy measured in approximately 69% of the datasets. The mean difference between RF and LR was 0.029 (95%-CI =[0.022,0.038]) for the accuracy, 0.041 (95%-CI =[0.031,0.053]) for the Area Under the Curve, and - 0.027 (95%-CI =[-0.034,-0.021]) for the Brier score, all measures thus suggesting a significantly better performance of RF. As a side-result of our benchmarking experiment, we observed that the results were noticeably dependent on the inclusion criteria used to select the example datasets, thus emphasizing the importance of clear statements regarding this dataset selection process. We also stress that neutral studies similar to ours, based on a high number of datasets and carefully designed, will be necessary in the future to evaluate further variants, implementations or parameters of random forests which may yield improved accuracy compared to the original version with default values.
Raphaël Couronné, Philipp Probst, Anne-Laure Boulesteix
BMC Bioinform.3
2018 Priority-Lasso: a simple hierarchical approach to the prediction of clinical outcome using multi-omics data
abstract
BACKGROUND: The inclusion of high-dimensional omics data in prediction models has become a well-studied topic in the last decades. Although most of these methods do not account for possibly different types of variables in the set of covariates available in the same dataset, there are many such scenarios where the variables can be structured in blocks of different types, e.g., clinical, transcriptomic, and methylation data. To date, there exist a few computationally intensive approaches that make use of block structures of this kind. RESULTS: In this paper we present priority-Lasso, an intuitive and practical analysis strategy for building prediction models based on Lasso that takes such block structures into account. It requires the definition of a priority order of blocks of data. Lasso models are calculated successively for every block and the fitted values of every step are included as an offset in the fit of the next step. We apply priority-Lasso in different settings on an acute myeloid leukemia (AML) dataset consisting of clinical variables, cytogenetics, gene mutations and expression variables, and compare its performance on an independent validation dataset to the performance of standard Lasso models. CONCLUSION: The results show that priority-Lasso is able to keep pace with Lasso in terms of prediction accuracy. Variables of blocks with higher priorities are favored over variables of blocks with lower priority, which results in easily usable and transportable models for clinical practice.
Simon Klau, Vindi Jurinovic, Roman Hornung, Tobias Herold, Anne-Laure Boulesteix
BMC Bioinform.5
2017 Improving cross-study prediction through addon batch effect adjustment or addon normalization
abstract
Motivation: To date most medical tests derived by applying classification methods to high-dimensional molecular data are hardly used in clinical practice. This is partly because the prediction error resulting when applying them to external data is usually much higher than internal error as evaluated through within-study validation procedures. We suggest the use of addon normalization and addon batch effect removal techniques in this context to reduce systematic differences between external data and the original dataset with the aim to improve prediction performance. Results: We evaluate the impact of addon normalization and seven batch effect removal methods on cross-study prediction performance for several common classifiers using a large collection of microarray gene expression datasets, showing that some of these techniques reduce prediction error. Availability and Implementation: All investigated addon methods are implemented in our R package bapred. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Roman Hornung, David Causeur, Christoph Bernau, Anne-Laure Boulesteix
Bioinform.4
2017 To Tune or Not to Tune the Number of Trees in Random Forest
Philipp Probst, Anne-Laure Boulesteix
J. Mach. Learn. Res.2
2016 Combining location-and-scale batch effect adjustment with data cleaning by latent factor adjustment
abstract
BACKGROUND: In the context of high-throughput molecular data analysis it is common that the observations included in a dataset form distinct groups; for example, measured at different times, under different conditions or even in different labs. These groups are generally denoted as batches. Systematic differences between these batches not attributable to the biological signal of interest are denoted as batch effects. If ignored when conducting analyses on the combined data, batch effects can lead to distortions in the results. In this paper we present FAbatch, a general, model-based method for correcting for such batch effects in the case of an analysis involving a binary target variable. It is a combination of two commonly used approaches: location-and-scale adjustment and data cleaning by adjustment for distortions due to latent factors. We compare FAbatch extensively to the most commonly applied competitors on the basis of several performance metrics. FAbatch can also be used in the context of prediction modelling to eliminate batch effects from new test data. This important application is illustrated using real and simulated data. We implemented FAbatch and various other functionalities in the R package bapred available online from CRAN. RESULTS: FAbatch is seen to be competitive in many cases and above average in others. In our analyses, the only cases where it failed to adequately preserve the biological signal were when there were extremely outlying batches and when the batch effects were very weak compared to the biological signal. CONCLUSIONS: As seen in this paper batch effect structures found in real datasets are diverse. Current batch effect adjustment methods are often either too simplistic or make restrictive assumptions, which can be violated in real datasets. Due to the generality of its underlying model and its ability to perform well FAbatch represents a reliable tool for batch effect adjustment for most situations found in practice.
Roman Hornung, Anne-Laure Boulesteix, David Causeur
BMC Bioinform.2
2015 Letter to the Editor: On Reviews and Papers on New Methods
abstract
Anne-Laure Boulesteix is a professor of computational molecular medicine at the University of Munich, Germany. She is working on statistical methods for the analysis of high-dimensional omics data with a particular focus on epistemological issues. Briefings in Bioinformatics is a high-impact journal, much appreciated by a wide audience ranging from (computational) biologists to statisticians and computer scientists. Its focus on reviews and tutorials has certainly much contributed to its popularity in a broad field characterized by a confusing profusion of new methods published in an increasing number of journals. The scope of the journal encompasses diverse biological topics as stated on the journal Web site and the state-of-the-art methods from statistics, mathematics and computer sciences necessary to address them. In this increasingly complex field reviews and tutorials are of extreme importance for scientists to stay in the loop also outside their specific area of expertise and for researchers and students starting to work on a new topic. As a high-impact journal with a broad audience, Briefings in Bioinformatics is attractive to authors. In my view, a problem is that some of them try—and sometimes manage—to have their work published in the journal even when it is an article introducing a new method rather than a review. More or less discreet tricks may be necessary to make the paper appear to be a review, for instance, avoiding words like ‘new’ and ‘novel’ (which are otherwise widely used), including an unusually long literature overview in a preliminary section, and presenting the illustration section as a ‘comparison study’, to cite only a few examples. We claim that an article suggesting a new method is however never a proper review and should not be published in a journal specialized in such. For authors, it is advantageous to have their new method published in a journal devoted to reviews because readers may feel that the promotion of the method is the result of an exhaustive comparison or consensus, and that the method is part of foundational knowledge everyone should have as part of their toolkit. This misunderstanding may in practice boost the popularity of the new method. But we claim that this is unfair to readers and detrimental to scientific progress. Unfortunately, such misunderstandings are likely, since briefing papers target semi-experts who may not be able to recognize whether a method is new. This debate raises the question of what makes a good review. Among other qualities, a good review should ideally be reasonably neutral, i.e. not strongly reflect the personal—subjective—preferences of the authors. This is not completely achievable, since authors are humans with their own experiences, but neutrality should be considered an ultimate goal. An important aspect of neutrality is that within the defined area, the weight of each reviewed bioinformatics approach should depend on aspects such as, for example, its performance as assessed in literature (or in a comparison study presented as part of the review), theoretical soundness, generalizability, frequency of use in the literature, historical considerations, availability of user-friendly implementations and so on, but not on the interests or publication records of the authors, see for instance Rule 9 of the ‘Ten Simple Rules for Writing a Literature Review’ [1]. A review which is not neutral may in some cases give readers a distorted picture of the state of the art, for instance, by giving them the impression that one of the approaches is ‘standard’ or more widely used than others, without this being the case. The criteria for including/excluding approaches in/from the review and determining their respective weights ideally should be transparent. Of course, the goal of an ideal neutral review is not attainable in practice. Authors will always have favorite methods based on subjective criteria, they will always be better able to describe methods they are familiar with than those they have never investigated in their own research (which is especially important for tutorial articles) and the choice of the specific topic of the review might itself be the result of a subjective decision. All these aspects will affect the weight of the approaches and the overall message propagated by their review, and possibly favor methods previously developed by the review’s authors. This is unavoidable, and by the way a reason why it may be interesting to read several reviews from different authors on the same or similar topics and why reviews written by teams are particularly valuable. An attainable minimum requirement for reviews, however, is that they should not suggest a new method, even—or especially—if the novel character of the considered method is masked by crude tricks. A paper suggesting a new method should include an overview of existing methods, but it cannot be considered a review because its goal—convincing readers of the usefulness of the new method—is inherently different from the goal of a review. Such a paper cannot be neutral in the sense described above. Further, authors of such papers can in all likelihood not devote as much time and attention to the overview of existing methods as required of a good review because they are focusing on the development of the new method [2]. Note, however, that integrative reviews might in some cases provide new insights and lead to novel hypotheses and priorities for further research/development. This is not in contradiction with our argument that a paper cannot both promote a new method and provide a neutral review. I personally believe that Briefings in Bioinformatics should remain faithful to its original aim of publishing only reviews or related. Although there are many journals requiring innovation as one of the main criteria for publication, there are not many explicitly welcoming reviews: in this regard Briefings in Bioinformatics remains distinct, and retains greater value. Finally, as stated, papers which appear to be reviews but actually promote a new method may mislead the readers. That is why I think that, if the editorial board decides to take the opposite strategy and to officially allow papers on new methods, these papers should be explicitly labeled as such and not disguised as reviews. The strength and originality of Briefings in Bioinformatics is that it is devoted to reviews. Good reviews are extremely important in scientific research. A good review should be reasonably ‘neutral’. Such neutrality cannot be achieved within a paper introducing a new method. Papers on new methods published in journals devoted to reviews are misleading for readers and give them a distorted picture of the state of the art. The author thanks Rory Wilson and the reviewers for helpful comments.
Anne-Laure Boulesteix
Briefings Bioinform.1
2015 Letter to the Editor: On the term 'interaction' and related phrases in the literature on Random Forests
abstract
In an interesting and quite exhaustive review on Random Forests (RF) methodology in bioinformatics Touw et al. address--among other topics--the problem of the detection of interactions between variables based on RF methodology. We feel that some important statistical concepts, such as 'interaction', 'conditional dependence' or 'correlation', are sometimes employed inconsistently in the bioinformatics literature in general and in the literature on RF in particular. In this letter to the Editor, we aim to clarify some of the central statistical concepts and point out some confusing interpretations concerning RF given by Touw et al. and other authors.
Anne-Laure Boulesteix, Silke Janitza, Alexander Hapfelmeier, Kristel Van Steen, Carolin Strobl
Briefings Bioinform.1
2015 Ten Simple Rules for Reducing Overoptimistic Reporting in Methodological Computational Research
abstract
In most scientific fields, and in biomedical research in particular, there have long been many discussions on how to improve research practices and methods. The trend has increased in recent years, as illustrated by the series on “reducing waste,” published in The Lancet in January 2014 [1], or by the recent essay by John Ioannidis on how to make published results more true [2], which echoes his earlier provocative paper entitled “Why most published research findings are false” [3]. One of the important aspects underlying these discussions is that biomedical literature is most often overoptimistic with respect to, for example, the superiority of a new therapy or the strength of association between a risk factor and an outcome. Published results appear more significant, more spectacular, or sometimes more intuitive—in a word, more “satisfactory”—to authors and readers than they actually would if they reflected the truth. Causes of this problem are diverse, numerous, and interrelated. The effects of “fishing for significance” strategies or selective/incomplete reporting are exacerbated by design issues (e.g., small sample sizes, many investigated features) [3] or publication bias [4], to cite only a few of the factors at work. Research and guidelines on how to reduce overoptimistic reporting in the context of computational research, including computational biology as an important special case, however, are surprisingly scarce. Many methodological articles published in computational literature report the (vastly) superior performance of new methods [5], too often in general terms and—directly or indirectly—implying that the presented positive results are generalizable to other settings. Such overoptimistic reporting confuses readers, makes literature less credible and more difficult to interpret, and might even ultimately lead to a waste of resources in some cases. Here I take advantage of the popular “ten-simple-rules” format [6] to address the problem of overoptimistic reporting in methodological computational biology research, that is papers—termed “methodological papers” here—devoted primarily to the development and testing of new computational methods (intended to be used by other researchers on other data in the future) rather than to the biological question itself or the specific dataset at hand.
Anne-Laure Boulesteix
PLoS Comput. Biol.1
2014 Cross-study validation for the assessment of prediction algorithms
abstract
MOTIVATION: Numerous competing algorithms for prediction in high-dimensional settings have been developed in the statistical and machine-learning literature. Learning algorithms and the prediction models they generate are typically evaluated on the basis of cross-validation error estimates in a few exemplary datasets. However, in most applications, the ultimate goal of prediction modeling is to provide accurate predictions for independent samples obtained in different settings. Cross-validation within exemplary datasets may not adequately reflect performance in the broader application context. METHODS: We develop and implement a systematic approach to 'cross-study validation', to replace or supplement conventional cross-validation when evaluating high-dimensional prediction models in independent datasets. We illustrate it via simulations and in a collection of eight estrogen-receptor positive breast cancer microarray gene-expression datasets, where the objective is predicting distant metastasis-free survival (DMFS). We computed the C-index for all pairwise combinations of training and validation datasets. We evaluate several alternatives for summarizing the pairwise validation statistics, and compare these to conventional cross-validation. RESULTS: Our data-driven simulations and our application to survival prediction with eight breast cancer microarray datasets, suggest that standard cross-validation produces inflated discrimination accuracy for all algorithms considered, when compared to cross-study validation. Furthermore, the ranking of learning algorithms differs, suggesting that algorithms performing best in cross-validation may be suboptimal when evaluated through independent validation. AVAILABILITY: The survHD: Survival in High Dimensions package (http://www.bitbucket.org/lwaldron/survhd) will be made available through Bioconductor.
Christoph Bernau, Markus Riester, Anne-Laure Boulesteix, Giovanni Parmigiani, Curtis Huttenhower, Levi Waldron, Lorenzo Trippa
Bioinform.3
2013 On representative and illustrative comparisons with real data in bioinformatics: response to the letter to the editor by Smith et al
abstract
Smith et al. (2013) recently published an interesting letter to the editors of Bioinformatics outlining the importance of validation of new proposed algorithms through a thorough comparison with existing approaches. We completely agree with the necessity of more comparison studies in general to help end users to make an informed choice based on objective criteria. We also agree that ‘The practical result [of the lack of comparison studies] is that practitioners stop short of exhaustively evaluating all the possible options and choose based on some other criteria (e.g. popularity, ease of use or familiarity)’, and that this may harm both algorithm developers and users—as also claimed in one of our previous publications on this topic (Boulesteix et al., 2013a). In a few words, we fully comply with the claim of the authors that everybody (users and algorithm makers) suffers from the trend observed in bioinformatics toward the publication of many new algorithms without proper comparison between these algorithms. We also agree with Smith et al. (2013) that articles suggesting new algorithms should always include a comparison with existing methods. Here we guess that the authors implicitly mean comparisons based on real datasets—as opposed to simulated data that are often investigated in other fields related to data analysis such as statistics. However, our opinion diverges when it comes to defining the goals of such comparison studies and their interpretation. Our major argument is that comparison studies conducted as a part of an article suggesting a new algorithm rarely reach the level of representativity and objectivity required to be used as guidance by other researchers when choosing their algorithm. The two underlying main ideas behind this claim are that (i) comparison studies comparing a new algorithm with existing algorithms are often severely biased for different reasons, as documented in our empirical study based on the example of supervised classification with high-dimensional data (Jelizarow et al., 2010); simply, when reading articles on new methods, most of us think ‘well, of course they say their method is better, but …’; and (ii) as stressed by Smith et al. (2013), an extensive and fair comparison study requires a lot of work, care and attention in itself (Boulesteix et al., 2013a, b) and can hardly be conducted within a study focused on another problem—the development of a new algorithm. To specify our point of view, we suggest a taxonomy of comparison studies based on real datasets while referring to the ideas presented by Smith et al. (2013). We essentially classify comparison studies based on real datasets into two categories: representative and illustrative comparisons. Representative comparisons aim to give conclusions on the new method for a certain field of application not limited to single datasets. These comparisons include several datasets chosen from a certain field of application that is ideally clearly defined. One hopes from a well-conducted representative comparative application that it will give information guiding the choice of the method in future similar applications from the same field. Note that we do not expect a particular method to outperform all other methods on all datasets from the considered field (Hand, 2006). We rather see representative comparative applications as giving information on the expectation of performance over the ‘distribution of distributions’ that is characteristic for this field (or, simply, over the datasets from this field); see Boulesteix et al. (2013b) for a formal probabilistic framework. In practice, however, representative comparisons presented in articles introducing new methods have three major pitfalls that often make them fail their goal. The first pitfall is that they are often substantially biased in favor of the new method for different reasons, including optimization of the datasets (Yousefi et al., 2010), of the settings or of the method’s characteristics, and publication bias (Boulesteix et al., 2013a; Jelizarow et al., 2010). Another issue that is probably even more difficult to address is the better expertise of the authors on the new method they have been working on for months. As stressed by Smith et al. (2013), applying algorithms designed and programmed by other researchers is often far from easy. The focus of a researcher developing new algorithms is on the new algorithms, and he/she can spend only a limited amount of time and energy trying to understand and optimize the use of other algorithms. It is also likely that researchers spend much time solving a problem occurring in their own algorithm, whereas they tend to spontaneously accept inferior results from another algorithm without trying to solve the problem. In this context, we also claim that it would be extremely difficult (if not impossible) for referees to detect all these issues in the articles they have to review. A bug in one of the competitive algorithms may perhaps be discovered if the referee spends a few hours or days checking the authors’ code. This requires that all codes are made available (which is currently not always the case; see Hothorn and Leisch (2011) for a recent study on this topic) and that the referee has much time to spend on this unpaid reviewing task (this is also unlikely). As far as other aspects such as optimization of the method’s characteristics are concerned, there is in our understanding no way even for the most careful referee to identify them with certitude. Based on all these facts, we stressed the importance of neutral comparison studies (i.e. comparisons that are not part of an article suggesting a new algorithm) in a previous publication (Boulesteix et al., 2013a). The second pitfall is that applications intended as representative comparisons are in practice often underpowered, as documented in Boulesteix et al. (2013b) in the special case of supervised learning. By underpowered, we mean that the performance (or the difference between the performances of two algorithms) is so variable across datasets that more datasets would be needed to draw conclusions from a statistical testing perspective. In a literature survey with focus on supervised classification, Boulesteix et al. (2013a) found that the median number of datasets considered in comparison studies included in articles on new algorithms was only five. In the statistical framework of Boulesteix et al. (2013b), assuming a significance level of 0.05, a power of 0.8 and a ‘typical value’ of 7% for the standard deviation of the performance difference over the datasets, this number of datasets would only allow to detect as significant a difference in error rates of —a big difference! Thus, simply conducting any comparison study is not sufficient to provide guidance: the comparison study also has to have enough power; a condition that is almost never fulfilled in practice. The third pitfall is that it is difficult to draw datasets at random within a defined field of interest. This problem is characterized by lack of literature and definition problems. It should be addressed in future research. Most importantly, it may be related to the overoptimistic bias mentioned earlier—both because authors might tend to underreport the results obtained with datasets that are not favorable to their new algorithm (Yousefi et al., 2010) and because they might choose datasets that are somehow interrelated and yield a distorted picture of the performance of the new algorithm in the whole field of interest. Ideally, representative comparisons that aim to yield general conclusions for an application field should roughly follow the rules considered as standard in many substantive fields, for example in clinical research, whereby in our metaphor datasets play the role of patients, methods play the role of treatments and applications fields play the role of populations. This implies, among others, that one (i) considers power issues while designing a comparison study, especially while determining the number of datasets (Boulesteix et al., 2013b); (ii) the selection of the datasets follows systematic and well-documented criteria in the vein of inclusion criteria used in trials; (iii) subgroup analyses are planned previously and interpreted cautiously; (iv) the main outcome (e.g. the error rate in the case of supervised learning) is clearly defined; (v) the datasets are selected at random or at least potential selection biases are discussed; and (vi) dropout (datasets that are disregarded in the course of the study) and its reasons are well-documented. Note that these recommendations would address the second pitfall but only partially the first one. Following our metaphor, the first pitfall in clinical research would be that inventors of a new treatment tend to be overoptimistic regarding its efficiency for many reasons, a fact that is widely recognized in medical research. In our context, the six recommendations outlined earlier should be completed by a seventh one to better address the first pitfall, say, ‘define the method at the beginning of the study and do not adapt it depending of its results on the considered datasets’, to avoid overfitting of the new algorithm on the considered datasets (Jelizarow et al., 2010; Rocke et al., 2009). Considering the three pitfalls discussed above, it is clear that designing a representative comparison study is an extremely difficult and time-intensive task that can and should probably not be performed in all articles presenting new methods. Indeed, many real data applications presented in articles on new methods are meant as examples. This is what we denote as illustrative comparisons. They may demonstrate the use of the new method and point to specific aspects, such as software implementation (possibly including exemplary code as additional file), preliminary data preparation, parameter/variant choice or computation time. They might also show in which form the results are obtained and how to interpret them. The main concern related to illustrative comparisons is that their results are sometimes (wrongly) interpreted in the literature and discussed as if they were representative. Typically, conclusions are drawn on the superiority of a method (most often the new method; see Boulesteix et al. (2013a) for an empirical study on this topic) based on a too small number of datasets that actually do not allow to draw such conclusions from a testing perspective. Coming back to our metaphor, it would be as if a team of medical scientists established the superiority of a new treatment based on only n = 2 or n = 3 patients. It just does not make sense. Similarly, it does not make sense to make conclusions such as ‘method A performs better than method B for datasets with the characteristics XYZ’ if the study is based on, say, one dataset with this characteristic and one without. It would be as if a medical team said that treatment A is more efficient than treatment B for male patients just because it was the case for the considered male patient but not for the considered female patient. Note that the term ‘illustrative’ does not necessarily imply that the comparison criteria are qualitative. But it implies that these comparisons and criteria are interpreted as examples, and not as representative of the considered field. In particular, differences in performance should not be interpreted in terms of guidance for the choice of the algorithm—even if the results of illustrative comparisons may provide information on the order of magnitude of the relative performance of the considered algorithms and tell us whether, roughly speaking, the ‘new algorithms behave as expected in real data settings’. The main difference between the two types of comparisons is the way in which datasets are selected that essentially affects their interpretation. In representative comparisons, selection is ideally performed at random within the defined field, and their number is chosen by taking power issues into consideration, whereas in illustrative comparisons datasets are selected because they are interesting to better present the new method. For example, in an illustrative application it is acceptable to select two extreme datasets, say (in the case of supervised learning), a dataset where the response is easy to predict and a more difficult dataset. Note that other important aspects of the new algorithms—such as ease of use, speed, conceptual simplicity or generalizability—can be adequately addressed in an article even in the absence of representative comparison study. However, it is important to note that (i) these aspects are of no consequence if the performance turns out to be bad, and (ii) some of them (such as ease of use and speed) also depend on the considered dataset. The above considerations on the selection of example datasets may thus also be relevant to aspects of the algorithms other than performance, even if our classification into illustrative and representative studies primarily focuses on performance. In conclusion, we agree that comparisons of new algorithms to existing algorithms are important. As Smith et al. (2013), we make a plea for more comparison studies on real datasets in bioinformatics literature. Suggesting an algorithm without even running it on a real dataset is just unacceptable. Going back to our parallel between medical and computational sciences, it would be as if a physician described a new therapy without even showing that it was successful on a few patients. The application of the new algorithm to, say, at least two distinct example datasets with different characteristics should be a minimum non-negotiable requirement for publication. Generally, we also think that in research articles applying bioinformatics algorithms to obtain substantive results, the results should not be based on a new novel algorithm that is scarcely described, applied only to the dataset of interest and not compared with any other method. We also believe that both approaches discussed in this letter—illustrative and representative—may make sense, but that they are completely different and should be reported differently. Each approach implies a different way to report the results and a different way to design the experiment. Reporting and design should be consistent. In practice, the most frequent violation of this principle is when algorithm developers ‘feel’ that their new method might be better than existing methods based on an underpowered comparison study and report their results as if the comparison was representative. The design of a representative comparison study is so difficult and time consuming and the risk of bias in favor of the new method is in practice so important that most applications presented in articles on new methods should probably be seen as illustrative applications. In our letter, we discussed a few conditions that a good representative application should in our view fulfill to be really representative. However, we believe that much more effort is needed to define precise criteria and guidelines. In particular, the neutrality issue addressed in Boulesteix et al. (2013a) should be carefully taken into account. Referring to the sentence ‘these evaluations ought to be primarily provided in the novel algorithm publications themselves’ (Smith et al., 2013), we again point out that representative comparisons (as defined in our letter) can hardly be performed within an article on a new method, especially because of the well-known optimistic bias in favor of the new method (see also the editorial by Rocke et al., 2009). Addressing this general issue related to scientific methodology and publication practice will probably need a long time and much coordinated effort from all parties—researchers, editors and reviewers. This problem is also connected with publication bias and publication of negative research findings (Boulesteix, 2010). In the meantime, we believe that authors and readers should interpret illustrative comparisons as such—without implicitly assuming that they give information about what would happen on other datasets, and that journals might give more attention to ‘neutral’ comparison studies entirely devoted to the comparison task itself. This would give a motivation to potential authors of such comparison studies: if they know that their work will not only provide useful information to other scientists but also have good chance to be published in a high-ranking journal, they will be more likely to conduct such a study. Motivating potential authors of comparison studies by publishing more of these studies is certainly easier than motivating overloaded reviewers to spend hours investigating the possible bias of a comparison study included in an article on a new method. To conclude, in this letter we tried to clarify some aspects left unaddressed by Smith et al. (2013). The exact definition of requirements for the publication of new algorithms, however, cannot be formulated within this modest framework. Such guidelines should be the result of coordinated efforts of consortia involving a large number of scientists, similarly to the teams working on reporting guidelines in medical sciences such as the EQUATOR network (Altman et al., 2008). I thank Manuel Eugster for helpful comments. Conflict of Interest: none declared.
Anne-Laure Boulesteix
Bioinform.1
2013 An AUC-based permutation variable importance measure for random forests
abstract
BACKGROUND: The random forest (RF) method is a commonly used tool for classification with high dimensional data as well as for ranking candidate predictors based on the so-called random forest variable importance measures (VIMs). However the classification performance of RF is known to be suboptimal in case of strongly unbalanced data, i.e. data where response class sizes differ considerably. Suggestions were made to obtain better classification performance based either on sampling procedures or on cost sensitivity analyses. However to our knowledge the performance of the VIMs has not yet been examined in the case of unbalanced response classes. In this paper we explore the performance of the permutation VIM for unbalanced data settings and introduce an alternative permutation VIM based on the area under the curve (AUC) that is expected to be more robust towards class imbalance. RESULTS: We investigated the performance of the standard permutation VIM and of our novel AUC-based permutation VIM for different class imbalance levels using simulated data and real data. The results suggest that the new AUC-based permutation VIM outperforms the standard permutation VIM for unbalanced data settings while both permutation VIMs have equal performance for balanced data settings. CONCLUSIONS: The standard permutation VIM loses its ability to discriminate between associated predictors and predictors not associated with the response for increasing class imbalance. It is outperformed by our new AUC-based permutation VIM for unbalanced data settings, while the performance of both VIMs is very similar in the case of balanced classes. The new AUC-based VIM is implemented in the R package party for the unbiased RF variant based on conditional inference trees. The codes implementing our study are available from the companion website: http://www.ibe.med.uni-muenchen.de/organisation/mitarbeiter/070_drittmittel/janitza/index.html.
Silke Janitza, Carolin Strobl, Anne-Laure Boulesteix
BMC Bioinform.3
2012 Random forest Gini importance favours SNPs with large minor allele frequency: impact, sources and recommendations
abstract
The use of random forests is increasingly common in genetic association studies. The variable importance measure (VIM) that is automatically calculated as a by-product of the algorithm is often used to rank polymorphisms with respect to their ability to predict the investigated phenotype. Here, we investigate a characteristic of this methodology that may be considered as an important pitfall, namely that common variants are systematically favoured by the widely used Gini VIM. As a consequence, researchers may overlook rare variants that contribute to the missing heritability. The goal of the present article is 3-fold: (i) to assess this effect quantitatively using simulation studies for different types of random forests (classical random forests and conditional inference forests, that employ unbiased variable selection criteria) as well as for different importance measures (Gini and permutation based); (ii) to explore the trees and to compare the behaviour of random forests and the standard logistic regression model in order to understand the statistical mechanisms behind the preference for common variants; and (iii) to summarize these results and previously investigated properties of random forest VIMs in the context of genetic association studies and to make practical recommendations regarding the choice of the random forest and variable importance type. All our analyses can be reproduced using R code available from the companion website: http://www.ibe.med.uni-muenchen.de/organisation/mitarbeiter/020_professuren/boulesteix/ginibias/.
Anne-Laure Boulesteix, Andreas Bender 0001, Justo Lorenzo Bermejo, Carolin Strobl
Briefings Bioinform.1
2011 Editorial
abstract
Featured in this special issue are 10 articles covering various aspects of validation strategies in bioinformatics and computational molecular medicine. The emergence of complex and high-dimensional ‘omics data’ and abundance of false research findings in this context [1, 2] has recently revived interest in validation issues. ‘Validation’ is a generic term embracing various procedures and concepts. The aim of this special issue is not to cover the full spectrum of such procedures in bioformatics but rather to highlight particularly relevant aspects of validation in different contexts including validation of molecular signatures estimated from high-dimensional omics data, computational reproducibility of bioinformatics analyses, validation of regulatory networks and validation issues in genetic association studies, phylogeny or next-generation sequencing. All scientists would probably agree that validation is important to all scientific activities and that a ‘validated’ research finding is more trustworthy than a finding that has not (yet) been validated. However, they probably understand very different things under the term ‘validation’ that may have various acceptations depending on the context. In bioinformatics, one has to distinguish validation of new bioinformatics algorithms from validation of biomedical research results obtained via bioinformatics algorithms on a specific data set. The first issue and its epistemiological background have surprisingly not focused much attention in the literature. However, a few recent articles (e.g. [3, 4]) demonstrate that optimistic reporting may noticeably affect methodological bioinformatics literature, hence pointing out the importance of the fair assessment of new algorithms and validation. Three of the articles of this special issue can be seen as contributing to this area of research in a broad sense (Dougherty, Stamatakis and Izquierdo-Carrasco, and Hothorn and Leisch). The second issue—validation of biomedical results obtained via bioinformatics methods—is addressed by the other papers of this special issue. These two perspectives on validation in bioinformatics are extensively discussed by Dougherty in the context of gene regulatory networks. The first three articles of the special issue deal with validation of molecular signatures derived from high-dimensional data sets, for instance transcriptomic data sets. Castaldi et al. present the results of an extensive survey of validation practices for validating molecular signatures, and examine the differences between cross-validation results and external validation in 28 published studies. Simon et al. review and illustrate the use of cross-validation techniques in the specific context of survival analysis based on high-dimensional molecular data, where censoring makes assessment of prediction accuracy more complex than in the case of non-censored outcomes. Boulesteix and Sauerbrei address the problem of added predictive value of high-dimensional molecular data to conventional clinical data and its validation. These three papers focus the assessment of prediction models derived from high-dimensional data sets: with cross-validation or related techniques in Simon et al.; with cross-validation or external data in Castaldi et al.; and with emphasis on added predictive value compared to classical predictors in Boulesteix and Sauerbrei. The two next papers of this special issue are devoted to the validation and assessment of biological networks. Mansmann and Jurinovic discuss validation of inferred networks from a practical point of view including an illustration through a leukaemia study. They address the problem of the biological interpretability of the output networks, while Dougherty discusses the validity of gene regulatory network models from an epistemiological point of view using a distance function between two networks. He examines the two perspectives of ‘network validation’: the ability to make predictions from a given network model, and the ability of a network inference procedure to infer a network from a sample. The four next papers address validation in different specific application areas. Thompson et al. and König review and critically discuss validation strategies in the context of genetic association studies. While König gives an overview of validation for genetic association studies in general including the definition of the term validation in this context and a review of the different types of approaches, Thompson et al. discuss meta-analysis techniques, which are in a broad sense related to validation. The paper by Stamatakis and Izquierdo-Carrasco focus on two issues related to validation and phylogeny software: the inference of support values on trees that provide some notion about the ‘correctness’ of the tree and more importantly, program verification, thus addressing the validation of bioinformatics algorithms (issue 1 discussed above). Finally, Fang & Cui review design and validation issues in the emerging field of next-generation sequencing experiments including sample size issues. The last paper of this special issue by Hothorn and Leisch is devoted to a topic relevant to both methodological bioinformatics research and biomedical research using bioinformatics tools. The publication of codes with the aim to make data analyses ‘reproducible’ may also greatly contribute to more transparent reporting and enable a more rapid verification and validation of published research results. Considering the large amount of research findings that fail to get validated, I believe that validation will become an increasingly important issue in bioinformatics, both from a methodological perspective (validation of new algorithms) and from the user’s point of view (validation of the results obtained with bioinformatics tools). It is up to all of us to ensure that omics data analysis does not yield too much ‘noise discovery’ [1, 2]. The problem of false research findings is inherent to any scientific activity and enhanced by the nature of omic data (small samples, many variables, missing standards and guidelines). In this context, well thought-out validation approaches are crucial. Important steps have been done. Important steps still have to be done. ALB was supported by the LMU–innovativ Project BioMed–S: Analysis and Modelling of Complex Systems in Biology and Medicine.
Anne-Laure Boulesteix
Briefings Bioinform.1
2011 Added predictive value of high-throughput molecular data to clinical data and its validation
abstract
Hundreds of 'molecular signatures' have been proposed in the literature to predict patient outcome in clinical settings from high-dimensional data, many of which eventually failed to get validated. Validation of such molecular research findings is thus becoming an increasingly important branch of clinical bioinformatics. Moreover, in practice well-known clinical predictors are often already available. From a statistical and bioinformatics point of view, poor attention has been given to the evaluation of the added predictive value of a molecular signature given that clinical predictors or an established index are available. This article reviews procedures that assess and validate the added predictive value of high-dimensional molecular data. It critically surveys various approaches for the construction of combined prediction models using both clinical and molecular data, for validating added predictive value based on independent data, and for assessing added predictive value using a single data set.
Anne-Laure Boulesteix, Willi Sauerbrei
Briefings Bioinform.1
2010 Over-optimism in bioinformatics research
abstract
Contact: [email protected] The problem of ‘false research findings’ in medical research has focused much attention in the last few years (Ioannidis, 2005). One of the main problems, termed as ‘fishing for significance’ in the present letter, is that researchers often (consciously or subconsciously) report results that are in fact the product of an intensive optimization, i.e. of multiple comparisons. Such results are typically unlikely to be reproduced in an independent study and have a high probability to be false (Ioannidis, 2005). The ‘fishing for significance’ problem is enhanced by the so-called ‘publication bias’: positive results have a much higher chance to get published than negative results, as already acknowledged 50 years ago (Sterling, 1959). In a word, many false positive results are produced through multiple comparisons, and false positives have higher chance to get published than true negatives. Moreover, the difficulty to publish negative results obviously encourages authors to find something positive in their study by performing numerous analyses until one of them yields positive results by chance, i.e. to fish for significance. Although this issue is by far less acknowledged and publicly admitted than in the medical context, the same types of problems occur in biostatistics and bioinformatics research. In a recent editorial of the journal Bioinformatics on ‘Papers on normalization, variable selection, classification or clustering of microarray data’, Rocke et al. (2009) state that ‘prediction methods enter a crowded area’. Indeed, hundreds of prediction algorithms for high-dimensional small-sample data have been proposed in statistics, bioinformatics and machine learning journals in the last few years. Rocke et al. (2009) further claim that ‘consciously or subconsciously, the developer of a new method optimizes its characteristics against the datasets to be used for evaluation’. This is because the development of a new prediction algorithm is often a trial-and-error learning task. The emerging method is successively adapted depending of the intermediate results. This problem can be paralleled to the ‘fishing for significance’ issue in medical research, except that in bioinformatics research the researcher fishes for an improvement (e.g. a decrease of the error rate) instead of fishing for a significant P-value. In statistical bioinformatics research, fishing for significance consists of two distinct components: (i) the sequential adaptation of the new methods to the considered datasets as described in the Bioinformatics editorial, and (ii) the search for a specific dataset or simulation setting for which the new method works better than existing approaches as quantitatively investigated in the recent paper by Yousefi et al. (2009). Both mechanisms lead to optimistic conclusions regarding the superiority of the new method. The first mechanism essentially affects all research fields related to data analysis such as statistics, machine learning or bioinformatics. The trial-and-error process is indeed an important component of data analysis research. One would not expect a statistician or bioinformatician to develop a method with pen and paper, try it once on a dataset and immediately write a paper. Most original good ideas have to be sequentially improved before reaching an acceptable maturity. The development of a new method is per se an unpredictable search process. Thus, the concept of analysis plan known from medical and pharmaceutical research cannot be easily accommodated to bioinformatics research. The problem is that, as stated by the Bioinformatics editorial team, this search process leads to an artificial optimization of the method's characteristics to the considered datasets. Hence, the superiority of the novel method over an existing method (as measured, e.g. through the difference between the cross-validation or bootstrap error rates) is sometimes considerably overestimated. This problem potentially occurs in all data analysis problems with a clearly defined and objective quality criteria. In a concrete medical prediction study, fitting a prediction model and estimating its error rate using the same training dataset yields a downwardly biased error estimate commonly termed as ‘apparent error’. Validation on independent fresh data is an important component of all prediction studies. Similarly, developing a new algorithm and evaluating it by comparison to existing methods using the same datasets may lead to optimistically biased results in the sense that the new algorithm's characteristics overfit the used datasets. The over-optimistic result is the superiority of the new algorithm compared with existing methods (for instance, in terms of prediction error) rather than (like in concrete prediction studies) the prediction error itself. In the same way as a prediction rule has to be validated using fresh data in applied research, one can try to validate the superiority of the new algorithm in methodological research. While this idea may appear at first glance as a trivial generalization of the validation concept from medical research, it raises its own methodological difficulties. Getting to relevant validation datasets is generally not a problem, in contrast to what happens in applied biomedical research. A plethora of data of all types can be found in the WWW on public repositories, journal web sites or homepages of the researchers. Completely new data types are an important exception. But in most cases, the main problem is rather the definition of eligibility criteria than the search for datasets. Authors may be consciously or subconsciously tempted to select a particular dataset because it is excepted to yield better results with the new method. This leads us to the second component of the fishing for significance mechanism: the biased selection of datasets that makes the new algorithm look better than it actually is. This biased selection may occur at different levels. For example, authors may deliberately omit to report the results obtained with a particular dataset just because existing methods outperform the new algorithm on these data. This kind of reporting bias is quantitatively investigated in the remarkable study by Yousefi et al. (2009) with striking results. Selecting the example datasets based on their results yields a substantial optimistic bias in error rate estimation. A less extreme variant of this scenario is when authors choose to ignore a dataset in their study because they suspect that, based on theoretical considerations or past experience from similar studies, their new algorithm will not perform well on this particular data structure. For instance, a prediction method for high-dimensional data may be expected to perform badly in the case of many very weak covariates or conversely in the case of few very strong covariates. In contrast with the deliberate omission of bad results, the omission of datasets that are expected to yield bad results may be correct as long as the authors state in their paper that the method is especially designed for a particular data situation but may be less appropriate in other cases. Conversely, it would be misleading to highlight the general character of the method but evaluate it only on a particular type of data which is expected to yield advantageous results. It is not wrong to focus on a particular data type, but these restrictions should be well documented and openly admitted. This problem is related to the theoretical assumptions underlying a method as extensively discussed by Mehta et al. (2004). It is unreasonable to expect that a new method performs universally better than existing methods in all settings: most methods perform well under certain assumptions only. It is thus important to specify the assumptions of a new method including both theoretically motivated assumptions and restrictions that were identified empirically. That said, evaluating the effect of violations may also be interesting, especially in the case of very restrictive assumptions (Mehta et al., 2004). To sum up, two mechanisms combine to yield over-optimistic results. First, the new method's characteristics often overfit the datasets which were used for its development. Secondly, when a new method is evaluated after its development, the biased selection of datasets also leads to over-optimistic results. The first problem is essentially inevitable. With respect to the second problem, the realistic definition and reporting of the area of application of the new method and the systematic selection of test datasets within this area should be given much attention in order to avoid over-optimistic results. Most importantly, we outline the importance of the two-stage approach: first, develop the method using example datasets and define its area of application as precisely as possible, then evaluate the developed method based on other datasets within the area of application (and report all the obtained results). This workflow, which is already consciously or subconsciously adopted by some authors, should perhaps be applied more stringently and consistently to ensure proper validation of research results. The difficulty to unbiasedly select eligible datasets and describe the field of application of the new method is certainly enhanced by the fact that journals accept mostly positive research results (a method that performs better), sometimes neutral results (a review or a comparison of existing methods), but almost never negative results (a promising and sensible method that finally does not fulfill the authors’ expectation), which leads us to the second major cause of false published research findings: the publication bias. Publication biases and the necessity to ‘accentuate the negative’ (PLoS Medicine Editors, 2009) are well-documented in the context of medical and pharmaceutical research. Much effort has been taken to reduce the publication bias, for instance the prepublication registration of trials—often without much success (Ross et al., 2009). Similarly, publication biases also affect studies on new bioinformatics or statistical methods, perhaps even more drastically than medical studies. A well-designed medical study with negative results has a reasonable chance to get published, at least in a low-impact journal. In contrast, a study on a new statistical method that turns out to be worse than existing methods in terms of objective performance will be rejected by almost all journals. In practice, the only way to present such negative results to the community is to include them in a comparison study. As an example, let us consider a novel method A that 10 independent researcher teams consider as promising, but that is in fact not better than existing approaches. Eight of the 10 researcher teams correctly find that the promising method is finally not better. They forget it and do not write any paper. The ninth one finds a particular dataset on which method A performs better and this positive result gets published. The 10th one fishes for significance and publishes a variant A′ that performs better than existing methods on his datasets. Note that in a subsequent validation based on an independent dataset, method A′ would perhaps have not been better than existing methods, but here we assume that the authors did not perform such a strict validation. Finally, two papers documenting the superiority of method A/A′ are published, although eight researcher teams found that it is not better than existing approaches. This example is probably caricatural, but similar things may occur in real life. Of course, one could argue that, if method A is really bad, other researchers will not further use it and the community will finally find out that the method is bad. However, it would certainly have been better to publish the negative studies for different reasons. The ability of the method to establish itself also depends on various factors (such as the availability of software, the personality of the authors, etc.), such that it may take much time to find out that method A is actually bad. Moreover, as stated in the instructions for authors of the recently initiated Journal of Interesting Negative Results in Natural Language Processing and Machine Learning, ‘knowing directions that lead to dead-ends in research can help others avoid replicating paths that take them nowhere’ and ‘much can be learned by analyzing why some ideas, while intuitive and plausible, do not work. […] Negative results may point to interesting and important open problems’. Hence, the publication of well-conducted studies with negative result may be more useful than commonly assumed by journals and reviewers. Moreover, publishing negative studies would potentially give more importance to qualitative aspects of new methods such as their conceptual simplicity, computational efficiency, interpretability, flexibility, ability to generalize or fit in a global framework, the absence of unplausible assumptions or, most importantly, the originality of the addressed research question. A negative aspect (an error rate that is slightly larger than those of existing methods) may be counterbalanced by positive aspects. For instance, in the context of prediction using high-dimensional data, a sparse method that can handle highly correlated variables may attract users even if its accuracy is not better (or even slightly worse) than existing methods. In the same vein, some studies that are negative in terms of the objective evaluation criterion (such as the error rate) may in fact contribute to the scientific progress as much as other studies with positive results because they suggest a completely new class of promising methods whose variants may eventually perform better in further studies. Last but not least, partially relaxing the requirement for quantitative improvement (e.g. in terms of error rate or power) may in the long-term encourage honest and unbiased reporting and reduce the temptation to fish for significance. One could argue that it would also give more space to subjectivity in the review process. The decision whether a method with disappointing results was originally promising and whether readers may be interested in the negative conclusion may indeed be quite subjective. Assessing whether a new method with negative results is worth publishing is anything but trivial and suitable criteria should be carefully defined. But if negative results are systematically excluded from publication, authors are virtually urged to make their results seem positive and the reviewers’ task is hindered by biased reporting, which also implies much subjectivity. Whether the referee should ‘believe’ the apparently positive results or not is a highly subjective question. Hence, the difficulty to objectively evaluate a negative study may be counterbalanced by an increased reporting transparency and a better application of appropriate validation procedures. On the whole, I believe that the occasional publication of well-designed studies on promising sensible ideas with disappointing quantitative results may in the long run contribute to a less optimistically biased literature. In this sense, the recently launched journals publishing negative results are an important step forward. The publication of analysis scripts with the aim of research reproducibility (Hothorn et al., 2009; Peng, 2009) may also greatly contribute to more transparent reporting and enable a rapid validation of research results. Note that the availability of computer codes for reproducing the obtained results does not ensure that the authors did not overfit their method to the analyzed datasets. However, reproducible analyses considerably simplify the post-publication unbiased validation of the research findings. If the method can be quickly tested by running a well-documented script, readers can easily find out whether it performs as well as claimed in the article with the considered data types, or e.g. identify parameters that were tuned consciously or subconsciously. In this sense, a stringent reproducibility policy may, among many other advantages, reduce the temptation to fish for significance and help researchers to adopt a somewhat more objective point of view on their own studies. In this perspective, the publication of computer codes for the purpose of reproducibility should be expressly encouraged by editors and referees. I thank Carolin Strobl and Nicole Krämer for helpful comments. Funding: LMU-innovativ Project BioMed-S: Analysis and Modelling of Complex Systems in Biology and Medicine. Conflict of Interest: none declared.
Anne-Laure Boulesteix
Bioinform.1
2010 Over-optimism in bioinformatics: an illustration
abstract
MOTIVATION: In statistical bioinformatics research, different optimization mechanisms potentially lead to 'over-optimism' in published papers. So far, however, a systematic critical study concerning the various sources underlying this over-optimism is lacking. RESULTS: We present an empirical study on over-optimism using high-dimensional classification as example. Specifically, we consider a 'promising' new classification algorithm, namely linear discriminant analysis incorporating prior knowledge on gene functional groups through an appropriate shrinkage of the within-group covariance matrix. While this approach yields poor results in terms of error rate, we quantitatively demonstrate that it can artificially seem superior to existing approaches if we 'fish for significance'. The investigated sources of over-optimism include the optimization of datasets, of settings, of competing methods and, most importantly, of the method's characteristics. We conclude that, if the improvement of a quantitative criterion such as the error rate is the main contribution of a paper, the superiority of new algorithms should always be demonstrated on independent validation data. AVAILABILITY: The R codes and relevant data can be downloaded from http://www.ibe.med.uni-muenchen.de/organisation/mitarbeiter/020_professuren/boulesteix/overoptimism/, such that the study is completely reproducible.
Monika Jelizarow, Vincent Guillemot, Arthur Tenenhaus, Korbinian Strimmer, Anne-Laure Boulesteix
Bioinform.5
2010 Testing the additional predictive value of high-dimensional molecular data
abstract
BACKGROUND: While high-dimensional molecular data such as microarray gene expression data have been used for disease outcome prediction or diagnosis purposes for about ten years in biomedical research, the question of the additional predictive value of such data given that classical predictors are already available has long been under-considered in the bioinformatics literature. RESULTS: We suggest an intuitive permutation-based testing procedure for assessing the additional predictive value of high-dimensional molecular data. Our method combines two well-known statistical tools: logistic regression and boosting regression. We give clear advice for the choice of the only method parameter (the number of boosting iterations). In simulations, our novel approach is found to have very good power in different settings, e.g. few strong predictors or many weak predictors. For illustrative purpose, it is applied to the two publicly available cancer data sets. CONCLUSIONS: Our simple and computationally efficient approach can be used to globally assess the additional predictive power of a large number of candidate predictors given that a few clinical covariates or a known prognostic index are already available. It is implemented in the R package "globalboosttest" which is publicly available from R-forge and will be sent to the CRAN as soon as possible.
Anne-Laure Boulesteix, Torsten Hothorn
BMC Bioinform.1
2009 Stability and aggregation of ranked gene lists
abstract
Ranked gene lists are highly instable in the sense that similar measures of differential gene expression may yield very different rankings, and that a small change of the data set usually affects the obtained gene list considerably. Stability issues have long been under-considered in the literature, but they have grown to a hot topic in the last few years, perhaps as a consequence of the increasing skepticism on the reproducibility and clinical applicability of molecular research findings. In this article, we review existing approaches for the assessment of stability of ranked gene lists and the related problem of aggregation, give some practical recommendations, and warn against potential misuse of these methods. This overview is illustrated through an application to a recent leukemia data set using the freely available Bioconductor package GeneSelector.
Anne-Laure Boulesteix, Martin Slawski
Briefings Bioinform.1
2009 Regularized estimation of large-scale gene association networks using graphical Gaussian models
abstract
BACKGROUND: Graphical Gaussian models are popular tools for the estimation of (undirected) gene association networks from microarray data. A key issue when the number of variables greatly exceeds the number of samples is the estimation of the matrix of partial correlations. Since the (Moore-Penrose) inverse of the sample covariance matrix leads to poor estimates in this scenario, standard methods are inappropriate and adequate regularization techniques are needed. Popular approaches include biased estimates of the covariance matrix and high-dimensional regression schemes, such as the Lasso and Partial Least Squares. RESULTS: In this article, we investigate a general framework for combining regularized regression methods with the estimation of Graphical Gaussian models. This framework includes various existing methods as well as two new approaches based on ridge regression and adaptive lasso, respectively. These methods are extensively compared both qualitatively and quantitatively within a simulation study and through an application to six diverse real data sets. In addition, all proposed algorithms are implemented in the R package "parcor", available from the R repository CRAN. CONCLUSION: In our simulation studies, the investigated non-sparse regression methods, i.e. Ridge Regression and Partial Least Squares, exhibit rather conservative behavior when combined with (local) false discovery rate multiple testing in order to decide whether or not an edge is present in the network. For networks with higher densities, the difference in performance of the methods decreases. For sparse networks, we confirm the Lasso's well known tendency towards selecting too many edges, whereas the two-stage adaptive Lasso is an interesting alternative that provides sparser solutions. In our simulations, both sparse and non-sparse methods are able to reconstruct networks with cluster structures. On six real data sets, we also clearly distinguish the results obtained using the non-sparse methods and those obtained using the sparse methods where specification of the regularization parameter automatically means model selection. In five out of six data sets, Partial Least Squares selects very dense networks. Furthermore, for data that violate the assumption of uncorrelated observations (due to replications), the Lasso and the adaptive Lasso yield very complex structures, indicating that they might not be suited under these conditions. The shrinkage approach is more stable than the regression based approaches when using subsampling.
Nicole Krämer 0002, Juliane Schäfer, Anne-Laure Boulesteix
BMC Bioinform.3
2008 Microarray-based classification and clinical predictors: on combined classifiers and additional predictive value
abstract
MOTIVATION: In the context of clinical bioinformatics methods are needed for assessing the additional predictive value of microarray data compared to simple clinical parameters alone. Such methods should also provide an optimal prediction rule making use of all potentialities of both types of data: they should ideally be able to catch subtypes which are not identified by clinical parameters alone. Moreover, they should address the question of the additional predictive value of microarray data in a fair framework. RESULTS: We propose a novel but simple two-step approach based on random forests and partial least squares (PLS) dimension reduction embedding the idea of pre-validation suggested by Tibshirani and colleagues, which is based on an internal cross-validation for avoiding overfitting. Our approach is fast, flexible and can be used both for assessing the overall additional significance of the microarray data and for building optimal hybrid classification rules. Its efficiency is demonstrated through simulations and an application to breast cancer and colorectal cancer data. AVAILABILITY: Our method is implemented in the freely available R package 'MAclinical' which can be downloaded from http://www.stat.uni-muenchen.de/~socher/MAclinical
Anne-Laure Boulesteix, Christine Porzelius, Martin Daumer
Bioinform.1
2008 CMA - a comprehensive Bioconductor package for supervised classification with high dimensional data
abstract
BACKGROUND: For the last eight years, microarray-based classification has been a major topic in statistics, bioinformatics and biomedicine research. Traditional methods often yield unsatisfactory results or may even be inapplicable in the so-called "p >> n" setting where the number of predictors p by far exceeds the number of observations n, hence the term "ill-posed-problem". Careful model selection and evaluation satisfying accepted good-practice standards is a very complex task for statisticians without experience in this area or for scientists with limited statistical background. The multiplicity of available methods for class prediction based on high-dimensional data is an additional practical challenge for inexperienced researchers. RESULTS: In this article, we introduce a new Bioconductor package called CMA (standing for "Classification for MicroArrays") for automatically performing variable selection, parameter tuning, classifier construction, and unbiased evaluation of the constructed classifiers using a large number of usual methods. Without much time and effort, users are provided with an overview of the unbiased accuracy of most top-performing classifiers. Furthermore, the standardized evaluation framework underlying CMA can also be beneficial in statistical research for comparison purposes, for instance if a new classifier has to be compared to existing approaches. CONCLUSION: CMA is a user-friendly comprehensive package for classifier construction and evaluation implementing most usual approaches. It is freely available from the Bioconductor website at (http://bioconductor.org/packages/2.3/bioc/html/CMA.html).
Martin Slawski, Martin Daumer, Anne-Laure Boulesteix
BMC Bioinform.3
2008 Conditional variable importance for random forests
abstract
BACKGROUND: Random forests are becoming increasingly popular in many scientific fields because they can cope with "small n large p" problems, complex interactions and even highly correlated predictor variables. Their variable importance measures have recently been suggested as screening tools for, e.g., gene expression studies. However, these variable importance measures show a bias towards correlated predictor variables. RESULTS: We identify two mechanisms responsible for this finding: (i) A preference for the selection of correlated predictors in the tree building process and (ii) an additional advantage for correlated predictor variables induced by the unconditional permutation scheme that is employed in the computation of the variable importance measure. Based on these considerations we develop a new, conditional permutation scheme for the computation of the variable importance measure. CONCLUSION: The resulting conditional variable importance reflects the true impact of each predictor variable more reliably than the original marginal approach.
Carolin Strobl, Anne-Laure Boulesteix, Thomas Kneib, Thomas Augustin 0001, Achim Zeileis
BMC Bioinform.2
2007 Partial least squares: a versatile tool for the analysis of high-dimensional genomic data
abstract
Partial least squares (PLS) is an efficient statistical regression technique that is highly suited for the analysis of genomic and proteomic data. In this article, we review both the theory underlying PLS as well as a host of bioinformatics applications of PLS. In particular, we provide a systematic comparison of the PLS approaches currently employed, and discuss analysis problems as diverse as, e.g. tumor classification from transcriptome data, identification of relevant genes, survival analysis and modeling of gene networks and transcription factor activities.
Anne-Laure Boulesteix, Korbinian Strimmer
Briefings Bioinform.1
2007 WilcoxCV: an R package for fast variable selection in cross-validation
abstract
UNLABELLED: In the last few years, numerous methods have been proposed for microarray-based class prediction. Although many of them have been designed especially for the case n << p (much more variables than observations), preliminary variable selection is almost always necessary when the number of genes reaches several tens of thousands, as usual in recent data sets. In the two-class setting, the Wilcoxon rank sum test statistic is, with the t-statistic, one of the standard approaches for variable selection. It is well known that the variable selection step must be seen as a part of classifier construction and, as such, be performed based on training data only. When classifier accuracy is evaluated via cross-validation or Monte-Carlo cross-validation, it means that we have to perform p Wilcoxon or t-tests for each iteration, which becomes a daunting task for increasing p. As a consequence, many authors often perform variable selection only once using all the available data, which can induce a dramatic underestimation of error rate and thus lead to misleadingly reporting predictive power. We propose a very fast implementation of variable selection based on the Wilcoxon test for use in cross-validation and Monte Carlo cross-validation (also known as random splitting into learning and test sets). This implementation is based on a simple mathematical formula using only the ranks calculated from the original data set. AVAILABILITY: Our method is implemented in the freely available R package WilcoxCV which can be downloaded from the Comprehensive R Archive Network at http://cran.r-project.org/src/contrib/Descriptions/WilcoxCV.html.
Anne-Laure Boulesteix
Bioinform.1
2007 Bias in random forest variable importance measures: Illustrations, sources and a solution
abstract
BACKGROUND: Variable importance measures for random forests have been receiving increased attention as a means of variable selection in many classification tasks in bioinformatics and related scientific fields, for instance to select a subset of genetic markers relevant for the prediction of a certain disease. We show that random forest variable importance measures are a sensible means for variable selection in many applications, but are not reliable in situations where potential predictor variables vary in their scale of measurement or their number of categories. This is particularly important in genomics and computational biology, where predictors often include variables of different types, for example when predictors include both sequence data and continuous variables such as folding energy, or when amino acid sequence data show different numbers of categories. RESULTS: Simulation studies are presented illustrating that, when random forest variable importance measures are used with data of varying types, the results are misleading because suboptimal predictor variables may be artificially preferred in variable selection. The two mechanisms underlying this deficiency are biased variable selection in the individual classification trees used to build the random forest on one hand, and effects induced by bootstrap sampling with replacement on the other hand. CONCLUSION: We propose to employ an alternative implementation of random forests, that provides unbiased variable selection in the individual classification trees. When this method is applied using subsampling without replacement, the resulting variable importance measures can be used reliably for variable selection even in situations where the potential predictor variables vary in their scale of measurement or their number of categories. The usage of both random forest algorithms and their variable importance measures in the R system for statistical computing is illustrated and documented thoroughly in an application re-analyzing data from a study on RNA editing. Therefore the suggested method can be applied straightforwardly by scientists in bioinformatics research.
Carolin Strobl, Anne-Laure Boulesteix, Achim Zeileis, Torsten Hothorn
BMC Bioinform.2
2003 A CART-based approach to discover emerging patterns in microarray data
abstract
MOTIVATION: Cancer diagnosis using gene expression profiles requires supervised learning and gene selection methods. Of the many suggested approaches, the method of emerging patterns (EPs) has the particular advantage of explicitly modeling interactions among genes, which improves classification accuracy. However, finding useful (i.e. short and statistically significant) EP is typically very hard. METHODS: Here we introduce a CART-based approach to discover EPs in microarray data. The method is based on growing decision trees from which the EPs are extracted. This approach combines pattern search with a statistical procedure based on Fisher's exact test to assess the significance of each EP. Subsequently, sample classification based on the inferred EPs is performed using maximum-likelihood linear discriminant analysis. RESULTS: Using simulated data as well as gene expression data from colon and leukemia cancer experiments we assessed the performance of our pattern search algorithm and classification procedure. In the simulations, our method recovers a large proportion of known EPs while for real data it is comparable in classification accuracy with three top-performing alternative classification algorithms. In addition, it assigns statistical significance to the inferred EPs and allows to rank the patterns while simultaneously avoiding overfit of the data. The new approach therefore provides a versatile and computationally fast tool for elucidating local gene interactions as well as for classification. AVAILABILITY: A computer program written in the statistical language R implementing the new approach is freely available from the web page http://www.stat.uni-muenchen.de/~socher/
Anne-Laure Boulesteix, Gerhard Tutz 0001, Korbinian Strimmer
Bioinform.1