Andreas Bender 0001

dblp:01/990-1 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0001-5628-8611ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 A large-scale neutral comparison study of survival models on low-dimensional data
abstract
MOTIVATION: This work presents the first large-scale neutral benchmark experiment focused on single-event, right-censored, low-dimensional survival data. Benchmark experiments are essential in methodological research to scientifically compare new and existing model classes through proper empirical evaluation. Existing benchmarks in the survival literature are smaller in scale regarding the number of used datasets and extent of empirical evaluation. They often lack appropriate tuning or evaluation procedures, while other comparison studies focus on qualitative reviews rather than quantitative comparisons. This comprehensive study aims to fill the gap by neutrally evaluating a broad range of methods and providing generalizable guidelines for practitioners. RESULTS: We benchmark 21 models, ranging from classical statistical approaches to many common machine learning methods, on 34 publicly available datasets. The benchmark tunes models using both a discrimination measure (Harrell's C-index) and a scoring rule (Integrated Survival Brier Score), and evaluates them across six metrics covering discrimination, calibration, and overall predictive performance. Despite superior average ranks in overall predictive performance from individual learners like oblique random survival forests and likelihood-based boosting, and better discrimination rankings from multiple boosting- and tree-based methods as well as parametric survival models, no method statistically significantly outperforms the commonly used Cox proportional hazards model for either tuning measure. We conclude that while the Cox Proportional Hazards model remains a robust default for low-dimensional, right-censored survival data, more flexible methods may be preferable for specific dataset characteristics. AVAILABILITY AND IMPLEMENTATION: All code, data, and results are publicly available on GitHub https://github.com/slds-lmu/paper_2023_survival_benchmark and archived on Zenodo https://doi.org/10.5281/zenodo.19075310.
Lukas Burk, John Zobolas, Bernd Bischl, Andreas Bender 0001, Marvin N. Wright, Raphael Sonabend
Bioinform.4
2025 On Training Survival Models with Scoring Rules
Philipp Kopper, David Rügamer, Raphael Sonabend, Bernd Bischl, Andreas Bender 0001
ECML/PKDD (7)5
2022 DeepPAMM: Deep Piecewise Exponential Additive Mixed Models for Complex Hazard Structures in Survival Analysis
Philipp Kopper, Simon Wiegrebe, Bernd Bischl, Andreas Bender 0001, David Rügamer
PAKDD (2)4
2022 Factorized Structured Regression for Large-Scale Varying Coefficient Models
David Rügamer, Andreas Bender 0001, Simon Wiegrebe, Daniel Racek, Bernd Bischl, Christian L. Müller, Clemens Stachl
ECML/PKDD (5)2
2022 Avoiding C-hacking when evaluating survival distribution predictions with discrimination measures
abstract
MOTIVATION: In this article, we consider how to evaluate survival distribution predictions with measures of discrimination. This is non-trivial as discrimination measures are the most commonly used in survival analysis and yet there is no clear method to derive a risk prediction from a distribution prediction. We survey methods proposed in literature and software and consider their respective advantages and disadvantages. RESULTS: Whilst distributions are frequently evaluated by discrimination measures, we find that the method for doing so is rarely described in the literature and often leads to unfair comparisons or 'C-hacking'. We demonstrate by example how simple it can be to manipulate results and use this to argue for better reporting guidelines and transparency in the literature. We recommend that machine learning survival analysis software implements clear transformations between distribution and risk predictions in order to allow more transparent and accessible model evaluation. AVAILABILITY AND IMPLEMENTATION: The code used in the final experiment is available at https://github.com/RaphaelS1/distribution_discrimination.
Raphael Sonabend, Andreas Bender 0001, Sebastian J. Vollmer
Bioinform.2
2021 mlr3proba: an R package for machine learning in survival analysis
abstract
SUMMARY: As machine learning has become increasingly popular over the last few decades, so too has the number of machine-learning interfaces for implementing these models. Whilst many R libraries exist for machine learning, very few offer extended support for survival analysis. This is problematic considering its importance in fields like medicine, bioinformatics, economics, engineering and more. mlr3proba provides a comprehensive machine-learning interface for survival analysis and connects with mlr3's general model tuning and benchmarking facilities to provide a systematic infrastructure for survival modelling and evaluation. AVAILABILITY AND IMPLEMENTATION: mlr3proba is available under an LGPL-3 licence on CRAN and at https://github.com/mlr-org/mlr3proba, with further documentation at https://mlr3book.mlr-org.com/survival.html.
Raphael Sonabend, Franz J. Király, Andreas Bender 0001, Bernd Bischl, Michel Lang
Bioinform.3
2020 A General Machine Learning Framework for Survival Analysis
Andreas Bender 0001, David Rügamer, Fabian Scheipl, Bernd Bischl
ECML/PKDD (3)1
2012 Random forest Gini importance favours SNPs with large minor allele frequency: impact, sources and recommendations
abstract
The use of random forests is increasingly common in genetic association studies. The variable importance measure (VIM) that is automatically calculated as a by-product of the algorithm is often used to rank polymorphisms with respect to their ability to predict the investigated phenotype. Here, we investigate a characteristic of this methodology that may be considered as an important pitfall, namely that common variants are systematically favoured by the widely used Gini VIM. As a consequence, researchers may overlook rare variants that contribute to the missing heritability. The goal of the present article is 3-fold: (i) to assess this effect quantitatively using simulation studies for different types of random forests (classical random forests and conditional inference forests, that employ unbiased variable selection criteria) as well as for different importance measures (Gini and permutation based); (ii) to explore the trees and to compare the behaviour of random forests and the standard logistic regression model in order to understand the statistical mechanisms behind the preference for common variants; and (iii) to summarize these results and previously investigated properties of random forest VIMs in the context of genetic association studies and to make practical recommendations regarding the choice of the random forest and variable importance type. All our analyses can be reproduced using R code available from the companion website: http://www.ibe.med.uni-muenchen.de/organisation/mitarbeiter/020_professuren/boulesteix/ginibias/.
Anne-Laure Boulesteix, Andreas Bender 0001, Justo Lorenzo Bermejo, Carolin Strobl
Briefings Bioinform.2