EDBT 2026 Demo / reviewers in the wild / expert
Thomas Augustin 0001
dblp:34/2415
· DBLP profile ↗
31ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-1854-6226ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 5 first-author · 13 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Empirical decision theory
Christoph Jansen, Georg Schollmeyer, Thomas Augustin 0001, Julian Rodemann |
Inf. Sci. | 3 |
| 2025 | Consensus in Motion: A Case of Dynamic Rationality of Sequential Learning in Probability Aggregation
Polina Gordienko, Christoph Jansen, Thomas Augustin 0001, Martin Rechenauer |
ECSQARU | 3 |
| 2025 | Explaining Bayesian Optimization by Shapley Values Facilitates Human-AI Collaboration for Exosuit Personalization
Julian Rodemann, Federico Croppi, Philipp Arens, Yusuf Sale, Julia Herbinger, Bernd Bischl, Eyke Hüllermeier, Thomas Augustin 0001, Conor J. Walsh, Giuseppe Casalicchio |
ECML/PKDD (8) | 8 |
| 2024 | Statistical Multicriteria Benchmarking via the GSD-FrontabstractGiven the vast number of classifiers that have been (and continue to be) proposed, reliable methods for comparing them are becoming increasingly important. The desire for reliability is broken down into three main aspects: (1) Comparisons should allow for different quality metrics simultaneously. (2) Comparisons should take into account the statistical uncertainty induced by the choice of benchmark suite. (3) The robustness of the comparisons under small deviations in the underlying assumptions should be verifiable. To address (1), we propose to compare classifiers using a generalized stochastic dominance ordering (GSD) and present the GSD-front as an information-efficient alternative to the classical Pareto-front. For (2), we propose a consistent statistical estimator for the GSD-front and construct a statistical test for whether a (potentially new) classifier lies in the GSD-front of a set of state-of-the-art classifiers. For (3), we relax our proposed test using techniques from robust statistics and imprecise probabilities. We illustrate our concepts on the benchmark suite PMLB and on the platform OpenML. Christoph Jansen, Georg Schollmeyer, Julian Rodemann, Hannah Blocher, Thomas Augustin 0001 |
NeurIPS | 5 |
| 2024 | Imprecise Bayesian optimizationabstractBayesian optimization (BO) with Gaussian processes (GPs) surrogate models is widely used to optimize analytically unknown and expensive-to-evaluate functions. In this paper, we propose a robust version of BO grounded in the theory of imprecise probabilities: Prior-mean-RObust Bayesian Optimization (PROBO). Our method is motivated by an empirical and theoretical analysis of the GP prior specifications’ effect on BO’s convergence. A thorough simulation study finds the prior’s mean parameters to have the highest influence on BO’s convergence among all prior components. We thus turn to this part of the prior GP in more detail. In particular, we prove regret bounds for BO under misspecification of GP prior’s mean parameters. We show that sublinear regret bounds become linear under GP misspecification but stay sublinear if the misspecification-induced error is bounded by the variance of the GP. In response to these empirical and theoretical findings, we introduce PROBO as a univariate generalization of BO that avoids prior mean parameter misspecification. This is achieved by explicitly accounting for prior GP mean imprecision via a prior near-ignorance model. We deploy our approach on graphene production, a real-world optimization problem in materials science, and observe PROBO to converge faster than classical BO.12 Julian Rodemann, Thomas Augustin 0001 |
Knowl. Based Syst. | 2 |
| 2024 | Learning de-biased regression trees and forests from complex samplesabstractAbstract Regression trees and forests are widely used due to their flexibility and predictive accuracy. Whereas typical tree induction assumes independently identically distributed (i.i.d.) data, in many applications the training sample follows a complex sampling structure. This includes unequal probability sampling, which is often found in survey data. Then, a ‘naive estimation’ that simply ignores the sampling weights may be substantially biased. This article analyzes the bias arising from a naive estimation of regression trees or forests under complex sample designs and proposes ways of de-biasing. This is achieved by bridging tree learning to survey statistics, due to the correspondence of the mean-squared-error criterion in regression trees and variance estimation. Transferring population variance estimation approaches from survey statistics to tree induction, indeed considerably reduces the bias in the resulting trees, both in predictions and the tree structure. The latter is particularly crucial if the trees are to be interpreted. Our methodology is extended to random forests, where we show on simulated data and a housing dataset that correcting for complex sample designs leads to overall much better predictive accuracy and more trustworthy interpretation. Interestingly, corrected forests can surpass forests learned on i.i.d. samples in terms of accuracy, which also has important implications for adaptive data collection approaches. Malte Nalenz, Julian Rodemann, Thomas Augustin 0001 |
Mach. Learn. | 3 |
| 2023 | Multi-target Decision Making Under Conditions of Severe Uncertainty
Christoph Jansen, Georg Schollmeyer, Thomas Augustin 0001 |
MDAI | 3 |
| 2023 | Robust statistical comparison of random variables with locally varying scale of measurementabstractSpaces with locally varying scale of measurement, like multidimensional structures with differently scaled dimensions, are pretty common in statistics and machine learning. Nevertheless, it is still understood as an open question how to exploit the entire information encoded in them properly. We address this problem by considering an order based on (sets of) expectations of random variables mapping into such non-standard spaces. This order contains stochastic dominance and expectation order as extreme cases when no, or respectively perfect, cardinal structure is given. We derive a (regularized) statistical test for our proposed generalized stochastic dominance (GSD) order, operationalize it by linear optimization, and robustify it by imprecise probability models. Our findings are illustrated with data from multidimensional poverty measurement, finance, and medicine. Christoph Jansen, Georg Schollmeyer, Hannah Blocher, Julian Rodemann, Thomas Augustin 0001 |
UAI | 5 |
| 2023 | Approximately Bayes-optimal pseudo-label selectionabstractSemi-supervised learning by self-training heavily relies on pseudo-label selection (PLS). This selection often depends on the initial model fit on labeled data. Early overfitting might thus be propagated to the final model by selecting instances with overconfident but erroneous predictions, often referred to as confirmation bias. This paper introduces BPLS, a Bayesian framework for PLS that aims to mitigate this issue. At its core lies a criterion for selecting instances to label: an analytical approximation of the posterior predictive of pseudo-samples. We derive this selection criterion by proving Bayes-optimality of the posterior predictive of pseudo-samples. We further overcome computational hurdles by approximating the criterion analytically. Its relation to the marginal likelihood allows us to come up with an approximation based on Laplace’s method and the Gaussian integral. We empirically assess BPLS on simulated and real-world data. When faced with high-dimensional data prone to overfitting, BPLS outperforms traditional PLS methods. Julian Rodemann, Jann Goschenhofer, Emilio Dorigatti, Thomas Nagler, Thomas Augustin 0001 |
UAI | 5 |
| 2023 | Statistical Comparisons of Classifiers by Generalized Stochastic DominanceabstractAlthough being a crucial question for the development of machine learning algorithms, there is still no consensus on how to compare classifiers over multiple data sets with respect to several criteria. Every comparison framework is confronted with (at least) three fundamental challenges: the multiplicity of quality criteria, the multiplicity of data sets and the randomness of the selection of data sets. In this paper, we add a fresh view to the vivid debate by adopting recent developments in decision theory. Based on so-called preference systems, our framework ranks classifiers by a generalized concept of stochastic dominance, which powerfully circumvents the cumbersome, and often even self-contradictory, reliance on aggregates. Moreover, we show that generalized stochastic dominance can be operationalized by solving easy-to-handle linear programs and moreover statistically tested employing an adapted two-sample observation-randomization test. This yields indeed a powerful framework for the statistical comparison of classifiers over multiple data sets with respect to multiple quality criteria simultaneously. We illustrate and investigate our framework in a simulation study and with a set of standard benchmark data sets. Christoph Jansen, Malte Nalenz, Georg Schollmeyer, Thomas Augustin 0001 |
J. Mach. Learn. Res. | 4 |
| 2022 | Compressed Rule Ensemble LearningabstractEnsembles of decision rules extracted from tree ensembles, like RuleFit, promise a good trade-off between predictive performance and model simplicity. However, they are affected by competing interests: While a sufficiently large number of binary, non-smooth rules is necessary to fit smooth, well generalizing decision boundaries, a too high number of rules in the ensemble severely jeopardizes interpretability. As a way out of this dilemma, we propose to take an extra step in the rule extraction step and compress clusters of similar rules into ensemble rules. The outputs of the individual rules in each cluster are pooled to produce a single soft output, reflecting the original ensemble’s marginal smoothing behaviour. The final model, that we call Compressed Rule Ensemble (CRE), fits a linear combination of ensemble rules. We empirically show that CRE is both sparse and accurate on various datasets, carrying over the ensemble behaviour while remaining interpretable. Malte Nalenz, Thomas Augustin 0001 |
AISTATS | 2 |
| 2022 | Decision Making with State-Dependent Preference Systems
Christoph Jansen, Thomas Augustin 0001 |
IPMU (1) | 2 |
| 2022 | Learning from Categorical Data Subject to Non-random Misclassification and Non-response Under Prior Quasi-Near-Ignorance Using an Imprecise Dirichlet Model
Aziz Omar, Timo von Oertzen, Thomas Augustin 0001 |
IPMU (2) | 3 |
| 2022 | Information efficient learning of complexly structured preferences: Elicitation procedures and their application to decision making under uncertainty
Christoph Jansen, Hannah Blocher, Thomas Augustin 0001, Georg Schollmeyer |
Int. J. Approx. Reason. | 3 |
| 2021 | Internal Validation of Unsupervised Clustering following an Association Accuracy HeuristicabstractOne challenge of unsupervised clustering is that its clustering results cannot be directly evaluated in terms of expected accuracy. In case of internal validation, the clustering is validated based on the compactness within a cluster as well as the separation of clusters. Especially in high dimensional settings, internal validation as well as user inspection, becomes more difficult and expensive the higher the dimension of the data. We therefore propose an association accuracy heuristic, relating the association of results obtained by different methods to their accuracy. This heuristic is based on an analogy to decision making where high homogeneity among the opinions of independent experts is a widely accepted indicator for having chosen the right decision. Analogous to expert opinions, we assess the groupings of different state-of-the art clustering methods. To measure the (dis)similarity of the clustering results, we propose method-association-measures, that are built on an adaption of $\chi^{2}-$based association measures. Our heuristic is investigated in a simulation study as well as on single-cell RNA sequencing data. Incorporating the ground truth allows a validation of the proposed association accuracy heuristic. Our results provide the opportunity to distinguish between situations where clustering results are expected to be trustworthy and settings where external intervention is indispensable to protect ourself against high risk of bad clustering results. Cornelia Fuetterer, Thomas Augustin 0001 |
BIBM | 2 |
| 2019 | Estimation of classification probabilities in small domains accounting for nonresponse relying on imprecise probability
Aziz Omar, Thomas Augustin 0001 |
Int. J. Approx. Reason. | 2 |
| 2018 | Kurt Weichselberger's contribution to imprecise probabilities and statistical inference
Thomas Augustin 0001, Rudolf Seising |
Int. J. Approx. Reason. | 1 |
| 2018 | Concepts for decision making under severe uncertainty with partial ordinal and partial cardinal preferences
Christoph Jansen, Georg Schollmeyer, Thomas Augustin 0001 |
Int. J. Approx. Reason. | 3 |
| 2017 | Decision Theory Meets Linear Optimization Beyond Computation
Christoph Jansen, Thomas Augustin 0001, Georg Schollmeyer |
ECSQARU | 2 |
| 2017 | Special Issue: Ninth International Symposium on Imprecise Probability: Theory and Applications (ISIPTA'15)
Thomas Augustin 0001, Serena Doria, Massimo Marinacci |
Int. J. Approx. Reason. | 1 |
| 2017 | Imprecise probability: Theories and applications
Thomas Augustin 0001, Serena Doria, Massimo Marinacci |
Int. J. Approx. Reason. | 1 |
| 2017 | On the testability of coarsening assumptions: A hypothesis test for subgroup independence
Julia Plass, Marco E. G. V. Cattaneo, Georg Schollmeyer, Thomas Augustin 0001 |
Int. J. Approx. Reason. | 4 |
| 2015 | Statistical modeling under partial identification: Distinguishing three types of identification regions in regression analysis with interval data
Georg Schollmeyer, Thomas Augustin 0001 |
Int. J. Approx. Reason. | 2 |
| 2013 | Information-based dissimilarity assessment in Dempster-Shafer theory
Atiye Sarabi-Jamab, Babak Nadjar Araabi, Thomas Augustin 0001 |
Knowl. Based Syst. | 3 |
| 2012 | Partially identified prevalence estimation under misclassification using the kappa coefficient
Helmut Küchenhoff, Thomas Augustin 0001, Anne Kunz |
Int. J. Approx. Reason. | 2 |
| 2010 | Imprecise probability in statistical inference and decision making
Thomas Augustin 0001, Frank P. A. Coolen, Serafín Moral, Matthias C. M. Troffaes |
Int. J. Approx. Reason. | 1 |
| 2009 | Imprecise probability models and their applications
Thomas Augustin 0001, Enrique Miranda 0001, Jirina Vejnarová |
Int. J. Approx. Reason. | 1 |
| 2009 | A nonparametric predictive alternative to the Imprecise Dirichlet Model: The case of a known number of categories
Frank P. A. Coolen, Thomas Augustin 0001 |
Int. J. Approx. Reason. | 2 |
| 2008 | Conditional variable importance for random forestsabstractBACKGROUND: Random forests are becoming increasingly popular in many scientific fields because they can cope with "small n large p" problems, complex interactions and even highly correlated predictor variables. Their variable importance measures have recently been suggested as screening tools for, e.g., gene expression studies. However, these variable importance measures show a bias towards correlated predictor variables. RESULTS: We identify two mechanisms responsible for this finding: (i) A preference for the selection of correlated predictors in the tree building process and (ii) an additional advantage for correlated predictor variables induced by the unconditional permutation scheme that is employed in the computation of the variable importance measure. Based on these considerations we develop a new, conditional permutation scheme for the computation of the variable importance measure. CONCLUSION: The resulting conditional variable importance reflects the true impact of each predictor variable more reliably than the original marginal approach. Carolin Strobl, Anne-Laure Boulesteix, Thomas Kneib, Thomas Augustin 0001, Achim Zeileis |
BMC Bioinform. | 4 |
| 2008 | Bayesian learning for a class of priors with prescribed marginals
Hermann Held, Thomas Augustin 0001, Elmar Kriegler |
Int. J. Approx. Reason. | 2 |
| 2007 | Decision making under incomplete data using the imprecise Dirichlet model
Lev V. Utkin, Thomas Augustin 0001 |
Int. J. Approx. Reason. | 2 |