Ludwig Bothmann

dblp:187/8625 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0002-1471-6582ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 50% Probabilistic and Bayesian machine learning · 30% Kernel, tree and ensemble methods · 15%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
0.912025
TabFairGDT: A Fast Fair Tabular Data Generator Using Autoregressive Decision Trees · ICDM 2025
Machine learning › Trustworthy machine learning › fairness › fairness in generative models
fair synthetic data generation
0.912025
TabFairGDT: A Fast Fair Tabular Data Generator Using Autoregressive Decision Trees · ICDM 2025
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks
0.812024
Connecting the Dots: Is Mode-Connectedness the Key to Feasible Sample-Based Inference in Bayesian Neural Networks? · ICML 2024
Machine learning › Kernel, tree and ensemble methods › ensemble learning
deep ensembles
0.812024
Connecting the Dots: Is Mode-Connectedness the Key to Feasible Sample-Based Inference in Bayesian Neural Networks? · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › sampling
sampling-based inference
0.812024
Connecting the Dots: Is Mode-Connectedness the Key to Feasible Sample-Based Inference in Bayesian Neural Networks? · ICML 2024
Machine learning › Trustworthy machine learning
uncertainty estimation
0.812024
Connecting the Dots: Is Mode-Connectedness the Key to Feasible Sample-Based Inference in Bayesian Neural Networks? · ICML 2024
Machine learning › Generative modeling › diffusion model
tabular data generation
0.312025
TabFairGDT: A Fast Fair Tabular Data Generator Using Autoregressive Decision Trees · ICDM 2025

Methods — techniques the papers use, named apart from their topics

soft leaf resampling · 0.9autoregressive decision trees · 0.9markov chain monte carlo · 0.8deep ensembles · 0.8
YearPublicationVenuePosition
2026 Invariance Pair Guidance: Robustness to Spurious Correlations via Corrective Gradients
abstract
Abstract Machine learning models are inherently bound to the distribution of the training data, often exploiting non-causal shortcuts. As a result, achieving robustness to spurious correlations remains a challenge. While existing approaches rely on data manipulation or re-weighting strategies to achieve robustness, they typically require dense group labels, multiple training domains, or specialized pre-processing. We propose Invariance Pair Guidance (IPG), a method to mitigate reliance on spurious correlations using a sparse set of counterfactual pairs. Unlike other methods demanding extensive supervision, IPG utilizes a novel dual-update mechanism to dynamically correct the optimization trajectory. We generate input pairs that isolate the spurious attribute to define the invariance, a characteristic that should not affect the outcome of the model. Based on these pairs, we define a corrective gradient that complements the traditional gradient descent approach. The correction adapts via a predefined invariance condition. Experiments on ColoredMNIST, Waterbirds-100, and CelebA datasets demonstrate the effectiveness of our approach and its robustness to group shifts, supported by a theoretical convergence analysis. IPG offers a data-efficient and theoretically grounded path to robustness.
Martin Surner, Abdelmajid Khelil, Ludwig Bothmann
Mach. Learn.3
2025 TabFairGDT: A Fast Fair Tabular Data Generator Using Autoregressive Decision Trees
abstract
Ensuring fairness in machine learning remains a significant challenge, as models often inherit biases from their training data. Generative models have recently emerged as a promising approach to mitigate bias at the data level while preserving utility. However, many rely on deep architectures, despite evidence that simpler models can be highly effective for tabular data. In this work, we introduce TabFairGDT, a novel method for generating fair synthetic tabular data using autoregressive decision trees. To enforce fairness, we propose a soft leaf resampling technique that adjusts decision tree outputs to reduce bias while preserving predictive performance. Our approach is non-parametric, effectively capturing complex relationships between mixed feature types, without relying on assumptions about the underlying data distributions. We evaluate TabFairGDT on benchmark fairness datasets and demonstrate that it outperforms state-of-the-art (SOTA) deep generative models, achieving better fairness-utility trade-off for downstream tasks, as well as higher synthetic data quality. Moreover, our method is lightweight, highly efficient, and CPU-compatible, requiring no data preprocessing. Remarkably, TabFairGDT achieves a 72% average speedup over the fastest SOTA baseline across various dataset sizes, and can generate fair synthetic data for medium-sized datasets (10 features, 10K samples) in just one second on a standard CPU, making it an ideal solution for realworld fairness-sensitive applications.
Emmanouil Panagiotou, Benoît Ronval, Arjun Roy 0001, Ludwig Bothmann, Bernd Bischl, Siegfried Nijssen, Eirini Ntoutsi
ICDM4
2024 Connecting the Dots: Is Mode-Connectedness the Key to Feasible Sample-Based Inference in Bayesian Neural Networks?
abstract
A major challenge in sample-based inference (SBI) for Bayesian neural networks is the size and structure of the networks’ parameter space. Our work shows that successful SBI is possible by embracing the characteristic relationship between weight and function space, uncovering a systematic link between overparameterization and the difficulty of the sampling problem. Through extensive experiments, we establish practical guidelines for sampling and convergence diagnosis. As a result, we present a deep ensemble initialized approach as an effective solution with competitive performance and uncertainty quantification.
Emanuel Sommer, Lisa Wimmer, Theodore Papamarkou, Ludwig Bothmann, Bernd Bischl, David Rügamer
ICML4
2023 Interpretable Regional Descriptors: Hyperbox-Based Local Explanations
Susanne Dandl, Giuseppe Casalicchio, Bernd Bischl, Ludwig Bothmann
ECML/PKDD (3)4