EDBT 2026 Demo / reviewers in the wild / expert
Sebastian Bordt
dblp:270/0462
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0001-8014-3594ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Trustworthy machine learning · 27% Language models and text generation · 24% Deep learning architectures and training · 24% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › large language model training
continual pre-training |
0.9 | 1 | 2025 | How Much Can We Forget about Data Contamination? · ICML 2025 |
Natural language and speech › Language models and text generation › large language model evaluation
data contamination |
0.9 | 1 | 2025 | How Much Can We Forget about Data Contamination? · ICML 2025 |
Machine learning › Learning theory › neural network theory
infinite-width limit |
0.9 | 1 | 2025 | On the Surprising Effectiveness of Large Learning Rates under Standard Width Scaling · NeurIPS 2025 |
Machine learning › Optimization for machine learning › learning rate
learning rate scaling |
0.9 | 1 | 2025 | On the Surprising Effectiveness of Large Learning Rates under Standard Width Scaling · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › weight initialization
parameterization and initialization |
0.9 | 1 | 2025 | On the Surprising Effectiveness of Large Learning Rates under Standard Width Scaling · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
scaling laws |
0.9 | 1 | 2025 | How Much Can We Forget about Data Contamination? · ICML 2025 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.7 | 1 | 2023 | Which Models have Perceptually-Aligned Gradients? An Explanation via Off-Manifold Robustness · NeurIPS 2023 |
Machine learning › Trustworthy machine learning
interpretability |
0.7 | 1 | 2023 | Which Models have Perceptually-Aligned Gradients? An Explanation via Off-Manifold Robustness · NeurIPS 2023 |
Machine learning › Trustworthy machine learning
robustness |
0.7 | 1 | 2023 | Which Models have Perceptually-Aligned Gradients? An Explanation via Off-Manifold Robustness · NeurIPS 2023 |
Machine learning › Generative modeling
image generation |
0.2 | 1 | 2023 | Which Models have Perceptually-Aligned Gradients? An Explanation via Off-Manifold Robustness · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
weight decay · 0.9mean squared error loss · 0.9cross-entropy loss · 0.9chinchilla scaling laws · 0.9randomized smoothing · 0.7gradient norm regularization · 0.7adversarial training · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | How Much Can We Forget about Data Contamination?abstractThe leakage of benchmark data into the training data has emerged as a significant challenge for evaluating the capabilities of large language models (LLMs). In this work, we challenge the common assumption that small-scale contamination renders benchmark evaluations invalid. First, we experimentally quantify the magnitude of benchmark overfitting based on scaling along three dimensions: The number of model parameters (up to 1.6B), the number of times an example is seen (up to 144), and the number of training tokens (up to 40B). If model and data follow the Chinchilla scaling laws, minor contamination indeed leads to overfitting. At the same time, even 144 times of contamination can be forgotten if the training data is scaled beyond five times Chinchilla, a regime characteristic of many modern LLMs. Continual pre-training of OLMo-7B corroborates these results. Next, we study the impact of the weight decay parameter on example forgetting, showing that empirical forgetting occurs faster than the cumulative weight decay. This allows us to gauge the degree of example forgetting in large-scale training runs, indicating that many LLMs, including Llama 3 405B, have forgotten the data seen at the beginning of training. Sebastian Bordt, Suraj Srinivas, Valentyn Boreiko, Ulrike von Luxburg |
ICML | 1 |
| 2025 | On the Surprising Effectiveness of Large Learning Rates under Standard Width ScalingabstractScaling limits, such as infinite-width limits, serve as promising theoretical tools to study large-scale models. However, it is widely believed that existing infinite-width theory does not faithfully explain the behavior of practical networks, especially those trained in *standard parameterization* (SP) meaning He initialization with a global learning rate. For instance, existing theory for SP predicts instability at large learning rates and vanishing feature learning at stable ones. In practice, however, optimal learning rates decay slower than theoretically predicted and networks exhibit both stable training and non-trivial feature learning, even at very large widths. Here, we show that this discrepancy is not fully explained by finite-width phenomena.
Instead, we find a resolution through a finer-grained analysis of the regime previously considered unstable and therefore uninteresting. In particular, we show that, under the cross-entropy (CE) loss, the unstable regime comprises two distinct sub-regimes: a catastrophically unstable regime and a more benign controlled divergence regime, where logits diverge but gradients and activations remain stable. Moreover, under large learning rates at the edge of the controlled divergence regime, there exists a well-defined infinite width limit where features continue to evolve in all the hidden layers. In experiments across optimizers, architectures, and data modalities, we validate that neural networks operate in this controlled divergence regime under CE loss but not under MSE loss. Our empirical evidence suggests that width-scaling considerations are surprisingly useful for predicting empirically maximal stable learning rate exponents which provide useful guidance on optimal learning rate exponents. Finally, our analysis clarifies the effectiveness and limitations of recently proposed layerwise learning rate scalings for standard initialization. Moritz Haas, Sebastian Bordt, Ulrike von Luxburg, Leena C. Vankadara |
NeurIPS | 2 |
| 2023 | From Shapley Values to Generalized Additive Models and backabstractIn explainable machine learning, local post-hoc explanation algorithms and inherently interpretable models are often seen as competing approaches. This work offers a partial reconciliation between the two by establishing a correspondence between Shapley Values and Generalized Additive Models (GAMs). We introduce $n$-Shapley Values, a parametric family of local post-hoc explanation algorithms that explain individual predictions with interaction terms up to order $n$. By varying the parameter $n$, we obtain a sequence of explanations that covers the entire range from Shapley Values up to a uniquely determined decomposition of the function we want to explain. The relationship between $n$-Shapley Values and this decomposition offers a functionally-grounded characterization of Shapley Values, which highlights their limitations. We then show that $n$-Shapley Values, as well as the Shapley Taylor- and Faith-Shap interaction indices, recover GAMs with interaction terms up to order $n$. This implies that the original Shapely Values recover GAMs without variable interactions. Taken together, our results provide a precise characterization of Shapley Values as they are being used in explainable machine learning. They also offer a principled interpretation of partial dependence plots of Shapley Values in terms of the underlying functional decomposition. A package for the estimation of different interaction indices is available at https://github.com/tml-tuebingen/nshap. Sebastian Bordt, Ulrike von Luxburg |
AISTATS | 1 |
| 2023 | Which Models have Perceptually-Aligned Gradients? An Explanation via Off-Manifold RobustnessabstractOne of the remarkable properties of robust computer vision models is that their input-gradients are often aligned with human perception, referred to in the literature as perceptually-aligned gradients (PAGs). Despite only being trained for classification, PAGs cause robust models to have rudimentary generative capabilities, including image generation, denoising, and in-painting. However, the underlying mechanisms behind these phenomena remain unknown. In this work, we provide a first explanation of PAGs via \emph{off-manifold robustness}, which states that models must be more robust off- the data manifold than they are on-manifold. We first demonstrate theoretically that off-manifold robustness leads input gradients to lie approximately on the data manifold, explaining their perceptual alignment. We then show that Bayes optimal models satisfy off-manifold robustness, and confirm the same empirically for robust models trained via gradient norm regularization, randomized smoothing, and adversarial training with projected gradient descent. Quantifying the perceptual alignment of model gradients via their similarity with the gradients of generative models, we show that off-manifold robustness correlates well with perceptual alignment. Finally, based on the levels of on- and off-manifold robustness, we identify three different regimes of robustness that affect both perceptual alignment and model accuracy: weak robustness, bayes-aligned robustness, and excessive robustness. Code is available at https://github.com/tml-tuebingen/pags. Suraj Srinivas, Sebastian Bordt, Himabindu Lakkaraju |
NeurIPS | 2 |
| 2022 | A Bandit Model for Human-Machine Decision Making with Private Information and OpacityabstractApplications of machine learning inform human decision makers in a broad range of tasks. The resulting problem is usually formulated in terms of a single decision maker. We argue that it should rather be described as a two-player learning problem where one player is the machine and the other the human. While both players try to optimize the final decision, the setup is often characterized by (1) the presence of private information and (2) opacity, that is imperfect understanding between the decision makers. We prove that both properties can complicate decision making considerably. A lower bound quantifies the worst-case hardness of optimally advising a decision maker who is opaque or has access to private information. An upper bound shows that a simple coordination strategy is nearly minimax optimal. More efficient learning is possible under certain assumptions on the problem, for example that both players learn to take actions independently. Such assumptions are implicit in existing literature, for example in medical applications of machine learning, but have not been described or justified theoretically. Sebastian Bordt, Ulrike von Luxburg |
AISTATS | 1 |
| 2021 | Recovery Guarantees for Kernel-based Clustering under Non-parametric Mixture ModelsabstractDespite the ubiquity of kernel-based clustering, surprisingly few statistical guarantees exist beyond settings that consider strong structural assumptions on the data generation process. In this work, we take a step towards bridging this gap by studying the statistical performance of kernel-based clustering algorithms under non-parametric mixture models. We provide necessary and sufficient separability conditions under which these algorithms can consistently recover the underlying true clustering. Our analysis provides guarantees for kernel clustering approaches without structural assumptions on the form of the component distributions. Additionally, we establish a key equivalence between kernel-based data-clustering and kernel density-based clustering. This enables us to provide consistency guarantees for kernel-based estimators of non-parametric mixture models. Along with theoretical implications, this connection could have practical implications, including in the systematic choice of the bandwidth of the Gaussian kernel in the context of clustering. Leena C. Vankadara, Sebastian Bordt, Ulrike von Luxburg, Debarghya Ghoshdastidar |
AISTATS | 2 |