Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sebastian Bordt

dblp:270/0462 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0001-8014-3594ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 27% Language models and text generation · 24% Deep learning architectures and training · 24%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › large language model training
continual pre-training
0.912025
How Much Can We Forget about Data Contamination? · ICML 2025
Natural language and speech › Language models and text generation › large language model evaluation
data contamination
0.912025
How Much Can We Forget about Data Contamination? · ICML 2025
Machine learning › Learning theory › neural network theory
infinite-width limit
0.912025
On the Surprising Effectiveness of Large Learning Rates under Standard Width Scaling · NeurIPS 2025
Machine learning › Optimization for machine learning › learning rate
learning rate scaling
0.912025
On the Surprising Effectiveness of Large Learning Rates under Standard Width Scaling · NeurIPS 2025
Machine learning › Deep learning architectures and training › weight initialization
parameterization and initialization
0.912025
On the Surprising Effectiveness of Large Learning Rates under Standard Width Scaling · NeurIPS 2025
Machine learning › Deep learning architectures and training
scaling laws
0.912025
How Much Can We Forget about Data Contamination? · ICML 2025
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.712023
Which Models have Perceptually-Aligned Gradients? An Explanation via Off-Manifold Robustness · NeurIPS 2023
Machine learning › Trustworthy machine learning
interpretability
0.712023
Which Models have Perceptually-Aligned Gradients? An Explanation via Off-Manifold Robustness · NeurIPS 2023
Machine learning › Trustworthy machine learning
robustness
0.712023
Which Models have Perceptually-Aligned Gradients? An Explanation via Off-Manifold Robustness · NeurIPS 2023
Machine learning › Generative modeling
image generation
0.212023
Which Models have Perceptually-Aligned Gradients? An Explanation via Off-Manifold Robustness · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

weight decay · 0.9mean squared error loss · 0.9cross-entropy loss · 0.9chinchilla scaling laws · 0.9randomized smoothing · 0.7gradient norm regularization · 0.7adversarial training · 0.7
YearPublicationVenuePosition
2025 How Much Can We Forget about Data Contamination?
abstract
The leakage of benchmark data into the training data has emerged as a significant challenge for evaluating the capabilities of large language models (LLMs). In this work, we challenge the common assumption that small-scale contamination renders benchmark evaluations invalid. First, we experimentally quantify the magnitude of benchmark overfitting based on scaling along three dimensions: The number of model parameters (up to 1.6B), the number of times an example is seen (up to 144), and the number of training tokens (up to 40B). If model and data follow the Chinchilla scaling laws, minor contamination indeed leads to overfitting. At the same time, even 144 times of contamination can be forgotten if the training data is scaled beyond five times Chinchilla, a regime characteristic of many modern LLMs. Continual pre-training of OLMo-7B corroborates these results. Next, we study the impact of the weight decay parameter on example forgetting, showing that empirical forgetting occurs faster than the cumulative weight decay. This allows us to gauge the degree of example forgetting in large-scale training runs, indicating that many LLMs, including Llama 3 405B, have forgotten the data seen at the beginning of training.
Sebastian Bordt, Suraj Srinivas, Valentyn Boreiko, Ulrike von Luxburg
ICML1
2025 On the Surprising Effectiveness of Large Learning Rates under Standard Width Scaling
abstract
Scaling limits, such as infinite-width limits, serve as promising theoretical tools to study large-scale models. However, it is widely believed that existing infinite-width theory does not faithfully explain the behavior of practical networks, especially those trained in *standard parameterization* (SP) meaning He initialization with a global learning rate. For instance, existing theory for SP predicts instability at large learning rates and vanishing feature learning at stable ones. In practice, however, optimal learning rates decay slower than theoretically predicted and networks exhibit both stable training and non-trivial feature learning, even at very large widths. Here, we show that this discrepancy is not fully explained by finite-width phenomena. Instead, we find a resolution through a finer-grained analysis of the regime previously considered unstable and therefore uninteresting. In particular, we show that, under the cross-entropy (CE) loss, the unstable regime comprises two distinct sub-regimes: a catastrophically unstable regime and a more benign controlled divergence regime, where logits diverge but gradients and activations remain stable. Moreover, under large learning rates at the edge of the controlled divergence regime, there exists a well-defined infinite width limit where features continue to evolve in all the hidden layers. In experiments across optimizers, architectures, and data modalities, we validate that neural networks operate in this controlled divergence regime under CE loss but not under MSE loss. Our empirical evidence suggests that width-scaling considerations are surprisingly useful for predicting empirically maximal stable learning rate exponents which provide useful guidance on optimal learning rate exponents. Finally, our analysis clarifies the effectiveness and limitations of recently proposed layerwise learning rate scalings for standard initialization.
Moritz Haas, Sebastian Bordt, Ulrike von Luxburg, Leena C. Vankadara
NeurIPS2
2023 From Shapley Values to Generalized Additive Models and back
abstract
In explainable machine learning, local post-hoc explanation algorithms and inherently interpretable models are often seen as competing approaches. This work offers a partial reconciliation between the two by establishing a correspondence between Shapley Values and Generalized Additive Models (GAMs). We introduce $n$-Shapley Values, a parametric family of local post-hoc explanation algorithms that explain individual predictions with interaction terms up to order $n$. By varying the parameter $n$, we obtain a sequence of explanations that covers the entire range from Shapley Values up to a uniquely determined decomposition of the function we want to explain. The relationship between $n$-Shapley Values and this decomposition offers a functionally-grounded characterization of Shapley Values, which highlights their limitations. We then show that $n$-Shapley Values, as well as the Shapley Taylor- and Faith-Shap interaction indices, recover GAMs with interaction terms up to order $n$. This implies that the original Shapely Values recover GAMs without variable interactions. Taken together, our results provide a precise characterization of Shapley Values as they are being used in explainable machine learning. They also offer a principled interpretation of partial dependence plots of Shapley Values in terms of the underlying functional decomposition. A package for the estimation of different interaction indices is available at https://github.com/tml-tuebingen/nshap.
Sebastian Bordt, Ulrike von Luxburg
AISTATS1
2023 Which Models have Perceptually-Aligned Gradients? An Explanation via Off-Manifold Robustness
abstract
One of the remarkable properties of robust computer vision models is that their input-gradients are often aligned with human perception, referred to in the literature as perceptually-aligned gradients (PAGs). Despite only being trained for classification, PAGs cause robust models to have rudimentary generative capabilities, including image generation, denoising, and in-painting. However, the underlying mechanisms behind these phenomena remain unknown. In this work, we provide a first explanation of PAGs via \emph{off-manifold robustness}, which states that models must be more robust off- the data manifold than they are on-manifold. We first demonstrate theoretically that off-manifold robustness leads input gradients to lie approximately on the data manifold, explaining their perceptual alignment. We then show that Bayes optimal models satisfy off-manifold robustness, and confirm the same empirically for robust models trained via gradient norm regularization, randomized smoothing, and adversarial training with projected gradient descent. Quantifying the perceptual alignment of model gradients via their similarity with the gradients of generative models, we show that off-manifold robustness correlates well with perceptual alignment. Finally, based on the levels of on- and off-manifold robustness, we identify three different regimes of robustness that affect both perceptual alignment and model accuracy: weak robustness, bayes-aligned robustness, and excessive robustness. Code is available at https://github.com/tml-tuebingen/pags.
Suraj Srinivas, Sebastian Bordt, Himabindu Lakkaraju
NeurIPS2
2022 A Bandit Model for Human-Machine Decision Making with Private Information and Opacity
abstract
Applications of machine learning inform human decision makers in a broad range of tasks. The resulting problem is usually formulated in terms of a single decision maker. We argue that it should rather be described as a two-player learning problem where one player is the machine and the other the human. While both players try to optimize the final decision, the setup is often characterized by (1) the presence of private information and (2) opacity, that is imperfect understanding between the decision makers. We prove that both properties can complicate decision making considerably. A lower bound quantifies the worst-case hardness of optimally advising a decision maker who is opaque or has access to private information. An upper bound shows that a simple coordination strategy is nearly minimax optimal. More efficient learning is possible under certain assumptions on the problem, for example that both players learn to take actions independently. Such assumptions are implicit in existing literature, for example in medical applications of machine learning, but have not been described or justified theoretically.
Sebastian Bordt, Ulrike von Luxburg
AISTATS1
2021 Recovery Guarantees for Kernel-based Clustering under Non-parametric Mixture Models
abstract
Despite the ubiquity of kernel-based clustering, surprisingly few statistical guarantees exist beyond settings that consider strong structural assumptions on the data generation process. In this work, we take a step towards bridging this gap by studying the statistical performance of kernel-based clustering algorithms under non-parametric mixture models. We provide necessary and sufficient separability conditions under which these algorithms can consistently recover the underlying true clustering. Our analysis provides guarantees for kernel clustering approaches without structural assumptions on the form of the component distributions. Additionally, we establish a key equivalence between kernel-based data-clustering and kernel density-based clustering. This enables us to provide consistency guarantees for kernel-based estimators of non-parametric mixture models. Along with theoretical implications, this connection could have practical implications, including in the systematic choice of the bandwidth of the Gaussian kernel in the context of clustering.
Leena C. Vankadara, Sebastian Bordt, Ulrike von Luxburg, Debarghya Ghoshdastidar
AISTATS2