EDBT 2026 Demo / reviewers in the wild / expert
Valentyn Boreiko
dblp:320/7840
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Trustworthy machine learning · 57% Language models and text generation · 25% Deep learning architectures and training · 8% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
robustness |
1.3 | 2 | 2023 | Spurious Features Everywhere - Large-Scale Detection of Harmful Spurious Features in ImageNet · ICCV 2023 Identification of Systematic Errors of Image Classifiers on Rare Subgroups · ICCV 2023 |
Natural language and speech › Language models and text generation › large language model training
continual pre-training |
0.9 | 1 | 2025 | How Much Can We Forget about Data Contamination? · ICML 2025 |
Natural language and speech › Language models and text generation › large language model evaluation
data contamination |
0.9 | 1 | 2025 | How Much Can We Forget about Data Contamination? · ICML 2025 |
Natural language and speech › Language models and text generation › language modeling
n-gram language model |
0.9 | 1 | 2025 | An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks · ICML 2025 |
Machine learning › Deep learning architectures and training
scaling laws |
0.9 | 1 | 2025 | How Much Can We Forget about Data Contamination? · ICML 2025 |
Machine learning › Trustworthy machine learning
threat model |
0.9 | 1 | 2025 | An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks · ICML 2025 |
Security and privacy of machine learning › adversarial attack
jailbreak attack |
0.9 | 1 | 2025 | An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks · ICML 2025 |
Machine learning › Trustworthy machine learning
fairness |
0.7 | 1 | 2023 | Identification of Systematic Errors of Image Classifiers on Rare Subgroups · ICCV 2023 |
Machine learning › Trustworthy machine learning › robustness › spurious correlation
spurious correlation mitigation |
0.7 | 1 | 2023 | Spurious Features Everywhere - Large-Scale Detection of Harmful Spurious Features in ImageNet · ICCV 2023 |
Machine learning › Trustworthy machine learning › interpretability
counterfactual explanation |
0.6 | 1 | 2022 | Diffusion Visual Counterfactual Explanations · NeurIPS 2022 |
Machine learning › Generative modeling
diffusion model |
0.6 | 1 | 2022 | Diffusion Visual Counterfactual Explanations · NeurIPS 2022 |
Machine learning › Trustworthy machine learning
interpretability |
0.6 | 1 | 2022 | Diffusion Visual Counterfactual Explanations · NeurIPS 2022 |
Machine learning › Trustworthy machine learning › interpretability › counterfactual explanation
visual counterfactual explanation |
0.6 | 1 | 2022 | Diffusion Visual Counterfactual Explanations · NeurIPS 2022 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.2 | 1 | 2023 | Identification of Systematic Errors of Image Classifiers on Rare Subgroups · ICCV 2023 |
Computer vision › Image recognition and object detection
image classification |
0.2 | 1 | 2022 | Diffusion Visual Counterfactual Explanations · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
n-gram perplexity · 1.7discrete optimization · 1.7weight decay · 0.9chinchilla scaling laws · 0.9text-to-image synthesis · 0.7spufix · 0.7prompt search · 0.7neural PCA · 0.7combinatorial testing · 0.7adaptive parameterization · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | How Much Can We Forget about Data Contamination?abstractThe leakage of benchmark data into the training data has emerged as a significant challenge for evaluating the capabilities of large language models (LLMs). In this work, we challenge the common assumption that small-scale contamination renders benchmark evaluations invalid. First, we experimentally quantify the magnitude of benchmark overfitting based on scaling along three dimensions: The number of model parameters (up to 1.6B), the number of times an example is seen (up to 144), and the number of training tokens (up to 40B). If model and data follow the Chinchilla scaling laws, minor contamination indeed leads to overfitting. At the same time, even 144 times of contamination can be forgotten if the training data is scaled beyond five times Chinchilla, a regime characteristic of many modern LLMs. Continual pre-training of OLMo-7B corroborates these results. Next, we study the impact of the weight decay parameter on example forgetting, showing that empirical forgetting occurs faster than the cumulative weight decay. This allows us to gauge the degree of example forgetting in large-scale training runs, indicating that many LLMs, including Llama 3 405B, have forgotten the data seen at the beginning of training. Sebastian Bordt, Suraj Srinivas, Valentyn Boreiko, Ulrike von Luxburg |
ICML | 3 |
| 2025 | An Interpretable N-gram Perplexity Threat Model for Large Language Model JailbreaksabstractA plethora of jailbreaking attacks have been proposed to obtain harmful responses from safety-tuned LLMs. These methods largely succeed in coercing the target output in their original settings, but their attacks vary substantially in fluency and computational effort. In this work, we propose a unified threat model for the principled comparison of these methods.
Our threat model checks if a given jailbreak is likely to occur in the distribution of text. For this, we build an N-gram language model on 1T tokens, which, unlike model-based perplexity, allows for an LLM-agnostic, nonparametric, and inherently interpretable evaluation. We adapt popular attacks to this threat model, and, for the first time, benchmark these attacks on equal footing with it. After an extensive comparison, we find attack success rates against safety-tuned modern models to be lower than previously presented and that attacks based on discrete optimization significantly outperform recent LLM-based attacks. Being inherently interpretable, our threat model allows for a comprehensive analysis and comparison of jailbreak attacks. We find that effective attacks exploit and abuse infrequent bigrams, either selecting the ones absent from real-world text or rare ones, e.g., specific to Reddit or code datasets. Valentyn Boreiko, Alexander Panfilov, Václav Vorácek, Matthias Hein 0001, Jonas Geiping |
ICML | 1 |
| 2023 | Identification of Systematic Errors of Image Classifiers on Rare SubgroupsabstractDespite excellent average-case performance of many image classifiers, their performance can substantially deteriorate on semantically coherent subgroups of the data that were under-represented in the training data. These systematic errors can impact both fairness for demographic minority groups as well as robustness and safety under domain shift. A major challenge is to identify such subgroups with subpar performance when the subgroups are not annotated and their occurrence is very rare. We leverage recent advances in text-to-image models and search in the space of textual descriptions of subgroups ("prompts") for sub-groups where the target model has low performance on the prompt-conditioned synthesized data. To tackle the exponentially growing number of subgroups, we employ combinatorial testing. We denote this procedure as PromptAttack as it can be interpreted as an adversarial attack in a prompt space. We study subgroup coverage and identifiability with PromptAttack in a controlled setting and find that it identifies systematic errors with high accuracy. Thereupon, we apply PromptAttack to ImageNet classifiers and identify novel systematic errors on rare subgroups. Jan Hendrik Metzen, Robin Hutmacher, N. Grace Hua, Valentyn Boreiko |
ICCV | 4 |
| 2023 | Spurious Features Everywhere - Large-Scale Detection of Harmful Spurious Features in ImageNetabstractBenchmark performance of deep learning classifiers alone is not a reliable predictor for the performance of a deployed model. In particular, if the image classifier has picked up spurious features in the training data, its predictions can fail in unexpected ways. In this paper, we develop a framework that allows us to systematically identify spurious features in large datasets like ImageNet. It is based on our neural PCA components and their visualization. Previous work on spurious features often operates in toy settings or requires costly pixel-wise annotations. In contrast, we work with ImageNet and validate our results by showing that presence of the harmful spurious feature of a class alone is sufficient to trigger the prediction of that class. We introduce the novel dataset "Spurious ImageNet" which allows to measure the reliance of any ImageNet classifier on harmful spurious features. Moreover, we introduce SpuFix as a simple mitigation method to reduce the dependence of any ImageNet classifier on previously identified harmful spurious features without requiring additional labels or retraining of the model. We provide code and data at https://github.com/YanNeu/spurious_imagenet. Yannic Neuhaus, Maximilian Augustin, Valentyn Boreiko, Matthias Hein 0001 |
ICCV | 3 |
| 2022 | Visual Explanations for the Detection of Diabetic Retinopathy from Retinal Fundus Images
Valentyn Boreiko, Indu Ilanchezian, Murat Seçkin Ayhan, Sarah Müller, Lisa M. Koch, Hanna Faber, Philipp Berens, Matthias Hein 0001 |
MICCAI (2) | 1 |
| 2022 | Diffusion Visual Counterfactual ExplanationsabstractVisual Counterfactual Explanations (VCEs) are an important tool to understand the decisions of an image classifier. They are “small” but “realistic” semantic changes of the image changing the classifier decision. Current approaches for the generation of VCEs are restricted to adversarially robust models and often contain non-realistic artefacts, or are limited to image classification problems with few classes. In this paper, we overcome this by generating Diffusion Visual Counterfactual Explanations (DVCEs) for arbitrary ImageNet classifiers via a diffusion process. Two modifications to the diffusion process are key for our DVCEs: first, an adaptive parameterization, whose hyperparameters generalize across images and models, together with distance regularization and late start of the diffusion process, allow us to generate images with minimal semantic changes to the original ones but different classification. Second, our cone regularization via an adversarially robust model ensures that the diffusion process does not converge to trivial non-semantic changes, but instead produces realistic images of the target class which achieve high confidence by the classifier. Maximilian Augustin, Valentyn Boreiko, Francesco Croce, Matthias Hein 0001 |
NeurIPS | 2 |