Evgenia Rusak

dblp:245/2556 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-6039-7781ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Trustworthy machine learning · 32% Vision and language · 31% Efficient and distributed learning · 16%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
CLIP
1.622025
In Search of Forgotten Domain Generalization · ICLR 2025
Does CLIP's generalization performance mainly stem from high train-test similarity? · ICLR 2024
Machine learning › Trustworthy machine learning
out-of-distribution generalization
1.622025
In Search of Forgotten Domain Generalization · ICLR 2025
Does CLIP's generalization performance mainly stem from high train-test similarity? · ICLR 2024
Computer vision › Vision and language
vision-language model
1.622025
In Search of Forgotten Domain Generalization · ICLR 2025
Does CLIP's generalization performance mainly stem from high train-test similarity? · ICLR 2024
Machine learning › Efficient and distributed learning › data selection
data pruning
1.022024
Effective pruning of web-scale datasets based on complexity of concept clusters · ICLR 2024
Does CLIP's generalization performance mainly stem from high train-test similarity? · ICLR 2024
Machine learning › Transfer learning and domain adaptation
domain generalization
0.912025
In Search of Forgotten Domain Generalization · ICLR 2025
Machine learning › Trustworthy machine learning
robustness
0.922020
Improving robustness against common corruptions by covariate shift adaptation · NeurIPS 2020
A Simple Way to Make Neural Networks Robust Against Diverse Image Corruptions · ECCV (3) 2020
Machine learning › Efficient and distributed learning
data-efficient learning
0.812024
Effective pruning of web-scale datasets based on complexity of concept clusters · ICLR 2024
Machine learning › Trustworthy machine learning › robustness › corruption robustness
common corruption robustness
0.412020
Improving robustness against common corruptions by covariate shift adaptation · NeurIPS 2020
Machine learning › Transfer learning and domain adaptation › domain adaptation › distribution adaptation
covariate shift adaptation
0.412020
Improving robustness against common corruptions by covariate shift adaptation · NeurIPS 2020
Machine learning › Trustworthy machine learning › robustness › corruption robustness
image corruption robustness
0.412020
A Simple Way to Make Neural Networks Robust Against Diverse Image Corruptions · ECCV (3) 2020
Machine learning › Transfer learning and domain adaptation
test-time adaptation
0.412020
Improving robustness against common corruptions by covariate shift adaptation · NeurIPS 2020
Natural language and speech › Language models and text generation › large language model training
data mixing
0.312025
In Search of Forgotten Domain Generalization · ICLR 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.212024
Effective pruning of web-scale datasets based on complexity of concept clusters · ICLR 2024

Methods — techniques the papers use, named apart from their topics

domain mixing · 0.9dataset subsampling · 0.9retraining · 0.8data pruning · 0.8concept clustering · 0.8complexity-based pruning · 0.8unsupervised online adaptation · 0.4data augmentation · 0.4batch normalization statistics adaptation · 0.4adversarial training · 0.4
YearPublicationVenuePosition
2025 InfoNCE: Identifying the Gap Between Theory and Practice
abstract
Prior theory work on Contrastive Learning via the InfoNCE loss showed that, under certain assumptions, the learned representations recover the ground-truth latent factors. We argue that these theories overlook crucial aspects of how CL is deployed in practice. Specifically, they either assume equal variance across all latents or that certain latents are kept invariant. However, in practice, positive pairs are often generated using augmentations such as strong cropping to just a few pixels. Hence, a more realistic assumption is that all latent factors change with a continuum of variability across all factors. We introduce AnInfoNCE, a generalization of InfoNCE that can provably uncover the latent factors in this anisotropic setting, broadly generalizing previous identifiability results in CL. We validate our identifiability results in controlled experiments and show that AnInfoNCE increases the recovery of previously collapsed information in CIFAR10 and ImageNet, albeit at the cost of downstream accuracy. Finally, we discuss the remaining mismatches between theoretical assumptions and practical implementations.
Evgenia Rusak, Patrik Reizinger, Attila Juhos, Oliver Bringmann 0001, Roland S. Zimmermann, Wieland Brendel
AISTATS1
2025 In Search of Forgotten Domain Generalization
abstract
Out-of-Domain (OOD) generalization is the ability of a model trained on one or more domains to generalize to unseen domains. In the ImageNet era of computer vision, evaluation sets for measuring a model's OOD performance were designed to be strictly OOD with respect to style. However, the emergence of foundation models and expansive web-scale datasets has obfuscated this evaluation process, as datasets cover a broad range of domains and risk test domain contamination. In search of the forgotten domain generalization, we create large-scale datasets subsampled from LAION---LAION-Natural and LAION-Rendition---that are strictly OOD to corresponding ImageNet and DomainNet test sets in terms of style. Training CLIP models on these datasets reveals that a significant portion of their performance is explained by in-domain examples. This indicates that the OOD generalization challenges from the ImageNet era still prevail and that training on web-scale data merely creates the illusion of OOD generalization. Furthermore, through a systematic exploration of combining natural and rendition datasets in varying proportions, we identify optimal mixing ratios for model generalization across these domains. Our datasets and results re-enable meaningful assessment of OOD robustness at scale---a crucial prerequisite for improving model robustness.
Prasanna Mayilvahanan, Roland S. Zimmermann, Thaddäus Wiedemer, Evgenia Rusak, Attila Juhos, Matthias Bethge, Wieland Brendel
ICLR4
2024 Effective pruning of web-scale datasets based on complexity of concept clusters
abstract
Utilizing massive web-scale datasets has led to unprecedented performance gains in machine learning models, but also imposes outlandish compute requirements for their training. In order to improve training and data efficiency, we here push the limits of pruning large-scale multimodal datasets for training CLIP-style models. Today’s most effective pruning method on ImageNet clusters data samples into separate concepts according to their embedding and prunes away the most proto- typical samples. We scale this approach to LAION and improve it by noting that the pruning rate should be concept-specific and adapted to the complexity of the concept. Using a simple and intuitive complexity measure, we are able to reduce the training cost to a quarter of regular training. More specifically, we are able to outperform the LAION-trained OpenCLIP-ViT-B/32 model on ImageNet zero-shot accuracy by 1.1p.p. while only using 27.7% of the data and training compute. On the DataComp Medium benchmark, we achieve a new state-of-the-art ImageNet zero-shot accuracy and a competitive average zero-shot accuracy on 38 evaluation tasks.
Amro Abbas, Evgenia Rusak, Kushal Tirumala, Wieland Brendel, Kamalika Chaudhuri, Ari S. Morcos
ICLR2
2024 Does CLIP's generalization performance mainly stem from high train-test similarity?
abstract
Foundation models like CLIP are trained on hundreds of millions of samples and effortlessly generalize to new tasks and inputs. Out of the box, CLIP shows stellar zero-shot and few-shot capabilities on a wide range of out-of-distribution (OOD) benchmarks, which prior works attribute mainly to today's large and comprehensive training dataset (like LAION). However, it is questionable how meaningful terms like out-of-distribution generalization are for CLIP as it seems likely that web-scale datasets like LAION simply contain many samples that are similar to common OOD benchmarks originally designed for ImageNet. To test this hypothesis, we retrain CLIP on pruned LAION splits that replicate ImageNet’s train-test similarity with respect to common OOD benchmarks. While we observe a performance drop on some benchmarks, surprisingly, CLIP’s overall performance remains high. This shows that high train-test similarity is insufficient to explain CLIP’s OOD performance, and other properties of the training data must drive CLIP to learn more generalizable representations. Additionally, by pruning data points that are dissimilar to the OOD benchmarks, we uncover a 100M split of LAION (¼ of its original size) on which CLIP can be trained to match its original OOD performance.
Prasanna Mayilvahanan, Thaddäus Wiedemer, Evgenia Rusak, Matthias Bethge, Wieland Brendel
ICLR3
2023 Robust deep learning object recognition models rely on low frequency information in natural images
abstract
Machine learning models have difficulty generalizing to data outside of the distribution they were trained on. In particular, vision models are usually vulnerable to adversarial attacks or common corruptions, to which the human visual system is robust. Recent studies have found that regularizing machine learning models to favor brain-like representations can improve model robustness, but it is unclear why. We hypothesize that the increased model robustness is partly due to the low spatial frequency preference inherited from the neural representation. We tested this simple hypothesis with several frequency-oriented analyses, including the design and use of hybrid images to probe model frequency sensitivity directly. We also examined many other publicly available robust models that were trained on adversarial images or with data augmentation, and found that all these robust models showed a greater preference to low spatial frequency information. We show that preprocessing by blurring can serve as a defense mechanism against both adversarial attacks and common corruptions, further confirming our hypothesis and demonstrating the utility of low spatial frequency information in robust object recognition.
Zhe Li 0002, Josue Ortega Caro, Evgenia Rusak, Wieland Brendel, Matthias Bethge, Fabio Anselmi, Ankit B. Patel, Andreas S. Tolias, Xaq Pitkow
PLoS Comput. Biol.3
2020 A Simple Way to Make Neural Networks Robust Against Diverse Image Corruptions
Evgenia Rusak, Lukas Schott, Roland S. Zimmermann, Julian Bitterwolf, Oliver Bringmann 0001, Matthias Bethge, Wieland Brendel
ECCV (3)1
2020 Improving robustness against common corruptions by covariate shift adaptation
abstract
Today’s state-of-the-art machine vision models are vulnerable to image corruptions like blurring or compression artefacts, limiting their performance in many real-world applications. We here argue that popular benchmarks to measure model robustness against common corruptions (like ImageNet-C) underestimate model robustness in many (but not all) application scenarios. The key insight is that in many scenarios, multiple unlabeled examples of the corruptions are available and can be used for unsupervised online adaptation. Replacing the activation statistics estimated by batch normalization on the training set with the statistics of the corrupted images consistently improves the robustness across 25 different popular computer vision models. Using the corrected statistics, ResNet-50 reaches 62.2% mCE on ImageNet-C compared to 76.7% without adaptation. With the more robust DeepAugment+AugMix model, we improve the state of the art achieved by a ResNet50 model up to date from 53.6% mCE to 45.4% mCE. Even adapting to a single sample improves robustness for the ResNet-50 and AugMix models, and 32 samples are sufficient to improve the current state of the art for a ResNet-50 architecture. We argue that results with adapted statistics should be included whenever reporting scores in corruption benchmarks and other out-of-distribution generalization settings.
Steffen Schneider 0001, Evgenia Rusak, Luisa Eck, Oliver Bringmann 0001, Wieland Brendel, Matthias Bethge
NeurIPS2