Angeline Pouget

dblp:292/7108 · also Angéline Pouget · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 50% Transfer learning and domain adaptation · 23% Vision and language · 20%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation › domain shift
covariate shift
0.912025
Suitability Filter: A Statistical Framework for Classifier Evaluation in Real-World Deployment Settings · ICML 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
Suitability Filter: A Statistical Framework for Classifier Evaluation in Real-World Deployment Settings · ICML 2025
Computer vision › Vision and language › vision-language model
contrastive vision-language model
0.812024
No Filter: Cultural and Socioeconomic Diversity in Contrastive Vision-Language Models · NeurIPS 2024
Machine learning › Trustworthy machine learning
fairness
0.812024
No Filter: Cultural and Socioeconomic Diversity in Contrastive Vision-Language Models · NeurIPS 2024
Machine learning › Trustworthy machine learning
uncertainty estimation
0.312025
Suitability Filter: A Statistical Framework for Classifier Evaluation in Real-World Deployment Settings · ICML 2025
Computer vision › 3D vision › visual localization
geo-localization
0.212024
No Filter: Cultural and Socioeconomic Diversity in Contrastive Vision-Language Models · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

suitability signal · 0.9statistical hypothesis testing · 0.9contrastive pre-training · 0.8benchmark evaluation · 0.8
YearPublicationVenuePosition
2025 Suitability Filter: A Statistical Framework for Classifier Evaluation in Real-World Deployment Settings
abstract
Deploying machine learning models in safety-critical domains poses a key challenge: ensuring reliable model performance on downstream user data without access to ground truth labels for direct validation. We propose the _suitability filter_, a novel framework designed to detect performance deterioration by utilizing _suitability signals_—model output features that are sensitive to covariate shifts and indicative of potential prediction errors. The suitability filter evaluates whether classifier accuracy on unlabeled user data shows significant degradation compared to the accuracy measured on the labeled test dataset. Specifically, it ensures that this degradation does not exceed a pre-specified margin, which represents the maximum acceptable drop in accuracy. To achieve reliable performance evaluation, we aggregate suitability signals for both test and user data and compare these empirical distributions using statistical hypothesis testing, thus providing insights into decision uncertainty. Our modular method adapts to various models and domains. Empirical evaluations across different classification tasks demonstrate that the suitability filter reliably detects performance deviations due to covariate shift. This enables proactive mitigation of potential failures in high-stakes applications.
Angeline Pouget, Mohammad Yaghini, Stephan Rabanser, Nicolas Papernot
ICML1
2024 No Filter: Cultural and Socioeconomic Diversity in Contrastive Vision-Language Models
abstract
We study cultural and socioeconomic diversity in contrastive vision-language models (VLMs). Using a broad range of benchmark datasets and evaluation metrics, we bring to attention several important findings. First, the common filtering of training data to English image-text pairs disadvantages communities of lower socioeconomic status and negatively impacts cultural understanding. Notably, this performance gap is not captured by - and even at odds with - the currently popular evaluation metrics derived from the Western-centric ImageNet and COCO datasets. Second, pretraining with global, unfiltered data before fine-tuning on English content can improve cultural understanding without sacrificing performance on said popular benchmarks. Third, we introduce the task of geo-localization as a novel evaluation metric to assess cultural diversity in VLMs. Our work underscores the value of using diverse data to create more inclusive multimodal systems and lays the groundwork for developing VLMs that better represent global perspectives.
Angeline Pouget, Lucas Beyer, Emanuele Bugliarello, Xiao Wang 0038, Andreas Steiner 0001, Xiaohua Zhai, Ibrahim Alabdulmohsin
NeurIPS1