Laura Cabello Piqueras

dblp:317/0387 · also Laura Cabello · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2023
0000-0003-4276-7114ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 82% Information extraction and text analysis · 11% Efficient and distributed learning · 3%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
1.322023
Evaluating Bias and Fairness in Gender-Neutral Pretrained Vision-and-Language Models · EMNLP 2023
Being Right for Whose Right Reasons? · ACL (1) 2023
Machine learning › Trustworthy machine learning
interpretability
1.322023
Rather a Nurse than a Physician - Contrastive Explanations under Investigation · EMNLP 2023
Being Right for Whose Right Reasons? · ACL (1) 2023
Machine learning › Trustworthy machine learning › interpretability › local explanation
contrastive explanation
0.712023
Rather a Nurse than a Physician - Contrastive Explanations under Investigation · EMNLP 2023
Machine learning › Trustworthy machine learning › interpretability
explanation evaluation
0.712023
Being Right for Whose Right Reasons? · ACL (1) 2023
Machine learning › Trustworthy machine learning › fairness › gender bias
gender bias in vision-language models
0.712023
Evaluating Bias and Fairness in Gender-Neutral Pretrained Vision-and-Language Models · EMNLP 2023
Machine learning › Trustworthy machine learning › interpretability
post-hoc explanation
0.712023
Rather a Nurse than a Physician - Contrastive Explanations under Investigation · EMNLP 2023
Machine learning › Trustworthy machine learning
robustness
0.712023
Evaluating Bias and Fairness in Gender-Neutral Pretrained Vision-and-Language Models · EMNLP 2023
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
model distillation
0.212023
Being Right for Whose Right Reasons? · ACL (1) 2023
Natural language and speech › Information extraction and text analysis
text classification
0.212023
Rather a Nurse than a Physician - Contrastive Explanations under Investigation · EMNLP 2023
Computer vision › Vision and language
vision-language pretraining
0.212023
Evaluating Bias and Fairness in Gender-Neutral Pretrained Vision-and-Language Models · EMNLP 2023
Natural language and speech › Language models and text generation
multilingual language models
0.212022
Challenges and Strategies in Cross-Cultural NLP · ACL (1) 2022

Methods — techniques the papers use, named apart from their topics

transformer models · 0.7human rationale annotation · 0.7human annotation · 0.7gradnorm · 0.7gradientxinput · 0.7fine-tuning · 0.7continued pretraining · 0.7LRP · 0.7
YearPublicationVenuePosition
2023 Being Right for Whose Right Reasons?
abstract
Explainability methods are used to benchmark the extent to which model predictions align with human rationales i.e., are 'right for the right reasons'.Previous work has failed to acknowledge, however, that what counts as a rationale is sometimes subjective.This paper presents what we think is a first of its kind, a collection of human rationale annotations augmented with the annotators demographic information.We cover three datasets spanning sentiment analysis and common-sense reasoning, and six demographic groups (balanced across age and ethnicity).Such data enables us to ask both what demographics our predictions align with and whose reasoning patterns our models' rationales align with.We find systematic inter-group annotator disagreement and show how 16 Transformer-based models align better with rationales provided by certain demographic groups: We find that models are biased towards aligning best with older and/or white annotators.We zoom in on the effects of model size and model distillation, finding -contrary to our expectations -negative correlations between model size and rationale agreement as well as no evidence that either model size or model distillation improves fairness.
Terne Sasha Thorn Jakobsen, Laura Cabello Piqueras, Anders Søgaard
ACL (1)2
2023 Evaluating Bias and Fairness in Gender-Neutral Pretrained Vision-and-Language Models
abstract
Pretrained machine learning models are known to perpetuate and even amplify existing biases in data, which can result in unfair outcomes that ultimately impact user experience.Therefore, it is crucial to understand the mechanisms behind those prejudicial biases to ensure that model performance does not result in discriminatory behaviour toward certain groups or populations.In this work, we define gender bias as our case study.We quantify bias amplification in pretraining and after fine-tuning on three families of vision-and-language models.We investigate the connection, if any, between the two learning stages, and evaluate how bias amplification reflects on model performance.Overall, we find that bias amplification in pretraining and after fine-tuning are independent.We then examine the effect of continued pretraining on gender-neutral data, finding that this reduces group disparities, i.e., promotes fairness, on VQAv2 and retrieval tasks without significantly compromising task performance.
Laura Cabello Piqueras, Emanuele Bugliarello, Stephanie Brandl, Desmond Elliott
EMNLP1
2023 Rather a Nurse than a Physician - Contrastive Explanations under Investigation
abstract
Contrastive explanations, where one decision is explained in contrast to another, are supposed to be closer to how humans explain a decision than non-contrastive explanations, where the decision is not necessarily referenced to an alternative.This claim has never been empirically validated.We analyze four English text-classification datasets (SST2, DynaSent, BIOS and DBpedia-Animals).We fine-tune and extract explanations from three different models (RoBERTa, GTP-2, and T5), each in three different sizes and apply three post-hoc explainability methods (LRP, GradientxInput, GradNorm).We furthermore collect and release human rationale annotations for a subset of 100 samples from the BIOS dataset for contrastive and non-contrastive settings.A crosscomparison between model-based rationales and human annotations, both in contrastive and non-contrastive settings, yields a high agreement between the two settings for models as well as for humans.Moreover, model-based explanations computed in both settings align equally well with human rationales.Thus, we empirically find that humans do not necessarily explain in a contrastive manner.
Oliver Eberle, Ilias Chalkidis, Laura Cabello Piqueras, Stephanie Brandl
EMNLP3
2022 Challenges and Strategies in Cross-Cultural NLP
abstract
Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, Anders Søgaard. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Daniel Hershcovich, Stella Frank, Heather C. Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Aikaterini Margatina, Phillip Rust, Anders Søgaard
ACL (1)8
2022 Are Pretrained Multilingual Models Equally Fair across Languages?
abstract
Pretrained multilingual language models can help bridge the digital language divide, enabling high-quality NLP models for lower-resourced languages. Studies of multilingual models have so far focused on performance, consistency, and cross-lingual generalisation. However, with their wide-spread application in the wild and downstream societal impact, it is important to put multilingual models under the same scrutiny as monolingual models. This work investigates the group fairness of multilingual models, asking whether these models are equally fair across languages. To this end, we create a new four-way multilingual dataset of parallel cloze test examples (MozArt), equipped with demographic information (balanced with regard to gender and native tongue) about the test participants. We evaluate three multilingual models on MozArt –mBERT, XLM-R, and mT5– and show that across the four target languages, the three models exhibit different levels of group disparity, e.g., exhibiting near-equal risk for Spanish, but high levels of disparity for German.
Laura Cabello Piqueras, Anders Søgaard
COLING1
2022 Date Recognition in Historical Parish Records
Laura Cabello Piqueras, Constanza Fierro, Jonas F. Lotz, Phillip Rust, Joen Rommedahl, Jeppe Klok Due, Christian Igel, Desmond Elliott, Carsten B. Pedersen, Israfel Salazar, Anders Søgaard
ICFHR1